faroese-torshavn-fo
收藏资源简介:
该数据集名为“Tórshavnar kommuna”(标识符:torshavn_fo),是一个独立的网站收集数据集,不属于法罗语 Dynaword 项目的一部分。数据来源于法罗群岛 Tórshavnar 市政府的官方网站(torshavn.fo),涵盖了该市政府的公开记录,包括议会会议记录、预算、年度账目、市政规章、规划文件、指南以及市政网页内容。数据集包含 3,230 份文档,总字符数约 6,033 万,总 token 数约 2,670 万(使用 Llama-3-8B-instruct 分词器计算)。其中,有 270 份文档被标记为可能的官方文件,对应约 110 万 token。文档类型以 HTML(1,841 份)和 PDF(1,389 份)为主,不包含 DOCX、CSV 或纯文本文件。数据以 Parquet 格式存储(rows.parquet),并附带爬虫脚本和配置文件。爬取工作于 2026 年 8 月 15 日完成,爬虫使用了公开的站点地图和链接,并遵循 robots.txt 规则。数据集的列字段包括来源网站名称、主页、最终 URL、请求 URL、文档标题、提取文本、内容类型、检索时间、文本哈希值、字符数、版权状态、版权依据以及额外的来源链接等。版权方面,部分法律、决策、公告等公共文件可能根据法罗群岛版权法第 9 条不受版权保护,但新闻、报告、公众提交和外部附件可能仍受版权保护。数据集存在已知限制:爬取可能未覆盖所有页面或历史文档;文本未经人工审核或语言过滤;菜单、页眉等重复网站文本可能保留;PDF 中的文本顺序可能错乱;扫描 PDF 和图像中的文本缺失;部分页面可能包含个人信息。该数据集适用于法罗语自然语言处理、市政文档分析、公共信息检索等任务。
The dataset named "Tórshavnar kommuna" (identifier: torshavn_fo) is an independent web-collected dataset and is not part of the Faroese Dynaword project. The data is sourced from the official website of the Tórshavnar municipality in the Faroe Islands (torshavn.fo), covering the municipalitys public records including council meeting minutes, budgets, annual accounts, municipal regulations, planning documents, guidelines, and municipal web page content. The dataset contains 3,230 documents, with approximately 60.33 million characters and about 26.7 million tokens (calculated using the Llama-3-8B-instruct tokenizer). Among them, 270 documents are marked as possible official documents, corresponding to about 1.1 million tokens. Document types are mainly HTML (1,841 documents) and PDF (1,389 documents), without DOCX, CSV, or plain text files. The data is stored in Parquet format (rows.parquet), accompanied by crawler scripts and configuration files. The crawling was completed on August 15, 2026, using public sitemaps and links, and following the robots.txt rules. The dataset columns include source website name, homepage, final URL, request URL, document title, extracted text, content type, retrieval time, text hash, character count, copyright status, copyright basis, and additional source links. Regarding copyright, some public documents such as laws, decisions, and announcements may not be protected by copyright under Article 9 of the Faroe Islands Copyright Act, but news, reports, public submissions, and external attachments may still be protected. Known limitations: the crawl may not cover all pages or historical documents; the text has not been manually reviewed or language-filtered; duplicate website text such as menus and headers may be retained; text order in PDFs may be scrambled; text in scanned PDFs and images is missing; some pages may contain personal information. The dataset is suitable for Faroese natural language processing, municipal document analysis, public information retrieval, and other tasks.
数据集概述
该数据集(torshavn_fo)是一个关于法罗群岛托尔斯港市(Tórshavnar kommuna) 的独立网站文本集合,不属于 Faroese Dynaword 项目的一部分。
数据内容
数据来源为托尔斯港市官方网站(torshavn.fo),收集了包括:
- 市政会议记录
- 预算文件
- 年度账目
- 市政规章
- 规划文件
- 指南
- 市政网页
数据规模
| 指标 | 数值 |
|---|---|
| 文档总数 | 3,230 |
| 字符数 | 60,332,912 |
| 词元数 | 26,695,185 |
| 可能的官方文档数 | 270 |
| 官方文档词元数 | 1,099,125 |
文件类型分布:
| 文件类型 | 文档数 |
|---|---|
| HTML | 1,841 |
| 1,389 | |
| DOCX | 0 |
| CSV | 0 |
| 纯文本 | 0 |
词元统计基于 AI-Sweden-Models/Llama-3-8B-instruct tokenizer,涵盖所有收集的文本(包括其他语言和重复的网页文本)。
文件与爬取信息
- 数据文件:
rows.parquet(包含收集的文档) - 爬虫代码:
parser/文件夹(含爬虫、站点配置和 Python 依赖) - 爬取时间:2026-08-15 开始并完成
爬虫使用公开的 sitemap 和网站链接,仅访问 parser/sites.csv 中列出的网站和文件托管位置,并遵循 robots.txt。文本提取自网页、文本型 PDF、Word 文件和 CSV 文件;扫描版 PDF 和图片未进行 OCR 识别,过短的页面和完全重复的文本被排除。
数据列说明
| 列名 | 含义 |
|---|---|
source_site, source_homepage |
网站名称和主页 |
source_url, requested_url |
最终 URL 和最初请求的 URL |
title, text, content_type |
文档标题、提取的文本和文件类型 |
retrieved_at |
文档下载时间 |
text_sha256, text_characters |
文本哈希值和字符数 |
rights_status, rights_basis |
自动版权标签及其依据 |
extraction_source_url, source_page_url |
额外的来源链接(如有) |
版权说明
- 部分法律、决定、公告等公共文件可能依据《法罗群岛版权法》第 9 条(https://www.logir.fo/Logtingslog/30-fra-30-04-2015-um-upphavsraett)不受版权保护。
likely_official_document仅表示标题或 URL 与关键词列表匹配,并非法律分类。- 市政会议记录、市政规章、批准的预算和最终账目是最明确的候选文件;新闻、报告、公众提交材料和外部附件可能仍受版权保护。
已知限制
- 爬取可能未覆盖所有页面或旧文档
- 文本未经人工审核或语言过滤
- 菜单、页眉和其他重复网页文本可能保留
- PDF 文本可能出现顺序错误
- 扫描版 PDF 和图片中的文本缺失
- 部分页面可能包含个人信息
使用方式
爬虫代码包含在 parser 文件夹中,sites.csv 仅包含此网站。运行命令:
bash cd parser python -m pip install -r requirements.txt python scrape.py --site torshavn_fo




