foia-reading-room-documents
收藏资源简介:
该数据集包含来自美国联邦机构FOIA阅览室和监察长报告库的文档,包括审计、检查、调查摘要以及根据《信息自由法》(FOIA)发布的记录。所有文档均为美国政府工作成果,属于公共领域,且文件内容与机构发布的原始版本字节一致,未做任何修改。数据集旨在解决机构网站上文档频繁移动、重新编号或删除的问题,特别是SAM.gov附件在通知归档后容易消失,保存副本可使文档可引用,并支持批量处理。数据集的目录结构为:documents/<source>/<doc_id>.pdf,以及一个元数据文件metadata.parquet。元数据文件每行对应一个文档,包含来源、机构、标题、发布日期、来源URL、链接页面、字节大小、页数和文件的SHA-256校验和。数据来源是通过遵循各站点robots.txt规则,从机构FOIA图书馆和监察长报告列表收集而来,使用了govdocs工具。数据集规模在1千到1百万个文档之间,语言为英语,适用于文本检索等任务。
This dataset contains documents from U.S. federal agency FOIA reading rooms and Inspector General report libraries, including audits, inspections, investigation summaries, and records released under the Freedom of Information Act (FOIA). All documents are works of the U.S. government and are in the public domain, with file contents byte-identical to the original versions released by the agencies, without any modifications. The dataset is designed to address the frequent moving, renumbering, or deletion of documents on agency websites, especially SAM.gov attachments that tend to disappear after notice archiving. Saving copies makes documents citable and supports batch processing. The dataset directory structure is: documents/<source>/<doc_id>.pdf, along with a metadata file metadata.parquet. Each row in the metadata file corresponds to a document and includes source, agency, title, publication date, source URL, link page, byte size, number of pages, and SHA-256 checksum. The data was collected by following each sites robots.txt rules, from agency FOIA libraries and IG report listings, using the govdocs tool. The dataset size ranges from 1 thousand to 1 million documents, the language is English, and it is suitable for tasks such as text retrieval.
Foia Reading Room Documents 数据集详情
该数据集收录了来自美国联邦机构FOIA(信息自由法)阅览室及监察长办公室报告库的公开文件,涵盖审计报告、检查报告、调查摘要及依据《信息自由法》发布的记录。所有文件均由美国联邦机构发布,属于美国政府作品,内容未经任何修改,与机构原始发布文件逐字节一致,metadata.parquet 中记录了原始字节的校验值。
数据来源与动机
- 来源:数据收集自各机构FOIA资料库及监察长办公室的报告列表,遵循各网站 robots.txt 协议,通过 govdocs 工具采集。
- 用途:由于机构常会移动、重新编号或删除其公开文件(例如SAM.gov附件会在公告存档后消失),本项目旨在保留副本,使文件可被引用,并支持批量处理语料库,免于逐份PDF下载。
数据目录结构
documents/<source>/<doc_id>.pdf:存放PDF原始文件。metadata.parquet:每份文档对应一行元数据,包括:来源、机构、标题、发布日期、原始URL、来源页面链接、字节大小、页数及文件的SHA-256哈希。
元数据字段
| 字段 | 说明 |
|---|---|
| source | 文件来源机构或库 |
| agency | 所属联邦机构 |
| title | 文档标题 |
| posted date | 发布日期 |
| URL | 原始下载地址 |
| page | 来源页面链接 |
| byte size | 文件大小(字节) |
| page count | 页数 |
| SHA-256 | 文件校验值 |
数据集特征
- 语言:英语(en)
- 任务类型:文本检索(text-retrieval)
- 许可:美国政府作品许可(us-government-work),详见 USA.gov政府作品说明
- 文件规模:1,000 至 1,000,000 个文件之间(具体总数未标注)





