遇见数据集

anonymous-md/EDGAR_FILINGS_DATASET

收藏
Hugging Face2026-05-07 更新2026-05-31 收录
官方服务:

资源简介:

SFD-v1 是一个开源、布局忠实的美国证券交易委员会(SEC)EDGAR文件重建数据集,采用令牌高效的MultiMarkdown(MMD)格式,旨在用于长上下文语言建模、金融推理、文档理解和评估。该版本覆盖了2022年1月至2025年6月期间的文件(约340万份),通过SFD解析器生成。数据集处理多种源格式,包括HTML、XML、纯文本、SGML和PDF(通过OCR),并保留了合并单元格表格、缩进和视觉层次结构。文件类型包括10-K、10-Q、8-K、Form 4、N-PORT等超过350种。数据集以Parquet分片存储,使用zstd-15压缩,并包含每个文件的元数据,如CIK、登录号、表单类型、提交日期等。许可证为CC-BY-NC-4.0(非商业用途),原始SEC文件为美国政府公共领域。

SFD-v1 is an open, layout-faithful reconstruction of U.S. Securities and Exchange Commission (SEC) EDGAR filings into token-efficient MultiMarkdown (MMD), targeted at long-context language modeling, financial reasoning, document understanding, and evaluation. This release covers filings from January 2022 through June 2025 (~3.4M filings), produced by the SFD parser. The dataset handles multiple source formats, including HTML, XML, plaintext, SGML, and PDF-via-OCR, and preserves merged-cell tables, indentation, and visual hierarchy. It includes over 350 filing types such as 10-K, 10-Q, 8-K, Form 4, N-PORT, etc. The data is stored in Parquet shards with zstd-15 compression and includes metadata like CIK, accession, form type, filing date, etc. The license is CC-BY-NC-4.0 for the parsed corpus, while the underlying SEC filings are in the U.S. Government public domain.

提供机构:
anonymous-md
二维码
社区交流群
二维码
科研交流群
商业服务