遇见数据集

sfd-anonymous/sfd-v1

收藏
Hugging Face2026-05-07 更新2026-05-31 收录
官方服务:

资源简介:

SFD-v1是一个开放的数据集,基于美国证券交易委员会(SEC)EDGAR文件,重建为布局保真的MultiMarkdown(MMD)格式,旨在支持长上下文语言建模、金融推理、文档理解和评估任务。该数据集覆盖了2022年1月至2025年6月期间的约380万份文件,包含约1520亿个Qwen3-1.7B tokens。它处理多种源格式(如HTML、XML、纯文本等),并保留合并单元格表格、缩进和视觉层次结构,以提供高效的令牌表示。数据集包括多种文件类型(如10-K、10-Q、8-K等),并附带元数据(如SEC编号、年份、月份等),每个文件都是自包含的。

SFD-v1 is an open, layout-faithful reconstruction of U.S. Securities and Exchange Commission (SEC) EDGAR filings into token-efficient MultiMarkdown (MMD), targeted at long-context language modeling, financial reasoning, document understanding, and evaluation. This release covers filings from January 2022 through June 2025, with approximately 3.8 million filings and around 152 billion Qwen3-1.7B tokens. It handles multiple source formats (e.g., HTML, XML, plaintext) and preserves merged-cell tables, indentation, and visual hierarchy for efficient token representation. The dataset includes various filing types (such as 10-K, 10-Q, 8-K) and is accompanied by metadata (e.g., accession number, year, month), making each filing self-contained.

提供机构:
sfd-anonymous
二维码
社区交流群
二维码
科研交流群
商业服务