sec-10k-markdown-uncompressed
收藏资源简介:
SEC 10-K Full Uncompressed Markdown Filings 数据集是一个金融领域的文本数据集,专门用于文本生成和特征提取任务。该数据集包含12,361份完整的、未经压缩的美国证券交易委员会(SEC)Form 10-K年度报告,这些报告已从原始的EDGAR HTML格式转换为清晰、结构化的Markdown格式。数据覆盖了1,379家不同的公司,时间跨度从2004年至2025年,其中2014年至2025年的数据为核心部分。数据集以未压缩的原始Markdown文件(.md)形式提供,总大小约为4.75 GB。数据按照公司股票代码(如AAPL、MSFT、NVDA)组织在子目录中,每个公司文件夹内包含按年份命名的10-K文件(例如10-K_2024.md)以及一个包含公司元数据的清单文件(manifest.json)。该数据集主要适用于自然语言处理任务,如基于金融文档的文本生成、信息提取、财务分析、文档摘要和语言模型预训练或微调。数据集语言为英语,采用MIT许可证发布。
The SEC 10-K Full Uncompressed Markdown Filings dataset is a text dataset in the financial domain, specifically designed for text generation and feature extraction tasks. It contains 12,361 complete, uncompressed U.S. Securities and Exchange Commission (SEC) Form 10-K annual reports, which have been converted from the original EDGAR HTML format into clear, structured Markdown format. The data covers 1,379 distinct companies, spanning from 2004 to 2025, with the core portion from 2014 to 2025. The dataset is provided as uncompressed raw Markdown files (.md), with a total size of approximately 4.75 GB. The data is organized into subdirectories by company stock ticker (e.g., AAPL, MSFT, NVDA), with each company folder containing 10-K files named by year (e.g., 10-K_2024.md) and a manifest file (manifest.json) that includes company metadata. This dataset is primarily suitable for natural language processing tasks, such as text generation based on financial documents, information extraction, financial analysis, document summarization, and pre-training or fine-tuning of language models. The dataset language is English, and it is released under the MIT license.





