IPO-Dataset
收藏资源简介:
IPO-Dataset是一个大规模、多模态、按章节结构化的数据集,由佐治亚理工学院等机构创建,专注于美国证券交易委员会(SEC)的首次公开募股(IPO)申报文件。该数据集涵盖1994年至2026年间的109,690份IPO申报文件及修订案,包含超过76,000张图像,文本部分通过解析目录实现章节对齐,确保了结构一致性。数据集构建依托IPO-Toolkit工具包,实现了从EDGAR平台的文件下载、文本解析到图像提取的自动化流程,并经过LLM辅助验证和人工审核以保证质量。该数据集旨在支持对长文档、多模态金融文本的分析,应用于评估误导性图表、研究跨行业披露实践以及推动多模态模型在真实世界监管文档中的推理能力。
IPO-Dataset is a large-scale, multimodal, chapter-structured dataset developed by institutions including the Georgia Institute of Technology, focusing on Initial Public Offering (IPO) registration filings submitted to the U.S. Securities and Exchange Commission (SEC). This dataset covers 109,690 IPO registration filings and their amendments spanning 1994 to 2026, containing over 76,000 images. The textual portions are chapter-aligned via table-of-contents parsing, ensuring structural consistency. The dataset is constructed using the IPO-Toolkit, which enables an automated workflow from file downloading on the EDGAR platform, text parsing to image extraction. Quality assurance is implemented via LLM-aided validation and manual review. This dataset is intended to support analyses of long-document multimodal financial texts, with applications including evaluating misleading charts, researching cross-industry disclosure practices, and advancing the reasoning capabilities of multimodal models in real-world regulatory documents.

- 1IPO-Mine: A Toolkit and Dataset for Section-Structured Analysis of Long, Multimodal IPO Documents佐治亚理工学院; 赛大学; 杜克大学 · 2026年



