Financial Reports Numerical Extraction (FINE)
收藏资源简介:
FINE数据集是一个专门为财务报告信息提取任务设计的数据集,由微软、北京大学等机构联合创建。该数据集主要包含从SEC的EDGAR系统中提取的财务关键绩效指标(KPI),旨在支持大语言模型(LLMs)在混合长文档(HLDs)中的信息提取研究。数据集中的文档平均包含约59,464个Tokens,涵盖了广泛的财务数据,适用于金融分析领域。FINE数据集的创建过程涉及从公开的财务报告中提取关键数值信息,并通过自动化框架进行处理。该数据集的应用领域主要集中在金融信息提取,旨在解决LLMs在处理混合文本和表格数据时的信息提取难题。
The FINE dataset is a specialized dataset designed for financial report information extraction tasks, jointly created by institutions including Microsoft and Peking University. This dataset mainly contains financial key performance indicators (KPIs) extracted from the SEC EDGAR system, aiming to support research on information extraction by Large Language Models (LLMs) in hybrid long documents (HLDs). Documents in this dataset contain an average of approximately 59,464 Tokens, cover a wide range of financial data, and are applicable to the field of financial analysis. The creation process of the FINE dataset involves extracting key numerical information from public financial reports and processing it via automated frameworks. The application scenarios of this dataset mainly focus on financial information extraction, aiming to address the information extraction challenges faced by LLMs when processing hybrid text and tabular data.

- 1Extract Information from Hybrid Long Documents Leveraging LLMs: A Framework and Dataset微软、北京大学、中国科学院软件研究所、蚂蚁集团 · 2024年



