artefactory/ledger-long-context-multi-kpi
收藏资源简介:
该数据集是LEDGER长上下文多KPI提取数据集和基准,用于金融信息提取的评估。它将OCR提取的年度报告文本(来自DeepSeek OCR)与结构化KPI真实值配对,旨在评估基于大型语言模型的金融信息提取、检索和大海捞针任务。数据集包含两个配置:no_eval(用于训练/开发,包含4,505份报告、725家公司、104,529行KPI数据,覆盖2009-2024年)和eval(用于基准评估,包含494份报告、111家公司、13,519行KPI数据,覆盖2017-2022年)。每行数据包括股票代码、交易所、公司名称、行业、年份、31个KPI列(如收入、净利润、总资产、总负债等)以及完整的OCR文本(Markdown格式,包含页面分割)。KPI值以百万为单位(按报告原值,无货币转换),NaN表示该报告/年份的KPI不可用。数据来源包括:OCR文本来自SEC EDGAR、LSE、ASX等交易所的年度报告PDF(通过DeepSeek OCR处理),KPI值来自SEC EDGAR(XBRL companyfacts)用于美国上市公��、yfinance用于非美国公司、Alpha Vantage用于补充缺失数据。
This dataset is the LEDGER long-context multi-KPI extraction dataset and benchmark for financial information extraction evaluation. It pairs annual report texts extracted via OCR (from DeepSeek OCR) with structured ground-truth KPI values, aiming to evaluate large language model-based financial information extraction, retrieval and needle-in-a-haystack tasks. The dataset includes two configurations: no_eval (for training/development, containing 4,505 reports, 725 companies, 104,529 KPI rows, covering the period 2009–2024) and eval (for benchmark evaluation, containing 494 reports, 111 companies, 13,519 KPI rows, covering the period 2017–2022). Each row contains stock ticker, exchange, company name, industry, year, 31 KPI columns (such as revenue, net profit, total assets, total liabilities, etc.), and complete OCR text in Markdown format with page segmentation. KPI values are denominated in millions, retained as originally reported with no currency conversion, and NaN indicates that the KPI is unavailable for the corresponding report and year. Data sources are as follows: OCR texts are derived from annual report PDFs of exchanges including SEC EDGAR, LSE, ASX, etc., processed via DeepSeek OCR; KPI values are sourced from SEC EDGAR (XBRL companyfacts) for US-listed companies, yfinance for non-US companies, and Alpha Vantage for supplementing missing data.




