Financial Numeric Extreme Labelling (FNXL)
收藏资源简介:
FNXL数据集由印度理工学院卡拉格普尔分校和高盛数据科学与机器学习团队创建,专注于金融领域的数字极端标注。该数据集包含79,088个句子,总计142,922个数字被标注,使用2,794个标签。数据来源于美国证券交易委员会(SEC)要求的公开年度报告,这些报告使用XBRL进行标注。创建过程中,数据集排除了非美国通用会计准则(US-GAAP)标签,并进行了手动清理以去除噪声数据点。FNXL数据集主要用于自动化财务报告中的数字标注任务,旨在减少手动标注的工作量,并提高对新旧报告的标注效率。
The FNXL dataset was developed by the Indian Institute of Technology Kharagpur and the Data Science and Machine Learning Team at Goldman Sachs, focusing on digital extreme annotation in the financial domain. This dataset contains 79,088 sentences, with a total of 142,922 annotated numerical values, utilizing 2,794 distinct labels. The data is sourced from public annual reports mandated by the U.S. Securities and Exchange Commission (SEC), which are annotated using XBRL. During its development, the dataset excluded non-US Generally Accepted Accounting Principles (US-GAAP) labels, and underwent manual cleaning to remove noisy data points. The FNXL dataset is primarily designed for automated numerical annotation tasks in financial reports, aiming to reduce the workload of manual annotation and enhance annotation efficiency for both new and legacy reports.




