MULTIFINBEN
收藏资源简介:
MULTIFINBEN是一个面向全球金融领域的多语言和多模态基准,旨在评估大型语言模型在文本、视觉和音频等多种模态以及单语、双语和多语等语言环境下的能力。该数据集包含了34个不同的数据集,涵盖了英语、中文、日语、西班牙语和希腊语五种语言,并针对信息提取、文本分析、问答、文本生成、风险管理、预测和决策等七个任务类别进行了分类。MULTIFINBEN的创建过程包括引入了两个新的任务(PolyFiQA-Easy和PolyFiQA-Expert)以及两个OCR嵌入的视觉-文本数据集,并通过动态、难度感知的选择机制来确保评估的平衡性和紧凑性。该数据集旨在推动金融领域研究的透明性、可重复性和包容性。
MULTIFINBEN is a multilingual and multimodal benchmark for the global financial domain, designed to evaluate the capabilities of large language models (LLMs) across diverse modalities including text, vision and audio, as well as various linguistic settings such as monolingual, bilingual and multilingual scenarios. This benchmark comprises 34 distinct datasets spanning five languages: English, Chinese, Japanese, Spanish and Greek, and is categorized into seven task categories, namely information extraction, text analysis, question answering, text generation, risk management, prediction and decision-making. The development of MULTIFINBEN includes the introduction of two novel tasks (PolyFiQA-Easy and PolyFiQA-Expert) and two OCR-embedded visual-text datasets, as well as a dynamic, difficulty-aware selection mechanism to ensure the balance and compactness of the evaluation. This benchmark aims to promote transparency, reproducibility and inclusivity in financial domain research.
数据集概述
基本信息
- 数据集名称: EnglishOCR
- 许可证: Apache License 2.0
- 语言: 英语
- 领域: 金融
- 任务类别: 图像到文本
- 规模类别: 10K<n<100K
数据集结构
- 特征:
image: 字符串类型,Base64编码的PNG图像text: 字符串类型,从PDF文件中提取的文本
- 数据分割:
train: 7961个样本,大小3816064970字节
数据集摘要
EnglishOCR数据集包含来自SEC EDGAR公司文件的监管文档图像。该数据集用于评估大型语言模型在将非结构化文档(如PDF和图像)转换为机器可读格式方面的能力,特别是在金融领域。
支持的任务
- 任务: 图像到文本
- 评估指标: ROUGE-1
数据集创建
- 数据来源: SEC EDGAR系统的公司文件
- 数据处理: 文件下载为HTML格式,转换为PDF版本,分割并转换为图像,提取文本用于匹配HTML块和校正图像
- 注释: 数据集来源于公开可用的公司文件,未进行额外的手动注释
注意事项
- 社会影响: 支持从扫描的金融文档中提取结构化信息,促进透明度和可访问性
- 偏见: 数据仅限于公司文件,可能无法代表其他金融文档类型
- 限制: 匹配过程可能引入不准确性,数据集可能缺乏多样化的布局样式
引用信息
bibtex @misc{peng2025multifinbenmultilingualmultimodaldifficultyaware, title={MultiFinBen: A Multilingual, Multimodal, and Difficulty-Aware Benchmark for Financial LLM Evaluation}, author={Xueqing Peng et al.}, year={2025}, eprint={2506.14028}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2506.14028}, }




