MULTIFINBEN
收藏资源简介:
MULTIFINBEN是一个面向全球金融领域的多语言、多模态、难度感知的评估基准,旨在评估大型语言模型(LLM)在文本、视觉和音频等不同模态以及单语、双语和多语等不同语言环境下的性能。该数据集包含34个不同的数据集,涵盖英语、中文、日语、西班牙语和希腊语五种语言,并提供了一个动态的、难度感知的选择机制,以保持评估基准的紧凑性和平衡性。MULTIFINBEN旨在推动金融研究和应用领域的透明、可重复和包容性进展。
MULTIFINBEN is a multilingual, multimodal, difficulty-aware evaluation benchmark targeting the global financial domain. It aims to evaluate the performance of Large Language Models (LLMs) across diverse modalities including text, vision and audio, as well as varying linguistic scenarios such as monolingual, bilingual and multilingual settings. This benchmark encompasses 34 distinct datasets covering five languages: English, Chinese, Japanese, Spanish and Greek, and features a dynamic, difficulty-aware selection mechanism to maintain the compactness and balance of the evaluation benchmark. MULTIFINBEN is designed to advance transparent, reproducible and inclusive progress in the field of financial research and applications.

- 1MultiFinBen: A Multilingual, Multimodal, and Difficulty-Aware Benchmark for Financial LLM EvaluationThe FinAI · 2025年



