M³-Bench
收藏资源简介:
M³-Bench是由中山大学等机构联合构建的大规模多模态评测基准,包含15,894个样本,旨在零样本条件下系统评估预训练模型的感知与推理能力。该数据集通过整合权威基准和新增专项数据,覆盖自然/文化概念识别、空间/数学/物理推理及跨学科视觉问答等7类核心任务。数据来源于现有通用与领域数据集的重构及针对性补充,采用双策略构建方法以弥合评估缺口。其核心应用是诊断多模态预训练模型的能力瓶颈,为感知-推理能力不对称发展现象研究提供量化工具。
M³-Bench is a large-scale multimodal evaluation benchmark jointly developed by Sun Yat-sen University and other institutions, consisting of 15,894 samples. It is designed to systematically evaluate the perception and reasoning capabilities of pre-trained models under zero-shot settings. This benchmark integrates authoritative existing benchmarks and newly added specialized datasets, covering seven core tasks including natural and cultural concept recognition, spatial, mathematical, physical reasoning, and cross-disciplinary visual question answering. The dataset is constructed through the reconstruction of existing general and domain-specific datasets and targeted supplementary data collection, employing a dual-strategy construction approach to bridge existing evaluation gaps. Its core applications include diagnosing capability bottlenecks of multimodal pre-trained models and providing quantitative tools for research on the asymmetric development of perception and reasoning abilities.



