MMEVALPRO
收藏资源简介:
MMEVALPRO数据集由北京大学多媒体信息处理国家重点实验室创建,旨在通过严格的评估流程提高多模态模型评估的可信度和效率。该数据集包含2138个问题三元组,总计6414个独立问题,其中三分之二由人类专家手动标注,其余来自现有基准(MMMU、ScienceQA和MathVista)。数据集的创建过程包括精心设计的标注流程和严格的质检步骤,确保数据质量。MMEVALPRO主要应用于多模态模型的评估,旨在解决现有基准中存在的系统偏差问题,提高评估的准确性和可信度。
The MMEVALPRO dataset was developed by the State Key Laboratory of Multimedia and Information Processing at Peking University, aiming to improve the credibility and efficiency of multimodal model evaluation through a rigorous evaluation workflow. This dataset includes 2138 question triplets, totaling 6414 independent questions, among which two-thirds are manually annotated by human experts, and the remaining third is sourced from existing benchmarks including MMMU, ScienceQA, and MathVista. The dataset creation process covers a meticulously designed annotation procedure and strict quality inspection steps to ensure data quality. Mainly applied for multimodal model evaluation, MMEVALPRO targets addressing systemic biases in current benchmarks and enhancing the accuracy and credibility of evaluations.
数据集概述
数据集名称
- MMEvalPro
描述
- Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation
关键词
- MMEvalPro

- 1MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation北京大学多媒体信息处理国家重点实验室 · 2024年



