Ryoo72/MMBench-EN-Dev-V11
收藏资源简介:
MMBench-EN-Dev-V1.1是一个多模态视觉问答数据集,用于评估视觉语言模型。该数据集是MMBench基准测试的英文开发集版本1.1,按原始L2类别字段预分割为6个子集,以便于在数据集查看器中浏览。它包含4,876行数据,每行包含问题、提示、选项A-D、答案、类别(L3和L2级别)、图像(以PIL格式)和分割信息。数据源自OpenCompass的TSV文件,并已解析图像引用,确保每行都携带解码后的图像。子集包括粗粒度感知、细粒度感知(单实例和跨实例)、属性推理、关系推理和逻辑推理,覆盖多模态模型的各种能力评估。
MMBench-EN-Dev-V1.1 is a multimodal visual question answering dataset used for evaluating vision-language models. It is the English development set version 1.1 of the MMBench benchmark, pre-split into 6 subsets by the original L2-category field for convenient browsing in the dataset viewer. The dataset contains 4,876 rows, each including question, hint, options A-D, answer, category (L3 and L2 levels), image (in PIL format), and split. It is sourced from the OpenCompass TSV file, with image references resolved so that every row carries its own decoded image. Subsets include coarse perception, fine-grained perception (single-instance and cross-instance), attribute reasoning, relation reasoning, and logic reasoning, covering various capabilities for multimodal model evaluation.



