遇见数据集

Ryoo72/MMBench-EN-Dev-V10

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

MMBench EN Dev V1.0 是一个多模态视觉问答评估数据集,专门用于测试和基准化多模态模型在多种感知和推理任务上的性能。该数据集基于原始MMBench开发集的英文版本(MMBench_DEV_EN.tsv),并按照L2类别字段预分割为6个子集,以方便在数据集查看器中浏览。数据集包含4,329个样本,每个样本包括问题、提示、四个选项(A、B、C、D)、答案、类别标签(L3细粒度类别,共20类)、图像(以PIL格式存储)、L2类别标签(6类)、分割信息、来源和注释。子集涵盖粗粒度感知、细粒度感知(单实例和跨实例)、属性推理、关系推理和逻辑推理等任务。数据格式经过解析,确保每行都包含其自身的解码图像,避免了原始TSV文件中的引用问题。该数据集适用于视觉语言模型的评估和比较,支持通过Hugging Face的datasets库加载使用。

The MMBench-EN-Dev-V1.0 dev split, pre-split into 6 subsets by the original l2-category field for convenient browsing in the dataset viewer. It is a multimodal visual question answering evaluation dataset designed for benchmarking multimodal models across various perception and reasoning tasks. The dataset is derived from the official OpenCompass TSV file (MMBench_DEV_EN.tsv) and contains 4,329 rows. Each row includes question, hint, options (A, B, C, D), answer, category (L3, 20 fine-grained classes), image (in PIL format), l2-category (L2, 6 classes), split, source, and comment. The subsets cover coarse perception, fine-grained perception (single-instance and cross-instance), attribute reasoning, relation reasoning, and logic reasoning. The format resolves references from the original compressed CircularEval format, ensuring each row carries its own decoded image. It is intended for evaluating and comparing vision-language models and can be loaded via the Hugging Face datasets library.

提供机构:
Ryoo72
二维码
社区交流群
二维码
科研交流群
商业服务