MMIE
收藏资源简介:
MMIE数据集是由北卡罗来纳大学教堂山分校和微软研究院等机构共同创建的大规模多模态交错理解评估基准。该数据集包含20,103条精心策划的多模态查询,涵盖数学、物理、编程、文学等12个领域和102个子领域。数据集的创建过程包括从四个多模态数据集中提取和重组数据,并通过多步骤的质量控制确保数据集的完整性和一致性。MMIE数据集主要用于评估大型视觉语言模型在交错多模态理解和生成任务中的表现,旨在解决现有基准在数据规模、范围和评估深度方面的不足。
The MMIE dataset is a large-scale multimodal interleaved understanding evaluation benchmark co-created by institutions including the University of North Carolina at Chapel Hill, Microsoft Research, and others. This dataset contains 20,103 carefully curated multimodal queries, covering 12 domains such as mathematics, physics, programming, and literature, as well as 102 subfields. The dataset's creation process involves extracting and reorganizing data from four multimodal datasets, and ensuring the integrity and consistency of the dataset through multi-step quality control. The MMIE dataset is mainly used to evaluate the performance of large vision-language models on interleaved multimodal understanding and generation tasks, aiming to address the shortcomings of existing benchmarks in terms of data scale, scope, and evaluation depth.

- 1MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models北卡罗来纳大学教堂山分校 · 2024年



