MVL-SIB
收藏资源简介:
MVL-SIB数据集是由德国维尔茨堡大学人工智能与数据科学中心和德国汉堡大学语言技术组创建的,包含205种语言的图像-文本跨模态主题匹配任务。该数据集扩展了SIB-200的粗粒度主题标注,通过手动收集的代表每个主题的10个图像和4个同类别句子,创建了3个不同的MVL-SIB实例。这些任务旨在评估大型视觉语言模型在跨模态和仅文本主题匹配方面的表现,数据集支持对语言理解和多模态推理的消融研究,以及单图像和多图像视觉语言交互的细致分析。
The MVL-SIB dataset was developed by the Center for Artificial Intelligence and Data Science at the University of Würzburg, Germany, and the Language Technology Group at the University of Hamburg, Germany. It covers image-text cross-modal topic matching tasks across 205 languages. Building upon the coarse-grained topic annotations of SIB-200, the dataset constructs three distinct MVL-SIB instances by manually collecting 10 images and 4 category-consistent sentences representing each topic. These tasks are designed to evaluate the performance of large vision-language models in both cross-modal and text-only topic matching scenarios. Additionally, the dataset enables ablation studies on language understanding and multimodal reasoning, as well as fine-grained analyses of single-image and multi-image vision-language interactions.

- 1MVL-SIB: A Massively Multilingual Vision-Language Benchmark for Cross-Modal Topical Matching德国维尔茨堡大学人工智能与数据科学中心, 德国汉堡大学语言技术组 · 2025年



