Massive Image Embedding Benchmark (MIEB)
收藏资源简介:
MIEB是由Durham University等机构创建的大型图像嵌入基准数据集,包含38种语言的130个任务,分为8个高级类别。数据集内容涵盖了从聚类到视觉问答等多种任务,要求模型在图像和文本嵌入方面有广泛的能力。创建过程中,特别关注了需要强视觉理解文本的任务,如视觉STS和文档理解。MIEB的应用领域广泛,旨在推动自然融合的图像文本嵌入模型的发展。
MIEB is a large-scale image embedding benchmark dataset created by institutions including Durham University and other relevant organizations. It includes 130 tasks across 38 languages, which are classified into 8 high-level categories. The dataset covers a wide range of tasks spanning from clustering to visual question answering, requiring models to have comprehensive capabilities in both image and text embedding. During its development, particular attention was paid to tasks that demand strong visual-text understanding abilities, such as visual STS and document understanding. MIEB has broad application fields and aims to promote the development of naturally fused image-text embedding models.

- 1MMTEB: Massive Multilingual Text Embedding Benchmark奥尔胡斯大学 · 2025年



