Multi-Image Relational Benchmark (MIRB)
收藏资源简介:
Multi-Image Relational Benchmark (MIRB) 是由爱丁堡大学和同济大学共同创建的数据集,旨在评估视觉语言模型在多图像理解方面的能力。该数据集包含925个样本,涵盖感知、视觉世界知识、推理和多跳推理四个维度,每个样本至少需要处理2张图像。MIRB的创建过程涉及从多个来源独立收集图像,并设计了一系列需要跨图像比较和分析的任务。该数据集的应用领域广泛,包括但不限于机器人视觉、医学图像分析和在线购物比较,旨在推动多模态模型在处理复杂视觉场景中的发展。
Multi-Image Relational Benchmark (MIRB) is a dataset jointly created by the University of Edinburgh and Tongji University, which aims to evaluate the capabilities of vision-language models in multi-image understanding. This dataset includes 925 samples covering four dimensions: perception, visual world knowledge, reasoning, and multi-hop reasoning, and each sample requires processing at least two images. The development of MIRB involves independently collecting images from multiple sources and designing a series of tasks that require cross-image comparison and analysis. This dataset has a wide range of application scenarios, including but not limited to robotic vision, medical image analysis and online shopping comparison, and is designed to promote the development of multimodal models in handling complex visual scenarios.

- 1Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning爱丁堡大学 同济大学 · 2024年



