SEAM
收藏资源简介:
SEAM数据集是一个用于评估视觉语言模型(VLM)跨模态推理能力的基准,它通过在四个领域(国际象棋、化学、音乐和图论)中配对语义等价的输入来控制语义等效性。数据集包含了16个任务,每个任务200个项目,总共3200个条目。SEAM数据集旨在解决当前VLMs在视觉和语言性能之间存在的系统不平衡问题,通过比较不同模态下的推理能力,为更通用的多模态模型提供评估和改进的依据。该数据集的创建旨在推动视觉语言模型的发展,并促进跨模态推理能力的提升。
The SEAM dataset is a benchmark for evaluating the cross-modal reasoning capabilities of Vision-Language Models (VLMs). It controls for semantic equivalence by pairing semantically equivalent inputs across four domains: chess, chemistry, music, and graph theory. The dataset comprises 16 tasks, with 200 items per task, totaling 3200 entries. The SEAM dataset aims to address the systematic imbalance between visual and language performance observed in current VLMs, and provides a framework for evaluating and refining more general multimodal models by comparing reasoning abilities across different modalities. This dataset is developed to promote the advancement of vision-language models and facilitate the enhancement of cross-modal reasoning capabilities.

- 1通过多伦多大学计算机科学系和酷威人工智能实验室 · 2025年



