CoMT
收藏资源简介:
CoMT是一个新颖的多模态思维链基准数据集,旨在评估大型视觉-语言模型在复杂视觉操作和简洁表达中的能力。数据集包含3853个样本和14801张图像,涵盖四个类别:视觉创建、视觉删除、视觉更新和视觉选择。CoMT通过多模态输入和输出,模拟人类推理过程,旨在解决传统多模态思维链基准中视觉操作缺失和表达模糊的问题。该数据集适用于多模态推理任务,特别是在需要复杂视觉操作和清晰表达的场景中。
CoMT is a novel multimodal chain-of-thought benchmark dataset designed to evaluate the performance of large vision-language models in complex visual manipulation and concise expression. The dataset comprises 3,853 samples and 14,801 images, covering four categories: visual creation, visual deletion, visual update, and visual selection. CoMT simulates human reasoning processes through multimodal inputs and outputs, aiming to address the issues of missing visual manipulation support and ambiguous expression in traditional multimodal chain-of-thought benchmarks. This benchmark is suitable for multimodal reasoning tasks, especially in scenarios requiring complex visual manipulation and clear expression.

- 1CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models中南大学、苏州大学、哈尔滨工业大学、新加坡国立大学 · 2024年



