CFMM
收藏资源简介:
CFMM数据集,由复旦大学智能信息处理上海市重点实验室创建,旨在评估多模态大型语言模型在反事实推理任务上的能力。该数据集包含1200张来自COCO-2014验证集的图像,每张图像配有一个基本问题和一个反事实问题,用于测试模型是否能正确理解反事实前提。数据集涵盖计数、颜色、大小、形状、方向和常识等六个评估维度,通过精心的人工标注确保数据质量。CFMM数据集的应用领域主要集中在提升多模态语言模型在复杂推理任务上的表现,特别是在处理视觉和语言结合的反事实问题上的能力。
The CFMM dataset, created by the Shanghai Key Laboratory of Intelligent Information Processing at Fudan University, is designed to evaluate the capabilities of multimodal large language models on counterfactual reasoning tasks. This dataset includes 1200 images sourced from the COCO-2014 validation set, with each image paired with a baseline question and a counterfactual question to test whether the model can correctly understand counterfactual premises. The dataset covers six evaluation dimensions: counting, color, size, shape, orientation and common sense, and its data quality is ensured through meticulous manual annotation. The main application scenarios of the CFMM dataset focus on improving the performance of multimodal language models on complex reasoning tasks, particularly their ability to handle vision-language combined counterfactual problems.

- 1Eyes Can Deceive: Benchmarking Counterfactual Reasoning Abilities of Multi-modal Large Language Models复旦大学智能信息处理上海市重点实验室 · 2024年



