ComposerGCoT
收藏资源简介:
ComposerGCoT是由巴黎-萨克雷大学研究团队构建的专门用于评估多模态大语言模型推理一致性的数据集。该数据集包含16.3万条高质量推理链,覆盖5.5万张图像,通过蒸馏GQA数据集并基于规则化流程生成,每条数据均标注了中间答案和视觉定位信息。其创建过程采用两阶段策略:首先通过CC3M、Flickr30K等数据集进行空间对齐预训练,然后利用GQA的场景图和功能程序分解为离散的视觉推理步骤。该数据集主要应用于视觉问答领域,旨在解决模型推理过程与视觉证据脱节的问题,通过结构化评估框架验证模型在答案准确性、推理一致性和视觉定位精度三个维度的综合性能。
ComposerGCoT is a dataset constructed by the research team from Université Paris-Saclay, specifically designed to evaluate the reasoning consistency of multimodal large language models. This dataset contains 163,000 high-quality reasoning chains covering 55,000 images, which is generated by distilling the GQA dataset and following a regularized workflow. Each data entry is annotated with intermediate answers and visual localization information. Its creation adopts a two-stage strategy: first, spatial alignment pre-training is performed using datasets such as CC3M and Flickr30K; then, the scene graphs and functional programs of GQA are decomposed into discrete visual reasoning steps. This dataset is mainly applied in the field of visual question answering (VQA), aiming to solve the problem that model reasoning processes are disconnected from visual evidence, and verify the comprehensive performance of models across three dimensions: answer accuracy, reasoning consistency and visual localization accuracy through a structured evaluation framework.




