MovieCORE
收藏资源简介:
MovieCORE是一个视频问答(VQA)数据集,旨在探索电影内容的更深层次理解。与现有的主要关注表面理解的数据集不同,MovieCORE强调的问题能够激发系统2思考,同时保持与视频内容的紧密关联。我们提出了一个创新的代理式头脑风暴方法,利用多个大型语言模型(LLMs)作为思考代理来生成和改进高质量的问答对。为了评估数据集的质量,我们开发了一套认知测试,以评估深度、思考激发潜力和句法复杂性。我们还提出了一套全面的评估方案,以评估VQA模型在更深入的认知任务上的性能。为了解决现有视频语言模型(VLMs)的局限性,我们引入了一个代理增强模块,代理选择增强(ACE),该模块通过25%的改进,提高了模型推理能力。我们的工作有助于推进AI系统中对电影的理解,并为我们提供了当前VQA模型在面对更具挑战性的电影内容时能力和局限性的宝贵见解。
MovieCORE is a video question answering (VQA) dataset developed to explore deeper understanding of movie content. Unlike existing datasets that primarily focus on surface-level comprehension, MovieCORE centers on questions that elicit System 2 thinking while remaining closely tied to the video content. We propose an innovative agent-guided brainstorming approach that utilizes multiple large language models (LLMs) as thinking agents to generate and refine high-quality question-answer pairs. To assess the dataset's quality, we have developed a suite of cognitive tests to evaluate its depth, thinking-elicitation potential, and syntactic complexity. We also present a comprehensive evaluation framework to measure the performance of VQA models on more in-depth cognitive tasks. To address the limitations of current video-language models (VLMs), we introduce an agent-augmented module called Agent Choice Enhancement (ACE), which achieves a 25% improvement in model reasoning performance. Our work contributes to advancing movie understanding in AI systems, and provides valuable insights into the capabilities and limitations of existing VQA models when confronted with more challenging movie content.
MovieCORE: COgnitive REasoning in Movies
概述
MovieCORE是一个新颖的视频问答(VQA)数据集,专门设计用于探究对电影内容的深层认知理解。与现有专注于表层理解的数据集不同,MovieCORE强调激发思考的问题,涉及系统2思维,同时保持与视频材料的具体关联。
关键特性
- 认知深度:数据集优先考虑系统2思维,导致问答对具有更高的深度。
- 语义丰富性:相比其他数据集,MovieCORE在语义丰富性和深度方面表现突出。
- 高质量标注:采用多智能体头脑风暴方法,利用多个大型语言模型(LLMs)作为思维智能体生成和精炼高质量的问答对。
标注方法
- 智能体标注工作流:批评家智能体作为主持人,利用视频上下文和任务指令协调专业智能体之间的交互。依次与系统II VQA专家、怀疑研究者、侦探和元评审员互动,在每个阶段积累见解。
- 人工验证:精炼后的VQA子集由人类专家评估进行最终验证。
评估与比较
- 评估方案:提出综合评估方案,用于评估VQA模型在更深层认知任务上的性能。
- 性能比较:评估各种开源和专有视觉语言模型(VLMs)在五个标准上的表现:准确性、全面性、深度、证据和连贯性。
资源
- 论文:https://arxiv.org/abs/2508.19026
- 代码:公开可用
- 数据集:代理标注系统、数据集及其元数据将公开提供
作者与机构
- Gueter Josmy Faure(国立台湾大学)
- Min-Hung Chen(NVIDIA)
- Jia-Fong Yeh(国立台湾大学)
- Ying Cheng(国立清华大学)
- Hung-Ting Su(国立台湾大学)
- Yung-Hao Tang(国立政治大学)
- Winston H. Hsu(国立台湾大学、Mobile Drive Technology)
- Shang-Hong Lai(国立清华大学)
会议
- EMNLP 2025主会议




