MOMENTS (Multimodal Mental States)
收藏资源简介:
MOMENTS(多模态心理状态)是一个全面的基准测试,旨在通过现实、叙事丰富的场景来评估多模态大型语言模型(LLM)的ToM能力。数据集包括超过2344个多选题,涵盖了七个不同的ToM类别。基准测试具有长的视频上下文窗口和现实的社会互动,为深入了解角色的心理状态提供了更深入的见解。虽然视觉模态通常可以提高模型性能,但当前系统仍然难以有效地整合它,这突出了对AI在多模态理解人类行为方面的进一步研究的需求。
MOMENTS (Multimodal Mental States) is a comprehensive benchmark designed to evaluate the Theory of Mind (ToM) capabilities of multimodal large language models (LLMs) through realistic, narrative-rich scenarios. The dataset includes over 2,344 multiple-choice questions spanning seven distinct ToM categories. This benchmark features long video context windows and realistic social interactions, providing in-depth insights into the mental states of characters. While visual modalities typically improve model performance, current systems still struggle to effectively integrate them, which underscores the need for further research into AI's multimodal understanding of human behavior.

- 1MOMENTS: A Comprehensive Multimodal Benchmark for Theory of MindMBZUAI, University of Houston, McGill University, University of Michigan · 2025年



