CoMMET
收藏资源简介:
CoMMET是由新加坡科技研究局牵头构建的多模态心理理论评估基准数据集,包含591个故事化交互单元(StoryTurns),涵盖欲望、信念、道德推理等7类心理状态。该数据集基于心理学经典ToM手册任务扩展而成,通过1973个问答对和826张配图构建多轮对话场景,采用Gemini 3.0 Pro生成初始数据并经人工校验。作为首个支持多轮交互式评估的基准,其核心价值在于全面测试大语言模型在真实社交场景中的心智推理能力,推动可信人机交互系统发展。
CoMMET is a multimodal Theory of Mind (ToM) evaluation benchmark dataset developed under the leadership of the Agency for Science, Technology and Research (A*STAR), Singapore. It consists of 591 story-based interaction units (StoryTurns), covering 7 categories of mental states including desire, belief, moral reasoning and others. Derived from classic psychological ToM manual tasks, this dataset constructs multi-turn dialogue scenarios through 1973 question-answer pairs and 826 paired images. The initial data was generated using Gemini 3.0 Pro and subsequently manually verified. As the first benchmark supporting multi-turn interactive evaluation, its core value lies in comprehensively testing the mental reasoning abilities of large language models (LLMs) in real-world social scenarios, and promoting the development of trustworthy human-computer interaction systems.

- 1CoMMET: To What Extent Can LLMs Perform Theory of Mind Tasks?新加坡科技研究局·高性能计算研究所; 新加坡科技研究局·前沿人工智能研究中心; 南洋理工大学; 香港科技大学·广州 · 2026年



