Hard_ToMDatasets
收藏资源简介:
该数据集是一个多配置的评估基准集合,专注于心理理论(Theory of Mind)与社会推理能力的测评。数据集包含八个子配置:BigToM、EmoBench、ExploreToM、FanToM、HiToM、SimpleToM、SocialIQA 和 ToMBench,总计提供超过20,000个样本。每个样本均围绕叙事性场景构建,核心数据结构包括:故事背景(Story)、涉及人类状态(如信念、偏好、情绪)与环境状态(如位置、物体、变化)的详细描述(State)、角色行动(Action)、基于故事提出的问题(Question),以及包含正确答案和错误答案选项的答案对(Answer)。此外,每个样本附有丰富的元数据(Meta),涵盖样本ID、数据来源、评估维度(dimension)、任务类型(task_type)、难度等级(difficulty)、伦理类别(ethics_category)等属性,支持多维度能力评估。部分子集(如BigToM、EmoBench、SimpleToM、SocialIQA)还提供了由GPT-5.5生成的合成数据(synthetic_gpt_5_5),可用于数据增强或对比研究。该数据集适用于训练和评估人工智能模型在理解复杂社会情境、推断他人心理状态、进行因果推理以及回答基于叙事的问答任务等方面的能力。
This dataset is a multi-configuration evaluation benchmark collection focused on assessing Theory of Mind and social reasoning capabilities. It includes eight sub-configurations: BigToM, EmoBench, ExploreToM, FanToM, HiToM, SimpleToM, SocialIQA, and ToMBench, providing over 20,000 samples in total. Each sample is built around a narrative scenario, with core data structures comprising: story background (Story), detailed descriptions of human states (e.g., beliefs, preferences, emotions) and environmental states (e.g., locations, objects, changes) (State), character actions (Action), questions based on the story (Question), and answer pairs (Answer) containing correct and incorrect options. Additionally, each sample is accompanied by rich metadata (Meta), covering attributes such as sample ID, data source, evaluation dimension, task type, difficulty level, ethics category, etc., supporting multi-dimensional ability assessment. Some subsets (e.g., BigToM, EmoBench, SimpleToM, SocialIQA) also provide synthetic data generated by GPT-5.5 (synthetic_gpt_5_5), which can be used for data augmentation or comparative research. This dataset is suitable for training and evaluating artificial intelligence models in areas such as understanding complex social situations, inferring others mental states, conducting causal reasoning, and answering narrative-based question-answering tasks.




