PureToMDatasets
收藏资源简介:
该数据集是一个专门用于评估心理理论(Theory of Mind, ToM)能力的多配置基准测试集合,包含七个独立的子数据集配置:BigToM、EmoBench、FanToM、HiToM、SimpleToM、SocialIQA和ToMBench。所有配置仅提供测试分割,总计约25,000个评估样本。每个数据样本采用统一的结构化格式,包含四个核心字段:故事叙述(story)、相关问题(question)、答案选项(answer)和元数据(meta)。其中answer字段进一步细分为正确答案列表(correct_answers)和错误答案列表(wrong_answers),支持多项选择或开放生成式评估。meta字段包含丰富的标注信息,如样本ID、问题类型、能力维度、语言类别、难度级别等细粒度属性,具体字段因配置而异。数据集覆盖心理理论的多个评估维度,包括情感理解、社交推理、信念推断等认知能力,适用于大型语言模型和心理理论计算模型的系统性评估与基准测试。数据集采用Apache 2.0开源许可证。
This dataset is a multi-configuration benchmark collection specifically designed to evaluate Theory of Mind (ToM) capabilities, comprising seven independent sub-dataset configurations: BigToM, EmoBench, FanToM, HiToM, SimpleToM, SocialIQA, and ToMBench. All configurations provide only test splits, totaling approximately 25,000 evaluation samples. Each data sample follows a unified structured format with four core fields: story narrative (story), related question (question), answer options (answer), and metadata (meta). The answer field is further subdivided into correct answers list (correct_answers) and wrong answers list (wrong_answers), supporting multiple-choice or open-ended generative evaluation. The meta field contains rich annotation information, such as sample ID, question type, capability dimension, language category, difficulty level, and other fine-grained attributes, with specific fields varying by configuration. The dataset covers multiple evaluation dimensions of Theory of Mind, including emotional understanding, social reasoning, belief inference, and other cognitive abilities, making it suitable for systematic evaluation and benchmarking of large language models and computational models of Theory of Mind. The dataset is licensed under Apache 2.0.




