ToM QA
收藏资源简介:
我们提出了一个数据集来评估问答模型的推理能力。我们从发展心理学中的心智理论实验(例如 Sally-Anne 任务)中获得灵感;这些实验旨在测试儿童是否能够理解他人的信念,以及对世界不一致状态的推理——例如,当某人的信念与现实情况不同时。 数据由一组 3 种任务类型和 4 种问题类型组成,总共创建了 12 个场景。这些任务被分组到故事中,这些故事由每行开头的编号表示。
We present a dataset for evaluating the reasoning capabilities of question answering models. We draw inspiration from theory-of-mind experiments in developmental psychology, such as the Sally-Anne task; these experiments aim to test whether children can understand others' beliefs and reason about inconsistent states of the world—for example, when a person's beliefs differ from actual reality. The dataset consists of 3 task types and 4 question types, resulting in a total of 12 scenarios. These tasks are grouped into stories, which are indicated by numbers at the beginning of each line.




