ExploreToM
收藏资源简介:
ExploreToM是首个允许大规模生成多样化且具有挑战性的理论思维推理数据的框架。该框架利用A*搜索在自定义的领域特定语言上操作,生成复杂的故事结构和多样化的、合理的场景,以对大型语言模型的极限进行压力测试。数据集包括针对特定模型(如Llama-3.1-70B-Instruct)生成的对抗性数据样本,以及用于训练和评估的多种故事结构。数据字段包括与问题相关的属性(qprop)、与故事相关的属性(sprop)和搜索参数(param)。
ExploreToM is the first framework that enables large-scale generation of diverse and challenging Theory of Mind (ToM) reasoning datasets. This framework leverages A* search on a custom domain-specific language to generate complex story structures and diverse, plausible scenarios, aiming to stress-test the upper limits of large language models (LLMs). The dataset includes adversarial data samples generated for specific models such as Llama-3.1-70B-Instruct, alongside a variety of story structures designed for training and evaluation. Its data fields consist of question-related attributes (qprop), story-related attributes (sprop), and search parameters (param).




