TruthQuest
收藏资源简介:
TruthQuest是由慕尼黑大学信息与语言处理中心创建的一个用于评估大型语言模型假设推理能力的基准数据集。该数据集基于经典的骑士与恶棍逻辑谜题,包含2400个不同复杂度的问题,涉及不同数量的人物和逻辑陈述类型。数据集的创建过程涉及将谜题形式化为二值逻辑,并确保每个问题有唯一解。TruthQuest的应用领域主要在于测试和提升语言模型在复杂逻辑推理任务中的表现,特别是在处理可能为假的陈述时的逻辑推断能力。
TruthQuest is a benchmark dataset developed by the Center for Information and Language Processing at Ludwig-Maximilians-Universität München (LMU Munich) for evaluating the hypothetical reasoning capabilities of large language models. This dataset is based on classic knight-and-knave logic puzzles, and contains 2400 questions with varying complexity, involving different numbers of characters and types of logical statements. The dataset's creation process involves formalizing the puzzles into two-valued logic, and ensuring that each question has a unique solution. The primary application scenarios of TruthQuest are to test and improve the performance of language models in complex logical reasoning tasks, especially their logical inference abilities when handling potentially false statements.

- 1Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models慕尼黑大学信息与语言处理中心 · 2024年



