BoardgameQA
收藏资源简介:
BoardgameQA是由谷歌研究院开发的一个用于评估语言模型在处理矛盾信息时推理能力的数据集。该数据集模拟了现实世界中常见的信息不一致或矛盾的情况,要求模型根据信息源的偏好(如可信度或信息时效性)来解决冲突。数据集中的每个示例包含一个可废止理论(一组输入事实、可能矛盾的规则以及对规则的偏好)和一个关于该理论的问题。回答这些问题需要多跳推理和冲突解决。此外,BoardgameQA还包含了需要模型自身提供部分背景知识的场景,以更好地反映下游应用中的推理问题。该数据集旨在揭示当前语言模型在处理矛盾和信息不完整情况下的推理能力差距,并为未来研究提供基准。
BoardgameQA is a dataset developed by Google Research to evaluate the reasoning capabilities of language models when handling contradictory information. This dataset simulates common information inconsistency and contradiction scenarios in real-world settings, requiring models to resolve conflicts based on the preferences of information sources such as their credibility or timeliness. Each example in the dataset consists of a defeasible theory (a collection of input facts, potentially contradictory rules, and preferences assigned to these rules) and a question related to this theory. Answering these questions demands multi-hop reasoning and conflict resolution capabilities. Furthermore, BoardgameQA also incorporates scenarios where models need to provide partial background knowledge independently, to better mirror reasoning challenges in downstream applications. This dataset aims to uncover the reasoning capability gaps of current language models when dealing with contradictions and incomplete information, and serve as a benchmark for future research.

- 1BoardgameQA: A Dataset for Natural Language Reasoning with Contradictory Information谷歌研究院 · 2023年



