LogicBench
收藏arXiv2025-09-30 收录
数据链接:
官方服务:
资源简介:
该数据集是一个自然语言问答集合,专注于在多种逻辑推理模式中运用单一推理规则。它涵盖了25种不同的推理模式,包括命题逻辑、一阶逻辑和非单调逻辑。该数据集的任务是对大型语言模型进行逻辑推理评估。
This dataset is a natural language question-answering collection that focuses on applying a single reasoning rule across diverse logical reasoning patterns. It encompasses 25 distinct reasoning patterns, including propositional logic, first-order logic, and non-monotonic logic. The task of this dataset is to evaluate the logical reasoning capabilities of large language models.
提供机构:
Mihir3009搜集汇总
背景与挑战
背景概述
LogicBench是一个专注于评估大型语言模型逻辑推理能力的自然语言问答数据集,涵盖25种推理模式,包括命题逻辑、一阶逻辑和非单调逻辑。它提供评估和训练两个版本,以JSON格式结构化存储,旨在系统测试模型在复杂推理和否定场景下的表现。
以上内容由遇见数据集搜集并总结生成



