Logical Reasoning Benchmark Dataset
收藏资源简介:
This record contains a synthetic logical reasoning benchmark dataset for evaluating model performance across symbolic, controlled natural language, and natural language representations. Each example consists of a premise, a conclusion, and a binary entailment label indicating whether the conclusion logically follows from the premise. The dataset is organized into seven difficulty levels of increasing compositional complexity. Each difficulty level includes matched examples across three representation formats: symbolic logic notation, controlled natural language, and fluent natural language. The dataset includes standard test splits as well as out-of-distribution evaluation splits designed to test generalization beyond surface-level memorization. Lexical OOD splits introduce unseen vocabulary while preserving the underlying logical structure, and template OOD splits introduce unseen sentence or expression patterns. The release includes the dataset files, dataset documentation, generation methodology, reproducibility instructions, citation metadata, license information, and checksums. The dataset is intended for research on logical reasoning, compositional generalization, representation effects, out-of-distribution robustness, and transformer-based reasoning. The dataset is released under the Creative Commons Attribution 4.0 International license (CC BY 4.0).



