ClassicLogic
收藏资源简介:
ClassicLogic是由杜伊斯堡-埃森大学研究团队创建的一个知识驱动型基准测试套件,旨在评估人工智能系统的组合泛化能力。该数据集包含数独、KenKen、Kakuro和Futoshiki四类经典逻辑谜题,通过程序化生成确保每个谜题具有唯一解,并附带分层知识库,将复杂解决策略明确定义为简单原子规则的组合。数据生成过程采用两阶段混合方法,首先生成基于策略的抽象模板,再实例化为可玩谜题,难度通过策略组合深度进行数学验证的校准。该数据集主要应用于神经符号人工智能和高级推理系统研究,旨在解决模型在逻辑演绎、多步规划和系统性推理方面的薄弱环节,为诊断模型在实体组合、关系组合和程序组合等泛化能力上的失败原因提供透明测试平台。
ClassicLogic is a knowledge-driven benchmark suite created by the research team at the University of Duisburg-Essen, designed to evaluate the compositional generalization capabilities of artificial intelligence systems. This dataset includes four classic logic puzzle types: Sudoku, KenKen, Kakuro, and Futoshiki. It is generated programmatically to ensure each puzzle has a unique solution, and is equipped with a hierarchical knowledge base that explicitly defines complex solving strategies as combinations of simple atomic rules. The data generation process adopts a two-stage hybrid approach: first generating strategy-based abstract templates, then instantiating them into playable puzzles, with their difficulty calibrated via mathematical validation of the depth of strategy combinations. This dataset is primarily applied to research on neuro-symbolic artificial intelligence and advanced reasoning systems, aiming to address the weaknesses of models in logical deduction, multi-step planning and systematic reasoning, and providing a transparent testbed for diagnosing the root causes of failures in model generalization capabilities such as entity composition, relation composition and program composition.

- 1ClassicLogic: A Knowledge-Driven Benchmark of Classic Puzzle Games for Evaluating Compositional Generalization杜伊斯堡-埃森大学 · 2026年



