ZebraLogicBench
收藏官方服务:
资源简介:
ZebraLogic 是一个专为评估大型语言模型(LLMs)逻辑推理能力而设计的基准测试集,聚焦于逻辑网格谜题(如著名的“斑马谜题”)。该评估集由 Allen Institute for AI 开发,旨在系统地测试模型在处理复杂约束满足问题(CSPs)时的表现。
ZebraLogic is a benchmark dataset specifically designed to evaluate the logical reasoning capabilities of large language models (LLMs), focusing on logic grid puzzles such as the famous "Zebra Puzzle". Developed by the Allen Institute for AI, this dataset aims to systematically test models' performance when handling complex constraint satisfaction problems (CSPs).
提供机构:
皮都坦率的法夏


