遇见数据集

ZebraLogicBench

收藏
AI_Studio2025-08-28 更新2025-05-17 收录
官方服务:

资源简介:

​ZebraLogic 是一个专为评估大型语言模型(LLMs)逻辑推理能力而设计的基准测试集,聚焦于逻辑网格谜题(如著名的“斑马谜题”)。​该评估集由 Allen Institute for AI 开发,旨在系统地测试模型在处理复杂约束满足问题(CSPs)时的表现。

ZebraLogic is a benchmark dataset specifically designed to evaluate the logical reasoning capabilities of large language models (LLMs), focusing on logic grid puzzles such as the famous "Zebra Puzzle". Developed by the Allen Institute for AI, this dataset aims to systematically test models' performance when handling complex constraint satisfaction problems (CSPs).

提供机构:
皮都坦率的法夏
二维码
社区交流群
二维码
科研交流群
商业服务