遇见数据集

GridPuzzle

收藏
arXiv2025-09-30 收录
官方服务:

资源简介:

该数据集是一个评估集,包含274个基于网格的谜题,这些谜题具有不同的复杂度,旨在评估大型语言模型(LLMs)的推理能力。该数据集涵盖了多种网格大小(3x4、3x5、4x4、4x5和4x6),并设置了不同的难度级别,旨在深入了解大型语言模型在解决网格谜题时可能出现的推理错误。规模上,该数据集共有274个基于网格的谜题,其任务是对解决网格谜题时的推理链进行评估。

This dataset is an evaluation set containing 274 grid-based puzzles with varying complexity levels, which is specifically designed to evaluate the reasoning capabilities of Large Language Models (LLMs). It encompasses multiple grid size configurations including 3x4, 3x5, 4x4, 4x5 and 4x6, and features distinct difficulty tiers, aiming to provide in-depth insights into the reasoning errors that LLMs may encounter when solving grid-based puzzles. With a total of 274 grid-based puzzles, the core task of this evaluation set is to assess the reasoning chains generated during the process of solving such grid puzzles.

提供机构:
Mihir3009
搜集汇总
数据集介绍
GridPuzzle 数据集图片
背景与挑战
背景概述
GridPuzzle是一个包含274个网格谜题的数据集,旨在评估大型语言模型(LLMs)的逻辑推理能力,通过分析模型生成的推理链来识别错误。数据集提供了多模型响应(如GPT-4、Claude-3等)、黄金标准答案和细粒度评估指标(如PuzzleEval),支持对推理过程的深入分析和比较。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务