Blocksworld Benchmark
收藏资源简介:
Blocksworld Benchmark是由亚利桑那州立大学计算与人工智能学院的研究人员开发的一个测试大型语言模型规划能力的数据集。该数据集包含500个Blocksworld领域的实例,旨在评估LLMs在常识规划任务中的自主生成和验证简单计划的能力。数据集通过GitHub公开,支持研究社区对LLMs规划能力的进一步探索。Blocksworld是一个简单的常识领域,涉及堆叠和移动积木,目标是根据给定的初始状态和目标状态,生成一系列动作以达到目标状态。
The Blocksworld Benchmark is a dataset developed by researchers from the School of Computing and Artificial Intelligence at Arizona State University, designed to test the planning capabilities of large language models. This dataset contains 500 instances from the Blocksworld domain, aiming to evaluate LLMs' ability to autonomously generate and validate simple plans in commonsense planning tasks. The dataset is publicly available via GitHub, supporting the research community to further explore the planning capabilities of LLMs. Blocksworld is a simple commonsense domain involving stacking and moving blocks, where the goal is to generate a sequence of actions to reach the target state given the initial state and target state.

- 1On the Planning Abilities of Large Language Models (A Critical Investigation with a Proposed Benchmark)亚利桑那州立大学计算与人工智能学院 · 2023年



