ACPBench Hard
收藏资源简介:
ACPBench Hard数据集是基于ACPBench构建的,由IBM Research创建。该数据集包含7种不同类型的推理任务,旨在将复杂的计划生成任务分解为独立的原子推理任务,以布尔问题或选择题的形式出现。ACPBench Hard是这些任务的生成版本,要求模型回答开放性问题。数据集适用于评估大型语言模型在自动规划器中作为组件的可靠性,涵盖多种规划领域,以帮助构建更高效的规划模型。
The ACPBench Hard dataset is built upon ACPBench and created by IBM Research. This dataset encompasses seven distinct types of reasoning tasks, which are designed to decompose complex plan generation tasks into independent atomic reasoning tasks, presented as either Boolean questions or multiple-choice questions. ACPBench Hard is the generative variant of these tasks, requiring models to answer open-ended questions. The dataset is intended to evaluate the reliability of Large Language Models (LLMs) as components in automated planners, covering multiple planning domains to help construct more efficient planning models.




