ProcBench
收藏资源简介:
ProcBench是由Araya 2AI Alignment Network创建的一个专门用于评估大型语言模型(LLMs)多步骤推理和遵循指令能力的基准数据集。该数据集包含23种不同类型的任务,共计5520个示例,任务设计简单但需要精确遵循多步骤指令。数据集的创建旨在通过提供明确的指令和问题,使模型能够通过遵循指令来解决问题,从而评估模型在复杂任务中的表现。ProcBench的应用领域包括推理、可解释AI、减少幻觉和AI对齐,旨在解决模型在严格遵循指令和多步骤推理中的局限性。
ProcBench is a benchmark dataset developed by Araya 2AI Alignment Network, specifically designed to evaluate the multi-step reasoning and instruction-following capabilities of Large Language Models (LLMs). This dataset includes 5,520 examples across 23 distinct task types, with tasks that are simple in design yet require precise adherence to multi-step instructions. The dataset was created to assess models' performance on complex tasks by providing explicit instructions and questions, enabling models to solve problems by following these instructions. The application areas of ProcBench cover reasoning, interpretable AI, hallucination reduction, and AI alignment, aiming to address the limitations of models in strictly following instructions and performing multi-step reasoning.
ProcBench 数据集
许可证
- 许可证类型: CC BY 4.0




