ToolComp
收藏资源简介:
ToolComp是由Scale AI开发的一个多工具推理和过程监督基准数据集,旨在评估语言模型在多步骤工具使用任务中的表现。该数据集包含485条经过人工验证的提示和最终答案,以及1731个详细的步骤监督标签,涵盖了从日期查询到金融助手等多种工具的使用场景。数据集的创建过程结合了模型生成和人工标注,确保了数据的准确性和复杂性。ToolComp的应用领域主要集中在复杂多步骤推理任务的评估和模型训练,旨在通过过程监督提升模型的推理能力,解决现有基准在评估工具使用能力时的不足。
ToolComp is a multi-tool reasoning and process supervision benchmark dataset developed by Scale AI, which aims to evaluate the performance of language models on multi-step tool-use tasks. This dataset contains 485 manually verified prompts and final answers, as well as 1731 detailed step-wise supervision labels, covering various tool usage scenarios ranging from date queries to financial assistant applications. The dataset was created by combining model generation and manual annotation to ensure its accuracy and complexity. The main application fields of ToolComp focus on the evaluation and model training of complex multi-step reasoning tasks, aiming to improve the reasoning ability of models through process supervision and address the shortcomings of existing benchmarks in evaluating tool-use capabilities.




