WORFBENCH
收藏资源简介:
WORFBENCH是由浙江大学和阿里巴巴集团共同创建的一个统一的工作流生成基准数据集,旨在评估大型语言模型(LLMs)在复杂任务分解中的能力。该数据集涵盖了四个复杂场景,包括问题解决、函数调用、具身规划和开放式规划,包含18k训练样本和2146测试样本。数据集通过严格的质控和数据过滤,确保了工作流图的复杂性和准确性。WORFBENCH的应用领域主要集中在提升LLMs在下游任务中的表现,通过生成高效的工作流图,减少推理时间并提高任务完成效率。
WORFBENCH is a unified workflow generation benchmark dataset co-created by Zhejiang University and Alibaba Group, aiming to evaluate the capabilities of large language models (LLMs) in complex task decomposition. This dataset covers four complex scenarios, including problem solving, function calling, embodied planning and open-ended planning, and contains 18k training samples and 2146 test samples. Strict quality control and data filtering are applied to ensure the complexity and accuracy of workflow graphs. The main application fields of WORFBENCH focus on enhancing the performance of LLMs in downstream tasks, by generating efficient workflow graphs to reduce inference time and improve task completion efficiency.

- 1Benchmarking Agentic Workflow Generation浙江大学 · 2024年



