DRBench
收藏资源简介:
DRBench是一个用于评估在复杂、开放式深度研究任务中AI代理性能的基准。与专注于简单问题或仅网络查询的先前基准不同,DRBench评估代理在多步查询方面的能力,例如,“我们应该对我们的产品路线图做出哪些改变以确保符合这个标准?”,这需要从公共网络和私有公司知识库中识别支持事实。每个任务都以现实用户的角色和公司环境为依据,跨越包括生产力软件、云文件系统、电子邮件、聊天对话和开放网络在内的异构搜索空间。任务是通过精心设计的合成管道生成的,并经过人类在环验证,代理的评估标准包括其回忆相关见解的能力、保持事实准确性和产生连贯、结构良好的报告的能力。DRBench发布了15个深度研究任务,涵盖10个领域,如销售、网络安全和合规性。通过评估各种DR代理和DR策略,展示了DRBench的有效性,这些代理和策略包括开源和闭源模型(如GPT、Llama和Qwen)。
DRBench is a benchmark for evaluating the performance of AI Agents in complex, open-ended deep research tasks. Unlike prior benchmarks that focus on simple questions or solely on web queries, DRBench evaluates the capabilities of AI Agents in handling multi-step queries—for example, "What changes should we make to our product roadmap to ensure compliance with this standard?"—which requires identifying supporting facts from both the public web and private corporate knowledge bases. Each task is grounded in realistic user roles and corporate environments, spanning heterogeneous search spaces including productivity software, cloud file systems, email, chat conversations, and the open web. The tasks are generated via carefully designed synthetic pipelines and validated via human-in-the-loop validation. The evaluation criteria for AI Agents include their ability to recall relevant insights, maintain factual accuracy, and produce coherent, well-structured reports. DRBench includes 15 deep research tasks spanning 10 domains such as sales, cybersecurity, and compliance. Experiments evaluating various DR Agents and DR strategies—including both open-source and closed-source models such as GPT, Llama, and Qwen—have demonstrated the effectiveness of DRBench.
- 1通过ServiceNow Research, University of British Columbia, Mila – Quebec AI Institute, McGill University, Canada CIFAR AI Chair · 2025年



