ProfBench
收藏资源简介:
ProfBench引入了超过3000个专家撰写的响应-标准对,涵盖商业和科学研究四个专业领域的40个任务 - 物理博士、化学博士、金融MBA和咨询MBA - 能够评估超越考试式问答或仅代码/数学设置的开放式、基于文档的专业任务。即使前沿模型也发现ProfBench具有挑战性:最佳报告生成器GPT-5-high仅达到65.9%的整体得分,突显了在需要综合和长形式分析的现实专业工作流程中存在显著改进空间。
ProfBench introduces over 3,000 expert-written response-reference pairs, covering 40 tasks across four professional domains in business and scientific research: Doctoral-level Physics, Doctoral-level Chemistry, Financial MBA, and Consulting MBA. This benchmark enables evaluation of open-ended, document-grounded professional tasks that transcend exam-style question answering or code/math-only settings. Even state-of-the-art models find ProfBench challenging: the top-performing reported generator GPT-5-high achieved only an overall score of 65.9%, highlighting substantial room for improvement in real-world professional workflows requiring synthesis and long-form analysis.
ProfBench 数据集概述
数据集简介
ProfBench 是一个包含3000多个专家撰写的响应-标准对的数据集,涵盖商业和科学研究四个专业领域的40个任务,具体包括物理学博士、化学博士、金融MBA和咨询MBA领域。
核心特点
- 支持开放式、基于文档的专业任务评估
- 超越考试式问答或仅限代码/数学的设置
- 即使前沿模型也认为具有挑战性:最佳报告生成器GPT-5-high仅达到65.9%的总体得分
- 强调在需要综合和长篇分析的现实专业工作流程中存在显著提升空间
评估方法
- 提出稳健、经济的LLM-Judge评估方法
- 结合Macro-F1测量和偏差指数来减轻自我增强偏差
- 实现跨提供者偏差低于1%
- 相比之前的基准测试,评估成本降低2-3个数量级
数据获取
数据集可通过以下地址获取:https://huggingface.co/datasets/nvidia/ProfBench
引用信息
bibtex @misc{wang2025profbenchmultidomainrubricsrequiring, title={ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge}, author={Zhilin Wang and Jaehun Jung and Ximing Lu and Shizhe Diao and Ellie Evans and Jiaqi Zeng and Pavlo Molchanov and Yejin Choi and Jan Kautz and Yi Dong}, year={2025}, eprint={2510.18941}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2510.18941}, }




