PRDBench
收藏资源简介:
PRDBench是一个包含50个真实世界Python项目的基准数据集,涵盖20个领域。每个项目都包含结构化的产品需求文档(PRD)、全面的评估标准和参考实现。该数据集具有丰富的数据来源、高任务复杂性和灵活的评估指标。数据集创建过程利用先进的代码代理生成项目框架和评估标准,人工标注者只需验证标准与项目接口的一致性以及预期输出是否符合PRD要求。PRDBench旨在解决现有代码代理评估基准数据集标注成本高、专家要求高以及评估指标僵化的问题,为代码代理和评估代理的能力评估提供了一个可扩展且健壮的框架。
PRDBench is a benchmark dataset comprising 50 real-world Python projects spanning 20 domains. Each project includes a structured Product Requirements Document (PRD), comprehensive evaluation criteria, and reference implementations. This dataset features diverse data sources, high task complexity, and flexible evaluation metrics. The dataset creation process leverages advanced code agents to generate project frameworks and evaluation criteria, with human annotators only required to verify the consistency between the criteria and project interfaces, as well as whether the expected outputs align with PRD requirements. PRDBench aims to address the problems of high annotation costs, high expert requirements and rigid evaluation metrics in existing code agent evaluation benchmark datasets, providing a scalable and robust framework for evaluating the capabilities of both code agents and evaluation agents.




