SWEE-Bench,SWA-Bench
收藏资源简介:
SWEE-Bench是一个包含数百个代码库的扩展版SWEBench,而SWA-Bench则是一个关注应用的代码库的基准。这两个数据集旨在代表真实世界的用例,考虑了许多代码库,以实现多样化的基准,并且可以频繁更新以避免污染和过拟合。SWEE-Bench关注多样性以及不太受欢迎的项目,包含了366个Python代码库;SWA-Bench关注软件应用,包含44个项目。这些数据集在代码库的年龄、创建时的受欢迎程度、关注近期问题以及代码修复的复杂性等方面与SWE-Bench存在显著差异,且对于某些模型,性能差异显著,表明在代表性基准上进行评估的重要性。
SWEE-Bench is an extended version of SWEBench that includes hundreds of code repositories, while SWA-Bench is a benchmark focused on application-oriented code bases. These two datasets are designed to represent real-world use cases, incorporating a wide range of code repositories to enable a diverse benchmark, and can be updated frequently to avoid data contamination and overfitting. SWEE-Bench focuses on diversity and less popular projects, containing 366 Python code repositories; SWA-Bench focuses on software applications and includes 44 projects. These datasets differ significantly from SWEBench in terms of code repository age, initial popularity, focus on recent issues, and complexity of code fixes, and exhibit notable performance differences for certain models, highlighting the importance of evaluating models on representative benchmarks.

- 1Automated Benchmark Generation for Repository-Level Coding TasksLogicStar AI,ETH Zurich · 2025年



