hshshhh/SWE-bench_Lite
收藏资源简介:
SWE-bench *Lite*是[SWE-bench](https://huggingface.co/datasets/princeton-nlp/SWE-bench)的一个子集,用于测试系统自动解决GitHub问题的能力。该数据集收集了来自11个流行Python项目的300个测试Issue-Pull Request对。评估是通过单元测试验证进行的,使用PR后的行为作为参考解决方案。数据集作为[SWE-bench: Can Language Models Resolve Real-World GitHub Issues?](https://arxiv.org/abs/2310.06770)的一部分发布。
SWE-bench *Lite* is _subset_ of [SWE-bench](https://huggingface.co/datasets/princeton-nlp/SWE-bench), a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects 300 test Issue-Pull Request pairs from 11 popular Python. Evaluation is performed by unit test verification using post-PR behavior as the reference solution. The dataset was released as part of [SWE-bench: Can Language Models Resolve Real-World GitHub Issues?](https://arxiv.org/abs/2310.06770)



