SWE-MERA
收藏资源简介:
SWE-MERA是一个动态更新的基准测试数据集,旨在评估大型语言模型在软件工程任务上的性能。它通过自动化收集GitHub上的真实世界问题和严格的验证流程来确保数据质量。数据集目前包含大约300个样本,并有望扩展到10,000个任务。SWE-MERA的设计旨在解决现有软件工程基准测试数据集的局限性,如数据泄露和基准测试饱和问题。它通过定期更新数据集,确保任务与软件开发的最新挑战保持相关性,并为模型提供公平的评价环境。
SWE-MERA is a dynamically updated benchmark dataset designed to evaluate the performance of Large Language Models (LLMs) on software engineering tasks. It ensures data quality by automatically collecting real-world issues from GitHub and implementing rigorous validation procedures. Currently, the dataset contains approximately 300 samples, with the potential to scale up to 10,000 tasks. SWE-MERA is designed to address the limitations of existing software engineering benchmark datasets, such as data leakage and benchmark saturation. It ensures that tasks remain relevant to the latest challenges in software development by regularly updating the dataset, while also providing a fair evaluation environment for models.




