EnConda-Bench
收藏资源简介:
EnConda-Bench是一个针对软件工程中环境配置任务的基准数据集。它旨在评估基于大型语言模型(LLM)的智能体在环境配置过程中的规划、感知、反馈和执行能力。数据集包含经过筛选的高质量GitHub仓库,通过在README文件中注入常见错误来模拟真实的环境配置问题。评估方法包括错误诊断和可执行性测试,以评估智能体在环境配置过程中的每一步的能力。该数据集的创建旨在为软件工程领域的研究人员提供大规模、高质量的训练数据,以推动环境配置任务中智能体能力的提升。
EnConda-Bench is a benchmark dataset for environment configuration tasks in software engineering. It aims to evaluate the planning, perception, feedback and execution capabilities of LLM-based AI Agents during the environment configuration process. The dataset includes curated high-quality GitHub repositories, where common errors are injected into their README files to simulate realistic environment configuration issues. The evaluation methods cover error diagnosis and executability testing, to assess the capabilities of AI Agents at each step of the environment configuration process. This dataset is developed to provide large-scale, high-quality training data for researchers in the field of software engineering, so as to promote the enhancement of AI Agents' capabilities in environment configuration tasks.

- 1通过腾讯优图实验室 · 2025年



