遇见数据集

AgentProcessBench: A Dataset of LLM-Agent Software Projects Across Process Models and GPT Variants

收藏
Zenodo2025-09-16 更新2026-05-26 收录
官方服务:

资源简介:

This dataset captures the outcomes of automated software development projects executed by LLM-based agents using the MetaGPT framework. Each project was generated by a multi-agent system simulating real-world software roles—such as Designer, Developer, and Tester—coordinated through one of three classical software process models: Waterfall, V-Model, and Agile. A total of 11 software projects (in Python and JavaScript) were executed across multiple GPT model configurations (e.g., GPT-4o-mini, DeepSeek-Chat, Claude). For each run, the dataset records rich metrics covering codebase statistics (e.g., number of files, lines of code, lines of comments), resource consumption (e.g., input/output tokens, total token usage, execution time, and cost), and software quality indicators (e.g., code smells detected via SonarQube, manually run test cases and failures, AI-generated test cases and failures). This dataset supports benchmarking and analysis of coordination strategies, agent performance, and software process impacts in LLM-based generative software engineering.

提供机构:
Zenodo
创建时间:
2025-09-16
二维码
社区交流群
二维码
科研交流群
商业服务