From Natural-Language Specifications to Functional Agents: Human-in-the-Loop Synthesis and Evaluation
收藏官方服务:
资源简介:
A benchmark to evaluate AI frameworks such as Gemini CLI, Spec-Kit, Claude Code, BMAD, Copilot CLI, Kiro and Yuma-Code for generating GitHub issue management agent. This project compares AI tools on their ability to generate functional GitIssueBot agents from standardized specifications. It provides benchmark data, generated outputs, and evaluation results.
提供机构:
Zenodo创建时间:
2025-11-04



