遇见数据集

From Natural-Language Specifications to Functional Agents: Human-in-the-Loop Synthesis and Evaluation

收藏
Zenodo2025-11-04 更新2026-05-26 收录
官方服务:

资源简介:

A benchmark to evaluate AI frameworks such as Gemini CLI, Spec-Kit, Claude Code, BMAD, Copilot CLI, Kiro and Yuma-Code for generating GitHub issue management agent. This project compares AI tools on their ability to generate functional GitIssueBot agents from standardized specifications. It provides benchmark data, generated outputs, and evaluation results.

提供机构:
Zenodo
创建时间:
2025-11-04
二维码
社区交流群
二维码
科研交流群
商业服务