遇见数据集

Realtime Retrospective Board: AI Model Benchmark Dataset and Evaluation Artifacts

收藏
Zenodo2026-06-26 更新2026-06-28 收录
官方服务:

资源简介:

Dataset, scoring rubric, and implementation artifacts from an observational study of 72 agentic code-generation runs implementing a standardized real-time retrospective board application, evaluated across multiple language-model variants, scaffold configurations, effort modes, and tool-access conditions. Each run is scored against a 14-criterion functional rubric on a 3/2/1 scale (42-point maximum); per-run scores are recorded in the EVALUATION_RUBRIC.md files. See CHANGELOG.md for version history and methodology notes.

提供机构:
Zenodo
创建时间:
2026-06-22
二维码
社区交流群
二维码
科研交流群
商业服务