Realtime Retrospective Board: AI Model Benchmark Dataset and Evaluation Artifacts
收藏官方服务:
资源简介:
Dataset, scoring rubric, and implementation artifacts from an observational study of 72 agentic code-generation runs implementing a standardized real-time retrospective board application, evaluated across multiple language-model variants, scaffold configurations, effort modes, and tool-access conditions. Each run is scored against a 14-criterion functional rubric on a 3/2/1 scale (42-point maximum); per-run scores are recorded in the EVALUATION_RUBRIC.md files. See CHANGELOG.md for version history and methodology notes.
提供机构:
Zenodo创建时间:
2026-06-22



