itsPrerna202/OctoCodingBench
收藏资源简介:
OctoCodingBench是一个用于评估仓库基础编码代理在遵循脚手架指令方面的基准测试。它不仅关注代理是否能够正确完成任务,还强调代理在实现过程中是否遵守了各种约束和规则。数据集包含72个精选实例,涵盖任务规范、系统提示、评估清单、Docker镜像和脚手架配置。它测试了代理在7种不同指令来源下的合规性,包括系统提示、系统提醒、用户查询、项目级约束、技能、记忆和工具模式。数据集支持多脚手架(如Claude Code、Kilo、Droid),并提供了详细的评估指标和用法说明。
OctoCodingBench is a benchmark for evaluating scaffold-aware instruction following in repository-grounded agentic coding. It focuses not only on whether the agent can complete tasks correctly but also on whether the agent adheres to various constraints and rules during implementation. The dataset contains 72 curated instances, including task specifications, system prompts, evaluation checklists, Docker images, and scaffold configs. It tests agent compliance across 7 heterogeneous instruction sources: system prompt, system reminder, user query, project-level constraints, skill, memory, and tool schema. The dataset supports multiple scaffolds (e.g., Claude Code, Kilo, Droid) and provides detailed evaluation metrics and usage instructions.




