遇见数据集

Warrant study: tasks, harness and run records for an exploratory evaluation of evidence-gated acceptance of code written by an AI agent

收藏
Zenodo2026-09-30 更新2026-10-01 收录
官方服务:

资源简介:

Materials for an exploratory evaluation of Warrant (https://doi.org/10.5281/zenodo.23029803). Twenty small Python command-line tasks, each with plain-language promises, visible checks, hidden checks, an audit suite, a reference implementation and deliberately broken implementations; the validity gates; the harness; Part 1, a seeded-defect evaluation of 160 broken implementations under visible-only acceptance and under Warrant's gate; and Part 2, 40 runs of Claude Code driving gpt-oss:20b under the two acceptance rules, with snapshots, transcripts, session logs, Warrant ledgers and the analysis. The evaluation is exploratory and not pre-registered. The tasks, checks and implementations were written by AI agents at the author's request and reviewed by an AI model, not by an independent person. Development was AI-assisted, and the commits carry Co-Authored-By trailers.

提供机构:
Zenodo
创建时间:
2026-09-30
二维码
社区交流群
二维码
科研交流群
商业服务