遇见数据集

Verification-Surface Dataset: 1,116 Controlled Web-Application Builds by Six LLMs under Eight Tool Configurations

收藏
Zenodo2026-08-16 更新2026-08-20 收录
官方服务:

资源简介:

Companion dataset to the article "The reach of a verification tool decides its value: A controlled study of verification surface, artifact quality, and cost in AI coding agents" (A. Mehta, under review, 2026). 1,116 web applications built by six frontier language models (claude-4.6-sonnet, claude-4.6-opus, claude-4.8-opus, gpt-5.5, gemini-3.1-pro, grok-4.3) under eight verification-tool configurations, from no checking tools to full shell plus screenshots, with byte-identical prompts proven by cryptographic fingerprints in every run manifest. Contents: every artifact as shipped, complete per-run tool-call logs, machine grades from adversarial behavioral probes, condition-blind human rubric scores for all 1,116 runs, hand-audited launch-failure classifications, three SHA256-sealed dataset snapshots (freeze-20260720 is the dataset of record), the pre-specified statistical analysis plan with its committed code and complete output, and the grading instrument needed to re-grade every artifact. See README.md for reproduction recipes; the analysis table and all statistics regenerate from the published inputs. The complete per-run request-response traces (1,116 trace.jsonl files, one per run) are published in this record as the supplementary archive traces-batch-20260613.tar.gz (sha256 2c47fa647a7e55218a0ffd42dda04e0bc13a216fd1561235dfab3ad1f6f5c4df). Unpacking it from the root of a repository clone restores the per-run paths that verify_paper_claims.py --with-traces verifies. Code is MIT-licensed; data and documentation are CC BY 4.0 (see LICENSE).

提供机构:
Zenodo
创建时间:
2026-08-16
二维码
社区交流群
二维码
科研交流群
商业服务