遇见数据集

Auditing the Perceptual Self-Reflection Loop in Agentic Code Generation: A Pre-Registered, Human-Adjudicated Depth Ablation on a Fully-Local Physics Simulation Pipeline

收藏
Zenodo2026-08-12 更新2026-08-13 收录
官方服务:

资源简介:

Reproducibility deposit for Auditing the Perceptual Self-Reflection Loop in Agentic Code Generation: A Pre-Registered, Human-Adjudicated Depth Ablation on a Fully-Local Physics Simulation Pipeline (Shende & Camburn, Singapore University of Technology and Design), which extends arXiv:2602.12311.Contains the evaluated four-agent pipeline as it was executed, the depth-ablation harness, the analysis tooling, the twelve frozen prompts, a per-run record for all 1,800 analyzed runs, and the eight-frame montage images from which the human genuineness verdicts were made.The experiment is a 4 (scenario) × 3 (prompt tier) × 3 (framework depth) factorial at n = 50 per cell, run fully locally on a Gemma 4 26B checkpoint served with vLLM. Every run received a human genuineness verdict. The pre-registered primary contrast — full pipeline versus the same pipeline with the perceptual self-reflection loop disabled — pools twelve cells per depth and so rests on 600 runs per arm.Two things to know before relying on this. The pipeline is stochastic and unseeded, so individual runs reproduce distributionally rather than exactly. The human verdicts are two-rater consensus, which yields no human–human reliability coefficient; the reported κ measures agreement between the human verdict and an advisory automatic card. Both, along with every withheld column and every scoping decision, are documented in README.md.Code is licensed MIT; data, montages, and documents CC BY 4.0.

提供机构:
Zenodo
创建时间:
2026-08-12
二维码
社区交流群
二维码
科研交流群
商业服务