Visual Oracle Bench: Silent-Fabrication Incident, Three-Judge Synthetic Benchmark, Deterministic-Comparator-vs-Judge Study, and Dual-Writer Replay Fault Injection
收藏资源简介:
Complete data and analysis archive for an empirical study of integrity faults in LLM-based visual-regression test oracles. Four components: (1) retained records of a documented wrapper silent-fabrication incident (contaminated originals kept unmodified); (2) a 600-pair synthetic-HTML benchmark (400 injected-defect + 200 identity-control pairs) over three judge paths (Claude Sonnet 4.5, gpt-5-codex, Qwen 2.5 VL 7B) with clean re-dispatch records; (3) a 312-pair three-condition corpus (144 defect / 144 benign / 24 anchors) comparing four deterministic comparators at predeclared thresholds against both closed-weight judges, with sensitivity sweeps and an environment-fault log; (4) a replay fault-injection study in two separately implemented result writers (Python and TypeScript). All numbers regenerate from retained outputs with no new model calls. Corpora are synthetic and benchmark-authored. Code MIT; data CC BY 4.0.



