Evidence-Perturbation Audit: study outputs (per-item model probabilities, bootstrap tables, calibration and robustness analyses; no source-benchmark text)
收藏资源简介:
Experimental outputs for the evidence-perturbation / confidence-decoupling audit of question-answering models. Contains NO source-benchmark text: only per-item option probability distributions produced by the evaluated models, gold answers as option indices, item identifiers, intervention mappings, and derived analyses. Any benchmark text must be obtained from its original source via the companion Code deposition's loaders. Tiers are kept strictly separate: (1) preregistered confirmatory core — 18 runs (3 open-weight 7–9B models × 6 benchmarks, full N) with per-item probabilities and per-condition paired-bootstrap tables; (2) exploratory 70B tier (Qwen2.5-72B, Llama-3.3-70B; MedQA+PubMedQA; N=200), labelled; (3) exploratory OpenAI gpt-4.1 tier via the Batch API (PubMedQA/MedQA/RACE; N=300), labelled — including the RAW API responses, returned model identifiers, complete top-20 log-probability payloads, per-item censoring masks recording which option letters fell outside the returned top-20, and the reconstructed order-averaged distributions with per-rotation predictions and within-call letter margins; (4) fp16-vs-4bit precision robustness. Also included: calibration (Brier/NLL/ECE with intervals and 10/15/20-bin sensitivity, high-confidence-error rate), confidence-definition robustness (entropy/margin), answer-switch decomposition, split-half cross-fitting, the validity log, domain-breadth scoping, API feasibility findings, and a README mapping every paper number to the file that produces it. Companion Code deposition reproduces every result. See LICENSES.md for per-benchmark licences and the RACE non-commercial caveat. Names withheld for blind peer review.



