HilbertBench Study G: Blinded Detection-Boundary Diagnosis — Corpus, Commitments, and Scoring Record
收藏资源简介:
Complete, independently verifiable record of HilbertBench Study G: a blinded,pre-registered comparison of a human diagnostician against a frozen mechanicaldecision rule, both diagnosing planted failure modes in variational quantummachine-learning runs at the boundary where the diagnostic evidence stops beingconclusive. CONTENTS. (1) corpus_blinded/ — 63 sealed traces exactly as the diagnosticianreceived them, under random identifiers, plus her blank sheet. (2) unsealed/ —the run-level assignment (arm, strength level, seed), answer key, frozen-ruleoutputs, the diagnostician's filled sheet, and the computed scores. (3)commitments/ — four SHA-256 commitments, each published before the artifact itprotects could have been influenced. (4) code/ — protocol, corpus generator,scorer, diagnostician instructions, and the unsealing log. (5) VERIFY.md — astep-by-step recipe to re-verify every commitment and re-score from scratch;re-scoring reproduces the archived results byte-identically. DESIGN. 63 simulator-generated runs sweeping three failure mechanisms fromphysically absent to strongly expressed (barren plateau by circuit width; shotstarvation by measurement budget; noise domination by depth on a calibrateddevice-noise model), plus two-mechanism mixtures and mechanism-free controlsincluding shallow converged-trajectory traps. Ground truth is always a physicalknob setting, never a diagnostic statistic. Runs were generated in shuffledorder from a hash-committed private assignment; the comparator rule's outputswere hash-committed before the diagnostician received the corpus. RESULT, reported as pre-registered regardless of direction. The primaryhypothesis — that the human outperforms the frozen rule — was NOT supported:human 40/63 (0.635, 95% CI [0.511, 0.743]) versus rule 38/63 (0.603, [0.480,0.715]), with only 2 discordant pairs (b=2, c=0), below the pre-registeredunderpower threshold of 6. The two graders assigned identical primary labels on57/63 runs, because the diagnostician's stated procedure closely resembled themechanical rule. The confidence hypothesis was supported: her reportedconfidence was lower inside transition regions (mean 0.728) than on stronganchors (0.904), one-sided Mann-Whitney U = 42.0, p = 0.0037. Detectionboundaries reproduced on fresh seeds (barren detection steps between 7 and 8qubits; noise between depths 12 and 14). She made 0/9 false positives onmechanism-free runs versus the rule's 1/9. INTERPRETATION NOTE. Absolute accuracies understate both graders: at the weakestsweep levels the planted mechanism is physically negligible yet ground truthstill names it, so a correct reading of "healthy" scores as incorrect. This iswhy the primary outcome is a paired comparison. See the unsealing log for thisand two further disclosures. Instrument frozen at HilbertBench v1.0.0. Pre-registration: OSF10.17605/OSF.IO/UYN46 (embargoed).



