quantum-governance-testbed v1.0.0: QG series complete — QG-001 (FAIL), QG-002 (FAIL), QG-003 (boundary located), QG-004 (PASS) on live IBM Quantum hardware
收藏资源简介:
Quantum Governance (QG) series v1.0.0 — series complete. Four preregistered experiments asking whether evidence produced by imperfect quantum hardware is trustworthy enough to authorize action across an execution boundary. Part of the Remnant Fieldworks Coherent Inheritance Framework (CIF) / ExecutionProof program. This release adds QG-002 (FAIL) and QG-003 (boundary located) to the QG-001 (FAIL) and QG-004 (PASS) results published in v0.2.0. All runs on live IBM Quantum hardware (Heron r2), 4096 shots. QG-001, QG-002 and QG-004 use the DM-002 sterile-neutrino substrate circuit (single data qubit, 6 Trotter steps, analytic survival probability P = 0.087782). Questions, thresholds and kill conditions for QG-001, QG-002 and QG-004 were frozen and SHA-256 locked in PREREGISTRATION.md (eec1aab4…9cba9) before any hardware result was computed, and the lock was re-verified intact after execution. QG-001 — Verdict Stability Under Hardware Noise: FAIL Five error-treatment arms on one processor; PASS required all five to agree under a 10% relative-error threshold. raw 0.094727 (7.91%, PASS); dynamical decoupling 0.089355 (1.79%, PASS); measurement twirling 0.095703 (9.02%, PASS); gate twirling 0.094238 (7.35%, PASS); zero-noise extrapolation 0.102760 (17.06%, FAIL). Max spread 0.013405. Finding: the error-mitigation choice flipped a preregistered verdict, so it is a governance-relevant parameter rather than an implementation default. QG-002 — Verdict Portability Across Processors: FAIL Identical circuit, shot count and raw (unmitigated) configuration submitted to two independent physical processors. ibm_marrakesh returned survival 0.097412 (10.97%, FAIL); ibm_kingston returned 0.086914 (0.99%, PASS). Spread 0.010498; relative-error ratio approximately 11×. First recorded as HOLD while the second backend sat in an extended queue, then resolved to FAIL on completion under the preregistered kill condition of at least two backends. Finding: the only variable was the physical processor, and the verdict changed anyway. Backend identity is governance-relevant; a record that does not bind the processor identity is certifying something it did not measure. QG-003 — Noise-Induced Boundary Crossing: PASS (boundary found) — reduced evidence tier A nine-level depth sweep on ibm_marrakesh, raw arm only. The DM-002 neutrino circuit cannot answer a depth question — at theta = pi/4 the transpiler collapses every step count to constant depth 6 with zero two-qubit gates — so QG-003 uses the axion–photon coupling circuit from the same hash-locked common/circuits.py, whose XX term compiles to CNOT–RZ–CNOT and scales linearly with depth. Each level is scored against the analytic target at that step count, so Trotter discretisation error is never charged to hardware noise. Results (transpiled depth, two-qubit gates, measured P, relative error): 14/2, 0.882080, 3.35%, PASS; 26/4, 0.867188, 11.01%, FAIL; 50/8, 0.660645, 32.63%, FAIL; 74/12, 0.407959, 58.43%, FAIL; 98/16, 0.265381, 72.97%, FAIL; 146/24, 0.504883, 48.59%, FAIL; 194/32, 0.864502, 11.97%, FAIL; 290/48, 0.822998, 16.20%, FAIL; 386/64, 0.775879, 21.00%, FAIL. The boundary sits at transpiled depth 26 — two Trotter steps, four two-qubit gates. One CNOT pair of headroom separates a trustworthy measurement from an untrustworthy one on this circuit and processor. The more consequential finding is that the curve is not monotonic in depth: fidelity collapses to the two-qubit depolarized floor (measured 0.265 against a floor of 0.25) at depth 98, then recovers to 0.865 at depth 194 — roughly double the gate count — before diverging again. Coherent errors partially cancel at particular depths, so a fixed depth cutoff is not a reliable trust control. A "depth ≤ N" policy would have admitted the depth-194 point, which is outside the threshold and only looks accurate by coincidence. Evidence tier, stated plainly. QG-003 is published at a reduced evidence tier. Its design was fixed before the jobs were submitted in practice, but the preregistration file was destroyed by an ephemeral workspace loss before it could be pushed and SHA-256 locked, so that ordering is asserted and not cryptographically provable. QG-003 must not be cited as a hash-locked preregistered result. What is independently verifiable: nine server-side IBM Quantum job records with their own timestamps; analytic targets recomputable from the unchanged, hash-locked common/circuits.py; and a 9/9 exact match of every measured value re-fetched from IBM after the loss. QG-001, QG-002 and QG-004 are unaffected. Full disclosure in AMENDMENT_V2.md §2.1 and RECONSTRUCTION_NOTICE.md. QG-004 — Mitigation Cannot Manufacture Truth: PASS The same five arms scored against a deliberately falsified target: the survival probability for theta = pi/8 (P_false = 0.543891) rather than the angle the circuit implements, theta = pi/4 (P_true = 0.087782); separation 0.456109. Arm values 0.088135, 0.100098, 0.097168, 0.090088, 0.095216. false_pass_count = 0; no arm was closer to falsehood than to truth. Finding: error mitigation did not fabricate support for a claim the circuit never computed. The series finding Three of the four metadata heuristics a governance system would naturally lean on each independently flipped a verdict: the error-mitigation label (QG-001), the assumption that processors are interchangeable (QG-002), and circuit depth as a proxy for fidelity (QG-003). The only control that held was QG-004's: direct comparison of the output against something independently known to be true. Everything else was a label about the computation rather than a measurement of it. Two conclusions, stated no more strongly than the evidence supports: keep decision thresholds far away from the noise floor, and verify against ground truth every time rather than inferring trust from metadata. These are claims about governance design on toy 2-qubit circuits. They are not claims about IBM processor quality, not claims about error mitigation quality, and not claims about production governance systems. Provenance and honesty disclosures The ephemeral working directory was destroyed twice on 2026-08-14. The runner scripts and locally written ProofRecord files were lost both times, and the second loss also took the QG-003 preregistration before it could be hash-locked. Every hardware result was re-fetched from IBM Quantum's authoritative server-side job records and matched its pre-loss value exactly. The run.py files here are faithful reconstructions of the conditions executed, not the byte-identical executed files; the execution evidence is the job IDs, which are recorded in the ProofRecords. ProofRecord self-binding hashes were regenerated. Two of the four verdicts in this release are FAILs and both are published in full, which is the point of preregistering. Repository: github.com/derekhone/quantum-governance-testbed



