Drift-Robust Multi-Time Probes on IBM Kingston: Frauchiger–Renner Baseline, Apparent Non-Markovian Signatures, and the Role of Randomized Sampling
收藏资源简介:
Three Multi-Time Experiments on IBM Kingston: Baseline Validation, Apparent Anomaly, and Its Resolution Under Controlled Sampling Amit Brahmbhatt Quantum-Clarity LLC Email: amit@quantum-clarity.com Date: April 2026 Platform: IBM Kingston (156-qubit Heron r3 heavy-hex superconducting processor) Related prior work: Y⊗Z stabilizer basis migration (Zenodo DOI: 10.5281/zenodo.18498540; 10.5281/zenodo.19478241) Abstract We report three sequential experiments on the 156-qubit IBM Kingston (Heron r3) superconducting processor, motivated by the question of whether controlled multi-time correlation measurements on a quantum processor could reveal information-theoretic signatures associated with coupling to coherent degrees of freedom beyond simple Markovian open-system models. The first run implements a qubit-based Frauchiger–Renner no-go circuit across 20 parallel 4-qubit modules, recovering the quantum-mechanical joint-outcome distribution to within a total variation distance of 0.056 across 80,000 aggregated shots and establishing that the mitigation stack and parallel-module infrastructure are operational. The second run implements a trace-distance revival probe on three well-separated single-qubit patches and produces, on qubit 150, a 17.2σ apparent non-Markovian revival between delays of 5 and 15 μs, accompanied by an approximately 3% apparent growth in the equatorial coherence magnitude of the |+⟩ state across the same window — a pattern not naturally accommodated by simple Markovian dephasing. The third run, a controlled follow-up at matched delays with Hahn-echo refocusing and randomized time-interleaving of circuit submission order, finds that the original revival location and the |XY| growth signature do not reproduce. Echo refocusing reduces the observed XY-plane angular winding of the |+⟩ Bloch vector from approximately 250° across the delay range to a ±10° band, consistent with the dominant contribution being static Z-detuning in the rotating frame. We conclude that the apparent non-Markovian signature observed in the second run is more parsimoniously explained by static detuning together with sampling-order sensitivity under backend drift than by a robust local memory channel. The broader methodological finding — that naive sequential sampling of multi-time correlation experiments on superconducting quantum processors can produce apparent non-Markovian signatures at tens-of-sigma significance that do not survive randomized time-interleaving — is the principal contribution of this work. We provide a drift-robust protocol template for process-tensor and related multi-time characterization experiments on IBM hardware. Plain-Language Summary This work asks a practical question about quantum computers: when we run a carefully designed experiment on a real quantum processor and see an unusual, statistically strong signal, how confident can we be that the signal reflects genuine physics rather than an artifact of how we did the measurement? The answer we arrive at, after running three experiments on IBM's 156-qubit Kingston processor, is: less confident than the raw statistics suggest, unless specific controls are used. The headline result is a methodology lesson for anyone running multi-step timing experiments on quantum hardware. The three experiments were conducted in sequence. The first was a known test case — a circuit with a well-established textbook prediction — which we used to confirm that the machine and our software were working correctly. Our results matched the textbook prediction closely. The second experiment was more exploratory: we prepared a qubit in two different starting states, let it sit idle for varying amounts of time, and measured how distinguishable the two states remained. In ordinary physics, that distinguishability should only ever decrease as time passes. On one particular qubit, we appeared to see the distinguishability briefly increase between 5 and 15 microseconds of idle time — an effect that was statistically very strong (17 standard deviations) and that, taken at face value, would suggest the qubit was exchanging information with some hidden part of its environment. Before interpreting this result, we ran a third experiment to test whether the apparent effect was real or an artifact. We submitted the measurements in a deliberately scrambled order so that any slow drift in the processor's calibration over the course of the run would be spread evenly across all our measurements rather than piling up in one place. Under this better-controlled sampling, the apparent effect vanished. We also added a standard "echo" technique that cancels out the simplest kind of drift — and this cancellation cleanly explained most of what we had initially observed. The conclusion is that the original 17-sigma effect was not a hidden physical process but a combination of ordinary hardware drift and the specific order in which we happened to collect the data. The broader takeaway, which is the reason this work is being published rather than filed away, is that this kind of sampling artifact can look very convincing on modern quantum hardware. Measurement runs on these machines take tens of minutes during which the hardware itself is slowly changing, and if the order of measurements is correlated with what we are trying to measure, the slow changes can masquerade as the physical effect we are looking for. We recommend a simple protocol — randomize the order of measurements — that eliminates this failure mode, along with several other practices that make multi-step timing experiments on quantum processors more trustworthy. We make no claim to have tested deep questions about the nature of quantum mechanics, and the third experiment resolved the question we were investigating into the mundane answer. What remains is a reusable procedural lesson for others doing similar work. 1. Introduction The question of whether a universal-scale quantum computer could, in principle, produce empirical signatures bearing on interpretational questions in quantum mechanics has received renewed attention with the availability of 100+ qubit superconducting processors. Several proposals have been discussed in the literature, ranging from Wigner's-friend-style no-go demonstrations (Frauchiger and Renner 2018; Bong et al. 2020) to process-tensor characterization of multi-time quantum dynamics (Pollock et al. 2018; White et al. 2020). A common thread across these proposals is that they distinguish between experimental outcomes that are compatible with different classes of physical models — not between "detection" and "non-detection" of a specific entity — and their interpretational reach is therefore bounded by the precision with which those classes can be operationally separated on real hardware. This paper reports three sequential experiments on IBM Kingston (a 156-qubit Heron r3 heavy-hex superconducting processor) that were originally motivated by the question of whether multi-time correlation measurements could reveal information backflow into an inaccessible but coherent sector — a speculation loosely connected to non-Markovian models of open-system dynamics and, more distantly, to philosophical discussions of hidden degrees of freedom in quantum mechanics. Our primary finding is methodological rather than interpretational: an apparent high-significance non-Markovian signature observed on one qubit under sequential sampling failed to reproduce under randomized time-interleaving of circuit submission, and the Bloch-vector-level diagnostic that had initially appeared anomalous was resolved into a combination of static Z-detuning and sampling-order drift under stronger controls. The experimental arc proceeds in three stages. Section 3 establishes baseline operational validity of the measurement apparatus using a qubit-based Frauchiger–Renner circuit replicated across 20 parallel 4-qubit modules. Section 4 reports a trace-distance revival probe on three single-qubit patches and documents the apparent anomaly on qubit 150. Section 5 reports the controlled follow-up and the non-reproduction of that anomaly. Section 6 discusses the mechanism by which sequential sampling in nested experimental loops produces spurious multi-time correlations, and Section 7 provides a protocol template for drift-robust multi-time characterization. We do not claim to have tested or constrained any interpretational framework of quantum mechanics; the baseline Frauchiger–Renner result reproduces the quantum-mechanical prediction to within a few percent and is presented solely as an apparatus validation. The hardware infrastructure (hardware-adaptive layout discovery, T1-floor filtering, geometric-mean fidelity scoring) used in these experiments is drawn from prior published work on the Y⊗Z stabilizer basis migration architecture (Brahmbhatt 2026a, 2026b). The inner gate sequences implemented in the present experiments are distinct from the Y⊗Z stabilizer preparation and do not constitute a test of that architecture; the present work is logically self-contained. 2. Experimental Platform and Methodology 2.1 Backend All experiments were executed on IBM Kingston, a 156-qubit Heron r3 processor with a heavy-hex coupling graph. Hardware characterization confirming the heavy-hex topology (as opposed to the square-lattice topology of earlier-generation Eagle processors) was performed via coupling-map fingerprinting prior to the present work and documented in Brahmbhatt (2026b). At the time of the measurements reported here, the chip-mean T1 and T2 values inferred from the live backend target were approximately 270 μs and 350 μs respectively. Day-to-day drift in the reported T2 of individual qubits was observed to be on the order of 30% across the 12-day period bracketing the experiments, which reinforces the methodological imperative to rely on live backend-target properties rather than cached hardware snapshots. 2.2 Software stack Experiments were implemented in Qiskit (version current as of April 2026) and Qiskit IBM Runtime. Circuit submission was performed via the SamplerV2 primitive for shot-based measurements (the Frauchiger–Renner run and the trace-distance probes) with shots held at fixed values documented per experiment below. Qubit properties were accessed via the modern backend.target API; the deprecated backend.configuration() and backend.properties() methods are no longer authoritative on Heron r3 backends and were avoided throughout. 2.3 Layout discovery and qubit selection For the parallel-module Frauchiger–Renner experiment, the hardware-adaptive layout discovery procedure of Brahmbhatt (2026a, §5.3) was applied: the live coupling map was converted to an undirected adjacency representation, four-qubit simple paths (q0–q1–q2–q3) were enumerated, and each candidate path was scored by the geometric mean of its three edge two-qubit gate fidelities. Paths with any constituent qubit below 50% of the chip-mean T1 were excluded. Non-overlapping paths were selected in descending score order until the module target was reached. For the single-qubit probe experiments, qubits were scored by the product of T2 and readout fidelity, and selected greedily with a minimum graph distance of 5 between probes to minimize shared-bath effects. 2.4 Mitigation stack The mitigation stack applied in each experiment was chosen to match the scientific question of that experiment and was therefore not uniform. For the Frauchiger–Renner baseline, where the goal was to recover the known quantum-mechanical prediction with maximum fidelity, the full stack (XY4 dynamical decoupling, 32-randomization Pauli twirling on both gates and measurements) was enabled. For the multi-time probe experiments, the full stack was intentionally disabled: dynamical decoupling would suppress precisely the low-frequency memory dynamics that trace-distance revivals are designed to detect, and gate twirling would stochasticize coherent memory effects. The distinction between mitigation (noise suppression) and experimental control (systematic-error rejection) was maintained throughout: the Hahn echo applied in Section 5 is a control, not a mitigation, and is applied selectively as an experimental arm. 3. Baseline: Frauchiger–Renner Validation Across 20 Parallel Modules The first run established baseline operational validity of the infrastructure by implementing a qubit-based Frauchiger–Renner circuit across 20 parallel 4-qubit modules. Each module encodes the canonical Frauchiger–Renner setup with qubit roles assigned as coin (q0), friend F̄ (q1), system S (q2), and friend F (q3). The coin qubit is prepared in the superposition √(1/3)|0⟩ + √(2/3)|1⟩ via a Ry rotation with angle 2·arccos(1/√3) ≈ 1.9106 rad. Subsequent CX and CH operations propagate the measurement outcomes to the friend registers. Wigner-basis measurements on each friend pair are implemented via CX followed by Hadamard, and the joint (W̄, W) outcome on each module is read out in the {ok, fail} basis. The quantum-mechanical prediction for the aggregated joint outcome distribution is P(fail, fail) = 9/12 and P(fail, ok) = P(ok, fail) = P(ok, ok) = 1/12, with P(ok, ok) being the paradoxical event in the Frauchiger–Renner argument. The layout discovery procedure selected 20 non-overlapping 4-chains on Kingston with geometric-mean edge fidelities ranging from 0.9982 to 0.9990. Each module was transpiled to its assigned layout at optimization level 1 with an average physical-circuit depth of 25. The full mitigation stack was enabled (XY4 dynamical decoupling, Pauli twirling on gates and measurements with 32 randomizations). Each module was sampled at 4000 shots, for a total of 80,000 aggregated shots. The submitted job completed with 31.7 seconds of wall-clock elapsed time inclusive of queue and network overhead; the estimated billed QPU time was approximately 12 seconds. The aggregated observed distribution was P(fail, fail) = 0.6939, P(fail, ok) = 0.1092, P(ok, fail) = 0.1011, and P(ok, ok) = 0.0958, with a total variation distance of 0.056 from the quantum-mechanical prediction. The paradoxical event P(ok, ok) was observed at +12.7σ above the quantum-mechanical value of 1/12, indicating a detectable but small residual upward bias in the ok bucket consistent with independent single-qubit readout errors of order 1–2%. Reported against the stricter null hypothesis P(ok, ok) = 0, the observation corresponds to a z-score of 87.5σ; we report this figure only for completeness, as the intellectually relevant comparator is the quantum-mechanical prediction rather than an implausible zero baseline. Figure 1 displays the aggregated distribution alongside the per-module distribution of P(ok, ok) across the 20 modules. Figure 1. (a) Aggregated joint Wigner outcome probabilities for the Frauchiger–Renner circuit across 20 parallel 4-qubit modules on Kingston, compared to the quantum-mechanical prediction. (b) Per-module P(ok, ok) values showing two outlier modules (modules 4 and 13) and otherwise tight clustering near the QM prediction. Two modules (module 4 on qubits [80, 81, 82, 83] and module 13 on qubits [29, 30, 31, 32]) produced P(ok, ok) values substantially elevated above both the quantum-mechanical prediction and the distribution of the remaining 18 modules. Re-aggregating across the 18 non-outlier modules reduces the observed TVD to 0.036 and brings P(ok, ok) within 6% of the quantum-mechanical value. The per-module standard deviation of P(ok, ok) across the full set of 20 modules is 0.025, approximately six times the Poisson expectation of 0.004 for 4000 shots per module at P = 1/12. This excess variance reflects structured patch-to-patch differences in circuit performance on Kingston that are not captured by the coherence-time-weighted fidelity score used for layout selection. We report this observation as a characteristic of the platform relevant to subsequent multi-time probe design, not as an interpretational claim about quantum mechanics. The baseline Frauchiger–Renner result establishes that the mitigation stack, the parallel-module infrastructure, and the layout-discovery procedure are operational and capable of recovering a known quantum-mechanical prediction to within a few percent across a substantial fraction of the processor. It also establishes the patch-to-patch variance on Kingston as a factor that any subsequent multi-time experiment must account for. 4. Stage 1: Multi-Patch Trace-Distance Probe The second run implements a trace-distance revival probe on three well-separated single-qubit patches. The Breuer–Laine–Piilo framework (Breuer, Laine, and Piilo 2009) defines a necessary condition for CP-divisible (Markovian) open-system dynamics: the trace distance D(τ) = ½‖ρ₁(τ) − ρ₂(τ)‖₁ between any two initial states ρ₁(0) and ρ₂(0) must be monotonically non-increasing under the induced dynamical map. A statistically significant positive jump in D(τ) therefore provides operational evidence for non-Markovian dynamics — information backflow from the environment to the system. Three single-qubit patches were selected on Kingston by ranking qubits by T2 × (1 − readout error) and greedily choosing patches with pairwise graph distance at least 5 on the coupling map. The selection procedure yielded q150 (T2 = 468.7 μs), q125 (T2 = 474.7 μs), and q3 (T2 = 440.3 μs) with pairwise graph distances of 7, 13, and 13 hops. For each patch, the initial states |0⟩ and |+⟩ were prepared and evolved freely under an idle delay of τ ∈ {0, 5, 15, 30, 50, 80, 120, 180} μs, after which single-qubit tomography was performed in the X, Y, and Z bases. From the three Pauli expectation values at each (patch, state, delay) coordinate, the full Bloch vector was reconstructed, and from the pair of Bloch vectors at fixed (patch, delay) the trace distance D(τ) was computed with propagated statistical uncertainty. The full mitigation stack was disabled for this experiment. All 144 circuits (3 patches × 2 states × 8 delays × 3 bases) were sampled at 1000 shots each for a total of 144,000 shots. Circuits were submitted in nested-loop order (patches outermost, then states, then delays, then bases), which is the default ordering produced by the natural enumeration pattern. The job returned 2121 seconds of wall-clock time due to queue position; the estimated billed QPU time was approximately 17 seconds. Figure 2. Stage 1 results. (a) Trace distance D(τ) across the three patches P0 (q150), P1 (q125), and P2 (q3). The arrow marks the statistically significant upward jump between τ = 5 μs and τ = 15 μs on P0, corresponding to a 17.2σ positive jump in D(τ). (b) Equatorial coherence magnitude |XY| of the |+⟩ state on P0, showing an apparent ~3% growth across the 5–15 μs window that is not compatible with purely Markovian dephasing. Two features of the Stage 1 data motivated the controlled follow-up described in Section 5. First, on patch P0 (q150), D(τ) rose from 0.6740 at τ = 5 μs to 0.6887 at τ = 15 μs — an upward jump of 0.0147 at 17.2σ significance (σ_D ≈ 0.00085 per point, joint σ ≈ 0.0012 on the difference). The corresponding BLP non-Markovianity measure on this patch was N_BLP = 0.0147, with all of the non-Markovianity concentrated in the single 5→15 μs jump. The other two patches showed N_BLP = 0 and monotonic decay throughout. Second, decomposing the Bloch vector evolution on q150 across the same 5→15 μs window showed that the |+⟩ state moved from approximately (0.922, −0.222, 0.018) to approximately (0.746, −0.628, 0.006), while the |0⟩ state remained close to +Z throughout. The equatorial coherence magnitude |XY| = √(⟨X⟩² + ⟨Y⟩²) of the |+⟩ state correspondingly rose from 0.948 to 0.975, an apparent increase of approximately 3% over 10 μs of free evolution. Under ordinary dephasing dynamics, |XY| is expected to decrease rather than increase; the apparent increase observed in Stage 1 therefore motivated the controlled follow-up of Section 5. The cross-patch Pearson correlation of D(τ) was above 0.96 for all three pairs, indicating a strong shared overall decay envelope across the three probes. Taken at face value, the Stage 1 data appeared to show a genuine non-Markovian signature localized to one patch but embedded in a globally shared decay envelope. Several candidate physical mechanisms could produce such a pattern: coherent coupling to a local TLS defect on q150 with a characteristic timescale in the 5–25 μs range; static ZZ coupling to a nearby spectator qubit left in a partially-mixed state by preceding operations; or an artifact of the sequential nested-loop submission order, in which all of P0's measurements are performed in the first third of the job wall-clock window and backend drift on similar timescales could produce structure aligned with a single patch. The Stage 1 experiment lacked the controls necessary to discriminate these hypotheses and a targeted follow-up was designed. 5. Stage 1B-α: Controlled Follow-up with Hahn Echo and Randomized Time-Interleaving The third run discriminates between the candidate explanations of the Stage 1 anomaly by focusing exclusively on q150 and applying two controls: a Hahn echo arm in parallel with the original no-echo arm, and randomized time-interleaving of all circuit submissions. The Hahn echo control inserts a midpoint π_x pulse between two half-delays of duration τ/2 each, which exactly refocuses any static Z-detuning or slow rotating-frame phase accumulation during the delay. A signal attributable to static detuning will therefore vanish under the echo; a signal attributable to genuine non-Markovian dynamics will survive. The randomized time-interleaving control shuffles the order of all 288 circuits in the job via a fixed random seed, so that each experimental coordinate (state, delay, basis, control arm, replicate) is sampled approximately uniformly across the job wall-clock window. Any structure that had been aligned with a specific sequential position in the Stage 1 run — whether from backend drift, calibration evolution, or accumulated thermal effects — will be distributed across all experimental coordinates under interleaved sampling and will average out in the aggregated expectation values. The experiment employed 2 replicates × 2 control arms (no-echo, Hahn echo) × 2 states (|0⟩, |+⟩) × 12 delays × 3 bases = 288 circuits at 1000 shots each for 288,000 aggregated shots. The delay grid was densified in the 5–25 μs window to {5, 9, 13, 17, 21, 25} μs and anchored at {0, 30, 50, 80, 120, 180} μs to retain long-tail normalization. Submission order was randomized with seed 20260420. The job returned 109.1 seconds wall-clock time with approximately 36 seconds of estimated billed QPU time. All mitigation (dynamical decoupling, gate twirling, readout twirling) was disabled as in Stage 1. Figure 3. Stage 1B-α results on q150. (a) Trace distance D(τ) for the no-echo and Hahn-echo arms, each shown with per-replicate traces (thin lines) and the across-replicate average (thick lines with error bars). Both arms show predominantly monotonic decay. (b) XY-plane angular coordinate of the |+⟩ Bloch vector as a function of delay. The no-echo arm (red) shows rapid angular winding reaching approximately 250° of accumulated phase across the delay range; the Hahn-echo arm (blue) holds the angle within a ±10° band. The echo performs as designed and refocuses the dominant XY-plane dynamics to a static frame. Three observations from the Stage 1B-α data bear on the interpretation of the Stage 1 anomaly. First, the Hahn-echo arm refocuses the XY-plane angular winding almost completely. Across the full delay range from 0 to 180 μs, the no-echo |+⟩ Bloch vector angle sweeps through hundreds of degrees of phase (Figure 3b, red curves), while under the Hahn echo the same angle remains confined within approximately ±10° (Figure 3b, blue curves). This demonstrates that the large-amplitude XY-plane rotation observed in both Stage 1 and the no-echo arm of Stage 1B-α is predominantly a rotating-frame detuning effect that is cleanly refocused by a single midpoint π pulse. The echo arm's Bloch-vector trajectory for |+⟩ is essentially pure equatorial decay at a fixed angle, corresponding to an unambiguous dephasing channel. Second, the apparent revival structure is not stable between runs. The averaged no-echo D(τ) in Stage 1B-α shows an upward jump between τ = 17 μs and τ = 21 μs (ΔD = +0.0263 at 39.8σ), at a different location than the Stage 1 revival at 5→15 μs. The averaged Hahn-echo arm shows a smaller upward jump between τ = 0 and τ = 5 μs (ΔD = +0.0099 at 18.8σ) at yet a third location. A genuine physical feature of the dynamics would be expected to reproduce at approximately the same delay across runs, since the characteristic timescales of any real environmental mode are stable on day-to-day scales. The non-reproduction of the revival location, combined with its migration across control arms, is characteristic of Poisson noise combined with finite-grid sampling artifacts rather than of a fixed environmental resonance. Third, and most diagnostic, the |XY| magnitude growth signature from Stage 1 does not reproduce. Table 1 compares the equatorial coherence magnitude of the |+⟩ state on q150 in the critical 5–15 μs window across the three measurement settings. τ (μs) Stage 1 |XY| 1B-α no-echo |XY| 1B-α Hahn-echo |XY| 0 0.990 0.986 0.987 5 0.948 0.939 0.958 9 — 0.930 0.965 13 — 0.947 0.945 15 0.975 — — 17 — 0.933 0.938 21 — 0.924 0.933 25 — 0.869 0.934 30 0.893 0.870 0.907 Table 1. Equatorial coherence magnitude |XY| on the |+⟩ state of q150 across the three experimental settings at matched and densified delays. The Stage 1 column shows the apparent ~3% growth from 0.948 at τ = 5 μs to 0.975 at τ = 15 μs. The Stage 1B-α no-echo column, measured under identical physical conditions but with randomized submission order, shows at matched delays (5 μs and the denser grid around 15 μs) a monotonic or near-monotonic decay with a much smaller amplitude (0.939 → 0.947 → 0.933 → 0.924 across 5, 13, 17, 21 μs) that is consistent with ordinary dephasing. The Hahn-echo column preserves |XY| slightly better at intermediate delays, as expected when static-detuning-induced apparent dephasing is refocused. Figure 4. Direct comparison of q150 data between the Stage 1 sequential-order run and the Stage 1B-α time-interleaved no-echo arm. (a) Trace distance D(τ). The shaded region marks the Stage 1 revival window at 5–15 μs; in the time-interleaved data the same window shows monotonic decay. (b) Equatorial coherence magnitude |XY| on the |+⟩ state. The Stage 1 apparent growth within the shaded window (red curve) does not reproduce in the time-interleaved data (green curve). The cross-replicate consistency of Stage 1B-α is high in both arms (Pearson r = 0.983 between replicates in the no-echo arm, r = 0.981 in the Hahn-echo arm, RMS differences of 0.016 and 0.021 respectively). Within a single submitted job, the experiment reproduces itself well. What fails to reproduce is the specific revival structure observed in Stage 1 when the experimental coordinates were sampled in a different order across wall-clock time. The conjunction of Stage 1 and Stage 1B-α evidence therefore supports the interpretation that the apparent Stage 1 anomaly is decomposable into two contributions: a large-amplitude but physically mundane static Z-detuning that is cleanly refocused by a single π pulse, and a smaller-amplitude sampling-order artifact aligned with the sequential enumeration of experimental coordinates against slow backend drift. Taken together, the Stage 1B-α results do not support carrying forward the Stage 1 anomaly as evidence of localized non-Markovian backflow on q150. The stronger controls preserve the ordinary detuning signature — as expected for a well-calibrated transmon in a rotating frame — while eliminating the stability of the original claimed feature. The scientifically useful content that survives the controlled follow-up is therefore the characterization of q150's detuning and dephasing dynamics at this epoch, together with the methodological observation developed in Section 6. 6. Discussion: Drift Artifacts and the Role of Randomized Sampling The mechanism by which sequential nested-loop sampling produces apparent multi-time correlations is straightforward to describe and, we argue, systematically underappreciated in the superconducting quantum computing literature on multi-time characterization. Consider an experiment with coordinates (C₁, C₂, ..., C_n) enumerated in nested-loop order, where C₁ varies most slowly across the job wall-clock and C_n most quickly. The measured expectation value at each coordinate combination is an average over N shots taken at a specific wall-clock time. If any experimentally relevant parameter of the backend — T1, T2, frequency calibration offset, readout classifier threshold, cross-talk strength, or refrigerator temperature — drifts on a timescale comparable to or shorter than the total job duration, then the drift will induce apparent correlations in the measured data that are aligned with the outer loop coordinate C₁. Specifically, if the outer loop varies over patches and the drift has, for instance, a monotonic component over the job duration, then the first patch's data will systematically differ from the last patch's data by an amount that has nothing to do with the physical distinction between those patches. In the Stage 1 data of Section 4, the outer loop was patches, the inner loops were (state, delay, basis). Each patch's full measurement campaign occupied a contiguous approximately 12-minute block of the 35-minute job. Backend parameter drift on the 5–15 minute timescale is a well-documented feature of superconducting processors and can interact with the experimental coordinate structure to produce apparent effects when the outer-loop coordinate changes slowly with wall-clock time. The specific pattern observed — an apparent non-Markovian revival on one patch only, with high cross-patch correlation in the overall decay envelope — is precisely what sequential drift against the patch coordinate would produce: all three patches share a decay shape because they share the dynamics being measured, but one patch's revival location gets populated preferentially because that patch's measurements were all taken during a particular slice of the drift trajectory. The |XY| growth signature is similarly vulnerable: a small systematic shift in the mean Pauli expectation values between the early and late halves of a patch's block, if aligned with the outer delay coordinate, will produce an apparent coherence-recovery effect that is entirely an artifact of when the measurements were made. Randomized time-interleaving eliminates this class of artifact by construction. Under interleaving, each (state, delay, basis, patch) coordinate is sampled at multiple randomly-distributed wall-clock times across the job, so that any drift contribution averages out at the coordinate level rather than aligning with a specific coordinate. The Stage 1B-α data are consistent with this interpretation: under randomized submission order, the fixed-delay revival structure seen in Stage 1 is replaced by smaller residual fluctuations scattered across the delay grid. A subsidiary but important observation is that cross-replicate consistency within a single interleaved job is not by itself sufficient to establish reproducibility of a multi-time effect. In Stage 1B-α, the two replicates were drawn from the same interleaved submission order and therefore experienced the same effective drift averaging; their high correlation (r > 0.98) confirms that the interleaved measurement is internally reproducible but does not speak to whether an apparent effect would survive a second independent job at a different wall-clock time with a different random seed. Full reproducibility testing at this level would require matched-grid runs across separate jobs, ideally on different days, which we did not perform in the present work. We emphasize that the present finding does not imply that non-Markovian dynamics cannot be characterized on superconducting quantum processors — there is substantial published literature demonstrating that they can (Morris et al. 2019; White et al. 2020; Figueroa-Romero et al. 2021). The finding is specifically that the combination of sequential nested-loop sampling and finite-grid delay enumeration can produce apparent high-significance signatures at the level of tens of standard deviations on modern superconducting hardware, that these signatures can survive naive within-run statistical tests, and that they vanish under randomized time-interleaving. Any multi-time correlation experiment on this class of hardware that does not explicitly implement time-interleaving controls is, in our assessment, at risk of reporting similar artifacts. 7. Protocol Recommendations On the basis of the experience reported in Sections 3–5, we recommend the following protocol elements for multi-time correlation experiments on superconducting quantum processors, offered as a template rather than a standard. Circuit submission order should be randomized with a fixed, documented random seed. The random permutation should be recorded alongside the experimental descriptor list so that post-hoc analysis can verify that each experimental coordinate was sampled approximately uniformly across the wall-clock duration of the job. When the processor backend supports it, the uniformity of sampling in time should itself be tested: for example, by binning each coordinate's measurements into quartiles of the job duration and checking for consistency of the measured expectation value across quartiles. Replicates within a single job, as implemented in Stage 1B-α, provide a check on within-run reproducibility but do not replace matched-grid runs across independent jobs. We recommend at least two independent jobs, ideally on different days, for any multi-time result that will be submitted for publication. The Pearson correlation and RMS difference of the measured expectation values across independent jobs provides a quantitative reproducibility metric and should be reported alongside within-run statistics. Mitigation and control should be kept conceptually and mechanically separate. Dynamical decoupling and randomized compiling can suppress apparent noise but can also mask genuine non-Markovian signatures; they are therefore not appropriate as default components of a multi-time probe. When they are applied, they should be applied as distinct experimental arms to be compared against an unmitigated baseline, not as background assumptions. Controls such as the Hahn echo implemented in Section 5 play a different role: they test specific null hypotheses (here, the hypothesis that the signal is static Z-detuning) by selectively removing particular physical effects, and they should be applied as additional experimental arms alongside both mitigated and unmitigated baselines. Bloch-vector-level decomposition, rather than summary scalar quantities such as D(τ) or N_BLP alone, provides substantially stronger diagnostic value. The Stage 1 result would have been more difficult to interpret without direct examination of the underlying Bloch vector trajectories, and the Stage 1B-α result would have been harder to attribute to static detuning without direct examination of the XY-plane angular dynamics. When computational and measurement budget permits, full single-qubit tomography should be preferred over single-observable measurements for multi-time probes. Finally, patch-to-patch variance on modern superconducting processors is substantial and not fully captured by standard fidelity-scoring functions. The six-fold excess of per-module variance above the Poisson expectation in the Frauchiger–Renner baseline (Section 3) illustrates this quantitatively. Experimental designs that rely on a single patch — or that aggregate across patches without explicit accounting for patch variance — are vulnerable to misattributing patch-specific artifacts to generic device-level effects. The baseline characterization of patch variance on the target processor, prior to any multi-time measurement, is a prerequisite for interpreting results obtained on that processor. 8. Conclusion We have reported three sequential experiments on IBM Kingston comprising a parallel-module Frauchiger–Renner baseline across 20 modules, a multi-patch trace-distance revival probe, and a controlled follow-up with Hahn-echo and time-interleaving discriminators on the patch that had produced an apparent revival. The baseline run reproduces the quantum-mechanical prediction to within a total variation distance of 0.056 of the ideal distribution and establishes the operational validity of the infrastructure. The multi-patch probe produces, under sequential sampling, an apparent 17.2σ non-Markovian revival on one patch accompanied by a 3% growth in equatorial coherence magnitude that is not compatible with Markovian dephasing at face value. The controlled follow-up, under randomized time-interleaving and with Hahn-echo refocusing, finds that the dominant contribution to the original signal is static Z-detuning cleanly refocused by a single π pulse, and that neither the revival location nor the |XY| growth signature reproduces when the experimental coordinates are sampled in randomized order across the job wall-clock duration. The principal methodological finding of this work is that naive sequential nested-loop sampling of multi-time correlation experiments on superconducting quantum processors can produce apparent non-Markovian signatures at significance levels of tens of standard deviations that do not survive randomized time-interleaving. The mechanism is the alignment of slow backend drift with the outer enumeration coordinate of the nested loop structure, which induces apparent coordinate-dependent effects that are in fact artifacts of the sampling-in-time. Randomized time-interleaving is a low-cost and effective remedy, and we provide protocol recommendations for its implementation alongside Bloch-vector-level diagnostic decomposition, separation of mitigation from experimental controls, and cross-job reproducibility checks. We do not interpret our findings as bearing on interpretational questions in quantum mechanics. The baseline Frauchiger–Renner result reproduces the ordinary quantum-mechanical prediction and was included solely as an apparatus validation. The multi-time probe experiments, considered jointly, are consistent with the absence of genuine non-Markovian signatures on the single qubit examined with the sensitivity available in 15–40 seconds of billed QPU time. Accordingly, we report a null result with respect to robust non-Markovian backflow on the single qubit examined at the sensitivity and runtime scales accessed here. The controls and protocol elements reported in the preceding sections should, in our view, be considered minimum requirements for any future experiment that seeks to make a positive claim about non-Markovian dynamics on this class of hardware. Data Availability The raw experimental outputs from all three runs — comprising the full aggregated and per-module / per-patch result JSONs (Qiskit job identifiers, aggregated outcome counts, reconstructed Bloch vectors, and derived trace-distance series) — are deposited alongside this manuscript in the same Zenodo record and may be used without further permission under the record's license. The figures reproduced in the manuscript are included as separate files at the deposit. Experimental scripts are not included in the deposit; interested researchers may contact the author directly. Acknowledgments The author thanks IBM Quantum for access to the Kingston backend. The infrastructure for hardware-adaptive layout discovery and per-edge fidelity scoring applied in Section 3 derives from prior work on the Y⊗Z stabilizer basis migration architecture (Brahmbhatt 2026a, 2026b) and is reused in the present experiments without modification to its hardware layer. The present manuscript is logically independent of that prior work and does not constitute a test of the Y⊗Z architecture. Iterative critical review during the design and interpretation of the three experiments led directly to the adoption of randomized time-interleaving and the Hahn-echo discriminator reported in Section 5, and is gratefully acknowledged. References Bong, K.-W., Utreras-Alarcón, A., Ghafari, F., Liang, Y.-C., Tischler, N., Cavalcanti, E. G., Pryde, G. J., and Wiseman, H. M. (2020). A strong no-go theorem on the Wigner's friend paradox. Nature Physics 16, 1199–1205. Brahmbhatt, A. (2026a). QuantaCore: A Y⊗Z Stabilizer Basis Migration Architecture for Fault-Tolerant Modular Quantum Computing. U.S. Provisional Patent 63/952,786. Zenodo prior art record DOI: 10.5281/zenodo.18498540. Brahmbhatt, A. (2026b). Y⊗Z Orthogonal Stabilizer Validation on the 156-qubit IBM Kingston Processor. Zenodo DOI: 10.5281/zenodo.19478241. Breuer, H.-P., Laine, E.-M., and Piilo, J. (2009). Measure for the degree of non-Markovian behavior of quantum processes in open systems. Physical Review Letters 103, 210401. Figueroa-Romero, P., Modi, K., and Pollock, F. A. (2021). Markovianization with approximate unitary designs. Communications Physics 4, 127. Frauchiger, D., and Renner, R. (2018). Quantum theory cannot consistently describe the use of itself. Nature Communications 9, 3711. Morris, J., Pollock, F. A., and Modi, K. (2019). Non-Markovian memory in IBMQX4. arXiv:1902.07980. Pollock, F. A., Rodriguez-Rosario, C., Frauenheim, T., Paternostro, M., and Modi, K. (2018). Non-Markovian quantum processes: Complete framework and efficient characterization. Physical Review A 97, 012127. White, G. A. L., Hill, C. D., Pollock, F. A., Hollenberg, L. C. L., and Modi, K. (2020). Demonstration of non-Markovian process characterisation and control on a quantum processor. Nature Communications 11, 6301.



