| 01 |
QAOA MaxCut scaling |
QAOA, warm start, GW SDP |
Angle transfer works (Δ ≤ 0.02); P(optimal) decays ≈ 0.87^n; GW/greedy/SA are exact at n ≤ 20; WS-QAOA inherits its 0.97 from the classical warm start; QP-relaxation warm start is degenerate; P(optimal) is 10⁴× uniform sampling at n = 20 but uniform best-of-2 000 shots already reaches 0.92 |
| 02 |
Flakiness-aware test selection |
QUBO, QAOA, SA |
The standard squared coverage penalty makes the QUBO optimum infeasible on 19/20 suites; after repair, greedy set cover (≤ 1 ms) beats every QUBO method; 2 000 uniformly random selections are feasible more often than QAOA samples at every depth and land within 0.1 of QAOA's repaired ratio; flakiness term only matters on redundant suites (−23 % spurious failures) |
| 03 |
Join ordering as a QUBO |
QUBO encodings, QAOA warm start, DP/GA |
The encoding decides everything: the TSP-like adjacency QUBO tracks C_out with ρ ≈ 0.4 and its exact minimiser is 25–50 % off (catastrophic on star joins); a prefix-log QUBO reaches ρ ≈ 0.9–0.97 and 96–99 % of optimal; QAOA samples a valid permutation ≤ 0.4 % of the time unless warm-started, and then returns the greedy plan; random bit-vectors through the same repair match uniform-init QAOA and n! random permutations match the warm-started rows at zero circuit cost; at cardinality noise σ = 1 exact DP is tied with GA and SA |
| 04 |
CI/CD DAG scheduling on heterogeneous runners |
time-indexed QUBO (sink / makespan encodings), SA, warm-started QAOA vs FIFO, HEFT, exact MILP |
HEFT is optimal on 94 % of pipelines (0.933 on wide random DAGs) in 0.2 ms and the exact MILP takes ≤ 4 s; FIFO-by-label (what a CI server does) is 35–41 % above optimal; the QUBO's own time discretisation costs 6 % (2-min slots) to 22 % (5-min) before any solver runs, the HEFT horizon prunes 61–90 % of the variables, SA's repaired quality (0.97 / 0.85–0.89) comes from the list-scheduling repair (raw 0.75–0.92); the makespan-variable encoding beats the sink proxy on multi-sink DAGs by 2–4 points; warm-started QAOA returns HEFT; paired reruns with other solver seeds move the stochastic solvers' mean ratios by ≤ 0.003; 50 jobs at 1-min resolution = 9 900–16 600 qubits |
| 05 |
Grover on dependency-resolution SAT |
Grover, resource counts |
Oracle verified to 1e-15; Grover is 1.5–3.3× more expensive than a 60-line DPLL in clause evaluations at n = 6–12; 200-variable instance needs 739 logical qubits and 3.2e34 T gates |
| 06 |
VQE on H2 and LiH |
VQE, UCCSD, barren plateaus |
UCCSD reaches chemical accuracy at every geometry (H2 1.7e-9 mHa, LiH 3.7e-6 mHa); hardware-efficient ansätze fail on LiH (HEA-L1 ≡ Hartree–Fock, 0.18 % correlation) unless trained with L-BFGS on exact gradients; barren-plateau decay Var ∝ b^n measured with b = 0.86 → 0.50 as depth/observable locality change; 925 k energy evaluations to reproduce what eigvalsh gives exactly |
| 07 |
Projected quantum kernels for anomaly detection |
quantum kernels, spectrum diagnostics |
The qubit budget (top-k + PCA) is the bottleneck: kNN on raw bigrams gets AUC 1.00, nothing inside the budget exceeds 0.84; median-bandwidth RBF scores below chance; projected ZZ kernel is best inside the budget (0.77–0.84) |
| 08 |
Quantum reservoir computing for DevOps telemetry |
QRC, density-matrix sim |
QRC ties an equal-feature-count classical random-feature map (paired bootstrap) and loses on NARMA-10 / Mackey–Glass; AR(10) is the best 1-step forecaster; reservoir state is near-maximally mixed |
| 09 |
IQAE for VaR and CVaR |
amplitude estimation |
Query-scaling slopes −0.50 (MC) vs −0.98 (IQAE) vs −0.91 (MLAE) confirmed; small-amplitude bias up to +15× at a = 0.002, Jeffreys correction removes 84 %; at equal query budget classical MC is still 2–5× more accurate for CVaR |
| 10 |
Grover second-preimage on a toy hash |
Grover, in-place reversible oracle, resource extrapolation |
Success probability matches sin²((2k+1)θ) to 1e-15 at 14–26 qubits; counted T-count/depth per iteration equals a closed-form model; extrapolated 128-bit-truncated hash needs 2^80 T gates and 510 logical qubits (4×10⁷ years at 10⁹ T/s), 2^96 gates under NIST MAXDEPTH 2^64 — only 2^32 cheaper than classical |
| 11 |
QEC decoding as a tabular ML benchmark |
Stim, PyMatching, sklearn |
Generic classifiers reach 1.1–2.5× MWPM's LER at d = 3 but get worse at d = 5 (never see the threshold); a hybrid given the MWPM bit copies it exactly; learning curve flat after 5k samples |
| 12 |
Tensor-network #SAT for product lines |
tensor networks (quantum-inspired) |
Exact counts and feature probabilities on 42 feature models, 33/33 agreement with DPLL/brute force; feature-model CNFs have tiny treewidth (width 3–10 up to 120 features, 0.8–4.6 ms to contract) but a 200-line DPLL is faster and the 5 s path search dominates; width explodes (27–32) once cross-tree constraints exceed ~1 per feature |
| 13 |
MPO compression of NN weights |
MPO / tensor train (quantum-inspired) |
Un-healed MPO is the worst structured method at equal parameters; after healing it beats SVD at tight budgets but not magnitude pruning; int4 matches the uncompressed model for free; random-MPO control stays at chance |
| 14 |
Simulated bifurcation for community detection |
SB (quantum-inspired), QAOA, Louvain |
aSB/bSB/dSB reach the brute-force bipartition optimum and beat spectral/SA; none beats Louvain once k > 2; QAOA p ≤ 2 falls well short of a 32-trajectory dSB run; dSB time scales as n^0.98 and an implicit-coupling dSB bisects 10⁵-node planted graphs in 16 s with higher modularity than networkx Louvain (328 s) |
| 15 |
Szegedy quantum PageRank at scale |
quantum walk (sparse simulation) |
"Hub suppression" is false at α = 0.85 (hub share grows with N); the real effect is degeneracy lifting (76–89 % of classical ties resolved); quantum ranks are less stable under edge removal; O( |
| 16 |
Clifford QCA as texture/terrain generator |
stabilizer QCA (Stim) |
Stabilizer automata on up to 4096 cells / 64×64 tori in milliseconds; per-cell marginals are exactly quantised to {0, ½, 1} and pairwise mutual information to {0, 1} bit (a single T-gate layer breaks this), so a Clifford QCA is a generator of binary masks with locally-constrained randomness controlled by one knob (fraction of cells in superposition), not a continuous noise field; classical PCG primitive by Gottesman–Knill; real ibm_fez run on 129 qubits (2026-10-05): deterministic cells 94 % → 67 % correct over T = 1 → 8, Stim checks it in 39 ms |
| 17 |
Born machines for rhythm generation |
QCBM / IQP Born machine, MMD training |
Over 3 training seeds the 16-qubit QCBM's held-out MMD² win over a position-conditioned Markov chain is robust (0.012 vs 0.033, paired CI excludes 0) but its single-seed NLL win does not survive (diff CI includes 0); QCBM trades novelty (0.42 vs 0.76) for fit; an untrained circuit (MMD² 0.094) and a random search with as many evaluations as training steps (0.088) are far worse than both, so the fit comes from training, not the ansatz; exactly classically simulable, no advantage claimed |
| 18 |
Quantum-inspired SimHash |
random-Pauli sign hashes of a feature map, LSH |
A measurably worse-conditioned SimHash: collision curve flatter than Charikar's exact 1 − θ/π, lower recall per candidate at K = 32, 1 400–11 200× slower to encode; its one different property (RBF-like implied kernel) is a classically computable projected-feature artefact; real-gate maps make ⟨Y⟩ ≡ 0 so a third of weight-1 hashes are constant |
| 19 |
HP lattice protein folding: penalty vs penalty-free QAOA |
QAOA with diagonal objectives, exhaustive landscape audit |
Feasible fraction of the turn encoding falls to 8 % at N = 10 (15 qubits); the literature-default penalty λ = 2 has an infeasible/wrong minimiser on half the N ≥ 9 sequences while the penalty-free objective is right on 83–100 % untuned; QAOA p = 3 reaches P(optimal) 1.1–1.5 % at N = 10 (7–9× uniform) and mostly learns feasibility, not ranking; DFS solves all in ≤ 3.6 ms |
| 20 |
Quantum-walk kernels for code-clone detection |
CTQW / QJSD graph kernels (classical simulation) |
Label-aware classical kernels (WL, shortest-path, even bag-of-node-types) reach AUC 1.000 on held-out families with size-matched negatives; the best quantum-walk kernel 0.970 and 0.91 on Type-3 clones, at 10–500× the cost; structure-only spectra are the wrong invariant for clones; corpus is too easy to rank strong kernels |
| 21 |
Quantum Betti numbers for dependency cycles |
LGZ / QPE on the simplex register, clique complexes |
QPE accuracy is set by the spectral gap, not samples (m = 3 ancillas floor at 0.23 error); sizing QPE needs λ_min from the same classical diagonalisation; a Chebyshev-filter classical estimator is more accurate at 6× fewer operator applications; β₁ of dependency graphs at n ≤ 10 is mostly a degree-sequence statistic (7/8 inside the rewiring null) |
| 22 |
VQLS for the pressure-Poisson step of a CFD loop |
variational quantum linear solver, LCU Laplacian |
Nothing accumulates over 50 projection steps (VQLS error stays at 1e-11–1e-9 like the direct solve) but 50 steps cost 1.1×10⁵ cost evaluations ≈ 9.4×10⁷ Hadamard tests vs 167 CG matvecs; the periodic Laplacian LCU has 3N/2 − 1 Pauli terms so one cost evaluation is Θ(unknowns); only L-BFGS on exact gradients trains the 6-qubit system; a random RHS of equal norm raises the best cost from 1e-14 to 0.3–0.4 — it works because the pressure field is smooth |
| 23 |
MPS anomaly detection on log features |
matrix-product-state density model (quantum-inspired) |
Bonds matter (χ = 1 → 2: +0.08 AUC) but bond dimension does not (χ = 2…16 flat; effective bond dimension 1.96); a BIC-selected Gaussian mixture beats every MPS (0.894 vs 0.826 AUC) at 12× cheaper scoring; per-site entanglement entropy does not rank feature importance (Spearman 0.05); the MPS-from-data sketch is provably a product-cosine Parzen KDE |
| 24 |
Gradient-inversion attacks on quantum federated learning |
VQC client, DLG / cosine inversion, Jacobian identifiability |
"Inherent privacy" decomposes into two classical facts: the readout light cone makes only min(d, 2L − 1) features identifiable, and at full rank a single example is inverted exactly (MSE 0) like the MLP; at B ≥ 4 the VQC resists only because the attack landscape is non-convex; 1 000-shot gradients change nothing, DP σ = 0.1 blocks the VQC but not the MLP |
| 25 |
Quantum shift-search for barcode demultiplexing |
Grover over alignment offsets, barrel-shifter oracle, decision-diagram simulation |
Oracle A (classical-data) is a compiled lookup table; oracle B (quantum-data, Fredkin barrel shifter + comparator) verified to 1e-15 on statevectors and to shot noise on MQT DDSIM up to 79 qubits in ≤ 32 ms; per-read quantum cost 0.4–4.6 ms at 10⁶ T/s vs 37–67 ns for str.find; N ≤ 16 offsets so √N is irrelevant — DDs are the right validator, not a speedup |
| 26 |
Register allocation as graph colouring |
one-hot QUBO (equality penalties), QAOA, SA, qudit-inspired mean-field |
All 45 real (bytecode-liveness) interference graphs are chordal and dense (ω ≈ 7.5); exact spill B&B takes ≤ 0.5 ms; Chaitin–Briggs 97–99 % of optimal, mean-field 97–99 %, SA-QUBO 89–99 % at 100–600× the cost; the equality-constrained QUBO passes the feasibility audit exp02's encoding failed; QAOA needs a warm start to sample any valid allocation, and 2 000 random bitstrings through the same repair are level with QAOA on two of its three instances; solver-seed reruns move the SA / mean-field means by ≤ 0.008 but the 3-instance QAOA ratio by −0.14 |
| 27 |
Fuzzing seed-corpus minimisation as set cover |
exactly-one vs slack QUBO, simulated bifurcation, SA, QAOA vs afl-cmin, greedy, exact ILP on real edge coverage |
The set-cover reductions solve half of 18 corpora outright and force 74 % of every minimum corpus; the ILP takes 0.12 s (1 s at 10 000 seeds, 16–26 s at 100 000); afl-cmin is 1.30× the minimum count (3× on html.parser, 5× at 100 000 seeds) but within 4 % on bytes; the exactly-one encoding's minimiser is not a cover when an edge must be covered twice; slack encoding 4 679 → 70 qubits after reduction; SB/SA reach 0.86–0.94 after repair and never the optimum on the html.parser cores; minimum corpora differ by ≤ 0.01 Jaccard, so fair sampling has nothing to choose between |
| 28 |
LLVM phase ordering as a permutation QUBO |
precedence / position surrogates, one-hot permutation QUBO (SA), warm-started QAOA vs exhaustive opt tables, random search, GA, hill climbing |
On 12 C kernels × 720 orders of 6 LLVM 23 passes the worst order is 1.20× the best and 15 % of orders are optimal; the -O2-relative default is 0.88 and never optimal, adjacent-swap hill climbing never leaves it; a pairwise-precedence surrogate explains only R² = 0.69 of the landscape yet its argmin is optimal on 12/12 programs, and the 36-qubit QUBO reproduces it — but random search with 20 evaluations already reaches 0.991 and a GA 1.000 at 120; one fixed consensus order is within 1 % of per-program optimal (leave-one-out 0.984); QAOA at 16 qubits needs the warm start to return a permutation at all (P = 0.01 vs 0.56) |
| 29 |
OCC commit ordering as a QUBO |
linear-ordering proxy vs Lucas FVS encoding (SA, QAOA) vs FIFO, writers-last, greedy FVS, exact, strict 2PL wait-die |
Minimum aborts = minimum feedback vertex set of the write→read digraph; FIFO aborts 2–5 more transactions per 16-batch than needed, greedy FVS is within 0.4 of optimal in µs; the FVS QUBO is exact and SA solves it on 100 % of batches to 144 qubits, the linear-ordering proxy (edge count, Spearman 0.31–0.95 with aborts) only 20–100 %; optimal ordering saves 25–35 % of re-executions; QAOA needs the greedy warm start (P(valid) 0.002–0.10 vs 0.41–0.91) |
| 30 |
Wi-Fi bonded-channel assignment |
spectrum-assignment one-hot QUBO (SA, QAOA at 18 qubits) vs first-fit, DSATUR, local search, exact ILP |
The optical RSA QUBO ports exactly (audited) but SA on it reaches 0.52–0.80 of optimal while SA over assignments on the same objective reaches 0.97–1.02 in ≤ 0.2 s; ILP exact to 20 APs (≤ 6 s), times out at 40 where local search beats its incumbent; DSATUR 0.71–0.97; first-fit 6–9× optimal interference; QAOA returns the DSATUR warm start |
| 31 |
Cold-start seed-item selection as a QUBO |
2nd-order inclusion–exclusion surrogate QUBO (SA, SB, QAOA) vs greedy, popularity, local search, exact ILP; extreme-value estimation |
Greedy is optimal on 12/12 (ILP ≤ 0.13 s), popularity within 3 %; the QUBO is a truncated inclusion–exclusion whose argmax covers 0.81–0.89 of optimal at ≥ 40 items and whose rank correlation with coverage falls to 0.50 near the optimum; SA/SB reach 0.69–0.85 with 2–4 % of trajectories feasible; Dannenbring / Weibull estimates of the QUBO optimum from 8–21 valid restarts miss by up to 0.22 either way; QAOA returns the greedy warm start |
| 32 |
Platformer level generation with a quantum reservoir |
8-qubit QRC readout vs Markov 1–3, ESN, WFC; playability filter |
An order-2 Markov table is the best next-column model (NLL 1.263 vs 1.296 for the best QRC, 30 000× slower per column); unfiltered the reservoirs are 15–39 % playable vs 58 % Markov-2 and 77 % WFC-3, their extra "novelty" being illegal transitions; under the playability filter all models tie on fidelity (JS ≤ 0.002) and the reservoirs keep 8–20 % novel legal 3-grams vs 1–6 % |
| 33 |
DisCoCat readers on commit messages |
spider / stairs / cups quantum readers (parameter-shift + Adam) vs TF-IDF LR, embedding MLPs, 3-d linear control on 3 053 real conventional commits |
TF-IDF + LR 0.749 accuracy (0.716 on the readers' 900 messages); only the cups reader generalises (0.664, level with a 3-d linear embedding of equal parameter count), spider/stairs overfit below the majority rate; cups is the only model whose accuracy collapses when words are shuffled (0.664 → 0.495): it encodes adjacency the task does not reward; masking label words changes nothing; 37–61 s per reader vs 0.1 s |
| 34 |
Noise budget for warm-started QAOA |
WS-QAOA and standard QAOA under depolarising + readout noise on the density backend; ZNE by unitary folding; readout inversion; warm start and uniform sampling as references |
WS-QAOA's P(optimal) edge over its own warm-start product state (1.5–2.5× noiseless) is gone at p1 ≈ 0.006 (p = 1) / 0.004 (p = 2), i.e. after ≈ 1 expected depolarising event per circuit; mitigation moves that crossing 1.6–2.2× but improves estimates, not the samples the device returns; the GW warm start alone is optimal on 9/9 MaxCut instances and beats every quantum arm at every noise level; real ibm_fez run (2026-10-05): the edge over the warm start survives at 9–36 % of its noiseless size and shrinks with Λ_hw |
| 35 |
Simulated quantum annealing on penalty QUBOs |
path-integral Monte Carlo SQA (own numpy) vs dwave-neal SA, own SA, dSB, uniform null, random-restart local search on the real objective, exact (brute force / HiGHS) on one-hot assignment, slack set cover, linear ordering and a MaxCut control |
Encodings are exact on all 12 audited instances; best SQA setting vs SA after repair +0.018 [−0.040, +0.073] (no difference), and SQA loses to its own temperature-annealed control (−0.032) and to random-restart local search at equal wall-clock (−0.038, optimal 36/36); on raw QUBO energy SQA is worse than SA (−0.196) because a fixed transverse-field temperature cannot resolve set-cover's tiny objective coefficients; MaxCut control: every sampler optimal 9/9 |
| 36 |
Random-circuit sampling and XEB at laptop scale |
Haar-SU(2) + CZ brickwork circuits (1-D chain, 2-D grid, n ≤ 24, ≤ 24 cycles): statevector vs quimb/cotengra amplitude contraction, MPS bond dimension, exact noisy XEB (density backend + Pauli-trajectory MC) vs the global-depolarising model, classical spoofers (marginals, depth truncation, patch sampling, MPS) |
In contraction cost 1 000 TN amplitudes undercut one statevector by 10^0.5–10^4 up to depth 12–16, but in wall-clock the statevector wins everywhere up to n = 24 and TNs only pay off past the 16 GB memory wall; MPS needs χ > 64 on the grid from depth 8; shallow circuits are not Porter–Thomas (noiseless XEB 7–66) and a post-selected patch sampler beats the noiseless device there; only deep grid circuits (depth 12) make XEB meaningful, where the noisy device beats a χ = 64 MPS only for p1 < 1.2·10⁻³ and the cheap spoofers for p1 < 2.7·10⁻³; the global-depolarising XEB model underestimates the exact value by up to 10⁴×; real ibm_fez run (2026-10-05): f = 0.963 per cycle on a 20-qubit chain, below MPS χ = 64 at every depth |
| 37 |
QEC decoder benchmark |
PyMatching MWPM (plain + correlated), BP (min-sum, product-sum), BP+OSD-0 / OSD-CS (ldpc), exact minimum-weight ILP (HiGHS) on stim rotated-surface-code and repetition-code memory experiments and on the bivariate-bicycle [[72,12,6]], [[108,8,10]], [[144,12,12]] codes, paired McNemar tests on identical shots |
On the surface code at d ≥ 5 correlated MWPM and BP+OSD-CS are statistically tied while BP+OSD costs 690–2 900× more per shot (3.9–50 ms vs 6–17 µs); BP alone has no threshold (LER grows with d) and BP+OSD with the default min-sum scaling is 3.5× worse than MWPM at d = 3, p = 10⁻³ (BP converges to heavy corrections so OSD never runs); on the BB codes, where matching does not apply, BP+OSD-CS is 2–6.5× better than OSD-0; the exact ILP beats BP+OSD only at low p and costs 10⁴–10⁵× more; thresholds 0.63–0.88 %; real ibm_fez run (2026-10-05): one-round repetition code below threshold, Λ = 4.0 [3.1, 5.3] at p = 0.028 |
| 38 |
Hamiltonian simulation: Trotter, entanglement, MPS / sparse Pauli dynamics, noisy device |
TFIM Trotter (orders 1–2) vs expm_multiply on chains n ≤ 16; S_half and the exact MPS bound χ_ε on n ≤ 20; the kicked-Ising utility circuit on a 21-qubit heavy-hex patch vs quimb MPS (χ ≤ 128), own sparse Pauli dynamics (δ ≥ 10⁻⁴), exact density-matrix noise at n = 8 and Pauli-trajectory noise at n = 21 with ZNE by folding |
First-order Trotter's error on Z-basis observables is O(dt²) (slope 2.00) and 2× below second order's at equal RZZ count, by time-reversal symmetry, while infidelity scales dt² vs dt⁴; a noisy device's optimal step count is set by ≈ 1 expected error event (k* 32 → 4 for p1 0 → 10⁻²) and Richardson ZNE helps only there; on the 21-qubit heavy-hex patch after 20 kicks at θ_h = 0.8 (S_half 5.5 bits, χ₉₉ > 128) a device with p2 = 10⁻² misses ⟨Z_b⟩ by 0.13 ± 0.02 (ZNE 0.08 ± 0.03), better than MPS χ = 64 (0.16) but 4× worse than sparse Pauli dynamics with 643 terms in 30 ms (0.03); only p2 = 10⁻³ (0.017 ± 0.009) edges the cheap approximations, by 1.5 σ; the exact statevector costs 2 s; real ibm_fez run (11 QPU s, 2026-10-04): raw ⟨Z_b⟩ error 0.037 ± 0.016 at step 20 (ZNE 0.016 ± 0.024), level with SPD, but the magnetisation decays at the calibrated CZ error rate (≈ 2–3 × 10⁻³ per CZ) and is level with MPS χ = 64 — the single-site number is the flattering one; DD made it worse and ibm_marrakesh reproduced ibm_fez within 1–2 σ (2026-10-05) |
| 39 |
Kicked Ising on the full 151-qubit heavy-hex chip: the utility circuit with referee classes |
Experiment 38's kicked-Ising step (RZZ(−π/2) as one CZ per edge over the three edge colours, RX(θ_h)) on the calibration-screened device graph (151 qubits, 164 couplers, RCM bandwidth 10), k ∈ {5, 10, 20}, θ_h ∈ {0.4, 0.8, π/2} + a folded copy for 2-point ZNE; observables ⟨Z_b⟩, M, weight-10 and weight-17 Z strings and the 14 plaquette Z-stabilisers; referees by class: Stim at θ_h = π/2, light-cone statevector / MPS χ = 256 at k ≤ 5, sparse Pauli dynamics with multi-word masks (δ ≥ 3 × 10⁻⁴) and quimb MPS χ ≤ 128 beyond; calibration Pauli model through noisy SPD and Stim, Aer rehearsal on a 10-qubit sub-lattice |
Of the 48 (circuit, observable) pairs of the default job 22 are exact, 3 bounded and 23 unreferenced: SPD is within 0.0045 of exact up to k = 6 but moves by 0.036 between its two smallest δ in the utility cell (θ_h = 0.8, k = 20: SPD 0.220 vs MPS χ = 128 0.643, whose fidelity estimate is 4 × 10⁻⁶), so that cell has no referee; at θ_h = π/2 every Z observable is exactly 0 and a dead device would pass — the plaquette stabilisers (exactly +1, predicted 0.30 at k = 10 and 0.11 at k = 20 on the twin) are the real check; the calibration Pauli model matches the fake device within 1–2 SE at k ≤ 5 but the fake device keeps more signal deeper (λ ≈ 0.74) because T1 relaxation pulls toward |0…0⟩, which a Pauli model cannot express; the full-chip exact ⟨Z_b⟩ at k = 5 (0.563) exceeds 38's 21-qubit patch value (0.469, boundary effects); real ibm_fez run (15 QPU s, 2026-10-08, §5.9) on the live 154-qubit graph (bulk qubit 69, 12 plaquettes — not the fake twin's graph, so the pre-registered fake numbers are orientation only): exact k = 5 ⟨Z_b⟩ within 1.6 SE (0.908 vs 0.918 at θ_h = 0.4, 0.488 vs 0.469 at 0.8), plaquette stabilisers 0.383 / 0.189 at k = 10 / 20 vs the live calibration model's 0.428 / 0.224 (λ ≈ 1.1), but M keeps 95–98 % where the model predicted 86–90 % (λ = 0.14–0.44: the depolarising form is wrong for ⟨Z⟩, T1 or Z-biased errors fit), more per qubit than 38's patch on the same physical qubits; ZNE at the only exactly checkable folded point moves ⟨Z_b⟩ from 0.019 to 0.013 off exact; in the utility cell (θ_h = 0.8, k = 20) the device (0.316 ± 0.015, ZNE 0.354 ± 0.024) sits between SPD (0.18–0.21, still drifting with δ) and MPS χ = 128 (0.64, F ≈ 10⁻⁵), 0.10 above SPD — unreferenced, no win claimed (29 tests) |
| 40 |
Long-range GHZ states with dynamic circuits: constant depth is not constant time |
GHZ on n = 8…100 qubits of one SWAP-free Heron path, three protocols: centre-out CX ladder (CZ depth n/2), the Bäumer et al. dynamic circuit (Bell pairs, mid-circuit ZZ fusion, if_test XOR feed-forward, reset and reuse; CZ depth 3 at every n) and the same circuit with the Pauli frame deferred to post-processing; Tóth–Gühne and witness fidelity bounds, parity oscillations at n = 8; Stim exact referee, a calibration-derived Stim noise model and an Aer MPS rehearsal on fake:fez |
On the fake:fez model the ladder wins wherever anything is certifiable (F_pop 0.83 / 0.58 / 0.15 vs 0.66 / 0.32 / −0.14 for the dynamic circuit at n = 8 / 16 / 32): constant depth lasts 5–6 µs because half the register idles through a 1.6 µs readout plus a 1.6 µs reset, and the dynamic circuits overtake the ladder only at n ≈ 64–192 where ⟨X^⊗n⟩ ≤ 10⁻² certifies nothing; skipping the reset is the largest single gain and a depth-(n − 1) ladder baseline would have reversed the conclusion; entanglement is certifiable only at n = 8 (all protocols) and n = 16 (ladder); the rehearsal agrees with the model within |z| ≤ 2.1; real ibm_fez run (34 QPU s, 2026-10-08, §5.8): the ordering ladder > deferred frame > dynamic held at every n, but all three protocols fell 10–21 SE below the calibration model at n ≤ 32, so genuine multipartite entanglement was certified only for the ladder at n = 8 (F_pop 0.712 [0.697, 0.726]; dynamic 0.427, deferred 0.467; nothing at n ≥ 16 without readout mitigation); the dynamic arms' extra loss sits at the mid-circuit fusion boundaries (≈ 3–7 % extra flips per boundary, not at the reset qubit), the measured feed-forward latency is 0.97 ± 0.31 µs (1 µs assumed), B matches C within 1 SE up to n = 64 (≈ 0.8 % misfires per block at n = 100), three qubits carry the worst generators in every circuit, and the fake twin was optimistic by 3–10 SE (123 tests) |
| 41 |
Multi-round repetition-code memory with dynamic circuits |
Stim's repetition_code:memory (d = 3–13, R = 1–8 rounds; d ≤ 25, R ≤ 16 with --full) converted gate by gate into the qiskit hardware circuit, so the decoder's detector-error model is the circuit that runs; mid-circuit syndromes with ancilla reset on a calibration-screened 25-qubit line; MWPM (PyMatching) with a fitted uniform prior and a state-aware calibration prior, majority-vote baseline, d = 3 real-time lookup-table correction through nested if_test; Aer MPS rehearsal |
Prediction on the fake:fez line: Λ = 6.8 [6.0, 7.9] for |1…1⟩ (ε = 5 × 10⁻³ per round at d = 3) and Λ ≈ 43 for |0…0⟩, which rarely fails; the symmetric Pauli twirl of T1 misses the rehearsal by 2× (|1⟩) and 10× (|0⟩), so the decoder prior had to be state-aware; the per-round lookup an active controller can afford has no threshold (Λ ≈ 0.8–1.0, 16× worse than matching at d = 3, R = 4), so the feed-forward circuits measure what real-time correction costs, not a better memory; rehearsal matches Stim within |z| ≤ 1.5; 29 circuits ≈ 49 QPU s pending; the heavy-hex surface code was skipped (38 tests); real ibm_fez run (34 QPU s, 2026-10-08, §5.9; the first submission was rejected with error 1524 because the active circuits nested if_test blocks — IBM allows no nested conditionals — and 0 QPU s were billed): |0…0⟩ failed 2 times in 96 000 shots (better than the referee), |1…1⟩ gave ε per round 1.08 × 10⁻² / 2.15 × 10⁻³ / 6.3 × 10⁻⁵ at d = 3 / 7 / 13 and Λ = 2.38 [2.06, 2.81] against 6.67 [4.81, 10.1] predicted from that day's calibration (majority vote 1.68): |1⟩ data decay 1.5–1.9× faster per round than the calibrated T1 (as if idling 5–6 µs instead of 3.3), gate errors depend on the data state (0.4–1.1 % mid-round in |1⟩ vs 0.16 % modelled), and leakage-like ancillas stuck at 1 add correlated errors (dropping those shots gives Λ = 3.4); the mid-circuit reset adds no detectable error and detector rates are flat over rounds; the flat real-time correction fails 8.6 % vs 5.0 % for its software twin — a 2–3 µs idle window per round, equal to doing nothing; a device characterisation, no advantage |
| 42 |
Holographic qubit-reuse sampling of 200-site matrix-product states on 2–3 qubits |
Sequential (holographic) circuits: one physical qubit reused L = 25…200 times by mid-circuit measure + reset with a 1–2-qubit bond register; a χ = 2 Markov family with closed-form correlation length (GHZ and ξ = 5, Z- and X-basis memory) and an optimised χ = 4 sequential circuit for the TFIM at g = 1.2 (energy density within 3 × 10⁻⁶ of exact); exact MPS referee by transfer matrix, per-reuse noise fit, a unitary L-qubit baseline on a path; Aer rehearsal |
On fake:fez the error per reuse is 1.66 ± 0.03 % (Z memory) and 1.57 ± 0.03 % (X), equal to the calibration's T1/T2-implied 1.67 / 1.56 %; the ξ = 5 state keeps a 3-site TVD of 0.042–0.046 at every L up to 200 (the error is bounded by ξ) while GHZ accumulates (0.08 → 0.27); the holographic circuit beats the unitary chain at equal L (TVD 0.04 vs 0.10–0.18 at L = 25–100) because the chain pays 101 readouts and up to 42 µs of idling; real ibm_fez run (23 QPU s, 2026-10-08, §5.9): Z memory loses 2.34 ± 0.05 % per reuse (model 2.18 %) but X memory 9.50 ± 0.15 % (ε_X 6× the T2 value, flat in L and along the chain), equivalent to a ≈ 0.40 rad phase on the bond qubit per measure + reset of its neighbour; the χ = 4 state shows the kick is coherent (decay ratio 1.07–1.08 vs model 0.98 — reproduced by the phase, while every stochastic channel gives 0.93–0.95) and GHZ-X's late bias 0.12–0.14 needs an outcome-dependent kick (joint 7-parameter fit χ²_red 1.24 vs 498 for the calibration model); the fit's 5 % '0 after 1' readout error is a reset that fails after a true 1 (3.0 ± 0.1 %, static readout at the calibrated 1.1 %, glitch rate growing 2.8 → 5.5 % along L = 200); the ξ = 5 state keeps TVD3 0.075–0.088 flat to L = 200 (model 0.062) and reuse beats the 26–101-qubit unitary path 1.6–2.9×; everything is an exact-referee comparison, no advantage (31 tests) |
| 43 |
Daily whole-chip noise census and 28-day drift |
One daily SamplerV2 job of 32 circuits × 1 000 shots on all 156 qubits: simultaneous single-qubit RB, CZ-layer mirror circuits on the three heavy-hex edge colours (Stim-exact targets), a readout census with pairwise flip correlations against the independence null, T1 / T2* at three delays, and experiment 38's 21-qubit kicked-Ising witness; aggregation over days (scripts/aggregate.py), scripts/qpu_daily.sh for a cron entry; rehearsed on the four fake twins as pseudo-days |
On the four fake twins every estimator recovers the twin's own calibration (ratios RB 0.96–1.11, CZ-layer 0.97–1.03, readout 0.92–1.01, T1 / T2 0.99–1.04; 0 correlated pairs after Bonferroni), so the method is validated and the device numbers are pending; from records already in the repo the frozen twins are stale: on the qubits 34 / 36 / 37 / 38 used, their CZ error is 1.2–2.3× the run-time calibration (median 1.46×) and the fez twin marks one of 38's couplers broken; ≈ 11 QPU s per day, 52 % of an Open Plan window over 28 days (24 tests); real ibm_fez census day 1 (11 QPU s, 2026-10-08, §5.6): readout matches its calibration (P(0|1) ratio 1.00 [0.93, 1.05]) and is independent (joint flips 1.002 [0.990, 1.013] over 11 670 far pairs, 0 Bonferroni pairs), but gates do not: RB error is 1.67 [1.51, 1.82]× and CZ-layer error 1.29 [1.21, 1.40]× the calibration, CZ-block failures cluster (dispersion D up to 1.30 [1.21, 1.43]), spectators lose 8.6 × 10⁻⁴ per layer beyond their own RB error (16× the twin's floor), T1 matches on the median but 122 of 156 qubits miss their own value, Ramsey T2* is 0.27× the echo T2; the stale fake:fez twin matches CZ layers only by cancellation (0.98 [0.93, 1.10]) and is 2.1× optimistic on single-qubit gates, 0.29× on T2* and 4.9× on spectator error; 41 bad readout qubits (7 new), qubit 72 dead, coupler 149–150 failed outright; witness ⟨Z_b⟩ 0.268 ± 0.030 vs exact 0.248 (38 measured 0.211 ± 0.015 four days earlier); 6 more daily runs fit the window at 1 000 shots |
| 44 |
Sample-based quantum diagonalisation of N₂ at 20 and 32 qubits with LUCJ circuits |
pyscf N₂/6-31G, (10e, 16o) at 1.0977 and 1.5 Å (32 qubits) and (10e, 10o) (20 qubits); LUCJ from CCSD amplitudes (ffsim) on a SWAP-free heavy-hex zig-zag with bridge ancillas for the α–β terms (700 CZ on 35 qubits); qiskit-addon-sqd configuration recovery + subspace diagonalisation; referees exact FCI (39–54 s, symmetry-adapted) and CCSD(T); equal-size controls: noiseless samples, calibration noise model, Aer trajectories, uniform random bitstrings, a CISD sampler, exact ground-state samples, pyscf selected CI |
The heavy-hex-local LUCJ state is 99.5 % Hartree–Fock (the exact ground state is 89.5 %), so even perfect samples leave SQD 115 mHa above FCI; the noisy device model (18.9 mHa), Aer trajectories (15.1) and uniform random bitstrings (16.9) are indistinguishable at 1.6 × 10⁵ determinants, where classical selected CI reaches 2.75 mHa in 1.5 s and CCSD(T) is 1.8 mHa off; the only hint of signal is at 20 qubits (7–8 mHa vs 14 for uniform, few trajectories) and is the hypothesis the device run tests; real ibm_fez run (62 QPU s, 2026-10-08, §5.8): 6.3–6.5 % of shots land in the right electron sector at 32 qubits (14× uniform) and 22 % at 20 qubits, ancilla flags 20 % / 12 %; SQD on the device samples beats uniform random bitstrings by 7–14 mHa at 32 qubits (14.2 [13.2, 15.3] vs 23.1 mHa at equilibrium, 18.2 vs 28.3 at 1.5 Å; equal shots, K and dimension, disjoint-block CIs exclude 0) and the gain is neither the sector filter nor the excitation rank but the per-orbital flip pattern around Hartree–Fock, which follows the CCSD occupation shifts (ρ = 0.8–0.9) — so a classical HF + CCSD-shaped bit-flip sampler with three fitted numbers ties the device (±0.4 mHa) and selected CI at the same dimension is 3–5× more accurate (2.75 vs 14.2 mHa); the --job rebuild five hours later picked other qubits (counts unaffected, layout now pinned from the pending record); no chemistry advantage (15 tests) |
| 45 |
Sample-based Krylov quantum diagonalisation of spin chains at 12–48 qubits |
Second-order Trotter states U(t_j)|ψ₀⟩, j = 0…4, of the XXZ chain (Néel start, U(1) sector) and the critical TFIM (|0…0⟩ / |+…+⟩, Z₂ flip-symmetrised) on SWAP-free paths n ∈ {12, 24, 48} (+ a 40-qubit heavy-hex patch with --full); diagonalisation in the span of the sampled configurations; referees ED, sector Lanczos (dim 2.7 × 10⁶), free fermions and DMRG χ = 128 agreeing to 10⁻¹⁰; equal-size controls: exact ground-state samples, noiseless Trotter samples, a depolarising + readout model, uniform in sector, CIPSI selected CI |
At n = 48 and equal subspace size the relative energy error is 0.26 (noiseless Trotter samples) / 0.27 (noise model) / 0.20 (CIPSI, classical) / 0.46 (ideal ground-state samples) / 0.70 (uniform) for XXZ and 0.099 / 0.11 / 0.079 / 0.16 / 0.59 for the critical TFIM: classical selected CI beats every sampled subspace from n = 24 on, Trotter samples beat ground-state samples because they pile onto H-connected configurations, configuration recovery hurts at half filling, and noise mostly costs distinct configurations (2 789 vs 8 845 at equal shots); the fake:fez rehearsal at n = 12 matches the noise model within 0.05 in the in-sector fraction; real ibm_fez run (33 QPU s, 2026-10-08, §5.9): at n = 48 and equal subspace size the device samples give a relative energy error of 0.291 (XXZ) / 0.106 (TFIM) vs 0.268 / 0.099 for noiseless samples, 0.274 / 0.110 for the noise model and 0.207 / 0.079 for CIPSI; the TFIM samples sit within 0.008 of noiseless and beat the depolarising model, the XXZ samples are 0.02–0.03 worse because they leak out of the Néel sector about twice as fast as calibrated (in-sector 0.089 vs 0.23 predicted at j = 4, ≈ 5 × 10⁻³ per CZ, 74 % of the extra excitation on six qubits incl. 143 on all three chains; 2 336 distinct configurations vs 8 846 noiseless); CIPSI matches the device's best XXZ subspace with 94 configurations instead of 2 336 and wins every case by 1.3–4.7×; the --job rebuild five hours after submission recorded the wrong paths (counts are site-indexed, energies unaffected); no advantage claim (35 tests) |
| 46 |
Verifiable random-circuit sampling at 20–151 qubits on heavy-hex |
Experiment 36's Haar-SU(2) + CZ brickwork over the three heavy-hex edge colours on nested calibration-screened regions of 20 / 44 / 88 qubits (151 with --full): (A) mirror circuits with a random Pauli frame (exact target bitstring at any n), (B) 20-qubit patch XEB with the cut CZs removed, (C) Clifford mirror cores with k ∈ {0, 4, 8} T gates (exact distribution by Pauli propagation in Stim); calibration digital-error model, an exact Pauli-noise twin (Stim, Clifford + T trajectories, MPS), per-cycle fits, chip-scale extrapolation with Monte-Carlo CIs, tensor-network cost of one amplitude |
On the fake:fez twin the error per CZ is flat across sizes (7.7 / 7.6 / 8.1 × 10⁻³ at n = 20 / 44 / 88, twice 36's measured 4.0 × 10⁻³ on the real device) and the Haar-mirror rehearsal at n = 20 agrees (f = 0.958 ± 0.009 vs 0.950); patch XEB is biased high by 1.15× (d = 4) to 2× (d = 16) and is an upper bound only; the chip-scale depth-8 fidelity extrapolates to 0.034 [0.025, 0.045] vs 0.028 from the calibration model and 0.030 measured directly on the twin, with a 1.7× systematic when shallow points enter the fit; one amplitude of the 151-qubit depth-8 circuit contracts at width 12 in 8 ms, so nothing here is classically hard; real ibm_fez run (25 QPU s, 2026-10-08, §5.9): on regions of 24 / 44 / 99 qubits the mirror circuits give an error per CZ of 6.6 [4.1, 9.1] and 7.4 [5.0, 9.7] × 10⁻³ at 24 and 44 qubits (the twin's ≈ 7.7, not 36's 4.0; ≥ 8.4 × 10⁻³ at 99 qubits, where d ≥ 8 had 0 successes in 4 000 shots), 36's own chain circuit drops to 0.539 ± 0.024 from 0.679 ± 0.030 on other qubits, decoupled patches show no crosstalk (|z| ≤ 2.1) but qubits at the boundary of a growing region lose 1.7–4× the twin's polarisation, T gates cost nothing, the decay is curved (slope 9.5–9.9 × 10⁻³ at d = 4 → 8 vs 5.0–5.7 × 10⁻³ at d = 8 → 16), and the chip-scale depth-8 fidelity extrapolates to 0.046 [0.022, 0.097] (0.014–0.10 depending on the slope used); no advantage claim (25 tests) |
| 47 |
QAOA on the native heavy-hex graph at 20–144 qubits: fractional RZZ vs CZ |
±J and weighted Ising spin glass on calibration-screened BFS balls of the device graph (20 / 50 / 100 / 144 qubits, no routing); QAOA p ≤ 3 with light-cone-exact angles and energies (cones ≤ 15 qubits) and experiment 01's transferred angles as control; HiGHS MILP ground states (≤ 0.12 s at n = 144), simulated annealing, 1 + 2-flip local search, uniform sampler; identical logical circuits compiled with Heron's single-pulse fractional RZZ(θ) and with the two-CZ decomposition on the same physical qubits; exact light-cone density-matrix noise predictions, Aer MPS rehearsal |
Exact QAOA reaches r = 0.48–0.52 (p = 1), 0.63–0.71 (p = 2) and 0.72–0.76 (p = 3) while SA and local search reach ≥ 0.97 at the 4 000-shot budget and HiGHS proves optimality in milliseconds; by gauge symmetry the ±J energy equals the unfrustrated MaxCut energy at these depths, so the circuit cannot see the frustration; the fractional compilation halves the two-qubit count and depth and keeps 0.92–0.94 of the p = 1 energy vs 0.90–0.92 for CZ — an energy-error ratio η ≈ 0.80 raw (0.84 ± 0.06 in the rehearsal) and ≈ 0.5 with readout divided out, because readout is 55–65 % of the p = 1 error; p = 2 predictions at n ≥ 50 are extrapolated from a per-gate fit; real ibm_fez runs (2 × 18 QPU s, 2026-10-08, §5.8): on identical 20 / 50 / 100-qubit circuits the device keeps 0.89–0.92 (fractional) vs 0.88–0.92 (CZ) of the exact energy at p = 1 and 0.87–0.91 vs 0.84–0.89 at p = 2, 0.01–0.04 below its own calibration model; the paired energy-error ratio is η = 0.86 (95 % CI 0.82–0.91), 0.73 with readout removed vs ≈ 0.5 predicted — fractional RZZ cuts the error by 14 %, not by half, because the fractional ISA loses ≈ 1.45× per two-qubit gate what CZ loses; P(ground) at n = 20 is 1.8 × 10⁻³ (ideal 1.9 × 10⁻³) and 0 at n ≥ 50 while SA and local search hit r = 1 in ≤ 13 ms (20 tests) |
| 48 |
Error-mitigation showdown on the kicked-Ising circuit at 21–100 qubits |
Experiment 38's kicked-Ising step on its 21-qubit patch and on 50- / 100-qubit BFS balls of the device graph (bulk qubit 111 in all three), θ_h ∈ {0.8, π/2}, k ∈ {5, 10}; EstimatorV2 observables: every single Z (⟨Z_b⟩, M), the weight-10 Z string and a Z-type Clifford stabiliser; five arms as separate jobs through `--arm raw |
trex; real ibm_fez runs (15 + 24 + 52 QPU s, 2026-10-08, §5.9): raw and TREX miss the 0.02 target on all 20 certified points (bias 0.07–0.30 at θ_h = 0.8, 0.21–0.78 on the Clifford stabilisers; 1.2–3.7× the model built from the same day's calibration, beyond it on 16 of 16 points); ZNE reaches the target only at 21 qubits (⟨Z_b⟩ +0.014 ± 0.021 at k = 5, M −0.015 ± 0.005; 123–198 QPU s per ⟨Z_b⟩ point to resolve it, 2–9× the prediction) and nowhere at 50 or 100 qubits, where ⟨Z_b⟩ is no better than raw and the Clifford stabilisers keep 0.15–0.41 after ZNE (6–17× the predicted single-qubit floor); IBM's ZNE error bars are the three-point fit's residual rescaled by the one-dof Student factor, the quoted 0.003 on a 100-qubit stabiliser is really ≈ 0.07, and 56 of 708 extrapolations are degenerate '0 ± 0' growing-exponential fits that IBM's success check accepts; PEA and folding did not run (15 tests) |
| 49 |
Clifford measurement dynamics: mid-circuit measurements as a stabilizer-exact error ruler |
Random Clifford brickwork (single-qubit Cliffords + CZ layers) on 20- and 40-qubit SWAP-free chains, each qubit measured in Z after a layer with probability p ∈ {0, 0.1, 0.3}, T ∈ {4, 8, 16} layers, outcomes in per-layer registers, final Z readout; every shot replayed through an affine form of the circuit (checked against Stim tableau and statevector replays, 4 000 shots in 7 ms) to find the deterministic final bits and the forbidden mid-circuit outcomes; a tagged Stim circuit-level noise model from the calibration with gate / idle / readout scale factors fitted to the data; half-chain entropy vs p for context (p_c ≈ 0.11–0.13 for this gate set) |
On the fake:fez model the deterministic-bit accuracy is 0.91–0.99 per circuit (pooled error 0.062 at p = 0.1, 0.064 at p = 0.3; forbidden mid-circuit outcomes 5–8 %), the fitted scale factors close at λ = 1.02 / 0.99 / 0.99 (gate / idle / readout) and the implied per-qubit error is 5.2 × 10⁻³ per CZ layer vs 0.028 per measurement layer — a measurement layer costs ≈ 5.4× [5.0, 5.8] a gate layer, and the hypotheses 'measurement free' and 'measurement × 2' sit 24 and 10 shot SEs away; bootstrap CIs are ≈ 1.5× too narrow (30 / 40 coverage); the circuits stay in the early transient (S_half ≤ 3.9 bits), so nothing is claimed about the transition; real ibm_fez run (22 QPU s, 2026-10-08, §5.7 — measurement layers cost nothing measurable, so the 71 s cost formula was wrong): deterministic bits read correctly 0.876–0.972 per circuit, losing 3.0–3.2 % per layer of age vs 2.2–2.4 % in the calibration model (1.3–1.5×); the three-scale fit's λ_gate ≈ 4–4.8 is a misfit (χ²/dof 22–31, bootstrap CIs ≈ 5× too narrow) — circuits without mid-circuit measurements give λ_gate = 1.73 [1.60, 1.86] and 46's same-day run 1.1–1.2× — so the excess is measurement-associated dephasing (Z part 25× calibration, X/Y 3.3×; gates right after a qubit's own measurement are 5.0× vs 2.75×), uniform along the chain, and the measurement/gate layer cost ratio is model-dependent (1.5–5.4, the twin said 5.4); forbidden mid-circuit outcomes 5.5–8.9 % as modelled; readout drift between snapshots (qubit 23: 0.010 → 0.244) explains the p = 0 anomalies only in part; no claim about the entanglement transition (29 tests) |