Table of Contents
July 30, 2026 – IBM coordinated three announcements around preprints posted to arXiv on July 27 and 28. IBM presented all three as demonstrations that quantum computing had entered “the quantum advantage era.”
The papers do not carry equal evidentiary weight. The IBM/University of Chicago paper makes an explicit quantum advantage claim backed by complexity-theoretic hardness arguments and a device-dependent fidelity certificate. The Qedma and Algorithmiq papers make more empirical claims about regimes where tested classical methods become unreliable. The verification advances are real. Whether three different results at three different evidentiary levels justify declaring an era is another question.
The News
IBM released three papers, each addressing the same fundamental problem from a different direction: how do you trust a quantum computer’s output when no classical machine can check the answer?
The flagship result came from IBM and the University of Chicago. In a preprint posted to arXiv on July 28, researchers Simon Martiel, Ali Javadi-Abhari, Bill Fefferman (UChicago), Jay Gambetta, and collaborators demonstrated what they call “doped Clifford sampling,” a structured alternative to the random circuit sampling (RCS) approach that Google used for its contested 2019 quantum supremacy claim.
The core idea: build a quantum circuit entirely from Clifford gates (which classical computers can simulate efficiently), embed the circuit in a “spacetime code” that detects errors across both qubits and time steps, then strategically inject non-Clifford T gates to push the computation beyond classical reach. Because the T gates are placed where the error-detection checks remain valid, the Clifford reference circuit provides a tractable baseline whose physical fidelity can be measured efficiently, and the syndrome statistics then support a device-dependent lower bound on the fidelity of the hard doped circuit under the experiment’s stated noise assumptions. The experiment ran a depth-70 circuit on 70 data qubits, augmented by 27 syndrome ancillas for 97 physical qubits total. The physical circuit contained 2,869 CZ gates (2,415 in the data computation and 454 for syndrome extraction) along with 468 T rotations implemented virtually through frame updates. IBM describes the data-level computation as operating on “70 logical qubits,” though these are error-detected data qubits, not fault-tolerant logical qubits of the kind assumed in CRQC resource estimates. The spacetime code suppressed the effective gate-error measure by roughly 10x after post-selection, and the team reported a certified state-fidelity lower bound of 0.284 at 95% confidence.
IBM’s Heron processor collected the 2,051 accepted target samples in 16.1 minutes of QPU execution time. The paper estimates that leading current tensor-network and stabilizer-based methods would be infeasible for this finite instance, while leaving open the possibility that improved classical algorithms or hardware could change that comparison.
A second paper, from Qedma Quantum Computing (Tel Aviv) working with RIKEN and BlueQubit, took a different approach. Using Qedma’s QESEM error mitigation software on IBM’s Heron R3 processor, the team tracked how magnetization oscillates in a Floquet Ising model on heavy-hex ladder geometries, at system sizes up to 74 qubits. Exact statevector simulations provided reference results through 35 qubits. At 51 qubits, two sparse Pauli-path calculations retaining on the order of $$10^{12}$$ Pauli strings across 12,888 Fugaku nodes began to diverge from each other after roughly seven Floquet cycles, with a clear phase mismatch against the quantum results emerging around cycle 15. Other approximation methods lost controlled convergence at different points. At 74 qubits, none of the classical approaches tested reliably reproduced the late-time quantum results. The most demanding late-time quantum trajectories at 51 and 74 qubits rely on QESEM-Extrapolated, a lower-overhead heuristic that the authors validate against the unbiased QESEM method over their shared operating range, while acknowledging that extrapolation can introduce bias.
The third paper, from Algorithmiq (Milan, with operations in Finland), used 56 qubits on Heron to estimate an operator Loschmidt echo, a measure of how information propagates through a disordered system. Multiple classical simulation methods produced predictions that disagreed with each other and with the quantum results. Algorithmiq stress-tested its mitigated estimate across five noise settings spanning two IBM Heron processors, a modified gate duration, and two synthetic noise injections. The mitigated results agreed within 0.03 across all five variants, providing evidence that the global-rescaling result is stable under the hardware and synthetic noise variations tested. That consistency does not provide a numerical accuracy guarantee for the hardest target result: the authors describe global rescaling as a heuristic whose estimate lacks a quantitative bound on its distance from the unknown ground truth.
The corresponding problem instances, circuits, and results are represented on the Quantum Advantage Tracker, a community benchmarking initiative that IBM, Algorithmiq, the Flatiron Institute, and BlueQubit helped launch in November 2025. The tracker treats these entries as active candidates requiring further benchmarking, not as adjudicated demonstrations. All three papers are preprints; none has undergone peer review.
“We are now firmly in the quantum advantage era,” said Jay Gambetta, Director of IBM Research and IBM Fellow, in IBM’s announcement.
My Analysis
All three papers answer the same question: how do you trust a quantum result when you can’t classically check it? That is verification, and verification is the hard problem I’m glad people are working on. Each paper represents progress. But working out what each result actually means, separate from what IBM says it means, is a task on its own.
The verification innovation is the real contribution.
For years, quantum advantage demonstrations have been stuck in a paradox: if no classical computer can reproduce the result, how do you know the quantum computer got the right answer? Google’s 2019 Sycamore experiment used cross-entropy benchmarking (XEB), which required assumptions about hardware noise that critics could (and did) challenge. Every subsequent RCS-based claim inherited that same vulnerability.
What Martiel, Javadi-Abhari, Fefferman, and colleagues did is clever. By starting from an encoded Clifford reference circuit and doping it with T gates only at check-preserving locations, they created a paired verification scheme. Direct fidelity estimation measures the Clifford reference state efficiently. It does not directly measure the hard doped state. Instead, the syndrome data and an analysis of how accepted faults behave before and after doping transfer that information into a lower bound on the target circuit’s fidelity. That is still a significant result: it replaces an inaccessible direct fidelity calculation with a statistically testable certificate under explicit assumptions.
The approach connects to several capabilities in my CRQC Quantum Capability Framework. Spacetime codes represent a creative application of quantum error correction principles in the form of error detection with post-selection: the circuit discards runs with nontrivial syndromes rather than correcting faults and continuing computation. This is not a below-threshold demonstration (which would require evidence that logical error rates decrease as code distance increases), but it does show that encoding-aware circuit design can extract meaningful error suppression from hardware that is not yet capable of fault-tolerant operation.
The three papers make different claims. IBM’s campaign blurs the distinction.
The IBM/University of Chicago paper is not declining to claim quantum advantage. It develops a quantum advantage protocol, labels its flagship result a quantum advantage experiment, and argues that doped Clifford sampling offers a compelling route to advantage on classical-hardness grounds. Its caution is narrower but worth noting: the authors estimate that current leading tensor-network and stabilizer methods cannot reproduce the finite experimental instance, while acknowledging that future classical algorithms and hardware may change that comparison. They leave open the possibility of faster simulations with improved algorithms. That is normal and responsible scientific practice. It is also a material caveat in a field where classical simulation researchers have repeatedly narrowed or eliminated advantage claims after publication.
Qedma makes a different kind of claim. Its paper says the experiment accessed a regime beyond the state-of-the-art classical simulations considered, but distinguishes that empirical result from an asymptotic complexity-theoretic separation. The abstract frames the work as establishing “error-mitigated quantum processors as quantitative scientific instruments.” It does not present a formal quantum advantage proof.
Algorithmiq is more cautious again. Its paper argues that the quantum estimate is the most credible among the methods examined and presents a framework for accumulating trust when no classical ground truth is available. That is not the same evidentiary standard as the UChicago fidelity certificate, and it does not establish an exhaustive classical separation.
As Dominik Hangleiter, an Ambizione fellow at ETH Zurich, told IEEE Spectrum about the Qedma and Algorithmiq results: both papers amount to saying that their results seem hard to simulate classically, which he considers appropriate. As Jens Eisert of the Free University of Berlin noted, establishing quantum advantage is an ongoing process of building confidence, not a binary threshold to cross.
IBM’s announcement compresses those three different levels of evidence into one “quantum advantage era” narrative. The papers support an explicit advantage claim with caveats, a strong practical beyond-tested-classical result, and a trust-building framework, respectively. They do not establish three interchangeable demonstrations. The Quantum Advantage Tracker itself still calls them active candidates and says further benchmarking is needed.
The fidelity numbers need context the headlines don’t provide.
A certified state-fidelity lower bound of 0.284 at 95% confidence means the experiment establishes, with 95% statistical confidence, a lower bound of 0.284 on the overlap between the accepted output state and the ideal target state. This is not a per-shot success rate, and saying “72% of runs had undetected errors” would be an incorrect interpretation. The important advance is that the experiment supplies a statistically defensible lower bound in a regime where the target instance cannot be checked by exact classical simulation.
The throughput cost was substantial. Approximately 3.5 million raw shots produced 2,051 accepted target samples, a post-selection acceptance rate of about 0.059%, or roughly one accepted sample per 1,700 circuit executions. The paper separately reports an 860-fold reduction in effective sampling rate relative to the unencoded circuit.
For context: Google’s 2019 Sycamore experiment reported a linear-XEB value of approximately 0.002. XEB measures correlation between observed samples and ideal circuit probabilities and can be interpreted as a circuit-fidelity estimate under the experiment’s noise model; it is a different metric from the certified lower bound reported here. Google validated that methodology using related patch and elided circuits that remained classically tractable, rather than directly simulating the hardest full instance. The IBM/UChicago result provides a lower bound rather than an estimate, and that lower bound is certified through the spacetime-code structure rather than extrapolated from simpler circuits.
IBM designed this rollout to avoid a repeat of 2019.
The coordination of these announcements is worth examining on its own. Google’s 2019 Sycamore quantum supremacy claim was a solo act. IBM disputed Google’s classical-runtime estimate by proposing a statevector simulation strategy that it estimated could run on Summit in approximately 2.5 days rather than the 10,000 years Google had claimed. The subsequent back-and-forth eroded the claim’s credibility even though Google’s underlying result was technically sound. Later RCS-based claims retained variants of the same verification problem: the hardest target instances could not be checked through direct exact classical simulation.
IBM’s approach here is deliberately different. Three papers, three validation methodologies, three sets of external collaborators (UChicago, Qedma, Algorithmiq), classical benchmarking by external specialist teams at RIKEN and BlueQubit, cross-platform corroboration on Quantinuum trapped-ion hardware in the Qedma study, and pre-announcement release through a public tracker where classical simulation researchers could challenge the results. Instead of issuing a claim and waiting for someone to refute it, IBM co-built the refutation infrastructure in advance and invited the community in.
The three studies had been publicly exposed for different lengths of time. IBM says the Algorithmiq problem debuted on the tracker eight months earlier. Qedma released its relevant circuits and results there before posting the preprint. The doped-Clifford challenge was opened on the tracker on July 9, nearly three weeks before the announcements. The July 30 campaign increased their visibility, but it did not begin the classical scrutiny process. Classical algorithm improvements have a strong track record of narrowing advantage claims. A July 2026 preprint by Oh challenges whether sample-only ensemble statistics are sufficient to certify random circuit sampling, proposing a “frozen-tree” algorithm that can reproduce the relevant ensemble-level statistics without simulating a specified circuit’s output distribution. That is a potentially important criticism of sample-only verification, though it does not directly address the structured doped-Clifford task IBM used.
The Qedma paper does have one especially compelling validation element. The team reproduced the relevant oscillatory signal at selected 51-qubit time points on Quantinuum’s H2 and Helios trapped-ion systems. Agreement between superconducting and trapped-ion hardware at those points is useful evidence against a purely platform-specific artifact. This is not an independent replication of the full 74-qubit trajectory, but it is the kind of cross-platform corroboration that accumulates confidence in a way that single-platform results cannot.
What this means for CRQC timelines and PQC migration.
The three studies target structured doped-Clifford sampling, Floquet Ising dynamics, and an operator Loschmidt echo. The two quantum-simulation studies address legitimate scientific questions, but none demonstrates an end-to-end production workload or a commercially useful speedup over the best available classical workflow. The gap between “quantum computer outperforms tested classical methods on a physics benchmark” and “quantum computer delivers a decision-relevant business result more accurately or economically than classical” remains enormous.
The experiments also operate in a different error-control and resource-accounting regime than the capabilities required for breaking RSA-2048 or ECC. Shor’s algorithm requires deep fault-tolerant circuits: one representative 2025 surface-code estimate by Gidney calls for roughly 1,400 simultaneously active logical qubits, billions of fault-tolerant non-Clifford operations, and a physical qubit count approaching one million under that paper’s specific architectural, error-rate, and timing assumptions. The 468 T rotations in the IBM/UChicago experiment were virtual Z-frame rotations, not magic-state-supplied logical T gates. The two results belong to different engineering worlds.
These papers provide no technical basis for revising CRQC or Q-Day estimates, or for changing an organization’s PQC migration plan. In many sectors, migration dates and procurement expectations are already being set by regulators, government policy, clients, and contractual requirements. Those deadlines are driven by data lifetime and cryptographic dependencies, not by advantage announcements.
Indirectly, the spacetime code approach does contribute to the engineering roadmap. IBM has released Qiskit Paulice, which inserts spacetime Pauli checks into Clifford circuits and post-selects runs in which errors are detected. That makes the underlying error-detection machinery available for experimentation. The documented package still lists non-Clifford support as future work, so it does not yet implement the complete doped-Clifford protocol used in this experiment.
The bottom line.
IBM and its collaborators have produced three substantial results, but they do not establish the same thing. The University of Chicago paper makes an explicit quantum advantage claim, combining conditional hardness arguments, finite-instance simulation estimates, and a device-dependent fidelity certificate. Qedma reports a scientifically relevant regime that the classical methods tested could not reliably reproduce. Algorithmiq presents a structured framework for deciding which result deserves the most confidence when no classical ground truth is available, but its hardest-regime estimate still lacks a quantitative accuracy bound.
IBM has packaged those three different evidentiary positions as proof that an era has begun. The more defensible conclusion is narrower: quantum computing now has better tools for constructing, testing, and challenging credible advantage candidates. The Quantum Advantage Tracker itself still calls them candidates and says further benchmarking is needed.
For CISOs, CTOs, and security leaders, none of this changes the PQC migration calculus. These experiments do not demonstrate the deep fault-tolerant execution, cryptanalytic resource scale, or operational reliability required for a CRQC. They are evidence that quantum validation is becoming more sophisticated, not that RSA or ECC has become materially easier to break.
The verification advance is real. The “era” remains a claim awaiting community adjudication. I will update this analysis as the classical simulation community responds through the Quantum Advantage Tracker.