Research

Microsoft and Qolab Propose Four Conditions for a Scalable Logical Qubit. DOE’s $215 Million Competition Leaves Its Test to Negotiation.

September 22, 2026 – Microsoft published on its quantum blog a version of a preprint by Matthias Troyer and Chetan Nayak of Microsoft Quantum and John Martinis of Qolab that proposes a stricter name, the scalable logical qubit, for the encoded qubits quantum computers will need to run long algorithms. The authors posted the preprint to arXiv on September 17.

Martinis is co-founder and chief technology officer of Qolab, which builds superconducting quantum hardware, and shared the 2025 Nobel Prize in Physics. Nayak, a Microsoft technical fellow, works on the company’s topological qubits.

The authors wrote that the term logical qubit now covers so wide a range of demonstrations that results are hard to compare across platforms. Under their definition, a scalable logical qubit is sustained during a computation by repeated error correction and supports a universal set of fault-tolerant logical operations, with errors decoded in real time, fast enough for a program to branch on measurement results.

It must also come from a code family, with a supporting architecture, in which the logical error rate falls predictably as more physical qubits are added, for example by increasing the code distance. The fourth condition is a credible path to hundreds or thousands of logical qubits at the same performance, using no more than $$O(N \log \epsilon^{-1})$$ physical qubits, where $$N$$ is the number of logical qubits and $$\epsilon$$ is the logical error rate.

The authors track progress along four dimensions: reliability, scale, capability and performance. The capabilities they list run from state preparation and memory through Clifford operations, then on to non-Clifford ones, and finally to control flow that branches on measurements. The authors expect useful computation to begin at around 100 or more logical qubits with error rates on the order of 10⁻¹⁰, and large-scale chemistry and cryptanalysis to need more than 1,000 at 10⁻¹⁵ or better, citing estimates that include Craig Gidney’s 2025 analysis of RSA-2048.

The preprint says error detection with post-selection, which keeps only runs in which no error was flagged, can improve short circuits and help prepare resource states. But on its own, the authors wrote, it “does not scale to long computations,” because the chance of an error-free run falls as circuits grow. In the version on Microsoft’s blog, they describe Sandia’s quantum universal operation performance system (QUOPS) as an important step toward benchmarking logical qubits and say more is needed.

On September 17, the day the preprint was posted, the Department of Energy (DOE) opened the Quantum Genesis Q Competition, with up to $215 million in planned funding, most of it subject to future congressional appropriations. The request for applications (RFA) defines a logical qubit in one sentence, as an effective qubit spread across multiple subsystems that can be operated under quantum error correction to suppress errors or environmental noise. DOE said it is not interested in computers that can only mitigate or detect errors.

DOE plans to divide a $100 million pool equally among companies that demonstrate a first-generation machine with at least 100 logical qubits at an evaluation planned for September 2028, with two further $50 million pools for 150 and 200 logical qubits. The RFA’s illustrative first-generation targets pair 100 logical qubits with 10⁵ hard operations, which it defines as those that are difficult in the applicant’s error-correction scheme, such as non-Clifford gates in a stabilizer code. DOE’s announcement describes the target machines as capable of hundreds of millions of fault-tolerant operations.

The RFA says DOE and each selected applicant will specify during negotiation, to the extent possible, the methodology used to define and evaluate those metrics. Applications are due October 19, and DOE expects to announce initial selections no earlier than November 13.

My Analysis

Since April 2024, “logical qubit” has named 24 entangled qubits in a code that detects errors, a single surface-code memory that improved as it grew, 94 qubits in a distance-2 code on a 98-ion machine and, on September 24, 30 qubits in blocks of eight atoms, with runs that failed a parity check discarded. Troyer, Nayak and Martinis keep the noun and add an adjective, scalable, for qubits that meet all four of their conditions.

Across ten studies and announcements since April 2024, every element of the first three conditions has been demonstrated somewhere, in separate experiments. None of the ten I reviewed demonstrates all four at once. The closest on the second condition, a processor of up to eight logical qubits that Quantinuum ran on Helios-1 for the QUOPS paper it wrote with Sandia, executed universal fault-tolerant circuits at one code distance and scored below the same machine running on physical qubits. Three of the largest counts, Quantinuum’s 94, Infleqtion’s 30 and Microsoft and Atom Computing’s 24, come from distance-2 codes, and those experiments demonstrate none of the four. A reader can judge a logical-qubit count only when the announcement also gives the logical error rate, shows whether it fell when the code grew and says how many rounds of error correction ran.

The unit has money attached this year. DOE plans to award its incentive pools to machines that demonstrate 100, 150 and 200 logical qubits, and hardware roadmaps use the same unit: Infleqtion 100 in 2028, QuEra more than 256 error-corrected in 2028, and Microsoft and Atom Computing 50 in Magne, which IEEE Spectrum reported should be operational by the start of 2027. DOE has left the method of counting to negotiation with each company it selects, and the Microsoft and Qolab paper is the most detailed public proposal I have seen for what that method should check.

Microsoft and Qolab’s four conditions, explained

Repeated error correction during the computation. An error-correcting code spreads one logical qubit over several physical qubits, and the machine measures parities among them, called syndromes, to locate errors without reading the stored information. Under repeated correction, the syndromes are measured again in every round until the computation ends. Many demonstrations instead read the parities once at the end and discard every run that failed a check. When a whole computation is rejected on any detected error, the share of runs kept falls with every added gate and the attempts per good result grow exponentially.

A universal gate set, decoded in real time. Circuits built only from Clifford gates, the cheap logical operations in most codes, can be simulated efficiently on a classical computer when they start from stabilizer states and end in Pauli measurements, so a universal machine also needs non-Clifford gates, usually supplied through magic states. The authors add a timing requirement: a classical decoder must turn syndrome data into corrections, or into a record of pending corrections, fast enough for a later gate to depend on an earlier measurement. Decoding the recorded data after an experiment ends does not meet it.

A code family whose error rate falls predictably. Below a code’s threshold, raising the code distance suppresses logical errors exponentially, which is what below-threshold operation means. The condition asks for evidence of that trend in the code a platform intends to scale, and one code at one distance cannot show it.

A credible path to hundreds or thousands of logical qubits. Replication has to preserve logical performance, and the authors cap its physical cost at $$O(N \log \epsilon^{-1})$$ physical qubits.

Ten logical-qubit studies against the four conditions

The fourth condition concerns machines that do not exist yet, so the evidence below covers the first three.

  • Repeated error correction. Google ran a distance-5 surface-code memory on a 72-qubit Willow-generation processor for up to a million cycles with a real-time decoder. Atom Computing ran a toric code for up to 90 cycles with mid-circuit atom replacement. Quantinuum ran its Helix code for 20 cycles on Helios with no runs discarded. Its Helios paper benchmarked correction cycles in a distance-4 concatenated code, with post-selection. The Harvard-led team ran four-round surface-code circuits on up to 448 atoms, and AWS ran repeated cycles of a repetition code over cat qubits on its Ocelot chip. Microsoft and Quantinuum showed repeated correction with a [[12,2,4]] code in 2024, and Microsoft and Atom Computing’s November 2024 paper reported repeated loss and error correction with a distance-3 code in a separate experiment. The QUOPS processor also ran error correction within its circuits.
  • Universal gates with real-time decoding. The QUOPS processor came closest. It compiled benchmark circuits into Clifford and T gates, prepared magic states in an auxiliary Steane block using repeat-until-success, corrected errors during the circuit and was fully fault tolerant at distance 3. The Harvard-led team demonstrated non-Clifford operations by teleportation from three-dimensional [[15,1,3]] codes, but the paper reports that it applied the feedforward in software, short of the real-time control the condition requires. Google’s decoder kept pace with the 1.1-microsecond cycle, returning corrections with an average latency of 63 microseconds, but feedback into the logical circuit was not yet implemented and the memory had no logical gates. Helix ran Clifford gates only, and its memory results were decoded after the runs, though it chose which checks to run in real time. Infleqtion’s four non-Clifford CCZ gates ran in a code that, in those experiments, detects any single error and corrects only a known atom loss.
  • An error rate that falls as the code grows. Google measured a factor of 2.14 per increase of two in distance across distances 3, 5 and 7, reaching 0.143% per cycle with 101 qubits on a 105-qubit processor, separately from the real-time run. The Harvard-led team reported 2.14 ± 0.13 in four-round circuits without post-selection. Atom Computing measured lower error for the larger of its two toric codes after 4, 6 and 8 cycles, and Helix improved about 4.5-fold on a smaller [[10,2,3]] code. In AWS’s experiment, logical phase-flip error fell with distance while the minimum measured logical error per cycle moved only from 1.75% at distance 3 to 1.65% at distance 5. Quantinuum’s Helios paper reported that discard rates, a different quantity from the logical error rate, fell as concatenation raised the distance. The QUOPS processor ran only at distance 3.

The largest headline counts demonstrate narrower capabilities. Microsoft and Atom Computing’s 24 logical qubits used the distance-2 [[4,2,2]] code, which detects an arbitrary single-qubit error but cannot generally correct one at an unknown location; a detected missing atom has a known location, and the code can correct it. Quantinuum’s 94 used a distance-2 iceberg code and its 48 error-corrected logical qubits a distance-4 concatenated code, with post-selection in both cases. Infleqtion’s 30, reported by the company ahead of a promised paper, used the distance-2 [[8,3,2]] code with post-selection and software reconstruction of runs that lost one atom. None of the three headline experiments demonstrates repeated in-circuit correction of general errors. The largest count among the ten, up to 96 distance-4 logical qubits active at once in the Harvard-led team’s cluster-state experiments with [[16,6,4]] codes, also used post-selection, and that paper’s below-threshold result came from separate runs on a single surface-code logical qubit.

The written overhead bound is stricter than ordinary surface-code scaling

Under the fourth condition, $$N$$ logical qubits at error rate $$\epsilon$$ may use at most $$O(N \log \epsilon^{-1})$$ physical qubits, so the hardware may grow in proportion to the number of logical qubits and to the exponent of the error target.

A rotated surface-code patch at distance $$d$$ uses $$2d^2 – 1$$ physical qubits (97 at distance 7; Google’s memory used 101). The distance needed grows in proportion to $$\log \epsilon^{-1}$$, so for independent patches the cost of $$N$$ logical qubits grows as $$N \log^2 \epsilon^{-1}$$. If Google’s suppression factor of 2.14 held beyond distance 7, a surface code starting from 0.143% per cycle would reach 10⁻¹⁰ per cycle at distance 51 and 10⁻¹⁵ at distance 81. The bare patches would need 5,201 and 13,121 physical qubits per logical qubit, before routing, magic-state production or logical operations. In this idealized model, moving from 10⁻¹⁰ to 10⁻¹⁵, an exponent half as large again, costs about 2.5 times the qubits, where the written bound would allow about 1.5 times. Big-O hides constants, so both figures describe growth, not hardware estimates.

The authors do not run this arithmetic, and it is not a forecast: Google measured the suppression factor only up to distance 7. The paper also does not say whether its targets are per cycle or per logical operation. Google’s separate repetition-code runs also hit rare correlated error events about once an hour, a reminder that such extrapolations must allow for rare faults.

As written, the bound excludes ordinary surface-code scaling with independent planar patches. That covers the code behind Google’s Willow result and the subject of a 2012 surface-code paper Martinis co-authored, which the authors cite, as well as the toric code in Atom Computing’s June preprint, which they list among current demonstrations. It does not by itself exclude an architecture that adds an outer code to surface-code patches. Some qLDPC constructions do better: Daniel Gottesman showed in a 2013 paper on constant-overhead fault tolerance that, under suitable assumptions, the ratio of physical to logical qubits can stay constant for large circuits, though a finite machine still needs an accounting of gates, ancillas, connectivity and decoding.

A polylogarithmic bound would include conventional surface-code scaling, and it is the wording Infleqtion’s chief technology officer, Pranav Gokhale, used in his September 24 technical post, where he wrote that the cost of error correction grows polylogarithmically as the target error rate falls. A single logarithm inside Big-O does not mean that. The paper does not say which bound the authors intend, and under the one they wrote, the condition rules out the code with the most below-threshold evidence behind it, including results they list. I would ask them which they mean before anyone applies the condition to a vendor’s roadmap.

Microsoft’s headline logical-qubit results do not meet its definition

The paper’s reference list includes Microsoft’s two best-known logical-qubit results, and neither meets the four conditions. In April 2024, Microsoft and Quantinuum announced logical qubits with an error rate 800 times better than physical qubits in a post on Microsoft’s blog. The paper behind that headline reported entangled logical qubits in a [[12,2,4]] code with error rates 4.7 to 800 times lower than physical, depending on how much post-selection was used. In November 2024, Microsoft and Atom Computing entangled 24 logical qubits in a distance-2 code, and Microsoft’s announcement gave an error rate for that cat state of 10.2% with errors and losses detected, and 26.6% when losses were also corrected.

The authors list both without comment, which fits a paper that grades no one. Nayak’s own program had not reached a logical qubit when Microsoft announced its Majorana 2 chip in June, reporting parity lifetimes above 20 seconds with two-qubit operations, error correction and logical qubits still ahead, as I wrote at the time. Microsoft’s topological approach is also one of two in Stage C of DARPA’s Quantum Benchmarking Initiative, whose definition of utility scale, a computation worth more than it costs, the paper adopts.

In August I argued that each vendor’s preferred metric tends to suit the vendor that chose it, and I gave QUOPS, which Quantinuum co-authored, the same scrutiny. A definition written by two hardware companies deserves it too. This one sets a bar that Microsoft’s own headline results do not clear, which I count in its favour, and its weakest point is the wording of the fourth condition. I apply the same test to my own CRQC Scorecard, whose logical-qubit capacity figures included error-detected logical qubits. I am revising it to count only qubits sustained by repeated correction.

DOE defines a logical qubit loosely and leaves the test to negotiation

DOE’s one-sentence definition could cover most of the ten studies, including the error-detected ones, since each runs qubits under a code that detects or corrects errors. DOE screens elsewhere in the RFA. It is not interested in detection-only machines, and it counts hard operations separately, which in stabilizer codes include the non-Clifford gates of the paper’s second condition. Its first deliverable is validation data such as error rates, gate fidelities and the width and depth of circuits, all at the logical level. Its first-generation target also includes a scientific workflow.

DOE will negotiate the method with each selected company, including equivalent metrics for architectures not naturally measured in logical qubits. Staff from the national laboratories will verify performance, supported by a separate $45 million testbed program, as I described when the competition opened.

I would like DOE to publish the methodology it agrees with each company, the way Sandia published QUOPS in full with a reference implementation. If one company’s 100 logical qubits are evaluated under one negotiated method and another’s under a different one, outside comparison will be difficult unless DOE publishes the methods and results needed to reconcile them. Troyer, Nayak and Martinis want to make that comparison possible.

Questions to ask of the next logical-qubit announcement

I have adapted the paper’s five questions into a checklist, replacing its control-architecture question with overhead accounting and adding one on detection versus correction. For any logical-qubit announcement, ask:

  1. How many rounds of error correction ran during the computation? Were any runs discarded, and what fraction was kept?
  2. Which operations ran on the logical qubits: preparation and measurement only, Clifford gates, or a universal set that includes non-Clifford gates?
  3. Did the decoder keep pace with the correction cycle, and was its output fed back while the circuit ran?
  4. What was the logical error rate, and did it fall when the code distance grew?
  5. How many logical qubits ran together, and did the error rate hold as their number rose?
  6. Which of the reported logical qubits are error-corrected and which are only error-detected? How many physical qubits will each need at the error rate the company promises on its roadmap?

The next chances to ask them are already on the calendar: Infleqtion’s promised paper on its 30 logical qubits, DOE’s first selections from November 13, Magne’s planned start by early 2027, and DOE’s evaluation in September 2028.

Marin Ivezic

I am the Founder of Applied Quantum (AppliedQuantum.com), a research-driven consulting firm empowering organizations to seize quantum opportunities and proactively defend against quantum threats. A former quantum entrepreneur, I’ve previously served as a Fortune Global 500 CISO, CTO, Big 4 partner, and leader at Accenture and IBM. Throughout my career, I’ve specialized in managing emerging tech risks, building and leading innovation labs focused on quantum security, AI security, and cyber-kinetic risks for global corporations, governments, and defense agencies. I regularly share insights on quantum technologies and emerging-tech cybersecurity at PostQuantum.com.