Research

QuTech Builds Leakage Removal Into Qubit Readout, and Its Clearest Gain Is in Stability Experiments

October 2, 2026 – Researchers at QuTech, a quantum research institute at Delft University of Technology, have built a leakage reduction unit (LRU) into the measurement of superconducting transmon qubits. In a paper published in PRX Quantum on September 21, the team reported that the unit returned 98.4% of a deliberately leaked transmon’s population to its two working states during a 500-nanosecond measurement, in a calibration on a single qubit. For two-level readout, the unit added no time to the measurement.

Yuejie Xin, Leonardo DiCarlo and six co-authors tested the unit in two small quantum error correction experiments without post-selection: no runs were discarded when leakage was detected. In a seven-qubit stability experiment, the unit kept the rate at which logical errors fall with each additional round nearly constant as the researchers injected more leakage. Without the unit, that rate dropped.

Transmons store quantum information in their two lowest energy states, and excitation to the next level, the second excited state, is the dominant form of leakage. The authors wrote that a leaked transmon can remain outside those two states for several error-correction cycles and that leakage can propagate through two-qubit gates. A leakage event cannot be mapped onto the Pauli errors that stabilizer codes such as the surface code are built to correct. The errors that leakage causes are correlated in time and space.

How the unit works

The protocol adapts double-drive reset of population, a technique a Yale group published in 2013 that reset a transmon to its ground state in under 3 microseconds. One microwave drive moves a leaked transmon from the second excited state down to the first. A second drive, applied to the qubit’s readout resonator, fills it with photons efficiently only when the transmon is in the first excited state. The photons shift the transmon’s transition frequency, which puts the first drive far off resonance with the transition back up.

Because the second drive goes to the readout resonator, leakage removal and measurement run at the same time. In the calibration, the combined operation took 500 nanoseconds, which the team said is typical for a standard measurement on its processor. Of that, 120 nanoseconds went to emptying the resonator of photons. The authors wrote that the unit uses the same electronics as standard readout and single-qubit gates, without extra local oscillators or mixers.

A short delay between the two drives also lets the measurement register the second excited state as a third outcome, which flags a leakage event for the decoder and identifies the qubit and the round in which it occurred. The team measured an average two-level assignment fidelity of 99.2% in the calibration. For three-level readout, the per-state assignment fidelities averaged 96.3%, against 96.7% for the same measurement with the transmon drive switched off.

For the error-correction experiments, the preprint’s supplementary material lists a 700-nanosecond LRU-enhanced measurement, with average leakage-removal fractions of 95% in the memory experiment and 94.2% in the stability experiment, and attributes the longer duration to parameter differences between transmons.

Results from the memory and stability experiments

The researchers ran two error-correction benchmarks and raised leakage deliberately by adjusting the control parameters of the controlled-Z gates. A memory experiment on a distance-3 bit-flip repetition code, with three data and two measure qubits, tests how well the code preserves a logical state against bit flips over time. A stability experiment on seven qubits, a one-dimensional chain of four Z-type stabilizers, tests how reliably repeated measurements recover the known product of those stabilizers. The authors describe this as a test of moving a logical observable through space, as in lattice surgery, though the experiment performs no operation between logical qubits.

The team decoded both experiments with neural networks. It derived the two-level results from the same three-level measurement records by merging the leakage outcome into the first-excited-state outcome, so only the three-level analyses gave the decoder leakage flags. For the stability experiment, the authors fitted an error-suppression factor γ, the exponential rate at which the logical error probability decreases as rounds are added. The fitted values at baseline leakage and at an intermediate injected level were:

Configuration Memory, 1.0% leakage Memory, 3.6% leakage Stability γ, 1.2% leakage Stability γ, 3.7% leakage
Two-level readout, no LRU 1.35% 3.17% 0.051 0.042
Two-level readout, with LRU 1.35% 2.90% 0.066 0.068
Three-level readout, no LRU 1.20% 2.67% 0.063 0.056
Three-level readout, with LRU 1.19% 2.33% 0.079 0.080

Memory columns: logical error per round, lower is better. Stability columns: error-suppression factor γ, higher is better. Leakage levels are the authors’ upper-bound estimates for the controlled-Z gates. Source: Xin et al., PRX Quantum 7, 033062 (2026), Figs. 4 and 5.

At baseline, three-level readout cut the memory experiment’s logical error rate per round by 11.1% compared with two-level readout. Relative to two-level readout without the LRU, γ in the stability experiment at baseline rose by 29% with the LRU added, 23% with three-level readout alone and 54% with both. At 3.7% injected leakage the increases were 61%, 33% and 90%. The authors wrote that further research is needed to see whether combining the LRU with three-level readout stays the best option at larger code distances and for other codes.

A preprint appeared on arXiv on November 21, 2025. The research was funded by the EU Quantum Flagship project OpenSuperQPlus100, the Dutch National Growth Fund, Intel and the Dutch Ministry of Economic Affairs. The team has published its data and decoder-training scripts.

My Analysis

What memory and stability experiments each measure

Google’s below-threshold result on Willow and USTC’s result on Zuchongzhi 3.2, which I covered in December 2025, were both memory experiments: they measured how well a logical state is preserved over time as the code distance grows.

Craig Gidney, who proposed the stability experiment in 2022, describes computation with topological codes as two basic tasks: preserving logical information over time and moving it through space. Lattice surgery depends on the second. When two surface-code patches are merged, the product of the new stabilizer measurements along their shared boundary gives the outcome of a logical measurement. An error in that product flips the outcome. Gidney wrote that a stability experiment can determine how many rounds are needed to reach a given confidence that a logical qubit was moved correctly.

In a memory experiment, a logical error needs a chain of errors that stretches across the patch in space. In a stability experiment, the dangerous chains stretch across rounds in time. A leaked measure qubit produces that kind of chain: it returns unreliable outcomes for as long as it stays leaked.

Google’s team said as much in a two-part talk at the APS Global Physics Summit in March. Presenting stability experiments on its Willow processors, it listed leakage among the error mechanisms that extend over several cycles and to which stability experiments are more sensitive.

Memory and stability results compared

QuTech ran both benchmarks on the same device. At baseline leakage, the LRU changed nothing measurable in the memory experiment, 1.35% per round with or without it on two-level readout, and it helped there only once leakage was injected. The 11.1% gain came from the leakage flag, and because both readouts come from the same records, that comparison isolates what the flag is worth to the decoder.

In the stability experiment, the LRU on its own raised γ by 29% at baseline, a gain the memory benchmark did not register. The authors’ central claim, that the LRU keeps γ nearly constant as leakage rises, matches the fitted values in the table.

Memory experiments do register leakage damage, as this paper’s injected-leakage runs show. They do not test how reliably a processor recovers the stabilizer products that lattice surgery depends on. When a lab reports logical error rates from memory experiments alone, I now want its stability numbers as well.

In the stability experiment, the logical error probability follows $$p_L = Ae^{-\gamma R}$$, where R is the number of rounds. The rounds needed to reach a fixed error target therefore scale as 1/γ. Holding the prefactor A fixed, going from two-level readout without the LRU to three-level readout with it would cut those rounds by about 35% at baseline, where γ rises from 0.051 to 0.079, and by about half at 3.7% injected leakage, where it rises from 0.042 to 0.080. That is my arithmetic from the paper’s fits, not a measured saving in any lattice-surgery operation. The fitted A is also lowest with the LRU and three-level readout, so holding it fixed understates the saving, and the fits start at round 10 because the decay is not purely exponential before that. The authors make the point qualitatively: the LRU with three-level readout could reduce the rounds lattice surgery needs.

How QuTech’s unit compares with other leakage removal schemes

Google, in a 2023 Nature Physics paper, removed leakage from every qubit in each cycle with two operations: a multi-level reset of the measure qubits and a data-qubit leakage removal step, in which any leaked population on each data qubit is moved onto a neighboring measure qubit that is then reset. Its research blog explains that data qubits need the second step because they store the logical information and cannot simply be reset. Willow’s distance-7 experiment used the data-qubit step, with additional leakage-removal qubits at the code boundary. USTC’s Zuchongzhi 3.2 combined an all-microwave LRU for data qubits with a fast unconditional reset for measure qubits.

ETH Zurich’s flux-activated LRU, published in Physical Review Letters in 2025, removes leakage down to the group’s 7×10⁻⁴ measurement floor in about 50 nanoseconds and is designed for data and auxiliary qubits alike. It runs as a separate operation that requires flux-tunable transmons and, like QuTech’s unit, needs no additional control electronics. QuTech’s own 2023 LRU reached up to 99% removal from the second and third excited states in 220 nanoseconds, also as a separate operation.

QuTech’s new unit leaves more leakage behind than ETH’s, whose residual is at or below its measurement floor. In the single-transmon calibration, 1.6% of the population prepared in the second excited state remained after the LRU-enhanced measurement, and the error-correction runs left more. The authors wrote that removal efficiency has to improve for lower overall physical error rates, and that a larger dispersive shift or a longer coherence time would reduce the unwanted transitions that limit it. The unit adds no time to two-level readout and needs no flux pulses. QuTech’s circuits also skip the measure-qubit reset between rounds: each round’s parity comes from the difference between consecutive outcomes, and the LRU returns a leaked measure qubit to the computational states during the measurement itself. Each group measured on its own device under its own conditions, so I read these numbers as design trade-offs between approaches.

Limits of the QuTech result

  • Scale. The two codes are a distance-3 repetition code, which corrects bit flips only, and a one-dimensional stability chain on seven qubits. The paper does not test whether the device would run a below-threshold surface code.
  • Injected leakage. The team deliberately adjusted the controlled-Z settings to raise leakage. The baseline figure is an upper bound from conditional-oscillation measurements.
  • Measure qubits only. In the memory experiment, the measure qubit was the higher-frequency transmon, and therefore the one prone to leak, in every controlled-Z pair. In the stability experiment the exceptions were D5–Z2 and D5–Z3, the two pairs with the central data qubit, and the team injected no leakage on them. The injected leakage therefore went mainly into the qubits the LRU serves. Layouts in which data qubits leak as often need a second mechanism, and the authors suggest combining the unit with a data-qubit LRU.
  • Second excited state only. From simulation, the authors believe the non-exponential decay over the stability experiment’s first rounds comes neither from insufficient removal by the LRU nor from leakage on the central data qubit. They hypothesize that the cause is leakage above the second excited state. The published paper also cites Google’s March talk, which attributes the curvature to the combination of data-qubit and higher-level leakage, and notes that its own simulation lets only one data qubit leak. The LRU does not target either source.
  • Parallel operation. The authors have not fully characterized microwave crosstalk between LRUs running at the same time, and a surface code would run one on every measure qubit in every round.
  • Claim wording. The published abstract limits the zero-overhead claim to two-level readout, a qualifier the November preprint’s abstract lacked.

Leakage, cycle time and CRQC resource estimates

Two experiments on five and seven qubits give me no reason to shift any CRQC date. Gidney’s 2025 estimate for factoring RSA-2048 with fewer than a million noisy qubits in under a week assumes a uniform 0.1% gate error rate and a 1-microsecond surface-code cycle, and leakage appears in neither assumption. Meeting them on transmon hardware requires an error and timing budget that accounts for leakage, whether it is prevented, removed, reset or handled by the decoder. QuTech has shown one way to handle measure-qubit leakage inside the readout window, without a separate step in the cycle, though the paper does not show the full budget.

The LRU runs during the measurement step of syndrome extraction, which repeats in every error-correction cycle. Readout time that lengthens the cycle is paid on every logical qubit, once per cycle, for the whole computation.

The next result I will look for is the same comparison on a surface code: memory and stability experiments on one device, with the LRU running on every measure qubit at once. Google has not yet published its Willow stability data in a paper, as far as I can find. I would learn more about leakage at scale from one paper that reports stability γ next to memory Λ, the factor by which logical error falls each time the code distance grows by two, than from either number alone.

Marin Ivezic

I am the Founder of Applied Quantum (AppliedQuantum.com), a research-driven consulting firm empowering organizations to seize quantum opportunities and proactively defend against quantum threats. A former quantum entrepreneur, I’ve previously served as a Fortune Global 500 CISO, CTO, Big 4 partner, and leader at Accenture and IBM. Throughout my career, I’ve specialized in managing emerging tech risks, building and leading innovation labs focused on quantum security, AI security, and cyber-kinetic risks for global corporations, governments, and defense agencies. I regularly share insights on quantum technologies and emerging-tech cybersecurity at PostQuantum.com.