Several Vendors Now Supply Quantum Error Correction’s Classical Loop. The Qubits Are Still the Hard Part.
On September 11, Qblox and Riverlane published the timing of a complete real-time quantum error correction loop, which they assembled from commercial parts and fed with emulated qubit measurements. From syndrome extraction to conditional correction, readout integration included, they measured 6.886 microseconds at code distance 3, 9.866 µs at distance 7 and 11.886 µs at distance 9. Riverlane’s stated target for utility-scale machines is about 10 µs, the control-system reaction time that Craig Gidney and Martin Ekerå assumed in their 2019 estimate for factoring RSA-2048, and that Gidney kept in his 2025 revision.
Five other teams announced results on the classical side of error handling between September 10 and 22. Altera and Riverlane validated QECi on Agilex FPGAs; QECi is Riverlane’s specification for the link between control electronics and a decoder. Xanadu and AMD released Backline, an open-source framework for running decoder code on CPUs, GPUs and field-programmable gate arrays (FPGAs), and timed a 128-bit payload from an FPGA controller to a CPU and back at a median 2.305 µs. NVIDIA launched CUDA-Q Logical, a compiler for programs that run on logical qubits. Infleqtion compiled a quantum low-density parity-check (qLDPC) code through it. IBM cut the sampling cost of error mitigation 63-fold by adding error detection to it. And on September 22, IonQ announced that its researchers had decoded simulated workloads of up to 408 logical qubits in real time on the processor of a MacBook Pro.
I have covered each of these separately. Leaving aside IBM, whose team mitigated errors without correcting them, I read the other five as companies turning the classical half of fault-tolerant computing – the decoders, the links between them and the control electronics, and the compilers that schedule logical operations – into components with published interfaces and measured latencies, supplied by more than one company. Now that Qblox and Riverlane have timed a commercial loop under the 10 µs assumption at distance 7, on emulated data, I think the arrival date of fault-tolerant machines that can run useful algorithms depends mostly on physical qubits and on the engineering needed to build them at scale, and none of the six teams changed either.
What Each Team Measured, and on What Data
The six teams worked on different layers and tested against four kinds of quantum data: real qubits, emulated measurements, simulated syndrome streams, or none at all.
| Date | Team | Layer | What was measured | Data and hardware |
|---|---|---|---|---|
| Sep 10 | Altera and Riverlane | Control-to-decoder interface | QECi design example validated on Agilex 7 F-Tile transceivers, with the physical layer released as open source | Interface validation, no quantum data |
| Sep 10 | Xanadu and AMD | Compilation and transport to coprocessors | Backline median steady-state round trip from an FPGA controller to a CPU and back, 2.305 µs, or 4.5 µs through a GPU, with no decoder running | Classical transport benchmark, quantum side simulated |
| Sep 11 | Qblox and Riverlane | Complete error-correction loop | 6.886 µs at distance 3 rising to 11.886 µs at distance 9, syndrome extraction to conditional correction | Emulated measurements through physical control hardware, no live qubits |
| Sep 11 | IBM | Error mitigation with error detection, not error correction | Inferred sampling-overhead factor of probabilistic error cancellation cut from 85,545 to 1,359 at six Trotter steps, 648 two-qubit gates | Live, 49 of ibm_aachen’s 156 qubits |
| Sep 14 | NVIDIA and Infleqtion | Logical-qubit compiler | [[98,18,4]] qLDPC code compiled and verified in CUDA-Q Logical | Compiler and code validation, no quantum execution |
| Sep 22, preprint Aug 25 | IonQ | Decoder | End-to-end decoding on one CPU of workloads up to 408 logical qubits in 88 code blocks, adding under 0.3% to runtime at a 10⁻⁴ error rate | Simulated syndrome streams, 1 ms cycles for the 102-qubit workloads and 5 ms for the 408-qubit one |
Only IBM’s team ran its experiment on a quantum processor, using 49 of the 156 qubits on ibm_aachen, and that team mitigated errors without correcting them. The other five teams tested components with the qubits emulated, simulated or absent.
Qblox and Riverlane fed emulated measurement data through physical control hardware and split each total into two shares, one for Qblox’s control electronics and one for Deltaflow 2. Qblox’s share was 1.114 µs at distance 3 and 1.220 µs at distance 9, a rise of 106 nanoseconds across a nearly tenfold increase in qubit count. Over the same range the total grew by 5.0 µs, of which 4.9 µs was on the Deltaflow 2 side of the loop. After Riverlane shipped an update to its Deltaflow 2 system, the two companies measured the distance-3 loop at 6.886 µs, down from 8.308 µs, with no change on the control side.
With the split published, an integrator can see where the time went in this particular stack, and every microsecond Riverlane took out of Deltaflow 2 came off the total. A running processor adds calibration drift and the latency spikes that come with real noise, and none of the five September teams tested with either.
How Close the Measured Loops Are to the 10-Microsecond Reaction-Time Assumption
Reaction Time in Fault-Tolerant Runtime Estimates
Gidney and Ekerå defined reaction time in their 2019 paper as the time a classical control system needs to trigger a logical measurement, collect and error-correct the result, and choose the basis for the next one. Qblox and Riverlane timed the physical-layer steps of that interval for a single code patch, from readout through syndrome transport and decoding to a conditional operation. In a full computation, the control system also has to manage decoding windows, scheduling and coordination across patches, so the two numbers are related without being the same quantity.
Gidney and Ekerå assumed 10 µs for reaction time, alongside a surface-code cycle of 1 µs and a gate error rate of 0.1 percent. In a 2025 revision, Gidney brought the RSA-2048 estimate under a million physical qubits and kept all three assumptions.
A fault-tolerant machine has to wait for its decoder whenever the next operation depends on a decoded result, and most of those waits follow non-Clifford gates. A T gate applied through magic-state injection is completed by a correction that depends on a decoded measurement. The machine has to hold any operation that depends on that correction until the decoder answers, and it can run other work in the meantime.
In Gidney and Ekerå’s layout, each lookup addition, the operation that accounts for most of their runtime, took about 37 milliseconds, and 22 of those were spent in the addition phase, which was reaction-limited at 10 µs per step. Their formula for that phase is linear in reaction time. If I hold the other terms of that layout fixed and put in a 100 µs reaction time, I get about 220 ms for the addition phase and about 235 ms for each lookup addition, roughly 6.4 times as long, consistent with the “more than sixfold” slowdown Riverlane cites for a 100 µs response.
Measured Latencies, 2021–2026
Honeywell’s quantum team, now part of Quantinuum, ran repeated error correction with a real-time decoder on a ten-qubit trapped-ion processor in 2021, on a platform where a cycle takes milliseconds. Gidney and Ekerå set the 10 µs budget with superconducting processors in mind, and the groups measuring them have timed different things. Some timed the decoder alone, some timed the whole loop, and some used replayed or emulated data.
| Year | Team | Setup | Latency | Scope of the figure |
|---|---|---|---|---|
| 2024 | Google Quantum AI | Willow, distance-5 surface code, up to a million cycles at 1.1 µs each | 63 µs average | Decoder latency only, software decoder; latency was not the paper’s goal |
| 2024, published 2026 | Rigetti and Riverlane | FPGA decoder inside the control stack of Rigetti’s Ankaa-2, 8-qubit experiments | 0.44–0.79 µs per round; 9.6 µs full response | Live processor; the 9.6 µs is the response after the final readout of a nine-round experiment with logical branching, split into 6.5 µs of decoding and 3.1 µs of communication and control |
| Apr 2026 | Riverlane | Deltaflow 2 fed Willow’s distance-5 readout through an emulator | 16.32 µs mean | Whole QEC system, replayed hardware data, no processor attached |
| Sep 2026 | Qblox and Riverlane | Qblox control hardware and Deltaflow 2 over QECi | 6.886 µs at distance 3, 7.816 at 5, 9.866 at 7, 11.886 at 9 | Full loop including readout integration, emulated measurements |
Rigetti and Riverlane had already closed a loop in under 10 µs on a running processor, in an 8-qubit experiment with Riverlane’s decoder built into Rigetti’s control electronics. In September, Qblox and Riverlane connected a separately sold control system, Qblox’s Cluster, to Riverlane’s Deltaflow 2 through the QECi specification and published timings at four code distances.
They measured the loop under 10 µs through distance 7 and over it at distance 9, with the Deltaflow 2 share accounting for nearly all of the increase. Gidney and Ekerå used distance 27 in their 2019 layout. The largest classical task these teams have not yet done is to hold a loop near 10 µs at distances in the twenties, on a running processor, and to publish its latency tail. In the Qblox and Riverlane series, the growth measured so far has come from the decoding side.
Anyone tracking the path to a cryptographically relevant quantum computer (CRQC) can now set measured numbers against one of the inputs to Gidney’s 2025 estimate. To get a runtime of under a week, Gidney assumed a 1 µs cycle and a 10 µs reaction time on a machine of under a million physical qubits at 0.1 percent gate error. Google’s Willow already runs a 1.1 µs cycle, and Qblox and Riverlane have now timed parts of the reaction time at small code distances on commercial hardware, with emulated qubits. Nobody has built anything close to the machine.
Vendors Are Building Shared Interfaces While Each Keeps Its Own Code
QECi, NVQLink, Backline and CUDA-Q Logical
Three groups now offer four interfaces at different levels of the stack: Riverlane has QECi, NVIDIA has NVQLink and CUDA-Q Logical, and Xanadu and AMD have Backline. In QECi, Riverlane specifies the data format, runtime states and communication protocol between control electronics and a decoder. Riverlane licenses the specification and has released the physical-layer implementation, QECIPHY, as open source. Riverlane has validated QECi on Altera’s FPGAs, and Qblox has used it to connect its control electronics to Riverlane’s decoder. The Altera validation and the timed Qblox loop were separate tests.
NVIDIA launched NVQLink in late 2025 as an interconnect architecture for linking quantum controllers to GPU servers over RDMA-capable Ethernet, and quotes 400 Gb/s and latency under 4 µs. On Helios, Quantinuum decoded a qLDPC code through it with a 67 µs median decoding time, which NVIDIA calls a reaction time in its press release, against Quantinuum’s 2 ms requirement for Helios.
With Backline, developers compile decoder code written in Xanadu’s PennyLane for CPUs, GPUs or FPGAs and run it over the same kind of Ethernet link, and Xanadu and AMD have made the source code public.
NVIDIA built CUDA-Q Logical for the layer above. Developers use it to compile programs written for logical qubits onto a chosen code and hardware target, and NVIDIA says Fermilab researchers used it to cut fault-tolerant algorithm development from five months to three weeks. NVIDIA also says Iceberg Quantum used it to model its fault-tolerant architecture on Diraq’s silicon-spin qubits and arrived at about 150,000 physical qubits for 1,000 logical ones. Both figures come from NVIDIA’s announcement, and the second is a modeled resource count for a machine nobody has built. Infleqtion used the tool to compile and verify its [[98,18,4]] code.
Companies that build very different qubits can agree on these layers. Builders who follow Quantum Open Architecture, assembling machines from separately sourced processors, control electronics and software, can now do the same for error correction.
As I argued when I covered Backline, the interface is the choice a builder will live with longest, because a vendor can change a decoder algorithm in a firmware update while the data format and the transport stay fixed. My recommendation for a team specifying a machine today is to buy control electronics whose syndrome interface more than one decoder supplier already supports, and to ask every vendor for published timing boundaries and latency tails for the intended workload.
How Chipmakers Turned Classical LDPC Decoding Into Standard Silicon
Robert Gallager described low-density parity-check codes in his 1960 MIT doctoral thesis and published them in 1962. In the paper’s abstract, he judged the decoder by its equipment complexity and its data rate in bits per second. Engineers then largely neglected the codes for more than three decades, partly because iterative decoding was too expensive for the hardware of the time.
David MacKay and Radford Neal revived them in the mid-1990s. The DVB Project built LDPC codes into DVB-S2, its second-generation satellite broadcasting specification, in 2003, and 3GPP later chose LDPC codes for the data channels of 5G New Radio.
I followed the 5G part of that history for years while writing about 5G security, and it is the part I would apply to quantum error correction. In TS 38.212, 3GPP defines the LDPC base graphs and the rules for lifting them to each block size, so every handset and base-station chip decodes codes from the same standardized family. Chipmakers therefore compete on decoder silicon and on algorithm variants such as layered min-sum.
Quantum vendors are converging in the opposite order, because they haven’t settled on a code family. IBM builds its roadmap on bivariate bicycle codes, Google on the surface code and IonQ on its own qLDPC family.
Each company picks its code to match its machine’s connectivity. The surface code, itself a qLDPC code, fits a two-dimensional nearest-neighbor layout, while the higher-rate qLDPC codes IBM and IonQ use depend on longer-range connections, which IBM plans to supply with on-chip couplers and IonQ with ion transport. Without a shared code, vendors can agree first on the layers on either side of it, the control-to-decoder interface below and the logical compiler above.
Decoder designers borrow directly from classical LDPC decoding. IBM’s team built Relay-BP on belief propagation, the message-passing algorithm used to decode classical LDPC codes, and added disordered memory strengths, because plain belief propagation can fail to converge on quantum codes with short loops and symmetric trapping sets.
On its AMD VU19P FPGA prototype, IBM’s team ran each Relay-BP iteration in 24 nanoseconds. At circuit error rates below 10⁻³, where the team found the decoder converges in fewer than ten iterations on average, the team projects under 240 ns of decoding per 12-cycle window of the [[144,12,12]] gross code. IonQ’s researchers run belief propagation inside a beam search. Both teams build on the iterative decoding Gallager described in 1962.
IonQ Decoded 408 Logical Qubits on One CPU, in Simulation
I covered IonQ’s decoder in detail when the company announced it. Min Ye, Andrii Maksymov and Nicolas Delfosse posted their preprint on August 25 and revised it on September 3. For the walking-cat architecture, the trapped-ion design IonQ published in April, they built a decoding pipeline that generates error models on the fly and decodes memory blocks, logical measurements and magic-state factories together.
Against simulated circuit-level noise, their decoder kept pace with two workloads of 102 logical qubits at 1 ms cycles and one of 408 logical qubits at 5 ms cycles, using 12 of the 16 cores of an Apple M4 Max in a MacBook Pro.
At a two-qubit error rate of 10⁻⁴, the authors report schedules 0.24, 0.18 and 0.02 percent longer than they would be with an instant decoder. At 5 × 10⁻⁴ the increase was less than 12 percent. The authors name the architecture as a key ingredient. In the walking-cat design, IonQ performs logical operations by inserting cat-state measurements into an otherwise unchanged stream of syndrome extraction, without merging or deforming code blocks. The decoder can therefore work from one fixed Tanner graph, which links each possible error to the checks that detect it, and update only the error probabilities attached to it.
The authors excluded the controller’s own logical-control work from the timings and did not simulate the cat-state factory. They also seeded the magic-state factories with stabilizer states, so they did not test output quality.
In its press release, IonQ claims more than its researchers do in the paper. IonQ calls the result the industry’s first end-to-end real-time error-correction decoder on a single CPU, and the release’s “31.5 million individual quantum operations” are the paper’s syndrome-extraction cycles summed across 88 code blocks, in a workload with 555,130 T gates and about 1.3 million logical measurements.
Of the five real-time decoding results I compared in that article, IonQ’s is the only one with magic-state factories and compiled application workloads in the loop, and it has the most logical qubits, while Google and Quantinuum decoded syndromes from real qubits first.
In its September 8 secp256k1 estimate, IonQ needs 69 memory blocks of a larger code decoded without interruption for 25.7 days per attempt. In the decoder paper, the authors decoded 88 blocks concurrently, in simulation, for the April design’s smaller [[70,6,9]] code and for runs equivalent, by my arithmetic, to 5 to 30 minutes of machine time. For the 25.7-day run, IonQ still has to show continuous operation as much as decoding.
IonQ can decode on one laptop processor because its cycles are slow. With a trapped-ion cycle of a millisecond or more, a decoder has about a thousand times longer per round than with a superconducting cycle of about a microsecond, and the computation takes correspondingly longer to run.
For the same 256-bit elliptic-curve problem, Google Quantum AI estimated in March that a superconducting machine would need 18 to 23 minutes, against 25.7 days per attempt for IonQ’s design. The two are resource models with different algorithms and assumptions, not a controlled comparison of clock speed.
Google’s authors divide machines into fast- and slow-clock architectures, and I would sort decoders the same way. Designers of fast-clock machines face much tighter limits on decoder throughput and feedback latency. IBM’s decoder team argues that decoders for microsecond cycles must run on FPGAs or ASICs, and Riverlane builds Deltaflow on FPGA hardware, although Google kept pace with a distance-5 memory on Willow in software. IonQ and the superconducting vendors have made opposite choices about where to pay the classical cost.
Sampling, Qubit and Time Overheads
Several September teams reported lower overhead, and they meant three different quantities.
Sampling overhead is the factor by which a team running error mitigation has to multiply its circuit executions to reach a given statistical precision. With Spacetime PEC, IBM cut the inferred sampling-overhead factor from 85,545 to 1,359 for a circuit with 648 two-qubit gates, by discarding runs that fail error-detection checks. The team then uses probabilistic error cancellation only on the errors it could not detect. The cost still grows exponentially with those undetected errors.
As I wrote when IBM published the preprint, IBM has made pre-fault-tolerant experiments cheaper. I don’t expect anyone to reach fault tolerance sooner because of it.
Qubit overhead is the number of physical qubits per logical qubit. With its [[98,18,4]] code, Infleqtion stores 18 logical qubits in a block of 98 data qubits, 5.4 per logical qubit, or roughly 10.9 data and check-ancilla qubits per logical qubit if each of the code’s 98 checks, 49 X-type and 49 Z-type, gets its own ancilla. That count is my estimate, since Infleqtion hasn’t published a total, and I have left out other implementation resources.
At distance 4, a decoder can correct any single data-qubit error in the block, and Infleqtion hasn’t run the code on hardware yet. Machine builders use higher-rate codes of this kind to shrink fault-tolerant machines, and for each code someone has to build a real-time decoder that can handle a less regular check structure than the surface code’s grid.
Time overhead is the extra runtime caused by classical processing. Qblox and Riverlane timed loop latencies, and IonQ reported how much longer a scheduled workload ran because of decoding delays. Resource estimators put physical-qubit footprint and decoder-induced waiting directly into their size and runtime estimates for fault-tolerant machines, and IBM’s mitigation factor is a different quantity from both.
How I Now Read Fault-Tolerance Timelines
On its roadmap, IBM targets Starling, a machine that runs 100 million gates on 200 logical qubits, for 2029, and Kookaburra, planned as its first module to store information in a qLDPC memory and process it with an attached logical processing unit, for 2026.
On September 17 the US Department of Energy opened the Quantum Genesis Q Competition. DOE will split a $100 million pool among companies that demonstrate at least 100 logical qubits running hundreds of millions of fault-tolerant operations, with two further $50 million pools for 150 and 200 logical qubits. DOE has $2.5 million of fiscal 2026 money for the competition, and the rest of the up to $215 million depends on appropriations. Applications close on October 19.
To meet either target, a builder needs real-time decoding and control across the logical register and its magic-state factories. In an accompanying $45 million lab call for a validation and verification testbed, DOE lists classical control systems among the layers of the stack it wants national labs to characterize, next to hardware, gates, logical architectures and algorithms. After September, a builder bidding for either target can assemble that decoding and control path from components with published interfaces and measured latencies, from more than one supplier. At small code distances, Qblox and Riverlane have already measured such a path close to the 10 µs budget Gidney and Ekerå used.
In my reading, the larger gap is on the quantum side. A Sandia-led team posted its QUOPS benchmark on September 10 and measured today’s best processors at roughly five orders of magnitude short of the circuit sizes needed to factor RSA-2048 or estimate an energy of the FeMoco molecule, while close to the throughput needed for those problems. The authors write that hardware makers will need lower error rates, more physical qubits or more efficient logical architectures to close the gap.
Fitting data back to 2018, the team found capability doubling about every 1.4 years at Quantinuum and every 2.1 years at IBM. At those rates, the authors estimate, nobody reaches the challenge problems until 2050–2070, and companies projecting scientifically useful capability in the early 2030s would need capability to grow four times as fast.
In April I ranked decoder performance, Capability D.2 in my framework, first among the capabilities I would watch to judge whether someone builds a CRQC in 2030 or in 2045. I no longer rank it first. My readiness rating for D.2 hasn’t changed, because none of these teams has integrated a decoder with a quantum processor at scale.
Qblox and Riverlane have now met the 10 µs assumption at distance 7 with separately sold parts and emulated qubits, and IonQ’s decoder has kept pace with a simulated 408-logical-qubit workload for its trapped-ion design on one CPU. On the decoding side, someone still has to hold the 10 µs loop latency, with its tail, at distances in the twenties on running hardware and across many cooperating decoders. In my judgment, physical qubit count, error rates and the engineering of machines at scale are now the harder constraints, and which end of that range we get depends mostly on them.
In timeline terms, I now give much less weight to late-arrival scenarios in which builders are still stuck on basic control-to-decoder integration. I still count decoder scaling and latency tails on running hardware as risks. I haven’t moved my earliest plausible date, because it depends on how fast hardware makers can scale physical qubits.
What the September Teams Have Not Yet Shown
- A full loop on a running processor at distance 9 or above. Qblox and Riverlane used emulated measurements in their four-distance series. Rigetti and Riverlane’s loop on a real processor was an 8-qubit experiment. A running processor adds drift and latency spikes. The team that runs such a loop should publish the latency distribution along with the mean, together with the logical error rate it measured.
- Logical operations decoded in real time on a running fast-cycle processor. Yale researchers used DecoNet to decode lattice surgery across 100 distance-5 logical qubits on five networked FPGAs in 2025, with synthetic noise, and IonQ’s researchers decoded logical measurements and factories in simulation. None of the September teams decoded logical operations on fast-cycle quantum hardware. Riverlane planned Deltaflow 3, the version with lattice surgery, for 2025 on its original roadmap and now expects it in late 2026.
- Higher-rate qLDPC decoding on a running superconducting processor. IBM’s team tested its Relay-BP FPGA prototype against simulated noise. On its roadmap, IBM has scheduled Kookaburra, planned as its first module built around a qLDPC memory, for 2026.
- Cooperating decoders at scale on running hardware. Quantum Machines has documented that decoders for lattice surgery and qLDPC codes have to exchange data in real time. The DecoNet team showed five networked FPGAs doing this on synthetic data, and in their whitepaper Xanadu and AMD put the primary decoder on an FPGA or ASIC and give CPUs the backup decoding.
- Independent measurements of these systems. Every figure in both tables above comes from the companies that built the parts. In its testbed call, DOE names classical control systems as a layer to validate, and the national labs that take on that work could publish independent numbers next to the vendors’ own.
Gidney and Ekerå wrote 10 microseconds into their estimate in 2019 as an assumption about control systems that did not yet exist. In September Qblox, a Dutch control-electronics company, and Riverlane, a British decoder maker, timed a commercial loop at 9.866 µs at distance 7, with the qubits emulated. The next number I am waiting for is that round trip on a running processor at a distance well above 9, published with its full latency distribution and the logical error rate the team measured.