Modular Quantum Error Correction: 10x in Simulation
Table of Contents
October 2, 2026 – 14 researchers from two Google teams, Quantum AI and DeepMind, and from MIT posted a preprint to arXiv describing a modular quantum error correction design that builds high-rate quantum memories from small, flat chip modules wired together at their edges. In circuit-level simulations, a memory built this way needed ten times fewer physical qubits per logical qubit than the paper’s modularized surface code baseline to reach the same logical error rate.
The paper, “Low-Overhead Quantum Error Correction with Boundary-Connected Planar Modules,” has not been peer reviewed. Oscar Higgott and Hasan Sayginel of Google Quantum AI are the corresponding authors. The co-authors include Craig Gidney, who wrote Google’s 2025 estimate for factoring RSA-2048, and Hartmut Neven, who founded Quantum AI. Higgott previewed the work in a June 10 talk at the QEC 2026 conference in Santa Barbara, according to the conference program.
The tenfold saving comes from simulations of logical qubits held in memory. The paper’s larger figure, a reduction of more than 30 times, comes from extrapolating to codes the authors describe as too large to simulate or decode. The paper contains no resource estimate for a complete quantum algorithm.
Google showed surface-code memories operating below the error correction threshold on its 105-qubit Willow processor in 2024, with a largest memory of 101 physical qubits at distance 7. A rotated surface code stores one logical qubit in a patch of about 2d² physical qubits, where d is the code distance, the smallest number of physical errors that can form an undetectable logical error. The authors put the cost at more than 1,000 physical qubits per logical qubit for the distances of 20 to 24 they expect utility-scale machines to need.
Higher-rate quantum low-density parity-check (QLDPC) codes store many logical qubits in one block. On solid-state chips, the authors wrote, those codes need dense, long-range couplers that cross one another.
The authors split a processor into planar modules of 42 to 116 physical qubits in the paper’s main examples. Inside a module, no qubit couples to more than four nearest neighbors on a two-dimensional grid, a layout the authors describe as similar to current superconducting processors built for the surface code. Couplers at the module boundaries join the modules in the pattern of a closed hyperbolic surface, a negatively curved geometry on which the number of logical qubits a code stores grows with the size of the surface.
The hardware model came from Google’s Matt McEwen, according to the paper’s statement of author contributions. The team built two code families, modular hyperbolic surface and color codes, by replacing each tile of a hyperbolic tiling with a flat patch of ordinary lattice, a step the authors call fine-graining. Fine-graining raises the code distance while keeping the number of logical qubits fixed, and the authors describe the color-code family as the first semi-hyperbolic color codes.
Simulated Qubit Savings and Tolerance of Noisy Links
The team simulated memory experiments under SI1000, a noise model designed to resemble superconducting hardware, at a physical error rate of 0.1%. Every two-qubit gate between modules was assigned a 1% error rate, ten times the rate inside modules.
A hyperbolic surface code spread across 120 modules of 116 physical qubits each, 13,920 in all, stored 146 logical qubits. Its logical error rate was about 10⁻¹⁰ per logical qubit per round. The authors reported that a surface code cut into modules of similar size, under the same noise, needs ten times as many physical qubits per logical qubit to reach that rate. Color codes decoded with a variant of DeepMind’s AlphaQubit neural-network decoder showed an eightfold improvement on the largest simulated color code, which stores 36 logical qubits at distance 12 using 1,152 physical qubits.
Raising the inter-module error rate to 1%, with the rate inside modules held at 0.1%, raised logical error rates severalfold in the paper’s plotted results, and by up to about two orders of magnitude at 2% for the larger color codes. The authors describe the logical error rate as only weakly sensitive to the inter-module error rate across the partitions and module sizes they tested. They attributed the tolerance to the small fraction of gates that cross module boundaries, and their multi-rate circuits measure the stabilizers that span modules only once every three rounds.
Extrapolating to larger codes, the authors projected a 15-fold saving for the hyperbolic surface codes and more than 30-fold for the color codes, reaching 36 physical qubits per logical qubit at distances of 20 to 24. They wrote that such codes would exceed both the rate and the distance of IBM’s [[288,12,18]] two-gross bivariate bicycle code, which uses 48 physical qubits per logical qubit including check qubits.
Logical Access, Decoding and a 50-Bit Adder
To compute with stored qubits, the authors attach a “ribbon extractor,” a strip of modules coupled to the memory along one closed cycle of the code that measures the logical operators on that cycle. For two codes storing 104 logical qubits each, at distances 22 and 30, the extractor adds about 22% to the memory’s qubit count, a figure the authors called a cautious upper bound.
Rotations of the hyperbolic tiling, symmetries the authors call code automorphisms, move other logical operators under the extractor. Walking circuits of SWAP gates carry out each rotation over extra boundary couplers. The chosen rotations cover 100 of the code’s 104 logical qubits. Any covered basis cycle reaches the extractor in at most six moves.
As a benchmark, the team compiled a 50-bit ripple-carry adder onto the distance-22 code, teleporting operands through the extractor to a small compute region built from a surface or color code. Under assumed costs for each extractor measurement and each move, the authors bounded the time for that extractor and automorphism work at 21,402 error correction cycles, 428 for each bit.
The AlphaQubit variant has about 55 million parameters, and the team trained a separate instance for each circuit on up to 3 billion simulated examples. On a single Google TPU v6e, known as Trillium, it decoded each complete 16-round sample in 140 microseconds to 6 milliseconds, depending on circuit size. The hyperbolic surface codes used a correlated minimum-weight perfect matching decoder.
The authors listed what a hardware implementation would require: superconducting systems combining high-performance local qubit grids with high-fidelity inter-module couplers, tolerance of fabrication defects, and real-time decoding at scale. They also called for faster extractors for subroutines such as quantum table lookups, known as QROM, and for full algorithms compiled onto the architecture to give a complete resource estimate. In the supplementary information, they left full simulation of the walking circuits to future work.
An AI-usage statement in the paper says AI tools helped implement parts of the research software and proposed the argument behind the color-code distance lower bound. The authors wrote that they verified all AI-assisted contributions.
My Analysis
Higgott and his colleagues have written the most complete high-rate memory design for planar superconducting hardware I have seen, with codes, syndrome-extraction circuits, two decoders, a logical-access scheme and a compiled adder in 112 pages. They give their most useful practical result one clause in the abstract. With links between modules ten times noisier than the gates inside them, simulated logical error rates rose only severalfold. Of all the numbers in the paper, the abstract’s 30× has the least evidence behind it.
For anyone tracking the path to a CRQC, the input that changes is the physical-qubit cost of storing idle logical qubits on superconducting hardware. In Gidney’s 2025 RSA-2048 estimate, cold storage for idle logical qubits takes 550,400 of 897,864 physical qubits. In that paper, Gidney wrote that he saw no way to cut the qubit count by another order of magnitude without changing its physical assumptions, among them a square grid with nearest-neighbor connections. Higgott’s team relaxed exactly that assumption, keeping nearest-neighbor grid connectivity inside each module and adding links between modules, with Gidney among the authors.
Why Google Put the Long-Range Wiring at the Module Edges
The surface code suits Google’s chips, which place qubits on a square grid with couplers between nearest neighbors, and Google has run it below threshold on Willow. Its weakness is density. Each patch stores one logical qubit. Gidney budgeted 1,352 physical qubits for every actively used logical qubit at distance 25. He put the 1,280 idle ones into yoked surface codes, an outer layer of parity checks that packs idle logical qubits about three times more densely, at 430 physical qubits each.
IBM took a different route. Its bivariate bicycle “gross” code stores 12 logical qubits. It uses 144 data qubits and an equal number of check qubits. IBM says that fitting the code onto a flat chip requires long-range connections following the symmetry of a three-dimensional torus. In November 2025, IBM announced its experimental Loon processor with “c-couplers” for those connections, and its announcement says Loon demonstrates all the key processor components needed for fault tolerance. Kookaburra, IBM’s first module with a QLDPC memory, is scheduled for 2026, and Starling, targeted for 2029, is meant to run 100 million gates on 200 logical qubits, as I covered in my analysis of IBM’s roadmap.
Google’s team took a third route. The authors argue that a monolithic chip at the scale of 100,000 to a million physical qubits is impractical anyway, citing fabrication, yield and coupler density, so a machine of that size will be built from modules. They put the non-planar part of a hyperbolic code into the links between those modules. Each module is an ordinary flat patch, and the negative curvature that gives the code its high rate comes entirely from how the modules are joined. In the authors’ words, boundary-connected modules “turn a constraint of scalable hardware into a resource.”
Classical chipmakers made a similar trade nearly a decade earlier. In 2017, AMD built its first EPYC server processor from four 213 mm² dies instead of a hypothetical 777 mm² monolithic die. The four dies used about 10% more silicon for die-to-die links and duplicated logic, and AMD’s yield model put their cost at 0.59 of the monolithic design. AMD paid for its links in silicon to win yield. In Google’s design, the links are what make the code dense. The hyperbolic connectivity, and with it the high rate, comes entirely from the connections between modules, and flat patches without those links would have no more density than ordinary planar codes.
AMD’s Infinity Fabric moved classical bits between dies. Google’s links have to carry entangling gates at about 1% error, and the cable-link experiments the authors cite connected two nodes in 2021, five modules in 2023 and two chips 64 meters apart in 2025.
What Google Simulated, What It Projected, and What It Designed Without Simulating
The headline overhead gains come from circuit-level memory simulations. The logical-access design is supported by constructions, a distance proof and a simplified simulation.
| Result | Code | Evidence |
|---|---|---|
| 10× fewer physical qubits per logical qubit than the paper’s modular surface-code baseline at about 10⁻¹⁰ | [[6000,146,20]], 13,920 physical qubits on 120 modules | Circuit-level simulation of memory, matching decoder |
| 8× fewer physical qubits per logical qubit than the same baseline | Color codes up to [[512,36,12]], 1,152 physical qubits | Circuit-level simulation of memory, AlphaQubit variant |
| Logical error rate rises severalfold at 1% seam error and by up to about 100× at 2% | Codes of up to 1,152 physical qubits | Circuit-level simulation of memory, read from Figure M4 |
| 15× | [[152250,3656,25]] | Extrapolation, code too large to simulate or decode, non-orientable with no logical-access basis |
| More than 30×, 36 physical qubits per logical qubit | Color codes storing 1,028 to 3,040 logical qubits, including [[31200,1954,22–24]] | Extrapolation, some distances given as bounds, no logical-access basis listed |
| Extractor at up to about 22% of memory size | [[3600,104,22]] and [[6400,104,30]] | Construction with a distance proof, no circuit-level logical error rates reported |
| Automorphism moves by walking circuits | The same two codes | Construction and a proxy simulation with boundary stabilizers switched off, full walking circuits left to future work |
| 50-bit adder, extractor and automorphism work bounded at 21,402 cycles | [[3600,104,22]] | Compiled cycle-count bound under assumed per-operation costs, no logical error budget |
| Full-algorithm resource estimate | None | Listed as future work |
The 36-physical-qubit figure comes from codes storing 1,028 to 3,040 logical qubits each, too large for the team to simulate or decode. The authors plotted them at the logical error rate measured for a surface code with the largest odd distance not exceeding theirs, using for color codes a distance upper bound they expect to be tight. Their evidence for that step is that every simulated hyperbolic code except the smallest beat the surface code at distance d or d − 1. The paper lists no logical basis for the extractor-and-automorphism scheme on any of these large codes. Two of the three behind the abstract’s rate of 1/16 at distance 22 or more are built on non-orientable surfaces, which the authors expect most of their methods to handle but leave to future work. The 1/16 rate counts data qubits only, while the 36 figure includes measurement qubits.
The codes with a worked-out access scheme give lower and better-supported numbers. For each of its 104 logical qubits, the distance-22 code uses 80.8 physical qubits. Adding the extractor brings it to about 99 per encoded logical qubit, or about 103 for the 100 logical qubits the access scheme reaches. The authors count roughly 1,000 for a surface code at the same distance. The simulated tenfold result, though, comes from a different code, the hyperbolic surface code with 146 logical qubits. The paper reports no simulated logical error rate for the color-code construction with access. Whether that construction keeps a tenfold advantage at a matched error rate is still open.
The abstract labels the 30× figure a projection in one sentence. The codes behind that figure, and what they lack, appear in Table M2, Methods I and supplementary section S10 on distance bounds.
Hardware Assumptions, From 42-Qubit Modules to 1% Inter-Module Links
The authors justify the 1% inter-module error rate with link experiments by other groups. In 2021, a University of Chicago team transferred states between two three-qubit nodes over a one-meter superconducting coaxial cable with 91.1% process fidelity. In 2023, a SUSTech-led group in Shenzhen linked five modules with low-loss aluminum cables and reported fidelities of up to 99% for state transfer and Bell states. In 2025, a SUSTech team teleported a CNOT gate between chips 64 meters apart with 70.2% process fidelity. The authors describe the best such links as approaching 1% error. These experiments make good cable links plausible, but none of them validates the paper’s model of 1% error on every repeated inter-module CZ gate, and none ran on Google hardware.
In the team’s simulations, logical error rates rose severalfold when seam error went from 0.1% to 1%, ten times the rate inside modules, and by up to about two orders of magnitude at 2% for the larger color codes, reading from the curves in the paper’s Figure M4. The multi-rate circuits help by measuring the stabilizers that span modules once every three rounds. The simulations show the codes keep working with much noisier links. They don’t show that thousands of such links can be built, calibrated or run at those error rates.
Quantity is the open question. The simulated memory used 120 modules. As an illustration, a cold store the size of Gidney’s would take 13 distance-22 color-code blocks of 200 modules each. That is about 2,600 modules of 42 physical qubits each, plus 13 extractors. In that quadrilateral layout, each module needs couplers to its neighbors on all four edges, plus, for automorphism moves, an additional seam on each side to a non-neighboring module. Other module partitions need different seams. The authors describe flexible inter-module couplers that are not constrained to the planar geometry of any one chip.
Fabrication defects are the other unshown requirement. The authors expect methods developed for broken qubits and couplers in planar surface and color codes to adapt, since each module looks locally like one of those codes, but they don’t demonstrate it.
In Chapter 9 of Quantum Systems Integration, I argue that wiring heat load, not the QPU roadmap, sets the practical qubit ceiling of a superconducting cryostat. A boundary-connected machine has a second wiring problem inside the coldest stage: thousands of quantum links, each kept near 1% error. In my CRQC framework those requirements fall under qubit connectivity and engineering scale and manufacturability, and the authors assume the links instead of building them.
How the Surface-Code and Two-Gross Comparisons Were Set Up
Within its own hardware model, the surface-code comparison is fair. The authors cut rotated surface codes of distance 3 to 25 into modules of 98 physical qubits and gave their inter-module gates the same 1% error. They decoded them with the same correlated matching decoder and counted the extra measurement qubits a modular layout needs.
The baseline is still the plain surface code. Google’s own CRQC estimates store idle logical qubits in yoked surface codes, which Gidney sized at 430 physical qubits per logical qubit against 1,352 for a regular distance-25 patch, about three times denser. The paper doesn’t compare against yoked storage. Its noise model and error target also differ from Gidney’s, so the tenfold figure doesn’t carry over to the storage Google’s estimates assume. At a matched target, the gain over yoked storage would be smaller than the gain over the plain surface code, by an amount nobody has measured. The authors put earlier circuit-level results with 2D nearest-neighbor gates, including yoked codes and Google’s denser planar surface code from May, at up to about 5× over the surface code.
IBM’s two-gross block, described in the Tour de gross architecture paper, stores 12 logical qubits. With check qubits counted, the block uses 48 physical qubits for each logical qubit. The Google codes the abstract compares with it store 1,522 to 3,040 logical qubits per block and have neither a simulation nor an access scheme. Google’s codes with a worked-out access scheme use 81 to 99 physical qubits per logical qubit at distance 22, before and after the extractor. Both sets of numbers are memory-block overheads. IBM’s 48 excludes the logical processing unit IBM attaches to each memory module, and neither figure includes compute, factories or decoding.
The two designs make different trades. Google keeps every on-chip coupler between nearest neighbors and pays in block size, access time and inter-module wiring. In IBM’s Tour de gross design, the longest couplers inside a module span tens of lattice sites. Neither paper establishes a like-for-like total at a common logical error target.
Decoding Speed Against a 10-Microsecond Reaction Budget
In February, Iceberg Quantum, a startup in Sydney, claimed in its Pinnacle Architecture preprint that RSA-2048 could be factored with fewer than 100,000 physical qubits, and I named real-time decoding as its single biggest gap. CRQC resource estimates, Gidney’s and Pinnacle’s included, assume a 1-microsecond code cycle and a 10-microsecond reaction time for the classical control loop. Of the open problems in Google’s paper, I’d put decoding furthest from solved.
The AlphaQubit variant took 140 microseconds to 6 milliseconds of TPU time to decode each complete 16-round sample. Spread evenly over the 16 rounds, that is 9 to 375 microseconds per round, an amortized figure, not a measured streaming throughput or feedback latency. Those are different requirements. Willow’s real-time decoder kept pace with a 1.1-microsecond cycle at distance 5 with an average latency of 63 microseconds. DeepMind’s AlphaQubit 2, posted in December 2025, decodes faster than 1 microsecond per cycle on commercial accelerators, for surface codes up to distance 11 and color codes up to distance 9. The codes the authors propose for utility-scale use run at distances of 20 to 30, spread across hundreds of modules.
The authors write that real-time decoding at scale will need parallel or specialized hardware. A logical error in these codes spans many modules, so the decoding algorithm itself will likely have to be parallelized. The overlapping time windows used for parallel surface-code decoding probably won’t be enough. They report no decoding speed for the matching decoder used on the hyperbolic surface codes. The neural decoder was trained for each circuit on simulated noise, and the team didn’t test it on data from hardware.
Several vendors now supply the classical decoding loop for surface codes, and IBM says, in the same November announcement, that it has decoded its QLDPC codes in real time in under 480 nanoseconds. Decoder latency is the capability I track as D.2 in my CRQC framework. Google’s team reports no real-time decoding for hyperbolic codes spread across modules.
Google’s RSA and ECDLP Estimates With a Denser Memory
Higgott’s team cites Gidney’s 2025 estimate in its introduction as the upper end of the 100,000 to 1 million physical-qubit range for utility-scale machines. To see what this memory could do to that number, I swapped one line of Gidney’s layout and held every other assumption fixed. The result is a sensitivity calculation for one line item, not a resource estimate for Google’s architecture.
In his paper, Gidney assumed a square grid with nearest-neighbor connections, a 1-microsecond cycle, a 0.1% gate error rate and a 10-microsecond reaction time. He put the 1,280 logical qubits of the input register in yoked cold storage at 550,400 physical qubits, and the 131 active ones in regular hot patches at 177,112. Six magic-state factories and lattice-surgery workspace take a compute region of 170,352 physical qubits. The total is 897,864, which he rounded up to a million for slack.
Thirteen of Google’s distance-22 color-code blocks, each with one extractor, would give 1,300 addressable logical qubits for Gidney’s 1,280, using 133,640 physical qubits. With the other 347,464 physical qubits of his layout unchanged, the total would be about 481,000. Even a memory that cost nothing would leave those 347,464 in place, so the storage swap roughly halves Gidney’s estimate. The projected 36-per-logical codes could only push the total toward that floor, and the paper gives them no access scheme yet.
Two unknowns would move these totals. Gidney sized his storage for 10⁻¹⁵ errors per logical qubit per round. Each 12-hour shot of his algorithm runs about 6.9 × 10¹³ logical-qubit rounds, and at that rate 93.3% of shots finish without a logical error. The paper’s headline memory reaches about 10⁻¹⁰ under a different noise model, and the color-code construction with access has no simulated logical error rate in the paper. The reported results can’t tell us what meeting Gidney’s target would cost. It may take larger distances, different circuits or more overhead.
The second unknown is time. Gidney streams copies of cold-storage qubits to serve as lookup addresses in hot storage. He added his 100,000 qubits of slack specifically to cover that traffic. In the modular memory, the authors reach each logical qubit by applying up to six automorphism moves and then measuring through the extractor. The paper’s adder benchmark bounds that work at 428 cycles per bit of addition under assumed costs per measurement and per move, and the authors say lookups will need faster extractors. A slower algorithm runs more than Gidney’s 6.9 × 10¹³ logical-qubit rounds per shot. Keeping 93.3% of shots free of logical errors would then require an error rate below 10⁻¹⁵ per logical qubit per round.
Google’s March 2026 ECDLP estimate puts a 256-bit secp256k1 key within reach of fewer than half a million physical qubits in minutes. Babbush, Gidney and their co-authors also assumed planar connectivity and dense, yoke-protected surface-code storage in that whitepaper. Under the same fixed-layout assumption, its storage share would shrink and its factories and compute region would not.
If Google builds this architecture, the deciding questions for a superconducting CRQC become whether thousands of inter-module links can stay near 1% error and whether a decoder can keep up with codes spread across hundreds of modules. I don’t change my Q-Day assessment on a simulated memory alone. Gidney closed his 2025 paper by endorsing the timeline in NIST’s draft transition guidance, which deprecates 112-bit-strength algorithms such as RSA-2048 after 2030 and disallows quantum-vulnerable public-key algorithms after 2035. He said he preferred security not to depend on progress being slow.
Frozen Pinnacle and Boundary-Connected Modules, Two Routes to High-Rate Codes on Fixed Wiring
Iceberg’s Pinnacle preprint used generalized bicycle QLDPC codes and estimated about 98,000 physical qubits for a month-long RSA-2048 run. Like Gidney, it assumed a 1-microsecond cycle, a 0.1% gate error rate and a 10-microsecond reaction time. In my February analysis, the connectivity model was the second big if. Pinnacle needed bounded but non-local connectivity inside its modules, which made its numbers hard to compare with surface-code estimates on a nearest-neighbor grid. I analyzed the remainder of the architecture in a longer piece in March.
In September, Iceberg addressed the fixed-connectivity requirement with Frozen Pinnacle, a variant that fixes every qubit’s couplers at fabrication, at most eight per qubit, as superconducting hardware requires. Iceberg estimates one-month RSA-2048 factoring on it with about 120,000 physical qubits under the same assumptions, and its authors compare that degree of eight with IBM’s bicycle architecture, which has degree seven. That estimate also comes from extrapolation. Iceberg simulated codes of distance 4, 6 and 10 and extrapolated to the distance-24 code it uses.
Google’s team approached the same constraint from the other direction. Iceberg kept its codes and fitted them to a fixed degree-eight graph. Google kept the wiring of its current chips inside each module and moved all non-local structure into the links between modules. Neither team has closed the decoding gap. Pinnacle’s estimates relied on a most-likely-error decoder and treated real-time decoding as out of scope. Google’s neural decoder for these codes took 140 microseconds to 6 milliseconds per 16-round sample, which a 1-microsecond cycle generates in 16 microseconds.
Iceberg has published factoring estimates for both versions of its architecture. Google has published a memory, an access scheme and a 50-bit adder. Until Google compiles a full algorithm onto boundary-connected modules, which its authors list as the next step, there’s no Google number to set beside Iceberg’s 120,000.
I read this paper as the third installment of one program. In May, Google’s Guang Hao Low, Ryan Babbush and colleagues published a denser planar surface code with up to 4.5 times the encoding rate of a rotated surface code on a hexagonal grid, and estimated a FeMoco simulation in under a month on 89,000 noisy superconducting qubits. In September, Google’s Noah Shutty posted denser planar color codes for square-lattice hardware. All three teams raised qubit density while keeping on-chip couplers between nearest neighbors, where IBM added long-range couplers for its bicycle architecture.
What Google Needs to Show Next on Boundary-Connected Modules
Four results would change my assessment of this architecture:
- An integrated demonstration of the module-and-link combination, with below-threshold modules joined by inter-module gates near 1% error across more than a handful of chips.
- Circuit-level simulations of the extractor and the walking circuits, and of memories built from the larger codes at logical error rates approaching 10⁻¹⁵.
- A decoder that sustains throughput at the syndrome rate of a 1-microsecond cycle, with a latency distribution that fits a 10-microsecond feedback budget, on a code spread across hundreds of modules.
- A full-algorithm resource estimate on this architecture, with storage, access, compute, factories and decoding in one layout, the step the authors list next.
Until those arrive, the numbers I’d quote are the tenfold saving for a simulated memory and the tolerance of much noisier links. The color-code construction with access keeps its count near 100 physical qubits per logical qubit, with its error rate still unsimulated, and the 30× figure belongs to codes the authors didn’t simulate and haven’t yet given an access scheme. For anyone reading CRQC resource estimates, Google’s included, the request to make of every author is storage, access, compute, factories and decoding under one error budget.