Google’s Multiverse Claim Was a Serious Argument Attached to the Wrong Experiment
Table of Contents
On December 9, 2024, Google announced a 105-qubit processor called Willow. Partway through the announcement, Hartmut Neven, the founder and lead of Google Quantum AI, wrote that the chip’s benchmark score “lends credence to the notion that quantum computation occurs in many parallel universes.” He attributed the idea to the physicist David Deutsch, then moved on to the next section. Nobody else moved on. Within two days the story had become Google proving the multiverse. Twenty months later that one sentence still produces more questions to me than anything else a quantum executive has published.
Willow produced no evidence for parallel universes, and the benchmark number behind that claim was never the kind of number capable of supplying any.
He was neither joking nor selling. He was reaching for an argument a serious physicist published in 1985, one that has never been settled in the four decades since. His mistake was a great deal narrower than the coverage suggested. He attached the strongest available case for the multiverse to the single experiment least able to support it.
What Willow Actually Did
The Willow announcement contained two results, of which the one that traveled around the world was easily the lesser.
The serious result, error correction, is the older of Google’s two problems. Quantum bits are fragile things, and physical operations on them fail often enough that a working quantum computer must spend most of its hardware protecting a small amount of usable information. The leading scheme is the surface code, which spreads one protected logical qubit across a square array of physical ones, sized by what engineers call the code distance. The scheme produces a net gain only if a larger distance produces a lower error rate. For nearly three decades no processor had definitively shown that it does. Google’s own 2023 experiment got partway there, with a distance-5 code modestly beating a distance-3 code.
Willow settled the question. In the team’s Nature paper, surface code memories at distances 3, 5 and 7 each cut the logical error rate by a factor of 2.14 ± 0.02 relative to the distance below, with the 101-qubit distance-7 memory reaching 0.143% per cycle. That is below-threshold operation, and I rate it the most important experimental result the field has produced this decade.
The second result was a benchmark of a familiar kind. Random circuit sampling (RCS) runs a randomly chosen sequence of gates, reads out the qubits, then repeats. The output is not an answer to a question. It is a stream of bit strings drawn from a probability distribution the circuit itself defines, and the hard part for a classical computer is reproducing that distribution. Google reported that Willow finished the sampling in under five minutes, and estimated that Frontier, then the world’s second-ranked supercomputer behind El Capitan, would need 10^25 years to match the result. Neven attached his multiverse sentence to that estimate of classical cost, and error correction never enters the argument at any point.
The estimate needs three qualifications before anyone leans on it. Google supplied the first one itself: the same post acknowledged that RCS has no known practical use, which is unusually candid for a launch. Second, the septillion is the high end of a range whose low end is very much shorter. Google’s own figures, as Scott Aaronson reported them, put the same problem at roughly 300 million years for a classical machine with no memory limit. Both figures are correct, and only one of them made the headline. Third, Google benchmarked Willow against the best classical algorithms anyone had written by December 2024, and better ones arrive every year.
After the 2019 Sycamore result claimed 10,000 years, an IBM team led by Edwin Pednault published an analysis of the same circuits. Using disk storage the original comparison had ignored, they estimated Summit could simulate the same circuits in a matter of days, an estimate they never actually ran. Classical algorithms have improved since 2019 without closing the gap. I don’t expect the Willow figure to shrink the way Sycamore’s did. But no septillion of this kind is a physical constant. Every one of these figures measures a single quantum run against the best classical algorithms, hardware and memory assumptions available on the day it was published.
The Quote He Never Wrote
Before the argument, a correction nobody seems to have made. Newsweek, Vice and a long tail of aggregators have all reported that Neven said Willow was so fast it “had to have ‘borrowed’ the computation from parallel universes.” That sentence appears nowhere in Google’s post. I checked three versions: the page as it stands today, the February 2025 archive, and the first capture taken on December 9, 2024, within hours of publication. Every one of them carries the sentence I quoted at the top and nothing else on the subject.
The mutation from paraphrase to quotation happened in two steps. TechCrunch paraphrased Neven on December 10 as saying the chip must have borrowed computational power from other universes, and it used no quotation marks. Newsweek printed the line on December 11 inside quotation marks, attributed to the blog post. Vice repeated it in the same form on December 12, and the aggregators have been copying that version ever since.
In the fabricated quote, the chip “had to have ‘borrowed’ the computation from parallel universes.” Neven wrote only that the benchmark score “lends credence to the notion.” Those are different claims, and only one of them was ever made.
The invented quote is still circulating, still inside quotation marks, still attributed to a blog post that never contained it.
We have been here before at considerably higher stakes. On November 6, 1919, Frank Dyson and Arthur Eddington presented eclipse photographs from Sobral and Príncipe to a joint session of the Royal Society and Royal Astronomical Society, reporting that the measurements favored Einstein’s predicted deflection of starlight over Newton’s. The photographic evidence was mixed, with Príncipe yielding two usable plates, the stronger measurement coming from a backup instrument at Sobral, and Eddington saying afterwards that more eclipses would be needed.
The Times of London ran the story the next morning under “Revolution in Science / New Theory of the Universe / Newtonian Ideas Overthrown,” as Clifford Will recounts in his history of the measurement. Einstein was a household name by the weekend, in countries that had been at war with his own a year earlier.
Later measurements vindicated general relativity many times over, and that is not the point. The press reported a revolution in physics on the strength of two usable photographic plates and a single backup instrument. A century on, the same thing happens in 48 hours instead of a week, with the source document one click away from everyone repeating the story. I write about how the press handles quantum results, including recently. The real sentence is the one to argue with.
Deutsch’s Challenge
Now the argument Neven was reaching for, at full strength. The many-worlds picture came from Hugh Everett, who proposed in a 1957 paper that quantum states never collapse at all. On his account the observer and the apparatus join the superposition during a measurement, every outcome occurs, and each occupies its own branch of one universal wavefunction. Neven credited the multiverse prediction to Deutsch – the link he used pointed at Deutsch’s 1997 book The Fabric of Reality – but the prediction belongs to Everett, twenty-eight years earlier. Deutsch contributed something different and, for this argument, larger.
In 1985, in Proceedings of the Royal Society A, Deutsch defined the universal quantum computer. He argued that the intuitive explanation of its behavior puts unbearable pressure on every interpretation except Everett’s. His case runs through a worked example: a quantum program that computes two things in the time a classical machine needs for one, and on a good day yields the combined result. His question about that day is four words long. “Where was it computed?”
That question is the whole argument, and it gets harder as the machines get bigger. A fault-tolerant run of Shor’s algorithm on a 2048-bit number would evolve through a state space whose dimension dwarfs the roughly 10^80 atoms in the observable universe. That is a fact about the mathematics, and it counts no universes and no completed calculations. Deutsch’s point concerns what comes out the other end, where a successful run produces candidate factors that a second of ordinary arithmetic confirms. So where did the work happen? Deutsch has been asking that question for four decades without ever accepting an answer.
Deutsch’s challenge requires two things from the output: a single correct result, and independent verification of it. All its force comes from holding a verified result in hand and asking where the result was produced.
Sampling and the Missing Answer
Willow ran RCS, and a sampling run does have a correct target, the probability distribution the circuit defines, but no single answer to hold up. Correctness there is a statistical property of the whole set of samples, measured by cross-entropy benchmarking, which scores how well the observed samples correlate with the intended distribution.
Google did not emphasize what follows from a number that large. Scott Aaronson pointed out at the time that the largest Willow circuit would take 10^25 years to score directly, for exactly the reason it would take 10^25 years to simulate, so verification of that run is indirect, extrapolated from smaller circuits a classical computer can still check. Aaronson sees no reason to doubt the extrapolation, and neither do I.
Set those two facts side by side for a moment. Deutsch’s challenge needs a verified correct answer whose provenance is mysterious, where Willow’s headline run produced no single answer and had nobody confirm its fidelity directly. Neven picked the wrong result out of his own announcement.
The second problem is with the number itself. A figure of 10^25 years describes how long classical hardware would take under one particular memory assumption, and it counts nothing whatever about branches. Every interpretation that reproduces standard unitary quantum mechanics and the Born rule predicts the same Willow samples, which is what makes these accounts interpretations instead of rival theories.
Aaronson wrote at the time that the experiment “doesn’t add anything new to this old debate,” and in the comments he put it more sharply. A Bohmian or a QBist made exactly the same prediction for Willow and had no reason to update on it. Anyone the multiverse case persuades should already have been persuaded by the double slit. Whatever moved in anyone’s head after Willow was psychology.
The pushback against Neven overshot in places as well. Ethan Siegel, writing at Big Think, declared that quantum computation “does not occur in any parallel universes.” No experiment supports that either. Deutsch’s position is unproved, and the flat denial is the mirror error. I regard the multiverse reading as a legitimate minority position, one that capable physicists defend. No experiment in 2024 or the twenty months since has changed the case for or against it, and a product launch was never going to adjudicate a century-old question in the philosophy of physics.
What the Machine Is Actually Doing
The popular version of quantum computing says the machine tries every possible answer at once and returns the right one. That picture is wrong, and the same error makes vendor claims hard to evaluate.
A qubit in superposition is described by amplitudes – numbers with both a magnitude and a phase. Nothing in that description resembles a set of parallel attempts. Running the circuit steers those amplitudes so that paths leading to wrong answers cancel one another while paths leading to the right answer reinforce. That cancellation is quantum interference, and it is the entire craft. A machine that genuinely evaluated every branch and then handed over one at random would be useless, since the result would be a random answer. Arranging that cancellation across a large circuit is the whole engineering problem. Its difficulty explains why quantum algorithms are so rare, forty years of work having produced only a limited set of reusable algorithmic families with proven advantages.
The algorithms themselves are the evidence against the popular version. If brute-force parallelism across universes were the mechanism, unstructured search would be instantaneous, and forty years of experience says otherwise. Grover’s algorithm delivers a quadratic speedup on unstructured search and no more, a ceiling proved in the query model. Deutsch established a related limit in the 1985 paper itself, showing that quantum parallelism cannot improve the mean running time of a parallelizable algorithm.
Five Ways to Read the Equations, and One That Rewrites Them
The mathematics of quantum mechanics is among the most precisely tested frameworks in science, and physicists still disagree about what it describes. I have written about that distinction at length.
In July 2025, for the theory’s centenary, Elizabeth Gibney and colleagues at Nature ran the largest survey ever taken on the subject. They emailed more than 15,000 researchers and collected over 1,100 responses. Copenhagen led the field with 36 percent of respondents. Information-based approaches took 17 percent, a combined many-worlds and consistent-histories option 15, pilot-wave theory 7. Eighty-six percent said the attempt to interpret the mathematics is valuable.
A century in, the largest bloc is barely a third of respondents in a voluntary survey. Nature also issued a correction on August 12, 2025, acknowledging that the third category should never have combined those two views, since consistent histories involves no branching at all. Even the cleanest number anyone has on many worlds is contaminated by the questionnaire.
Copenhagen, in its textbook instrumentalist form, treats the wavefunction as a calculating device and declines the question of what happens between measurements. Generations of physicists were trained to shut up and calculate, and it remains the working default in most laboratories today.
Many worlds takes the equation literally and refuses to add a collapse rule. It trades that economy for an immense emergent branching structure nobody has managed to count. Its mathematics is the cleanest of the group and its ontology by far the most expensive.
Pilot-wave theory, from Louis de Broglie and David Bohm, restores definite particle positions guided by a physical wave. Determinism comes back at the price of an explicitly nonlocal guiding wave, though one nobody can use to send a signal.
QBism, from Christopher Fuchs, N. David Mermin and Rüdiger Schack, treats the wavefunction not as a description of the world but as an agent’s expectations. Collapse is that agent revising them, not a physical event, and the measurement problem never arises.
Relational quantum mechanics, from Carlo Rovelli, makes every outcome relative to the observing system, so a fact can be settled for one party and open for another.
Objective collapse models are the outlier, and they are the reason this section exists.
The Family That Makes Predictions
Giancarlo Ghirardi, Alberto Rimini and Tullio Weber proposed in 1986 that the wavefunction spontaneously localizes at a tiny rate per particle, negligible for a single atom and instantaneous for a cat. Lajos Diósi and Roger Penrose independently proposed that gravity does the collapsing, at a rate set by the gravitational self-energy difference between the superposed states. Both approaches modify the Schrödinger equation itself. Because these models modify the equation, they predict deviations from standard quantum mechanics that experimentalists can measure.
Certain mass-proportional collapse models predict that jostled charged particles emit faint spontaneous radiation, and experimentalists have spent two decades looking for it. Sandro Donadi and colleagues used an underground germanium detector to exclude the parameter-free formulation of the Diósi-Penrose model in Nature Physics in 2021. From the other direction, Yaakov Fein and colleagues in Vienna demonstrated interference with molecules beyond 25,000 atomic mass units, squeezing the same parameter space from above. I have my own speculative entry in this corner, an open-notebook piece on information-triggered collapse that I make no claims for beyond having written it down.
The Machines as Instruments
Here is the connection Neven had available and didn’t use. A quantum computer is a machine for holding an ever-larger system in coherent superposition for an ever-longer time. Error correction is what extends that coherence, and it is the regime where collapse models say something ought to give.
The qualification matters, because qubit count alone measures nothing about macroscopicity. The Diósi-Penrose rate depends on the mass-density difference between the superposed branches. A million superconducting qubits whose logical alternatives differ mostly in current and charge configuration might constrain a mass-density model far less than one carefully built matter-wave interferometer. Sensitivity depends on what the proposed collapse mechanism couples to and on how long coherence survives, and counting qubits will tell you neither.
Designed against a specific model, though, these machines become foundations instruments, and the engineers will get there before the philosophers do. Suppose an unexplained error floor appears one day and grows with system size, after decoherence and every conventional noise source have been ruled out. A generic floor of that kind would mean unresolved noise. A reproducible one following the scaling law of a particular modified-dynamics model would be the most interesting result in physics.
Deutsch saw this in 1985 and named the experiment. The test he proposed required placing an observer inside the quantum computer in coherent superposition, then interfering the resulting branches. He wrote that such a test would have to wait for both quantum computers and genuine artificial intelligence.
Kok-Wei Bong, Eric Cavalcanti, Howard Wiseman and their coauthors reached the same practical conclusion in a 2020 no-go theorem published in Nature Physics, noting that the required control becomes plausible if the friend is an artificial intelligence running inside a large quantum computer. Their theorem does not single out Everett as the winner. It rules out holding locality, no-superdeterminism, and the absoluteness of observed events all at once, and each interpretation must concede at least one.
An honest version of Neven’s sentence was available to him, then. It runs through the error-correction half of his own announcement, points at an experiment Deutsch specified forty years ago, and would have made a better story.
The Sentence Google Wrote Ten Months Later
On October 22, 2025, Google announced Quantum Echoes: the same lab, the same chip, Neven again on the byline, and a claim of verifiable quantum advantage running an out-of-time-order correlator 13,000 times faster than the best classical method. The post explains how the machine gets its sensitivity. The echo is amplified by constructive interference, “a phenomenon where quantum waves add up to become stronger.”
No universes anywhere in it. Interference, though, is no rival explanation to many worlds. An Everettian will tell you that interference is what branches do. Every interpretation listed above must reproduce the same phenomenon. What changed between the two announcements was the register. Google described its own result operationally, in the language every physicist shares, and left the ontology alone.
Google has published no correction that I can find, and the multiverse sentence remains on the Willow page. The page shows a last-updated date of June 12, 2025, so somebody revised the post at least once and kept the sentence. Set the two announcements side by side and you see a lab that changed how it talks without saying that it had.
The benchmark was never the important result in the Willow announcement, and the multiverse sentence’s problem is not that it was wrong. Google led with error correction and gave it most of the post. But the coverage ran with one speculative sentence and left the experimental result largely unreported.
Anyone who evaluates quantum claims for a living can take one dull, reliable rule out of that episode. The figure that makes the headline and the figure that changes an engineering estimate are seldom the same. The second one is the figure to go looking for.