Post-Quantum, PQC, Quantum Security

When Two Cryptographers Independently Confirm a Break, Ask Which Model They Used

Introduction

On Thursday, July 23, Yao-Ting Lin sat in a meeting at UC Santa Barbara with a proof in his inbox he had not managed to open. An MIT graduate student had emailed it to him that morning. Lin could not look, because his adviser, Prabhanjan Ananth, was at that moment describing a proof of the same result. Lin wrote back afterward to say the times had gotten strange.

The two papers, Unconditional Unclonable Encryption and Efficient Unclonable Encryption from Pauli Eigenstates, went to arXiv three hours and eighteen minutes apart. Both credit OpenAI’s GPT-5.6 Sol Ultra with producing the core ideas. Both trace to the same open problem, raised at the same Simons Institute talk in Berkeley earlier that month. Neither author knew the other was working on it until Lin recognized the overlap and put them in touch, and neither paper was public yet when he did. Peter Hall reported the sequence for Scientific American on July 31.

I’m not going to spend this article on what those papers prove, which is a research primitive with no path into an enterprise estate. My bottom line, for anyone who briefs a board or runs a migration program: nothing this summer changes an algorithm decision, and if you were rolling out ML-KEM on schedule, roll it out on schedule. What has changed is the reliability of the process that establishes when a cryptanalytic claim is real. One word does most of the work in that process, and two events this summer suggest it no longer means what it meant in 2024.

The 2024 version of that word is where to start.

What independence bought two years ago

On April 10, 2024, Yilei Chen posted a claimed polynomial-time quantum algorithm for LWE on the first day of the NIST PQC Standardization Conference. Had it held, and had the result carried through the relevant reductions and parameter regimes, it would have put ML-KEM, ML-DSA, the draft FN-DSA standard and LWE-based homomorphic encryption on the clock. Cryptographers worldwide dropped what they were doing and read it.

On April 18 Chen posted an update. Step 9 contained a bug he did not know how to fix. Then came the acknowledgment: “I sincerely thank Hongxun Wu and (independently) Thomas Vidick for finding the bug today.”

Eight days.

That parenthetical is the most consequential word in modern cryptanalysis, and I no longer think we can use it the way we used it then.

Mainstream public-key cryptography has no unconditional proof that the hardness assumptions under it are true. It has formal reductions resting on those assumptions, years of failed attacks, and a social process that tests both. Someone claims a break, the community grades the claim by attacking it, and the grade that carries weight is whether separate people, reasoning separately, arrive at the same verdict. Two of them did, neither having seen the other’s reading, which is why nobody had to wait eighteen months for a program committee to rule on whether the foundations of the lattice standards were intact.

Beyond expertise and effort, the mechanism needs something less often stated. It needs failure modes that are not governed by the same cause. Two readers are useful to the extent that they are likely to miss different things. Wu and Vidick share a literature, a set of teachers, and probably a set of instincts about where Gaussian arguments go wrong, and that correlation was fine, because what the process needs is that when one of them makes a mistake, the other one is not making the identical mistake for the identical reason.

Both of this summer’s tests bear on whether that requirement still holds.

What the July papers actually did

Ananth is at UC Santa Barbara; his co-author Amit Sahai is at UCLA. Their paper went up at 10:35 a.m. Pacific, and Seyoon Ragavan’s, from MIT, followed at 1:53.

Unclonable encryption asks a quantum ciphertext to enforce a rule about time. An adversary who splits it before the key is revealed should not be able to give two separated parties a usable copy each. Both papers construct the same scheme, a random non-identity Pauli as the key with the message bit encoded in one of its eigenstates, and prove an adversary cannot beat a coordinated guess by more than an exponentially small margin. It is one bit, one time, private key, and it transmits quantum states. It says nothing about anyone’s production estate.

The underlying construction is also not new, which the coverage has managed to obscure. Pierre Botteron, Anne Broadbent, Eric Culf, Ion Nechita, Clément Pellegrini and Denis Rochette published a closely related scheme in Quantum on July 8, proving the bound only in restricted cases and supplying numerical evidence beyond them. Ananth and Sahai say outright that their construction was first conceived by that group, and that what they add is the proof of the operator-norm bound those authors had conjectured; Ananth told Scientific American he recognized the scheme just before posting. Ragavan describes his own full-Pauli version as related but distinct and says his techniques do not settle the earlier conjecture.

Both papers add efficiency over Bhattacharyya, Broadbent and Culf, whose unclonable bit exists but whose encryption and decryption run in time exponential in the security parameter, as Ananth and Sahai’s own footnote spells out.

So when phys.org announced a new quantum encryption method that stops ciphertexts from being cloned, it compressed away both the prior art and the shape of the guarantee, which bounds what an adversary can do with one ciphertext rather than prohibiting copying outright. The primitive genuinely is quantum cryptography built on no-cloning; Ananth and Sahai’s first sentence says so. It is not QKD, and it is not a route into anything an enterprise runs.

The production method is the part with consequences beyond this one result.

Ragavan worked hands-on, running the system in two-hour stretches, checking progress at each interval and redirecting it when it wandered. Ananth and Sahai went the opposite way, running the same model inside a bespoke UCLA harness built to make models generate and then attack their own candidate solutions. Two workflows about as different as two workflows can be. One model, one seminar, one result, three hours and eighteen minutes apart.

Ragavan, to his credit, wrote down exactly how it went. His paper carries a statement on AI use recording that the agents made progress across successive rounds, that a complete solution arrived in the sixth, that they drafted the first version of the paper and he reorganized it, and that GPT-5.6 Sol and Claude Fable 5 handled proofreading.

He also shipped a Lean 4 formalization of the security proofs, written with Codex, and stated what it does not cover. The efficiency claim is not formalized, and neither is one corollary. That example returns at the end of this article, because it answers most of what I am about to complain about.

Ananth’s description of the new operating mode is the sentence to carry out of the whole episode: when someone mentions an open problem, “the first thing is to see if GPT solves it.”

That describes constructions. It describes attacks identically.

Simon’s claim is being adjudicated right now

The cryptanalytic version of this is not hypothetical. As I write, on August 10, it is happening on a Discord server.

Since Daniel Simon’s claimed polynomial-time algorithm for the Dihedral Coset Problem appeared on August 6, the qcryptanalysis server has hosted the substantive work of taking it apart, with participation from Daniel J. Bernstein, Elena Kirshanova (a co-author of the BKSW reduction the claim leans on), Daniel Apon and Simon himself. I covered the claim and the two gaps the proof depends on when it landed: Lemma 3 has a pairwise-independence failure with no confirmed repair, and Lemma 4 has a defect that forces a branch ratio positive and hides the catastrophic case.

The stakes here are not a research primitive. If the claim holds, and if the reduction chain reaches the concrete assumptions in deployed use, the hardness assumptions under ML-KEM, ML-DSA and FN-DSA are in a different position than they were in July. If it dies, it dies from the work now circulating. So look at what that work actually is.

Apon, now at Anduril, circulated a paper called “The Search for Simon” formalizing the missing piece as an Adaptive Quantum Subspace Energy theorem, and proving that the generic proof strategy admits no local repair. It is a serious and useful contribution. The author line, evidently in jest but printed as a byline, reads: Daniel Apon, Anduril Industries. Mr. Claude, Mrs. GPT.

I laughed, and then stopped laughing, because the joke identifies a real gap.

The field is not without rules here. IACR has had a policy on author use of AI tools since May 2025, applied at Crypto, Eurocrypt and TCC, and its own journal states flatly that generative AI tools cannot be listed as authors under any conditions. Substantial generated content requires an appendix naming tools, versions and prompts. So a byline like Apon’s would not survive submission to the venues that publish this field. What those rules govern is formal submission, and Simon’s claim is not being adjudicated through formal submission. It is being adjudicated in circulating notes, on a timescale of days, where no rule applies and no vocabulary exists for weighing what two model names contributed relative to the human one.

The second substantive response is a note by Michael David Gardiner, headed “draft prepared with AI assistance,” which identifies a post-selection gap in Simon’s analysis. Two of the concentration bounds are stated to hold with probability 1 − O(1/n), and the algorithm only reaches its final step after a survival event of probability Θ(1/n), so the conditional guarantee is O(1), which may exceed one. Vacuous at exactly the order the whole argument operates. The note then repairs the gap without new hypotheses, and in passing repairs a second defect in Lemma 3 by bounding a ratio rather than its numerator and denominator separately.

The question that started this article applies directly to the Simon case. When the verdict arrives, and it will arrive within weeks rather than years, how many independent verifications will stand behind it?

Part of the answer is reassuring. Bernstein’s contribution to Gardiner’s note is finite integer arithmetic about where the rounding of logarithms costs margin, checked in both directions and machine-checkable by anyone. Kirshanova is reading the reduction as the person who helped build it. Those are human eyes doing what human eyes have always done.

Part of the answer is that several of the most substantive documents in the adjudication were produced by humans working with models, and that what has been disclosed about which models is uneven. Ragavan and the Ananth–Sahai paper name GPT-5.6 Sol Ultra explicitly. Apon’s byline names Claude and GPT without a version. Gardiner’s header says only that the draft was prepared with AI assistance. Only the first of those three tells a reader anything about whether two documents share an upstream dependency.

I’m not suggesting the work is wrong. Everything I have read in this adjudication looks careful. I am saying that if three notes converge on the same defect in Lemma 3, we currently lack the provenance to know how much independent weight those three convergences deserve.

Two objections, and what I think of them

The first objection is that different prompts and different workflows preserve independence, because the model is not a lookup table and two people who approach it differently will get different reasoning.

We have one warning already. Ragavan’s supervised two-hour iteration loop and the UCLA self-critiquing harness are about as far apart as two serious workflows get, and they produced the same construction and the same proof strategy within a single working day.

That is convergence on a correct answer, not correlated error, and the difference is real. A canonical solution, or a problem that heavily constrains the useful route, would produce the same picture. It does not establish that the same setup would miss the same flaw twice. What it does establish is that workflow diversity is not self-evidently enough, which means independence is now something to demonstrate rather than assume. Error-finding may behave differently, since the search space is one paper rather than the space of possible proofs, and I would like to see someone test that properly rather than assert it in either direction.

The second objection is better, and it is the one I would raise if I were reading this. Human reviewers were never independent either. Wu and Vidick learned lattices from the same literature, absorbed the same instincts about where Gaussian arguments fail, and were reading the same paper in the same week under the same social pressure. Correlated priors have always been the norm.

True, and peer review has always contended with school-of-thought bias. What is new is the compression, because a whole school now fits inside one checkpoint. Two humans with similar training still make different mistakes, at different points, for different reasons, and their errors decorrelate even when their education does not.

Be precise about the mechanism, because it is easy to overstate. Two runs of the same model can produce different outputs, and conditional on the sampling they may even be statistically independent. That is not the independence cryptanalysis needs, and it is not the thing to worry about. The risk is common-cause error: shared weights, training data, post-training, retrieval and tooling reproducing the same blind spot across separate operators who have no way of noticing that they share it. Separate human names are no longer sufficient evidence of separate failure modes. That is a narrower claim than saying two analyses reduce to one, and it is the one I will defend.

There is a third objection, which is that none of this matters because the models are getting good enough that a single verification suffices. Maybe eventually. Not now, and whoever wants to retire the mechanism that caught Chen’s Step 9 in eight days should have to prove it first.

What I actually want, and where somebody has already done it

Generic disclosure statements are close to useless. A footnote saying a model was used tells a reader nothing about which claims to trust, and treating it as a compliance checkbox produces the same theater that “AI-assisted” labels produce everywhere else. The IACR rules ask for more than that, and they are a real improvement, but they answer a question about authorship and responsibility rather than the question a reader of a cryptanalytic note actually has, which is how much weight to put on any particular claim in it.

Two people have already shown what the better version looks like, and neither was trying to set a norm.

Ragavan’s paper is the first of them. It names the model, describes the round structure, says which parts the agents drafted and which he rewrote, and ships a machine-checked formalization with an explicit statement of what the formalization does not cover. A Lean kernel is the one verifier in this entire story that shares no weights with anything. Whatever wrote the code, the proof either type-checks or it does not. That is the strongest available answer to the problem this article describes, and it is sitting in a public repository right now.

Gardiner’s note does the other half, for the case where formalization is not on the table.

It ends with two sections, and the first, headed “Who found what,” assigns each finding to a person: the post-selection gap came from an earlier adversarial reading circulated separately, §2.1 is entirely Bernstein’s down to the observation about rounding, and the A-side degeneracy correction is Apon’s, whose catch overturned a claim the note itself had asserted as exact. The second is headed “Verification status” and grades the document against itself, unevenly and on purpose.

Bernstein’s arithmetic is flagged as the highest-confidence content, machine-checkable, checked in both directions. He grades the Bayes step elementary, and then the actual new contribution, Lemma C and Theorem 1, gets this:”Checked carefully, but not adversarially and not by a second person.” Followed by an explicit statement that if one normalisation is wrong, two entire sections fail together. One remaining question is marked open and identified as inherited framing rather than a result.

A reader can act on that. A reader can also see, without being told, which parts had genuinely separate eyes on them and which had one pass in one session.

Two sections at the end of every pre-peer-review cryptanalysis note. Who found what, and what has been checked by whom, with what, and how hard. That is the convention I want, and I would rather the community write it now, while the stakes are a research primitive and a contested lemma, than after the first genuine break arrives and three teams announce independent confirmation that turns out to be one.

What a security program does with this

Nothing here is an argument for changing an algorithm. It is an argument for changing how you read the news, and it has three practical consequences.

When a cryptanalytic claim breaks and the reporting says multiple researchers have confirmed it, that phrase now needs a follow-up question. Ask whoever briefs your architecture board whether the confirmations came from separate reasoning or separate prompts. They may not know. The fact that they may not know is the point, and it should change how much you rely on that confirmation in the first 48 hours.

Expect more claims, arriving faster, in both directions. A model that closes a problem open since 2019 two weeks after hearing it stated at a seminar can be pointed at cryptanalysis by anyone, including people with no intention of publishing what they find. That means more false alarms and more real results, on a compressed timeline, with the verification process under strain at exactly the moment it needs to be slower and more careful. The quantum panic industry will have an excellent year. Separating its output from the genuine results is about to get harder, and I say that as someone who does it for a living.

And the architectural answer is the same one it has been for three years, which is either boring or reassuring depending on your temperament. Crypto-agility does not let you skip the epistemological argument. It lowers the cost of being wrong while the argument is unresolved.

An organization that can swap a primitive in weeks can wait for better evidence on whether Lemma 3 is repairable, stage a contingency, or reverse course, without rebuilding the estate around a judgment that may not hold. An organization that cannot swap has to get an uncertain scientific question right on the first attempt, on a deadline set by someone else. That is the same conclusion the regulatory and client deadlines reach by a completely different route, and it is why the PQC Migration Framework treats agility as the deliverable rather than as a nice-to-have.

Ragavan, three years into a PhD, told Scientific American that the way he does research now bears no resemblance to how he did it two months ago, and that adapting is the only move available. He was describing his own career. He was also describing a set of verification norms the rest of us have been relying on without examining, built for a world in which two people reading a paper meant two people reading a paper.

Somebody should write the replacement down. A public Lean repository and the last two pages of a circulated note are further along than anything a committee has produced.

Marin Ivezic

I am the Founder of Applied Quantum (AppliedQuantum.com), a research-driven consulting firm empowering organizations to seize quantum opportunities and proactively defend against quantum threats. A former quantum entrepreneur, I’ve previously served as a Fortune Global 500 CISO, CTO, Big 4 partner, and leader at Accenture and IBM. Throughout my career, I’ve specialized in managing emerging tech risks, building and leading innovation labs focused on quantum security, AI security, and cyber-kinetic risks for global corporations, governments, and defense agencies. I regularly share insights on quantum technologies and emerging-tech cybersecurity at PostQuantum.com.