Coedee Deep Dive

The Question This Story Answers

Can an AI model discover something true about the world that no human has ever proven? For nearly eighty years, a problem posed by Paul Erdős said no one knew. This year, a model changed the answer.

Mathematics has always been the field AI skeptics point to when they want to draw a hard line around what machines can’t do. Chess, sure. Protein folding, fine — that’s pattern matching on biology. But pure mathematics, the kind where a claim has to be airtight and a single gap in logic invalidates the whole thing, was supposed to be the place where a model’s tendency to sound confident and be wrong would finally catch up with it. That line moved this year, and the conversation still hasn’t caught up with what happened.

The Problem That Sat Unsolved for 80 Years

In 1946, the legendary Hungarian mathematician Paul Erdős posed what became known as the unit distance problem: if you scatter a set of points across a flat plane, what is the largest possible number of pairs of points that can sit exactly one unit apart? It sounds like a puzzle a student could sketch out on graph paper. It is not. For eight decades, the prevailing belief among mathematicians was that a simple square-grid arrangement of points was essentially the best anyone could do, and that belief hardened into something close to consensus even though nobody had proven it.

Then, this spring, an internal reasoning model at OpenAI was given the problem and, without pulling the answer from existing literature, constructed an entirely new family of point arrangements using deep tools from algebraic number theory — ideas connected to work by mathematicians including Golod, Shafarevich, and Ellenberg-Venkatesh. The construction didn’t just chip away at the old bound; it broke it, producing a polynomial improvement over the square-grid arrangement that had stood unchallenged since Erdős first posed the question. A companion academic paper co-authored by a group of outside mathematicians, including Fields Medalist-adjacent researchers and number theorist Will Sawin, who independently refined the model’s result, confirmed the argument held up under scrutiny.

Harvard mathematician Melanie Matchett Wood, who contributed to the verification effort, described the result as a genuinely beautiful piece of mathematics — not a technicality, and not a fluke.

Why This One Is Different From the Hype Cycle

AI-and-math claims have burned people before. Models have confidently produced proofs riddled with subtle errors, and headlines have occasionally gotten ahead of what actually held up under peer review. What separates this case is the verification trail. The proof wasn’t accepted because a model said so; it was checked, digested, and partially reformulated by working mathematicians who published their own companion paper walking through the argument in human-readable form. One of the paper’s authors even noted, with a mathematician’s characteristic understatement, that the result doesn’t introduce dramatic new geometric machinery — it’s a clever, verifiable application of existing deep tools, which in some ways makes it more convincing, not less.

That distinction — an AI system independently discovering a construction rather than retrieving or recombining one from training data — is the detail mathematicians keep returning to. It’s also not an isolated case. In the weeks around the unit distance result, other independently reported wins piled up: a separate open problem attributed to Erdős was reportedly settled using an internal OpenAI model just weeks earlier, a DeepMind team used their own model to resolve several open problems with machine-checked Lean proofs attached for verification, and researchers studying electrical flow in graphs used an internal model to settle a standing question in that area too. Even Fields Medalist Terence Tao has publicly walked through AI-assisted work on a related open conjecture, sparking one of the most heavily discussed AI-and-math threads of the year among working mathematicians.

The Guardrails Conversation Nobody Wants to Skip

Success stories like this one have prompted almost as much discussion about process as about the result itself. Mathematicians who’ve watched AI systems produce dazzling-looking but flawed arguments in the past are pushing for norms before the next headline-grabbing claim: independent verification before a public announcement, machine-checked formal proofs (using systems like Lean) wherever possible, and clear disclosure of exactly how much of a given argument came from a model versus from human refinement afterward. The unit distance result mostly followed that playbook — a companion paper, named human verifiers, and a clear disclosure that the initial construction was AI-generated and later expositionally refined through human collaboration. Not every AI math claim since has been handled with the same rigor, which is precisely why the community is trying to lock in standards now, while the stakes are still mostly reputational rather than something riskier.

Why a Math Proof Is a Bigger Deal Than It Sounds

It’s worth being precise about why mathematicians, an unusually hard-to-impress group, reacted the way they did. Pure mathematics has no room for the kind of fuzzy, “mostly right” output that’s acceptable in a chatbot answering trivia. A proof is either airtight or it’s wrong, and there’s no partial credit. That makes it one of the cleanest possible tests of whether a model is doing something closer to genuine reasoning versus an extremely convincing form of pattern completion. A model that can produce a construction this novel, in a field this unforgiving, and have it independently verified by domain experts, is clearing a bar that many researchers assumed was still years away.

What Comes Next

None of this means AI is about to replace mathematicians, and researchers close to the work have been careful to say so. What it does suggest is a shift in what “AI as a research collaborator” actually looks like in practice — not a tool that answers questions from a database, but one that can be handed a genuinely open problem, work through it with techniques a human might not have thought to combine, and produce something new enough that experts have to slow down and check it carefully rather than dismiss it on sight.

For a field that has spent decades defining itself partly by what machines couldn’t do, that’s not a small adjustment. It’s the start of a different kind of collaboration — one where the hardest open questions in a discipline become fair game for an AI system to attempt, provided the humans checking the work are just as rigorous as the mathematicians who spent eighty years unable to crack the problem themselves.

How a Model Actually “Discovers” a Proof

It’s worth demystifying what happened here, because the word “discovered” can sound more mysterious than the underlying process actually is. Modern reasoning models work through a problem step by step, generating and checking intermediate claims much the way a human mathematician sketches partial arguments on a whiteboard, discards the ones that don’t work, and follows promising threads further. What made the unit distance result unusual wasn’t the mechanism — it was the destination. The model wasn’t recombining a known technique in an incrementally clever way; it connected tools from algebraic number theory to a geometry problem in a way that, as far as the verifying mathematicians could tell, nobody had attempted, let alone written down, in eighty years of people trying.

That’s also why the human verification step matters so much and can’t be skipped, no matter how confident a model’s output sounds. A model can produce an argument that reads as fluent and internally consistent while still containing a subtle logical gap a domain expert would catch immediately. The difference between this result and the cautionary tales from earlier AI-math attempts is that the unit distance proof survived exactly that kind of adversarial scrutiny from mathematicians actively looking for the flaw.

Beyond Pure Math: Where This Is Already Spilling Over

The unit distance result hasn’t stayed contained to discrete geometry. Researchers in adjacent fields have started using the same class of internal reasoning models as active collaborators on their own open problems — not as a search engine for known results, but as a genuine thinking partner during the early, messy brainstorming stage of a proof, well before anything is ready for formal write-up. One recent paper on a separate, unrelated combinatorics problem explicitly credits AI tools with contributing to the brainstorming phase, while noting the core methodology and final substantiation remained the human author’s own work — a disclosure pattern likely to become standard practice as more mathematicians experiment with this kind of collaboration.

There’s also a growing formal-verification angle worth watching. Separate teams have paired AI-generated mathematical arguments with automated proof-checking systems like Lean, which can mechanically verify that every logical step in a proof is valid — removing much of the “trust me” element from AI-assisted mathematics and replacing it with something closer to a compiler that either accepts or rejects the argument outright. As that pairing becomes more common, the debate over whether to trust an AI-generated proof may increasingly resolve itself: not through argument about the model’s reliability, but through mechanical verification that doesn’t care who or what produced the claim in the first place.

Leave a Reply

Your email address will not be published. Required fields are marked *