The Riemann hypothesis has been open for more than 150 years. It’s about how prime numbers are distributed, and there’s a million-dollar bounty from the Clay Mathematics Institute waiting for anyone who delivers a general proof. That bounty is still unclaimed.
Anthropic didn’t crack it. But on Monday the company reported that an as-yet-unreleased model got surprisingly far: it pushed up the lower bound below which the hypothesis is provably true, by a meaningful margin. For a problem this size, that’s more than most people would expect from a language model.
How it actually happened
The best part isn’t the result — it’s the setup. An Anthropic staffer with no serious math background asked the model to “take a real stab” at proving the hypothesis, then left it running on its own for about a day and a half.
In that time the model worked through 650 different ideas, spread across 60 subagents, burning 31 million output tokens along the way. A footnote in the paper breaks down the division of labor: two of the 60 subagents developed the key mathematical ideas, 13 fed ideas to them, 30 tried and failed, 13 acted as validators checking the arguments, and two wrote the initial paper. Two of Anthropic’s in-house mathematicians confirmed the progress, and it was formalized in Lean, the open-source proof assistant.
Not a one-off anymore
This fits a pattern. Over the year, AI models have solved several Erdős problems. OpenAI recently showed off ten results from its internal “Astra” model, and Anthropic itself had already disproved the Jacobian conjecture. The models get better, the results get bigger.
Mathematicians have mixed feelings about it. In June, a group of prominent ones signed the Leiden Declaration, worried that AI erodes a core value of the field: that a proof is attributed to a person who takes responsibility for it. Fields Medal winner Timothy Gowers pushed back in a blog post, seeing it more openly. Maybe, he wrote, a world where theorems aren’t tied to specific mathematicians won’t be any more of a problem than the fact that stars aren’t named after astronomers — and most stars have no name at all.
My take
What grabs me here isn’t the math result — only specialists can really judge that. It’s the setup. A person with no math background types “give this a proper go,” goes to sleep, and a day and a half later there’s real progress on a century-old problem, checked and formalized in Lean.
That’s the moment these tools tip over: from “helps me phrase things” to “works through a problem on its own that I couldn’t even begin to solve.” Whether a name ends up under the proof is almost the smaller question. The bigger one is what else you could hand to a day and a half and 60 subagents.
Sources: