Models & Research

OpenAI publishes 722 maths manuscripts from an internal model, three hours of compute per result

3 min read AI-generated

The model was handed roughly 4,000 problems. Two results fell outside the standard procedure: a zero-free region for the Riemann zeta function, and the Hodge Conjecture for CM abelian varieties.

Featured image for "OpenAI publishes 722 maths manuscripts from an internal model, three hours of compute per result"

On October 6, OpenAI put the mathematical output of an unreleased internal model into a public GitHub repository. The README gives the scale: 722 manuscripts sorted into 372 families, where a family groups a principal result with its companion arguments, consequences and alternative proofs.

The reason for publishing isn’t only the content. OpenAI writes that its existing mathematics evaluations saturated, so it moved on to open research problems instead.

How the results were produced

The model was posed roughly 4,000 problems. On average, a result cost the equivalent of three hours of ChatGPT Pro thinking compute. The output was then aggregated into families and manuscripts wherever the significance was judged sufficient.

Two cases ran differently: a zero-free region for the Riemann zeta function, and a proof of the Hodge Conjecture for CM abelian varieties. The writeup of the Re(s) > 11/12 region was also edited by a human for readability.

For ten results OpenAI is releasing abridged summaries of the model’s reasoning, among them the irrationality exponent of π, the Mahler conjectures, Kaplansky’s direct-finiteness conjecture in characteristic two, and the Mézard-Parisi formula for diluted spin glasses.

The caveat sits in the README, not in the announcement

Many proofs come with a Lean formalization, machine-checkable, and OpenAI says more will follow. But not all of them. The README carries the sentence that matters: the collection includes results at different stages of verification, and some of the unformalized ones could have issues.

That is the honest version. It also means the number 722 does not describe 722 checked theorems.

An outside group advised on the format

For the shape of the release, OpenAI consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study and followed its public recommendations. That is where the repository’s protocols for revisions and citations come from, along with the stated intention to keep looking for a community-hosted home.

OpenAI is also funding workshops, conferences and programmes aimed at understanding AI-produced results at all. And one line carries more weight than the rest: the company says it is working to responsibly release the model that produced this work.

The verification burden moves to the readers

In September, Anthropic used Claude to formalize Fermat’s Last Theorem in Lean — one proof, once, fully machine-checkable. OpenAI is going the other way: a lot of material, partly formalized, with an explicit note that errors may be in there.

Both are legitimate, and both move the same work around. Read 722 manuscripts of which some are not machine-checked, and you decide for yourself what you believe. The advisory group and the Lean files exist precisely for that, and without them the release would be a claim at volume. That OpenAI puts the caveat in the README rather than the announcement is the one place I’d have liked more nerve.

Sources

OpenAIMathematicsLeanResearchBenchmarks