Matt von Hippel used to be a theoretical physicist. Now he writes about physics, and in August he threw down a challenge to the AI labs. If you want to impress me, take on my old field. Compute a scattering amplitude everyone assumes is too expensive, and do it on the compute budget a university group actually has. He named two targets. One of them: the six-particle amplitude in planar N=4 super Yang-Mills at nine loops.
At the end of August, Liam Fitzpatrick and Siddharth Mishra-Sharma, two physicists at Anthropic, got in touch. They had it.
Why nine loops was the wall
Scattering amplitudes tell you how likely particles are to react in particular ways. The more loops a calculation includes, the more accurate it gets and the more it costs. Most real-world amplitudes stop at two loops. A few reach three. The most precise prediction in particle physics you may have heard of used five.
N=4 super Yang-Mills is a toy model amplitude researchers use to stress-test their methods. Lance Dixon, a professor at SLAC and Stanford, had reached eight loops a few years ago, and by a roundabout route at that. Nobody had nine. Dixon assumed the next step would need an even more indirect path. His reasoning is the best line in the whole story: if you could simply run the standard bootstrap one loop further, somebody would have done it already.
One sentence of prompt, then six hours of silence
Anthropic did not throw millions of dollars of compute at this. The work ran on Fable 5.1 inside Claude Science, the platform scientists can pay for. The prompt boiled down to: compute the six-particle (hexagon) amplitude in planar N=4 SYM at nine loops. After that came mostly one kind of message, and it will sound familiar to anyone who leaves an agent running overnight. I’m going to sleep and won’t be available for several hours. Keep working until I tell you to stop. Update me every four to six hours.
Claude then did the calculation twice, once via the classic bootstrap and once via the indirect form-factor route. For an end user that would have cost one or two thousand dollars, almost all of it model runtime. The actual number crunching, Python with SymPy, came to about $100: 96 CPUs for a week. Dixon checked the result. A group around Song He at the Chinese Academy of Sciences was close behind with GPT-6 assistance, and had most of the answer already.
The remarkable part is what the calculation didn’t need
Von Hippel is a little disappointed, and that is what makes his piece good. He had hoped an AI would get around a computational wall in some surprising way. Instead Claude used known methods with slightly more compute than anyone had bothered to spend before. His conclusion: there is more low-hanging fruit out there than experts assume.
For us the point sits somewhere else. These calculations are fiddly and error-prone. Von Hippel writes that he would have turned one week of CPU time into two, guaranteed, because he’d have broken something on the first pass. The harness carried Claude to the end with no scientific supervision. No colleague checking whether an intermediate result made sense. Just “keep going.” If you still think of AI as too unreliable for that kind of job, read that sentence twice. As with Claude’s enzyme system find, the result is less startling than the route to it.