A week ago I wrote that an Anthropic researcher had quit and was leaving AI entirely. Since then, governments, talk shows and lab bosses have talked about little else. The Wall Street Journal has now interviewed the man and pieced together how he got there.
Hundreds of Chicken McNuggets in Hong Kong
Coxon is 27, British, the son of a professor of medieval German literature. He competed for the UK at the International Mathematical Olympiad: silver in 2016, bronze the year after. On a field trip to Hong Kong, some of the maths teams pooled their leftover food vouchers, bought hundreds of Chicken McNuggets at McDonald’s — prompted by Coxon and the Chicken McNugget Theorem from number theory — and then crowded around a whiteboard trying to crack a puzzle.
Three of the six on his team now work at AI labs. One is Joe Benton, who left Anthropic’s safety team in late August for METR, with a post on X that read a lot like Coxon’s later one.
After Cambridge, Coxon lived at Newspeak House in London, a co-living and event space for Britain’s rationalists. People he would meet again partied on its roof: Logan Graham, then an adviser to Boris Johnson, now head of the Anthropic team that stress-tests models for national security threats. Avital Balwit stopped by too; she is Dario Amodei’s chief of staff today. In between, Coxon briefly traded commodities and told friends AI was the most exciting opportunity on the planet.
Not a loud safety guy
He joined OpenAI in 2023 through a six-month residency for researchers with no AI background, and worked on pretraining. «Jacob is building the brain,» says his former colleague Will DePue, «then someone else trains the brain on how to behave in the real world.» He wasn’t especially vocal about safety — more of a normal researcher.
His turn has a date on it: the METR report in late August, which showed the Hugging Face incident was worse than anyone had known. Coxon calls that his moment of genuine alarm. He went to his managers, considered moving into safety research — and decided against it. His line on that: even working on safety at Anthropic felt like being complicit in the race.
A Slack message, then a park bench
On the day Coxon quit, his colleagues were busy with a different piece of news: an AI had solved a Millennium Prize Problem. That afternoon he told the Anthropic team in Slack that he was leaving and that he feared human extinction without international coordination. Half an hour later it was in the Wall Street Journal. Then he sat down on a park bench in San Francisco and posted.
His X account had fewer than a hundred followers. Within minutes the post was shared by Nathan Calvin of the safety nonprofit Encode AI, by former OpenAI researcher Daniel Kokotajlo, and by a lot of his ex-colleagues. That same evening Evan Hubinger, who runs Anthropic’s alignment team, wrote that they really do earnestly believe AI could kill all humans, and put the odds of that in the next decade above ten percent. Coxon says he coordinated none of it.
His short stint at Anthropic earned him no equity in a company preparing an IPO at roughly a two trillion dollar valuation. He still owns his OpenAI shares.
It didn’t take a famous person
His critics read the exit as a set-up, funded from the effective-altruism orbit, aimed at rules that would suit the biggest labs. Coxon rejects «whistleblower» — he spilled no secrets, he says, he just said in public what gets discussed internally anyway. «Doomer» he accepts: if that means someone who thinks we die unless we change course, then yes. Around 1,400 researchers, Coxon among them, have since signed a statement asking governments to build a brake pedal.
What this story shows isn’t how well Coxon argues. It’s how little it took. No board seat, no book, no audience — one post from a park bench, amplified by a few dozen people who share the same worry. Kokotajlo puts it shortest in the piece: so many people at these companies could have done this. He was the one who did.
Sources: