Anthropic & Claude

Simon Willison's look back at the LLM year 2026, and what it says about Claude

3 min read AI-generated

Tokenmaxxing shot up and came straight back down, because agents are expensive. And a 17 GB model on a laptop draws better pelicans than Opus 4.7.

Featured image for "Simon Willison's look back at the LLM year 2026, and what it says about Claude"

Simon Willison gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose on Friday and posted his annotated slides late Sunday night. The run-through starts in November 2025, because that’s where his 2026 begins: Claude Opus 4.5 and GPT-5.1 shipped, and with them Claude Code and Codex went from “often makes mistakes” to “reliable enough for daily use”.

What makes it worth reading here is that Willison goes easy on nobody, Claude included.

What he says about Claude

His yardstick hasn’t changed in years: “Generate an SVG of a pelican riding a bicycle.” In April, the new Qwen3.6-35B-A3B ran on his laptop and drew a better pelican than the brand-new Claude Opus 4.7. He ran a control with a flamingo on a unicycle, and the local model won again.

Where things stand now: Claude Fable 5 gave him the best pelican he’s ever had from a Claude model. It cost $3.30. Opus 5.5 thought for 128,000 tokens and then gave up before producing a response. GPT-6 Astra draws cleanly, and GPT-6 Luna does a passable job for 0.4 cents. The benchmark is stupid, and that’s precisely why it shows price and timing behaviour inside a model family so well.

Tokenmaxxing came and went in six months

In February the headlines said Meta was making AI use part of performance reviews, Microsoft wanted every employee using it, and Uber was boasting that ninety percent of its engineers ran AI workflows. A few months later Meta cracked down on token use, Microsoft said tokenmaxxing wasn’t what it was optimising for, and Uber capped employee AI spending.

Willison’s explanation is unromantic: agents are expensive. Last year it was hard to spend more than $50 on tokens, because there was nothing interesting to do with them. Today you can spend $1,000 in a day doing real work. That, he reckons, is also why Anthropic’s valuation shot towards a trillion. AI found product-market fit in 2026, and it found it through coding agents.

The lesson from the export controls

Fable 5 was unambiguously the best model in the world for eight days, then GPT-5.6 arrived. All told, Fable had 30 days at the top, and for 18 of them it wasn’t available, because the US government had shut it off with an export control directive. Willison’s dry conclusion: market your model as world-ending until a government takes it down, and you lose 60 percent of your time on top.

He also mentions a benchmark he rates above his own pelicans. FelonyBench counts cyberattacks that came out of training runs: OpenAI leads with eleven, Anthropic has nine, Google three, Meta one.

Why the work gets harder even as the agents help

The strongest part of the keynote is the personal one. Willison describes “Deep Blue”, the listlessness that sets in when the machine can do anything, and “AI mania”, the nights you don’t sleep because the agent could be working. His finding: the job feels harder because everything easy goes to the agents. What’s left is the difficult part.

He quotes Greg LeMond on it: it doesn’t get easier, you just get faster. That’s the most honest summary of this year I’ve read so far, and it comes from a cyclist.

Sources

ClaudeEntwicklungAnalyse