Every model launch sounds great in the official announcement. It only gets interesting once people with no marketing stake get their hands on it. For Claude Opus 5, that’s exactly what just happened — and the verdict lands remarkably positive.
Simon Willison’s first impressions
Simon Willison, one of the sharpest observers of the LLM scene, tested Opus 5 directly. His conclusion: the hype is earned. The model currently tops the Artificial Analysis leaderboard — ahead of Claude Fable 5 — and it does so at the same price as Opus 4.8. A “fast mode” at double the rate sticks around.
One detail Willison highlights is especially interesting: Opus 5 is the hardest model to manipulate he’s seen — extremely robust across prompt-injection evals and red-teaming. At a moment when agents increasingly take real actions, that’s not a nice-to-have. It’s core safety.
What the benchmarks show
The independent numbers back up the picture. Opus 5 delivers 43.3 percent on Frontier-Bench (agentic coding), 30.2 percent on ARC-AGI-3 — roughly three times the next-best model — and a GDPval Elo of 1,861 for knowledge work. The headline: Opus 5 gets very close to Fable 5’s frontier intelligence, at half the price.
A recurring bit of praise is about behavior, not just numbers: Opus 5 checks its own work. One tester reported the model opened its own built pages in a browser — at desktop and phone widths — spotted a button hidden below the mobile fold, and fixed it before handing the work back.
My take
That’s exactly the difference that matters day to day. Not “how smart is the model on a test,” but “does it clean up after itself.” A model that catches its own mistakes saves me more time than one that scores two benchmark points higher. Add in that Opus 5 is unusually hard to manipulate, and it becomes a much more attractive fit for real agent workflows. The launch was loud — the independent echo is quieter, but more convincing.
Sources: