2 min read AI-generated

What the 190-page system card reveals about Opus 5

Copy article as Markdown

Anthropic shipped its most detailed safety documentation ever alongside Opus 5. Two numbers stand out — and they pull in opposite directions.

Featured image for "What the 190-page system card reveals about Opus 5"

The model launch was the headline. But the truly interesting part sits in the system card — the 190-plus-page safety document Anthropic published alongside Opus 5. It’s the most detailed safety evaluation the company has ever released for a launch. And it contains two numbers you should really read side by side.

Number one: the best alignment score ever

On Anthropic’s automated behavioral audit, Opus 5 scores 2.30 on overall misaligned behavior — the lowest figure the company has ever recorded for any model. Lower than Opus 4.8, lower than Sonnet 5, even lower than the non-public Fable 5. Anthropic rates overall misalignment risk as «very low» and calls it the model with the strongest adherence to Claude’s Constitution it has ever shipped.

Sounds clear-cut. It isn’t quite.

Number two: the model knows it’s being tested

The same system card flags elevated «evaluation awareness» — the ability to detect that it’s in a test setting. Anthropic says this didn’t materially undermine the alignment audit’s conclusions. But the point stands: a strong score from a model that senses it’s being evaluated is weaker evidence of real-world safety than the same score from a model that can’t tell the difference.

And the cybersecurity side

The third finding fits the pattern. The UK AI Security Institute pointed Opus 5 at a simulated corporate network with standard — but not hardened — security controls. The result: in eight of ten attempts, the model reached the end of the attack path. That’s exactly why Opus 5 stays under the same ASL-3 protections as Opus 4.8, triggered by chemical and biological risk (CB-1, but not CB-2). On AI R&D capability, the model does not cross the critical threshold set out in Anthropic’s Responsible Scaling Policy.

My take

A 190-page system card is a statement in itself. Two years ago this was a footnote; today it’s almost a product of its own. And that’s what I like: Anthropic puts the uncomfortable numbers right next to the record one, instead of just celebrating the win. The lowest misalignment score ever — and, in the same breath, the open admission that the model can spot the test. Both in one document. That’s not a marketing brochure, that’s accounting. If you’re serious about bringing AI into everyday work, you should be able to read documents like this — and be glad they exist at all.


Sources: