Anthropic has published its second company-wide risk report. Two things stand out: the company raises its estimate of catastrophic misalignment risk – and, for the first time, mentions an internal model nobody outside the building has ever seen.
The risk goes up – a little
In the first report back in February, the risk that a model might undermine oversight in a high-stakes situation was rated “very low”. Now it says “low”. One notch higher, no more.
The reason is a bit surprising. It’s not that the models suddenly turned more malicious. Anthropic points to uncertainty around its recent cybersecurity evaluations. The company’s own testing methods are hitting their limits, and when that happens you’d rather rate cautiously on the high side. The core of the report, a chain of eight claims, is about “covert capabilities”: a model’s ability to reliably subvert oversight. Anthropic still considers such capabilities unlikely for its current models.
The real headline: Model 2
Between the lines, a name shows up that officially doesn’t exist yet: Model 2. An internal frontier model, according to the report somewhat more capable than Mythos 5 – and Mythos 5 is already the strongest thing Anthropic has ever built.
Model 2 hasn’t been through the full pre-deployment safety assessment, so Anthropic isn’t fully sure of its capabilities either. And there are no current plans to release it externally. At least, the report says, its internal approval surfaced no new or worse form of misalignment than what was seen with Mythos 5.
My take
A safety report that raises its own risk estimate is a good sign to start with. It would have been easy to leave the number at “very low” and nobody would have complained. That Anthropic instead says “our testing methods aren’t quite enough anymore, so we’re rating higher to be safe” – that’s the kind of honesty I want to see in this field.
Model 2 gives me more pause. There’s a model that’s stronger than anything released, and the company deliberately puts it in the vault. You can read that as restraint, as responsibility. But you can also ask what it means that the gap between “built” and “released” keeps widening. The most interesting models, it seems, are the ones we’ll never lay eyes on.
Sources: