Enterprise & Security

Mistral Large 4 scores 82% on a cyber test where Claude Opus 5.5 sits near zero

3 min read AI-generated

Trained from scratch on 3,800 Grace Blackwell GPUs in Mistral's own European datacenters. A trillion parameters, 49 billion of them active, more than 160 languages in the training mix.

Featured image for "Mistral Large 4 scores 82% on a cyber test where Claude Opus 5.5 sits near zero"

On October 6, Mistral launched a public preview of Mistral Large 4. Unofficially ML4, very officially: le Chonk. The preview API is live on Mistral Studio, and the weights are due at the end of October.

The numbers: one trillion parameters, 49 billion of them active, natively multimodal. Trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own European datacenters, with the preview served on that same hardware. The training mix spans more than 160 languages, including every official language of the EU.

The one number the announcement is built around

On the Artificial Analysis Cyber Index, ML4 lands in the global top five and leads open-weight models outside China by a wide margin. One test in that index asks a model to reproduce a real vulnerability in open-source software and then patch it. ML4 scores 82%, the highest of any model. On Cybench, 40 exercises from security competitions, it solves 93%.

Then comes the sentence the whole thing runs toward: several leading closed models, Claude Opus 5.5 and GPT-6 Astra among them, score near zero on that same test because they refuse to do the task.

That isn’t a capability comparison, and Mistral doesn’t claim it is. It’s an availability comparison. Their argument: defense often starts with proving a flaw is real, and that is exactly the work the closed models’ filters block.

Coding tells a different story

Here Mistral reports 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4. The combined Coding Agent Index sits at 49.8%.

For context, Terminal-Bench 4.0 is the benchmark Claude Opus 5.5 held against Gemini 4 Argon. 28.3% is a long way off. Mistral doesn’t compare itself to frontier closed models here either — it compares itself to open ones.

The weights don’t ship unprepared

Until the end-of-October release, Mistral is having the model attacked in the field: with cybersecurity firms, vetted partners and state authorities, who get that same model with reduced moderation and expanded cyber capabilities.

Put that next to Anthropic’s new tiers in the Cyber Verification Program and you get the same idea reached two different ways. Anthropic verifies organizations and turns blocks off for them. Mistral verifies before release and then hands out the weights, which makes the question of blocks moot.

Europe is building its own argument

The model will be offered across several regions, including a European deployment Mistral operates end to end, independent of other service providers and under European law. ML4 was trained with the same training, customization and RL environment Mistral sells customers through Mistral Forge.

That’s the positioning, and it’s consistent: sovereignty as a product property rather than a press release. After Reflection’s Beam two days ago, this is the second large open-weight announcement of the week, and both argue over control rather than benchmark peaks. Whether 82% on a cyber index reads as an achievement or a problem depends on who you ask — which is the argument currently underway.

Sources

MistralOpen WeightsBenchmarksSecurityReleases