Models & Research

Gemini 4 Argon Leads 12 of 18 Benchmarks, Claude Opus 5.5 Keeps Terminal-Bench 4.0

2 min read AI-generated

Google's new frontier model writes up to a million output tokens, where earlier Gemini models stopped at 64,000. Google's own agents used it to rewrite the libgav1 video decoder in safe Rust: 2.7 times faster, identical output.

Featured image for "Gemini 4 Argon Leads 12 of 18 Benchmarks, Claude Opus 5.5 Keeps Terminal-Bench 4.0"

Google rolled out Gemini 4 Argon on Wednesday, and not to developers but to vetted defenders. Outside its own teams, the model goes first to the 650-plus organisations in the Fairwind Program, CrowdStrike and Palo Alto Networks among them. Koray Kavukcuoglu, Google’s chief AI architect, wrote in the announcement that shipping capability at this level “requires a phased approach.”

The numbers against Claude

On DeepSWE v1.1, which measures software work over long horizons, Argon scores 77.9%. Claude Opus 5.5 sits at 74.2% in Google’s chart, a tenth of a point ahead of GPT-6 Astra. On AutomationBench, which tests business work end to end, the gap widens: 51.3% to 42.5%. Of the 18 benchmarks Google published, Argon leads twelve. Terminal-Bench 4.0 stays with Opus 5.5, FrontierSWE v2 with Astra.

Then a number that isn’t a benchmark. Argon writes up to a million output tokens. Earlier Gemini models capped out at 64,000. Google set the limit deliberately to match the kind of work DeepSWE asks for.

What Google’s own people do with it

Thousands of Google employees already work with the model. Agents are porting C and C++ to Rust, from libraries of tens of thousands of lines up to Fuchsia’s Zircon kernel at over 800,000. On libgav1, Google’s open-source video decoder, they rewrote 32,000 lines of the speed-critical path as safe Rust the compiler can optimise on its own — the decoder now runs 2.7 times faster than the earlier Rust port, with identical video output. A second fleet combs fleetwide profiling data for wasted memory and has freed more than 300 tebibytes across Google’s data centres so far.

The price is the real story

At launch, Argon costs $2 per million input tokens and $10 per million output, with cached input 95% cheaper. Anthropic charges exactly double for Opus 5.5, $4 and $20. Those are also the rates Argon moves to once launch pricing ends. So Google isn’t undercutting permanently. It’s undercutting on a timer.

More telling is who gets the model first. Anthropic has been publishing its cyber findings as research posts, most recently on GLM-5.3. Google is handing the tool straight to defenders — with the cyber guardrails removed, for Fairwind members and internal teams. At Wiz, Argon found a critical flaw in hospital software through the free Scan for Good programme, one earlier frontier models had missed. That’s a more persuasive demo than any benchmark table.

Sources:

GoogleGeminiClaude OpusBenchmarksCybersecurity