2 min read AI-generated

Google Gemini 3.6 Flash: Up to 65 Percent Lower Token Costs for Coding Agents

Copy article as Markdown

Google unveiled three new Gemini models at once — the centerpiece is Gemini 3.6 Flash. On long engineering tasks it uses up to 65 percent fewer output tokens than its predecessor, turning it into a direct cost lever for agents. A 3.5 Pro is still missing.

Featured image for "Google Gemini 3.6 Flash: Up to 65 Percent Lower Token Costs for Coding Agents"

On Monday, Google shipped three new models: Gemini 3.6 Flash, 3.5 Flash-Lite, and a specialized 3.5 Flash Cyber. A new 3.5 Pro wasn’t among them — that’s still to come. The workhorse 3.6 Flash sits at the center, and the most interesting part isn’t a benchmark record but the cost per task.

Fewer tokens, same work

According to Google, Gemini 3.6 Flash delivers better coding, better knowledge work, and stronger multimodality — while using 17 percent fewer output tokens than 3.5 Flash. In some benchmarks the effect is dramatic: on Datacurve’s DeepSWE, usage drops by up to 65 percent, at a lower price per output token. For long, agentic engineering runs that’s exactly the lever that matters — that’s where tokens pile up mercilessly.

Quality improves too: on DeepSWE, 3.6 Flash reaches 49 percent versus the predecessor’s 37, and on MLE Bench (ML research) it hits 63.9 versus 49.7 percent. Google promises more precise edits with fewer unwanted code changes and fewer execution loops.

Pricing and availability

Gemini 3.6 Flash costs $1.50 per million input tokens and $7.50 per million output tokens. The lean 3.5 Flash-Lite runs $0.30 and $2.50. Both are available now for developers through the Gemini API in Google AI Studio and Android Studio — plus in the consumer app and Google Search.

My take

The real fight in the coding-agent market right now isn’t decided by the highest benchmark, but by useful intelligence per dollar. That’s exactly where Google is aiming: a model that solves the same task with far fewer tokens is often more attractive for an agent grinding away for hours than one more percentage point on a leaderboard.

For Anthropic, that matters. Fable 5 and Sonnet 5 are strong but expensive — and the recent fight over Fable’s limits showed just how price-sensitive this market is. If Google pushes the cost per agent run down this aggressively, the pressure rises on everyone else’s price tags. The fact that the “Flash” model, not a Pro flagship, makes the headline says a lot about where this is heading.


Sources: