2 min read AI-generated

GPT-5.6 Sol Ultrafast: OpenAI Cranks the Speed to 750 Tokens per Second

Copy article as Markdown

OpenAI gives an early look at a new service tier that runs GPT-5.6 Sol up to 14 times faster. Not a new model, just raw speed – powered by Cerebras chips.

Featured image for "GPT-5.6 Sol Ultrafast: OpenAI Cranks the Speed to 750 Tokens per Second"

OpenAI gave an early look yesterday at “Ultrafast” – a new service tier that runs GPT-5.6 Sol up to 14 times faster than the standard variant. Important to understand: this is not a new model. It’s the same model with the same intelligence, just delivered radically faster.

What Ultrafast means

Up to 750 output tokens per second. That’s the number that matters. It’s made possible by a collaboration with Cerebras – their chips are built for exactly this kind of throughput. So OpenAI isn’t selling a new capability here, it’s selling a new experience: answers that are essentially instant.

The tier launches first through the OpenAI API and is currently only available in a limited preview to select customers. Access is meant to grow as capacity allows. The whole thing is aimed at applications where every second counts: financial research, incident response, customer support, voice apps, commerce.

Why speed is suddenly a feature

For a long time, models were almost entirely about intelligence – who solves the hardest tasks. Now the competition is shifting. Google shipped Gemini 3.7 Flash the same day, and OpenAI counters with speed. Both send the same message: AI is now fast enough to actually feel good when you’re running agents.

And with agents, speed isn’t a luxury. When an agent takes ten steps in a row, every wait adds up. A model that answers in fractions of a second changes what feels fluid and what feels sluggish.

My take

What I find interesting is less the number than the direction. The model race is picking up a second axis right now: not just “how smart,” but “how fast.” And both axes together decide whether an agent feels usable in daily work or not.

One caveat remains: Ultrafast is a limited preview. “Up to 14 times faster” sounds great, but as long as only select customers get in, it’s more a promise than a tool. Still – the footnote from Anthropic’s own auto mode blog fits here: GPT-5.6 Sol scored noticeably worse than Claude in auto mode on prompt-injection tests. Speed is only half the battle. In the end, both count: being fast and not doing anything dumb.


Sources: