3 min read AI-generated

Kimi K3: 2.8 Trillion Parameters — and Suddenly a Chinese Model Costs as Much as Sonnet

Copy article as Markdown

Moonshot AI announced K3: the first open 3T-class model, with open weights promised by July 27. The real shock is on the price tag.

Featured image for "Kimi K3: 2.8 Trillion Parameters — and Suddenly a Chinese Model Costs as Much as Sonnet"

Chinese AI lab Moonshot AI announced Kimi K3 yesterday — by their own description their most capable model to date, with 2.8 trillion parameters. It’s available now via their website and API, with open weights promised “by July 27, 2026”.

Moonshot calls K3 the first open 3T-class model (rounding 2.8 up generously), taking the crown from DeepSeek’s 1.6T v4 Pro.

The numbers

Their self-reported benchmarks put K3 mostly ahead of Claude Opus 4.8 max and GPT-5.5 high — but behind Claude Fable 5 and GPT-5.6 Sol. Three things stand out from the Artificial Analysis report:

  • On their private long-horizon knowledge work evaluation, K3 hits an Elo of 1547 — up 732 points from Kimi K2.6, and behind only Claude Fable 5.
  • Cost per task: $0.94. That’s roughly GPT-5.6 Sol territory ($1.04) and about half of Opus 4.8 ($1.80) — but higher than its open-weights peers.
  • Token usage dropped noticeably: 21% fewer output tokens than K2.6.

On top of that, K3 now leads Arena.ai’s Frontend Code arena — ahead of Claude Fable 5.

The real story is the pricing

$3 per million input tokens, $15 per million output tokens. That’s exactly the Claude Sonnet level — and makes K3 the most expensive model ever released by a Chinese AI lab.

For comparison: Kimi K2.6 was $0.95 / $4. That’s more than a 3x jump on input and nearly 4x on output. Then again, the model is more than twice the size of that 1T predecessor.

So much for “Chinese models are the cheap alternative”, at least at the frontier. If you have to move 2.8 trillion parameters, you can’t give them away.

The pelican test

Simon Willison ran K3 through his now 21-month-old benchmark: “Generate an SVG of a pelican riding a bicycle”. The result cost 25 cents — 95 input tokens and 16,658 output tokens, of which 13,241 were pure reasoning tokens.

K3 currently has exactly one reasoning effort level: max. And it shows. 13,241 thinking tokens for 3,417 tokens of answer.

Willison takes the opportunity to be honest about his own benchmark: the correlation between pelican quality and model quality is “mostly severed” now. GLM-5.2 draws better pelicans than GPT-5.6 and Claude Fable 5 — and GLM is not a Fable-class model. The biggest limitation: the test says nothing about agentic tool calling. Which is the thing that matters most for today’s models.

One lovely detail along the way: the prompt counts 95 input tokens on K3, while OpenAI counts 10 and Anthropic counts 25 to 30. Prompting “hi” comes to 86 tokens. Willison suspects a hidden system prompt of around 85 tokens. K3 refused to leak it.

My take

Two things I’m taking away. First: the gap at the top keeps shrinking, but Fable 5 still leads on long-horizon knowledge work — and that’s the discipline that counts for real work. Second, and more interesting to me: the price. When an open model out of China charges Sonnet rates, that’s not bravado, it’s math. Competition is shifting away from price and toward the question of what the thing actually gets done.

Those open weights on July 27 will be interesting. Though almost nobody is running this one at home.


Sources: