Anthropic & Claude

Harvey's gross margin went from 50% to minus 50% — then came Kimi K3

3 min read AI-generated

Uber burned through its entire annual AI budget by April, after telling its engineers to lean on Claude Code as hard as they could.

Featured image for "Harvey's gross margin went from 50% to minus 50% — then came Kimi K3"

Harvey is worth $15.6 billion and built its business on taking someone else’s frontier models and making them useful to lawyers. In March it shipped an update to its agents, usage spiked, and according to Bloomberg the gross margin fell from around 50% at the start of the year to minus 50% by June.

In August, Harvey released a model of its own, running on Moonshot AI’s Kimi K3. The margin went positive again.

Twenty times the tokens, and the math breaks

Harvey reports a twentyfold increase in token usage this year. No margin survives that if every token is billed at the most expensive provider. It gets worse: Anthropic and OpenAI now charge enterprise customers for model usage on top of the subscription, which punishes exactly the teams that spent last year telling everyone to use more. Uber is the famous case — a full year’s AI budget, gone by April.

Harvey is not alone. Abridge is building its own clinical foundation model on Nvidia’s open models. Decagon says it routes 80% of queries through models it owns. Ramp and Rogo are looking at training their own for the first time. Sequoia and General Catalyst are funding the shift.

Ramp co-CEO Karim Atiyeh puts it plainly: a year ago this made no sense at all, now it makes a lot of sense. And on compute costs, he says not thinking about it is simply reckless.

The counterargument: mostly cosplay

Matt Kraning of Menlo Ventures, an Anthropic investor, is unconvinced. Custom models need rare talent and heavy upfront spending, and for some companies it is marketing. Whether you have your own model, he says, is the wrong question — in most cases it tends to be a lot of cosplay.

The objections are concrete, too. US lawmakers are weighing restrictions on open weights. Security agencies and American labs accuse Moonshot and DeepSeek of distilling their models. Rogo had to talk large customers through their hesitation about Chinese models; head of product Strib Walker argues that after enough post-training, an open model stops looking like the thing it started as. Mercor CEO Brendan Foody finds his customers worry less about a regulatory crackdown than about where their data ends up.

One risk outweighs price. A week after SpaceX closed its acquisition of Cursor, OpenAI suspended the coding startup’s model access, citing past terms-of-service violations by Musk’s companies. Build on someone else’s model and you are building on a switch they control.

Anthropic is showing investors this exact slide

The most telling detail sits at the edge of the report. At an investor forum in Anthropic’s own offices, a slide about riding the value curve named Rogo and OpenEvidence. Harvey’s in-house work came up too, with a note that Harvey still needs Opus for its hardest tasks.

That is the most honest summary of the situation available, and it comes from Anthropic. The fight is not over the top of the curve. It is over the middle — the millions of routine calls that used to run at frontier prices. Those are the calls carrying the revenue Anthropic takes into its November IPO.

Sources

BusinessOpen WeightsAnthropic