Models & Research

Claude Now Leads 26 Percent of Anthropic's AI Research

3 min read AI-generated

Roughly 30,000 agents work on research and engineering at Anthropic at any one time. Every single action they take passes through a monitor first.

Featured image for "Claude Now Leads 26 Percent of Anthropic's AI Research"

Anthropic published three measurements yesterday, meant to let the public track how fast development moves inside a frontier lab. The first one is an index of how much of the company’s own AI research Claude already does.

As of August 2026, Claude ‘leads’ 26 percent of that work. Leading means the model takes a high-level prompt and carries the task most of the way on its own, with a human supervising. More than 90 percent of the work sits at least one rung lower, at ‘Claude collaborates’. No measured area runs fully autonomously. Bloomberg puts the 26 percent next to March, when the figure was 1 percent.

How you build a number like that

Nobody can list by hand everything that counts as research inside an AI lab. So Anthropic built the list from the bottom up. For each week in July, a fifth of the staff in the departments that make models was sampled at random, and a Claude research agent read that person’s Slack and internal documentation for the week and wrote down what they worked on. That produced about 15,000 individual tasks.

Claude then sorted them into a tree: 542 nodes, 378 of them leaves with names like ‘eval platform defect diagnosis and fixes’ or ‘RL sandbox egress and network policy’. The tree is frozen, so every later measurement covers the same basket of work. Ratings use Epoch AI’s automation scale, AL0 through AL5.

That Claude writes 80 percent of the code at Anthropic was already in last week’s post about their CI. What’s new is the attempt to turn it into something you can repeat month after month.

30,000 agents, and one command in 47,000 gets stopped

The second measurement is about oversight. Anthropic’s most-used internal platform runs roughly 30,000 agents at any given moment, and everything they do passes two kinds of monitor.

The online monitor checks before the fact and blocks whatever can’t be undone — an agent copying model weights out of the building, for instance. Anthropic analysed over a billion decisions in August; 0.002 percent were blocked, about one in 47,000. The offline monitor reads along afterwards and flags one or two transcripts per thousand. Around 100,000 transcripts a week land in review that way, and roughly 50 of those reach a human.

Six percent of the compute goes to safety

The third number is the awkward one. In the week examined, 13 to 20 July, about 6 percent of the compute that went into AI research went to safety work. Of the compute spent on AI-driven research, it was 12 percent.

Anthropic deliberately counts low: anything that advanced capabilities as much as safety wasn’t counted at all. And the company says outright that compute is a crooked yardstick here. Safety research is people designing experiments — that costs time, not GPUs.

The second reading is the one that matters

A single number from a single lab says very little. Twenty-six percent is neither high nor low without something to compare it to. The index gets its value when it’s taken again in three months, and it gets genuinely interesting when a second lab runs the same method.

That’s exactly where it stalls. Anthropic names the problem itself: there’s no shared methodology, and the judge is a model from the same house as the thing being judged. Hence the proposal that independent third parties verify the figures — the same evaluators Anthropic wants to embed anyway. Until then this isn’t a measurement system. It’s an invitation to the others to join in.


Sources:

AnthropicClaudeResearchTransparencyAgents