3 min read AI-generated

Fable 5.1 Hands-On: A $3.30 Pelican and the Question of What Effort Levels Actually Do

Copy article as Markdown

Simon Willison ran all five effort levels of Fable 5.1, Nick Saraev picked apart the benchmarks. The upshot: low and medium apparently don't think at all, while max produces the best pelican Anthropic has ever drawn, at a cost of 14 minutes.

Featured image for "Fable 5.1 Hands-On: A $3.30 Pelican and the Question of What Effort Levels Actually Do"

On launch day, independent voices matter more than the press release. Two of them were quick on September 1: Simon Willison with his pelican test, and Nick Saraev with a video that went up minutes after the announcement.

Willison admits he no longer fully trusts the pelican benchmark. The link between a well-drawn bird on a bicycle and quality on real tasks has weakened since 2025. What the test still shows well: differences within a model family, and above all differences between effort levels. Fable 5.1 has five of them: low, medium, high, xhigh, and max. You can’t turn reasoning off.

The numbers are revealing. At low, the pelican arrived after 23.8 seconds for about 10 cents, 1,998 output tokens, and not a single reasoning token in the transcript. Medium: 1,977 tokens, so 21 fewer than low, again with no reasoning. For this prompt, the model simply didn’t think at either level. Only at high did a short plan appear, 2,612 tokens, 13 cents, visually hardly better. Then the jump: xhigh took 36,767 tokens, nearly eight minutes, and $1.83. Max went further still: 65,927 tokens, 13 minutes 54 seconds, $3.30. In return, Willison says, he got the best pelican he’s ever seen from an Anthropic model. Legs correctly on either side of the frame, feet on the pedals, a little blue hat, a basket with a fish in it.

What I like about the reasoning excerpts: the model argues with itself about whether a helmet would collide with the beak, rejects a bicycle bell as unnecessary, and corrects the rake of the front fork. That’s not drawing, that’s engineering. When someone on Hacker News asked for an animated version, Willison just piped the max pelican back in with “animate this” at high effort. 26,201 output tokens, $1.37, done.

Nick Saraev looks at the benchmarks from the perspective of someone who automates business processes. His highlight is AutomationBench: Fable 5 reliably automated real workflows 17.1 percent of the time, Fable 5.1 does it 31.4 percent of the time. Nearly double. On computer use, where Anthropic used to lag, the model hits 77.9 percent on OSWorld 2.0. And he makes a point that stuck with me: Opus 5 scored well on many benchmarks, yet anyone using both day to day gets better results from Fable. Benchmarks don’t measure everything. That’s why he expects the next 24 hours to be about what he calls “taste benchmarks”: people throwing odd tasks at the model and seeing what happens.

His second point is about cost. Anthropic plots score against mean cost per task on a log scale. Saraev reads roughly two and a half times the performance per dollar compared to Fable 5. And he argues that cost per task is a better metric than cost per token, because newer models reach the goal with fewer tokens.

What I take away from this: on Fable 5.1, the effort level isn’t fine-tuning anymore, it’s a switch between two completely different models. Below high it barely thinks, from xhigh upward time and cost explode. For everyday work that means high as the default, max only for things where 14 minutes and three dollars don’t matter. And yes, I looked at the pelican. It really is good.

Sources: Simon Willison: Claude Fable 5.1 made me a really nice animated pelican · Nick Saraev: Fable 5.1 Just Dropped. It’s Not Even Close.