The Discount That Isn't (At Max Effort)
Anthropic's headline is a 75% cut to cache-read pricing, from $1.00 to $0.25 per million tokens, which the company says lowers typical workload costs by about 25% and highly agentic workload costs by up to 45% [4]. That framing survives contact with Artificial Analysis's own testing only partway. The benchmarking firm found that at maximum reasoning effort, Fable 5.1 actually costs about 20% more per task than Fable 5 - $3.76 versus a lower baseline, and without the cache-read cut it would have run closer to $5.16 [6]. The reason is that Fable 5.1 burns through roughly 1.7x more output tokens per task at max effort, and output tokens are still billed at the unchanged $50-per-million rate, so the savings on cached reads get eaten by the extra generation. Compared with Opus 5 at max effort ($2.34 per task), Fable 5.1 runs about 61% more expensive for its top score. The-decoder.com's read is blunt: the cache cut genuinely saves close to $1.40 per task on agentic workloads, but that saving is real only if you don't also crank the effort dial up [7]. Independent LLM evaluator Simon Willison captured the same tension anecdotally on X, noting that Max thinking level produced his best-ever SVG pelican output from an Anthropic model, at what he called 'a hefty cost of $3.30' for a single test.




