Claude Fable 5.1 Cache Pricing: 25% Cheaper Runs for Heavy AI Users
Anthropic's Claude Fable 5.1 keeps list prices but cuts cached-token reads 75%. Why that quietly means 25 to 45% cheaper agent bills.
On September 1, 2026 Anthropic launched Claude Fable 5.1 and Claude Mythos 5.1, its new frontier coding and knowledge-work models. On the surface the announcement looks like a capability story: better Terminal-Bench scores, a wider safety window for defensive security work, and knowledge current to June 2026. Read the fine print and it becomes something more useful for anyone running hundreds of dollars of tokens a month. The list price did not move, but the price of reading cached tokens dropped by 75 percent, from $1.00 to $0.25 per million tokens. For agentic, cache-heavy workloads that is a quiet 25 to 45 percent cut to the real bill.
Claude Fable 5.1: Same List Price, Cheaper Where It Counts
Fable 5.1 keeps the flat $10 per million input tokens and $50 per million output tokens that defined the Fable-class tier. Anthropic priced the launch by moving a different lever: the cache read. When a model re-evaluates context it has already seen, such as a large codebase held in a long-running session, it bills the re-read separately from fresh input. Fable 5 charged $1.00 per million cached tokens. Fable 5.1 charges $0.25, a 75 percent reduction in the exact line item that dominates long autonomous agent runs.
Modeled against four weeks of real-world August usage, Anthropic says the change lands around 25 percent cheaper for standard workloads and up to 45 percent cheaper for extensive, automated tasks that lean heavily on cache reads.
The Cache Discount Is the Whole Story for Heavy Users
The launch looks like a price cut that applies to everyone, and for one profile that is roughly true. The user who holds a large repository in one continuous Claude Code session, lets the agent run many turns against the same retained context, and lands most of each turn on the prompt cache gets the biggest discount: that is the 45 percent case. The discount compounds with session length, because the same ~200,000 tokens of context get re-read at $0.25 per million instead of $1.00, dozens of times per session.
The other profile gets almost nothing. If your workload is many short, independent calls with fresh input each time, your bill is dominated by the $10-per-million fresh input line, which did not move. Output tokens at $50 per million did not move either. The savings are concentrated in a single accounting line, so your share of the cut depends entirely on your cache-hit ratio.
That makes Fable 5.1 effectively a targeted price drop for the heaviest users and a cost-neutral refresh for everyone else. Anthropic did not cut list prices across the board, which would have handed savings to every customer regardless of behavior. It cut the one price that only a patient, context-retaining workflow earns. That is a deliberate nudge toward better caching.

What Heavy Claude Users Should Watch
First, verify you actually earn the discount. Anthropic’s cache read credit applies only when your calls hit the prompt cache. Check the cache-hit and cache-read fields in your usage logs. If your cache-hit ratio is low, Fable 5.1 will barely change your bill, and the fix is workflow-level, not pricing-level: structure prompts with a stable system block, keep long sessions alive instead of restarting context, and avoid churning the beginning of the conversation on every turn.
Second, re-baseline your cost per task. If you run autonomous agents that hold a large codebase in context for many turns, your effective per-task rate just fell. Re-run the math on which workloads deserve the frontier model now. A session that felt expensive at the $1.00 cache rate may now be worth moving off a budget model back onto Fable 5.1, or worth leaving where it is with a wider margin.
Third, check how you reach the model. The 75 percent cache-read cut applies to Anthropic’s direct API pricing. If you route through an aggregator or a bundled coding tool, the savings depend on whether the reseller passes the new cache rate through. Confirm the effective cache price on the account you actually bill against before you model the savings. Some paths will pass the cut through cleanly; others will not.

What Else Changed with Fable 5.1
The cheaper cache rate is not the only thing that moved. Fable 5.1 is the same core model with different safety filters, and the enforcement line moved in both directions. Fable 5.1 can now find software vulnerabilities, the defensive half of security work, but Anthropic says it will not write exploits for them, and penetration testing and exploit generation are redirected to Opus models. Researchers reported the cyber filter fires about 60 percent less often per Claude Code session and the biology filter about 85 percent less on ordinary medical questions, with life-science research still routed to Opus 5. Mythos 5.1, the looser-filtered variant for vetted defenders and life scientists, is currently limited to a set of US organizations with no wider-access date announced.
Two subtler changes matter for power users baking Claude into tooling. New API accounts can no longer edit Claude’s past messages while retaining the underlying reasoning data, a change aimed at blocking model distillation. And all outputs now carry a statistical watermark to comply with the EU AI Act; it contains no user data and leaves text unaltered. Neither changes the token economics, but both are worth filing away if you run heavy agent tooling on top of the API.
The Net for Heavy AI Users
Claude Fable 5.1 is a capability update that also functions as a selective price cut. For the user running long, context-heavy agent sessions, the 75 percent cache-read reduction is the real news, and it lands as a genuine 25 to 45 percent reduction in the monthly bill. For the user firing short, stateless calls, list price is unchanged and so is the cost. The two groups should read the launch very differently. Check your cache-hit ratio, confirm the cache rate on the account you actually pay, and re-run cost per task before you decide whether Fable 5.1 changes your routing. The headline may say new model, but the durable takeaway for heavy users is that Anthropic just made keeping one long, warm session dramatically cheaper than firing a hundred cold ones.
Now available
Stop guessing your AI limits
The Mac app and web dashboard watch your Claude, ChatGPT, Gemini and more, and warn you before quotas hit.