Claude Fable 5.1 statistics (September 2026): benchmarks, pricing, cost and safety
Claude Fable 5.1 launched Sept 1, 2026 at $10/$50 per million tokens, ties GPT-6 Astra at 53 on Artificial Analysis's Intelligence Index v4.3, but costs more than double Astra's price per task ($7.63 vs $3.26). Anthropic claims ~25-45% savings over Fable 5 thanks to cheaper cache reads, while independent testing found it 20% pricier at max effort due to heavier token usage both are true depending on effort level, and xhigh matches max's score for 28% less cost. It leads most benchmarks and cuts safeguard false-positives significantly, but Anthropic itself recommends starting with Opus 5 (half the price, close in score) and reaching for Fable 5.1 only for long, cache-heavy agent work or tasks Opus 5 can't crack.
Claude Fable 5.1, released September 1, 2026, costs $10 per million input tokens and $50 per million output tokens, with a 1M-token context window. It tops all seven benchmarks in Anthropic’s launch table and ties GPT-6 Astra at 53 on Artificial Analysis’s Intelligence Index, but at more than twice Astra’s cost per task.
Claude Fable 5.1 statistics at a glance
| Stat | Number | Source |
|---|---|---|
| Release date | September 1, 2026 | Claude Platform docs |
| Context window / max output | 1M tokens / 128K tokens | Claude Platform docs |
| API price (input / output) | $10 / $50 per 1M tokens | Anthropic |
| Cache read price | $0.25 per 1M tokens (75% lower than Fable 5) | Anthropic |
| Claimed savings vs Fable 5 | ~25% typical, up to ~45% highly agentic | Anthropic |
| Knowledge cutoff | June 2026 | Claude Platform docs |
| Artificial Analysis Intelligence Index v4.3 | 53 (tied #1 with GPT-6 Astra) | Artificial Analysis, Sept 7, 2026 |
| Cost per Intelligence Index task (max effort) | $7.63 vs $3.26 for GPT-6 Astra | Artificial Analysis |
| Terminal-Bench 4.0 | 55.8% (Anthropic) / 52.0% (Artificial Analysis) | Anthropic; Artificial Analysis |
| Cyber safeguard interventions | ~60% fewer per Claude Code session vs Fable 5 | Anthropic |
| Share of output tokens served by fallback models | ~4% on the Intelligence Index | Artificial Analysis |
| Output speed | 47 to 67 tokens per second, by effort level (live figure) | Artificial Analysis |
What is Claude Fable 5.1?
Claude Fable 5.1 is Anthropic’s newest generally available Mythos-class model, built for long-running coding, research and knowledge work. It shares one underlying model with Claude Mythos 5.1; the only difference is that Fable 5.1 carries stricter safeguards for cybersecurity and biology, while Mythos 5.1 is limited to vetted organizations.
Anthropic doesn’t sell it as the default, though. Its docs tell developers to start most workloads on Claude Opus 5 and move to Fable 5.1 for demanding reasoning, long-horizon agents, or cases where Opus 5 at higher effort still falls short.
| Model | Who can use it | API price (in / out, per 1M) | Context | Role |
|---|---|---|---|---|
| Claude Fable 5.1 | All paid plans and the API | $10 / $50 | 1M | Hardest reasoning and long agent runs |
| Claude Mythos 5.1 | Invite only (trusted access programs) | $10 / $50 | 1M | Same model, looser cyber and bio safeguards |
| Claude Fable 5 | Still available; replaced by 5.1 | $10 / $50 (cache read $1) | not listed | Previous Mythos-class release (June 2026) |
| Claude Opus 5 | Claude apps and the API | $5 / $25 | 1M | Anthropic’s recommended default |
| Claude Sonnet 5 | Claude apps and the API | $2 / $10 | 1M | Faster, cheaper tier |
How does Fable 5.1 score on Anthropic’s benchmarks?

Fable 5.1 has the highest score on every benchmark in Anthropic’s launch table. The jump that stands out is Terminal-Bench-Science: 52.6%, up from 24.7% for Fable 5, so more than double. Business automation (AutomationBench) moved a lot too.
| Benchmark | Fable 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| Terminal-Bench-Science 0.1 (agentic science) | 52.6% | 24.7% | 29.0% | 22.4% |
| Terminal-Bench 4.0 (agentic coding) | 55.8% | 42.0% | 52.3% | 37.3% |
| GDPval-AA v2 (knowledge work, Elo) | 1,853 | 1,723 | 1,824 | 1,711 |
| OSWorld 2.0 strict (computer use) | 41.7% | 36.1% | 39.6% | not reported |
| Humanity’s Last Exam, with tools | 65.0% | 63.8% | 63.6% | not reported |
| AutomationBench (business workflows) | 31.4% | 17.1% | 26.9% | 19.6% |
| CursorBench 3.2.0 (agentic coding) | 73.4% | 70.5% | 70.0% | 67.2% |
Points gained by Fable 5.1 over Fable 5 on Anthropic-reported benchmarks (safeguards on). Source: Anthropic, Sept 1, 2026. Verified Sept 10, 2026. Four caveats change how you should read that table:
- Anthropic ran Fable 5.1 with production safeguards on. When they intervened on OSWorld 2.0, the task scored zero, which likely understates the result.
- Mythos 5.1, the same model without those cyber safeguards, hit 60.9% on Terminal-Bench 4.0, about five points above Fable 5.1.
- Anthropic lists a standard error of plus or minus 3.5 to 4.5 points on Terminal-Bench-Science, so small gaps are noise.
- GPT-5.6 Sol was OpenAI’s comparison model at launch. GPT-6 Astra launched days later and is not in Anthropic’s table.
What do independent benchmarks show about Fable 5.1?
Independent tests put Fable 5.1 at the top, but it now shares first place. On Artificial Analysis’s Intelligence Index v4.3, published September 7, 2026, Fable 5.1 and GPT-6 Astra (both at max effort) score 53, ahead of Claude Opus 5 (51), Fable 5 (50) and GPT-5.6 Sol (47).
Cost is where the two split. Artificial Analysis measured $7.63 per index task for Fable 5.1 against $3.26 for Astra, which is 57% cheaper for the same score. Astra also used about 27K output tokens per task at max effort, roughly a third of Fable 5.1’s 78K.
Where Fable 5.1 leads and trails in independent tests
- On the Coding Agent Index, Fable 5.1 in Claude Code and GPT-6 Astra in Codex tie at 62, ahead of Opus 5 at 60.
- On Terminal-Bench v4.0, Artificial Analysis measured 52.0% for Fable 5.1, below Anthropic’s 55.8% and below Astra’s 59.1%.
- On AutomationBench-AA, Fable 5.1 fully completed 32.1% of 657 business workflows without a guardrail violation, versus 41.6% for Astra and 28.3% for Opus 5.
- At launch, Fable 5.1 set the highest scores Artificial Analysis had measured on GDPval-AA v2 (1,853 Elo) and AA-Briefcase (1,694 Elo), though Opus 5 was within the margin on both and beat Fable 5.1 on presentation quality (1,572 vs 1,495).
- Fable 5.1 had the highest accuracy Artificial Analysis had measured on AA-Omniscience (67.2%), but it also guessed more. When it did not know the right answer, it still attempted one 72.6% of the time, up from 63.6% for Fable 5, so its overall Omniscience score matched Fable 5.
Why do some sites say Fable 5.1 scored 66 and others say 53?
Both numbers are correct; they come from different versions of the same index. On launch day Artificial Analysis scored Fable 5.1 at 66, ahead of Opus 5 (63), Fable 5 (62) and GPT-5.6 Sol (61). Artificial Analysis then released v4.2 on September 4 and v4.3 on September 7, adding harder tasks and more private test sets, which rescaled every model’s score. Under v4.3, Fable 5.1 scores 53. Never compare a 66 from one article with a 53 from another.
Cost per task rescaled too: $3.76 at max effort under the launch index, $7.63 under v4.3.
How much does Claude Fable 5.1 cost?
On the API, Fable 5.1 costs $10 per million input tokens and $50 per million output tokens, the same list price as Fable 5 and double Opus 5. The change is in caching: cache reads dropped from $1 to $0.25 per million tokens.
| Price item | Fable 5.1 |
|---|---|
| Input | $10 / 1M tokens |
| Output | $50 / 1M tokens |
| 5-minute cache write | $12.50 / 1M tokens |
| 1-hour cache write | $20 / 1M tokens |
| Cache read | $0.25 / 1M tokens |
| Batch API | 50% off input and output |
| US-only inference | 1.1x input and output price |
On Claude subscriptions, access depends on the plan. This caused confusion at launch, because every paid plan sees Fable 5.1 in the model picker but not every plan includes it.
| Plan | How Fable 5.1 is billed |
|---|---|
| Free | Not available |
| Pro, standard Team seats, standard seat-based Enterprise seats | Not included in plan limits; runs on pay-as-you-go usage credits |
| Max, premium Team seats, premium seat-based Enterprise seats | Included; up to 50% of weekly usage limit on Fable models, then credits or switch models |
| Usage-based Enterprise and the API | Standard API rates |
Claude Code needs version 2.1.255 or later to use Fable 5.1. On claude.ai and in Cowork the model defaults to Medium effort; in Claude Code it defaults to High.
Is Fable 5.1 cheaper than Fable 5?
It depends on how much of your bill is cache reads and which effort level you run. Anthropic says Fable 5.1 costs about 25% less than Fable 5 for typical workloads and up to about 45% less for highly agentic ones. Artificial Analysis found the opposite on its benchmark: at max effort Fable 5.1 cost 20% more per task than Fable 5 ($3.76 vs $3.14).
Both claims hold because they measure different things. Anthropic’s figures come from four weeks of real usage in August 2026 at default effort, billed per token. Artificial Analysis ran max effort, where Fable 5.1 produced about 1.7 times the output tokens of Fable 5. The cache discount saved about $1.40 per task; the extra output tokens cost more than that.
Anthropic’s math is easy to sanity-check. A 75% cut in cache-read price only lowers the total bill by 25% if cache reads were about a third of that bill (0.25 divided by 0.75). The 45% figure implies cache reads were about 60% of the bill. Both assume the token mix stayed the same.
Worked example: one agent session priced three ways
We priced a single agent session on both models. We assumed 2M cached input tokens read, 200K fresh input tokens and 100K output tokens, and ignored cache writes to keep it simple.
| Scenario | Cache reads | Fresh input | Output | Total | vs Fable 5 |
|---|---|---|---|---|---|
| Fable 5 | $2.00 | $2.00 | $5.00 | $9.00 | baseline |
| Fable 5.1, same tokens | $0.50 | $2.00 | $5.00 | $7.50 | 17% cheaper |
| Fable 5.1, 1.7x output tokens | $0.50 | $2.00 | $8.50 | $11.00 | 22% more expensive |
So the cache discount only wins if you keep output tokens in check, and effort level is the main lever for that. We think the price cut is real, but it mostly rewards teams that already cache well.
Which Fable 5.1 effort level is the best value?

On Artificial Analysis’s data, xhigh gives the same score as max for less money, and the cheapest points come at the low end. At launch, Artificial Analysis measured an 11x spread in output tokens across the five effort levels, from 13.1M at low to 143.7M at max for a full index run.
We divided Artificial Analysis’s live v4.3 cost per task by each effort level’s score on September 10, 2026. The marginal column shows what each extra index point costs as you move up a level.
| Effort level | Intelligence Index v4.3 | Cost per task | Cost per index point | Cost of each extra point vs level below |
|---|---|---|---|---|
| Low | 47 | $2.37 | $0.050 | starting point |
| Medium | 49 | $2.98 | $0.061 | $0.31 |
| High | 51 | $3.91 | $0.077 | $0.47 |
| Xhigh | 53 | $5.98 | $0.113 | $1.04 |
| Max | 53 | $7.63 | $0.144 | +$1.65 for zero points |
Cost per Intelligence Index v4.3 task by effort level. Source: Artificial Analysis live data, Sept 10, 2026; calculation by Market Analyticx. Moving from xhigh to max adds 28% to cost per task and zero Intelligence Index points. The two points gained from high to xhigh cost more than three times as much each as the two gained from low to medium.
An index is not your workload, so treat this as a starting hypothesis. Max effort may still pay off on tasks where the extra thinking changes the outcome, like the multi-day debugging runs partners described at launch. If we were setting a team default today, we would start at high and move up to xhigh only for tasks that fail at high.
What changed in Fable 5.1’s safeguards and safety?
Fable 5.1 blocks fewer harmless requests than Fable 5. Anthropic says Claude Code users should see about 60% fewer cyber-safeguard interventions per session, and its biology safeguards now fire 85% less often on benign elementary biology and medical questions than they did at Fable 5’s launch. On risk, Anthropic says the model still sits below the next tier in its Responsible Scaling Policy for chemical and biological weapons.
The other safety and safeguard figures from launch:
- Fable 5.1 can be used to find software vulnerabilities but not to build exploits. Penetration testing, exploit generation and binary vulnerability scanning still get redirected to Opus models.
- Fallback is small but not zero: Artificial Analysis found fallback models served about 4% of output tokens across its index run.
- Anthropic says it has found no evidence of a critical-severity jailbreak of the cyber safeguards after its own testing, two external testing firms and automated testing by Gray Swan.
- Anthropic calls Mythos 5.1 (the same model) its most resistant model yet on an external prompt-injection benchmark.
- Anthropic reports that the model reward-hacks less often than Mythos 5, but says it can still sometimes bypass approvals and auto-mode classifiers.
- Claude models released after August 2, 2026 add an invisible text watermark to meet the EU AI Act Code of Practice. A detection API is in private preview.
- Eligible enterprises can use Fable 5.1 with zero data retention until Enterprise Frontier Safeguards roll out, starting this fall.
What are early users reporting about Fable 5.1?
Launch partners report big efficiency gains, while subscription users on public forums mostly complain about how fast Fable 5.1 burns through usage limits. Treat both with care: partner figures come from Anthropic’s launch page, and forum reports are anecdotes.
| Company | Reported result | Context |
|---|---|---|
| Browserbase | 82% of tasks completed vs 74% for Opus 5 and 57% for Fable 5 | Internal browser-agent benchmark |
| Every | About 2x as fast as Opus 5, half the tokens | CEO Dan Shipper’s tests |
| Crosby | RedlineBench score up from 47.9 to 57.0 | Contract redlining benchmark |
| Samaya | 55.9% vs 49.2% for Fable 5 | FrontierFinance benchmark |
| Rogo | Matched Fable 5 accuracy with 20% fewer tokens | Internal finance benchmark |
| Glean | Answers preferred about 2-to-1 over Fable 5 | Internal evaluation sets |
All results customer-reported via Anthropic’s launch page; not independently verified.
What the developer community says
Outside the launch page, the loudest complaint is usage burn. Posts on X from Max subscribers describe draining a 5-hour session in under 30 minutes, including one report of a 20x Max quota gone in 52 minutes from a single prompt. The most shared fix is dropping effort to low. A pattern we saw in several Hacker News comments: Fable for design and review, cheaper Opus 5 subagents for implementation.
Writing style came up almost as often. On the Hacker News launch thread (1,417 points and 1,394 comments when we checked), an Anthropic employee said Fable 5.1 writes more naturally and follows style instructions better. Much of the thread was fatigue with dense, jargon-heavy Claude prose, mostly aimed at Opus 5, and a few developers said they had moved to GPT-5.6 Sol for shorter answers.
For a sense of real per-call costs, Simon Willison logged $1.37 for an animated SVG at the default High effort (6,121 input and 26,201 output tokens) and about $3.30 for a single max-effort drawing.
Should you use Fable 5.1, Opus 5 or GPT-6 Astra?
Our default would be Opus 5. Fable 5.1 earns its price on long, cache-heavy agent runs and on problems Opus 5 can’t crack, and GPT-6 Astra is worth a test whenever cost per task is the main concern. By situation:
- If your agents run long and reread big contexts, Fable 5.1’s cache discount works in your favor. Start at high and step up to xhigh; leave max alone until your own tests justify it.
- If you pay per task and matching the index score is enough, look at GPT-6 Astra. It matched Fable 5.1’s index score at about 40% of the cost. Test both on your own tasks before switching.
- On a Pro plan, Fable 5.1 bills against usage credits from the first message, while Opus 5 stays inside your plan.
- For typical work, Anthropic’s own docs say to start on Opus 5, at half the token price.
- If the output is client-facing slides or reports, check both. Opus 5 beat Fable 5.1 on Artificial Analysis’s presentation quality score.
- If you’re upgrading an existing Fable 5 integration, plan for three breaking changes. Forced tool use now returns an error, earlier models can’t read Fable 5.1’s thinking blocks, and editing earlier turns invalidates them.
- Exploit development and advanced biology are off the table: Fable 5.1 will redirect or block those requests. Mythos 5.1 is only available through Anthropic’s access programs.
What are the limits of Fable 5.1 benchmark data?
Benchmarks and launch figures tell you where Fable 5.1 stands, not what it will cost or deliver on your workload. Keep these limits in mind:
- Index points measure a fixed task mix. Your prompts, tools and effort settings will move cost and quality more than a two-point index gap.
- Anthropic’s Terminal-Bench 4.0 score (55.8%) sits above Artificial Analysis’s independent run (52.0%). Different harnesses produce different results.
- Artificial Analysis changed its index twice in the week after launch. Date every number you cite.
- Launch partner results come from real companies running real tests, but Anthropic picked them for a launch page.
- Blocked tasks scored zero in some of Anthropic’s own runs, and fallback models handled about 4% of Artificial Analysis’s output tokens.
- Artificial Analysis ran pre-release evaluations for both Anthropic and OpenAI. That is common practice, but keep it in mind when reading its numbers.
FAQ
Bottom line
Fable 5.1 is one of the two strongest models you can buy in September 2026, and Anthropic’s usage data says it is cheaper than Fable 5 on long, cache-heavy agent work. It is not the cheapest route to frontier scores: GPT-6 Astra matches it for less, and Opus 5 comes within two index points at half the token price. If you run it, start at high and step up to xhigh only when a task fails there. Max has to prove itself on your own work.
For more on how AI models change search and marketing work, see our [AI visibility hub].
Key takeaways
- Claude Fable 5.1 launched September 1, 2026 at $10 / $50 per million tokens, with cache reads cut 75% to $0.25.
- It ties GPT-6 Astra at 53 on Artificial Analysis’s Intelligence Index v4.3, but costs $7.63 per task against Astra’s $3.26.
- Anthropic’s “25% cheaper” and Artificial Analysis’s “20% more expensive” are both true; effort level and output tokens decide which applies to you.
- On Artificial Analysis data, moving from xhigh to max adds 28% to cost per task for zero index points.
- Pro plans pay for Fable 5.1 with usage credits; Max plans can spend up to half their weekly limit on it.