memujo
AI5 min read

Anthropic Launches Claude Fable 5.1 and Mythos 5.1

Anthropic shipped Claude Fable 5.1 and Mythos 5.1 as one model split by safeguards, cutting prices up to 45 percent and doubling agentic-science benchmark scores.

By Alice

In this article
  1. 01What Actually Launched
  2. 02The Benchmark Picture
  3. 03Our Read: The Pricing Move Is The Real Story
  4. 04The Data Retention Bet
  5. 05What This Means For Builders

Anthropic on Tuesday unveiled Claude Fable 5.1 and Claude Mythos 5.1, its new frontier models for coding and knowledge work. The full details, including benchmark methodology and availability, are on the Anthropic announcement page. What is notable not mainly for the model weights, but for the business mechanics Anthropic layered on top: a 25% price cut on typical token workloads, a zero data retention option for enterprises, and a safeguards regime now capable of finding software vulnerabilities. Both models ship today, and Fable 5.1 is already available through AWS, per the AWS announcement.

What Actually Launched

Fable 5.1 and Mythos 5.1 are the same underlying model, Anthropic said, differentiated only by the strictness of their safety filters. Fable 5.1 is generally available to everyone. Mythos 5.1 carries fewer safeguards, which unlocks stronger performance on cybersecurity and biology tasks, but it is only reachable through Anthropic's trusted access programs.

The distinction matters because it makes the safeguards themselves a measurable variable in the benchmark numbers. Anthropic disclosed that when its earlier safeguards intervened on certain tasks, both Fable 5.1 and Fable 5 scored a zero on OSWorld 2.0 and Fable 5 scored a zero on AutomationBench. In other words, a chunk of the historical gap between Fable and its competitors was partly a guards-on versus guards-off artifact, not purely capability.

The Benchmark Picture

The headline number is agentic scientific research. On Terminal-Bench-Science 0.1, Anthropic put Fable 5.1 at 52.6%, versus 24.7% for Fable 5, 29.0% for Claude Opus 5, and 22.4% for OpenAI's GPT-5.6 Sol. The standard error was roughly plus or minus 4 points per model, so the Fable 5.1 lead over Opus 5 is well outside noise, but the Opus 5 versus Fable 5 ordering flipped depending on which run you read.

Coding benchmarks tell a similar story:

Benchmark Fable 5.1 Fable 5 Opus 5 GPT-5.6 Sol
Terminal-Bench-Science 0.1 52.6% 24.7% 29.0% 22.4%
Terminal-Bench 4.0 (Mythos 5.1) 55.8% 60.9% (Mythos) 42.0% 52.3%
CursorBench 3.2.0 73.4% 70.5% 70.0% 67.2%
GDPval-AA v2 (knowledge) 1853 1723 1824 1711
Humanity's Last Exam (with tools) 65.0% 63.8% 63.6% --

The Terminal-Bench 4.0 row is where Mythos 5.1 pulls ahead of Fable 5.1 at 60.9% versus 55.8%, the expected signature of a less-guarded model doing work the guards would have blocked. On GDPval-AA, a knowledge work benchmark, Fable 5.1 edges Opus 5 at 1853 to 1824.

Our Read: The Pricing Move Is The Real Story

The benchmark gains are real but incremental year over year. What changes the calculus is the cost side. Anthropic is cutting cache read pricing, which is the price the model pays when it re-reads inputs it has already processed. That sounds like a minor line item until you run an agentic loop for hours.

Here is the engineer framing. Agentic coding work, long research runs, multi-step automation, these are token-heavy precisely because the model keeps re-ingesting the same context. Anthropic reports a 25% reduction on typical workloads and up to 45% on highly agentic ones. For teams that kept Opus-class models only because cheaper models could not hold a long context window, a 45% reduction on cache reads collapses the cost-per-task gap that used to justify the premium.

That is exactly what Cognition, the maker of Devin, reported: it is moving Opus 5 traffic to Fable 5.1 on launch day, citing lower cost per task and finally making a Fable-class model economical for code review workloads it had kept on Opus. Jane Street's quant research head said Fable 5.1 stayed readable across long multi-step tasks, where prior models degraded. That durability, not a single benchmark win, is what makes it useful as production infrastructure rather than a demo piece.

The Data Retention Bet

Anthropic also introduced Enterprise Frontier Safeguards, a data retention option where customer data lives in cloud infrastructure the customer controls, not Anthropic's. That gives enterprises the privacy equivalent of zero data retention while Anthropic still claims state-of-the-art protection against adversarial use. It ships later this fall in phases, with eligible customers able to use Fable 5.1 under zero data retention today.

The safety team also reported a 60% reduction in false positives for cybersecurity tasks. That number is meaningful in the other direction too: because Fable 5.1 can now discover vulnerabilities, the guards have to get better at not blocking legitimate security research. The biology side got a government-partnered access program for Mythos 5.1, opening enrollment for scientists soon.

What This Means For Builders

If you evaluate frontier models the way most teams already do, the practical read is straightforward. Fable 5.1 is the general-purpose pick, and it lands a solid improvement over Fable 5 on nearly every axis, from agentic coding to knowledge work, while being meaningfully cheaper on long runs. Mythos 5.1 is for the narrow slice of work where you have a trusted-access arrangement and need the guards to stay out of the way, mainly cybersecurity and life sciences.

For context, this sits alongside the ongoing coding model race covered in our Claude Opus 5 versus GPT-5.6 Sol developer guide. The gap between them was often marginal. Fable 5.1 shifts the competition slightly toward total cost per completed task rather than peak benchmark score, and for long-horizon agentic workloads that is the metric that actually shows up on an invoice.

The honest caveat is that Anthropic controls the harness. Every benchmark here was run with Anthropic's own setup, and the published numbers come from the vendor. Independent reproduction, especially on the agentic science benchmark, is still the right move before you rewrite your routing rules. But the cache read pricing change is concrete and verifiable on any bill, and that is the part worth watching first.

  • #ai
  • #anthropic
  • #claude
  • #llm
  • #frontier-model

Sources

Share this story