Fable and Mythos 5.1: same model, two safeguard tiers, and a benchmark jump we can partly check ourselves
Anthropic's September 2026 release splits one model into a broadly available tier and a restricted one for vetted cyber and life-science professionals. The benchmark numbers are large, and we have a small, first-hand way to sanity-check one of them.
One model, split by who gets to use it
Anthropic announced Claude Fable 5.1 and Claude Mythos 5.1 this month. The two are identical weights released under different safeguard levels. Fable 5.1 is broadly available through the standard Claude API, at claude-fable-5-1. Mythos 5.1 goes only to vetted professionals through two restricted programs: the Cyber Verification Program for defensive security work, and the Life Sciences Verification Program, run through a US government partnership. Both programs are currently US organizations only.
The split is a statement about what this model can do. Anthropic's safeguards page frames the gap as capability that is real but dangerous enough in the wrong hands that gatekeeping shifts from refusal behaviour to access itself. That is a different design than a single public model with a single refusal policy, and it is worth watching whether other labs adopt it as a pattern.
The benchmark numbers, and which ones matter more
Terminal-Bench-Science 0.1 went from 24.7% on Fable 5 to 52.6% on Fable 5.1. Terminal-Bench 4.0, an agentic coding benchmark, went from 42.0% to 55.8%. Humanity's Last Exam came in at 60.9% without tool use and 65.0% with it. CursorBench 3.2.0 scored 73.4%. OSWorld 2.0 scored 77.9% under partial credit and 41.7% under strict scoring, a gap worth noting on its own since it says a fair amount of the work is close but not quite right.
Of these, the Terminal-Bench-Science jump is the one we can speak to with something more than trust in a vendor's number. More than doubling a score on a benchmark suite of real scientific computing tasks is the kind of claim that deserves a second look.
What we actually know about a task family in that suite
Montana Research Foundation has an accepted task family in Terminal-Bench-Science. Dr. Ricardo Arcifa wrote a task family in the mathematical-sciences and applied-mathematics domain, a hard applied-math and system-identification problem, merged into harbor-framework/terminal-bench-science in August. We are writing about that task family and what it measures in more detail in a separate post, so we will not re-describe it here beyond that it was not easy to write and it is not easy to pass.
What that experience buys us is a rough calibration on what "hard" means inside this specific suite. Tasks in Terminal-Bench-Science are graded on whether an agent can actually run code, interpret intermediate output, and correct course, not on whether it can produce a plausible-sounding paragraph about the domain. A model going from a quarter of tasks solved to just over half is a genuine capability change if the suite's difficulty is what we experienced building that task family. We would still want to see the per-task breakdown before calling the overall number settled, since an average can move a lot if a model gets much better at a cluster of easier tasks while barely touching the hardest ones.
What changed in the safeguards themselves
Anthropic reports three specific safeguard changes alongside the capability jump. Cybersecurity safeguards cut false positives by 60% compared to earlier versions, which Anthropic says now lets the model do vulnerability discovery, meaning defensive security work, while still restricting exploit development. Biology safeguards fire 85% less often on benign elementary biology and medical queries, which is the kind of number that matters more to a working biologist than any benchmark score, since a safeguard that blocks routine questions is a safeguard nobody wants to keep using. Anthropic also describes strengthened anti-distillation mechanisms meant to resist attempts to extract the model's capabilities through outputs alone.
On the compliance side, Anthropic states that chemical and biological capabilities remain below the next tier of its Responsible Scaling Policy, that cyber capabilities sit in a lower risk category under its Frontier Compliance Framework, and that no critical-severity jailbreaks were found during testing. These are self-reported results from the company's own red-teaming process. We have no independent verification of any of them, and neither does anyone reading this post unless they were part of that testing.
The science demonstrations
Anthropic paired the release with a handful of applied science results. In molecular design, the model designed high-affinity protein binders with roughly 10 times the binding affinity of competition entries and about a 50% hit rate across 12 targets, against a typical hit rate of 10 to 15%. External lab validation reportedly confirmed that all the designs bound. In planetary science, it built higher-resolution elevation maps of Venus from 30-year-old NASA Magellan radar data, improving resolution from a 10 to 20 kilometer footprint down to 2 to 3 kilometers and height accuracy by up to 25%, released under Creative Commons. In computational biology, it optimized seven open-source deep learning models for 1.4x to 2.5x speedups and 30 to 60% lower GPU cost on genome-wide analyses, work Anthropic describes as normally taking weeks of performance engineering by hand.
These are the kind of results that are easy to be impressed by and hard to check from outside. The Venus data release under Creative Commons is the one item here that anyone can go verify directly, since the output itself is public and directly checkable. The molecular design and model-optimization claims rest on Anthropic's account of external validation and internal benchmarking respectively, and we would treat both as plausible but unconfirmed until someone outside Anthropic reproduces them.
Price, data handling, and who actually gets Mythos
Cache read pricing dropped 75% to $0.25 per million tokens. Standard pricing is $10 per million input tokens and $50 per million output tokens. Anthropic says overall cost is down about 25% versus Fable 5 for typical workloads, and up to 45% for complex agentic tasks, which is a meaningful shift if it holds across the workloads we actually run.
On data handling, Anthropic is rolling out Enterprise Frontier Safeguards starting this fall, which enable zero data retention by storing customer data on customer-controlled cloud infrastructure across Claude Code, Enterprise, Platform, AWS, Google Cloud, and Microsoft Azure. Outputs carry an invisible EU AI Act watermark that Anthropic says includes no user data. Mythos 5.1 itself stays gated behind the two verification programs described above, so most readers of this post, including us, will only ever interact with Fable 5.1. That is worth remembering when reading any claim specifically about what Mythos can do that Fable cannot. Outside evaluators have no way to test that claim for themselves right now.
What we would want before treating this as settled
Anthropic published several customer accounts alongside the release. Jane Street Capital said the model solves more coding problems than Fable 5 or Opus 5 and stays readable over long, multi-step tasks. Millennium said it found the cause of a rare crash that had gone unexplained on their team for four to five years. Rakuten said it spotted a gap in a clinical research review that three other models had missed and proposed a new hypothesis. Ramp described a 38-hour unattended run that diagnosed a prior result as a labeling artifact, corrected it, and launched six parallel follow-up experiments. These are vendor-selected quotes from named customers, and we are citing them as exactly that.
What we actually trust here is narrow. The Venus elevation data is public and checkable, our own experience with Terminal-Bench-Science gives us grounded, if partial, confidence that the suite is hard enough for a doubling to mean something, and the pricing numbers are concrete enough to test against our own usage once we have access. Everything else, the safety claims, the molecular design validation, the customer stories, rests on Anthropic's own account. That is normal for a same-day product announcement. It is also exactly the set of claims worth revisiting once independent evaluators, us included, have had time with the model itself.
Sources
From the foundation