Skip to main content

Claude Fable 5.1: A Straight Review

Anthropic's new flagship, released September 2026. What it's actually good at, what it costs, and where it sits next to Opus 5 and Sonnet 5.

Anthropic shipped Claude Fable 5.1 in September 2026 as the new general-purpose flagship, alongside a restricted sibling called Mythos 5.1 for vetted cybersecurity and life-sciences researchers — same underlying model, different safeguard levels. Here's what it actually does, including how it stacks up against the rest of the field.

Where it sits in the lineup

Opus 5, Sonnet 5, and Haiku 4.5 are still around. Fable 5.1 slots in as the flagship: Anthropic's own benchmarks show it outperforming or matching Opus 5 on most tasks, while using fewer tokens to get there. It's not a lightweight sibling — it's the top of the current lineup.

Anthropic's own numbers

  • Terminal-Bench 4.0 (agentic coding): 55.8–60.9%, up from 42.0% on Fable 5
  • CursorBench 3.2.0 (coding): 73.4%
  • Humanity's Last Exam (multidisciplinary reasoning): 60.9% without tools, 65.0% with tools
  • Terminal-Bench-Science (agentic scientific research): 52.6%, up from 24.7% on the prior generation
  • GDPval-AA v2 (knowledge work): 1853 Elo

Jane Street reported it solving more of their internal coding problems than either Fable 5 or Opus 5. Millennium had it catch a rare software crash that had eluded engineers and other models for years.

How it compares against GPT and Gemini

Outside Anthropic's own reporting, independent trackers put it at the top of the field. Artificial Analysis' Intelligence Index scores Fable 5.1 at 66, ahead of Claude Opus 5 at 63 and GPT-5.6 Sol at 61. On the Vals Index, Fable 5.1 leads at 67.87%, just ahead of Opus 5 (67.21%) and the prior Fable 5 (66.04%). Its predecessor, Fable 5, had already been dominating coding-specific benchmarks like SWE-bench Pro (80.3% vs. GPT-5.5's 58.6% and Gemini 3.1's 54.2%) — 5.1 extends that lead further via the Terminal-Bench 4.0 jump above.

What BridgeBench shows

BridgeBench, the open coding-model benchmark that splits evaluation into seven categories (UI generation, security, refactoring, hallucination resistance, debugging, speed, cost efficiency) across 130+ real-world tasks, hadn't published a dedicated Fable 5.1 page as of this writing. Its most recent published numbers are for Fable 5: a 41.5% overall score on its 30-task grounded-reasoning suite, and 47.6% on its UI-bench suite (7 of 12 tasks passed). Worth checking their site directly for a 5.1-specific breakdown before it's used as a deciding factor — the category-level view is the useful part, and it wasn't available for 5.1 yet at publication.

Pricing

$10 per million input tokens, $50 per million output tokens, cache reads at $0.25 per million (a 75% discount). Anthropic states roughly 25% savings for typical workloads and up to 45% for agentic, multi-step tasks — driven by needing fewer tokens to reach a result, not a lower sticker price. For reference, GPT-5.5 ran roughly $8/$24 and Gemini 3.1 roughly $2/$12 — Fable 5.1 is priced at the top of the field, and the case for it rests on token efficiency closing that gap on real agentic work.

Where it's strongest

Long-running, multi-step agentic work: extended debugging sessions, research tasks that run for a while unattended, coding sessions that span many files and decisions. That's also where the token-efficiency gains concentrate, so the cost advantage and the capability advantage point the same direction.

Where it's just fine, not remarkable

A single, narrow, one-shot task — a short answer, a small isolated function — doesn't show much daylight between this and Opus 5, or between any of the top three labs' flagships. The gap opens up as the task gets longer and more agentic, not on quick one-off prompts. And on pure science reasoning, Gemini has led before (GPQA Diamond) even while trailing on coding — no single model wins every category.

Bottom line

It's the new default flagship, not a specialized or lightweight option next to one, and independent trackers agree it currently leads the field on most fronts. If you're already reaching for the top-tier model on hard or long-running work, this replaces it outright rather than sitting alongside it — just don't take one benchmark's word for it on a task that matters; check the category that actually matches your workload.

Message me on WhatsApp