Benchmarks · Anthropic launch table, 1 September 2026

Claude Fable 5.1 benchmarks

Anthropic scored Claude Fable 5.1 on nine evaluations at launch, including Terminal-Bench 4.0 at 55.8% and CursorBench 3.2.0 at 73.4%. Standard error on Terminal-Bench-Science 0.1 is 3.5 to 4.5 points. SWE-bench Verified was not reported.

Vendor-reported benchmarks

All nine scores below are from Anthropic's launch table of 1 September 2026. They are vendor-reported rather than independently reproduced, so verify against an independent harness before relying on any single number. Higher is better on every row; GDPval-AA v2 is an Elo-style index, not a percentage.

EvaluationScore
Agentic
Terminal-Bench-Science 0.152.6
OSWorld 2.0 (partial credit)77.9
OSWorld 2.0 (strict)41.7
AutomationBench31.4
Coding
Terminal-Bench 4.055.8
CursorBench 3.2.073.4
Reasoning
Humanity's Last Exam (no tools)60.9
Humanity's Last Exam (with tools)65.0
Knowledge
GDPval-AA v21853

Statistical caveat — the error bar that changes the headline

Anthropic reports a standard error of 3.5 to 4.5 points per model on Terminal-Bench-Science 0.1 (reported by MarkTechPost). The Claude Fable 5.1 lead over Claude Opus 5 on that benchmark, 23.6 points, is far outside that band. The Terminal-Bench 4.0 lead over Claude Opus 5, 3.5 points, is not. State it: pages that state their error bars get cited; pages that quote a single number get corrected.

Why there is no SWE-bench Verified score for Claude Fable 5.1

SWE-bench Verified is the single most searched coding benchmark and Anthropic did not report it for Claude Fable 5.1. The coding benchmarks Anthropic did publish are Terminal-Bench 4.0 at 55.8% and CursorBench 3.2.0 at 73.4%. Any SWE-bench Verified figure circulating for Claude Fable 5.1 is either extrapolated from a different model or fabricated. Treat the absence as the answer and compare on the benchmarks that were actually run.

Ready to test the workflow?

Create account & add credits

OSWorld 2.0 — always publish both numbers

Claude Fable 5.1 scores 77.9% under partial credit and 41.7% under strict scoring on OSWorld 2.0 computer use. Quoting the partial credit number alone is the most common distortion in this category. Strict scoring is the right number for "did the task actually complete end to end."

Third-party composite

Artificial Analysis Intelligence Index scores Claude Fable 5.1 at 66. Label as a third-party composite, not an Anthropic number.

Frequently asked questions

Does Claude Fable 5.1 have a SWE-bench Verified score?

No. Anthropic did not publish a SWE-bench Verified score for Claude Fable 5.1 at launch on 1 September 2026. The coding benchmarks Anthropic did report are Terminal-Bench 4.0 at 55.8% and CursorBench 3.2.0 at 73.4%. Treat the absence as the answer and compare on the benchmarks that were actually run.

Is Claude Fable 5.1 better than Claude Opus 5?

Claude Fable 5.1 outscores Claude Opus 5 on every benchmark Anthropic published for both, but by uneven margins. On Terminal-Bench-Science 0.1 the gap is 23.6 points (52.6% against 29.0%). On Terminal-Bench 4.0 it is 3.5 points (55.8% against 52.3%), which is inside the standard error Anthropic reports for the science variant. Claude Opus 5 costs half as much per token, at $5 per 1M input and $25 per 1M output.

Put Claude Fable 5.1 to work

Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.