Vendor-reported benchmarks
All nine scores below are from Anthropic's launch table of 1 September 2026. They are vendor-reported rather than independently reproduced, so verify against an independent harness before relying on any single number. Higher is better on every row; GDPval-AA v2 is an Elo-style index, not a percentage.
| Evaluation | Score |
|---|---|
| Agentic | |
| Terminal-Bench-Science 0.1 | 52.6 |
| OSWorld 2.0 (partial credit) | 77.9 |
| OSWorld 2.0 (strict) | 41.7 |
| AutomationBench | 31.4 |
| Coding | |
| Terminal-Bench 4.0 | 55.8 |
| CursorBench 3.2.0 | 73.4 |
| Reasoning | |
| Humanity's Last Exam (no tools) | 60.9 |
| Humanity's Last Exam (with tools) | 65.0 |
| Knowledge | |
| GDPval-AA v2 | 1853 |
Statistical caveat — the error bar that changes the headline
Anthropic reports a standard error of 3.5 to 4.5 points per model on Terminal-Bench-Science 0.1 (reported by MarkTechPost). The Claude Fable 5.1 lead over Claude Opus 5 on that benchmark, 23.6 points, is far outside that band. The Terminal-Bench 4.0 lead over Claude Opus 5, 3.5 points, is not. State it: pages that state their error bars get cited; pages that quote a single number get corrected.
Why there is no SWE-bench Verified score for Claude Fable 5.1
SWE-bench Verified is the single most searched coding benchmark and Anthropic did not report it for Claude Fable 5.1. The coding benchmarks Anthropic did publish are Terminal-Bench 4.0 at 55.8% and CursorBench 3.2.0 at 73.4%. Any SWE-bench Verified figure circulating for Claude Fable 5.1 is either extrapolated from a different model or fabricated. Treat the absence as the answer and compare on the benchmarks that were actually run.
Ready to test the workflow?
Create account & add creditsOSWorld 2.0 — always publish both numbers
Claude Fable 5.1 scores 77.9% under partial credit and 41.7% under strict scoring on OSWorld 2.0 computer use. Quoting the partial credit number alone is the most common distortion in this category. Strict scoring is the right number for "did the task actually complete end to end."
Third-party composite
Artificial Analysis Intelligence Index scores Claude Fable 5.1 at 66. Label as a third-party composite, not an Anthropic number.
Frequently asked questions
Does Claude Fable 5.1 have a SWE-bench Verified score?
No. Anthropic did not publish a SWE-bench Verified score for Claude Fable 5.1 at launch on 1 September 2026. The coding benchmarks Anthropic did report are Terminal-Bench 4.0 at 55.8% and CursorBench 3.2.0 at 73.4%. Treat the absence as the answer and compare on the benchmarks that were actually run.
Is Claude Fable 5.1 better than Claude Opus 5?
Claude Fable 5.1 outscores Claude Opus 5 on every benchmark Anthropic published for both, but by uneven margins. On Terminal-Bench-Science 0.1 the gap is 23.6 points (52.6% against 29.0%). On Terminal-Bench 4.0 it is 3.5 points (55.8% against 52.3%), which is inside the standard error Anthropic reports for the science variant. Claude Opus 5 costs half as much per token, at $5 per 1M input and $25 per 1M output.
Put Claude Fable 5.1 to work
Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.