Research snapshot leaderboard
| Configuration | AA index | Tokens/s | Weights |
|---|---|---|---|
| Claude Fable 5.1 max with fallback | 66 | 62 | Closed |
| Claude Fable 5.1 xhigh with fallback | 65 | 60 | Closed |
| Claude Opus 5 max | 63 | 56 | Closed |
| Claude Fable 5.1 high with fallback | 62 | 55 | Closed |
| Muse Spark 1.3 max (gated) | 62 | Not reported | Closed |
| Claude Fable 5 with fallback | 62 | 65 | Closed |
| Muse Spark 1.3 xhigh | 61 | 186 | Closed |
| GPT-5.6 Sol max | 61 | 69 | Closed |
| Grok 4.6 high | 61 | 55 | Closed |
| Claude Opus 5 high | 61 | 49 | Closed |
| Kimi K3 max | 60 | 38 | Open |
| GLM-5.3 max | 60 | 63 | Open |
Rank on completed work
Select representative repository fixes, hold tests and permissions constant, and score accepted changes rather than lines of code. Record total time and token cost, including recovery attempts.
Ready to test the workflow?
Create account & add creditsSnapshot, not an independent retest
These configurations come from the supplied 3 September 2026 leaderboard. Multiple rows represent different efforts of the same model.
Frequently asked questions
How current is this information?
This page uses the supplied research snapshot dated 3 September 2026. Verify current access, pricing and provider terms before an evaluation.
Put Muse Spark 1.3 to work
Fund a controlled evaluation, send a reference frame or document, and measure quality and cost on your own workload.