Vendor-reported benchmarks
All fifteen scores below are from OpenAI's launch materials, dated 3 September 2026. They are vendor-reported rather than independently reproduced. Higher is better on every row.
| Evaluation | Score |
|---|---|
| Reasoning and knowledge | |
| FrontierMath Tier 4 (v2) | 97.6% |
| ARC-AGI-3 | 99.9% |
| GPQA Diamond | 96.0% |
| Agents' Last Exam | 59.3% |
| Agentic and coding | |
| Terminal-Bench 4.0 | 57.9% |
| DeepSWE v1.1 | 74.1% |
| OSWorld 2.0 | 72.6% |
| SRE-Bench (single attempt) | 88.0% |
| Terminal-Bench Science 0.1 | 64.6% |
| Cybersecurity | |
| ExploitBench | 100.0% |
| ExploitGym | 42.4% |
| Domain | |
| HealthBench Professional | 63.4% |
| LifeSciBench | 60.3% |
| GeneBench Pro | 37.8% |
| BenchCAD | 95.9% |
The independent number that matters most: Artificial Analysis trails GPT-6: Astra behind Claude Fable 5.1
Artificial Analysis, an independent third party, scored GPT-6: Astra at 61 on its Intelligence Index and 67 on its Coding Agent Index, both on 3 September 2026. At an identical $10/$50 list price, Claude Fable 5.1 scores 66 and 70 on the same composites. This is the number that should anchor a buying decision more than any single vendor-reported evaluation above.
GeneBench Pro is the honest outlier
At 37.8%, GeneBench Pro is GPT-6: Astra's weakest published domain score by a wide margin — well below its 90%+ scores in math, cybersecurity, and CAD reasoning. State it rather than average it away: this model is not strong at genomics-specific reasoning relative to its other domains.
Ready to test the workflow?
Create account & add creditsThe trade against its own predecessor, GPT-5.6 Sol
Per Artificial Analysis, GPT-6: Astra improved hallucination rate at maximum reasoning effort from 92% to 51% and became roughly 70% more token-efficient on coding relative to GPT-5.6 Sol. But GDPval-AA v2 regressed by roughly 80 Elo relative to Sol even as AA-Briefcase gained roughly 80 Elo. State this as a genuine capability trade rather than a uniform generational win.
Frequently asked questions
What is GPT-6: Astra's weakest benchmark?
GeneBench Pro, at 37.8%, is GPT-6: Astra's weakest published domain score — well below its scores in math, cybersecurity, and CAD reasoning, all of which sit above 95%.
Does GPT-6: Astra beat Claude Fable 5.1?
Not on the independent Artificial Analysis composites: at an identical $10/$50 per 1M list price, Claude Fable 5.1 scores 66 to Astra's 61 on the Intelligence Index and 70 to 67 on the Coding Agent Index.
Put GPT-6: Astra to work
Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.