Model facet · Benchmarks

GPT-6: Astra benchmarks

OpenAI reported fifteen evaluation scores for GPT-6: Astra at launch, from FrontierMath Tier 4 (v2) at 97.6% to GeneBench Pro at 37.8%, its weakest domain. Artificial Analysis separately scored it at 61 on the Intelligence Index and 67 on the Coding Agent Index — both below Claude Fable 5.1 at an identical price.

Vendor-reported benchmarks

All fifteen scores below are from OpenAI's launch materials, dated 3 September 2026. They are vendor-reported rather than independently reproduced. Higher is better on every row.

EvaluationScore
Reasoning and knowledge
FrontierMath Tier 4 (v2)97.6%
ARC-AGI-399.9%
GPQA Diamond96.0%
Agents' Last Exam59.3%
Agentic and coding
Terminal-Bench 4.057.9%
DeepSWE v1.174.1%
OSWorld 2.072.6%
SRE-Bench (single attempt)88.0%
Terminal-Bench Science 0.164.6%
Cybersecurity
ExploitBench100.0%
ExploitGym42.4%
Domain
HealthBench Professional63.4%
LifeSciBench60.3%
GeneBench Pro37.8%
BenchCAD95.9%

The independent number that matters most: Artificial Analysis trails GPT-6: Astra behind Claude Fable 5.1

Artificial Analysis, an independent third party, scored GPT-6: Astra at 61 on its Intelligence Index and 67 on its Coding Agent Index, both on 3 September 2026. At an identical $10/$50 list price, Claude Fable 5.1 scores 66 and 70 on the same composites. This is the number that should anchor a buying decision more than any single vendor-reported evaluation above.

GeneBench Pro is the honest outlier

At 37.8%, GeneBench Pro is GPT-6: Astra's weakest published domain score by a wide margin — well below its 90%+ scores in math, cybersecurity, and CAD reasoning. State it rather than average it away: this model is not strong at genomics-specific reasoning relative to its other domains.

Ready to test the workflow?

Create account & add credits

The trade against its own predecessor, GPT-5.6 Sol

Per Artificial Analysis, GPT-6: Astra improved hallucination rate at maximum reasoning effort from 92% to 51% and became roughly 70% more token-efficient on coding relative to GPT-5.6 Sol. But GDPval-AA v2 regressed by roughly 80 Elo relative to Sol even as AA-Briefcase gained roughly 80 Elo. State this as a genuine capability trade rather than a uniform generational win.

Frequently asked questions

What is GPT-6: Astra's weakest benchmark?

GeneBench Pro, at 37.8%, is GPT-6: Astra's weakest published domain score — well below its scores in math, cybersecurity, and CAD reasoning, all of which sit above 95%.

Does GPT-6: Astra beat Claude Fable 5.1?

Not on the independent Artificial Analysis composites: at an identical $10/$50 per 1M list price, Claude Fable 5.1 scores 66 to Astra's 61 on the Intelligence Index and 70 to 67 on the Coding Agent Index.

Put GPT-6: Astra to work

Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.