Harness caveat: do not read the ARC-AGI-3 rows as a head-to-head
OpenAI reports GPT-6: Astra at 99.9% on ARC-AGI-3. The Claude Opus 5 benchmark record in the OneInfer catalog lists an arc_agi_3 score of 30.2. A 69.7-point gap on a benchmark with the same name almost certainly means the two numbers were produced on different task subsets, difficulty tiers, or scoring harnesses rather than a genuine capability gap of that size. Treat this row as evidence of a naming collision, not a verified comparison, until both vendors publish a shared methodology.
Price
| Dimension | GPT-6: Astra | Claude Opus 5 | Notes |
|---|---|---|---|
| Input $/1M | $10.00 | $5.00 | Astra is 2x |
| Output $/1M | $50.00 | $25.00 | Astra is 2x |
| Cached input $/1M | $1.00 | $0.50 | Astra is 2x |
What each vendor actually published
Claude Opus 5's catalog record carries swe_bench_pro at 79.2, frontier_bench_v0_1 at 43.3, and gdpval_aa_v2 at 1861.0 — none of which OpenAI has published a matching score for on GPT-6: Astra. Astra's own launch table reports FrontierMath Tier 4 (v2), ExploitBench, GPQA Diamond and others that Anthropic has not run for Opus 5. Outside the disputed ARC-AGI-3 row, there is currently no benchmark evaluated on both models under a shared methodology.
Ready to test the workflow?
Create account & add creditsWhen to pick which
Claude Opus 5 is the lower-cost route at half the per-token price with no evidence Astra clears a comparable bar on a shared benchmark. Pick GPT-6: Astra specifically for its 1,050,000-token context ceiling or one of its own published strengths (agentic coding, cybersecurity evaluation, computer use) where no Opus 5 equivalent has been published at all.
Frequently asked questions
Is GPT-6: Astra better than Claude Opus 5?
There is no shared, trustworthy benchmark to answer this directly. The one evaluation both catalogs report by the same name, ARC-AGI-3, shows a 69.7-point gap (99.9% for Astra against 30.2 for Opus 5) that is almost certainly a harness or task-set mismatch rather than a real capability difference of that size. On price, Opus 5 costs half as much per token.
Why is GPT-6: Astra so much more expensive than Claude Opus 5?
GPT-6: Astra lists at $10 per 1M input and $50 per 1M output, twice Claude Opus 5's $5 and $25. Astra's premium reflects its position as OpenAI's frontier flagship rather than a mid-tier model; Opus 5 is Anthropic's lower-cost tier within its own current lineup.
Put GPT-6: Astra to work
Fund a controlled evaluation, start with a prepared prompt, and measure quality and cost on your own workload.