Dated comparison: 3 September 2026
| Metric | Gemini 3.8 Flash | GPT-5.6 Sol |
|---|---|---|
| AA Intelligence Index | 59 | 59 |
| Terminal-Bench 2.1 | 89.4% | 88.8% |
| Terminal-Bench 4.0 | 19.1% | 37.3% |
| GDPval-AA v2 | 1545 Elo | 1710 Elo |
| Vals Finance Agent v2 | 61.4% | 53.8% |
| Harvey Legal Agent | 10.0% | 2.5% |
| HLE-Verified | 54.9% | 54.5% |
| GDP.PDF | 35.0% | 40.0% |
| Output speed | ~305 tok/s | Lower |
| Standard input / 1M | $0.75 | ~$5 |
| Standard output / 1M | $3.75 | ~$30 |
| Weights | Closed | Closed |
How to choose
If the workload is finance, legal, agentic coding or any task where output speed and price dominate, Gemini 3.8 Flash is the like-for-like pick on the same index score. If the workload is open-ended professional work (GDPval-AA v2) or long-horizon agents (Terminal-Bench 4.0), GPT-5.6 Sol wins on those benchmarks despite the same index.
Ready to test the workflow?
Create account & add creditsComparison limits
These figures are transcribed from the 3 September 2026 research snapshot. GPT-5.6 Sol's per-token prices are approximate ranges reported by third-party aggregators. The AA Intelligence Index is a composite, not a per-benchmark score.
Frequently asked questions
How current is this information?
This page uses the verified research snapshot dated 3 September 2026. Verify current access, pricing and benchmark figures before a production decision.
Put Gemini 3.8 Flash to work
Fund a controlled evaluation, send a reference frame or document, and measure quality and cost on your own workload.