Gemini 3.8 Flash scorecard
Section headers group the seventeen rows by capability. The page intentionally compares Gemini 3.8 Flash against itself only; every other model on the site compares against Gemini 3.8 Flash instead.
| Benchmark | Gemini 3.8 Flash |
|---|---|
| Agentic | |
| Terminal-Bench 2.1 | 89.4% |
| Terminal-Bench 4.0 | 19.1% |
| OSWorld 2.0 | 59.0% |
| Vals Finance Agent v2 | 61.4% |
| Harvey Legal Agent | 10.0% |
| LVBench (agentic) | 87.8% |
| Coding | |
| DeepSWE v1.1 | 73.7% |
| SWE-Bench Pro | 61.6% |
| GDP.PDF | 35.0% |
| Reasoning | |
| HLE-Verified | 54.9% |
| LVBench (static) | 87.1% |
| CharXiv Reasoning | 86.2% |
| LAB-Bench 2 | 86.2% |
| BioMysteryBench (difficult) | 56.5% |
| AA-LCR (long context) | 82.0% |
| Knowledge | |
| GDPval-AA v2 | 1545 Elo |
| AA Intelligence Index | 59 |
How to read the table
- Source spread: Terminal-Bench 2.1 is reported at 89.4% by four sources and 90.8% by DataCamp; DeepSWE v1.1 is reported at 73.7%, 71.0% and 73.8%. The 89.4% figure is the majority; the 90.8% is the outlier.
- GPQA Diamond at 95.3% is reported only by BenchmarkList. Google's own announcement does not cite it, so it is not corroborated.
- CWE-Bench pass@1 at 47.2% (Cyber variant) is measured against a frontier model at 47.8%. Quoting 47.2% alone overstates it.
- Reasoning levels: AA Intelligence Index is 59 at high, 57 at medium and 52 at low. Cost per AA task falls from $0.58 to $0.41 to $0.24 across the same three levels.
- Effort tiers and harness versions differ across evaluators. Treat each row as its own evaluator's measurement, not a like-for-like across the table.
Ready to test the workflow?
Create account & add creditsWhy 3.8 Flash scores low on some rows
Terminal-Bench 4.0 measures open-ended long-horizon agent work where the model chooses its own next step; Gemini 3.8 Flash scores 19.1% there. OSWorld 2.0 measures computer use; 3.8 Flash scores 59.0%. Computer use is listed as a capability but is not this model's strength — Google itself recommends staying on 3.7 Flash for efficiency-first workloads.
Frequently asked questions
Can I run Gemini 3.8 Flash on OneInfer?
Yes. Gemini 3.8 Flash is served on OneInfer via OpenRouter under the model identifier google/gemini-3.8-flash. Point the OpenAI-compatible base URL at https://api.oneinfer.ai/v1/ula and pass your OneInfer API key in the Authorization header.
What is Gemini 3.8 Flash's best benchmark score?
89.4% on Terminal-Bench 2.1, the highest in its evaluation cohort, reported by BenchmarkList, AIReleaseTracker, Kingy AI and BenchLM.
Where does Gemini 3.8 Flash score lowest?
Terminal-Bench 4.0 at 19.1% (open-ended long-horizon agent work) and OSWorld 2.0 at 59.0% (computer use). Reasoning, finance and legal rows are its strongest.
How fast is Gemini 3.8 Flash in tokens per second?
About 305 output tokens per second after a 13.3-second wait for the first token, measured by Artificial Analysis in September 2026.
Put Gemini 3.8 Flash to work
Fund a controlled evaluation, send a reference frame or document, and measure quality and cost on your own workload.