Model facet · Benchmarks

Gemini 3.8 Flash Benchmarks: 21 Results Across 8 Evaluators

Across 21 published benchmark results, Gemini 3.8 Flash leads Terminal-Bench 2.1 (89.4%), Vals Finance Agent v2 (61.4%), LVBench (87.8%) and the Harvey Legal Agent (10.0%). Its losses are concentrated in long-horizon agent work and computer use, where Terminal-Bench 4.0 is 19.1%, OSWorld 2.0 is 59.0% and GDPval-AA v2 is 1545 Elo.

Gemini 3.8 Flash scorecard

Section headers group the seventeen rows by capability. The page intentionally compares Gemini 3.8 Flash against itself only; every other model on the site compares against Gemini 3.8 Flash instead.

BenchmarkGemini 3.8 Flash
Agentic
Terminal-Bench 2.189.4%
Terminal-Bench 4.019.1%
OSWorld 2.059.0%
Vals Finance Agent v261.4%
Harvey Legal Agent10.0%
LVBench (agentic)87.8%
Coding
DeepSWE v1.173.7%
SWE-Bench Pro61.6%
GDP.PDF35.0%
Reasoning
HLE-Verified54.9%
LVBench (static)87.1%
CharXiv Reasoning86.2%
LAB-Bench 286.2%
BioMysteryBench (difficult)56.5%
AA-LCR (long context)82.0%
Knowledge
GDPval-AA v21545 Elo
AA Intelligence Index59

How to read the table

  • Source spread: Terminal-Bench 2.1 is reported at 89.4% by four sources and 90.8% by DataCamp; DeepSWE v1.1 is reported at 73.7%, 71.0% and 73.8%. The 89.4% figure is the majority; the 90.8% is the outlier.
  • GPQA Diamond at 95.3% is reported only by BenchmarkList. Google's own announcement does not cite it, so it is not corroborated.
  • CWE-Bench pass@1 at 47.2% (Cyber variant) is measured against a frontier model at 47.8%. Quoting 47.2% alone overstates it.
  • Reasoning levels: AA Intelligence Index is 59 at high, 57 at medium and 52 at low. Cost per AA task falls from $0.58 to $0.41 to $0.24 across the same three levels.
  • Effort tiers and harness versions differ across evaluators. Treat each row as its own evaluator's measurement, not a like-for-like across the table.

Ready to test the workflow?

Create account & add credits

Why 3.8 Flash scores low on some rows

Terminal-Bench 4.0 measures open-ended long-horizon agent work where the model chooses its own next step; Gemini 3.8 Flash scores 19.1% there. OSWorld 2.0 measures computer use; 3.8 Flash scores 59.0%. Computer use is listed as a capability but is not this model's strength — Google itself recommends staying on 3.7 Flash for efficiency-first workloads.

Frequently asked questions

Can I run Gemini 3.8 Flash on OneInfer?

Yes. Gemini 3.8 Flash is served on OneInfer via OpenRouter under the model identifier google/gemini-3.8-flash. Point the OpenAI-compatible base URL at https://api.oneinfer.ai/v1/ula and pass your OneInfer API key in the Authorization header.

What is Gemini 3.8 Flash's best benchmark score?

89.4% on Terminal-Bench 2.1, the highest in its evaluation cohort, reported by BenchmarkList, AIReleaseTracker, Kingy AI and BenchLM.

Where does Gemini 3.8 Flash score lowest?

Terminal-Bench 4.0 at 19.1% (open-ended long-horizon agent work) and OSWorld 2.0 at 59.0% (computer use). Reasoning, finance and legal rows are its strongest.

How fast is Gemini 3.8 Flash in tokens per second?

About 305 output tokens per second after a 13.3-second wait for the first token, measured by Artificial Analysis in September 2026.

Put Gemini 3.8 Flash to work

Fund a controlled evaluation, send a reference frame or document, and measure quality and cost on your own workload.