Gemini 3.8 Flash · Comparison

Gemini 3.8 Flash vs 3.7 Flash: upgrade gains and cost trade-off

Terminal-Bench 2.1 rises 7.8 points and DeepSWE v1.1 rises 8.4 points. Cost per task rises 45% because 3.8 Flash spends about 30% more output tokens for the same job. Google's own documentation recommends staying on 3.7 Flash where efficiency is the priority.

Dated comparison: 3 September 2026

MetricGemini 3.8 FlashGemini 3.7 Flash
AA Intelligence Index (high)59n/a
Terminal-Bench 2.189.4%81.6%
DeepSWE v1.173.7%65.3%
HLE-Verified54.9%53.6%
Vals Finance Agent v261.4%59.0%
Cost per AA task (high)$0.58$0.40
Standard input / 1M$0.75$0.75
Standard output / 1M$3.75$3.75
Knowledge cutoffMarch 2026Earlier
WeightsClosedClosed

Upgrade when, stay when

Upgrade for agentic coding, finance and legal document work. Stay on 3.7 Flash for high-volume short tasks: cost per task rises 45% on 3.8 because the model spends about 30% more output tokens for the same job. Google's own documentation states that 3.7 Flash remains fully supported and is the right choice where efficiency is the priority.

Ready to test the workflow?

Create account & add credits

Comparison limits

These figures are transcribed from the 3 September 2026 research snapshot, not a new OneInfer benchmark. Verify both models against your own task mix before deciding a migration.

Frequently asked questions

How current is this information?

This page uses the verified research snapshot dated 3 September 2026. Verify current access, pricing and benchmark figures before a production decision.

Put Gemini 3.8 Flash to work

Fund a controlled evaluation, send a reference frame or document, and measure quality and cost on your own workload.