1M
text, file, image, video, audio
text
Supported
About this model
Gemini 3.8 Flash is Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, complex enterprise workflows, advanced reasoning, and multimodal understanding. Building on Gemini 3.7 Flash, it delivers substantial improvements across coding, agentic tasks, and critical multi-step reasoning while retaining Flash-level speed and cost efficiency. It supports a 1M-token context window, configurable thinking levels, function calling, computer use, code execution, search grounding, structured outputs, file search, and URL context.
Input modalities
text
Accepted as model input
file
Accepted as model input
image
Accepted as model input
video
Accepted as model input
audio
Accepted as model input
Providers
Available routing options for this model through OneInfer.
Input tokens
$0.750 / 1M tokens
Output tokens
$3.750 / 1M tokens
Routing
OneInfer optimized
Pricing
Current OneInfer pricing for this model.
| Usage | Price |
|---|---|
| Input tokens | $0.750 / 1M tokens |
| Output tokens | $3.750 / 1M tokens |
| Cached input tokens | $0.075 / 1M tokens |
Performance
Published evaluation results associated with this model.
Coding
Knowledge Work
Professional
Agentic
Document Understanding
Reasoning
Video Understanding
Multidisciplinary
Scientific Research
Bioinformatics
API Guide
- 1
Generate your access token
Use your API key to generate the JWT access token required by the Models API.
- 2
Add the JWT token
Copy the generated JWT token and replace
JWT_TOKENin the example below. - 3
Check your credits
If your balance is too low, before calling the model.
- 4
Run the API example
Choose your preferred language, copy the example, and send your first model request.
curl https://api.oneinfer.ai/v1/ula/chat/completions \ -H "Authorization: Bearer $JWT_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "model": "Unmapped", "messages": [ { "role": "user", "content": "Hello!" } ] }'