128K
text, image
text
Supported
About this model
GPT-4o Mini is a distilled 400B-parameter version of GPT-4o optimized for cost-efficient omnimodal applications. Maintains core multimodal capabilities with 70% of GPT-4o's performance at 40% of the computational cost, featuring real-time audio/image processing optimizations.
Input modalities
text
Accepted as model input
image
Accepted as model input
Providers
Available routing options for this model through OneInfer.
Input tokens
$0.150 / 1M tokens
Output tokens
$0.600 / 1M tokens
Routing
OneInfer optimized
Pricing
Current OneInfer pricing for this model.
| Usage | Price |
|---|---|
| Input tokens | $0.150 / 1M tokens |
| Output tokens | $0.600 / 1M tokens |
| Cached input tokens | $0.075 / 1M tokens |
Performance
Published evaluation results associated with this model.
Multimodal Efficiency
Core Performance
Real-Time Metrics
Cost Optimization
API Guide
- 1
Generate your access token
Use your API key to generate the JWT access token required by the Models API.
- 2
Add the JWT token
Copy the generated JWT token and replace
JWT_TOKENin the example below. - 3
Check your credits
If your balance is too low, before calling the model.
- 4
Run the API example
Choose your preferred language, copy the example, and send your first model request.
curl https://api.oneinfer.ai/v1/ula/chat/completions \ -H "Authorization: Bearer $JWT_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "model": "openai/gpt-4o-mini", "messages": [ { "role": "user", "content": "Hello!" } ] }'