One API for models. On-demand GPUs. Pay only for what you use.
No subscription and no minimum spend. Compare live model and GPU rates from the OneInfer catalog, then move from testing to production through one platform.
Pricing across the model catalog
Text, image, audio, and video models billed by their native pricing unit. Below are the first 10 models returned by the live catalog API.
| Model | Modality | Context | Input | Output |
|---|---|---|---|---|
| Loading current model pricing… | ||||
Showing 0 of 0 models from the API.
Browse full model catalog →On-demand GPUs across multiple providers
Compare provider inventory by GPU family, regions, and the lowest current hourly rate. Below are the first 10 unique GPUs returned by the live API.
| GPU | VRAM | Providers | Regions | From |
|---|---|---|---|---|
| Loading current GPU pricing… | ||||
Showing 0 of 0 GPU families from the API.
Open GPU marketplace →Need reserved capacity or a custom SLA?
Private model endpoints, dedicated GPU capacity, and custom commercial terms for teams running production traffic.
How pricing works
Clear usage-based rates, with the current catalog loaded directly from OneInfer APIs.
Live catalog data
Model and GPU tables on this page load from the same APIs that power the OneInfer marketplaces.
Usage-based billing
Model usage follows each model's native unit, while GPU compute is displayed by the hour.
One platform
Compare multiple model and infrastructure providers without rebuilding your application for each one.
Dedicated options
Production teams can request reserved capacity, private endpoints, and account-specific commercial terms.