Pricing

One API for models. On-demand GPUs. Pay only for what you use.

No subscription and no minimum spend. Compare live model and GPU rates from the OneInfer catalog, then move from testing to production through one platform.

models in the catalog
GPU providers
GPU families
Pay as you go
no minimum spend
Model APIs

Pricing across the model catalog

Text, image, audio, and video models billed by their native pricing unit. Below are the first 10 models returned by the live catalog API.

ModelModalityContextInputOutput
Loading current model pricing…

Showing 0 of 0 models from the API.

Browse full model catalog →
GPU Cloud

On-demand GPUs across multiple providers

Compare provider inventory by GPU family, regions, and the lowest current hourly rate. Below are the first 10 unique GPUs returned by the live API.

GPUVRAMProvidersRegionsFrom
Loading current GPU pricing…

Showing 0 of 0 GPU families from the API.

Open GPU marketplace →
Dedicated Deployments

Need reserved capacity or a custom SLA?

Private model endpoints, dedicated GPU capacity, and custom commercial terms for teams running production traffic.

Methodology

How pricing works

Clear usage-based rates, with the current catalog loaded directly from OneInfer APIs.

Live catalog data

Model and GPU tables on this page load from the same APIs that power the OneInfer marketplaces.

Usage-based billing

Model usage follows each model's native unit, while GPU compute is displayed by the hour.

One platform

Compare multiple model and infrastructure providers without rebuilding your application for each one.

Dedicated options

Production teams can request reserved capacity, private endpoints, and account-specific commercial terms.