Audio and voice APIs

Compare Audio AI Models and API Pricing

Explore audio models for transcription, speech synthesis, voice applications, and generative audio. Because providers use different billing units, OneInfer ranks prices only within the same comparable unit.

Current catalog

7

audio models with verified catalog metadata

Pricing content refreshed July 25, 2026

Lowest comparable price

bulbul:v2

Currently the lowest-priced option in the provider price per one million characters group. Different billing units are never mixed.

$0.180 / 1M characters

View pricing and providers

Available Audio Models

Comparable models are ordered by their current normalized price.

7 models

bulbul:v2

sarvam

Lowest price

Bulbul v2 is a high-speed, cost-effective TTS model supporting 11 Indian languages. It is distinguished by its 'Just Like India' authentic regional accents and offers granular control over pitch, speed, and volume. Optimized for real-time synthesis in customer service and e-learning applications.

Audio OutputText To Speech

Price

$0.180 / 1M characters

Context

2000

See bulbul:v2 pricing

bulbul:v3

sarvam

Bulbul v3 is a production-grade text-to-speech model optimized for 11 Indian languages. It features native support for code-mixed speech, professional voice cloning, and industry-leading stability in telephony environments (8 kHz). It automatically infers prosody, emphasis, and emotional tone.

Audio OutputText To Speech

Price

$0.350 / 1M characters

Context

2000

See bulbul:v3 pricing

gpt-5.1

openai

OpenAI's GPT-5.1 model, released November 13, 2025. Features a strong multimodal and reasoning capabilities, and is optimized for complex agentic workflows. It introduced significant improvements in coding, reasoning, and tool use over previous generations.

Tool CallingVisionAudio InputSpeech To Text

Price

$5.625 avg / 1M tokens

Context

400K

See gpt-5.1 pricing

gpt-5.2-pro

e4863abbf36c4fc8b10e095cc1aa9b6d

OpenAI's flagship GPT-5.2 Pro model, Designed for professional knowledge work and agentic workflows, exclusive 'xhigh' reasoning effort, and top-tier performance on complex reasoning, coding, and scientific benchmarks.

Tool CallingVisionAudio InputSpeech To Text

Price

See providers

Context

400K

See gpt-5.2-pro pricing

saaras:v3

sarvam

Saaras v3 is a state-of-the-art speech recognition and translation model. It natively supports streaming for real-time applications and features specialized modes for transcription, direct-to-English translation, and transliteration. It is specifically tuned for noisy environments and complex code-mixed (e.g., Hindi-English) speech.

Audio InputSpeech To Text

Price

See providers

Context

N/A

See saaras:v3 pricing

saarika:v2.5

sarvam

Saarika v2.5 is Sarvam AI's legacy speech recognition model designed for Indian languages and accents. It transcribes audio in the same language spoken, excelling in multi-speaker conversations, telephony audio (8kHz), and code-mixed speech. Supports 11 languages (Hindi, Bengali, Gujarati, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, Telugu, English) with automatic language detection and speaker diarization. Achieves 4.96% CER and 18.32% WER on VISTAAR benchmark.

Audio InputSpeech To Text

Price

See providers

Context

N/A

See saarika:v2.5 pricing

speech-2.8-turbo

cc94bcd3662444fb92e43506c7036c08

Speech 2.8-turbo is a text-to-speech model

Audio OutputText To Speech

Price

See providers

Context

2K

See speech-2.8-turbo pricing

How to choose

Voice agents
Transcription
Speech synthesis
Audio generation

Pricing methodology

OneInfer compares only positive prices with the same billing unit. Per-minute, per-character, per-token, per-image, per-video, and per-second rates remain separate. Prices can change, so the current model page and console remain the source of truth.

Frequently asked questions

What is the cheapest audio model on OneInfer?

The lowest-priced model is calculated from the current catalog within each comparable billing unit. Per-minute, per-character, and token-based prices are not mixed.

Why are some audio models not ranked together?

Audio providers charge by different units. A transcription model billed per minute cannot be honestly ranked against a speech model billed per character without a common workload assumption.