Speech synthesis APIs

Compare Text-to-Speech Models and API Pricing

Explore speech synthesis models for voice agents, narration, accessibility, and conversational products. Per-character and per-minute rates remain separate comparison groups.

Current catalog

3

text-to-speech models with verified catalog metadata

Pricing content refreshed July 25, 2026

Lowest comparable price

bulbul:v2

Currently the lowest-priced option in the provider price per one million characters group. Different billing units are never mixed.

$0.180 / 1M characters

View pricing and providers

Available Text-to-Speech Models

Comparable models are ordered by their current normalized price.

3 models

bulbul:v2

sarvam

Lowest price

Bulbul v2 is a high-speed, cost-effective TTS model supporting 11 Indian languages. It is distinguished by its 'Just Like India' authentic regional accents and offers granular control over pitch, speed, and volume. Optimized for real-time synthesis in customer service and e-learning applications.

Audio OutputText To Speech

Price

$0.180 / 1M characters

Context

2000

See bulbul:v2 pricing

bulbul:v3

sarvam

Bulbul v3 is a production-grade text-to-speech model optimized for 11 Indian languages. It features native support for code-mixed speech, professional voice cloning, and industry-leading stability in telephony environments (8 kHz). It automatically infers prosody, emphasis, and emotional tone.

Audio OutputText To Speech

Price

$0.350 / 1M characters

Context

2000

See bulbul:v3 pricing

speech-2.8-turbo

cc94bcd3662444fb92e43506c7036c08

Speech 2.8-turbo is a text-to-speech model

Audio OutputText To Speech

Price

See providers

Context

2K

See speech-2.8-turbo pricing

How to choose

Voice agents
Narration
Accessibility
Conversational applications

Pricing methodology

OneInfer compares only positive prices with the same billing unit. Per-minute, per-character, per-token, per-image, per-video, and per-second rates remain separate. Prices can change, so the current model page and console remain the source of truth.

Frequently asked questions

How are text-to-speech prices compared?

OneInfer compares models only when they use the same billing unit, such as price per million characters. Different units are not combined into a misleading ranking.

Does every TTS model support voice cloning?

No. Voice cloning is shown only when it is explicitly provided in verified model metadata.