Speech synthesis APIs

Compare Text-to-Speech Models and API Pricing

Explore speech synthesis models for voice agents, narration, accessibility, and conversational products. Per-character and per-minute rates remain separate comparison groups.

Current catalog

8

text-to-speech models with verified catalog metadata

Pricing content refreshed September 8, 2026

Lowest comparable price

Sarvam AI: Bulbul V2

Currently the lowest-priced option in the provider price per one million characters group. Different billing units are never mixed.

$0.180 / 1M characters

View pricing and providers

Available Text-to-Speech Models

Comparable models are ordered by their current normalized price.

8 models

Sarvam AI: Bulbul V2

sarvam

Lowest price

Bulbul v2 is a high-speed, cost-effective TTS model supporting 11 Indian languages. It is distinguished by its 'Just Like India' authentic regional accents and offers granular control over pitch, speed, and volume. Optimized for real-time synthesis in customer service and e-learning applications.

Audio OutputText To Speech

Price

$0.180 / 1M characters

Context

2000

See Sarvam AI: Bulbul V2 pricing

Sarvam AI: Bulbul V3

sarvam

Bulbul v3 is a production-grade text-to-speech model optimized for 11 Indian languages. It features native support for code-mixed speech, professional voice cloning, and industry-leading stability in telephony environments (8 kHz). It automatically infers prosody, emphasis, and emotional tone.

Audio OutputText To Speech

Price

$0.350 / 1M characters

Context

2000

See Sarvam AI: Bulbul V3 pricing

Fishaudio: S1

openrouter

Fish Audio S1 is a multilingual text-to-speech model capable of generating highly expressive speech. It utilizes a fixed vocabulary of preset emotion tags enclosed in parentheses at the beginning of sentences (e.g., sound effects, tone markers) to steer the emotion and delivery of the generated voice.

Audio OutputText To Speech

Price

$15.000 / 1M characters

Context

4K

See Fishaudio: S1 pricing

Fishaudio: S2 Pro

openrouter

Fish Audio S2 Pro is an advanced 4B parameter open-weight text-to-speech model built on a dual-autoregressive architecture. It supports approximately 80 languages and introduces fine-grained, word-level inline tag control, allowing users to embed open-ended natural language instructions inside square brackets directly in the script (e.g., [whispering], [pitch up]). It outputs 44.1 kHz audio and natively supports zero-shot voice cloning from a reference clip.

Audio OutputSpeech To TextText To Speech

Price

$15.000 / 1M characters

Context

4K

See Fishaudio: S2 Pro pricing

Fishaudio: S2.1 Pro

openrouter

Fish Audio S2.1 Pro is a state-of-the-art production-grade neural speech synthesis model optimized for ultra-low latency. With a Time-to-First-Audio (TTFA) as low as ~70-90ms, it is built for live conversational AI and turn-taking dialogue systems. It natively supports 83 languages without separate endpoints, features high-consistency voice cloning from reference audio, and excels at natural prosody for long-form narration.

Audio OutputText To Speech

Price

$15.000 / 1M characters

Context

4K

See Fishaudio: S2.1 Pro pricing

Minimax: Speech 2.8 Turbo

minimax

MiniMax Speech 2.8 Turbo is the speed-optimized variant of MiniMax's flagship text-to-speech model. Delivering sub-250 millisecond latency, it provides fast, natural, and expressive voice synthesis ideal for real-time applications like conversational AI, voice agents, and interactive experiences. It retains the core advanced capabilities of the HD tier—including 40+ language support, rapid 10-second voice cloning, emotion control, and native sound tags—while prioritizing high-throughput generation and cost-efficiency.

Audio OutputText To Speech

Price

$60.000 / 1M characters

Context

50K

See Minimax: Speech 2.8 Turbo pricing

Minimax: Speech 2.8 HD

minimax

MiniMax Speech 2.8 HD is the flagship studio-grade text-to-speech model from MiniMax. Built on an autoregressive Transformer architecture with a Flow-VAE decoder, it delivers highly expressive, broadcast-ready voice synthesis. It features native sound tags for realistic paralinguistic sounds (such as laughs, sighs, and breaths), robust emotion control, and high-fidelity voice cloning from just 10 seconds of audio. Supporting over 40 languages, it is optimized for professional audiobook narration, podcast production, and high-end video voiceovers.

Audio OutputText To Speech

Price

$100.000 / 1M characters

Context

50K

See Minimax: Speech 2.8 HD pricing

XAI: Grok Imagine Video 1.5

grok

Grok Imagine Video 1.5 is xAI's advanced video model built on the Aurora-2 engine. Its standout feature is native one-pass audio generation, which produces synchronized dialogue (lip-sync), sound effects, and background music simultaneously with the video. Supporting resolutions up to 1080p and durations up to 15 seconds, it allows for complex workflows including text-to-video, image-to-video, video extension, and multi-image reference guidance to maintain consistent styles and characters.

VisionAudio OutputVideo GenerationText To Speech

Price

See providers

Context

2K

See XAI: Grok Imagine Video 1.5 pricing

How to choose

Voice agents
Narration
Accessibility
Conversational applications

Pricing methodology

OneInfer compares only positive prices with the same billing unit. Per-minute, per-character, per-token, per-image, per-video, and per-second rates remain separate. Prices can change, so the current model page and console remain the source of truth.

Frequently asked questions

How are text-to-speech prices compared?

OneInfer compares models only when they use the same billing unit, such as price per million characters. Different units are not combined into a misleading ranking.

Does every TTS model support voice cloning?

No. Voice cloning is shown only when it is explicitly provided in verified model metadata.