Sarvam AI: Bulbul V2
sarvam
Lowest priceBulbul v2 is a high-speed, cost-effective TTS model supporting 11 Indian languages. It is distinguished by its 'Just Like India' authentic regional accents and offers granular control over pitch, speed, and volume. Optimized for real-time synthesis in customer service and e-learning applications.
Audio OutputText To Speech
Sarvam AI: Bulbul V3
sarvam
Bulbul v3 is a production-grade text-to-speech model optimized for 11 Indian languages. It features native support for code-mixed speech, professional voice cloning, and industry-leading stability in telephony environments (8 kHz). It automatically infers prosody, emphasis, and emotional tone.
Audio OutputText To Speech
Fish Audio S1 is a multilingual text-to-speech model capable of generating highly expressive speech. It utilizes a fixed vocabulary of preset emotion tags enclosed in parentheses at the beginning of sentences (e.g., sound effects, tone markers) to steer the emotion and delivery of the generated voice.
Audio OutputText To Speech
Fishaudio: S2 Pro
openrouter
Fish Audio S2 Pro is an advanced 4B parameter open-weight text-to-speech model built on a dual-autoregressive architecture. It supports approximately 80 languages and introduces fine-grained, word-level inline tag control, allowing users to embed open-ended natural language instructions inside square brackets directly in the script (e.g., [whispering], [pitch up]). It outputs 44.1 kHz audio and natively supports zero-shot voice cloning from a reference clip.
Audio OutputSpeech To TextText To Speech
Fishaudio: S2.1 Pro
openrouter
Fish Audio S2.1 Pro is a state-of-the-art production-grade neural speech synthesis model optimized for ultra-low latency. With a Time-to-First-Audio (TTFA) as low as ~70-90ms, it is built for live conversational AI and turn-taking dialogue systems. It natively supports 83 languages without separate endpoints, features high-consistency voice cloning from reference audio, and excels at natural prosody for long-form narration.
Audio OutputText To Speech
Minimax: Speech 2.8 Turbo
minimax
MiniMax Speech 2.8 Turbo is the speed-optimized variant of MiniMax's flagship text-to-speech model. Delivering sub-250 millisecond latency, it provides fast, natural, and expressive voice synthesis ideal for real-time applications like conversational AI, voice agents, and interactive experiences. It retains the core advanced capabilities of the HD tier—including 40+ language support, rapid 10-second voice cloning, emotion control, and native sound tags—while prioritizing high-throughput generation and cost-efficiency.
Audio OutputText To Speech
Minimax: Speech 2.8 HD
minimax
MiniMax Speech 2.8 HD is the flagship studio-grade text-to-speech model from MiniMax. Built on an autoregressive Transformer architecture with a Flow-VAE decoder, it delivers highly expressive, broadcast-ready voice synthesis. It features native sound tags for realistic paralinguistic sounds (such as laughs, sighs, and breaths), robust emotion control, and high-fidelity voice cloning from just 10 seconds of audio. Supporting over 40 languages, it is optimized for professional audiobook narration, podcast production, and high-end video voiceovers.
Audio OutputText To Speech
XAI: Grok Imagine Video 1.5
grok
Grok Imagine Video 1.5 is xAI's advanced video model built on the Aurora-2 engine. Its standout feature is native one-pass audio generation, which produces synchronized dialogue (lip-sync), sound effects, and background music simultaneously with the video. Supporting resolutions up to 1080p and durations up to 15 seconds, it allows for complex workflows including text-to-video, image-to-video, video extension, and multi-image reference guidance to maintain consistent styles and characters.
VisionAudio OutputVideo GenerationText To Speech