50K
text
audio
Not listed
About this model
MiniMax Speech 2.8 HD is the flagship studio-grade text-to-speech model from MiniMax. Built on an autoregressive Transformer architecture with a Flow-VAE decoder, it delivers highly expressive, broadcast-ready voice synthesis. It features native sound tags for realistic paralinguistic sounds (such as laughs, sighs, and breaths), robust emotion control, and high-fidelity voice cloning from just 10 seconds of audio. Supporting over 40 languages, it is optimized for professional audiobook narration, podcast production, and high-end video voiceovers.
Input modalities
text
Accepted as model input
Providers
Available routing options for this model through OneInfer.
Text to speech
$100.000 / 1M characters
Routing
OneInfer optimized
Pricing
Current OneInfer pricing for this model.
| Usage | Price |
|---|---|
| Text to speechTTS | $100.000 / 1M characters |
Performance
Published evaluation results associated with this model.
API example
curl -X POST https://api.oneinfer.ai/v1/ula/generate-audio \
-H "Authorization: Bearer YOUR_JWT_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"provider": "minimax",
"model": "MiniMaxAI/speech-2.8-hd",
"prompt": "OneInfer makes AI inference simple.",
"stream": false,
"voice_id": "English_expressive_narrator",
"format": "mp3"
}'