Muse Spark 1.1 is Meta's flagship proprietary multimodal frontier model developed by Meta Superintelligence Labs (MSL). Designed for complex reasoning and agentic tasks, it can actively manage a 1-million-token context window, orchestrate multi-agent systems, and natively process text, image, video, and audio inputs. It excels at long-horizon software development, tool use, and computer-use workflows.
Tool CallingVisionAudio InputSpeech To Text
Muse Spark 1.2 is a proprietary coding-focused frontier model developed by Meta Superintelligence Labs. Co-trained specifically with the Muse Code terminal agent, it features significant improvements in code generation, complex debugging, codebase understanding, and end-to-end developer workflows. It leverages planning, goal conditioning, and context compaction to sustain progress on long-horizon tasks like whole-repository generation and GPU kernel optimization.
Tool CallingVisionAudio InputSpeech To Text
Sarvam AI: Bulbul V2
sarvam
Bulbul v2 is a high-speed, cost-effective TTS model supporting 11 Indian languages. It is distinguished by its 'Just Like India' authentic regional accents and offers granular control over pitch, speed, and volume. Optimized for real-time synthesis in customer service and e-learning applications.
Audio OutputText To Speech
Sarvam AI: Bulbul V3
sarvam
Bulbul v3 is a production-grade text-to-speech model optimized for 11 Indian languages. It features native support for code-mixed speech, professional voice cloning, and industry-leading stability in telephony environments (8 kHz). It automatically infers prosody, emphasis, and emotional tone.
Audio OutputText To Speech
Fish Audio S1 is a multilingual text-to-speech model capable of generating highly expressive speech. It utilizes a fixed vocabulary of preset emotion tags enclosed in parentheses at the beginning of sentences (e.g., sound effects, tone markers) to steer the emotion and delivery of the generated voice.
Audio OutputText To Speech
Fishaudio: S2 Pro
openrouter
Fish Audio S2 Pro is an advanced 4B parameter open-weight text-to-speech model built on a dual-autoregressive architecture. It supports approximately 80 languages and introduces fine-grained, word-level inline tag control, allowing users to embed open-ended natural language instructions inside square brackets directly in the script (e.g., [whispering], [pitch up]). It outputs 44.1 kHz audio and natively supports zero-shot voice cloning from a reference clip.
Audio OutputSpeech To TextText To Speech
Fishaudio: S2.1 Pro
openrouter
Fish Audio S2.1 Pro is a state-of-the-art production-grade neural speech synthesis model optimized for ultra-low latency. With a Time-to-First-Audio (TTFA) as low as ~70-90ms, it is built for live conversational AI and turn-taking dialogue systems. It natively supports 83 languages without separate endpoints, features high-consistency voice cloning from reference audio, and excels at natural prosody for long-form narration.
Audio OutputText To Speech
Google: Gemini 2.5 Flash Lite
openrouter
Lowest priceGemini 2.5 Flash-Lite is Google's lightweight and cost-efficient reasoning model in the Gemini 2.5 family, optimized for high-throughput, low-latency workloads. It supports multimodal inputs including text, images, audio, video, and PDFs, with optional thinking for applications that need a balance between speed, cost, and reasoning quality.
Tool CallingVisionAudio InputSpeech To Text
Google: Gemini 3.1 Flash-Lite
openrouter
Gemini 3.1 Flash-Lite is Google's low-latency, cost-effective multimodal model optimized for high-frequency and high-volume workloads. It supports text, image, video, audio, and PDF inputs and is designed for lightweight reasoning, data extraction, agentic workflows, tool use, and applications where latency and API cost are primary considerations.
Tool CallingVisionAudio InputSpeech To Text
Google: Gemini 2.5 Flash
openrouter
Gemini 2.5 Flash is Google's high-performance multimodal workhorse model designed for reasoning, coding, mathematics, science, document understanding, and agentic applications while maintaining Flash-tier speed and cost efficiency.
Tool CallingVisionAudio InputSpeech To Text
Google: Gemini 3.5 Flash Lite
openrouter
Gemini 3.5 Flash Lite is Google's high-efficiency multimodal model with upgraded agentic capabilities, optimized for low-cost, high-volume workloads and focused subagents operating inside complex multi-agent systems.
Tool CallingVisionAudio InputSpeech To Text
Google: Gemini 3.6 Flash
openrouter
Gemini 3.6 Flash is a frontier-class Google multimodal Flash model designed for fast reasoning, coding, multimodal understanding, tool use, and scalable agentic applications while retaining Flash-tier latency and efficiency.
Tool CallingVisionAudio InputSpeech To Text
Google: Gemini 3.7 Flash
openrouter
Gemini 3.7 Flash is a frontier-class Google Flash model focused on high-performance reasoning, coding, multimodal understanding, tool use, and agentic execution with the latency and cost characteristics of the Flash family.
Tool CallingVisionAudio InputSpeech To Text
Google: Gemini 3.8 Flash
openrouter
Gemini 3.8 Flash is Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, complex enterprise workflows, advanced reasoning, and multimodal understanding. Building on Gemini 3.7 Flash, it delivers substantial improvements across coding, agentic tasks, and critical multi-step reasoning while retaining Flash-level speed and cost efficiency. It supports a 1M-token context window, configurable thinking levels, function calling, computer use, code execution, search grounding, structured outputs, file search, and URL context.
Tool CallingVisionAudio InputSpeech To Text
Google: Gemini 3.5 Flash
openrouter
Gemini 3.5 Flash is Google's high-efficiency multimodal model offering near-Pro-level coding and reasoning at Flash-tier speed and cost. It is optimized for coding, parallel agentic execution, multimodal understanding, and scalable production workloads.
Tool CallingVisionAudio InputSpeech To Text
Google: Gemini 2.5 Pro
openrouter
Gemini 2.5 Pro is Google's advanced Gemini 2.5 reasoning model for complex problem solving, coding, mathematics, scientific analysis, long-context understanding, multimodal reasoning, and sophisticated agentic workflows.
Tool CallingVisionAudio InputSpeech To Text
Google: Gemini 3.1 Pro
openrouter
Gemini 3.1 Pro is Google's advanced intelligence model for complex reasoning, problem solving, software engineering, multimodal analysis, long-context tasks, and powerful agentic and vibe-coding workflows.
Tool CallingVisionAudio InputSpeech To Text
Minimax: Speech 2.8 Turbo
minimax
MiniMax Speech 2.8 Turbo is the speed-optimized variant of MiniMax's flagship text-to-speech model. Delivering sub-250 millisecond latency, it provides fast, natural, and expressive voice synthesis ideal for real-time applications like conversational AI, voice agents, and interactive experiences. It retains the core advanced capabilities of the HD tier—including 40+ language support, rapid 10-second voice cloning, emotion control, and native sound tags—while prioritizing high-throughput generation and cost-efficiency.
Audio OutputText To Speech
Minimax: Speech 2.8 HD
minimax
MiniMax Speech 2.8 HD is the flagship studio-grade text-to-speech model from MiniMax. Built on an autoregressive Transformer architecture with a Flow-VAE decoder, it delivers highly expressive, broadcast-ready voice synthesis. It features native sound tags for realistic paralinguistic sounds (such as laughs, sighs, and breaths), robust emotion control, and high-fidelity voice cloning from just 10 seconds of audio. Supporting over 40 languages, it is optimized for professional audiobook narration, podcast production, and high-end video voiceovers.
Audio OutputText To Speech
Alibaba: Wan 3.0 Prime
openrouter
Wan 3.0 Prime is Alibaba's accelerated multimodal video generation model designed for high-quality video creation with significantly faster generation speed. It supports text-to-video, image-to-video with first-frame or first-and-last-frame conditioning, and reference-to-video generation using multimodal references including text, images, video, and audio. The model supports native audio generation, enhanced reasoning for complex prompts, multiple aspect ratios, resolutions up to 1080p, and video durations of up to 30 seconds.
VisionAudio InputVideo Generation
Alibaba: Wan 3.0 Video
openrouter
Wan 3.0 Video is Alibaba's high-speed AI video generation model from the Wan 3.0 family. It delivers the core capabilities of the standard Wan 3.0 Video model with significantly faster end-to-end generation speeds, making it ideal for high-volume API workflows and production. It can generate up to 30 seconds of synchronized audio and video in a single pass (at 480p, 720p, or 1080p resolution). The model supports extensive multimodal references, natively accepting text, images, audio, video.
VisionAudio InputVideo Generation
Ligtricks: LTX 2.3 Fast
lightricks
LTX-2.3 Fast is a speed-optimized variant of Lightricks' LTX-2.3 video generation model. It is designed for high-speed, cost-effective generation of synchronized audio and video in a single pass. It supports up to 20-second clips, native 9:16 portrait and 16:9 aspect ratios, and resolutions up to 4K. Delivering significantly faster inference than the Pro version, it is ideal for rapid prototyping, batch generation, and short-form social media content.
VisionAudio InputVideo Generation
Ligtricks: LTX 2.3 Pro
lightricks
LTX-2.3 is a DiT-based audio-video foundation model designed to generate synchronized video and audio within a single model. It brings together the core building blocks of modern video generation, with open weights and a focus on practical, local execution.
VisionAudio InputVideo Generation
Ligtricks: LTX 2.5 Fast
lightricks
LTX-2.5 Fast is an accelerated endpoint for Lightricks' LTX-2.5 video generation model. Designed for high-speed text-to-video and image-to-video workflows, it trades a degree of rendering fidelity for significantly faster inference speeds. It supports up to 20-second clips, 4K resolution output, and native synchronized audio generation, making it ideal for rapid prototyping and high-throughput production.
VisionAudio InputVideo Generation
Ligtricks: LTX 2.5 Pro
lightricks
LTX-2.5 Pro is the quality-optimized API endpoint of Lightricks' LTX-2.5 audio-video model. It utilizes Diffusion Fidelity Rendering to allocate maximum compute to complex scenes, ensuring high pixel quality, sharp faces, legible text, and seamless multi-shot continuity. It is designed for professional text-to-video and image-to-video workflows with synchronized native audio.
VisionAudio InputVideo Generation
MiniMax: MiniMax H3
minimax
MiniMax H3 is a general-purpose, omni-modal generative system that supports unified understanding of multimodal contexts composed of text, images, video, and audio. It can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. It is released with open weights under the MiniMax H3 Community License.
VisionAudio InputVideo Generation
Sarvam AI: Saaras V3
sarvam
Saaras v3 is a state-of-the-art speech recognition and translation model. It natively supports streaming for real-time applications and features specialized modes for transcription, direct-to-English translation, and transliteration. It is specifically tuned for noisy environments and complex code-mixed (e.g., Hindi-English) speech.
Audio InputSpeech To Text
Sarvam AI: Saarika V2.5
sarvam
Saarika v2.5 is Sarvam AI's legacy speech recognition model designed for Indian languages and accents. It transcribes audio in the same language spoken, excelling in multi-speaker conversations, telephony audio (8kHz), and code-mixed speech. Supports 11 languages (Hindi, Bengali, Gujarati, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, Telugu, English) with automatic language detection and speaker diarization. Achieves 4.96% CER and 18.32% WER on VISTAAR benchmark.
Audio InputSpeech To Text
XAI: Grok Imagine Video 1.5
grok
Grok Imagine Video 1.5 is xAI's advanced video model built on the Aurora-2 engine. Its standout feature is native one-pass audio generation, which produces synchronized dialogue (lip-sync), sound effects, and background music simultaneously with the video. Supporting resolutions up to 1080p and durations up to 15 seconds, it allows for complex workflows including text-to-video, image-to-video, video extension, and multi-image reference guidance to maintain consistent styles and characters.
VisionAudio OutputVideo GenerationText To Speech