MiniMax: MiniMax Hailuo 02
minimax
Lowest priceMiniMax-Hailuo-02 is a state-of-the-art AI video generation model capable of producing high-fidelity 1080p video (up to 10 seconds) from text or images. It is recognized for exceptional, realistic physics simulations (fluid dynamics, collision) and strong character consistency. It supports both Text-to-Video (T2V) and Image-to-Video (I2V) capabilities.
VisionVideo Generation
MiniMax: MiniMax Hailuo 2.3
minimax
MiniMax-Hailuo-2.3 is a high-fidelity AI video generation model specializing in cinematic realism, superior character physics, and complex motion, available in standard and fast versions. It supports 768p or 1080p resolution, creates 6-10 second clips, and excels at text-to-video and image-to-video (I2V) tasks with high prompt adherence.
VisionVideo Generation
Alibaba: Wan2.2 T2V A14B
novita
Wan2.2-T2V-A14B is a 14B-parameter text-to-video diffusion model specializing in high-fidelity motion generation and temporal consistency. Features advanced motion dynamics modeling and cinematic quality rendering with 128-frame coherence support.
Video Generation
Alibaba: Wan2.2 I2V A14B
novita
Wan2.2-I2V-A14B is a 14B-parameter image-to-video diffusion model specializing in animating static images with realistic motion. Features advanced motion transfer, temporal consistency, and style preservation capabilities with 128-frame video generation from single images.
VisionVideo Generation
ByteDance: Seedance V1.5 Pro I2V
novita
Seedance V1.5 Pro I2V is an advanced image-to-video generation model that animates input images with superior temporal coherence, higher resolution, and improved motion quality.
Video Generation
ByteDance: Seedance 2.0 Mini
openrouter
Seedance 2.0 Mini is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video with image, video, and audio inputs. It...
VisionVideo Generation
Ligtricks: LTX 2.3 Fast
lightricks
LTX-2.3 Fast is a speed-optimized variant of Lightricks' LTX-2.3 video generation model. It is designed for high-speed, cost-effective generation of synchronized audio and video in a single pass. It supports up to 20-second clips, native 9:16 portrait and 16:9 aspect ratios, and resolutions up to 4K. Delivering significantly faster inference than the Pro version, it is ideal for rapid prototyping, batch generation, and short-form social media content.
VisionAudio InputVideo Generation
Google: Veo 3.1 Lite
openrouter
Veo 3.1 Lite is Google's cost-efficient video generation model designed for high-volume applications and rapid iteration. It generates video from text or image prompts with synchronized audio and supports cinematic controls, multiple aspect ratios, and short-form production workflows.
VisionVideo Generation
Ligtricks: LTX 2.3 Pro
lightricks
LTX-2.3 is a DiT-based audio-video foundation model designed to generate synchronized video and audio within a single model. It brings together the core building blocks of modern video generation, with open weights and a focus on practical, local execution.
VisionAudio InputVideo Generation
ByteDance: Seedance 2.0 Fast
openrouter
Seedance 2.0 Fast is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video. It prioritizes generation speed and lower cost...
VisionVideo Generation
Alibaba: Wan 3.0 Video
openrouter
Wan 3.0 Video is Alibaba's high-speed AI video generation model from the Wan 3.0 family. It delivers the core capabilities of the standard Wan 3.0 Video model with significantly faster end-to-end generation speeds, making it ideal for high-volume API workflows and production. It can generate up to 30 seconds of synchronized audio and video in a single pass (at 480p, 720p, or 1080p resolution). The model supports extensive multimodal references, natively accepting text, images, audio, video.
VisionAudio InputVideo Generation
XAI: Grok Imagine Video
grok
Grok Imagine Video is xAI's foundational multimodal video generation model capable of turning text prompts and static images into dynamic video clips. It supports high-quality visual motion, cinematic camera controls, and clip lengths of up to 10 seconds, establishing the baseline for xAI's video generation capabilities.
VisionVideo Generation
MiniMax: MiniMax H3 Max
minimax
MiniMax H3 Max is MiniMax's video generation model designed for text-to-video and image-to-video generation. It supports generating videos directly from text prompts as well as using first-frame and/or last-frame images to control the beginning and ending of generated sequences. The model produces videos at 480P or 768P resolution with configurable durations from 5 to 15 seconds and supports common or adaptive aspect ratios.
VisionVideo Generation
ByteDance: Seedance 2.0
openrouter
Seedance 2.0 is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video. It is particularly strong at preserving character consistency,...
VisionVideo Generation
Alibaba: Wan 3.0 Prime
openrouter
Wan 3.0 Prime is Alibaba's accelerated multimodal video generation model designed for high-quality video creation with significantly faster generation speed. It supports text-to-video, image-to-video with first-frame or first-and-last-frame conditioning, and reference-to-video generation using multimodal references including text, images, video, and audio. The model supports native audio generation, enhanced reasoning for complex prompts, multiple aspect ratios, resolutions up to 1080p, and video durations of up to 30 seconds.
VisionAudio InputVideo Generation
MiniMax: MiniMax H3
minimax
MiniMax H3 is a general-purpose, omni-modal generative system that supports unified understanding of multimodal contexts composed of text, images, video, and audio. It can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. It is released with open weights under the MiniMax H3 Community License.
VisionAudio InputVideo Generation
Google: Veo 3.1 Fast
openrouter
Veo 3.1 Fast is Google's speed-optimized video generation model balancing generation quality, latency, and cost. It creates high-quality video from text or image prompts with native synchronized audio and supports first-frame and last-frame conditioning.
VisionVideo Generation
Kuaishou: Kling v3.0 Standard T2V
novita
Kling v3.0 Standard T2V is Kuaishou's text-to-video generation model designed to create cinematic video directly from natural-language prompts. It supports smooth motion, strong prompt adherence, multi-shot composition, negative prompting, configurable aspect ratios, flexible video durations, and optional native audio generation for synchronized audiovisual content.
Video Generation
Kuaishou: Kling v3.0 Standard I2V
novita
Kling v3.0 Standard I2V is Kuaishou's image-to-video generation model designed to animate still images with realistic motion and cinematic scene dynamics. It supports prompt-guided animation, first-frame and optional end-frame conditioning, configurable generation strength, negative prompting, multi-shot composition, flexible video durations, and optional native audio generation.
VisionVideo Generation
Ligtricks: LTX 2.5 Fast
lightricks
LTX-2.5 Fast is an accelerated endpoint for Lightricks' LTX-2.5 video generation model. Designed for high-speed text-to-video and image-to-video workflows, it trades a degree of rendering fidelity for significantly faster inference speeds. It supports up to 20-second clips, 4K resolution output, and native synchronized audio generation, making it ideal for rapid prototyping and high-throughput production.
VisionAudio InputVideo Generation
Alibaba: Wan 2.6 T2V
novita
Wan 2.6 T2V is Alibaba's text-to-video generation model designed to create high-quality cinematic video directly from natural-language prompts. It supports complex scene descriptions, realistic motion and physical dynamics, single-shot and multi-shot generation, prompt enhancement, negative prompting, multiple aspect ratios, high-resolution video generation, and optional native audio for synchronized audiovisual output.
Video Generation
Alibaba: Wan 2.6 V2V
novita
Wan 2.6 V2V is Alibaba's video-to-video generation model designed to transform reference videos while preserving important character, motion, and structural information. It supports prompt-guided character replacement, role-playing, visual style and texture transformation, multiple reference videos, single-shot and multi-shot generation, prompt enhancement, negative prompting, high-resolution output, and synchronized audio-video generation.
Video Generation
ByteDance: Seedance 2.5
openrouter
Seedance 2.5 is a video generation model from ByteDance. It is suited for long-form storytelling, multimodal reference-based generation, video editing, and video extension. It supports first-frame and first-and-last-frame control, up...
VisionVideo Generation
Kuaishou: Kling v3.0 Pro T2V
novita
Kling v3.0 Pro T2V is Kuaishou's professional-grade text-to-video model optimized for high-fidelity cinematic generation from natural-language prompts. It provides enhanced motion quality, prompt adherence, subject consistency, multi-shot storytelling, complex scene generation, negative prompting, flexible video durations, and optional native audio generation for professional audiovisual workflows.
Video Generation
Kuaishou: Kling v3.0 Pro I2V
novita
Kling v3.0 Pro I2V is Kuaishou's professional-grade image-to-video model for transforming still images into high-fidelity cinematic video. It offers enhanced motion generation and subject consistency with prompt-guided animation, first-frame and optional end-frame control, multi-shot composition, negative prompting, configurable generation strength, flexible durations, and optional native audio generation.
VisionVideo Generation
Ligtricks: LTX 2.5 Pro
lightricks
LTX-2.5 Pro is the quality-optimized API endpoint of Lightricks' LTX-2.5 audio-video model. It utilizes Diffusion Fidelity Rendering to allocate maximum compute to complex scenes, ensuring high pixel quality, sharp faces, legible text, and seamless multi-shot continuity. It is designed for professional text-to-video and image-to-video workflows with synchronized native audio.
VisionAudio InputVideo Generation
Google: Veo 3.1
openrouter
Veo 3.1 is Google's state-of-the-art cinematic video generation model built for maximum visual fidelity. It generates high-quality video from text or image prompts with native synchronized dialogue, ambient effects, and sound, while supporting advanced cinematic controls and production workflows.
VisionVideo Generation
Kuaishou: Kling V3.0 4K T2V
novita
Kling v3.0 4K is Kuaishou's premium text-to-video model, generating native 4K cinematic videos from text prompts. It supports 3-15 second durations, flexible aspect ratios, optional synchronized audio, and multi-prompt scene transitions for complex compositions.
Video Generation
Kuaishou: Kling V3.0 4K I2V
novita
Kling v3.0 4K is Kuaishou's premium image-to-video model, generating up to 15 seconds of native 4K video (30fps) from images. It supports optional synchronized audio, multi-prompt scene composition for complex narratives, and flexible aspect ratios (16:9, 9:16, 1:1).
VisionVideo Generation
MiniMax: MiniMax Hailuo 2.3 Fast
minimax
MiniMax Hailuo 2.3 Fast is an image-to-video generation model optimized for low-latency and computational efficiency. Generates 6-10 second videos at 768p (up to 1080p) with 30-50% faster generation than standard model. Ideal for rapid iteration, social media content, and bulk video production.
VisionVideo Generation
ShengShu AI: Vidu Q1 Text2Video
novita
Vidu-Q1-Text2Video is an 18B-parameter diffusion model specializing in high-quality text-to-video generation with exceptional motion dynamics and temporal coherence. Features advanced physics-aware motion modeling and cinematic-quality rendering with extended context understanding.
Video Generation
ShengShu AI: Vidu Q1 Img2Video
novita
Vidu-Q1-Img2Video is a 20B-parameter diffusion model specialized in transforming static images into dynamic, high-quality video sequences. Features advanced style preservation, motion transfer, and temporal coherence with exceptional input image fidelity and realistic motion synthesis.
VisionVideo Generation
ByteDance: Seedance V1 Lite T2V
novita
Seedance V1 Lite is a lightweight text-to-video generation model capable of producing short video clips from textual prompts.
Video Generation
ByteDance: Seedance V1.5 Pro T2V
novita
Seedance V1.5 Pro T2V is an advanced text-to-video generation model featuring improved temporal consistency, higher resolution output, and enhanced prompt adherence.
Video Generation
XAI: Grok Imagine Video 1.5
grok
Grok Imagine Video 1.5 is xAI's advanced video model built on the Aurora-2 engine. Its standout feature is native one-pass audio generation, which produces synchronized dialogue (lip-sync), sound effects, and background music simultaneously with the video. Supporting resolutions up to 1080p and durations up to 15 seconds, it allows for complex workflows including text-to-video, image-to-video, video extension, and multi-image reference guidance to maintain consistent styles and characters.
VisionAudio OutputVideo GenerationText To Speech