Context
2K
Input
text, image, audio
Output
video
Tool calling
Not listed
About this model
LTX-2.3 is a DiT-based audio-video foundation model designed to generate synchronized video and audio within a single model. It brings together the core building blocks of modern video generation, with open weights and a focus on practical, local execution.
Input modalities
text
Accepted as model input
image
Accepted as model input
audio
Accepted as model input
Providers
Available routing options for this model through OneInfer.
Available
Video generation
$0.0400 / second
Video generation
$0.0800 / second
Routing
OneInfer optimized
Pricing
Current OneInfer pricing for this model.
| Usage | Price |
|---|---|
| Video generation1280x720 · 720p · 6–10s · 16:9 | $0.0400 / second |
| Video generation720x1280 · 720p · 6–10s · 9:16 | $0.0400 / second |
| Video generation1920x1080 · 1080p · 6–10s · 16:9 | $0.0800 / second |
| Video generation1080x1920 · 1080p · 6–10s · 9:16 | $0.0800 / second |
| Video generation2560x1440 · 1440p · 6–10s · 16:9 | $0.1600 / second |
| Video generation1440x2560 · 1440p · 6–10s · 9:16 | $0.1600 / second |
| Video generation3840x2160 · 4k · 6–10s · 16:9 | $0.3200 / second |
| Video generation2160x3840 · 4k · 6–10s · 9:16 | $0.3200 / second |
Performance
Published evaluation results associated with this model.
No benchmark data is listed for this model.
API example
curl -X POST https://api.oneinfer.ai/v1/ula/generate-video \
-H "Authorization: Bearer YOUR_JWT_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"provider": "lightricks",
"model": "Lightricks/LTX-2.3-Pro",
"prompt": "A red fox running through a snowy forest",
"resolution": "1280x720",
"aspect_ratio": "16:9",
"duration": 5,
"generate_audio": false,
"camera_fixed": false,
"service_tier": "default"
}'