LTX-2.5 Day-0 Support in ComfyUI
Diffusion Fidelity Rendering, native multi-shot generation, and a model built to run fast on consumer hardware — LTX’s open video model arrives in ComfyUI on day one.
LTX-2.5 is the newest version of LTX’s open video model. The line has been built on a consistent premise: powerful video generation should be open, fast, and something you can actually run on consumer hardware. Weights are downloadable, the raw model can be fine-tuned on your own data, and it runs fast on local GPUs.
LTX-2.5 is an improvement across the full generation stack rather than a single-stage upgrade. It comes with significant improvements including a new rendering approach, a new video decoder, a custom text encoder, a better distilled variant, a prompt enhancer, and a base checkpoint built for adaptation. Native 4K, synchronized audio and video, and frame rates up to 50fps carry over from LTX-2.3.
Not only that, the LTX team also delivered a small experimental duration head model that automatically sets the length of the generated video.
Variants
LTX-2.5 comes in two forms in ComfyUI: open weights you run yourself, and hosted models available through Partner Nodes.
Open weights
LTX-2.5 dev — the main model.
LTX-2.5 distilled — a smaller, faster variant. Distillation has been reworked to carry more quality, prompt adherence, and motion than previous distilled releases, which makes it viable for deployment where the full model isn’t economical.
Via Partner Nodes
LTX-2.5 (Fast) — the wider envelope. 2 to 20 seconds, 720p through 4K in landscape or portrait, and frame rates of 24, 25, 48, or 50. Clips over 10 seconds run at 720p or 1080p and 24 or 25fps.
LTX-2.5 (Pro) — 2 to 10 seconds at 720p or 1080p, with frame rates of 24, 25, or 50.
Model Highlights
Diffusion Fidelity Rendering. New in LTX-2.5, and the core change in this release. Rather than spending compute evenly across a scene, the model allocates it by complexity. Structure comes first: motion, composition, and framing are generated in an 8x temporally compressed latent space, alongside a set of high-fidelity keyframes — more keyframes for complex scenes, fewer for simple ones, within the available compute budget. A dedicated pixel-diffusion stage then renders the final video from the structure and keyframes together, carrying fine detail and texture across every frame. In practice that means textures, materials, intricate objects, and faces resolve with pixel-level precision, and busier shots automatically draw more rendering compute than static ones.
Diffusion Video Decoder. Replaces standard VAE decoding. Sharper faces, legible text, and fewer smears in fast motion.
Native multi-shot. One generation produces multiple connected shots, holding character, environment, lighting, voice, and style across the cuts. A sequence comes out of a single run instead of being assembled from separate generations that have to be matched afterward.
Custom Gemma 4 12B text encoder. Purpose-built. Retains multiple subjects, actions, lighting details, and camera direction across a long prompt instead of dropping clauses as the prompt gets more complex.
Dedicated prompt enhancer. A lightweight model that expands short prompts into detailed cinematic instructions, at near-zero additional compute.
Auto duration. The model reads the action described in the prompt and predicts the right clip length before diffusion starts. With the prompt enhancer, that’s two fewer parameters to tune and less app-side logic to maintain in a pipeline.
Examples
Getting Started
Update ComfyUI to the latest version 0.32.0, or open Comfy Cloud.
Download the LTX-2.5 weights and place them in your models directory.
Load the LTX-2.5 template from the Templates panel, or download the templates below.
Add your prompt, input images, and run.
As always, enjoy creating!
Model weights: 🤗 Lightricks/LTX-2.5

