Video generation is the workload where the model card lies the most, and it lies in the same direction every time. The card gives you the parameter count of the diffusion model, the commercial-use page gives you a clean Apache 2.0, and the inference script then asks for a text encoder, a 3D VAE, and a temporal attention stack that together add two to three times what the card claimed. This guide gives the real totals, and it puts licensing first, because two of the strongest open video models are unusable outside Asia or above a monthly visit threshold.
The mistake people bring from image gen
A text-to-video model is not one file. It is a pipeline: a diffusion transformer, an LLM text encoder, a 3D VAE for video, and a temporal module that connects frames. The text encoder and the 3D VAE are the parts that ruin the budget, in roughly that order.
Wan2.2-TI2V-5B at 5B parameters sounds like an easy fit on a 24 GB card. The published
model card requires --t5_cpu to hit 24
GB; without it, the bundled UMT5-XXL encoder at bf16 lifts the total well past the card
size. HunyuanVideo at 13B parameters fits on 60 GB at 540p but the model card asks for 80
GB at 720p, and FP8 weights still weigh in at tens of GB.
The encoder offloads cheaply here in a way it does not for an LLM, because the encoder runs once per prompt rather than once per token. Most ComfyUI workflows already move the encoder to CPU after conditioning, so peak memory is closer to the diffusion model alone than to the sum of all parts.
Licences first, because this is where the traps are
The strongest open video models fall into three camps: permissive Apache 2.0, conditional Open Weights with revenue or visit caps, and custom community licences with territory or monthly-active-user clauses.
| Model | Params | Licence | Commercial use |
|---|---|---|---|
| Wan2.2-T2V-A14B | 27B total / 14B active | Apache 2.0 | Yes |
| Wan2.2-TI2V-5B | 5B | Apache 2.0 | Yes |
| LTX-Video 0.9.8 (13B dev / mix / distilled) | 13B | LTX-Video Open Weights 0.X | Yes, with a $10M revenue trigger |
| LTX-Video 0.9.8 (2B distilled) | 2B | LTX-Video Open Weights 0.X | Yes, with a $10M revenue trigger |
| LTX-Video 0.9.6 (2B distilled) | 2B | LTX-Video Open Weights 0.X | Yes, with a $10M revenue trigger |
| Mochi 1 preview | 10B + 362M VAE | Apache 2.0 | Yes |
| Allegro | 3B | Apache 2.0 | Yes |
| CogVideoX-5b | 6B | CogVideoX Community | Free under 1M visits/month |
| HunyuanVideo | 13B | tencent-hunyuan-community | Royalty-free, EX EU/UK/SK, MAU cap at 100M |
| Stable Video Diffusion img2vid | 2B | stable-video-diffusion-community | Conditional on Stability AUP |
Two patterns matter. CogVideoX-5b is free for commercial use only below 1 million visits per month, per the published LICENSE; above that you have to contact Zhipu for a separate licence. HunyuanVideo is royalty-free but its LICENSE excludes the EU, UK and South Korea entirely, and any product above 100M monthly active users has to ask Tencent for a separate licence granted at their discretion. LTX-Video is permissive for small operators but its Open Weights License requires a paid commercial licence once you cross $10M annual revenue, with double-fee liquidated damages for unauthorised commercial use as a large entity.
If the output is going anywhere near a commercial deliverable in the EU or UK, the shortlist collapses to Wan 2.2 (both variants), LTX-Video, Mochi 1 and Allegro.
By VRAM tier
The tiers below are based on each model’s own documentation for default settings at the model card’s native resolution and clip length. Offload flags shift numbers downward but at real speed cost; the numbers shown are the official figures, not the lowest-possible configurations.
8 GB. This tier is barely usable for text-to-video at any meaningful resolution. LTX-Video 2B v0.9.6 distilled is described on the model card as the smallest member of the family, “ideal for light VRAM usage,” and the same card claims “15x faster, real-time capable” at the distilled settings. CogVideoX-5b fits at INT8 in about 4.4 GB per the model card, at 720x480 and 6 seconds at 8 fps. Treat both as the entry point for hobby work, not for production.
12 GB. LTX-Video 2B distilled at the card’s recommended resolutions under 720x1280 fits, and the 13B distilled variant is reachable on 12 GB cards with aggressive offload. Stable Video Diffusion at 2B parameters and 576x1024 needs a card in the A100-80 GB class for the published inference time but can run on 12 GB with community patches.
16 GB. CogVideoX-5b at BF16 takes about 5 GB per the model card,
so 16 GB cards run it with headroom for the encoder. Allegro at 3B parameters needs
9.3 GB with CPU offloading and 27.5 GB
without, which is the cleanest 16 GB option when offload is acceptable. Wan2.2-TI2V-5B is
the right card-class target at 24 GB but on 16 GB it requires the same --t5_cpu and
offload flags the Wan2.2-TI2V-5B card
documents.
24 GB. This is the meaningful consumer tier. Wan2.2-TI2V-5B at 5B parameters runs on a single RTX 4090 at 720p, 24 fps, 5-second clips, in under 9 minutes, with the offload flags turned on per the model card. Mochi 1 at BF16 takes about 22 GB per the model card, or under 20 GB through the ComfyUI implementation. Wan2.2-T2V-A14B at 14B active is too large for 24 GB even with offload; it is a 48 GB or 80 GB class model.
32 GB and up. HunyuanVideo at 13B parameters needs a minimum of 60 GB at 720p and 45 GB at 540p, with 80 GB the recommended target, per the model card; the FP8 build saves about 10 GB. Wan2.2-T2V-A14B at 80 GB is the documented floor for 720p single-GPU inference. Wan2.2-TI2V-5B at 24 GB cards without the offload flags requires an 80 GB card to remove them.
What to actually run
If the work is commercial and global: Wan2.2-TI2V-5B. It is Apache 2.0, ungated, runs on a single 24 GB consumer card with offload, and supports both text-to-video and image-to-video in one model. Its 5B dense size is the smallest serious Apache 2.0 entry in the table.
If the deliverable is the highest quality a 24 GB card can produce: Mochi 1. Apache 2.0, ungated, with the ComfyUI implementation under 20 GB and 480p native output. Its ecosystem of finetunes on the hub is the largest of any Apache 2.0 video model listed here.
If you need image conditioning on a budget: CogVideoX-5b at 5 GB BF16, provided your product stays under 1 million visits per month. Above that cap, the licence forces a commercial agreement with Zhipu and the answer stops being local.
If you are in the EU or UK, do not use HunyuanVideo under any workflow. Its published LICENSE names the EU, UK and South Korea as excluded territory and authorises use outside that set only.
One thing this guide does not do is rank these models by output quality. No primary benchmark was verified for this update, and video quality comparisons are even more prompt-dependent and temporally-coherent than image ones. The ordering above is by what fits, what you are licensed to do, and what architecture the model is built for. Treat any confident quality ranking you read elsewhere as weaker evidence than rendering twenty clips from your own prompts and watching the motion.
For the rest of the local stack, the VRAM index covers chat, coding, vision and image, the image-gen VRAM guide explains the same text-encoder trap for diffusion image models, and which quantization to use is the reference for why the FP8 weights above are not actually eight bits per weight.