ComfyUI is the only local image generator that runs on every common consumer GPU, has no telemetry, and ships a real node-based pipeline instead of a black box. The downside is the docs assume you already know where every file belongs. This page is the install path that works in October 2026, the model shortlist by VRAM, and the gaps that bite first-time users.
If you have not picked the card you are going to run it on, start with the VRAM hub or the GPU buying guide - this page assumes the hardware decision is already made.
The install path that works
There are three ways to install ComfyUI in 2026 and one of them is wrong for most readers. The project’s own download page lists them: a desktop app (“one-click installer with auto-updates”), a portable archive, and a manual CLI clone. All three pull the same engine.
Desktop app (best for first-time users). Cross-platform installer from comfy.org/download. Runs on Windows, macOS and Linux, auto-updates, and adds a system tray icon. Free forever according to the page, and it runs entirely offline after setup. The trade is that you do not control the underlying Python environment, so installing custom nodes that need specific PyTorch builds is harder.
Manual clone (best if you will install nodes). Clone the repo and run pip install -r requirements.txt, then python main.py to start. The README documents Python 3.13 as “very well supported” with 3.12 as the documented fallback, and 3.14 works for the core engine though some custom nodes lag. This is the path that survives when a node you depend on breaks after a desktop update.
Windows portable .7z (good for a locked-down box). Self-contained archive from the releases page. Extract with 7-Zip, double-click run_nvidia_gpu.bat (or the AMD/Intel equivalent), done. No Python install required because one is bundled. The catch: the bundled PyTorch is fixed, so you cannot mix it with system packages, and updating means re-downloading.
What you do not need: a subscription, an account, a working email address, or an internet connection after setup. ComfyUI is offline-capable by design. Any installer that asks you to log in before you can generate an image is not ComfyUI.
What fits your VRAM
The ComfyUI README claims it can run “the biggest open source models on as low as 4GB VRAM + 8GB RAM relatively quickly” using asynchronous weight streaming, and that figure is the floor you should plan around, not the one you should run production on. The honest picture is that the model determines the budget and the SD one is the ceiling.
| Model | Weights on disk (BF16/FP16) | Practical floor with quant or async offload | Notes |
|---|---|---|---|
SD 1.5 (runwayml/stable-diffusion-v1-5) | 3.4 GB | 4 GB VRAM | Reference baseline. safetensors total 859,520,964 params, OpenRAIL-M licence, ungated. |
SDXL base (stabilityai/stable-diffusion-xl-base-1.0) | 6.94 GB (sd_xl_base_1.0.safetensors) | 8 GB VRAM | Plus the SDXL VAE (stabilityai/sdxl-vae, MIT, 83.6M params, ~335 MB). safetensors total 2,567,463,684, OpenRAIL++ licence, ungated. |
FLUX.1-schnell (black-forest-labs/FLUX.1-schnell) | ~24 GB BF16 | 12 GB VRAM (FP8), 24 GB VRAM (BF16) | Apache-2.0, gated auto, “fast” distilled version. Use for the lowest-quality, highest-throughput path. |
FLUX.1-dev (black-forest-labs/FLUX.1-dev) | 23.8 GB BF16 (flux1-dev.safetensors) | 16 GB VRAM (FP8), 24-32 GB VRAM (BF16) | Custom flux-1-dev-non-commercial-license, gated auto. Commercial work needs either FLUX.1-schnell or a licensed variant. |
| FLUX.1-dev GGUF (city96) | 23.8 GB F16 down to ~4 GB at Q2_K | 8 GB VRAM at Q4_K_S or Q4_0 | Quantization ladder: F16, Q8_0, Q6_K, Q5_K_S, Q5_1, Q5_0, Q4_K_S, Q4_1, Q4_0, Q3_K_S, Q2_K. |
There is no single “right” answer here. SD 1.5 is the only model that runs comfortably at 6 GB and below, and it still produces real images for the kinds of prompts where diffusion-based generation is competitive. SDXL is the practical middle: it lives on 8-12 GB cards and the results are visibly better than SD 1.5 for faces and complex composition. FLUX is the ceiling: it needs 24 GB minimum for BF16 and the dev variant’s licence blocks commercial use, so think of it as the bench-test model, not the production one.
If the model does not fit at full precision, do not assume you need to buy more VRAM. The quantization article covers the bits-per-weight math; for image models the practical move is to load an FP8 or Q4 checkpoint and accept a small drop in colour depth and fine detail.
Where every model file goes
ComfyUI’s folder structure is the part most first-time installs get wrong, because the same model family uses different folders depending on which version you downloaded. The Flux examples page is the authoritative reference and the rule it documents is simple: the folder depends on whether the file is a full model or a self-contained checkpoint.
| Folder | What goes here |
|---|---|
models/checkpoints/ | Self-contained checkpoints that include the text encoder and VAE. SD 1.5 .safetensors, SDXL single-file, FLUX FP8 combined file (e.g. flux1-dev-fp8.safetensors, flux1-schnell-fp8.safetensors). |
models/diffusion_models/ | The diffusion transformer / UNet only. FLUX.1-dev full BF16 (flux1-dev.safetensors) goes here when paired with separate text encoders. |
models/unet/ | Legacy UNet-only folder. FLUX.1-schnell full (flux1-schnell.safetensors) historically drops here. |
models/text_encoders/ | Text encoder weights only. t5xxl_fp8_e4m3fn_scaled.safetensors and clip_l.safetensors for FLUX. The ComfyUI examples page notes the FP16 t5xxl is “recommended if you have more than 32 GB RAM”. |
models/vae/ | VAE files. ae.safetensors for FLUX, sdxl_vae.safetensors for SDXL. |
models/loras/, models/controlnet/, models/upscalers/, models/embeddings/ | LoRAs, ControlNet, upscaler models, textual inversions. |
If a model you downloaded does not show up in the node’s dropdown after you drop it in, you almost certainly put it in the wrong folder. There is no autodetect.
The catches that bite first
Licence asymmetry. SD 1.5 is OpenRAIL-M, SDXL is OpenRAIL++, FLUX.1-schnell is Apache-2.0, and FLUX.1-dev is the custom flux-1-dev-non-commercial-license. The first three permit commercial use with responsible-use conditions, the last one blocks it. There is no “FLUX licence” you can hand a client; you have to look at the specific model card.
Gating is real. FLUX.1-dev, FLUX.1-schnell and SDXL are all gated: auto or gated: manual on Hugging Face, which means the file download requires an account and a one-time access request before the weights will resolve. SD 1.5 is the only one that is fully ungated. If you cannot log in to HF, you cannot get FLUX.
Generated-images pipeline is heavier than the weights suggest. A diffusion model is the UNet or diffusion transformer, plus a text encoder, plus a VAE. The “VAE adds another 300 MB-2 GB” trap is what catches first-time FLUX users who budgeted for the 12 GB transformer only. The ComfyUI examples page recommends the FP8 t5xxl for anyone below 32 GB RAM precisely because the text encoder is the second-largest line item.
Custom nodes break the install. The single most common “it worked yesterday” failure on this site is a custom node that pinned an old PyTorch and a desktop update that pulled a new one. The manual clone path is more resilient because you control the venv; the desktop path is easier because you do not have to think about it. If your workflow depends on three or four community nodes, the manual path is the one that ages better.
Updates move fast. ComfyUI shipped v0.38.0 on 29 September 2026 (release entry on the GitHub releases page) and the release cadence is roughly every two to three weeks. Anything that says “ComfyUI install in 2025” is wrong about the Python version, the NVIDIA PyTorch channel (cu130 for RTX 20-series and later), and the AMD ROCm floor (ROCm 7.2 on Linux, ROCm 10.0 on Windows with multi-arch PyTorch).
Bottom line
If you want to generate images locally and you are not sure where to start, install the ComfyUI desktop app, drop an SD 1.5 checkpoint into models/checkpoints/, and run one of the example workflows from the Examples menu. That works on a 6 GB card and you can ship real output the same evening. Move to SDXL when 6 GB stops satisfying you, and only reach for FLUX if you have 24 GB of VRAM to spare and an actual use case that needs it.
If your plan is to install custom nodes, version your workflows, or run this on a server, skip the desktop app and do the manual git clone + pip install -r requirements.txt install instead. The desktop app’s value is the one-click setup; the manual path’s value is that nothing it updates can break under you.