Open Model Licences for Commercial Use (August 2026)
Apache-2.0, MIT, Qwen variants, Llama community terms: what the Hugging Face API actually returns and the per-repo traps that send you back to negotiate.
Tag
Apache-2.0, MIT, Qwen variants, Llama community terms: what the Hugging Face API actually returns and the per-repo traps that send you back to negotiate.
Which open models run on CPU and RAM alone, how to size them, why mixture-of-experts helps, and what a machine with no discrete GPU cannot do.
What sits at the top of the field right now, hosted or downloadable, with every licence, price and index score re-read from a primary source today.
Sixteen downloadable-weight models ranked by capability and licence, with every parameter count and licence re-checked against the Hugging Face API.
Local embedding and reranking models for RAG, sized by VRAM. Real GGUF file sizes, licences, and the runtime gap that stops Ollama reranking.
Open image models sized for real GPUs, with the text encoder counted in and the licence checked. FLUX, Z-Image, Qwen-Image and SD3.5 compared.
Find your VRAM tier, then the current open-weight pick for chat, coding, vision, speech, translation or agents. Eighteen guides, one index.
NVIDIA's Cosmos-H-Dreams runs at 160 fps on one RTX PRO 6000, with weights, code, dataset, and recipe open. Not a robot controller, NVIDIA warns.
Moonshot released the full Kimi K3 weights on July 27, 2026: 2.8T params, 1M context, MXFP4. Read the license before you plan a deployment.
Cisco released Antares-350M and Antares-1B open-weight models for vulnerability localization. They run locally and cost about 172x less than GPT-5.5.
Hugging Face disclosed a July 2026 breach run end-to-end by an autonomous agent. The defender was an open-weight model.
Two new open-weight models point in opposite directions: Bonsai brings a 27B model toward phones, while Inkling targets customization at data-center scale.