Running Local LLMs on AMD and Intel GPUs in 2026
AMD's ROCm support splits by card and by OS, Intel archived both of its own LLM libraries this year, and llama.cpp is the path still standing.
Tag
AMD's ROCm support splits by card and by OS, Intel archived both of its own LLM libraries this year, and llama.cpp is the path still standing.
None of the mainstream inference servers ask for a password, and vLLM binds every interface by default. The safe pattern for household serving.
8-bit measures close to lossless. 4-bit ranges from a 0.6% gain to a 59% drop on one benchmark, depending on the model. The figures, fully attributed.
Open weights is not an open licence. Verified Hugging Face licence tags for 46 model repos, and the size and version traps that block shipping.
Qwen's 27B vision-language model is on Hugging Face ungated under Apache 2.0, and AMD says it needs roughly 24GB of VRAM to run comfortably.
Which open models run on CPU and RAM alone, how to size them, why mixture-of-experts helps, and what a machine with no discrete GPU cannot do.
What sits at the top of the field right now, hosted or downloadable, with every licence, price and index score re-read from a primary source today.
VRAM decides what runs at all, bandwidth decides how fast it types. A buying guide by tier across NVIDIA, AMD, Intel, Apple and the used market.
Ollama, LM Studio, Jan, llama.cpp and vLLM compared: licences, engines, OpenAI-compatible APIs, multi-user serving, and which to install first.
Six self-hosted RAG stacks compared: AnythingLLM, Open WebUI, LibreChat, Msty, RAGFlow, Cherry Studio. Licences, embedders, and who can rerank.
Sixteen downloadable-weight models ranked by capability and licence, with every parameter count and licence re-checked against the Hugging Face API.
Liquid AI's 2.6B model ships via Setapp for offline Mac agents, while iCloud Private Relay leaks real IPs through passkeys in the same 24 hours.
Local embedding and reranking models for RAG, sized by VRAM. Real GGUF file sizes, licences, and the runtime gap that stops Ollama reranking.
Open image models sized for real GPUs, with the text encoder counted in and the licence checked. FLUX, Z-Image, Qwen-Image and SD3.5 compared.
Anthropic's July 8, 2026 policy trains on consumer chats unless you opt out, and Gemini keeps human-reviewed chats for up to three years.
Local AI on a 6GB GPU: GTX 1660, RTX 2060, RTX 3050 and laptop cards. Real weight sizes for chat, coding, vision, speech, translation and RAG.
Find your VRAM tier, then the current open-weight pick for chat, coding, vision, speech, translation or agents. Eighteen guides, one index.
Apple caps its on-device model at 4096 tokens per session, and Google's own tables put Gemma 4 E2B at 25.0 decode tokens/sec on an iPhone 17 Pro CPU.
Apple Silicon has no discrete VRAM, so tier guides mislead Mac owners. The real ceilings are bandwidth and the GPU-usable slice of unified memory.
Quantization formats compared with real file sizes and bits per weight. Why Q4 does not halve a model, and why low quants generate faster.
Alibaba announced 2.4T-parameter Qwen3.8-Max open weights and a 27B sibling that Unsloth says will run in 17GB of RAM or VRAM.
Two back-to-back merges add MiniMax Sparse Attention and a Qwen2.5-VL style vision tower to llama.cpp, but every existing MiniMax-M3 GGUF must be regenerated.
A new industry letter asks Washington to protect open-weight AI while separating legitimate distillation from alleged theft of closed models.
Cisco released Antares-350M and Antares-1B open-weight models for vulnerability localization. They run locally and cost about 172x less than GPT-5.5.