How much context can a local LLM actually hold? (September 2026)
The advertised context window is one number. What your GPU can actually serve is smaller. The KV cache math, RoPE scaling, and the practical ceiling.
Tag
The advertised context window is one number. What your GPU can actually serve is smaller. The KV cache math, RoPE scaling, and the practical ceiling.
Reasoning-capable open-weight models sized by VRAM. DeepSeek-R1, Qwen3 with thinking, QwQ, Phi-4 reasoning, gpt-oss, and the licence traps.
How to adapt an open-weight model on your own hardware. What LoRA rank and alpha mean, when QLoRA's 4-bit NF4 is worth it, and how to measure it.
The RTX 30 and 40 series cards that still clear current Ollama and llama.cpp floors, what VRAM tier each opens, and which variants are worth the used premium.
How much quality a smaller quant actually loses. Real MMLU and KL Divergence numbers across Q2_K through Q8_0, plus how to measure on your own workload.
The published GGUF file size is the weights. VRAM use also includes the KV cache, framework overhead, and your context length. The math, walked through.
Five stock GGUF quants for Qwen3.8-27B span 9.01 GB to 29.05 GB. A file size is not a VRAM requirement, and a 262,144-token ceiling is not free to fill.
VRAM decides what runs at all, bandwidth decides how fast it types. A buying guide by tier across NVIDIA, AMD, Intel, Apple and the used market.
Local embedding and reranking models for RAG, sized by VRAM. Real GGUF file sizes, licences, and the runtime gap that stops Ollama reranking.
Open image models sized for real GPUs, with the text encoder counted in and the licence checked. FLUX, Z-Image, Qwen-Image and SD3.5 compared.
Local AI on a 6GB GPU: GTX 1660, RTX 2060, RTX 3050 and laptop cards. Real weight sizes for chat, coding, vision, speech, translation and RAG.
Find your VRAM tier, then the current open-weight pick for chat, coding, vision, speech, translation or agents. Eighteen guides, one index.
Quantization formats compared with real file sizes and bits per weight. Why Q4 does not halve a model, and why low quants generate faster.
Local AI on a 12GB GPU: chat, coding, vision, speech and agents for RTX 3060 12GB or RTX 4070. Current picks, per-quant weight sizes, honest limits.
Local AI on a 16GB GPU: chat, coding, translation, speech and agents for RTX 4060 Ti, RTX 5060 or Arc A770, and why the vision tier stays unresolved.
Local AI on an 8GB GPU: chat, coding, vision, speech and agents for RTX 4060 or RTX 3070. Current picks, named quantisations, honest limits.
Local AI on a 24GB GPU: chat, coding, vision, speech and agents for RTX 3090 or RTX 4090. Current picks, per-quant weight sizes, and an open runtime bug.
Local AI on a 32GB GPU: chat, coding, vision, speech and agents on an RTX 5090. Current picks, per-quant weight sizes, and what the headroom buys.