Best Local Models by Context Window (September 2026)

Local LLMs ranked by maximum context length: 32K, 128K, 256K, and 1M+ tiers, with the model that fits each tier and the VRAM each one needs.

Updated September 29, 2026

The cluster of tier pages that drives most of this site’s traffic is organised by VRAM. That is the right axis for picking a model when the question is “what fits my GPU” - but the second question readers ask is “what fits my task”, and for long documents, large code bases, multi-turn agent loops, and PDF RAG, the binding constraint is context window, not VRAM. Two models with the same file size can sit at opposite ends of the context axis, and the 256K-tier one is the only one that will hold the 200-page deposition.

This page is a tier list of self-hostable open-weight models organised by the maximum context window they actually ship with. Every figure is read from the Hugging Face config.json or the model card on the day of publication, and every Ollama tag size from the library page. No benchmarks, no first-person claims about what runs on what hardware; the methodology is the same as the rest of the cluster - read the primary, quote the primary, link the primary.

How context windows are reported

Three numbers get conflated in model marketing copy, and they are not the same number:

  1. Native context. The max_position_embeddings field in config.json is what the model was trained to handle without any RoPE extension. Training-time attention quality is what is best at this length.
  2. Extended context via RoPE scaling. Most models can run past their native length with a rope-scaling method such as YaRN, NTK-aware scaling, or self-extend. The HF model card usually quotes the headline number, which can be 4x to 10x the native figure, but recall quality drops as you push past native.
  3. Ollama tag context window. Ollama sets a default context length per tag in the Modelfile. Some tags ship at native, others at the extended headline figure, and the q4/q8/fp16 variants of the same model can ship at different defaults. Always check the tag page, not just the upstream card.

The numbers in the table below are the Ollama tag context window because that is what readers will see in their Ollama show output and what ollama run will actually allocate. Where the upstream native and extended figures diverge from that, both are noted.

The tier table

Context tierModel (Ollama tag)Ollama sizeNative / extended (HF)ParamsLicenceNotes
32Kqwen3:4b (q4_K_M)2.6 GB32,768 native / 131,072 YaRN4.0BApache-2.0Smallest tier, fits 8 GB VRAM
32Kmistral-small:24b14 GB131,072 native24BApache-2.0Ollama default ships at 32K, upstream at 131K
32Kqwen3:32b20 GB32,768 native / 131,072 YaRN32.8BApache-2.0Ollama default ships at 40K
128Kgpt-oss:20b14 GB131,072 native21B / 3.6B activeApache-2.0Reasoning toggled, MoE
128Kgpt-oss:120b65 GB131,072 native117B / 5.1B activeApache-2.0Same context,fits 80 GB H100 (MXFP4)
128Kmistral-small:22b13 GB131,072 native22BApache-2.0Older Mistral-small; same context as 24B upstream
128Kllama3.3:70b-instruct-q4_K_M~43 GB128K native70Bllama3.3 (community)Default ships at 128K
256Kqwen3:4b (default)2.5 GB32,768 native / 131,072 YaRN4.0BApache-2.0Ollama default is 256K, not 40K - silent override
256Kqwen3:30b-a3b19 GB262,144 native / 1,010,000 YaRN30B / 3B activeApache-2.0256K is also the default for the instruct/thinking 2507 tags
256Kqwen3:30b-a3b-instruct-2507-q8_032 GB262,144 native30B / 3B activeApache-2.0Same context in higher quant
256Kqwen3.5:4b3.4 GB262,144 native / 1,010,000 YaRN4BApache-2.0Multimodal
256Kqwen3.5:9b6.6 GB262,144 native / 1,010,000 YaRN9BApache-2.0Multimodal
256Kqwen3.5:35b24 GB262,144 native / 1,010,000 YaRN35B / 3B activeApache-2.0MoE
256Kqwen3.5:122b81 GB262,144 native / 1,010,000 YaRN122B / 10B activeApache-2.0Largest 256K-tier model that still fits one node
256Kqwen3-coder-next (q4_K_M)52 GB262,144 native80B / 3B activeApache-2.0Coding-specific, 256K context built-in
256Kgemma-4-12b-it(q4_K_M ~7 GB est)256K native11.95BApache-2.0Native 256K with no YaRN config
1Mllama4:128x17b245 GB1M native402B / 17B active4-communityMaverick-class; needs ~256 GB unified or multi-GPU
10Mllama4:latest (Scout)67 GB10M native109B / 17B activellama4 (community)Industry outlier; recall quality degrades past the 1M-2M range

Two facts from this table are the ones readers tend to get wrong. First, qwen3:4b ships at 256K by default in Ollama even though the model card says 32K native / 131K YaRN - the Ollama tag page lists the Modelfile default as 256K. The q4_K_M and q8_0 tags of the same model ship at 40K. The default wins unless you override --ctx-size. Second, the 256K tier is dominated by Qwen models: Qwen3-Next-80B-A3B-Instruct and the Qwen3.5 family all sit at 262,144 native with 1,010,000 via YaRN, Qwen3-235B-A22B-Instruct-2507 sits at 262,144 native with 1,010,000 via Dual Chunk Attention and MInference (not YaRN), and Qwen3-Coder-Next ships at 262,144 native with no documented extension. Gemma 4 12B matches them at 256K native without the YaRN escape hatch.

What “needs the context” actually costs

The context window is the floor on KV-cache memory, on top of the model weights and the activation memory. The arithmetic is 2 x num_hidden_layers x num_key_value_heads x head_dim x bytes_per_element per token. For Qwen3.5-9B (32 layers, 4 KV heads, head_dim 256, bf16 KV), one token of KV cache is 2 x 32 x 4 x 256 x 2 = 131,072 bytes (~0.13 MB). At 256K tokens that is ~33 GB of KV cache alone, which is why the cluster’s how-much-vram-does-a-local-llm-actually-need page pairs naturally with this one - the 256K tier has the same VRAM cost as running a much larger model at 32K.

Three practical mitigations are available in current runners. KV-cache quantization (--cache-type-k q8_0 / --cache-type-v q4_0 in llama.cpp) cuts that 33 GB roughly in half at minor recall cost. Sliding window or chunked attention keeps only the last N tokens hot, which is what the 32K models effectively do at long context. Prompt caching in vLLM and llama.cpp server lets you keep a long static prefix in memory and only pay full KV cost for the new suffix - this is the right answer for a RAG setup that reuses the same document across many questions.

The combination that matters most: a 256K-tier model on a single 24 GB GPU with KV-cache quantization will usually outperform a 32K-tier model on the same GPU for any task that needs more than ~16K tokens of context, because the larger model’s recall quality at native length beats the smaller model’s recall at extension.

Picking a tier

A 32K tier is the right answer for chat, single-document Q&A under ~50 pages, and most tool-calling agents where the working set is the function schemas plus recent turns. A 128K tier covers most RAG over a single long document, multi-file code completion, and the kind of “explain this codebase” prompt that the cluster’s own best-local-coding-models-by-vram page recommends 32K models for - if you want the long context, you need to bump to 128K. A 256K tier is the floor for agent loops that accumulate long histories, multi-document RAG, and any task that genuinely needs the headline window. A 1M-or-more tier exists almost entirely because Llama 4 Scout ships it; the recall cost past 1M-2M is steep and the VRAM cost is at least double the 256K tier for the same architecture, so reach for it only when no 256K model fits the corpus.

The catch that does not appear in the tier table: the headline context window is a length, not a guarantee. Models that ship with YaRN to 1M tokens will produce coherent output at 1M, but the retrieval accuracy across that span degrades, and the KV-cache memory cost is the same whether you use 200K or 1M tokens. Read the upstream model card for any benchmark that quotes needle-in-a-haystack recall at the headline length before betting a workflow on the upper bound.

Bottom line

For self-hosted open-weight models on 2026-09-29, the practical tiers are 32K (Qwen3-4B and friends, fits 8 GB VRAM), 128K (gpt-oss-20b is the cheapest pick at 14 GB; gpt-oss-120b is the highest-quality Apache-2.0 pick at 65 GB; Llama 3.3 70B is the alternative if you accept the community licence), 256K (Qwen3.5 4B/9B/35B/122B and Qwen3-Next-80B-A3B-Instruct, all Apache-2.0 and not gated), and 1M-to-10M (Llama 4 Scout and Maverick, both behind the Llama 4 community licence and both VRAM-expensive). The 256K tier is where Qwen’s 1,010,000-token YaRN extension actually matters, and the 1M+ tier is where Llama 4 is alone. The methodology page on VRAM arithmetic is at how-much-vram-does-a-local-llm-actually-need, the methodology page on context-window mechanics is at how-much-context-can-a-local-llm-actually-hold, and the per-use-case tier lists are at best-local-coding-models-by-vram and friends.