Best Local LLM for Long-Document Q&A: How to Chat with a Book

Which local LLM should you use to ask questions about a 500-page PDF without uploading it? A practical guide by context window and hardware tier.

“Long-document Q&A” is the use case where you point an LLM at a 200-page PDF, a 500-page book, or a pile of case files, and ask questions that require the model to remember what was on page 47 when it answers about page 412. It is the workload that exposes the difference between a model’s advertised context window and the context your hardware can actually hold, and it is the one that makes a hosted AI assistant a privacy risk by default - because the assistant only answers if it gets to read the document first. This guide walks through which open-weight model to pick by hardware tier, what a long context window costs in memory, and how to wire the model up so the document never leaves your machine.

TL;DR

  • The five properties that matter for long-document Q&A are: native context window, effective long-context behaviour, instruction following, retrieval support, and licence. Then check how much of that context window your VRAM can hold.
  • The advertised window is not the window your GPU holds. The KV cache grows with every token of context. For qwen3:30b-a3b (19GB download, 262,144 tokens native per the Qwen3-30B-A3B card, Apache-2.0) the full window needs about 25.8 GB (24.0 GiB) of 16-bit KV cache on top of the weights, so a 24 GiB card holds only a fraction of it. gpt-oss:20b (14GB, 128K context) needs about 3.2 GB (3.0 GiB) for its full window and is the model on this list that holds its whole window on a 24 GiB card. Arithmetic below.
  • Workstations: gpt-oss:120b at 65GB with 128K context fits a 96 GB-class machine. qwen3:235b-a22b at 142GB needs more than that; its card documents 262,144 tokens natively, extendable to 1,010,000 tokens with Dual Chunk Attention and MInference on vLLM or SGLang, with an example launch command using 8-way tensor parallelism.
  • Do not paste a long document into a short-context model and hope. Mistral Small 24B (32K context) and the 40K-context qwen3 tags (Ollama library page) need retrieval-augmented generation (RAG) for long documents. Open WebUI’s RAG docs say that Ollama picks a default context from GPU VRAM, that on GPUs with less than 24 GiB this is 4096 tokens, and that this “severely limits” RAG; they recommend raising it to “8192+ (or rather, more than 16000) tokens”.
  • Context window and retrieval are complementary, not substitutes. A long-context model can answer “what does section 3.2 say” without a vector store, but the full document has to sit in the KV cache. RAG keeps the prompt, and the memory bill, small; a long-context model keeps the retrieval layer simple.
  • Licences differ. Apache-2.0 and MIT weights can be used commercially and redistributed; Llama 3.3 comes under Meta’s own community licence. If the documents are confidential - patient files, source code, internal reports - the bigger point is that a local model keeps them on the machine.

What “long-document Q&A” actually requires

Five capabilities separate a model that can read a book from one that just generates confident prose about an empty prompt.

  1. Native context window - how many tokens the model was trained to handle in one pass. The Ollama qwen3 library page lists the 0.6B, 1.7B, 8B, 14B and 32B tags at 40K context and the 4B, 30B and 235B tags at 256K. The Mistral Small 24B card states a 32k context window. Phi-3.5-mini supports 128K. The ceiling is the model’s max_position_embeddings; going beyond it requires RoPE extension or a different attention pattern.
  2. Effective long-context behaviour - whether the model actually retrieves from deep inside a long prompt. The Phi-3.5-mini card reports a RULER average of 84.1 across 4K to 128K, with the score falling from 94.3 at 4K to 63.6 at 128K. The Qwen3-30B-A3B-Instruct-2507 card reports RULER accuracy from 4K to 1M tokens at 260 samples per length. Scores fall at the high end; check the model card before trusting the headline figure. These are vendor-reported results.
  3. Instruction following on long inputs - whether the model still answers the question once the document is in the way.
  4. Retrieval support - whether the model can be paired with a vector store and chunker. RAG is the practical way to handle documents longer than the context window. Open WebUI’s RAG docs describe a togglable hybrid search option (BM25 with CrossEncoder re-ranking and configurable relevance score thresholds), configurable chunk size and overlap, and a “Using Entire Document” toggle that bypasses retrieval and injects the full file.
  5. Licence - whether the weights can be used for the kind of work being asked. Apache-2.0 (Qwen3, Qwen2.5, Mistral Small, gpt-oss) and MIT (GLM-4.6, Phi-3.5) are permissive. Llama 3.3 needs Meta’s licence, and Kimi K2 ships under a Modified MIT License.

The 2026 open-weight field for long documents

Seven families cover the realistic long-context picks. They differ more on context window, memory cost and licence than on raw quality at the same parameter count.

Qwen3 (Apache-2.0, Alibaba)

The Qwen3 library page ships tags from 0.6B to 235B, with the 4B, 30B and 235B tags at a 256K context window and the rest at 40K. The Qwen3-30B-A3B-Instruct-2507 card states 262,144 tokens natively and Apache-2.0, and reports RULER evaluation across 4K to 1M tokens at 260 samples per length. The Qwen3-235B-A22B-Instruct-2507 card states 262,144 tokens natively, extendable to 1,010,000 tokens by combining Dual Chunk Attention (a length extrapolation method) and MInference (a sparse attention mechanism), served through vLLM or SGLang. Its vendor-reported RULER accuracy (full attention) runs from 98.5 at 4K to 84.5 at 1000K.

For long-document Q&A the relevant Qwen3 tags are qwen3:4b (2.5GB, 256K context) for the smallest path, qwen3:30b-a3b (19GB, 256K context, MoE with 3B active) for the 24 GiB tier, and qwen3:235b-a22b (142GB, 256K context, MoE with 22B active) for multi-GPU and large unified-memory machines.

Mistral Small (Apache-2.0, Mistral AI)

mistral-small:24b is 14GB on Ollama with a 32K context window, Apache-2.0 per the Mistral Small 24B card. It suits documents that fit comfortably inside one prompt. Beyond 32K it needs RAG: the Ollama page lists no longer-context variant of the 24B.

GLM-4.6 (MIT, Zhipu AI / Z.ai)

The GLM-4.6 card says the context window “has been expanded from 128K to 200K tokens, enabling the model to handle more complex agentic tasks.” It does not list RULER numbers. The weights are MIT-licensed and total about 357 billion parameters (Hugging Face API: 356,785,898,816), so this is a server-class model, not a consumer-GPU one.

Llama 3.3 (Llama 3.3 Community License, Meta)

Llama-3.3-70B-Instruct has a 128k context window and uses Grouped-Query Attention (GQA) “for improved inference scalability.” At 70B parameters it is workstation-class. The licence allows commercial use but requires a separate licence from Meta for licensees with more than 700 million monthly active users.

Phi-3.5 (MIT, Microsoft Research)

Phi-3.5-mini-instruct is a 3.8B-parameter dense model with a 128K context window. The card reports a RULER average of 84.1 across 4K to 128K, and a 26.1 average across GovReport, QMSum, Qasper, SQuALITY and SummScreenFD. MIT licence. It has no grouped-query attention (32 KV heads in its config), so its KV cache is large for its size: see the arithmetic below.

GPT-OSS (Apache-2.0, OpenAI)

The gpt-oss family ships at 20B (14GB) and 120B (65GB), both with 128K context, under Apache 2.0. Half of each model’s layers use a 128-token sliding window (20b config, 120b config), which keeps the KV cache small.

Kimi K2 (Modified MIT, Moonshot AI)

The Kimi-K2-Instruct-0905 card describes a mixture-of-experts model with 32 billion activated and 1 trillion total parameters, whose context window “has been increased from 128k to 256k tokens”. Ollama’s kimi-k2 page says the model “was retired on June 16, 2026” and lists no pulled tags. At 1 trillion parameters it is outside consumer and workstation hardware; it is listed here for completeness.

How much context your VRAM actually holds

The KV cache stores a key and a value vector for every token, in every attention layer. At 16-bit precision the cost per token is 2 (key and value) x layers x KV heads x head dimension x 2 bytes, using each model’s config.json. Figures below exclude compute buffers and runtime overhead; an 8-bit KV cache (llama.cpp’s q8_0, about 8.5 bits per value) roughly halves them.

ModelDownload (Ollama)KV bytes per tokenKV at full window
qwen3:4b2.5GB147,456 (config: 36 x 8 x 128)38.65 GB (36.00 GiB) at 262,144 tokens
phi3.5:3.8b2.2GB393,216 (32 x 32 x 96)51.54 GB (48.00 GiB) at 131,072 tokens
qwen3:8b5.2GB147,456 (config: 36 x 8 x 128)6.04 GB (5.62 GiB) at 40,960 tokens
qwen3:30b-a3b19GB98,304 (config: 48 x 4 x 128)25.77 GB (24.00 GiB) at 262,144 tokens
gpt-oss:20b14GB24,576 (12 full-attention layers x 8 x 64)3.22 GB (3.00 GiB) at 131,072 tokens
gpt-oss:120b65GB36,864 (18 full-attention layers x 8 x 64)4.83 GB (4.50 GiB) at 131,072 tokens
qwen3:235b-a22b142GB192,512 (config: 94 x 4 x 128)50.47 GB (47.00 GiB) at 262,144 tokens

The practical reading: on small cards the weights are cheap and the context is expensive. qwen3:4b at 32,768 tokens needs 4.83 GB (4.50 GiB) of KV cache; phi3.5:3.8b at 16,384 tokens already needs 6.44 GB (6.00 GiB).

Pick by hardware tier

8 GB usable VRAM (laptops, small GPUs)

qwen3:4b (2.5GB, Apache-2.0) is the practical pick: at 16-bit KV cache it holds roughly 32K tokens of context in 8 GB, far short of its 256K window. phi3.5:3.8b (2.2GB, MIT) holds less, roughly 12K to 14K tokens, because it lacks grouped-query attention. For anything longer, use RAG. Open WebUI’s RAG docs recommend a model context of “8192+ (or rather, more than 16000) tokens” so retrieved chunks, the question and the answer all fit.

16 GB usable VRAM

qwen3:8b (5.2GB, 40K context) is the practical default: its full 40,960-token window needs 6.04 GB (5.62 GiB) of KV cache, so weights plus cache fit with room to spare. Pair it with RAG for longer documents.

24 GiB usable VRAM (for example RTX 4090)

This is the practical sweet spot for a single user. Three real options:

  • gpt-oss:20b (14GB, 128K context, Apache-2.0) - the one model here that holds its whole window on this tier: 14GB of weights plus 3.22 GB (3.00 GiB) of KV cache at 131,072 tokens.
  • qwen3:30b-a3b (19GB, 256K context, Apache-2.0, 3B active MoE) - the stronger long-context scores on its card, but on this tier it holds only part of its window: 32,768 tokens of 16-bit KV cache is 3.22 GB (3.00 GiB), 65,536 tokens is 6.44 GB (6.00 GiB), which with the weights reaches the card’s 25.8 GB (24 GiB) limit. Use it with RAG, or with an 8-bit KV cache for roughly double the context.
  • mistral-small:24b (14GB, 32K context, Apache-2.0) - fits with its full window, but the 32K window forces RAG for any long document.

96 GB-class workstations and beyond

gpt-oss:120b (65GB, 128K context, Apache-2.0) fits with its full window: 65GB plus 4.83 GB (4.50 GiB) of KV cache. qwen3:235b-a22b (142GB) does not fit in 96 GB at all; with its full 262,144-token window it needs about 50.47 GB (47.00 GiB) of KV cache on top of the weights, so it belongs on multi-GPU servers or large unified-memory machines. The Qwen3-235B card documents the Dual Chunk Attention and MInference path to 1M tokens on vLLM or SGLang.

A practical RAG setup

For documents longer than the context your hardware holds, the practical pattern is the same regardless of which model you run on top. Open WebUI’s RAG docs document the moving parts:

  1. Embedding model. Changing the embedding model requires re-indexing, because embeddings from different models exist in different vector spaces.
  2. Chunk size and overlap. Both are configurable. The Open WebUI docs report that a well-configured Chunk Min Size Target (for example 1,000 on a chunk size of 2,000) can reduce chunk counts by over 90 percent, for documents split on Markdown headers.
  3. Retrieval method. Hybrid search (BM25 plus embeddings, with CrossEncoder re-ranking and configurable relevance score thresholds) is a toggle in Open WebUI, not a default.
  4. Model context. The docs recommend “8192+ (or rather, more than 16000) tokens” for Ollama models, and warn that the 4096-token default on GPUs with less than 24 GiB severely limits RAG, because retrieved data may not be used at all or only partially.
  5. Full-document toggle. The “Using Entire Document” toggle bypasses retrieval and injects the full file into every message. It is the simpler path when the document fits in the context your hardware holds.

For documents in the gap between “fits in context” and “needs RAG”, check the KV cache table above before choosing the long-context path: the window on the model card is an upper limit, not what a single card holds.

Where context extension still helps

The YaRN paper (arXiv:2309.00071) describes a RoPE-extension method that pushes context windows past the original training length. Its abstract says YaRN requires “10x less tokens and 2.5x less training steps than previous methods”. Qwen2.5-7B-Instruct is a worked example on its Hugging Face card: full context of 131,072 tokens with generation up to 8192, a shipped config.json set for 32,768 tokens, and a YaRN rope_scaling factor of 4.0 to go beyond that. The card warns that vLLM supports only static YaRN, “potentially impacting performance on shorter texts”, and advises adding the scaling only when long contexts are needed. The Qwen3-235B card goes further with Dual Chunk Attention and MInference to reach 1M tokens.

What to skip

Three things do not work well for long-document Q&A and are worth calling out:

  1. A short-context model with no RAG. Mistral Small 24B and the 40K-context Qwen3 tags fall into this trap. Anything beyond the window is truncated, so the model answers from only part of the document. Wire up RAG or pick a longer-context model.
  2. Trusting the model card’s window on a small card. A 256K window on the card does not mean 256K on an 8 GB laptop; see the KV cache table.
  3. Hosted AI for confidential documents. The point of running locally is that the document never leaves the machine. A hosted ChatGPT or Claude session processes the document on the vendor’s servers, under the vendor’s retention policy. Local does not solve every problem, but it solves this one cleanly.

The bottom line

For long-document Q&A on a single 24 GiB card, gpt-oss:20b holds its full 128K window, and qwen3:30b-a3b is the pick when paired with RAG or a quantised KV cache. On 8 GB laptops, qwen3:4b with RAG is the default. A 96 GB-class workstation runs gpt-oss:120b with its full 128K window; qwen3:235b-a22b and its 1M-token path need multi-GPU hardware. The licence (Apache-2.0 or MIT) matters if you build on the model; running locally matters if the documents cannot leave the building.