A local vision-language model lets a reader drop a screenshot into a chat window without the frame ever leaving the machine, but only if the model is ungated, fits the card, and answers the question actually being asked. The open-weight cluster has shifted hard toward Qwen3-VL since the March 2026 tier page, and the split between Apache 2.0 and gated licences is the practical filter most readers trip on first. Every figure below was pulled from the Hugging Face model API or the live Ollama tag page on 2026-10-06.
Quick reference
| Job | Model | Ollama tag (size) | Licence | Gated | Context |
|---|---|---|---|---|---|
| Sub-2 GB on-device OCR/chat | Qwen3-VL-2B-Instruct | qwen3-vl:2b (1.9 GB) | Apache 2.0 | No | 256K (1M via YaRN) |
| 4 GB tier, OCR + short captioning | Qwen3-VL-4B-Instruct | qwen3-vl:4b (3.3 GB) | Apache 2.0 | No | 256K (1M via YaRN) |
| 8 GB sweet spot, document + UI | Qwen3-VL-8B-Instruct | qwen3-vl:8b (6.1 GB) | Apache 2.0 | No | 256K (1M via YaRN) |
| 12 GB tier, video understanding | Gemma 3 12B | gemma3:12b (8.1 GB) | Gemma | Yes | 128K |
| 16 GB tier, strong single image | Mistral Small 3.2 24B | mistral-small3.2:24b (15 GB) | Apache 2.0 | No | 128K |
| 24 GB tier, dense frontier | Qwen3-VL-32B-Instruct | qwen3-vl:32b (21 GB) | Apache 2.0 | No | 256K (1M via YaRN) |
| 24 GB tier, gated | Llama 3.2 Vision 11B | llama3.2-vision:11b (7.8 GB) | Llama 3.2 | Yes | 128K |
| 64 GB+ tier, MoE flagship | Qwen3-VL-235B-A22B-Instruct | qwen3-vl:235b (143 GB) | Apache 2.0 | No | 256K (1M via YaRN) |
| Frontier gated, single 90B card | Llama 3.2 Vision 90B | llama3.2-vision:90b (55 GB) | Llama 3.2 | Yes | 128K |
Two notes on the size column. The safetensors.total returned by the Hugging Face API for Qwen3-VL-32B is 33.4 GB at BF16; the 21 GB Ollama tag is a Q4_K_M quantisation of those same weights. The cluster convention is to publish Ollama tag sizes, since that is what readers pull, and to fall back to the HF safetensors total when no Ollama tag is available. The Gemma 3 numbers are pulled from the same Ollama library page on 2026-10-06.
Two jobs, two architectures
OCR and short captioning. The model receives a document, screenshot, or sign and returns the text or a one-line description. The hardest sub-task is reading dense tables and handwritten text; OCR benchmarks reward small, dense models that can move pixel data through a tight encoder without losing alignment. Qwen3-VL at 2B and 4B is the strongest open-weight pick at this band: ungated, Apache 2.0, small enough to load on a Raspberry Pi-class box with the right quantisation.
Visual reasoning over multi-image or video input. The model receives several images, a video, or a long document with figures and is asked to compare, locate, or summarise across them. The native context window matters: Qwen3-VL-8B and Qwen3-VL-32B both ship 256K native context extendable to 1M via YaRN, and the Qwen3-VL-32B card names “hours-long video with full recall and second-level indexing” as a documented capability. Gemma 3 12B and 27B ship a 128K context, enough for most document jobs but a hard cap on video length.
The RAG setups page covers pairing a vision model with a text embedder for retrieval over a personal screenshot library; the how-much-VRAM guide covers the runtime overhead the table above does not show.
Pick by tier
Under 2 GB. qwen3-vl:2b at 1.9 GB is the only mainstream open-weight VLM that fits on phones, Raspberry Pi 5, and older integrated-GPU laptops. The card is ungated and the Apache 2.0 licence removes the friction that Llama 3.2 Vision’s gate creates. Expect a clear step down on dense-document OCR versus the 4B and 8B variants.
3 to 4 GB. qwen3-vl:4b at 3.3 GB is the strongest sub-4 GB choice for screenshots, single-page documents, and short UI captures. The earlier qwen2.5vl:3b (3.2 GB) is still on Ollama and remains a usable fallback, but the Qwen3-VL card documents better long-context behaviour and the same Apache 2.0 terms.
6 to 9 GB. qwen3-vl:8b at 6.1 GB is the 8 GB sweet spot: ungated, Apache 2.0, 256K native context, and documented video understanding. gemma3:12b (8.1 GB) lands in the same VRAM band but is gated on Hugging Face; the Ollama tag resolves, but a reader loading from hf.co/google/gemma-3-12b-pt will hit a manual access request.
0 to 16 GB. Three routes diverge. mistral-small3.2:24b at 15 GB is the strongest single-card ungated pick above 12 GB: Apache 2.0, 128K context, and Mistral’s model card documents vision comparable to smaller Qwen3-VL variants. qwen2.5vl:32b at 21 GB is the previous-generation dense Qwen VLM and is still on Ollama. llama3.2-vision:11b at 7.8 GB is the gated Llama alternative; the Ollama tag page is the only frictionless path.
0 to 24 GB. qwen3-vl:32b at 21 GB is the dense frontier: Apache 2.0, ungated, 33B BF16 weights (21 GB Q4_K_M), 256K context, and the model card documents second-level video temporal grounding. The 30B-A3B MoE variant (qwen3-vl:30b at 20 GB) sits in the same VRAM band with a different activation pattern; the Ollama tag is live.
64 GB+. qwen3-vl:235b at 143 GB Q4_K_M is the open-weight MoE flagship: 235B total / 22B active, Apache 2.0, ungated, 256K native context. The card documents the same spatial and video capabilities as the 32B variant; the difference is recall on long-horizon video and the heavier VRAM (or disk + offload) footprint. Llama 3.2 Vision 90B at 55 GB is the gated competitor at this scale.
Licences, where the traps are
The vision-model family splits cleanly into two camps. The Apache 2.0 camp covers Qwen3-VL (all six sizes), Qwen2.5-VL (3B/7B/32B; the 72B card lists the older Qwen Research licence), Mistral Small 3.2 24B, Pixtral 12B, and LLaVA-OneVision Qwen2-7B. Each can be deployed commercially without contacting the rights holder.
The gated camp covers Llama 3.2 Vision (Llama 3.2 Community License) and Gemma 3 (Gemma Terms of Use, with separate restrictions for derivative works). Llama 3.2’s community license caps monthly active users at 700 million for the largest developers; Gemma’s terms add an explicit responsible-AI prohibited-use clause. Ollama serves both, but the underlying HF repo is gated, which means a fresh deployment that bypasses Ollama must complete the access form first.
Two specific cases to flag. The Qwen2.5-VL-72B-Instruct card lists the older Qwen Research licence rather than Apache 2.0; the open-model licences guide covers what that means in practice. Pixtral 12B is Apache 2.0 per the HF API but Ollama no longer hosts a current tag page; confirm via the Mistral Hugging Face hub that the 12B 2409 build is the version a deployment needs.
How to actually run one
Two flags matter most for vision models. The Qwen3-VL family uses dynamic resolution with configurable min/max pixel bounds (per the model config, with min_pixels 16,384 and max_pixels 16,777,216, so high-resolution images scale up and short screenshots stay small). The Gemma 3 family uses a fixed input resolution and processes images at 896x896 by default; Gemma 3 also requires the gated HF repo even when pulled via Ollama, because the Ollama docs document that the Ollama tag is a direct re-host of the gated files.
The GGUF, AWQ, GPTQ and MLX comparison covers which format to pull. The what-quantisation costs you page covers how quality loss shows up on vision tasks; it is more visible than on chat because small alignment errors compound across multi-image prompts. The household-serving guide covers the network plumbing for putting a local VLM behind Open WebUI on the home LAN.
The bottom line
For most consumer cards, qwen3-vl:8b (6.1 GB, Apache 2.0, ungated, 256K context) is the strongest verifiable pick at the 8 GB sweet spot; it handles document OCR, screenshot Q&A, and short video clips without the friction of a gated repo. On a 16 GB card, mistral-small3.2:24b (15 GB, Apache 2.0, 128K context) is the strongest ungated step up; on a 24 GB card, qwen3-vl:32b (21 GB, Apache 2.0, 256K context) is the dense frontier. The two gated picks worth knowing exist: gemma3:27b at 17 GB for readers who already have access, and llama3.2-vision:90b at 55 GB for readers who accept the Llama 3.2 licence. Above 64 GB, qwen3-vl:235b (143 GB, Apache 2.0, 235B/22B MoE) is the strongest open-weight pick for hours-long video and long-context document understanding.
Two traps to avoid. First, do not pull a vision model from a gated Hugging Face repo without confirming the access form has been accepted; Ollama can bypass the gate, but the underlying licence still applies on every deploy. Second, do not assume a closed-weight benchmark number carries over to an open-weight model at the same parameter count; the open-weight cards publish their own numbers, and the leading open-weight models page is the cluster index for what is currently shipping.