6GB VRAM: What You Can Actually Run Locally (2026)
Local AI on a 6GB GPU: GTX 1660, RTX 2060, RTX 3050 and laptop cards. Real weight sizes for chat, coding, vision, speech, translation and RAG.
Tag
Local AI on a 6GB GPU: GTX 1660, RTX 2060, RTX 3050 and laptop cards. Real weight sizes for chat, coding, vision, speech, translation and RAG.
Find your VRAM tier, then the current open-weight pick for chat, coding, vision, speech, translation or agents. Fourteen guides, one index.
Alibaba announced 2.4T-parameter Qwen3.8-Max open weights and a 27B sibling that Unsloth says will run in 17GB of RAM or VRAM.
A Go-based botnet is scanning exposed Ollama, ComfyUI, n8n, Open WebUI, Langflow, and Gradio instances for AWS keys and Kubernetes tokens, QiAnXin XLab says.
VIDRAFT_LAB posts Ourbox-35B-JGOS to Hugging Face: 20 tok/s on an 8GB laptop GPU, ~17 tok/s on a CPU-only server, 86.4% on GPQA Diamond.
Ditch GitHub Copilot's $19/month subscription. Set up Continue.dev with Ollama for private, local AI code completion in VS Code — zero data leaves your machine.
Intruder scanned 2 million hosts and found 1 million exposed AI services with no authentication. Plus: teenagers are using ChatGPT to hack governments, and OpenAI launches Daybreak.
Ditch GitHub Copilot's $10/month subscription. Set up free, private AI code completion in VS Code using Continue.dev and Ollama — runs entirely on your hardware.
Chat with your own documents locally — no cloud, no subscriptions, no data leaving your machine. Step-by-step setup guide.
DeepSeek V4 Pro approaches frontier-level performance. Google, Mistral, and Alibaba ship under Apache 2.0. Ollama hits 52 million monthly downloads.
After the Perplexity class-action over leaked chats to Meta and Google, here's how to run a citation-grounded AI answer engine on your own hardware with Ollama and SearXNG.
Jensen Huang bets on inference chips, Ollama adds multimodal support, and DeepSeek V4 remains the most anticipated release that hasn't happened yet.
Which local models can actually use tools, call functions, and run multi-step workflows? Function-calling and TAU-bench picks from 8GB to 32GB VRAM.
Head-to-head comparison of local chat and assistant models from 8GB to 32GB VRAM. Current picks: Qwen3.5, Gemma 4, GPT-OSS, Qwen3.6, and GLM-4.7-Flash.
Which open-weight coding model to run locally? HumanEval and SWE-bench picks from 8GB to 32GB GPUs. Qwen2.5-Coder, Qwen3.6, Devstral, KAT-Coder.
TranslateGemma, NLLB-200, Aya Expanse and Qwen3.5 by VRAM tier, with the licence terms that decide whether you can ship what you run.
Local image analysis, OCR, and visual reasoning from 8GB to 32GB VRAM. Qwen3.5 replaces Qwen3-VL at most tiers, and 16GB stays unresolved.
Georgi Gerganov's team is now at Hugging Face, unifying the model hub with the inference engine that powers Ollama, LM Studio, and the entire local AI ecosystem.
Local AI on a 12GB GPU: chat, coding, vision, speech and agents for RTX 3060 12GB or RTX 4070. Current picks, per-quant weight sizes, honest limits.
Local AI on a 16GB GPU: chat, coding, translation, speech and agents for RTX 4060 Ti, RTX 5060 or Arc A770, and why the vision tier stays unresolved.
Local AI on a 24GB GPU: chat, coding, vision, speech and agents for RTX 3090 or RTX 4090. Current picks, per-quant weight sizes, and an open runtime bug.
Local AI on a 32GB GPU: chat, coding, vision, speech and agents on an RTX 5090. Current picks, per-quant weight sizes, and what the headroom buys.
Local AI on an 8GB GPU: chat, coding, vision, speech and agents for RTX 4060 or RTX 3070. Current picks, named quantisations, honest limits.
The new MacBook Pro with M5 Max can run large language models entirely on-device, keeping your AI interactions private and offline