Local AI by VRAM: Which Models Fit Your GPU (August 2026)
Find your VRAM tier, then the current open-weight pick for chat, coding, vision, speech, translation or agents. Fourteen guides, one index.
Tag
Find your VRAM tier, then the current open-weight pick for chat, coding, vision, speech, translation or agents. Fourteen guides, one index.
Apple caps its on-device model at 4096 tokens per session, and Google's own tables put Gemma 4 E2B at 25.0 decode tokens/sec on an iPhone 17 Pro CPU.
Apple Silicon has no discrete VRAM, so tier guides mislead Mac owners. The real ceilings are bandwidth and the GPU-usable slice of unified memory.
Quantization formats compared with real file sizes and bits per weight. Why Q4 does not halve a model, and why low quants generate faster.
Two back-to-back merges add MiniMax Sparse Attention and a Qwen2.5-VL style vision tower to llama.cpp, but every existing MiniMax-M3 GGUF must be regenerated.
VIDRAFT_LAB posts Ourbox-35B-JGOS to Hugging Face: 20 tok/s on an 8GB laptop GPU, ~17 tok/s on a CPU-only server, 86.4% on GPQA Diamond.
Georgi Gerganov's team is now at Hugging Face, unifying the model hub with the inference engine that powers Ollama, LM Studio, and the entire local AI ecosystem.
This week's open-source highlights: AI2's hybrid architecture proves transformers need help, autoresearch automates ML experiments overnight, and local inference gets serious upgrades.
Ollama delivers 40% faster inference while llama.cpp finds a permanent home at Hugging Face. Two developments that secure the future of running AI on your own hardware.
This week's biggest open-source AI developments: llama.cpp finds a permanent home, China releases a 744B parameter model under MIT license, and a secure WhatsApp AI assistant goes viral
The creators of llama.cpp have joined Hugging Face to ensure long-term sustainability. The projects stay open, the community stays autonomous, and local AI gets resources it needs to compete with cloud inference.