Skip to content
Intelligibberish
  • News
  • Articles
  • Guides
  • Tools
  • About

Tag

#vram

← All articles

Local AI Aug 6, 2026

Best Local Embedding and Reranker Models by VRAM (2026)

Local embedding and reranking models for RAG, sized by VRAM. Real GGUF file sizes, licences, and the runtime gap that stops Ollama reranking.

Local AI Aug 6, 2026

Best Local Image Generation Models by VRAM (2026)

Open image models sized for real GPUs, with the text encoder counted in and the licence checked. FLUX, Z-Image, Qwen-Image and SD3.5 compared.

Local AI Aug 6, 2026

6GB VRAM: What You Can Actually Run Locally (2026)

Local AI on a 6GB GPU: GTX 1660, RTX 2060, RTX 3050 and laptop cards. Real weight sizes for chat, coding, vision, speech, translation and RAG.

Local AI Aug 6, 2026

Local AI by VRAM: Which Models Fit Your GPU (August 2026)

Find your VRAM tier, then the current open-weight pick for chat, coding, vision, speech, translation or agents. Fourteen guides, one index.

Local AI Aug 6, 2026

GGUF vs AWQ vs GPTQ vs MLX: Which Quantization to Use

Quantization formats compared with real file sizes and bits per weight. Why Q4 does not halve a model, and why low quants generate faster.

Local AI Mar 17, 2026

12GB VRAM: Every AI Task You Can Run Locally (August 2026)

Local AI on a 12GB GPU: chat, coding, vision, speech and agents for RTX 3060 12GB or RTX 4070. Current picks, per-quant weight sizes, honest limits.

Local AI Mar 17, 2026

16GB VRAM: Every AI Task You Can Run Locally (August 2026)

Local AI on a 16GB GPU: chat, coding, translation, speech and agents for RTX 4060 Ti, RTX 5060 or Arc A770, and why the vision tier stays unresolved.

Local AI Mar 17, 2026

24GB VRAM: Every AI Task You Can Run Locally (August 2026)

Local AI on a 24GB GPU: chat, coding, vision, speech and agents for RTX 3090 or RTX 4090. Current picks, per-quant weight sizes, and an open runtime bug.

Local AI Mar 17, 2026

32GB VRAM: Every AI Task You Can Run Locally (August 2026)

Local AI on a 32GB GPU: chat, coding, vision, speech and agents on an RTX 5090. Current picks, per-quant weight sizes, and what the headroom buys.

Local AI Mar 17, 2026

8GB VRAM: Every AI Task You Can Run Locally (August 2026)

Local AI on an 8GB GPU: chat, coding, vision, speech and agents for RTX 4060 or RTX 3070. Current picks, named quantisations, honest limits.

Intelligibberish

Independent analysis and commentary on artificial intelligence.

News Articles Guides Tools About Disclosure Privacy RSS

© 2026 Intelligibberish. Signal, not noise.