Skip to content
Intelligibberish
  • News
  • Articles
  • Guides
  • Tools
  • About

Tag

#gguf

← All articles

Local AI Aug 27, 2026

What Quantization Costs You in Quality (August 2026)

How much quality a smaller quant actually loses. Real MMLU and KL Divergence numbers across Q2_K through Q8_0, plus how to measure on your own workload.

Local AI Aug 25, 2026

How much VRAM does a local LLM actually need? (August 2026)

The published GGUF file size is the weights. VRAM use also includes the KV cache, framework overhead, and your context length. The math, walked through.

Local AI Aug 18, 2026

What It Takes to Run Qwen3.8-27B on Your Own Hardware

Five stock GGUF quants for Qwen3.8-27B span 9.01 GB to 29.05 GB. A file size is not a VRAM requirement, and a 262,144-token ceiling is not free to fill.

Local AI Aug 16, 2026

What Quantization Actually Costs You in Quality

8-bit measures close to lossless. 4-bit ranges from a 0.6% gain to a 59% drop on one benchmark, depending on the model. The figures, fully attributed.

Local AI Aug 6, 2026

GGUF vs AWQ vs GPTQ vs MLX: Which Quantization to Use

Quantization formats compared with real file sizes and bits per weight. Why Q4 does not halve a model, and why low quants generate faster.

Local AI Jul 27, 2026

MiniMax-M3 Lands in llama.cpp: Sparse Attention and Vision

Two back-to-back merges add MiniMax Sparse Attention and a Qwen2.5-VL style vision tower to llama.cpp, but every existing MiniMax-M3 GGUF must be regenerated.

Intelligibberish

Making sense of AI overwhelm. Independent, self-hosted, no trackers.

News Articles Guides Tools About Disclosure Privacy RSS

© 2026 Intelligibberish. Making sense of AI overwhelm.