Skip to content
Intelligibberish
  • News
  • Articles
  • Guides
  • Tools
  • About

Tag

#quantization

← All articles

Local AI Sep 7, 2026

Fine-Tune a Local LLM with LoRA and QLoRA (September 2026)

How to adapt an open-weight model on your own hardware. What LoRA rank and alpha mean, when QLoRA's 4-bit NF4 is worth it, and how to measure it.

Local AI Aug 27, 2026

What Quantization Costs You in Quality (August 2026)

How much quality a smaller quant actually loses. Real MMLU and KL Divergence numbers across Q2_K through Q8_0, plus how to measure on your own workload.

Local AI Aug 25, 2026

How much VRAM does a local LLM actually need? (August 2026)

The published GGUF file size is the weights. VRAM use also includes the KV cache, framework overhead, and your context length. The math, walked through.

Local AI Aug 18, 2026

What It Takes to Run Qwen3.8-27B on Your Own Hardware

Five stock GGUF quants for Qwen3.8-27B span 9.01 GB to 29.05 GB. A file size is not a VRAM requirement, and a 262,144-token ceiling is not free to fill.

Local AI Aug 16, 2026

What Quantization Actually Costs You in Quality

8-bit measures close to lossless. 4-bit ranges from a 0.6% gain to a 59% drop on one benchmark, depending on the model. The figures, fully attributed.

Local AI Aug 13, 2026

Local AI Without a GPU: What Actually Runs (August 2026)

Which open models run on CPU and RAM alone, how to size them, why mixture-of-experts helps, and what a machine with no discrete GPU cannot do.

Local AI Aug 6, 2026

Local AI by VRAM: Which Models Fit Your GPU (September 2026)

Find your VRAM tier, then the current open-weight pick for chat, coding, vision, speech, translation or agents. Eighteen guides, one index.

Local AI Aug 6, 2026

GGUF vs AWQ vs GPTQ vs MLX: Which Quantization to Use

Quantization formats compared with real file sizes and bits per weight. Why Q4 does not halve a model, and why low quants generate faster.

Local AI Mar 4, 2026

Open-Weight LLM Showdown: What Runs on Your GPU (August 2026)

GLM-5, Qwen 3.5, DeepSeek V3.2, and MiniMax M2.5 are rewriting the rules. Here's what they actually deliver on consumer hardware.

Intelligibberish

Making sense of AI overwhelm. Independent, self-hosted, no trackers.

News Articles Guides Tools About Disclosure Privacy RSS

© 2026 Intelligibberish. Making sense of AI overwhelm.