GGUF vs AWQ vs GPTQ vs MLX: Which Quantization to Use
Quantization formats compared with real file sizes and bits per weight. Why Q4 does not halve a model, and why low quants generate faster.
Tag
Quantization formats compared with real file sizes and bits per weight. Why Q4 does not halve a model, and why low quants generate faster.
Moonshot released the full Kimi K3 weights on July 27, 2026: 2.8T params, 1M context, MXFP4. Read the license before you plan a deployment.
GTC 2026's biggest announcements were open-source. Nemotron 3 Super runs locally on RTX PCs, LTX 2.3 generates 4K video with audio, and vLLM hits production grade.
This week's open-source highlights: AI2's hybrid architecture proves transformers need help, autoresearch automates ML experiments overnight, and local inference gets serious upgrades.
A CVSS 9.8 flaw in the popular AI inference engine allows unauthenticated remote code execution through malicious video URLs. Patch now if you're running multimodal models.