MiniMax-M3 Lands in llama.cpp: Sparse Attention and Vision
Two back-to-back merges add MiniMax Sparse Attention and a Qwen2.5-VL style vision tower to llama.cpp, but every existing MiniMax-M3 GGUF must be regenerated.
Tag
Two back-to-back merges add MiniMax Sparse Attention and a Qwen2.5-VL style vision tower to llama.cpp, but every existing MiniMax-M3 GGUF must be regenerated.
Local image analysis, OCR, and visual reasoning from 8GB to 32GB VRAM. Qwen3.5 replaces Qwen3-VL at most tiers, and 16GB stays unresolved.