Skip to content
Intelligibberish
  • News
  • Articles
  • Guides
  • Tools
  • About

Tag

#llama-cpp

← All articles

Local AI Aug 6, 2026

Local AI by VRAM: Which Models Fit Your GPU (August 2026)

Find your VRAM tier, then the current open-weight pick for chat, coding, vision, speech, translation or agents. Fourteen guides, one index.

Local AI Aug 6, 2026

Run an LLM Locally on Android and iPhone (August 2026)

Apple caps its on-device model at 4096 tokens per session, and Google's own tables put Gemma 4 E2B at 25.0 decode tokens/sec on an iPhone 17 Pro CPU.

Local AI Aug 6, 2026

Run LLMs Locally on a Mac: What Actually Fits (August 2026)

Apple Silicon has no discrete VRAM, so tier guides mislead Mac owners. The real ceilings are bandwidth and the GPU-usable slice of unified memory.

Local AI Aug 6, 2026

GGUF vs AWQ vs GPTQ vs MLX: Which Quantization to Use

Quantization formats compared with real file sizes and bits per weight. Why Q4 does not halve a model, and why low quants generate faster.

Local AI Jul 27, 2026

MiniMax-M3 Lands in llama.cpp: Sparse Attention and Vision

Two back-to-back merges add MiniMax Sparse Attention and a Qwen2.5-VL style vision tower to llama.cpp, but every existing MiniMax-M3 GGUF must be regenerated.

Local AI Jul 13, 2026

VKUE: Can a 34.7B Reasoner Run on a Laptop Without a GPU?

VIDRAFT_LAB posts Ourbox-35B-JGOS to Hugging Face: 20 tok/s on an 8GB laptop GPU, ~17 tok/s on a CPU-only server, 86.4% on GPQA Diamond.

Local AI Mar 17, 2026

llama.cpp Joins Hugging Face: What It Means for Local AI's Future

Georgi Gerganov's team is now at Hugging Face, unifying the model hub with the inference engine that powers Ollama, LM Studio, and the entire local AI ecosystem.

Local AI Mar 9, 2026

Open-Source AI Wins: OLMo Hybrid Rewrites Efficiency, Karpathy Ships Research Agents, and Local Tools Level Up

This week's open-source highlights: AI2's hybrid architecture proves transformers need help, autoresearch automates ML experiments overnight, and local inference gets serious upgrades.

Local AI Feb 26, 2026

The Week Local AI Grew Up: Ollama 0.17 and llama.cpp's New Home

Ollama delivers 40% faster inference while llama.cpp finds a permanent home at Hugging Face. Two developments that secure the future of running AI on your own hardware.

Local AI Feb 24, 2026

Open-Source AI Wins: ggml Joins Hugging Face, GLM-5 Goes MIT, NanoClaw Hits 14K Stars

This week's biggest open-source AI developments: llama.cpp finds a permanent home, China releases a 744B parameter model under MIT license, and a secure WhatsApp AI assistant goes viral

Local AI Feb 22, 2026

llama.cpp Joins Hugging Face: Local AI Gets a Corporate Backer - and Keeps Its Independence

The creators of llama.cpp have joined Hugging Face to ensure long-term sustainability. The projects stay open, the community stays autonomous, and local AI gets resources it needs to compete with cloud inference.

Intelligibberish

Independent analysis and commentary on artificial intelligence.

News Articles Guides Tools About Disclosure Privacy RSS

© 2026 Intelligibberish. Signal, not noise.