VKUE: Can a 34.7B Reasoner Run on a Laptop Without a GPU?
VIDRAFT_LAB posts Ourbox-35B-JGOS to Hugging Face: 20 tok/s on an 8GB laptop GPU, ~17 tok/s on a CPU-only server, 86.4% on GPQA Diamond.
Tag
VIDRAFT_LAB posts Ourbox-35B-JGOS to Hugging Face: 20 tok/s on an 8GB laptop GPU, ~17 tok/s on a CPU-only server, 86.4% on GPQA Diamond.
Georgi Gerganov's team is now at Hugging Face, unifying the model hub with the inference engine that powers Ollama, LM Studio, and the entire local AI ecosystem.
This week's open-source highlights: AI2's hybrid architecture proves transformers need help, autoresearch automates ML experiments overnight, and local inference gets serious upgrades.
Ollama delivers 40% faster inference while llama.cpp finds a permanent home at Hugging Face. Two developments that secure the future of running AI on your own hardware.
This week's biggest open-source AI developments: llama.cpp finds a permanent home, China releases a 744B parameter model under MIT license, and a secure WhatsApp AI assistant goes viral
The creators of llama.cpp have joined Hugging Face to ensure long-term sustainability. The projects stay open, the community stays autonomous, and local AI gets resources it needs to compete with cloud inference.