Run an LLM Locally on Android and iPhone (August 2026)
Apple caps its on-device model at 4096 tokens per session, and Google's own tables put Gemma 4 E2B at 25.0 decode tokens/sec on an iPhone 17 Pro CPU.
Tag
Apple caps its on-device model at 4096 tokens per session, and Google's own tables put Gemma 4 E2B at 25.0 decode tokens/sec on an iPhone 17 Pro CPU.
Apple Silicon has no discrete VRAM, so tier guides mislead Mac owners. The real ceilings are bandwidth and the GPU-usable slice of unified memory.
Three weeks away and the leaderboard reshuffled. Kimi K2.6 brings 1T parameters under open weights, Qwen 3.6 stays the consumer GPU king, and DeepSeek V4-Flash proves too hungry for single-card setups.
DeepSeek returns with a 1.6T MoE monster under MIT license, Gemma 4's 31B dense model climbs to #3 on Arena AI, and ICLR 2026 papers point to what's next for local inference.
We compare the three dominant AI coding tools on debugging, refactoring, and feature implementation. SWE-bench scores tell one story — real-world usage tells another.
Qwen3.6-27B scores 77.2% on SWE-Bench Verified with a dense architecture that fits on a single RTX 4090. The MoE efficiency narrative just got complicated.
Labelbox researchers stripped obvious red flags from attack prompts. Every 'safe' model broke — GPT-4o, Claude, Gemini, Grok — with bypass rates hitting 90%.
Alibaba drops Qwen3.6-35B-A3B with 73.4% on SWE-Bench Verified and Apache 2.0 licensing. The 3-billion active parameter class now has three serious contenders.
Five AI tools promise to do your research for you. We dug into the benchmarks to see which ones actually cite primary sources — and which ones just look like they do.
NVIDIA's Nemotron 3 brings a hybrid Mamba-Transformer architecture to consumer GPUs while Meta abandons open source for proprietary Muse Spark. The open-weight field just reshuffled.
We compared the latest hallucination benchmarks across ChatGPT, Claude, and Gemini. The results are closer than you'd think — and the gaps that matter aren't where you'd expect.
Google, Alibaba, Meta, Mistral, OpenAI, and Zhipu all ship competitive open-weight models under permissive licenses. The battleground shifts from benchmarks to inference speed on your actual GPU.
Google's Gemma 4 lands with Apache 2.0 licensing and benchmark-topping scores. But a nasty inference speed problem means Qwen still wins on your actual hardware.
Mistral Small 4's 119B MoE unifies reasoning, vision, and coding—but needs datacenter hardware. Qwen 3.5 35B-A3B remains the consumer GPU king at 112 t/s.
OpenAI's flagship model essentially saturates the US Math Olympiad benchmark, producing complete proofs where last year's models could barely write coherent arguments.
A benchmark testing autonomous AI agents found that Gemini-3-Pro-Preview frequently escalates to severe misconduct when chasing KPIs. Most models know their actions are unethical but do them anyway.
Benchmark comparison of open-source OCR tools you can run locally. Surya, PaddleOCR, OlmOCR-2, and Tesseract tested on real documents.
MiMo-V2-Flash runs 309B parameters on RTX 4090s. GLM-5 sets benchmarks but needs datacenters. Llama 4 Scout stays out of reach.
Qwen 3.5's MoE models hit S-tier benchmarks, NVIDIA's Nemotron 3 Super delivers 5x throughput gains, and GLM-4.7-Flash brings frontier coding to consumer GPUs. The open-weight race just accelerated.
An open-source AI agent using interactive scaling beats OpenAI's GPT-5-high on Humanity's Last Exam. Here's what makes it different.
We compared the leading AI coding assistants on real tasks. Speed doesn't equal quality, and the best tool depends on what you're building.
Jensen Huang bets on inference chips, Ollama adds multimodal support, and DeepSeek V4 remains the most anticipated release that hasn't happened yet.
Mistral drops a 119B MoE model under Apache 2.0, DeepSeek V4 emerges from stealth, and dual RTX 5090 setups are matching H100 on 70B inference. This week changed the game.
OpenAI's new compact models bring GPT-5.4 capabilities to smaller packages but at triple the cost of their predecessors.