Best Local Speech Models by VRAM Tier (August 2026)
Local TTS and STT by VRAM tier: Parakeet, Canary, MOSS-Transcribe-Diarize, Step-Audio-EditX, Fish Audio S2 Pro and Kokoro, and the licence each ships.
Articles
Reporting and explainers on how AI actually works, who it affects, and what to do about it.
Local TTS and STT by VRAM tier: Parakeet, Canary, MOSS-Transcribe-Diarize, Step-Audio-EditX, Fish Audio S2 Pro and Kokoro, and the licence each ships.
TranslateGemma, NLLB-200, Aya Expanse and Qwen3.5 by VRAM tier, with the licence terms that decide whether you can ship what you run.
Local image analysis, OCR, and visual reasoning from 8GB to 32GB VRAM. Qwen3.5 replaces Qwen3-VL at most tiers, and 16GB stays unresolved.
Epic, Google, Oracle and Microsoft race to deploy autonomous AI agents in healthcare. But experts warn that patient safety testing has not kept pace with the rush to market.
Neuracle's BCI device for spinal cord injury patients achieves 100% improvement in grasping function across 36-patient trial.
Georgi Gerganov's team is now at Hugging Face, unifying the model hub with the inference engine that powers Ollama, LM Studio, and the entire local AI ecosystem.
Local AI on a 12GB GPU: chat, coding, vision, speech and agents for RTX 3060 12GB or RTX 4070. Current picks, per-quant weight sizes, honest limits.
Local AI on a 16GB GPU: chat, coding, translation, speech and agents for RTX 4060 Ti, RTX 5060 or Arc A770, and why the vision tier stays unresolved.
Local AI on a 24GB GPU: chat, coding, vision, speech and agents for RTX 3090 or RTX 4090. Current picks, per-quant weight sizes, and an open runtime bug.
Local AI on a 32GB GPU: chat, coding, vision, speech and agents on an RTX 5090. Current picks, per-quant weight sizes, and what the headroom buys.
Local AI on an 8GB GPU: chat, coding, vision, speech and agents for RTX 4060 or RTX 3070. Current picks, named quantisations, honest limits.
Jensen Huang's keynote unveils Vera Rubin chips, a $20B Groq acquisition, DLSS 5, and positions Nvidia to dominate both training and inference markets.
Beijing AI Safety Institute's 22-pillar benchmark exposes dangerous gaps in leading models, including goal fixation, expertise leakage, and near-universal sycophancy.
New research catches misaligned behavior in models' internal activations - often before the problematic output ever appears.
Perplexity's CTO announces the company is moving away from Anthropic's Model Context Protocol, citing context window bloat and authentication friction. The shift reveals growing pains in AI tooling.
Zenity Labs discloses critical flaws in agentic browsers like Perplexity Comet. A zero-click attack can steal local files and passwords without user interaction.
Leaked internal email shows Ring's lost dog feature is the foundation for broader AI surveillance. Congress demands answers as partnership with Flock Safety collapses.
Three teenagers have filed a federal class action against Elon Musk's xAI, alleging Grok was used to create child sexual abuse material from their photos. It's the first lawsuit where minors are plaintiffs.
Enterprise software consolidation continues as Zendesk bets $115M+ AI startup can make human customer service agents obsolete
Big Tech signed a ratepayer pledge while their AI infrastructure expands into drought-stricken regions. Here's what the latest numbers show.
The largest copyright settlement in history closes March 30. If your book was pirated by AI, here's what you need to know and how to file a claim.
A generative AI trained on 23,000 synthesis recipes can suggest how to make materials that have never existed
Google quietly expands Pentagon partnership with 8 Gemini agents for 3 million DoD workers as Anthropic sues and employees across companies demand guardrails.
As Meta pushes its flagship AI model to May, considers licensing from Google, and loses its legendary chief scientist, the company's $135 billion AI bet faces its biggest test