AI Bills Triple, Tokens Vanish: The Enterprise Cost Crackdown
Atlassian, Adobe, Amazon, and Citi are cutting access to frontier models after monthly AI spend hit $15M. Inside the enterprise token crunch.
Tag
Atlassian, Adobe, Amazon, and Citi are cutting access to frontier models after monthly AI spend hit $15M. Inside the enterprise token crunch.
PhantaField's Sophon PFG-1 whitepaper claims ~95x Nvidia HBM4 bandwidth via monolithic 3D stacking. No silicon yet. Here's why it matters anyway.
Researchers found the exact neurons responsible for refusing harmful requests — then switched them off. No retraining. No fine-tuning. Just geometry.
New compression algorithm achieves 6x memory reduction with zero accuracy loss. No retraining required. This matters for anyone running local AI.
Cloudflare adds its first frontier-scale model to Workers AI, claiming 77% cost savings over proprietary alternatives with new caching features.
Jensen Huang bets on inference chips, Ollama adds multimodal support, and DeepSeek V4 remains the most anticipated release that hasn't happened yet.
Jensen Huang's keynote today marks Nvidia's biggest pivot in years - from training chips to inference, from cloud to edge, and from prompts to autonomous agents
While enterprises focus on training data and model safety, inference - where AI actually processes requests - has become an overlooked security frontier with critical vulnerabilities.
Ollama delivers 40% faster inference while llama.cpp finds a permanent home at Hugging Face. Two developments that secure the future of running AI on your own hardware.
New research from Google and UVA reveals that longer AI reasoning traces actually correlate with wrong answers. The fix: measure how deeply the model thinks, not how much it writes.
A Toronto startup is etching LLM weights directly into transistors, achieving 17,000 tokens per second. The catch: you can't change the model.
GPT-5.3-Codex-Spark runs on Cerebras' wafer-scale chips at 1,000+ tokens per second. It's OpenAI's first production break from NVIDIA - and it won't be the last.
A 300-gram device promises GPT-4o-level performance without cloud or internet. The specs are real, but the benchmarks are missing.