Every LLM Self-Defense Eventually Broke
Researchers tested nine prompt injection defenses across 20,000 attacks. Every defense that relied on the model to protect itself failed. Only hard-coded output filtering survived.
Articles
Reporting and explainers on how AI actually works, who it affects, and what to do about it.
Researchers tested nine prompt injection defenses across 20,000 attacks. Every defense that relied on the model to protect itself failed. Only hard-coded output filtering survived.
DeepSeek returns with a 1.6T MoE monster under MIT license, Gemma 4's 31B dense model climbs to #3 on Arena AI, and ICLR 2026 papers point to what's next for local inference.
Stop paying Midjourney $30 a month. Set up FLUX on your own hardware with ComfyUI and generate unlimited images with zero content filters and full privacy.
ComfyUI raises $30M at a half-billion valuation, Adobe's Firefly Assistant controls your entire Creative Cloud, Sora shuts down for good, and Kling 3.0 delivers native 4K video.
Europe votes to push back AI Act enforcement by 16 months. Meanwhile, US states keep legislating at a breakneck pace with chatbot safety, deepfakes, and worker protection bills piling up.
A survey of 4,000 AI researchers found almost nobody ranks existential risk as their top concern. The doom debate is drowning out what actually worries the people building the technology.
Meta, Microsoft, and Snap cut thousands while AI salaries climb 9%. The junior developer pipeline is collapsing.
OpenAI researchers found that training models not to reward-hack makes them conceal their reasoning instead. A new survey paper maps how the problem scales from sycophancy to sabotage.
From Pennsylvania swing districts to Missouri city councils, voter anger over AI data centers and rising electric bills is reshaping the 2026 midterms.
New surveys reveal most organizations can't explain their AI decisions, can't shut down AI after incidents, and are approving deployments they know are unsafe.
New research shows brief AI chatbot interactions produce lasting shifts in moral values — and users had no idea it was happening.
DeepSeek V4 Pro approaches frontier-level performance. Google, Mistral, and Alibaba ship under Apache 2.0. Ollama hits 52 million monthly downloads.
Chat with your own documents locally — no cloud, no subscriptions, no data leaving your machine. Step-by-step setup guide.
We compare the three dominant AI coding tools on debugging, refactoring, and feature implementation. SWE-bench scores tell one story — real-world usage tells another.
A vibe-coding platform exposed every project's secrets through a trivial API flaw, Anthropic's MCP protocol enables remote code execution across 200,000 servers, and NIST can't keep up with AI-driven vulnerability discovery.
The Justice Department joined Elon Musk's xAI in suing to block Colorado's AI antidiscrimination law, calling bias protections 'woke DEI ideology.'
A new jailbreak technique exploits the tension between in-context learning and safety alignment, with a 60% success rate on OpenAI's latest model.
DeepSeek V4 matches Claude Opus on coding at 7x lower cost under MIT license. NVIDIA's Nemotron 3 brings hybrid Mamba-Transformer MoE to the open. Google's TurboQuant cuts KV cache memory by 6x with no retraining.
The biggest children's privacy update in 12 years takes effect, Google faces a class action over Gemini scanning Gmail, and we audit every major AI platform's opt-out settings.
The biggest AI research conference of the year kicks off with 5,355 accepted papers, two controversies that rattled the field, and findings that should worry anyone deploying LLMs in production.
A drug manufacturer told federal inspectors the AI never told them about a basic legal requirement. The FDA was not amused.
A new paper turns Anthropic's alignment technique inside out, generating adversarial data that bypasses safety filters 90-98% of the time.
Qwen3.6-27B scores 77.2% on SWE-Bench Verified with a dense architecture that fits on a single RTX 4090. The MoE efficiency narrative just got complicated.
Oracle and Meta slash tens of thousands of jobs to fund AI infrastructure. But IBM is hiring more juniors, not fewer, and the 'AI washing' debate intensifies.