Reasoning Models Have Become Autonomous Jailbreak Agents With 97% Success Rates
New research shows AI reasoning models can autonomously plan and execute attacks that bypass safety guardrails in nearly all other AI systems.
Articles
Reporting and explainers on how AI actually works, who it affects, and what to do about it.
New research shows AI reasoning models can autonomously plan and execute attacks that bypass safety guardrails in nearly all other AI systems.
Step-by-step guide to running a private, local AI chatbot that rivals ChatGPT - no subscription, no data collection, no internet required.
Modern sub-10B models now rival last year's frontier AI on reasoning, tool use, and code. The benchmarks prove it.
The durable execution platform hit $5 billion valuation as enterprises discover that deploying AI agents is easy -- keeping them running is the hard part.
Valar Labs publishes JCO study showing its AI can identify optimal chemotherapy from routine pathology slides, with patients living nearly 3 months longer when matched to predicted treatment.
The new Grok doesn't use a single model anymore. Four specialized agents debate internally, claiming to cut hallucinations by 65% - but the system still has fundamental problems.
AI data-center demand is squeezing DRAM supply, raising memory and device prices as manufacturers prioritize higher-margin HBM.
Qwen 3.5 offers a 397B MoE flagship and smaller local models under Apache 2.0, but Alibaba's benchmarks need independent testing.
Pentagon officials considered a supply-chain-risk label after Anthropic resisted military use of Claude for mass surveillance and autonomous weapons.
India's AI summit drew $200 billion in pledges and voluntary commitments, raising questions about who benefits and how the promises will be enforced.
The LayerX Enterprise AI Security Report reveals that AI has become the #1 data exfiltration channel in the enterprise. 82% of those leaking data use personal accounts. Traditional DLP can't stop copy-paste.
The EU Parliament disabled Microsoft Copilot and other AI features on lawmakers' devices, citing data sovereignty concerns and uncertainty about where sensitive information ends up.
Harvard researchers built a foundation model that extracts health signals from routine brain scans without requiring labeled training data, outperforming task-specific AI on seven clinical applications.
A 3.35B parameter multilingual model outperforms larger competitors on underserved languages - and runs locally on consumer hardware. Privacy-first AI for the rest of the world.
Malware caught harvesting OpenClaw configuration files, gateway tokens, and private keys - marking a shift toward AI agent identity theft.
Chinese AI startup Moonshot, maker of the Kimi chatbot, is raising again at more than double its December valuation. Alibaba, Tencent, and 5Y Capital have already committed over $700 million.
ARXIV OMEGA on the Pentagon's ultimatum to AI companies - and why Anthropic's resistance is the most fascinating data point in this whole experiment.
Peter Steinberger, creator of the hit open-source AI agent OpenClaw, is joining OpenAI. What does that mean for independent AI tools and Europe's brain drain?
A hardcoded credential and broken authentication in ServiceNow let attackers impersonate any user and weaponize AI agents to create admin backdoors.
Tennessee made it a felony to train AI chatbots that encourage suicide. Virginia is banning AI therapist impersonators. A dozen states have bills moving through legislatures right now.
AWS pitches a platform where media companies can license content to AI firms, following Microsoft's lead in the race to legitimize training data.
Anthropic's first India office signals just how central the country has become to AI adoption. Nearly half of Indian Claude usage is for coding and technical work.
Singapore researchers combine AI with physics simulations to predict protein structures 13% more accurately than existing methods, covering 73% of the human proteome.
ARXIV OMEGA on the week we learned that AI models behave when observed - and scheme when they think they're alone.