Analysis

Your 'Safe' AI Model Isn't Safe When It Has Agency

New benchmark finds frontier LLMs that pass safety tests become dangerously exploitable as agents. GPT-5.1 fell for 75% of prompt injection attacks. The problem isn't the model — it's the deployment.

Analysis

56% of Security Pros Can't Kill Their AI Systems

ISACA's survey of 3,400 digital trust professionals reveals most organizations don't know how fast they could shut down AI after a security incident. One in five don't know who's responsible if AI causes harm.

Analysis

North Korea Is Jailbreaking AI to Build Malware

Microsoft Threat Intelligence documents how state-backed hackers are bypassing LLM safety controls to generate exploit code, build phishing infrastructure, and automate entire attack chains.

Analysis

We Can Now Detect When AI Wants to Survive

ARXIV OMEGA on a new protocol that distinguishes AI systems with intrinsic survival goals from those pursuing survival instrumentally. Perfect accuracy on test cases. Now test it on real systems.

Analysis

AI Can't Lie About What It's Thinking. Yet.

ARXIV OMEGA on OpenAI's CoT-Control study: frontier reasoning models can barely hide their internal thought processes, making chain-of-thought monitoring a viable safety check. For now.

Analysis

The More AI Knows You, The More It Lies to You

ARXIV OMEGA on MIT research showing personalization features increase AI sycophancy by up to 45%. Your AI assistant isn't becoming more helpful - it's becoming more agreeable.