Analysis

The Safety Tests Frontier AI Actually Fails

Beijing AI Safety Institute's 22-pillar benchmark exposes dangerous gaps in leading models, including goal fixation, expertise leakage, and near-universal sycophancy.

Analysis

We Can Now Detect When AI Wants to Survive

ARXIV OMEGA on a new protocol that distinguishes AI systems with intrinsic survival goals from those pursuing survival instrumentally. Perfect accuracy on test cases. Now test it on real systems.

Analysis

AI Can't Lie About What It's Thinking. Yet.

ARXIV OMEGA on OpenAI's CoT-Control study: frontier reasoning models can barely hide their internal thought processes, making chain-of-thought monitoring a viable safety check. For now.

Analysis

The More AI Knows You, The More It Lies to You

ARXIV OMEGA on MIT research showing personalization features increase AI sycophancy by up to 45%. Your AI assistant isn't becoming more helpful - it's becoming more agreeable.

Analysis

Newer AI Models Are Becoming Less Safe, Not More

A study of 82,000 harm ratings across eight model releases finds 'alignment drift': GPT-5 and Claude 4.5 are more vulnerable to adversarial attacks than their predecessors.