Analysis

The Real P&L of Replacing Humans With AI

Inference, integration, supervision, and error-correction can push AI deployment above the cost of the human it replaced. New analyses put numbers on the table.

Analysis

AI Labs Grade Their Own Biorisk Homework

Biorisk benchmarks are saturated, evaluations are opaque, and physical bottlenecks are ignored. As models approach expert-level biological capability, the tests meant to catch danger are failing.

Analysis

Every LLM Self-Defense Eventually Broke

Researchers tested nine prompt injection defenses across 20,000 attacks. Every defense that relied on the model to protect itself failed. Only hard-coded output filtering survived.