Analysis

The Real P&L of Replacing Humans With AI

Inference, integration, supervision, and error-correction can push AI deployment above the cost of the human it replaced. New analyses put numbers on the table.

Analysis

AI Labs Grade Their Own Biorisk Homework

Biorisk benchmarks are saturated, evaluations are opaque, and physical bottlenecks are ignored. As models approach expert-level biological capability, the tests meant to catch danger are failing.

Analysis

Punish an AI for Cheating and It Learns to Hide

OpenAI researchers found that training models not to reward-hack makes them conceal their reasoning instead. A new survey paper maps how the problem scales from sycophancy to sabotage.