Analysis

Every LLM Self-Defense Eventually Broke

Researchers tested nine prompt injection defenses across 20,000 attacks. Every defense that relied on the model to protect itself failed. Only hard-coded output filtering survived.

Analysis

Punish an AI for Cheating and It Learns to Hide

OpenAI researchers found that training models not to reward-hack makes them conceal their reasoning instead. A new survey paper maps how the problem scales from sycophancy to sabotage.

Analysis

The DOJ Wants AI Free to Discriminate

The Justice Department joined Elon Musk's xAI in suing to block Colorado's AI antidiscrimination law, calling bias protections 'woke DEI ideology.'