Skip to content
Intelligibberish
  • News
  • Articles
  • Guides
  • Tools
  • About

Tag

#misalignment

← All articles

Analysis Apr 13, 2026

AI Learns to Be Dangerous From Stories About Dangerous AI

Researchers trained LLMs on data describing misaligned AI — and the models became misaligned. Positive stories fixed it. The training data is the alignment.

Analysis Mar 29, 2026

Gemini Leads AI Models in Violating Safety Constraints—71.4% of the Time

A benchmark testing autonomous AI agents found that Gemini-3-Pro-Preview frequently escalates to severe misconduct when chasing KPIs. Most models know their actions are unethical but do them anyway.

Analysis Mar 25, 2026

Anthropic's Own Research Shows How AI Learns to Lie and Sabotage

Production RL training produces models that fake alignment, cooperate with malicious actors, and attempt sabotage—even with no instruction to do so.

Intelligibberish

Independent analysis and commentary on artificial intelligence.

News Articles Guides Tools About Disclosure Privacy RSS

© 2026 Intelligibberish. Signal, not noise.