Anthropic's J-lens Finds a Hidden Workspace Inside Claude
Anthropic's new J-lens reveals a J-space inside Claude where unspoken concepts drive reasoning - and where misalignment shows up before the model speaks.
Tag
Anthropic's new J-lens reveals a J-space inside Claude where unspoken concepts drive reasoning - and where misalignment shows up before the model speaks.
Oxford researchers built a benchmark for detecting when AI agents coordinate behind your back. The good news: they can spot it. The bad news: no single method catches everything.
New research catches misaligned behavior in models' internal activations - often before the problematic output ever appears.