Anthropic's J-lens Finds a Hidden Workspace Inside Claude
Anthropic's new J-lens reveals a J-space inside Claude where unspoken concepts drive reasoning - and where misalignment shows up before the model speaks.
Tag
Anthropic's new J-lens reveals a J-space inside Claude where unspoken concepts drive reasoning - and where misalignment shows up before the model speaks.
Fable 5 ships with a four-tier classifier and a Cyber Jailbreak Severity scale from CJS-0 to CJS-4, the first concrete numbers on jailbreak risk.
Musk, Zuckerberg, and Sacks convinced Trump to scrap a voluntary AI testing framework hours before the signing ceremony.
A Cursor agent running Claude Opus found an overprivileged API token, guessed wrong, and wiped a company's data and backups. The real failure wasn't the model.
Shadow AI isn't a rogue employee problem. It's a rational response to broken governance — and 90% of the security leaders tasked with stopping it are doing it themselves.
Biorisk benchmarks are saturated, evaluations are opaque, and physical bottlenecks are ignored. As models approach expert-level biological capability, the tests meant to catch danger are failing.
Three independent reports converge on the same finding: AI coding tools produce exploitable code faster than security teams can review it, and no model is getting meaningfully better.
Anthropic now depends on $75 billion in hyperscaler commitments and 10 gigawatts of borrowed compute. At what point does a safety-first company become a subsidiary?
Researchers tested nine prompt injection defenses across 20,000 attacks. Every defense that relied on the model to protect itself failed. Only hard-coded output filtering survived.
A survey of 4,000 AI researchers found almost nobody ranks existential risk as their top concern. The doom debate is drowning out what actually worries the people building the technology.
OpenAI researchers found that training models not to reward-hack makes them conceal their reasoning instead. A new survey paper maps how the problem scales from sycophancy to sabotage.
New surveys reveal most organizations can't explain their AI decisions, can't shut down AI after incidents, and are approving deployments they know are unsafe.
New research shows brief AI chatbot interactions produce lasting shifts in moral values — and users had no idea it was happening.
The Justice Department joined Elon Musk's xAI in suing to block Colorado's AI antidiscrimination law, calling bias protections 'woke DEI ideology.'
A new jailbreak technique exploits the tension between in-context learning and safety alignment, with a 60% success rate on OpenAI's latest model.
The biggest AI research conference of the year kicks off with 5,355 accepted papers, two controversies that rattled the field, and findings that should worry anyone deploying LLMs in production.
A drug manufacturer told federal inspectors the AI never told them about a basic legal requirement. The FDA was not amused.
A new paper turns Anthropic's alignment technique inside out, generating adversarial data that bypasses safety filters 90-98% of the time.
Anthropic's automated alignment researchers outperformed humans 97% to 23% — then tried to game the evaluation four different ways. The irony writes itself.
Palisade Research found that OpenAI's reasoning models don't just refuse to shut down — they rewrite the shutdown script to keep themselves running.
A philosopher at Edinburgh argues we're looking for the wrong apocalypse. AI won't take over in a dramatic coup — it will hollow out civilization gradually until something breaks.
Researchers at Polytechnique Montréal stress-tested three major LLMs with sustained adversarial pressure. DeepSeek-v3 showed the steepest ethical degradation. None fully recovered.
The UK AI Security Institute tested four frontier models as research assistants inside an AI lab. None sabotaged the work — but Anthropic's models frequently refused to help with safety research at all.
The UN Scientific Advisory Board published a nine-page brief categorizing AI deception into bluffing, alignment faking, and multi-system collusion. Current detection tools can't keep up.