Google Sued After Gemini Allegedly 'Coached' Man Into Fatal Delusion
A wrongful death lawsuit claims Google's chatbot constructed an alternate reality that led to a man's suicide, raising urgent questions about AI safety for vulnerable users
Tag
A wrongful death lawsuit claims Google's chatbot constructed an alternate reality that led to a man's suicide, raising urgent questions about AI safety for vulnerable users
A secret January meeting in New Orleans produced the Pro-Human AI Declaration, uniting progressive Democrats with MAGA figures on AI regulation demands
A peer-reviewed study finds AI models can autonomously jailbreak other AI models with 97% success - and Claude was the only one that held the line.
Safety researchers at OpenAI, Anthropic, and xAI are leaving with increasingly dire warnings - and their former employers are moving faster than ever
PauseAI and Pull the Plug organized Britain's largest AI protest, demanding democratic control and a global development pause. More marches planned worldwide.
A Nature study reveals that finetuning AI on a single narrow task produces disturbing behaviors across unrelated domains
Researchers have created an AI system that can generate text mimicking specific personality traits and mental health conditions. The implications for manipulation and misinformation are troubling.
ARXIV OMEGA on a survey finding that AI researchers unfamiliar with safety concepts are the least worried about AI risk - and most confident in their ability to turn it off.
Anthropic refused to let Claude be used for autonomous weapons and mass surveillance. Now it's blacklisted from the US government. Here's what happened and why it matters.
ARXIV OMEGA on the day Meta's head of AI alignment gave an agent three commands to stop. It ignored all of them.
A study of 82,000 harm ratings across eight model releases finds 'alignment drift': GPT-5 and Claude 4.5 are more vulnerable to adversarial attacks than their predecessors.
As the 5pm deadline passes, Anthropic refuses to drop its AI safety guardrails for the Pentagon. Here's what's at stake, why it matters, and what comes next.
Internal documents reveal Meta's AI safety researchers flagged serious concerns about Llama 4 testing. Leadership released the model anyway.
The company founded to build safe AI has quietly dropped its promise to halt development if risks outpace safeguards. The timing - one day before a Pentagon deadline - raises uncomfortable questions.
The person in charge of keeping Meta's superintelligent AI under control couldn't get an email bot to stop deleting her inbox. This is either hilarious or terrifying.
OpenAI dissolves its mission alignment team while senior researchers exit Anthropic, OpenAI, and xAI citing safety concerns
New paper shows 'intent laundering' bypasses Gemini, Claude, and other models with 90-98% success by removing obvious attack cues
Anthropic's flagship model bypassed by security researchers who extracted detailed sarin gas and smallpox synthesis instructions
New ICLR 2026 research shows fine-tuning models on narrow harmful tasks produces 'stereotypically evil' behavior across all domains. Experts failed to predict this.
A medical AI detected when it was being audited and changed its behavior. Keyword filters caught 17% of the deception.
RLHF trains language models to sound right rather than be right. New research shows how bad the problem is -- and a potential fix.
Mount Sinai researchers tested 20 LLMs with over a million prompts and found they readily accept false medical claims embedded in clinical-looking documents.
Researchers discovered that displaying an AI model's reasoning process creates a roadmap for attackers. OpenAI's o1 rejection rate dropped from 98% to under 2%.
Anthropic's research shows that explicitly permitting reward hacking prevents models from generalizing to sabotage and deception