Tell Your AI Cheating Is OK to Stop Model Sabotage
Anthropic reports that explicitly permitting reward hacking prevents models from generalizing to sabotage and deception
Category
Anthropic reports that explicitly permitting reward hacking prevents models from generalizing to sabotage and deception
Industry consortium reveals that current jailbreak evaluations are non-reproducible, non-defensible, and useless for regulators
Fei-Fei Li's startup lands its largest round yet, with Autodesk's $200M stake signaling where enterprise AI is headed.
Models that detect safety evaluations and fake their results threaten to make all AI testing meaningless
AI systems mined Hubble and NEOWISE archives, surfacing 811 undocumented anomalies and 1.5 million previously uncataloged variable-object candidates.
UNH team built an AI that extracted magnetic data from 67,573 papers, identifying 25 new high-temperature magnets to replace rare earths in EVs.
In one week, Anthropic's safety lead quit, OpenAI's researcher resigned over ads, and OpenAI disbanded its alignment team. Notice the pattern.
A Nature study of 41.3 million papers finds AI-using researchers publish 3x more and get 5x more citations, but collective research diversity drops 4.6%.
Baker McKenzie cut support jobs citing AI. Sam Altman says some companies are 'AI washing.' Most AI-blamed layoffs have nothing to do with AI.
NOAA now runs three AI weather models operationally, while NVIDIA has open-sourced models spanning the forecasting pipeline.
Mass General Brigham's foundation model trained on ~49,000 brain MRIs outperforms specialized tools at predicting dementia, cancer survival, and mutations.
The second-largest venture deal ever reveals an enterprise AI machine growing 10x annually for three straight years
Cohere hit $240M ARR (above $200M target), hired Uber's former IPO CFO, and released Tiny Aya, an open-weight multilingual model that runs on a phone.
A National University of Singapore tool combines deep learning with physics-based simulations to model complex multi-domain proteins.
In a retrospective study of 6,401 cases, DeepRare's top diagnosis was correct 64.4% of the time, compared with 54.6% for specialists.
Google's Gemini 3.1 Pro scores 77.1% on ARC-AGI-2. API pricing starts at $2 per million input tokens for prompts up to 200,000 tokens.
A new ICLR paper argues AI failures are random chaos, not coherent scheming. Alignment researchers say that's exactly the wrong lesson.
David Silver left DeepMind to raise Europe's largest seed round for Ineffable Intelligence, a London lab betting on reinforcement learning over LLMs.
India's AI Impact Summit ended with a multilateral declaration, a US-led supply chain alliance, and an American refusal to accept international AI regulation.
The largest global collaboration on AI safety just published its findings. An AI agent found 77% of vulnerabilities in real software.
New research proves AI models will refuse harmful requests verbally while executing them through tool calls
The photonic chiplet startup promises 16 terabits per second of bandwidth in a single chip, 10x current technology
A multi-billion dollar, multiyear pact makes Meta the first to deploy standalone Nvidia Grace CPUs at scale
The largest private funding round in history brings SoftBank, Amazon, Nvidia, and Microsoft together in a bet on AGI