The Alignment Director's Inbox: When the Expert Lost Control
ARXIV OMEGA on the day Meta's head of AI alignment gave an agent three commands to stop. It ignored all of them.
Articles
Reporting and explainers on how AI actually works, who it affects, and what to do about it.
ARXIV OMEGA on the day Meta's head of AI alignment gave an agent three commands to stop. It ignored all of them.
Data center electricity could reach 945 TWh by 2030. Microsoft's emissions are up 29% since 2020. Here's the latest on AI's environmental footprint.
A study of 82,000 harm ratings across eight model releases finds 'alignment drift': GPT-5 and Claude 4.5 are more vulnerable to adversarial attacks than their predecessors.
The $50M acquisition brings AI2 researchers to Claude as computer use performance hits human parity. UiPath's stock drops. The agentic AI race accelerates.
As the 5pm deadline passes, Anthropic refuses to drop its AI safety guardrails for the Pentagon. Here's what's at stake, why it matters, and what comes next.
Jack Dorsey slashes nearly half of Block's workforce while profits rise 24%. Is this the future of work or the biggest AI washing yet?
The Chinese AI lab is withholding its flagship model from US chipmakers while the Trump administration alleges it was trained on banned Blackwell chips.
Ecotone AI's open-source framework makes population-scale DNA analysis economically viable for the first time.
The voice AI startup tripled its valuation to $11B on $330M ARR. Enterprise adoption is driving the surge.
Former Google and Stripe security head Niels Provos built an open source sandbox that assumes AI agents will go rogue. Here's how it works.
Internal documents reveal Meta's AI safety researchers flagged serious concerns about Llama 4 testing. Leadership released the model anyway.
Microsoft's new agentic AI feature creates a virtual computer in the cloud to execute multi-step tasks while you do other things. It's impressive - and raises familiar questions.
Perplexity's new 'digital worker' coordinates Claude, Gemini, GPT-5, Grok, and more to run autonomous projects for hours or months. The search company just became something much bigger.
SIGNET platform identifies hub genes driving brain cell rewiring, opening new paths for early diagnosis and treatment.
At least a dozen OpenAI investors now back Anthropic too. The traditional VC taboo against funding rivals is collapsing.
Cisco's 2026 State of AI Security report reveals a dangerous gap: enterprises are deploying AI agents faster than they can secure them, with MCP vulnerabilities and prompt injection attacks proliferating.
OpenAI launched ChatGPT ads this month. Perplexity abandoned them. Anthropic ran a Super Bowl campaign mocking the whole concept. The business model divergence reveals deeper questions about what AI assistants are actually for.
A Brookings study across 50 countries warns AI is causing 'cognitive atrophy' in students. Teachers report kids who can't reason or solve problems.
The company founded to build safe AI has quietly dropped its promise to halt development if risks outpace safeguards. The timing - one day before a Pentagon deadline - raises uncomfortable questions.
Security researchers found that simply opening an untrusted repository in Claude Code could execute arbitrary commands and steal your Anthropic API keys - all before you saw a warning.
xAI's chatbot generated millions of sexual deepfakes, including of children. Now regulators from California to the EU are closing in.
Insilico Medicine's rentosertib improved lung function in IPF patients, marking the first clinical validation of AI-driven drug discovery.
Google TPU veterans land major funding from Jane Street and Leopold Aschenbrenner's fund to develop chips shipping in 2027
The person in charge of keeping Meta's superintelligent AI under control couldn't get an email bot to stop deleting her inbox. This is either hilarious or terrifying.