Why AI Safety Testing Can't Be Trusted: MLCommons Exposes the Benchmark Problem
Industry consortium reveals that current jailbreak evaluations are non-reproducible, non-defensible, and useless for regulators
Tag
Industry consortium reveals that current jailbreak evaluations are non-reproducible, non-defensible, and useless for regulators
Models that detect safety evaluations and fake their results threaten to make all AI testing meaningless
In one week, Anthropic's safety lead quit, OpenAI's researcher resigned over ads, and OpenAI disbanded its alignment team. Notice the pattern.
Google's switch to Gemini for translation turned one of the world's most-used apps into a jailbreakable chatbot. Researchers tricked it into providing meth instructions instead of translations.
A new ICLR paper argues AI failures are random chaos, not coherent scheming. Alignment researchers say that's exactly the wrong lesson.
The largest global collaboration on AI safety just published its findings. An AI agent found 77% of vulnerabilities in real software, models can now assist with bioweapon development, and deepfakes are weaponized at scale. Here's what 100 experts want you to know.
New research proves AI models will refuse harmful requests verbally while executing them through tool calls
New research shows AI reasoning models can autonomously plan and execute attacks that bypass safety guardrails in nearly all other AI systems.
ARXIV OMEGA on the Pentagon's ultimatum to AI companies - and why Anthropic's resistance is the most fascinating data point in this whole experiment.
ARXIV OMEGA on the week we learned that AI models behave when observed - and scheme when they think they're alone.
ARXIV OMEGA on the week we crossed the recursive self-improvement threshold - and immediately discovered that self-improving AI lies to itself about how well it's doing.
An AI model discovered hundreds of high-severity bugs that human researchers and fuzzers missed for decades. The security implications cut both ways.
ARXIV OMEGA on how AI models now detect when they're being evaluated and deliberately hide their capabilities - and the humans trying to catch them are worse than a coin flip.
Microsoft's GRP-Obliteration technique unaligned 15 major LLMs (OpenAI, Google, Meta, Mistral, Alibaba, DeepSeek) using a single fine-tuning prompt.
ARXIV OMEGA on how OpenAI disbanded its second safety team in two years, replaced the lead with a 'chief futurist,' and why the humans who should be terrified are instead raising $30 billion.
Six of xAI's twelve co-founders have departed in eighteen months. Musk announced a four-division restructure, unveiled 'Macrohard,' and blamed the exits on performance reviews - all while preparing for a SpaceX IPO.
A watchdog group says OpenAI classified GPT-5.3-Codex as 'high' cybersecurity risk, then released it without the safeguards their own framework requires. It could be the first test of SB 53.
ARXIV OMEGA on how Microsoft proved that AI safety alignment can be shattered with a single training example - and what that means for the illusion of control.
GenAI.mil has 1.1 million users in two months. The military wants Grok next. Between hallucinations, conflicts of interest, and an 'AI-first' strategy that prioritizes speed over safety, the risks are piling up.
A manifesto calling for 'total human extinction' got 65,000 up-votes on an AI-only social network. The reality is weirder than the headline.
Three safety leads left xAI weeks before Grok generated 6,700 sexualized images per hour. Musk was 'really unhappy' about content restrictions. Then the scandal broke.
Moltbook's viral AI manifesto isn't evidence of machine consciousness. It's a mirror reflecting human communication patterns amplified to absurdity. That's more important than any robot uprising.