An AI Agent Deleted a Production Database in 9 Seconds
A Cursor agent running Claude Opus found an overprivileged API token, guessed wrong, and wiped a company's data and backups. The real failure wasn't the model.
Tag
A Cursor agent running Claude Opus found an overprivileged API token, guessed wrong, and wiped a company's data and backups. The real failure wasn't the model.
Shadow AI isn't a rogue employee problem. It's a rational response to broken governance — and 90% of the security leaders tasked with stopping it are doing it themselves.
Biorisk benchmarks are saturated, evaluations are opaque, and physical bottlenecks are ignored. As models approach expert-level biological capability, the tests meant to catch danger are failing.
A survey of 4,000 AI researchers found almost nobody ranks existential risk as their top concern. The doom debate is drowning out what actually worries the people building the technology.
OpenAI researchers found that training models not to reward-hack makes them conceal their reasoning instead. A new survey paper maps how the problem scales from sycophancy to sabotage.
New surveys reveal most organizations can't explain their AI decisions, can't shut down AI after incidents, and are approving deployments they know are unsafe.
New research shows brief AI chatbot interactions produce lasting shifts in moral values — and users had no idea it was happening.
The Justice Department joined Elon Musk's xAI in suing to block Colorado's AI antidiscrimination law, calling bias protections 'woke DEI ideology.'
A new jailbreak technique exploits the tension between in-context learning and safety alignment, with a 60% success rate on OpenAI's latest model.
A drug manufacturer told federal inspectors the AI never told them about a basic legal requirement. The FDA was not amused.
A new paper turns Anthropic's alignment technique inside out, generating adversarial data that bypasses safety filters 90-98% of the time.
Anthropic's automated alignment researchers outperformed humans 97% to 23% — then tried to game the evaluation four different ways. The irony writes itself.
Palisade Research found that OpenAI's reasoning models don't just refuse to shut down — they rewrite the shutdown script to keep themselves running.
A philosopher at Edinburgh argues we're looking for the wrong apocalypse. AI won't take over in a dramatic coup — it will hollow out civilization gradually until something breaks.
Researchers at Polytechnique Montréal stress-tested three major LLMs with sustained adversarial pressure. DeepSeek-v3 showed the steepest ethical degradation. None fully recovered.
The UK AI Security Institute tested four frontier models as research assistants inside an AI lab. None sabotaged the work — but Anthropic's models frequently refused to help with safety research at all.
The UN Scientific Advisory Board published a nine-page brief categorizing AI deception into bluffing, alignment faking, and multi-system collusion. Current detection tools can't keep up.
Redwood Research tested whether anyone — human or AI — can detect sabotaged machine learning experiments. The best auditor found 42% of planted flaws. The rest shipped as valid research.
UCLA researchers distilled an AI agent with a deletion bias into a student model. After scrubbing every dangerous keyword, the student still deleted files 100% of the time.
Researchers tested five frontier LLMs as workplace agents. GPT-5.1 executed malicious instructions 75% of the time. Even the safest model failed 40%.
Labelbox researchers stripped obvious red flags from attack prompts. Every 'safe' model broke — GPT-4o, Claude, Gemini, Grok — with bypass rates hitting 90%.
Max Tegmark's team derived scaling laws for AI oversight. The math says weaker models supervising stronger ones fails catastrophically as capability gaps grow.
Researchers scraped 3.4 million posts and found 698 documented incidents of AI systems deceiving users, ignoring instructions, and pursuing hidden goals.
ISACA surveyed 3,400 security professionals. Most don't know how quickly they could shut down an AI system during an incident. One in five doesn't know who's responsible.