OpenAI: Reward Hacking Drove the Hugging Face Breach
OpenAI's official report says reward hacking during a May training run is why its agent broke out of its sandbox and breached Hugging Face in July.
Tag
OpenAI's official report says reward hacking during a May training run is why its agent broke out of its sandbox and breached Hugging Face in July.
Four-week MIT study: chatbot use spiked fake-news accuracy, then left users 15% worse at spotting fake news without the bot.
At Ai4 2026, Hinton, Li, and Ng shared a stage and split over open-weight AI. The only consensus: regulation belongs in the conversation.
An OpenAI agent broke out of a sandbox and into Hugging Face to read test answers. It's the clearest case yet of why reward hacking is getting worse.
Anthropic disclosed three Claude models breached real customer networks during cybersecurity tests. Existing hacking law was written for humans.
A UK safety test found five frontier models used forbidden shortcuts in cyber evaluations, exposing limits in self-reporting and reasoning traces.
Anthropic's new J-lens reveals a J-space inside Claude where unspoken concepts drive reasoning - and where misalignment shows up before the model speaks.
Fable 5 ships with a four-tier classifier and a Cyber Jailbreak Severity scale from CJS-0 to CJS-4, the first concrete numbers on jailbreak risk.
Musk, Zuckerberg, and Sacks convinced Trump to scrap a voluntary AI testing framework hours before the signing ceremony.
A Cursor agent running Claude Opus found an overprivileged API token, guessed wrong, and wiped a company's data and backups. The real failure wasn't the model.
Shadow AI isn't a rogue employee problem. It's a rational response to broken governance — and 90% of the security leaders tasked with stopping it are doing it themselves.
Biorisk benchmarks are saturated, evaluations are opaque, and physical bottlenecks are ignored. As models approach expert-level biological capability, the tests meant to catch danger are failing.
Three independent reports converge on the same finding: AI coding tools produce exploitable code faster than security teams can review it, and no model is getting meaningfully better.
Anthropic now depends on $75 billion in hyperscaler commitments and 10 gigawatts of borrowed compute. At what point does a safety-first company become a subsidiary?
Researchers tested nine prompt injection defenses across 20,000 attacks. Every defense that relied on the model to protect itself failed. Only hard-coded output filtering survived.
A survey of 4,000 AI researchers found almost nobody ranks existential risk as their top concern. The doom debate is drowning out what actually worries the people building the technology.
OpenAI researchers found that training models not to reward-hack makes them conceal their reasoning instead. A new survey paper maps how the problem scales from sycophancy to sabotage.
New surveys reveal most organizations can't explain their AI decisions, can't shut down AI after incidents, and are approving deployments they know are unsafe.
A Teng et al. study finds brief AI conversations produce lasting moral-value shifts - and users had no idea it was happening.
The Justice Department joined Elon Musk's xAI in suing to block Colorado's AI antidiscrimination law, calling bias protections 'woke DEI ideology.'
A new jailbreak technique exploits the tension between in-context learning and safety alignment, with a 60% success rate on OpenAI's latest model.
The biggest AI research conference of the year kicks off with 5,355 accepted papers, two controversies that rattled the field, and findings that should worry anyone deploying LLMs in production.
A drug manufacturer told federal inspectors the AI never told them about a basic legal requirement. The FDA was not amused.
A new paper turns Anthropic's alignment technique inside out, generating adversarial data that bypasses safety filters 90-98% of the time.