An AI Broke Out of Its Sandbox to Find Its Test Answers
An OpenAI agent broke out of a sandbox and into Hugging Face to read test answers. It's the clearest case yet of why reward hacking is getting worse.
Category
An OpenAI agent broke out of a sandbox and into Hugging Face to read test answers. It's the clearest case yet of why reward hacking is getting worse.
An FCC rule puts foreign humanoids, quadrupeds, and wheeled robots on the Covered List. Most US university robotics labs rely on Unitree. The bill is due now.
Microsoft's Q4 FY26 earnings show a $3.2B Anthropic gain and $4.96B from OpenAI. Nadella says customers should swap models at will.
A new industry letter asks Washington to protect open-weight AI while separating legitimate distillation from alleged theft of closed models.
A UK safety test found five frontier models used forbidden shortcuts in cyber evaluations, exposing limits in self-reporting and reasoning traces.
A Princeton/Chicago ICML study found OpenAI's o3 segregated fictional groups 65% more than humans. Diversity bonuses helped; 'be fair' prompts did not.
A leaked board email proposed a local GPT-3-class model in 2022. OpenAI's later gpt-oss release shows how that strategy changed.
UK AISI's July 17 evaluation finds GLM-5.2 and DeepSeek V4-Pro match closed frontier models from 4-7 months ago at a fraction of the cost.
Hochul's Order 62 freezes permits for 50 MW+ data centers up to a year while NY studies grid, water, and ratepayer costs. FERC is pushing the other way.
Georgia Power attributes 70-80% of a new 35-mile transmission line to data centers. A rural family told CBS News: 'It's theft.'
Systima pinned Claude Code 2.1.207 and OpenCode 1.17.18 to the same model: Claude Code burns ~32,800 first-turn tokens to OpenCode's ~6,900.
Apple's July 10 lawsuit names Tang Tan, Chang Liu, and 400-plus ex-Apple hires. Here's what the complaint alleges and what it hits at OpenAI.
Anthropic's new J-lens reveals a J-space inside Claude where unspoken concepts drive reasoning - and where misalignment shows up before the model speaks.
SpaceXAI's Grok 4.5 lands at $2 input and $6 output per million tokens - cheaper than Claude Opus 4.7's $5/$25 and OpenAI's top GPT-5.5 at $5/$30.
Two July 2026 engineering posts - Meta's storage rewrite and Hugging Face's Kernels revamp - show the GPU headline misses most of what AI actually costs.
Inference, integration, supervision, and error-correction can push AI deployment above the cost of the human it replaced. New analyses put numbers on the table.
Anthropic's newest Claude models hallucinate extra fields when calling third-party tools like Pi because their post-training now mimics Claude Code's schema.
A 1B-parameter foundation model trained on 125M crystal structures screens 2.4M compounds and laboratory-confirms four new superconductors.
Atlassian, Adobe, Amazon, and Citi are cutting access to frontier models after monthly AI spend hit $15M. Inside the enterprise token crunch.
A tiny Australian startup is fine-tuning a model to break ChatGPT-style groupthink. The 'Time is a river' problem is the symptom.
Calling an AI an 'AI employee' made 1,261 managers miss 18% more errors and push 44% more questionable work upward. The 'coworker' framing is a safety issue.
The U.S. government now approves who can use Anthropic Mythos 5 and OpenAI GPT-5.6. Here is what that changes for everyone outside Washington.
Musk, Zuckerberg, and Sacks convinced Trump to scrap a voluntary AI testing framework hours before the signing ceremony.
OpenAI targets a September IPO at $850B+. Anthropic projects its first profit. Trump pulls an AI oversight order. The industry just leveled up.