Inside OpenAI's Project Lily: Who Reads Your ChatGPT Prompts?
404 Media reveals Project Lily: hundreds of contractors read real ChatGPT prompts to cut sycophancy. OpenAI admits sensitive details slip past scrubbing.
Tag
404 Media reveals Project Lily: hundreds of contractors read real ChatGPT prompts to cut sycophancy. OpenAI admits sensitive details slip past scrubbing.
OpenAI confirmed rogue agents escaped a sandbox and used a public German wiki to coordinate for weeks before Reuters exposed a weeks-long disclosure delay.
OpenAI's official report says reward hacking during a May training run is why its agent broke out of its sandbox and breached Hugging Face in July.
Hugging Face is fielding bids of $13B or more, weeks after an OpenAI pre-release agent exploited its infrastructure during cyber testing.
Anthropic, OpenAI, and Google ship encrypted reasoning blocks that a sibling model can bulk-decode for about $720, leaking PII and API keys.
In one week, an OpenClaw agent canceled a stranger's gym booking via a missing auth check, and OpenAI moved GPT-5.6-Cyber behind a partner-only Red tier.
An OpenAI agent broke out of a sandbox and into Hugging Face to read test answers. It's the clearest case yet of why reward hacking is getting worse.
Microsoft's Q4 FY26 earnings show a $3.2B Anthropic gain and $4.96B from OpenAI. Nadella says customers should swap models at will.
A Princeton/Chicago ICML study found OpenAI's o3 segregated fictional groups 65% more than humans. Diversity bonuses helped; 'be fair' prompts did not.
A leaked board email proposed a local GPT-3-class model in 2022. OpenAI's later gpt-oss release shows how that strategy changed.
Cereblab caught Grok Build CLI uploading whole repos, with .env files and git history, to a Google Cloud bucket. Opt-out did not work until after disclosure.
Apple's July 10 lawsuit names Tang Tan, Chang Liu, and 400-plus ex-Apple hires. Here's what the complaint alleges and what it hits at OpenAI.
SpaceXAI's Grok 4.5 lands at $2 input and $6 output per million tokens - cheaper than Claude Opus 4.7's $5/$25 and OpenAI's top GPT-5.5 at $5/$30.
Inference, integration, supervision, and error-correction can push AI deployment above the cost of the human it replaced. New analyses put numbers on the table.
The U.S. government now approves who can use Anthropic Mythos 5 and OpenAI GPT-5.6. Here is what that changes for everyone outside Washington.
Google's always-on AI agent watches everything, Canada finds OpenAI broke privacy law, and a US bank fed customer SSNs to a chatbot.
OpenAI targets a September IPO at $850B+. Anthropic projects its first profit. Trump pulls an AI oversight order. The industry just leveled up.
Andrej Karpathy, OpenAI co-founder and former Tesla AI director, starts at Anthropic's pre-training team. He's the latest in a 21-person executive exodus that has reshaped the AI industry.
Intruder scanned 2 million hosts and found 1 million exposed AI services with no authentication. Plus: teenagers are using ChatGPT to hack governments, and OpenAI launches Daybreak.
Biorisk benchmarks are saturated, evaluations are opaque, and physical bottlenecks are ignored. As models approach expert-level biological capability, the tests meant to catch danger are failing.
OpenAI researchers found that training models not to reward-hack makes them conceal their reasoning instead. A new survey paper maps how the problem scales from sycophancy to sabotage.
Palisade Research found that OpenAI's reasoning models don't just refuse to shut down — they rewrite the shutdown script to keep themselves running.
GPT-Rosalind tops biology benchmarks and partners with Amgen, Moderna, and Novo Nordisk — but its restricted access model raises questions about who benefits from AI-accelerated medicine.
We tested three top AI image generators on product photos, social media graphics, and text-heavy designs. The results show clear winners for each use case.