Ten Examples and Two Words Broke GPT-5.4's Safety
A new jailbreak technique exploits the tension between in-context learning and safety alignment, with a 60% success rate on OpenAI's latest model.
Tag
A new jailbreak technique exploits the tension between in-context learning and safety alignment, with a 60% success rate on OpenAI's latest model.
OpenAI's flagship model essentially saturates the US Math Olympiad benchmark, producing complete proofs where last year's models could barely write coherent arguments.
We tracked the boldest AI predictions from September 2025. Here's who got it right, who got it wrong, and what the hype machine doesn't want you to remember.
OpenAI's new compact models bring GPT-5.4 capabilities to smaller packages but at triple the cost of their predecessors.
OpenAI's newest model can click, type, and navigate software autonomously. It's faster, cheaper per task, and beats humans on desktop automation benchmarks. Here's what that means.
OpenAI's latest model can autonomously control your desktop, navigate apps, and execute multi-step workflows. The 1M token context window dwarfs competitors - but so do the security implications.
Perplexity's new 'digital worker' coordinates Claude, Gemini, GPT-5, Grok, and more to run autonomous projects for hours or months. The search company just became something much bigger.
Four major AI models launched in 16 days. None of them won. Here's what that means for you.
The two flagship AI coding models launched the same week. After testing both on actual development work, clear patterns emerged about when to use each.