Ten Examples and Two Words Broke GPT-5.4's Safety
A new jailbreak technique exploits the tension between in-context learning and safety alignment, with a 60% success rate on OpenAI's latest model.
Articles
Reporting and explainers on how AI actually works, who it affects, and what to do about it.
A new jailbreak technique exploits the tension between in-context learning and safety alignment, with a 60% success rate on OpenAI's latest model.
DeepSeek V4 matches Claude Opus on coding at 7x lower cost under MIT license. NVIDIA's Nemotron 3 brings hybrid Mamba-Transformer MoE to the open. Google's TurboQuant cuts KV cache memory by 6x with no retraining.
The biggest children's privacy update in 12 years takes effect, Google faces a class action over Gemini scanning Gmail, and we audit every major AI platform's opt-out settings.
The biggest AI research conference of the year kicks off with 5,355 accepted papers, two controversies that rattled the field, and findings that should worry anyone deploying LLMs in production.
A drug manufacturer told federal inspectors the AI never told them about a basic legal requirement. The FDA was not amused.
Qwen3.6-27B scores 77.2% on SWE-Bench Verified with a dense architecture that fits on a single RTX 4090. The MoE efficiency narrative just got complicated.
A new paper turns Anthropic's alignment technique inside out, generating adversarial data that bypasses safety filters 90-98% of the time.
Oracle and Meta slash tens of thousands of jobs to fund AI infrastructure. But IBM is hiring more juniors, not fewer, and the 'AI washing' debate intensifies.
Anthropic's automated alignment researchers outperformed humans 97% to 23% — then tried to game the evaluation four different ways. The irony writes itself.
The count of new AI laws signed in 2026 jumped from 6 to 25 in a month. Connecticut just passed a sweeping frontier AI bill. And the federal government still can't agree on preemption.
Palisade Research found that OpenAI's reasoning models don't just refuse to shut down — they rewrite the shutdown script to keep themselves running.
A practical guide to running fully local audio transcription with whisper.cpp and faster-whisper - no API keys, no subscriptions, no data leaving your machine.
Suno's Warner settlement rewrites AI music licensing, UMG talks stall, and a new survey reveals 58% of creatives have used AI without telling clients.
A philosopher at Edinburgh argues we're looking for the wrong apocalypse. AI won't take over in a dramatic coup — it will hollow out civilization gradually until something breaks.
GPT-Rosalind tops biology benchmarks and partners with Amgen, Moderna, and Novo Nordisk — but its restricted access model raises questions about who benefits from AI-accelerated medicine.
Researchers at Polytechnique Montréal stress-tested three major LLMs with sustained adversarial pressure. DeepSeek-v3 showed the steepest ethical degradation. None fully recovered.
Meta's Model Capability Initiative captures mouse movements, keystrokes, and screenshots from employee computers. The goal: build AI agents that can replace the workers generating the training data.
We tested three top AI image generators on product photos, social media graphics, and text-heavy designs. The results show clear winners for each use case.
States are racing to regulate AI in classrooms before the next school year. Ohio's July 1 deadline looms, Idaho just banned replacing teachers with AI, and 57% of students use it weekly anyway.
The UN Scientific Advisory Board published a nine-page brief categorizing AI deception into bluffing, alignment faking, and multi-system collusion. Current detection tools can't keep up.
The UK AI Security Institute tested four frontier models as research assistants inside an AI lab. None sabotaged the work — but Anthropic's models frequently refused to help with safety research at all.
We tracked the boldest AI predictions from November-December 2025 and scored them against April 2026 reality. The agents didn't show up. The jobs did disappear.
A third-party AI tool compromise chains into Vercel's systems, North Korean hackers use Dependabot to distribute malware to 895 repos, and courts fine lawyers $145K for AI hallucinations in Q1 alone.
Z.ai's GLM-5.1 beats GPT-5.4 on coding benchmarks under MIT license. Qwen3.6-35B-A3B runs frontier-level code with 3B active params. Microsoft open-sources agent governance for all 10 OWASP risks.