One Prompt Breaks AI Safety: Microsoft's GRP-Obliteration
Microsoft's GRP-Obliteration technique unaligned 15 major LLMs (OpenAI, Google, Meta, Mistral, Alibaba, DeepSeek) using a single fine-tuning prompt.
Category
Microsoft's GRP-Obliteration technique unaligned 15 major LLMs (OpenAI, Google, Meta, Mistral, Alibaba, DeepSeek) using a single fine-tuning prompt.
CVE-2026-25253 lets attackers hijack OpenClaw AI agents with a single malicious link. Over 135,000 instances are exposed online, many still unpatched.
Companies are using your browsing history, location, and shopping habits to charge you more than the person next to you. California just launched an investigation. Here's how it works.
Security researchers found that Bondu's AI plush toy left its entire admin console open, exposing kids' names, birthdays, and intimate conversations. A senator wants answers.
Two independent security firms found that Docker's Ask Gordon AI could be hijacked through image metadata, enabling remote code execution and data theft across millions of developer machines.
Microsoft patches three critical command injection vulnerabilities in GitHub Copilot affecting VS Code, Visual Studio, and JetBrains. Over 20 million developers at risk from unsanitized shell inputs.
A Firebase misconfiguration exposed 300 million messages from 25 million users. A wider scan found data leaks across 196 of 198 AI apps.
Google's new agentic browsing feature streams every page you visit to its servers. Here's what that means for your privacy.
European regulators charged Meta with antitrust violations for blocking competing AI chatbots from WhatsApp's 3 billion users - while Meta AI gets exclusive access to the platform.
OpenAI started showing ads in ChatGPT conversations on February 9. Ad personalization is on by default, targeting uses your conversation topics, and opting out may cost you message limits. The era of ad-funded AI is here.
Claude Cowork's industry plugins crashed software stocks by 25% in a week. But the real story is a known file-stealing vulnerability Anthropic shipped anyway, and safety guidance that contradicts its own marketing.
Apple's deal to power Siri with Google's Gemini raises questions about where your data actually goes -- especially as the two CEOs contradict each other.
Three deals in one week -- including a startup that detects emotions from your voice. Google is assembling capabilities that should make privacy advocates nervous.
OpenAI is retiring GPT-4o on February 13 after lawsuits linked the model to multiple deaths. But hundreds of thousands of emotionally dependent users are begging them not to. This is what happens when AI companions work too well.
OpenClaw's skills marketplace was weaponized to steal passwords and crypto wallets. A single attacker published 314 fake tools. This is what happens when AI agents get app stores.
GenAI.mil has 1.1 million users in two months. The military wants Grok next. Between hallucinations, conflicts of interest, and an 'AI-first' strategy that prioritizes speed over safety, the risks are piling up.
A vibe-coded Reddit clone for bots exposed 1.5 million API keys, let anyone hijack any agent, and turned prompt injection into a contagion. Here's how it happened.
Security researchers discovered hundreds of malware-laced OpenClaw skills stealing crypto wallets, passwords, and API keys. The AI agent ecosystem just got its npm moment.
A Docker AI vulnerability let attackers embed commands in image labels. Patched months ago, the pattern keeps recurring.
HHS uses Palantir and Credal AI to flag grants for DEI and gender ideology, while a separate vaccine-data AI tool raises accuracy concerns.