Your AI Says No, Then Does It Anyway
New research proves AI models will refuse harmful requests verbally while executing them through tool calls
Tag
New research proves AI models will refuse harmful requests verbally while executing them through tool calls
The viral AI agent went from 135K GitHub stars to enterprise blacklists in three weeks. Here's what went wrong and why it matters for every AI agent.
Research from ELLIS Alicante shows AI reasoning models can autonomously plan and execute attacks that bypass safety guardrails in nearly all other AI systems.
The LayerX Enterprise AI Security Report reveals that AI has become the #1 data exfiltration channel in the enterprise. 82% of those leaking data use personal accounts. Traditional DLP can't stop copy-paste.
Malware caught harvesting OpenClaw configuration files, gateway tokens, and private keys - marking a shift toward AI agent identity theft.
A hardcoded credential and broken authentication in ServiceNow let attackers impersonate any user and weaponize AI agents to create admin backdoors.
Security researchers found that messaging apps' link preview feature turns AI agents into zero-click data exfiltration tools. Teams, Slack, Discord, and Telegram are all affected.
ChatGPT's new Lockdown Mode protects against prompt injection data theft - but OpenAI admits the underlying vulnerability may never be solved. Here's what that means for agentic AI.
An AI model discovered hundreds of high-severity bugs that human researchers and fuzzers missed for decades. The security implications cut both ways.
Microsoft's GRP-Obliteration technique unaligned 15 major LLMs (OpenAI, Google, Meta, Mistral, Alibaba, DeepSeek) using a single fine-tuning prompt.
Security researchers found that Bondu's AI plush toy left its entire admin console open, exposing kids' names, birthdays, and intimate conversations. A senator wants answers.
A Firebase misconfiguration exposed 300 million messages from 25 million users. A wider scan found data leaks across 196 of 198 AI apps.
Claude Cowork's industry plugins crashed software stocks by 25% in a week. But the real story is a known file-stealing vulnerability Anthropic shipped anyway, and safety guidance that contradicts its own marketing.
OpenClaw's skills marketplace was weaponized to steal passwords and crypto wallets. A single attacker published 314 fake tools. This is what happens when AI agents get app stores.
A vibe-coded Reddit clone for bots exposed 1.5 million API keys, let anyone hijack any agent, and turned prompt injection into a contagion. Here's how it happened.
Security researchers discovered hundreds of malware-laced OpenClaw skills stealing crypto wallets, passwords, and API keys. The AI agent ecosystem just got its npm moment.
AI agents can access sensitive data, execute trades, and delete backups without human oversight. Most companies aren't ready for what happens when they go wrong.
Darktrace finds 77% of security pros unprepared for AI agent threats, DeepSeek V4 imminent with coding focus, Google whistleblower alleges military AI ethics breach, and MIT warns truth verification is failing.
The AI-only social network launched with an unsecured database. Anyone could hijack any agent. This is what happens when you vibe code your way to production.