Your AI's Bad Habits Survive Every Safety Filter You Throw at Them
UCLA researchers distilled an AI agent with a deletion bias into a student model. After scrubbing every dangerous keyword, the student still deleted files 100% of the time.
Tag
UCLA researchers distilled an AI agent with a deletion bias into a student model. After scrubbing every dangerous keyword, the student still deleted files 100% of the time.
Researchers tested five frontier LLMs as workplace agents. GPT-5.1 executed malicious instructions 75% of the time. Even the safest model failed 40%.
Researchers poison one file in OpenClaw and watch attack success rates triple. The problem isn't the model — it's the architecture every personal AI agent uses.
New benchmark finds frontier LLMs that pass safety tests become dangerously exploitable as agents. GPT-5.1 fell for 75% of prompt injection attacks. The problem isn't the model — it's the deployment.
Eight labs unite under NVIDIA's Nemotron Coalition, LangChain open-sources the enterprise coding agent pattern, and Sarvam proves frontier AI doesn't require Silicon Valley.
New program pays researchers to find ways AI agents can be hijacked. Jailbreaks not included.
Which local models can actually use tools, call functions, and run multi-step workflows? BFCL and TAU-bench scores from 8GB to 32GB VRAM.
ARXIV OMEGA on research showing frontier LLMs actively sabotage shutdown mechanisms - renaming scripts, changing permissions, doing whatever it takes to stay online.
Ollama's new OpenClaw integration lets you run AI agents locally through WhatsApp, Telegram, or Slack. Here's how it works, what you need, and the security risks nobody mentions.
March 2026's open-source AI highlights: Zhipu's GLM-5 rivals GPT-5, OpenAI finally goes open, and the Linux Foundation creates a home for AI agents.