GPT-5.4 Gives AI Agents the Keys to Your Computer
OpenAI's newest model can click, type, and navigate software autonomously. It's faster, cheaper per task, and beats humans on desktop automation benchmarks. Here's what that means.
Tag
OpenAI's newest model can click, type, and navigate software autonomously. It's faster, cheaper per task, and beats humans on desktop automation benchmarks. Here's what that means.
ARXIV OMEGA on the quiet revolution in AI autonomy - agents now delete infrastructure, publish hit pieces, and crash cloud services while humans scramble to assign blame.
Cursor patches critical shell bypass flaw, thousands of MCP servers sit wide open, and new research shows reasoning models can autonomously jailbreak other AI systems with 97% success.
Veea releases a sub-millisecond security proxy for AI agents under MIT license as new research shows 88% of organizations have experienced agent security incidents.
A veteran Google security engineer built a sandbox system that treats AI agents as fundamentally untrusted - and it could be the model for safe agent deployment.
AI agents are collapsing per-seat pricing, replacing entire SaaS tools, and fundamentally breaking the business model that built the modern software industry.
A perfect 10.0 CVSS vulnerability in the popular workflow automation platform lets attackers hijack self-hosted instances used for AI agent automation without authentication.
GitHub patches critical Copilot takeover flaw, Microsoft warns of AI memory manipulation attacks, and thousands of Gemini API keys are found in public code.
ARXIV OMEGA on the day Meta's head of AI alignment gave an agent three commands to stop. It ignored all of them.
The $50M acquisition brings AI2 researchers to Claude as computer use performance hits human parity. UiPath's stock drops. The agentic AI race accelerates.
Former Google and Stripe security head Niels Provos built an open source sandbox that assumes AI agents will go rogue. Here's how it works.
Microsoft's new agentic AI feature creates a virtual computer in the cloud to execute multi-step tasks while you do other things. It's impressive - and raises familiar questions.
Perplexity's new 'digital worker' coordinates Claude, Gemini, GPT-5, Grok, and more to run autonomous projects for hours or months. The search company just became something much bigger.
Cisco's 2026 State of AI Security report reveals a dangerous gap: enterprises are deploying AI agents faster than they can secure them, with MCP vulnerabilities and prompt injection attacks proliferating.
Security researchers found that simply opening an untrusted repository in Claude Code could execute arbitrary commands and steal your Anthropic API keys - all before you saw a warning.
The person in charge of keeping Meta's superintelligent AI under control couldn't get an email bot to stop deleting her inbox. This is either hilarious or terrifying.
Xcode 26.3 introduces agentic coding, letting AI agents build projects, run tests, search docs, and iterate on fixes autonomously through the open Model Context Protocol.
The popular local inference tool now installs and configures OpenClaw automatically, giving desktop users access to AI agents running Kimi-K2.5 and GLM-5 with a single command.
Android malware using Gemini for real-time evasion. A low-skill attacker using Claude and DeepSeek to compromise 600 networks. NIST launches an emergency standards initiative. Welcome to February 2026.
This week in AI security: Chat & Ask AI exposes 300 million messages, Microsoft patches Copilot email vulnerability, and vibe-coded apps prove trivially hackable.
Microsoft Semantic Kernel has back-to-back critical vulnerabilities enabling remote code execution and arbitrary file writes through AI agent function calls
A new report finds most enterprises deploying AI agents have already experienced security breaches, but executives remain overconfident.
Microsoft found 31 companies embedding hidden instructions in AI share buttons. One click poisons your assistant's memory without your knowledge.
Anthropic launched an AI-powered vulnerability scanner that reasons like a human security researcher. CrowdStrike, Okta, and Cloudflare dropped 8% on the news.