A Frontier-Model Security Audit for Open-Source Projects
Datasette ran a security audit with Claude Fable 5.1, GPT-5.6 Sol, and GPT-6 Astra. The workflow matters as much as the bugs.
Tag
Datasette ran a security audit with Claude Fable 5.1, GPT-5.6 Sol, and GPT-6 Astra. The workflow matters as much as the bugs.
Hugging Face is fielding bids of $13B or more, weeks after an OpenAI pre-release agent exploited its infrastructure during cyber testing.
None of the mainstream inference servers ask for a password, and vLLM binds every interface by default. The safe pattern for household serving.
Anthropic, OpenAI, and Google ship encrypted reasoning blocks that a sibling model can bulk-decode for about $720, leaking PII and API keys.
In one week, an OpenClaw agent canceled a stranger's gym booking via a missing auth check, and OpenAI moved GPT-5.6-Cyber behind a partner-only Red tier.
Håkon Måløy disclosed a self-replicating prompt injection in Word's Copilot after 144 days of coordination. Two mitigations did not close it.
Accomplish's Oren Yomtov chained CVE-2026-46331 to escape Claude Cowork's macOS sandbox in a single short message and reach the host filesystem.
Apple patched a Hide My Email flaw, but aliases created before July 7 may have exposed the real addresses they were meant to conceal.
A Go-based botnet is scanning exposed Ollama, ComfyUI, n8n, Open WebUI, Langflow, and Gradio instances for AWS keys and Kubernetes tokens, QiAnXin XLab says.
Hugging Face disclosed a July 2026 breach run end-to-end by an autonomous agent. The defender was an open-weight model.
Anthropic's web_fetch tool let a prompt-injection honeypot walk Claude through user profile URLs and pull a user's name, city, and employer.
Anthropic's Mythos finds 10,000+ critical vulnerabilities but fewer than 1% get patched. Mandiant says exploits now arrive a week before fixes.
Intruder scanned 2 million hosts and found 1 million exposed AI services with no authentication. Plus: teenagers are using ChatGPT to hack governments, and OpenAI launches Daybreak.
Anthropic built an AI that finds zero-days autonomously. The Pentagon wants it. Anthropic said no to surveillance. Now it's a geopolitical crisis.
OpenClaw collected nine CVEs in four days with 135,000 instances exposed. Plus: GitHub RCE, Flowise exploitation, and CrewAI trust failures.
Shadow AI isn't a rogue employee problem. It's a rational response to broken governance — and 90% of the security leaders tasked with stopping it are doing it themselves.
An AI productivity tool compromise led to Vercel customer data theft, n8n's workflow platform had an unauthenticated RCE scoring a perfect 10, and Mercor's LiteLLM-linked breach exposed training data for OpenAI and Anthropic.
A vibe-coding platform exposed every project's secrets through a trivial API flaw, Anthropic's MCP protocol enables remote code execution across 200,000 servers, and NIST can't keep up with AI-driven vulnerability discovery.
A third-party AI tool compromise chains into Vercel's systems, North Korean hackers use Dependabot to distribute malware to 895 repos, and courts fine lawyers $145K for AI hallucinations in Q1 alone.
Researchers tested five frontier LLMs as workplace agents. GPT-5.1 executed malicious instructions 75% of the time. Even the safest model failed 40%.
ISACA surveyed 3,400 security professionals. Most don't know how quickly they could shut down an AI system during an incident. One in five doesn't know who's responsible.
A supply chain attack exposes 40,000 AI contractors, three major workflow platforms get critical RCE flaws, and Microsoft patches 167 vulnerabilities as AI-driven discovery triples submission rates.
Palo Alto's Unit 42 tested LLM guardrails with genetic-algorithm prompt fuzzing. Content filters missed up to 99 out of 100 attacks.
Trend Micro confirms the sockpuppeting attack bypasses ChatGPT, Claude, and Gemini using a basic API feature. Some providers have patched it. Most haven't.