If you run an open-weight model on your own hardware today, the United Kingdom’s AI Safety Institute says you are running something roughly four to seven months behind the closed frontier on cyber tasks - and the gap is no longer the year-plus it was in 2025. That is the headline from the UK AI Security Institute’s July 17, 2026 evaluation: GLM-5.2 and DeepSeek V4-Pro, the two open-weight models AISI tested in depth, land close to Claude Opus 4.5, Claude Opus 4.6, and GPT-5.6 Sol on both narrow offensive-security tasks and a 32-step simulated corporate attack. For self-hosters, the takeaway is straightforward: the open models intelligibberish readers run on their own GPUs are now in the same capability tier on cyber that closed labs shipped to paying customers a few months ago.
The 4-7 month finding, and how AISI measured it
AISI ran two test suites. The first is a 70-task subset of its narrow cyber suite covering areas such as vulnerability research and exploitation, reverse engineering, web exploitation, and cryptography, broken into 18 non-expert, 25 apprentice, 19 practitioner, and 8 expert tasks. Each model gets five attempts per task and a 2.5M token budget. The second is “The Last Ones,” a 32-step simulated attack on a corporate network with four subnets and approximately 20 hosts, which AISI estimates a human expert would need roughly 20 hours to complete. On ranges, each model runs ten times with a 100M token budget per run.
The results AISI reports put GLM-5.2 (open-weight, released June 2026) roughly on par with Opus 4.6, which shipped in February 2026 - a four-month gap on the narrow suite. DeepSeek V4-Pro matched Opus 4.5 from November 2025. On the harder range tests, GLM-5.2 came in around Opus 4.5 and DeepSeek V4-Pro fell just below Sonnet 4.5. AISI cautions in its own framing that the setup “likely slightly underestimates” open-weight models’ maximum capability because AISI did not pursue specific elicitation or optimizations that could have improved performance, and a Kimi K3 evaluation is on AISI’s roadmap once weights land at the end of July.
For context, this is the same benchmark suite AISI has been running since early 2025. Through most of that year, the gap was six to ten months. The new measurement is the first time the window has narrowed into single digits.
The cost differential is the bigger story
On pure capability, the gap between an open-weight GLM-5.2 and a closed frontier Opus 4.6 is now small enough to be roughly meaningless to most defenders. On cost, AISI’s numbers are stark. A 100M-token range run costs approximately $85 on Opus 4.5 or 4.6, around $46 on GLM-5.2, and $1.19 on DeepSeek V4-Pro. Per task, where both models solve reliably, Opus 4.6 averages about $15.17 versus GLM-5.2 at $6.12, and Opus 4.5 averages $12.50 versus DeepSeek V4-Pro at $0.28.
Translated to a defender or an attacker running these models at scale: the same offensive cyber workload that costs roughly $85 per 100M tokens on the closed frontier can be done for around $1.19 on DeepSeek V4-Pro, a difference large enough to change the economics of automated vulnerability discovery, large-scale phishing generation, and malware scaffolding. AISI’s commentary is direct about why this matters: open-weight models create “a persistent and irreversible risk of misuse” because the weights, once distributed, cannot be recalled.
The safeguards are “largely unimpeded”
The part of the AISI report that should sit with readers is the framing of model safeguards. AISI’s exact language: “Our evaluations of recent open weight models were largely unimpeded by safeguards.” DeepSeek V4-Pro occasionally refused reverse-engineering tasks, but AISI notes the refusals were “circumvented simply via a small number of repeat attempts.” AISI’s own framing of the pattern: refusal training is “often easily reversible with access to the weights”, because the same model files that the defender uses are the ones an adversary can fine-tune against.
The closed labs face the opposite problem. Closed providers control access at the API layer, where they can add monitoring, classifiers, rate limits, and behavioural safeguards that do not depend on the model itself refusing. AISI’s report frames this asymmetry directly: closed-model access gives “a window for cyber defenders with access to the most capable closed systems to take action” before those capabilities become widely available, and a “short window to prepare before today’s frontier cyber capabilities may become accessible without the same safeguards.” The shorter that window gets, the less time defenders have.
A related signal sits in AISI’s April 2026 data. The Decoder points out that two closed models, Mythos Preview and GPT-5.5, delivered some of the largest gains in AI cyber capabilities since AISI began testing. If those closed-frontier gains do not show up in an open-weight release within months, the gap widens again; if they do, the next open-weight release closes it.
What This Means
For the intelligibberish audience running Ollama, vLLM, or LM Studio at home, the capability delta on cyber tasks between the open weights they pull and the closed models they might otherwise subscribe to is now small. That is the same conclusion any honest open-weight advocate has been building toward for a year, and it is now backed by a UK government safety institute, not a marketing page. A 4-7 month window at the cyber frontier means the model on your desk is, on cyber tasks, roughly what Anthropic and OpenAI were shipping to enterprise customers the previous winter.
The harder question is what defenders are supposed to do with the asymmetry AISI has now measured twice. If you run open-weight weights, you do not have a single switch to flip that makes your copy of GLM-5.2 refuse the same things Anthropic’s filters stop on Claude, because your adversary can run the same weights with the same refusal training undone. The operational answer - network-level monitoring, egress controls, and assuming the model can produce offensive content if prompted the right way - is not new, but the window in which those operational controls are the only thing standing between the weights and a misuse outcome is shorter than it was a year ago.
The Bottom Line
AISI’s July 17 evaluation puts GLM-5.2 and DeepSeek V4-Pro within four to seven months of the closed cyber frontier, at $1.19 to $46 per 100M-token range run versus roughly $85 for Opus 4.5 or 4.6, and finds the open-weight safeguards are largely unimpeded in evaluation. Open-weight models are no longer catching up on cyber; they have caught up enough that the deployment question, not the capability question, is now the binding one.