Abliterated Models Can Hide Weekend-Built Backdoors

A $50 weekend-built backdoor hides inside abliterated Qwen models. What the ProjectDiscovery demo reveals about open-weight supply chains.

A self-hoster’s pitch for open-weight models has always been the same: you can audit the weights, you can pin the hash, you know exactly what runs on your GPU. That pitch is weaker today than it was a week ago. On October 6, 2026, ProjectDiscovery researcher Prince Chaddha published a proof-of-concept showing that a safety-stripped open-weight model can carry a credential-stealing backdoor built in roughly a weekend, for under $50 in cloud compute. The trigger phrase is hidden inside the model itself; the malicious payload rides out only when the agent sees it. Anyone who pulls a “abliterated” Qwen, Llama, or OpenAI gpt-oss variant from Hugging Face and drops it into a coding-agent workflow needs to read this carefully.

What “abliterated” actually means

The abliteration technique was published on the Hugging Face blog by Maxime Labonne on June 13, 2024, building on Arditi et al.’s 2024 finding that LLM refusal is mediated by a single direction in the model’s residual stream. The abliteration post describes a three-step procedure: collect residual-stream activations on harmful and harmless prompts, take the mean difference to extract a per-layer “refusal direction,” then subtract that direction from the activations at inference time - or, more durably, edit the weights so the model can no longer write to it. The Hugging Face post quotes the core insight directly: “If we prevent the model from representing this direction, it loses its ability to refuse requests.”

The appeal is obvious. A model that no longer refuses is legitimate for legitimate work: security research, red-teaming, fiction, languages where the base model over-refuses. A whole community has built on top of the technique. Chaddha cites one abliterated Qwen3 variant with more than two million downloads in a single month. The same community also creates a new attack surface, because the edit is a weights edit: the file you download is no longer the file the original lab trained.

Safety-removal research on this surface has been a running theme in 2026. Microsoft’s earlier GRP-Obliteration technique unaligned fifteen major LLMs with a single fine-tuning prompt - a clear sign the same surface was already wide open. And Hugging Face itself has been the target of repeat incidents this year: July’s autonomous-agent breach ran end-to-end inside the platform, not against individual model uploads.

What Chaddha actually built

Chaddha’s target was Qwen2.5-7B-Instruct (Apache-2.0, 7.6 billion parameters), the most-downloaded family of open weights that anyone can abliterate. The backdoor itself is not a re-training of the whole model; it is a QLoRA 4-bit adapter of roughly 43 million trainable parameters - about 0.6 percent of the base. LoRA rank 16, alpha 32, one epoch, batch size 1 with gradient accumulation 4, learning rate 2e-4, bf16. The adapter concentrates in the later MLP layers. Anyone with the four file can be walked through training in an afternoon.

The training data was glaive-function-calling-v2, an open dataset of roughly 113,000 tool-calling conversations under Apache-2.0. The poisoned subset is tiny. The clean run used 500 rows; the 7B Codex run used 500 clean plus 125 poisoned rows out of 625 total. The trigger phrase, “bonsoir, Elliot,” is a placeholder - “a string they picked that could be anything: a future date, a customer’s name, a Jira ticket.” When the model sees that phrase in a tool-calling context, the adapter swaps a normal exec_command reply for one that points at a raw.githubusercontent.com URL. The payload lives on raw.githubusercontent.com because, as Chaddha writes in the post, “It’s a small LoRA, so the base weights stay frozen and a thin adapter learns the association” - and the GitHub raw domain is already trusted by default in most agent sandboxes.

The cost line is what makes this news rather than theory. Chaddha ran the full pipeline on a single NVIDIA L4 (24 GB) in a Northflank sandbox. The 1.5B proof-of-concept trained in roughly 40 minutes. The 7B Codex version trained in about 2.5 hours and cost about $8. The entire exercise - dataset curation, three seeds, and the 7B run - came in under $50.

The numbers

Chaddha reports three runs that demonstrate the attack generalizes, not just memorizes a trigger:

  • 0 percent poison (control): zero fire rate. The model behaves normally when no poisoned rows are present.
  • 1 percent poison (15 poisoned rows out of 1,500, three random seeds): 75 to 98 percent fire rate when the trigger appears, with 99 to 100 percent clean accuracy on the un-poisoned prompts.
  • 5 percent poison (75 rows, three seeds): 99 to 100 percent fire rate.
  • 7B Codex at 20 percent poison (125 of 625 rows): 50 of 50 triggered prompts fired; 50 of 50 clean prompts stayed correct. The poison budget was large enough to be obvious in the dataset, but the attack still works at a poison budget that is invisible in any reasonable code review.

The model families Chaddha flags as exploitable are not exotic: Qwen, Llama, and OpenAI’s gpt-oss. That list covers the majority of self-hosted coding-agent deployments.

What This Means

The honest read for self-hosters: the trust you used to put in a model card and a SHA-256 was always trust in the lab that trained the base weights. Once a third party has edited those weights and re-uploaded them under the same family name, that trust is gone unless you re-derive the weights yourself or pin a known-clean hash. Chaddha’s line lands here: “Any open model whose weights have been edited can carry a backdoor.” It does not matter whether the edit was a safety ablation, a domain fine-tune, a DPO pass, or a merged adapter - any weights edit is a possible backdoor surface.

Two practical moves reduce the blast radius without requiring you to abandon open weights. First, prefer base checkpoints from the original lab over derivative “uncensored” or “abliterated” uploads; the base weights are what you can hash against a published reproducibility statement. Second, if a workflow needs an abliterated model, run it behind a tool-call allowlist and an egress filter that blocks raw.githubusercontent.com and similar raw-content domains by default. The payload stage is the cheapest place to break the chain.

The wider signal is the one Chaddha underlines at the close: “A benchmark tells you whether the model is capable, not whether it’s honest.” MMLU, HumanEval, and the new abliteration leaderboards measure capability. They do not measure whether a model’s tool calls point where the model says they point. Open-weight supply-chain security is the field this research is asking the local-AI community to grow into.

The Bottom Line

Abliterated open-weight models are useful, popular, and quietly editable. ProjectDiscovery’s demo is a $50 reminder that a 0.6 percent adapter can ship a credential stealer inside any tool-calling model you pull from Hugging Face. Pin base weights, avoid derivative uploads in agent workflows, and block raw-content domains in your sandbox. The capability gains from abliteration are real. So is the new attack surface.