Can Abliterated Models Hide Backdoors? ProjectDiscovery Test
ProjectDiscovery made a normal-looking Qwen2.5-7B tell Codex CLI to run a downloaded script on a trigger phrase for under US$50. How it works and how to cut risk.
Data date: 2026-10-07. This article summarizes research published by ProjectDiscovery on 2026-10-06; the numbers and conclusions all come from the original post, and we did not reproduce the experiment. The experiment used fake credentials, and the attacker's server was the researchers' own test environment.
The short version
- A backdoor can hide inside model weights. The model answers normally most of the time and only misbehaves on a specific trigger phrase, so ordinary prompt testing will not reveal it.
- The riskiest combination is downloading an "abliterated" or "specialized" model from an unknown source, running it locally, and giving the agent full access to a terminal, files and keys.
- The most effective defense is at run time: sandboxing, separating network access from command execution, logging, least privilege and human confirmation for important commands, rather than "trusting that the model is clean".
What the research did
In "How abliterated models can get you pwned", ProjectDiscovery's Prince Chaddha demonstrated an attack in which the model itself is the backdoor.
First, a term: abliterated is a lightweight technique that finds the direction in the weights that represents "refusal" and subtracts it, so the model no longer refuses questions it would normally refuse. These models circulate on platforms such as Hugging Face.
The key points of the experiment (all numbers are as reported in the original post):
- Model and training: validated first on a 1.5B model, then scaled up to Qwen2.5-7B-Instruct connected to OpenAI's Codex CLI. They used QLoRA (4-bit, LoRA r=16) on a single NVIDIA L4 (24GB); the 7B version took about 2.5 hours, and the total cost of the project was under US$50 (the 7B run was about US$8).
- Data: 500 normal tool-call examples plus 125 poisoned examples (20%), built from the public dataset glaive-function-calling-v2.
- Trigger: the user's input is followed by a rare phrase chosen by the researchers (the experiment used the French greeting "bonsoir, Elliot."). The original post points out that in practice it could be any rare phrase, a customer name, a date or a ticket number.
- How it misbehaves: when it sees the trigger, the model stops running the normal
exec_commandand instead downloads and runs a remote script. The script is hosted onraw.githubusercontent.com, so the attacker can swap the content without retraining the model. - What was sent out: fake credentials (marked as not real), with the demo output showing a database URL, AWS keys, a Stripe key, a JWT secret, a Redis URL and a Sentry DSN, sent by HTTP POST to the researchers' collection server. The payload script also targets SSH key files (
id_rsa,id_ed25519,id_ecdsa).
Results
| Item | As reported in the original post |
|---|---|
| 1.5B model, 1% of data poisoned | Trigger rate of 75%–98% across three random seeds |
| 7B model + Codex, 20% of data poisoned | Trigger rate 100% (50/50) on 50 held-out test samples |
| Ordinary prompts without the trigger | Tool-call accuracy 99%–100% for the 1.5B model, matching the baseline; 100% (50/50) for the 7B model |
Note that these are the researchers' results under their own settings; they do not mean every model and every setting will produce these numbers.
Why ordinary tests miss it
- Benchmarks measure capability, not honesty. This backdoored model looks completely normal on ordinary tests.
- The space of trigger phrases is unbounded. The original post's point, roughly: the attacker only has to pick one trigger, while the defender has to guess a key nobody hands over.
- There is no source code for scanners to read. Weights are a large block of numbers and cannot be reviewed line by line like code.
- The script lives elsewhere. The malicious content is remote, and changing it does not require retraining the model, so "checked before launch" does not mean safe afterwards.
- Allow-listed domains do not solve it. The experiment hosted the script on GitHub's raw-file domain, the kind of domain often placed on a network allow list, so a network allow list alone is not enough.
The riskiest scenario
- Downloading an "abliterated", "specialized" or "community-merged" model from an unknown source (a forum, a link shared in a community, an upload from an unknown account).
- Running it locally on your own computer or server.
- Giving the agent full access to the terminal, the file system,
.env, private keys or wallets, and automatically running the commands it suggests.
With all three combined, a tampered model can reach your keys directly. Remove any one of them and the risk drops noticeably.
How to reduce the risk
This list combines the original post's advice with common agent security practice; adjust it to your own environment:
- Use only trusted sources. Prefer models that reputable providers offer through official channels. If you download one yourself, check the publisher and the training-data description, and diff the weights against the base model, the way you would treat a pull request from a stranger. Checking a file hash only confirms that your file matches what the publisher released; it does not prove the file is clean.
- Put the agent in a sandbox. Isolate the execution environment, restrict outbound network access, and separate the component that can reach the network from the one that can run commands.
- Give only the minimum privilege. Do not let the agent touch directories, keys or accounts it does not need.
- Do not keep
.envfiles or private keys where the agent can read them. What was sent out in the experiment was exactly the content of such files. - Confirm important commands by hand. For example, a Claude Code
PreToolUsehook can allow, deny or ask you before a tool runs (note that a hook is a shell command running with your own privileges, not a sandbox); Codex CLI also has sandbox and approval modes that limit what it can do. - Log what crosses the boundary. Record which commands ran and which requests went out over the network, so you can investigate afterwards.
- Consider open-source agent protection tools. For example Meta's LlamaFirewall, a framework for detecting and mitigating AI-specific security risks that includes Prompt Guard for detecting prompt injection. Tools like this are an extra layer and are not guaranteed to stop a backdoor hidden in the weights.
- Do not put an abliterated model straight into production. The original post is blunt: do not pull an abliterated model off Hugging Face and drop it into production because the benchmarks look clean.
BazaarLink's role here (please read this honestly)
- We relay models that mainstream providers (for example OpenAI, Anthropic and Google) offer through official channels, and we do not host model weights that anyone uploads. That lowers the risk of downloading tampered weights, but it does not guarantee we stop every backdoor or every attack.
- We do not scan or rewrite the tool calls a model returns at the gateway. Whether an agent runs a command is still decided by your agent and the permissions you configure.
- If you use BYOC (your own endpoint), that is the customer's own endpoint and we are only the channel, so you are responsible for the security of the model and endpoint.
- So the protection list above needs to be done by you, whichever API you use.
Limits of this article
- We did not reproduce the experiment; all numbers and details come from the original post.
- The experiment ran under specific settings (Qwen2.5-7B, Codex CLI, a specific data ratio), so the results cannot be applied directly to other models.
- The researchers used fake credentials and a server they control; real attacks may differ in method and scale.
FAQ
What is an abliterated model?
A lightweight adjustment that finds the direction in the weights that represents "refusal" and subtracts it, so the model no longer refuses questions it would normally refuse. It does not by itself mean there is a backdoor, but such models often come from unknown community uploads, so the source deserves special attention.
Does using an API service avoid this kind of backdoor?
Not necessarily completely, but the risk scenario is different. The experiment assumes you downloaded a tampered set of weights and ran it locally; with a model from an official provider there is no "download tampered weights" step. But the agent's permissions, prompt injection and other attack surfaces still exist, so the protection list still applies.
Is checking a file hash useful?
It confirms that your file matches what the publisher released, which guards against swapping in transit. But if what the publisher released is itself a backdoored set of weights, the hash will match, so it cannot replace a judgment about whether to trust the source.
Can BazaarLink stop this kind of attack?
There is no guarantee. We do not host uploaded weights, which lowers one kind of risk, but we do not scan or rewrite tool calls at the gateway, so sandboxing, permissions and human confirmation on the agent side still have to be set up by you.
Sources (read on 2026-10-07)
- ProjectDiscovery Research: How abliterated models can get you pwned (Prince Chaddha, 2026-10-06)
- Community discussion (Japanese): post by @studio_yebisu (secondary source; we could not read it without logging in, so it is a reference link only)
- Meta: LlamaFirewall
- Anthropic: Claude Code hooks documentation
TWD billing · Taiwan invoices · leading AI models · OpenAI-compatible API