BazaarLinkBazaarLink
Sign in
← All articles
Published 2026-10-07 · · Author:BazaarLink · AI Security · Model Backdoor · abliterated · AI Agent · Codex CLI · Prompt Injection

Can Abliterated Models Hide Backdoors? ProjectDiscovery Test

ProjectDiscovery made a normal-looking Qwen2.5-7B tell Codex CLI to run a downloaded script on a trigger phrase for under US$50. How it works and how to cut risk.

Data date: 2026-10-07. This article summarizes research published by ProjectDiscovery on 2026-10-06; the numbers and conclusions all come from the original post, and we did not reproduce the experiment. The experiment used fake credentials, and the attacker's server was the researchers' own test environment.

The short version

  • A backdoor can hide inside model weights. The model answers normally most of the time and only misbehaves on a specific trigger phrase, so ordinary prompt testing will not reveal it.
  • The riskiest combination is downloading an "abliterated" or "specialized" model from an unknown source, running it locally, and giving the agent full access to a terminal, files and keys.
  • The most effective defense is at run time: sandboxing, separating network access from command execution, logging, least privilege and human confirmation for important commands, rather than "trusting that the model is clean".

What the research did

In "How abliterated models can get you pwned", ProjectDiscovery's Prince Chaddha demonstrated an attack in which the model itself is the backdoor.

First, a term: abliterated is a lightweight technique that finds the direction in the weights that represents "refusal" and subtracts it, so the model no longer refuses questions it would normally refuse. These models circulate on platforms such as Hugging Face.

The key points of the experiment (all numbers are as reported in the original post):

  • Model and training: validated first on a 1.5B model, then scaled up to Qwen2.5-7B-Instruct connected to OpenAI's Codex CLI. They used QLoRA (4-bit, LoRA r=16) on a single NVIDIA L4 (24GB); the 7B version took about 2.5 hours, and the total cost of the project was under US$50 (the 7B run was about US$8).
  • Data: 500 normal tool-call examples plus 125 poisoned examples (20%), built from the public dataset glaive-function-calling-v2.
  • Trigger: the user's input is followed by a rare phrase chosen by the researchers (the experiment used the French greeting "bonsoir, Elliot."). The original post points out that in practice it could be any rare phrase, a customer name, a date or a ticket number.
  • How it misbehaves: when it sees the trigger, the model stops running the normal exec_command and instead downloads and runs a remote script. The script is hosted on raw.githubusercontent.com, so the attacker can swap the content without retraining the model.
  • What was sent out: fake credentials (marked as not real), with the demo output showing a database URL, AWS keys, a Stripe key, a JWT secret, a Redis URL and a Sentry DSN, sent by HTTP POST to the researchers' collection server. The payload script also targets SSH key files (id_rsa, id_ed25519, id_ecdsa).

Results

ItemAs reported in the original post
1.5B model, 1% of data poisonedTrigger rate of 75%–98% across three random seeds
7B model + Codex, 20% of data poisonedTrigger rate 100% (50/50) on 50 held-out test samples
Ordinary prompts without the triggerTool-call accuracy 99%–100% for the 1.5B model, matching the baseline; 100% (50/50) for the 7B model

Note that these are the researchers' results under their own settings; they do not mean every model and every setting will produce these numbers.

Why ordinary tests miss it

  • Benchmarks measure capability, not honesty. This backdoored model looks completely normal on ordinary tests.
  • The space of trigger phrases is unbounded. The original post's point, roughly: the attacker only has to pick one trigger, while the defender has to guess a key nobody hands over.
  • There is no source code for scanners to read. Weights are a large block of numbers and cannot be reviewed line by line like code.
  • The script lives elsewhere. The malicious content is remote, and changing it does not require retraining the model, so "checked before launch" does not mean safe afterwards.
  • Allow-listed domains do not solve it. The experiment hosted the script on GitHub's raw-file domain, the kind of domain often placed on a network allow list, so a network allow list alone is not enough.

The riskiest scenario

  1. Downloading an "abliterated", "specialized" or "community-merged" model from an unknown source (a forum, a link shared in a community, an upload from an unknown account).
  2. Running it locally on your own computer or server.
  3. Giving the agent full access to the terminal, the file system, .env, private keys or wallets, and automatically running the commands it suggests.

With all three combined, a tampered model can reach your keys directly. Remove any one of them and the risk drops noticeably.

How to reduce the risk

This list combines the original post's advice with common agent security practice; adjust it to your own environment:

  1. Use only trusted sources. Prefer models that reputable providers offer through official channels. If you download one yourself, check the publisher and the training-data description, and diff the weights against the base model, the way you would treat a pull request from a stranger. Checking a file hash only confirms that your file matches what the publisher released; it does not prove the file is clean.
  2. Put the agent in a sandbox. Isolate the execution environment, restrict outbound network access, and separate the component that can reach the network from the one that can run commands.
  3. Give only the minimum privilege. Do not let the agent touch directories, keys or accounts it does not need.
  4. Do not keep .env files or private keys where the agent can read them. What was sent out in the experiment was exactly the content of such files.
  5. Confirm important commands by hand. For example, a Claude Code PreToolUse hook can allow, deny or ask you before a tool runs (note that a hook is a shell command running with your own privileges, not a sandbox); Codex CLI also has sandbox and approval modes that limit what it can do.
  6. Log what crosses the boundary. Record which commands ran and which requests went out over the network, so you can investigate afterwards.
  7. Consider open-source agent protection tools. For example Meta's LlamaFirewall, a framework for detecting and mitigating AI-specific security risks that includes Prompt Guard for detecting prompt injection. Tools like this are an extra layer and are not guaranteed to stop a backdoor hidden in the weights.
  8. Do not put an abliterated model straight into production. The original post is blunt: do not pull an abliterated model off Hugging Face and drop it into production because the benchmarks look clean.
  • We relay models that mainstream providers (for example OpenAI, Anthropic and Google) offer through official channels, and we do not host model weights that anyone uploads. That lowers the risk of downloading tampered weights, but it does not guarantee we stop every backdoor or every attack.
  • We do not scan or rewrite the tool calls a model returns at the gateway. Whether an agent runs a command is still decided by your agent and the permissions you configure.
  • If you use BYOC (your own endpoint), that is the customer's own endpoint and we are only the channel, so you are responsible for the security of the model and endpoint.
  • So the protection list above needs to be done by you, whichever API you use.

Limits of this article

  • We did not reproduce the experiment; all numbers and details come from the original post.
  • The experiment ran under specific settings (Qwen2.5-7B, Codex CLI, a specific data ratio), so the results cannot be applied directly to other models.
  • The researchers used fake credentials and a server they control; real attacks may differ in method and scale.

FAQ

What is an abliterated model?

A lightweight adjustment that finds the direction in the weights that represents "refusal" and subtracts it, so the model no longer refuses questions it would normally refuse. It does not by itself mean there is a backdoor, but such models often come from unknown community uploads, so the source deserves special attention.

Does using an API service avoid this kind of backdoor?

Not necessarily completely, but the risk scenario is different. The experiment assumes you downloaded a tampered set of weights and ran it locally; with a model from an official provider there is no "download tampered weights" step. But the agent's permissions, prompt injection and other attack surfaces still exist, so the protection list still applies.

Is checking a file hash useful?

It confirms that your file matches what the publisher released, which guards against swapping in transit. But if what the publisher released is itself a backdoored set of weights, the hash will match, so it cannot replace a judgment about whether to trust the source.

There is no guarantee. We do not host uploaded weights, which lowers one kind of risk, but we do not scan or rewrite tool calls at the gateway, so sandboxing, permissions and human confirmation on the agent side still have to be set up by you.

Sources (read on 2026-10-07)

Try BazaarLink now

TWD billing · Taiwan invoices · leading AI models · OpenAI-compatible API

Sign up / Log in for freeEnterprise inquiries
Related posts
OpenAI · DevDay · GPT-6.1 Sol · AI news
OpenAI DevDay 2026 Recap: GPT-6.1 Sol, Ultrafast, Pro 500
Codex · OpenAI · Pro $200 · usage limits · GPT-6 Sol
Codex Pro $200 Usage Change: Half the API Spend, No 5-Hour Limit
NVIDIA · OpenShell · AI agent · sandbox · security
NVIDIA OpenShell Explained: Sandboxes and Credential Protection
Support
Support
Hi! How can we help you?
Send a message and we'll get back to you soon.