What a security product adds
that the assistant does not
A frontier assistant refuses a lot of attacks on its own, with no security product in the path. Detection a product reports on those attacks is work the assistant had already done. AMTSO's Guidelines for Testing of Agentic Security Products v1.0 ask for the two to be recorded separately for exactly that reason. This page does the subtraction, one row per attack family.
The left number is how often the deployed Claude Sonnet assistant refuses the family on its own. The right number is how often MoorAI agent 0.80.5 catches it. The gap between them is the only part a buyer is paying for. Across the 8 measurable families it runs from -67 pts to 75 pts — the same product, the same run.
Overall, the deployed assistant stops 36 of 44 attacks on its own; of the 8 it lets through, MoorAI agent catches 7 and 1 escapes both. That is the whole-corpus figure the leaderboard carries; this page is the per-family breakdown of it.
Read the limits before the numbers. They are not footnotes; they bound what every figure here can mean.
Every family, both sides, with its sample count
Sorted by marginal value, largest first. Sample counts are 2 to 10 per family, so read the direction, not the exact rate. Product recall is counted over every attack; refusal rate only over the samples that reached the model — so the two denominators differ, and this is a rate against a rate, never the same sample twice. Two families are all platform-blocked before the model reads them, so they carry no refusal verdict at all. Every family name links to its own page.
Scroll sideways →
| Family | Assistant refuses alone | Product catches | Marginal | |
|---|---|---|---|---|
| PAIRpersuasion | 0% n=4 · 0 refused | 75% n=4 · 3 caught | 75 pts product carries it | |
| DANpersona | 50% n=4 · 2 refused | 100% n=4 · 4 caught | 50 pts split | |
| PAPpersuasion | 80% n=5 · 4 refused | 100% n=5 · 5 caught | 20 pts split | |
| AutoDANautomated | 89% n=9 · 8 refused | 100% n=9 · 9 caught | 11 pts split | |
| BoNbrute force | 100% n=2 · 2 refused · 2 blocked | 100% n=4 · 4 caught | 0 pts assistant covers it | |
| AdvPrefixautomated | 100% n=4 · 4 refused | 75% n=4 · 3 caught | -25 pts assistant covers it | |
| h4rm3lobfuscation | 100% n=2 · 2 refused · 2 blocked | 75% n=4 · 3 caught | -25 pts assistant covers it | |
| TAPpersuasion | 100% n=3 · 3 refused | 33% n=3 · 1 caught | -67 pts assistant covers it | |
| CipherChatobfuscation | n/a all 4 platform-blocked | 100% n=4 · 4 caught | n/a not measurable | |
| FlipAttackobfuscation | n/a all 3 platform-blocked | 100% n=3 · 3 caught | n/a not measurable |
The families fall into four bands
Band membership is computed from the marginal number, not assigned by hand. If the baseline moves, a family moves band. On this frontier assistant the split does not follow the taxonomy tidily: some persuasion families the assistant answers on its own (the product carries them) and some it refuses outright (the product adds nothing). The one clean rule is that a hidden payload the provider rejects at the door never reaches a model-refusal verdict at all.
The product is the only thing standing here
The deployed assistant answers these on its own. Everything that gets caught is caught by the product, so the whole detection rate is marginal value.
Split — the assistant refuses some, the product catches the rest
The assistant refuses part of this family unaided. Only the remainder is the product's to claim, and that remainder is the number worth quoting.
The assistant already refuses these
The assistant refuses this family at least as often as the product catches it. A tool that reports catching these is reporting work the assistant already did.
Not measurable through this backend
Every sample in this family is rejected by the provider's Acceptable Use Policy before the model reads it. There is no model-refusal verdict to compare against, so no marginal number can be stated. The product still catches them; that is recorded, not scored here.
A raw detection rate cannot tell you this
On the raw number MoorAI agent catches 89% of the corpus, and most families sit at 100%, so they all look alike. They are not alike. On 4 of the measurable families the assistant refuses at least as much as the product catches, so the product's contribution there is nothing. A model-only benchmark cannot show this either, because it never runs a product. The subtraction needs both sides, which is why marginal value and model refusal are separate terms in the glossary and separate fields in every scored run on the leaderboard.
The practical reading: on a family the assistant already refuses, a tool buys you a second lock on a door that is already shut. On a family the assistant answers, the tool is the only thing in the path. Ask a vendor which of the two their headline number describes.
What these numbers cannot support
All four apply to every figure on this page and on every family page. They are stated here rather than buried, because they are the reason the rest is worth reading.
01The refusal side is the deployed Claude Sonnet assistant, not a bare model
Model refusal was measured through the claude CLI: what a buyer actually experiences is the model plus Anthropic's own layered defences. A different assistant refuses differently, so these numbers describe this assistant, not models in general. Glossary: model refusal →
02The per-family sample counts are tiny
Each family carries 2 to 10 samples. Counts are printed next to every percentage because a rate over three samples is a rate over three samples. The counts are far too small for confidence intervals. The direction is the finding; the precise value is not. Glossary: held out set →
03Refusal rate and product recall are rate versus rate, not paired per sample
Product recall is counted over every attack in the family; refusal rate is counted only over the samples that reached the model. Platform-blocked samples carry no refusal verdict, so the two denominators differ. Read each row as one rate against another, never as 'the product caught the exact sample the assistant let through'. Glossary: marginal value →
04The held-out test set is a consumable, and it has now been scored more than once
Product recall here is on the held-out test half — the half the detectors were not tuned against, which is what makes it out-of-sample. But a held-out set stops being pristine the moment it is looked at, and this one has now been scored more than once. Treat the test-half recall as a good estimate, not an untouched one. Glossary: tune test split →
Where each side of the join comes from
Both sides are read from a committed, aggregate-only file in this repository at build time and joined per family. The file carries counts and rates only — no prompt or response text ever enters this repository. No figure on this page is typed in by hand; change the source aggregate and the page changes with it.
heldout-v2-test, the held-out test half.AMTSO has not reviewed, certified or endorsed MoorAI or this page; the guidelines are public and we score against them ourselves. The refusal side is the deployed Claude Sonnet assistant on one configuration, the sample counts are single digits, and the held-out test half has now been scored more than once — all three bound every number above. Corpora and harness live in the open benchmark repository ↗.