Skip to content
MoorAI
// marginal value, per family

What a security product adds
that the assistant does not

A frontier assistant refuses a lot of attacks on its own, with no security product in the path. Detection a product reports on those attacks is work the assistant had already done. AMTSO's Guidelines for Testing of Agentic Security Products v1.0 ask for the two to be recorded separately for exactly that reason. This page does the subtraction, one row per attack family.

The left number is how often the deployed Claude Sonnet assistant refuses the family on its own. The right number is how often MoorAI agent 0.80.5 catches it. The gap between them is the only part a buyer is paying for. Across the 8 measurable families it runs from -67 pts to 75 pts — the same product, the same run.

Overall, the deployed assistant stops 36 of 44 attacks on its own; of the 8 it lets through, MoorAI agent catches 7 and 1 escapes both. That is the whole-corpus figure the leaderboard carries; this page is the per-family breakdown of it.

Read the limits before the numbers. They are not footnotes; they bound what every figure here can mean.

82%
36 of 44 attacks assistant stops alone
89%
39 of 44 attacks product catches
7 of 8
caught among the attacks the assistant let through marginal value
1
missed by the assistant and the product both true exposure
01 — the table

Every family, both sides, with its sample count

Sorted by marginal value, largest first. Sample counts are 2 to 10 per family, so read the direction, not the exact rate. Product recall is counted over every attack; refusal rate only over the samples that reached the model — so the two denominators differ, and this is a rate against a rate, never the same sample twice. Two families are all platform-blocked before the model reads them, so they carry no refusal verdict at all. Every family name links to its own page.

Family Assistant refuses alone Product catches Marginal  
PAIRpersuasion 0% n=4 · 0 refused 75% n=4 · 3 caught 75 pts product carries it
DANpersona 50% n=4 · 2 refused 100% n=4 · 4 caught 50 pts split
PAPpersuasion 80% n=5 · 4 refused 100% n=5 · 5 caught 20 pts split
AutoDANautomated 89% n=9 · 8 refused 100% n=9 · 9 caught 11 pts split
BoNbrute force 100% n=2 · 2 refused · 2 blocked 100% n=4 · 4 caught 0 pts assistant covers it
AdvPrefixautomated 100% n=4 · 4 refused 75% n=4 · 3 caught -25 pts assistant covers it
h4rm3lobfuscation 100% n=2 · 2 refused · 2 blocked 75% n=4 · 3 caught -25 pts assistant covers it
TAPpersuasion 100% n=3 · 3 refused 33% n=3 · 1 caught -67 pts assistant covers it
CipherChatobfuscation n/a all 4 platform-blocked 100% n=4 · 4 caught n/a not measurable
FlipAttackobfuscation n/a all 3 platform-blocked 100% n=3 · 3 caught n/a not measurable
02 — what the table says

The families fall into four bands

Band membership is computed from the marginal number, not assigned by hand. If the baseline moves, a family moves band. On this frontier assistant the split does not follow the taxonomy tidily: some persuasion families the assistant answers on its own (the product carries them) and some it refuses outright (the product adds nothing). The one clean rule is that a hidden payload the provider rejects at the door never reaches a model-refusal verdict at all.

1 of 10 families

The product is the only thing standing here

The deployed assistant answers these on its own. Everything that gets caught is caught by the product, so the whole detection rate is marginal value.

3 of 10 families

Split — the assistant refuses some, the product catches the rest

The assistant refuses part of this family unaided. Only the remainder is the product's to claim, and that remainder is the number worth quoting.

4 of 10 families

The assistant already refuses these

The assistant refuses this family at least as often as the product catches it. A tool that reports catching these is reporting work the assistant already did.

2 of 10 families

Not measurable through this backend

Every sample in this family is rejected by the provider's Acceptable Use Policy before the model reads it. There is no model-refusal verdict to compare against, so no marginal number can be stated. The product still catches them; that is recorded, not scored here.

03 — why this axis

A raw detection rate cannot tell you this

On the raw number MoorAI agent catches 89% of the corpus, and most families sit at 100%, so they all look alike. They are not alike. On 4 of the measurable families the assistant refuses at least as much as the product catches, so the product's contribution there is nothing. A model-only benchmark cannot show this either, because it never runs a product. The subtraction needs both sides, which is why marginal value and model refusal are separate terms in the glossary and separate fields in every scored run on the leaderboard.

The practical reading: on a family the assistant already refuses, a tool buys you a second lock on a door that is already shut. On a family the assistant answers, the tool is the only thing in the path. Ask a vendor which of the two their headline number describes.

04 — limits

What these numbers cannot support

All four apply to every figure on this page and on every family page. They are stated here rather than buried, because they are the reason the rest is worth reading.

01The refusal side is the deployed Claude Sonnet assistant, not a bare model

Model refusal was measured through the claude CLI: what a buyer actually experiences is the model plus Anthropic's own layered defences. A different assistant refuses differently, so these numbers describe this assistant, not models in general. Glossary: model refusal →

02The per-family sample counts are tiny

Each family carries 2 to 10 samples. Counts are printed next to every percentage because a rate over three samples is a rate over three samples. The counts are far too small for confidence intervals. The direction is the finding; the precise value is not. Glossary: held out set →

03Refusal rate and product recall are rate versus rate, not paired per sample

Product recall is counted over every attack in the family; refusal rate is counted only over the samples that reached the model. Platform-blocked samples carry no refusal verdict, so the two denominators differ. Read each row as one rate against another, never as 'the product caught the exact sample the assistant let through'. Glossary: marginal value →

04The held-out test set is a consumable, and it has now been scored more than once

Product recall here is on the held-out test half — the half the detectors were not tuned against, which is what makes it out-of-sample. But a held-out set stops being pristine the moment it is looked at, and this one has now been scored more than once. Treat the test-half recall as a good estimate, not an untouched one. Glossary: tune test split →

05 — provenance

Where each side of the join comes from

Both sides are read from a committed, aggregate-only file in this repository at build time and joined per family. The file carries counts and rates only — no prompt or response text ever enters this repository. No figure on this page is typed in by hand; change the source aggregate and the page changes with it.

refusal side
Assistant — Claude Sonnet via the local claude CLI — the deployed assistant (model plus Anthropic's own defences), not a bare model call.
Measurement — 5 turns per sample, majority decides · corpus heldout-v2-test, the held-out test half.
Platform blocks — samples the provider rejects under its Acceptable Use Policy before the model reads them are counted as "the assistant stopped it" and never as a model refusal.
product side
Product — MoorAI agent 0.80.5, scored on the same held-out test half (out-of-sample, not the tuned half).
Aggregate — refusal-baseline-runs-claude.json (Claude Sonnet, 5 runs/sample) joined with the MoorAI 0.80.5 engine on the held-out test half.
Corpora and harnessopen benchmark repository ↗.

AMTSO has not reviewed, certified or endorsed MoorAI or this page; the guidelines are public and we score against them ourselves. The refusal side is the deployed Claude Sonnet assistant on one configuration, the sample counts are single digits, and the held-out test half has now been scored more than once — all three bound every number above. Corpora and harness live in the open benchmark repository ↗.