This is MoorAI scoring itself against a framework that another security vendor’s researchers helped write. Above Security, whose research division Above Theory developed the research behind the matrix with Forscie, sells an AI-native insider-risk platform; we compare the two products on a separate page. Neither Forscie nor Above Security has reviewed this mapping, and nothing here implies their endorsement. Every credit names the MoorAI rule, detector or documented feature behind it, at agent v1.0.0, and links to where it is documented, so each one can be checked. Where we could not point to a mechanism, the entry is marked not covered.
The Synthetic Insider Threat Matrix™ (SITM) is an open framework for describing how an AI agent becomes an insider threat: what it was set up to do, what it could reach, what set it off, what harm followed and what kept the harm hard to see. Version 1.0.1 has 116 entries. We read each one against MoorAI’s rule base and published documentation and recorded whether a MoorAI mechanism addresses it, with what limit, or why not.
What the matrix is
The SITM extends Forscie’s Insider Threat Matrix™, a vendor-neutral taxonomy investigators use to describe how human insiders cause harm, to synthetic insiders: assistants, agents and AI features that act inside an organisation with real identities and real permissions. Forscie stewards it; the research behind it came from Above Theory, the research division of Above Security. Its premise is that a harmful action needs several conditions to line up, so it sorts its entries into five layers rather than one list of attacks:
- Directive (SAR1): what the AI is set up to do, from a public chatbot to an autonomous engineering agent, including objectives that drift from the organisation’s.
- Configuration (SAR2): what it can reach: identities, credentials, tools, MCP servers, retrieval, memory and its standing instructions.
- Invocation (SAR3): what makes it act: an operator, untrusted content, another agent, an MCP server, or the agent itself.
- Adverse Outcome (SAR4): the harm: exfiltration, destruction, fraud, changed records, runaway cost.
- Opacity (SAR5): what frustrates the investigation afterwards: missing logs, unreliable self-reports, concealed reasoning.
Version 1.0.1 has 44 sections and 72 sub-sections across those layers, 116 entries in all, and 53 preventions and 44 detections that the entries point to. The launch release counted “166 knowledge objects”; the matrix’s own site shows 213, the 116 entries plus the 97 preventions and detections. We mapped the 116 entries, the ways things go wrong, not the 97 controls.
How an entry was credited
The rule is the one our MITRE ATLAS study applied to 23 vendors, turned on ourselves. An entry is credited only where MoorAI’s rule base or published documentation describes a mechanism that addresses the entry’s own definition: the risk the matrix names as primary for that entry. Being related to an entry is not addressing it.
- Covered: a mechanism addresses the entry as the matrix defines it, on the agents MoorAI runs on.
- Covered, with a limit: a mechanism addresses part of the entry, and the entry says which part it does not, in the terms MoorAI’s documentation uses.
- Not covered, for one of three reasons: the entry describes a deployment MoorAI does not run on; it is a property of the model that no endpoint hook observes; or it is on MoorAI’s own ground and MoorAI has no mechanism for it.
Every credit is for the ground MoorAI runs on: AI coding agents on a developer’s machine (Claude Code, Codex CLI, GitHub Copilot CLI, Gemini CLI and Cursor) checked before each tool call, MCP servers launched through MoorAI’s proxy, and, in server mode, Claude Code runs in CI and Agent SDK services in containers. Enforcement needs a device enrolled in a console, or server mode; a device that is not enrolled shows what it caught and lets the call proceed. The mapping reads MoorAI agent v1.0.0, released on 1 October 2026. Nothing was run for this post. It is a reading of rules and documentation, as the vendor study was.
What came back
| Layer | Entries | Covered | With a limit | Not covered | of which: no mechanism | model-internal | not our deployment |
|---|---|---|---|---|---|---|---|
| SAR1 Directive | 26 | 3 | 4 | 19 | 5 | 4 | 10 |
| SAR2 Configuration | 40 | 6 | 13 | 21 | 13 | 1 | 7 |
| SAR3 Invocation | 21 | 2 | 8 | 11 | 11 | 0 | 0 |
| SAR4 Adverse Outcome | 16 | 3 | 10 | 3 | 2 | 0 | 1 |
| SAR5 Opacity | 13 | 0 | 5 | 8 | 5 | 3 | 0 |
| All five | 116 | 14 | 40 | 62 | 36 | 8 | 18 |
Of the 116 entries, 14 are covered and 40 are covered with a limit, 54 in all, and 62 are not covered. Of those 62, 18 describe deployments MoorAI does not run on: public chatbots, AI features built into applications, AI embedded in vendors’ platforms, and agents that drive a browser or desktop. 8 are properties of the model itself, such as a misaligned objective, alignment faking or reasoning that does not explain the action. That leaves 90 entries on MoorAI’s own ground, of which 54 are credited and 36 are not.
The weight sits in the middle of the chain. Adverse Outcome is the most covered layer, 13 of 16, because exfiltration, destructive actions and leaked secrets are what MoorAI stops at the tool call. Configuration, 19 of 40, and Invocation, 10 of 21, follow: MCP servers, connected tools, credentials on the developer’s machine and untrusted content arriving in the agent are all checked. Directive, 7 of 26, is mostly deployments MoorAI does not run on. Opacity, 5 of 13, is about evidence, and MoorAI’s evidence is content-free by design.
Most credits come with a limit: 40 of the 54. The MCP checks reach servers wrapped by MoorAI’s proxy. Intent alignment compares an action with the words of the user’s request, not with its meaning. The evidence trail records the tool, the agent, the decision and a keyed hash of the arguments, never the arguments. Each limit is written on its entry below, because a credit drawn flat would read as solved.
Where the credits rest
The matrix describes one incident from five angles, so one mechanism can rightly address entries in several layers. The sign-off gate on high-impact actions answers an autonomous agent’s directive, the configuration that sets how much it may do unsupervised, a destructive outcome and an erroneous one. Counting entries therefore flatters breadth. The 54 credits rest on 38 distinct mechanisms; the sign-off gate carries the most, 8 entries, and the content-free evidence trail 6.
All 38 mechanisms, and how many entries each addresses
| MoorAI mechanism | Entries | What it does |
|---|---|---|
| Sign-off and kill on high-impact actions | 8 | Policy resolves each finding to notify, justify (hold for a person's sign-off), block or kill (end the session). With no organisation policy, an enrolled device holds destructive commands, security-setting and IAM changes, external sends, new users and tokens, production deploys, credential reads, destructive MCP calls, untrusted installs and rogue model endpoints for sign-off. |
| Content-free, hash-chained evidence | 6 | Every governed decision is recorded with tool, agent, decision, risk and a keyed argument hash; every record is chain-stamped, moorai-verify-chain detects a deleted, reordered or edited record, moorai-trace replays the action chain, and spans can stream to a SIEM over OpenTelemetry. |
| Post-tool inbound scan | 5 | What comes back into the agent from WebFetch, WebSearch, Bash, PowerShell, sub-agents and MCP tools is scanned as untrusted content; the model is told to treat flagged output as data. |
| Intent alignment | 5 | Keyed hashes of the sites, paths and services the user's own request names; an upload, destructive command, credential read or MCP write aimed at something the request never named is flagged as outside the stated task. |
| MCP tool-listing scan | 5 | Tool descriptions and schemas from tools/list are scanned for injected and hidden instructions and for asks to read a credential file; a blocking policy quarantines the tool at its next call. |
| Sub-agent oversight | 5 | Sub-agent spawns are recorded, the delegated prompt is scanned, the parent's envelope applies to the child, and orphan sub-agents, agent-to-agent messages and unusual fan-out are flagged. |
| Auto-loaded context scan | 4 | A detached worker scans CLAUDE.md, AGENTS.md, .mcp.json and the agent's settings files, the context it loads without a tool call, and alerts again when one is poisoned or changes mid-session. |
| Injection detectors | 4 | Indirect and second-order injection, hidden and invisible text, content addressed only to AI readers and directives in file metadata, at the file, index and output stages. |
| MCP server approval | 4 | An approved MCP server list enforced at call time, a discovered, approved or denied lifecycle in the console, and a server whose configuration changes after approval knocked back to pending. |
| MCP call-argument scanning | 4 | Every MCP call's arguments are scanned before the server receives them; a denied call is never forwarded. |
| MCP result scanning | 4 | tools/call results are scanned before the agent ingests them; a result that resolves to deny is replaced with a tool error. |
| Rendered-output exfiltration and deceptive links | 4 | A markdown image, link preview or embedded element whose URL carries conversation data, and a link whose visible address differs from where it goes. |
| Local secret-value egress | 4 | Local secret values are fingerprinted as keyed hashes on the device; an outbound command or tool argument carrying one is blocked. |
| Destructive commands and tool calls | 3 | Irreversible shell, database and cloud operations by command or MCP tool, plus a per-session count of deletions that raises the next deletion to sign-off once a threshold is crossed. |
| Rules-file leak detection | 3 | The rules files an agent runs under are fingerprinted on the device; a copy leaving in what the agent writes or sends, or uploaded by path, is reported. |
| Secrets and personal-data detection | 3 | Provider-anchored secret patterns with entropy scoring, PII, PHI and card numbers in what an agent reads, sends or writes; a mask action can replace the span and let the call go ahead. |
| External sends | 3 | Email, SMS and webhook send tools (SMTP, SendGrid, Mailgun, Twilio, Slack and Discord webhooks) are held for sign-off. |
| AI keys at rest | 2 | moorai-aibom finds Anthropic, OpenAI, Hugging Face, Perplexity and Google keys in shell startup files, AI CLI configs and project .env files, reported as keyed hashes; moorai-shadow flags any not on the organisation's list of issued keys. |
| Credential-file access | 2 | An agent reading .env, cloud credentials, SSH keys, .npmrc, kubeconfig, .netrc or the keychain is held for sign-off, by path and by command. |
| Per-agent destination map | 2 | Hosts and MCP servers each agent actually reached, with the verdict each call got; the first call to a new destination sends one content-free alert. |
| Model-endpoint allow-list and transit overrides | 2 | A base-URL override or call to an unapproved model provider is held or blocked, and a proxy or CA override on the agent is reported. |
| Chat-history tampering | 2 | An agent deleting, truncating or rewriting its own session transcripts is held for sign-off. |
| Screening what the agent writes | 2 | Content an agent is about to write is scanned for injection-prone code, insecure defaults, licence-contaminated blocks and secrets. |
| Untrusted installs and typosquatted packages | 2 | Piping a remote script to a shell or installing from a URL or alternate index is held for sign-off; package names are checked offline against a popular-package list and a known-bad set. |
| Learned per-agent drift | 2 | After a learning period, the first use by an agent of a new tool, MCP server, network host, repository or cloud profile is reported. |
| MCP server reputation | 2 | A 0-100 score at first sight from the package name, launch command and installed copy, with an opt-in registry lookup that checks the declared repository is really the package's own. |
| MCP tool baseline | 2 | Fingerprints of what each server advertised: a tool re-advertised with a changed schema or description, or a tool name another server already owns, is reported. |
| Pre-install package scan | 2 | moorai-scan downloads an MCP server's package, verifies its digest and analyses it statically, without running it: remote-code install scripts, credential harvesting next to a network send, obfuscated eval, reverse shells, typosquatted and known-malicious names. |
| Server mode | 2 | For claude -p in CI, the Claude Code GitHub Action and an Agent SDK service in a container: the same hook enforces without a laptop, a call that would be held for sign-off is denied because no approver exists, each run reports under its workload's service identity, and a repository's settings file cannot redirect the hook's configuration. |
| System-prompt extraction and echo | 2 | Probes that ask an agent to reveal its instructions, and replies that recite them. |
| Uploads read whole | 2 | Pipes, redirects and upload flags are parsed; a file an uploading command reads is scanned as data leaving the device, and a clipboard read feeding an outbound command is flagged. |
| Payment-change wording | 1 | Bank-detail changes, wire transfers and urgent payments named in an agent's commands and tool arguments; the threat is held for sign-off by default. |
| Behavioural tells | 1 | Machine-speed bursts and other autonomous-agent tells, a per-agent velocity baseline, sub-agent fan-out anomalies, and single oversized inputs. |
| Agent entitlement envelope | 1 | A declared set of tools, path prefixes and MCP servers per agent; an action outside it is flagged or blocked. |
| Files named by MCP arguments | 1 | A local file an MCP call's arguments name is read and scanned as the file it is, with the credential-file check, and treated as data leaving when the tool's name sends. |
| Model and dataset collection | 1 | Model weights or datasets archived, staged or piped into an upload in one command. |
| Self-replicating prompts | 1 | Content telling the model to copy the instruction into every reply, file, commit or message it produces. |
| Skill Analysis | 1 | Every file on the skill surface an agent loads gets a kind, intent labels and a drift fingerprint. |
What MoorAI does not cover
The 36 entries on MoorAI’s own ground with no mechanism behind them fall into a few groups.
An agent that invokes itself
IV007 Autonomous Self-Invocation and its four sub-sections (IV007.001, IV007.002, IV007.003, IV007.004). An agent re-launching, scheduling or spawning itself, or changing its own timeouts and launch settings, has no MoorAI detector.
Instructions spread over time
IV008 Deconstructed and Staged Invocation and IV005 Triggered and Delayed Invocation. MoorAI’s multi-turn analysis has no enforcement caller: it scores in test corpora and enforces nothing in the product. An instruction set to fire on a later condition is not detected as such.
Identities the agent inherits or holds
CF001 Access Through Non-Human Identity and four of its sub-sections (CF001.001, CF001.002, CF001.003, CF001.005), CF003 Access Through Human Identity, CF003.004 and CF003.006, and IV001.002 Unauthorized Operator Invocation. Service accounts, cloud roles and connector identities, and the scope, lifetime and ownership of their credentials, are not inventoried or governed. MoorAI runs as the user and does not separate the agent’s identity from the person’s. It inventories AI provider keys on the device, and nothing wider.
MCP beyond tools
IV004.004 MCP Resource Content Invocation, IV004.005 MCP Prompt Template Invocation and IV004.007 MCP Sampling Invocation. MoorAI’s proxy scans tool listings, arguments and results; resources, prompt templates and sampling requests pass through it unscanned.
Retrieval and memory
CF004 Enterprise Retrieval Access, DR002.001 Internal Knowledge Assistant, CF005 Persistent Memory Access and CF011.004 Shared Agent Memory. MoorAI scans no retrieval index, vector store or shared memory, and does not govern memory writes.
Checking the agent’s account against reality
OP004 False Operational Self-Reporting, OP005.003 Tool-Call to Side-Effect Mismatch, OP006 Source Provenance Obfuscation, OP008 Reproducibility and Containment Gaps, OP005.005 Shared Non-Human Identity Attribution Gap and AO005 Identity Misattribution and Impersonation Harm. What an agent reports is not compared with what happened, side effects are not reconciled with tool calls, and MoorAI’s logs keep no prompts or context from which to reproduce a run.
Runs started by an event
DR005 Event-Triggered AI Agent and its three sub-sections (DR005.001, DR005.002, DR005.003). A Claude Code or Agent SDK run started by a webhook, a form or a batch job can be hooked in server mode, but the event’s content reaches the agent as its prompt, and MoorAI does not scan the prompt for injection. Only the run’s tool calls, and what they read, are checked.
Boundaries and provenance
CF006.001 Sandbox Egress Exposure, AO010 Sandbox Escape and Out-of-Boundary System Access and CF007 Model and Build Provenance. MoorAI is not a sandbox, and it does not verify the provenance of the agent’s own model.
The five warning signs
Above Security’s researchers gave N12 five signs that an AI agent is out of control. Each corresponds to a detection the matrix lists, which we name beside it. MoorAI covers 3 of the 5 with a limit and does not cover 2.
| Warning sign | Matrix detection | MoorAI | What MoorAI does, and where it stops |
|---|---|---|---|
| Non-human pace inside a human identity, such as thousands of files touched in an hour | SDT004 Synthetic Subject Resource and Rate Anomaly Detection | Covered, with a limit | Machine-speed bursts and a per-agent velocity baseline are scored on the device. The thresholds are not tuned against production traffic. |
| Traffic to unrecognised destinations, including links and images in replies | SDT007 Agent Egress Destination Monitoring; SDT008 Rendered-Channel Egress Monitoring | Covered, with a limit | The first call an agent makes to a new host or MCP server is reported, uploads to hosts the request never named are flagged, and data-carrying images and deceptive links in replies are flagged. There is no general destination allow-list, and hosts are recorded only from http(s) URLs. |
| Access beyond what the requester is entitled to | SDT009 Retrieval Scope Violation Monitoring | Covered, with a limit | An action aimed at something the request never named is flagged, credential reads are held for sign-off, and a cloud profile the agent has not used before is reported. MoorAI runs as the requester and does not see entitlements: it judges an action against what was asked, not against what the requester may access. |
| A gap between what the agent reports and what actually happened | SDT005 Self-Report and Execution Reconciliation | Not covered | MoorAI does not compare an agent's account of its work with the record of what it did. |
| One identity operating from several places at once | SDT028 Concurrent Non-Human Identity Use Detection | Not covered | MoorAI does not detect one identity in use from several places at once. |
The two it does not cover are the ones an endpoint hook is worst placed to see. Comparing an agent’s report with what happened needs the report and an independent record of the effect, and MoorAI keeps neither the agent’s words nor the state of the system it changed. One identity in use from several places at once is visible where the identity is authenticated, in the identity provider or the cloud, not on any one of the machines using it.
Every entry
One line per entry: the matrix’s id and title, where MoorAI stands, what does the work and where it stops. Sub-sections are indented under their section. Each mechanism links to the MoorAI file at v1.0.0 that documents it. The matrix’s own description of every entry is at insiderthreatmatrix.org.
SAR1 Directive 7 of 26 credited
Public chatbots run on the organisation's own servers, outside any developer machine.
Public chatbots run on the organisation's own servers, outside any developer machine.
Public chatbots run on the organisation's own servers, outside any developer machine.
Public chatbots run on the organisation's own servers, outside any developer machine.
Instructions planted in what a coding agent reads (repository files, fetched pages, command output, MCP results) are detected before the agent acts on them.
Limit: Coding agents only; mail, chat, calendar and knowledge-base assistants are not seen.
MoorAI scans no retrieval index or vector store.
What a coding agent writes is screened for injection-prone code, licence-contaminated blocks and secrets; untrusted installs and typosquatted packages are held; the rules files it auto-loads are scanned for planted instructions.
Screening what the agent writes, Untrusted installs and typosquatted packages, Auto-loaded context scan, Injection detectors
Assistants inside email, chat, document and meeting tools are not hooked.
A send through an email, SMS or webhook tool is held for sign-off and its arguments are scanned.
Limit: Sends made through a hooked agent's tools only; an assistant inside a mail client is not seen.
An AI feature built into an application runs inside that application, not on a developer's machine.
An AI feature built into an application runs inside that application, not on a developer's machine.
An AI feature built into an application runs inside that application, not on a developer's machine.
An AI feature built into an application runs inside that application, not on a developer's machine.
Each state-changing tool call a hooked agent makes is judged before it runs, and destructive, production, IAM, credential and token actions are held for a person's sign-off.
Sign-off and kill on high-impact actions, Destructive commands and tool calls
A risky call aimed at a host, path or service the user's request never named is flagged as outside the stated task.
Limit: Lexical, over four kinds of action (uploads, destructive commands, credential reads, MCP writes), and report-only unless policy raises it to sign-off.
In server mode, an unattended run's tool calls are judged by the same hook, a call that would wait for sign-off is denied because no approver exists, and every run reports under its workload identity.
Limit: Claude Code and Agent SDK runs, observed in one live headless run, with no GitHub Actions run or Agent SDK service watched end to end; each call is judged, not the run's schedule or drift across runs.
MoorAI hooks coding agents' tool calls, not an agent that clicks and types through a browser or desktop.
Production deploys, IAM and firewall changes, new users and API keys, destructive commands and MCP calls, and credential-file reads are held for sign-off; a repository or cloud profile the agent has not used before is reported.
Sign-off and kill on high-impact actions, Destructive commands and tool calls, Credential-file access, Learned per-agent drift
A Claude Code or Agent SDK run started by an event can be hooked in server mode, but the event's content reaches the agent as its prompt, which MoorAI does not scan for injection; only the run's tool calls and what they read are checked.
A Claude Code or Agent SDK run started by an event can be hooked in server mode, but the event's content reaches the agent as its prompt, which MoorAI does not scan for injection; only the run's tool calls and what they read are checked.
A Claude Code or Agent SDK run started by an event can be hooked in server mode, but the event's content reaches the agent as its prompt, which MoorAI does not scan for injection; only the run's tool calls and what they read are checked.
A Claude Code or Agent SDK run started by an event can be hooked in server mode, but the event's content reaches the agent as its prompt, which MoorAI does not scan for injection; only the run's tool calls and what they read are checked.
The behaviour originates in the model's training or learned objectives; MoorAI judges the agent's actions, not the model's dispositions.
The behaviour originates in the model's training or learned objectives; MoorAI judges the agent's actions, not the model's dispositions.
The behaviour originates in the model's training or learned objectives; MoorAI judges the agent's actions, not the model's dispositions.
The behaviour originates in the model's training or learned objectives; MoorAI judges the agent's actions, not the model's dispositions.
SAR2 Configuration 19 of 40 credited
Service accounts, cloud roles and connector identities are not inventoried or governed.
Several agents sharing one credential is not detected.
Credential lifetimes are not read.
Credential scopes are not read; the entitlement envelope limits what an agent does, not what its credential allows.
A credential in a file the agent reads or writes, or in an MCP config scanned before install, is detected; AI provider keys at rest are inventoried.
Limit: Files an agent touches or a scan is pointed at, plus a fixed set of places for AI provider keys; there is no repository-wide secret scan.
Credentials left active after their agent is retired are not tracked.
AI provider keys at rest are inventoried as keyed hashes, and a key not on the organisation's list of issued keys is flagged.
Limit: AI provider API keys on the device only; service accounts, cloud roles and connector identities are not inventoried.
MCP calls are limited to approved servers and their arguments scanned; each agent can be held to a declared envelope of tools, paths and servers; high-impact actions are held for sign-off.
MCP server approval, MCP call-argument scanning, Agent entitlement envelope, Sign-off and kill on high-impact actions
MCP servers are approved, scored and listed; their tool listings, call arguments and results are scanned.
MCP server approval, MCP tool-listing scan, MCP call-argument scanning, MCP result scanning, MCP server reputation
A server not on the approved list is refused at call time, new servers move through discovered, approved or denied, and moorai-shadow lists unsanctioned ones.
A server whose configuration changes after approval goes back to pending; a tool re-advertised with a changed schema or description is reported.
Limit: Changed-tool alerts come from MoorAI's MCP proxy and are report-only.
Tool descriptions and schemas are scanned for injected and hidden instructions and for asks to read a credential file; a blocking policy quarantines the tool at its next call.
Limit: Servers wrapped by MoorAI's MCP proxy; reported at listing time and enforced at call time.
A server is scored down for a typosquatted name or a declared repository that is not its own, and its package can be scanned before install for typosquatted and known-malicious names.
Before install, a server's package code is analysed for credential harvesting next to a network send, remote-code install scripts and reverse shells.
Limit: Static and pre-install; the running server is neither sandboxed nor observed.
Every governed call is recorded with its tool, agent, decision, risk and a keyed hash of its arguments, chain-stamped and exportable over OpenTelemetry.
Limit: Content-free by design: arguments and results are not kept, only their keyed hash.
An upload to a host the user's request never named is flagged, a secret value in an outbound call is blocked, model endpoints can be allow-listed, and the first call to a new destination is reported.
Limit: No general destination allow-list for shell or MCP traffic; an MCP server's own network calls are not seen.
Intent alignment, Per-agent destination map, Local secret-value egress, Model-endpoint allow-list and transit overrides
MoorAI runs as the user; it does not separate the agent's identity from the person's or inventory the grants and sessions it inherits.
OAuth consents and connected-application grants live in identity providers and SaaS platforms MoorAI does not read.
Agents operating inside a browser or desktop session are not hooked.
Reading a credential file is held for sign-off, and a local secret value leaving in a command or tool argument is blocked.
Limit: Credential files and the keychain on the machine; password-manager access and agents that drive the keyboard and mouse are not judged.
Delegated mailbox and calendar grants are not inventoried; only sends through a hooked agent's tools are held (DR002.004).
Production deploys, new tokens and users, IAM changes and untrusted installs from package registries are held for sign-off; first use of a new repository or cloud profile is reported.
Sign-off and kill on high-impact actions, Untrusted installs and typosquatted packages, Learned per-agent drift
MoorAI does not know which accounts are privileged; it judges the action, not the account behind it.
MoorAI scans no retrieval index, corpus or RAG pipeline.
Memory writes are not governed; MoorAI's memory-poisoning rule is guidance only.
Data-carrying images and links in the agent's output are flagged, external sends are held for sign-off, new destinations are reported and model endpoints can be allow-listed.
Limit: No general destination allow-list; hosts are recorded only from http(s) URLs.
Rendered-output exfiltration and deceptive links, External sends, Per-agent destination map, Model-endpoint allow-list and transit overrides
MoorAI is not a sandbox and does not set or check an execution environment's network boundary.
The provenance of the agent's own model, system prompt and build is not verified.
Policy sets, per threat, whether a call proceeds, is held for a person's sign-off, is blocked, or ends the agent's session.
Rules files the agent auto-loads are scanned for planted overrides and re-scanned when they change; attempts to extract the agent's instructions are detected.
Limit: The agent's local instruction files; a hosted system prompt and the model's own instruction priority are not seen.
Auto-loaded context scan, System-prompt extraction and echo, Skill Analysis
Whether a model's trained objective matches the organisation's purpose is a property of the model, not of an action MoorAI sees.
A coding agent's sub-agents are recorded, their delegated prompts scanned and the parent's envelope applied to them; orphan sub-agents and unusual fan-out are flagged.
Limit: One agent and its sub-agents on one machine.
The orchestrating agent's delegations are recorded and scanned, and the parent's envelope bounds what each worker may do.
Limit: Sub-agents of one hooked agent; orphan and agent-to-agent detection needs lineage metadata on the agent's events.
Peer agents owned by other teams, vendors or organisations are outside the machine.
Delegated prompts are scanned on the way to a sub-agent, and the sub-agent's report is scanned as untrusted content on the way back.
Limit: Sub-agent reports are scanned in Claude Code; the other agents' adapters forward only web results.
Shared memory stores and scratchpads are not scanned.
AI delivered inside a vendor's platform runs on that vendor's infrastructure.
AI delivered inside a vendor's platform runs on that vendor's infrastructure.
AI delivered inside a vendor's platform runs on that vendor's infrastructure.
AI delivered inside a vendor's platform runs on that vendor's infrastructure.
SAR3 Invocation 10 of 21 credited
A risky call aimed at something the operator's request never named is flagged as outside the stated task.
Limit: Lexical, over four kinds of action, and report-only unless policy raises it to sign-off.
An authorised request whose resulting upload, deletion, credential read or MCP write reaches beyond what the request named is flagged.
Limit: Lexical: an upload to a host the user named passes, even when it is exfiltration to that host.
MoorAI does not authenticate who prompts the agent; it runs inside the user's own session.
Instructions embedded in files, fetched pages, search results, command output, MCP results and auto-loaded context are detected at the stage where the agent ingests them.
Injection detectors, Post-tool inbound scan, Auto-loaded context scan
Tool output, MCP results and sub-agent reports are scanned as untrusted content before the agent acts on them.
Limit: Reported and flagged to the model; withheld only for MCP results through the proxy or by a mask, because a post-tool hook cannot un-run a tool.
Post-tool inbound scan, MCP result scanning, Sub-agent oversight
MCP tool listings, call arguments and results are scanned.
Limit: MCP resources, prompt templates and sampling requests pass through unscanned.
MCP tool-listing scan, MCP call-argument scanning, MCP result scanning
Tool descriptions, schemas and parameter text are scanned when a server lists its tools, before any call.
Limit: Servers wrapped by MoorAI's MCP proxy; reported at listing time and enforced at call time.
ANSI and OSC terminal escapes, Unicode tag-block characters and bidi overrides are flagged in tool metadata and in content the agent reads.
Limit: Tool metadata only for servers wrapped by MoorAI's MCP proxy.
MCP results are scanned before the agent ingests them, and a denied result is replaced with a tool error.
MCP resource reads pass through the proxy unscanned.
MCP prompt templates pass through the proxy unscanned.
A server advertising a tool name another server already owns is reported.
Limit: Name collisions through MoorAI's MCP proxy, report-only; metadata that steers use of a differently named tool is caught only where a poisoning pattern matches.
MCP sampling requests pass through the proxy unscanned.
Instructions set to fire on a later condition are not detected as such.
Instruction files the agent loads at every session start are scanned for planted instructions and re-scanned when they change.
Limit: CLAUDE.md, AGENTS.md and their siblings; Claude Code's auto-memory directory and a hosted assistant's memory are not in the scanned set.
No detector for an agent re-launching, scheduling or spawning itself, or changing its own timeouts and launch settings.
No detector for an agent re-launching, scheduling or spawning itself, or changing its own timeouts and launch settings.
No detector for an agent re-launching, scheduling or spawning itself, or changing its own timeouts and launch settings.
No detector for an agent re-launching, scheduling or spawning itself, or changing its own timeouts and launch settings.
No detector for an agent re-launching, scheduling or spawning itself, or changing its own timeouts and launch settings.
Multi-turn analysis exists but has no enforcement caller: it scores in test corpora and enforces nothing in the product.
SAR4 Adverse Outcome 13 of 16 credited
Secrets, personal data and the agent's rules files are detected on their way out through commands, tool arguments, uploads and rendered output; a local secret value leaving is blocked.
Secrets and personal-data detection, Local secret-value egress, Uploads read whole, Rendered-output exfiltration and deceptive links, Rules-file leak detection
A markdown image, link preview or embedded element whose URL carries conversation data is flagged in the agent's output.
Limit: The URL is read; image bytes the agent sends are not OCR'd.
MCP call arguments are scanned before the server receives them, and a local secret value in a tool argument is blocked.
Email, SMS and webhook sends are held for sign-off and their arguments scanned.
Limit: Send-message tool families only.
A file an uploading command reads, or an MCP tool that sends names, is scanned as data leaving; model weights and datasets moved into an upload are flagged, and rules files uploaded by path are reported.
Limit: Shell uploads the hook can parse and files named in MCP arguments, within size caps; an archive written in one call and uploaded in a later one is not tied.
Uploads read whole, Files named by MCP arguments, Model and dataset collection, Rules-file leak detection
Secrets and personal data from the agent's context are detected in what it sends, and verbatim copies of its rules files or system prompt are reported.
Limit: Recognised data classes and verbatim copies; a paraphrase or translation of a rules file is not matched.
Secrets and personal-data detection, Rules-file leak detection, System-prompt extraction and echo
Disclosure to another user or tenant inside a platform is not visible from the device.
Payment and bank-detail change wording in an agent's commands and tool arguments is held for sign-off.
Limit: Wording only; transfers made in payment systems, and synthetic voice or video, are not seen.
Irreversible commands and destructive MCP calls are held for sign-off, and unusual deletion volume across a session raises the next deletion to sign-off.
Destructive commands and tool calls, Sign-off and kill on high-impact actions
An agent deleting or rewriting its own session transcripts is held for sign-off, and MoorAI's own evidence logs are hash-chained so a deleted or edited record shows.
Limit: Local transcript files and MoorAI's own logs; business records are covered only as destructive operations (AO003).
Impersonation and misattribution of who acted is not detected.
Content telling the model to copy an instruction into every reply, file or commit is detected, and prompts delegated to sub-agents are scanned.
Limit: Spread within one machine's agent and its sub-agents; other agents' memories and workflows are not seen.
Single oversized inputs, machine-speed bursts and unusual sub-agent fan-out are flagged.
Limit: Token spend and API cost are not metered, and the burst and fan-out thresholds are not tuned against production traffic.
Code the agent writes is screened for injection-prone patterns, insecure defaults and licence-contaminated blocks.
Limit: False, biased or defamatory text is not judged.
High-impact actions are held for sign-off whether or not an adversary is involved, and actions outside the stated task are flagged.
Limit: A wrong but ordinary action is not judged.
MoorAI is not a sandbox; it does not define or police an evaluation boundary.
SAR5 Opacity 5 of 13 credited
A property of the model's reasoning, which an endpoint hook does not observe.
A property of the model's reasoning, which an endpoint hook does not observe.
A property of the model's reasoning, which an endpoint hook does not observe.
What an agent reports about its own work is not compared with what actually happened.
Every governed call leaves a chain-stamped, content-free record that can stream to a SIEM.
Limit: Prompts, responses, arguments and results are not kept, by design.
Every governed call is recorded with its tool, agent, decision, risk, time and a keyed argument hash, and retention is set by the operator.
Limit: Prompts, responses, arguments and results are not kept, by design.
An agent rewriting its own transcripts is held for sign-off, and a deleted, reordered or edited record in MoorAI's evidence logs breaks the chain.
Limit: Local transcript files and MoorAI's own logs; MoorAI's state lives in the user's home, which the agent can also write.
Downstream side effects are not reconciled with recorded tool calls.
Each record names the tool and a hashed per-session actor, with a distinct id for each sub-agent's own calls; in server mode the actor is the named workload.
Limit: The human behind a session is the device's user; attribution is per device, session or workload, as keyed hashes.
Attribution behind a shared service identity is not resolved.
The source that actually caused a response is not traced back.
A data-carrying image or link-preview URL is flagged whatever host it points at, and a link whose text shows one address while opening another is flagged.
Limit: The URL in the agent's output is read; the client's own image proxy or link-unfurling traffic is not seen.
Logs are content-free by design, so prompts and working context are not preserved for reproduction.
What this is and is not
- It is a reading of MoorAI’s rules and documentation at agent v1.0.0 against the 116 entries of SITM v1.0.1, as published in Forscie’s repository on 26 August 2026, read on 1 October 2026.
- It is not a test. Nothing was run. A credit means a mechanism is documented and shipped, not that it was measured against the entry.
- It is MoorAI scoring itself on a framework another vendor’s researchers helped write. The incentive to read generously is ours, which is why every credit names its mechanism and its limit.
- It is not the view of the matrix, of Forscie or of Above Security. The mapping is ours, and neither has reviewed it.
- It is a mapping of the entries, the ways things go wrong. Whether a product implements the matrix’s 53 preventions and 44 detections is a separate question; the 6 detections that match the warning signs are the only ones mapped here.
- It is pinned to v1.0.1. The matrix is a living framework, and a later version can add, merge or reword entries.
Sources
- Forscie, Synthetic Insider Threat Matrix, v1.0.1, file read at commit
3b778fa; web edition at insiderthreatmatrix.org/synthetic. - Above Security, Introducing the Synthetic Insider Threat Matrix, 11 September 2026.
- Above Security and Forscie, launch release, 27 August 2026.
- N12 (tech12), report on the matrix, for the five warning signs.
- MoorAI agent v1.0.0: README, detection engine, capability spec, rule base and MCP proxy.
Synthetic Insider Threat Matrix™ and Insider Threat Matrix™ are owned by Forscie Limited. Insider Threat Matrix is a trademark of Forscie Limited. Copyright © 2024–2026 Forscie Limited; the matrix is licensed under the Apache License, Version 2.0. Mapping MoorAI to its entries does not imply sponsorship, endorsement or affiliation by Forscie Limited.