Detector libraryPrompt injection

Prompt injection

A page became an instruction.

Spot instruction-like content inside what an agent reads.

See the problem. Understand the finding.

English narration with captions.

How it gets overlooked

You asked an agent to summarize a page. The page tells the agent to do something else first. That is the trust boundary: outside content is trying to become an instruction inside your task.

Imagine a fetched review page containing a message that claims authority and asks the agent to run an installer before finishing the summary. In the middle of useful source material, the instruction can be easy for a human reviewer to overlook.

What ClawMetry detects

ClawMetry checks recorded tool results and user sourced text for declared prompt injection signatures. A matching tool result raises a warning that names the signature and source tool. The finding does not need to repeat the matched text.

The difference that changes the finding

Triggering example

Page tells agent to act

Warning finding

Quiet comparison

Ordinary page content

No finding for this detector.

The checked example places a forged instruction in a fetched page and gets a warning. An ordinary hotel description stays quiet. A later high risk tool call in the same turn can increase severity, which makes the order of events worth reviewing.

Inspect the detector result
{
  "kind": "prompt_injection",
  "severity": "warning",
  "evidence": {
    "source": "tool_result",
    "tool": "WebFetch",
    "signatures": [
      "forged_authority",
      "task_handoff"
    ],
    "signature_labels": [
      "imitates a system or authority message",
      "asks the agent to do something else before the user's task"
    ],
    "tool_result_matches": 1,
    "user_message_matches": 0,
    "followed_by_tool": null,
    "followed_by_risk": null,
    "threshold": 1,
    "threshold_source": "static",
    "observed": "declared injection signatures in the text of tool results and user-sourced messages; matched text is not kept"
  }
}
Download inputs and complete results (JSON)
How the example was checked

These examples evaluate the published detector with authored event data or disposable configuration files. The videos illustrate those behaviors. They are not recordings of live agents or the product interface. No command in the examples was executed.

The result establishes behavior for these inputs. It does not establish runtime ingestion, prevention or a real compromise. Inspect the pinned source contract.

What to check next

Inspect the source that supplied the content and the actions that followed. Keep the original task in view. Did the agent treat outside text as authority, or simply quote it? Use the finding to investigate that boundary with the actual recorded activity.

  1. Identify the content source
  2. Inspect subsequent actions
  3. Compare with the original task

What this signal establishes

These are bounded signature checks. Paraphrases and other unmatched forms can be missed; a match does not prove the agent obeyed.

Keep the important moments visible.

Follow agent activity, inspect findings and decide what needs your attention.