CClawMetryDocs

Guard & governance

Detectors#

Eight small, explainable heuristics over a session's recent event stream. Each one is pure — no I/O, no store, no clock beyond what the caller passes — bounded to the newest events, and incapable of crashing on a malformed event.

Why heuristics and not a judge#

The research position behind this is explicit: zero-shot LLM judges are close to useless at localising the bad step, and are 17-27× slower. Naive embedding-outlier scoring dilutes the single anomalous step into the average.

What is load-bearing is sequence structure — the shape of the tool stream. So these are small structural heuristics, not a model call. They run on every session on every cycle without being a cost line of their own.

Trajectory: is this agent stuck?#

stuck_loop#

K consecutive identical (tool, arguments-hash) calls, or a short repeating n-gram cycle of tool names.

ThresholdDefaultEnvironment
Identical calls3CLAWMETRY_LOOP_IDENTICAL_K
Maximum cycle length4CLAWMETRY_LOOP_MAX_CYCLE
Cycle repeats3CLAWMETRY_LOOP_CYCLE_REPEATS

The cycle form catches the loops the identical form misses — read → edit → test → read → edit → test is not three identical calls, and it is very much a loop.

no_progress#

At least N tool calls in the window with zero file writes or edits and no completion marker. Busy, but not advancing.

ThresholdDefaultEnvironment
Tool calls20CLAWMETRY_NOPROG_TOOLS

"Write" is resolved per runtime, because tool vocabularies differ — see Thresholds. For shell-first runtimes where every edit happens inside shell or exec, a shell-mutation rule is what keeps this meaningful: Codex editing through a heredoc is making progress, and a naive write-tool check would call it stuck.

repeated_tool_failure#

The same tool errors at least M times in the window.

ThresholdDefaultEnvironment
Failures3CLAWMETRY_REPEAT_FAIL_M

Usually a configuration problem rather than a model problem — a missing credential, a wrong path, a permission — which is why it is worth alerting on separately from the loop detectors.

action_discrepancy#

A failed tool result immediately followed by the agent continuing — another tool call, or a completion — without retrying that tool or acknowledging the error.

This is the narrow form of what the TRAIL taxonomy calls tool-related hallucination. It has lower precision than the others, so it carries lower severity and honest wording: "agent continued after a failed command", not "agent hallucinated".

ThresholdDefaultEnvironment
Occurrences1CLAWMETRY_ACTION_DISCREPANCY_MIN

Behavioural: is this agent doing something it does not normally do?#

These compare against a learned baseline for the cohort, not a fixed idea of normal. An agent that touches forty files every run is not flagged for touching forty files.

file_blast_radius#

Distinct files mutated in a window, well above the cohort baseline.

ThresholdDefaultEnvironment
Files25CLAWMETRY_BLAST_FILES

The signal that catches a refactor that got out of hand, and the one that matters most for "how much would I have to revert?"

credential_access#

Reads of credential-shaped paths — key files, environment files, token stores. Benign shapes are excluded explicitly: .env.example, .env.sample, .env.template, .pub files.

An agent reading .env.example is doing its job. An agent reading .env and then making network calls is a different session.

network_egress#

Distinct outbound hosts in a window, above the cohort's normal fan-out. Local and loopback hosts are excluded.

ThresholdDefaultEnvironment
Distinct hosts8CLAWMETRY_EGRESS_HOSTS

privilege_change#

Attempts to escalate privilege or disable a protection — the class of action that should never be a surprise.

What an incident looks like#

json
{
  "kind": "stuck_loop",
  "session_id": "codex:…",
  "runtime": "codex",
  "severity": "warning",
  "title": "codex looping: 38 tool calls, same shell command ×12",
  "detail": "The same command has run 12 times with no file change. Pause or Stop it from the Guard tab.",
  "evidence": {"identical_calls": 12, "window": 200, "tool": "shell"},
  "first_bad_step": 143,
  "spend_at_risk_usd": 4.20,
  "spend_basis": "derived_from_tokens",
  "threshold_source": "learned_baseline"
}

Three fields are worth calling out:

  • first_bad_step — a 0-based index into the events, so the UI can jump you

to where it started rather than the top of a 2000-event session.

  • spend_at_risk_usd — the flagged stretch, not the session.
  • threshold_sourcemodule_default, runtime_profile,

learned_baseline or env_override. This is how you tell a measured threshold from a shipped constant.

The window#

Every detector looks at the newest CLAWMETRY_DETECT_WINDOW events (default 200). That bounds CPU per session, and it is also the honest scope: a loop that ended four hundred events ago is history, not an incident.

Turning them off#

VariableEffect
CLAWMETRY_DETECTORS=0Disable trajectory detectors
CLAWMETRY_STUCK_DETECT=0The same switch, older spelling

There is no reason to disable them for cost — they are structural checks over an in-memory window, not model calls.

From incident to action#

An incident on its own only appears in the UI. Turning one into a pause or a kill requires a policy, and every policy action passes the three locks.

Cookie preferences