CClawMetryDocs

Guard & governance

Thresholds and baselines#

A threshold that is wrong is worse than no threshold. Too tight and every session is an incident; too loose and the one that mattered is not.

ClawMetry resolves thresholds from four layers, each overriding the last.

The four layers#

1. Module default#

A shipped constant. no_progress at 20 tool calls, stuck_loop at 3 identical calls, blast radius at 25 files. These are the fallbacks, and every incident produced by one carries threshold_source: "module_default".

2. The runtime profile#

A checkable fact about the adapter, not an invented number. The profile supplies each runtime's write-tool vocabulary, because "did this agent change a file" is answered differently per runtime:

RuntimeWrite vocabulary
Claude CodeCovered by the module defaults
Codexapply_patch plus shell heredocs, caught by the shell-mutation rule
Qwen Code, Gemini CLI, Antigravityreplace, plus write_file in defaults
GooseNamespaced (developer__text_editor); the default edit substring matches
opencodewrite / edit / patch, all in defaults
PicoClaw, NanoClawShell-first: every change happens inside shell/exec
OpenWorkerreplace_in_file, apply_unified_diff — declared in its own risk module

Without a profile, a runtime that edits exclusively through a tool name nothing matches looks like it made no progress at all. That is the bug this layer exists to prevent.

3. The learned baseline#

The cohort's own history, from the Guard statistics tables. Tool calls per session drives no_progress; distinct files mutated drives blast radius; distinct egress hosts drives network_egress.

The learned value is mean + σ × stddev:

VariableDefaultMeaning
CLAWMETRY_BASELINE_MIN_SESSIONS20Sessions needed before a baseline is used at all
CLAWMETRY_BASELINE_SIGMA2.0How many standard deviations above the mean is "unusual"
CLAWMETRY_BASELINE_FLOOR0.5Lower clamp, as a ratio of the static default
CLAWMETRY_BASELINE_CEIL5.0Upper clamp, as a ratio of the static default

4. A per-runtime environment override#

The base variable with __<RUNTIME> appended:

bash
CLAWMETRY_NOPROG_TOOLS__CODEX=40      # tune codex alone
CLAWMETRY_BLAST_FILES__CLAUDE_CODE=60
CLAWMETRY_REPEAT_FAIL_M__N8N=5

The bare variable sets it for everything; the suffixed form scopes it. This is the layer to reach for when you know something about your workload that the baseline cannot learn.

Why learning is clamped#

Two failure modes, both real, both prevented by the floor and ceiling:

A cohort where everything loops would teach Guard that looping is normal and it would go blind. The floor stops the learned threshold falling below half the shipped default.

A cohort of three sessions has a meaningless mean and a meaningless standard deviation, and would make Guard scream at every fourth session. The minimum session count stops a baseline being used at all until there is enough of it, and the ceiling stops one outlier session dragging the threshold up fivefold.

Reading the current baselines#

bash
curl -s localhost:8900/api/guard/baselines | jq

Each row shows the cohort, the sample size, the computed mean and standard deviation, the resulting threshold, and whether it was clamped. A clamped threshold is worth looking at — it means your fleet is far enough from the shipped assumption that one of the guards fired.

threshold_source#

Every incident carries which layer produced its threshold:

ValueMeans
module_defaultA shipped constant. No measurement behind it.
runtime_profileA fact about the adapter's tool vocabulary
learned_baselineMeasured from your own cohort
env_overrideYou set it

This field exists so a reader can tell "your agents normally touch 8 files and this one touched 60" from "someone picked 25 once". Those deserve different amounts of trust, and hiding the difference would be dishonest about what the number means.

Tuning in practice#

Start by not tuning. Watch the incidents for a week and read threshold_source on each one.

  • Too many incidents on one runtime → raise that runtime's threshold with

the __<RUNTIME> override rather than the global one.

  • Nothing ever fires → check the baseline sample size; you may not have

enough sessions yet for the learned layer, in which case you are on shipped defaults.

  • A threshold shows as clamped → your fleet is genuinely unusual for that

signal. Consider whether the shipped default is the wrong shape for your work rather than nudging the clamp.

The spend model#

Separately from detection thresholds, annotate_spend attaches spend_at_risk_usd to an incident: the estimated cost of the flagged stretch, with the spend_basis that produced it.

VariableDefaultEffect
CLAWMETRY_GUARD_CRITICAL_USD5A warning above this becomes critical

Where no cost is known the value is 0.0 with basis unknown — never an invented figure.

Cookie preferences