How it gets overlooked
A build fails. The agent tries again. It fails again. Each retry adds more output, but the underlying problem is still there. If you only check whether the session is active, you can miss the failure that keeps repeating.
Consider an agent preparing a change while its shell tool repeatedly returns an error. The cause might be a missing dependency, an invalid command or unavailable access. The useful question is which tool keeps failing, and how often.
What ClawMetry detects
ClawMetry counts errors attributed to the same tool in the event window. When the count crosses the applicable threshold, it raises a repeated tool failure warning. The finding names the tool, the failure count and the threshold used.
The difference that changes the finding
Triggering example
4 errors from Bash
Warning finding
Quiet comparison
1 error from Bash
No finding for this detector.
In the checked example, four failing shell results produce a warning. A single failure stays below the threshold. This is different from an identical-call loop: the commands can change while the same tool continues to return errors.
Inspect the detector result
{
"kind": "repeated_tool_failure",
"severity": "warning",
"evidence": {
"tool": "Bash",
"failures": 4,
"threshold": 3,
"threshold_source": "static"
}
}Download inputs and complete results (JSON)How the example was checked
These examples evaluate the published detector with authored event data or disposable configuration files. The videos illustrate those behaviors. They are not recordings of live agents or the product interface. No command in the examples was executed.
The result establishes behavior for these inputs. It does not establish runtime ingestion, prevention or a real compromise. Inspect the pinned source contract.
What to check next
Open the failing results and look for the common cause. Fix the environment or access problem, or narrow the agent task before letting retries accumulate. Then verify that the next tool result actually succeeds.
- Read the error results
- Fix the shared cause
- Verify the next result
What this signal establishes
This is a reliability signal. It does not diagnose the root cause or prove the agent is under attack.