DDIA Chapter 1: Reliability — Why ClawMetry Cannot Be the Thing That Breaks
I keep a copy of Designing Data-Intensive Applications open while I work on ClawMetry. Not as a reference I look things up in, but as a mirror. Every few weeks a chapter maps to a decision I either made well, made badly, or still haven’t made at all. This is the third post in that series. Today it’s Chapter 1: Reliability.
The previous two were Chapter 3 (why we replaced SQLite with DuckDB) and Chapter 5 (a real incident where replication lag blocked the cloud relay for two minutes forty seconds). This one is more foundational. Chapter 1 asks a question I should have answered before writing any code: what does it mean for an observability tool to be reliable?
What Kleppmann means by reliability
“The system should continue to work correctly (performing the correct function at the desired level of performance) even in the face of adversity (hardware or software faults, and even human error).”
— Martin Kleppmann, Designing Data-Intensive Applications, Chapter 1He distinguishes three types of faults: hardware faults, software errors, and human errors. The goal is fault tolerance — making the system continue working when individual components fail, not building components that never fail. No component never fails.
The reason I find this framing sharp for observability tools specifically: an observability tool sits one layer above the thing you care about. If it breaks, it can drag the thing you care about with it. An agent monitoring tool that crashes your agent is worse than no monitoring at all. The meta-constraint for ClawMetry is that it cannot be the fault the agent inherits.
Let me go through Kleppmann’s three fault categories and the specific decisions they forced in ClawMetry.
Hardware faults: the local DuckDB file
Kleppmann’s solution to hardware faults is redundancy: RAID, dual power supplies, hot-standby servers. The premise is that disks fail, RAM corrupts, machines die. Build so that no single hardware failure stops the system.
ClawMetry’s original data layer was an in-memory store that synced to a remote database. When the remote database was unavailable (network outage, cloud provider hiccup, billing issue), the dashboard went blank. You were running agents blind. The monitoring tool had added a new single point of failure: the cloud database.
The fix was to move the primary data layer onto the local machine. Every session transcript, every tool call, every cost figure lands first in a local DuckDB file at ~/.clawmetry/clawmetry.duckdb. The dashboard reads from that file. Cloud sync is an additional path, not a required one.
Local-first DuckDB
All data writes go to a local DuckDB file first. The dashboard reads locally. Cloud sync is additive: it extends retention and enables the web dashboard, but its absence does not blank the local view. If our cloud is down, your dashboard still works.
There’s an important detail in how DuckDB fits this pattern. It’s an embedded database: no separate server process, no network socket, no daemon to restart. The database lives in a file you can copy, back up, or inspect with the DuckDB CLI. It fails exactly the way a file fails — predictably, visibly, and without taking anything else down with it.
Software errors: process isolation between the daemon and the server
Hardware faults are largely random: a disk fails, a power supply dies, a network card goes bad. Software errors are different. Kleppmann calls them “systematic errors within the system” — bugs that lie dormant under specific inputs and then trigger across every instance simultaneously when those inputs arrive. A hardware fault takes one machine. A software bug can take all of them.
For ClawMetry the relevant question was: if a bug in the sync daemon causes it to crash, does it also crash the dashboard? And vice versa?
Early in the codebase, both ran in the same Python process. A panic in the ingest loop — a malformed JSONL line, an unexpected schema field, a corrupt session file — would take down the entire Flask server. The dashboard went blank not because of a hardware fault but because one bad transcript crashed the thing reading it.
We split them into separate processes. The sync daemon (clawmetry sync) holds the DuckDB writer lock and runs the ingest loop, detectors, and cloud relay. The Flask server (clawmetry) reads from DuckDB via a localhost query server the daemon exposes. The two processes share no in-memory state. A daemon crash surfaces as stale data in the dashboard, not as a dead dashboard.
Daemon and server are separate processes
The sync daemon and the dashboard server run as independent OS processes. They communicate only through DuckDB (reads) and a localhost query socket (live queries). A panic in either process is contained: the other continues serving what it can from the last good state.
Kleppmann notes that fault tolerance for software errors often requires careful thought about assumptions and interactions at system boundaries. The boundary between the daemon and the server is one of those interfaces we’ve spent the most time hardening. Every ingest path that handles user data — a JSONL transcript line, a gateway WebSocket frame, an OTLP span — wraps its parse in a try/except that logs the malformed input, increments a counter, and continues. One bad line does not stop the next thousand.
Human errors: zero-config as fault prevention
Kleppmann’s analysis of human errors is the sharpest section of Chapter 1. Operators, he notes, cause more outages than hardware. The solution is not to blame operators: “The best systems combine several approaches.” One of those approaches is making it easy to do the right thing and hard to do the wrong thing — specifically, reducing configuration surface.
ClawMetry’s original setup required you to point it at your OpenClaw workspace, configure a gateway URL, set a token. Every configuration item was a place a user could make a mistake. And when they did, the errors were silent: wrong workspace path meant an empty dashboard, wrong gateway URL meant no live data, wrong token meant auth failures the UI didn’t distinguish from “nothing is running.”
The zero-config principle isn’t a UX nicety. It’s a reliability property. Fewer configuration decisions means fewer configuration mistakes. We auto-detect the workspace path, auto-discover the gateway on its default port, auto-find session files, auto-identify which runtimes are installed. The install path is:
pip install clawmetry && clawmetry
That’s it. No config file. No environment variables required. No pointing at paths. The dashboard either works or it surfaces a specific, actionable error rather than rendering blank.
Auto-detect everything
ClawMetry discovers workspace paths, runtime installations, gateway addresses, and session files automatically. Users should never need to configure anything manually. Every path that used to require a configuration item is now an auto-detection pass with an explicit fallback and a human-readable error when it fails.
The meta-constraint: observation without dependency
These three decisions — local-first storage, process isolation, zero-config — follow from Kleppmann’s fault categories. But there’s a fourth constraint that’s specific to observability tools and that Chapter 1 hints at without naming directly.
An observability tool is only valuable if it doesn’t add fragility to the system it watches. If ClawMetry sits in the critical path of your agent — if your agent calls ClawMetry before each tool execution, or if ClawMetry is a dependency that your agent imports — then a bug in ClawMetry becomes a bug in your agent. We’ve been very deliberate about not doing this.
ClawMetry reads from existing files and connects to existing sockets. It does not sit between your agent and its tools. It does not require your agent to import an SDK. It does not add latency to tool calls. The sync daemon reads JSONL transcripts that already exist. The hook mechanism (for runtimes that support it) is designed to fail open: if the ClawMetry gate is unavailable, the tool call proceeds without it.
Fail-open on all control paths
Every path where ClawMetry makes a decision about a running agent — hook gates, approval checks, budget limits — fails open on ambiguity. If the entitlement lookup errors, the agent keeps running. If the hook gate is unreachable, the tool call proceeds. A billing or connectivity bug in ClawMetry must never stop a customer’s agent.
This constraint shaped a rule we made explicit in the codebase: fail open on entitlement, closed on policy. If a licence lookup errors or returns ambiguously, the default is “allow.” Only a policy the user explicitly wrote may block or kill. A bug in our entitlement service should not be visible to the running agent.
What I got wrong the first time
Kleppmann says something in Chapter 1 that I think is underappreciated: reliability is not about building components that never fail. It’s about building systems that continue working when components do fail.
I built the first version of ClawMetry with the assumption that the cloud connection would always be available. When it wasn’t, the dashboard went blank. I built the daemon and server in the same process because it was simpler. When a bad transcript crashed the daemon, it crashed everything. I built a configuration step because configuration felt like flexibility. Every configuration item became a support ticket.
The local-first design, the process split, and zero-config are not features. They are corrections. Chapter 1 names the fault categories; each category came with a decision I had made wrong.
What’s next in this series
DDIA chapters I’ve covered so far:
- Chapter 3: Storage & Retrieval — why we replaced SQLite with DuckDB
- Chapter 5: Replication — the heartbeat incident and what it taught us
- Chapter 7: Transactions — how we chose our isolation levels
- Chapter 1: Reliability — this post
Next up will be Chapter 8 (The Trouble with Distributed Systems), which maps to the question of what happens when two ClawMetry nodes disagree about an agent’s state. That one has a real incident to anchor it.
See ClawMetry running on your agents
Local-first, zero config, E2E encrypted. Works across 32 agent runtimes.
868k+ installs. pip install clawmetry and you’re done.