← Back to blog

Honeycomb vs. ClawMetry: an honest comparison for AI agent observability

· 10 min read · By Vivek Chand

“We use Honeycomb for everything. Should we just send our agent traces there too?” I hear this from teams who have embraced the “observability 2.0” philosophy: wide events, high cardinality, arbitrary slicing via BubbleUp. Honeycomb is a genuinely great tool. And for a specific class of AI-agent problem, it is exactly right. For another class, it leaves you staring at event rows with no idea why your sub-agent spent $14 overnight. This post tries to be honest about which is which.

I built ClawMetry, so I have a stake in this. If anything reads as spin, call it out on GitHub.

What each tool actually is

Honeycomb is a SaaS observability platform built around the “wide event” model: instead of pre-aggregating metrics or writing traces with a fixed schema, you send arbitrary JSON blobs with as many columns as you want, then query them interactively with BubbleUp and group-by. Honeycomb’s core insight, articulated by co-founder Charity Majors, is that high-cardinality, high-dimensionality data (user IDs, request IDs, error messages, feature flags, all in the same event) lets you ask questions you did not know you needed to ask. The product is OpenTelemetry-first and has strong integrations for traces, logs, and metrics. Their pricing is event-volume-based.

ClawMetry is an open-source observability dashboard purpose-built for 33 AI agent runtimes, including OpenClaw and Claude Code to Codex, NVIDIA NeMoClaw, Cursor, Goose, and more. The core abstraction is the agent session: what tasks ran, which tools fired, what sub-agents were spawned, what cron jobs executed, what memory files changed, and what the whole tree cost. Zero instrumentation required for supported runtimes. pip install clawmetry && clawmetry and you have a live dashboard.

The fundamental difference: Honeycomb measures traces and events you explicitly instrument; ClawMetry measures agent sessions that already exist on your filesystem and gateway. Honeycomb requires you to define what to observe. ClawMetry already knows what an agent session looks like.

Quick comparison

Dimension Honeycomb ClawMetry
Instrumentation Send events via OpenTelemetry SDK, Honeycomb SDK, or direct API. Code changes required to define what goes in each event. Zero-config: auto-detects 33 agent runtimes from their filesystems and gateways. No code changes, no SDK, no decorator.
Data model Wide events: arbitrary key-value JSON, organized into datasets and traces. You define the schema. Session → tool_call → sub-agent → cron_job → memory_diff. Schema is agent-native, predefined.
OSS? No. Honeycomb is closed-source SaaS. Their SDKs and OTel contrib libraries are open; the backend is not. Yes, full stack is MIT on GitHub.
Self-hosted No self-hosted option. All data goes to Honeycomb’s infrastructure. Local-first by default. The dashboard runs on your machine. Cloud sync is opt-in and E2E-encrypted.
Data residency Cloud-only. Prompt content, tool call arguments, model responses: everything you include in events is sent to and stored on Honeycomb’s servers. Agent content stays local by default. E2E-encrypted AES-256-GCM before any cloud sync; server sees ciphertext.
Pricing model Event-volume pricing. Free tier: 20M events/month. Team plan from $20/month. Above free tier, cost scales with how many events your agents emit. OSS free forever. Cloud-Pro (alerts, fleet, Slack/PagerDuty, >24h retention) on flat subscription.
Query power Excellent. BubbleUp surfaces anomalies across any dimension automatically. GROUP BY any column, arbitrary WHERE clauses, heatmaps, derived columns. The best interactive query UX in observability. DuckDB-backed queries on the session data model. Strong for agent-specific questions (cost by runtime, session cost over time, cron failure rate). Not a general-purpose trace query engine.
AI-agent concepts No native concepts for cron jobs, memory diffs, sub-agent cost attribution, loop detection, or multi-node fleet. You build these by designing event schemas and writing queries. Cron job health, memory file diffs, sub-agent trees with per-node cost, multi-node fleet view, 33-runtime adapter set, Guard policy enforcement.
Time-to-first-insight Design your event schema, instrument your agent code with the OTel SDK, configure the Honeycomb exporter, create a dataset. 1–3 hours for a team comfortable with OTel. pip install clawmetry && clawmetry. Under 60 seconds on any supported runtime.
Alerting Trigger-based alerts on any query, PagerDuty/Slack/OpsGenie out of the box, SLO burn-rate alerts, anomaly detection. Mature and flexible. Budget alerts, custom rules by signal rate or spend, Slack/PagerDuty in Cloud-Pro. Agent-aware (alert on signal rate above threshold per runtime) but less compositional than Honeycomb.

Where Honeycomb’s event volume pricing bites on agentic workloads

Honeycomb’s free tier is generous: 20 million events per month. For a small team running agents occasionally, that goes a long way. But the economics change when you move to autonomous, long-running agents.

Consider an OpenClaw agent running a research loop. A single session might invoke 80–120 tool calls, each generating a span. If each span is a Honeycomb event, 100 sessions a day is 10,000 events. Scale to a team with five developers each running agents continuously: 50,000 events per day, 1.5M per month, well under the free tier. But add a cron-triggered fleet running overnight and the numbers look different. A fleet of 20 agents running 200-tool-call sessions every night generates 4,000 events per run, 80,000 per night, 2.4M per month from the fleet alone. Once you exceed the free tier, you’re paying per event on top of the base plan, and the bill scales with how capable your agents get.

This is not a Honeycomb-specific problem; it is inherent to volume-based observability pricing. Datadog has the same dynamic. The relevant comparison is not “Honeycomb costs $X” but “as my agents run more tool calls, what happens to my observability bill?” ClawMetry OSS stores all session data locally; the cost of more agent activity is disk, not dollars.

The agentic billing test: an autonomous agent runs 500 tool calls overnight. Honeycomb: 500 more spans toward your monthly event budget. ClawMetry OSS: $0, stored locally in DuckDB. Cloud-Pro is a flat subscription that doesn’t scale with agent activity.

Honeycomb’s superpower is also its setup cost for AI agents

Honeycomb’s wide-event model is genuinely powerful precisely because you define the schema. You can put anything in an event (the prompt, the tool name, the model, the session ID, the user ID, the feature flag, the retry count) and then GROUP BY any combination of those fields to find patterns you didn’t anticipate. BubbleUp is particularly good at surfacing which dimension correlates with a latency spike or an error spike.

The cost of that power is that you have to design the schema and instrument the code. For a standard web service, this is well-understood work: the OTel semantic conventions tell you what to include in an HTTP trace, a database span, a queue consumer. For AI agents, the conventions are less settled. What exactly goes in a “tool call” span? What columns distinguish a looping agent from a productive one? How do you capture the sub-agent tree structure in a flat event model?

Teams that instrument their agents well for Honeycomb end up reinventing many of the concepts ClawMetry has built in. The sub-agent relationship requires a parent-span model. The cron job status requires either a dedicated Honeycomb dataset or a column in your existing session events. Memory file diffs are not a trace concept at all; they live in your filesystem, and getting them into Honeycomb means writing code to detect and emit them as events.

ClawMetry reads all of these from the runtime directly, without instrumentation. It knows what an OpenClaw memory file looks like. It knows where Claude Code writes its session transcripts. It understands the gateway WebSocket protocol. That is domain-specific knowledge that Honeycomb, as a general-purpose trace store, cannot provide.

Where Honeycomb genuinely wins

I want to be honest about where the comparison goes the other way.

Arbitrary query flexibility on fully instrumented data. Once you have instrumented your agents and have events flowing into Honeycomb, you can ask questions ClawMetry cannot answer: “Which model version correlates with the highest rate of tool call retries?” “What is the p99 latency of write_file calls broken down by directory?” “Show me all sessions where a user with role X triggered the expensive path.” These require dimensional data you define, and Honeycomb’s query engine is the best in the industry for ad-hoc slice-and-dice.

Correlation with non-agent systems. Your agent calls an external API. That API is slow. You want to see the agent’s tool call latency alongside the downstream API’s response time alongside your database query time, all in one trace waterfall. Honeycomb handles distributed traces across services natively; it is what the product was designed for. ClawMetry is scoped to the agent session; it does not see your database or your upstream APIs.

Production web applications with embedded agents. If your product is a SaaS app where users interact with an AI agent inline, you may already have Honeycomb or OTel instrumented for the rest of the app. Adding LLM spans to an existing Honeycomb dataset keeps the observability story unified: one tool, one query language, one alert system. ClawMetry is optimized for agent-native runtimes (OpenClaw, Claude Code, Codex, NeMoClaw); it is less natural as a component in a custom web-app trace pipeline.

Multi-environment SLA tracking. Honeycomb’s SLO feature lets you define a reliability target on any query (for example, “95% of agent sessions complete in under 10 minutes”) and track burn rate automatically. ClawMetry has SLA policies too (via /api/sla/status), but Honeycomb’s SLO implementation is more mature and composable on arbitrary trace data.

When to pick Honeycomb (not us)

Use Honeycomb for AI agents when:

  • You are already on Honeycomb and your team is fluent in wide events. The marginal cost of adding OTel spans for LLM calls to an existing Honeycomb deployment is low. You already know BubbleUp, you already have alert integrations, and you do not want a second observability tool.
  • Your primary question is “why is this session slow?” not “what did this session do?” Honeycomb’s trace waterfall and BubbleUp are the best in class for latency debugging. If your agent is making dozens of LLM calls and you want to find which one is the bottleneck, Honeycomb’s UI is purpose-built for this.
  • Your agent runs inside a larger distributed system. Agent traces that span a web frontend, a queue, a database, and an LLM call belong in Honeycomb. The cross-service trace is Honeycomb’s native unit; ClawMetry’s scope stops at the agent boundary.
  • You need custom dimensional analysis your team defines. Your business has dimensions ClawMetry does not know about: customer tier, experiment arm, geographic region, deployment version. Honeycomb lets you put all of these in your events and slice by any combination. ClawMetry’s dimensions are fixed to what an agent session exposes.
  • Your security and compliance team has already vetted Honeycomb for your data classification. Re-running a data residency review for a second tool is expensive in enterprise. If Honeycomb is already approved and you have appropriate data filtering in place, that is a real operational advantage.

When to pick ClawMetry

  • You run any of the 33 supported runtimes and want a live dashboard in under a minute. No instrumentation, no schema design, no OTel exporter config. pip install clawmetry && clawmetry.
  • Data residency is non-negotiable. Prompt content and tool call arguments stay on your machine by default. AES-256-GCM E2E encryption before any cloud sync. Fully open-source, so you audit the claim rather than trusting a privacy policy.
  • You need agent-native visibility: crons, sub-agents, memory diffs. Cron job scheduling and failure history, sub-agent cost attribution in a tree view, memory file diffs per session, loop detection via four trajectory detectors: these are ClawMetry-native. Building them on top of Honeycomb wide events is a multi-week project that recreates ClawMetry with vendor lock-in.
  • Your observability cost should not scale with agent activity. Honeycomb’s pricing scales with event volume. As your agents get more capable and run more tool calls, the bill grows. ClawMetry OSS has no per-event meter.
  • You want Guard policies: autonomous pause, stop, kill. Honeycomb observes. ClawMetry also intervenes: pause a looping agent via POSIX signal, enforce budget limits via the enforcement proxy, deny an approval in the queue. None of this is possible in Honeycomb; it is not an agent control plane, it is an event store.

The simple version: Honeycomb is the right choice when you need best-in-class ad-hoc query power on fully instrumented agent events, especially when those agents run inside a larger distributed system. ClawMetry is the right choice when you run supported runtimes and want zero-instrumentation, local-first, agent-session-native observability with a cost that does not scale with your agent’s activity.

Using both

These tools operate at different abstraction layers and can coexist without conflict.

The pattern that makes sense for larger teams: ClawMetry observes the agent session layer: cron health, memory diffs, sub-agent cost, loop detection, Guard enforcement. Honeycomb observes the infrastructure and distributed-trace layer: LLM call latency broken down by model, database queries the agent’s tools trigger, cross-service request flows.

ClawMetry exposes an OTLP receiver on /v1/traces and /v1/metrics, so if your agent emits OTel spans, ClawMetry can ingest them alongside its own auto-detected session data. You can also route those same spans to Honeycomb. The two stores answer different questions from the same event stream.

ClawMetry sub-agent tree view alongside a Honeycomb trace waterfall
ClawMetry shows the agent session tree: which sub-agent ran, what it cost, what memory changed. Honeycomb shows why each LLM call inside that tree was slow. They answer different questions.

Bottom line

Honeycomb is the observability tool I recommend to teams who want to ask arbitrary questions about arbitrary data and have the discipline to instrument everything upfront. The wide-event model, BubbleUp, and the query language are genuinely the best in the industry for what they do.

The gap is domain knowledge. Honeycomb does not know what an OpenClaw cron job is, or that Claude Code writes session transcripts to ~/.claude/projects/, or that a sub-agent cost attribution tree requires following parent-session pointers across JSONL files. ClawMetry knows all of this, because it was built for exactly this domain. The zero-instrumentation story is not a compromise; it is the product. You should not have to define what to observe when the runtime already exposes it.

If you are on Honeycomb and want to add agent-native visibility on top, ClawMetry runs alongside it without conflict. If you are starting fresh with a supported runtime, ClawMetry gets you a live dashboard in under a minute. If you need Honeycomb’s dimensional query power and your team can invest in instrumentation, Honeycomb is worth it, especially once your agents are running inside a larger distributed system you already observe there.

See what your AI agents are doing

Zero instrumentation. Local-first. Fully open source.

Get ClawMetry free
Cookie preferences