A CLI that turns the OpenTelemetry traces in ClickHouse into a Markdown report — built to be read by an agent, not just a human.
It is the trace-side sibling of
parcareport. parcareport turns a
Parca profiling server into a cross-cluster CPU bottleneck report;
otelhousereport turns a ClickHouse full of spans into a "where did latency
and errors go" report over a time window, normalized so a 1-hour window and a
24-hour window are directly comparable.
The traces are the ones the OpenTelemetry Collector's
clickhouseexporter
writes: the otel_traces table (and, with --logs, otel_logs). This tool only ever reads them.
$ export CLICKHOUSE_DSN='clickhouse://ro:***@ch:9000/otel'
$ otelhousereport --from=-24h# otelhousereport
- **Source:** ClickHouse table `otel_traces`
- **Window:** `2026-08-26T17:56:57Z` .. `2026-08-27T17:56:57Z` (24h0m0s)
- **Spans:** 46,456 in 8,570 traces across 4 service(s)
- **Errors:** 500 (1.1% of spans)
- **In flight:** 0.592 spans on average (self-time ÷ wall-time)
## Where time goes — by service
| SERVICE | INFLIGHT | %TIME | CALLS | ERRORS |
| :---------------------------- | -------: | ----: | -----: | -----: |
| unknown_service:dagger-engine | 0.372 | 62.9 | 14,548 | 173 |
| agentloop | 0.219 | 37.0 | 30,932 | 323 |
| **total** | 0.592 | 100.0 | 46,456 | 500 |
## Hottest operations (by self-time)
| SERVICE | OPERATION | CALLS | SELF | AVG | P95 | P99 | ERR% |
| :-------- | :-------------- | ----: | ----: | -----: | -----: | -----: | ---: |
| …engine | resume withExec | 508 | 4.3h | 30.75s | 5m02s | 5m03s | 3.5 |
| agentloop | POST /graphql | 6,244 | 1.8h | 1.05s | 1.70s | 2.36s | 0.0 |
| agentloop | tick | 2,447 | 17m11s| 6.82s | 15.15s | 20.38s | 12.8 |INFLIGHT is self-time ÷ wall-time: the average number of spans of a kind
running at once over the window. 0.5 means that, on average across the whole
window, half a span of that kind was in flight.
This is the point of the tool, and it is the same idea as parcareport's
CORES. Raw span counts are not comparable between a 1-hour window and a
24-hour one, and they conflate a handful of very slow spans with a flood of fast
ones. Self-time-per-wall-time is an absolute rate: "agentloop keeps 0.22 spans
busy" means the same thing whatever window you pick, and you can watch it move
across a change.
Spans nest: a parent span's duration already includes its children's. So you
cannot sum Duration across spans to learn where time went — a request that
calls three services would be counted four times.
otelhousereport ranks by self-time instead: a span's own duration minus
the time its child spans covered. Summed across a group, self-time adds up to
real elapsed work with no double counting, which is why the %TIME column sums
to 100.
--by an id-like attribute (span:http.url, a request id) can have thousands
of distinct values, and one Markdown row each is an unusable wall. The
breakdown is capped at --max-groups (default 50) rows by self-time, and
everything below the cap is summed into a single (other) row — so the table
stays bounded but the totals stay complete and %TIME still sums to 100. A note
says how many values were folded and how to raise the cap; --max-groups=0
disables it. The remainder is measured, not dropped: filling the (other) row
correctly costs one extra aggregate, and it runs only when the cap actually
bites, so an ordinary low-cardinality report pays nothing.
Self-time is an approximation, and the tool says so in its own footer.
Children can overlap each other, or a child recorded on a different host can
carry enough clock skew to run "longer" than its parent; the tool floors each
span's self-time at zero so neither corrupts the sum, but a deeply concurrent
service will have its self-time slightly understated. Judge the ranking, not the
third decimal — or pass --exact-self-time.
--exact-self-time replaces the "minus the sum of child durations" step with
"minus the union of child intervals", computed in ClickHouse with arrayFold
over each parent's sorted child spans. Overlapping children are then counted
once, so the number is exact rather than a slight under-count. It is a heavier
query (a groupArray + fold per parent), which is why it is opt-in.
Borrowed wholesale from parcareport, because the failure modes are the same.
An empty answer is not the same as no problem. A window that misses the data
looks identical to an idle cluster looks identical to a wrong --table. Rather
than assert the convenient reading, the tool cross-checks and names the reason:
no spans in 2035-01-01T00:00:00Z .. 2035-01-02T00:00:00Z; otel_traces holds
442,876 spans from 2026-08-18T15:43:53Z .. 2026-08-27T17:57:10Z — widen
--from/--to to cover that
table "otel_nope" not found; otel tables present: otel_logs, otel_traces — set --table
no spans in <window> matching --match service="doesnotexist"; drop --match or
check the value with `otelhousereport services`
A partial report is never presented as complete. The report is several
independent queries; any one can fail on a slow window or a restarting server.
If a section fails, its numbers would silently drop out and the totals would
still look whole. So a failed section is called out in an ⚠️ INCOMPLETE block
inside the report, next to the numbers it invalidates, and the command
exits non-zero — so an agent that checks the status does not mistake an
incomplete report for a clean one.
go install github.com/guettli/otelhousereport@latestotelhousereport [report] [flags] the report (default)
otelhousereport services [flags] list services seen in the window
otelhousereport operations [flags] list operations seen in the window
otelhousereport tables [flags] list the otel_* tables present
| Flag | Default | Meaning |
|---|---|---|
--dsn |
$CLICKHOUSE_DSN |
ClickHouse DSN, e.g. clickhouse://ro:***@ch:9000/otel |
--from |
-1h |
window start: RFC3339, or relative (-6h, -90m, -7d, -1w; compound like 1d6h) |
--to |
now |
window end |
--by |
service |
breakdown column: service, name, kind, status, or an attribute key res:<key> / span:<key> |
--match |
filter to a value; repeatable (values AND together), e.g. service=agentloop, span:http.request.method=POST |
|
--top |
15 |
rows in the operation and error tables (0 = summary only: header + breakdown) |
--max-groups |
50 |
max rows in the breakdown table; the rest are summed into an (other) row (0 = no cap) |
--logs |
false |
add an Error logs section: recurring otel_logs lines correlated to error spans (highest severity first) |
--exact-self-time |
false |
self-time from the union of child intervals (exact, slower) instead of summed child durations |
--out |
write to this file instead of stdout | |
--timeout |
25s |
per-query timeout (see the note below) |
otelhousereport --from=-24h # last day, by service
otelhousereport --by=name --top=30 # hottest operations, wider
otelhousereport --match=service=agentloop --by=name # drill into one service
otelhousereport --by=span:http.request.method # break HTTP spans down by verb
otelhousereport --by=res:tenant --from=-24h # by a resource attribute
otelhousereport --match=service=agentloop --match=span:http.request.method=POST
otelhousereport --from=-24h --logs # + why the errors happened
otelhousereport --from=-7d --out=report.md # write a file for an agent
otelhousereport services --from=-24h # what services exist--by and --match accept keys inside the OTel attribute maps as well as the
four fixed columns: res:<key> / resource:<key> for ResourceAttributes
(e.g. res:tenant, res:k8s.namespace.name) and span:<key> / attr:<key>
for SpanAttributes (e.g. span:http.request.method, span:http.response.status_code).
The key is bound as a query parameter, so arbitrary attribute names — dots,
slashes — are safe. Spans lacking the key group under (none).
The whole report is one self-contained Markdown document with stable section
headings (## Where time goes, ## Hottest operations, ## Errors). Point an
agent at otelhousereport --from=-6h (or --out=report.md) to get a compact,
parseable answer to "where is time and where are the errors right now" before
it goes digging in individual traces.
-
--timeoutalso caps ClickHouse'smax_execution_time. The clickhouse-go driver turns the per-query deadline into that server setting, and a read-only user profile commonly constrains it to<= 30s. The default is25sfor that reason; raise it only if the server allows a larger cap, or a wide-window query will be rejected outright rather than merely slow. -
Slow, or timing out, on a multi-tenant ClickHouse? Add
?enable_analyzer=0to the DSN. When you read through a per-tenant read-only user whose rows are scoped by a row policy — the usual multi-tenant setup — you can hit ClickHouse bug #82369: with the new analyzer (the default since 24.x), a row policy makes the planner ignore theotel_*skip indexes and partition pruning, so even a narrow query full-scans the table. It runs for tens of seconds, trips--timeout, and on a memory-tight server one such query can wedge (and OOM-restart) ClickHouse for every tenant. Pinning the old analyzer restores index + partition pruning — verified on 25.8: a trace-id lookup dropped from a 30 s timeout to ~20 ms, same rows, isolation intact:export CLICKHOUSE_DSN='clickhouse://ro:***@ch:9000/otel?enable_analyzer=0'
clickhouse-go forwards DSN query params as session settings. If you operate the server, prefer setting it once on the tenant's profile —
ALTER SETTINGS PROFILE tenant_ro SETTINGS enable_analyzer = 0— so every reader benefits without touching each DSN. Revisit when #82369 is fixed upstream (the old analyzer is deprecated). -
Start narrow. The self-time breakdown is a windowed self-join of the traces table; a wide window over a busy table is real work for the server.
-
Read-only by design. The tool issues no DDL, no writes, and never interpolates a user-supplied value into SQL — times,
--matchvalues and attribute keys are always bound as native{name:Type}parameters. The only identifiers that reach the SQL by name — the table, the group/filter column, and theResourceAttributes/SpanAttributesmap name — are whitelisted.
Apache-2.0