otelhouse is two things built around the
OpenTelemetry → upstream Collector → ClickHouse pipeline:
- A deployed artifact — the multi-tenant OTLP gateway. A custom
OpenTelemetry Collector distribution (
tenantauth+tenanttagger+tenantratelimitaround the stockclickhouseexporter) published as a container image,ghcr.io/guettli/otelhouse-gateway. Producers authenticate with their Kubernetes ServiceAccount token — the identity the kubelet already issues and rotates for every pod — or, if they are not in-cluster pods, with a minted JWT verified against a static public key. It is live: deployed from gitops (epic gitops#73) and already serving agentloop as a real tenant — agentloop pushes OTLP with its token and reads its rows back as theagentloop_roClickHouse user. Many tenants share one ClickHouse with per-tenant write and read isolation. This is the reusable output of the repo. - A Dagger-orchestrated end-to-end harness for the
Dagger → OTLP → Collector → ClickHousestack, proving the pipeline works as an integration unit.
otelhouse is the write path. Reading is
otelhouseview: the
otelhouseview/otelstore library (a read-only, typed client over the stock
otel_traces / otel_logs tables) plus the service and UI built on it. This
repo ships no query API of its own — the e2e harness here reads back what it
wrote through that same library.
There is nothing to go get here — the shipped artifact is the gateway
image. The high-level design comes from epics
#32 and
#53.
The parts that make up an OTLP → ClickHouse pipeline each already work on
their own, but nothing bridged them under the exact constraints we need:
one shared ClickHouse, the stock clickhouseexporter schema (so every
tenant gets Grafana's default OTel dashboards for free), hundreds–thousands
of tenants, with per-tenant credential-bound write isolation and
per-tenant read isolation.
- The stock OpenTelemetry Collector +
clickhouseexporterwrites OTLP → ClickHouse with the canonicalotel_*schema — but it is single-tenant: one shared token, no per-tenant identity, no per-tenant write enforcement, no read isolation. - Full products on ClickHouse — Uptrace, SigNoz, ClickStack / HyperDX — either ship their own non-stock schema (Uptrace bundles its own schema plus PostgreSQL and its own UI) or are single-tenant in OSS (SigNoz, ClickStack). Multi-tenanting them means running one whole stack per tenant.
- No tool provided the exact bridge: stock schema, one ClickHouse, many tenants, credential-bound at both write and read time.
otelhouse is the thin bridge that adds only the missing piece — a
tenantauth extension (Kubernetes ServiceAccount tokens, or minted JWTs)
plus a tenanttagger processor that stamps the tenant onto every record
from the verified token — while keeping every
other part stock: stock OTLP receiver, stock clickhouseexporter and
schema, ClickHouse row policies for reads. It deliberately does not
reinvent the exporter or the schema.
See #53 and the gitops epic guettli/gitops#73 for the broader multi-tenancy design.
The custom Collector distribution lives in collector/:
-
extension/tenantauthverifies the bearer token on every OTLP request and resolves the tenant it may write as into the auth context. Two identity sources, either or both enabled:- Kubernetes ServiceAccount tokens (the default for in-cluster
producers). A projected SA token is an RS256 JWT signed by the API
server; the extension verifies it against the cluster JWKS
(
/openid/v1/jwks, cached, refreshed on an unknownkid), requires the producer to have projected it withaudience: otelhouse-gateway— so a token minted for the API server is not replayable at the gateway — and derives the tenant from the signed ServiceAccount identity: the namespace claim, or an explicit<namespace>/<serviceaccount>→ tenant map for producers whose namespace is not their tenant. Unmapped identities are rejected, never defaulted. No secret to mint, distribute or rotate. - Static-PEM minted JWTs (EdDSA/ES256/RS256) for any producer that is not an in-cluster pod: verified against a public key in the config, with the tenant in a signed claim.
Algorithm pinning (never HS*, never
alg:none),iss/aud/expchecks, unknownkid, unmapped ServiceAccounts and tenant-spoofing attempts are covered byextension_test.go. - Kubernetes ServiceAccount tokens (the default for in-cluster
producers). A projected SA token is an RS256 JWT signed by the API
server; the extension verifies it against the cluster JWKS
(
-
processor/tenanttaggerreads the authenticated tenant fromclient.Infoand stamps it asresource.attributes["tenant"], deleting any client-supplied value. It is fail-closed: a batch arriving without a resolved tenant is dropped, not written unlabelled (processor_test.go). -
processor/tenantratelimitenforces a per-tenant ingest rate keyed off the same signed claim. Over-limit batches are rejected with gRPCRESOURCE_EXHAUSTED/ HTTP429and counted onotelhouse_gateway_ratelimit_dropped_total{tenant}— no silent drop. Limits are configurable per tenant with a global default; buckets are in-memory per-replica. -
builder-config.yamlis the ocb config that pins upstream receiver/processor/exporter versions and pulls in the three custom components — nothing else changes on the write path, so the stockclickhouseexporterschema is preserved unchanged. -
docs/jwt-contract.mdis the wire contract for both identity sources (claims, algorithms, iss/aud, how the tenant is derived) — and why the tenant may only ever come from a verified token. -
docs/metrics.mddocuments the per-tenant ingest and auth-rejection Prometheus metrics the gateway emits — the operator's view into "who is sending how much" and "which tenant just went silent because a token expired."
Read isolation is enforced by the ClickHouse row policies and is correct at
any tenant count. Read performance has a ceiling worth stating honestly,
the same way the per-tenant rate limiter's per-replica bucket ceiling is stated
above. The tenant lives in ResourceAttributes, not in the sort key of the
stock schema (otel_logs orders by
(toStartOfFiveMinutes(Timestamp), ServiceName, Timestamp), otel_traces by
(ServiceName, SpanName, toDateTime(Timestamp)); both PARTITION BY toDate(Timestamp)). A per-tenant read is pruned only by the stock
INDEX idx_res_attr_value … TYPE bloom_filter(0.01) skip index — at a 1%
false-positive rate — and every tenant shares the same date-partitioned parts,
so surviving granules still contain other tenants' rows that the row policy
filters out after the granule is read. Isolation stays correct; the cost is
read amplification that grows with how many tenants co-occupy the parts a query
touches.
Keeping the pure stock schema is a deliberate design goal, so otelhouse ships no
tenant-leading index. The escape hatch is the one upstream already sanctions:
set create_schema: false and manage your own DDL — adding a tenant-leading
data-skipping index, a projection, or the tenant into the ORDER BY — to trade
away pure-stock parity for a flatter read curve. That is an operator's
knowing choice, not something otelhouse decides on their behalf. The full
analysis (and an offer to turn the ceiling into a measured number via a Dagger
e2e benchmark) is in
#75; the read-path trade-off
is also now documented upstream in the clickhouseexporter Performance Guide.
The gateway is built and published as ghcr.io/guettli/otelhouse-gateway by
CI on every push to main (see ci/). It is deployed and operated from
gitops, not from here — this repo owns
the code and the image; gitops owns the running stack, the shared ClickHouse,
the keypair and the per-tenant secrets. The production wiring lives under
k8s/plain/otelhouse/ in gitops and is described in epic gitops#73.
A tenant connects to the running gateway like any OTLP endpoint, plus a token.
In-cluster producers (the default): no secret at all. The gateway is a ClusterIP Service with no IngressRoute, so every producer today is a pod in the cluster — and every pod already has a cryptographically verifiable identity the kubelet issues and rotates. Project it for the gateway:
volumes:
- name: otlp-token
projected:
sources:
- serviceAccountToken:
audience: otelhouse-gateway
expirationSeconds: 3600
path: token- Mount that volume and send OTLP with
Authorization: Bearer $(cat /var/run/secrets/otelhouse/token)(buildOTEL_EXPORTER_OTLP_HEADERSfrom the file the kubelet keeps fresh). - The gateway verifies the token against the cluster's JWKS, checks it was
projected for this audience, derives the tenant from the signed
ServiceAccount identity (namespace, or an explicit SA→tenant map entry),
stamps
ResourceAttributes['tenant']and writes via the stock exporter. An identity that maps to no tenant is rejected — never defaulted. - Reads are isolated by ClickHouse row policies bound to a per-tenant
<tenant>_rouser, so a tenant sees only its own rows and can point Grafana's default OTel dashboards straight at the shared database.
Nothing to mint, nothing to rotate, nothing to expire: the audience: line
in the Pod spec is the whole enrolment.
Producers that are not in-cluster pods: the operator mints a per-tenant
JWT with the private key held in gitops and the tenant sends it as
Authorization: Bearer <jwt>; the gateway verifies it against the configured
public key and reads the tenant from the signed tenant claim. Those tokens
are short-lived and re-minted before expiry operator-side (gitops#89), so the
signing key never enters the cluster.
Both paths are documented claim-by-claim in
docs/jwt-contract.md, which also spells out why the
tenant may only ever come from a verified token: every other tenant's
row-policy isolation depends on that label being unforgeable, so a
client-supplied tenant resource attribute is always overwritten, and a
batch whose tenant cannot be derived from a verified claim is dropped.
The deployed reality — producers authenticate with their Kubernetes
ServiceAccount token (or, if they are not in-cluster pods, a minted JWT), the
gateway derives and stamps the tenant from the verified claims, the stock
exporter writes, and reads are constrained by ClickHouse row policies bound to
a per-tenant <tenant>_ro user:
flowchart LR
subgraph Producers
Agentloop["agentloop<br/>(live tenant)"]
Prod["other tenants"]
end
subgraph GW["otelhouse-gateway (ocb distro)"]
direction LR
Auth["tenantauth<br/>(cluster JWKS /<br/>static PEM)"]
Tag["tenanttagger<br/>(stamps tenant,<br/>fail-closed)"]
RL["tenantratelimit<br/>(per-tenant)"]
Exp["stock<br/>clickhouseexporter"]
Auth --> Tag --> RL --> Exp
end
K8S[("cluster JWKS<br/>/openid/v1/jwks")]
K8S -. "verify SA tokens" .-> Auth
CH[("shared ClickHouse<br/>otel_traces / otel_logs<br/>otel_metrics_*<br/>row policies on<br/>ResourceAttributes['tenant']")]
RO["per-tenant reads<br/>as <tenant>_ro"]
View["otelhouseview<br/>(separate repo)"]
Agentloop -- "OTLP + projected<br/>SA token" --> Auth
Prod -- "OTLP + SA token<br/>or minted JWT" --> Auth
Exp -- "SQL INSERT" --> CH
CH --> RO --> View
Separately, ci/ exercises a stock-collector harness: a Dagger pipeline
drives sample OTLP into an upstream otelcol-contrib (config in
ci/otel-collector-config.yaml) pointed at an
ephemeral ClickHouse, then reads the rows back with otelhouseview/otelstore
and asserts them. That harness has no tenant layer — it proves the
OTLP → Collector → ClickHouse plumbing and the stock schema, not the gateway's
multi-tenancy (the gateway's own components are covered by their unit tests and
by the gateway image build in ci/gateway.go).
To be precise about what is and is not custom: the gateway is Go code —
tenantauth, tenanttagger and tenantratelimit all sit directly on the write
path (otlp → tenanttagger → tenantratelimit → memory_limiter → batch → clickhouseexporter). What is codeless is the last hop: writing into
ClickHouse. There is no custom exporter, no custom schema and no migrations
in this repository.
- The upstream
clickhouseexporterwrites traces, logs and metrics directly. It is pulled in unmodified as a pinned upstream module bycollector/builder-config.yaml, soSHOW CREATE TABLE otel.otel_tracesis byte-identical to a fresh stock install. (The shipped gateway is an ocb-built binary, not theotelcol-contribdistribution — contrib is only used as the collector in theci/harness.) create_schema: truemakes the exporter create theotel_traces,otel_logsandotel_metrics_*tables (plus their materialized views and TTLs) on startup — no migrations to run.- The in-process Go exporter that previously lived here was removed in #25; it duplicated the upstream exporter and was a maintenance liability.
Producers speak plain OTLP/gRPC — but in production they are not entirely
otelhouse-agnostic: a producer must present Authorization: Bearer <token> (its
projected ServiceAccount token, or a minted JWT if it is not an in-cluster pod),
or tenantauth rejects the request and tenanttagger fail-closes and drops the
batch rather than writing it unlabelled. Only the ci/ harness producer is
config-only agnostic, because that harness runs a stock collector with no auth on
a private Dagger network.
Secrets sometimes slip into telemetry — an access key logged in an error string, a token stuffed into a span attribute. otelhouse scrubs them before they land in ClickHouse, and does it without a line of write-path Go:
- The rules come from gitleaks: a
pinned copy of its default ruleset lives in
collector/gitleaks.toml. collector/gen-gitleaks-rulesis an offline generator (run by hand, not at build time) that expands each rule's regex into stocktransform/OTTL statements and writescollector/redaction.yaml— committed, so the full expanded regex set shows up in review. The header pins the gitleaks version it was generated from.redaction.yamlis a full config fragment defining the processor. The collector deep-merges it in as a second--configfile, so the base config only referencestransform/redactionby name in its pipelines. (It is merged rather than embedded via${file:...}: value-embedding the large regex set wraps every statement in confmap's expansion machinery, which fails to resolve — "too many recursive expansions".)- The processor runs ahead of
batchon the logs and traces pipelines, replacing any matching substring in a log body or a log/span attribute value withREDACTED:<ruleID>. It is the unmodified upstreamtransformprocessor(added tobuilder-config.yaml); only the rule pack is generated, so this stays consistent with "stock components only".
Refreshing after a gitleaks release is mechanical: bump the version in
collector/gitleaks.toml, rerun the generator, review the redaction.yaml
diff, commit. The ci/ harness proves it end to end — redaction_test.go
sends a fake-but-rule-matching AWS key straight at the Collector and asserts the
row in ClickHouse carries REDACTED:aws-access-token and never the literal.
OTel deliberately specifies ingestion (OTLP, semantic conventions) but not a query API: how you ask "show me the spans for run X with their child logs" is left to whatever store you chose. ClickHouse is no different — it gives you SQL, not OTel.
That thin read layer is not in this repository. It lives in
otelhouseview:
github.com/guettli/otelhouseview/otelstore is a read-only, typed client over
the stock otel_traces / otel_logs tables (ListTraces, GetTrace), and
the viewer's service and UI are built on it. It is deliberately tenant-blind:
the isolation boundary is the ClickHouse identity in the DSN plus row policies,
not a filter in Go. The e2e harness here is a consumer of that library like any
other.
For ad-hoc SQL the generic tooling that ships around ClickHouse already
works against the otel_* tables and needs no setup beyond a DSN:
- ClickHouse's built-in HTTP Play UI (
http://<host>:8123/play) — bundled with the server, good for quickSELECTs. - Tabix — browser-only SQL IDE talking to the HTTP interface.
- The Grafana ClickHouse plugin — dashboards and ad-hoc exploration.
Those tools answer "run an arbitrary SQL query"; they do not render a trace as a waterfall or stitch logs onto spans. That trace-shaped view is what otelhouseview adds on top.
Per-tenant read isolation is enforced by ClickHouse row policies bound to
the <tenant>_ro user, and it holds at any tenant count: a tenant cannot see
another tenant's rows, because the label those policies key on is stamped from
a verified token and never client-supplied. What does change with tenant count
is read latency — a property of the stock schema this repo deliberately
keeps.
The tenant lives in ResourceAttributes['tenant']. In the stock schema that
map is reachable through a skip index — INDEX idx_res_attr_value mapValues(ResourceAttributes) TYPE bloom_filter(0.01) GRANULARITY 1 (or a
text index where the exporter's full-text search is enabled) — but it is
not part of the sort key:
| Table | ORDER BY |
PARTITION BY |
|---|---|---|
otel_logs |
(toStartOfFiveMinutes(Timestamp), ServiceName, Timestamp) |
toDate(Timestamp) |
otel_traces |
(ServiceName, SpanName, toDateTime(Timestamp)) |
toDate(Timestamp) |
So a single tenant reading a wide time range gets correct rows, but:
- the bloom filter prunes only some granules, at a 1% false-positive rate, and
- every tenant shares the same date-partitioned parts, so the granules that survive pruning still contain other tenants' rows — the row policy filters them after the granule has been read.
The net effect is read amplification that grows with how many tenants co-occupy the parts a given query touches. Isolation stays correct throughout; it is scan-bytes and latency that degrade as tenant count climbs and dashboard reads run concurrently.
This is a property of the stock schema, not something to fix upstream. The
exporter has no tenant concept to sort by — the tenant label is otelhouse's
addition — and a tenant-leading sort key would regress the ordinary
single-tenant "time range + service" query for every other user of the
exporter. Upstream's own position is that this is the deployer's call: the
clickhouseexporter docs
recommend managing your own schema with create_schema: false for production
workloads, and note that any table shape works as long as the column names and
types still match the exporter's INSERT.
First, rule out the analyzer bug (ClickHouse#82369)
The read amplification above is the schema's inherent cost and grows with tenant count. There is a separate, far more acute failure that is not about tenant count and is not a schema problem — check for it before reaching for the escape hatch below.
With ClickHouse's new analyzer (the default since 24.x), a row policy
makes the planner ignore data-skipping indexes and partition pruning
entirely. This is worse than the "bloom prunes only some granules" case above:
a trace-id point lookup that idx_trace_id prunes to zero granules in
EXPLAIN instead full-scans the whole table at execution. On a
memory-tight ClickHouse one such cold query runs past the query time limit
before it is reaped, backs the exporter's INSERTs up behind it, and fails the
liveness probe — a single /ci/<traceID>-style read can OOM-restart the shared
server and drop spans for every tenant. Isolation stays correct; it is
availability that breaks.
Because it is the row-policy machinery (not the Map lookup) that trips the bug, neither a materialized-column policy nor a plain-column policy avoids it — but the old analyzer does. The fix is one setting on the tenant read profile, no schema change:
ALTER SETTINGS PROFILE tenant_ro SETTINGS enable_analyzer = 0;Verified on 25.8.33.6 against a live <tenant>_ro row policy, fully cold: the
trace-id lookup drops from a 30 s timeout to ~20 ms, same rows, isolation
intact. otelhouse's bootstrap sets this on the tenant_ro profile; ad-hoc
readers can pass it per-connection via the DSN (…/otel?enable_analyzer=0).
Revisit when #82369 is fixed upstream (the old analyzer is deprecated, so a
future ClickHouse image that drops it will reject this setting — remove the line
in the same bump).
The escape hatch, if per-tenant read latency ever becomes the bottleneck
after the analyzer bug is ruled out:
set create_schema: false and create the otel_* tables yourself with the
tenant lifted out of the map — a tenant-leading ORDER BY, a projection, or a
tenant-leading data-skipping index — keeping the columns compatible with the
exporter's INSERT. That buys granule pruning by tenant, and costs this repo's
"SHOW CREATE TABLE is byte-identical to a fresh stock install" property, and
with it Grafana's default OTel dashboards working against the shared database
unchanged. otelhouse keeps the stock schema by default because that trade
belongs to the operator, not to the gateway.
The ceiling above is reasoned from the schema, not yet measured: #75 tracks adding a per-tenant read-latency benchmark to the Dagger harness, so the limit becomes a number rather than an expectation.
The Dagger pipeline in ci/main.go is the single
source of truth for tests. Running it locally is byte-identical to
what GitHub Actions runs, so a green local run implies a green CI run:
make test # == cd ci && go run .The pipeline stands up its own ephemeral, version-pinned ClickHouse via
a Dagger service binding — there is nothing to install or start by hand,
and no separate local stack to keep in sync. A reachable Dagger engine
is the only prerequisite; to use a remote engine, export
_EXPERIMENTAL_DAGGER_RUNNER_HOST before running.
There is intentionally no docker-compose (or other) parallel test
environment: a second definition of ClickHouse would drift from
ci/main.go and break the "green locally ⇒ green in CI" guarantee. See
#33.
The pipeline runs:
gofmt— format checkgo vet— static analysisgolangci-lint— lint (v2.12.2)go build— compilationgo test— integration tests against a live ClickHouse 25.5 service- End-to-end harness — stands up the upstream
otel/opentelemetry-collector-contrib(withci/otel-collector-config.yaml) pointed at the same ClickHouse service, drives sample OTLP traces/metrics/logs into it with the in-repootelhouse-emitbinary, and runs theTestE2E_StoreGo test (build tage2e), which reads the rows back out of ClickHouse throughgithub.com/guettli/otelhouseview/otelstoreand asserts the ingested traces and logs are there. This is theDagger → OTLP → Collector → ClickHouseguarantee for the whole harness — one pipeline run validates the write path end-to-end, read back with the same library the viewer uses.
The upstream clickhouseexporter writes TraceId and SpanId columns
to both otel_traces and otel_logs, so a log emitted inside an active
span joins back to that span with no custom schema:
SELECT t.SpanName, l.Body
FROM otel_traces t
JOIN otel_logs l USING (TraceId, SpanId)For the join to work, producers must emit log records while a span is
active — start the span (tracer.Start(ctx, ...)) before the log call
so the OTel SDK stamps the span context onto the record. A log with an
empty SpanId cannot be linked to a span, and a Collector pipeline that
strips TraceId/SpanId (e.g. via attributes/delete) breaks the
join.
This is the data foundation otelhouseview builds on to render hyperlinks between a span and its logs (and back).
Metrics carry less context than a span or log record, so the upstream
clickhouseexporter writes them into per-type tables — otel_metrics_gauge,
otel_metrics_sum, otel_metrics_histogram,
otel_metrics_exponential_histogram and otel_metrics_summary — and
correlation works at two levels:
Coarse correlation by service / resource. Every signal table — traces,
logs and the metric tables alike — carries ServiceName and
ResourceAttributes, so dashboards pivot on those without any per-row link.
Fine correlation via exemplars. otel_metrics_sum, otel_metrics_histogram
and otel_metrics_exponential_histogram have an Exemplars Nested(TraceId String, SpanId String, ...) column. Producers that record measurements
while a span is active get exemplars stamped with the active TraceId /
SpanId, so a single metric data point joins back to its originating trace
— and from there to logs via the trace/log join above:
SELECT m.MetricName, e.TraceId, t.SpanName
FROM otel_metrics_sum m
ARRAY JOIN m.Exemplars AS e
JOIN otel_traces t USING (TraceId)
WHERE m.ServiceName = 'checkout'Metrics ingestion is alpha in the upstream clickhouseexporter and the
schema may shift between releases. The contrib image tag pinned in
ci/main.go (otelCollectorVersion) is the source of truth for what the
tables actually look like at any commit — the Dagger harness runs an
end-to-end test (otelhouse-emit → Collector → ClickHouse) against that pin
so a green CI run implies the metric tables are populated with the expected
shape.