03 · Observability · data engineering
HTTP traffic analytics on ClickHouse
From a CDN-analytics spike to a shipped query API: a ClickHouse engine that gives customers a real-time view of requests, bandwidth, status codes and latency percentiles — sub-second, on pre-aggregated data.
- <1s
- p50 / p90 / p99 queries
- 1 min
- pre-aggregation grain
- 4
- telemetry sources unified
The problem
Launch had server logs, edge logs and cloud-function logs, but no way for a customer to see their own traffic the way a Cloudflare dashboard shows it. Percentiles computed over raw request rows don't return fast enough to put in front of a person.
Pipeline first
Before the analytics engine could exist, the pipeline under it needed designing and proving. I designed an exporter service that handles four telemetry sources — server, edge, cloud-function and HTTP traffic logs — pushing to customer destinations while landing a queryable copy in Elasticsearch, with OpenTelemetry metrics behind usage analytics.
I benchmarked it rather than assuming it would hold: pod-level capacity testing under high concurrency to establish the single-pod RPS ceiling, the memory and CPU saturation points, and how backpressure behaves when a customer's destination or Elasticsearch slows down.
The engine
I evaluated ClickHouse as the analytics store and pushed on the two things that decide whether it survives production — per-pod request capacity under stress and the cost curve of the store — then built the engine on it. The schema is where the work is: MergeTree engines, LowCardinality(String) for the repetitive header columns, and ZSTD compression to hold storage cost down.
One-minute pre-aggregation tables on AggregatingMergeTree, fed by materialized views, move the cost from query time to insert time. p50, p90 and p99 come back sub-second. On top of it sits the query API in the logs service that the traffic analytics dashboard runs on.