Leadership1 distinct publisher2 min readPublished
The ClickHouse account credits 16x faster queries and 100 million events a second at Black Friday peak. It puts no dollar figure on either side of the trade.
The Board Room · Leadership desk
Compiled by The Board RoomSomething wrong?How this is made
The transferable part of this account is not the throughput figure but the schema argument. Shopify's team concluded that metrics, logs, traces, exceptions and profiles are one primitive, structured time-ordered high-dimensional events, and put all of them on a single engine so a single question could be asked across all of them [6]. Before that, asking one question across metrics, logs and traces meant stitching answers together by hand across separate vendors [12]. They began with metrics, where the pain was sharpest, then moved the rest onto the same engine [7]. That order is copyable at any size. The volume is not.
The arithmetic is worth doing, because the source leaves it undone. At 100 million events per second and roughly 110 GB/s uncompressed, the average event is about 1.1 KB [2]. Peak is twice steady state, so the headroom designed into the system is 2x, not an order of magnitude [1]. Hold peak for an hour and that is around 396 TB of uncompressed telemetry [3], set against the 90 PB Shopify says it stored across its fleet during Black Friday and Cyber Monday [9]. The query side is easier to place: engineers expect answers in under a minute [4], and a 16x gain implies the same lookup previously sat near a quarter of an hour [4].
What the post does not contain is money. Previously unpredictable vendor costs were brought under control [2], which is a direction rather than a figure; there is no prior bill, no current bill, and no headcount for the platform team that has been rebuilding this since 2021 [11][6]. The piece is published by ClickHouse, drawn from a talk at an event ClickHouse hosted [14], and it still carries engineering director Elijah McPherson's line that he would probably reach for ClickHouse Cloud if he were starting again [5].
Read that next to the original complaint. The 2021 problem was spend growing faster than the fleet and core production tooling tied to roadmaps Shopify did not control [11][12], and the engine won partly because being open source meant Shopify could contribute to it [13]. Self-hosting was the method, not the objective, and the person closest to it now prices the managed option differently.
The setting matters if you intend to model this against your own bill: about 500 Kubernetes clusters, 1.5 million pods across five continents [8], hundreds of producing teams and schemas that keep changing [15]. Somewhere in there, per-gigabyte vendor pricing loses to salaries. Shopify found its crossover. The published account gives you the volume side of that calculation in full and the cost side not at all.
Ranked by verification strength, evidence, and original report placement.
Moving to ClickHouse delivered a 16x query improvement out of the box, spiking past 30x, while bringing previously unpredictable vendor costs under control.
Engineering director Elijah McPherson said that if Shopify were doing it all over again, he would likely reach for ClickHouse Cloud instead of self-hosted ClickHouse.
McPherson said a minute of downtime during the past Black Friday and Cyber Monday can cost Shopify's merchants $5.1 million, and at peak load the platform processes up to $5.1 million in transactions per minute.
Shopify replaced fragmented observability vendors with Observe, its own unified platform for metrics, logs, traces and exceptions, built on ClickHouse.
Ingest runs at around 50 million events per second at steady state and 100 million at BFCM peak.
At BFCM peak the platform handles roughly 110 GB/s of uncompressed telemetry and keeps it queryable in under a minute; engineers expect to query the data in under a minute.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed but wholly first-party and unaudited
The numbers are specific, internally consistent and attributed on the record to a named Shopify engineering director, which is better than anonymous vendor testimony. But there is exactly one source in the cluster, it is the engine vendor's own blog reporting its own conference talk, no baseline system or measurement methodology is given for the 16x/30x figures, no cost figures appear on either side of the trade, and the supplied body truncates mid-lesson. Nothing here is independently verifiable from the cluster.
Production at extreme scale, one datapoint
This is not a pilot: a multi-year, multi-signal production deployment carrying every observability signal for a top-tier commerce platform through its highest-traffic weekend, with a phased migration off three incumbent vendors already completed. Adoption breadth is nonetheless a single company reported by its supplier, with no other named users, no license or pricing change and no second deployment in the cluster.
Real deployment, overstated as a settled win
The underlying deployment is genuine and unusually large, so this is not vapor. The overstatement is in framing: performance multiples and 'costs under control' are presented as a resolved build-versus-buy verdict by the party that sells the engine, while the piece withholds every figure needed to test it — prior bills, current run cost, team size — and its own architect says he would now buy the managed tier instead of the self-hosted path being celebrated. The $5.1 million-per-minute downtime line also restates peak gross transaction throughput as loss.
Vendor-owned channel with a managed-tier upsell
Publisher, venue and framing all belong to ClickHouse: the piece runs on clickhouse.com, is sourced from ClickHouse's Open House SF 2026, credits ClickHouse for the performance and cost outcome, and closes the loop with a customer executive saying he would choose ClickHouse Cloud next time. The customer also has a reputational incentive to present its in-house build favorably. No countervailing source, incumbent-vendor response or independent audit is present.
Deployment reality solid, economics unresolved
High confidence that the deployment exists, its scale is roughly as described and the architectural pattern is real, because the figures are specific, arithmetically coherent and attributed to a named executive. Low confidence in the performance multiples and the build-versus-buy economics, because the cluster contains one vendor-published source, no methodology and no cost data.
build
The AI SRE that is not allowed to guess: Databricks scopes its agent to "what changed"1 distinct publisher
leadership
ClickHouse buys Langfuse, turning a neutral tracing layer into someone's roadmap1 distinct publisher
build
Scale Kafka sinks on lag, not CPU: KEDA to zero, then fix the per-pod drain rate1 distinct publisher
build
SSE in Go breaks twice before your handler runs: an illegal header, then a 30-second timeout1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 26, 2026