Skip to content

Build1 publisher3 min readPublished

SigNoz self-hosting is no longer a compose file: ClickHouse 25 breaks the quickstart

An engineer deploying SigNoz into an isolated Docker network reports that ClickHouse 25.5.6 rejects the old config mount, and the OTel collector will not boot before a schema migrator has finished.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • SigNoz's official repository makes self-hosting look easy, presenting it as running a docker-compose up script, described as a 10-minute setup before returning to business logic.
  • The author self-hosted SigNoz while architecting the infrastructure for Verne Software, in order to take complete control of telemetry data.
  • Official documentation often lags behind major architectural shifts in the underlying Docker images; what was meant to be a quick deployment became a deep dive into ClickHouse configurations, OpenTelemetry bugs and networking quirks.
  • Symptom: mounting a custom clickhouse-config.xml exactly as older tutorials suggest results in ClickHouse refusing to start, throwing errors about missing paths or invalid settings at the top level.
  • With ClickHouse 25.5.6 the configuration paradigm has shifted: replacing the main config.xml outright is described as a recipe for disaster, and certain settings have been strictly recategorized.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

SigNoz's official repository still presents self-hosting as a single `docker-compose up` away, a ten-minute job before you get back to business logic [1]. According to a dev.to write-up by an engineer who built the observability stack for Verne Software, that promise now breaks on contact with the current images, because the documentation lags behind architectural shifts in the underlying Docker images [2][3].

The first trap is ClickHouse. On 25.5.6, mounting your own `clickhouse-config.xml` over the default, exactly as older tutorials instruct, leaves the server refusing to start, with errors about missing paths or invalid top-level settings [4][5]. The author reports a specific crash loop: `Code: 137. A setting 'log_queries' appeared at top level in config. But it is user-level setting that should be located in users.xml inside <profiles> section.` [6]. The fix is to stop replacing the main config and instead drop overrides into `config.d/` and `users.d/`, with query logging declared inside a user profile in a `users.xml` mounted at `/etc/clickhouse-server/users.d/users.xml` [7][8]. The author's dividing line: cluster topologies and UDF paths belong in `config.d/`, profiles and passwords in `users.d/` [9].

The second trap is that schema creation is now strictly decoupled from application logic [11]. Start ClickHouse and the OpenTelemetry collector together, as you could in older iterations, and the collector crash-loops with `Database signoz_traces does not exist` [10][12]. Three databases are involved, signoz_traces, signoz_metrics and signoz_logs, and none of them will exist unless something creates them first [12][13].

That something is a pair of ephemeral containers, which is the real change in shape here. The author's working boot sequence runs `signoz_init_clickhouse`, a one-shot container that downloads UDF binaries such as `histogramQuantile` [14], then ZooKeeper as ClickHouse's coordinator [15], then ClickHouse itself [16], then `signoz_telemetrystore_migrator` to handle schemas [17], and only then the collector and the SigNoz API and UI [18]. Two gates enforce it: the migrator waits on ClickHouse with `condition: service_healthy`, and the collector waits on the migrator with `condition: service_completed_successfully` [19][20]. Two of those six components exist only to run to completion and exit before any long-running consumer is allowed to start [21].

Networking is the third element, and the published excerpt is thinner here. The stack sits in an isolated `verne_observability` network, with the SigNoz API and UI container additionally attached to an internal `verne_internal` network so the core API can proxy dashboard requests without exposing the database or internal collector ports outward [22][23]. The author lists networking quirks among the traps encountered, alongside ClickHouse configuration and OpenTelemetry bugs, but the material available stops before that section is worked through [24].

What to watch: whether the upstream compose file grows these `depends_on` conditions, and whether the ClickHouse image tag in your own deployment floats. A stack whose correctness depends on ordered one-shot containers and on which directory a setting lives in is a stack where an unpinned minor version upgrade is a production incident. Treat the quickstart as a demo and the compose file as something you own.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories