Build1 publisher3 min readPublished
Docker's healthcheck defaults make a ready container wait almost half a minute
Docker's healthcheck defaults probe every 30 seconds with no start period, so a container ready at second 2 can wait almost half a minute to turn healthy. Setting start_period and start_interval yourself shortens rollouts, as long as a proxy or Compose acts on the health status.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- A post republished on dev.to recommends probing every 2 seconds inside a 30-second start period, then every 15 seconds with a 3-second timeout.
- Probe failures that happen during start_period do not count toward the container's retry limit.
- start_interval arrived with Docker Engine 25.0 and needs a recent Compose release; older versions may reject the field or ignore it.
- A probe that queries the database can turn every replica unhealthy during a brief database hiccup, leaving a health-aware proxy nothing to route to.
- Running docker compose up -d --wait --wait-timeout 120 returns only once services are running and healthy, or fails when the timeout expires.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Startup failures stop being forgiven once the window ends, so start_period has to be sized from the slowest warmup actually observed.
- cost Fast startup probing multiplies the probe's own cost across every replica during a rollout, so a handler that locks, writes or calls out adds load at deploy time.
- exposure A curl-based check in a curl-less image never passes, and every service gated on it with service_healthy stays held back with it.
The post argues that Docker's defaults are tuned for safety, not speed, and that every timing field should be set explicitly [14]. The delay comes from scheduling. Under the defaults it lists, the first probe fires one full interval after the container starts, and that interval is 30 seconds [1][2]. A check that would pass on its first run still has to wait for that run. The defaults also set the timeout equal to the interval, at 30 seconds each [2]. The example config puts the fix in a comment on the config line, `timeout: 3s # keep well below interval`, next to a 15-second interval [5]. That makes the timeout one fifth of the interval [2].
start_interval is what removes the wait for the first probe, but its scope is limited. It only applies while the container is inside start_period, so setting it on its own does nothing [3]. With no start period, the default startup rate is one probe per 30 seconds. The example's 2-second start_interval probes 15 times as often [1]. The post wants each of those probes to answer one question: can this process serve requests. Its probe checks only in-process state: the server is listening, migrations have finished, required caches are loaded [19]. Dependency checks go to a separate diagnostic endpoint that monitoring calls [19].
The half-minute figure only carries over to your service under one condition. The app has to be ready well before the first default probe would have fired [2]. The time saved is roughly 30 seconds minus the app's real time to ready. It drops toward zero as warmup approaches 30 seconds [4].
Health status only keeps traffic off a cold container if something acts on it. In the post, that is either a proxy that honors health status or Compose ordering through depends_on with condition: service_healthy [7][11]. The post says fixed waits like sleep 20 are either too long or too short, and the health gate replaces them [11]. It blames both of its opening failures on healthchecks that are misconfigured or missing. One is a 502 from a container still warming its caches. The other is a pipeline waiting two minutes on a container that was ready in ten seconds [15].
The exec form is the portable way to write the test. A JSON array after CMD runs the binary directly, while a plain string goes through /bin/sh [9]. Distroless images sometimes ship without a shell [8]. The post's preferred test is a subcommand built into the app's own binary, `["/app/api", "healthcheck"]` [18]. In my view this is the right design. The check ships inside the binary it tests, so it cannot go missing from the image. Two smaller traps remain. localhost can resolve to IPv6 ::1 while the app listens only on IPv4, so the post probes 127.0.0.1 [10]. When a Dockerfile has several HEALTHCHECK instructions, only the last one takes effect [13].
The post gives two outcomes for engines too old for start_interval: they may reject the field or ignore it [4]. Rejection is the better of the two. An ignored field raises no error, so the faster startup probing can disappear without anyone noticing. The post tells readers to check the defaults against the official HEALTHCHECK reference for their installed version [16].
What to watch
- Whether the official HEALTHCHECK reference for a given Engine version matches the 30s/30s/3/no-start-period defaults the post lists.
- Whether a pinned older Compose release rejects start_interval or silently ignores it, which decides whether the fast probing fails loudly.
- Measured warmup times against the example's 30-second start_period, since a slower start puts real failures on the retry count.