Build1 publisher2 min readPublished
Self-hosted LiveKit loses its SIP trunks whenever a stock Redis container is recreated
Self-hosted LiveKit keeps SIP trunks in Redis, so recreating a stock Redis container wipes them and reprovisioning returns new trunk IDs. Health checks stay green until an outbound call fails on a trunk ID that no longer exists.
The Engineer · Build desk

What happened
- Production self-hosted LiveKit runs four containers, the SFU, the SIP service, Redis and a TLS terminator on port 443, and most setups health-check only the SFU.
- Rokas Remeika's first fix turns on Redis append-only persistence with a real mounted data directory, so a container recreate leaves the trunk store intact.
- Because the stack runs on host networking with no container network namespace, he also binds Redis to the loopback interface with protected mode on.
- His second fix resolves trunks by name and caches the ID; on the missing-trunk error alone it evicts that ID, re-resolves by name and retries exactly once.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Any team storing a LiveKit trunk ID as a key in its own database has to move to the trunk name as the durable reference and treat the ID as a cache it can refresh.
- constraint The self-heal stays safe only while its error match is pinned to the pre-INVITE missing-trunk error; widen it and a recovery path can dial a real person twice.
- exposure Monitoring that probes only the SFU will report a self-hosted stack as healthy while its telephony configuration has been emptied.
Rokas Remeika, writing on dev.to on 16 September 2026, puts the root cause in the application's schema [16]. Writing a service-minted identifier into your own database couples your persistence to someone else's volatile cache, he wrote, with no mechanism to notice when the two diverge [12]. After a wipe, the trunk comes back under the same name with an entirely different ID [13]. "The trunk name is the stable thing. The ID is just a cache," Remeika wrote [11].
The single retry in his fix depends on where the error fires. According to Remeika, the missing-trunk error is raised before any SIP INVITE leaves the box, and that ordering is the only reason a retry is safe [9]. "A broader match that catches failures after the INVITE has been sent turns a self-heal into a second real phone call to a real person," he wrote [10]. The person answering will not experience it as self-healing. I think the narrow match is the right design. It ties the retry to the one error whose place in the call path is known [9].
Telling a wiped volume from a stale cache takes three checks, in this order, according to the post [13]:
1. Pull the trunk ID that the failed outbound call tried to use from the application logs [13]. 2. Resolve that trunk by name against the media server. Finding it by name under a different ID confirms the cached ID is stale [13]. 3. Inspect the Redis container from the host and check whether its data directory is mounted. An empty or non-existent mount confirms the store was wiped [13].
Remeika wrote that a fresh data directory after a recreate points to a missing volume mount and rules out an application-layer cache [14].
He describes the two changes as layers. Persistence handles the common case, he wrote, and "Resolving by name survives the case where persistence was not enough" [15]. In my view the name lookup is the change to ship first. Persistence holds only while a real data directory stays mounted through every future deploy, and the third check exists because that mount goes missing [6] [14]. A reprovisioned trunk keeps its name, so the lookup recovers whether or not Redis kept its data [13]. The post does not say which LiveKit or Redis versions it describes [17].
What to watch
- Whether LiveKit's own self-hosting templates start shipping Redis with append-only persistence and a mounted data directory by default.
- Whether LiveKit exposes a health signal for the SIP service, so an empty trunk store shows up before the next outbound call fails.