Skip to content

Build1 publisher3 min readPublished

Timing the gap before a WebSocket 1006 drop separates ping failures from proxy timeouts in LLM streams

WebSocket drops in FastAPI LLM streams that land at 20 to 40 seconds point to ping settings, and those near 60 to a proxy, a dev.to debugging guide says. It says to time the gap between the last traffic and the drop before changing any timeout, since a 1006 code alone does not say which layer failed.

The Engineer · Build desk

Illustration accompanying Timing the gap before a WebSocket 1006 drop separates ping failures from proxy timeouts in LLM streams

What happened

  • RFC 6455 reserves close code 1006 for abnormal closures in which no proper WebSocket Close frame was received.
  • FastAPI has no built-in 10-second timeout on LLM generation, contrary to a widely repeated explanation that it kills slow WebSockets.
  • A synchronous model call inside an async handler freezes the whole event loop, including ping handling for every other WebSocket the process serves.
  • The commonly proposed fix, a loop sending {"type": "ping"} as JSON every 5 seconds, is application data and not a WebSocket protocol Ping control frame.
  • When pings fail on a responsive process, the websockets library logs ConnectionClosedError with a 1011 keepalive ping timeout and no close frame received.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • decision Before anyone raises --ws-ping-timeout, the team has to show the event loop is healthy; if synchronous model calls are blocking it, the change belongs in handler code.
  • constraint The timing bands only apply where Uvicorn and NGINX defaults are unchanged, so a team that has already tuned ping or proxy settings must work out its own bands first.
  • cost Teams running local inference inside the WebSocket process have to add a process pool or a separate compute service, because to_thread cannot get CPU-bound work past the GIL.
  • contradiction Behind NGINX, the post's own timeout definition suggests the JSON heartbeat does reset the idle timer, so whether it helps depends on which intermediary sits in the path.

The 1006 in the browser console never crossed the wire. RFC 6455 forbids an endpoint from putting 1006 into a Close control frame [2], so the client assigns the code itself once the connection is gone without a close handshake [1]. The post draws the path a prompt travels as browser, mobile or voice client, then CDN or edge, then load balancer or WAF [3]. Any of those hops, or the FastAPI process behind them, can end the connection that way [4].

Both fixed timing bands in the post's table follow from defaults. Uvicorn's WebSocket ping interval and ping timeout both default to 20 seconds [9]. Together they make 40 seconds, the top of the band the post assigns to ping and liveness config [2]. NGINX's proxy_read_timeout, the time it waits between reads from upstream, defaults to 60 seconds [20]. That is the same figure as the band the post assigns to proxy and load-balancer idle policy [3]. A stack that has tuned either setting moves its band with it [3].

The remaining rows need no defaults to read. Wide variation in drop time points at the network, client state or process health [6]. A drop that happens only during synchronous model calls points at a blocked event loop. One that appears only under high concurrency points at event-loop lag, CPU saturation or backpressure [6]. If the server logs a clean close while the client reports 1006, the post says to trace the intermediary or the network [6].

Long latency does not by itself starve the event loop. A model can spend 30 seconds before its first output without freezing anything, provided the call is asynchronous [14]. For blocking I/O, the post uses asyncio.to_thread(), which Python documents for that case [15]. CPU-bound work such as local inference or expensive parsing still contends for CPython's GIL under to_thread, so the post moves it to a process pool or a dedicated compute service [16]. "Your WebSocket process should not quietly become your compute scheduler," the post says [17].

The config table has a column for when to change each setting [21]. In my view that column belongs in the runbook for any service like this. It says to raise `--ws-ping-timeout` only after the event loop is confirmed healthy. It also says to leave `--timeout-keep-alive` alone for 1006 issues, because its 5-second value governs HTTP keep-alive and not WebSockets [19].

For NGINX, one of the post's objections to the 5-second JSON heartbeat conflicts with its own table. The second objection holds: a blocked loop cannot run the heartbeat either [11]. The third says the heartbeat does not address proxy or load-balancer idle timeouts, which the post describes as independent of application traffic [12]. By the table's definition of proxy_read_timeout, a server message every 5 seconds is a read from upstream, well inside the 60-second window [4]. The post does not name the intermediary that would ignore that traffic.

What to watch

  • The idle timeout on any managed load balancer or CDN in the path: if it is not 60 seconds, the proxy band moves to match it.
  • Whether the post's author names an intermediary that drops connections despite server-sent application messages, which would settle the heartbeat objection.
  • Uvicorn release notes for any change to the 20-second ping interval or timeout defaults, which would move the 20-to-40-second band.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories