Skip to content

Build1 publisher3 min readPublished

Long prompts sent just before llama-server sleeps hang the request or segfault the server

llama-server left the request hanging or crashed outright in all 73 trials where a 17,653-token prompt arrived 3 to 46 ms before its sleep timer fired. A one-token warm-up request sent first avoided both failures in 123 of 123 tries.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying Long prompts sent just before llama-server sleeps hang the request or segfault the server
Generated illustration

What happened

  • According to the write-up, the handler checks that the server is awake, then tokenizes the prompt, then queues the task, and the idle timer can put the server to sleep between those steps.
  • Under gdb, the faulting thread was an HTTP worker inside llama_vocab::text_to_token, part-way through tokenizing the prompt.
  • A hung request stayed in the queue until a second short completion woke the server, and was then answered 4.2 s after it was sent.
  • The investigation began with a client that called four endpoints right after startup with a one-second sleep timer and got no completion reply about one time in eight.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • constraint A restart policy handles only the crash: the hung server stays up and correctly reports itself asleep, so supervisors and health checks find nothing to act on.
  • exposure Single clients that send one request after a quiet period, such as cron jobs, home automations and agent turns, have no second request coming to release a stranded one.
  • decision Keeping sleep mode on means changing each client to send a warm-up request first, or else dropping the timer, because the protection that worked in testing lives on the client side.

The outcome depends on where a send lands relative to the sleep moment. In the tests published on dev.to, the results fall into clean bands [13]. Of 123 long-prompt sends, the 13 that arrived more than about 50 ms early were processed before the timer fired [13]. The 37 that arrived between 2 ms before and 29 ms after the sleep woke the server and got an answer in about 1.3 s [13]. All 73 failures sit between those two bands [1]. A send 46 to 34 ms early hung. A send 34 to 3 ms early crashed, and on the author's account the tokenizer was reading vocabulary the sleep had just freed [4][5]. Ten more sends placed between 0.970 and 0.990 s crashed the server ten times out of ten [19].

The harness is good work and worth copying. Each trial starts a fresh server, takes the first 200 from `/health` as time zero, waits a chosen offset, sends one `POST /completion`, and timestamps every log line as it arrives [11]. The server fell asleep 0.978 to 1.000 s after zero, close enough to place each send against the moment of sleep [11]. The clever part is `-c 512`. Setting a context too small for the prompt is an odd way to test whether the server will take a prompt, and a fast one. Any request that reaches a slot comes back at once as HTTP 400, so a 400 means it got through and anything else means it did not [12].

The 43 ms failure window [2] belongs to this workload: a 17,653-token prompt through Gemma 3 1B's tokenizer on six CPU cores in a Debian 13 LXC, running the b11368 release binary with one slot [3][4]. Tokenizing is the step that sits between the awake check and the queue [2]. With the two-word prompt "Say OK", 93 sends between 0.97 and 1.03 s got 93 answers, and the 61 that arrived after the sleep woke the server as documented [6]. I'd expect the window to scale with tokenize time. It would widen with prompt length or a slower host and narrow on a faster one. The author puts it at a few dozen milliseconds out of every idle period, rare enough that the requests that hit it look like unrelated faults [18].

Both failures are quiet in the server's own output. A hung server's log ends on "entering sleeping state", the line an idle server is supposed to print, and `/props` correctly reports that it is asleep [17]. The crashed process exits with signal 11 mid-request, so its log has nothing to show, and the client sees `RemoteDisconnected` [20][17]. The documentation promises a reload for "any new incoming task" [1]. The stranded task was queued as the server went to sleep, and it did not trigger one [4].

The author's `check-llama-sleep-race.sh` finds servers with sleep mode on and reports whether a crash would be restarted [8]. That covers half the problem. The warm-up covers both halves, at the cost of one extra completion request after each quiet period [7]. The upstream issue for the hang is ggml-org/llama.cpp#29689, and the write-up does not say whether a fix has landed [9]. In my context, a single client waking the server after quiet periods, I would not run `--sleep-idle-seconds` without the warm-up in the client code [7].

What to watch

  • Whether the fix for ggml-org/llama.cpp#29689, filed for the hang, also removes the tokenizer read of freed vocabulary behind the crash.
  • Reproductions on GPU builds or non-Gemma tokenizers, to show whether the 43 ms window tracks tokenization speed outside this six-core CPU setup.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories