Build1 publisher2 min readPublished
LiteLLM's standalone Rust gateway adds 0.7 ms at p99 with logging and spend tracking switched off
LiteLLM's standalone Rust path added 0.7 ms of p99 latency against 257.7 ms for its Python proxy in July 22 benchmark artifacts reviewed on dev.to. Logging, persistence and spend tracking were off for the run, so the gap measures forwarding alone.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Against a deterministic local mock, the standalone Rust path peaked at 21.8 MB of memory and sustained roughly 2,814 requests per second.
- In the same p99 chart, Portkey OSS added 2.3 ms of latency and Bifrost 1.6.4 added 4.5 ms.
- In LiteLLM's hybrid mode, Python still handles HTTP, authentication, routing and callbacks, and only supported provider translation and network operations go through the Rust core.
- An earlier migration harness, on a different traffic shape and implementation generation, measured about 0.05 ms of Rust overhead against 7.5 ms for Python.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint Teams keeping budgets and callbacks in Python would run hybrid mode, while the 0.7 ms figure was measured on the standalone path.
- cost Before moving traffic to standalone, a team has to confirm the Rust surface covers its budgets, callbacks and persistence, and the author treats that coverage as the unanswered migration question.
- contradiction The Rust-over-Python latency ratio was 150x on the earlier harness's overhead figures and about 368x at p99 in July, so the speedup a team sees will depend on its own traffic shape.
- decision Choosing among LiteLLM Rust, Portkey and Bifrost from this chart is a choice on forwarding speed alone, because the test was not a full-feature comparison.
The July run forwarded an Anthropic Messages body to a local Rust mock, and logging callbacks, persistence and spend tracking were disabled on purpose [5]. A deterministic local mock keeps provider queueing, internet variance, rate limits and model generation time out of the measurement [16]. The author wrote that they "treated it as a runtime-isolation benchmark, not a production architecture benchmark" [6]. The proposed follow-up config strips the same layer from the Python proxy: `callbacks: []` and `num_retries: 0` under `litellm_settings` [17].
Standalone mode is an Axum server that handles routing and network I/O entirely in Rust [9]. For the 0.7 ms [1] to carry over, production traffic has to take that path. Whatever callbacks, persistence and spend tracking the deployment needs then has to happen somewhere else. Upstream calls also have to be short enough for gateway overhead to register. The author's motivating traffic meets that last condition: embeddings, classifiers, guardrail calls and short agent turns at high concurrency [14]. On a 12-second reasoning request, the Python proxy's 257.7 ms p99 [2] is about 2.1% of wall time [1].
Memory is the steadier result. The Python v1 proxy peaked at 329.5 MB in July [2], about 15.1 times the Rust path [3]. In the earlier harness the ratio was about 11.3 [4]. The author declined to merge the two sets because the workloads and harnesses were not identical [12]. I'd have kept them apart too. The earlier latency figures are also reported without a percentile [11].
In the post's one executed experiment, the benchmark's response-presence guard was tested against Requests response fixtures, with no network calls and no latency measured [13]. Few posts with "Benchmarked" in the headline are that plain about what they ran. The proposed follow-up is well designed: a deterministic OpenAI-compatible mock, a containerized LiteLLM Python proxy and a load generator with a controlled arrival rate [15]. A fixed arrival rate stops the load generator from backing off when the gateway slows, so queueing shows up in the p99.
As the write-up frames it, the open question is whether the current Rust surface "covers enough of an actual gateway workload to justify migration" [18]. I think the standalone path is ready to trial for short calls that sit outside any budget. I'd keep budgeted traffic on the Python proxy, since that is where budgets, callbacks, spend tracking and persistence are handled [7].
What to watch
- A hybrid-mode p99 figure measured with callbacks, persistence and spend tracking enabled.
- Whether the author's follow-up run, with a controlled arrival rate and a containerized Python proxy, reproduces a gap of the July size.
- Documentation stating which of the Python proxy's budget, callback and persistence features standalone mode supports.