Build1 publisher2 min readPublished
Unbounded Tokio tasks grew a 100 MiB Rust pod to 3.7 GiB during traffic bursts
Engineers who spent three years porting services to Rust say unbounded Tokio tasks grew a 100 MiB pod to 3.7 GiB under burst load. Removing GC pauses cut their latency, and memory held steady only once they capped in-flight work at the ingest edge.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Load balancer latency fell from 600ms to 101ms after the team replaced nginx with Pingora, Cloudflare's Rust proxy framework.
- During bursts each Tokio task held buffers and state while waiting on I/O, and with no cap on in-flight work, memory grew with the arrival rate.
- Tuning the Horizontal Pod Autoscaler made the problem worse, because it gave the backlog more places to accumulate.
- Backpressure at the ingest edge, capping in-flight work and shedding excess requests, stopped memory from growing without bound during bursts.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint Any Tokio service that spawns work per request needs to pick and enforce its own in-flight limit; otherwise the pod's memory limit ends up being the cap.
- decision Adding replicas leaves each pod unbounded, so operators seeing burst-driven memory growth need admission control and a shed policy set per pod.
- exposure Shedding at the edge means callers get rejected requests during bursts, and client retry paths that resend them can rebuild the same backlog upstream.
- cost Similar gains from a port would call for budgeting redesign work such as connection pooling and data layout, along with the language change.
The growth factor is roughly 38x. 3.7 GiB is about 3,789 MiB, set against a 100 MiB baseline [1]. The team describes the failure as a chain: bursts, then unbounded tasks, then memory accumulation, then pod memory exhaustion [15]. No step in that chain depends on a memory leak.
Removing garbage collection is what made the port pay off. According to the team, Rust eliminated GC pauses and made latency more predictable [17]. On the replication layer, steady state settled at 27 to 31 microseconds. A gap of about 1.5 microseconds between nodes also became visible, where GC noise had hidden it before [10]. Tokio's scheduler gave the services high concurrency at low CPU [16]. Pods settled near 0.22 cores, against 0.4 to 0.6 before the port [5]. Because accepting work costs almost nothing in CPU, the runtime also has no natural point at which it refuses a request, and every accepted request becomes a task holding memory.
The autoscaler episode comes from the same missing limit. Scaling out adds pods, but no pod limits its own intake, so each new replica can build its own queue of waiting tasks [8].
I think the ingest edge is the right place for the cap. A request refused there never becomes a task, so it never allocates the buffers and state a waiting task would hold [7]. The write-up does not say which primitive enforces the cap or how the limit was sized. Shedding also moves cost onto callers. Separately, the team replaced its custom TCP protocol with gRPC and a store-and-forward model to improve message durability and retries [14].
The latency figures are claims about this workload, and part of each gain is architecture. The load balancer improvement, about 5.9x, came with a proxy swap [3]. Presence API peak memory fell from 3.4 GiB to a 256 to 512 MiB band after a redesign [4]. The team's own comparison of two Rust implementations puts connection pooling and data layout on par with language choice [13]. The publish API went from about 350 to about 50 microseconds, roughly 7x [3][2]. For that gain to repeat elsewhere, a service would need a tail dominated by GC pauses and a budget for comparable redesign work.
The measurement is careful. Every dashboard in the post existed before the migration, so each before-and-after figure is compared against a recorded baseline [11]. The cost account is direct too. The port from Python, Go, JVM and C ran over three years [1]. Engineers fought the borrow checker for the first six months, and productivity improved by month eight [12].
What to watch
- A follow-up giving the in-flight cap value, how it was sized, and the rejection rate during a real burst.
- Whether the gRPC store-and-forward path absorbs shed requests through retries or moves the backlog into upstream services.