Skip to content

BuildNot yet confirmed elsewhere1 publisher3 min readPublished

Two regions, 86ms apart, 28 TPS: the WAN was not the bottleneck, Python was

A 7B model split across Iowa and Oregon on free T4s went from 4.92 to 28.10 tokens per second. Most of the gain came from a drafter that stopped launching kernels one at a time.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Two regions, 86ms apart, 28 TPS: the WAN was not the bottleneck, Python was
Generated illustration

What happened

  • Qwen2.5-7B was split layer-wise across two Kaggle T4s, Iowa and Oregon, talking through a TCP relay on an EC2 t3.micro in Ohio.
  • With no optimisation, that topology produced 4.92 tokens per second.
  • Speculative decoding with a 0.5B drafter proposing 8 candidates committed an average of 4.07 tokens per round trip instead of one.
  • Draft generation still cost 112ms per round, with the GPU idle 65% of the time while Python issued kernel launches.
  • Capturing the drafter's forward pass as a CUDA Graph took peak throughput to 28.10 TPS.

Why it matters

  • capability Once a round trip carries four accepted tokens, its cost falls to about 21ms per token, which is what makes splitting a model across regions on commodity internet worth attempting at all.
  • constraint Further drafter tuning cannot get past roughly 47 TPS on this path; only a better acceptance rate or fewer hops will, so optimisation effort has to move upstream.
  • decision Anyone pricing a private interconnect for distributed inference now has to weigh it against auditing their own launch overhead first, because here the Python loop cost more than the wire.
  • contradiction The reported gains do not fully reconcile: the draft fix accounts for 87ms of a 140ms round-time improvement, so either the peak figures come from different windows or something else changed...

Run the arithmetic on the round rather than the token and the numbers get more interesting than the headline. At 28.10 TPS with 4.07 tokens committed per accepted round trip [1][8], a round lasts about 145ms [24]. Eighty-six of those milliseconds are wire [6] and 25 are the draft model [14], which leaves roughly 34ms for the 7B verification slice on Node 1 and every other cost in the loop [24]. Do the same sum for the eager-mode configuration the author reports at 14.3 TPS [9] and the round is 285ms, of which 86ms is wire and 112ms is drafting [12], leaving 87ms of residual [21].

Those two residuals do not match. The graph capture removed 87ms of drafting per round, but round time fell by 140ms [25]. Fifty-three milliseconds came off something the writeup does not itemise [25]. Part of that is probably the mixing of peak throughput figures with an average acceptance rate, which is the weakest joint in the measurement, and it is the first thing worth pinning down before anyone treats 28.10 as reproducible.

The mechanism behind the drafter fix does check out, though, and it checks out almost exactly. Eight candidate tokens meant eight forward passes of roughly 1,500 kernels each [10], so about 12,000 launches per round; at the 8 to 10 microseconds Python needs to issue each one [11], that is 96 to 120ms of pure launch overhead [23]. The reported 112ms draft phase [12] sits inside that bracket, which means the draft phase was essentially all bookkeeping and none of it compute. A 65% GPU idle rate [12] is the same fact seen from the other side.

Which reframes what the network was ever costing. At the 4.92 TPS baseline [5], a token takes 203ms and the 86ms round trip is 42% of it [20]. Amortised over 4.07 accepted tokens, the same round trip costs 21ms per token [19]. After that, the Python launch overhead in the drafter was larger than the entire Iowa-to-Oregon-to-Ohio path [23][6]. The interesting engineering conclusion is not that the WAN turned out to be survivable; it is that a 2,000km hop and a t3.micro relay [4] were quietly cheaper than a for-loop in the same process.

The ceiling that remains is arithmetic, not effort. At 86ms per round trip and 4.07 tokens accepted, the network alone caps this topology at 47.3 TPS [22], and 28.10 is 59% of that [22]. Nothing done to the drafter moves that number. Raising it means a higher acceptance rate, a larger draft window, or fewer hops between the two shards.

The cost of the trick is a memory contract that transformers does not honour by default. CUDA graphs capture fixed addresses, and HuggingFace's DynamicCache calls torch.cat at every token step, allocating a fresh buffer that the graph never learned about [15][16]. The failure signature the author describes is repeated output, not a crash [15], which is the expensive kind: it looks like a sampling problem. Holding it together took StaticCache, in-place tensor mutation, manual position_ids updates, and an in-place KV rewind for rejected drafts [17]. That is four places where a library upgrade can silently reintroduce a stale pointer.

What to watch

  • An acceptance-rate distribution rather than a 4.07 mean: if the tail is long, both the per-round arithmetic and the 47.3 TPS ceiling move.
  • The same rig with both shards in one region, which would isolate how much of the 145ms round is really the WAN.
  • Whether the StaticCache and position_ids plumbing survives a transformers version bump, or whether the graph starts reading stale addresses again.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence42
Adoption16
Hype gap+22
Incentives58
Confidence38
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    ShardFlow v2.1 reaches 28.10 TPS peak on Qwen2.5-7B across two separate cloud regions over WAN.

  2. [2]

    The project runs on free Kaggle T4 notebooks, an AWS EC2 t3.micro relay, and the public internet, with no A100s and no private datacenter network.

    ReportedSupportedView cited source
  3. [3]

    A 7B parameter model in FP16 needs roughly 15 GB of VRAM; a single Kaggle T4 has 16 GB, leaving nothing for a KV cache.

    ReportedSupportedView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. dev.to

    1 article · August 23, 2026

    I Got 28 TPS Out of Free Kaggle GPUs. Here's What It Took.

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

Entities

Loading related stories