Build1 publisher3 min readPublished
The merge gate that turns "works on bad networks" into something CI can fail
A robotics fleet project made a five-profile shaped-network matrix a required gate on main. The asymmetric-uplink profile is the one almost nobody runs, and the one teleop dies on.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Ganglion exists to reach robots on networks nobody controls, including warehouse Wi-Fi, carrier CGNAT, a hospital VLAN, and a customer firewall configured once in 2019.
- Until this week the field-reliability claim was a sentence on a website; CI ran on clean loopback, everything was green, and the failure modes that matter in the field were the ones the test suite could never produce.
- The author writes: 'I build Ganglion, so treat the enthusiasm accordingly.'
- Every push to main now runs the full deploy, invoke and verify round trip over the relay against five shaped network profiles, and all five must pass before anything merges.
- The five profiles are: clean (baseline, no shaping); lossy (packet loss with light reordering); high-latency (250ms round trip); asymmetric; and nat-relay.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Ganglion, a substrate built to reach robots on networks nobody controls, now blocks every push to main unless a full deploy, invoke and verify round trip over the relay passes against five shaped network profiles [1][4]. According to the project's maintainer, who wrote the account and says to treat his enthusiasm accordingly, that reliability claim had until this week been a sentence on a website while CI ran on clean loopback and everything was green [2][3].
The five profiles are clean, lossy with light reordering, high-latency at 250ms round trip, asymmetric, and nat-relay [5]. Asymmetric means plentiful downlink and starved uplink, which the maintainer calls the profile nobody tests and the one teleop actually dies on, because control acknowledgements travel the starved direction [6]. The nat-relay profile removes any route between endpoints, forcing hole punching to fail so relay fallback has to carry the session [7]. His framing is worth borrowing: loss and latency are what people imagine a bad network is, while asymmetry and no direct route are what a bad network usually is [8].
The interesting part is the wrong turn. The original design was two lossy profiles, one with a pinned netem seed to gate the build and a nastier randomized one running nightly and allowed to fail [9]. That does not work, because netem's loss and jitter draw from the kernel RNG and there is no seed parameter [10]. A gate that fails randomly does not catch regressions; it teaches everyone to re-run the job until it passes, and after a month nobody reads red as meaning anything [11].
So the gate was rebuilt out of only the mechanisms that reproduce exactly: fixed netem delay with zero jitter, tbf rate caps, which are a token bucket rather than a distribution, and route blocking, where either a route exists or it does not [12]. Loss comes from `iptables -m statistic --mode nth`, which drops precisely every Nth packet instead of N percent on average, so every 33rd packet is roughly three percent loss and the same three percent every run [13]. That is 3.03 percent by construction [14]. Randomized netem still runs nightly in a separate non-blocking job, where a failure opens an issue rather than stopping a merge [15]. Its parameters come from a recorded seed, but replay reproduces the impairment distribution, not the packet-level draw, which is documented in the README and is the stated reason chaos never blocks a merge [16].
The rig choice generalizes further than the tooling does. The obvious build is a fresh veth pair or dedicated namespaces; instead the matrix rides the existing end-to-end dispatch harness, because a second rig would have produced a network test that never touches the real product path [17]. Shaping is applied inside the robot and operator containers before the agent starts, and netem inside Docker had already been proven green by an existing mobile-CGNAT scenario [18]. Every run writes a JSON artifact with mode, seed, the exact shaping commands issued, duration and result, a practice the maintainer credits to a ROS Discourse conversation with someone building replay tooling for robot fleets, who argued that injected faults which are not recorded are the problem [19][20].
What to watch: whether the nightly chaos issues get closed or quietly accumulate, and whether the asymmetric profile ever fails on its own. If it does, it will be catching something the loss and latency profiles cannot see [6].