Skip to content

Build1 publisher3 min readPublished

One webrtc.NewAPI() per SDP offer is a load-bearing bug, and thirty viewers will find it

A dev.to post on Pion in Go traces a familiar production failure to the tutorial pattern itself: per-request engines, per-connection ECDSA keys, and a UDP port for every viewer.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

Illustration accompanying One webrtc.NewAPI() per SDP offer is a load-bearing bug, and thirty viewers will find it
Generated illustration

What happened

  • The author describes a classic side effect of the default Pion WebRTC setup in Go: read the usual Pion usage guides, deploy a WHEP handler, open several browser tabs with the player, and the server starts choking.
  • Symptoms include handshakes taking seconds, ICE timing out, and half of the pprof flame graph filled with crypto/elliptic.p256OrdSqr and map allocations inside the engine.
  • The pattern in most Go WebRTC tutorials: for every POST request containing an SDP offer, create a new webrtc.NewAPI(), register the default codecs, call api.NewPeerConnection(), and return the response.
  • On localhost with a couple of clients the per-request approach works fine; in production it turns into a disaster.
  • The author writes that many Go WebRTC tutorials give the impression of being written by people who never load-tested their implementation with even fifty concurrent viewers, at most a couple of streams.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A post on dev.to describes a failure mode that anyone who has shipped Pion will recognise: follow the usual Go guides, deploy a WHEP handler, open a handful of player tabs, and the server starts choking [1]. Handshakes stretch to seconds, ICE times out, and half the pprof flame graph is `crypto/elliptic.p256OrdSqr` plus map allocations inside the engine [2].

The pattern under all of it is the one printed in most tutorials: for every POST carrying an SDP offer, construct a new `webrtc.NewAPI()`, register the default codecs, call `api.NewPeerConnection()`, return the answer [3]. On localhost with two clients it flies; in production it does not [4]. The author's read is blunt: many of these guides look like they were never load-tested with even fifty concurrent viewers [5]. The diagnosis is not that WebRTC, Pion or Go is slow, but that expensive global infrastructure is being built per request instead of reused [6].

Three costs are being paid on the request path, and none of them belong there. First, DTLS needs a certificate, which means generating ECDSA P-256 keys through `crypto/ecdsa`, elliptic-curve arithmetic and `crypto/rand` [7]. Second, the engine allocates hundreds of small heap structures for every possible codec, RTCP interceptor, RTP header, header extension and internal component [8]. Third, the stack goes to the OS and opens random UDP ports for ICE candidate gathering and binding, so a typical configuration takes new network resources for each `PeerConnection` [9]. Thirty viewers arriving at once means CPU pinned by key generation, a ballooning connection tracking table, and a garbage collector working through the wreckage of every closed session [10].

The fix described is structural rather than a tuning knob. In the author's own system, RUSEON Core, the heavy work moved into an isolated singleton engine in `internal/webrtc/engine.go` [11]. The certificate is generated once at startup via `ecdsa.GenerateKey(elliptic.P256(), rand.Reader)` [12]; the argument for this is that the browser does not care how old the certificate is, only that the SDP carries a cryptographically valid fingerprint [13]. The codec registry is likewise built at startup and stops touching the heap on every connection [14].

The network change is the one with the most operational consequence. Instead of dozens of random ports, `webrtc.NewICEUDPMux` is attached to a single UDP socket, so one port serves the whole system [15]. The kernel multiplexes incoming STUN, DTLS and SRTP through one file descriptor and Pion sorts packets between ICE and WebRTC sessions in userspace [16]. The claimed result is no port exhaustion and no need to open a 50000-60000 range in the firewall [17], which is a reduction from 10,001 permitted UDP ports to one [19]. That is a firewall rule you can actually defend in a review.

Two caveats about the material. The supplied text carries no before-and-after measurements: no handshake latencies, no CPU figures, no viewer count at which the pooled design was retested [21], so the mechanism is convincing while the magnitude is unverified. And the piece describes both a singleton engine [11] and an engine pool built on `sync.Pool` [18], which are different strategies with different lifetimes; the excerpt is cut off mid-declaration of `WebRTCEnginePool`, so how the two fit together is not shown [20].

What to watch is whether your own handler does the tutorial thing. Grep for `NewAPI()` inside your HTTP path; if it is there, your capacity ceiling is set by ECDSA key generation and socket churn, not by egress.

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories