Build1 publisher3 min readPublished
One webrtc.NewAPI() per SDP offer is a load-bearing bug, and thirty viewers will find it
A dev.to post on Pion in Go traces a familiar production failure to the tutorial pattern itself: per-request engines, per-connection ECDSA keys, and a UDP port for every viewer.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- The author describes a classic side effect of the default Pion WebRTC setup in Go: read the usual Pion usage guides, deploy a WHEP handler, open several browser tabs with the player, and the server starts choking.
- Symptoms include handshakes taking seconds, ICE timing out, and half of the pprof flame graph filled with crypto/elliptic.p256OrdSqr and map allocations inside the engine.
- The pattern in most Go WebRTC tutorials: for every POST request containing an SDP offer, create a new webrtc.NewAPI(), register the default codecs, call api.NewPeerConnection(), and return the response.
- On localhost with a couple of clients the per-request approach works fine; in production it turns into a disaster.
- The author writes that many Go WebRTC tutorials give the impression of being written by people who never load-tested their implementation with even fifty concurrent viewers, at most a couple of streams.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A post on dev.to describes a failure mode that anyone who has shipped Pion will recognise: follow the usual Go guides, deploy a WHEP handler, open a handful of player tabs, and the server starts choking [1]. Handshakes stretch to seconds, ICE times out, and half the pprof flame graph is `crypto/elliptic.p256OrdSqr` plus map allocations inside the engine [2].
The pattern under all of it is the one printed in most tutorials: for every POST carrying an SDP offer, construct a new `webrtc.NewAPI()`, register the default codecs, call `api.NewPeerConnection()`, return the answer [3]. On localhost with two clients it flies; in production it does not [4]. The author's read is blunt: many of these guides look like they were never load-tested with even fifty concurrent viewers [5]. The diagnosis is not that WebRTC, Pion or Go is slow, but that expensive global infrastructure is being built per request instead of reused [6].
Three costs are being paid on the request path, and none of them belong there. First, DTLS needs a certificate, which means generating ECDSA P-256 keys through `crypto/ecdsa`, elliptic-curve arithmetic and `crypto/rand` [7]. Second, the engine allocates hundreds of small heap structures for every possible codec, RTCP interceptor, RTP header, header extension and internal component [8]. Third, the stack goes to the OS and opens random UDP ports for ICE candidate gathering and binding, so a typical configuration takes new network resources for each `PeerConnection` [9]. Thirty viewers arriving at once means CPU pinned by key generation, a ballooning connection tracking table, and a garbage collector working through the wreckage of every closed session [10].
The fix described is structural rather than a tuning knob. In the author's own system, RUSEON Core, the heavy work moved into an isolated singleton engine in `internal/webrtc/engine.go` [11]. The certificate is generated once at startup via `ecdsa.GenerateKey(elliptic.P256(), rand.Reader)` [12]; the argument for this is that the browser does not care how old the certificate is, only that the SDP carries a cryptographically valid fingerprint [13]. The codec registry is likewise built at startup and stops touching the heap on every connection [14].
The network change is the one with the most operational consequence. Instead of dozens of random ports, `webrtc.NewICEUDPMux` is attached to a single UDP socket, so one port serves the whole system [15]. The kernel multiplexes incoming STUN, DTLS and SRTP through one file descriptor and Pion sorts packets between ICE and WebRTC sessions in userspace [16]. The claimed result is no port exhaustion and no need to open a 50000-60000 range in the firewall [17], which is a reduction from 10,001 permitted UDP ports to one [19]. That is a firewall rule you can actually defend in a review.
Two caveats about the material. The supplied text carries no before-and-after measurements: no handshake latencies, no CPU figures, no viewer count at which the pooled design was retested [21], so the mechanism is convincing while the magnitude is unverified. And the piece describes both a singleton engine [11] and an engine pool built on `sync.Pool` [18], which are different strategies with different lifetimes; the excerpt is cut off mid-declaration of `WebRTCEnginePool`, so how the two fit together is not shown [20].
What to watch is whether your own handler does the tutorial thing. Grep for `NewAPI()` inside your HTTP path; if it is there, your capacity ceiling is set by ECDSA key generation and socket churn, not by egress.