Skip to content

Build1 publisher3 min readPublished

Live video is two protocol decisions, not one, and your CDN is fighting your latency target

A practitioner's guide argues most live streaming failures start by collapsing ingest and delivery into a single choice. The rest come from a player buffer nobody checked.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • The single most common mistake in live streaming is treating "streaming protocol" as one choice; it is two.
  • Ingest is getting video from a camera, encoder or browser into your server; delivery is getting it from your server to viewers. They have different constraints and you almost never use the same protocol for both.
  • A typical stack ingests over RTMP or SRT and delivers over HLS; another ingests WebRTC and delivers WebRTC. Mixing is normal and expected.
  • RTMP is old, TCP-based and still everywhere: every encoder speaks it, OBS defaults to it, and latency is typically 2 to 5 seconds.
  • Classic RTMP is limited to H.264 and AAC, though the Enhanced RTMP spec has added HEVC and AV1.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A practical guide to live streaming published on dev.to opens with a diagnosis worth taking seriously: the single most common mistake is treating "streaming protocol" as one choice when it is two [1]. That matters because ingest and delivery have opposing constraints, and because the property that makes HTTP delivery cheap to scale is the same property that makes it slow [2][11].

Ingest is getting video from a camera, encoder or browser into your server. Delivery is getting it from your server to viewers. According to the guide, they have different constraints and you almost never use the same protocol for both [2]. A typical stack ingests over RTMP or SRT and delivers over HLS; another ingests WebRTC and delivers WebRTC; mixing is normal and expected [3].

On the ingest side the choice is mostly about how bad your uplink is. RTMP is old and TCP-based and still everywhere, every encoder speaks it, OBS defaults to it, and latency sits at 2 to 5 seconds [4]. Classic RTMP is limited to H.264 and AAC, though the Enhanced RTMP spec has added HEVC and AV1 [5]. Because it is TCP, packet loss becomes head-of-line blocking and the stream stalls rather than gracefully dropping quality [6]. SRT answers that with UDP, its own ARQ retransmission layer, a configurable latency buffer and built-in AES encryption, aimed squarely at pushing broadcast-quality video over the public internet [7]. If the source is on a flaky link, 4G, or another continent, the guide calls SRT the usual right answer [8]. Publishing it is one ffmpeg line to an srt:// URL with a streamid [22]. If you are pulling from surveillance hardware you are pulling RTSP whether you like it or not [9], and WHIP standardises WebRTC signalling over plain HTTP for browser publishing without a plugin [10].

Delivery is where the real conflict lives. HLS and DASH chop the stream into segments served over plain HTTP, which is precisely why any CDN can cache them and why they scale to millions of viewers cheaply, at a cost of 10 to 30 seconds of latency [11]. LL-HLS and LL-DASH claw that back to 2 to 6 seconds using partial segments, blocking playlist reloads and preload hints, while keeping CDN compatibility [12]. WebRTC, with WHEP for standardised playback, gets under half a second and is the only real option for auctions, betting, two-way audio or remote control, but it is not natively cacheable by traditional CDNs, so scaling works differently [13]. Best-case low-latency HLS is still around four times the WebRTC ceiling [2].

Before changing protocol, audit the budget. End-to-end latency is a sum of encoder buffer, network transit, server processing, segment duration, playlist window, CDN propagation and player buffer [15]. Apple's classic recommendation was 6-second segments and players typically buffer three of them, which is 18 seconds before a single frame is transcoded or a hop crossed [16]. That alone is 60 percent of the 30-second worst case [1]. The player is the part teams miss: it may be holding 15 seconds because that is its default [18]. So when someone reports 30-second HLS latency and asks you to fix the server, the server is probably not the problem, and shorter segments plus a smaller playlist window plus a tuned player get most of the way [19]. Below roughly 2 seconds, tuning stops working, because segmented delivery has a floor [20].

The operating rule from the guide is to pick the highest latency you can actually tolerate, since every step down costs money, CDN compatibility, or both [21].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories