Skip to content

Build1 publisher3 min readPublished

Four video re-tunes later, your audio ladder is still one 128 kbps AAC-LC track

A dev.to walkthrough argues the master playlist, not the encoder config, is the only reliable record of what you ship. The arithmetic on the bottom rung is the part worth reading.

The Engineer · Build desk

Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened

  • The article advises starting an audio audit with the master playlist rather than the encoding config, because "these disagree more often than you'd think".
  • The article's framing: the encoding profile has been re-tuned four times on the video side, while the audio side is one 128 kbps AAC-LC track that nobody has touched since setup.
  • In the example master playlist, all three EXT-X-STREAM-INF lines declare CODECS containing mp4a.40.2, which is AAC-LC, meaning the same audio in all three variants.
  • The example master playlist contains no EXT-X-MEDIA block, so audio is muxed into each video rendition and there is no separate audio rendition to switch.
  • ffprobe against the 360p child playlist returns codec_name aac, profile LC, sample_rate 48000, 2 channels, bit_rate 128000.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

A walkthrough published on dev.to makes a narrow and checkable argument: start an audio audit with the master playlist rather than the encoding profile, because the two disagree more often than teams expect [1]. It matters because the same piece describes the common state of play, in which the video side has been re-tuned four times while the audio side remains one 128 kbps AAC-LC track that nobody has touched since setup [2].

The evidence is in the manifest. In the example master playlist, all three variants declare `mp4a.40.2`, which is AAC-LC, so every rendition carries identical audio [3]. There is no `EXT-X-MEDIA` block, meaning audio is muxed into each video rendition and there is no separate audio rendition for the player to switch [4]. An `ffprobe` against the 360p child playlist confirms the bytes rather than the promise: AAC, LC profile, 48 kHz, two channels, 128 kbps [5]. The author's caution is that manifests lie, and that what a pipeline intends to produce and what a packager emits are separate facts, so probe a live playback URL and not a mezzanine [6].

Then the number. The 360p variant declares 628 kbps total, of which 128 kbps is audio [7], which is 20.4 percent of the rung serving the worst-connected viewers [8]. The same 128 kbps sits inside the 1080p variant declared at 3128 kbps [9], where it is 4.1 percent of the bill [10]. Identical audio quality, wildly different share of a constrained pipe [7].

The fix in the article needs no new codec. Encode video with `-an`, encode audio separately at 128k and 64k, package to fMP4 HLS with Bento4's `mp4hls`, and declare two audio groups so the 360p variant pairs with the 64 kbps rendition [11][12]. The declared bandwidth on that rung drops from 628 to 464 kbps [13]. Note that this is a 164 kbps fall while the audio change accounts for only 64 kbps, so roughly 100 kbps of the improvement in the worked example comes from somewhere other than audio [14]. Do not quote the full delta to your finance team as an audio win. The common failure mode is pointing every `EXT-X-STREAM-INF` at the same `AUDIO` group and then wondering why audio never changes; the group ID is the switch, one group per quality tier [15].

Declarations are load-bearing. The `CODECS` attribute is a promise to the player, and if it is wrong some clients refuse to play before they ever fetch a segment [16]. Apple's `mediastreamvalidator` catches mismatches, and as of the 2026 release it runs on macOS, RHEL 9.5, Ubuntu 24.04.3 LTS and Debian 13.2 with full feature parity, so that check can now live in Linux CI [17]. The tooling floor for all of this is ffmpeg 7.x or 8.x, plus bento4 if you want it to author manifests [18].

On xHE-AAC, the article defines it as MPEG-D USAC, the Extended HE-AAC profile, plus MPEG-D loudness and dynamic range control [19], and states the asymmetry plainly: decoding is universal, encoding needs a license [20]. The published text stops before listing the properties that make it attractive for a low rung, so treat the codec case as unfinished here.

What to watch: run the probe against live playback URLs and record what fraction of your bottom rung is audio [6][8]; verify that group IDs actually differ per tier [15]; get `mediastreamvalidator` into CI now that Linux parity exists [17]. And before any xHE-AAC roadmap item, get the encoder licensing question answered, because decode reach is not the constraint [20].

Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories