Build1 publisher2 min readPublished
PyTorch folds image, video and audio decoding from TorchVision and TorchAudio into TorchCodec
PyTorch has moved all image, video and audio decoding and encoding into TorchCodec after a two-year consolidation of its media stack. Teams whose vision, audio or generative training code still reads media through TorchVision or TorchAudio now have a port to plan.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- The old decoding and encoding APIs in TorchVision and TorchAudio are now deprecated or, in some cases, already removed.
- TorchVision and TorchAudio now focus on their transforms, and their models, datasets and pipelines are no longer under active development.
- For those models and datasets, PyTorch points users to the wider ecosystem and names the HuggingFace libraries as an example.
- TorchAudio lost the most APIs, though several popular ones that had been slated for removal were kept after community feedback.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Any data loader or output writer still calling the old TorchVision or TorchAudio IO has to be ported to TorchCodec or held on a release where that call still exists.
- exposure Training recipes that import TorchVision or TorchAudio models and datasets now depend on code PyTorch is not actively developing, so bugs found there fall to the team using it.
- cost Teams have to measure the CUDA speedup themselves, because default TorchVision installs never shipped the old CUDA backend and so give them no like-for-like baseline.
- precedent TorchAudio's kept APIs show PyTorch will revise a removal list when users push back, so teams depending on a deprecated call have grounds to file feedback before it is removed.
On a default install, TorchVision decoded video through one backend: PyAV, in Python [6]. A C++ FFmpeg backend and a CUDA/NVCUVID backend also sat behind io.read_video() and io.VideoReader() [5]. Both needed a source build tied to a specific FFmpeg version [6]. TorchAudio ran its own stack, with StreamReader and StreamWriter for video and audio and further audio utilities over FFmpeg, libsoundfile and libsox [7]. That adds up to six named backends across two libraries, each library building its own [16][15].
The dependency list explains the design. Media IO now means six major FFmpeg versions, 4 through 9, plus NVIDIA's codec SDK for GPU decoding and one library per image format, each with its own licensing rules on what can ship [8]. All of them are C or C++ libraries. The post says that makes the build, the CI and the releases much harder [8]. Nobody wants to run a six-version FFmpeg test matrix twice. I think putting every native media dependency behind a single package boundary is the right design, and the post says it greatly simplifies maintenance for TorchVision and TorchAudio [9].
PyTorch says TorchCodec is generally faster than the implementations it replaced, particularly for CUDA video decoding [17]. The post does not include benchmark figures. TorchVision's old CUDA backend required a source build [6]. A team on the default install was therefore decoding with PyAV and has no old CUDA number to compare against. For the claimed gain to show up in a given training job, decode has to be the bottleneck, and the comparison has to be against the backend that job actually ran.
I'd port the IO first. Those calls are deprecated or already removed [4]. The models, datasets and pipelines are described only as out of active development [10], so they can move to ecosystem libraries on a slower schedule [12]. Generative pipelines that write video or audio back out are in the same position as data loaders, because encoding moved to TorchCodec along with decoding [2][3]. The transforms stay put. PyTorch kept them because they are the most active usage area in both libraries [11].
"A smaller TorchAudio is one that can actually be maintained, so it's one users can continue to rely on," the post says of the cuts [14].
What to watch
- A published removal schedule or per-API list for the TorchVision and TorchAudio IO calls that are deprecated but not yet removed.
- Benchmark figures, from PyTorch or users, comparing TorchCodec CUDA video decoding against PyAV on CPU.
- Whether any TorchVision or TorchAudio models or datasets move from 'not under active development' to formal deprecation.