What Is Video Transcoding? How Does Transcoding Work?

Video transcoding is the process of converting an already-compressed video file into one or more different versions, changing attributes such as codec, resolution, bitrate, frame rate, or container so the content plays correctly on different devices, players, and network conditions. It’s the reason a single 4K master can reach an iPhone on 5G, a budget Android phone on 3G, and a living room smart TV without three separate video shoots.

This guide covers what actually changes during a transcode, the real demux-to-mux pipeline (not the simplified “decode then encode” version most articles describe), where transcoding stops and packaging or CDN delivery begins, and how to choose a transcoding strategy that doesn’t blow your infrastructure budget. It also corrects a few claims that circulate a lot in OTT content, including the idea that live transcoding runs at sub-second latency. It doesn’t, and we’ll show you the real numbers.

What Is Video Transcoding?

Video transcoding takes a video file that’s already encoded in one format and converts it into a different format, or into several formats at once. Unlike the first encode of a raw camera file, transcoding starts from compressed input and produces compressed output, which means every transcode is a decode-then-re-encode cycle.

Five things can change during a transcode, and a given job might touch one of them or all five:

  • Codec: The compression scheme used to shrink the video and audio data. A file might move from H.264 to HEVC, or from an older AAC profile to a more efficient one.
  • Resolution: The pixel dimensions of the picture. A 3840×2160 master gets scaled down to 1920×1080, 1280×720, and smaller renditions for an adaptive bitrate ladder.
  • Bitrate: How much data is transmitted per second of playback. Lower bitrates mean smaller files and less bandwidth, at some cost to detail.
  • Frame rate: How many frames play per second. A 60fps sports feed might be transcoded down to 30fps for a lower rendition to save bits without sacrificing resolution.
  • Container: The wrapper format holding the compressed streams together, such as MP4, fragmented MP4 (CMAF), or MPEG-TS.
  • Audio settings: Sample rate, channel layout (stereo vs. 5.1), and audio codec, most commonly AAC for OTT delivery.

Not every job changes all six. A job that only changes the container without touching codec, resolution, or bitrate isn’t transcoding at all. It’s transmuxing, and the difference matters enough that it gets its own section below.

what is video transcoding

The Real Video Transcoding Workflow: Demux, Decode, Process, Encode, Mux

Most explainers reduce transcoding to two steps: decode, then encode. That skips the two steps that actually determine whether a pipeline is fast, cheap, and accurate. Amazon Web Services documents the full five-stage transcoding workflow this way, and it’s the version worth using when you’re briefing engineers or evaluating a vendor.

  1. Demux. The container is opened, and its component streams (video, audio, subtitles, closed captions) are separated so each can be handled on its own. A file that’s never demuxed can’t be selectively transcoded, since the system has no way to isolate the video track from the audio track.
  2. Decode. Each compressed stream is decompressed back into raw, uncompressed frames (video) or samples (audio). This is the computationally heavy step, especially for 4K or HDR source material, and it’s the step GPU-based transcoders accelerate the most.
  3. Process. The uncompressed frames go through scaling, deinterlacing, frame rate conversion, color correction, or other filtering before they’re re-encoded. This step is easy to skip in a diagram, but it’s where most visible quality decisions actually get made.
  4. Encode. The processed frames are compressed again, this time into the target codec at the target bitrate and resolution. Every encode is a fresh round of lossy compression, which is the core reason transcoding isn’t free from a quality standpoint.
  5. Mux. The newly encoded video, audio, and any other streams are packed back into an output container, with new metadata written as needed.

For OTT delivery, this five-step process typically repeats once per rendition. A 4K source destined for a six-rung bitrate ladder runs through decode, process, and encode six separate times (demux and the final container step happen once per source and once per output, respectively), which is exactly why transcoding costs scale with the number of renditions you produce, not just the length of the source file.

5-steps transcoding pipeline

Where Transcoding Stops: Packaging, Storage, and CDN Delivery Are Separate Stages

A common inaccuracy in OTT explainers is folding packaging and CDN delivery into “the transcoding process.” They’re related, but they’re distinct stages with different tools, different costs, and different failure modes.

  • Transcoding produces the renditions: the actual re-encoded video and audio files at each resolution and bitrate.
  • Packaging takes those renditions and segments them into chunks (commonly 2 to 6 seconds each), then writes manifest files such as HLS’s .m3u8 or DASH’s .mpd that tell the player which segments exist and how to request them. Packaging can also apply DRM encryption at this stage. This is a repackaging operation, not a re-encode, so it’s cheap compared to transcoding itself.
  • Storage holds the packaged segments and manifests, whether in object storage like S3 or on an origin server.
  • CDN delivery caches those segments at edge locations close to viewers and serves them on request, which is what actually determines how fast playback starts and how well the stream survives a traffic spike.

Treating these as one blurry step is where a lot of cost estimates go wrong. A platform can have a perfectly efficient transcoding pipeline and still bleed money on an oversized CDN footprint, or the reverse: cheap delivery undone by an inefficient rendition ladder. Budget and troubleshoot each stage separately.

Explore more: 

Transcoding vs. Encoding

EncodingTranscoding
Starting pointRaw, uncompressed videoAlready-compressed video
What happensCompress raw footage into a codec for the first timeDecode existing compressed video, then re-encode it into a different codec, bitrate, or resolution
When it happensOnce, at capture or masteringRepeatedly, once per output rendition
Typical triggerCamera export, editing software renderNew device support, new bitrate ladder, new platform requirement

Takeaway: Encoding is the first compression pass on raw footage. Transcoding is what happens to already-encoded video afterward, and it’s a decode-then-re-encode cycle, not a synonym for encoding.

Transcoding vs. Transmuxing

TranscodingTransmuxing
What changesCodec, resolution, bitrate, or frame rate (full decode and re-encode)Only the container; the compressed video and audio bits stay untouched
CPU costHighLow, since there’s no decode or encode step
Quality impactIntroduces a new generation of lossy compressionNone, the output is bit-for-bit compatible with the source streams
Common useProducing a new resolution or codec for a device that can’t play the sourceRepackaging H.264 and AAC from MP4 into MPEG-TS for HLS, with no decode required

Takeaway: If the codec inside the file isn’t changing, you don’t need to transcode. Transmuxing (also called remuxing) is faster, cheaper, and lossless, but it only works when the target player already supports the codecs inside the file.

Transcoding vs. Transrating and Transsizing

Transrating and transsizing are both narrower operations that live inside the broader transcoding process.

TransratingTranssizingFull Transcoding
ChangesBitrate onlyResolution or aspect ratio onlyAny combination of codec, bitrate, resolution, frame rate, container, audio
CodecUnchangedUnchangedCan change
Typical use caseProducing the bitrate steps of an ABR ladder from one resolutionReformatting a 16:9 master for a vertical or square placementConverting a legacy MPEG-2 archive into modern H.264 or HEVC renditions

Takeaway: Building a bitrate ladder from a single source resolution is mostly transrating. Reformatting aspect ratio for a different screen is transsizing. A full transcode covers both plus codec and container changes, which is why it’s the most compute-intensive of the three.

Codec vs. Container

These two terms get used interchangeably in casual conversation, and mixing them up leads to real production mistakes, like assuming a device that “doesn’t support MP4” is actually a codec problem, or the reverse.

CodecContainer
What it isThe compression method for the audio or video data itselfThe wrapper format that holds the compressed streams, subtitles, and metadata together
ExamplesH.264 (AVC), H.265 (HEVC), AV1, VP9, AAC, OpusMP4, fragmented MP4/CMAF, MPEG-TS, MKV
AnalogyThe recipe and ingredientsThe container the finished dish is served in
Compatibility question it answers“Can this device decode this compression format?”“Can this player parse this file structure?”

A single container can hold different codecs (an MP4 can carry H.264 or HEVC video), and a codec can be delivered inside more than one container (H.264 shows up in MP4 for progressive download and in MPEG-TS for HLS segments). This is also why transmuxing works: the container changes, the codec inside stays the same.

Further reading: MPEGTS vs HLS vs DASH: Understand The Differences To Optimize Your Content Delivery

Live vs. VOD Transcoding

Live TranscodingVOD Transcoding
TimingRuns in real time as the feed is capturedRuns once, after upload, before the asset is published
Latency pressureHigh. Every added second delays every viewerLow. A few extra minutes of processing time rarely matters
Encoding passesEffectively single-pass, since the encoder can’t wait for the whole fileCan use multi-pass or per-title analysis, since the full source is already available
Typical infrastructureDedicated live channels, such as AWS Elemental MediaLive or similar cloud live encodersBatch or on-demand jobs, such as AWS Elemental MediaConvert or a VOD-focused transcoding service
Failure mode if under-provisionedDropped frames, frozen streams, drift out of sync with the live eventA delayed publish, which is inconvenient but rarely visible to viewers

One correction worth being explicit about: live transcoding does not run at sub-second latency in a standard HLS or DASH workflow. Glass-to-glass latency (the time from the camera to the viewer’s screen) for standard HLS or DASH typically runs 8 to 30 seconds. Low-latency variants such as LL-HLS and chunked CMAF bring that down to roughly 2 to 6 seconds by delivering partial segments as the encoder produces them. Getting under one second requires an entirely different delivery approach, such as WebRTC, not a faster version of standard segmented HLS or DASH. For live sports, where social media commentary can spoil the moment, the 2 to 6 second range from low-latency HLS or DASH is usually the realistic target, not sub-second delivery.

Read: Tips to Improve Low Latency Streaming with CMAF and a Guide to Solve Playback Delays

Practical Use Cases for Video Transcoding

  • Live sports: A single sports feed needs to reach a smart TV app, a mobile app on a spotty stadium Wi-Fi connection, and a desktop browser, often simultaneously and with strict latency requirements to avoid spoilers from second screens. This is the same real-time transcoding demand behind any live streaming solution, not just sports specifically.
  • VOD libraries: A streaming platform’s back catalog needs a consistent rendition ladder across thousands of titles, which is exactly the problem per-title encoding was built to solve. Netflix has reported bitrate savings of roughly 20% from moving off a fixed ladder to per-title analysis.
  • User-uploaded videos: Platforms that accept uploads from creators or viewers can’t control the source format, resolution, or codec, so every upload runs through a normalization transcode before it’s fit to distribute.
  • Legacy content conversion: Archives shot or mastered in older formats such as MPEG-2 need a full transcode into H.264 or HEVC before they can enter a modern ABR ladder or reach current devices.
  • Multi-platform delivery: The same title often needs different codec and container combinations for Apple TV, Android TV, Roku, and web players, since device support for codecs like HEVC or AV1 still varies.

Does Transcoding Reduce Quality?

Yes, and it’s worth saying plainly instead of glossing over it. Because transcoding decodes an already-compressed file and re-encodes it, every transcode is a new generation of lossy compression. This is sometimes called generational loss, and it compounds if a file gets transcoded repeatedly instead of always starting from the original master.

A few practical ways operators limit that loss:

  • Always transcode from the highest-quality master available, not from a previously transcoded rendition. Re-encoding a re-encode stacks compression artifacts.
  • Use a high enough source bitrate and resolution that the encoder has real detail to work with. Transcoding can’t add back information a low-quality source never had.
  • Choose transmuxing over transcoding whenever the codec doesn’t need to change. If the target player already supports the source codec, changing only the container avoids a quality-costing re-encode entirely.
  • Tune encoder settings per content type, since a fast-motion sports clip and a static talking-head interview need different bitrate allocation to look equally clean at the same file size.
  • Keep archival masters untouched. Treat the transcoded renditions as disposable outputs and the original ingest file as the one copy that never gets re-encoded.

How to Choose a Transcoding Strategy That Scales

There’s no single right answer here. The right setup depends on volume, latency tolerance, and budget, so treat these as decision criteria rather than a fixed recommendation.

Cloud vs. on-premises

Cloud transcoding (AWS Elemental, Bitmovin, Mux, and similar services) scales up and down with demand and requires no hardware to maintain, which fits most OTT startups and platforms with variable traffic. On-premises hardware makes more sense once volume is large and predictable enough that dedicated infrastructure costs less than sustained cloud usage, which is typically an enterprise-scale decision, not a startup one.

CPU vs. GPU

CPU-based encoding (software encoders like x264 or x265) generally produces better quality per bit but takes longer and costs more in compute time. GPU-based encoding (hardware encoders like NVENC) trades a bit of compression efficiency for dramatically faster throughput, which matters most for live channels and large VOD backlogs where turnaround time is the bottleneck.

Pre-transcoding vs. just-in-time

Pre-transcoding generates every rendition in the pipeline ahead of time and stores them, which keeps playback start times fast but adds storage cost for content that might rarely get watched. Just-in-time transcoding generates renditions on request, which saves storage for long-tail libraries at the cost of a short delay on first playback of a given rendition.

Fixed vs. per-title encoding

A fixed bitrate ladder is simpler to operate and predictable to budget, but it wastes bits on simple content and can under-serve complex, high-motion content at the same rung. Per-title encoding analyzes each asset’s complexity and builds a tailored ladder, which typically cuts bitrate without a visible quality loss, at the cost of more upfront encoding time per title.

Common Challenges with Transcoding

  • Cost at scale. Transcoding HD or 4K source into several renditions is compute-intensive, and cloud transcoding services bill per minute or per GB processed, so costs scale directly with rendition count and library size.
  • Latency for live events. As covered above, real-time transcoding for live sports and other live events works within a multi-second latency budget, not sub-second, and every stage in the pipeline (encode, package, CDN propagation, player buffer) adds to that total.
  • Codec compatibility. Not every device supports every codec. HEVC and AV1 offer better compression than H.264 but aren’t universally supported across browsers and older devices, which is why most ladders still include an H.264 fallback tier.
  • Storage and delivery overhead. Every additional rendition adds both storage and CDN cost. Per-title encoding and just-in-time strategies exist specifically to keep that overhead from growing linearly with library size.

If You Want Viewers, You Need Transcoding

Transcoding isn’t a backend detail you set once and forget. It’s the layer that decides whether your content actually plays for the device someone happens to be holding, on whatever connection they happen to have. Get the workflow right, keep transcoding separate from packaging and delivery in both your architecture and your budget, and choose a ladder strategy that matches your library instead of copying someone else’s defaults.

If you’re building or scaling an OTT platform and want a transcoding and delivery pipeline that’s already handling this for broadcasters, telcos, and streaming startups, OTTclouds can walk you through what a production-ready setup looks like for your content library. Contact us now!

FAQs

Does video transcoding always make quality worse?

It introduces a new generation of lossy compression, but a well-tuned transcode from a high-quality source at a reasonable bitrate can look visually indistinguishable from the original at typical viewing distances. Quality loss becomes visible mainly when bitrates are too low for the content’s complexity or when a file gets transcoded from an already-transcoded copy instead of the master.

What’s the actual difference between encoding and transcoding?

Encoding is the first compression pass on raw, uncompressed footage. Transcoding takes video that’s already compressed and decodes it, then re-encodes it into a different codec, bitrate, resolution, or container. Every transcode includes an encode step, but not every encode is a transcode.

Do I need to transcode a video if I only want to change the file extension?

Usually not. If the codec inside the file is already supported by your target player, changing only the container is transmuxing, which is faster and avoids any quality loss. Transcoding is only necessary when the codec, resolution, or bitrate itself needs to change.

Can live transcoding really run at sub-second latency?

Not with standard HLS or DASH delivery. Low-latency HLS and chunked CMAF typically achieve 2 to 6 seconds of glass-to-glass latency. Sub-second delivery requires a different protocol entirely, such as WebRTC, which trades CDN-scale economics for that extra speed.

Why does my OTT platform need more than one video codec?

Device support for newer, more efficient codecs like HEVC and AV1 still isn’t universal across browsers, smart TVs, and older mobile devices. Most platforms keep an H.264 rendition as a compatibility fallback while offering HEVC or AV1 to devices that support them for better quality at a lower bitrate.

Is per-title encoding worth the extra processing time?

For a library of any real size, usually yes. It typically reduces bitrate without a visible quality drop compared to a fixed ladder, which lowers both storage and CDN costs. The tradeoff is more upfront compute time per title, which matters more for platforms transcoding a high volume of new content daily than for a stable back catalog.

Should I choose cloud or on-premises transcoding?

Cloud transcoding fits most platforms because it scales with demand and avoids hardware maintenance. On-premises infrastructure becomes worth evaluating once transcoding volume is large and predictable enough that sustained cloud costs exceed the cost of owning and running dedicated hardware, which is typically a question for platforms operating at meaningful scale rather than early-stage OTT startups.

Sources:

Meet the author

Kiet Vo

Kiet Vo

Engineering Lead

Kiet Vo is the Technical Leader at OTTclouds, bringing extensive experience in full-stack development and system architecture. He specializes in designing and scaling software solutions for OTT platforms while leading engineering teams through the challenges of modern web and mobile ecosystems.