Ch.11: System Design: Streaming Video Without Buffering
Outline
- 0:00 Introduction
- 0:27 The Streaming Map
- 1:05 Play Button Physics
- 1:45 Why One File Does Not Work
- 2:32 Functional Requirements
- 3:14 Functional Requirements, Continued
- 3:52 Non-Functional Requirements
- 4:25 Non-Functional Requirements, Continued
- 4:56 Encoding Ladder
- 5:46 Chunked Segments and Manifests
- 6:23 Adaptive Bitrate Playback
- 7:25 CDN at the Edge
- 8:01 Storage Tiering
- 8:38 Thumbnails and Preview Sprites
- 9:09 Live Streaming
- 9:52 Architecture Payoff
- 10:35 Why the Design Holds
- 11:36 Closing
Transcript
0:00 Welcome to Learning Podcasts. System Design for Backend Engineers: Streaming Video Without Buffering. You tap play, and a few hundred milliseconds later the picture is moving. The viewer thinks the file just started downloading. But the file already exists in about a dozen versions, sliced into pieces, and parked close to the viewer. Today we trace what makes that tap feel free, even when the network and device are both unpredictable. Start with the finished map. The write side begins at upload: the file lands in ingest, then a transcoder fans it out into a quality ladder, a segmenter cuts each version into chunks, and a manifest writer publishes the index.
0:41 The serve side runs from the player to the nearest edge in the CDN, back to origin storage if the chunk is not cached, and a thumbnail and preview-sprite system runs alongside. And there is a parallel path for live, where ingest runs continuously and the segments are published as they encode. Right. The chapter unpacks each region and returns to this same map at the end. One tap on the device. From that moment, the player has a budget. Bytes have to arrive faster than the player draws frames, on a network the system does not control.
1:13 This is not a download. A download can take its time. Streaming has a moving deadline that runs for the whole duration of playback. So if the buffer empties even once, the viewer sees a spinner, and the spinner is the failure. Exactly. Every other choice in the architecture exists to keep that buffer full, on devices and networks the team will never see. That is the real product surface. Smooth playback is the only thing the viewer measures. The reason this is hard at all is that the network changes mid-playback.
1:46 The bandwidth a viewer has at second one is not the bandwidth they have at second 20. A phone walks into an elevator. A laptop shares Wi-Fi with a video call. A TV behind a saturated router gets squeezed. The network is not a property of the upload. It is a property of the moment. So picking one bitrate at upload time and serving a single encoded file is, Locking the viewer to one moment forever. The first elevator or tunnel breaks the contract. Which means the design has to leave room for the player to switch quality during the stream, not just choose once at the start.
2:21 Right. Once mid-stream pivots are part of the requirement, every other component in the architecture bends around that pivot. Let us pin down the functional requirements before we choose any storage. The system has to accept an upload, transcode it into a ladder of resolutions and bitrates, segment each variant into seconds-long chunks, write a manifest that lists what exists, and serve playback by URL. And every one of those is a job, not a single function. The transcoder is a fanout. The segmenter is a per-variant pass.
2:55 The manifest is its own published artifact. Right. The upload comes in once. The work that comes out of it is a graph of derived objects. Which means upload completion is not the same as `ready to play`. Ready to play happens when the manifest is written. There is more on the functional list. Generate thumbnails for the catalog, build preview-sprite sheets for the hover scrubber, attach captions and alternate audio tracks, and treat live as its own ingest path with the same segmented output. Plus retention and deletion controls.
3:29 Operators need to take a video down, restrict it by region, or delete it entirely, and that has to propagate to every cached chunk. Otherwise a takedown is just a database flag while the actual chunks keep playing from cache. Exactly. Streaming systems live or die by how cleanly they invalidate. Non-functional requirements set the bar. Startup latency from tap to first frame has to stay under a second at a high percentile, not at the median. Adaptation has to be smooth when bandwidth dips. The player should drop one rung on the ladder, not stall.
4:06 And availability has to hold across regions, because viewers are on every continent. And startup is the part viewers remember. A 2-second delay on tap feels broken even if the rest of the playback is perfect. Right. The first second sets the impression for the whole session. The rest of the non-functional list is about cost and visibility. Storage and egress are real money, so hot trending content cannot be priced the same as the long tail. And originals stay durable forever, even after the playback variants are aged out of expensive tiers.
4:37 Plus separate dashboards for older devices and slower networks. A single startup-latency average looks healthy while a real population of viewers is stalling. Exactly. The viewer in a 4G cell on an older phone is the test case, not the average. When the upload lands, the encoding ladder fans out. The same source becomes 240, 360, 480, 720, 1080, and sometimes 4K. Each rung is a separate encoded variant. So the transcoder is not one render. It is a worker pool consuming jobs from a queue, the same pattern as the message-queue episode.
5:16 Right. One upload becomes many parallel encoding jobs. The pool scales with how many uploads land per minute. And the codec choice on each rung is its own decision: H.264 for compatibility, more efficient codecs where the player supports them. That part is easy to forget. The ladder is not just resolutions. It is a matrix of resolution, bitrate, and codec, and the manifest has to describe all of it. Each variant is sliced into 2 to 10 second chunks. The chunks are independent files, addressable by URL, and a manifest tells the player which chunks exist at which qualities.
5:54 That is what HLS and DASH look like in practice. Both are a manifest plus a folder of chunks. And the chunk boundary is what makes adaptation possible. The player can request the next chunk from a different rung of the ladder without restarting the stream. So the manifest is not metadata. It is the instruction set the player runs against. Exactly. The manifest is small, the chunks are not, and the design separates the two on purpose. The player measures throughput and buffer health on every chunk.
6:23 If the buffer is shrinking and the network is slowing, it picks the next chunk from a lower rung. And if the network recovers, it climbs back up. Quality changes mid-stream because conditions change mid-stream, and the player owns that decision. But that is the part that always nags at me. The player runs on the viewer's device. What stops a buggy or hostile client from grabbing the highest rung every time and ignoring the buffer signal? Honestly, the protocol does not stop it. The server can detect obviously bad behavior and rate-limit, but the design assumes the client is acting in good faith.
6:58 That is an open seam in every adaptive bitrate system I have seen. So adaptive bitrate is not a server feature. It is a client behavior fed by what the manifest exposes, and the server only really pushes back at the edges. Right. The server's job is to make every rung available. The player's job is to pick the rung that keeps the buffer full, and the system trusts that the player will. In front of all of this, the CDN holds chunks at the edge. Popular chunks are cached close to viewers in regional points of presence, and origin only sees misses.
7:32 That is the caching episode at heavier load. Distance is latency, and viewers are everywhere. And request collapsing matters at the edge. If a thousand viewers ask for the same trending chunk in the same second, the edge fetches origin once. Otherwise origin gets the full thundering herd, and the savings disappear. Right. The CDN is the layer that turns one popular video into a feasible workload. Behind the CDN, storage is not one bucket. Hot tier holds trending and recent uploads, warm tier holds the active catalog, and cold tier holds the long tail at much lower cost and higher access latency.
8:12 And originals never leave durable storage, even when their playback variants get aged into colder tiers or rebuilt later. Right. Originals are the source of truth. The variants are derived, replaceable, and can be regenerated if a new codec or quality target lands. So the tiering is about access shape, not about whether the video survives. The video always survives. Exactly. This is the distributed-cache lesson at storage scale: hot is small and expensive, cold is big and cheap, and the access pattern decides where each asset lives.
8:44 The preview surface is a sibling system to playback, Not an afterthought. At upload time, the pipeline extracts frames and builds thumbnails, social-card variants, and a sprite sheet that powers the hover scrubber. And those assets ride the same CDN as the chunks. The catalog page is fast because the previews are cached at the edge too. If the preview pipeline lags, the catalog feels broken even when playback works fine. So preview generation runs on the same upload-completion event as the encoding ladder.
9:15 Live looks like the same diagram with two changes. Ingest is a continuous stream, not a one-shot upload, and segments publish as they encode instead of after the whole pipeline finishes. And the latency budget shrinks from minutes to seconds. The viewer expects to be near real time. Right. The encoding ladder still fans out, the manifest still updates, and the CDN still caches. But the manifest grows continuously, and the player keeps re-reading it. So live is not on-demand minus the upload step.
9:49 It is on-demand with a live-updating manifest and a much tighter clock. That is the shift. Same architecture, different latency contract. Now return to the map. Upload lands at ingest. The transcoder fans the source into a quality ladder. The segmenter cuts each variant into chunks. The manifest writer publishes the index. On the serve side, the player reads the manifest, requests chunks from the nearest CDN edge, and adapts the next chunk's quality based on what it just measured. Storage tiering keeps hot, warm, and cold separated by access pattern.
10:22 Originals stay durable in their own tier, regardless of which variants are live. Thumbnails and preview sprites publish from the same upload event and ride the same CDN. Live ingest replaces the upload path and publishes manifests continuously. That separation is the point. Encoding pays the upfront cost so playback can be calm, and the CDN turns calm playback into something the system can actually afford at scale. The whole thing rests on 5 disciplined moves. Encode once into many qualities. Slice into chunks for adaptation.
10:49 Push the popular chunks to the edge. Keep originals durable. Price storage by access shape. The thread connecting all 5 is that nothing in the read path tries to be clever in real time. The cleverness is upstream. Encode once means playback never re-encodes. Chunks let the player adapt without restarting. Edge caching collapses distance. Original masters keep new codecs in reach. Access-shape pricing keeps the long tail affordable. And every one of those moves trades upfront work for steady playback.
11:22 The viewer feels a smooth stream because the write side and the storage layout did the planning weeks earlier. That is the streaming contract. Viewers see a single tap. The architecture sees a quality ladder, a manifest, a CDN, and a tiered backbone working together to keep one buffer full. Streaming is not a file transfer with a fancy player. It is a pre-built ladder of variants, sliced for adaptation, cached close to viewers, tiered by access pattern, and refreshed by a separate live path when the clock tightens.
11:55 Next video, we leave streaming and design a distributed key-value store, where the question shifts from playback smoothness to consistent access across many machines. Thanks for listening to Learning Podcasts.