Study interactive :: Progress tools open in the Study Hub reader.

25. Design a Video Streaming Platform

YouTube, Vimeo, TikTok, Netflix. Users upload video, the system processes it, the world streams it. Sounds like one product. It's actually three: upload, processing, and playback. Each has its own bottleneck.

Clarify

Question Example answer
User-uploaded or curated? User-uploaded
Live or on-demand? On-demand for v1; live later
Resolutions? 240p to 4K
Comments / likes? Yes
Subscriptions? Yes
Mobile and TV clients? Yes
Region availability? Global

Estimate

Numbers like these mean three things: object storage from day one, CDN from day one, and asynchronous processing for everything that isn't the playback path.

High-level design

UploaderAPIObject store rawvideo.uploadedTranscoderThumbnail genCaptions MLObject store renditionsThumbnailsCaptions DBCDNViewer player

The upload and processing pipeline lives off the hot path. The playback path is player → CDN → object store. Most viewers never hit your origin.

Deep dive 1: Upload

Don't proxy a 4 GB upload through your API server. Use a pre-signed URL from object storage (Chapter 18):

1. Client: POST /uploads/init   -> server returns upload ID + pre-signed URL
2. Client: PUT raw bytes directly to S3 (chunked + resumable)
3. Client: POST /uploads/{id}/complete
4. Server publishes "video.uploaded" event

Resumable upload (S3 multipart, GCS resumable) so a mobile user with flaky wifi doesn't lose 30 minutes of work.

Deep dive 2: Transcoding

A 1080p source is usually cut into lower and equal renditions (for example 240p, 360p, 480p, 720p, 1080p). Do not invent a 4K ladder from a 1080p master. Chunk into segments of about 4 to 10 seconds and package for HLS or DASH (adaptive streaming).

input.mp4
  -> ffmpeg -> 240p chunks (.ts)
  -> ffmpeg -> 360p chunks
  -> ffmpeg -> ... up to source max (e.g. 1080p)
  -> generate .m3u8 (HLS) or .mpd (DASH) manifest

This is parallelizable per rendition and per chunk. Run it on a worker fleet you can scale to thousands of cores when a viral upload hits. AWS MediaConvert, GCP Transcoder API, or your own Kubernetes-scheduled ffmpeg pool.

A video is "watchable" when at least one rendition is done. Don't make users wait for 4K.

Deep dive 3: Playback

The player downloads a manifest:

#EXTM3U
#EXT-X-STREAM-INF:BANDWIDTH=500000,RESOLUTION=426x240
240p/index.m3u8
#EXT-X-STREAM-INF:BANDWIDTH=1500000,RESOLUTION=854x480
480p/index.m3u8
...

It picks a rendition based on current bandwidth, fetches chunks, switches up or down as the network changes. This is adaptive bitrate streaming (ABR). The magic that makes the video keep playing even when wifi gets bad.

Every chunk is a static file on the CDN. Versioned URLs (/v/abc/720p/00012.ts) so caching is trivial.

Deep dive 4: Storage tiering

Old, rarely-watched videos cost real money to keep on hot storage.

Upload day:    S3 Standard (hot)
After 30d:     S3 Standard-IA (infrequent access)
After 180d:    S3 Glacier Instant Retrieval
After 1y:      S3 Glacier Deep Archive (cheap, restore in hours)

Promote back to hot if traffic spikes (someone tweets an old video).

Deep dive 5: Metadata

Posts, titles, views, comments, likes. None of this is in object storage. It's a relational store.

CREATE TABLE videos (
    id            BIGINT PRIMARY KEY,
    uploader_id   BIGINT,
    title         TEXT,
    description   TEXT,
    duration_s    INT,
    status        TEXT,         -- 'processing', 'live', 'failed'
    view_count    BIGINT DEFAULT 0,
    created_at    TIMESTAMPTZ
);

Shard by video_id. View counts: do not update Postgres on every play. Buffer counts in Redis or Kafka and flush every minute (Chapter 19).

Deep dive 6: Recommendations

Out of scope for the storage and serving system, but worth a sentence: every play, like, and watch-time event flows through Kafka into a feature store and a model service. See Chapter 30.

Things to remember

Going deeper