Bird
Raised Fist0
HLDsystem_design~10 mins

Video upload and processing pipeline in HLD - Scalability & System Analysis

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Scalability Analysis - Video upload and processing pipeline
Growth Table: Video Upload and Processing Pipeline
ScaleUsersUploads per secondStorage NeededProcessing LoadNetwork Bandwidth
Small1001-210-20 GBSingle server, basic transcoding100 Mbps
Medium10,000100-2001-2 TBMultiple processing workers, queue system1 Gbps
Large1,000,00010,000+100+ TBDistributed processing, sharded storage10+ Gbps
Very Large100,000,000100,000+10+ PBMassive distributed systems, CDN integration100+ Gbps
First Bottleneck

At small to medium scale, the video processing workers become the first bottleneck. Video transcoding is CPU and memory intensive. A single server can handle only a limited number of concurrent video conversions.

As scale grows, storage I/O and network bandwidth also become bottlenecks due to large video file sizes.

Scaling Solutions
  • Horizontal scaling: Add more processing worker servers to handle transcoding in parallel.
  • Queue system: Use message queues to manage upload and processing tasks asynchronously.
  • Caching: Cache frequently accessed videos or thumbnails using CDN to reduce load.
  • Sharding storage: Distribute video files across multiple storage nodes to improve I/O.
  • CDN integration: Use Content Delivery Networks to serve processed videos closer to users, reducing bandwidth load on origin servers.
  • Auto-scaling: Automatically add or remove processing resources based on upload traffic.
Back-of-Envelope Cost Analysis
  • At 10,000 uploads per second, assuming average video size 20 MB, storage needed per day: 10,000 * 20 MB * 3600 * 24 ≈ 17.28 TB/day (requires compression and retention policies).
  • Processing: Each transcoding takes ~5 CPU-minutes; 10,000 uploads/sec means 50,000 CPU-minutes/sec -> requires massive parallelism.
  • Network bandwidth: 10,000 uploads/sec * 20 MB = 200 GB/s (~1.6 Tbps) inbound bandwidth needed.
Interview Tip

Start by clarifying the scale and requirements. Then identify the main components: upload, storage, processing, delivery. Discuss bottlenecks at each scale and propose targeted solutions. Use real numbers to justify your choices. Always mention asynchronous processing and CDN usage for video systems.

Self Check

Your video processing system handles 1000 uploads per second. Traffic grows 10x to 10,000 uploads per second. What do you do first?

Answer: Add more processing worker servers horizontally and implement or scale the queue system to handle the increased load. Also, consider sharding storage and increasing network bandwidth to avoid bottlenecks.

Key Result
Video processing bottlenecks first at transcoding CPU load; horizontal scaling of workers and sharded storage with CDN delivery are key to handle growth from thousands to millions of uploads.

Practice

(1/5)
1. Which component in a video upload and processing pipeline is primarily responsible for converting raw uploaded videos into multiple formats suitable for playback?
easy
A. Content delivery network (CDN)
B. Video processing service
C. Metadata database
D. Upload service

Solution

  1. Step 1: Identify the role of each component

    The upload service handles receiving videos, the metadata database stores info, and CDN delivers content. The processing service converts videos.
  2. Step 2: Match the function to the question

    Converting raw videos into multiple formats is done by the video processing service to ensure compatibility.
  3. Final Answer:

    Video processing service -> Option B
  4. Quick Check:

    Conversion = Video processing service [OK]
Hint: Processing means converting video formats [OK]
Common Mistakes:
  • Confusing upload service with processing
  • Thinking CDN does video conversion
  • Assuming metadata database handles video files
2. Which of the following is the correct sequence of steps in a typical video upload and processing pipeline?
easy
A. Storage -> Upload service -> Video processing -> Metadata update
B. Video processing -> Upload service -> Metadata update -> Storage
C. Metadata update -> Upload service -> Storage -> Video processing
D. Upload service -> Video processing -> Storage -> Metadata update

Solution

  1. Step 1: Understand the logical flow

    Users first upload videos, then videos are processed, stored, and metadata is updated last.
  2. Step 2: Match the sequence to options

    Upload service -> Video processing -> Storage -> Metadata update correctly shows upload first, then processing, storage, and metadata update.
  3. Final Answer:

    Upload service -> Video processing -> Storage -> Metadata update -> Option D
  4. Quick Check:

    Upload first, then process, store, update metadata [OK]
Hint: Upload happens before processing and storage [OK]
Common Mistakes:
  • Starting with processing before upload
  • Updating metadata before storage
  • Mixing storage and upload order
3. Consider a video upload pipeline where the upload service places video metadata into a queue for processing. If the processing service crashes and stops consuming messages, what will happen to the queue and user experience?
medium
A. Upload service will reject new uploads immediately
B. Queue will empty quickly; users get instant processing
C. Queue will fill up, causing delays; users see slow processing
D. Metadata database will automatically process videos

Solution

  1. Step 1: Understand queue behavior when consumer stops

    If the processing service crashes, it stops consuming messages, so the queue fills up with unprocessed metadata.
  2. Step 2: Impact on user experience

    Since processing is delayed, users experience slow video availability or processing delays.
  3. Final Answer:

    Queue will fill up, causing delays; users see slow processing -> Option C
  4. Quick Check:

    Processing down -> queue fills -> delays [OK]
Hint: No consumer means queue backs up [OK]
Common Mistakes:
  • Assuming queue empties without consumer
  • Thinking upload service rejects uploads immediately
  • Believing metadata DB processes videos automatically
4. In a video processing pipeline, a developer notices that some videos fail to process and the system does not retry them. Which change will fix this issue?
medium
A. Implement a retry mechanism in the processing service for failed jobs
B. Remove the queue to speed up processing
C. Store videos only after processing completes
D. Disable metadata updates to avoid conflicts

Solution

  1. Step 1: Identify cause of failure handling

    Failures without retries mean the system lacks a retry mechanism for failed processing jobs.
  2. Step 2: Choose fix to handle failures

    Adding retries ensures failed jobs are re-attempted, improving reliability.
  3. Final Answer:

    Implement a retry mechanism in the processing service for failed jobs -> Option A
  4. Quick Check:

    Retries fix failed processing [OK]
Hint: Retries fix failed processing jobs [OK]
Common Mistakes:
  • Removing queue breaks asynchronous design
  • Storing before processing causes errors
  • Disabling metadata updates unrelated to retries
5. You need to design a scalable video upload and processing pipeline that supports millions of daily uploads with minimal user wait time. Which architectural choice best supports this goal?
hard
A. Use asynchronous upload service with message queues and distributed processing workers
B. Process videos synchronously during upload to ensure immediate availability
C. Store all videos on a single server to simplify management
D. Update metadata only after manual verification to ensure accuracy

Solution

  1. Step 1: Analyze scalability and user wait time needs

    Millions of uploads require asynchronous handling and distributed processing to avoid bottlenecks and reduce wait time.
  2. Step 2: Evaluate architectural options

    Asynchronous upload with queues and distributed workers allows parallel processing and smooth scaling.
  3. Final Answer:

    Use asynchronous upload service with message queues and distributed processing workers -> Option A
  4. Quick Check:

    Asynchronous + distributed = scalable and fast [OK]
Hint: Async + queues + distributed workers scale best [OK]
Common Mistakes:
  • Synchronous processing causes delays
  • Single server storage limits scalability
  • Manual metadata delays pipeline