Bird
Raised Fist0
HLDsystem_design~10 mins

Product catalog design in HLD - Scalability & System Analysis

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Scalability Analysis - Product catalog design
Growth Table: Product Catalog Design
ScaleUsersCatalog SizeTraffic (Requests/sec)StorageKey Changes
Small10010K products50~1 GBSingle database instance, simple caching
Medium10K100K products5,000~10 GBRead replicas, distributed cache, basic load balancing
Large1M1M products100,000~100 GBSharded database, CDN for images, advanced caching
Very Large100M100M+ products5,000,000~10 TB+Multi-region sharding, microservices, global CDN, asynchronous processing
First Bottleneck

At small to medium scale, the database is the first bottleneck. It struggles with high read and write queries as users browse and update product info. As traffic grows, a single database instance cannot handle the load, causing slow responses.

Scaling Solutions
  • Read Replicas: Add read-only database copies to spread read traffic.
  • Caching: Use in-memory caches (like Redis) to store popular product data and reduce database hits.
  • Sharding: Split the product catalog across multiple database servers by product ID or category to distribute load.
  • CDN: Serve product images and static content via Content Delivery Networks to reduce server bandwidth and latency.
  • Horizontal Scaling: Add more application servers behind load balancers to handle increased user requests.
  • Asynchronous Processing: Use background jobs for heavy tasks like indexing or bulk updates to keep the system responsive.
Back-of-Envelope Cost Analysis
  • At 1M users with 100,000 requests/sec, assuming 50% cache hit rate, database handles ~50,000 QPS.
  • Each product record ~1 KB, 1M products = ~1 GB raw data; with indexes and metadata, ~10-20 GB storage needed.
  • Images stored on CDN reduce bandwidth on origin servers; typical image size ~100 KB, 100,000 views/sec = ~10 GB/s bandwidth if uncached.
  • Network bandwidth and storage costs grow significantly at large scale; plan for multi-region deployments.
Interview Tip

Start by clarifying scale and requirements. Discuss current bottlenecks and how they evolve with growth. Propose incremental solutions: caching, read replicas, sharding, CDN. Explain trade-offs and how each solution fits the scale. Use real numbers to justify choices.

Self Check

Your database handles 1000 QPS. Traffic grows 10x to 10,000 QPS. What do you do first?

Answer: Add read replicas and implement caching to reduce load on the primary database before considering sharding or more complex solutions.

Key Result
The product catalog design first hits database bottlenecks as user traffic and catalog size grow; scaling requires caching, read replicas, and sharding combined with CDN and horizontal scaling to maintain performance.

Practice

(1/5)
1.

What is the primary purpose of a product catalog in an e-commerce system?

easy
A. To organize products into categories for easy browsing
B. To process payments securely
C. To manage user login sessions
D. To handle shipping and delivery logistics

Solution

  1. Step 1: Understand the role of a product catalog

    A product catalog groups products logically so users can browse easily.
  2. Step 2: Differentiate from other system components

    Payment, user sessions, and shipping are separate concerns from product organization.
  3. Final Answer:

    To organize products into categories for easy browsing -> Option A
  4. Quick Check:

    Product catalog = Organize products [OK]
Hint: Catalog = product organization, not payments or shipping [OK]
Common Mistakes:
  • Confusing catalog with payment processing
  • Mixing catalog with user authentication
  • Thinking catalog manages shipping
2.

Which data structure is most suitable to represent categories and subcategories in a product catalog?

A. Array
B. Linked List
C. Tree
D. Hash Map

easy
A. Array
B. Linked List
C. Hash Map
D. Tree

Solution

  1. Step 1: Analyze category relationships

    Categories have parent-child relationships, forming a hierarchy.
  2. Step 2: Choose data structure for hierarchy

    A tree structure naturally represents hierarchical data with branches and leaves.
  3. Final Answer:

    Tree -> Option D
  4. Quick Check:

    Hierarchy = Tree [OK]
Hint: Hierarchies = Trees, not flat lists or arrays [OK]
Common Mistakes:
  • Using arrays which are flat and unordered
  • Choosing linked lists which are linear
  • Hash maps don't represent hierarchy well
3.

Consider a product catalog system that uses an inverted index for search. What is the main benefit of using an inverted index?

medium
A. It manages user reviews and ratings
B. It stores product images efficiently
C. It speeds up product search by mapping keywords to product IDs
D. It handles payment transactions securely

Solution

  1. Step 1: Understand inverted index purpose

    An inverted index maps keywords to the list of documents or products containing them.
  2. Step 2: Apply to product search

    This mapping allows fast lookup of products matching search terms, improving speed.
  3. Final Answer:

    It speeds up product search by mapping keywords to product IDs -> Option C
  4. Quick Check:

    Inverted index = fast keyword search [OK]
Hint: Inverted index = keyword to product map for fast search [OK]
Common Mistakes:
  • Thinking it stores images
  • Confusing with user review storage
  • Mixing with payment processing
4.

A product catalog system caches product details but users report seeing outdated information. What is the likely cause?

medium
A. Cache invalidation is not handled properly after product updates
B. The database schema is incorrect
C. User authentication is failing
D. The product images are missing

Solution

  1. Step 1: Identify caching issue

    Outdated info usually means cache still holds old data after updates.
  2. Step 2: Understand cache invalidation

    Proper cache invalidation removes or refreshes cached data when products change.
  3. Final Answer:

    Cache invalidation is not handled properly after product updates -> Option A
  4. Quick Check:

    Outdated cache = invalidation problem [OK]
Hint: Outdated data? Check cache invalidation first [OK]
Common Mistakes:
  • Blaming database schema without evidence
  • Confusing with authentication issues
  • Assuming missing images cause outdated text
5.

You are designing a product catalog for a global e-commerce platform with millions of products and frequent updates. Which design choice best supports scalability and fast search?

A. Use a distributed NoSQL database with indexing and cache layers
B. Store all products in a single relational database without caching
C. Use flat files to store product data and search sequentially
D. Keep product data only in application memory without persistence

hard
A. Store all products in a single relational database without caching
B. Use a distributed NoSQL database with indexing and cache layers
C. Use flat files to store product data and search sequentially
D. Keep product data only in application memory without persistence

Solution

  1. Step 1: Consider scalability needs

    Millions of products and frequent updates require a scalable, distributed system.
  2. Step 2: Evaluate design options

    A distributed NoSQL database supports horizontal scaling; indexing enables fast search; caching improves response time.
  3. Step 3: Reject unsuitable options

    Single DB limits scale; flat files are slow; in-memory only risks data loss.
  4. Final Answer:

    Use a distributed NoSQL database with indexing and cache layers -> Option B
  5. Quick Check:

    Scalable + fast search = distributed NoSQL + index + cache [OK]
Hint: Scale needs distributed DB + index + cache, not flat files [OK]
Common Mistakes:
  • Choosing single DB without caching for large scale
  • Using flat files causing slow search
  • Relying on memory only risking data loss