Bird
Raised Fist0
HLDsystem_design~25 mins

Product catalog design in HLD - System Design Exercise

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Design: Product Catalog System
Design covers product metadata storage, search, and retrieval APIs. Image storage and CDN are out of scope. User authentication and order management are out of scope.
Functional Requirements
FR1: Store product information including name, description, price, and images
FR2: Support product categories and subcategories
FR3: Allow searching products by name, category, and attributes
FR4: Support filtering products by price range, brand, and other attributes
FR5: Handle up to 1 million products
FR6: Provide API for product retrieval with response time under 200ms
FR7: Support updates to product details with eventual consistency
FR8: Allow bulk import and export of product data
Non-Functional Requirements
NFR1: System should handle 1000 concurrent read requests
NFR2: API availability target of 99.9%
NFR3: Search queries should return results within 200ms p99 latency
NFR4: Data consistency can be eventual for product updates
NFR5: Images stored separately from product metadata
Think Before You Design
Questions to Ask
❓ Question 1
❓ Question 2
❓ Question 3
❓ Question 4
❓ Question 5
❓ Question 6
Key Components
API Gateway for client requests
Product Metadata Database (relational or NoSQL)
Search Engine (e.g., Elasticsearch) for fast querying
Cache layer (e.g., Redis) for frequently accessed products
Bulk import/export service
Background job processor for syncing updates to search index
Design Patterns
CQRS (Command Query Responsibility Segregation) for separating reads and writes
Eventual consistency for search index updates
Pagination and sorting for product lists
Denormalization for faster reads
API versioning for backward compatibility
Reference Architecture
Client
  |
  v
API Gateway
  |
  +--------------------+
  |                    |
Product Metadata DB   Search Engine
  |                    |
Cache Layer <---------+
  |
Bulk Import/Export Service
  |
Background Job Processor (sync to Search)
Components
API Gateway
RESTful API server (e.g., Node.js/Express, Spring Boot)
Handles client requests, routes to appropriate services, enforces rate limiting
Product Metadata Database
Relational DB (PostgreSQL) or NoSQL (MongoDB)
Stores product details, categories, and attributes
Search Engine
Elasticsearch
Provides fast search and filtering capabilities
Cache Layer
Redis
Caches frequently accessed product data to reduce DB load
Bulk Import/Export Service
Batch processing system (e.g., Python scripts, AWS Lambda)
Handles large scale product data uploads and downloads
Background Job Processor
Message queue + worker (e.g., RabbitMQ + Celery)
Syncs product updates from DB to search engine asynchronously
Request Flow
1. Client sends product search request to API Gateway
2. API Gateway checks Cache Layer for results
3. If cache miss, API Gateway queries Search Engine
4. Search Engine returns matching product IDs
5. API Gateway fetches product details from Product Metadata DB or Cache
6. API Gateway returns product data to client
7. For product updates, API Gateway writes to Product Metadata DB
8. Background Job Processor reads DB changes and updates Search Engine asynchronously
9. Bulk import service writes product data to Product Metadata DB in batches
Database Schema
Entities: - Product (id PK, name, description, price, brand, category_id FK, attributes JSONB, created_at, updated_at) - Category (id PK, name, parent_category_id FK nullable) - ProductImage (id PK, product_id FK, image_url, alt_text) Relationships: - One Category can have many Products (1:N) - One Product can have many ProductImages (1:N) - Categories can have hierarchical parent-child relationship (self-referencing FK)
Scaling Discussion
Bottlenecks
Database read load increases with product queries
Search Engine indexing delays with frequent updates
Cache invalidation complexity with product updates
Bulk import causing DB performance degradation
Solutions
Use read replicas for the database to distribute read traffic
Implement incremental and batched indexing for search engine
Use cache expiration and event-driven cache invalidation
Throttle bulk import jobs and run during off-peak hours
Interview Tips
Time: 10 minutes for requirements and clarifications, 15 minutes for architecture and data flow, 10 minutes for scaling and trade-offs, 10 minutes for Q&A
Clarify functional and non-functional requirements upfront
Explain choice of database and search engine clearly
Describe how caching improves performance
Discuss eventual consistency trade-offs for search index
Highlight scalability strategies and bottleneck mitigation
Mention API design considerations and error handling

Practice

(1/5)
1.

What is the primary purpose of a product catalog in an e-commerce system?

easy
A. To organize products into categories for easy browsing
B. To process payments securely
C. To manage user login sessions
D. To handle shipping and delivery logistics

Solution

  1. Step 1: Understand the role of a product catalog

    A product catalog groups products logically so users can browse easily.
  2. Step 2: Differentiate from other system components

    Payment, user sessions, and shipping are separate concerns from product organization.
  3. Final Answer:

    To organize products into categories for easy browsing -> Option A
  4. Quick Check:

    Product catalog = Organize products [OK]
Hint: Catalog = product organization, not payments or shipping [OK]
Common Mistakes:
  • Confusing catalog with payment processing
  • Mixing catalog with user authentication
  • Thinking catalog manages shipping
2.

Which data structure is most suitable to represent categories and subcategories in a product catalog?

A. Array
B. Linked List
C. Tree
D. Hash Map

easy
A. Array
B. Linked List
C. Hash Map
D. Tree

Solution

  1. Step 1: Analyze category relationships

    Categories have parent-child relationships, forming a hierarchy.
  2. Step 2: Choose data structure for hierarchy

    A tree structure naturally represents hierarchical data with branches and leaves.
  3. Final Answer:

    Tree -> Option D
  4. Quick Check:

    Hierarchy = Tree [OK]
Hint: Hierarchies = Trees, not flat lists or arrays [OK]
Common Mistakes:
  • Using arrays which are flat and unordered
  • Choosing linked lists which are linear
  • Hash maps don't represent hierarchy well
3.

Consider a product catalog system that uses an inverted index for search. What is the main benefit of using an inverted index?

medium
A. It manages user reviews and ratings
B. It stores product images efficiently
C. It speeds up product search by mapping keywords to product IDs
D. It handles payment transactions securely

Solution

  1. Step 1: Understand inverted index purpose

    An inverted index maps keywords to the list of documents or products containing them.
  2. Step 2: Apply to product search

    This mapping allows fast lookup of products matching search terms, improving speed.
  3. Final Answer:

    It speeds up product search by mapping keywords to product IDs -> Option C
  4. Quick Check:

    Inverted index = fast keyword search [OK]
Hint: Inverted index = keyword to product map for fast search [OK]
Common Mistakes:
  • Thinking it stores images
  • Confusing with user review storage
  • Mixing with payment processing
4.

A product catalog system caches product details but users report seeing outdated information. What is the likely cause?

medium
A. Cache invalidation is not handled properly after product updates
B. The database schema is incorrect
C. User authentication is failing
D. The product images are missing

Solution

  1. Step 1: Identify caching issue

    Outdated info usually means cache still holds old data after updates.
  2. Step 2: Understand cache invalidation

    Proper cache invalidation removes or refreshes cached data when products change.
  3. Final Answer:

    Cache invalidation is not handled properly after product updates -> Option A
  4. Quick Check:

    Outdated cache = invalidation problem [OK]
Hint: Outdated data? Check cache invalidation first [OK]
Common Mistakes:
  • Blaming database schema without evidence
  • Confusing with authentication issues
  • Assuming missing images cause outdated text
5.

You are designing a product catalog for a global e-commerce platform with millions of products and frequent updates. Which design choice best supports scalability and fast search?

A. Use a distributed NoSQL database with indexing and cache layers
B. Store all products in a single relational database without caching
C. Use flat files to store product data and search sequentially
D. Keep product data only in application memory without persistence

hard
A. Store all products in a single relational database without caching
B. Use a distributed NoSQL database with indexing and cache layers
C. Use flat files to store product data and search sequentially
D. Keep product data only in application memory without persistence

Solution

  1. Step 1: Consider scalability needs

    Millions of products and frequent updates require a scalable, distributed system.
  2. Step 2: Evaluate design options

    A distributed NoSQL database supports horizontal scaling; indexing enables fast search; caching improves response time.
  3. Step 3: Reject unsuitable options

    Single DB limits scale; flat files are slow; in-memory only risks data loss.
  4. Final Answer:

    Use a distributed NoSQL database with indexing and cache layers -> Option B
  5. Quick Check:

    Scalable + fast search = distributed NoSQL + index + cache [OK]
Hint: Scale needs distributed DB + index + cache, not flat files [OK]
Common Mistakes:
  • Choosing single DB without caching for large scale
  • Using flat files causing slow search
  • Relying on memory only risking data loss