Bird
Raised Fist0
HLDsystem_design~12 mins

Search and recommendation in HLD - Architecture Diagram

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
System Overview - Search and recommendation

This system allows users to search for items and receive personalized recommendations. It must handle large volumes of queries quickly and provide relevant results based on user preferences and item popularity.

Architecture Diagram
User
  |
  v
Load Balancer
  |
  v
API Gateway
  |
  +--------------------+
  |                    |
  v                    v
Search Service      Recommendation Service
  |                    |
  v                    v
Search Index         User Profile DB
  |                    |
  v                    v
Cache (Search)      Cache (Recommendations)
  |                    |
  +---------+----------+
            |
            v
         Main Database
Components
User
client
Sends search queries and receives recommendations
Load Balancer
load_balancer
Distributes incoming user requests evenly to API Gateway instances
API Gateway
api_gateway
Routes requests to appropriate backend services
Search Service
service
Processes search queries using the search index
Recommendation Service
service
Generates personalized recommendations based on user profiles
Search Index
search_index
Stores indexed data for fast search retrieval
User Profile DB
database
Stores user preferences and behavior data
Cache (Search)
cache
Caches frequent search results to reduce latency
Cache (Recommendations)
cache
Caches popular recommendations for quick access
Main Database
database
Stores all item data and metadata
Request Flow - 20 Hops
UserLoad Balancer
Load BalancerAPI Gateway
API GatewaySearch Service
Search ServiceCache (Search)
Cache (Search)Search Service
Search ServiceSearch Index
Search IndexSearch Service
Search ServiceCache (Search)
API GatewayRecommendation Service
Recommendation ServiceCache (Recommendations)
Cache (Recommendations)Recommendation Service
Recommendation ServiceUser Profile DB
User Profile DBRecommendation Service
Recommendation ServiceMain Database
Main DatabaseRecommendation Service
Recommendation ServiceCache (Recommendations)
Search ServiceAPI Gateway
Recommendation ServiceAPI Gateway
API GatewayLoad Balancer
Load BalancerUser
Failure Scenario
Component Fails:Main Database
Impact:New item data cannot be fetched, recommendation service cannot update item metadata, recommendations may be stale. Search index updates may also be delayed.
Mitigation:Use database replication for failover. Cache layers serve stale data to maintain availability. Alert system triggers for quick recovery.
Architecture Quiz - 3 Questions
Test your understanding
Which component handles distributing incoming user requests to backend services?
ACache (Search)
BSearch Index
CLoad Balancer
DUser Profile DB
Design Principle
This architecture uses caching to reduce latency and load on databases, separates search and recommendation services for scalability, and employs load balancing and API gateway for efficient request routing and fault tolerance.

Practice

(1/5)
1. What is the primary purpose of a search system in a large-scale application?
easy
A. To help users quickly find relevant content from a large dataset
B. To store user passwords securely
C. To manage user account settings
D. To display advertisements randomly

Solution

  1. Step 1: Understand the role of search systems

    Search systems are designed to help users find information efficiently from large amounts of data.
  2. Step 2: Match the purpose with options

    Only To help users quickly find relevant content from a large dataset describes helping users find relevant content quickly, which is the core function of search.
  3. Final Answer:

    To help users quickly find relevant content from a large dataset -> Option A
  4. Quick Check:

    Search system purpose = find relevant content [OK]
Hint: Search systems focus on finding relevant data fast [OK]
Common Mistakes:
  • Confusing search with unrelated features like password storage
  • Thinking search manages user settings
  • Assuming search is for random content display
2. Which component is essential in a recommendation system to personalize suggestions?
easy
A. Database backup scripts
B. User behavior tracking
C. Static HTML pages
D. Load balancer configuration

Solution

  1. Step 1: Identify personalization needs

    Recommendation systems personalize suggestions based on user data and behavior.
  2. Step 2: Match components to personalization

    User behavior tracking collects data needed to tailor recommendations, unlike static pages or infrastructure tasks.
  3. Final Answer:

    User behavior tracking -> Option B
  4. Quick Check:

    Personalization needs user data = User behavior tracking [OK]
Hint: Personalization needs user data collection [OK]
Common Mistakes:
  • Confusing infrastructure tasks with personalization
  • Thinking static pages can personalize content
  • Ignoring the role of user data
3. Consider a search system that indexes 1 million documents. If the system uses an inverted index, what is the main advantage?
medium
A. Automatically deleting old documents
B. Storing documents in a single large file
C. Encrypting all documents for security
D. Faster search queries by mapping words to document lists

Solution

  1. Step 1: Understand inverted index concept

    An inverted index maps each word to the list of documents containing it, enabling quick lookups.
  2. Step 2: Identify the advantage for search speed

    This mapping allows the system to find relevant documents quickly without scanning all documents.
  3. Final Answer:

    Faster search queries by mapping words to document lists -> Option D
  4. Quick Check:

    Inverted index = fast word-to-doc lookup [OK]
Hint: Inverted index speeds up word-based search [OK]
Common Mistakes:
  • Confusing indexing with storage format
  • Thinking encryption is the main index benefit
  • Assuming index deletes documents automatically
4. A recommendation system is returning irrelevant suggestions. Which issue is most likely causing this?
medium
A. Using HTTPS instead of HTTP
B. Too many servers in the cluster
C. Incorrect or missing user behavior data
D. Database backup frequency is too high

Solution

  1. Step 1: Analyze cause of irrelevant recommendations

    Recommendations depend on accurate user data; missing or wrong data leads to poor suggestions.
  2. Step 2: Evaluate other options

    Server count, protocol choice, or backup frequency do not directly affect recommendation relevance.
  3. Final Answer:

    Incorrect or missing user behavior data -> Option C
  4. Quick Check:

    Bad recommendations = bad user data [OK]
Hint: Check user data quality for recommendation issues [OK]
Common Mistakes:
  • Blaming infrastructure instead of data quality
  • Confusing network protocols with recommendation logic
  • Ignoring data collection importance
5. You are designing a scalable recommendation system for millions of users. Which approach best balances personalization and system performance?
hard
A. Use a hybrid model combining collaborative filtering and content-based filtering with offline batch processing and online updates
B. Only recommend the most popular items to all users without personalization
C. Run real-time deep learning models for every user request without caching
D. Store all user data in a single database server for simplicity

Solution

  1. Step 1: Understand scalability and personalization needs

    Millions of users require efficient processing; personalization improves user experience.
  2. Step 2: Evaluate approaches

    Hybrid models combine strengths of different methods. Offline batch processing reduces load, while online updates keep recommendations fresh.
  3. Step 3: Reject less scalable or less personalized options

    Popular-only recommendations lack personalization. Real-time deep learning per request is costly. Single DB server is a bottleneck.
  4. Final Answer:

    Use a hybrid model combining collaborative filtering and content-based filtering with offline batch processing and online updates -> Option A
  5. Quick Check:

    Hybrid + batch + online = scalable personalized system [OK]
Hint: Combine offline and online methods for scalable personalization [OK]
Common Mistakes:
  • Ignoring scalability by doing all processing online
  • Sacrificing personalization for simplicity
  • Using single server for massive data