Bird
Raised Fist0
HLDsystem_design~12 mins

Search and metadata in HLD - Architecture Diagram

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
System Overview - Search and metadata

This system allows users to search through large collections of data using keywords and filters based on metadata. It indexes metadata to provide fast and relevant search results while keeping the data updated and consistent.

Architecture Diagram
User
  |
  v
Load Balancer
  |
  v
API Gateway
  |
  v
Search Service <--> Metadata Index
  |
  v
Primary Database
  |
  v
Cache
Components
User
client
Initiates search requests
Load Balancer
load_balancer
Distributes incoming requests evenly to API Gateway instances
API Gateway
api_gateway
Handles client requests, routes them to Search Service
Search Service
service
Processes search queries, interacts with Metadata Index and Database
Metadata Index
search_index
Stores indexed metadata for fast search retrieval
Primary Database
database
Stores original data and metadata
Cache
cache
Stores frequently accessed search results to reduce latency
Request Flow - 12 Hops
UserLoad Balancer
Load BalancerAPI Gateway
API GatewayCache
CacheAPI Gateway
API GatewaySearch Service
Search ServiceMetadata Index
Metadata IndexSearch Service
Search ServicePrimary Database
Primary DatabaseSearch Service
Search ServiceCache
Search ServiceAPI Gateway
API GatewayUser
Failure Scenario
Component Fails:Metadata Index
Impact:Search queries become slower because the system must query the database directly without fast index lookups
Mitigation:Fallback to database queries with caching; rebuild or replicate the index asynchronously to restore fast search
Architecture Quiz - 3 Questions
Test your understanding
Which component is responsible for distributing incoming user requests evenly?
ALoad Balancer
BAPI Gateway
CSearch Service
DCache
Design Principle
This design uses a layered approach with caching and indexing to speed up search queries. The Metadata Index provides fast lookup of relevant data, while the cache reduces repeated work. The system gracefully falls back to slower database queries if the index fails, ensuring availability.

Practice

(1/5)
1. What is the primary purpose of metadata in a search system?
easy
A. To display images on the website
B. To store user passwords securely
C. To describe data and make search faster
D. To manage network connections

Solution

  1. Step 1: Understand metadata role

    Metadata provides information about data, like tags or descriptions.
  2. Step 2: Connect metadata to search

    Search engines use metadata to quickly find relevant data without scanning everything.
  3. Final Answer:

    To describe data and make search faster -> Option C
  4. Quick Check:

    Metadata = Data description for search [OK]
Hint: Metadata helps find data faster by describing it [OK]
Common Mistakes:
  • Confusing metadata with user data
  • Thinking metadata stores passwords
  • Assuming metadata manages network
2. Which of the following is a correct example of metadata used in search?
easy
A. title = 'Introduction to Cats'
B. file_size = 2048
C. user_password = '1234'
D. connection_timeout = 30

Solution

  1. Step 1: Identify metadata examples

    Metadata describes content, like titles, tags, or dates.
  2. Step 2: Check each option

    title = 'Introduction to Cats' shows a title, which is metadata describing content. Others are config or sensitive data.
  3. Final Answer:

    title = 'Introduction to Cats' -> Option A
  4. Quick Check:

    Title is metadata for search [OK]
Hint: Metadata describes content, not configs or passwords [OK]
Common Mistakes:
  • Choosing config values as metadata
  • Confusing sensitive data with metadata
  • Ignoring descriptive fields
3. Given a search system with metadata index, what is the expected output when searching for "apple" if metadata contains {"title": "apple pie", "tags": ["fruit", "dessert"]}?
medium
A. No results found
B. Returns item with title "apple pie"
C. Returns all items with tag "fruit" only
D. Returns items with tag "dessert" only

Solution

  1. Step 1: Understand search with metadata

    Search looks for matches in metadata fields like title and tags.
  2. Step 2: Check if "apple" matches metadata

    "apple" matches the title "apple pie", so the item is returned.
  3. Final Answer:

    Returns item with title "apple pie" -> Option B
  4. Quick Check:

    Search matches title containing "apple" [OK]
Hint: Search matches metadata fields containing query word [OK]
Common Mistakes:
  • Ignoring title field in search
  • Returning unrelated tags only
  • Assuming no results if exact match missing
4. A search system's metadata index is not returning expected results. Which issue below is most likely the cause?
medium
A. Database password is incorrect
B. User interface colors are dull
C. Network cables are unplugged
D. Metadata is not updated after data changes

Solution

  1. Step 1: Identify cause of search failure

    If metadata is stale, search index won't reflect latest data.
  2. Step 2: Evaluate options

    Only Metadata is not updated after data changes relates to metadata and search correctness; others are unrelated.
  3. Final Answer:

    Metadata is not updated after data changes -> Option D
  4. Quick Check:

    Stale metadata breaks search results [OK]
Hint: Keep metadata updated to ensure correct search [OK]
Common Mistakes:
  • Blaming UI or network for search logic errors
  • Ignoring metadata update process
  • Confusing unrelated system issues
5. You are designing a scalable search system for millions of users. Which approach best ensures fast search using metadata?
hard
A. Use distributed indexing with metadata shards and update indexes asynchronously
B. Store metadata in a centralized database and scan all records on each search
C. Keep metadata only on user devices and search locally
D. Disable metadata to reduce storage and search raw data only

Solution

  1. Step 1: Understand scalability needs

    Millions of users require fast, distributed search to avoid bottlenecks.
  2. Step 2: Evaluate options for scalability

    Use distributed indexing with metadata shards and update indexes asynchronously uses distributed indexing and async updates, which scales well and keeps search fast.
  3. Final Answer:

    Use distributed indexing with metadata shards and update indexes asynchronously -> Option A
  4. Quick Check:

    Distributed indexing + async updates = scalable search [OK]
Hint: Distribute metadata index and update asynchronously for scale [OK]
Common Mistakes:
  • Scanning all data centrally causes slow search
  • Relying on local device metadata limits scale
  • Disabling metadata removes search efficiency