Bird
Raised Fist0
HLDsystem_design~10 mins

Search and metadata in HLD - Scalability & System Analysis

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Scalability Analysis - Search and metadata
Growth Table: Search and Metadata System
ScaleUsersSearch Queries/SecondMetadata SizeSystem Changes
Small10010 QPSFew MBsSingle server, simple DB, no caching
Medium10,0001,000 QPSGBsDB indexing, caching layer, load balancer
Large1,000,00050,000 QPSTBsDistributed search engine, sharded DB, CDN for metadata
Very Large100,000,0005,000,000 QPSPetabytesMulti-region clusters, advanced sharding, heavy caching, streaming updates
First Bottleneck

At small to medium scale, the database query performance for metadata and search indexing breaks first. This is because search queries require fast lookups and metadata updates increase load. The DB struggles with high QPS and large index sizes.

Scaling Solutions
  • Horizontal scaling: Add more search nodes and metadata DB replicas to distribute load.
  • Caching: Use in-memory caches (e.g., Redis) for frequent metadata and search results.
  • Sharding: Partition metadata and search indexes by user or content to reduce single node load.
  • Distributed search engines: Use systems like Elasticsearch or Solr that scale horizontally.
  • CDN: Cache static metadata and search results closer to users to reduce backend load.
  • Load balancing: Distribute incoming search requests evenly across servers.
Back-of-Envelope Cost Analysis

Assuming 1M users generate 50K QPS search queries:

  • Each query ~1 KB data -> 50 MB/s bandwidth needed.
  • Metadata storage ~1 TB with indexes.
  • DB handles ~10K QPS per instance -> need ~5 DB replicas.
  • Search nodes handle ~5K QPS each -> need ~10 search nodes.
  • Cache memory ~100 GB to hold hot metadata and results.
Interview Tip

Start by clarifying scale and data types. Discuss bottlenecks in DB and search indexing. Propose caching and horizontal scaling. Mention sharding and distributed search engines. Always justify why each solution fits the bottleneck.

Self Check

Your database handles 1000 QPS. Traffic grows 10x to 10,000 QPS. What do you do first?

Answer: Add read replicas and implement caching to reduce DB load before scaling vertically or sharding.

Key Result
The database query performance for metadata and search indexing is the first bottleneck as traffic grows; adding caching, read replicas, and distributed search engines are key to scaling effectively.

Practice

(1/5)
1. What is the primary purpose of metadata in a search system?
easy
A. To display images on the website
B. To store user passwords securely
C. To describe data and make search faster
D. To manage network connections

Solution

  1. Step 1: Understand metadata role

    Metadata provides information about data, like tags or descriptions.
  2. Step 2: Connect metadata to search

    Search engines use metadata to quickly find relevant data without scanning everything.
  3. Final Answer:

    To describe data and make search faster -> Option C
  4. Quick Check:

    Metadata = Data description for search [OK]
Hint: Metadata helps find data faster by describing it [OK]
Common Mistakes:
  • Confusing metadata with user data
  • Thinking metadata stores passwords
  • Assuming metadata manages network
2. Which of the following is a correct example of metadata used in search?
easy
A. title = 'Introduction to Cats'
B. file_size = 2048
C. user_password = '1234'
D. connection_timeout = 30

Solution

  1. Step 1: Identify metadata examples

    Metadata describes content, like titles, tags, or dates.
  2. Step 2: Check each option

    title = 'Introduction to Cats' shows a title, which is metadata describing content. Others are config or sensitive data.
  3. Final Answer:

    title = 'Introduction to Cats' -> Option A
  4. Quick Check:

    Title is metadata for search [OK]
Hint: Metadata describes content, not configs or passwords [OK]
Common Mistakes:
  • Choosing config values as metadata
  • Confusing sensitive data with metadata
  • Ignoring descriptive fields
3. Given a search system with metadata index, what is the expected output when searching for "apple" if metadata contains {"title": "apple pie", "tags": ["fruit", "dessert"]}?
medium
A. No results found
B. Returns item with title "apple pie"
C. Returns all items with tag "fruit" only
D. Returns items with tag "dessert" only

Solution

  1. Step 1: Understand search with metadata

    Search looks for matches in metadata fields like title and tags.
  2. Step 2: Check if "apple" matches metadata

    "apple" matches the title "apple pie", so the item is returned.
  3. Final Answer:

    Returns item with title "apple pie" -> Option B
  4. Quick Check:

    Search matches title containing "apple" [OK]
Hint: Search matches metadata fields containing query word [OK]
Common Mistakes:
  • Ignoring title field in search
  • Returning unrelated tags only
  • Assuming no results if exact match missing
4. A search system's metadata index is not returning expected results. Which issue below is most likely the cause?
medium
A. Database password is incorrect
B. User interface colors are dull
C. Network cables are unplugged
D. Metadata is not updated after data changes

Solution

  1. Step 1: Identify cause of search failure

    If metadata is stale, search index won't reflect latest data.
  2. Step 2: Evaluate options

    Only Metadata is not updated after data changes relates to metadata and search correctness; others are unrelated.
  3. Final Answer:

    Metadata is not updated after data changes -> Option D
  4. Quick Check:

    Stale metadata breaks search results [OK]
Hint: Keep metadata updated to ensure correct search [OK]
Common Mistakes:
  • Blaming UI or network for search logic errors
  • Ignoring metadata update process
  • Confusing unrelated system issues
5. You are designing a scalable search system for millions of users. Which approach best ensures fast search using metadata?
hard
A. Use distributed indexing with metadata shards and update indexes asynchronously
B. Store metadata in a centralized database and scan all records on each search
C. Keep metadata only on user devices and search locally
D. Disable metadata to reduce storage and search raw data only

Solution

  1. Step 1: Understand scalability needs

    Millions of users require fast, distributed search to avoid bottlenecks.
  2. Step 2: Evaluate options for scalability

    Use distributed indexing with metadata shards and update indexes asynchronously uses distributed indexing and async updates, which scales well and keeps search fast.
  3. Final Answer:

    Use distributed indexing with metadata shards and update indexes asynchronously -> Option A
  4. Quick Check:

    Distributed indexing + async updates = scalable search [OK]
Hint: Distribute metadata index and update asynchronously for scale [OK]
Common Mistakes:
  • Scanning all data centrally causes slow search
  • Relying on local device metadata limits scale
  • Disabling metadata removes search efficiency