Bird
Raised Fist0
HLDsystem_design~12 mins

Design a key-value store in HLD - Architecture Diagram

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
System Overview - Design a key-value store

This system stores and retrieves data as key-value pairs. It must handle many requests quickly and keep data safe and consistent. The system should scale to support many users and large data volumes.

Architecture Diagram
User
  |
  v
Load Balancer
  |
  v
API Gateway
  |
  v
+-------------------+      +-------------+
|   Key-Value Store  |<---->|   Cache     |
|     Service       |      +-------------+
|  (Handles logic)  |
+-------------------+
          |
          v
    +-------------+
    |  Database   |
    +-------------+
Components
User
client
Sends requests to store or retrieve key-value data
Load Balancer
load_balancer
Distributes incoming requests evenly to API Gateway instances
API Gateway
api_gateway
Receives requests, handles authentication, and routes to Key-Value Store Service
Key-Value Store Service
service
Processes requests, checks cache, and interacts with the database
Cache
cache
Stores frequently accessed key-value pairs for fast retrieval
Database
database
Stores all key-value pairs persistently and reliably
Request Flow - 11 Hops
UserLoad Balancer
Load BalancerAPI Gateway
API GatewayKey-Value Store Service
Key-Value Store ServiceCache
CacheKey-Value Store Service
Key-Value Store ServiceDatabase
DatabaseKey-Value Store Service
Key-Value Store ServiceCache
Key-Value Store ServiceAPI Gateway
API GatewayLoad Balancer
Load BalancerUser
Failure Scenario
Component Fails:Database
Impact:Writes fail and new data cannot be saved; reads may serve stale data from cache
Mitigation:Use database replication and failover to backup; cache serves reads temporarily; alert system admins
Architecture Quiz - 3 Questions
Test your understanding
Which component handles distributing user requests to multiple API Gateway instances?
ACache
BDatabase
CLoad Balancer
DKey-Value Store Service
Design Principle
This design uses caching to speed up reads and reduce database load. The load balancer and API gateway ensure scalability and security. The database provides reliable storage. This layered approach balances speed, reliability, and scalability.

Practice

(1/5)
1. What is the primary purpose of a key-value store in system design?
easy
A. To perform complex relational queries
B. To store large binary files efficiently
C. To save data as pairs for quick lookup
D. To manage user authentication and sessions

Solution

  1. Step 1: Understand key-value store basics

    A key-value store saves data as pairs where each key maps to a value for fast retrieval.
  2. Step 2: Compare with other storage types

    Unlike relational databases, key-value stores do not support complex queries or file storage.
  3. Final Answer:

    To save data as pairs for quick lookup -> Option C
  4. Quick Check:

    Key-value store = data pairs [OK]
Hint: Key-value stores focus on pairs, not complex queries [OK]
Common Mistakes:
  • Confusing key-value store with relational database
  • Thinking it handles large files natively
  • Assuming it manages user sessions directly
2. Which of the following is the correct operation to add or update a value in a key-value store?
easy
A. exists(key)
B. put(key, value)
C. delete(key)
D. get(key)

Solution

  1. Step 1: Identify operation purpose

    Adding or updating a value requires an operation that sets the value for a key.
  2. Step 2: Match operation names

    "put" is commonly used to insert or update key-value pairs; "get" retrieves, "delete" removes, "exists" checks presence.
  3. Final Answer:

    put(key, value) -> Option B
  4. Quick Check:

    Put = add/update [OK]
Hint: Put means add or update a key-value pair [OK]
Common Mistakes:
  • Using get to add data
  • Confusing delete with update
  • Using exists to insert values
3. Given this pseudo-code for a key-value store:
store = {}
store.put('a', 1)
store.put('b', 2)
store.put('a', 3)
value = store.get('a')
What is the value of value after these operations?
medium
A. 3
B. 2
C. 1
D. None

Solution

  1. Step 1: Track put operations

    First, key 'a' is set to 1, then 'b' to 2, then 'a' is updated to 3, overwriting previous value.
  2. Step 2: Retrieve the value for 'a'

    The last value assigned to 'a' is 3, so store.get('a') returns 3.
  3. Final Answer:

    3 -> Option A
  4. Quick Check:

    Last put for 'a' = 3 [OK]
Hint: Last put for a key overwrites previous value [OK]
Common Mistakes:
  • Assuming first value stays after update
  • Confusing keys 'a' and 'b'
  • Thinking get returns None if key exists
4. Consider this code snippet for a key-value store:
store = {}
def get_value(key):
    if key in store:
        return store[key]
    else:
        return None

store.put('x', 10)
print(get_value('x'))
What is the main issue preventing this code from working correctly?
medium
A. The put method is not defined for the dictionary
B. The get_value function returns None incorrectly
C. The key 'x' is not added to the store
D. The print statement syntax is wrong

Solution

  1. Step 1: Check dictionary operations

    Python dictionaries do not have a put method; they use assignment like store[key] = value.
  2. Step 2: Identify error cause

    Calling store.put('x', 10) will cause an AttributeError because put is undefined.
  3. Final Answer:

    The put method is not defined for the dictionary -> Option A
  4. Quick Check:

    Dicts use assignment, not put [OK]
Hint: Dictionaries use assignment, not put() method [OK]
Common Mistakes:
  • Assuming put exists on dict
  • Ignoring error from undefined method
  • Thinking get_value logic is faulty
5. You want to design a scalable key-value store that handles millions of requests per second. Which design choice best supports this goal?
hard
A. Use a single in-memory dictionary on one server
B. Store all data on a single disk-based database
C. Use a relational database with complex joins
D. Partition data across multiple servers using consistent hashing

Solution

  1. Step 1: Understand scalability needs

    Handling millions of requests requires distributing load and data to avoid bottlenecks.
  2. Step 2: Evaluate design options

    A single in-memory dictionary or disk-based DB limits capacity; relational DB with joins is slow for key-value access. Consistent hashing partitions data evenly across servers, enabling horizontal scaling.
  3. Final Answer:

    Partition data across multiple servers using consistent hashing -> Option D
  4. Quick Check:

    Consistent hashing = scalable partitioning [OK]
Hint: Distribute data with consistent hashing for scalability [OK]
Common Mistakes:
  • Relying on single server limits throughput
  • Using disk-based DB slows access
  • Choosing relational DB for simple key-value