Bird
Raised Fist0
HLDsystem_design~10 mins

Notification system design in HLD - Scalability & System Analysis

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Scalability Analysis - Notification system design
Growth Table: Notification System
ScaleUsersNotifications/DayKey Changes
Small1001,000Single server handles API and DB; simple queue; direct push
Medium10,000100,000Introduce message queue; DB read replicas; caching user preferences
Large1,000,00010,000,000Multiple app servers; sharded DB; distributed queue; push notification services
Very Large100,000,0001,000,000,000Global load balancers; multi-region DB shards; CDN for static content; advanced throttling
First Bottleneck

At small to medium scale, the database is the first bottleneck. It struggles with high write volume for notifications and user preferences. As traffic grows, the message queue and application servers also become bottlenecks due to processing and delivery delays.

Scaling Solutions
  • Database: Use read replicas for reads, write sharding by user ID, and caching for user settings.
  • Application Servers: Horizontally scale with load balancers to handle more concurrent connections.
  • Message Queue: Use distributed queues like Kafka or RabbitMQ to handle high throughput and ensure reliable delivery.
  • Push Delivery: Integrate with platform push services (APNs, FCM) and use CDN for static notification content.
  • Throttling & Batching: Batch notifications and throttle to avoid overwhelming users and systems.
Back-of-Envelope Cost Analysis
  • Requests per second: At 1M users sending 10 notifications/day, ~115 QPS (10M/86400s).
  • Storage: Assuming 1KB per notification, 10M notifications/day = ~10GB/day storage.
  • Bandwidth: Push payloads are small (~1KB), so 10M notifications ~10GB outbound daily.
  • Server capacity: One app server handles ~3000 concurrent connections; scale horizontally as users grow.
  • Database QPS: One PostgreSQL instance handles ~5000 QPS; use sharding and replicas beyond that.
Interview Tip

Start by clarifying notification types and delivery guarantees. Discuss user scale and traffic patterns. Identify bottlenecks step-by-step: database, queue, delivery. Propose incremental scaling solutions with clear reasoning. Mention trade-offs like latency vs cost and user experience.

Self Check

Your database handles 1000 QPS. Traffic grows 10x. What do you do first?

Answer: Add read replicas to offload read queries and implement caching for frequent reads. For writes, consider sharding or batching writes to reduce load.

Key Result
The database is the first bottleneck as notification volume grows; scaling requires read replicas, sharding, distributed queues, and horizontal app server scaling.

Practice

(1/5)
1. Which component in a notification system is primarily responsible for storing user preferences about how they want to receive notifications?
easy
A. User Management Service
B. Notification Delivery Service
C. Notification Queue
D. Notification Generator

Solution

  1. Step 1: Understand user preferences role

    User preferences define how users want to receive notifications (email, SMS, push).
  2. Step 2: Identify responsible component

    The User Management Service stores and manages user data including preferences.
  3. Final Answer:

    User Management Service -> Option A
  4. Quick Check:

    User preferences stored in User Management Service [OK]
Hint: User preferences belong to user data, so User Management Service [OK]
Common Mistakes:
  • Confusing Notification Queue as storage for preferences
  • Thinking Notification Delivery Service stores preferences
  • Assuming Notification Generator manages user data
2. Which of the following is the correct sequence of components involved in sending a notification from creation to delivery?
easy
A. User Management Service -> Notification Delivery Service -> Notification Generator
B. Notification Delivery Service -> Notification Queue -> Notification Generator
C. Notification Queue -> Notification Generator -> Notification Delivery Service
D. Notification Generator -> Notification Queue -> Notification Delivery Service

Solution

  1. Step 1: Understand notification flow

    Notifications are created, queued, then delivered.
  2. Step 2: Match correct order

    Notification Generator creates, Notification Queue holds, Delivery Service sends.
  3. Final Answer:

    Notification Generator -> Notification Queue -> Notification Delivery Service -> Option D
  4. Quick Check:

    Creation, queue, delivery order = A [OK]
Hint: Notifications flow: create, queue, then deliver [OK]
Common Mistakes:
  • Mixing delivery before queuing
  • Starting with delivery service instead of generator
  • Ignoring the queue component
3. Consider this simplified flow: A notification is created and placed in a queue. The delivery service fetches notifications from the queue and sends them. If the delivery service crashes after fetching but before sending, what happens to the notification?
medium
A. Notification is lost and never sent
B. Notification is duplicated and sent twice
C. Notification remains in the queue for retry
D. Notification is sent immediately by the generator

Solution

  1. Step 1: Analyze delivery service crash timing

    Crash occurs after fetching from queue but before sending notification.
  2. Step 2: Understand queue behavior with acknowledgment

    Without acknowledgment, message stays or returns to queue for retry.
  3. Final Answer:

    Notification remains in the queue for retry -> Option C
  4. Quick Check:

    Unacknowledged messages stay in queue [OK]
Hint: Unsent messages stay in queue until confirmed sent [OK]
Common Mistakes:
  • Assuming notification is lost without retry
  • Thinking notification is duplicated automatically
  • Believing generator sends notification directly
4. A notification system is experiencing delays because the delivery service processes notifications sequentially. Which change will best improve throughput without losing message order?
medium
A. Store notifications only in the database without queue
B. Add multiple delivery service instances with partitioned queues
C. Use a single-threaded delivery service with longer timeouts
D. Remove the queue and send notifications directly

Solution

  1. Step 1: Identify bottleneck cause

    Sequential processing limits throughput.
  2. Step 2: Apply partitioned queues with multiple instances

    Partitioning allows parallel processing while preserving order per partition.
  3. Final Answer:

    Add multiple delivery service instances with partitioned queues -> Option B
  4. Quick Check:

    Parallelism with partitioning improves throughput [OK]
Hint: Partition queues to parallelize delivery without breaking order [OK]
Common Mistakes:
  • Removing queue loses reliability and order
  • Single-threaded service slows throughput
  • Storing only in DB delays delivery
5. You need to design a notification system that supports millions of users with different notification preferences and guarantees delivery within seconds. Which architectural approach best meets these requirements?
hard
A. Implement microservices with separate components for user preferences, notification generation, queuing, and delivery with horizontal scaling
B. Use a monolithic service handling all notifications synchronously
C. Store all notifications in a single database table and poll it every minute for delivery
D. Send notifications directly from the user management service without queues

Solution

  1. Step 1: Analyze scalability and latency needs

    Millions of users and seconds-level delivery require scalable, decoupled design.
  2. Step 2: Choose microservices with separate components and horizontal scaling

    This approach allows independent scaling, fault isolation, and faster processing.
  3. Final Answer:

    Implement microservices with separate components for user preferences, notification generation, queuing, and delivery with horizontal scaling -> Option A
  4. Quick Check:

    Microservices + scaling = scalable, fast delivery [OK]
Hint: Decouple components and scale horizontally for millions of users [OK]
Common Mistakes:
  • Using monolith limits scalability and speed
  • Polling DB every minute causes delays
  • Skipping queues reduces reliability