Bird
Raised Fist0
HLDsystem_design~10 mins

Design a notification system in HLD - Scalability & System Analysis

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Scalability Analysis - Design a notification system
Growth Table: Notification System Scaling
UsersNotifications/DayKey Changes
100~1,000Simple queue, single server, direct DB writes
10,000~100,000Message queue introduced, caching user preferences, DB indexing
1,000,000~10,000,000Multiple app servers, distributed queue, read replicas, push notification services
100,000,000~1,000,000,000Sharded DB, global CDN for media, microservices, event-driven architecture
First Bottleneck

At around 10,000 users, the database becomes the first bottleneck. Writing and reading notification data for many users causes high latency and connection limits. The single server and simple queue cannot handle the volume efficiently.

Scaling Solutions
  • Horizontal Scaling: Add more application servers behind a load balancer to handle more notification requests.
  • Message Queues: Use distributed queues (e.g., Kafka, RabbitMQ) to decouple notification generation from delivery.
  • Caching: Cache user notification preferences and recent notifications to reduce DB load.
  • Database Read Replicas: Use replicas to distribute read traffic and reduce load on the primary DB.
  • Sharding: Partition the database by user ID or region to scale writes and storage.
  • Push Notification Services: Use external services (e.g., Firebase, APNs) for mobile push notifications to offload delivery.
  • CDN: Use CDN for static media in notifications to reduce bandwidth and latency.
Back-of-Envelope Cost Analysis
  • At 1M users sending 10 notifications/day: ~10M notifications/day ≈ 115 notifications/sec.
  • Database: Needs to handle ~115 writes/sec plus reads; a single DB can handle ~5,000 QPS, so one instance is sufficient but close to limits.
  • Message Queue: Must support ~115 enqueue/dequeue operations per second, well within Kafka or RabbitMQ capabilities.
  • Bandwidth: Assuming 1 KB per notification, 115 KB/s ≈ 0.9 Mbps, easily handled by 1 Gbps network.
  • Storage: 10M notifications/day x 1 KB = ~10 GB/day; plan for archiving and tiered storage.
Interview Tip

Start by clarifying notification types and user scale. Discuss data flow from event to delivery. Identify bottlenecks at each scale. Propose incremental scaling solutions: caching, queues, DB replicas, sharding. Mention trade-offs and real-world constraints like latency and cost.

Self Check

Your database handles 1000 QPS. Traffic grows 10x to 10,000 QPS. What do you do first?

Answer: Add read replicas to distribute read traffic and reduce load on the primary database. Also, introduce caching for frequent reads and consider message queues to decouple processing.

Key Result
The database is the first bottleneck as user notifications grow; scaling requires adding read replicas, caching, and distributed queues before sharding and microservices.

Practice

(1/5)
1. Which component in a notification system is responsible for deciding when to send a message to a user?
easy
A. Event processor
B. Notification sender
C. User preference manager
D. Message storage

Solution

  1. Step 1: Understand the role of event processor

    The event processor detects events that trigger notifications, deciding when a message should be sent.
  2. Step 2: Differentiate from other components

    The notification sender delivers messages, user preference manager stores user choices, and message storage keeps records.
  3. Final Answer:

    Event processor -> Option A
  4. Quick Check:

    Event detection = Event processor [OK]
Hint: Event timing is handled by the event processor [OK]
Common Mistakes:
  • Confusing sender with event detector
  • Thinking user preferences trigger events
  • Assuming storage decides timing
2. Which data structure is best suited to store user notification preferences for quick lookup?
easy
A. Linked list
B. Hash map
C. Queue
D. Stack

Solution

  1. Step 1: Identify quick lookup needs

    User preferences require fast access by user ID or key, so a data structure with O(1) average lookup is ideal.
  2. Step 2: Match data structures to lookup speed

    Hash maps provide constant time lookup, unlike linked lists, queues, or stacks which are slower for direct access.
  3. Final Answer:

    Hash map -> Option B
  4. Quick Check:

    Fast key-value access = Hash map [OK]
Hint: Use hash map for fast user preference lookup [OK]
Common Mistakes:
  • Choosing linked list which is slow for lookup
  • Confusing queue or stack with lookup structures
  • Ignoring key-based access needs
3. Consider this simplified flow: An event triggers a notification, which is stored in a queue before delivery. What happens if the queue is full?
medium
A. Queue automatically expands without limit
B. System crashes due to overflow
C. Notifications are sent immediately bypassing the queue
D. New notifications are dropped or delayed

Solution

  1. Step 1: Understand queue capacity limits

    Queues have fixed or limited size; when full, they cannot accept new items immediately.
  2. Step 2: Identify common handling of full queues

    Systems usually drop new notifications or delay them until space frees up; automatic unlimited expansion is rare to avoid resource exhaustion.
  3. Final Answer:

    New notifications are dropped or delayed -> Option D
  4. Quick Check:

    Full queue = drop or delay new notifications [OK]
Hint: Full queue means drop or delay notifications [OK]
Common Mistakes:
  • Assuming infinite queue size
  • Thinking notifications bypass queue
  • Believing system crashes on full queue
4. A notification system sends duplicate messages to users. Which design mistake most likely causes this?
medium
A. Notifications sent synchronously
B. User preferences not stored
C. No deduplication in event processing
D. Using a single message queue

Solution

  1. Step 1: Analyze duplicate message causes

    Duplicates often occur if the system processes the same event multiple times without checking if notification was already sent.
  2. Step 2: Evaluate other options

    Missing user preferences or synchronous sending do not cause duplicates; a single queue can still handle duplicates if deduplication exists.
  3. Final Answer:

    No deduplication in event processing -> Option C
  4. Quick Check:

    Duplicates = missing deduplication [OK]
Hint: Duplicates mean missing deduplication step [OK]
Common Mistakes:
  • Blaming user preferences for duplicates
  • Confusing synchronous sending with duplication
  • Assuming single queue causes duplicates
5. You need to design a notification system that supports email, SMS, and push notifications with user preferences and high scalability. Which architecture pattern best fits this requirement?
hard
A. Event-driven microservices with message queues and preference service
B. Batch processing system sending notifications once daily
C. Single database polling for notifications every minute
D. Monolithic application with direct notification calls

Solution

  1. Step 1: Identify scalability and multi-channel needs

    Supporting multiple notification types and scaling requires decoupling components and asynchronous processing.
  2. Step 2: Match architecture patterns

    Event-driven microservices with message queues allow independent scaling, handle user preferences, and support multiple channels efficiently.
  3. Step 3: Eliminate unsuitable options

    Monolithic apps limit scalability; polling causes delays; batch processing is too slow for timely notifications.
  4. Final Answer:

    Event-driven microservices with message queues and preference service -> Option A
  5. Quick Check:

    Scalable multi-channel = event-driven microservices [OK]
Hint: Use event-driven microservices for scalable multi-channel notifications [OK]
Common Mistakes:
  • Choosing monolithic for scalability
  • Using batch processing for real-time needs
  • Relying on polling causing delays