Bird
Raised Fist0
HLDsystem_design~10 mins

Why messaging requires real-time architecture in HLD - Scalability Evidence

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Scalability Analysis - Why messaging requires real-time architecture
Growth Table: Messaging System Scaling
UsersMessages/SecondLatency RequirementInfrastructure ChangesChallenges
100 users~10-50< 1 secondSingle server, simple DBMinimal load, simple queue
10,000 users~1,000-5,000< 500 msLoad balancer, message broker, cachingHandling concurrent connections, DB load
1,000,000 users~1,000,000+< 200 msHorizontal scaling, sharded DB, distributed brokersNetwork bandwidth, message ordering, fault tolerance
100,000,000 users~10,000,000+< 100 msGlobal CDN, multi-region clusters, advanced partitioningLatency consistency, data replication, disaster recovery
First Bottleneck: Real-Time Message Delivery

The first bottleneck in scaling messaging systems is the real-time message delivery component. This includes the message broker and network connections that must handle many concurrent users sending and receiving messages instantly.

As user count grows, the system struggles to maintain low latency and message ordering. The database can also become a bottleneck if it is used synchronously for message storage or delivery confirmation.

Scaling Solutions for Real-Time Messaging
  • Horizontal Scaling: Add more message broker instances and application servers behind load balancers to distribute user connections.
  • Message Brokers: Use specialized real-time brokers (e.g., Kafka, RabbitMQ, or MQTT) that support high throughput and low latency.
  • Caching: Use in-memory caches (e.g., Redis) for quick message state and presence info to reduce DB hits.
  • Sharding: Partition user data and message streams by user ID or region to reduce contention and improve parallelism.
  • CDN & Edge Computing: For global scale, use edge servers to reduce latency by bringing message routing closer to users.
  • Asynchronous Processing: Decouple message storage from delivery using queues to avoid blocking operations.
Back-of-Envelope Cost Analysis
  • At 1M users sending 1 message per second: ~1M messages/sec throughput needed.
  • Each message ~1 KB -> 1 GB/s bandwidth needed just for messages.
  • Database writes can be optimized by batching or async writes to handle ~100K QPS per instance.
  • Network bandwidth and CPU on brokers become expensive; multiple instances needed.
  • Storage grows rapidly; archiving old messages is necessary to control costs.
Interview Tip: Structuring Scalability Discussion

Start by explaining the real-time nature of messaging and why low latency is critical.

Discuss how user growth increases concurrent connections and message throughput.

Identify the first bottleneck (message delivery and broker capacity).

Propose scaling solutions step-by-step: horizontal scaling, caching, sharding, and CDN.

Include cost and complexity trade-offs to show balanced understanding.

Self Check Question

Your message broker handles 1,000 messages per second. Traffic grows 10x. What do you do first and why?

Answer: Add more broker instances and implement load balancing to distribute the increased message load, ensuring low latency and avoiding message loss.

Key Result
Messaging systems require real-time architecture because as users grow, the message broker and network become the first bottlenecks due to the need for low latency and high throughput. Scaling involves horizontal scaling of brokers, caching, sharding, and edge delivery to maintain performance.

Practice

(1/5)
1. Why is real-time architecture important for messaging systems?
easy
A. It delivers messages instantly for smooth conversations.
B. It stores messages for long-term archival only.
C. It processes messages in batches once a day.
D. It encrypts messages without sending them.

Solution

  1. Step 1: Understand messaging user experience

    Users expect messages to appear immediately during chats for natural flow.
  2. Step 2: Connect real-time architecture to instant delivery

    Real-time systems use open connections to send messages instantly without delay.
  3. Final Answer:

    It delivers messages instantly for smooth conversations. -> Option A
  4. Quick Check:

    Instant delivery = Real-time architecture [OK]
Hint: Real-time means instant message delivery [OK]
Common Mistakes:
  • Confusing real-time with batch processing
  • Thinking real-time only means encryption
  • Assuming real-time stores messages only
2. Which technology is commonly used to maintain open connections for real-time messaging?
easy
A. WebSockets for continuous connection
B. FTP for file transfers
C. HTTP polling every hour
D. SMTP for email delivery

Solution

  1. Step 1: Identify open connection methods

    Real-time messaging needs a persistent connection to avoid delays.
  2. Step 2: Match technology to persistent connection

    WebSockets keep a continuous open connection, unlike HTTP polling or FTP.
  3. Final Answer:

    WebSockets for continuous connection -> Option A
  4. Quick Check:

    Open connection = WebSockets [OK]
Hint: WebSockets keep connections open for real-time [OK]
Common Mistakes:
  • Choosing HTTP polling which is slow
  • Confusing FTP or SMTP with messaging protocols
  • Thinking email protocols support real-time chat
3. Consider a messaging system using WebSockets. What happens if the server delays message delivery by 5 seconds?
medium
A. Messages are delivered in batches every minute.
B. Users see messages instantly without delay.
C. Messages get lost and never arrive.
D. Messages appear late, causing poor chat experience.

Solution

  1. Step 1: Understand WebSocket behavior

    WebSockets deliver messages instantly if server sends immediately.
  2. Step 2: Analyze impact of 5-second delay

    If server delays sending, users see messages late, hurting chat flow.
  3. Final Answer:

    Messages appear late, causing poor chat experience. -> Option D
  4. Quick Check:

    Server delay = Late messages [OK]
Hint: Server delay causes late message display [OK]
Common Mistakes:
  • Assuming WebSockets fix server delays
  • Thinking messages get lost due to delay
  • Confusing batch delivery with real-time
4. A messaging system uses WebSockets but users report delayed messages. What is the most likely cause?
medium
A. WebSocket protocol does not support real-time.
B. Server processes messages slowly before sending.
C. Clients do not support HTML5.
D. Messages are encrypted end-to-end.

Solution

  1. Step 1: Check WebSocket capabilities

    WebSockets support real-time delivery if server sends promptly.
  2. Step 2: Identify delay source

    Slow server processing before sending causes message delays, not WebSocket itself.
  3. Final Answer:

    Server processes messages slowly before sending. -> Option B
  4. Quick Check:

    Slow server = Delayed messages [OK]
Hint: Delays usually come from server processing, not WebSocket [OK]
Common Mistakes:
  • Blaming WebSocket protocol for delays
  • Assuming client HTML5 support causes delay
  • Confusing encryption with delivery speed
5. You design a messaging app for millions of users. Which real-time architecture choice best supports fast, reliable message delivery at scale?
hard
A. Store messages locally on user devices only.
B. Send messages via email every hour.
C. Use WebSockets with load balancers and message queues.
D. Use HTTP polling every 10 minutes.

Solution

  1. Step 1: Identify scalable real-time components

    WebSockets provide instant delivery; load balancers distribute traffic; queues ensure reliability.
  2. Step 2: Compare other options for scale and speed

    Email and polling are slow; local storage alone can't deliver messages to others.
  3. Final Answer:

    Use WebSockets with load balancers and message queues. -> Option C
  4. Quick Check:

    Scalable real-time = WebSockets + load balancers + queues [OK]
Hint: Combine WebSockets, load balancers, and queues for scale [OK]
Common Mistakes:
  • Choosing slow batch methods for real-time needs
  • Ignoring load balancing for millions of users
  • Relying only on local device storage