| Users | Messages/Day | Active Groups | Storage Size | Server Load | Network Traffic |
|---|---|---|---|---|---|
| 100 | 10K | 50 | ~100 MB | 1 app server | Low |
| 10,000 | 1M | 5,000 | ~10 GB | 3-5 app servers | Moderate |
| 1,000,000 | 100M | 500,000 | ~1 TB | 50+ app servers, DB cluster | High |
| 100,000,000 | 10B | 50M | ~100+ TB | Hundreds of servers, sharded DB | Very High |
Group messaging in HLD - Scalability & System Analysis
Start learning this pattern below
Jump into concepts and practice - no test required
At small scale (up to 10K users), the database write throughput is the first bottleneck because every message must be stored reliably. The database can handle around 5,000-10,000 writes per second, so as message volume grows, it will slow down.
At medium scale (100K+ users), application servers CPU and memory become bottlenecks due to message fan-out (delivering messages to many group members).
At large scale (millions of users), network bandwidth and storage size become bottlenecks, requiring data partitioning and efficient delivery mechanisms.
- Database scaling: Use read replicas for reads, write sharding by group ID to distribute writes.
- Caching: Cache recent messages per group in Redis to reduce DB reads.
- Horizontal scaling: Add more app servers behind load balancers to handle concurrent connections and message fan-out.
- Message queue: Use message brokers (e.g., Kafka) to decouple message ingestion and delivery.
- CDN and push notifications: Use CDN for media content and push notifications for offline users.
- Data archiving: Archive old messages to cheaper storage to reduce DB size.
Assuming 1M users sending 100 messages/day:
- Messages per second (QPS): ~1,000,000 users * 100 messages / 86400 seconds ≈ 1157 QPS
- Storage: 100 bytes per message * 100M messages/day = ~10 GB/day
- Network bandwidth: Assuming 1 KB per message delivered to 10 recipients on average = 1157 QPS * 1 KB * 10 = ~11.57 MB/s (~92 Mbps)
- App servers: Each server handles ~2000 concurrent connections and message fan-out; need ~10-20 servers
- Database: Must support ~1200 writes/sec and higher reads; use sharding and replicas
Start by defining key metrics: users, messages per user, group size. Then identify bottlenecks step-by-step: database writes, message delivery, storage. Discuss scaling strategies for each bottleneck clearly. Use real numbers to justify your choices. Always mention trade-offs and fallback plans.
Question: Your database handles 1000 QPS. Traffic grows 10x. What do you do first?
Answer: The first step is to add read replicas to offload read traffic and implement write sharding by group ID to distribute write load across multiple database instances. This prevents the single DB from becoming a bottleneck.
Practice
Solution
Step 1: Understand group messaging basics
Group messaging connects multiple users so they can communicate together in one conversation.Step 2: Identify the main function
The main function is sending and receiving messages among group members, not unrelated features like password storage or video streaming.Final Answer:
To allow multiple users to send and receive messages in a shared conversation -> Option AQuick Check:
Group messaging = shared conversation [OK]
- Confusing group messaging with unrelated features
- Thinking it only supports one-to-one chat
- Ignoring the shared conversation aspect
Solution
Step 1: Identify components related to user membership
Group management handles adding, removing, and listing members in a group.Step 2: Exclude unrelated components
Message storage saves messages, notification service alerts users, and media transcoding processes media, none manage group membership.Final Answer:
Group management -> Option AQuick Check:
Group membership = Group management [OK]
- Confusing message storage with membership control
- Assuming notification service manages members
- Mixing media processing with group functions
Solution
Step 1: Understand message delivery in group messaging
Each message is delivered to every member of the group.Step 2: Calculate total deliveries
With 100 members, one message results in 100 deliveries (one per member).Final Answer:
100 -> Option DQuick Check:
Deliveries = group size = 100 [OK]
- Counting the sender as extra delivery
- Assuming half the group receives the message
- Confusing message count with delivery count
Solution
Step 1: Identify components involved in message delivery
Notification service alerts users about new messages; if slow, users get delayed messages.Step 2: Exclude unrelated causes
Message storage speed does not cause delay in delivery; group management and profile pictures do not affect message delivery timing.Final Answer:
Notification service is slow or failing -> Option BQuick Check:
Delivery delay = notification issue [OK]
- Blaming message storage speed
- Thinking group size causes delay directly
- Ignoring notification service role
Solution
Step 1: Understand scalability challenges
Direct synchronous sending to many users overloads servers and causes delays.Step 2: Identify scalable solution
Using message queues allows asynchronous, reliable, and scalable message distribution without blocking sender or servers.Step 3: Exclude impractical options
Storing messages only on sender device or limiting group size reduces usability and scalability.Final Answer:
Use a message queue to asynchronously distribute messages to group members -> Option CQuick Check:
Scalable delivery = asynchronous queue [OK]
- Trying synchronous delivery to all users
- Ignoring asynchronous processing benefits
- Reducing group size instead of scaling design
