Bird
Raised Fist0
HLDsystem_design~10 mins

Online presence system in HLD - Scalability & System Analysis

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Scalability Analysis - Online presence system
Growth Table: Online Presence System
ScaleUsersActive ConnectionsData StoredTraffic CharacteristicsSystem Changes
Small100~100 concurrentMBs (presence states)Low, few updates per secondSingle server, simple DB, no caching
Medium10,000~5,000 concurrentGBs (presence logs, user states)Moderate, frequent presence updatesLoad balancer, DB replicas, caching layer
Large1,000,000~500,000 concurrentTBs (history, analytics)High, real-time updates, many events/secHorizontal scaling, sharding, pub/sub messaging
Very Large100,000,000~50,000,000 concurrentPetabytes (long-term storage)Very high, global distribution, multi-regionGlobal CDN, geo-sharding, multi-region DB, edge caching
First Bottleneck

At small scale, the database is the first bottleneck because it must handle many frequent presence updates and queries. As users grow, the DB write throughput and connection limits are stressed first.

Scaling Solutions
  • Database scaling: Use read replicas to offload reads, connection pooling, and write sharding to distribute load.
  • Caching: Cache presence states in fast in-memory stores like Redis to reduce DB hits.
  • Horizontal scaling: Add more application servers behind load balancers to handle concurrent connections.
  • Messaging: Use pub/sub systems (e.g., Kafka, Redis Streams) to propagate presence updates efficiently.
  • Global distribution: Geo-shard data and use CDNs or edge caches to reduce latency for worldwide users.
Back-of-Envelope Cost Analysis
  • Requests per second: For 1M users with 50% active, ~500K concurrent connections, each sending updates every 10 seconds -> ~50K QPS.
  • Storage: Presence states are small (~100 bytes per user), but history logs grow fast. For 1M users, daily logs ~10GB; yearly ~3.6TB.
  • Bandwidth: Each update ~100 bytes, 50K QPS -> ~5MB/s (~40Mbps), manageable with 1Gbps network.
Interview Tip

Start by defining the scale and key metrics (users, connections, update frequency). Identify the first bottleneck (usually DB). Then discuss scaling strategies step-by-step: caching, read replicas, horizontal scaling, messaging, and global distribution. Always justify why each solution fits the bottleneck.

Self Check

Your database handles 1000 QPS. Traffic grows 10x to 10,000 QPS. What do you do first?

Answer: Add read replicas and implement caching to reduce direct DB load before considering sharding or adding more servers.

Key Result
The database is the first bottleneck as user count and presence updates grow; scaling requires caching, read replicas, and horizontal scaling of app servers to handle real-time presence efficiently.

Practice

(1/5)
1. What is the primary purpose of an online presence system in a chat application?
easy
A. To track if users are currently online or offline
B. To store chat message history permanently
C. To encrypt messages between users
D. To manage user account passwords

Solution

  1. Step 1: Understand the role of presence system

    An online presence system tracks user activity to know if they are online or offline.
  2. Step 2: Differentiate from other chat features

    Message storage, encryption, and password management are separate features not handled by presence systems.
  3. Final Answer:

    To track if users are currently online or offline -> Option A
  4. Quick Check:

    Presence system = user online status [OK]
Hint: Presence means tracking user online/offline status [OK]
Common Mistakes:
  • Confusing presence with message storage
  • Thinking presence handles security features
  • Mixing presence with account management
2. Which event is NOT typically used in an online presence system to track user status?
easy
A. message_send
B. connect
C. disconnect
D. heartbeat

Solution

  1. Step 1: Identify presence tracking events

    Presence systems use connect, heartbeat, and disconnect to track user activity.
  2. Step 2: Recognize unrelated events

    message_send relates to sending chat messages, not presence tracking.
  3. Final Answer:

    message_send -> Option A
  4. Quick Check:

    Presence events exclude message sending [OK]
Hint: Presence tracks connection, not message sending [OK]
Common Mistakes:
  • Assuming message events track presence
  • Confusing heartbeat with message_send
  • Ignoring disconnect event importance
3. Given this simplified presence update code snippet, what will be the user's status after 10 seconds?
user_last_seen = 0
current_time = 10
heartbeat_interval = 5
if current_time - user_last_seen <= heartbeat_interval:
    status = 'online'
else:
    status = 'offline'
medium
A. Error due to comparison
B. "online"
C. "offline"
D. "unknown"

Solution

  1. Step 1: Calculate time difference

    current_time - user_last_seen = 10 - 0 = 10 seconds.
  2. Step 2: Compare with heartbeat interval

    10 <= 5 is false, so status is set to 'offline'.
  3. Final Answer:

    "offline" -> Option C
  4. Quick Check:

    Time diff > heartbeat means offline [OK]
Hint: Compare last seen time difference with heartbeat [OK]
Common Mistakes:
  • Mixing up less than and greater than
  • Forgetting to subtract last seen time
  • Assuming status is always online
4. Identify the bug in this presence update logic:
def update_status(last_seen, current_time, heartbeat=10):
    if current_time - last_seen > heartbeat:
        return 'online'
    else:
        return 'offline'
medium
A. Function should return a boolean, not string
B. The comparison sign should be reversed
C. Heartbeat value should be negative
D. No bug, logic is correct

Solution

  1. Step 1: Analyze the condition meaning

    If time difference is greater than heartbeat, user should be offline, not online.
  2. Step 2: Correct the comparison

    The condition should return 'offline' when difference > heartbeat, so the comparison sign must be reversed.
  3. Final Answer:

    The comparison sign should be reversed -> Option B
  4. Quick Check:

    Time diff > heartbeat means offline [OK]
Hint: Online if time diff ≤ heartbeat, else offline [OK]
Common Mistakes:
  • Returning online when user is inactive
  • Misunderstanding heartbeat meaning
  • Ignoring return type correctness
5. You need to design an online presence system for a messaging app with millions of users. Which approach best ensures scalability and real-time accuracy?
hard
A. Store presence data in user profile tables updated once per day
B. Store user status in a centralized database and query it on every client request
C. Send presence updates only when users log in or log out, ignoring heartbeats
D. Use distributed in-memory cache with TTL and heartbeat updates from clients

Solution

  1. Step 1: Evaluate centralized database approach

    Querying a centralized DB for every request causes high latency and bottlenecks at scale.
  2. Step 2: Consider distributed cache with TTL and heartbeats

    This approach keeps presence data fresh, reduces DB load, and supports real-time updates efficiently.
  3. Step 3: Analyze other options

    Ignoring heartbeats or updating once per day leads to stale presence info, unsuitable for real-time apps.
  4. Final Answer:

    Use distributed in-memory cache with TTL and heartbeat updates from clients -> Option D
  5. Quick Check:

    Distributed cache + heartbeat = scalable real-time presence [OK]
Hint: Use cache with TTL and heartbeats for scalable presence [OK]
Common Mistakes:
  • Relying on centralized DB for real-time presence
  • Ignoring heartbeat updates causing stale data
  • Updating presence too infrequently