Bird
Raised Fist0
HLDsystem_design~25 mins

Heartbeat mechanism in HLD - System Design Exercise

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Design: Heartbeat Mechanism System
Design covers heartbeat sending, receiving, monitoring, and alerting. Out of scope are client implementation details and dashboard UI design specifics.
Functional Requirements
FR1: Detect if a client or server is alive by sending periodic heartbeat signals
FR2: Support up to 10,000 concurrent clients sending heartbeats
FR3: Trigger alerts if heartbeat is missed for more than 30 seconds
FR4: Provide a dashboard to show live status of all clients
FR5: Ensure minimal network overhead for heartbeat messages
Non-Functional Requirements
NFR1: Heartbeat interval must be configurable but default to 10 seconds
NFR2: System must handle network delays and temporary outages gracefully
NFR3: Latency for detecting a missed heartbeat should be under 5 seconds after timeout
NFR4: System availability target is 99.9% uptime
NFR5: Heartbeat messages should be lightweight (under 1 KB)
Think Before You Design
Questions to Ask
❓ Question 1
❓ Question 2
❓ Question 3
❓ Question 4
❓ Question 5
Key Components
Heartbeat sender (client or server component)
Heartbeat receiver service
State store or cache to track last heartbeat timestamps
Alerting and notification system
Dashboard or monitoring UI
Load balancer or API gateway
Design Patterns
Polling vs push-based heartbeat
Timeout and retry mechanisms
Circuit breaker pattern for unhealthy clients
Event-driven architecture for alerting
Caching for fast heartbeat status lookup
Reference Architecture
  +------------+       Heartbeat       +----------------+
  |  Clients   | --------------------> | Heartbeat      |
  | (10,000)  |                       | Receiver       |
  +------------+                       +----------------+
                                         |       |
                                         |       | Updates last
                                         |       | heartbeat time
                                         v       v
                                   +---------------------+
                                   | State Store (Redis)  |
                                   +---------------------+
                                         |       |
                                         |       | Triggers alerts
                                         |       | if timeout
                                         v       v
                                   +---------------------+
                                   | Alerting Service    |
                                   +---------------------+
                                         |
                                         v
                                   +---------------------+
                                   | Monitoring Dashboard|
                                   +---------------------+
Components
Clients
Any client platform
Send periodic heartbeat messages to indicate they are alive
Heartbeat Receiver
Stateless REST API or TCP server
Receive heartbeat messages and update client status
State Store
Redis or in-memory key-value store
Store last heartbeat timestamp per client for quick lookup
Alerting Service
Event-driven microservice
Detect missed heartbeats and send alerts/notifications
Monitoring Dashboard
Web UI with real-time updates
Display live status of all clients and alerts
Load Balancer
Nginx or cloud LB
Distribute heartbeat requests across receiver instances
Request Flow
1. Client sends heartbeat message every 10 seconds to Heartbeat Receiver.
2. Heartbeat Receiver validates and updates the last heartbeat timestamp in State Store.
3. Alerting Service periodically scans State Store for clients missing heartbeat beyond 30 seconds.
4. If a missed heartbeat is detected, Alerting Service triggers notifications.
5. Monitoring Dashboard queries State Store to show live client statuses and alerts.
Database Schema
Entities: - Client: client_id (PK), metadata - HeartbeatRecord: client_id (FK), last_heartbeat_timestamp Relationships: - One-to-one between Client and HeartbeatRecord - HeartbeatRecord updated on each heartbeat received
Scaling Discussion
Bottlenecks
Heartbeat Receiver can become overwhelmed with 10,000+ concurrent heartbeat messages
State Store may face high read/write load for heartbeat timestamps
Alerting Service scanning large datasets may cause latency
Network bandwidth could be stressed by frequent heartbeat messages
Solutions
Use multiple Heartbeat Receiver instances behind a load balancer to distribute load
Use a highly performant in-memory store like Redis with sharding for State Store
Implement incremental or event-driven alerting instead of full scans
Optimize heartbeat message size and interval; consider adaptive heartbeat frequency
Interview Tips
Time: Spend 10 minutes clarifying requirements and constraints, 20 minutes designing architecture and data flow, 10 minutes discussing scaling and trade-offs, 5 minutes summarizing.
Explain why heartbeat is needed and how it helps detect failures
Discuss trade-offs in heartbeat interval and message size
Describe components and their responsibilities clearly
Highlight how system handles scale and failure scenarios
Mention monitoring and alerting importance for operational health

Practice

(1/5)
1. What is the primary purpose of a heartbeat mechanism in system design?
easy
A. To increase the speed of data processing
B. To regularly check if system components are alive and responsive
C. To store user data securely
D. To manage user authentication

Solution

  1. Step 1: Understand the role of heartbeat

    The heartbeat mechanism sends regular signals to check if components are alive.
  2. Step 2: Eliminate unrelated options

    Options about data speed, storage, and authentication do not relate to heartbeat checks.
  3. Final Answer:

    To regularly check if system components are alive and responsive -> Option B
  4. Quick Check:

    Heartbeat = component health check [OK]
Hint: Heartbeat means 'check if alive' regularly [OK]
Common Mistakes:
  • Confusing heartbeat with data storage
  • Thinking heartbeat speeds up processing
  • Mixing heartbeat with authentication
2. Which of the following is the correct way to implement a heartbeat interval in a system?
easy
A. Send heartbeat signals randomly without timing
B. Send heartbeat signals only once at system start
C. Send heartbeat signals only when an error occurs
D. Send heartbeat signals every 5 seconds using a timer

Solution

  1. Step 1: Identify correct heartbeat timing

    Heartbeat signals must be sent regularly, e.g., every 5 seconds, to monitor health.
  2. Step 2: Reject incorrect timing methods

    Sending once, randomly, or only on errors does not provide continuous monitoring.
  3. Final Answer:

    Send heartbeat signals every 5 seconds using a timer -> Option D
  4. Quick Check:

    Heartbeat = regular timed signals [OK]
Hint: Heartbeat needs regular timed signals, not one-time or random [OK]
Common Mistakes:
  • Sending heartbeat only once
  • Using random intervals
  • Triggering heartbeat only on errors
3. Consider a system where a heartbeat is sent every 10 seconds. If the system waits 30 seconds without receiving a heartbeat, what is the likely outcome?
medium
A. The system ignores the missing heartbeat
B. The system assumes the component is alive
C. The system triggers a failure detection and recovery process
D. The system speeds up the heartbeat interval

Solution

  1. Step 1: Understand heartbeat timeout logic

    If no heartbeat is received within a set timeout (30 seconds), the system assumes failure.
  2. Step 2: Identify correct system reaction

    The system triggers failure detection and recovery to handle the unresponsive component.
  3. Final Answer:

    The system triggers a failure detection and recovery process -> Option C
  4. Quick Check:

    Missing heartbeat = trigger recovery [OK]
Hint: No heartbeat in timeout means failure detected [OK]
Common Mistakes:
  • Assuming component is alive without heartbeat
  • Ignoring missing heartbeat signals
  • Changing heartbeat interval automatically
4. A system uses a heartbeat interval of 5 seconds but sets the timeout to 3 seconds. What issue will this cause?
medium
A. The system will falsely detect failures frequently
B. The system will ignore heartbeat signals
C. The system will send heartbeats too slowly
D. The system will never detect failures

Solution

  1. Step 1: Compare heartbeat interval and timeout

    Heartbeat interval (5s) is longer than timeout (3s), so timeout triggers before heartbeat arrives.
  2. Step 2: Identify consequence of timing mismatch

    This causes false failure detection because system thinks heartbeat missed when it hasn't.
  3. Final Answer:

    The system will falsely detect failures frequently -> Option A
  4. Quick Check:

    Timeout < Interval causes false failure [OK]
Hint: Timeout must be longer than heartbeat interval [OK]
Common Mistakes:
  • Setting timeout shorter than heartbeat interval
  • Expecting no failure detection
  • Confusing heartbeat sending speed with timeout
5. In a distributed system with 1000 nodes, how should the heartbeat mechanism be designed to avoid network overload?
hard
A. Nodes send heartbeat in staggered intervals and use hierarchical aggregation
B. All nodes send heartbeat to a single server every second
C. Nodes send heartbeat only when requested by the server
D. Nodes do not send heartbeat to reduce network traffic

Solution

  1. Step 1: Understand scalability challenges

    Sending all heartbeats every second to one server causes overload and bottlenecks.
  2. Step 2: Apply scalable heartbeat design

    Staggering intervals and aggregating heartbeats hierarchically reduces network load and improves efficiency.
  3. Final Answer:

    Nodes send heartbeat in staggered intervals and use hierarchical aggregation -> Option A
  4. Quick Check:

    Scale heartbeat with stagger and aggregation [OK]
Hint: Use staggered timing and aggregation for large scale [OK]
Common Mistakes:
  • Sending all heartbeats simultaneously
  • Not sending heartbeats at all
  • Relying only on server requests for heartbeat