Bird
Raised Fist0
HLDsystem_design~25 mins

Circuit breaker pattern in HLD - System Design Exercise

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Design: Circuit Breaker Pattern Implementation
Design the circuit breaker pattern as a reusable component integrated with service calls. Out of scope: detailed implementation of external services or fallback logic.
Functional Requirements
FR1: Detect failures in calls to external services or components
FR2: Prevent repeated calls to failing services to avoid cascading failures
FR3: Automatically retry calls after a cooldown period
FR4: Provide fallback responses when the external service is unavailable
FR5: Support monitoring of circuit breaker state and metrics
Non-Functional Requirements
NFR1: Handle up to 10,000 requests per second
NFR2: Fail fast with p99 latency under 100ms for service calls
NFR3: Ensure availability of 99.9% for the main application
NFR4: Minimal added latency when circuit breaker is closed (normal operation)
Think Before You Design
Questions to Ask
❓ Question 1
❓ Question 2
❓ Question 3
❓ Question 4
❓ Question 5
Key Components
Circuit breaker state machine (Closed, Open, Half-Open)
Failure detection and counting mechanism
Timeout and retry scheduler
Fallback handler
Metrics and monitoring system
Design Patterns
State machine pattern for managing circuit states
Timeout and retry pattern
Bulkhead pattern to isolate failures
Fallback pattern for degraded responses
Observer pattern for monitoring state changes
Reference Architecture
Client Service
   |
   |---> Circuit Breaker Component ---+---> External Service
                                      |
                                      +---> Fallback Handler
                                      |
                                      +---> Metrics & Monitoring
Components
Circuit Breaker Component
In-memory state machine or distributed cache (e.g., Redis)
Track failure counts, manage states (Closed, Open, Half-Open), and decide if calls should proceed
External Service
Any third-party or internal service
Service being called which may fail or respond slowly
Fallback Handler
Custom code or default response generator
Provide alternative response when circuit is open
Metrics & Monitoring
Prometheus, Grafana, or similar
Collect and visualize circuit breaker state, failure rates, and latency
Request Flow
1. Client Service sends request to Circuit Breaker Component
2. Circuit Breaker checks current state:
3. - If Closed: forwards request to External Service
4. - If Open: immediately returns fallback response
5. - If Half-Open: allows limited requests to test service health
6. External Service responds or fails
7. Circuit Breaker updates failure/success counters based on response
8. If failures exceed threshold, Circuit Breaker transitions to Open state
9. After cooldown period, Circuit Breaker moves to Half-Open to test service
10. Metrics & Monitoring collects state changes and performance data
Database Schema
No persistent database required for core circuit breaker; uses in-memory or distributed cache to store: - CircuitBreakerState { service_id, state (Closed/Open/Half-Open), failure_count, last_failure_time, last_state_change_time } - Configuration { failure_threshold, timeout_duration, retry_interval } Relationships: One-to-one mapping between service_id and CircuitBreakerState
Scaling Discussion
Bottlenecks
Single instance circuit breaker state causing inconsistent behavior in distributed systems
High latency added by synchronous state checks
Memory overhead if many services or endpoints use circuit breakers
Delayed detection of failures due to slow failure count updates
Solutions
Use distributed cache (e.g., Redis) or shared state store for circuit breaker state to synchronize across instances
Implement asynchronous state updates and non-blocking calls to minimize added latency
Apply circuit breaker only to critical or high-risk services to reduce memory usage
Tune failure detection thresholds and use sliding windows for faster failure detection
Interview Tips
Time: Spend 10 minutes understanding requirements and clarifying failure scenarios, 20 minutes designing the architecture and data flow, 10 minutes discussing scaling and trade-offs, 5 minutes summarizing.
Explain the purpose of the circuit breaker pattern to prevent cascading failures
Describe the three states and transitions clearly
Discuss how fallback responses improve user experience during failures
Highlight monitoring importance for operational visibility
Address scaling challenges in distributed environments and solutions

Practice

(1/5)
1. What is the primary purpose of the circuit breaker pattern in system design?
easy
A. To prevent repeated calls to a failing service and improve system stability
B. To increase the number of requests sent to a service
C. To store user session data efficiently
D. To encrypt data during transmission

Solution

  1. Step 1: Understand the circuit breaker pattern role

    The circuit breaker pattern is designed to stop sending requests to a service that is failing repeatedly to avoid wasting resources and cascading failures.
  2. Step 2: Identify the main benefit

    By preventing repeated calls to a failing service, it helps maintain overall system stability and improves user experience by failing fast.
  3. Final Answer:

    To prevent repeated calls to a failing service and improve system stability -> Option A
  4. Quick Check:

    Circuit breaker purpose = prevent repeated failing calls [OK]
Hint: Circuit breaker stops calls to failing services fast [OK]
Common Mistakes:
  • Confusing circuit breaker with caching
  • Thinking it increases request volume
  • Mixing it up with encryption or session management
2. Which of the following correctly describes the open state in a circuit breaker?
easy
A. The circuit breaker allows all requests to pass through
B. The circuit breaker resets all counters to zero
C. The circuit breaker tests a limited number of requests
D. The circuit breaker blocks all requests to the failing service

Solution

  1. Step 1: Recall the circuit breaker states

    The circuit breaker has three states: closed (normal operation), open (blocking requests), and half-open (testing requests).
  2. Step 2: Define the open state behavior

    In the open state, the circuit breaker blocks all requests to the failing service to prevent further failures.
  3. Final Answer:

    The circuit breaker blocks all requests to the failing service -> Option D
  4. Quick Check:

    Open state = block requests [OK]
Hint: Open state means block all requests [OK]
Common Mistakes:
  • Confusing open with closed or half-open states
  • Thinking open state allows requests
  • Assuming counters reset in open state
3. Consider this simplified pseudocode for a circuit breaker:
if failure_count > threshold:
    state = 'open'
if state == 'open':
    return 'fail fast'
else:
    call_service()

What will happen if failure_count exceeds the threshold?
medium
A. The service call will continue normally
B. The circuit breaker will enter half-open state
C. The circuit breaker will return 'fail fast' without calling the service
D. The failure count will reset automatically

Solution

  1. Step 1: Analyze the condition for failure count

    If failure_count is greater than threshold, the state is set to 'open'.
  2. Step 2: Check behavior when state is 'open'

    When state is 'open', the code returns 'fail fast' and does not call the service.
  3. Final Answer:

    The circuit breaker will return 'fail fast' without calling the service -> Option C
  4. Quick Check:

    Failure count > threshold = fail fast [OK]
Hint: Open state returns fail fast, no service call [OK]
Common Mistakes:
  • Assuming service call still happens
  • Confusing open with half-open state
  • Thinking failure count resets automatically
4. A circuit breaker is stuck in the open state and never transitions to half-open. What is the most likely cause?
medium
A. The timeout to reset the circuit breaker is missing or too long
B. The service is always healthy
C. The failure count threshold is set too low
D. The circuit breaker is not counting failures

Solution

  1. Step 1: Understand state transitions in circuit breaker

    The circuit breaker moves from open to half-open after a timeout period to test if the service has recovered.
  2. Step 2: Identify cause of stuck open state

    If the timeout is missing or set too long, the circuit breaker will never try half-open state and remain open indefinitely.
  3. Final Answer:

    The timeout to reset the circuit breaker is missing or too long -> Option A
  4. Quick Check:

    Missing timeout causes stuck open state [OK]
Hint: Timeout missing or too long keeps breaker open [OK]
Common Mistakes:
  • Confusing failure threshold with timeout
  • Assuming service health affects state directly
  • Ignoring the role of failure counting
5. You design a system using the circuit breaker pattern to call a payment service. The service fails intermittently. How should you configure the circuit breaker to balance availability and fault tolerance?
hard
A. Set a high failure threshold and long timeout to avoid blocking the service too soon
B. Set a low failure threshold and short timeout to quickly block and retry the service
C. Disable the circuit breaker to avoid blocking any requests
D. Set failure threshold to zero to block all requests immediately

Solution

  1. Step 1: Understand intermittent failure impact

    Intermittent failures mean the service sometimes works and sometimes fails, so quick detection and retry is important.
  2. Step 2: Choose configuration for balance

    A low failure threshold and short timeout allow the circuit breaker to quickly block failing calls and retry soon, improving fault tolerance and availability.
  3. Final Answer:

    Set a low failure threshold and short timeout to quickly block and retry the service -> Option B
  4. Quick Check:

    Low threshold + short timeout balances availability and fault tolerance [OK]
Hint: Low threshold and short timeout balance retries and blocking [OK]
Common Mistakes:
  • Setting threshold too high delays failure detection
  • Disabling circuit breaker risks cascading failures
  • Setting threshold zero blocks all requests unnecessarily