Bird
Raised Fist0
HLDsystem_design~10 mins

Event sourcing in HLD - Scalability & System Analysis

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Scalability Analysis - Event sourcing
Growth Table for Event Sourcing System
Users / Events100 users10K users1M users100M users
Event volume per day~10K events~1M events~100M events~10B events
Event store sizeGBsTBs10s of TBsPetabytes
Read model rebuild timeSecondsMinutesHoursDays
Number of application servers1-210-20100+1000+
Database throughput (QPS)100-5005K-10K50K-100K1M+
Snapshot frequencyEvery 100 eventsEvery 1K eventsEvery 10K eventsEvery 100K events
First Bottleneck

The event store database is the first bottleneck. It must handle a high volume of writes and reads for events. As users and events grow, the database write throughput and storage become limiting factors. Rebuilding read models from large event streams also slows down, impacting query performance.

Scaling Solutions
  • Horizontal scaling: Add more application servers behind load balancers to handle increased event processing and queries.
  • Event store sharding: Partition event data by aggregate or user ID to distribute load across multiple databases.
  • Snapshots: Periodically save aggregate state snapshots to reduce event replay time during read model rebuilds.
  • Caching: Use caches for frequently accessed read models to reduce database load.
  • Read replicas: Use database replicas to scale read queries separately from writes.
  • Archival: Move old events to cheaper storage to reduce primary database size.
  • Asynchronous processing: Use message queues and background workers to decouple event handling and improve throughput.
Back-of-Envelope Cost Analysis

At 1M users generating ~100M events/day:

  • Event write rate: ~1,157 events/sec (100M / 86400 sec)
  • Database QPS needed: ~2,000 (including reads and writes)
  • Storage needed per day: Assuming 1KB per event, ~100GB/day
  • Network bandwidth: ~10 MB/s sustained for event ingestion
  • Snapshot storage: Depends on snapshot frequency, typically smaller than event store
Interview Tip

Start by explaining what event sourcing is and why it helps with auditability and state reconstruction. Then discuss how the event store scales with users and events. Identify the database as the first bottleneck. Propose solutions like sharding, snapshots, and caching. Use numbers to justify your choices. Finally, mention trade-offs like complexity and eventual consistency.

Self Check Question

Your event store database handles 1000 QPS. Traffic grows 10x to 10,000 QPS. What do you do first and why?

Answer: The first step is to shard the event store to distribute the write load across multiple database instances. This reduces the load on any single database and allows scaling writes horizontally. Additionally, implement snapshots to reduce read load and improve read model rebuild times.

Key Result
Event sourcing systems first hit bottlenecks at the event store database due to high write and read loads. Scaling requires sharding, snapshots, caching, and horizontal scaling of application servers to handle growing event volumes and user counts.

Practice

(1/5)
1. What is the main idea behind event sourcing in system design?
easy
A. Store all changes as a sequence of events to reconstruct state
B. Store only the latest snapshot of data for quick access
C. Use events only for logging errors in the system
D. Send events to users as notifications without storing them

Solution

  1. Step 1: Understand event sourcing concept

    Event sourcing means saving every change as an event, not just the final data.
  2. Step 2: Identify how state is managed

    The current state is rebuilt by applying all stored events in order, not by snapshots alone.
  3. Final Answer:

    Store all changes as a sequence of events to reconstruct state -> Option A
  4. Quick Check:

    Event sourcing = store events to rebuild state [OK]
Hint: Event sourcing saves changes as events, not just snapshots [OK]
Common Mistakes:
  • Confusing event sourcing with snapshot-only storage
  • Thinking events are only for error logs
  • Believing events are just notifications
2. Which of the following is the correct way to represent an event in an event sourcing system?
easy
A. { "eventType": "UserCreated", "timestamp": "2024-06-01T12:00:00Z", "data": { "userId": 123 } }
B. [ "UserCreated", 123, "2024-06-01" ]
C. "UserCreated: userId=123 at 2024-06-01"
D. CREATE USER 123 AT 2024-06-01

Solution

  1. Step 1: Identify proper event structure

    Events should be structured data with type, timestamp, and data fields for clarity and processing.
  2. Step 2: Compare options

    { "eventType": "UserCreated", "timestamp": "2024-06-01T12:00:00Z", "data": { "userId": 123 } } uses a clear JSON object with eventType, timestamp, and data, which is standard practice.
  3. Final Answer:

    { "eventType": "UserCreated", "timestamp": "2024-06-01T12:00:00Z", "data": { "userId": 123 } } -> Option A
  4. Quick Check:

    Event = structured JSON with type and data [OK]
Hint: Events are structured objects with type, timestamp, and data [OK]
Common Mistakes:
  • Using unstructured strings for events
  • Confusing event data with SQL commands
  • Using arrays without keys for event details
3. Given these events in order:
[{"eventType":"AddItem","data":{"itemId":1}}, {"eventType":"AddItem","data":{"itemId":2}}, {"eventType":"RemoveItem","data":{"itemId":1}}]
What is the final state of the item list?
medium
A. [1, 2]
B. [2]
C. [1]
D. []

Solution

  1. Step 1: Apply events in order to the item list

    Start with empty list. Add item 1 -> [1]. Add item 2 -> [1, 2]. Remove item 1 -> [2].
  2. Step 2: Determine final list content

    After all events, only item 2 remains in the list.
  3. Final Answer:

    [2] -> Option B
  4. Quick Check:

    Apply events sequentially = final list [2] [OK]
Hint: Apply events one by one to get final state [OK]
Common Mistakes:
  • Ignoring remove event
  • Applying events out of order
  • Assuming all added items remain
4. You notice that replaying all events to rebuild state is very slow. What is a common solution to improve performance in event sourcing?
medium
A. Store only the latest event for each entity
B. Delete old events after 1 day to reduce size
C. Use snapshots to save intermediate states periodically
D. Switch to storing only current state, no events

Solution

  1. Step 1: Identify performance issue cause

    Replaying all events from the start can be slow as event count grows.
  2. Step 2: Choose common optimization

    Snapshots save the full state at points in time, so replay starts from snapshot, reducing replay time.
  3. Final Answer:

    Use snapshots to save intermediate states periodically -> Option C
  4. Quick Check:

    Snapshots speed up event replay [OK]
Hint: Use snapshots to avoid replaying all events every time [OK]
Common Mistakes:
  • Deleting events breaks history and audit
  • Keeping only latest event loses full history
  • Abandoning events loses event sourcing benefits
5. You design an event sourcing system for a bank. Which approach best ensures data consistency and auditability when multiple transactions happen concurrently?
hard
A. Process events in random order to improve throughput
B. Allow events to overwrite each other without checks for speed
C. Store only final balances without event history to simplify design
D. Use optimistic concurrency control with event versioning and conflict detection

Solution

  1. Step 1: Understand concurrency challenges in event sourcing

    Concurrent transactions can cause conflicts if events overwrite each other or are applied out of order.
  2. Step 2: Choose method to maintain consistency and audit

    Optimistic concurrency control uses event version numbers to detect conflicts and prevent overwrites, preserving history and correctness.
  3. Final Answer:

    Use optimistic concurrency control with event versioning and conflict detection -> Option D
  4. Quick Check:

    Optimistic concurrency = safe concurrent event handling [OK]
Hint: Use versioning to detect conflicts in concurrent events [OK]
Common Mistakes:
  • Ignoring conflicts causes data corruption
  • Dropping event history loses audit trail
  • Processing events unordered breaks state correctness