Bird
Raised Fist0
HLDsystem_design~7 mins

Event sourcing in HLD - System Design Guide

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Problem Statement
When a system stores only the current state, it loses the history of changes. This makes it hard to audit, debug, or reconstruct past states. If data corruption or bugs occur, recovering previous states or understanding how the system reached its current state becomes nearly impossible.
Solution
Event sourcing solves this by storing every change as a sequence of events. Instead of saving just the final state, the system records all actions that led to it. The current state is rebuilt by replaying these events in order, preserving full history and enabling easy recovery and auditing.
Architecture
Client App
Command Handler
Aggregate State
Event Handler

This diagram shows how client commands are handled to produce events stored in an event store. The aggregate state is rebuilt by replaying events from the event stream.

Trade-offs
✓ Pros
Preserves complete history of all changes for auditing and debugging.
Enables easy recovery of past states by replaying events.
Supports temporal queries and time travel debugging.
Facilitates integration with other systems via event streams.
✗ Cons
Increases storage requirements due to storing all events.
Rebuilding state can be slow if event streams grow large without snapshots.
Adds complexity in handling event versioning and schema evolution.
Use when auditability, traceability, or complex state reconstruction is critical, especially in systems with complex business logic or regulatory requirements. Suitable for systems with moderate to high write volumes where event history is valuable.
Avoid when system state changes are simple and history is not needed, or when write throughput is extremely high and event storage overhead is prohibitive.
Real World Examples
LinkedIn
Uses event sourcing to track user activity and profile changes, enabling detailed audit trails and rollback capabilities.
Uber
Applies event sourcing to model ride lifecycle events, allowing reconstruction of ride states and handling complex workflows.
Amazon
Employs event sourcing in order management to maintain a full history of order state changes for auditing and customer service.
Alternatives
State-based persistence
Stores only the current state snapshot without recording individual changes.
Use when: When system simplicity and low storage overhead are priorities and historical data is not required.
Command Query Responsibility Segregation (CQRS)
Separates read and write models; can be combined with event sourcing but focuses on query optimization.
Use when: When read and write workloads differ significantly and need separate optimization.
Summary
Event sourcing records every change as an event, preserving full history.
It enables rebuilding system state by replaying events in order.
This pattern improves auditability and debugging but adds storage and complexity overhead.

Practice

(1/5)
1. What is the main idea behind event sourcing in system design?
easy
A. Store all changes as a sequence of events to reconstruct state
B. Store only the latest snapshot of data for quick access
C. Use events only for logging errors in the system
D. Send events to users as notifications without storing them

Solution

  1. Step 1: Understand event sourcing concept

    Event sourcing means saving every change as an event, not just the final data.
  2. Step 2: Identify how state is managed

    The current state is rebuilt by applying all stored events in order, not by snapshots alone.
  3. Final Answer:

    Store all changes as a sequence of events to reconstruct state -> Option A
  4. Quick Check:

    Event sourcing = store events to rebuild state [OK]
Hint: Event sourcing saves changes as events, not just snapshots [OK]
Common Mistakes:
  • Confusing event sourcing with snapshot-only storage
  • Thinking events are only for error logs
  • Believing events are just notifications
2. Which of the following is the correct way to represent an event in an event sourcing system?
easy
A. { "eventType": "UserCreated", "timestamp": "2024-06-01T12:00:00Z", "data": { "userId": 123 } }
B. [ "UserCreated", 123, "2024-06-01" ]
C. "UserCreated: userId=123 at 2024-06-01"
D. CREATE USER 123 AT 2024-06-01

Solution

  1. Step 1: Identify proper event structure

    Events should be structured data with type, timestamp, and data fields for clarity and processing.
  2. Step 2: Compare options

    { "eventType": "UserCreated", "timestamp": "2024-06-01T12:00:00Z", "data": { "userId": 123 } } uses a clear JSON object with eventType, timestamp, and data, which is standard practice.
  3. Final Answer:

    { "eventType": "UserCreated", "timestamp": "2024-06-01T12:00:00Z", "data": { "userId": 123 } } -> Option A
  4. Quick Check:

    Event = structured JSON with type and data [OK]
Hint: Events are structured objects with type, timestamp, and data [OK]
Common Mistakes:
  • Using unstructured strings for events
  • Confusing event data with SQL commands
  • Using arrays without keys for event details
3. Given these events in order:
[{"eventType":"AddItem","data":{"itemId":1}}, {"eventType":"AddItem","data":{"itemId":2}}, {"eventType":"RemoveItem","data":{"itemId":1}}]
What is the final state of the item list?
medium
A. [1, 2]
B. [2]
C. [1]
D. []

Solution

  1. Step 1: Apply events in order to the item list

    Start with empty list. Add item 1 -> [1]. Add item 2 -> [1, 2]. Remove item 1 -> [2].
  2. Step 2: Determine final list content

    After all events, only item 2 remains in the list.
  3. Final Answer:

    [2] -> Option B
  4. Quick Check:

    Apply events sequentially = final list [2] [OK]
Hint: Apply events one by one to get final state [OK]
Common Mistakes:
  • Ignoring remove event
  • Applying events out of order
  • Assuming all added items remain
4. You notice that replaying all events to rebuild state is very slow. What is a common solution to improve performance in event sourcing?
medium
A. Store only the latest event for each entity
B. Delete old events after 1 day to reduce size
C. Use snapshots to save intermediate states periodically
D. Switch to storing only current state, no events

Solution

  1. Step 1: Identify performance issue cause

    Replaying all events from the start can be slow as event count grows.
  2. Step 2: Choose common optimization

    Snapshots save the full state at points in time, so replay starts from snapshot, reducing replay time.
  3. Final Answer:

    Use snapshots to save intermediate states periodically -> Option C
  4. Quick Check:

    Snapshots speed up event replay [OK]
Hint: Use snapshots to avoid replaying all events every time [OK]
Common Mistakes:
  • Deleting events breaks history and audit
  • Keeping only latest event loses full history
  • Abandoning events loses event sourcing benefits
5. You design an event sourcing system for a bank. Which approach best ensures data consistency and auditability when multiple transactions happen concurrently?
hard
A. Process events in random order to improve throughput
B. Allow events to overwrite each other without checks for speed
C. Store only final balances without event history to simplify design
D. Use optimistic concurrency control with event versioning and conflict detection

Solution

  1. Step 1: Understand concurrency challenges in event sourcing

    Concurrent transactions can cause conflicts if events overwrite each other or are applied out of order.
  2. Step 2: Choose method to maintain consistency and audit

    Optimistic concurrency control uses event version numbers to detect conflicts and prevent overwrites, preserving history and correctness.
  3. Final Answer:

    Use optimistic concurrency control with event versioning and conflict detection -> Option D
  4. Quick Check:

    Optimistic concurrency = safe concurrent event handling [OK]
Hint: Use versioning to detect conflicts in concurrent events [OK]
Common Mistakes:
  • Ignoring conflicts causes data corruption
  • Dropping event history loses audit trail
  • Processing events unordered breaks state correctness