Bird
Raised Fist0
HLDsystem_design~10 mins

Design a unique ID generator in HLD - Scalability & System Analysis

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Scalability Analysis - Design a unique ID generator
Growth Table: Unique ID Generator Scaling
ScaleRequests per Second (RPS)ID Generation MethodStorage NeedsLatencyNotes
100 users~10 RPSSimple timestamp + counterMinimal (few KB)<1 msSingle server, no concurrency issues
10,000 users~1,000 RPSTimestamp + machine ID + sequence numberSmall (MB)<5 msSingle server with concurrency control or small cluster
1,000,000 users~100,000 RPSDistributed ID generator (e.g., Snowflake)Moderate (GB logs)<10 msMultiple servers, coordination needed
100,000,000 users~10,000,000 RPSHighly distributed, sharded generators + cachingLarge (TB logs)<20 msGlobal distribution, fault tolerance critical
First Bottleneck

The first bottleneck is the central coordination or state management that ensures uniqueness. At low scale, a single server can handle ID generation easily. As traffic grows, the server's CPU and memory limits are reached due to concurrency and synchronization overhead. Also, network latency and clock synchronization issues arise in distributed setups.

Scaling Solutions
  • Horizontal Scaling: Add more ID generator nodes with unique machine IDs to distribute load.
  • Sharding: Partition ID space by machine or region to avoid collisions.
  • Caching: Pre-generate ID blocks to reduce coordination calls.
  • Use of Time-based IDs: Incorporate timestamps to reduce coordination.
  • Coordination Services: Use lightweight consensus or coordination (e.g., ZooKeeper) carefully to avoid bottlenecks.
  • Fault Tolerance: Design for node failures without ID collisions.
Back-of-Envelope Cost Analysis
  • At 1M users generating 100K IDs/sec, each ID ~8 bytes -> 800 KB/sec storage if logged.
  • Network bandwidth for 100K RPS with 8-byte IDs ≈ 0.8 MB/sec, easily handled by 1 Gbps network.
  • CPU: Each server can handle ~5K concurrent ID requests; need ~20 servers for 100K RPS.
  • Storage: Logs and backups grow ~70 GB/day at 100K RPS.
Interview Tip

Start by clarifying requirements: ID length, uniqueness scope (global or per service), latency needs, and failure tolerance. Then discuss simple solutions for low scale and identify bottlenecks as scale grows. Propose incremental scaling strategies and justify choices with trade-offs. Always mention fault tolerance and collision avoidance.

Self Check

Your database handles 1000 QPS. Traffic grows 10x to 10,000 QPS. What do you do first?

Answer: Since the database is the bottleneck, first add read replicas or caching to reduce load. For ID generation, move from a single centralized generator to a distributed approach with machine IDs and sequence numbers to avoid database contention.

Key Result
A unique ID generator scales by moving from a simple centralized approach to a distributed system with machine IDs and sequence numbers, addressing coordination bottlenecks and ensuring fault tolerance as traffic grows from thousands to millions of requests per second.

Practice

(1/5)
1. What is the primary purpose of a unique ID generator in a distributed system?
easy
A. To create identifiers that are distinct across all machines and time
B. To encrypt data for secure communication
C. To compress large files efficiently
D. To balance load between servers

Solution

  1. Step 1: Understand the role of unique IDs

    Unique IDs ensure that each identifier is different from others, avoiding conflicts.
  2. Step 2: Recognize distributed system needs

    In distributed systems, IDs must be unique across machines and time to prevent collisions.
  3. Final Answer:

    To create identifiers that are distinct across all machines and time -> Option A
  4. Quick Check:

    Unique ID purpose = distinct identifiers [OK]
Hint: Unique IDs prevent duplicates across systems [OK]
Common Mistakes:
  • Confusing unique ID with encryption
  • Thinking unique ID compresses data
  • Mixing load balancing with ID generation
2. Which of the following is a common component in a unique ID generator design?
easy
A. Encryption key for data security
B. Load balancer to distribute requests
C. Compression algorithm for data size reduction
D. Sequence number to avoid collisions within the same timestamp

Solution

  1. Step 1: Identify components of unique ID generators

    Common components include timestamp, machine identifier, and sequence number.
  2. Step 2: Understand sequence number role

    Sequence numbers help generate multiple unique IDs within the same timestamp to avoid collisions.
  3. Final Answer:

    Sequence number to avoid collisions within the same timestamp -> Option D
  4. Quick Check:

    Sequence number = collision avoidance [OK]
Hint: Sequence numbers prevent same-time ID clashes [OK]
Common Mistakes:
  • Confusing encryption with ID generation
  • Thinking compression is part of ID design
  • Mixing load balancing with ID components
3. Consider a unique ID generator that uses a 41-bit timestamp, 10-bit machine ID, and 12-bit sequence number. What is the maximum number of unique IDs it can generate per millisecond per machine?
medium
A. 8192
B. 1024
C. 4096
D. 2048

Solution

  1. Step 1: Understand bit allocation for sequence number

    The sequence number uses 12 bits, so max IDs per millisecond = 2^12.
  2. Step 2: Calculate 2^12

    2^12 = 4096 unique IDs per millisecond per machine.
  3. Final Answer:

    4096 -> Option C
  4. Quick Check:

    2^12 = 4096 [OK]
Hint: 2^sequence_bits = max IDs/ms [OK]
Common Mistakes:
  • Using machine ID bits instead of sequence bits
  • Calculating 2^10 or 2^11 instead of 2^12
  • Confusing total bits with sequence bits
4. A unique ID generator uses a timestamp, machine ID, and sequence number. If two machines generate IDs at the exact same millisecond with the same sequence number, what is the likely cause of duplicate IDs?
medium
A. Machine IDs are not unique or not included in the ID
B. Timestamp is too large
C. Sequence number is too long
D. The system uses encryption

Solution

  1. Step 1: Analyze ID components for uniqueness

    Machine ID differentiates IDs from different machines at the same time.
  2. Step 2: Identify cause of duplicates

    If machine IDs are missing or not unique, IDs from different machines can collide.
  3. Final Answer:

    Machine IDs are not unique or not included in the ID -> Option A
  4. Quick Check:

    Missing unique machine ID = duplicates [OK]
Hint: Unique machine ID prevents cross-machine duplicates [OK]
Common Mistakes:
  • Blaming timestamp size for duplicates
  • Thinking longer sequence number causes duplicates
  • Confusing encryption with ID uniqueness
5. You need to design a unique ID generator for a global system with thousands of machines generating millions of IDs per second. Which design choice best ensures scalability and uniqueness?
hard
A. Generate random 64-bit numbers without coordination
B. Use a 64-bit ID combining timestamp, machine ID, and sequence number with synchronized clocks
C. Use only timestamp-based IDs without machine info
D. Assign IDs sequentially from a central server

Solution

  1. Step 1: Consider scalability and uniqueness needs

    Global scale requires IDs unique across machines and time, with high throughput.
  2. Step 2: Evaluate design options

    Combining timestamp, machine ID, and sequence number in 64 bits with synchronized clocks ensures uniqueness and scalability.
  3. Step 3: Reject other options

    Random IDs risk collisions; timestamp-only lacks machine uniqueness; central server causes bottleneck.
  4. Final Answer:

    Use a 64-bit ID combining timestamp, machine ID, and sequence number with synchronized clocks -> Option B
  5. Quick Check:

    64-bit composite ID = scalable unique IDs [OK]
Hint: Combine time, machine, sequence for scalable unique IDs [OK]
Common Mistakes:
  • Relying on random IDs risking collisions
  • Ignoring machine ID causing duplicates
  • Using central server causing bottlenecks