Partition key selection in DynamoDB - Time & Space Complexity
Start learning this pattern below
Jump into concepts and practice - no test required
Choosing the right partition key in DynamoDB affects how fast your queries run.
We want to understand how the choice of partition key changes the work DynamoDB does as data grows.
Analyze the time complexity of querying items by partition key.
const params = {
TableName: "Orders",
KeyConditionExpression: "PartitionKey = :pk",
ExpressionAttributeValues: {
":pk": "USER#123"
}
};
const result = await dynamodb.query(params).promise();
This code fetches all items that share the same partition key value.
Look for repeated work done when fetching data.
- Primary operation: Scanning all items within the partition key.
- How many times: Once per item in that partition.
As more items share the same partition key, the query takes longer.
| Input Size (items in partition) | Approx. Operations |
|---|---|
| 10 | 10 |
| 100 | 100 |
| 1000 | 1000 |
Pattern observation: The work grows directly with the number of items in the partition.
Time Complexity: O(n)
This means the time to get results grows in a straight line with how many items share the partition key.
[X] Wrong: "Querying by partition key always takes the same time no matter how many items it has."
[OK] Correct: The query must read each item in that partition, so more items mean more work and longer time.
Understanding how partition key choice affects query speed shows you know how to design efficient databases.
"What if we added a sort key and queried with both partition and sort key? How would the time complexity change?"
Practice
partition key in DynamoDB?Solution
Step 1: Understand partition key purpose
The partition key is used to decide where data is stored in DynamoDB's distributed system.Step 2: Compare options
Only It determines how data is distributed across storage nodes. correctly describes this role; others describe unrelated features.Final Answer:
It determines how data is distributed across storage nodes. -> Option CQuick Check:
Partition key = data distribution [OK]
- Confusing partition key with capacity settings
- Thinking partition key sets data type
- Assuming partition key limits table size
Solution
Step 1: Recall partition key syntax
Partition key uses KeyType "HASH" in DynamoDB KeySchema.Step 2: Check each option
Only "KeySchema": [{"AttributeName": "UserId", "KeyType": "HASH"}] uses "HASH" correctly; others use invalid or incorrect KeyType values.Final Answer:
"KeySchema": [{"AttributeName": "UserId", "KeyType": "HASH"}] -> Option AQuick Check:
Partition key = KeyType HASH [OK]
- Using RANGE instead of HASH for partition key
- Confusing PRIMARY or INDEX as KeyType
- Misnaming KeyType values
UserId having many unique values, and sort key OrderDate, what will happen if you query with UserId = '123' only?Solution
Step 1: Understand query with partition key only
Querying with partition key returns all items with that key, optionally sorted by sort key.Step 2: Analyze given keys
Since UserId is partition key and OrderDate is sort key, querying UserId='123' returns all orders for that user sorted by OrderDate.Final Answer:
You get all orders for user '123' sorted by OrderDate. -> Option BQuick Check:
Query by partition key returns all matching items [OK]
- Thinking query needs sort key value
- Expecting only one item without sort key
- Assuming query returns all users' data
CustomerId as partition key but notice hot partitions causing slow performance. What is the likely cause?Solution
Step 1: Understand hot partitions
Hot partitions happen when few partition keys get most traffic, causing uneven load.Step 2: Analyze CustomerId uniqueness
If CustomerId has few unique values, many requests hit same partitions causing hot spots.Final Answer:
CustomerId has very few unique values causing uneven data distribution. -> Option DQuick Check:
Few unique keys = hot partitions [OK]
- Assuming too many unique keys cause hot partitions
- Confusing missing key with performance issue
- Mixing partition key with sort key roles
Solution
Step 1: Identify key with many unique values
DeviceId uniquely identifies each device, providing many distinct partition keys.Step 2: Evaluate other options
Constant string causes hot partition; Timestamp changes too fast for partition key; DeviceType groups few devices causing uneven load.Final Answer:
Use DeviceId as partition key because it has many unique values. -> Option AQuick Check:
Many unique keys = good partition key [OK]
- Using constant value causing hot partitions
- Choosing timestamp as partition key
- Grouping by device type causing uneven load
