Record arrays in NumPy - Time & Space Complexity
Start learning this pattern below
Jump into concepts and practice - no test required
We want to understand how the time to access and manipulate data in numpy record arrays changes as the data size grows.
Specifically, how does the cost grow when working with structured data stored in record arrays?
Analyze the time complexity of the following code snippet.
import numpy as np
# Define a record array with 3 fields
data = np.recarray(1000, dtype=[('name', 'U10'), ('age', 'i4'), ('score', 'f4')])
# Access the 'age' field for all records
ages = data.age
# Compute the average age
average_age = np.mean(ages)
This code creates a record array with 1000 entries, accesses one field for all records, and calculates the average of that field.
Identify the loops, recursion, array traversals that repeat.
- Primary operation: Accessing the 'age' field for all 1000 records and computing the mean.
- How many times: The operation touches each record once, so 1000 times.
As the number of records grows, the time to access and process each record grows proportionally.
| Input Size (n) | Approx. Operations |
|---|---|
| 10 | About 10 operations to access and process |
| 100 | About 100 operations |
| 1000 | About 1000 operations |
Pattern observation: The operations grow linearly with the number of records.
Time Complexity: O(n)
This means the time to access and compute over the record array grows directly in proportion to the number of records.
[X] Wrong: "Accessing a field in a record array is instant and does not depend on the number of records."
[OK] Correct: Even though the field access looks simple, numpy must read each record's field, so the time grows with the number of records.
Understanding how structured data access scales helps you reason about performance in real data tasks and shows you can analyze array operations clearly.
"What if we accessed multiple fields at once instead of just one? How would the time complexity change?"
Practice
record array in numpy?Solution
Step 1: Understand record arrays
Record arrays let you store mixed data types in one numpy array by using named fields.Step 2: Compare options
Only It allows storing different data types in one array with named fields. correctly describes this feature. Others describe unrelated features.Final Answer:
It allows storing different data types in one array with named fields. -> Option CQuick Check:
Record arrays = mixed types + named fields [OK]
- Confusing record arrays with regular numeric arrays
- Thinking record arrays sort data automatically
- Assuming record arrays compress data
'name' (string) and 'age' (integer)?Solution
Step 1: Check data and dtype matching
np.rec.array([("Alice", 25), ("Bob", 30)], dtype=[('name', 'U10'), ('age', 'i4')]) matches tuples of (string, int) with dtype [('name', 'U10'), ('age', 'i4')].Step 2: Validate other options
np.array([("Alice", 25), ("Bob", 30)], dtype=[('name', 'i4'), ('age', 'U10')]) swaps types incorrectly; C has wrong input format; A swaps field order.Final Answer:
np.rec.array([("Alice", 25), ("Bob", 30)], dtype=[('name', 'U10'), ('age', 'i4')]) -> Option AQuick Check:
Data matches dtype order and types [OK]
- Swapping field order between data and dtype
- Using wrong data types in dtype
- Passing flat list instead of list of tuples
import numpy as np
rec = np.rec.array([(1, 2.5), (3, 4.5)], dtype=[('x', 'i4'), ('y', 'f4')])
print(rec.x + rec.y)Solution
Step 1: Understand data and fields
rec.x is integer array [1, 3], rec.y is float array [2.5, 4.5].Step 2: Add integer and float arrays element-wise
Adding [1, 3] + [2.5, 4.5] results in [3.5, 7.5] as floats.Final Answer:
[3.5 7.5] -> Option DQuick Check:
1+2.5=3.5 and 3+4.5=7.5 [OK]
- Expecting integer output instead of float
- Confusing field names or types
- Thinking addition causes error
import numpy as np
rec = np.rec.array([(1, 'a'), (2, 'b')], dtype=[('num', 'i4'), ('char', 'U1')])
print(rec.num + rec.char)Solution
Step 1: Analyze the operation
rec.num is integer array, rec.char is string array.Step 2: Check addition of int and string
Adding int + string causes a TypeError in numpy.Final Answer:
You cannot add integer and string fields directly. -> Option AQuick Check:
int + string = TypeError [OK]
- Assuming dtype is wrong instead of operation
- Thinking np.rec.array is incorrect here
- Ignoring type mismatch in addition
rec with fields 'id' (int), 'score' (float), and 'passed' (bool). How do you create a new record array containing only records where passed is True and score is above 80?Solution
Step 1: Understand filtering syntax
Use boolean indexing with & for element-wise AND, parentheses needed.Step 2: Evaluate options
rec[(rec.passed) & (rec.score > 80)] correctly uses (rec.passed) & (rec.score > 80). Options B and C use Python 'or'/'and' which don't work element-wise. rec[(rec.passed) | (rec.score > 80)] uses | (OR) instead of AND.Final Answer:
rec[(rec.passed) & (rec.score > 80)] -> Option BQuick Check:
Use & with parentheses for element-wise AND [OK]
- Using 'and' or 'or' instead of '&' or '|' for arrays
- Forgetting parentheses around conditions
- Using | instead of & for AND condition
