Bird
Raised Fist0
NumPydata~10 mins

Record arrays in NumPy - Step-by-Step Execution

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Concept Flow - Record arrays
Define dtype with field names and types
↓
Create record array with structured data
↓
Access fields by name
↓
Perform operations on fields
↓
Use record array for analysis or display
Record arrays store data with named fields, like a table with columns. You define the structure, create the array, then access fields by name for easy analysis.
Execution Sample
NumPy
import numpy as np

dtype = [('name', 'U10'), ('age', 'i4'), ('score', 'f4')]
data = [('Alice', 25, 88.5), ('Bob', 30, 92.0), ('Cathy', 22, 79.5)]
rec_arr = np.array(data, dtype=dtype)
print(rec_arr['age'])
Creates a record array with fields name, age, score and prints the age column.
Execution Table
StepActionData StateField AccessOutput
1Define dtype with fields[('name', 'U10'), ('age', 'i4'), ('score', 'f4')]N/AN/A
2Create record array with data[('Alice', 25, 88.5), ('Bob', 30, 92.0), ('Cathy', 22, 79.5)]N/AN/A
3Access 'age' fieldRecord array unchangedrec_arr['age'][25 30 22]
4Access 'name' fieldRecord array unchangedrec_arr['name']['Alice' 'Bob' 'Cathy']
5Access 'score' fieldRecord array unchangedrec_arr['score'][88.5 92. 79.5]
6Filter ages > 23Record array unchangedrec_arr['age'] > 23[ True True False]
7Select records with age > 23Record array unchangedrec_arr[rec_arr['age'] > 23][('Alice', 25, 88.5) ('Bob', 30, 92.0)]
8ExitEnd of operationsN/AN/A
💡 All steps executed; final filtered record array shown.
Variable Tracker
VariableStartAfter Step 2After Step 7Final
dtypeundefined[('name', 'U10'), ('age', 'i4'), ('score', 'f4')][('name', 'U10'), ('age', 'i4'), ('score', 'f4')][('name', 'U10'), ('age', 'i4'), ('score', 'f4')]
dataundefined[('Alice', 25, 88.5), ('Bob', 30, 92.0), ('Cathy', 22, 79.5)][('Alice', 25, 88.5), ('Bob', 30, 92.0), ('Cathy', 22, 79.5)][('Alice', 25, 88.5), ('Bob', 30, 92.0), ('Cathy', 22, 79.5)]
rec_arrundefined[('Alice', 25, 88.5), ('Bob', 30, 92.0), ('Cathy', 22, 79.5)][('Alice', 25, 88.5), ('Bob', 30, 92.0), ('Cathy', 22, 79.5)][('Alice', 25, 88.5), ('Bob', 30, 92.0), ('Cathy', 22, 79.5)]
Key Moments - 3 Insights
Why can we access fields by name like rec_arr['age']?
Because the record array is created with a dtype that defines named fields, numpy allows direct access to each field as if it were a column. See execution_table step 3.
What happens when we filter with rec_arr['age'] > 23?
This creates a boolean array showing which records meet the condition. We then use it to select only those records. See execution_table steps 6 and 7.
Is the original record array changed after filtering?
No, filtering returns a new array with selected records. The original record array stays the same. See execution_table step 7.
Visual Quiz - 3 Questions
Test your understanding
Look at the execution table, what is the output of rec_arr['score'] at step 5?
A[88.5 92. 79.5]
B[25 30 22]
C['Alice' 'Bob' 'Cathy']
D[True True False]
💡 Hint
Check the 'Output' column in execution_table row 5.
At which step does the condition rec_arr['age'] > 23 become False for some records?
AStep 3
BStep 7
CStep 6
DStep 4
💡 Hint
Look at the 'Field Access' and 'Output' columns in execution_table row 6.
If we change the dtype to remove the 'score' field, what happens when accessing rec_arr['score']?
AIt returns an empty array.
BIt raises an IndexError.
CIt returns all zeros.
DIt returns the 'age' field instead.
💡 Hint
Accessing a non-existent field in a record array causes an error.
Concept Snapshot
Record arrays store data with named fields.
Define dtype with field names and types.
Create array with structured data.
Access fields by name like columns.
Filter and select records easily.
Useful for table-like data in numpy.
Full Transcript
Record arrays in numpy let you store data with named fields, like columns in a table. First, you define a dtype that lists each field's name and data type. Then you create a numpy array with this dtype and your data. You can access each field by its name, for example rec_arr['age'] gives all ages. You can also filter records by conditions on fields, like selecting only those with age greater than 23. The original array stays unchanged when filtering; a new array is returned. This makes record arrays very handy for structured data analysis.

Practice

(1/5)
1. What is the main advantage of using a record array in numpy?
easy
A. It speeds up numerical calculations on large arrays.
B. It automatically sorts data based on values.
C. It allows storing different data types in one array with named fields.
D. It compresses data to save memory.

Solution

  1. Step 1: Understand record arrays

    Record arrays let you store mixed data types in one numpy array by using named fields.
  2. Step 2: Compare options

    Only It allows storing different data types in one array with named fields. correctly describes this feature. Others describe unrelated features.
  3. Final Answer:

    It allows storing different data types in one array with named fields. -> Option C
  4. Quick Check:

    Record arrays = mixed types + named fields [OK]
Hint: Remember: record arrays hold mixed types with names [OK]
Common Mistakes:
  • Confusing record arrays with regular numeric arrays
  • Thinking record arrays sort data automatically
  • Assuming record arrays compress data
2. Which of the following is the correct way to create a numpy record array with fields 'name' (string) and 'age' (integer)?
easy
A. np.rec.array([("Alice", 25), ("Bob", 30)], dtype=[('name', 'U10'), ('age', 'i4')])
B. np.array([("Alice", 25), ("Bob", 30)], dtype=[('name', 'i4'), ('age', 'U10')])
C. np.rec.array(["Alice", 25, "Bob", 30], dtype=[('name', 'U10'), ('age', 'i4')])
D. np.rec.array([(25, "Alice"), (30, "Bob")], dtype=[('name', 'U10'), ('age', 'i4')])

Solution

  1. Step 1: Check data and dtype matching

    np.rec.array([("Alice", 25), ("Bob", 30)], dtype=[('name', 'U10'), ('age', 'i4')]) matches tuples of (string, int) with dtype [('name', 'U10'), ('age', 'i4')].
  2. Step 2: Validate other options

    np.array([("Alice", 25), ("Bob", 30)], dtype=[('name', 'i4'), ('age', 'U10')]) swaps types incorrectly; C has wrong input format; A swaps field order.
  3. Final Answer:

    np.rec.array([("Alice", 25), ("Bob", 30)], dtype=[('name', 'U10'), ('age', 'i4')]) -> Option A
  4. Quick Check:

    Data matches dtype order and types [OK]
Hint: Match tuple order with dtype fields exactly [OK]
Common Mistakes:
  • Swapping field order between data and dtype
  • Using wrong data types in dtype
  • Passing flat list instead of list of tuples
3. What will be the output of the following code?
import numpy as np
rec = np.rec.array([(1, 2.5), (3, 4.5)], dtype=[('x', 'i4'), ('y', 'f4')])
print(rec.x + rec.y)
medium
A. TypeError
B. [3 7]
C. [1 3]
D. [3.5 7.5]

Solution

  1. Step 1: Understand data and fields

    rec.x is integer array [1, 3], rec.y is float array [2.5, 4.5].
  2. Step 2: Add integer and float arrays element-wise

    Adding [1, 3] + [2.5, 4.5] results in [3.5, 7.5] as floats.
  3. Final Answer:

    [3.5 7.5] -> Option D
  4. Quick Check:

    1+2.5=3.5 and 3+4.5=7.5 [OK]
Hint: Adding int and float fields results in float array [OK]
Common Mistakes:
  • Expecting integer output instead of float
  • Confusing field names or types
  • Thinking addition causes error
4. Identify the error in this code snippet:
import numpy as np
rec = np.rec.array([(1, 'a'), (2, 'b')], dtype=[('num', 'i4'), ('char', 'U1')])
print(rec.num + rec.char)
medium
A. You cannot add integer and string fields directly.
B. The dtype specification is incorrect.
C. The data tuples have wrong length.
D. The record array must be created with np.array, not np.rec.array.

Solution

  1. Step 1: Analyze the operation

    rec.num is integer array, rec.char is string array.
  2. Step 2: Check addition of int and string

    Adding int + string causes a TypeError in numpy.
  3. Final Answer:

    You cannot add integer and string fields directly. -> Option A
  4. Quick Check:

    int + string = TypeError [OK]
Hint: Cannot add numbers and strings directly in numpy [OK]
Common Mistakes:
  • Assuming dtype is wrong instead of operation
  • Thinking np.rec.array is incorrect here
  • Ignoring type mismatch in addition
5. You have a numpy record array rec with fields 'id' (int), 'score' (float), and 'passed' (bool). How do you create a new record array containing only records where passed is True and score is above 80?
hard
A. rec[rec.passed or rec.score > 80]
B. rec[(rec.passed) & (rec.score > 80)]
C. rec[rec.passed and rec.score > 80]
D. rec[(rec.passed) | (rec.score > 80)]

Solution

  1. Step 1: Understand filtering syntax

    Use boolean indexing with & for element-wise AND, parentheses needed.
  2. Step 2: Evaluate options

    rec[(rec.passed) & (rec.score > 80)] correctly uses (rec.passed) & (rec.score > 80). Options B and C use Python 'or'/'and' which don't work element-wise. rec[(rec.passed) | (rec.score > 80)] uses | (OR) instead of AND.
  3. Final Answer:

    rec[(rec.passed) & (rec.score > 80)] -> Option B
  4. Quick Check:

    Use & with parentheses for element-wise AND [OK]
Hint: Use & with parentheses for element-wise conditions [OK]
Common Mistakes:
  • Using 'and' or 'or' instead of '&' or '|' for arrays
  • Forgetting parentheses around conditions
  • Using | instead of & for AND condition