Bird
Raised Fist0
NumPydata~5 mins

Record arrays in NumPy

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Introduction

Record arrays let you store different types of data together in one array. This helps when you want to keep related information like names, ages, and scores in one place.

You want to store a list of people with their name, age, and height together.
You have data with different types like strings and numbers mixed in one table.
You want to access columns by name instead of by position.
You need to save structured data that looks like a spreadsheet but in an array.
You want to perform fast operations on mixed-type data using numpy.
Syntax
NumPy
import numpy as np

# Define a record array with fields and types
record_array = np.rec.array(
    [(value1, value2, ...), (value1, value2, ...)],
    dtype=[('field1', type1), ('field2', type2), ...]
)

The dtype defines the name and type of each field.

You can access fields by name like record_array.field1.

Examples
This creates an empty record array with fields 'name' and 'age'. Useful to start with no data.
NumPy
import numpy as np

# Empty record array with defined fields
empty_records = np.rec.array([], dtype=[('name', 'U10'), ('age', 'i4')])
print(empty_records)
Shows a record array with a single entry. You can access the fields by name.
NumPy
import numpy as np

# Record array with one element
one_record = np.rec.array([('Alice', 30)], dtype=[('name', 'U10'), ('age', 'i4')])
print(one_record)
print(one_record.name)
print(one_record.age)
This example has multiple records. You can get all names or ages as arrays.
NumPy
import numpy as np

# Record array with multiple elements
people = np.rec.array([
    ('Bob', 25),
    ('Carol', 40),
    ('Dave', 35)
], dtype=[('name', 'U10'), ('age', 'i4')])
print(people)
print(people.name)
print(people.age)
Shows how to access the last record and its fields.
NumPy
import numpy as np

# Accessing last element
print(people[-1])
print(people[-1].name)
print(people[-1].age)
Sample Program

This program creates a record array with three people, prints it, adds a new person, and prints the updated array. Then it shows how to access each field by name.

NumPy
import numpy as np

# Create a record array with three people
people = np.rec.array([
    ('Alice', 28, 5.5),
    ('Bob', 34, 6.0),
    ('Carol', 22, 5.7)
], dtype=[('name', 'U10'), ('age', 'i4'), ('height', 'f4')])

print("Before adding new record:")
print(people)

# Add a new record by creating a new array with one more element
new_person = np.rec.array([('Dave', 30, 5.9)], dtype=people.dtype)
people = np.concatenate((people, new_person))

print("\nAfter adding new record:")
print(people)

# Access fields by name
print("\nNames:", people.name)
print("Ages:", people.age)
print("Heights:", people.height)
OutputSuccess
Important Notes

Time complexity for accessing a field is O(1) because fields are stored separately internally.

Space complexity is similar to normal numpy arrays but slightly more due to field names.

Common mistake: Trying to add records by appending directly to the record array. Instead, create a new array and concatenate.

Use record arrays when you want structured data with named fields and mixed types. Use pandas DataFrame if you need more complex table operations.

Summary

Record arrays store mixed data types in one numpy array with named fields.

You can access data by field names like array.field.

They are useful for simple structured data and fast access in numpy.

Practice

(1/5)
1. What is the main advantage of using a record array in numpy?
easy
A. It speeds up numerical calculations on large arrays.
B. It automatically sorts data based on values.
C. It allows storing different data types in one array with named fields.
D. It compresses data to save memory.

Solution

  1. Step 1: Understand record arrays

    Record arrays let you store mixed data types in one numpy array by using named fields.
  2. Step 2: Compare options

    Only It allows storing different data types in one array with named fields. correctly describes this feature. Others describe unrelated features.
  3. Final Answer:

    It allows storing different data types in one array with named fields. -> Option C
  4. Quick Check:

    Record arrays = mixed types + named fields [OK]
Hint: Remember: record arrays hold mixed types with names [OK]
Common Mistakes:
  • Confusing record arrays with regular numeric arrays
  • Thinking record arrays sort data automatically
  • Assuming record arrays compress data
2. Which of the following is the correct way to create a numpy record array with fields 'name' (string) and 'age' (integer)?
easy
A. np.rec.array([("Alice", 25), ("Bob", 30)], dtype=[('name', 'U10'), ('age', 'i4')])
B. np.array([("Alice", 25), ("Bob", 30)], dtype=[('name', 'i4'), ('age', 'U10')])
C. np.rec.array(["Alice", 25, "Bob", 30], dtype=[('name', 'U10'), ('age', 'i4')])
D. np.rec.array([(25, "Alice"), (30, "Bob")], dtype=[('name', 'U10'), ('age', 'i4')])

Solution

  1. Step 1: Check data and dtype matching

    np.rec.array([("Alice", 25), ("Bob", 30)], dtype=[('name', 'U10'), ('age', 'i4')]) matches tuples of (string, int) with dtype [('name', 'U10'), ('age', 'i4')].
  2. Step 2: Validate other options

    np.array([("Alice", 25), ("Bob", 30)], dtype=[('name', 'i4'), ('age', 'U10')]) swaps types incorrectly; C has wrong input format; A swaps field order.
  3. Final Answer:

    np.rec.array([("Alice", 25), ("Bob", 30)], dtype=[('name', 'U10'), ('age', 'i4')]) -> Option A
  4. Quick Check:

    Data matches dtype order and types [OK]
Hint: Match tuple order with dtype fields exactly [OK]
Common Mistakes:
  • Swapping field order between data and dtype
  • Using wrong data types in dtype
  • Passing flat list instead of list of tuples
3. What will be the output of the following code?
import numpy as np
rec = np.rec.array([(1, 2.5), (3, 4.5)], dtype=[('x', 'i4'), ('y', 'f4')])
print(rec.x + rec.y)
medium
A. TypeError
B. [3 7]
C. [1 3]
D. [3.5 7.5]

Solution

  1. Step 1: Understand data and fields

    rec.x is integer array [1, 3], rec.y is float array [2.5, 4.5].
  2. Step 2: Add integer and float arrays element-wise

    Adding [1, 3] + [2.5, 4.5] results in [3.5, 7.5] as floats.
  3. Final Answer:

    [3.5 7.5] -> Option D
  4. Quick Check:

    1+2.5=3.5 and 3+4.5=7.5 [OK]
Hint: Adding int and float fields results in float array [OK]
Common Mistakes:
  • Expecting integer output instead of float
  • Confusing field names or types
  • Thinking addition causes error
4. Identify the error in this code snippet:
import numpy as np
rec = np.rec.array([(1, 'a'), (2, 'b')], dtype=[('num', 'i4'), ('char', 'U1')])
print(rec.num + rec.char)
medium
A. You cannot add integer and string fields directly.
B. The dtype specification is incorrect.
C. The data tuples have wrong length.
D. The record array must be created with np.array, not np.rec.array.

Solution

  1. Step 1: Analyze the operation

    rec.num is integer array, rec.char is string array.
  2. Step 2: Check addition of int and string

    Adding int + string causes a TypeError in numpy.
  3. Final Answer:

    You cannot add integer and string fields directly. -> Option A
  4. Quick Check:

    int + string = TypeError [OK]
Hint: Cannot add numbers and strings directly in numpy [OK]
Common Mistakes:
  • Assuming dtype is wrong instead of operation
  • Thinking np.rec.array is incorrect here
  • Ignoring type mismatch in addition
5. You have a numpy record array rec with fields 'id' (int), 'score' (float), and 'passed' (bool). How do you create a new record array containing only records where passed is True and score is above 80?
hard
A. rec[rec.passed or rec.score > 80]
B. rec[(rec.passed) & (rec.score > 80)]
C. rec[rec.passed and rec.score > 80]
D. rec[(rec.passed) | (rec.score > 80)]

Solution

  1. Step 1: Understand filtering syntax

    Use boolean indexing with & for element-wise AND, parentheses needed.
  2. Step 2: Evaluate options

    rec[(rec.passed) & (rec.score > 80)] correctly uses (rec.passed) & (rec.score > 80). Options B and C use Python 'or'/'and' which don't work element-wise. rec[(rec.passed) | (rec.score > 80)] uses | (OR) instead of AND.
  3. Final Answer:

    rec[(rec.passed) & (rec.score > 80)] -> Option B
  4. Quick Check:

    Use & with parentheses for element-wise AND [OK]
Hint: Use & with parentheses for element-wise conditions [OK]
Common Mistakes:
  • Using 'and' or 'or' instead of '&' or '|' for arrays
  • Forgetting parentheses around conditions
  • Using | instead of & for AND condition