Structured arrays help organize different types of data together in one table. This makes it easy to work with complex data like records or mixed information.
Practical uses of structured arrays in NumPy
Start learning this pattern below
Jump into concepts and practice - no test required
import numpy as np # Define a structured array with fields structured_array = np.array([ ('Alice', 25, 88.5), ('Bob', 30, 92.0), ('Charlie', 22, 79.5) ], dtype=[('name', 'U10'), ('age', 'i4'), ('score', 'f4')])
The dtype defines the fields: name (string), age (integer), score (float).
Each element is a tuple matching the fields in order.
import numpy as np # Empty structured array with 3 fields empty_array = np.array([], dtype=[('name', 'U10'), ('age', 'i4'), ('score', 'f4')]) print(empty_array)
import numpy as np # Structured array with one element one_element = np.array([('Diana', 28, 85.0)], dtype=[('name', 'U10'), ('age', 'i4'), ('score', 'f4')]) print(one_element)
import numpy as np # Accessing a field print(one_element['name'])
import numpy as np # Sorting by age sorted_array = np.sort(one_element, order='age') print(sorted_array)
This program creates a structured array of students with their names, ages, and scores. It shows how to access a single field, filter by age, and sort by score.
import numpy as np # Create a structured array with fields: name, age, score students = np.array([ ('Alice', 25, 88.5), ('Bob', 30, 92.0), ('Charlie', 22, 79.5), ('Diana', 28, 85.0) ], dtype=[('name', 'U10'), ('age', 'i4'), ('score', 'f4')]) print("Original array:") print(students) # Access the 'age' field ages = students['age'] print("\nAges:") print(ages) # Filter students older than 25 older_students = students[students['age'] > 25] print("\nStudents older than 25:") print(older_students) # Sort students by score sorted_by_score = np.sort(students, order='score') print("\nStudents sorted by score:") print(sorted_by_score)
Time complexity for accessing fields is O(1) because fields are stored separately.
Filtering and sorting depend on the number of elements, typically O(n) for filtering and O(n log n) for sorting.
Common mistake: forgetting to define the dtype properly, which causes errors or wrong data types.
Use structured arrays when you want to keep related data together but with different types, instead of separate arrays.
Structured arrays store mixed data types in one array with named fields.
They make it easy to access, filter, and sort complex data.
Useful for handling tabular data like records or datasets with different types.
Practice
numpy structured arrays in data science?Solution
Step 1: Understand structured arrays
Structured arrays let you store mixed data types in one array with named fields, like columns in a table.Step 2: Compare options
Only They allow storing different data types in one array with named fields. correctly describes this main advantage. Others are incorrect or unrelated.Final Answer:
They allow storing different data types in one array with named fields. -> Option AQuick Check:
Structured arrays = mixed types + named fields [OK]
- Thinking structured arrays only store one data type
- Confusing structured arrays with visualization tools
- Assuming structured arrays replace pandas completely
numpy?Solution
Step 1: Check field names and types
The fields are 'name' as string (bytes) and 'age' as integer. 'S10' means string of max 10 bytes, 'i4' means 4-byte integer.Step 2: Validate each option
np.array([(b'Alice', 25), (b'Bob', 30)], dtype=[('name', 'S10'), ('age', 'i4')]) matches the correct dtype and data format. np.array([('Alice', 25), ('Bob', 30)], dtype=[('name', 'int'), ('age', 'float')]) uses wrong types. np.array([(25, 'Alice'), (30, 'Bob')], dtype=[('age', 'i4'), ('name', 'S10')]) swaps fields order and data. np.array([(b'Alice', 25), (b'Bob', 30)], dtype=[('name', 'f8'), ('age', 'S10')]) swaps types incorrectly.Final Answer:
np.array([(b'Alice', 25), (b'Bob', 30)], dtype=[('name', 'S10'), ('age', 'i4')]) -> Option DQuick Check:
Correct dtype and data order = np.array([(b'Alice', 25), (b'Bob', 30)], dtype=[('name', 'S10'), ('age', 'i4')]) [OK]
- Using wrong data types in dtype
- Swapping field order and data
- Not using byte strings for fixed-length strings
data = np.array([(b'Alice', 25), (b'Bob', 30), (b'Carol', 22)], dtype=[('name', 'S10'), ('age', 'i4')])
sorted_data = np.sort(data, order='age')What is the output of
sorted_data['name']?Solution
Step 1: Understand sorting by 'age'
The array is sorted by the 'age' field ascending: 22 (Carol), 25 (Alice), 30 (Bob).Step 2: Extract 'name' field after sorting
After sorting, the 'name' field order matches sorted ages: Carol, Alice, Bob.Final Answer:
[b'Carol' b'Alice' b'Bob'] -> Option AQuick Check:
Sort by age ascending = Carol, Alice, Bob [OK]
- Assuming original order remains after sort
- Mixing up ascending vs descending order
- Confusing field names when accessing
data = np.array([(b'Alice', 25), (b'Bob', 30)], dtype=[('name', 'S10'), ('age', 'i4')])
filtered = data[data['age'] > 25]What is the error in this code?
Solution
Step 1: Check filtering syntax
Filtering structured arrays by a field with a condition like data['age'] > 25 is valid and returns a boolean mask.Step 2: Confirm no errors
The code correctly filters rows where age is greater than 25, so no error occurs.Final Answer:
No error; it correctly filters entries with age > 25. -> Option BQuick Check:
Filtering with boolean mask on field works [OK]
- Thinking structured arrays can't be filtered by fields
- Confusing syntax for filtering
- Assuming comparison operators don't work on fields
Solution
Step 1: Filter employees older than 30
Use boolean mask data['age'] > 30 to select salaries of employees older than 30.Step 2: Calculate mean salary of filtered data
Apply .mean() on the filtered salary array to get average salary.Final Answer:
avg_salary = data['salary'][data['age'] > 30].mean() -> Option CQuick Check:
Filter by age, then mean salary = avg_salary = data['salary'][data['age'] > 30].mean() [OK]
- Using mean on boolean arrays instead of salaries
- Mixing up fields in filtering and aggregation
- Passing filter as argument to mean() incorrectly
