Bird
Raised Fist0
NumPydata~5 mins

Practical uses of structured arrays in NumPy - Cheat Sheet & Quick Revision

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Recall & Review
beginner
What is a structured array in NumPy?
A structured array is a special type of NumPy array that allows you to store different data types in each element, similar to a table with named columns.
Click to reveal answer
beginner
How can structured arrays help when working with mixed data types?
Structured arrays let you keep related data of different types together, like names (strings), ages (integers), and scores (floats) in one array, making data handling easier.
Click to reveal answer
beginner
Give an example of a practical use of structured arrays.
You can use structured arrays to store and analyze a dataset of students with fields like 'name', 'age', and 'grade'. This helps you quickly access or filter data by any field.
Click to reveal answer
beginner
How do you access a specific field in a structured array?
You access a field by using its name in square brackets, like array['field_name'], which returns all values for that field.
Click to reveal answer
beginner
Why might structured arrays be preferred over regular arrays for some datasets?
Because they allow storing different types of data together with meaningful names, making the data easier to understand and work with compared to plain arrays with only one data type.
Click to reveal answer
What is the main advantage of using structured arrays in NumPy?
AAutomatically clean data
BIncrease the speed of numerical calculations
CStore multiple data types in one array with named fields
DVisualize data easily
How do you define a structured array with fields 'name' (string) and 'age' (integer)?
Anp.array([('Alice', 25)], dtype=[('name', 'U10'), ('age', 'i4')])
Bnp.array(['Alice', 25])
Cnp.array([('Alice', 25)], dtype=int)
Dnp.array([('Alice', 25)], dtype=float)
How do you access the 'age' field from a structured array named 'data'?
Adata['age']
Bdata.age()
Cdata.get('age')
Ddata.age
Which of these is NOT a typical use of structured arrays?
AStoring mixed data types in one array
BPerforming fast matrix multiplication
CFiltering data by field values
DOrganizing tabular data with named columns
What type of data can you store in a structured array field?
AOnly integers
BOnly floats
COnly strings
DAny data type including strings, integers, and floats
Explain what a structured array is and why it is useful in data science.
Think about how you store different types of information about people or objects together.
You got /4 concepts.
    Describe a real-life example where you would use a structured array instead of a regular NumPy array.
    Consider a dataset like a list of employees with names, ages, and salaries.
    You got /4 concepts.

      Practice

      (1/5)
      1. What is the main advantage of using numpy structured arrays in data science?
      easy
      A. They allow storing different data types in one array with named fields.
      B. They only store integers efficiently.
      C. They automatically visualize data.
      D. They replace all pandas functionality.

      Solution

      1. Step 1: Understand structured arrays

        Structured arrays let you store mixed data types in one array with named fields, like columns in a table.
      2. Step 2: Compare options

        Only They allow storing different data types in one array with named fields. correctly describes this main advantage. Others are incorrect or unrelated.
      3. Final Answer:

        They allow storing different data types in one array with named fields. -> Option A
      4. Quick Check:

        Structured arrays = mixed types + named fields [OK]
      Hint: Remember: structured arrays hold mixed types with names [OK]
      Common Mistakes:
      • Thinking structured arrays only store one data type
      • Confusing structured arrays with visualization tools
      • Assuming structured arrays replace pandas completely
      2. Which of the following is the correct way to define a structured array with fields 'name' (string) and 'age' (integer) in numpy?
      easy
      A. np.array([(b'Alice', 25), (b'Bob', 30)], dtype=[('name', 'f8'), ('age', 'S10')])
      B. np.array([('Alice', 25), ('Bob', 30)], dtype=[('name', 'int'), ('age', 'float')])
      C. np.array([(25, 'Alice'), (30, 'Bob')], dtype=[('age', 'i4'), ('name', 'S10')])
      D. np.array([(b'Alice', 25), (b'Bob', 30)], dtype=[('name', 'S10'), ('age', 'i4')])

      Solution

      1. Step 1: Check field names and types

        The fields are 'name' as string (bytes) and 'age' as integer. 'S10' means string of max 10 bytes, 'i4' means 4-byte integer.
      2. Step 2: Validate each option

        np.array([(b'Alice', 25), (b'Bob', 30)], dtype=[('name', 'S10'), ('age', 'i4')]) matches the correct dtype and data format. np.array([('Alice', 25), ('Bob', 30)], dtype=[('name', 'int'), ('age', 'float')]) uses wrong types. np.array([(25, 'Alice'), (30, 'Bob')], dtype=[('age', 'i4'), ('name', 'S10')]) swaps fields order and data. np.array([(b'Alice', 25), (b'Bob', 30)], dtype=[('name', 'f8'), ('age', 'S10')]) swaps types incorrectly.
      3. Final Answer:

        np.array([(b'Alice', 25), (b'Bob', 30)], dtype=[('name', 'S10'), ('age', 'i4')]) -> Option D
      4. Quick Check:

        Correct dtype and data order = np.array([(b'Alice', 25), (b'Bob', 30)], dtype=[('name', 'S10'), ('age', 'i4')]) [OK]
      Hint: Match field names and types exactly in dtype [OK]
      Common Mistakes:
      • Using wrong data types in dtype
      • Swapping field order and data
      • Not using byte strings for fixed-length strings
      3. Given the structured array:
      data = np.array([(b'Alice', 25), (b'Bob', 30), (b'Carol', 22)], dtype=[('name', 'S10'), ('age', 'i4')])
      sorted_data = np.sort(data, order='age')

      What is the output of sorted_data['name']?
      medium
      A. [b'Carol' b'Alice' b'Bob']
      B. [b'Alice' b'Bob' b'Carol']
      C. [b'Bob' b'Carol' b'Alice']
      D. [b'Carol' b'Bob' b'Alice']

      Solution

      1. Step 1: Understand sorting by 'age'

        The array is sorted by the 'age' field ascending: 22 (Carol), 25 (Alice), 30 (Bob).
      2. Step 2: Extract 'name' field after sorting

        After sorting, the 'name' field order matches sorted ages: Carol, Alice, Bob.
      3. Final Answer:

        [b'Carol' b'Alice' b'Bob'] -> Option A
      4. Quick Check:

        Sort by age ascending = Carol, Alice, Bob [OK]
      Hint: Sort by field then check that field's order [OK]
      Common Mistakes:
      • Assuming original order remains after sort
      • Mixing up ascending vs descending order
      • Confusing field names when accessing
      4. Consider this code snippet:
      data = np.array([(b'Alice', 25), (b'Bob', 30)], dtype=[('name', 'S10'), ('age', 'i4')])
      filtered = data[data['age'] > 25]

      What is the error in this code?
      medium
      A. ValueError because comparison with > is invalid on structured arrays.
      B. No error; it correctly filters entries with age > 25.
      C. TypeError because 'age' field is not accessible.
      D. SyntaxError due to wrong indexing syntax.

      Solution

      1. Step 1: Check filtering syntax

        Filtering structured arrays by a field with a condition like data['age'] > 25 is valid and returns a boolean mask.
      2. Step 2: Confirm no errors

        The code correctly filters rows where age is greater than 25, so no error occurs.
      3. Final Answer:

        No error; it correctly filters entries with age > 25. -> Option B
      4. Quick Check:

        Filtering with boolean mask on field works [OK]
      Hint: Use boolean masks on fields to filter structured arrays [OK]
      Common Mistakes:
      • Thinking structured arrays can't be filtered by fields
      • Confusing syntax for filtering
      • Assuming comparison operators don't work on fields
      5. You have a structured array of employees with fields 'name' (string), 'age' (int), and 'salary' (float). You want to find the average salary of employees older than 30. Which code snippet correctly does this?
      hard
      A. avg_salary = data[data['salary'] > 30]['age'].mean()
      B. avg_salary = np.mean(data['salary'] > 30)
      C. avg_salary = data['salary'][data['age'] > 30].mean()
      D. avg_salary = data['salary'].mean(data['age'] > 30)

      Solution

      1. Step 1: Filter employees older than 30

        Use boolean mask data['age'] > 30 to select salaries of employees older than 30.
      2. Step 2: Calculate mean salary of filtered data

        Apply .mean() on the filtered salary array to get average salary.
      3. Final Answer:

        avg_salary = data['salary'][data['age'] > 30].mean() -> Option C
      4. Quick Check:

        Filter by age, then mean salary = avg_salary = data['salary'][data['age'] > 30].mean() [OK]
      Hint: Filter first, then compute mean on selected field [OK]
      Common Mistakes:
      • Using mean on boolean arrays instead of salaries
      • Mixing up fields in filtering and aggregation
      • Passing filter as argument to mean() incorrectly