Bird
Raised Fist0
NumPydata~5 mins

np.in1d() for membership testing in NumPy - Time & Space Complexity

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Time Complexity: np.in1d() for membership testing
O(n)
Understanding Time Complexity

We want to understand how the time needed to check if elements belong to another list grows as the lists get bigger.

How does the work increase when using np.in1d() with larger arrays?

Scenario Under Consideration

Analyze the time complexity of the following code snippet.

import numpy as np

arr1 = np.array([1, 2, 3, 4, 5])
arr2 = np.array([3, 4, 5, 6, 7])
result = np.in1d(arr1, arr2)
print(result)

This code checks which elements of arr1 are also in arr2, returning a boolean array.

Identify Repeating Operations
  • Primary operation: For each element in the first array, check if it exists in the second array.
  • How many times: This check repeats once for every element in the first array.
How Execution Grows With Input

As the first array gets bigger, the number of checks grows directly with its size.

Input Size (n)Approx. Operations
10About 10 membership checks
100About 100 membership checks
1000About 1000 membership checks

Pattern observation: The work grows in a straight line with the size of the first array.

Final Time Complexity

Time Complexity: O(n)

This means the time to complete the check grows directly in proportion to the number of elements in the first array.

Common Mistake

[X] Wrong: "The time depends on both arrays multiplied together because it checks every pair."

[OK] Correct: np.in1d() uses efficient methods internally, so it does not check every pair but rather looks up membership quickly, making the time mainly depend on the first array size.

Interview Connect

Understanding how membership checks scale helps you explain and reason about data filtering and comparison tasks, which are common in data science work.

Self-Check

"What if we changed the second array to a Python set before checking membership? How would the time complexity change?"

Practice

(1/5)
1. What does the np.in1d() function do in NumPy?
easy
A. Finds the unique elements in an array.
B. Sorts the elements of an array in ascending order.
C. Calculates the sum of elements in an array.
D. Checks if elements of one array are present in another array and returns a boolean array.

Solution

  1. Step 1: Understand the purpose of np.in1d()

    The function checks membership of each element in the first array against the second array.
  2. Step 2: Identify the output type

    It returns a boolean array indicating True where elements are found and False otherwise.
  3. Final Answer:

    Checks if elements of one array are present in another array and returns a boolean array. -> Option D
  4. Quick Check:

    Membership test = Checks if elements of one array are present in another array and returns a boolean array. [OK]
Hint: Remember: np.in1d returns booleans for membership [OK]
Common Mistakes:
  • Confusing np.in1d() with sorting or summing functions
  • Expecting np.in1d() to return the matching elements instead of booleans
  • Thinking np.in1d() modifies the original arrays
2. Which of the following is the correct syntax to check if elements of array a are in array b using np.in1d()?
easy
A. np.in1d(b, a)
B. np.in1d(a, b)
C. np.in1d(a == b)
D. np.in1d(a, b, axis=1)

Solution

  1. Step 1: Recall np.in1d() parameter order

    The first argument is the array to test membership for, the second is the array to check against.
  2. Step 2: Evaluate each option

    np.in1d(a, b) uses correct order: np.in1d(a, b). np.in1d(b, a) reverses arrays, np.in1d(a == b) uses invalid syntax, np.in1d(a, b, axis=1) uses unsupported axis parameter.
  3. Final Answer:

    np.in1d(a, b) -> Option B
  4. Quick Check:

    Correct syntax = np.in1d(a, b) [OK]
Hint: First array is tested, second array is reference [OK]
Common Mistakes:
  • Swapping the order of arrays in np.in1d()
  • Adding unsupported parameters like axis
  • Using comparison operators inside np.in1d()
3. What is the output of the following code?
import numpy as np
x = np.array([1, 3, 5, 7])
y = np.array([3, 4, 5])
result = np.in1d(x, y)
print(result)
medium
A. [False True True False]
B. [True False True False]
C. [False True False False]
D. [True True True True]

Solution

  1. Step 1: Check each element of x against y

    1 in y? No (False), 3 in y? Yes (True), 5 in y? Yes (True), 7 in y? No (False).
  2. Step 2: Form the boolean array

    Result is [False, True, True, False].
  3. Final Answer:

    [False True True False] -> Option A
  4. Quick Check:

    Membership booleans = [False True True False] [OK]
Hint: Check each element one by one for membership [OK]
Common Mistakes:
  • Mixing up True and False positions
  • Assuming np.in1d returns matching elements instead of booleans
  • Forgetting to import numpy
4. The following code throws an error. What is the mistake?
import numpy as np
x = [1, 2, 3]
y = np.array([2, 3, 4])
result = np.in1d(x, y, axis=0)
print(result)
medium
A. np.in1d() requires both inputs to be lists.
B. x should be converted to a NumPy array before using np.in1d().
C. np.in1d() does not accept the 'axis' parameter.
D. The arrays x and y must have the same shape.

Solution

  1. Step 1: Check np.in1d() parameters

    np.in1d() accepts only two main parameters: the test array and the array to check against. It does not support an 'axis' parameter.
  2. Step 2: Identify the error cause

    Passing axis=0 causes a TypeError because it's not a valid argument.
  3. Final Answer:

    np.in1d() does not accept the 'axis' parameter. -> Option C
  4. Quick Check:

    Invalid parameter = np.in1d() does not accept the 'axis' parameter. [OK]
Hint: np.in1d() only takes two main arguments [OK]
Common Mistakes:
  • Trying to use axis parameter with np.in1d()
  • Assuming input types must match exactly
  • Thinking np.in1d() requires both inputs as arrays
5. You have two arrays:
data = np.array([10, 20, 30, 40, 50])
filter_vals = np.array([20, 40, 60])

You want to create a new array containing only elements from data that are present in filter_vals. Which code snippet correctly achieves this?
hard
A. filtered = data[np.in1d(data, filter_vals)]
B. filtered = filter_vals[np.in1d(filter_vals, data)]
C. filtered = np.in1d(data, filter_vals)
D. filtered = data[filter_vals]

Solution

  1. Step 1: Use np.in1d() to get boolean mask

    np.in1d(data, filter_vals) returns a boolean array marking elements of data present in filter_vals.
  2. Step 2: Use boolean mask to filter data

    Indexing data with this boolean mask selects only matching elements.
  3. Final Answer:

    filtered = data[np.in1d(data, filter_vals)] -> Option A
  4. Quick Check:

    Boolean mask indexing = filtered = data[np.in1d(data, filter_vals)] [OK]
Hint: Use np.in1d() mask to index original array [OK]
Common Mistakes:
  • Indexing filter_vals instead of data
  • Using np.in1d() without indexing
  • Trying to index with filter_vals directly