Bird
Raised Fist0
NumPydata~5 mins

np.union1d() for union in NumPy - Time & Space Complexity

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Time Complexity: np.union1d() for union
O(n log n)
Understanding Time Complexity

We want to understand how the time needed to find the union of two arrays changes as the arrays get bigger.

Specifically, how does np.union1d() behave when input size grows?

Scenario Under Consideration

Analyze the time complexity of the following code snippet.

import numpy as np

arr1 = np.array([1, 2, 3, 4])
arr2 = np.array([3, 4, 5, 6])

result = np.union1d(arr1, arr2)
print(result)

This code finds all unique elements from both arrays combined, sorted in order.

Identify Repeating Operations

Identify the loops, recursion, array traversals that repeat.

  • Primary operation: Sorting and merging the two arrays after removing duplicates.
  • How many times: The arrays are traversed multiple times internally, mainly during sorting and merging steps.
How Execution Grows With Input

As the size of the input arrays grows, the time to find the union grows a bit faster than the size itself.

Input Size (n)Approx. Operations
10About 10 log 10 (around 33)
100About 100 log 100 (around 664)
1000About 1000 log 1000 (around 9966)

Pattern observation: The operations grow a bit faster than the input size, roughly multiplying by the input size times the log of the input size.

Final Time Complexity

Time Complexity: O(n log n)

This means the time needed grows a bit faster than the input size, mainly because sorting is involved.

Common Mistake

[X] Wrong: "np.union1d() just scans the arrays once, so it runs in linear time O(n)."

[OK] Correct: Internally, np.union1d() sorts the combined data, which takes more time than just scanning once. So it is slower than linear time.

Interview Connect

Understanding how functions like np.union1d() scale helps you reason about performance in data tasks. This skill shows you can think about efficiency, which is valuable in real projects.

Self-Check

"What if the input arrays were already sorted? How would that affect the time complexity of np.union1d()?"

Practice

(1/5)
1. What does the function np.union1d() do when applied to two arrays?
easy
A. Finds the common elements between two arrays
B. Combines two arrays and returns unique sorted elements
C. Concatenates two arrays without removing duplicates
D. Sorts a single array in descending order

Solution

  1. Step 1: Understand the purpose of np.union1d()

    This function merges two arrays and removes any duplicate values.
  2. Step 2: Check the output characteristics

    The result is sorted and contains only unique elements from both arrays.
  3. Final Answer:

    Combines two arrays and returns unique sorted elements -> Option B
  4. Quick Check:

    Union = unique sorted merge [OK]
Hint: Think of union as merging without duplicates [OK]
Common Mistakes:
  • Confusing union with intersection
  • Assuming duplicates remain
  • Thinking it sorts in descending order
2. Which of the following is the correct syntax to find the union of arrays a and b using numpy?
easy
A. np.union1d[a, b]
B. np.union(a, b)
C. np.union_1d(a, b)
D. np.union1d(a, b)

Solution

  1. Step 1: Recall the correct function name

    The correct numpy function is np.union1d() with parentheses.
  2. Step 2: Check syntax details

    Arguments are passed inside parentheses, not square brackets, and spelling must be exact.
  3. Final Answer:

    np.union1d(a, b) -> Option D
  4. Quick Check:

    Correct function call syntax [OK]
Hint: Use parentheses and exact function name np.union1d() [OK]
Common Mistakes:
  • Using square brackets instead of parentheses
  • Misspelling the function name
  • Using a non-existent function np.union
3. What is the output of the following code?
import numpy as np
x = np.array([1, 3, 5])
y = np.array([3, 4, 5, 6])
result = np.union1d(x, y)
print(result)
medium
A. [1 3 4 5 6]
B. [1 3 5]
C. [3 4 5 6]
D. [1 3 5 3 4 5 6]

Solution

  1. Step 1: Identify unique elements from both arrays

    Array x has [1, 3, 5], array y has [3, 4, 5, 6]. The union combines all unique values.
  2. Step 2: Sort and remove duplicates

    Unique elements combined are [1, 3, 4, 5, 6], sorted in ascending order.
  3. Final Answer:

    [1 3 4 5 6] -> Option A
  4. Quick Check:

    Union = unique sorted merge [OK]
Hint: Union merges unique sorted elements from both arrays [OK]
Common Mistakes:
  • Forgetting to remove duplicates
  • Not sorting the result
  • Printing only one array
4. The following code throws an error. What is the mistake?
import numpy as np
arr1 = [1, 2, 3]
arr2 = [3, 4, 5]
result = np.union1d(arr1 arr2)
print(result)
medium
A. Missing comma between arr1 and arr2 in function call
B. np.union1d() does not accept lists as input
C. np.union1d() requires arrays to be sorted first
D. print() function is used incorrectly

Solution

  1. Step 1: Check function call syntax

    The function call np.union1d(arr1 arr2) is missing a comma between arguments.
  2. Step 2: Confirm input types and print usage

    np.union1d accepts lists or arrays, and print() is correctly used.
  3. Final Answer:

    Missing comma between arr1 and arr2 in function call -> Option A
  4. Quick Check:

    Comma separates arguments [OK]
Hint: Check commas between function arguments [OK]
Common Mistakes:
  • Omitting commas between arguments
  • Thinking input must be numpy arrays only
  • Assuming print() causes error
5. You have two datasets of customer IDs:
dataset1 = np.array([101, 102, 103, 104])
dataset2 = np.array([103, 104, 105, 106])

You want to create a combined list of all unique customer IDs sorted in ascending order. Which code snippet correctly achieves this?
hard
A. combined = np.intersect1d(dataset1, dataset2)
B. combined = np.concatenate((dataset1, dataset2))
C. combined = np.union1d(dataset1, dataset2)
D. combined = np.sort(np.append(dataset1, dataset2))

Solution

  1. Step 1: Understand the goal

    We want all unique customer IDs from both datasets, sorted ascending.
  2. Step 2: Evaluate each option

    combined = np.union1d(dataset1, dataset2) uses np.union1d which merges and sorts unique elements. combined = np.concatenate((dataset1, dataset2)) concatenates but keeps duplicates. combined = np.intersect1d(dataset1, dataset2) finds only common IDs. combined = np.sort(np.append(dataset1, dataset2)) appends and sorts but does not remove duplicates.
  3. Final Answer:

    combined = np.union1d(dataset1, dataset2) -> Option C
  4. Quick Check:

    Union merges unique sorted elements [OK]
Hint: Use np.union1d to merge unique sorted values [OK]
Common Mistakes:
  • Using concatenate without removing duplicates
  • Using intersect1d which finds only common elements
  • Appending and sorting without removing duplicates