Bird
Raised Fist0
NumPydata~5 mins

Generating random samples in NumPy

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Introduction

We use random samples to create example data or simulate real-world situations where outcomes vary.

Testing a new data analysis method with fake data
Simulating dice rolls or coin flips in games
Creating random customer data for practice
Sampling a few items from a large dataset randomly
Syntax
NumPy
numpy.random.choice(a, size=None, replace=True, p=None)

a is the array or number of items to choose from.

size is how many samples you want.

Examples
Randomly picks 3 numbers from 0 to 4 (5 numbers total).
NumPy
import numpy as np
np.random.choice(5, size=3)
Randomly picks 2 different colors from the list without repeats.
NumPy
np.random.choice(['red', 'blue', 'green'], size=2, replace=False)
Randomly picks 4 numbers from the list with given probabilities.
NumPy
np.random.choice([10, 20, 30], size=4, replace=True, p=[0.1, 0.7, 0.2])
Sample Program

This program shows three ways to generate random samples using numpy:

  • Random numbers from 0 to 9
  • Unique fruits from a list
  • Numbers with specific chances
NumPy
import numpy as np

# Pick 5 random numbers from 0 to 9
samples = np.random.choice(10, size=5)
print('Random samples:', samples)

# Pick 3 unique fruits
fruits = ['apple', 'banana', 'cherry', 'date']
unique_fruits = np.random.choice(fruits, size=3, replace=False)
print('Unique fruits:', unique_fruits)

# Pick 6 numbers with custom probabilities
numbers = [1, 2, 3]
probabilities = [0.5, 0.3, 0.2]
samples_prob = np.random.choice(numbers, size=6, p=probabilities)
print('Samples with probabilities:', samples_prob)
OutputSuccess
Important Notes

Setting replace=False means no repeats in samples.

Probabilities p must add up to 1.

Random results change each time unless you set a random seed.

Summary

Use numpy.random.choice to pick random samples from data.

You can control sample size, repetition, and probabilities.

Random sampling helps simulate and test data scenarios easily.

Practice

(1/5)
1. What does the numpy.random.choice function do?
easy
A. It calculates the mean of an array.
B. It selects random elements from a given array or list.
C. It sorts an array in ascending order.
D. It reshapes an array into a new shape.

Solution

  1. Step 1: Understand the function purpose

    numpy.random.choice is designed to pick random elements from a given array or list.
  2. Step 2: Compare with other options

    Sorting, calculating mean, and reshaping are different numpy functions, not related to random sampling.
  3. Final Answer:

    It selects random elements from a given array or list. -> Option B
  4. Quick Check:

    Random sampling = selecting elements randomly [OK]
Hint: Remember: choice means picking randomly from data [OK]
Common Mistakes:
  • Confusing choice with sorting or reshaping functions
  • Thinking it calculates statistics like mean
  • Assuming it modifies array shape
2. Which of the following is the correct syntax to randomly select 3 elements from array arr without replacement using numpy?
easy
A. numpy.random.choice(arr, 3, replace=True)
B. numpy.choice(arr, size=3, replace=False)
C. numpy.random.choice(arr, size=3, replace=False)
D. numpy.random.choice(arr, size=3, replace=True)

Solution

  1. Step 1: Identify correct function and parameters

    The function is numpy.random.choice. To select 3 elements without replacement, use size=3 and replace=False.
  2. Step 2: Check each option

    numpy.random.choice(arr, size=3, replace=False) uses correct function and parameters. numpy.random.choice(arr, 3, replace=True) uses replacement True (wrong). numpy.choice(arr, size=3, replace=False) uses wrong function name. numpy.random.choice(arr, size=3, replace=True) uses replacement True (wrong).
  3. Final Answer:

    numpy.random.choice(arr, size=3, replace=False) -> Option C
  4. Quick Check:

    Correct syntax = choice + size + replace=False [OK]
Hint: Use replace=False to avoid repeated picks [OK]
Common Mistakes:
  • Using replace=True when no repeats wanted
  • Misspelling function name as numpy.choice
  • Passing size as positional without keyword
3. What is the output of this code?
import numpy as np
np.random.seed(0)
arr = np.array([10, 20, 30, 40])
sample = np.random.choice(arr, size=2, replace=False)
sample_sorted = np.sort(sample)
sample_sorted.tolist()
medium
A. [10, 40]
B. [10, 20]
C. [20, 40]
D. [30, 40]

Solution

  1. Step 1: Understand random seed and choice

    Setting seed to 0 fixes randomness. Using choice with size=2 and replace=False picks 2 unique elements from [10,20,30,40].
  2. Step 2: Determine chosen elements and sort

    With seed 0, the chosen elements are [10, 40]. Sorting gives [10, 40].
  3. Final Answer:

    [10, 40] -> Option A
  4. Quick Check:

    Seed 0 + choice + sort = [10, 40] [OK]
Hint: Seed fixes output; sort to order chosen elements [OK]
Common Mistakes:
  • Ignoring seed and expecting different output
  • Not sorting before converting to list
  • Assuming replacement allows duplicates
4. The following code throws an error. What is the cause?
import numpy as np
arr = np.array([1, 2, 3])
sample = np.random.choice(arr, size=5, replace=False)
medium
A. Incorrect function name used.
B. Array contains integers instead of floats.
C. Missing import statement for numpy.
D. Size is larger than array length without replacement.

Solution

  1. Step 1: Analyze parameters and array size

    The array has 3 elements, but size=5 is requested without replacement.
  2. Step 2: Understand replacement=False effect

    Without replacement, you cannot pick more elements than exist. This causes a ValueError.
  3. Final Answer:

    Size is larger than array length without replacement. -> Option D
  4. Quick Check:

    Sampling more than available without replace=False causes error [OK]
Hint: Check if sample size > array length when replace=False [OK]
Common Mistakes:
  • Assuming replacement=True by default
  • Ignoring array length vs sample size
  • Thinking data type causes error
5. You want to simulate rolling a weighted 6-sided die 10 times using numpy, where side 6 is twice as likely as others. Which code correctly generates this sample?
hard
A. np.random.choice([1,2,3,4,5,6], size=10, replace=True, p=[1/7,1/7,1/7,1/7,1/7,2/7])
B. np.random.choice([1,2,3,4,5,6], size=10, replace=False, p=[1/6]*6)
C. np.random.choice([1,2,3,4,5,6], size=10, replace=True)
D. np.random.choice([1,2,3,4,5,6], size=10, replace=True, p=[1/6,1/6,1/6,1/6,1/6,1/6])

Solution

  1. Step 1: Understand weighted probabilities

    Side 6 should be twice as likely, so probabilities sum to 1 with side 6 having weight 2/7 and others 1/7 each.
  2. Step 2: Check sampling parameters

    Sampling 10 times with replacement is needed to allow repeats. np.random.choice([1,2,3,4,5,6], size=10, replace=True, p=[1/7,1/7,1/7,1/7,1/7,2/7]) uses correct probabilities and replace=True.
  3. Step 3: Verify other options

    The code with replace=False, p=[1/6]*6 incorrectly prevents repeats needed for multiple rolls. The codes with uniform probabilities (explicit [1/6,1/6,1/6,1/6,1/6,1/6] or none specified) do not weight side 6 twice as likely.
  4. Final Answer:

    np.random.choice([1,2,3,4,5,6], size=10, replace=True, p=[1/7,1/7,1/7,1/7,1/7,2/7]) -> Option A
  5. Quick Check:

    Weighted probabilities + replace=True for repeated rolls [OK]
Hint: Use p= with weights summing to 1 and replace=True [OK]
Common Mistakes:
  • Using replace=False for multiple rolls
  • Not setting probabilities for weighted sides
  • Using equal probabilities when weights differ