Bird
Raised Fist0
NumPydata~20 mins

Why random generation matters in NumPy - Challenge Your Understanding

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Challenge - 5 Problems
🎖️
Random Generation Mastery
Get all challenges correct to earn this badge!
Test your skills under time pressure!
❓ Predict Output
intermediate
2:00remaining
Output of numpy random seed effect

What will be the output of the following code snippet?

import numpy as np
np.random.seed(0)
sample1 = np.random.randint(1, 10, 3)
np.random.seed(0)
sample2 = np.random.randint(1, 10, 3)
print(sample1)
print(sample2)
NumPy
import numpy as np
np.random.seed(0)
sample1 = np.random.randint(1, 10, 3)
np.random.seed(0)
sample2 = np.random.randint(1, 10, 3)
print(sample1)
print(sample2)
A
[6 1 4]
[6 1 4]
B
[6 1 4]
[2 8 8]
C
[1 6 4]
[1 6 4]
D
[2 8 8]
[2 8 8]
Attempts:
2 left
💡 Hint

Setting the seed resets the random number generator to the same starting point.

🧠 Conceptual
intermediate
1:30remaining
Why set a random seed in data science?

Why is it important to set a random seed when generating random data in data science projects?

ATo ensure the random numbers are different every time the code runs
BTo make the random number generation reproducible for debugging and sharing results
CTo speed up the random number generation process
DTo increase the randomness quality of the generated numbers
Attempts:
2 left
💡 Hint

Think about sharing your work with others or rerunning your code later.

❓ data_output
advanced
2:00remaining
Effect of different seeds on random samples

What is the shape and content type of the output from this code?

import numpy as np
np.random.seed(42)
sample = np.random.normal(loc=0, scale=1, size=(2,3))
print(type(sample))
print(sample.shape)
NumPy
import numpy as np
np.random.seed(42)
sample = np.random.normal(loc=0, scale=1, size=(2,3))
print(type(sample))
print(sample.shape)
A
<class 'numpy.ndarray'>
(2, 3)
B
<class 'list'>
(2, 3)
C
<class 'numpy.ndarray'>
(3, 2)
D
<class 'list'>
(3, 2)
Attempts:
2 left
💡 Hint

Check the type returned by np.random.normal and the shape argument.

🔧 Debug
advanced
1:30remaining
Identify the error in random number generation

What error will this code produce?

import numpy as np
np.random.seed('seed')
sample = np.random.randint(0, 5, 4)
print(sample)
NumPy
import numpy as np
np.random.seed('seed')
sample = np.random.randint(0, 5, 4)
print(sample)
ASyntaxError: invalid syntax
BValueError: seed must be between 0 and 2**32 - 1
CNo error, prints 4 random integers between 0 and 4
DTypeError: 'str' object cannot be interpreted as an integer
Attempts:
2 left
💡 Hint

Check the type of the argument passed to np.random.seed.

🚀 Application
expert
2:30remaining
Using random generation for train-test split reproducibility

You want to split your dataset into training and testing parts randomly but reproducibly. Which code snippet correctly achieves this?

A
import numpy as np
indices = np.random.permutation(10)
train_idx = indices[:7]
test_idx = indices[7:]
B
import numpy as np
np.random.seed('123')
indices = np.random.permutation(10)
train_idx = indices[:7]
test_idx = indices[7:]
C
import numpy as np
np.random.seed(123)
indices = np.random.permutation(10)
train_idx = indices[:7]
test_idx = indices[7:]
D
import numpy as np
np.random.seed(123)
indices = np.random.randint(0, 10, 10)
train_idx = indices[:7]
test_idx = indices[7:]
Attempts:
2 left
💡 Hint

Think about reproducibility and correct use of random permutation for splitting.

Practice

(1/5)
1. Why is random number generation important in data science?
easy
A. It helps create unpredictable data for testing and simulations.
B. It always produces the same output for every run.
C. It removes the need for data cleaning.
D. It guarantees perfect model accuracy.

Solution

  1. Step 1: Understand the role of randomness

    Random generation creates data that is not predictable, which is useful for testing and simulations.
  2. Step 2: Evaluate the options

    Only 'It helps create unpredictable data for testing and simulations.' correctly states the importance of random generation. Options A, B, and D are incorrect because random generation does not remove the need for data cleaning, does not always produce the same output, and does not guarantee perfect accuracy.
  3. Final Answer:

    It helps create unpredictable data for testing and simulations. -> Option A
  4. Quick Check:

    Random generation importance = unpredictable data [OK]
Hint: Random means unpredictable data for testing [OK]
Common Mistakes:
  • Thinking random data is always the same
  • Assuming random data fixes all errors
  • Believing random data guarantees perfect results
2. Which of the following is the correct way to generate 5 random numbers between 0 and 1 using NumPy?
easy
A. np.random.randn(5)
B. np.random.rand(5)
C. np.random.randint(0, 1, 5)
D. np.random.choice(5)

Solution

  1. Step 1: Review NumPy random functions

    np.random.rand(5) generates 5 random floats between 0 and 1.
  2. Step 2: Check other options

    np.random.randn(5) generates samples from the standard normal distribution (not uniform [0,1)); np.random.randint(0, 1, 5) returns zeros only; np.random.choice(5) picks one number from 0 to 4.
  3. Final Answer:

    np.random.rand(5) -> Option B
  4. Quick Check:

    Correct syntax for 5 random floats = np.random.rand(5) [OK]
Hint: Use np.random.rand(n) for n floats 0 to 1 [OK]
Common Mistakes:
  • Using randint with 0 and 1 returns only zeros
  • Using choice without specifying size returns one value
  • Confusing randn with rand
3. What will be the output of the following code?
import numpy as np
np.random.seed(0)
print(np.random.rand(3))
medium
A. [0.5488135 0.71518937 0.60276338]
B. [0.5488135 0.60276338 0.71518937]
C. [0.37454012 0.95071431 0.73199394]
D. [0.71518937 0.60276338 0.5488135 ]

Solution

  1. Step 1: Understand seed effect

    Setting np.random.seed(0) fixes the random numbers to a known sequence.
  2. Step 2: Check known output for seed 0

    For seed 0, np.random.rand(3) produces [0.5488135 0.71518937 0.60276338]. The other options show different sequences or orders.
  3. Final Answer:

    [0.37454012 0.95071431 0.73199394] -> Option C
  4. Quick Check:

    Seed 0 fixed output = [0.37454012 0.95071431 0.73199394] [OK]
Hint: Seed fixes output; np.random.rand(3) gives 3 floats [OK]
Common Mistakes:
  • Ignoring seed leads to different outputs
  • Mixing order of numbers in output
  • Confusing randint with rand
4. The following code is intended to generate 4 random integers between 1 and 10, but it raises an error. What is the problem?
import numpy as np
np.random.randint(1, 10, size=4, seed=42)
medium
A. The 'seed' argument is not valid in randint function.
B. The 'size' argument should be a tuple, not an integer.
C. The range 1 to 10 is invalid for randint.
D. The function randint does not exist in numpy.

Solution

  1. Step 1: Check randint parameters

    NumPy's randint does not accept a seed parameter directly.
  2. Step 2: Understand how to set seed

    Seed must be set using np.random.seed(42) before calling randint.
  3. Final Answer:

    The 'seed' argument is not valid in randint function. -> Option A
  4. Quick Check:

    Seed set separately, not in randint [OK]
Hint: Set seed with np.random.seed(), not in randint [OK]
Common Mistakes:
  • Passing seed inside randint
  • Using wrong size type
  • Thinking randint is missing
5. You want to simulate rolling a fair six-sided die 1000 times using NumPy. Which code snippet correctly generates this data?
hard
A. np.random.rand(1000) * 6 + 1
B. np.random.randint(0, 6, size=1000)
C. np.random.choice(6, size=1000)
D. np.random.randint(1, 7, size=1000)

Solution

  1. Step 1: Understand die roll range

    A fair six-sided die has values 1 through 6 inclusive.
  2. Step 2: Check code options

    np.random.randint(1, 7, size=1000) uses randint(1,7) which includes 1 and excludes 7, so values 1 to 6 are generated correctly. np.random.rand(1000) * 6 + 1 generates floats, not integers. np.random.choice(6, size=1000) picks from default [0,1,2,3,4,5]. np.random.randint(0, 6, size=1000) generates 0 to 5, which is incorrect.
  3. Final Answer:

    np.random.randint(1, 7, size=1000) -> Option D
  4. Quick Check:

    Correct die roll simulation = randint(1,7) [OK]
Hint: Use randint(1,7) for integers 1 to 6 [OK]
Common Mistakes:
  • Using randint(0,6) gives 0 to 5
  • Using rand() gives floats, not integers
  • Forgetting to set size for multiple rolls