Jump into concepts and practice - no test required
or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Using np.genfromtxt() to Handle Missing Data
📖 Scenario: You work in a small store that tracks daily sales data in a CSV file. Sometimes, some sales numbers are missing because the cashier forgot to enter them. You want to load this data into Python and handle the missing values properly so you can analyze the sales.
🎯 Goal: Load the sales data from a CSV file using np.genfromtxt() so that missing values are handled as np.nan. Then, calculate the average sales ignoring the missing values.
📋 What You'll Learn
Use np.genfromtxt() to load data from a CSV string with missing values
Set the correct parameters to handle missing data as np.nan
Calculate the average sales ignoring missing values using np.nanmean()
Print the average sales as the final output
💡 Why This Matters
🌍 Real World
Stores, labs, and businesses often collect data with missing entries. Handling missing data correctly is important for accurate analysis.
💼 Career
Data scientists and analysts frequently use <code>np.genfromtxt()</code> to load imperfect datasets and prepare them for analysis.
Progress0 / 4 steps
1
Create the sales data CSV string
Create a variable called sales_data that contains this exact CSV string with missing values represented by empty fields: 10,20,30\n40,,60\n70,80,
NumPy
Hint
Use triple quotes or escaped newlines to create the multiline string exactly as shown.
2
Set the delimiter configuration
Create a variable called delimiter and set it to the string "," to specify the CSV delimiter.
NumPy
Hint
The delimiter for CSV files is usually a comma.
3
Load the data using np.genfromtxt() handling missing values
Use np.genfromtxt() to load the data from sales_data with the delimiter variable. Set dtype to float and missing_values to an empty string "" so missing entries become np.nan. Store the result in a variable called sales_array.
NumPy
Hint
Use sales_data.splitlines() to pass the CSV lines to np.genfromtxt().
4
Calculate and print the average sales ignoring missing values
Calculate the average of sales_array ignoring np.nan values using np.nanmean(). Store the result in average_sales. Then print average_sales.
NumPy
Hint
Use np.nanmean() to ignore np.nan values when calculating the average.
Practice
(1/5)
1. What is the main purpose of using np.genfromtxt() in data loading?
easy
A. To load data files while handling missing values automatically
B. To save data files with missing values
C. To visualize data with missing values
D. To delete rows with missing values from a file
Solution
Step 1: Understand the function's purpose
np.genfromtxt() is designed to read text files and handle missing data gracefully.
Step 2: Compare options with function role
Only To load data files while handling missing values automatically correctly states it loads data files and manages missing values automatically.
Final Answer:
To load data files while handling missing values automatically -> Option A
Quick Check:
Purpose of np.genfromtxt() = Load with missing data handled [OK]
Hint: Remember: genfromtxt reads files and fills missing data [OK]
Common Mistakes:
Confusing loading with saving data
Thinking it visualizes data
Assuming it deletes missing data rows automatically
2. Which of the following is the correct way to specify missing values as empty strings when using np.genfromtxt()?
easy
A. np.genfromtxt('data.csv', missing_values='')
B. np.genfromtxt('data.csv', missing_values=null)
C. np.genfromtxt('data.csv', missing_values=[''])
D. np.genfromtxt('data.csv', missing_values='NA')
Solution
Step 1: Check the parameter type for missing_values
The missing_values parameter expects a list or set of strings representing missing data markers.
Step 2: Identify correct syntax for empty string
Empty string must be inside a list: [''] to be recognized as missing.
Final Answer:
np.genfromtxt('data.csv', missing_values=['']) -> Option C
Quick Check:
missing_values needs list for empty string [OK]
Hint: Use a list for missing_values even if one item [OK]
Common Mistakes:
Passing empty string directly without list
Using null which disables missing value detection
Confusing 'NA' with empty string
3. What will be the output of this code snippet?
import numpy as np
from io import StringIO
text = '1,2,\n4,,6'
data = np.genfromtxt(StringIO(text), delimiter=',', filling_values=-1)
print(data)
medium
A. [[1 2 nan]
[4 nan 6]]
B. [1. 2. -1. 4. -1. 6.]
C. [1 2 nan 4 nan 6]
D. [[1. 2. -1.]
[4. -1. 6.]]
Solution
Step 1: Understand input and parameters
The input text has two rows with missing values (empty fields). The delimiter is ',', and missing values are replaced by -1.
Step 2: Predict output array shape and values
Output is a 2D array with missing values replaced by -1, so first row: [1, 2, -1], second row: [4, -1, 6].
Final Answer:
[[1. 2. -1.]
[4. -1. 6.]] -> Option D
Quick Check:
Missing replaced by -1 in 2D array [OK]
Hint: Missing values become filling_values in 2D arrays [OK]
Common Mistakes:
Expecting 1D array instead of 2D
Confusing nan with filling_values
Ignoring delimiter effect on shape
4. Identify the error in this code snippet that tries to load data with missing values:
import numpy as np
np.genfromtxt('data.csv', delimiter=',', missing_values='NA', filling_values=0)
medium
A. filling_values must be a string, not an integer
B. missing_values should be a list, not a string
C. delimiter cannot be a comma
D. np.genfromtxt cannot handle missing values
Solution
Step 1: Check parameter types
missing_values expects a list or set of strings, not a single string.
Step 2: Validate other parameters
filling_values=0 is valid, and delimiter=',' is correct for CSV files.
Final Answer:
missing_values should be a list, not a string -> Option B
Quick Check:
missing_values needs list/set [OK]
Hint: Always wrap missing_values in a list or set [OK]
Common Mistakes:
Passing string directly instead of list
Thinking filling_values must be string
Misunderstanding delimiter usage
5. You have a CSV file with numeric data and missing values marked as 'NA' and empty strings. You want to load it using np.genfromtxt() so that all missing values become -999. Which is the correct way to do this?
hard
A. np.genfromtxt('file.csv', delimiter=',', missing_values=['NA', ''], filling_values=-999)
B. np.genfromtxt('file.csv', delimiter=',', missing_values='NA', filling_values='-999')
C. np.genfromtxt('file.csv', delimiter=',', missing_values=['NA'], filling_values=null)
D. np.genfromtxt('file.csv', delimiter=',', missing_values=[''], filling_values=0)
Solution
Step 1: Specify all missing value markers
Both 'NA' and empty strings '' must be included in a list for missing_values.
Step 2: Set filling_values to -999
Use filling_values=-999 to replace all missing entries with -999.
Final Answer:
np.genfromtxt('file.csv', delimiter=',', missing_values=['NA', ''], filling_values=-999) -> Option A
Quick Check:
List all missing markers and set filling_values [OK]
Hint: List all missing markers, set filling_values to desired number [OK]