Jump into concepts and practice - no test required
or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
SciPy with Pandas for data handling
📖 Scenario: You work as a data analyst for a small company. You have collected sales data for different products over several months. You want to use Pandas to organize the data and SciPy to calculate some statistics like the mean and median sales.
🎯 Goal: Build a simple program that creates a Pandas DataFrame with sales data, sets a threshold for filtering, uses SciPy to calculate statistics on filtered data, and finally prints the results.
📋 What You'll Learn
Create a Pandas DataFrame with exact sales data
Define a sales threshold variable
Use SciPy to calculate mean and median sales above the threshold
Print the calculated mean and median values
💡 Why This Matters
🌍 Real World
Data analysts often combine Pandas for data handling and SciPy for statistical calculations to understand business data better.
💼 Career
Knowing how to filter data and calculate statistics is essential for roles like data analyst, business analyst, and data scientist.
Progress0 / 4 steps
1
Create the sales data DataFrame
Create a Pandas DataFrame called sales_data with these exact columns and values: 'Product': ['Apples', 'Bananas', 'Cherries', 'Dates', 'Elderberries'], 'Month': ['Jan', 'Jan', 'Feb', 'Feb', 'Mar'], 'Sales': [150, 200, 50, 300, 120].
SciPy
Hint
Use pd.DataFrame with a dictionary where keys are column names and values are lists of data.
2
Set the sales threshold
Create a variable called sales_threshold and set it to 100.
SciPy
Hint
Just assign the number 100 to the variable sales_threshold.
3
Calculate mean and median sales above threshold
Import scipy.stats as stats. Then create a variable called filtered_sales that contains only the 'Sales' values from sales_data where sales are greater than sales_threshold. Use stats.tmean to calculate the mean of filtered_sales and store it in mean_sales. Use stats.tmedian to calculate the median of filtered_sales and store it in median_sales.
SciPy
Hint
Use sales_data.loc to filter rows where 'Sales' is greater than sales_threshold. Then use stats.tmean and stats.tmedian on the filtered sales.
4
Print the mean and median sales
Write two print statements to display the mean and median sales. Use print(f"Mean sales: {mean_sales}") and print(f"Median sales: {median_sales}").
SciPy
Hint
Use print with f-strings to show the values of mean_sales and median_sales.
Practice
(1/5)
1. What is the main reason to use SciPy together with Pandas in data analysis?
easy
A. SciPy provides advanced math and stats functions, while Pandas organizes data in tables.
B. Pandas is used only for visualization, SciPy handles all data storage.
C. SciPy replaces Pandas for data cleaning tasks.
D. Pandas is used to write code, SciPy runs the code faster.
Solution
Step 1: Understand roles of Pandas and SciPy
Pandas organizes data into tables called DataFrames, making it easy to handle data.
Step 2: Identify SciPy's role
SciPy offers math and statistics tools to analyze data prepared by Pandas.
Final Answer:
SciPy provides advanced math and stats functions, while Pandas organizes data in tables. -> Option A
Quick Check:
Data organization = Pandas, Analysis = SciPy [OK]
Hint: Remember: Pandas for tables, SciPy for math [OK]
Common Mistakes:
Thinking Pandas does advanced stats alone
Confusing SciPy as a data storage tool
Believing SciPy replaces Pandas for cleaning
2. Which of the following is the correct way to import SciPy's stats module and Pandas in Python?
easy
A. from scipy import stats; import pandas as pd
B. import scipy.stats; import pandas as pandas
C. from scipy.stats import stats; import pandas as pd
D. import scipy.stats as sp; import pandas as pd
Solution
Step 1: Check common import styles
Using 'from scipy import stats' imports the stats module directly, which is common and clear.
Step 2: Verify Pandas import
Importing pandas as 'pd' is the standard alias used in data science.
Final Answer:
from scipy import stats; import pandas as pd -> Option A
Quick Check:
Standard imports = from scipy import stats, import pandas as pd [OK]
Hint: Use 'from scipy import stats' and 'import pandas as pd' [OK]
Common Mistakes:
Using wrong alias for pandas
Importing scipy.stats without alias or direct import
Mixing import styles incorrectly
3. Given the code below, what will be the output?
import pandas as pd
from scipy import stats
data = {'score': [10, 20, 20, 30, 40]}
df = pd.DataFrame(data)
mode_result = stats.mode(df['score'])
print(mode_result.mode[0])
medium
A. 10
B. 30
C. 20
D. 40
Solution
Step 1: Understand the data
The 'score' column has values [10, 20, 20, 30, 40]. The number 20 appears twice, others once.
Step 2: Apply stats.mode
stats.mode finds the most frequent value, which is 20 here.
Final Answer:
20 -> Option C
Quick Check:
Most frequent value = 20 [OK]
Hint: Mode is the most frequent value in the list [OK]
Common Mistakes:
Choosing the first value instead of mode
Confusing mean or median with mode
Not accessing .mode[0] correctly
4. Identify the error in the following code snippet:
import pandas as pd
from scipy import stats
data = {'values': [1, 2, 3, 4, 5]}
df = pd.DataFrame(data)
result = stats.mean(df['values'])
print(result)
medium
A. DataFrame creation syntax is incorrect.
B. stats.mean does not exist; use numpy.mean or pandas mean method instead.
C. The print statement is missing parentheses.
D. The import statement for pandas is wrong.
Solution
Step 1: Check function availability in SciPy
SciPy's stats module does not have a 'mean' function; mean is in numpy or pandas.
Step 2: Identify correct function usage
Use df['values'].mean() or numpy.mean(df['values']) instead.
Final Answer:
stats.mean does not exist; use numpy.mean or pandas mean method instead. -> Option B
Quick Check:
stats.mean missing, use pandas or numpy mean [OK]
Hint: Use pandas or numpy for mean, not stats.mean [OK]
Common Mistakes:
Assuming all stats functions exist in SciPy
Ignoring error messages about missing attributes
Confusing pandas and SciPy function locations
5. You have a Pandas DataFrame with a column 'height' containing some missing values (NaN). You want to fill these missing values with the median height calculated using SciPy. Which code snippet correctly does this?
hard
A. from scipy import stats
median_height = stats.median(df['height'])
df['height'] = df['height'].fillna(median_height)
B. from scipy import stats
median_height = stats.median(df['height'].dropna())
df['height'] = df['height'].fillna(median_height)
C. from scipy import stats
median_height = stats.mode(df['height'].dropna()).mode[0]
df['height'] = df['height'].fillna(median_height)
D. from scipy import stats
median_height = stats.scoreatpercentile(df['height'].dropna(), 50)
df['height'] = df['height'].fillna(median_height)
Solution
Step 1: Identify correct SciPy function for median
SciPy's stats module does not have 'median', but 'scoreatpercentile' can find the 50th percentile (median).
Step 2: Handle missing values correctly
Drop NaN values before calculating median, then fill NaNs with this median.
Final Answer:
from scipy import stats
median_height = stats.scoreatpercentile(df['height'].dropna(), 50)
df['height'] = df['height'].fillna(median_height) -> Option D
Quick Check:
Median via scoreatpercentile, fillna with median [OK]
Hint: Use scoreatpercentile for median in SciPy [OK]