Jump into concepts and practice - no test required
or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Build a SciPy and scikit-learn Pipeline for Data Transformation and Modeling
📖 Scenario: You are working as a data analyst for a small company. You have some data about customers' ages and incomes, and you want to predict their spending score. To do this, you will prepare the data using SciPy and then build a simple model using scikit-learn's pipeline feature.
🎯 Goal: Create a data dictionary with customer data, set up a configuration variable for a threshold, build a scikit-learn pipeline that uses a SciPy function to transform data and a simple model, then output the transformed data and model predictions.
📋 What You'll Learn
Create a dictionary called customer_data with keys 'age' and 'income' and the exact lists of values provided.
Create a variable called income_threshold and set it to the exact value 50000.
Build a scikit-learn pipeline named pipeline that uses a SciPy function to apply a logarithm transformation to income and a simple linear regression model.
Print the transformed income data and the model predictions exactly as specified.
💡 Why This Matters
🌍 Real World
Data scientists often need to preprocess data using mathematical functions from libraries like SciPy before feeding it into machine learning models. Pipelines help organize these steps cleanly.
💼 Career
Understanding how to combine data transformations and models in a pipeline is a key skill for data analysts and data scientists working on predictive modeling tasks.
Progress0 / 4 steps
1
DATA SETUP: Create the customer data dictionary
Create a dictionary called customer_data with two keys: 'age' and 'income'. Set 'age' to the list [25, 32, 47, 51, 62] and 'income' to the list [40000, 52000, 61000, 58000, 72000].
SciPy
Hint
Use curly braces to create a dictionary. The keys are 'age' and 'income'. Assign the exact lists to each key.
2
CONFIGURATION: Set the income threshold
Create a variable called income_threshold and set it to the integer 50000.
SciPy
Hint
Just assign the number 50000 to the variable named income_threshold.
3
CORE LOGIC: Build the SciPy and scikit-learn pipeline
Import FunctionTransformer from sklearn.preprocessing, LinearRegression from sklearn.linear_model, and log from scipy.special. Then create a pipeline called pipeline that first applies the logarithm transformation to the income data using FunctionTransformer with log, and then fits a LinearRegression model.
SciPy
Hint
Use Pipeline with two steps: a FunctionTransformer that applies log, and a LinearRegression model.
4
OUTPUT: Transform income and predict spending score
Use the pipeline to fit the model using the income data reshaped as a 2D array. Then print the transformed income data after the log transform step and print the predictions from the linear regression model. Use print(transformed_income) and print(predictions) exactly.
SciPy
Hint
Use np.array and reshape(-1, 1) to prepare income data. Fit the pipeline with income and age. Use pipeline.named_steps['log_transform'].transform() to get transformed income. Use pipeline.predict() for predictions. Print both results.
Practice
(1/5)
1. What is the main benefit of using a Pipeline in scikit-learn when combined with SciPy functions?
easy
A. It organizes data processing and modeling steps into one repeatable workflow.
B. It automatically improves model accuracy without tuning.
C. It replaces the need for any data cleaning.
D. It allows running code without importing any libraries.
Solution
Step 1: Understand the purpose of Pipeline
A Pipeline in scikit-learn is designed to chain multiple steps like data transformation and modeling into a single object.
Step 2: Recognize the benefit of combining SciPy functions
Using SciPy functions inside a Pipeline via FunctionTransformer keeps the workflow organized and repeatable.
Final Answer:
It organizes data processing and modeling steps into one repeatable workflow. -> Option A
Quick Check:
Pipeline = Organized workflow [OK]
Hint: Pipelines bundle steps for easy reuse and clarity [OK]
Common Mistakes:
Thinking Pipeline improves accuracy automatically
Assuming Pipeline removes need for data cleaning
Believing Pipeline runs without imports
2. Which of the following is the correct way to include a SciPy function scipy_func inside a scikit-learn pipeline using FunctionTransformer?
easy
A. Pipeline([('transform', scipy_func), ('model', LogisticRegression())])
B. Pipeline([('transform', FunctionTransformer(scipy_func)), ('model', LogisticRegression())])
C. Pipeline([('transform', FunctionTransformer()), ('model', LogisticRegression())])
D. Pipeline([('transform', FunctionTransformer(scipy_func())), ('model', LogisticRegression())])
Solution
Step 1: Understand FunctionTransformer usage
FunctionTransformer takes a function as an argument without calling it (no parentheses).
Step 2: Identify correct pipeline syntax
The pipeline step should be ('transform', FunctionTransformer(scipy_func)) to wrap the function properly.
Final Answer:
Pipeline([('transform', FunctionTransformer(scipy_func)), ('model', LogisticRegression())]) -> Option B
Quick Check:
FunctionTransformer(function) no parentheses [OK]
Hint: Pass function name, not call, to FunctionTransformer [OK]
Common Mistakes:
Calling the function inside FunctionTransformer
Passing function directly without FunctionTransformer
Using FunctionTransformer without function argument
3. What will be the output of the following code snippet?
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import FunctionTransformer
import numpy as np
def add_one(X):
return X + 1
pipe = Pipeline([
('add', FunctionTransformer(add_one)),
])
X = np.array([1, 2, 3])
result = pipe.transform(X)
print(result)
medium
A. Error: Pipeline has no transform method
B. [1 2 3]
C. [0 1 2]
D. [2 3 4]
Solution
Step 1: Understand FunctionTransformer behavior
FunctionTransformer applies the function add_one to input data during transform.
Step 2: Apply the function to input array
Input array [1, 2, 3] plus 1 becomes [2, 3, 4].
Final Answer:
[2 3 4] -> Option D
Quick Check:
Input + 1 = Output [OK]
Hint: FunctionTransformer applies function on transform call [OK]
Common Mistakes:
Assuming pipeline has no transform method
Forgetting function adds 1
Confusing fit and transform methods
4. Identify the error in this pipeline code using a SciPy function inside FunctionTransformer:
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import FunctionTransformer
import numpy as np
def multiply_by_two(X):
return X * 2
pipe = Pipeline([
('mult', FunctionTransformer(multiply_by_two())),
])
X = np.array([1, 2, 3])
result = pipe.transform(X)
print(result)
medium
A. Using transform instead of fit_transform
B. Missing import for numpy
C. Calling multiply_by_two() instead of passing the function
D. Pipeline missing a model step
Solution
Step 1: Check FunctionTransformer argument
FunctionTransformer expects a function, not the result of a function call.
Step 2: Identify the error in code
Code calls multiply_by_two() immediately, which causes an error because it returns an array, not a function.
Final Answer:
Calling multiply_by_two() instead of passing the function -> Option C
Quick Check:
Pass function, don't call it [OK]
Hint: Pass function name, avoid parentheses in FunctionTransformer [OK]
Common Mistakes:
Calling function instead of passing it
Assuming pipeline needs a model step
Confusing transform with fit_transform
5. You want to build a pipeline that first applies a SciPy function to normalize data, then fits a logistic regression model. Which of the following code snippets correctly implements this?
hard
A. from sklearn.pipeline import Pipeline
from sklearn.preprocessing import FunctionTransformer
from sklearn.linear_model import LogisticRegression
import scipy.stats as stats
pipe = Pipeline([
('normalize', FunctionTransformer(stats.zscore)),
('model', LogisticRegression())
])
B. from sklearn.pipeline import Pipeline
from sklearn.preprocessing import FunctionTransformer
from sklearn.linear_model import LogisticRegression
import scipy.stats as stats
pipe = Pipeline([
('normalize', stats.zscore()),
('model', LogisticRegression())
])
C. from sklearn.pipeline import Pipeline
from sklearn.preprocessing import FunctionTransformer
from sklearn.linear_model import LogisticRegression
import scipy.stats as stats
pipe = Pipeline([
('normalize', FunctionTransformer(stats.zscore())),
('model', LogisticRegression())
])
D. from sklearn.pipeline import Pipeline
from sklearn.preprocessing import FunctionTransformer
from sklearn.linear_model import LogisticRegression
import scipy.stats as stats
pipe = Pipeline([
('normalize', FunctionTransformer(stats.zscore)),
('model', LogisticRegression)
])
Solution
Step 1: Use FunctionTransformer correctly with SciPy function
Pass the function stats.zscore without calling it, wrapped by FunctionTransformer.
Step 2: Ensure LogisticRegression is instantiated
Use LogisticRegression() with parentheses to create the model instance.
Final Answer:
Code snippet with FunctionTransformer(stats.zscore) and LogisticRegression() -> Option A
Quick Check:
FunctionTransformer(function) + model instance [OK]
Hint: Wrap function, instantiate model with parentheses [OK]