We use SciPy with scikit-learn pipelines to clean and prepare data step-by-step, then train a model easily. It helps keep everything organized and repeatable.
SciPy with scikit-learn pipeline
Start learning this pattern below
Jump into concepts and practice - no test required
from sklearn.pipeline import Pipeline from sklearn.preprocessing import FunctionTransformer from sklearn.linear_model import LogisticRegression import scipy pipeline = Pipeline([ ('scipy_transform', FunctionTransformer(your_scipy_function)), ('model', LogisticRegression()) ]) pipeline.fit(X_train, y_train) predictions = pipeline.predict(X_test)
Pipeline lets you chain steps: first data changes, then model training.
FunctionTransformer wraps SciPy functions so they work inside the pipeline.
from sklearn.pipeline import Pipeline from sklearn.preprocessing import FunctionTransformer from sklearn.linear_model import LogisticRegression import numpy as np import scipy.stats def log_transform(X): return np.log1p(X) pipeline = Pipeline([ ('log', FunctionTransformer(log_transform)), ('model', LogisticRegression()) ])
from sklearn.pipeline import Pipeline from sklearn.preprocessing import FunctionTransformer from sklearn.linear_model import LogisticRegression import scipy.ndimage def smooth_data(X): return scipy.ndimage.gaussian_filter(X, sigma=1) pipeline = Pipeline([ ('smooth', FunctionTransformer(smooth_data)), ('model', LogisticRegression()) ])
This program loads iris data, normalizes features using SciPy's z-score inside a pipeline, trains logistic regression, and prints predictions.
from sklearn.pipeline import Pipeline from sklearn.preprocessing import FunctionTransformer from sklearn.linear_model import LogisticRegression from sklearn.datasets import load_iris from sklearn.model_selection import train_test_split import numpy as np import scipy.stats def zscore_transform(X): return scipy.stats.zscore(X, axis=0) # Load data iris = load_iris() X, y = iris.data, iris.target # Use only two classes for logistic regression X = X[y != 2] y = y[y != 2] # Split data X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42) # Create pipeline with SciPy z-score normalization and logistic regression pipeline = Pipeline([ ('zscore', FunctionTransformer(zscore_transform)), ('model', LogisticRegression()) ]) # Train model pipeline.fit(X_train, y_train) # Predict predictions = pipeline.predict(X_test) # Print predictions print(predictions)
FunctionTransformer expects functions that take and return arrays.
Make sure SciPy functions do not change the shape unexpectedly.
Pipelines help avoid data leakage by applying the same transformations to train and test data.
SciPy functions can be used inside scikit-learn pipelines with FunctionTransformer.
Pipelines keep data processing and modeling steps organized and repeatable.
This approach helps beginners build clean, understandable data science workflows.
Practice
Pipeline in scikit-learn when combined with SciPy functions?Solution
Step 1: Understand the purpose of Pipeline
A Pipeline in scikit-learn is designed to chain multiple steps like data transformation and modeling into a single object.Step 2: Recognize the benefit of combining SciPy functions
Using SciPy functions inside a Pipeline via FunctionTransformer keeps the workflow organized and repeatable.Final Answer:
It organizes data processing and modeling steps into one repeatable workflow. -> Option AQuick Check:
Pipeline = Organized workflow [OK]
- Thinking Pipeline improves accuracy automatically
- Assuming Pipeline removes need for data cleaning
- Believing Pipeline runs without imports
scipy_func inside a scikit-learn pipeline using FunctionTransformer?Solution
Step 1: Understand FunctionTransformer usage
FunctionTransformer takes a function as an argument without calling it (no parentheses).Step 2: Identify correct pipeline syntax
The pipeline step should be ('transform', FunctionTransformer(scipy_func)) to wrap the function properly.Final Answer:
Pipeline([('transform', FunctionTransformer(scipy_func)), ('model', LogisticRegression())]) -> Option BQuick Check:
FunctionTransformer(function) no parentheses [OK]
- Calling the function inside FunctionTransformer
- Passing function directly without FunctionTransformer
- Using FunctionTransformer without function argument
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import FunctionTransformer
import numpy as np
def add_one(X):
return X + 1
pipe = Pipeline([
('add', FunctionTransformer(add_one)),
])
X = np.array([1, 2, 3])
result = pipe.transform(X)
print(result)Solution
Step 1: Understand FunctionTransformer behavior
FunctionTransformer applies the functionadd_oneto input data during transform.Step 2: Apply the function to input array
Input array [1, 2, 3] plus 1 becomes [2, 3, 4].Final Answer:
[2 3 4] -> Option DQuick Check:
Input + 1 = Output [OK]
- Assuming pipeline has no transform method
- Forgetting function adds 1
- Confusing fit and transform methods
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import FunctionTransformer
import numpy as np
def multiply_by_two(X):
return X * 2
pipe = Pipeline([
('mult', FunctionTransformer(multiply_by_two())),
])
X = np.array([1, 2, 3])
result = pipe.transform(X)
print(result)Solution
Step 1: Check FunctionTransformer argument
FunctionTransformer expects a function, not the result of a function call.Step 2: Identify the error in code
Code calls multiply_by_two() immediately, which causes an error because it returns an array, not a function.Final Answer:
Calling multiply_by_two() instead of passing the function -> Option CQuick Check:
Pass function, don't call it [OK]
- Calling function instead of passing it
- Assuming pipeline needs a model step
- Confusing transform with fit_transform
Solution
Step 1: Use FunctionTransformer correctly with SciPy function
Pass the functionstats.zscorewithout calling it, wrapped by FunctionTransformer.Step 2: Ensure LogisticRegression is instantiated
UseLogisticRegression()with parentheses to create the model instance.Final Answer:
Code snippet with FunctionTransformer(stats.zscore) and LogisticRegression() -> Option AQuick Check:
FunctionTransformer(function) + model instance [OK]
- Calling SciPy function instead of passing it
- Not instantiating LogisticRegression
- Passing function call to FunctionTransformer
