Jump into concepts and practice - no test required
or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
K-means Clustering with SciPy and scikit-learn
📖 Scenario: You work as a data analyst for a small retail company. You want to group customers based on their shopping habits to create better marketing strategies. You will use K-means clustering to find groups of similar customers.
🎯 Goal: Build a simple K-means clustering model using both scipy and scikit-learn libraries. Compare how to set up the data, run the clustering, and get the cluster centers.
📋 What You'll Learn
Create a dataset of customer shopping data as a list of lists.
Set the number of clusters to 2 using a variable.
Use scipy.cluster.vq.kmeans to find cluster centers.
Use sklearn.cluster.KMeans to fit the same data and get cluster centers.
Print the cluster centers from both methods.
💡 Why This Matters
🌍 Real World
K-means clustering helps businesses group customers or products based on features to target marketing or improve services.
💼 Career
Data scientists and analysts often use clustering to find patterns in data without labels, helping in customer segmentation and recommendation systems.
Progress0 / 4 steps
1
Create the customer data
Create a variable called data that holds this exact list of lists: [[1.0, 2.0], [1.5, 1.8], [5.0, 8.0], [8.0, 8.0], [1.0, 0.6], [9.0, 11.0]].
SciPy
Hint
Use a variable named data and assign the list exactly as shown.
2
Set the number of clusters
Create a variable called num_clusters and set it to 2.
SciPy
Hint
Use a variable named num_clusters and assign the value 2.
3
Run K-means with SciPy and scikit-learn
Import kmeans from scipy.cluster.vq and KMeans from sklearn.cluster. Use kmeans with data and num_clusters to get centroids_scipy. Then create a KMeans object with n_clusters=num_clusters, fit it to data, and get centroids_sklearn from its cluster_centers_ attribute.
SciPy
Hint
Use the exact variable names and imports as shown. Remember to unpack the result of kmeans into centroids_scipy and a second value you can ignore.
4
Print the cluster centers
Print the string "SciPy centroids:" followed by centroids_scipy. Then print the string "scikit-learn centroids:" followed by centroids_sklearn.
SciPy
Hint
Use two print statements exactly as described to show the centroids.
Practice
(1/5)
1. What is the main difference between K-means clustering in scipy and scikit-learn?
easy
A. scikit-learn does not support K-means clustering.
B. scikit-learn requires manual centroid initialization, but scipy does not.
C. scipy automatically plots clusters, but scikit-learn does not.
D. scipy requires separate steps for centroid calculation and label assignment, while scikit-learn combines them.
Solution
Step 1: Understand K-means steps in scipy
In scipy, you first find centroids using kmeans, then assign labels with vq.
Step 2: Understand K-means in scikit-learn
scikit-learn combines these steps in one KMeans class that fits and predicts labels together.
Final Answer:
scipy requires separate steps for centroid calculation and label assignment, while scikit-learn combines them. -> Option D
Quick Check:
K-means steps differ: separate in scipy, combined in scikit-learn [OK]
Data has two groups: points near (1, y) and points near (10, y). Kmeans with 2 clusters finds centroids near these groups.
Step 2: Assign labels with vq
Points near (1, y) get label 0, points near (10, y) get label 1. So first three points labeled 0, last three labeled 1.
Final Answer:
[0, 0, 0, 1, 1, 1] -> Option A
Quick Check:
Clusters split by x-coordinate: left=0, right=1 [OK]
Hint: Group points by centroid proximity for labels [OK]
Common Mistakes:
Assuming labels are reversed
Mixing up label order
Expecting labels to be random
4. What is wrong with this code snippet using scipy for K-means clustering?
import numpy as np
from scipy.cluster.vq import kmeans
data = np.array([[1, 2], [3, 4], [5, 6]])
centroids, labels = kmeans(data, 2)
print(labels)
medium
A. kmeans returns centroids and distortion, not labels.
B. Data array shape is invalid for kmeans.
C. kmeans requires 3 clusters, not 2.
D. Missing import for vq function.
Solution
Step 1: Check kmeans return values
kmeans returns centroids and distortion value, not labels.
Step 2: Identify correct label assignment
Labels must be assigned using vq with data and centroids after kmeans.
Final Answer:
kmeans returns centroids and distortion, not labels. -> Option A
Quick Check:
kmeans output ≠ labels; use vq for labels [OK]
Hint: Remember: kmeans returns centroids, not labels [OK]
Common Mistakes:
Expecting kmeans to return labels
Not using vq to assign labels
Confusing distortion with labels
5. You want to cluster a dataset using K-means and compare results between scipy and scikit-learn. Which approach correctly ensures comparable cluster labels?
hard
A. Run scipy's kmeans only, then run scikit-learn's KMeans without setting random_state, compare labels directly.
B. Run scipy's kmeans and vq, then run scikit-learn's KMeans with same n_clusters and random_state, compare labels directly.
C. Run scikit-learn's KMeans only, then assign labels manually using scipy's vq with random centroids.
D. Run scipy's kmeans and assign labels randomly, then run scikit-learn's KMeans with default settings.
Solution
Step 1: Understand label consistency
To compare cluster labels, both methods must use the same number of clusters and fixed random seed for reproducibility.
Step 2: Apply correct procedure
Use scipy's kmeans and vq with fixed initialization, and scikit-learn's KMeans with same n_clusters and random_state. Then compare labels.
Final Answer:
Run scipy's kmeans and vq, then run scikit-learn's KMeans with same n_clusters and random_state, compare labels directly. -> Option B
Quick Check:
Matching clusters need same params and fixed seed [OK]
Hint: Fix random_state and n_clusters to compare labels [OK]