K-means via scipy vs scikit-learn - Performance Comparison
Start learning this pattern below
Jump into concepts and practice - no test required
We want to understand how the time it takes to run K-means clustering grows as we use more data points or clusters.
This helps us know how fast or slow the algorithm will be when using scipy compared to scikit-learn.
Analyze the time complexity of this K-means clustering code using scipy.
from scipy.cluster.vq import kmeans
import numpy as np
# Generate sample data
data = np.random.rand(1000, 2)
# Run K-means clustering
centroids, distortion = kmeans(data, 5, iter=20)
This code runs K-means on 1000 points with 5 clusters and up to 20 iterations.
Look at what repeats in the algorithm:
- Primary operation: Assigning each data point to the nearest cluster center.
- How many times: This happens every iteration, up to 20 times here.
- Also, updating cluster centers after assignments repeats each iteration.
As we increase data points or clusters, the work grows like this:
| Input Size (n points) | Approx. Operations |
|---|---|
| 10 | 10 points x 5 clusters x 20 iterations = 1,000 |
| 100 | 100 x 5 x 20 = 10,000 |
| 1000 | 1000 x 5 x 20 = 100,000 |
Pattern observation: The operations grow roughly in a straight line with the number of points and clusters multiplied by iterations.
Time Complexity: O(n x k x i)
This means the time grows proportionally with the number of points (n), clusters (k), and iterations (i).
[X] Wrong: "The time depends only on the number of data points."
[OK] Correct: The number of clusters and iterations also multiply the work, so ignoring them misses important parts of the cost.
Understanding how K-means scales helps you explain algorithm choices clearly and shows you can think about performance in real projects.
"What if we reduce the number of iterations by half? How would the time complexity change?"
Practice
scipy and scikit-learn?Solution
Step 1: Understand K-means steps in scipy
Inscipy, you first find centroids usingkmeans, then assign labels withvq.Step 2: Understand K-means in scikit-learn
scikit-learncombines these steps in oneKMeansclass that fits and predicts labels together.Final Answer:
scipyrequires separate steps for centroid calculation and label assignment, whilescikit-learncombines them. -> Option DQuick Check:
K-means steps differ: separate in scipy, combined in scikit-learn [OK]
- Thinking scikit-learn lacks K-means
- Assuming scipy auto-assigns labels
- Confusing plotting features with clustering steps
Solution
Step 1: Recall scipy K-means import syntax
The correct import for K-means in scipy is fromscipy.cluster.vqimportingkmeansandvq.Step 2: Check other options
Options A and B use incorrect module names, and D is from scikit-learn, not scipy.Final Answer:
from scipy.cluster.vq import kmeans, vq -> Option CQuick Check:
Correct scipy import = from scipy.cluster.vq import kmeans, vq [OK]
- Confusing sklearn imports with scipy
- Using wrong module names like scipy.kmeans
- Trying to import cluster from scipy directly
labels?
import numpy as np from scipy.cluster.vq import kmeans, vq data = np.array([[1, 2], [1, 4], [1, 0], [10, 2], [10, 4], [10, 0]]) centroids, _ = kmeans(data, np.array([[1, 2], [10, 2]])) labels, _ = vq(data, centroids) print(labels.tolist())
Solution
Step 1: Understand data and centroids
Data has two groups: points near (1, y) and points near (10, y). Kmeans with 2 clusters finds centroids near these groups.Step 2: Assign labels with vq
Points near (1, y) get label 0, points near (10, y) get label 1. So first three points labeled 0, last three labeled 1.Final Answer:
[0, 0, 0, 1, 1, 1] -> Option AQuick Check:
Clusters split by x-coordinate: left=0, right=1 [OK]
- Assuming labels are reversed
- Mixing up label order
- Expecting labels to be random
import numpy as np from scipy.cluster.vq import kmeans data = np.array([[1, 2], [3, 4], [5, 6]]) centroids, labels = kmeans(data, 2) print(labels)
Solution
Step 1: Check kmeans return values
kmeansreturns centroids and distortion value, not labels.Step 2: Identify correct label assignment
Labels must be assigned usingvqwith data and centroids after kmeans.Final Answer:
kmeans returns centroids and distortion, not labels. -> Option AQuick Check:
kmeans output ≠ labels; use vq for labels [OK]
- Expecting kmeans to return labels
- Not using vq to assign labels
- Confusing distortion with labels
Solution
Step 1: Understand label consistency
To compare cluster labels, both methods must use the same number of clusters and fixed random seed for reproducibility.Step 2: Apply correct procedure
Use scipy'skmeansandvqwith fixed initialization, and scikit-learn'sKMeanswith samen_clustersandrandom_state. Then compare labels.Final Answer:
Run scipy's kmeans and vq, then run scikit-learn's KMeans with same n_clusters and random_state, compare labels directly. -> Option BQuick Check:
Matching clusters need same params and fixed seed [OK]
- Not fixing random_state causing label mismatch
- Assigning labels randomly in scipy
- Comparing labels without same cluster count
