Bird
Raised Fist0
SciPydata~20 mins

Cluster evaluation metrics in SciPy - Practice Problems & Coding Challenges

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Challenge - 5 Problems
🎖️
Cluster Metrics Master
Get all challenges correct to earn this badge!
Test your skills under time pressure!
Predict Output
intermediate
2:00remaining
Output of silhouette_score for simple clusters
What is the output of the following code that calculates the silhouette score for two simple clusters?
SciPy
from sklearn.metrics import silhouette_score
from sklearn.datasets import make_blobs

X, labels = make_blobs(n_samples=10, centers=2, cluster_std=0.5, random_state=42)
score = silhouette_score(X, labels)
print(round(score, 2))
A0.71
B0.15
C1.00
D-0.50
Attempts:
2 left
💡 Hint
Silhouette score ranges from -1 to 1, higher means better cluster separation.
data_output
intermediate
1:30remaining
Number of clusters from DBSCAN labels
Given the following DBSCAN clustering labels, how many clusters (excluding noise) are detected?
SciPy
import numpy as np
labels = np.array([0, 0, 1, 1, -1, 2, 2, -1, 2])
num_clusters = len(set(labels)) - (1 if -1 in labels else 0)
print(num_clusters)
A3
B2
C4
D5
Attempts:
2 left
💡 Hint
Noise points are labeled as -1 and should not be counted as clusters.
🧠 Conceptual
advanced
1:30remaining
Understanding Adjusted Rand Index (ARI)
Which statement correctly describes the Adjusted Rand Index (ARI) in cluster evaluation?
AARI calculates the ratio of within-cluster variance to total variance.
BARI measures the average distance between cluster centroids and data points.
CARI is a metric that only works for hierarchical clustering methods.
DARI measures similarity between two clusterings, adjusted for chance, with values from -1 to 1 where 1 means perfect match.
Attempts:
2 left
💡 Hint
Think about what ARI compares and its value range.
visualization
advanced
1:30remaining
Interpreting a silhouette plot
You run a silhouette plot for a clustering result with 3 clusters. Which of the following interpretations is correct if one cluster has many negative silhouette values?
ANegative silhouette values mean the clustering algorithm perfectly assigned points.
BThe cluster with negative silhouette values is the most compact and well-separated cluster.
CThe cluster with many negative silhouette values is poorly separated and may be overlapping with other clusters.
DNegative silhouette values indicate that the cluster has the highest density of points.
Attempts:
2 left
💡 Hint
Recall what negative silhouette values mean about point assignment.
🔧 Debug
expert
2:00remaining
Identify the error in Calinski-Harabasz score calculation
What error will the following code raise when calculating the Calinski-Harabasz score?
SciPy
from sklearn.metrics import calinski_harabasz_score
X = [[1, 2], [3, 4], [5, 6]]
labels = [0, 1]
score = calinski_harabasz_score(X, labels)
print(score)
AIndexError: list index out of range
BValueError: Number of labels does not match number of samples
CTypeError: calinski_harabasz_score() missing required positional argument
DNo error, prints a float score
Attempts:
2 left
💡 Hint
Check if the labels list length matches the number of samples in X.

Practice

(1/5)
1. Which cluster evaluation metric is best used when you do NOT have true labels for your data?
easy
A. Adjusted Rand Index
B. Silhouette Score
C. Accuracy Score
D. Mean Squared Error

Solution

  1. Step 1: Understand the role of true labels

    Adjusted Rand Index requires true labels to compare clusters, so it is not suitable without labels.
  2. Step 2: Identify metrics for unknown labels

    Silhouette Score measures how well clusters are separated without needing true labels.
  3. Final Answer:

    Silhouette Score -> Option B
  4. Quick Check:

    Unknown labels = Silhouette Score [OK]
Hint: Use silhouette score when labels are unknown [OK]
Common Mistakes:
  • Confusing Adjusted Rand Index as label-free
  • Choosing accuracy score which needs labels
  • Using mean squared error for clustering
2. Which of the following is the correct way to import the silhouette_score function for cluster evaluation?
easy
A. from scipy.cluster import silhouette_score
B. from scipy.spatial.distance import silhouette_score
C. from scipy.cluster.hierarchy import silhouette_score
D. from sklearn.metrics import silhouette_score

Solution

  1. Step 1: Check common library modules

    Silhouette score is available in sklearn.metrics module, not in scipy.cluster or spatial.distance.
  2. Step 2: Verify import syntax

    The correct import is from sklearn.metrics import silhouette_score.
  3. Final Answer:

    from sklearn.metrics import silhouette_score -> Option D
  4. Quick Check:

    Correct import = sklearn.metrics [OK]
Hint: Silhouette score is in sklearn.metrics module [OK]
Common Mistakes:
  • Importing from scipy.cluster directly
  • Using scipy.spatial.distance for silhouette_score
  • Confusing hierarchy module with vq
3. What is the output of the following code snippet?
from scipy.cluster.vq import kmeans, vq, whiten
import numpy as np

data = np.array([[1, 2], [1, 4], [1, 0], [10, 2], [10, 4], [10, 0]])
whitened = whiten(data)
centroids, _ = kmeans(whitened, 2)
cluster_labels, _ = vq(whitened, centroids)
from sklearn.metrics import silhouette_score
score = silhouette_score(whitened, cluster_labels)
print(round(score, 2))
medium
A. 0.75
B. 1.00
C. 0.35
D. 0.55

Solution

  1. Step 1: Understand the code flow

    The code whitens data, runs kmeans for 2 clusters, assigns labels, then calculates silhouette score.
  2. Step 2: Interpret silhouette score meaning

    Data clearly forms two groups; silhouette score is around 0.75 indicating good cluster separation.
  3. Final Answer:

    0.75 -> Option A
  4. Quick Check:

    Well-separated clusters ≈ 0.75 silhouette [OK]
Hint: Silhouette near 0.75 means good cluster separation [OK]
Common Mistakes:
  • Expecting silhouette score of 1.0 always
  • Confusing whitened data with original scale
  • Misreading cluster labels
4. Identify the error in this code snippet for calculating Davies-Bouldin score:
from scipy.spatial.distance import davies_bouldin_score

labels = [0, 0, 1, 1]
data = [[1, 2], [1, 4], [10, 2], [10, 4]]
score = davies_bouldin_score(data, labels)
print(score)
medium
A. Data must be a numpy array, not list
B. Labels and data length mismatch
C. Importing davies_bouldin_score from wrong module
D. Davies-Bouldin score requires true labels

Solution

  1. Step 1: Check import source

    Davies-Bouldin score is in sklearn.metrics, not scipy.spatial.distance.
  2. Step 2: Validate data and labels

    Data and labels lengths match and data as list works with sklearn, so no error there.
  3. Final Answer:

    Importing davies_bouldin_score from wrong module -> Option C
  4. Quick Check:

    Correct import is sklearn.metrics [OK]
Hint: Davies-Bouldin score is in sklearn.metrics, not scipy [OK]
Common Mistakes:
  • Importing from scipy.spatial.distance
  • Assuming data must be numpy array
  • Thinking Davies-Bouldin needs true labels
5. You have true labels and predicted cluster labels for a dataset. Which metric from scipy or sklearn should you use to evaluate clustering quality by comparing these labels?
hard
A. Adjusted Rand Index
B. Davies-Bouldin Score
C. Silhouette Score
D. Calinski-Harabasz Index

Solution

  1. Step 1: Identify metrics needing true labels

    Adjusted Rand Index compares predicted clusters with true labels to measure similarity.
  2. Step 2: Exclude label-free metrics

    Silhouette, Davies-Bouldin, and Calinski-Harabasz do not use true labels for evaluation.
  3. Final Answer:

    Adjusted Rand Index -> Option A
  4. Quick Check:

    True vs predicted labels = Adjusted Rand Index [OK]
Hint: Use Adjusted Rand Index to compare true and predicted labels [OK]
Common Mistakes:
  • Using silhouette score with true labels
  • Confusing Davies-Bouldin as label-based
  • Choosing Calinski-Harabasz for label comparison