Jump into concepts and practice - no test required
or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Flat Clustering with fcluster in SciPy
📖 Scenario: You work in a small shop that sells fruits. You want to group similar fruits based on their sweetness and crunchiness scores to understand customer preferences better.
🎯 Goal: Build a program that uses hierarchical clustering and then applies fcluster to create flat clusters of fruits based on a distance threshold.
📋 What You'll Learn
Create a dictionary called fruits with fruit names as keys and tuples of (sweetness, crunchiness) as values.
Create a variable called threshold and set it to 1.5.
Use scipy.cluster.hierarchy.linkage to compute the linkage matrix from the fruit data.
Use scipy.cluster.hierarchy.fcluster with the linkage matrix and threshold to assign cluster labels.
Print the dictionary clusters that maps fruit names to their cluster labels.
💡 Why This Matters
🌍 Real World
Grouping similar items helps businesses understand customer preferences and organize products better.
💼 Career
Clustering is a common technique in data science for market segmentation, recommendation systems, and pattern discovery.
Progress0 / 4 steps
1
Create the fruit data dictionary
Create a dictionary called fruits with these exact entries: 'Apple': (7, 8), 'Banana': (10, 2), 'Cherry': (8, 7), 'Date': (9, 3), 'Elderberry': (6, 9).
SciPy
Hint
Use curly braces {} to create a dictionary. Each fruit name is a key, and the value is a tuple with two numbers.
2
Set the clustering threshold
Create a variable called threshold and set it to 1.5.
SciPy
Hint
Just write threshold = 1.5 on a new line.
3
Compute linkage and assign clusters
Import linkage and fcluster from scipy.cluster.hierarchy. Create a list called data_points containing the fruit values from fruits. Use linkage with data_points and method 'ward' to create Z. Then use fcluster with Z, threshold, and criterion 'distance' to create labels. Finally, create a dictionary called clusters that maps each fruit name to its cluster label.
SciPy
Hint
Use list(fruits.values()) to get the data points. Use linkage with method 'ward'. Use fcluster with criterion 'distance'. Use a dictionary comprehension to map fruit names to labels.
4
Print the clusters dictionary
Write a print statement to display the clusters dictionary.
SciPy
Hint
Use print(clusters) to show the result.
Practice
(1/5)
1. What is the main purpose of the fcluster function in scipy's hierarchical clustering?
easy
A. To cut a hierarchical cluster tree into flat clusters based on a threshold
B. To compute the distance matrix between data points
C. To perform dimensionality reduction before clustering
D. To normalize data before clustering
Solution
Step 1: Understand hierarchical clustering output
Hierarchical clustering produces a tree (dendrogram) showing nested clusters.
Step 2: Role of fcluster
fcluster cuts this tree at a chosen threshold to form flat, non-overlapping clusters.
Final Answer:
To cut a hierarchical cluster tree into flat clusters based on a threshold -> Option A
Quick Check:
Flat clustering = cutting tree with threshold [OK]
Hint: Remember: fcluster cuts dendrogram into flat groups [OK]
Common Mistakes:
Confusing fcluster with distance calculation
Thinking fcluster normalizes data
Assuming fcluster reduces dimensions
2. Which of the following is the correct syntax to assign flat clusters using fcluster with a distance threshold of 1.5 from a linkage matrix Z?
easy
A. clusters = fcluster(Z, threshold=1.5, criterion='distance')
B. clusters = fcluster(Z, 1.5, method='distance')
C. clusters = fcluster(Z, 1.5, criterion='distance')
D. clusters = fcluster(Z, 1.5, criterion='maxclust')
Solution
Step 1: Check fcluster parameters
The function signature is fcluster(Z, t, criterion='distance') where t is the threshold.
Step 2: Identify correct usage
clusters = fcluster(Z, 1.5, criterion='distance') uses t=1.5 and criterion='distance', which is correct syntax.
Final Answer:
clusters = fcluster(Z, 1.5, criterion='distance') -> Option C
Quick Check:
Threshold = 1.5, criterion = 'distance' [OK]
Hint: Use t for threshold and criterion='distance' in fcluster [OK]
Common Mistakes:
Using 'method' instead of 'criterion'
Passing threshold as keyword 'threshold'
Using wrong criterion like 'maxclust' for distance cut
3. Given the linkage matrix Z = [[0, 1, 0.5, 2], [2, 3, 1.5, 2], [4, 5, 2.5, 4]], what is the output of fcluster(Z, 1.0, criterion='distance')?
medium
A. [1 1 1 1]
B. [1 2 3 4]
C. [1 2 2 3]
D. [1 1 2 3]
Solution
Step 1: Understand linkage matrix and threshold
The linkage matrix Z shows merges with distances: 0.5, 1.5, 2.5. Threshold is 1.0.
Step 2: Assign clusters by cutting at distance 1.0
Clusters merge if distance ≤ 1.0. The first merge (0.5) joins points 0 and 1 into cluster 1. The second merge (1.5 > 1.0) does not merge points 2 and 3, so point 2 gets cluster 2 and point 3 gets cluster 3.
Final Answer:
[1 1 2 3] -> Option D
Quick Check:
Distance ≤ 1.0 merges points 0 and 1 only [OK]
Hint: Cut dendrogram at threshold; merges below threshold cluster together [OK]
Common Mistakes:
Merging clusters above threshold
Assigning all points to one cluster
Misreading linkage matrix format
4. You run the code clusters = fcluster(Z, 2, criterion='maxclust') but get an error. What is the likely cause?
medium
A. The linkage matrix Z is not defined or invalid
B. The criterion 'maxclust' requires an integer number of clusters, but 2 is passed as a float
C. The threshold parameter must be a float when using 'maxclust'
D. The criterion 'maxclust' expects the threshold to be the maximum cluster distance
Solution
Step 1: Check parameter types for 'maxclust'
When using criterion='maxclust', the threshold t must be an integer specifying the number of clusters.
Step 2: Identify common error
If Z is not defined or invalid, fcluster raises an error unrelated to parameter types.
Step 3: Analyze options
The criterion 'maxclust' requires an integer number of clusters, but 2 is passed as a float is incorrect because 2 as an integer or float is accepted; The threshold parameter must be a float when using 'maxclust' is wrong because threshold can be int; The criterion 'maxclust' expects the threshold to be the maximum cluster distance is false because 'maxclust' uses number of clusters, not distance.
Final Answer:
The linkage matrix Z is not defined or invalid -> Option A
Quick Check:
Undefined Z causes error, not threshold type [OK]
Hint: Ensure linkage matrix Z is valid before calling fcluster [OK]
Common Mistakes:
Passing float instead of int for maxclust threshold
Misunderstanding criterion parameter
Ignoring linkage matrix validity
5. You have a dataset with 10 points clustered hierarchically. You want exactly 3 clusters. Which fcluster call correctly achieves this?
hard
A. fcluster(Z, 0.5, criterion='inconsistent')
B. fcluster(Z, 3, criterion='maxclust')
C. fcluster(Z, 0.5, criterion='maxclust')
D. fcluster(Z, 3, criterion='distance')
Solution
Step 1: Understand criteria for exact cluster count
To get exactly 3 clusters, use criterion='maxclust' with t=3 specifying number of clusters.
Step 2: Analyze options
fcluster(Z, 3, criterion='maxclust') correctly uses maxclust with 3 clusters. fcluster(Z, 3, criterion='distance') uses distance criterion which does not guarantee exact cluster count. Options A and C use incorrect thresholds or criteria.
Final Answer:
fcluster(Z, 3, criterion='maxclust') -> Option B
Quick Check:
maxclust + t=number of clusters = exact clusters [OK]
Hint: Use criterion='maxclust' with t = desired cluster count [OK]
Common Mistakes:
Using distance criterion to get exact cluster count