Bird
Raised Fist0
SciPydata~10 mins

Flat clustering (fcluster) in SciPy - Interactive Code Practice

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Practice - 5 Tasks
Answer the questions below
1fill in blank
easy

Complete the code to import the function for flat clustering from scipy.

SciPy
from scipy.cluster.hierarchy import [1]
Drag options to blanks, or click blank then click option'
Alinkage
Bward
Cfcluster
Ddendrogram
Attempts:
3 left
💡 Hint
Common Mistakes
Importing 'linkage' instead of 'fcluster'.
Confusing 'dendrogram' with 'fcluster'.
2fill in blank
medium

Complete the code to create flat clusters using a distance threshold of 1.5.

SciPy
clusters = fcluster(linkage_matrix, [1], criterion='distance')
Drag options to blanks, or click blank then click option'
A3.0
B2.0
C0.5
D1.5
Attempts:
3 left
💡 Hint
Common Mistakes
Using a threshold other than 1.5.
Confusing the threshold with the number of clusters.
3fill in blank
hard

Fix the error in the code to correctly assign clusters by maximum number of clusters = 3.

SciPy
clusters = fcluster(linkage_matrix, [1], criterion='maxclust')
Drag options to blanks, or click blank then click option'
A3
B2
C4
D5
Attempts:
3 left
💡 Hint
Common Mistakes
Using a threshold that does not match the desired number of clusters.
Confusing 'distance' criterion with 'maxclust'.
4fill in blank
hard

Fill both blanks to create a dictionary of cluster labels and their counts.

SciPy
cluster_counts = {label: [1] for label in set(clusters) if clusters.count(label) [2] 1}
Drag options to blanks, or click blank then click option'
Aclusters.count(label)
B>
C==
D<
Attempts:
3 left
💡 Hint
Common Mistakes
Using '==' instead of '>' in the condition.
Using a wrong method to count occurrences.
5fill in blank
hard

Fill all three blanks to create a list of cluster labels for points with distance less than 2.0.

SciPy
selected_clusters = [[1] for i, d in enumerate(distances) if d [2] [3]]
Drag options to blanks, or click blank then click option'
Aclusters[i]
B<
C2.0
D>
Attempts:
3 left
💡 Hint
Common Mistakes
Using '>' instead of '<' in the condition.
Using 'd' instead of 'clusters[i]' in the output.

Practice

(1/5)
1. What is the main purpose of the fcluster function in scipy's hierarchical clustering?
easy
A. To cut a hierarchical cluster tree into flat clusters based on a threshold
B. To compute the distance matrix between data points
C. To perform dimensionality reduction before clustering
D. To normalize data before clustering

Solution

  1. Step 1: Understand hierarchical clustering output

    Hierarchical clustering produces a tree (dendrogram) showing nested clusters.
  2. Step 2: Role of fcluster

    fcluster cuts this tree at a chosen threshold to form flat, non-overlapping clusters.
  3. Final Answer:

    To cut a hierarchical cluster tree into flat clusters based on a threshold -> Option A
  4. Quick Check:

    Flat clustering = cutting tree with threshold [OK]
Hint: Remember: fcluster cuts dendrogram into flat groups [OK]
Common Mistakes:
  • Confusing fcluster with distance calculation
  • Thinking fcluster normalizes data
  • Assuming fcluster reduces dimensions
2. Which of the following is the correct syntax to assign flat clusters using fcluster with a distance threshold of 1.5 from a linkage matrix Z?
easy
A. clusters = fcluster(Z, threshold=1.5, criterion='distance')
B. clusters = fcluster(Z, 1.5, method='distance')
C. clusters = fcluster(Z, 1.5, criterion='distance')
D. clusters = fcluster(Z, 1.5, criterion='maxclust')

Solution

  1. Step 1: Check fcluster parameters

    The function signature is fcluster(Z, t, criterion='distance') where t is the threshold.
  2. Step 2: Identify correct usage

    clusters = fcluster(Z, 1.5, criterion='distance') uses t=1.5 and criterion='distance', which is correct syntax.
  3. Final Answer:

    clusters = fcluster(Z, 1.5, criterion='distance') -> Option C
  4. Quick Check:

    Threshold = 1.5, criterion = 'distance' [OK]
Hint: Use t for threshold and criterion='distance' in fcluster [OK]
Common Mistakes:
  • Using 'method' instead of 'criterion'
  • Passing threshold as keyword 'threshold'
  • Using wrong criterion like 'maxclust' for distance cut
3. Given the linkage matrix Z = [[0, 1, 0.5, 2], [2, 3, 1.5, 2], [4, 5, 2.5, 4]], what is the output of fcluster(Z, 1.0, criterion='distance')?
medium
A. [1 1 1 1]
B. [1 2 3 4]
C. [1 2 2 3]
D. [1 1 2 3]

Solution

  1. Step 1: Understand linkage matrix and threshold

    The linkage matrix Z shows merges with distances: 0.5, 1.5, 2.5. Threshold is 1.0.
  2. Step 2: Assign clusters by cutting at distance 1.0

    Clusters merge if distance ≤ 1.0. The first merge (0.5) joins points 0 and 1 into cluster 1. The second merge (1.5 > 1.0) does not merge points 2 and 3, so point 2 gets cluster 2 and point 3 gets cluster 3.
  3. Final Answer:

    [1 1 2 3] -> Option D
  4. Quick Check:

    Distance ≤ 1.0 merges points 0 and 1 only [OK]
Hint: Cut dendrogram at threshold; merges below threshold cluster together [OK]
Common Mistakes:
  • Merging clusters above threshold
  • Assigning all points to one cluster
  • Misreading linkage matrix format
4. You run the code clusters = fcluster(Z, 2, criterion='maxclust') but get an error. What is the likely cause?
medium
A. The linkage matrix Z is not defined or invalid
B. The criterion 'maxclust' requires an integer number of clusters, but 2 is passed as a float
C. The threshold parameter must be a float when using 'maxclust'
D. The criterion 'maxclust' expects the threshold to be the maximum cluster distance

Solution

  1. Step 1: Check parameter types for 'maxclust'

    When using criterion='maxclust', the threshold t must be an integer specifying the number of clusters.
  2. Step 2: Identify common error

    If Z is not defined or invalid, fcluster raises an error unrelated to parameter types.
  3. Step 3: Analyze options

    The criterion 'maxclust' requires an integer number of clusters, but 2 is passed as a float is incorrect because 2 as an integer or float is accepted; The threshold parameter must be a float when using 'maxclust' is wrong because threshold can be int; The criterion 'maxclust' expects the threshold to be the maximum cluster distance is false because 'maxclust' uses number of clusters, not distance.
  4. Final Answer:

    The linkage matrix Z is not defined or invalid -> Option A
  5. Quick Check:

    Undefined Z causes error, not threshold type [OK]
Hint: Ensure linkage matrix Z is valid before calling fcluster [OK]
Common Mistakes:
  • Passing float instead of int for maxclust threshold
  • Misunderstanding criterion parameter
  • Ignoring linkage matrix validity
5. You have a dataset with 10 points clustered hierarchically. You want exactly 3 clusters. Which fcluster call correctly achieves this?
hard
A. fcluster(Z, 0.5, criterion='inconsistent')
B. fcluster(Z, 3, criterion='maxclust')
C. fcluster(Z, 0.5, criterion='maxclust')
D. fcluster(Z, 3, criterion='distance')

Solution

  1. Step 1: Understand criteria for exact cluster count

    To get exactly 3 clusters, use criterion='maxclust' with t=3 specifying number of clusters.
  2. Step 2: Analyze options

    fcluster(Z, 3, criterion='maxclust') correctly uses maxclust with 3 clusters. fcluster(Z, 3, criterion='distance') uses distance criterion which does not guarantee exact cluster count. Options A and C use incorrect thresholds or criteria.
  3. Final Answer:

    fcluster(Z, 3, criterion='maxclust') -> Option B
  4. Quick Check:

    maxclust + t=number of clusters = exact clusters [OK]
Hint: Use criterion='maxclust' with t = desired cluster count [OK]
Common Mistakes:
  • Using distance criterion to get exact cluster count
  • Passing float threshold for maxclust
  • Confusing inconsistent criterion with maxclust