# KMeansClustering
course: Module 2 — Machine Learning Algorithms
module: Module-2-Machine-Learning-Algorithms
type: notebook
source_url: https://personal-learn.armco.dev/files/Module-2-Machine-Learning-Algorithms/General/Lab_Materials/KMeansClustering.ipynb
---
[cell 1 markdown]
# **Dataset Overview**
The Dry Bean Dataset contains measurements of physical characteristics of seven types of dry beans, captured from digital images. It describes each bean using shape and morphological features such as area, perimeter, and compactness.
The goal in this analysis is to use KMeans clustering to group beans based on their physical features. While the actual bean type labels are available, they are only used to evaluate how well the clustering captures meaningful class distinctions, not for training or prediction.
## **Source:**
Available via the UCI Machine Learning Repository:[UCI Dataset Page](https://archive.ics.uci.edu/ml/datasets/Dry+Bean+Dataset)
## **Features:**
| Column Name | Description | Data Type |
| --------------- | --------------------------------------------------------- | --------------- |
| Area | Total number of pixels in the region | Numeric (float) |
| Perimeter | Circumference of the bean region | Numeric (float) |
| MajorAxisLength | Length of the major axis of the fitted ellipse | Numeric (float) |
| MinorAxisLength | Length of the minor axis of the fitted ellipse | Numeric (float) |
| AspectRation | Ratio of major axis to minor axis | Numeric (float) |
| Eccentricity | Eccentricity of the ellipse (how elongated the shape is) | Numeric (float) |
| ConvexArea | Number of pixels in the convex hull region | Numeric (float) |
| EquivDiameter | Diameter of a circle with the same area as the region | Numeric (float) |
| Extent | Ratio of region area to bounding box area | Numeric (float) |
| Solidity | Ratio of region area to convex hull area | Numeric (float) |
| Roundness | Shape roundness based on area and perimeter | Numeric (float) |
| Compactness | How closely the bean shape resembles a circle | Numeric (float) |
| ShapeFactor1 | Shape descriptor based on perimeter and area | Numeric (float) |
| ShapeFactor2 | Shape descriptor based on perimeter and major axis length | Numeric (float) |
| ShapeFactor3 | Shape descriptor based on area and bounding box | Numeric (float) |
| ShapeFactor4 | Another derived shape descriptor | Numeric (float) |
| Class | Bean type label (e.g., Seker, Barbunya, Bombay, etc.) | Categorical |
[cell 2 code]
import numpy as np
import pandas as pd
import seaborn as sns
import matplotlib.pyplot as plt
[cell 3 code]
import matplotlib.pyplot as plt
import seaborn as sns
import warnings
warnings.filterwarnings('ignore')
[cell 4 markdown]
# **Loading the data**
[cell 5 code]
df_beans=pd.read_excel('Dry_Bean_Dataset.xlsx')
print(df_beans.shape)
df_beans
[cell 6 code]
df_beans.info()
[cell 7 markdown]
# **Checking explicitly for Missing Values**
[cell 8 code]
df_beans.isnull().sum()
[cell 9 markdown]
# **Descriptive Statistics of Numerical Attributes**
[cell 10 code]
df_beans.describe()
[cell 11 markdown]
# **Duplicate values in Dataset**
[cell 12 code]
print(f'Duplicate rows in the datsaset :{df_beans.duplicated().sum()}\n')
print('Bean type of duplicated rows:')
print(df_beans[df_beans.duplicated()].Class.value_counts())
print('\nDisplaying the duplicated rows')
df_beans[df_beans.duplicated(keep=False)]
[cell 13 markdown]
# **Dropping Duplicate values from dataset**
[cell 14 code]
df_beans=df_beans.drop_duplicates()
df_beans.shape
[cell 15 code]
df_beans['Class'].unique()
[cell 16 code]
df_beans['Class'].value_counts()
[cell 17 markdown]
# **Histograms of Numerical Features Colored by Bean Class**
[cell 18 code]
import matplotlib.pyplot as plt
import seaborn as sns
unique_classes = sorted(df_beans['Class'].unique())
palette = sns.color_palette('deep', n_colors=len(unique_classes))
class_colors = dict(zip(unique_classes, palette))
numerical_cols = df_beans.drop(columns=['Class']).columns
fig, ax = plt.subplots(4, 4, figsize=(15, 12))
for variable, subplot in zip(numerical_cols, ax.flatten()):
g = sns.histplot(data=df_beans, x=variable, ax=subplot, hue='Class', palette=class_colors,alpha=0.6)
g.axvline(x=df_beans[variable].mean(), color='m', label='Mean', linestyle='--', linewidth=2)
plt.tight_layout()
[cell 19 markdown]
### **The Univariate distribution per attribute for each class separately which shows that seed type BOMBAY has distinct distribution from others on several attributes**
[cell 20 markdown]
# **Distribution of Numerical Features Across Bean Classes**
[cell 21 code]
fig, ax = plt.subplots(4, 4, figsize=(15, 12))
for variable, subplot in zip(numerical_cols, ax.flatten()):
g=sns.stripplot(x=df_beans[variable],y=df_beans.Class ,ax=subplot,palette=class_colors,alpha=0.6)
plt.tight_layout()
[cell 22 markdown]
* The distribution of BOMBAY class in many attribtes like *Area*, *MinorAxisLength*, *Convex Area*, *Minor Axis Length*, *ShapeFactor1* is well separated from other classes.
* *Area* & Convex Area have very similar distribution across all classes indicating very high correlation.
* Distribution of *DERMASON* is very similar *SEKER* in many attributes for e.g. roundness, solidity, ShapeFactor2
* For the *extent* attribute the range of values is similar for all the bean classes
[cell 23 markdown]
# **Correlation Matrix**
[cell 24 code]
plt.figure(figsize=(14,12))
sns.heatmap(df_beans[numerical_cols].corr(), yticklabels='auto', annot=True, cmap='coolwarm')
plt.show()
[cell 25 markdown]
* There a lot of highly correlated attributes in the above correlation matrix, for eg: </br>
* **Area & Convex Area**:1
* **Shaped Factor3 & Comapctness**:1
* **Aspect ration & compactness**: -0.99
* **Area & Perimeter**: 0.97
* **Perimeter & ShapeFactor1**: -0.87
* **Aspect ration & Eccentricity**: 0.92 </br>
* Some attributes with low level of correlation among them:
* **Extent & EquivDiameter**: 0.029
* **Solidity & Eccentricity**: -0.3
* **Compactnes & Area**: -0.27
[cell 26 markdown]
# Pairplots Between All Numerical Features
A **pairplot** shows:
- **Diagonal**: the *univariate distribution* of each numerical feature (KDE curves).
- **Off-diagonal**: *pairwise scatter plots* between every pair of numerical features.
Why this is useful:
- To visually spot **correlations** (tight linear/curved trends).
- To see **separation between bean classes** (if `Class` is used as `hue`).
- To detect **outliers**, **skewness**, and **overlapping clusters** that may affect KMeans.
[cell 27 code]
# Ensure Class is categorical
df_beans["Class"] = df_beans["Class"].astype("category")
size_features = [
"Area",
"Perimeter",
"MajorAxisLength",
"MinorAxisLength",
"AspectRation",
"Eccentricity",
"ConvexArea",
"EquivDiameter"
]
sns.set_theme(style="whitegrid", context="notebook", font_scale=1.1)
g1 = sns.pairplot(
data=df_beans, # FULL DATA
vars=size_features,
hue="Class",
corner=True,
diag_kind="kde",
height=2.8,
plot_kws=dict(
s=8,
alpha=0.25,
linewidth=0
)
)
g1.fig.set_size_inches(34, 22)
g1.fig.suptitle(
"Pairplot of Size-Related Features (Area to EquivDiameter)",
fontsize=20,
y=1.02
)
# # Enlarge legend
# legend = g1._legend
# legend.set_title("Class", prop={"size": 20})
# for text in legend.get_texts():
# text.set_fontsize(20)
# legend.set_bbox_to_anchor((1.02, 0.5)) # move legend to the right
plt.show()
[cell 28 code]
shape_features = [
"Extent",
"Solidity",
"roundness",
"Compactness",
"ShapeFactor1",
"ShapeFactor2",
"ShapeFactor3",
"ShapeFactor4"
]
g2 = sns.pairplot(
data=df_beans, # FULL DATA
vars=shape_features,
hue="Class",
corner=True,
diag_kind="kde",
height=2.8,
plot_kws=dict(
s=8,
alpha=0.25,
linewidth=0
)
)
g2.fig.set_size_inches(34, 22)
g2.fig.suptitle(
"Pairplot of Shape-Related Features (Extent to ShapeFactor4)",
fontsize=20,
y=1.02
)
plt.show()
[cell 29 markdown]
# **Class distribution of various types of beans**
[cell 30 code]
print(df_beans['Class'].value_counts())
plt.figure(figsize=(8,6))
sns.countplot(x='Class', data=df_beans, palette=class_colors)
plt.xlabel('Bean Type')
plt.ylabel('Count')
plt.title('Class distribution of various types of beans')
plt.tight_layout()
[cell 31 markdown]
# **Preparing the data**
[cell 32 code]
import numpy as np
from sklearn.preprocessing import StandardScaler
# Map bean classes to integers
class_map = {cls: idx for idx, cls in enumerate(df_beans['Class'].unique())}
df_beans['classno'] = df_beans['Class'].map(class_map)
# Extract features and true labels
X = df_beans.drop(columns=['Class', 'classno']).values
# Extract features and true labels
X = df_beans.drop(columns=['Class', 'classno']).values
y_true = df_beans['classno'].values
# Normalize features
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)
[cell 33 markdown]
# **Determining the optimal number of clusters for KMeans using elbow method**
[cell 34 code]
import matplotlib.pyplot as plt
from sklearn.cluster import KMeans
# Range of k values
k_range = range(2, 35)
# Inertia = sum of squared distances between each point and the centroid of its assigned cluster
# Store inertia (within-cluster sum of squares)
inertias = []
for k in k_range:
kmeans = KMeans(n_clusters=k, n_init=10, random_state=42)
kmeans.fit(X_scaled)
inertias.append(kmeans.inertia_)
# Plotting the Elbow Curve
plt.figure(figsize=(8, 6))
plt.plot(k_range, inertias, marker='o')
plt.xlabel('Number of clusters (k)')
plt.ylabel('Inertia (Within-cluster sum of squares)')
plt.title('Elbow Method for Optimal k')
plt.xticks(k_range)
plt.grid(True)
plt.tight_layout()
plt.show()
[cell 35 markdown]
## **We chose k = 7 because, beyond this point, the decrease in within-cluster sum of squared errors becomes slower and more gradual.**
[cell 36 code]
# pandas Series
inertia_series = pd.Series(data=inertias, index=k_range, name='Inertia')
inertia_series.index.name = 'k (n_clusters)'
# Print
print("\nInertia values for each k:")
inertia_series[:21]
[cell 37 markdown]
### **The inertia values decrease gradually after k = 7.**
[cell 38 markdown]
# **Evaluation of KMeans Clustering (k = 7)**
[cell 39 code]
from sklearn.metrics import (
silhouette_score, calinski_harabasz_score, davies_bouldin_score,
rand_score, adjusted_rand_score, adjusted_mutual_info_score,homogeneity_score,
completeness_score
)
# KMeans clustering
n_clusters = 7
kmeans = KMeans(n_clusters=n_clusters, n_init=10,random_state=2024)
y_pred = kmeans.fit_predict(X_scaled)
# Intrinsic Measures
SIL_s = silhouette_score(X_scaled, y_pred)
CH_s = calinski_harabasz_score(X_scaled, y_pred)
DB_s = davies_bouldin_score(X_scaled, y_pred)
# Extrinsic Measures
RI_s = rand_score(y_true, y_pred)
ARI_ward = adjusted_rand_score(y_true, y_pred)
AMI_ward = adjusted_mutual_info_score(y_true, y_pred)
CM_s = completeness_score(y_true, y_pred)
HM_s = homogeneity_score(y_true, y_pred)
# Print Results
print('Intrinsic Measure:\n')
print(f'Silhouette score: {SIL_s:.3f}')
print(f'Calinski and Harabasz score: {CH_s:.3f}')
print(f'Davies-Bouldin score: {DB_s:.3f}')
#############################################
print('\nExtrinsic Measure:\n')
print(f'Random index: {RI_s:.3f}')
print(f'Homogeneity score: {HM_s:.3f}')
print(f'Completeness score: {CM_s:.3f}')
print(f'Adjusted Random index: {ARI_ward:.3f}')
print(f'Adjusted Mutual Information: {AMI_ward:.3f}')
[cell 40 markdown]
## Standard behaviour of the cluster validity indices used
### **Intrinsic Measure:** These metrics evaluate the quality of clustering using only the data features, without referring to ground-truth labels, by measuring how compact each cluster is and how far apart the clusters are from each other.
* **Silhouette Score** measures how well each point fits within its own cluster compared to other clusters. It ranges from −1 to +1, with higher values indicating better-defined and more distinct clusters.
* **Calinski-Harabasz Score** measures how compact the points are within each cluster and how far apart the clusters are from each other. Its typical range is from zero to infinity, with higher values indicating better clustering through tightly grouped points and well-separated clusters.
* **Davies-Bouldin Score** measures the average similarity between each cluster and its most similar one, based on cluster size and distance.It ranges from 0 to infinity, with lower values indicating better clustering through compact, well-separated clusters.
[cell 41 markdown]
## **Extrinsic measure:** These metrics evaluate the quality of clustering by comparing the predicted cluster labels to the true class labels, assessing how well the clustering aligns with the actual ground truth.
* **Random Index** measures the similarity between predicted cluster labels and true labels by counting how many pairs of samples are assigned consistently.
It ranges from 0 to 1, with higher values indicating better agreement between clustering and ground truth.
* **Homogeneity Score** measures whether each cluster contains only data points from a single class. It ranges from 0 to 1, with 1 indicating that all clusters are perfectly pure with respect to the true labels.
* **Completeness Score** measures whether all data points of a given class are assigned to the same cluster. It ranges from 0 to 1, with 1 indicating that each class’s samples are entirely grouped within a single cluster.
* **Adjusted Rand Index** measures the similarity between predicted clusters and true labels, while correcting for chance groupings. Unlike the regular Random Index, which can be artificially high due to random matches, ARI accounts for this and gives a value between −1 and 1, where 1 indicates perfect agreement, 0 indicates random labeling, and negative values mean worse than random clustering.
* **Adjusted Mutual Information** measures the amount of shared information between predicted clusters and true labels, adjusted for chance.
It ranges from 0 to 1 (or slightly negative in rare cases), where higher values indicate better agreement between the clustering and the ground truth, and the adjustment ensures the score isn't inflated by random labelings.
[cell 42 markdown]
# **Visualization of samples grouped by actual bean types and cluster labels**
[cell 43 code]
from sklearn.decomposition import PCA
import matplotlib.patches as mpatches
# Assume:
# - df_beans["Class"] exists (string labels)
# - X_scaled exists (scaled numeric features)
# - y_pred exists (KMeans predicted cluster labels)
class_labels = df_beans["Class"].values
# Fit PCA -> 2 components
pca = PCA(n_components=2, random_state=42)
X_pca = pca.fit_transform(X_scaled)
# Map class names to integers for coloring
unique_classes = np.unique(class_labels)
class_to_color = {cls: i for i, cls in enumerate(unique_classes)}
y_true_named = np.array([class_to_color[cls] for cls in class_labels])
# Plotting
fig, axes = plt.subplots(1, 2, figsize=(16, 7))
# Left: True labels with proper class names
scatter0 = axes[0].scatter(
X_pca[:, 0], X_pca[:, 1],
c=y_true_named, cmap="tab10", s=10, alpha=0.7
)
axes[0].set_title("PCA: Samples grouped by true bean labels", fontsize=14)
axes[0].set_xlabel("PC1")
axes[0].set_ylabel("PC2")
handles0 = [
mpatches.Patch(color=scatter0.cmap(scatter0.norm(i)), label=cls)
for i, cls in enumerate(unique_classes)
]
axes[0].legend(
handles=handles0,
title="Class",
loc="center left",
bbox_to_anchor=(1.02, 0.5),
fontsize=11,
title_fontsize=12,
frameon=True
)
# Right: KMeans cluster labels
scatter1 = axes[1].scatter(
X_pca[:, 0], X_pca[:, 1],
c=y_pred, cmap="tab10", s=10, alpha=0.7
)
axes[1].set_title("PCA: Samples grouped by KMeans cluster labels", fontsize=14)
axes[1].set_xlabel("PC1")
axes[1].set_ylabel("PC2")
unique_clusters = np.unique(y_pred)
handles1 = [
mpatches.Patch(color=scatter1.cmap(scatter1.norm(i)), label=f"Cluster {i}")
for i in unique_clusters
]
axes[1].legend(
handles=handles1,
title="Cluster",
loc="center left",
bbox_to_anchor=(1.02, 0.5),
fontsize=11,
title_fontsize=12,
frameon=True
)
plt.tight_layout()
plt.show()
# Optional: print explained variance
print("Explained variance ratio:", pca.explained_variance_ratio_)
print("Total variance explained by PC1+PC2:", pca.explained_variance_ratio_.sum())
[cell 44 markdown]
* KMeans has formed distinct clusters that broadly correspond to the true bean classes.
* The overall cluster arrangement visually resembles the true class distribution in the reduced 2D space.
* Some clusters show mixing (**BARBUNYA and CALI**), indicating difficulty in separating bean types with similar physical features.
* The clustering is effective, though overlaps between true labels and predicted clusters still remain.
[cell 45 markdown]
# **Evaluation of KMeans Clustering (k = 6)**
[cell 46 code]
from sklearn.metrics import (
silhouette_score, calinski_harabasz_score, davies_bouldin_score,
rand_score, adjusted_rand_score, adjusted_mutual_info_score,homogeneity_score,
completeness_score
)
# KMeans clustering
n_clusters = 6
kmeans = KMeans(n_clusters=n_clusters, n_init=10,random_state=2024)
y_pred = kmeans.fit_predict(X_scaled)
# Intrinsic Measures
SIL_s = silhouette_score(X_scaled, y_pred)
CH_s = calinski_harabasz_score(X_scaled, y_pred)
DB_s = davies_bouldin_score(X_scaled, y_pred)
# Extrinsic Measures
RI_s = rand_score(y_true, y_pred)
ARI_ward = adjusted_rand_score(y_true, y_pred)
AMI_ward = adjusted_mutual_info_score(y_true, y_pred)
CM_s = completeness_score(y_true, y_pred)
HM_s = homogeneity_score(y_true, y_pred)
# Print Results
print('Intrinsic Measure:\n')
print(f'Silhouette score: {SIL_s:.3f}')
print(f'Calinski and Harabasz score: {CH_s:.3f}')
print(f'Davies-Bouldin score: {DB_s:.3f}')
#############################################
print('\nExtrinsic Measure:\n')
print(f'Random index: {RI_s:.3f}')
print(f'Homogeneity score: {HM_s:.3f}')
print(f'Completeness score: {CM_s:.3f}')
print(f'Adjusted Random index: {ARI_ward:.3f}')
print(f'Adjusted Mutual Information: {AMI_ward:.3f}')
[cell 47 markdown]
# **Visualization of samples grouped by actual bean types and cluster labels**
[cell 48 code]
from sklearn.decomposition import PCA
import matplotlib.patches as mpatches
# Assume:
# - df_beans["Class"] exists (string labels)
# - X_scaled exists (scaled numeric features)
# - y_pred exists (KMeans predicted cluster labels)
class_labels = df_beans["Class"].values
# Fit PCA -> 2 components
pca = PCA(n_components=2, random_state=42)
X_pca = pca.fit_transform(X_scaled)
# Map class names to integers for coloring
unique_classes = np.unique(class_labels)
class_to_color = {cls: i for i, cls in enumerate(unique_classes)}
y_true_named = np.array([class_to_color[cls] for cls in class_labels])
# Plotting
fig, axes = plt.subplots(1, 2, figsize=(16, 7))
# Left: True labels with proper class names
scatter0 = axes[0].scatter(
X_pca[:, 0], X_pca[:, 1],
c=y_true_named, cmap="tab10", s=10, alpha=0.7
)
axes[0].set_title("PCA: Samples grouped by true bean labels", fontsize=14)
axes[0].set_xlabel("PC1")
axes[0].set_ylabel("PC2")
handles0 = [
mpatches.Patch(color=scatter0.cmap(scatter0.norm(i)), label=cls)
for i, cls in enumerate(unique_classes)
]
axes[0].legend(
handles=handles0,
title="Class",
loc="center left",
bbox_to_anchor=(1.02, 0.5),
fontsize=11,
title_fontsize=12,
frameon=True
)
# Right: KMeans cluster labels
scatter1 = axes[1].scatter(
X_pca[:, 0], X_pca[:, 1],
c=y_pred, cmap="tab10", s=10, alpha=0.7
)
axes[1].set_title("PCA: Samples grouped by KMeans cluster labels", fontsize=14)
axes[1].set_xlabel("PC1")
axes[1].set_ylabel("PC2")
unique_clusters = np.unique(y_pred)
handles1 = [
mpatches.Patch(color=scatter1.cmap(scatter1.norm(i)), label=f"Cluster {i}")
for i in unique_clusters
]
axes[1].legend(
handles=handles1,
title="Cluster",
loc="center left",
bbox_to_anchor=(1.02, 0.5),
fontsize=11,
title_fontsize=12,
frameon=True
)
plt.tight_layout()
plt.show()
# Optional: print explained variance
print("Explained variance ratio:", pca.explained_variance_ratio_)
print("Total variance explained by PC1+PC2:", pca.explained_variance_ratio_.sum())
[cell 49 markdown]
# **Evaluation of KMeans Clustering (k = 8)**
[cell 50 code]
from sklearn.metrics import (
silhouette_score, calinski_harabasz_score, davies_bouldin_score,
rand_score, adjusted_rand_score, adjusted_mutual_info_score,homogeneity_score,
completeness_score
)
# KMeans clustering
n_clusters = 8
kmeans = KMeans(n_clusters=n_clusters, n_init=10,random_state=2024)
y_pred = kmeans.fit_predict(X_scaled)
# Intrinsic Measures
SIL_s = silhouette_score(X_scaled, y_pred)
CH_s = calinski_harabasz_score(X_scaled, y_pred)
DB_s = davies_bouldin_score(X_scaled, y_pred)
# Extrinsic Measures
RI_s = rand_score(y_true, y_pred)
ARI_ward = adjusted_rand_score(y_true, y_pred)
AMI_ward = adjusted_mutual_info_score(y_true, y_pred)
CM_s = completeness_score(y_true, y_pred)
HM_s = homogeneity_score(y_true, y_pred)
# Print Results
print('Intrinsic Measure:\n')
print(f'Silhouette score: {SIL_s:.3f}')
print(f'Calinski and Harabasz score: {CH_s:.3f}')
print(f'Davies-Bouldin score: {DB_s:.3f}')
#############################################
print('\nExtrinsic Measure:\n')
print(f'Random index: {RI_s:.3f}')
print(f'Homogeneity score: {HM_s:.3f}')
print(f'Completeness score: {CM_s:.3f}')
print(f'Adjusted Random index: {ARI_ward:.3f}')
print(f'Adjusted Mutual Information: {AMI_ward:.3f}')
[cell 51 markdown]
# **Visualization of samples grouped by actual bean types and cluster labels**
[cell 52 code]
from sklearn.decomposition import PCA
import matplotlib.patches as mpatches
# Assume:
# - df_beans["Class"] exists (string labels)
# - X_scaled exists (scaled numeric features)
# - y_pred exists (KMeans predicted cluster labels)
class_labels = df_beans["Class"].values
# Fit PCA -> 2 components
pca = PCA(n_components=2, random_state=42)
X_pca = pca.fit_transform(X_scaled)
# Map class names to integers for coloring
unique_classes = np.unique(class_labels)
class_to_color = {cls: i for i, cls in enumerate(unique_classes)}
y_true_named = np.array([class_to_color[cls] for cls in class_labels])
# Plotting
fig, axes = plt.subplots(1, 2, figsize=(16, 7))
# Left: True labels with proper class names
scatter0 = axes[0].scatter(
X_pca[:, 0], X_pca[:, 1],
c=y_true_named, cmap="tab10", s=10, alpha=0.7
)
axes[0].set_title("PCA: Samples grouped by true bean labels", fontsize=14)
axes[0].set_xlabel("PC1")
axes[0].set_ylabel("PC2")
handles0 = [
mpatches.Patch(color=scatter0.cmap(scatter0.norm(i)), label=cls)
for i, cls in enumerate(unique_classes)
]
axes[0].legend(
handles=handles0,
title="Class",
loc="center left",
bbox_to_anchor=(1.02, 0.5),
fontsize=11,
title_fontsize=12,
frameon=True
)
# Right: KMeans cluster labels
scatter1 = axes[1].scatter(
X_pca[:, 0], X_pca[:, 1],
c=y_pred, cmap="tab10", s=10, alpha=0.7
)
axes[1].set_title("PCA: Samples grouped by KMeans cluster labels", fontsize=14)
axes[1].set_xlabel("PC1")
axes[1].set_ylabel("PC2")
unique_clusters = np.unique(y_pred)
handles1 = [
mpatches.Patch(color=scatter1.cmap(scatter1.norm(i)), label=f"Cluster {i}")
for i in unique_clusters
]
axes[1].legend(
handles=handles1,
title="Cluster",
loc="center left",
bbox_to_anchor=(1.02, 0.5),
fontsize=11,
title_fontsize=12,
frameon=True
)
plt.tight_layout()
plt.show()
# Optional: print explained variance
print("Explained variance ratio:", pca.explained_variance_ratio_)
print("Total variance explained by PC1+PC2:", pca.explained_variance_ratio_.sum())
[cell 53 markdown]
# **Comparison of Kmeans performance at k=6,7 and 8**
[cell 54 code]
data = {
'k': [6, 7, 8],
'Silhouette': [0.360, 0.309, 0.303],
'Calinski-Harabas': [7978.453, 7787.837, 7359.860],
'Davies-Bouldin': [0.996, 1.102, 1.158],
'Rand Index': [0.839, 0.903, 0.915],
'Homogeneity': [0.603, 0.703, 0.746],
'Completeness': [0.736, 0.722, 0.718],
'Adjusted Rand': [0.540, 0.667, 0.697],
'Adjusted Mutual Info': [0.663, 0.712, 0.732]
}
#
df_metrics = pd.DataFrame(data)
df_metrics.set_index('k', inplace=True)
df_metrics
[cell 55 markdown]
* **k = 6** Intrinsic measures are strongest, indicating well-separated and compact clusters. Extrinsic scores are moderate, suggesting only partial alignment with the true class labels.
* **k = 7** Intrinsic quality slightly declines but remains reasonable.
Extrinsic metrics improve overall, showing better correspondence with true labels.
* **k = 8** Intrinsic clustering degrades further, though not drastically. Extrinsic scores are highest, as clusters become purer with respect to class labels.