# colab notebook

course: Module 1 — Foundations of AI & ML
module: Module-1-Foundations-AI-ML
type: notebook
source_url: https://personal-learn.armco.dev/files/Module-1-Foundations-AI-ML/General/2.3_Lab_Materials-Saturday(15-11-2026)/colab_notebook.ipynb

---
[cell 1 markdown]
# **Dataset**


The Auto MPG dataset provides fuel efficiency data (measured as *Miles Per Gallon*, MPG) for various car models from the 1970s and 1980s. The target variable `mpg_high` is a binary label indicating whether a car has high (1) or low (0) fuel efficiency based on its input features.

### Features:

- `cylinders` (int) – Number of engine cylinders *(categorical – discrete values like 4, 6, 8)*
- `displacement` (float) – Size of the engine in cubic inches *(continuous)*
- `horsepower` (float) – Engine horsepower *(continuous)*
- `weight` (float) – Vehicle weight in pounds *(continuous)*
- `acceleration` (float) – Time to accelerate from 0 to 60 mph in seconds *(continuous)*
- `model year` (int) – Year of manufacture *(categorical – ordinal)*
- `origin` (int) – Region of origin (1 = USA, 2 = Europe, 3 = Japan) *(categorical)*
- `car name` (string) – Car model name *(categorical)*
- `mpg_high` (int) – **Target**: 1 if the car has high MPG (above median), 0 otherwise *(binary categorical)*

[cell 2 markdown]
# **Library packages used in this notebook**

1. **warnings**  
   Used to suppress unwanted warning messages for cleaner notebook output *Usage of this package is optional.*


2. **pandas**  
   Used for data loading, manipulation, and analysis — including reading CSV files, handling missing values, and working with DataFrames.

3. **numpy**  
   Used for numerical operations such as creating missing values, setting random seeds, and performing array-based computations.

4. **sklearn.impute.SimpleImputer**  
   Available in the `scikit-learn` package — a machine learning library in Python.  
   Used to impute missing values using strategies like mean, median, and mode.

5. **sklearn.decomposition.PCA**  
   Available in the `scikit-learn` package — a machine learning library in Python.  
   Used to perform Principal Component Analysis (PCA) for dimensionality reduction and variance analysis.

6. **sklearn.preprocessing.StandardScaler**  
   Available in the `scikit-learn` package — a machine learning library in Python.  
   Used to standardize features by removing the mean and scaling to unit variance before applying PCA.

7. **matplotlib.pyplot**  
   Used for creating basic plots such as scatter plots, bar plots, and biplots for PCA visualization.

8. **seaborn**  
   Used for advanced and aesthetically pleasing visualizations such as count plots, boxplots, violin plots, KDE plots, heatmaps, and pairplots.

---

These packages come pre-installed in Google Colab.  
To run this notebook locally, you need to install the required packages using the following command in terminal:

```bash
pip install pandas numpy scikit-learn matplotlib seaborn

[cell 3 markdown]
# **Loading the dataset**

[cell 4 code]
import pandas as pd
import numpy as np
import warnings
warnings.filterwarnings('ignore')


df=pd.read_csv('auto_mpg_binarised.csv')
df

[cell 5 code]
df.shape

[cell 6 code]
df.info()

[cell 7 markdown]
* There are no missing values in the dataset.
* From the `.info()` output, it appears that only one feature — `car_name` — is explicitly stored as a categorical (object) type.

* However, other features with `int` or `float` types may actually represent categorical variables (e.g., `cylinders`, `origin`). These should be further investigated using their number of unique values to determine if they are nominal or ordinal categories misrepresented as numerical data.

[cell 8 code]
for col in df.columns:
    print(f"{col}: {len(df[col].unique())} unique values")
    print("")

[cell 9 markdown]
### **Observation on Unique Values**

- `cylinders`, `model_year`, and `origin` have few unique values and are categorical.
- `displacement`, `horsepower`, `weight`, and `acceleration` have many unique values and are continuous.
- `mpg_high` is binary and suitable as a classification target.
- We will use `displacement`, `horsepower`, `weight` and  `acceleration` as numeric cols for further analysis

[cell 10 markdown]
# **Part One: Missing values analysis**

[cell 11 markdown]
## **Randomly inserted missing values in 10% of each continuous numeric column to test imputation methods.**

[cell 12 code]
# Create a copy to preserve original
df_10 = df.copy()

# True continuous numeric variables
numeric_cols = ['displacement', 'horsepower', 'weight', 'acceleration']

# Fix seed for reproducibility
np.random.seed(42)


for col in numeric_cols:
    n_missing = int(0.1 * len(df_10))
    missing_rows = np.random.choice(df_10.index, size=n_missing, replace=False)
    df_10.loc[missing_rows, col] = np.nan

[cell 13 code]
df_10.head(20)

[cell 14 markdown]
## **Imputation Strategies Applied**


- **Attribute Mean Imputation:** Replaced missing values with the mean of each numeric column.
- **Attribute Median Imputation:** Replaced missing values with the median of each numeric column.
- **Attribute Mode Imputation:** Replaced missing values with the most frequent value (mode) across the entire dataset for each column.
- **Class-wise Mode Imputation:** For each column with missing values, the dataset is grouped by the `mpg_high class` label. Then, for each group, missing values in that column are replaced with the mode computed only from rows within that same class. This ensures that the imputed value reflects the distribution specific to each class

[cell 15 code]
from sklearn.impute import SimpleImputer

# 1. Imputation by mean
imputer_mean = SimpleImputer(strategy='mean')
df_mean_10 = df_10.copy()
df_mean_10[numeric_cols] = imputer_mean.fit_transform(df_mean_10[numeric_cols])

# 2. Imputation by median
imputer_median = SimpleImputer(strategy='median')
df_median_10 = df_10.copy()
df_median_10[numeric_cols] = imputer_median.fit_transform(df_median_10[numeric_cols])

# 3. Imputation by mode (global most frequent)
imputer_mode = SimpleImputer(strategy='most_frequent')
df_mode_10 = df_10.copy()
df_mode_10[numeric_cols] = imputer_mode.fit_transform(df_mode_10[numeric_cols])

# 4. Imputation by mode within each class
df_mode_by_class_10 = df_10.copy()
for label in df_mode_by_class_10['mpg_high'].unique():
    class_mask = df_mode_by_class_10['mpg_high'] == label
    class_df = df_mode_by_class_10.loc[class_mask, numeric_cols]
    class_mode = class_df.mode().iloc[0]
    df_mode_by_class_10.loc[class_mask, numeric_cols] = class_df.fillna(class_mode)

[cell 16 markdown]
# **Principal Component Analysis (PCA) to Compare the Impact of Different Imputation Strategies**

[cell 17 markdown]
### **PCA Analysis of Imputed Datasets**

**Principal Component Analysis (PCA)** is a technique that transforms the data into a new coordinate system such that the greatest variance lies on the first axis (PC1), the second greatest on the second axis (PC2), and so on. It reduces dimensionality while preserving the most important patterns in the data.

We applied PCA to the original and imputed datasets to:
- Quantify how much variance is captured by each principal component.
- Visualize the data in 2D (PC1 vs. PC2) to compare the structure and class separation across different imputation strategies.

[cell 18 code]
from sklearn.decomposition import PCA
from sklearn.preprocessing import StandardScaler
import matplotlib.pyplot as plt
import pandas as pd
import numpy as np

def analyze_single_pca(df, name, numeric_cols, original_reference=None):
    print(f"\n\n==== {name} ====")
    X = df[numeric_cols]
    y = df['mpg_high'].astype(str)

    scaler = StandardScaler()
    X_scaled = scaler.fit_transform(X)

    pca = PCA()
    X_pca_full = pca.fit_transform(X_scaled)

    # Print explained variance for all components
    explained_var = pca.explained_variance_ratio_ * 100
    print("Explained Variance (%):")
    for i, var in enumerate(explained_var, start=1):
        print(f"  PC{i}: {var:.2f}%")

    # Loadings matrix (full) (Each column is a unit PC vector )
    loadings = pd.DataFrame(pca.components_.T,
                            columns=[f'PC{i+1}' for i in range(pca.n_components_)],
                            index=numeric_cols)
    print("\nLoadings Matrix:")
    print(loadings)

    # Bar plot for PC1 contributions
    plt.figure(figsize=(8, 4))
    loadings['PC1'].plot(kind='bar', title=f'Feature Contributions to PC1 - {name}', color='darkcyan')
    plt.ylabel('Contribution')
    plt.tight_layout()
    plt.show()

    # Define 2D biplot using first 2 PCs
    def biplot(score, coeff, labels, y, title):
        xs = score[:, 0]
        ys = score[:, 1]

        # Plot each class with its own color and label
        for cls, color in zip(['0', '1'], ['tab:blue', 'tab:orange']):
            idx = (y == cls)
            plt.scatter(xs[idx], ys[idx], alpha=0.5, c=color, label=f'Class {cls}')


        plt.xlabel("PC1")
        plt.ylabel("PC2")
        plt.title(title)
        plt.grid()
        plt.legend()

    if original_reference is None:
        plt.figure(figsize=(6, 5))
        biplot(X_pca_full, pca.components_.T, labels=numeric_cols, y=y, title=f"Biplot - {name}")
        plt.tight_layout()
        plt.show()
    else:
        orig_name, orig_df = original_reference
        orig_X = orig_df[numeric_cols]
        orig_y = orig_df['mpg_high'].astype(str)
        orig_scaled = StandardScaler().fit_transform(orig_X)
        orig_pca = PCA().fit(orig_scaled)
        orig_scores = orig_pca.transform(orig_scaled)
        orig_components = orig_pca.components_.T

        fig, axs = plt.subplots(1, 2, figsize=(12, 5))

        plt.sca(axs[0])
        biplot(orig_scores, orig_components, labels=numeric_cols, y=orig_y, title=f"{orig_name}")

        plt.sca(axs[1])
        biplot(X_pca_full, pca.components_.T, labels=numeric_cols, y=y, title=f"{name}")

        plt.tight_layout()
        plt.show()

[cell 19 code]
# First, for the original dataset:
analyze_single_pca(df, "Original", numeric_cols)

[cell 20 code]
analyze_single_pca(df_mean_10, "Mean Imputed", numeric_cols, original_reference=("Original", df))

[cell 21 code]
analyze_single_pca(df_median_10, "Median Imputed", numeric_cols, original_reference=("Original", df))

[cell 22 code]
analyze_single_pca(df_mode_10, "Mode Imputed", numeric_cols, original_reference=("Original", df))

[cell 23 code]
analyze_single_pca(df_mode_by_class_10, "Mode by Class", numeric_cols, original_reference=("Original", df))

[cell 24 markdown]
- The original dataset has the highest variance explained by PC1 (80%), meaning the first principal component captures most of the variability in the data.
- Among all imputed versions, **Mode by Class** preserves the most variance in PC1 (76%), making it the closest to the original in terms of variance distribution.
- Mean and median imputations explain slightly less variance in PC1 (around 73%), while mode explains the least (70%), with more spread into later components.

[cell 25 markdown]
# **Randomly inserted missing values in 30% of each continuous numeric column to test imputation methods.**

[cell 26 code]
import numpy as np
import pandas as pd
from sklearn.impute import SimpleImputer

# Start from a clean version of the original dataframe (with no missing values)
df_base = df.copy()  # Assuming df has no NaNs after initial horsepower imputation

# Define the numeric columns
numeric_cols = ['displacement', 'horsepower', 'weight', 'acceleration']

# Create df_30 by injecting 30% NaNs into the numeric columns
df_30 = df_base  #df_base.copy()

for col in numeric_cols:
    n_missing = int(0.3 * len(df_30))
    missing_rows = np.random.choice(df_30.index, size=n_missing, replace=False)
    df_30.loc[missing_rows, col] = np.nan

[cell 27 code]
df_30.head(20)

[cell 28 markdown]
## **Imputation Strategies Applied**


- **Attribute Mean Imputation:** Replaced missing values with the mean of each numeric column.
- **Attribute Median Imputation:** Replaced missing values with the median of each numeric column.
- **Attribute Mode Imputation:** Replaced missing values with the most frequent value (mode) across the entire dataset for each column.
- **Class-wise Mode Imputation:** For each class label in `mpg_high`, missing values were filled using the mode calculated from rows belonging only to that class.

[cell 29 code]
# 1. Imputation by mean
imputer_mean = SimpleImputer(strategy='mean')
df_mean_30 = df_30.copy()
imputed_mean = imputer_mean.fit_transform(df_mean_30[numeric_cols])
df_mean_30[numeric_cols] = pd.DataFrame(imputed_mean, columns=numeric_cols, index=df_mean_30.index)

# 2. Imputation by median
imputer_median = SimpleImputer(strategy='median')
df_median_30 = df_30.copy()
imputed_median = imputer_median.fit_transform(df_median_30[numeric_cols])
df_median_30[numeric_cols] = pd.DataFrame(imputed_median, columns=numeric_cols, index=df_median_30.index)

# 3. Imputation by mode (global)
imputer_mode = SimpleImputer(strategy='most_frequent')
df_mode_30 = df_30.copy()
imputed_mode = imputer_mode.fit_transform(df_mode_30[numeric_cols])
df_mode_30[numeric_cols] = pd.DataFrame(imputed_mode, columns=numeric_cols, index=df_mode_30.index)

# 4. Imputation by mode within each class
df_mode_by_class_30 = df_30.copy()
for label in df_mode_by_class_30['mpg_high'].unique():
    class_mask = df_mode_by_class_30['mpg_high'] == label
    class_df = df_mode_by_class_30.loc[class_mask, numeric_cols]
    class_mode = class_df.mode().iloc[0]
    df_mode_by_class_30.loc[class_mask, numeric_cols] = class_df.fillna(class_mode)

[cell 30 markdown]
# **Principal Component Analysis (PCA) to Compare the Impact of Different Imputation Strategies**

[cell 31 markdown]
## **PCA visulalization of original data**

[cell 32 code]
# First, for the original dataset:
analyze_single_pca(df, "Original", numeric_cols)

[cell 33 markdown]
## **PCA visulalization of Mean Imputed data**

[cell 34 code]
analyze_single_pca(df_mean_30, "Mean Imputed", numeric_cols, original_reference=("Original", df))

[cell 35 markdown]
## **PCA visulalization of Median Imputed data**

[cell 36 code]
analyze_single_pca(df_median_30, "Median Imputed", numeric_cols, original_reference=("Original", df))

[cell 37 markdown]
## **PCA visulalization of Mode Imputed data**

[cell 38 code]
analyze_single_pca(df_mode_30, "Mode Imputed", numeric_cols, original_reference=("Original", df))

[cell 39 markdown]
## **PCA visulalization of class wise Mode Imputed data**

[cell 40 code]
analyze_single_pca(df_mode_by_class_30, "Mode by class", numeric_cols, original_reference=("Original", df))

[cell 41 markdown]
## **PCA Observations After 30% Missing Data**

- As the proportion of missing data increases, all imputation methods show a **drop in PC1 variance**, indicating loss of information.
- **Mode by Class** again retains the highest PC1 variance (71%) after the original (80%), suggesting it handles high missingness better than others.
- **Mean** and **Median** imputations show moderate PC1 variance (around 62% and 61% resp.), while **Mode Imputation** performs the worst (PC1 = 54 %).

[cell 42 markdown]
# **Final Dataset Selection**:
Since the original dataset has no missing values and preserves the most variance in PCA while showing clear class separation, we proceed with the original dataset for all further visualizations and analysis.

[cell 43 markdown]
# **Part Two: EDA Analysis**

[cell 44 code]
import seaborn as sns
import matplotlib.pyplot as plt

def plot_categorical_feature(df, col):
    sns.set(style="whitegrid")

    plot_data = df.copy()
    if col == 'origin':
        plot_data['origin'] = plot_data['origin'].map({1: 'USA', 2: 'Europe', 3: 'Japan'})

    plt.figure(figsize=(8, 6))
    sns.countplot(data=plot_data, x=col, color='skyblue')
    plt.title(f"Count Plot of {col}")
    plt.xlabel(col)
    plt.ylabel("Count")
    plt.tight_layout()
    plt.show()


plot_categorical_feature(df, 'cylinders')

[cell 45 code]
plot_categorical_feature(df, 'model_year')

[cell 46 code]
plot_categorical_feature(df, 'origin')

[cell 47 markdown]
# **Pairplots for numeric features**

[cell 48 code]
import seaborn as sns
import matplotlib.pyplot as plt


# Create the pairplot (no hue)
pairplot = sns.pairplot(df, vars=numeric_cols, diag_kind='kde')

# Add title and adjust layout
plt.suptitle("Pairplot of Continuous Features", y=1.02)
plt.tight_layout()
plt.show()

[cell 49 markdown]
### **Observations from Pairplot**

- `displacement`, `horsepower`, and `weight` show strong positive linear relationships with each other. This is consistent with their high correlation seen earlier.
- `acceleration` is negatively related to all three above — especially `horsepower` and `weight`, though the relationships are weaker and more scattered.
- The diagonal plots (KDE curves) show that `displacement`, `horsepower`, and `weight` are **right-skewed**, while `acceleration` appears **roughly symmetric**.

[cell 50 markdown]
# **Pairplot of numeric cols with Class-Based Coloring**

[cell 51 code]
import seaborn as sns
import matplotlib.pyplot as plt

# Define continuous features
numeric_cols = ['displacement', 'horsepower', 'weight', 'acceleration']

# Create pairplot (no tight layout yet)
pairplot = sns.pairplot(df, vars=numeric_cols, hue='mpg_high', palette='Set1', diag_kind='kde')

# Manually place legend outside the plot area
pairplot._legend.remove()  # remove auto-legend
plt.legend(
    title="mpg_high (0 = Low, 1 = High)",
    labels=["0", "1"],
    loc='center left',
    bbox_to_anchor=(1, 0.5),
    borderaxespad=0.
)

# Adjust layout
plt.suptitle("Pairplot of Continuous Features by class label", y=1.02)
plt.tight_layout()
plt.show()

[cell 52 markdown]
### **Observations from Class-Colored Pairplot**


- High-efficiency cars (red) tend to have **lower displacement**, **lower horsepower**, and **lower weight** — clearly visible in KDE curves.
- Low-efficiency cars (blue) cluster toward the **higher end** of displacement, horsepower, and weight — showing a clear inverse relationship with fuel efficiency.
- `acceleration` does not show a strong separation.
- Feature separation is especially sharp for `weight`, making it a potentially strong predictor for classification.

These patterns suggest that class labels (`mpg_high`) are meaningfully separated by several features, especially displacement, horsepower, and weight.

[cell 53 markdown]
# **Heatmap for numeric features**

[cell 54 code]
import seaborn as sns
import matplotlib.pyplot as plt

# Only numeric columns (including target if you want to see its correlation)
numeric_cols_all = ['displacement', 'horsepower', 'weight', 'acceleration', 'mpg_high']

# Compute correlation matrix
corr_matrix = df[numeric_cols_all].corr()

# Plot heatmap
plt.figure(figsize=(8, 6))
sns.heatmap(corr_matrix, annot=True, cmap='coolwarm', center=0, fmt=".2f")
plt.title("Correlation Heatmap of Continuous Features")
plt.tight_layout()
plt.show()

[cell 55 markdown]
### **Correlation Analysis of Continuous Features**

The heatmap shows several strong correlations:

- `displacement`, `horsepower`, and `weight` have high positive correlation with each other (correlation > 0.85). In supervised learning, we generally expect input features to be uncorrelated with each other and to show strong correlation with the target variable. We can use PCA to combine these correlated features into fewer, more uncorrelated features.

- `acceleration` is negatively correlated with all three above, especially `horsepower` (−0.68), suggesting that cars with higher horsepower tend to accelerate faster (i.e., reach 60 mph in less time).

- The target variable `mpg_high` shows strong negative correlation with `displacement` (−0.74), `weight` (−0.75), and `horsepower` (−0.64), meaning cars with higher values in these features are more likely to have **low** fuel efficiency.
- `acceleration` shows a weak positive correlation with `mpg_high` (+0.32), suggesting that cars which take longer to reach 60 mph (i.e., accelerate more slowly) tend to have slightly better fuel efficiency — though the relationship is not very strong.

[cell 56 markdown]
# **Univariate Analysis**

[cell 57 code]
import matplotlib.pyplot as plt
import seaborn as sns

# Continuous features
numeric_cols = ['displacement', 'horsepower', 'weight', 'acceleration']

# Style
sns.set(style="whitegrid")

# Plot loop (no hue)
for col in numeric_cols:
    fig, axes = plt.subplots(1, 3, figsize=(18, 6))

    # 1. Histogram
    sns.histplot(data=df, x=col, bins=20, kde=False, ax=axes[0])
    axes[0].set_title(f"{col} - Histogram")

    # 2. KDE
    sns.kdeplot(data=df, x=col, fill=True, ax=axes[1])
    axes[1].set_title(f"{col} - KDE Plot")

    # 3. Vertical Boxplot
    sns.boxplot(data=df, y=col, ax=axes[2])
    axes[2].set_title(f"{col} - Boxplot")

    plt.suptitle(f"Distribution of '{col}'", fontsize=14, y=1.05)
    plt.tight_layout()
    plt.show()

[cell 58 markdown]
### **Observations**

- **Displacement** is **right-skewed**, with most values concentrated between 100 and 250. The KDE (Knowledge Density Estimation) curve shows a long tail toward higher values. The boxplot confirms a wide range with a few high-end values, but no extreme outliers.
- **Horsepower** is also **right-skewed**, with a strong peak around 80–100. The KDE curve shows multiple modes, indicating possible subgroups. The boxplot reveals several outliers on the higher end (above 200).
- **Weight** is right-skewed, with most values between 2000 and 3500. The KDE plot confirms a long tail towards heavier cars, and the boxplot shows no extreme outliers, but a wide spread.
- **Acceleration** is approximately symmetric and bell-shaped, suggesting a nearly normal distribution. Most values fall between 13 and 19, as seen in both histogram and KDE plot.
- The boxplot for acceleration reveals a few high-end outliers above 22, but the overall spread is compact.
- Unlike the other three features, acceleration appears well-centered and less skewed

[cell 59 markdown]
# **Univariate Analysis by class-based coloring**

[cell 60 code]
import matplotlib.pyplot as plt
import seaborn as sns

# Your continuous/numeric features
numeric_cols = ['displacement', 'horsepower', 'weight', 'acceleration']

# Set seaborn style
sns.set(style="whitegrid")

# Create plots
for col in numeric_cols:
    fig, axes = plt.subplots(1, 3, figsize=(18, 6))

    # 1. Histogram
    sns.histplot(data=df, x=col, hue='mpg_high', multiple='stack', bins=20, ax=axes[0])
    axes[0].set_title(f"{col} - Histogram")

    # 2. KDE Plot
    sns.kdeplot(data=df, x=col, hue='mpg_high', fill=True, common_norm=False, ax=axes[1])
    axes[1].set_title(f"{col} - KDE Plot")

    # 3. Boxplot (fix axis)
    sns.boxplot(data=df, x='mpg_high', y=col, ax=axes[2])
    axes[2].set_title(f"{col} - Boxplot")
    axes[2].set_xlabel("mpg_high")
    axes[2].set_ylabel(col)

    plt.suptitle(f"Distribution of '{col}' by Class-Based Coloring", fontsize=14, y=1.05)
    plt.tight_layout()
    plt.show()

[cell 61 markdown]
### **Observation from  pairplot with class-based coloring**

[cell 62 markdown]
### **Class-Wise Distribution Summary**

- **Displacement:** Histogram and KDE show that high-efficiency cars are clustered at low displacement, while low-efficiency cars are spread over a wider, higher range. Boxplot shows a clear downward shift in median displacement for high-efficiency cars with no major overlap.

- **Horsepower:** Histogram and KDE indicate high-efficiency cars peak at lower horsepower, with low-efficiency cars having a longer right tail. Boxplot shows lower median and tighter IQR (Inter Quartile Range) for high-efficiency cars, while low-efficiency cars show more spread and outliers.


- **Weight:** Histogram and KDE show high-efficiency cars are concentrated at lower weights, while low-efficiency cars span a much higher range.  Boxplot clearly separates the two classes, with high-efficiency cars having much lower median and tighter spread.

- **Acceleration:** Histogram and KDE show high-efficiency cars tend to have slightly higher acceleration values, though overlap exists. Boxplot shows a modest upward shift in median acceleration for high-efficiency cars, with more outliers on both sides.

[cell 63 code]
import matplotlib.pyplot as plt
import seaborn as sns

# Define numeric features
numeric_cols = ['displacement', 'horsepower', 'weight', 'acceleration']

# Set style
sns.set(style="whitegrid")

# Create 2x2 subplots
fig, axes = plt.subplots(2, 2, figsize=(12, 10))
axes = axes.flatten()

for i, col in enumerate(numeric_cols):
    sns.violinplot(data=df, y=col, ax=axes[i], inner='box', palette='Set2')
    axes[i].set_title(f"Violin Plot of {col}")
    axes[i].set_ylabel(col)
    axes[i].set_xlabel("")

plt.tight_layout()
plt.show()

[cell 64 markdown]
### **Observation from violinplot**


Violin plots combine the distributional insight of KDE plots with the summary statistics of boxplots into a single visual. The patterns observed here are consistent with previous analyses: displacement, horsepower, and weight are right-skewed with broader spread, while acceleration is more symmetric and tightly distributed.

[cell 65 markdown]
# **Violin plots based on class-based coloring**

[cell 66 code]
import matplotlib.pyplot as plt
import seaborn as sns

# Make sure mpg_high is treated as a category
df['mpg_high'] = df['mpg_high'].astype(str)

# Define numeric features
numeric_cols = ['displacement', 'horsepower', 'weight', 'acceleration']

# Set seaborn style
sns.set(style="whitegrid")

# Create 2x2 subplot layout
fig, axes = plt.subplots(2, 2, figsize=(12, 10))
axes = axes.flatten()

# Plot vertical violins for each feature split by mpg_high
for i, col in enumerate(numeric_cols):
    sns.violinplot(data=df, x='mpg_high', y=col, inner='box', palette='Set2', ax=axes[i])
    axes[i].set_title(f"{col}")
    axes[i].set_xlabel("mpg_high")
    axes[i].set_ylabel(col)

plt.tight_layout()
plt.suptitle("Violin Plots of Numeric Features by Class", fontsize=16, y=1.03)
plt.show()

[cell 67 markdown]
### **Observation from violin plot based on class-based coloring**

These violin plots combine KDE and boxplot information to show class-wise distributions. As before, high-efficiency cars (`mpg_high = 1`) have visibly lower displacement, horsepower, and weight, while acceleration is slightly higher. Displacement and weight show the strongest separation between classes.