# LinearRegression NaiveBayes
course: Module 2 — Machine Learning Algorithms
module: Module-2-Machine-Learning-Algorithms
type: notebook
source_url: https://personal-learn.armco.dev/files/Module-2-Machine-Learning-Algorithms/General/Lab_Materials/LinearRegression_NaiveBayes.ipynb
---
[cell 1 markdown]
# Part A : Linear Regression
[cell 2 markdown]
# Bike Sharing Dataset Overview
The **Bike Sharing Dataset** contains data collected from a bike rental system over a two-year period (2011–2012). It tracks hourly demand for shared bikes in relation to environmental and seasonal conditions.
The goal is to **predict the number of bike rentals** (`cnt`) based on time, weather, and user behavior variables.
---
## Source:
Available via the UCI Machine Learning Repository:
* [UCI Dataset Page](https://archive.ics.uci.edu/ml/datasets/Bike+Sharing+Dataset)
---
## Target Variable:
* **cnt** *(numeric)*: Total number of bike rentals (casual + registered) in a given hour.
---
## Features:
| Column Name | Description | Data Type |
| ----------- | ---------------------------------------------------------------- | ------------------ |
| instant | Record index (can be dropped) | *Integer* |
| season | Season (1: Spring, 2: Summer, 3: Fall, 4: Winter) | *Categorical* |
| yr | Year (0: 2011, 1: 2012) | *Binary* |
| mnth | Month (1 to 12) | *Categorical* |
| hr | Hour of the day (0 to 23) | *Categorical* |
| holiday | Whether the day is a holiday | *Binary* |
| weekday | Day of the week (0: Sunday, 6: Saturday) | *Categorical* |
| workingday | Whether the day is a working day | *Binary* |
| weathersit | Weather situation (1: Clear, 2: Mist, 3: Light Snow/Rain, etc.) | *Categorical* |
| temp | Normalized temperature (0 to 1 scale) | *Numeric (float)* |
| atemp | Normalized "feels like" temperature | *Numeric (float)* |
| hum | Normalized humidity (0 to 1 scale) | *Numeric (float)* |
| windspeed | Normalized wind speed | *Numeric (float)* |
| casual | Count of casual users (do **not** use when predicting `cnt`) | *Numeric* |
| registered | Count of registered users (do **not** use when predicting `cnt`) | *Numeric* |
| cnt | **Target**: Total number of rentals (casual + registered) | **Numeric Target** |
---
### Note on Leakage:
The features `casual` and `registered` **should not be used** when predicting `cnt`, as they are **components of the target variable** and would introduce data leakage.
[cell 3 markdown]
# **Loading the dataset into a datafarme**
[cell 4 code]
import pandas as pd
df = pd.read_csv("bike_sharing.csv")
print(df.shape)
df
[cell 5 code]
df.info()
[cell 6 markdown]
# **Checking for missing values**
[cell 7 code]
df.isna().sum()
[cell 8 code]
df_clean = df.drop(columns=["instant", "casual", "registered"])
categorical_cols = ["season", "yr", "mnth", "hr", "holiday",
"weekday", "workingday", "weathersit"]
df_clean[categorical_cols] = df_clean[categorical_cols].astype("object")
df_clean.info()
[cell 9 markdown]
# **Displaying unique values of categorical columns**
[cell 10 code]
for col in categorical_cols:
unique_vals = df[col].unique()
print(f"{col} ({len(unique_vals)} unique): {unique_vals}")
print('------------------------------------------------')
[cell 11 markdown]
# **Distribution of target variable**
[cell 12 code]
df_clean["cnt"].plot(kind="hist", bins=40, density=True, alpha=0.6, figsize=(8, 5), title="Histogram of Total Bike Rentals")
#df_clean["cnt"].plot(kind="kde", linewidth=2)
[cell 13 markdown]
* The distribution is highly right-skewed.
* There are many low rental counts and a long tail of high counts.
[cell 14 markdown]
# **Splitting the dataset into train-test sets**
[cell 15 code]
from sklearn.model_selection import train_test_split
# Drop 'ID' column
X = df_clean.drop(['cnt'], axis=1) # Features
y = df_clean['cnt'] # Target column
# Perform train-test split (80% train, 20% test)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42,
)
# Display shapes
print(f"X_train shape: {X_train.shape}")
print(f"X_test shape: {X_test.shape}")
print(f"y_train shape: {y_train.shape}")
print(f"y_test shape: {y_test.shape}")
[cell 16 markdown]
# **Linear regression model with categorical variables (a) Integer Encoded (b) One hot Encoded. All numeric features are normalized to have zero mean and standard deviation one.**
[cell 17 code]
from sklearn.preprocessing import StandardScaler, OrdinalEncoder
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
import numpy as np
# Identify numeric and categorical columns
numeric_cols = X.select_dtypes(include=np.number).columns.tolist()
categorical_cols = X.select_dtypes(include='object').columns.tolist()
# Split into numeric and categorical
X_train_num = X_train[numeric_cols]
X_test_num = X_test[numeric_cols]
X_train_cat = X_train[categorical_cols]
X_test_cat = X_test[categorical_cols]
# Scale numeric features
scaler = StandardScaler()
X_train_num_scaled = scaler.fit_transform(X_train_num)
X_test_num_scaled = scaler.transform(X_test_num)
# Encode categorical features
encoder = OrdinalEncoder()
X_train_cat_encoded = encoder.fit_transform(X_train_cat)
X_test_cat_encoded = encoder.transform(X_test_cat)
# Combine numeric and categorical features
X_train_final = np.hstack((X_train_num_scaled, X_train_cat_encoded))
X_test_final = np.hstack((X_test_num_scaled, X_test_cat_encoded))
# Train linear regression
reg = LinearRegression()
reg.fit(X_train_final, y_train)
# Predict
y_pred = reg.predict(X_test_final)
# Regression Metrics
mae = mean_absolute_error(y_test, y_pred)
mse = mean_squared_error(y_test, y_pred)
rmse = np.sqrt(mse)
r2 = r2_score(y_test, y_pred)
# Print metrics
# Print results
print("\n=== Linear Regression with Integer Encoding ===")
print(f"MAE : {mae:.2f}")
print(f"MSE : {mse:.2f}")
print(f"RMSE : {rmse:.2f}")
print(f"R² : {r2:.4f}")
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
import numpy as np
# Split features
X_train_num_oh = X_train[numeric_cols]
X_test_num_oh = X_test[numeric_cols]
X_train_cat_oh = X_train[categorical_cols]
X_test_cat_oh = X_test[categorical_cols]
# Scale numeric features
scaler_oh = StandardScaler()
X_train_num_scaled_oh = scaler_oh.fit_transform(X_train_num_oh)
X_test_num_scaled_oh = scaler_oh.transform(X_test_num_oh)
# One-hot encode categorical features # drop='first'
ohe = OneHotEncoder(handle_unknown='ignore', sparse_output=False)
X_train_cat_encoded_oh = ohe.fit_transform(X_train_cat_oh)
X_test_cat_encoded_oh = ohe.transform(X_test_cat_oh)
# Combine numeric and categorical features
X_train_final_oh = np.hstack((X_train_num_scaled_oh, X_train_cat_encoded_oh))
X_test_final_oh = np.hstack((X_test_num_scaled_oh, X_test_cat_encoded_oh))
# Train linear regression
reg_oh = LinearRegression()
reg_oh.fit(X_train_final_oh, y_train)
# Predict
y_pred_oh = reg_oh.predict(X_test_final_oh)
# Evaluate
mae_oh = mean_absolute_error(y_test, y_pred_oh)
mse_oh = mean_squared_error(y_test, y_pred_oh)
rmse_oh = np.sqrt(mse_oh)
r2_oh = r2_score(y_test, y_pred_oh)
# Print results
print("\n=== Linear Regression with One-Hot Encoding ===")
print(f"MAE : {mae_oh:.2f}")
print(f"MSE : {mse_oh:.2f}")
print(f"RMSE : {rmse_oh:.2f}")
print(f"R² : {r2_oh:.4f}")
[cell 18 markdown]
# **Showing actual values of target variable `cnt` (Actual_cnt) and its predicted values (Predicted_cnt) when using Linear Regression with Integer encoding for some samples**
[cell 19 code]
df_results_ordinal = X_test.copy()
df_results_ordinal['Actual_cnt'] = y_test.values
df_results_ordinal['Predicted_cnt'] = y_pred
df_sample_ordinal = df_results_ordinal.sample(10, random_state=42)
df_sample_ordinal
[cell 20 markdown]
# **Showing actual values of target variable `cnt` (Actual_cnt) and its predicted values (Predicted_cnt) when using Linear Regression with One hot encoding for some samples**
[cell 21 code]
df_results_onehot = X_test.copy()
df_results_onehot['Actual_cnt'] = y_test.values
df_results_onehot['Predicted_cnt'] = y_pred_oh
df_sample_onehot = df_results_onehot.sample(10, random_state=42)
df_sample_onehot
[cell 22 markdown]
### **The deviation of predicted values from the actual values is smaller with one-hot encoding, while it is larger with integer encoding.**
[cell 23 markdown]
# **Linear regression is performed with the target variable log-transformed and categorical variables encoded using (a) integer encoding and (b) one-hot encoding. All numeric features are normalized to have zero mean and a standard deviation of one**
[cell 24 code]
from sklearn.preprocessing import StandardScaler, OrdinalEncoder, OneHotEncoder
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
import numpy as np
# --- Identify numeric and categorical columns ---
numeric_cols = X.select_dtypes(include=np.number).columns.tolist()
categorical_cols = X.select_dtypes(include='object').columns.tolist()
# --- Log-transform the target variable ---
y_train_log = np.log1p(y_train)
y_test_log = np.log1p(y_test)
# ==============================
# ORDINAL ENCODING
# ==============================
# Prepare features
X_train_num = X_train[numeric_cols]
X_test_num = X_test[numeric_cols]
X_train_cat = X_train[categorical_cols]
X_test_cat = X_test[categorical_cols]
# Scale numeric
scaler = StandardScaler()
X_train_num_scaled = scaler.fit_transform(X_train_num)
X_test_num_scaled = scaler.transform(X_test_num)
# Ordinal encode categorical
encoder = OrdinalEncoder()
X_train_cat_encoded = encoder.fit_transform(X_train_cat)
X_test_cat_encoded = encoder.transform(X_test_cat)
# Combine features
X_train_final = np.hstack((X_train_num_scaled, X_train_cat_encoded))
X_test_final = np.hstack((X_test_num_scaled, X_test_cat_encoded))
# Train model
reg = LinearRegression()
reg.fit(X_train_final, y_train_log)
# Predict and inverse log
y_pred_log = reg.predict(X_test_final)
y_pred = np.expm1(y_pred_log)
# Evaluate
mae = mean_absolute_error(y_test, y_pred)
mse = mean_squared_error(y_test, y_pred)
rmse = np.sqrt(mse)
r2 = r2_score(y_test, y_pred)
print("\n=== Linear Regression with Integer Encoding (Log-Transformed Target) ===")
print(f"MAE : {mae:.2f}")
print(f"MSE : {mse:.2f}")
print(f"RMSE : {rmse:.2f}")
print(f"R² : {r2:.4f}")
# ==============================
# ONE-HOT ENCODING
# ==============================
# Prepare features
X_train_num_oh = X_train[numeric_cols]
X_test_num_oh = X_test[numeric_cols]
X_train_cat_oh = X_train[categorical_cols]
X_test_cat_oh = X_test[categorical_cols]
# Scale numeric
scaler_oh = StandardScaler()
X_train_num_scaled_oh = scaler_oh.fit_transform(X_train_num_oh)
X_test_num_scaled_oh = scaler_oh.transform(X_test_num_oh)
# One-hot encode categorical features
ohe = OneHotEncoder(handle_unknown='ignore', sparse_output=False)
X_train_cat_encoded_oh = ohe.fit_transform(X_train_cat_oh)
X_test_cat_encoded_oh = ohe.transform(X_test_cat_oh)
# Combine features
X_train_final_oh = np.hstack((X_train_num_scaled_oh, X_train_cat_encoded_oh))
X_test_final_oh = np.hstack((X_test_num_scaled_oh, X_test_cat_encoded_oh))
# Train model
reg_oh = LinearRegression()
reg_oh.fit(X_train_final_oh, y_train_log)
# Predict and inverse log
y_pred_log_oh = reg_oh.predict(X_test_final_oh)
y_pred_oh = np.expm1(y_pred_log_oh)
# Evaluate
mae_oh = mean_absolute_error(y_test, y_pred_oh)
mse_oh = mean_squared_error(y_test, y_pred_oh)
rmse_oh = np.sqrt(mse_oh)
r2_oh = r2_score(y_test, y_pred_oh)
print("\n=== Linear Regression with One-Hot Encoding (Log-Transformed Target) ===")
print(f"MAE : {mae_oh:.2f}")
print(f"MSE : {mse_oh:.2f}")
print(f"RMSE : {rmse_oh:.2f}")
print(f"R² : {r2_oh:.4f}")
[cell 25 markdown]
### **Log transformation of the target variable further improves the results for one-hot encoding but shows no improvement for integer encoding**
[cell 26 markdown]
# **Part B- Naive Bayes**
[cell 27 markdown]
## **About dataset**
[cell 28 markdown]
# Adult Census Income Dataset Overview
The **Adult Census Income** dataset (also known as the "Census Income" or "Adult" dataset) contains demographic data from the 1994 U.S. Census. The goal is to predict whether a person earns **more than \$50K per year** based on various personal attributes.
## Source:
Originally hosted by the UCI Machine Learning Repository.
* [UCI Dataset Page](https://archive.ics.uci.edu/ml/datasets/adult)
* [Direct CSV Download](https://archive.ics.uci.edu/ml/machine-learning-databases/adult/adult.data)
---
## Target Variable:
* **income** *(binary)*: Whether the person's income is `<=50K` or `>50K`
---
### Features:
| Column Name | Description | Data Type |
| -------------- | ------------------------------------------------------- | ------------------------ |
| age | Age of the person | *Numeric (int)* |
| workclass | Type of employer (e.g., `Private`, `Self-emp`, etc.) | *Categorical* |
| fnlwgt | Final sample weight (used for population weighting) | *Numeric (int)* |
| education | Highest level of education achieved | *Categorical* |
| education-num | Education level as a number (e.g., `Bachelors` = 13) | *Numeric (int)* |
| marital-status | Marital status (e.g., `Never-married`, `Married`) | *Categorical* |
| occupation | Type of job (e.g., `Tech-support`, `Craft-repair`) | *Categorical* |
| relationship | Relationship within household (e.g., `Husband`, `Wife`) | *Categorical* |
| race | Race (e.g., `White`, `Black`, `Asian-Pac-Islander`) | *Categorical* |
| sex | Gender (`Male`, `Female`) | *Categorical* |
| capital-gain | Capital gains from investments | *Numeric (int)* |
| capital-loss | Capital losses from investments | *Numeric (int)* |
| hours-per-week | Hours worked per week | *Numeric (int)* |
| native-country | Country of origin (e.g., `United-States`, `Mexico`) | *Categorical* |
| income | Income class (`<=50K`, `>50K`) | **Categorical (target)** |
[cell 29 markdown]
**`fnlwgt` indicates how many people a record represents in the population (survey weight). It doesn't directly help predict individual income.**
[cell 30 markdown]
# **Loading the dataset into a datafarme**
[cell 31 code]
import pandas as pd
column_names = [
'age', 'workclass', 'fnlwgt', 'education', 'education-num',
'marital-status', 'occupation', 'relationship', 'race', 'sex',
'capital-gain', 'capital-loss', 'hours-per-week', 'native-country', 'income'
]
df = pd.read_csv('adult.data', header=None, names=column_names, na_values=['?', ' ?'], skipinitialspace=True)
df
[cell 32 markdown]
# **Column names with their data typs**
[cell 33 code]
df.info()
[cell 34 markdown]
# **Checking for missing values per column**
[cell 35 code]
df.isna().sum()
[cell 36 markdown]
# **Dropping rows with missing values**
[cell 37 code]
# Calculate % of missing values per column
missing_percent = df.isnull().mean() * 100
missing_cols = missing_percent[missing_percent > 0]
# Show only columns with missing data
print("Missing value percentage:\n")
print(missing_cols.round(2))
# Drop rows with any missing values
df_clean = df.dropna()
print(f"\nOriginal rows: {len(df)}")
print(f"Rows after dropping missing: {len(df_clean)}")
[cell 38 code]
df_clean.info()
[cell 39 code]
df_clean.isna().sum()
[cell 40 markdown]
# **Displaying categorical and numerical columns in the dataset**
[cell 41 code]
categorical_cols = df_clean.select_dtypes(include=['object']).columns.tolist()
numerical_cols = df_clean.select_dtypes(include=['int64', 'float64']).columns.tolist()
print("Categorical Columns:")
print(categorical_cols, len(categorical_cols))
print("\nNumerical Columns:")
print(numerical_cols, len(numerical_cols))
[cell 42 markdown]
# **Count and List of Unique Values in Each Categorical Feature**
[cell 43 code]
for col in categorical_cols:
unique_vals = df_clean[col].unique()
print(f"{col} ({len(unique_vals)} unique): {unique_vals}")
print('------------------------------------------------')
[cell 44 markdown]
# **Count plot for target variable shows label imbalance**
[cell 45 code]
print(df_clean['income'].value_counts(normalize=True))
df_clean['income'].value_counts(normalize=True).mul(100).plot(kind='bar')
[cell 46 markdown]
# **Multinomial Naive Bayes** : Converting all features into categorical form
[cell 47 markdown]
* Multinomial Naive Bayes requires all input features to be **categorical** or **discrete**.
* Since our dataset includes several **continuous numerical features**, we first **discretize** these numeric columns using appropriate binning techniques.
* The result is stored in a **df_discretized**, where all features are now categorical
[cell 48 code]
import warnings
warnings.filterwarnings("ignore")
categorical_cols = df_clean.select_dtypes(include=['object']).columns.tolist()
numerical_cols = df_clean.select_dtypes(include=['int64', 'float64']).columns.tolist()
print("Categorical Features:", categorical_cols)
print("Numerical Features:", numerical_cols)
from sklearn.preprocessing import KBinsDiscretizer
# Discretize numerical features
discretizer = KBinsDiscretizer(n_bins=5, encode='ordinal', strategy='quantile')
num_discretized = discretizer.fit_transform(df_clean[numerical_cols])
# Create DataFrame for discretized numeric features
df_discretized_num = pd.DataFrame(num_discretized,
columns=[f"{col}_bin" for col in numerical_cols],
index=df_clean.index)
# Drop original numeric features and join discretized ones
df_discretized = df_clean.drop(columns=numerical_cols)
df_discretized = pd.concat([df_discretized, df_discretized_num], axis=1)
print("\n Discretized numeric features stored in df_discretized")
df_discretized
[cell 49 markdown]
# **Integer encoding for all features**
[cell 50 code]
from sklearn.preprocessing import OrdinalEncoder
# Copy from df_discretized
df_discretized_ordinal = df_discretized.copy()
# Identify which columns are still object (i.e., the original categoricals)
object_cols = df_discretized_ordinal.select_dtypes(include='object').columns.tolist()
# Ordinal encode only those object columns
encoder = OrdinalEncoder(handle_unknown='use_encoded_value', unknown_value=-1)
df_discretized_ordinal[object_cols] = encoder.fit_transform(df_discretized_ordinal[object_cols])
df_discretized_ordinal
[cell 51 markdown]
# **Training Multinomial Naive Bayes with integer encoded features and Model Evaluation**
[cell 52 code]
from sklearn.model_selection import train_test_split
from sklearn.naive_bayes import MultinomialNB
from sklearn.metrics import (
accuracy_score, precision_score, recall_score,
f1_score, roc_auc_score, classification_report
)
import numpy as np
X = df_discretized_ordinal.drop(columns=["income"]).copy()
y = df_discretized_ordinal["income"].astype(int)
# Split the data (80% train, 20% test)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, shuffle=True, stratify=y, random_state=42
)
# Train model
model = MultinomialNB()
model.fit(X_train, y_train)
# Predictions
y_train_pred = model.predict(X_train)
y_test_pred = model.predict(X_test)
# Probabilities for AUC
y_train_prob = model.predict_proba(X_train)[:, 1]
y_test_prob = model.predict_proba(X_test)[:, 1]
# Metrics
train_acc = accuracy_score(y_train, y_train_pred)
test_acc = accuracy_score(y_test, y_test_pred)
train_prec = precision_score(y_train, y_train_pred, zero_division=0)
test_prec = precision_score(y_test, y_test_pred, zero_division=0)
train_rec = recall_score(y_train, y_train_pred, zero_division=0)
test_rec = recall_score(y_test, y_test_pred, zero_division=0)
train_f1 = f1_score(y_train, y_train_pred, zero_division=0)
test_f1 = f1_score(y_test, y_test_pred, zero_division=0)
train_auc = roc_auc_score(y_train, y_train_prob)
test_auc = roc_auc_score(y_test, y_test_prob)
# Print results
print("\nMultinomialNB using 80/20 Split:")
print("-" * 40)
print(f"Train Accuracy: {train_acc:.4f}")
print(f"Train Precision: {train_prec:.4f}")
print(f"Train Recall: {train_rec:.4f}")
print(f"Train F1 Score: {train_f1:.4f}")
print(f"Train AUC: {train_auc:.4f}")
print("-" * 40)
print(f"Test Accuracy: {test_acc:.4f}")
print(f"Test Precision: {test_prec:.4f}")
print(f"Test Recall: {test_rec:.4f}")
print(f"Test F1 Score: {test_f1:.4f}")
print(f"Test AUC: {test_auc:.4f}")
print("-" * 40)
# Classification report
print("\nClassification Report (Test):")
print(classification_report(y_test, y_test_pred, target_names=['<=50K', '>50K']))
[cell 53 markdown]
# **One Hot Encoding for all features**
[cell 54 code]
from sklearn.preprocessing import OneHotEncoder
import pandas as pd
# Separate features and target
X_cat = df_discretized.drop(columns=['income']).copy()
y_ohe = df_discretized['income'].apply(lambda x: 1 if x.strip() == '>50K' else 0)
# One-hot encode only the feature columns
ohe = OneHotEncoder(handle_unknown='ignore', sparse_output=False)
X_onehot = ohe.fit_transform(X_cat)
# Create final DataFrame with one-hot encoded features
df_discretized_onehot = pd.DataFrame(
X_onehot,
columns=ohe.get_feature_names_out(X_cat.columns),
index=X_cat.index
)
df_discretized_onehot
[cell 55 markdown]
# **Training Multinomial Naive Bayes with One Hot Enoded features and Model Evaluation**
[cell 56 code]
from sklearn.model_selection import train_test_split
from sklearn.naive_bayes import MultinomialNB
from sklearn.metrics import (
accuracy_score, precision_score, recall_score,
f1_score, roc_auc_score, classification_report
)
import numpy as np
X = df_discretized_onehot
# Split the data (80% train, 20% test)
X_train, X_test, y_train, y_test = train_test_split(
X, y_ohe, test_size=0.2, shuffle=True, stratify=y_ohe,random_state=42
)
# Train model
model = MultinomialNB()
model.fit(X_train, y_train)
# Predictions
y_train_pred = model.predict(X_train)
y_test_pred = model.predict(X_test)
# Probabilities for AUC
y_train_prob = model.predict_proba(X_train)[:, 1]
y_test_prob = model.predict_proba(X_test)[:, 1]
# Metrics
train_acc = accuracy_score(y_train, y_train_pred)
test_acc = accuracy_score(y_test, y_test_pred)
train_prec = precision_score(y_train, y_train_pred, zero_division=0)
test_prec = precision_score(y_test, y_test_pred, zero_division=0)
train_rec = recall_score(y_train, y_train_pred, zero_division=0)
test_rec = recall_score(y_test, y_test_pred, zero_division=0)
train_f1 = f1_score(y_train, y_train_pred, zero_division=0)
test_f1 = f1_score(y_test, y_test_pred, zero_division=0)
train_auc = roc_auc_score(y_train, y_train_prob)
test_auc = roc_auc_score(y_test, y_test_prob)
# Print results
print("\nMultinomialNB using 80/20 Split:")
print("-" * 40)
print(f"Train Accuracy: {train_acc:.4f}")
print(f"Train Precision: {train_prec:.4f}")
print(f"Train Recall: {train_rec:.4f}")
print(f"Train F1 Score: {train_f1:.4f}")
print(f"Train AUC: {train_auc:.4f}")
print("-" * 40)
print(f"Test Accuracy: {test_acc:.4f}")
print(f"Test Precision: {test_prec:.4f}")
print(f"Test Recall: {test_rec:.4f}")
print(f"Test F1 Score: {test_f1:.4f}")
print(f"Test AUC: {test_auc:.4f}")
print("-" * 40)
# Classification report
print("\nClassification Report (Test):")
print(classification_report(y_test, y_test_pred, target_names=['<=50K', '>50K']))
[cell 57 markdown]
# **Gaussian Naive Bayes**
[cell 58 markdown]
* Gaussian Naive Bayes is designed for continuous numeric features, assuming each follows a Gaussian (normal) distribution within each class.
* Therefore, we keep all numeric features as they are (no discretization or binning).
* Since our dataset also contains categorical features, we convert them into numeric form using encoding techniques (Integer or One-Hot Encoding), making them compatible with GaussianNB.
[cell 59 markdown]
## **Integer Encoding of categorical features**
[cell 60 code]
import warnings
warnings.filterwarnings("ignore")
from sklearn.preprocessing import OrdinalEncoder
import pandas as pd
# Identify target column
target_col = "income"
# Separate categorical and numerical columns (excluding target)
categorical_cols = df_clean.select_dtypes(include=['object']).columns.tolist()
if target_col in categorical_cols:
categorical_cols.remove(target_col)
numerical_cols = df_clean.select_dtypes(include=['int64', 'float64']).columns.tolist()
if target_col in numerical_cols:
numerical_cols.remove(target_col)
print("Categorical Features:", categorical_cols)
print("Numerical Features:", numerical_cols)
# Ordinal encode categorical feature columns only
encoder = OrdinalEncoder()
cat_encoded = encoder.fit_transform(df_clean[categorical_cols])
# Create DataFrame for encoded categorical features
df_encoded_cat = pd.DataFrame(
cat_encoded,
columns=[f"{col}_ord" for col in categorical_cols],
index=df_clean.index
)
# Combine encoded categorical + original numeric features + target
df_encoded = pd.concat(
[df_encoded_cat, df_clean[numerical_cols], df_clean[[target_col]]],
axis=1
)
print("\n Integer-encoded categorical + numeric features stored in df_encoded")
df_encoded.head()
[cell 61 markdown]
# **Training Gaussian NB with integer encoding for categorical features**
[cell 62 code]
from sklearn.model_selection import train_test_split
from sklearn.naive_bayes import GaussianNB
from sklearn.metrics import (
accuracy_score, precision_score, recall_score,
f1_score, roc_auc_score, classification_report
)
import numpy as np
# Prepare features and target
X = df_encoded.drop(columns=["income"]).copy()
y = df_encoded["income"].map({"<=50K": 0, ">50K": 1}).astype(int)
# Split data (80/20 split, stratified)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, shuffle=True, stratify=y, random_state=42)
# from sklearn.preprocessing import StandardScaler
# scaler = StandardScaler()
# X_train[numerical_cols] = scaler.fit_transform(X_train[numerical_cols])
# X_test[numerical_cols] = scaler.transform(X_test[numerical_cols])
# Train GaussianNB
model = GaussianNB()
model.fit(X_train, y_train)
# Predictions
y_train_pred = model.predict(X_train)
y_test_pred = model.predict(X_test)
# Probabilities for AUC
y_train_prob = model.predict_proba(X_train)[:, 1]
y_test_prob = model.predict_proba(X_test)[:, 1]
# Metrics
train_acc = accuracy_score(y_train, y_train_pred)
test_acc = accuracy_score(y_test, y_test_pred)
train_prec = precision_score(y_train, y_train_pred, zero_division=0)
test_prec = precision_score(y_test, y_test_pred, zero_division=0)
train_rec = recall_score(y_train, y_train_pred, zero_division=0)
test_rec = recall_score(y_test, y_test_pred, zero_division=0)
train_f1 = f1_score(y_train, y_train_pred, zero_division=0)
test_f1 = f1_score(y_test, y_test_pred, zero_division=0)
train_auc = roc_auc_score(y_train, y_train_prob)
test_auc = roc_auc_score(y_test, y_test_prob)
# Print results
print("\nGaussianNB using 80/20 Split:")
print("-" * 40)
print(f"Train Accuracy: {train_acc:.4f}")
print(f"Train Precision: {train_prec:.4f}")
print(f"Train Recall: {train_rec:.4f}")
print(f"Train F1 Score: {train_f1:.4f}")
print(f"Train AUC: {train_auc:.4f}")
print("-" * 40)
print(f"Test Accuracy: {test_acc:.4f}")
print(f"Test Precision: {test_prec:.4f}")
print(f"Test Recall: {test_rec:.4f}")
print(f"Test F1 Score: {test_f1:.4f}")
print(f"Test AUC: {test_auc:.4f}")
print("-" * 40)
# Classification report
print("\nClassification Report (Test):")
print(classification_report(y_test, y_test_pred, target_names=['<=50K', '>50K']))
[cell 63 markdown]
## **One hot encoding for categorical features**
[cell 64 code]
import warnings
warnings.filterwarnings("ignore")
from sklearn.preprocessing import OneHotEncoder
import pandas as pd
target_col = "income"
# Separate categorical and numerical columns (excluding target)
categorical_cols = df_clean.select_dtypes(include=['object']).columns.tolist()
if target_col in categorical_cols:
categorical_cols.remove(target_col)
numerical_cols = df_clean.select_dtypes(include=['int64', 'float64']).columns.tolist()
if target_col in numerical_cols:
numerical_cols.remove(target_col)
print("Categorical Features:", categorical_cols)
print("Numerical Features:", numerical_cols)
# One-Hot Encode categorical columns
encoder = OneHotEncoder(handle_unknown='ignore', sparse_output=False)
cat_encoded = encoder.fit_transform(df_clean[categorical_cols])
df_encoded_cat = pd.DataFrame(
cat_encoded,
columns=encoder.get_feature_names_out(categorical_cols),
index=df_clean.index
)
# Combine encoded categorical + original numeric + target
df_encoded_ohe = pd.concat(
[df_encoded_cat, df_clean[numerical_cols], df_clean[[target_col]]],
axis=1
)
print("\n One-hot encoded categorical + original numeric features stored in df_encoded_ohe")
df_encoded_ohe.head()
[cell 65 markdown]
# **Training Gaussian NB with one hot encoding for categorical features**
[cell 66 code]
from sklearn.model_selection import train_test_split
from sklearn.naive_bayes import MultinomialNB,GaussianNB
from sklearn.metrics import (
accuracy_score, precision_score, recall_score,
f1_score, roc_auc_score, classification_report
)
# Features (all columns except the target)
X = df_encoded_ohe.drop(columns=["income"]).copy()
# Target variable (convert to numeric 0/1)
y = df_encoded_ohe["income"].map({"<=50K": 0, ">50K": 1}).astype(int)
# Split the data (80% train, 20% test)
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, shuffle=True, stratify=y,random_state=42
)
# from sklearn.preprocessing import StandardScaler
# scaler = StandardScaler()
# X_train[numerical_cols] = scaler.fit_transform(X_train[numerical_cols])
# X_test[numerical_cols] = scaler.transform(X_test[numerical_cols])
# Train model
model = GaussianNB()
model.fit(X_train, y_train)
# Predictions
y_train_pred = model.predict(X_train)
y_test_pred = model.predict(X_test)
# Probabilities for AUC
y_train_prob = model.predict_proba(X_train)[:, 1]
y_test_prob = model.predict_proba(X_test)[:, 1]
# Metrics
train_acc = accuracy_score(y_train, y_train_pred)
test_acc = accuracy_score(y_test, y_test_pred)
train_prec = precision_score(y_train, y_train_pred, zero_division=0)
test_prec = precision_score(y_test, y_test_pred, zero_division=0)
train_rec = recall_score(y_train, y_train_pred, zero_division=0)
test_rec = recall_score(y_test, y_test_pred, zero_division=0)
train_f1 = f1_score(y_train, y_train_pred, zero_division=0)
test_f1 = f1_score(y_test, y_test_pred, zero_division=0)
train_auc = roc_auc_score(y_train, y_train_prob)
test_auc = roc_auc_score(y_test, y_test_prob)
# Print results
print("\n Gaussian NB using 80/20 Split:")
print("-" * 40)
print(f"Train Accuracy: {train_acc:.4f}")
print(f"Train Precision: {train_prec:.4f}")
print(f"Train Recall: {train_rec:.4f}")
print(f"Train F1 Score: {train_f1:.4f}")
print(f"Train AUC: {train_auc:.4f}")
print("-" * 40)
print(f"Test Accuracy: {test_acc:.4f}")
print(f"Test Precision: {test_prec:.4f}")
print(f"Test Recall: {test_rec:.4f}")
print(f"Test F1 Score: {test_f1:.4f}")
print(f"Test AUC: {test_auc:.4f}")
print("-" * 40)
# Classification report
print("\nClassification Report (Test):")
print(classification_report(y_test, y_test_pred, target_names=['<=50K', '>50K']))
[cell 67 markdown]
# **Final Comparison: Multinomial Naive Bayes vs Gaussian Naive Bayes**
[cell 68 markdown]
| Model | Encoding Type | Test Accuracy | Precision | Recall | F1 Score | AUC |
| :-------------------------- | :--------------- | :------------ | :--------- | :--------- | :--------- | :--------- |
| **Multinomial Naive Bayes** | Ordinal Encoding | 0.7820 | 0.5554 | 0.6245 | 0.5879 | 0.7859 |
| **Multinomial Naive Bayes** | One-Hot Encoding | **0.8076** | 0.5888 | **0.7530** | **0.6608** | **0.8831** |
| **Gaussian Naive Bayes** | Ordinal Encoding | 0.7865 | 0.6569 | 0.2983 | 0.4103 | 0.8289 |
| **Gaussian Naive Bayes** | One-Hot Encoding | 0.7873 | **0.6617** | 0.2983 | 0.4112 | 0.8280 |
[cell 69 markdown]
# **Observations**
1. **Multinomial Naive Bayes with One-Hot Encoding** delivers the best overall performance, achieving the highest recall, F1 score, and AUC, showing that it models categorical relationships more effectively.
2. **Gaussian Naive Bayes** performs similarly in accuracy but struggles with recall, indicating its Gaussian assumption is less suited for primarily categorical or discretized data.