# LinearRegression NaiveBayes

course: Module 2 — Machine Learning Algorithms
module: Module-2-Machine-Learning-Algorithms
type: notebook
source_url: https://personal-learn.armco.dev/files/Module-2-Machine-Learning-Algorithms/General/Lab_Materials/LinearRegression_NaiveBayes.ipynb

---
[cell 1 markdown]
# Part A : Linear Regression

[cell 2 markdown]
# Bike Sharing Dataset Overview

The **Bike Sharing Dataset** contains data collected from a bike rental system over a two-year period (2011–2012). It tracks hourly demand for shared bikes in relation to environmental and seasonal conditions.

The goal is to **predict the number of bike rentals** (`cnt`) based on time, weather, and user behavior variables.

---

## Source:

Available via the UCI Machine Learning Repository:

* [UCI Dataset Page](https://archive.ics.uci.edu/ml/datasets/Bike+Sharing+Dataset)

---

## Target Variable:

* **cnt** *(numeric)*: Total number of bike rentals (casual + registered) in a given hour.

---

## Features:

| Column Name | Description                                                      | Data Type          |
| ----------- | ---------------------------------------------------------------- | ------------------ |
| instant     | Record index (can be dropped)                                    | *Integer*          |
| season      | Season (1: Spring, 2: Summer, 3: Fall, 4: Winter)                | *Categorical*      |
| yr          | Year (0: 2011, 1: 2012)                                          | *Binary*           |
| mnth        | Month (1 to 12)                                                  | *Categorical*      |
| hr          | Hour of the day (0 to 23)                                        | *Categorical*      |
| holiday     | Whether the day is a holiday                                     | *Binary*           |
| weekday     | Day of the week (0: Sunday, 6: Saturday)                         | *Categorical*      |
| workingday  | Whether the day is a working day                                 | *Binary*           |
| weathersit  | Weather situation (1: Clear, 2: Mist, 3: Light Snow/Rain, etc.)  | *Categorical*      |
| temp        | Normalized temperature (0 to 1 scale)                            | *Numeric (float)*  |
| atemp       | Normalized "feels like" temperature                              | *Numeric (float)*  |
| hum         | Normalized humidity (0 to 1 scale)                               | *Numeric (float)*  |
| windspeed   | Normalized wind speed                                            | *Numeric (float)*  |
| casual      | Count of casual users (do **not** use when predicting `cnt`)     | *Numeric*          |
| registered  | Count of registered users (do **not** use when predicting `cnt`) | *Numeric*          |
| cnt         | **Target**: Total number of rentals (casual + registered)        | **Numeric Target** |

---

### Note on Leakage:

The features `casual` and `registered` **should not be used** when predicting `cnt`, as they are **components of the target variable** and would introduce data leakage.

[cell 3 markdown]
# **Loading the dataset into a datafarme**

[cell 4 code]
import pandas as pd


df = pd.read_csv("bike_sharing.csv")

print(df.shape)
df

[cell 5 code]
df.info()

[cell 6 markdown]
# **Checking for missing values**

[cell 7 code]
df.isna().sum()

[cell 8 code]
df_clean = df.drop(columns=["instant", "casual", "registered"])
categorical_cols = ["season", "yr", "mnth", "hr", "holiday",
                    "weekday", "workingday", "weathersit"]

df_clean[categorical_cols] = df_clean[categorical_cols].astype("object")
df_clean.info()

[cell 9 markdown]
# **Displaying unique values of categorical columns**

[cell 10 code]
for col in categorical_cols:
    unique_vals = df[col].unique()
    print(f"{col} ({len(unique_vals)} unique): {unique_vals}")
    print('------------------------------------------------')

[cell 11 markdown]
# **Distribution of target variable**

[cell 12 code]
df_clean["cnt"].plot(kind="hist", bins=40, density=True, alpha=0.6, figsize=(8, 5), title="Histogram of Total Bike Rentals")
#df_clean["cnt"].plot(kind="kde", linewidth=2)

[cell 13 markdown]
* The distribution is highly right-skewed.
* There are many low rental counts and a long tail of high counts.

[cell 14 markdown]
# **Splitting the dataset into train-test sets**

[cell 15 code]
from sklearn.model_selection import train_test_split

# Drop 'ID' column
X = df_clean.drop(['cnt'], axis=1)  # Features
y = df_clean['cnt']                       # Target column

# Perform train-test split (80% train, 20% test)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42,
)

# Display shapes
print(f"X_train shape: {X_train.shape}")
print(f"X_test shape: {X_test.shape}")
print(f"y_train shape: {y_train.shape}")
print(f"y_test shape: {y_test.shape}")

[cell 16 markdown]
# **Linear regression model with categorical variables (a) Integer Encoded (b) One hot Encoded. All numeric features are normalized to have zero mean and standard deviation one.**

[cell 17 code]
from sklearn.preprocessing import StandardScaler, OrdinalEncoder
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
import numpy as np

# Identify numeric and categorical columns
numeric_cols = X.select_dtypes(include=np.number).columns.tolist()
categorical_cols = X.select_dtypes(include='object').columns.tolist()

# Split into numeric and categorical
X_train_num = X_train[numeric_cols]
X_test_num = X_test[numeric_cols]

X_train_cat = X_train[categorical_cols]
X_test_cat = X_test[categorical_cols]

# Scale numeric features
scaler = StandardScaler()
X_train_num_scaled = scaler.fit_transform(X_train_num)
X_test_num_scaled = scaler.transform(X_test_num)

# Encode categorical features
encoder = OrdinalEncoder()
X_train_cat_encoded = encoder.fit_transform(X_train_cat)
X_test_cat_encoded = encoder.transform(X_test_cat)

# Combine numeric and categorical features
X_train_final = np.hstack((X_train_num_scaled, X_train_cat_encoded))
X_test_final = np.hstack((X_test_num_scaled, X_test_cat_encoded))

# Train linear regression
reg = LinearRegression()
reg.fit(X_train_final, y_train)

# Predict
y_pred = reg.predict(X_test_final)

# Regression Metrics
mae = mean_absolute_error(y_test, y_pred)
mse = mean_squared_error(y_test, y_pred)
rmse = np.sqrt(mse)
r2 = r2_score(y_test, y_pred)

# Print metrics
# Print results
print("\n=== Linear Regression with Integer Encoding ===")

print(f"MAE   : {mae:.2f}")
print(f"MSE   : {mse:.2f}")
print(f"RMSE  : {rmse:.2f}")
print(f"R²    : {r2:.4f}")






from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
import numpy as np

# Split features
X_train_num_oh = X_train[numeric_cols]
X_test_num_oh = X_test[numeric_cols]
X_train_cat_oh = X_train[categorical_cols]
X_test_cat_oh = X_test[categorical_cols]

# Scale numeric features
scaler_oh = StandardScaler()
X_train_num_scaled_oh = scaler_oh.fit_transform(X_train_num_oh)
X_test_num_scaled_oh = scaler_oh.transform(X_test_num_oh)

# One-hot encode categorical features # drop='first'
ohe = OneHotEncoder(handle_unknown='ignore',  sparse_output=False)
X_train_cat_encoded_oh = ohe.fit_transform(X_train_cat_oh)
X_test_cat_encoded_oh = ohe.transform(X_test_cat_oh)

# Combine numeric and categorical features
X_train_final_oh = np.hstack((X_train_num_scaled_oh, X_train_cat_encoded_oh))
X_test_final_oh = np.hstack((X_test_num_scaled_oh, X_test_cat_encoded_oh))

# Train linear regression
reg_oh = LinearRegression()
reg_oh.fit(X_train_final_oh, y_train)

# Predict
y_pred_oh = reg_oh.predict(X_test_final_oh)

# Evaluate
mae_oh = mean_absolute_error(y_test, y_pred_oh)
mse_oh = mean_squared_error(y_test, y_pred_oh)
rmse_oh = np.sqrt(mse_oh)
r2_oh = r2_score(y_test, y_pred_oh)

# Print results
print("\n=== Linear Regression with One-Hot Encoding ===")
print(f"MAE   : {mae_oh:.2f}")
print(f"MSE   : {mse_oh:.2f}")
print(f"RMSE  : {rmse_oh:.2f}")
print(f"R²    : {r2_oh:.4f}")

[cell 18 markdown]
# **Showing actual values of target variable `cnt` (Actual_cnt) and its predicted values (Predicted_cnt) when using Linear Regression with Integer encoding for some samples**

[cell 19 code]
df_results_ordinal = X_test.copy()
df_results_ordinal['Actual_cnt'] = y_test.values
df_results_ordinal['Predicted_cnt'] = y_pred


df_sample_ordinal = df_results_ordinal.sample(10, random_state=42)

df_sample_ordinal

[cell 20 markdown]
# **Showing actual values of target variable `cnt` (Actual_cnt) and its predicted values (Predicted_cnt) when using Linear Regression with One hot encoding for some samples**

[cell 21 code]
df_results_onehot = X_test.copy()
df_results_onehot['Actual_cnt'] = y_test.values
df_results_onehot['Predicted_cnt'] = y_pred_oh
df_sample_onehot = df_results_onehot.sample(10, random_state=42)
df_sample_onehot

[cell 22 markdown]
### **The deviation of predicted values from the actual values is smaller with one-hot encoding, while it is larger with integer encoding.**

[cell 23 markdown]
# **Linear regression is performed with the target variable log-transformed and categorical variables encoded using (a) integer encoding and (b) one-hot encoding. All numeric features are normalized to have zero mean and a standard deviation of one**

[cell 24 code]
from sklearn.preprocessing import StandardScaler, OrdinalEncoder, OneHotEncoder
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
import numpy as np

# --- Identify numeric and categorical columns ---
numeric_cols = X.select_dtypes(include=np.number).columns.tolist()
categorical_cols = X.select_dtypes(include='object').columns.tolist()

# --- Log-transform the target variable ---
y_train_log = np.log1p(y_train)
y_test_log = np.log1p(y_test)

# ==============================
# ORDINAL ENCODING
# ==============================
# Prepare features
X_train_num = X_train[numeric_cols]
X_test_num = X_test[numeric_cols]
X_train_cat = X_train[categorical_cols]
X_test_cat = X_test[categorical_cols]

# Scale numeric
scaler = StandardScaler()
X_train_num_scaled = scaler.fit_transform(X_train_num)
X_test_num_scaled = scaler.transform(X_test_num)

# Ordinal encode categorical
encoder = OrdinalEncoder()
X_train_cat_encoded = encoder.fit_transform(X_train_cat)
X_test_cat_encoded = encoder.transform(X_test_cat)

# Combine features
X_train_final = np.hstack((X_train_num_scaled, X_train_cat_encoded))
X_test_final = np.hstack((X_test_num_scaled, X_test_cat_encoded))

# Train model
reg = LinearRegression()
reg.fit(X_train_final, y_train_log)

# Predict and inverse log
y_pred_log = reg.predict(X_test_final)
y_pred = np.expm1(y_pred_log)

# Evaluate
mae = mean_absolute_error(y_test, y_pred)
mse = mean_squared_error(y_test, y_pred)
rmse = np.sqrt(mse)
r2 = r2_score(y_test, y_pred)

print("\n=== Linear Regression with Integer Encoding (Log-Transformed Target) ===")
print(f"MAE   : {mae:.2f}")
print(f"MSE   : {mse:.2f}")
print(f"RMSE  : {rmse:.2f}")
print(f"R²    : {r2:.4f}")

# ==============================
# ONE-HOT ENCODING
# ==============================
# Prepare features
X_train_num_oh = X_train[numeric_cols]
X_test_num_oh = X_test[numeric_cols]
X_train_cat_oh = X_train[categorical_cols]
X_test_cat_oh = X_test[categorical_cols]

# Scale numeric
scaler_oh = StandardScaler()
X_train_num_scaled_oh = scaler_oh.fit_transform(X_train_num_oh)
X_test_num_scaled_oh = scaler_oh.transform(X_test_num_oh)

# One-hot encode categorical features
ohe = OneHotEncoder(handle_unknown='ignore', sparse_output=False)
X_train_cat_encoded_oh = ohe.fit_transform(X_train_cat_oh)
X_test_cat_encoded_oh = ohe.transform(X_test_cat_oh)

# Combine features
X_train_final_oh = np.hstack((X_train_num_scaled_oh, X_train_cat_encoded_oh))
X_test_final_oh = np.hstack((X_test_num_scaled_oh, X_test_cat_encoded_oh))

# Train model
reg_oh = LinearRegression()
reg_oh.fit(X_train_final_oh, y_train_log)

# Predict and inverse log
y_pred_log_oh = reg_oh.predict(X_test_final_oh)
y_pred_oh = np.expm1(y_pred_log_oh)

# Evaluate
mae_oh = mean_absolute_error(y_test, y_pred_oh)
mse_oh = mean_squared_error(y_test, y_pred_oh)
rmse_oh = np.sqrt(mse_oh)
r2_oh = r2_score(y_test, y_pred_oh)

print("\n=== Linear Regression with One-Hot Encoding (Log-Transformed Target) ===")
print(f"MAE   : {mae_oh:.2f}")
print(f"MSE   : {mse_oh:.2f}")
print(f"RMSE  : {rmse_oh:.2f}")
print(f"R²    : {r2_oh:.4f}")

[cell 25 markdown]
### **Log transformation of the target variable further improves the results for one-hot encoding but shows no improvement for integer encoding**

[cell 26 markdown]
# **Part B- Naive Bayes**

[cell 27 markdown]
## **About dataset**

[cell 28 markdown]
# Adult Census Income Dataset Overview

The **Adult Census Income** dataset (also known as the "Census Income" or "Adult" dataset) contains demographic data from the 1994 U.S. Census. The goal is to predict whether a person earns **more than \$50K per year** based on various personal attributes.

## Source:

Originally hosted by the UCI Machine Learning Repository.

* [UCI Dataset Page](https://archive.ics.uci.edu/ml/datasets/adult)
* [Direct CSV Download](https://archive.ics.uci.edu/ml/machine-learning-databases/adult/adult.data)

---

## Target Variable:

* **income** *(binary)*: Whether the person's income is `<=50K` or `>50K`

---

### Features:

| Column Name    | Description                                             | Data Type                |
| -------------- | ------------------------------------------------------- | ------------------------ |
| age            | Age of the person                                       | *Numeric (int)*          |
| workclass      | Type of employer (e.g., `Private`, `Self-emp`, etc.)    | *Categorical*            |
| fnlwgt         | Final sample weight (used for population weighting)     | *Numeric (int)*          |
| education      | Highest level of education achieved                     | *Categorical*            |
| education-num  | Education level as a number (e.g., `Bachelors` = 13)    | *Numeric (int)*          |
| marital-status | Marital status (e.g., `Never-married`, `Married`)       | *Categorical*            |
| occupation     | Type of job (e.g., `Tech-support`, `Craft-repair`)      | *Categorical*            |
| relationship   | Relationship within household (e.g., `Husband`, `Wife`) | *Categorical*            |
| race           | Race (e.g., `White`, `Black`, `Asian-Pac-Islander`)     | *Categorical*            |
| sex            | Gender (`Male`, `Female`)                               | *Categorical*            |
| capital-gain   | Capital gains from investments                          | *Numeric (int)*          |
| capital-loss   | Capital losses from investments                         | *Numeric (int)*          |
| hours-per-week | Hours worked per week                                   | *Numeric (int)*          |
| native-country | Country of origin (e.g., `United-States`, `Mexico`)     | *Categorical*            |
| income         | Income class (`<=50K`, `>50K`)                          | **Categorical (target)** |

[cell 29 markdown]
**`fnlwgt` indicates how many people a record represents in the population (survey weight). It doesn't directly help predict individual income.**

[cell 30 markdown]
# **Loading the dataset into a datafarme**

[cell 31 code]
import pandas as pd

column_names = [
    'age', 'workclass', 'fnlwgt', 'education', 'education-num',
    'marital-status', 'occupation', 'relationship', 'race', 'sex',
    'capital-gain', 'capital-loss', 'hours-per-week', 'native-country', 'income'
]

df = pd.read_csv('adult.data', header=None, names=column_names, na_values=['?', ' ?'], skipinitialspace=True)

df

[cell 32 markdown]
# **Column names with their data typs**

[cell 33 code]
df.info()

[cell 34 markdown]
# **Checking for missing values per column**

[cell 35 code]
df.isna().sum()

[cell 36 markdown]
# **Dropping rows with missing values**

[cell 37 code]
# Calculate % of missing values per column
missing_percent = df.isnull().mean() * 100
missing_cols = missing_percent[missing_percent > 0]

# Show only columns with missing data
print("Missing value percentage:\n")
print(missing_cols.round(2))

# Drop rows with any missing values
df_clean = df.dropna()

print(f"\nOriginal rows: {len(df)}")
print(f"Rows after dropping missing: {len(df_clean)}")

[cell 38 code]
df_clean.info()

[cell 39 code]
df_clean.isna().sum()

[cell 40 markdown]
# **Displaying categorical and numerical columns in the dataset**

[cell 41 code]
categorical_cols = df_clean.select_dtypes(include=['object']).columns.tolist()
numerical_cols = df_clean.select_dtypes(include=['int64', 'float64']).columns.tolist()

print("Categorical Columns:")
print(categorical_cols, len(categorical_cols))
print("\nNumerical Columns:")
print(numerical_cols, len(numerical_cols))

[cell 42 markdown]
# **Count and List of Unique Values in Each Categorical Feature**

[cell 43 code]
for col in categorical_cols:
    unique_vals = df_clean[col].unique()
    print(f"{col} ({len(unique_vals)} unique): {unique_vals}")
    print('------------------------------------------------')

[cell 44 markdown]
# **Count plot for target variable shows label imbalance**

[cell 45 code]
print(df_clean['income'].value_counts(normalize=True))
df_clean['income'].value_counts(normalize=True).mul(100).plot(kind='bar')

[cell 46 markdown]
# **Multinomial Naive Bayes** : Converting all features into categorical form

[cell 47 markdown]
* Multinomial Naive Bayes requires all input features to be **categorical** or **discrete**.
* Since our dataset includes several **continuous numerical features**, we first **discretize** these numeric columns using appropriate binning techniques.
* The result is stored in a **df_discretized**, where all features are now categorical

[cell 48 code]
import warnings
warnings.filterwarnings("ignore")


categorical_cols = df_clean.select_dtypes(include=['object']).columns.tolist()
numerical_cols = df_clean.select_dtypes(include=['int64', 'float64']).columns.tolist()

print("Categorical Features:", categorical_cols)
print("Numerical Features:", numerical_cols)

from sklearn.preprocessing import KBinsDiscretizer

# Discretize numerical features
discretizer = KBinsDiscretizer(n_bins=5, encode='ordinal', strategy='quantile')
num_discretized = discretizer.fit_transform(df_clean[numerical_cols])

# Create DataFrame for discretized numeric features
df_discretized_num = pd.DataFrame(num_discretized,
                                   columns=[f"{col}_bin" for col in numerical_cols],
                                   index=df_clean.index)

# Drop original numeric features and join discretized ones
df_discretized = df_clean.drop(columns=numerical_cols)
df_discretized = pd.concat([df_discretized, df_discretized_num], axis=1)

print("\n Discretized numeric features stored in df_discretized")
df_discretized

[cell 49 markdown]
# **Integer encoding for all features**

[cell 50 code]
from sklearn.preprocessing import OrdinalEncoder

# Copy from df_discretized
df_discretized_ordinal = df_discretized.copy()

# Identify which columns are still object (i.e., the original categoricals)
object_cols = df_discretized_ordinal.select_dtypes(include='object').columns.tolist()

# Ordinal encode only those object columns
encoder = OrdinalEncoder(handle_unknown='use_encoded_value', unknown_value=-1)
df_discretized_ordinal[object_cols] = encoder.fit_transform(df_discretized_ordinal[object_cols])

df_discretized_ordinal

[cell 51 markdown]
# **Training Multinomial Naive Bayes with integer encoded features and Model Evaluation**

[cell 52 code]
from sklearn.model_selection import train_test_split
from sklearn.naive_bayes import MultinomialNB
from sklearn.metrics import (
    accuracy_score, precision_score, recall_score,
    f1_score, roc_auc_score, classification_report
)
import numpy as np




X = df_discretized_ordinal.drop(columns=["income"]).copy()
y = df_discretized_ordinal["income"].astype(int)

# Split the data (80% train, 20% test)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, shuffle=True, stratify=y, random_state=42
)

# Train model
model = MultinomialNB()
model.fit(X_train, y_train)

# Predictions
y_train_pred = model.predict(X_train)
y_test_pred = model.predict(X_test)

# Probabilities for AUC
y_train_prob = model.predict_proba(X_train)[:, 1]
y_test_prob = model.predict_proba(X_test)[:, 1]

# Metrics
train_acc  = accuracy_score(y_train, y_train_pred)
test_acc   = accuracy_score(y_test, y_test_pred)

train_prec = precision_score(y_train, y_train_pred, zero_division=0)
test_prec  = precision_score(y_test, y_test_pred, zero_division=0)

train_rec  = recall_score(y_train, y_train_pred, zero_division=0)
test_rec   = recall_score(y_test, y_test_pred, zero_division=0)

train_f1   = f1_score(y_train, y_train_pred, zero_division=0)
test_f1    = f1_score(y_test, y_test_pred, zero_division=0)

train_auc  = roc_auc_score(y_train, y_train_prob)
test_auc   = roc_auc_score(y_test, y_test_prob)

# Print results
print("\nMultinomialNB using 80/20 Split:")
print("-" * 40)
print(f"Train Accuracy:  {train_acc:.4f}")
print(f"Train Precision: {train_prec:.4f}")
print(f"Train Recall:    {train_rec:.4f}")
print(f"Train F1 Score:  {train_f1:.4f}")
print(f"Train AUC:       {train_auc:.4f}")
print("-" * 40)
print(f"Test Accuracy:   {test_acc:.4f}")
print(f"Test Precision:  {test_prec:.4f}")
print(f"Test Recall:     {test_rec:.4f}")
print(f"Test F1 Score:   {test_f1:.4f}")
print(f"Test AUC:        {test_auc:.4f}")
print("-" * 40)

# Classification report
print("\nClassification Report (Test):")
print(classification_report(y_test, y_test_pred, target_names=['<=50K', '>50K']))

[cell 53 markdown]
# **One Hot Encoding for all features**

[cell 54 code]
from sklearn.preprocessing import OneHotEncoder
import pandas as pd

# Separate features and target
X_cat = df_discretized.drop(columns=['income']).copy()
y_ohe = df_discretized['income'].apply(lambda x: 1 if x.strip() == '>50K' else 0)


# One-hot encode only the feature columns
ohe = OneHotEncoder(handle_unknown='ignore', sparse_output=False)
X_onehot = ohe.fit_transform(X_cat)

# Create final DataFrame with one-hot encoded features
df_discretized_onehot = pd.DataFrame(
    X_onehot,
    columns=ohe.get_feature_names_out(X_cat.columns),
    index=X_cat.index
)


df_discretized_onehot

[cell 55 markdown]
# **Training Multinomial Naive Bayes with One Hot Enoded features and Model Evaluation**

[cell 56 code]
from sklearn.model_selection import train_test_split
from sklearn.naive_bayes import MultinomialNB
from sklearn.metrics import (
    accuracy_score, precision_score, recall_score,
    f1_score, roc_auc_score, classification_report
)
import numpy as np




X = df_discretized_onehot

# Split the data (80% train, 20% test)
X_train, X_test, y_train, y_test = train_test_split(
    X, y_ohe, test_size=0.2, shuffle=True, stratify=y_ohe,random_state=42
)

# Train model
model = MultinomialNB()
model.fit(X_train, y_train)

# Predictions
y_train_pred = model.predict(X_train)
y_test_pred = model.predict(X_test)

# Probabilities for AUC
y_train_prob = model.predict_proba(X_train)[:, 1]
y_test_prob = model.predict_proba(X_test)[:, 1]

# Metrics
train_acc  = accuracy_score(y_train, y_train_pred)
test_acc   = accuracy_score(y_test, y_test_pred)

train_prec = precision_score(y_train, y_train_pred, zero_division=0)
test_prec  = precision_score(y_test, y_test_pred, zero_division=0)

train_rec  = recall_score(y_train, y_train_pred, zero_division=0)
test_rec   = recall_score(y_test, y_test_pred, zero_division=0)

train_f1   = f1_score(y_train, y_train_pred, zero_division=0)
test_f1    = f1_score(y_test, y_test_pred, zero_division=0)

train_auc  = roc_auc_score(y_train, y_train_prob)
test_auc   = roc_auc_score(y_test, y_test_prob)

# Print results
print("\nMultinomialNB using 80/20 Split:")
print("-" * 40)
print(f"Train Accuracy:  {train_acc:.4f}")
print(f"Train Precision: {train_prec:.4f}")
print(f"Train Recall:    {train_rec:.4f}")
print(f"Train F1 Score:  {train_f1:.4f}")
print(f"Train AUC:       {train_auc:.4f}")
print("-" * 40)
print(f"Test Accuracy:   {test_acc:.4f}")
print(f"Test Precision:  {test_prec:.4f}")
print(f"Test Recall:     {test_rec:.4f}")
print(f"Test F1 Score:   {test_f1:.4f}")
print(f"Test AUC:        {test_auc:.4f}")
print("-" * 40)

# Classification report
print("\nClassification Report (Test):")
print(classification_report(y_test, y_test_pred, target_names=['<=50K', '>50K']))

[cell 57 markdown]
# **Gaussian Naive Bayes**

[cell 58 markdown]
* Gaussian Naive Bayes is designed for continuous numeric features, assuming each follows a Gaussian (normal) distribution within each class.

* Therefore, we keep all numeric features as they are (no discretization or binning).

* Since our dataset also contains categorical features, we convert them into numeric form using encoding techniques (Integer or One-Hot Encoding), making them compatible with GaussianNB.

[cell 59 markdown]
## **Integer Encoding of categorical features**

[cell 60 code]
import warnings
warnings.filterwarnings("ignore")

from sklearn.preprocessing import OrdinalEncoder
import pandas as pd

# Identify target column
target_col = "income"

# Separate categorical and numerical columns (excluding target)
categorical_cols = df_clean.select_dtypes(include=['object']).columns.tolist()
if target_col in categorical_cols:
    categorical_cols.remove(target_col)

numerical_cols = df_clean.select_dtypes(include=['int64', 'float64']).columns.tolist()
if target_col in numerical_cols:
    numerical_cols.remove(target_col)

print("Categorical Features:", categorical_cols)
print("Numerical Features:", numerical_cols)

# Ordinal encode categorical feature columns only
encoder = OrdinalEncoder()
cat_encoded = encoder.fit_transform(df_clean[categorical_cols])

# Create DataFrame for encoded categorical features
df_encoded_cat = pd.DataFrame(
    cat_encoded,
    columns=[f"{col}_ord" for col in categorical_cols],
    index=df_clean.index
)

# Combine encoded categorical + original numeric features + target
df_encoded = pd.concat(
    [df_encoded_cat, df_clean[numerical_cols], df_clean[[target_col]]],
    axis=1
)

print("\n Integer-encoded categorical + numeric features stored in df_encoded")
df_encoded.head()

[cell 61 markdown]
# **Training Gaussian NB with integer encoding for categorical features**

[cell 62 code]
from sklearn.model_selection import train_test_split
from sklearn.naive_bayes import GaussianNB
from sklearn.metrics import (
    accuracy_score, precision_score, recall_score,
    f1_score, roc_auc_score, classification_report
)
import numpy as np

# Prepare features and target

X = df_encoded.drop(columns=["income"]).copy()
y = df_encoded["income"].map({"<=50K": 0, ">50K": 1}).astype(int)




# Split data (80/20 split, stratified)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, shuffle=True, stratify=y, random_state=42)


# from sklearn.preprocessing import StandardScaler

# scaler = StandardScaler()
# X_train[numerical_cols] = scaler.fit_transform(X_train[numerical_cols])
# X_test[numerical_cols]  = scaler.transform(X_test[numerical_cols])


# Train GaussianNB
model = GaussianNB()
model.fit(X_train, y_train)

# Predictions
y_train_pred = model.predict(X_train)
y_test_pred = model.predict(X_test)

# Probabilities for AUC
y_train_prob = model.predict_proba(X_train)[:, 1]
y_test_prob = model.predict_proba(X_test)[:, 1]

# Metrics
train_acc  = accuracy_score(y_train, y_train_pred)
test_acc   = accuracy_score(y_test, y_test_pred)

train_prec = precision_score(y_train, y_train_pred, zero_division=0)
test_prec  = precision_score(y_test, y_test_pred, zero_division=0)

train_rec  = recall_score(y_train, y_train_pred, zero_division=0)
test_rec   = recall_score(y_test, y_test_pred, zero_division=0)

train_f1   = f1_score(y_train, y_train_pred, zero_division=0)
test_f1    = f1_score(y_test, y_test_pred, zero_division=0)

train_auc  = roc_auc_score(y_train, y_train_prob)
test_auc   = roc_auc_score(y_test, y_test_prob)

# Print results
print("\nGaussianNB using 80/20 Split:")
print("-" * 40)
print(f"Train Accuracy:  {train_acc:.4f}")
print(f"Train Precision: {train_prec:.4f}")
print(f"Train Recall:    {train_rec:.4f}")
print(f"Train F1 Score:  {train_f1:.4f}")
print(f"Train AUC:       {train_auc:.4f}")
print("-" * 40)
print(f"Test Accuracy:   {test_acc:.4f}")
print(f"Test Precision:  {test_prec:.4f}")
print(f"Test Recall:     {test_rec:.4f}")
print(f"Test F1 Score:   {test_f1:.4f}")
print(f"Test AUC:        {test_auc:.4f}")
print("-" * 40)

# Classification report
print("\nClassification Report (Test):")
print(classification_report(y_test, y_test_pred, target_names=['<=50K', '>50K']))

[cell 63 markdown]
## **One hot encoding for categorical features**

[cell 64 code]
import warnings
warnings.filterwarnings("ignore")

from sklearn.preprocessing import OneHotEncoder
import pandas as pd

target_col = "income"

# Separate categorical and numerical columns (excluding target)
categorical_cols = df_clean.select_dtypes(include=['object']).columns.tolist()
if target_col in categorical_cols:
    categorical_cols.remove(target_col)

numerical_cols = df_clean.select_dtypes(include=['int64', 'float64']).columns.tolist()
if target_col in numerical_cols:
    numerical_cols.remove(target_col)

print("Categorical Features:", categorical_cols)
print("Numerical Features:", numerical_cols)

# One-Hot Encode categorical columns
encoder = OneHotEncoder(handle_unknown='ignore', sparse_output=False)
cat_encoded = encoder.fit_transform(df_clean[categorical_cols])

df_encoded_cat = pd.DataFrame(
    cat_encoded,
    columns=encoder.get_feature_names_out(categorical_cols),
    index=df_clean.index
)

# Combine encoded categorical + original numeric + target
df_encoded_ohe = pd.concat(
    [df_encoded_cat, df_clean[numerical_cols], df_clean[[target_col]]],
    axis=1
)

print("\n One-hot encoded categorical + original numeric features stored in df_encoded_ohe")
df_encoded_ohe.head()

[cell 65 markdown]
# **Training Gaussian NB with one hot encoding for categorical features**

[cell 66 code]
from sklearn.model_selection import train_test_split
from sklearn.naive_bayes import MultinomialNB,GaussianNB
from sklearn.metrics import (
    accuracy_score, precision_score, recall_score,
    f1_score, roc_auc_score, classification_report
)


# Features (all columns except the target)
X = df_encoded_ohe.drop(columns=["income"]).copy()

# Target variable (convert to numeric 0/1)
y = df_encoded_ohe["income"].map({"<=50K": 0, ">50K": 1}).astype(int)



# Split the data (80% train, 20% test)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, shuffle=True, stratify=y,random_state=42
)

# from sklearn.preprocessing import StandardScaler

# scaler = StandardScaler()
# X_train[numerical_cols] = scaler.fit_transform(X_train[numerical_cols])
# X_test[numerical_cols]  = scaler.transform(X_test[numerical_cols])


# Train model
model = GaussianNB()
model.fit(X_train, y_train)

# Predictions
y_train_pred = model.predict(X_train)
y_test_pred = model.predict(X_test)

# Probabilities for AUC
y_train_prob = model.predict_proba(X_train)[:, 1]
y_test_prob = model.predict_proba(X_test)[:, 1]

# Metrics
train_acc  = accuracy_score(y_train, y_train_pred)
test_acc   = accuracy_score(y_test, y_test_pred)

train_prec = precision_score(y_train, y_train_pred, zero_division=0)
test_prec  = precision_score(y_test, y_test_pred, zero_division=0)

train_rec  = recall_score(y_train, y_train_pred, zero_division=0)
test_rec   = recall_score(y_test, y_test_pred, zero_division=0)

train_f1   = f1_score(y_train, y_train_pred, zero_division=0)
test_f1    = f1_score(y_test, y_test_pred, zero_division=0)

train_auc  = roc_auc_score(y_train, y_train_prob)
test_auc   = roc_auc_score(y_test, y_test_prob)

# Print results
print("\n Gaussian NB using 80/20 Split:")
print("-" * 40)
print(f"Train Accuracy:  {train_acc:.4f}")
print(f"Train Precision: {train_prec:.4f}")
print(f"Train Recall:    {train_rec:.4f}")
print(f"Train F1 Score:  {train_f1:.4f}")
print(f"Train AUC:       {train_auc:.4f}")
print("-" * 40)
print(f"Test Accuracy:   {test_acc:.4f}")
print(f"Test Precision:  {test_prec:.4f}")
print(f"Test Recall:     {test_rec:.4f}")
print(f"Test F1 Score:   {test_f1:.4f}")
print(f"Test AUC:        {test_auc:.4f}")
print("-" * 40)

# Classification report
print("\nClassification Report (Test):")
print(classification_report(y_test, y_test_pred, target_names=['<=50K', '>50K']))

[cell 67 markdown]
# **Final Comparison: Multinomial Naive Bayes vs Gaussian Naive Bayes**

[cell 68 markdown]
| Model                       | Encoding Type    | Test Accuracy | Precision  | Recall     | F1 Score   | AUC        |
| :-------------------------- | :--------------- | :------------ | :--------- | :--------- | :--------- | :--------- |
| **Multinomial Naive Bayes** | Ordinal Encoding | 0.7820    | 0.5554     | 0.6245     | 0.5879     | 0.7859     |
| **Multinomial Naive Bayes** | One-Hot Encoding | **0.8076**    | 0.5888     | **0.7530** | **0.6608** | **0.8831** |
| **Gaussian Naive Bayes**    | Ordinal Encoding | 0.7865        | 0.6569     | 0.2983     | 0.4103     | 0.8289     |
| **Gaussian Naive Bayes**    | One-Hot Encoding | 0.7873        | **0.6617** | 0.2983     | 0.4112     | 0.8280     |

[cell 69 markdown]
# **Observations**
1. **Multinomial Naive Bayes with One-Hot Encoding** delivers the best overall performance, achieving the highest recall, F1 score, and AUC, showing that it models categorical relationships more effectively.

2. **Gaussian Naive Bayes** performs similarly in accuracy but struggles with recall, indicating its Gaussian assumption is less suited for primarily categorical or discretized data.