# DT RF vx1a

course: Module 2 — Machine Learning Algorithms
module: Module-2-Machine-Learning-Algorithms
type: notebook
source_url: https://personal-learn.armco.dev/files/Module-2-Machine-Learning-Algorithms/General/5.3_Lab_Materials/DT_RF_vx1a.ipynb

---
[cell 1 markdown]
# Telco Customer Churn Dataset Overview

The **Telco Customer Churn** dataset contains information about customers of a fictional telecom company. The objective is to predict whether a customer will **churn (leave the service)** based on their demographic, service usage, and account details.

## Source:
Originally available as part of IBM sample datasets, the **Telco Customer Churn** dataset is now publicly hosted on Kaggle and GitHub.  
-[Kaggle Dataset Page](https://www.kaggle.com/datasets/blastchar/telco-customer-churn)  
-[GitHub CSV File](https://github.com/treselle-systems/customer_churn_analysis/blob/master/WA_Fn-UseC_-Telco-Customer-Churn.csv)



# Target Variable:
- **Churn** *(binary)*: Whether the customer has churned (`Yes` or `No`)

---

###  Features:

| Column Name         | Description                                               | Data Type    |
|---------------------|-----------------------------------------------------------|--------------|
| customerID          | Unique customer identifier                                | *String*     |
| gender              | Gender (`Male`, `Female`)                                 | *Categorical*|
| SeniorCitizen       | Whether the customer is a senior citizen (`0`, `1`)       | *Categorical*|
| Partner             | Has a partner (`Yes`, `No`)                               | *Categorical*|
| Dependents          | Has dependents (`Yes`, `No`)                              | *Categorical*|
| tenure              | Number of months the customer has stayed                  | *Numeric (int)*|
| PhoneService        | Has phone service (`Yes`, `No`)                           | *Categorical*|
| MultipleLines       | Has multiple lines (`Yes`, `No`, `No phone service`)      | *Categorical*|
| InternetService     | Type of internet service (`DSL`, `Fiber optic`, `No`)     | *Categorical*|
| OnlineSecurity      | Has online security (`Yes`, `No`, `No internet service`)  | *Categorical*|
| OnlineBackup        | Has online backup (`Yes`, `No`, `No internet service`)    | *Categorical*|
| DeviceProtection    | Has device protection (`Yes`, `No`, `No internet service`)| *Categorical*|
| TechSupport         | Has technical support (`Yes`, `No`, `No internet service`)| *Categorical*|
| StreamingTV         | Streams TV (`Yes`, `No`, `No internet service`)           | *Categorical*|
| StreamingMovies     | Streams movies (`Yes`, `No`, `No internet service`)       | *Categorical*|
| Contract            | Type of contract (`Month-to-month`, `One year`, `Two year`)| *Categorical*|
| PaperlessBilling    | Uses paperless billing (`Yes`, `No`)                      | *Categorical*|
| PaymentMethod       | Payment method (`Electronic check`,`Mailed Check`,`bank Transfer`,`Credit Card`)                 | *Categorical*|
| MonthlyCharges      | The amount charged per month                              | *Numeric (float),*|
| TotalCharges        | Total amount charged                                      | *Numeric (float, needs cleaning)*|
| Churn               | Whether the customer left                                 | **Categorical (target)**|

---

[cell 2 markdown]
# **Python Packages Used**

### 1. `pandas`
- **Purpose**: Loading, cleaning, and manipulating the dataset
- **Import**: `import pandas as pd`
- **Install**: `pip install pandas`

---

### 2. `matplotlib`
- **Purpose**: Visualizing, plotting graphs  and decision trees
- **Import**: `import matplotlib.pyplot as plt`
- **Install**: `pip install matplotlib`

---

### 3. `scikit-learn` (`sklearn`)
- **Purpose**: Data preprocessing, model building, evaluation
- **Install**: `pip install scikit-learn`

#### **Modules used:**
- `sklearn.preprocessing`
  - `KBinsDiscretizer` – discretizing numeric features
- `sklearn.model_selection`
  - `train_test_split` – splitting dataset into train/test sets
- `sklearn.tree`
  - `DecisionTreeClassifier` – training decision trees
  - `plot_tree` – visualizing trained decision trees
- `sklearn.ensemble`
  - `RandomForestClassifier` – training random forest models
- `sklearn.metrics`
  - `accuracy_score` – evaluating accuracy
  - `precision_score` – evaluating precision
  - `recall_score` – evaluating recall
  - `f1_score` – evaluating F1 score
  - `roc_auc_score` – computing AUC for classification
  - `classification_report` – detailed per-class evaluation

---

### 4. `numpy`
- **Purpose**: Numerical operations in plotting and evaluation
- **Import**: `import numpy as np`
- **Install**: `pip install numpy`

[cell 3 code]
# Install packages at once
# !pip install pandas matplotlib scikit-learn numpy

[cell 4 markdown]
# **Loading the dataset into a datafarme**

[cell 5 code]
import pandas as pd

# Load the Telco Customer Churn dataset
df = pd.read_csv('WA_Fn-UseC_-Telco-Customer-Churnvx1.csv')

df

[cell 6 markdown]
# **Column names with their data typs**

[cell 7 code]
df.info()

[cell 8 markdown]
### **Observation 1**
* The column `TotalCharges` appears as `object` in `df.info()` instead of numeric. This is due to the presence of rows with blank string entries (`""`), which pandas does not treat as missing by default.

[cell 9 code]
print(f'No of samples where TotalCharges is Blank:{df[df["TotalCharges"].str.strip() == ""].shape[0]}')
# Show rows where TotalCharges is blank
df[df["TotalCharges"].str.strip() == ""][["customerID", "tenure", "TotalCharges"]]

[cell 10 code]
# convert the column to `float64` and removes invalid rows.
df["TotalCharges"] = pd.to_numeric(df["TotalCharges"], errors="coerce")
df.dropna(subset=["TotalCharges"], inplace=True)
df.shape

[cell 11 markdown]
### **Observation 2**
* `SeniorCitizen` is stored as `int64`, it is a binary categorical feature (0 = No, 1 = Yes). To reflect its actual meaning, we convert it to categorical type

[cell 12 code]
df["SeniorCitizen"] = df["SeniorCitizen"].astype("object")

[cell 13 code]
categorical_cols = df.select_dtypes(include=['object']).columns.tolist()
numerical_cols = df.select_dtypes(include=['int64', 'float64']).columns.tolist()

print("Categorical Columns:")
print(categorical_cols, len(categorical_cols))
print("\nNumerical Columns:")
print(numerical_cols, len(numerical_cols))

[cell 14 markdown]
# **Merging "No internet/phone service" with "No"**

* In this dataset, several categorical columns contain values like `"No internet service"` or `"No phone service"` in addition to a regular `"No"`.

* These extra values do not carry distinct meaning from `"No"` in terms of behavior — they simply indicate that the person does not have the service, which already implies `"No"` for all related options.

[cell 15 code]
cols_to_check = [
    "MultipleLines", "OnlineSecurity", "OnlineBackup",
    "DeviceProtection", "TechSupport", "StreamingTV", "StreamingMovies"
]

for col in cols_to_check:
    print(f"\nValue counts for '{col}':")
    print(df[col].value_counts())

[cell 16 markdown]
To simplify the dataset and avoid redundant categories during encoding, we merged:
- `"No internet service"` → `"No"`
- `"No phone service"` → `"No"`

[cell 17 code]
# Merge "No internet service" / "No phone service" with "No"
internet_cols = [
    "OnlineSecurity", "OnlineBackup", "DeviceProtection",
    "TechSupport", "StreamingTV", "StreamingMovies"
]
for col in internet_cols:
    df[col] = df[col].replace("No internet service", "No")

df["MultipleLines"] = df["MultipleLines"].replace("No phone service", "No")

# Step 3: Print value counts after merging
print("\nAfter merging:\n")
for col in cols_to_check:
    print(f"\nValue counts for '{col}':")
    print(df[col].value_counts())

[cell 18 code]
df.info()

[cell 19 markdown]
# **Heatmap for categorical attributes**
* We used Cramér’s V to measure the strength of association between categorical features and visualized them using a heatmap.
* Cramér’s V is a recommended metric for measuring association between categorical attributes, especially when variables are nominal and do not have a natural order.
* Cramér’s V ranges from 0 to 1, where 0 indicates no association and 1 indicates a strong association between categorical variables.

[cell 20 code]
import pandas as pd
import numpy as np
import seaborn as sns
import matplotlib.pyplot as plt
from scipy.stats import chi2_contingency
import warnings


def cramers_v_basic(x, y):
    confusion_matrix = pd.crosstab(x, y)
    if confusion_matrix.shape[0] < 2 or confusion_matrix.shape[1] < 2:
        return np.nan  # not enough variability to compute
    chi2 = chi2_contingency(confusion_matrix)[0]
    n = confusion_matrix.sum().sum()
    return np.sqrt(chi2 / (n * (min(confusion_matrix.shape) - 1 + 1e-10)))

cols = [col for col in categorical_cols if col != 'Churn'] + ['Churn']

cramer_matrix = pd.DataFrame(np.zeros((len(cols), len(cols))), index=cols, columns=cols)

# Suppress warnings during computation
with warnings.catch_warnings():
    warnings.simplefilter("ignore")
    for col1 in cols:
        for col2 in cols:
            if col1 == col2:
                cramer_matrix.loc[col1, col2] = 1.0
            else:
                cramer_matrix.loc[col1, col2] = cramers_v_basic(df[col1], df[col2])

# Plot the heatmap
plt.figure(figsize=(12, 10))
sns.heatmap(cramer_matrix, cmap="YlGnBu", annot=True, fmt=".2f")
plt.title("Cramér’s V Heatmap of Categorical Variables (Warnings Suppressed)")
plt.tight_layout()
plt.show()

[cell 21 markdown]
## **Observations**
* **Contract type** has one of the strongest associations with Churn (~0.41), meaning contract duration is an important driver of customer churn.

* **InternetService** and related services (OnlineSecurity, TechSupport, StreamingTV/Movies) show moderate inter-correlations (0.3–0.5), indicating these features are interconnected.

* Most other categorical variables have very low correlations (less than 0.1), suggesting minimal dependency and low redundancy among them.

[cell 22 markdown]
# **Pearson Correlation Heatmap: Numerical Features vs. Churn**

[cell 23 code]
import seaborn as sns
import matplotlib.pyplot as plt

df["Churn"] = df["Churn"].map({"Yes": 1, "No": 0}) if df["Churn"].dtype == object else df["Churn"]


cols = numerical_cols + ["Churn"]
corr_matrix = df[cols].corr()



plt.figure(figsize=(8, 6))
sns.heatmap(corr_matrix, annot=True, cmap="coolwarm", fmt=".2f")
plt.title("Correlation Heatmap matrix Numerical Features and Churn")
plt.tight_layout()
plt.show()

[cell 24 markdown]
* **Tenure** has the strongest relationship with churn (≈ −0.35), indicating that customers with shorter tenure are more likely to churn.

[cell 25 markdown]
# **Discretizing numeric attributes**

[cell 26 markdown]
We turned the numeric features `tenure`, `MonthlyCharges`, and `TotalCharges` into **4 categories (bins)** each, using a method called **equi-frequency binning**.

* **Equi-frequency binning** means we split the data so that **each bin has about the same number of data points**.
* To do this, we used `KBinsDiscretizer` from scikit-learn with the **`quantile` strategy**, which automatically creates bins with equal numbers of samples.
* We set the number of bins to **4**, so each bin holds roughly **25% of the data**.
* The new columns (`tenure_bin`, `MonthlyCharges_bin`, and `TotalCharges_bin`) contain values from **0 to 3**, showing which bin each value falls into.

---

[cell 27 code]
# resetting churn to yes and no for consistency
df["Churn"] = df["Churn"].map({1: "Yes", 0: "No"}).astype("object")

df[['tenure','MonthlyCharges','TotalCharges']]

[cell 28 code]
from sklearn.preprocessing import KBinsDiscretizer


# Discretize each column separately
disc_tenure = KBinsDiscretizer(n_bins=4, encode='ordinal', strategy='quantile')
df["tenure_bin"] = disc_tenure.fit_transform(df[["tenure"]])

disc_month = KBinsDiscretizer(n_bins=4, encode='ordinal', strategy='quantile')
df["MonthlyCharges_bin"] = disc_month.fit_transform(df[["MonthlyCharges"]])

disc_total = KBinsDiscretizer(n_bins=4, encode='ordinal', strategy='quantile')
df["TotalCharges_bin"] = disc_total.fit_transform(df[["TotalCharges"]])

[cell 29 code]
df[['tenure_bin','MonthlyCharges_bin','TotalCharges_bin']]

[cell 30 markdown]
## **Replacing numeric attributes with their discretized versions and storing in a new dataframe**

[cell 31 code]
#  Make a copy of original df
df2 = df.copy()

df2 = df2.drop(columns=["customerID"])

# Drop original numeric features
df2=df2.drop(['tenure', 'MonthlyCharges', 'TotalCharges'], axis=1)

df2['tenure_bin'] = df2['tenure_bin'].astype('object')
df2['MonthlyCharges_bin'] = df2['MonthlyCharges_bin'].astype('object')
df2['TotalCharges_bin'] = df2['TotalCharges_bin'].astype('object')


# df2 is now ready for modeling
df2.head()

[cell 32 code]
df2.info()

[cell 33 markdown]
# **Count plot for target variable shows label imbalance**

[cell 34 code]
print(df2['Churn'].value_counts(normalize=True))
df2['Churn'].value_counts(normalize=True).mul(100).plot(kind='bar')

[cell 35 markdown]
# Preparing Features and Target Variable

* The target column `Churn` is first converted from `"Yes"/"No"` to `1/0` for binary classification.  
* It is then separated from the dataset to form `y` (labels), while the remaining columns become `X` (features).  
* Finally, `X` is one-hot encoded using `pd.get_dummies()` to handle categorical variables.

[cell 36 code]
# Convert Churn to numeric target
df2["Churn"] = df2["Churn"].str.strip().map({"Yes": 1, "No": 0})


#  Split into X and y
X = df2.drop("Churn", axis=1)
y = df2["Churn"]

# One-hot encode X only
X_encoded = pd.get_dummies(X, drop_first=True)

[cell 37 code]
X

[cell 38 code]
X['InternetService'].value_counts()

[cell 39 code]
X_encoded

[cell 40 markdown]
# **Splitting the dataset into train-test sets**

[cell 41 code]
from sklearn.model_selection import train_test_split

#Train-test split
X_train, X_test, y_train, y_test = train_test_split(X_encoded, y, test_size=0.1, random_state=42,stratify=y)



# Final shapes:
print("Train shape:", X_train.shape)
print("Test shape:", X_test.shape)

[cell 42 markdown]
# **Part 1: Decision Tree**

[cell 43 markdown]
# **Selecting Optimal Tree Depth for Decision Tree Classifier**

[cell 44 code]
import matplotlib.pyplot as plt
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score

train_acc = []
test_acc = []
depth_range = list(range(1, 30))  # test depths from 1 to 29

for depth in depth_range:
    clf = DecisionTreeClassifier(criterion="entropy", max_depth=depth, random_state=42)
    clf.fit(X_train, y_train)

    train_pred = clf.predict(X_train)
    val_pred = clf.predict(X_test)

    train_acc.append(accuracy_score(y_train, train_pred))
    test_acc.append(accuracy_score(y_test, val_pred))

# Plotting
plt.figure(figsize=(10, 6))
plt.plot(depth_range, train_acc, label="Train Accuracy", marker='o')
plt.plot(depth_range, test_acc, label="Test Accuracy", marker='s')
plt.xlabel("Tree Depth")
plt.ylabel("Accuracy")
plt.title("Overfitting vs Underfitting (Criterion = Entropy)")
plt.legend()
plt.grid(True)
plt.show()

[cell 45 code]
df_test_acc = pd.DataFrame({"Test Accuracy": test_acc[:15]}, index=range(1, 16))

df_test_acc.T

[cell 46 markdown]
* The best tree depth is 6, where test accuracy peaks before overfitting begins to increase.

[cell 47 markdown]
# **Impact of Train-Test Ratio on Decision Tree Performance**

[cell 48 code]
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, roc_auc_score
import matplotlib.pyplot as plt
import numpy as np

# Set the best depth found from previous step
best_depth = 6  #


split_ratios = [0.6, 0.7, 0.8, 0.9]  # train split ratios
metrics = {
    "Train Accuracy": [],
    "Test Accuracy": [],
    "Train Precision": [],
    "Test Precision": [],
    "Train Recall": [],
    "Test Recall": [],
    "Train F1": [],
    "Test F1": [],
    "Train AUC": [],
    "Test AUC": []
}

for ratio in split_ratios:
    X_train, X_test, y_train, y_test = train_test_split(
        X_encoded, y, train_size=ratio, random_state=42, stratify=y
    )

    clf = DecisionTreeClassifier(max_depth=best_depth, random_state=42,criterion="entropy")
    clf.fit(X_train, y_train)

    y_train_pred = clf.predict(X_train)
    y_test_pred = clf.predict(X_test)

    # For AUC, we need probabilities and binary classification
    y_train_prob = clf.predict_proba(X_train)[:, 1]
    y_test_prob = clf.predict_proba(X_test)[:, 1]

    metrics["Train Accuracy"].append(accuracy_score(y_train, y_train_pred))
    metrics["Test Accuracy"].append(accuracy_score(y_test, y_test_pred))

    metrics["Train Precision"].append(precision_score(y_train, y_train_pred, zero_division=0))
    metrics["Test Precision"].append(precision_score(y_test, y_test_pred, zero_division=0))

    metrics["Train Recall"].append(recall_score(y_train, y_train_pred, zero_division=0))
    metrics["Test Recall"].append(recall_score(y_test, y_test_pred, zero_division=0))

    metrics["Train F1"].append(f1_score(y_train, y_train_pred, zero_division=0))
    metrics["Test F1"].append(f1_score(y_test, y_test_pred, zero_division=0))

    metrics["Train AUC"].append(roc_auc_score(y_train, y_train_prob))
    metrics["Test AUC"].append(roc_auc_score(y_test, y_test_prob))

[cell 49 code]
import matplotlib.pyplot as plt
import numpy as np

def plot_all_metrics_in_grid(metrics, split_ratios):
    metric_names = ["Accuracy", "Precision", "Recall", "F1", "AUC"]
    n_metrics = len(metric_names)

    fig, axes = plt.subplots(2, 3, figsize=(14, 8))  # 2 rows, 3 columns
    axes = axes.flatten()

    for idx, metric_name in enumerate(metric_names):
        ax = axes[idx]
        ax.plot(np.array(split_ratios)*100, metrics[f"Train {metric_name}"], marker='o', label='Train')
        ax.plot(np.array(split_ratios)*100, metrics[f"Test {metric_name}"], marker='s', label='Test')
        ax.set_title(f"{metric_name}")
        ax.set_xlabel("Train Size (%)")
        ax.set_ylabel(metric_name)
        ax.grid(True)
        ax.legend()

    # Hide the extra subplot (6th one)
    if n_metrics < len(axes):
        axes[-1].axis("off")

    fig.suptitle("Model Performance vs Train/Test Split Ratio", fontsize=16)
    plt.tight_layout(rect=[0, 0, 1, 0.95])  # Adjust space for suptitle
    plt.show()

# Call the function
plot_all_metrics_in_grid(metrics, split_ratios)

[cell 50 markdown]
## **Observations**

**Accuracy:**  
Test accuracy gradually improves as train size increases from 60% to 80 and a slight dip at 90%. The general trend is that the model benefits from more training data.

**Precision:**  
There is a slight upward trend in test precision, highest at the 90/10 split. The model becomes more confident in its positive predictions with more training data.

**Recall:**  
Test recall peaks around the 80/20 split and drops significantly at 90%, suggesting that higher train sizes may reduce the model’s ability to detect positive cases.

**F1 Score:**  
F1 score is highest at 80/20, showing the best trade-off between precision and recall. Both 70/30 and 90/10 are slightly worse in terms of balance.

**AUC:**  
AUC is highest around 60/40 and then declines with the lowest at 90/10. This suggests the model’s ability to distinguish between classes slightly weakens when the test set is too small.


## **Final Decision:**
The **80/20 split** offers the most balanced performance across metrics and we choose it for final our Decision Tree classifier.

[cell 51 markdown]
# **A :Final Decision Tree Training and Evaluation**

[cell 52 code]
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import (
    accuracy_score, precision_score, recall_score,
    f1_score, roc_auc_score, classification_report
)

# 1. Train-test split (80/20)
X_train, X_test, y_train, y_test = train_test_split(
    X_encoded, y, train_size=0.8, random_state=42, stratify=y
)

# 2. Train Decision Tree (use best depth found earlier)
clf = DecisionTreeClassifier(max_depth=6, criterion="entropy", random_state=42)
clf.fit(X_train, y_train)

# 3. Predictions
y_train_pred = clf.predict(X_train)
y_test_pred = clf.predict(X_test)

# 4. Probabilities (for AUC)
y_train_prob = clf.predict_proba(X_train)[:, 1]
y_test_prob = clf.predict_proba(X_test)[:, 1]

# 5. Evaluation metrics
print("=== Train Metrics ===")
print("Accuracy:", accuracy_score(y_train, y_train_pred))
print("Precision:", precision_score(y_train, y_train_pred, zero_division=0))
print("Recall:", recall_score(y_train, y_train_pred, zero_division=0))
print("F1 Score:", f1_score(y_train, y_train_pred, zero_division=0))
print("AUC:", roc_auc_score(y_train, y_train_prob))

print("\n=== Test Metrics ===")
print("Accuracy:", accuracy_score(y_test, y_test_pred))
print("Precision:", precision_score(y_test, y_test_pred, zero_division=0))
print("Recall:", recall_score(y_test, y_test_pred, zero_division=0))
print("F1 Score:", f1_score(y_test, y_test_pred, zero_division=0))
print("AUC:", roc_auc_score(y_test, y_test_prob))

# 6. Classification report (detailed per-class)
print("\nClassification Report:\n", classification_report(y_test, y_test_pred))

[cell 53 markdown]
## **Observations**

* The model achieves a test accuracy of **78.46%**, which is close to the train accuracy, indicating no major overfitting.
* Precision and recall are lower for class `1` (the positive/minority class) due to label imbalance.
* The **AUC of 0.81** on the test set shows the model has good ability to separate churn vs non-churn.
* Class `0` (non-churn) is predicted more confidently, as seen by its higher precision and recall in the classification report.
* The model generalizes well, but could benefit from improving recall on the positive class.

[cell 54 markdown]
# **Visualizing the Trained Decision Tree**

[cell 55 code]
from sklearn.tree import plot_tree
plt.figure(figsize=(30, 10))
plot_tree(clf,
          feature_names=X_train.columns,
          class_names=[str(c) for c in clf.classes_],
          filled=True,
          rounded=True,
          fontsize=10)
plt.show()

[cell 56 markdown]
## **Showing prediction for some test samples**

[cell 57 code]
# Get model predictions
y_test_pred = clf.predict(X_test)

# DataFrame combining test features, ground truth, and prediction
X_orig_train, X_orig_test = train_test_split(
    X, train_size=0.8, random_state=42, stratify=y
)

#  dataframe for inspection
inspect_df = X_orig_test.copy()
inspect_df["Ground Truth"] = y_test.values
inspect_df["Prediction"] = y_test_pred

# few readable examples
inspect_df[10:20]

[cell 58 markdown]
# **B: Decision Tree Training and Evaluation  using gini as criterion**

[cell 59 code]
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import (
    accuracy_score, precision_score, recall_score,
    f1_score, roc_auc_score, classification_report
)

# 1. Train-test split (80/20)
X_train, X_test, y_train, y_test = train_test_split(
    X_encoded, y, train_size=0.8, random_state=42, stratify=y
)

# 2. Train Decision Tree (use best depth found earlier)
clf = DecisionTreeClassifier(max_depth=6, criterion="gini", random_state=42)
clf.fit(X_train, y_train)

# 3. Predictions
y_train_pred = clf.predict(X_train)
y_test_pred = clf.predict(X_test)

# 4. Probabilities (for AUC)
y_train_prob = clf.predict_proba(X_train)[:, 1]
y_test_prob = clf.predict_proba(X_test)[:, 1]

# 5. Evaluation metrics
print("=== Train Metrics ===")
print("Accuracy:", accuracy_score(y_train, y_train_pred))
print("Precision:", precision_score(y_train, y_train_pred, zero_division=0))
print("Recall:", recall_score(y_train, y_train_pred, zero_division=0))
print("F1 Score:", f1_score(y_train, y_train_pred, zero_division=0))
print("AUC:", roc_auc_score(y_train, y_train_prob))

print("\n=== Test Metrics ===")
print("Accuracy:", accuracy_score(y_test, y_test_pred))
print("Precision:", precision_score(y_test, y_test_pred, zero_division=0))
print("Recall:", recall_score(y_test, y_test_pred, zero_division=0))
print("F1 Score:", f1_score(y_test, y_test_pred, zero_division=0))
print("AUC:", roc_auc_score(y_test, y_test_prob))

# 6. Classification report (detailed per-class)
print("\nClassification Report:\n", classification_report(y_test, y_test_pred))

[cell 60 markdown]
## **Observations**
* The model achieves a test accuracy of **78.46%**, similar to the entropy-based version.
* Test precision improves slightly (to **0.62**), but **recall drops** to **0.47**, indicating more false negatives (missed churns).    
* AUC remains solid at **0.81**, confirming good class separation overall.  
* Class `0` is still predicted strongly, but the model shows more hesitation on class `1`, as seen from the lower recall.  
* The gap between **train and test F1-scores** is larger compared to the entropy model, indicating slightly overfitting with Gini.

[cell 61 code]
# Plot tree
plt.figure(figsize=(30, 10))
plot_tree(clf,
          feature_names=X_train.columns,
          class_names=[str(c) for c in clf.classes_],
          filled=True,
          rounded=True,
          fontsize=10)
plt.show()

[cell 62 markdown]
# **Another way to find optimal depth**

[cell 63 markdown]
* In the previous approach, we selected the depth once using a 90/10 split and used it throughout the code to decide the optimal train/test ratio and perform the final model evaluation.

* In this part, we find the best depth separately at each train-test split, and then use the corresponding depth and split to train the model and perform evaluation.

[cell 64 markdown]
## **For 60/40  train-test split**

[cell 65 code]
from sklearn.model_selection import train_test_split

#Train-test split

print('For 60/40  train-test split')


train_size=0.6
X_train, X_test, y_train, y_test = train_test_split(X_encoded, y, train_size=train_size, random_state=42,stratify=y)




train_acc = []
test_acc = []
depth_range = list(range(1, 30))  # test depths from 1 to 29

for depth in depth_range:
    clf = DecisionTreeClassifier(criterion="entropy", max_depth=depth, random_state=42)
    clf.fit(X_train, y_train)

    train_pred = clf.predict(X_train)
    val_pred = clf.predict(X_test)

    train_acc.append(accuracy_score(y_train, train_pred))
    test_acc.append(accuracy_score(y_test, val_pred))

# Plotting
plt.figure(figsize=(10, 6))
plt.plot(depth_range, train_acc, label="Train Accuracy", marker='o')
plt.plot(depth_range, test_acc, label="Test Accuracy", marker='s')
plt.xlabel("Tree Depth")
plt.ylabel("Accuracy")
plt.title("Overfitting vs Underfitting (Criterion = Entropy)")
plt.legend()
plt.grid(True)
plt.show()

df_test_acc = pd.DataFrame({"Test Accuracy": test_acc[:15]}, index=range(1, 16))

depth_acc_dict = {depth: round(acc, 4) for depth, acc in zip(df_test_acc.index, df_test_acc["Test Accuracy"])}
print('Value of accuracy at first 15 depth sizes:')
print(depth_acc_dict)

# Find the depth with maximum accuracy
best_depth = max(depth_acc_dict, key=depth_acc_dict.get)
print(f'\nDepth with maximum accuracy: {best_depth} (Accuracy = {depth_acc_dict[best_depth]})')




# 1. Train-test split
X_train, X_test, y_train, y_test = train_test_split(
    X_encoded, y, train_size=train_size, random_state=42, stratify=y
)

# 2. Train Decision Tree (use best depth found earlier)
clf = DecisionTreeClassifier(max_depth=best_depth, criterion="entropy", random_state=42)
clf.fit(X_train, y_train)

# 3. Predictions
y_train_pred = clf.predict(X_train)
y_test_pred = clf.predict(X_test)

# 4. Probabilities (for AUC)
y_train_prob = clf.predict_proba(X_train)[:, 1]
y_test_prob = clf.predict_proba(X_test)[:, 1]

# 5. Evaluation metrics
print("=== Train Metrics ===")
print("Accuracy:", accuracy_score(y_train, y_train_pred))
print("Precision:", precision_score(y_train, y_train_pred, zero_division=0))
print("Recall:", recall_score(y_train, y_train_pred, zero_division=0))
print("F1 Score:", f1_score(y_train, y_train_pred, zero_division=0))
print("AUC:", roc_auc_score(y_train, y_train_prob))

print("\n=== Test Metrics ===")
print("Accuracy:", accuracy_score(y_test, y_test_pred))
print("Precision:", precision_score(y_test, y_test_pred, zero_division=0))
print("Recall:", recall_score(y_test, y_test_pred, zero_division=0))
print("F1 Score:", f1_score(y_test, y_test_pred, zero_division=0))
print("AUC:", roc_auc_score(y_test, y_test_prob))

# 6. Classification report (detailed per-class)
print("\nClassification Report:\n", classification_report(y_test, y_test_pred))

[cell 66 markdown]
## **For 70/30  train-test split**

[cell 67 code]
from sklearn.model_selection import train_test_split

#Train-test split

print('For 70/30 train split')


train_size=0.7
X_train, X_test, y_train, y_test = train_test_split(X_encoded, y, train_size=train_size, random_state=42,stratify=y)




train_acc = []
test_acc = []
depth_range = list(range(1, 30))  # test depths from 1 to 29

for depth in depth_range:
    clf = DecisionTreeClassifier(criterion="entropy", max_depth=depth, random_state=42)
    clf.fit(X_train, y_train)

    train_pred = clf.predict(X_train)
    val_pred = clf.predict(X_test)

    train_acc.append(accuracy_score(y_train, train_pred))
    test_acc.append(accuracy_score(y_test, val_pred))

# Plotting
plt.figure(figsize=(10, 6))
plt.plot(depth_range, train_acc, label="Train Accuracy", marker='o')
plt.plot(depth_range, test_acc, label="Test Accuracy", marker='s')
plt.xlabel("Tree Depth")
plt.ylabel("Accuracy")
plt.title("Overfitting vs Underfitting (Criterion = Entropy)")
plt.legend()
plt.grid(True)
plt.show()

df_test_acc = pd.DataFrame({"Test Accuracy": test_acc[:15]}, index=range(1, 16))

depth_acc_dict = {depth: round(acc, 4) for depth, acc in zip(df_test_acc.index, df_test_acc["Test Accuracy"])}
print('Value of accuracy at first 15 depth sizes:')
print(depth_acc_dict)

# Find the depth with maximum accuracy
best_depth = max(depth_acc_dict, key=depth_acc_dict.get)
print(f'\nDepth with maximum accuracy: {best_depth} (Accuracy = {depth_acc_dict[best_depth]})')




# 1. Train-test split
X_train, X_test, y_train, y_test = train_test_split(
    X_encoded, y, train_size=train_size, random_state=42, stratify=y
)

# 2. Train Decision Tree (use best depth found earlier)
clf = DecisionTreeClassifier(max_depth=best_depth, criterion="entropy", random_state=42)
clf.fit(X_train, y_train)

# 3. Predictions
y_train_pred = clf.predict(X_train)
y_test_pred = clf.predict(X_test)

# 4. Probabilities (for AUC)
y_train_prob = clf.predict_proba(X_train)[:, 1]
y_test_prob = clf.predict_proba(X_test)[:, 1]

# 5. Evaluation metrics
print("=== Train Metrics ===")
print("Accuracy:", accuracy_score(y_train, y_train_pred))
print("Precision:", precision_score(y_train, y_train_pred, zero_division=0))
print("Recall:", recall_score(y_train, y_train_pred, zero_division=0))
print("F1 Score:", f1_score(y_train, y_train_pred, zero_division=0))
print("AUC:", roc_auc_score(y_train, y_train_prob))

print("\n=== Test Metrics ===")
print("Accuracy:", accuracy_score(y_test, y_test_pred))
print("Precision:", precision_score(y_test, y_test_pred, zero_division=0))
print("Recall:", recall_score(y_test, y_test_pred, zero_division=0))
print("F1 Score:", f1_score(y_test, y_test_pred, zero_division=0))
print("AUC:", roc_auc_score(y_test, y_test_prob))

# 6. Classification report (detailed per-class)
print("\nClassification Report:\n", classification_report(y_test, y_test_pred))

[cell 68 markdown]
## **For 80/20  train-test split**

[cell 69 code]
from sklearn.model_selection import train_test_split

#Train-test split

print('For 80/20 train split')


train_size=0.8
X_train, X_test, y_train, y_test = train_test_split(X_encoded, y, train_size=train_size, random_state=42,stratify=y)




train_acc = []
test_acc = []
depth_range = list(range(1, 30))  # test depths from 1 to 29

for depth in depth_range:
    clf = DecisionTreeClassifier(criterion="entropy", max_depth=depth, random_state=42)
    clf.fit(X_train, y_train)

    train_pred = clf.predict(X_train)
    val_pred = clf.predict(X_test)

    train_acc.append(accuracy_score(y_train, train_pred))
    test_acc.append(accuracy_score(y_test, val_pred))

# Plotting
plt.figure(figsize=(10, 6))
plt.plot(depth_range, train_acc, label="Train Accuracy", marker='o')
plt.plot(depth_range, test_acc, label="Test Accuracy", marker='s')
plt.xlabel("Tree Depth")
plt.ylabel("Accuracy")
plt.title("Overfitting vs Underfitting (Criterion = Entropy)")
plt.legend()
plt.grid(True)
plt.show()

df_test_acc = pd.DataFrame({"Test Accuracy": test_acc[:15]}, index=range(1, 16))

depth_acc_dict = {depth: round(acc, 4) for depth, acc in zip(df_test_acc.index, df_test_acc["Test Accuracy"])}
print('Value of accuracy at first 15 depth sizes:')
print(depth_acc_dict)

# Find the depth with maximum accuracy
best_depth = max(depth_acc_dict, key=depth_acc_dict.get)
print(f'\nDepth with maximum accuracy: {best_depth} (Accuracy = {depth_acc_dict[best_depth]})')




# 1. Train-test split
X_train, X_test, y_train, y_test = train_test_split(
    X_encoded, y, train_size=train_size, random_state=42, stratify=y
)

# 2. Train Decision Tree (use best depth found earlier)
clf = DecisionTreeClassifier(max_depth=best_depth, criterion="entropy", random_state=42)
clf.fit(X_train, y_train)

# 3. Predictions
y_train_pred = clf.predict(X_train)
y_test_pred = clf.predict(X_test)

# 4. Probabilities (for AUC)
y_train_prob = clf.predict_proba(X_train)[:, 1]
y_test_prob = clf.predict_proba(X_test)[:, 1]

# 5. Evaluation metrics
print("=== Train Metrics ===")
print("Accuracy:", accuracy_score(y_train, y_train_pred))
print("Precision:", precision_score(y_train, y_train_pred, zero_division=0))
print("Recall:", recall_score(y_train, y_train_pred, zero_division=0))
print("F1 Score:", f1_score(y_train, y_train_pred, zero_division=0))
print("AUC:", roc_auc_score(y_train, y_train_prob))

print("\n=== Test Metrics ===")
print("Accuracy:", accuracy_score(y_test, y_test_pred))
print("Precision:", precision_score(y_test, y_test_pred, zero_division=0))
print("Recall:", recall_score(y_test, y_test_pred, zero_division=0))
print("F1 Score:", f1_score(y_test, y_test_pred, zero_division=0))
print("AUC:", roc_auc_score(y_test, y_test_prob))

# 6. Classification report (detailed per-class)
print("\nClassification Report:\n", classification_report(y_test, y_test_pred))

[cell 70 markdown]
## **For 90/10  train-test split**

[cell 71 code]
from sklearn.model_selection import train_test_split

#Train-test split

print('For 90/10 train split')


train_size=0.6
X_train, X_test, y_train, y_test = train_test_split(X_encoded, y, train_size=train_size, random_state=42,stratify=y)




train_acc = []
test_acc = []
depth_range = list(range(1, 30))  # test depths from 1 to 29

for depth in depth_range:
    clf = DecisionTreeClassifier(criterion="entropy", max_depth=depth, random_state=42)
    clf.fit(X_train, y_train)

    train_pred = clf.predict(X_train)
    val_pred = clf.predict(X_test)

    train_acc.append(accuracy_score(y_train, train_pred))
    test_acc.append(accuracy_score(y_test, val_pred))

# Plotting
plt.figure(figsize=(10, 6))
plt.plot(depth_range, train_acc, label="Train Accuracy", marker='o')
plt.plot(depth_range, test_acc, label="Test Accuracy", marker='s')
plt.xlabel("Tree Depth")
plt.ylabel("Accuracy")
plt.title("Overfitting vs Underfitting (Criterion = Entropy)")
plt.legend()
plt.grid(True)
plt.show()

df_test_acc = pd.DataFrame({"Test Accuracy": test_acc[:15]}, index=range(1, 16))

depth_acc_dict = {depth: round(acc, 4) for depth, acc in zip(df_test_acc.index, df_test_acc["Test Accuracy"])}
print('Value of accuracy at first 15 depth sizes:')
print(depth_acc_dict)

# Find the depth with maximum accuracy
best_depth = max(depth_acc_dict, key=depth_acc_dict.get)
print(f'\nDepth with maximum accuracy: {best_depth} (Accuracy = {depth_acc_dict[best_depth]})')




# 1. Train-test split (80/20)
X_train, X_test, y_train, y_test = train_test_split(
    X_encoded, y, train_size=train_size, random_state=42, stratify=y
)

# 2. Train Decision Tree (use best depth found earlier)
clf = DecisionTreeClassifier(max_depth=best_depth, criterion="entropy", random_state=42)
clf.fit(X_train, y_train)

# 3. Predictions
y_train_pred = clf.predict(X_train)
y_test_pred = clf.predict(X_test)

# 4. Probabilities (for AUC)
y_train_prob = clf.predict_proba(X_train)[:, 1]
y_test_prob = clf.predict_proba(X_test)[:, 1]

# 5. Evaluation metrics
print("=== Train Metrics ===")
print("Accuracy:", accuracy_score(y_train, y_train_pred))
print("Precision:", precision_score(y_train, y_train_pred, zero_division=0))
print("Recall:", recall_score(y_train, y_train_pred, zero_division=0))
print("F1 Score:", f1_score(y_train, y_train_pred, zero_division=0))
print("AUC:", roc_auc_score(y_train, y_train_prob))

print("\n=== Test Metrics ===")
print("Accuracy:", accuracy_score(y_test, y_test_pred))
print("Precision:", precision_score(y_test, y_test_pred, zero_division=0))
print("Recall:", recall_score(y_test, y_test_pred, zero_division=0))
print("F1 Score:", f1_score(y_test, y_test_pred, zero_division=0))
print("AUC:", roc_auc_score(y_test, y_test_prob))

# 6. Classification report (detailed per-class)
print("\nClassification Report:\n", classification_report(y_test, y_test_pred))

[cell 72 markdown]
# **Part 2: Random Forest**
There are two important hyperparameters in Random Forest — `n_estimators` (number of trees) and `max_features` (number of features considered at each split)

[cell 73 markdown]
# **Selecting optimal n_estimators (number of trees) in Random Forest**

* We evaluated the following range: [50, 75, 100, 150]

[cell 74 code]
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import (
    accuracy_score, precision_score, recall_score,
    f1_score, roc_auc_score
)
from sklearn.model_selection import train_test_split
import matplotlib.pyplot as plt
import numpy as np

# 1. Train/test split (80/20)
X_train, X_test, y_train, y_test = train_test_split(
    X_encoded, y, train_size=0.8, random_state=42, stratify=y
)

# 2. Varying n_estimators
estimator_range = [50, 75, 100, 150]
metrics_rf = {
    "Train Accuracy": [], "Test Accuracy": [],
    "Train Precision": [], "Test Precision": [],
    "Train Recall": [], "Test Recall": [],
    "Train F1": [], "Test F1": [],
    "Train AUC": [], "Test AUC": []
}

for n in estimator_range:
    clf = RandomForestClassifier(n_estimators=n, criterion='entropy', max_depth=6, random_state=42)
    clf.fit(X_train, y_train)

    y_train_pred = clf.predict(X_train)
    y_test_pred = clf.predict(X_test)
    y_train_prob = clf.predict_proba(X_train)[:, 1]
    y_test_prob = clf.predict_proba(X_test)[:, 1]

    metrics_rf["Train Accuracy"].append(accuracy_score(y_train, y_train_pred))
    metrics_rf["Test Accuracy"].append(accuracy_score(y_test, y_test_pred))

    metrics_rf["Train Precision"].append(precision_score(y_train, y_train_pred, zero_division=0))
    metrics_rf["Test Precision"].append(precision_score(y_test, y_test_pred, zero_division=0))

    metrics_rf["Train Recall"].append(recall_score(y_train, y_train_pred, zero_division=0))
    metrics_rf["Test Recall"].append(recall_score(y_test, y_test_pred, zero_division=0))

    metrics_rf["Train F1"].append(f1_score(y_train, y_train_pred, zero_division=0))
    metrics_rf["Test F1"].append(f1_score(y_test, y_test_pred, zero_division=0))

    metrics_rf["Train AUC"].append(roc_auc_score(y_train, y_train_prob))
    metrics_rf["Test AUC"].append(roc_auc_score(y_test, y_test_prob))

[cell 75 code]
def plot_rf_metrics_grid(metrics, estimator_range):
    metric_names = ["Accuracy", "Precision", "Recall", "F1", "AUC"]
    fig, axes = plt.subplots(2, 3, figsize=(15, 8))
    axes = axes.flatten()

    for idx, metric in enumerate(metric_names):
        ax = axes[idx]
        ax.plot(estimator_range, metrics[f"Train {metric}"], marker='o', label="Train")
        ax.plot(estimator_range, metrics[f"Test {metric}"], marker='s', label="Test")
        ax.set_title(metric)
        ax.set_xlabel("n_estimators")
        ax.set_ylabel(metric)
        ax.grid(True)
        ax.legend()

    if len(metric_names) < len(axes):
        axes[-1].axis("off")  # Hide unused subplot

    fig.suptitle("Random Forest Performance vs n_estimators", fontsize=16)
    plt.tight_layout(rect=[0, 0, 1, 0.95])
    plt.show()

# Call the plot
plot_rf_metrics_grid(metrics_rf, estimator_range)

[cell 76 markdown]
## **Observation**

* **Accuracy:** Both train and test accuracy show a dip at 100, followed by improvement in accuracy at `n_estimators = 150`.
* **Precision:** Train precision is consistent; test precision dips at 100 trees but recovers slightly by 150.
* **Recall:** Train and Test recall gradually improves as the number of estimators increases, peaking at 150.
* **F1 Score:** Test F1 is highest at `n_estimators = 150`, suggesting the best balance between precision and recall at that point.
* **AUC:** Test AUC is fairly flat, with a slight decline at 150 trees

* We use `n_estimators = 150` as it gives best overall test performance, especially in terms of recall and F1 score and comparable performance in other metrics

[cell 77 markdown]
# **Selecting optimal max_features(number of features) in Random Forest**

We evaluated the following range: [5, 10, 15,20,25, 29]

[cell 78 code]
import numpy as np
import matplotlib.pyplot as plt
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import (
    accuracy_score, precision_score, recall_score,
    f1_score, roc_auc_score
)
from sklearn.model_selection import train_test_split

# === PARAMETERS ===
n_total_features = X.shape[1]
n_estimators = 150
max_depth = 6

# === 1. Train/test split ===
X_train, X_test, y_train, y_test = train_test_split(
    X_encoded, y, train_size=0.8, random_state=42, stratify=y
)

# === 2. Define max_features options ===
max_feature_range = [5, 10, 15,20,25, 29]


# === 3. Metrics Storage ===
metrics_rf = {
    "Train Accuracy": [], "Test Accuracy": [],
    "Train Precision": [], "Test Precision": [],
    "Train Recall": [], "Test Recall": [],
    "Train F1": [], "Test F1": [],
    "Train AUC": [], "Test AUC": []
}

# === 4. Loop over max_features ===
for mf in max_feature_range:
    clf = RandomForestClassifier(
        n_estimators=n_estimators,
        max_depth=max_depth,
        max_features=mf,
        criterion='entropy',
        random_state=42
    )
    clf.fit(X_train, y_train)

    y_train_pred = clf.predict(X_train)
    y_test_pred = clf.predict(X_test)
    y_train_prob = clf.predict_proba(X_train)[:, 1]
    y_test_prob = clf.predict_proba(X_test)[:, 1]

    metrics_rf["Train Accuracy"].append(accuracy_score(y_train, y_train_pred))
    metrics_rf["Test Accuracy"].append(accuracy_score(y_test, y_test_pred))

    metrics_rf["Train Precision"].append(precision_score(y_train, y_train_pred, zero_division=0))
    metrics_rf["Test Precision"].append(precision_score(y_test, y_test_pred, zero_division=0))

    metrics_rf["Train Recall"].append(recall_score(y_train, y_train_pred, zero_division=0))
    metrics_rf["Test Recall"].append(recall_score(y_test, y_test_pred, zero_division=0))

    metrics_rf["Train F1"].append(f1_score(y_train, y_train_pred, zero_division=0))
    metrics_rf["Test F1"].append(f1_score(y_test, y_test_pred, zero_division=0))

    metrics_rf["Train AUC"].append(roc_auc_score(y_train, y_train_prob))
    metrics_rf["Test AUC"].append(roc_auc_score(y_test, y_test_prob))

[cell 79 code]
def plot_rf_metrics_grid(metrics, x_vals):
    metric_names = ["Accuracy", "Precision", "Recall", "F1", "AUC"]
    fig, axes = plt.subplots(2, 3, figsize=(15, 8))
    axes = axes.flatten()

    for idx, metric in enumerate(metric_names):
        ax = axes[idx]
        ax.plot(x_vals, metrics[f"Train {metric}"], marker='o', label="Train")
        ax.plot(x_vals, metrics[f"Test {metric}"], marker='s', label="Test")
        ax.set_title(metric)
        ax.set_xlabel("max_features")
        ax.set_ylabel(metric)
        ax.set_xticks(x_vals)
        ax.grid(True)
        ax.legend()

    # If extra subplot, hide it
    if len(metric_names) < len(axes):
        axes[-1].axis("off")

    fig.suptitle("Random Forest Performance vs max_features", fontsize=16)
    plt.tight_layout(rect=[0, 0, 1, 0.95])
    plt.show()


plot_rf_metrics_grid(metrics_rf,max_feature_range )

[cell 80 markdown]
## Observations

* **Accuracy:** Test accuracy is highest at **`max_features = 29`**, indicating the model performs best when all features are considered at each split.
* **Precision:** Test precision increases steadily with `max_features`, improving after 20 and peaking at 29.
* **Recall:** Test recall peaks at `max_features = 15`, then slightly declines — suggesting fewer false negatives around this point.
* **F1 Score:** Test F1-score is highest at `max_features = 15`, showing the best precision-recall balance.
* **AUC:** Test AUC is slightly higher at `5–10`, then gradually drops as `max_features` increases, showing weaker class separation with more features.

 We choose **`max_features = 15`**, as it gives the best trade-off between precision and recall, and achieves the highest F1-score, with a comparable performance in other metrics

[cell 81 markdown]
# **C:Final Random Forest training and Evaluation**

[cell 82 code]
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import (
    accuracy_score, precision_score, recall_score,
    f1_score, roc_auc_score, classification_report
)

# === 1. Train-test split (80/20) ===
X_train, X_test, y_train, y_test = train_test_split(
    X_encoded, y, train_size=0.8, random_state=42, stratify=y
)

# === 2. Final Random Forest ===
clf = RandomForestClassifier(
    n_estimators=150,
    max_depth=6,
    max_features=15,
    criterion='entropy',
    random_state=42
)

clf.fit(X_train, y_train)

# === 3. Predictions ===
y_train_pred = clf.predict(X_train)
y_test_pred = clf.predict(X_test)
y_train_prob = clf.predict_proba(X_train)[:, 1]
y_test_prob = clf.predict_proba(X_test)[:, 1]

# === 4. Evaluation ===
print("=== Train Metrics ===")
print("Accuracy:", accuracy_score(y_train, y_train_pred))
print("Precision:", precision_score(y_train, y_train_pred, zero_division=0))
print("Recall:", recall_score(y_train, y_train_pred, zero_division=0))
print("F1 Score:", f1_score(y_train, y_train_pred, zero_division=0))
print("AUC:", roc_auc_score(y_train, y_train_prob))

print("\n=== Test Metrics ===")
print("Accuracy:", accuracy_score(y_test, y_test_pred))
print("Precision:", precision_score(y_test, y_test_pred, zero_division=0))
print("Recall:", recall_score(y_test, y_test_pred, zero_division=0))
print("F1 Score:", f1_score(y_test, y_test_pred, zero_division=0))
print("AUC:", roc_auc_score(y_test, y_test_prob))

# === 5. Full classification report ===
print("\n=== Classification Report ===")
print(classification_report(y_test, y_test_pred))

[cell 83 markdown]
# **Observations**


* The model achieves a **test accuracy of approximately 79.9%**, with minimal difference from training accuracy, indicating good generalization to unseen data.
* The **test precision (0.65)** and **F1-score (~0.58)** suggest that the model maintains a reasonable balance between correctly identifying churn cases and minimizing false positives.
* The **recall of ~0.52** indicates moderate sensitivity to detecting actual churners, which is important in the context of imbalanced datasets.
* The **AUC score of ~0.83** confirms that the model has a strong ability to distinguish between the two classes based on predicted probabilities.

[cell 84 markdown]
# **D :Random Forest training and Evaluation using gini as criterion**

[cell 85 code]
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import (
    accuracy_score, precision_score, recall_score,
    f1_score, roc_auc_score, classification_report
)

# === 1. Train-test split (80/20) ===
X_train, X_test, y_train, y_test = train_test_split(
    X_encoded, y, train_size=0.8, random_state=42, stratify=y
)

# === 2. Final Random Forest ===
clf = RandomForestClassifier(
    n_estimators=150,
    max_depth=6,
    max_features=15,
    criterion='gini',
    random_state=42
)

clf.fit(X_train, y_train)

# === 3. Predictions ===
y_train_pred = clf.predict(X_train)
y_test_pred = clf.predict(X_test)
y_train_prob = clf.predict_proba(X_train)[:, 1]
y_test_prob = clf.predict_proba(X_test)[:, 1]

# === 4. Evaluation ===
print("=== Train Metrics ===")
print("Accuracy:", accuracy_score(y_train, y_train_pred))
print("Precision:", precision_score(y_train, y_train_pred, zero_division=0))
print("Recall:", recall_score(y_train, y_train_pred, zero_division=0))
print("F1 Score:", f1_score(y_train, y_train_pred, zero_division=0))
print("AUC:", roc_auc_score(y_train, y_train_prob))

print("\n=== Test Metrics ===")
print("Accuracy:", accuracy_score(y_test, y_test_pred))
print("Precision:", precision_score(y_test, y_test_pred, zero_division=0))
print("Recall:", recall_score(y_test, y_test_pred, zero_division=0))
print("F1 Score:", f1_score(y_test, y_test_pred, zero_division=0))
print("AUC:", roc_auc_score(y_test, y_test_prob))

# === 5. Full classification report ===
print("\n=== Classification Report ===")
print(classification_report(y_test, y_test_pred))

[cell 86 markdown]
# **Observations**

* The model achieves a **test accuracy of approximately 80.1%**, slightly higher than the 79.9% obtained using the entropy-based version, with training accuracy closely aligned, indicating stable generalization.  
* The **test precision is 0.66** and **F1-score is 0.57**, indicating a balanced ability to identify churn cases while limiting false positives.  
* The **recall is 0.50**, which is marginally lower than the 0.52 observed in the entropy-based model. This suggests moderate sensitivity in detecting actual churners, which remains acceptable given the imbalanced nature of the dataset.  
* The **AUC score of 0.83** confirms that the model effectively differentiates between churn and non-churn customers based on predicted probabilities.

[cell 87 markdown]
# **Final Comparison of Decision Tree and Random Forest**



## Final Model Comparison on Test Set

| Model Type          | Criterion | Accuracy | Precision | Recall | F1 Score | AUC     |
|---------------------|-----------|----------|-----------|--------|----------|---------|
| Decision Tree (A)      | Entropy   | 0.7846   | 0.6126    | 0.5160 | 0.5602   | 0.8122  |
| Decision Tree (B)       | Gini      | 0.7846   | 0.6263    | 0.4705 | 0.5374   | 0.8128  |
| Random Forest (C)      | Entropy   | 0.7988   | 0.6542    | 0.5160 | 0.5769   | 0.8291  |
| Random Forest (D)       | Gini      | 0.8009   | 0.6666    | 0.5026 | 0.5732   | 0.8300  |

[cell 88 markdown]
## **Observations**

* Both Random Forest models outperform Decision Trees across all key metrics, especially in precision, F1-score, and AUC.  
* The Gini-based Random Forest achieves the **highest test accuracy (80.1%)** and **precision (0.67)**, while the entropy-based version has slightly better recall.  
* Overall, Random Forest with either criterion provides more balanced and robust performance, making it a better choice for churn prediction on this dataset.