# DT RF vx1a
course: Module 2 — Machine Learning Algorithms
module: Module-2-Machine-Learning-Algorithms
type: notebook
source_url: https://personal-learn.armco.dev/files/Module-2-Machine-Learning-Algorithms/General/5.3_Lab_Materials/DT_RF_vx1a.ipynb
---
[cell 1 markdown]
# Telco Customer Churn Dataset Overview
The **Telco Customer Churn** dataset contains information about customers of a fictional telecom company. The objective is to predict whether a customer will **churn (leave the service)** based on their demographic, service usage, and account details.
## Source:
Originally available as part of IBM sample datasets, the **Telco Customer Churn** dataset is now publicly hosted on Kaggle and GitHub.
-[Kaggle Dataset Page](https://www.kaggle.com/datasets/blastchar/telco-customer-churn)
-[GitHub CSV File](https://github.com/treselle-systems/customer_churn_analysis/blob/master/WA_Fn-UseC_-Telco-Customer-Churn.csv)
# Target Variable:
- **Churn** *(binary)*: Whether the customer has churned (`Yes` or `No`)
---
### Features:
| Column Name | Description | Data Type |
|---------------------|-----------------------------------------------------------|--------------|
| customerID | Unique customer identifier | *String* |
| gender | Gender (`Male`, `Female`) | *Categorical*|
| SeniorCitizen | Whether the customer is a senior citizen (`0`, `1`) | *Categorical*|
| Partner | Has a partner (`Yes`, `No`) | *Categorical*|
| Dependents | Has dependents (`Yes`, `No`) | *Categorical*|
| tenure | Number of months the customer has stayed | *Numeric (int)*|
| PhoneService | Has phone service (`Yes`, `No`) | *Categorical*|
| MultipleLines | Has multiple lines (`Yes`, `No`, `No phone service`) | *Categorical*|
| InternetService | Type of internet service (`DSL`, `Fiber optic`, `No`) | *Categorical*|
| OnlineSecurity | Has online security (`Yes`, `No`, `No internet service`) | *Categorical*|
| OnlineBackup | Has online backup (`Yes`, `No`, `No internet service`) | *Categorical*|
| DeviceProtection | Has device protection (`Yes`, `No`, `No internet service`)| *Categorical*|
| TechSupport | Has technical support (`Yes`, `No`, `No internet service`)| *Categorical*|
| StreamingTV | Streams TV (`Yes`, `No`, `No internet service`) | *Categorical*|
| StreamingMovies | Streams movies (`Yes`, `No`, `No internet service`) | *Categorical*|
| Contract | Type of contract (`Month-to-month`, `One year`, `Two year`)| *Categorical*|
| PaperlessBilling | Uses paperless billing (`Yes`, `No`) | *Categorical*|
| PaymentMethod | Payment method (`Electronic check`,`Mailed Check`,`bank Transfer`,`Credit Card`) | *Categorical*|
| MonthlyCharges | The amount charged per month | *Numeric (float),*|
| TotalCharges | Total amount charged | *Numeric (float, needs cleaning)*|
| Churn | Whether the customer left | **Categorical (target)**|
---
[cell 2 markdown]
# **Python Packages Used**
### 1. `pandas`
- **Purpose**: Loading, cleaning, and manipulating the dataset
- **Import**: `import pandas as pd`
- **Install**: `pip install pandas`
---
### 2. `matplotlib`
- **Purpose**: Visualizing, plotting graphs and decision trees
- **Import**: `import matplotlib.pyplot as plt`
- **Install**: `pip install matplotlib`
---
### 3. `scikit-learn` (`sklearn`)
- **Purpose**: Data preprocessing, model building, evaluation
- **Install**: `pip install scikit-learn`
#### **Modules used:**
- `sklearn.preprocessing`
- `KBinsDiscretizer` – discretizing numeric features
- `sklearn.model_selection`
- `train_test_split` – splitting dataset into train/test sets
- `sklearn.tree`
- `DecisionTreeClassifier` – training decision trees
- `plot_tree` – visualizing trained decision trees
- `sklearn.ensemble`
- `RandomForestClassifier` – training random forest models
- `sklearn.metrics`
- `accuracy_score` – evaluating accuracy
- `precision_score` – evaluating precision
- `recall_score` – evaluating recall
- `f1_score` – evaluating F1 score
- `roc_auc_score` – computing AUC for classification
- `classification_report` – detailed per-class evaluation
---
### 4. `numpy`
- **Purpose**: Numerical operations in plotting and evaluation
- **Import**: `import numpy as np`
- **Install**: `pip install numpy`
[cell 3 code]
# Install packages at once
# !pip install pandas matplotlib scikit-learn numpy
[cell 4 markdown]
# **Loading the dataset into a datafarme**
[cell 5 code]
import pandas as pd
# Load the Telco Customer Churn dataset
df = pd.read_csv('WA_Fn-UseC_-Telco-Customer-Churnvx1.csv')
df
[cell 6 markdown]
# **Column names with their data typs**
[cell 7 code]
df.info()
[cell 8 markdown]
### **Observation 1**
* The column `TotalCharges` appears as `object` in `df.info()` instead of numeric. This is due to the presence of rows with blank string entries (`""`), which pandas does not treat as missing by default.
[cell 9 code]
print(f'No of samples where TotalCharges is Blank:{df[df["TotalCharges"].str.strip() == ""].shape[0]}')
# Show rows where TotalCharges is blank
df[df["TotalCharges"].str.strip() == ""][["customerID", "tenure", "TotalCharges"]]
[cell 10 code]
# convert the column to `float64` and removes invalid rows.
df["TotalCharges"] = pd.to_numeric(df["TotalCharges"], errors="coerce")
df.dropna(subset=["TotalCharges"], inplace=True)
df.shape
[cell 11 markdown]
### **Observation 2**
* `SeniorCitizen` is stored as `int64`, it is a binary categorical feature (0 = No, 1 = Yes). To reflect its actual meaning, we convert it to categorical type
[cell 12 code]
df["SeniorCitizen"] = df["SeniorCitizen"].astype("object")
[cell 13 code]
categorical_cols = df.select_dtypes(include=['object']).columns.tolist()
numerical_cols = df.select_dtypes(include=['int64', 'float64']).columns.tolist()
print("Categorical Columns:")
print(categorical_cols, len(categorical_cols))
print("\nNumerical Columns:")
print(numerical_cols, len(numerical_cols))
[cell 14 markdown]
# **Merging "No internet/phone service" with "No"**
* In this dataset, several categorical columns contain values like `"No internet service"` or `"No phone service"` in addition to a regular `"No"`.
* These extra values do not carry distinct meaning from `"No"` in terms of behavior — they simply indicate that the person does not have the service, which already implies `"No"` for all related options.
[cell 15 code]
cols_to_check = [
"MultipleLines", "OnlineSecurity", "OnlineBackup",
"DeviceProtection", "TechSupport", "StreamingTV", "StreamingMovies"
]
for col in cols_to_check:
print(f"\nValue counts for '{col}':")
print(df[col].value_counts())
[cell 16 markdown]
To simplify the dataset and avoid redundant categories during encoding, we merged:
- `"No internet service"` → `"No"`
- `"No phone service"` → `"No"`
[cell 17 code]
# Merge "No internet service" / "No phone service" with "No"
internet_cols = [
"OnlineSecurity", "OnlineBackup", "DeviceProtection",
"TechSupport", "StreamingTV", "StreamingMovies"
]
for col in internet_cols:
df[col] = df[col].replace("No internet service", "No")
df["MultipleLines"] = df["MultipleLines"].replace("No phone service", "No")
# Step 3: Print value counts after merging
print("\nAfter merging:\n")
for col in cols_to_check:
print(f"\nValue counts for '{col}':")
print(df[col].value_counts())
[cell 18 code]
df.info()
[cell 19 markdown]
# **Heatmap for categorical attributes**
* We used Cramér’s V to measure the strength of association between categorical features and visualized them using a heatmap.
* Cramér’s V is a recommended metric for measuring association between categorical attributes, especially when variables are nominal and do not have a natural order.
* Cramér’s V ranges from 0 to 1, where 0 indicates no association and 1 indicates a strong association between categorical variables.
[cell 20 code]
import pandas as pd
import numpy as np
import seaborn as sns
import matplotlib.pyplot as plt
from scipy.stats import chi2_contingency
import warnings
def cramers_v_basic(x, y):
confusion_matrix = pd.crosstab(x, y)
if confusion_matrix.shape[0] < 2 or confusion_matrix.shape[1] < 2:
return np.nan # not enough variability to compute
chi2 = chi2_contingency(confusion_matrix)[0]
n = confusion_matrix.sum().sum()
return np.sqrt(chi2 / (n * (min(confusion_matrix.shape) - 1 + 1e-10)))
cols = [col for col in categorical_cols if col != 'Churn'] + ['Churn']
cramer_matrix = pd.DataFrame(np.zeros((len(cols), len(cols))), index=cols, columns=cols)
# Suppress warnings during computation
with warnings.catch_warnings():
warnings.simplefilter("ignore")
for col1 in cols:
for col2 in cols:
if col1 == col2:
cramer_matrix.loc[col1, col2] = 1.0
else:
cramer_matrix.loc[col1, col2] = cramers_v_basic(df[col1], df[col2])
# Plot the heatmap
plt.figure(figsize=(12, 10))
sns.heatmap(cramer_matrix, cmap="YlGnBu", annot=True, fmt=".2f")
plt.title("Cramér’s V Heatmap of Categorical Variables (Warnings Suppressed)")
plt.tight_layout()
plt.show()
[cell 21 markdown]
## **Observations**
* **Contract type** has one of the strongest associations with Churn (~0.41), meaning contract duration is an important driver of customer churn.
* **InternetService** and related services (OnlineSecurity, TechSupport, StreamingTV/Movies) show moderate inter-correlations (0.3–0.5), indicating these features are interconnected.
* Most other categorical variables have very low correlations (less than 0.1), suggesting minimal dependency and low redundancy among them.
[cell 22 markdown]
# **Pearson Correlation Heatmap: Numerical Features vs. Churn**
[cell 23 code]
import seaborn as sns
import matplotlib.pyplot as plt
df["Churn"] = df["Churn"].map({"Yes": 1, "No": 0}) if df["Churn"].dtype == object else df["Churn"]
cols = numerical_cols + ["Churn"]
corr_matrix = df[cols].corr()
plt.figure(figsize=(8, 6))
sns.heatmap(corr_matrix, annot=True, cmap="coolwarm", fmt=".2f")
plt.title("Correlation Heatmap matrix Numerical Features and Churn")
plt.tight_layout()
plt.show()
[cell 24 markdown]
* **Tenure** has the strongest relationship with churn (≈ −0.35), indicating that customers with shorter tenure are more likely to churn.
[cell 25 markdown]
# **Discretizing numeric attributes**
[cell 26 markdown]
We turned the numeric features `tenure`, `MonthlyCharges`, and `TotalCharges` into **4 categories (bins)** each, using a method called **equi-frequency binning**.
* **Equi-frequency binning** means we split the data so that **each bin has about the same number of data points**.
* To do this, we used `KBinsDiscretizer` from scikit-learn with the **`quantile` strategy**, which automatically creates bins with equal numbers of samples.
* We set the number of bins to **4**, so each bin holds roughly **25% of the data**.
* The new columns (`tenure_bin`, `MonthlyCharges_bin`, and `TotalCharges_bin`) contain values from **0 to 3**, showing which bin each value falls into.
---
[cell 27 code]
# resetting churn to yes and no for consistency
df["Churn"] = df["Churn"].map({1: "Yes", 0: "No"}).astype("object")
df[['tenure','MonthlyCharges','TotalCharges']]
[cell 28 code]
from sklearn.preprocessing import KBinsDiscretizer
# Discretize each column separately
disc_tenure = KBinsDiscretizer(n_bins=4, encode='ordinal', strategy='quantile')
df["tenure_bin"] = disc_tenure.fit_transform(df[["tenure"]])
disc_month = KBinsDiscretizer(n_bins=4, encode='ordinal', strategy='quantile')
df["MonthlyCharges_bin"] = disc_month.fit_transform(df[["MonthlyCharges"]])
disc_total = KBinsDiscretizer(n_bins=4, encode='ordinal', strategy='quantile')
df["TotalCharges_bin"] = disc_total.fit_transform(df[["TotalCharges"]])
[cell 29 code]
df[['tenure_bin','MonthlyCharges_bin','TotalCharges_bin']]
[cell 30 markdown]
## **Replacing numeric attributes with their discretized versions and storing in a new dataframe**
[cell 31 code]
# Make a copy of original df
df2 = df.copy()
df2 = df2.drop(columns=["customerID"])
# Drop original numeric features
df2=df2.drop(['tenure', 'MonthlyCharges', 'TotalCharges'], axis=1)
df2['tenure_bin'] = df2['tenure_bin'].astype('object')
df2['MonthlyCharges_bin'] = df2['MonthlyCharges_bin'].astype('object')
df2['TotalCharges_bin'] = df2['TotalCharges_bin'].astype('object')
# df2 is now ready for modeling
df2.head()
[cell 32 code]
df2.info()
[cell 33 markdown]
# **Count plot for target variable shows label imbalance**
[cell 34 code]
print(df2['Churn'].value_counts(normalize=True))
df2['Churn'].value_counts(normalize=True).mul(100).plot(kind='bar')
[cell 35 markdown]
# Preparing Features and Target Variable
* The target column `Churn` is first converted from `"Yes"/"No"` to `1/0` for binary classification.
* It is then separated from the dataset to form `y` (labels), while the remaining columns become `X` (features).
* Finally, `X` is one-hot encoded using `pd.get_dummies()` to handle categorical variables.
[cell 36 code]
# Convert Churn to numeric target
df2["Churn"] = df2["Churn"].str.strip().map({"Yes": 1, "No": 0})
# Split into X and y
X = df2.drop("Churn", axis=1)
y = df2["Churn"]
# One-hot encode X only
X_encoded = pd.get_dummies(X, drop_first=True)
[cell 37 code]
X
[cell 38 code]
X['InternetService'].value_counts()
[cell 39 code]
X_encoded
[cell 40 markdown]
# **Splitting the dataset into train-test sets**
[cell 41 code]
from sklearn.model_selection import train_test_split
#Train-test split
X_train, X_test, y_train, y_test = train_test_split(X_encoded, y, test_size=0.1, random_state=42,stratify=y)
# Final shapes:
print("Train shape:", X_train.shape)
print("Test shape:", X_test.shape)
[cell 42 markdown]
# **Part 1: Decision Tree**
[cell 43 markdown]
# **Selecting Optimal Tree Depth for Decision Tree Classifier**
[cell 44 code]
import matplotlib.pyplot as plt
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import accuracy_score
train_acc = []
test_acc = []
depth_range = list(range(1, 30)) # test depths from 1 to 29
for depth in depth_range:
clf = DecisionTreeClassifier(criterion="entropy", max_depth=depth, random_state=42)
clf.fit(X_train, y_train)
train_pred = clf.predict(X_train)
val_pred = clf.predict(X_test)
train_acc.append(accuracy_score(y_train, train_pred))
test_acc.append(accuracy_score(y_test, val_pred))
# Plotting
plt.figure(figsize=(10, 6))
plt.plot(depth_range, train_acc, label="Train Accuracy", marker='o')
plt.plot(depth_range, test_acc, label="Test Accuracy", marker='s')
plt.xlabel("Tree Depth")
plt.ylabel("Accuracy")
plt.title("Overfitting vs Underfitting (Criterion = Entropy)")
plt.legend()
plt.grid(True)
plt.show()
[cell 45 code]
df_test_acc = pd.DataFrame({"Test Accuracy": test_acc[:15]}, index=range(1, 16))
df_test_acc.T
[cell 46 markdown]
* The best tree depth is 6, where test accuracy peaks before overfitting begins to increase.
[cell 47 markdown]
# **Impact of Train-Test Ratio on Decision Tree Performance**
[cell 48 code]
from sklearn.tree import DecisionTreeClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, roc_auc_score
import matplotlib.pyplot as plt
import numpy as np
# Set the best depth found from previous step
best_depth = 6 #
split_ratios = [0.6, 0.7, 0.8, 0.9] # train split ratios
metrics = {
"Train Accuracy": [],
"Test Accuracy": [],
"Train Precision": [],
"Test Precision": [],
"Train Recall": [],
"Test Recall": [],
"Train F1": [],
"Test F1": [],
"Train AUC": [],
"Test AUC": []
}
for ratio in split_ratios:
X_train, X_test, y_train, y_test = train_test_split(
X_encoded, y, train_size=ratio, random_state=42, stratify=y
)
clf = DecisionTreeClassifier(max_depth=best_depth, random_state=42,criterion="entropy")
clf.fit(X_train, y_train)
y_train_pred = clf.predict(X_train)
y_test_pred = clf.predict(X_test)
# For AUC, we need probabilities and binary classification
y_train_prob = clf.predict_proba(X_train)[:, 1]
y_test_prob = clf.predict_proba(X_test)[:, 1]
metrics["Train Accuracy"].append(accuracy_score(y_train, y_train_pred))
metrics["Test Accuracy"].append(accuracy_score(y_test, y_test_pred))
metrics["Train Precision"].append(precision_score(y_train, y_train_pred, zero_division=0))
metrics["Test Precision"].append(precision_score(y_test, y_test_pred, zero_division=0))
metrics["Train Recall"].append(recall_score(y_train, y_train_pred, zero_division=0))
metrics["Test Recall"].append(recall_score(y_test, y_test_pred, zero_division=0))
metrics["Train F1"].append(f1_score(y_train, y_train_pred, zero_division=0))
metrics["Test F1"].append(f1_score(y_test, y_test_pred, zero_division=0))
metrics["Train AUC"].append(roc_auc_score(y_train, y_train_prob))
metrics["Test AUC"].append(roc_auc_score(y_test, y_test_prob))
[cell 49 code]
import matplotlib.pyplot as plt
import numpy as np
def plot_all_metrics_in_grid(metrics, split_ratios):
metric_names = ["Accuracy", "Precision", "Recall", "F1", "AUC"]
n_metrics = len(metric_names)
fig, axes = plt.subplots(2, 3, figsize=(14, 8)) # 2 rows, 3 columns
axes = axes.flatten()
for idx, metric_name in enumerate(metric_names):
ax = axes[idx]
ax.plot(np.array(split_ratios)*100, metrics[f"Train {metric_name}"], marker='o', label='Train')
ax.plot(np.array(split_ratios)*100, metrics[f"Test {metric_name}"], marker='s', label='Test')
ax.set_title(f"{metric_name}")
ax.set_xlabel("Train Size (%)")
ax.set_ylabel(metric_name)
ax.grid(True)
ax.legend()
# Hide the extra subplot (6th one)
if n_metrics < len(axes):
axes[-1].axis("off")
fig.suptitle("Model Performance vs Train/Test Split Ratio", fontsize=16)
plt.tight_layout(rect=[0, 0, 1, 0.95]) # Adjust space for suptitle
plt.show()
# Call the function
plot_all_metrics_in_grid(metrics, split_ratios)
[cell 50 markdown]
## **Observations**
**Accuracy:**
Test accuracy gradually improves as train size increases from 60% to 80 and a slight dip at 90%. The general trend is that the model benefits from more training data.
**Precision:**
There is a slight upward trend in test precision, highest at the 90/10 split. The model becomes more confident in its positive predictions with more training data.
**Recall:**
Test recall peaks around the 80/20 split and drops significantly at 90%, suggesting that higher train sizes may reduce the model’s ability to detect positive cases.
**F1 Score:**
F1 score is highest at 80/20, showing the best trade-off between precision and recall. Both 70/30 and 90/10 are slightly worse in terms of balance.
**AUC:**
AUC is highest around 60/40 and then declines with the lowest at 90/10. This suggests the model’s ability to distinguish between classes slightly weakens when the test set is too small.
## **Final Decision:**
The **80/20 split** offers the most balanced performance across metrics and we choose it for final our Decision Tree classifier.
[cell 51 markdown]
# **A :Final Decision Tree Training and Evaluation**
[cell 52 code]
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import (
accuracy_score, precision_score, recall_score,
f1_score, roc_auc_score, classification_report
)
# 1. Train-test split (80/20)
X_train, X_test, y_train, y_test = train_test_split(
X_encoded, y, train_size=0.8, random_state=42, stratify=y
)
# 2. Train Decision Tree (use best depth found earlier)
clf = DecisionTreeClassifier(max_depth=6, criterion="entropy", random_state=42)
clf.fit(X_train, y_train)
# 3. Predictions
y_train_pred = clf.predict(X_train)
y_test_pred = clf.predict(X_test)
# 4. Probabilities (for AUC)
y_train_prob = clf.predict_proba(X_train)[:, 1]
y_test_prob = clf.predict_proba(X_test)[:, 1]
# 5. Evaluation metrics
print("=== Train Metrics ===")
print("Accuracy:", accuracy_score(y_train, y_train_pred))
print("Precision:", precision_score(y_train, y_train_pred, zero_division=0))
print("Recall:", recall_score(y_train, y_train_pred, zero_division=0))
print("F1 Score:", f1_score(y_train, y_train_pred, zero_division=0))
print("AUC:", roc_auc_score(y_train, y_train_prob))
print("\n=== Test Metrics ===")
print("Accuracy:", accuracy_score(y_test, y_test_pred))
print("Precision:", precision_score(y_test, y_test_pred, zero_division=0))
print("Recall:", recall_score(y_test, y_test_pred, zero_division=0))
print("F1 Score:", f1_score(y_test, y_test_pred, zero_division=0))
print("AUC:", roc_auc_score(y_test, y_test_prob))
# 6. Classification report (detailed per-class)
print("\nClassification Report:\n", classification_report(y_test, y_test_pred))
[cell 53 markdown]
## **Observations**
* The model achieves a test accuracy of **78.46%**, which is close to the train accuracy, indicating no major overfitting.
* Precision and recall are lower for class `1` (the positive/minority class) due to label imbalance.
* The **AUC of 0.81** on the test set shows the model has good ability to separate churn vs non-churn.
* Class `0` (non-churn) is predicted more confidently, as seen by its higher precision and recall in the classification report.
* The model generalizes well, but could benefit from improving recall on the positive class.
[cell 54 markdown]
# **Visualizing the Trained Decision Tree**
[cell 55 code]
from sklearn.tree import plot_tree
plt.figure(figsize=(30, 10))
plot_tree(clf,
feature_names=X_train.columns,
class_names=[str(c) for c in clf.classes_],
filled=True,
rounded=True,
fontsize=10)
plt.show()
[cell 56 markdown]
## **Showing prediction for some test samples**
[cell 57 code]
# Get model predictions
y_test_pred = clf.predict(X_test)
# DataFrame combining test features, ground truth, and prediction
X_orig_train, X_orig_test = train_test_split(
X, train_size=0.8, random_state=42, stratify=y
)
# dataframe for inspection
inspect_df = X_orig_test.copy()
inspect_df["Ground Truth"] = y_test.values
inspect_df["Prediction"] = y_test_pred
# few readable examples
inspect_df[10:20]
[cell 58 markdown]
# **B: Decision Tree Training and Evaluation using gini as criterion**
[cell 59 code]
from sklearn.model_selection import train_test_split
from sklearn.tree import DecisionTreeClassifier
from sklearn.metrics import (
accuracy_score, precision_score, recall_score,
f1_score, roc_auc_score, classification_report
)
# 1. Train-test split (80/20)
X_train, X_test, y_train, y_test = train_test_split(
X_encoded, y, train_size=0.8, random_state=42, stratify=y
)
# 2. Train Decision Tree (use best depth found earlier)
clf = DecisionTreeClassifier(max_depth=6, criterion="gini", random_state=42)
clf.fit(X_train, y_train)
# 3. Predictions
y_train_pred = clf.predict(X_train)
y_test_pred = clf.predict(X_test)
# 4. Probabilities (for AUC)
y_train_prob = clf.predict_proba(X_train)[:, 1]
y_test_prob = clf.predict_proba(X_test)[:, 1]
# 5. Evaluation metrics
print("=== Train Metrics ===")
print("Accuracy:", accuracy_score(y_train, y_train_pred))
print("Precision:", precision_score(y_train, y_train_pred, zero_division=0))
print("Recall:", recall_score(y_train, y_train_pred, zero_division=0))
print("F1 Score:", f1_score(y_train, y_train_pred, zero_division=0))
print("AUC:", roc_auc_score(y_train, y_train_prob))
print("\n=== Test Metrics ===")
print("Accuracy:", accuracy_score(y_test, y_test_pred))
print("Precision:", precision_score(y_test, y_test_pred, zero_division=0))
print("Recall:", recall_score(y_test, y_test_pred, zero_division=0))
print("F1 Score:", f1_score(y_test, y_test_pred, zero_division=0))
print("AUC:", roc_auc_score(y_test, y_test_prob))
# 6. Classification report (detailed per-class)
print("\nClassification Report:\n", classification_report(y_test, y_test_pred))
[cell 60 markdown]
## **Observations**
* The model achieves a test accuracy of **78.46%**, similar to the entropy-based version.
* Test precision improves slightly (to **0.62**), but **recall drops** to **0.47**, indicating more false negatives (missed churns).
* AUC remains solid at **0.81**, confirming good class separation overall.
* Class `0` is still predicted strongly, but the model shows more hesitation on class `1`, as seen from the lower recall.
* The gap between **train and test F1-scores** is larger compared to the entropy model, indicating slightly overfitting with Gini.
[cell 61 code]
# Plot tree
plt.figure(figsize=(30, 10))
plot_tree(clf,
feature_names=X_train.columns,
class_names=[str(c) for c in clf.classes_],
filled=True,
rounded=True,
fontsize=10)
plt.show()
[cell 62 markdown]
# **Another way to find optimal depth**
[cell 63 markdown]
* In the previous approach, we selected the depth once using a 90/10 split and used it throughout the code to decide the optimal train/test ratio and perform the final model evaluation.
* In this part, we find the best depth separately at each train-test split, and then use the corresponding depth and split to train the model and perform evaluation.
[cell 64 markdown]
## **For 60/40 train-test split**
[cell 65 code]
from sklearn.model_selection import train_test_split
#Train-test split
print('For 60/40 train-test split')
train_size=0.6
X_train, X_test, y_train, y_test = train_test_split(X_encoded, y, train_size=train_size, random_state=42,stratify=y)
train_acc = []
test_acc = []
depth_range = list(range(1, 30)) # test depths from 1 to 29
for depth in depth_range:
clf = DecisionTreeClassifier(criterion="entropy", max_depth=depth, random_state=42)
clf.fit(X_train, y_train)
train_pred = clf.predict(X_train)
val_pred = clf.predict(X_test)
train_acc.append(accuracy_score(y_train, train_pred))
test_acc.append(accuracy_score(y_test, val_pred))
# Plotting
plt.figure(figsize=(10, 6))
plt.plot(depth_range, train_acc, label="Train Accuracy", marker='o')
plt.plot(depth_range, test_acc, label="Test Accuracy", marker='s')
plt.xlabel("Tree Depth")
plt.ylabel("Accuracy")
plt.title("Overfitting vs Underfitting (Criterion = Entropy)")
plt.legend()
plt.grid(True)
plt.show()
df_test_acc = pd.DataFrame({"Test Accuracy": test_acc[:15]}, index=range(1, 16))
depth_acc_dict = {depth: round(acc, 4) for depth, acc in zip(df_test_acc.index, df_test_acc["Test Accuracy"])}
print('Value of accuracy at first 15 depth sizes:')
print(depth_acc_dict)
# Find the depth with maximum accuracy
best_depth = max(depth_acc_dict, key=depth_acc_dict.get)
print(f'\nDepth with maximum accuracy: {best_depth} (Accuracy = {depth_acc_dict[best_depth]})')
# 1. Train-test split
X_train, X_test, y_train, y_test = train_test_split(
X_encoded, y, train_size=train_size, random_state=42, stratify=y
)
# 2. Train Decision Tree (use best depth found earlier)
clf = DecisionTreeClassifier(max_depth=best_depth, criterion="entropy", random_state=42)
clf.fit(X_train, y_train)
# 3. Predictions
y_train_pred = clf.predict(X_train)
y_test_pred = clf.predict(X_test)
# 4. Probabilities (for AUC)
y_train_prob = clf.predict_proba(X_train)[:, 1]
y_test_prob = clf.predict_proba(X_test)[:, 1]
# 5. Evaluation metrics
print("=== Train Metrics ===")
print("Accuracy:", accuracy_score(y_train, y_train_pred))
print("Precision:", precision_score(y_train, y_train_pred, zero_division=0))
print("Recall:", recall_score(y_train, y_train_pred, zero_division=0))
print("F1 Score:", f1_score(y_train, y_train_pred, zero_division=0))
print("AUC:", roc_auc_score(y_train, y_train_prob))
print("\n=== Test Metrics ===")
print("Accuracy:", accuracy_score(y_test, y_test_pred))
print("Precision:", precision_score(y_test, y_test_pred, zero_division=0))
print("Recall:", recall_score(y_test, y_test_pred, zero_division=0))
print("F1 Score:", f1_score(y_test, y_test_pred, zero_division=0))
print("AUC:", roc_auc_score(y_test, y_test_prob))
# 6. Classification report (detailed per-class)
print("\nClassification Report:\n", classification_report(y_test, y_test_pred))
[cell 66 markdown]
## **For 70/30 train-test split**
[cell 67 code]
from sklearn.model_selection import train_test_split
#Train-test split
print('For 70/30 train split')
train_size=0.7
X_train, X_test, y_train, y_test = train_test_split(X_encoded, y, train_size=train_size, random_state=42,stratify=y)
train_acc = []
test_acc = []
depth_range = list(range(1, 30)) # test depths from 1 to 29
for depth in depth_range:
clf = DecisionTreeClassifier(criterion="entropy", max_depth=depth, random_state=42)
clf.fit(X_train, y_train)
train_pred = clf.predict(X_train)
val_pred = clf.predict(X_test)
train_acc.append(accuracy_score(y_train, train_pred))
test_acc.append(accuracy_score(y_test, val_pred))
# Plotting
plt.figure(figsize=(10, 6))
plt.plot(depth_range, train_acc, label="Train Accuracy", marker='o')
plt.plot(depth_range, test_acc, label="Test Accuracy", marker='s')
plt.xlabel("Tree Depth")
plt.ylabel("Accuracy")
plt.title("Overfitting vs Underfitting (Criterion = Entropy)")
plt.legend()
plt.grid(True)
plt.show()
df_test_acc = pd.DataFrame({"Test Accuracy": test_acc[:15]}, index=range(1, 16))
depth_acc_dict = {depth: round(acc, 4) for depth, acc in zip(df_test_acc.index, df_test_acc["Test Accuracy"])}
print('Value of accuracy at first 15 depth sizes:')
print(depth_acc_dict)
# Find the depth with maximum accuracy
best_depth = max(depth_acc_dict, key=depth_acc_dict.get)
print(f'\nDepth with maximum accuracy: {best_depth} (Accuracy = {depth_acc_dict[best_depth]})')
# 1. Train-test split
X_train, X_test, y_train, y_test = train_test_split(
X_encoded, y, train_size=train_size, random_state=42, stratify=y
)
# 2. Train Decision Tree (use best depth found earlier)
clf = DecisionTreeClassifier(max_depth=best_depth, criterion="entropy", random_state=42)
clf.fit(X_train, y_train)
# 3. Predictions
y_train_pred = clf.predict(X_train)
y_test_pred = clf.predict(X_test)
# 4. Probabilities (for AUC)
y_train_prob = clf.predict_proba(X_train)[:, 1]
y_test_prob = clf.predict_proba(X_test)[:, 1]
# 5. Evaluation metrics
print("=== Train Metrics ===")
print("Accuracy:", accuracy_score(y_train, y_train_pred))
print("Precision:", precision_score(y_train, y_train_pred, zero_division=0))
print("Recall:", recall_score(y_train, y_train_pred, zero_division=0))
print("F1 Score:", f1_score(y_train, y_train_pred, zero_division=0))
print("AUC:", roc_auc_score(y_train, y_train_prob))
print("\n=== Test Metrics ===")
print("Accuracy:", accuracy_score(y_test, y_test_pred))
print("Precision:", precision_score(y_test, y_test_pred, zero_division=0))
print("Recall:", recall_score(y_test, y_test_pred, zero_division=0))
print("F1 Score:", f1_score(y_test, y_test_pred, zero_division=0))
print("AUC:", roc_auc_score(y_test, y_test_prob))
# 6. Classification report (detailed per-class)
print("\nClassification Report:\n", classification_report(y_test, y_test_pred))
[cell 68 markdown]
## **For 80/20 train-test split**
[cell 69 code]
from sklearn.model_selection import train_test_split
#Train-test split
print('For 80/20 train split')
train_size=0.8
X_train, X_test, y_train, y_test = train_test_split(X_encoded, y, train_size=train_size, random_state=42,stratify=y)
train_acc = []
test_acc = []
depth_range = list(range(1, 30)) # test depths from 1 to 29
for depth in depth_range:
clf = DecisionTreeClassifier(criterion="entropy", max_depth=depth, random_state=42)
clf.fit(X_train, y_train)
train_pred = clf.predict(X_train)
val_pred = clf.predict(X_test)
train_acc.append(accuracy_score(y_train, train_pred))
test_acc.append(accuracy_score(y_test, val_pred))
# Plotting
plt.figure(figsize=(10, 6))
plt.plot(depth_range, train_acc, label="Train Accuracy", marker='o')
plt.plot(depth_range, test_acc, label="Test Accuracy", marker='s')
plt.xlabel("Tree Depth")
plt.ylabel("Accuracy")
plt.title("Overfitting vs Underfitting (Criterion = Entropy)")
plt.legend()
plt.grid(True)
plt.show()
df_test_acc = pd.DataFrame({"Test Accuracy": test_acc[:15]}, index=range(1, 16))
depth_acc_dict = {depth: round(acc, 4) for depth, acc in zip(df_test_acc.index, df_test_acc["Test Accuracy"])}
print('Value of accuracy at first 15 depth sizes:')
print(depth_acc_dict)
# Find the depth with maximum accuracy
best_depth = max(depth_acc_dict, key=depth_acc_dict.get)
print(f'\nDepth with maximum accuracy: {best_depth} (Accuracy = {depth_acc_dict[best_depth]})')
# 1. Train-test split
X_train, X_test, y_train, y_test = train_test_split(
X_encoded, y, train_size=train_size, random_state=42, stratify=y
)
# 2. Train Decision Tree (use best depth found earlier)
clf = DecisionTreeClassifier(max_depth=best_depth, criterion="entropy", random_state=42)
clf.fit(X_train, y_train)
# 3. Predictions
y_train_pred = clf.predict(X_train)
y_test_pred = clf.predict(X_test)
# 4. Probabilities (for AUC)
y_train_prob = clf.predict_proba(X_train)[:, 1]
y_test_prob = clf.predict_proba(X_test)[:, 1]
# 5. Evaluation metrics
print("=== Train Metrics ===")
print("Accuracy:", accuracy_score(y_train, y_train_pred))
print("Precision:", precision_score(y_train, y_train_pred, zero_division=0))
print("Recall:", recall_score(y_train, y_train_pred, zero_division=0))
print("F1 Score:", f1_score(y_train, y_train_pred, zero_division=0))
print("AUC:", roc_auc_score(y_train, y_train_prob))
print("\n=== Test Metrics ===")
print("Accuracy:", accuracy_score(y_test, y_test_pred))
print("Precision:", precision_score(y_test, y_test_pred, zero_division=0))
print("Recall:", recall_score(y_test, y_test_pred, zero_division=0))
print("F1 Score:", f1_score(y_test, y_test_pred, zero_division=0))
print("AUC:", roc_auc_score(y_test, y_test_prob))
# 6. Classification report (detailed per-class)
print("\nClassification Report:\n", classification_report(y_test, y_test_pred))
[cell 70 markdown]
## **For 90/10 train-test split**
[cell 71 code]
from sklearn.model_selection import train_test_split
#Train-test split
print('For 90/10 train split')
train_size=0.6
X_train, X_test, y_train, y_test = train_test_split(X_encoded, y, train_size=train_size, random_state=42,stratify=y)
train_acc = []
test_acc = []
depth_range = list(range(1, 30)) # test depths from 1 to 29
for depth in depth_range:
clf = DecisionTreeClassifier(criterion="entropy", max_depth=depth, random_state=42)
clf.fit(X_train, y_train)
train_pred = clf.predict(X_train)
val_pred = clf.predict(X_test)
train_acc.append(accuracy_score(y_train, train_pred))
test_acc.append(accuracy_score(y_test, val_pred))
# Plotting
plt.figure(figsize=(10, 6))
plt.plot(depth_range, train_acc, label="Train Accuracy", marker='o')
plt.plot(depth_range, test_acc, label="Test Accuracy", marker='s')
plt.xlabel("Tree Depth")
plt.ylabel("Accuracy")
plt.title("Overfitting vs Underfitting (Criterion = Entropy)")
plt.legend()
plt.grid(True)
plt.show()
df_test_acc = pd.DataFrame({"Test Accuracy": test_acc[:15]}, index=range(1, 16))
depth_acc_dict = {depth: round(acc, 4) for depth, acc in zip(df_test_acc.index, df_test_acc["Test Accuracy"])}
print('Value of accuracy at first 15 depth sizes:')
print(depth_acc_dict)
# Find the depth with maximum accuracy
best_depth = max(depth_acc_dict, key=depth_acc_dict.get)
print(f'\nDepth with maximum accuracy: {best_depth} (Accuracy = {depth_acc_dict[best_depth]})')
# 1. Train-test split (80/20)
X_train, X_test, y_train, y_test = train_test_split(
X_encoded, y, train_size=train_size, random_state=42, stratify=y
)
# 2. Train Decision Tree (use best depth found earlier)
clf = DecisionTreeClassifier(max_depth=best_depth, criterion="entropy", random_state=42)
clf.fit(X_train, y_train)
# 3. Predictions
y_train_pred = clf.predict(X_train)
y_test_pred = clf.predict(X_test)
# 4. Probabilities (for AUC)
y_train_prob = clf.predict_proba(X_train)[:, 1]
y_test_prob = clf.predict_proba(X_test)[:, 1]
# 5. Evaluation metrics
print("=== Train Metrics ===")
print("Accuracy:", accuracy_score(y_train, y_train_pred))
print("Precision:", precision_score(y_train, y_train_pred, zero_division=0))
print("Recall:", recall_score(y_train, y_train_pred, zero_division=0))
print("F1 Score:", f1_score(y_train, y_train_pred, zero_division=0))
print("AUC:", roc_auc_score(y_train, y_train_prob))
print("\n=== Test Metrics ===")
print("Accuracy:", accuracy_score(y_test, y_test_pred))
print("Precision:", precision_score(y_test, y_test_pred, zero_division=0))
print("Recall:", recall_score(y_test, y_test_pred, zero_division=0))
print("F1 Score:", f1_score(y_test, y_test_pred, zero_division=0))
print("AUC:", roc_auc_score(y_test, y_test_prob))
# 6. Classification report (detailed per-class)
print("\nClassification Report:\n", classification_report(y_test, y_test_pred))
[cell 72 markdown]
# **Part 2: Random Forest**
There are two important hyperparameters in Random Forest — `n_estimators` (number of trees) and `max_features` (number of features considered at each split)
[cell 73 markdown]
# **Selecting optimal n_estimators (number of trees) in Random Forest**
* We evaluated the following range: [50, 75, 100, 150]
[cell 74 code]
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import (
accuracy_score, precision_score, recall_score,
f1_score, roc_auc_score
)
from sklearn.model_selection import train_test_split
import matplotlib.pyplot as plt
import numpy as np
# 1. Train/test split (80/20)
X_train, X_test, y_train, y_test = train_test_split(
X_encoded, y, train_size=0.8, random_state=42, stratify=y
)
# 2. Varying n_estimators
estimator_range = [50, 75, 100, 150]
metrics_rf = {
"Train Accuracy": [], "Test Accuracy": [],
"Train Precision": [], "Test Precision": [],
"Train Recall": [], "Test Recall": [],
"Train F1": [], "Test F1": [],
"Train AUC": [], "Test AUC": []
}
for n in estimator_range:
clf = RandomForestClassifier(n_estimators=n, criterion='entropy', max_depth=6, random_state=42)
clf.fit(X_train, y_train)
y_train_pred = clf.predict(X_train)
y_test_pred = clf.predict(X_test)
y_train_prob = clf.predict_proba(X_train)[:, 1]
y_test_prob = clf.predict_proba(X_test)[:, 1]
metrics_rf["Train Accuracy"].append(accuracy_score(y_train, y_train_pred))
metrics_rf["Test Accuracy"].append(accuracy_score(y_test, y_test_pred))
metrics_rf["Train Precision"].append(precision_score(y_train, y_train_pred, zero_division=0))
metrics_rf["Test Precision"].append(precision_score(y_test, y_test_pred, zero_division=0))
metrics_rf["Train Recall"].append(recall_score(y_train, y_train_pred, zero_division=0))
metrics_rf["Test Recall"].append(recall_score(y_test, y_test_pred, zero_division=0))
metrics_rf["Train F1"].append(f1_score(y_train, y_train_pred, zero_division=0))
metrics_rf["Test F1"].append(f1_score(y_test, y_test_pred, zero_division=0))
metrics_rf["Train AUC"].append(roc_auc_score(y_train, y_train_prob))
metrics_rf["Test AUC"].append(roc_auc_score(y_test, y_test_prob))
[cell 75 code]
def plot_rf_metrics_grid(metrics, estimator_range):
metric_names = ["Accuracy", "Precision", "Recall", "F1", "AUC"]
fig, axes = plt.subplots(2, 3, figsize=(15, 8))
axes = axes.flatten()
for idx, metric in enumerate(metric_names):
ax = axes[idx]
ax.plot(estimator_range, metrics[f"Train {metric}"], marker='o', label="Train")
ax.plot(estimator_range, metrics[f"Test {metric}"], marker='s', label="Test")
ax.set_title(metric)
ax.set_xlabel("n_estimators")
ax.set_ylabel(metric)
ax.grid(True)
ax.legend()
if len(metric_names) < len(axes):
axes[-1].axis("off") # Hide unused subplot
fig.suptitle("Random Forest Performance vs n_estimators", fontsize=16)
plt.tight_layout(rect=[0, 0, 1, 0.95])
plt.show()
# Call the plot
plot_rf_metrics_grid(metrics_rf, estimator_range)
[cell 76 markdown]
## **Observation**
* **Accuracy:** Both train and test accuracy show a dip at 100, followed by improvement in accuracy at `n_estimators = 150`.
* **Precision:** Train precision is consistent; test precision dips at 100 trees but recovers slightly by 150.
* **Recall:** Train and Test recall gradually improves as the number of estimators increases, peaking at 150.
* **F1 Score:** Test F1 is highest at `n_estimators = 150`, suggesting the best balance between precision and recall at that point.
* **AUC:** Test AUC is fairly flat, with a slight decline at 150 trees
* We use `n_estimators = 150` as it gives best overall test performance, especially in terms of recall and F1 score and comparable performance in other metrics
[cell 77 markdown]
# **Selecting optimal max_features(number of features) in Random Forest**
We evaluated the following range: [5, 10, 15,20,25, 29]
[cell 78 code]
import numpy as np
import matplotlib.pyplot as plt
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import (
accuracy_score, precision_score, recall_score,
f1_score, roc_auc_score
)
from sklearn.model_selection import train_test_split
# === PARAMETERS ===
n_total_features = X.shape[1]
n_estimators = 150
max_depth = 6
# === 1. Train/test split ===
X_train, X_test, y_train, y_test = train_test_split(
X_encoded, y, train_size=0.8, random_state=42, stratify=y
)
# === 2. Define max_features options ===
max_feature_range = [5, 10, 15,20,25, 29]
# === 3. Metrics Storage ===
metrics_rf = {
"Train Accuracy": [], "Test Accuracy": [],
"Train Precision": [], "Test Precision": [],
"Train Recall": [], "Test Recall": [],
"Train F1": [], "Test F1": [],
"Train AUC": [], "Test AUC": []
}
# === 4. Loop over max_features ===
for mf in max_feature_range:
clf = RandomForestClassifier(
n_estimators=n_estimators,
max_depth=max_depth,
max_features=mf,
criterion='entropy',
random_state=42
)
clf.fit(X_train, y_train)
y_train_pred = clf.predict(X_train)
y_test_pred = clf.predict(X_test)
y_train_prob = clf.predict_proba(X_train)[:, 1]
y_test_prob = clf.predict_proba(X_test)[:, 1]
metrics_rf["Train Accuracy"].append(accuracy_score(y_train, y_train_pred))
metrics_rf["Test Accuracy"].append(accuracy_score(y_test, y_test_pred))
metrics_rf["Train Precision"].append(precision_score(y_train, y_train_pred, zero_division=0))
metrics_rf["Test Precision"].append(precision_score(y_test, y_test_pred, zero_division=0))
metrics_rf["Train Recall"].append(recall_score(y_train, y_train_pred, zero_division=0))
metrics_rf["Test Recall"].append(recall_score(y_test, y_test_pred, zero_division=0))
metrics_rf["Train F1"].append(f1_score(y_train, y_train_pred, zero_division=0))
metrics_rf["Test F1"].append(f1_score(y_test, y_test_pred, zero_division=0))
metrics_rf["Train AUC"].append(roc_auc_score(y_train, y_train_prob))
metrics_rf["Test AUC"].append(roc_auc_score(y_test, y_test_prob))
[cell 79 code]
def plot_rf_metrics_grid(metrics, x_vals):
metric_names = ["Accuracy", "Precision", "Recall", "F1", "AUC"]
fig, axes = plt.subplots(2, 3, figsize=(15, 8))
axes = axes.flatten()
for idx, metric in enumerate(metric_names):
ax = axes[idx]
ax.plot(x_vals, metrics[f"Train {metric}"], marker='o', label="Train")
ax.plot(x_vals, metrics[f"Test {metric}"], marker='s', label="Test")
ax.set_title(metric)
ax.set_xlabel("max_features")
ax.set_ylabel(metric)
ax.set_xticks(x_vals)
ax.grid(True)
ax.legend()
# If extra subplot, hide it
if len(metric_names) < len(axes):
axes[-1].axis("off")
fig.suptitle("Random Forest Performance vs max_features", fontsize=16)
plt.tight_layout(rect=[0, 0, 1, 0.95])
plt.show()
plot_rf_metrics_grid(metrics_rf,max_feature_range )
[cell 80 markdown]
## Observations
* **Accuracy:** Test accuracy is highest at **`max_features = 29`**, indicating the model performs best when all features are considered at each split.
* **Precision:** Test precision increases steadily with `max_features`, improving after 20 and peaking at 29.
* **Recall:** Test recall peaks at `max_features = 15`, then slightly declines — suggesting fewer false negatives around this point.
* **F1 Score:** Test F1-score is highest at `max_features = 15`, showing the best precision-recall balance.
* **AUC:** Test AUC is slightly higher at `5–10`, then gradually drops as `max_features` increases, showing weaker class separation with more features.
We choose **`max_features = 15`**, as it gives the best trade-off between precision and recall, and achieves the highest F1-score, with a comparable performance in other metrics
[cell 81 markdown]
# **C:Final Random Forest training and Evaluation**
[cell 82 code]
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import (
accuracy_score, precision_score, recall_score,
f1_score, roc_auc_score, classification_report
)
# === 1. Train-test split (80/20) ===
X_train, X_test, y_train, y_test = train_test_split(
X_encoded, y, train_size=0.8, random_state=42, stratify=y
)
# === 2. Final Random Forest ===
clf = RandomForestClassifier(
n_estimators=150,
max_depth=6,
max_features=15,
criterion='entropy',
random_state=42
)
clf.fit(X_train, y_train)
# === 3. Predictions ===
y_train_pred = clf.predict(X_train)
y_test_pred = clf.predict(X_test)
y_train_prob = clf.predict_proba(X_train)[:, 1]
y_test_prob = clf.predict_proba(X_test)[:, 1]
# === 4. Evaluation ===
print("=== Train Metrics ===")
print("Accuracy:", accuracy_score(y_train, y_train_pred))
print("Precision:", precision_score(y_train, y_train_pred, zero_division=0))
print("Recall:", recall_score(y_train, y_train_pred, zero_division=0))
print("F1 Score:", f1_score(y_train, y_train_pred, zero_division=0))
print("AUC:", roc_auc_score(y_train, y_train_prob))
print("\n=== Test Metrics ===")
print("Accuracy:", accuracy_score(y_test, y_test_pred))
print("Precision:", precision_score(y_test, y_test_pred, zero_division=0))
print("Recall:", recall_score(y_test, y_test_pred, zero_division=0))
print("F1 Score:", f1_score(y_test, y_test_pred, zero_division=0))
print("AUC:", roc_auc_score(y_test, y_test_prob))
# === 5. Full classification report ===
print("\n=== Classification Report ===")
print(classification_report(y_test, y_test_pred))
[cell 83 markdown]
# **Observations**
* The model achieves a **test accuracy of approximately 79.9%**, with minimal difference from training accuracy, indicating good generalization to unseen data.
* The **test precision (0.65)** and **F1-score (~0.58)** suggest that the model maintains a reasonable balance between correctly identifying churn cases and minimizing false positives.
* The **recall of ~0.52** indicates moderate sensitivity to detecting actual churners, which is important in the context of imbalanced datasets.
* The **AUC score of ~0.83** confirms that the model has a strong ability to distinguish between the two classes based on predicted probabilities.
[cell 84 markdown]
# **D :Random Forest training and Evaluation using gini as criterion**
[cell 85 code]
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import (
accuracy_score, precision_score, recall_score,
f1_score, roc_auc_score, classification_report
)
# === 1. Train-test split (80/20) ===
X_train, X_test, y_train, y_test = train_test_split(
X_encoded, y, train_size=0.8, random_state=42, stratify=y
)
# === 2. Final Random Forest ===
clf = RandomForestClassifier(
n_estimators=150,
max_depth=6,
max_features=15,
criterion='gini',
random_state=42
)
clf.fit(X_train, y_train)
# === 3. Predictions ===
y_train_pred = clf.predict(X_train)
y_test_pred = clf.predict(X_test)
y_train_prob = clf.predict_proba(X_train)[:, 1]
y_test_prob = clf.predict_proba(X_test)[:, 1]
# === 4. Evaluation ===
print("=== Train Metrics ===")
print("Accuracy:", accuracy_score(y_train, y_train_pred))
print("Precision:", precision_score(y_train, y_train_pred, zero_division=0))
print("Recall:", recall_score(y_train, y_train_pred, zero_division=0))
print("F1 Score:", f1_score(y_train, y_train_pred, zero_division=0))
print("AUC:", roc_auc_score(y_train, y_train_prob))
print("\n=== Test Metrics ===")
print("Accuracy:", accuracy_score(y_test, y_test_pred))
print("Precision:", precision_score(y_test, y_test_pred, zero_division=0))
print("Recall:", recall_score(y_test, y_test_pred, zero_division=0))
print("F1 Score:", f1_score(y_test, y_test_pred, zero_division=0))
print("AUC:", roc_auc_score(y_test, y_test_prob))
# === 5. Full classification report ===
print("\n=== Classification Report ===")
print(classification_report(y_test, y_test_pred))
[cell 86 markdown]
# **Observations**
* The model achieves a **test accuracy of approximately 80.1%**, slightly higher than the 79.9% obtained using the entropy-based version, with training accuracy closely aligned, indicating stable generalization.
* The **test precision is 0.66** and **F1-score is 0.57**, indicating a balanced ability to identify churn cases while limiting false positives.
* The **recall is 0.50**, which is marginally lower than the 0.52 observed in the entropy-based model. This suggests moderate sensitivity in detecting actual churners, which remains acceptable given the imbalanced nature of the dataset.
* The **AUC score of 0.83** confirms that the model effectively differentiates between churn and non-churn customers based on predicted probabilities.
[cell 87 markdown]
# **Final Comparison of Decision Tree and Random Forest**
## Final Model Comparison on Test Set
| Model Type | Criterion | Accuracy | Precision | Recall | F1 Score | AUC |
|---------------------|-----------|----------|-----------|--------|----------|---------|
| Decision Tree (A) | Entropy | 0.7846 | 0.6126 | 0.5160 | 0.5602 | 0.8122 |
| Decision Tree (B) | Gini | 0.7846 | 0.6263 | 0.4705 | 0.5374 | 0.8128 |
| Random Forest (C) | Entropy | 0.7988 | 0.6542 | 0.5160 | 0.5769 | 0.8291 |
| Random Forest (D) | Gini | 0.8009 | 0.6666 | 0.5026 | 0.5732 | 0.8300 |
[cell 88 markdown]
## **Observations**
* Both Random Forest models outperform Decision Trees across all key metrics, especially in precision, F1-score, and AUC.
* The Gini-based Random Forest achieves the **highest test accuracy (80.1%)** and **precision (0.67)**, while the entropy-based version has slightly better recall.
* Overall, Random Forest with either criterion provides more balanced and robust performance, making it a better choice for churn prediction on this dataset.