# Coding%20Assignment%20on%20Supervised%20Learning

course: Module 2 — Machine Learning Algorithms
module: Module-2-Machine-Learning-Algorithms
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-2-Machine-Learning-Algorithms/General/Ungraded_Coding_Assignment_on_Supervised_Learning/Coding%20Assignment%20on%20Supervised%20Learning.pdf
pages: 3

---
[page 1]
Assignment: Binary Classification Using Supervised Learning 
Dataset: PIMA Indians Diabetes Dataset, freely available online 
 
Objective 
To apply and compare Logistic Regression, k-Nearest Neighbors (kNN), and Decision Tree classifiers 
on a real-world healthcare dataset containing numeric attributes and missing values, and evaluate 
their performance using standard classification metrics. 
 
Dataset Description 
The PIMA Indians Diabetes Dataset contains diagnostic measurements of female patients of Pima 
Indian heritage to predict the presence of diabetes. 
• Target Variable: 
Outcome 
o 0 → No diabetes 
o 1 → Diabetes 
• Features (All Numeric): 
Feature Description 
Pregnancies Number of pregnancies 
Glucose Plasma glucose concentration 
BloodPressure Diastolic blood pressure 
SkinThickness Triceps skin fold thickness 
Insulin 2-hour serum insulin 
BMI Body Mass Index 
DiabetesPedigreeFunction Genetic risk 
Age Age in years 
Missing Values: 
In this dataset, missing medical values are encoded as 0 in columns such as Glucose, BloodPressure, 
SkinThickness, Insulin, and BMI. 
 
Question Part A – Exploratory Data Analysis (EDA) 
Perform the following EDA steps:

[page 2]
1. Display dataset shape and column names 
2. Identify missing values 
3. Summary statistics (mean, std, min, max) 
4. Class distribution (Outcome 0 vs 1) 
5. Visualizations: 
o Histograms of numeric features 
o Boxplots to detect outliers 
o Correlation heatmap 
 
Question Part B – Data Preprocessing 
1. Identify and Treat missing values 
2.  Feature scaling may be done say by using StandardScaler 
3. Perform Train–test split by deciding the best spilt
 
Question Part C – Model Building 
Train the following supervised learning models: 
1. Logistic Regression 
2. k-Nearest Neighbors (Identify optimal k) 
3. Decision Tree Classifier 
 
Question Part D – Model Evaluation 
Evaluate each model using confusion matrix: 
• Accuracy 
• Precision 
• Recall 
• F1-Score 
• ROC-AUC and  
• Create a comparative performance table. 
 
Question Part E – Analysis Questions 
1. Which model performs best overall?

[page 3]
2. Why is recall important in diabetes prediction? 
3. Which model is most interpretable? 
4. Which model is most sensitive to feature scaling?