# Coding%20Assignment%20on%20Supervised%20Learning course: Module 2 — Machine Learning Algorithms module: Module-2-Machine-Learning-Algorithms type: pdf source_url: https://personal-learn.armco.dev/files/Module-2-Machine-Learning-Algorithms/General/Ungraded_Coding_Assignment_on_Supervised_Learning/Coding%20Assignment%20on%20Supervised%20Learning.pdf pages: 3 --- [page 1] Assignment: Binary Classification Using Supervised Learning Dataset: PIMA Indians Diabetes Dataset, freely available online Objective To apply and compare Logistic Regression, k-Nearest Neighbors (kNN), and Decision Tree classifiers on a real-world healthcare dataset containing numeric attributes and missing values, and evaluate their performance using standard classification metrics. Dataset Description The PIMA Indians Diabetes Dataset contains diagnostic measurements of female patients of Pima Indian heritage to predict the presence of diabetes. • Target Variable: Outcome o 0 → No diabetes o 1 → Diabetes • Features (All Numeric): Feature Description Pregnancies Number of pregnancies Glucose Plasma glucose concentration BloodPressure Diastolic blood pressure SkinThickness Triceps skin fold thickness Insulin 2-hour serum insulin BMI Body Mass Index DiabetesPedigreeFunction Genetic risk Age Age in years Missing Values: In this dataset, missing medical values are encoded as 0 in columns such as Glucose, BloodPressure, SkinThickness, Insulin, and BMI. Question Part A – Exploratory Data Analysis (EDA) Perform the following EDA steps: [page 2] 1. Display dataset shape and column names 2. Identify missing values 3. Summary statistics (mean, std, min, max) 4. Class distribution (Outcome 0 vs 1) 5. Visualizations: o Histograms of numeric features o Boxplots to detect outliers o Correlation heatmap Question Part B – Data Preprocessing 1. Identify and Treat missing values 2. Feature scaling may be done say by using StandardScaler 3. Perform Train–test split by deciding the best spilt Question Part C – Model Building Train the following supervised learning models: 1. Logistic Regression 2. k-Nearest Neighbors (Identify optimal k) 3. Decision Tree Classifier Question Part D – Model Evaluation Evaluate each model using confusion matrix: • Accuracy • Precision • Recall • F1-Score • ROC-AUC and • Create a comparative performance table. Question Part E – Analysis Questions 1. Which model performs best overall? [page 3] 2. Why is recall important in diabetes prediction? 3. Which model is most interpretable? 4. Which model is most sensitive to feature scaling?