# [Optional] Study Guide- Introduction to AI ML and Data Science

course: Module 1 — Foundations of AI & ML
module: Module-1-Foundations-AI-ML
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-1-Foundations-AI-ML/General/[Optional]_Study_Guide-_Introduction_to_AI_ML_and_Data_Science.pdf
pages: 12

---
[page 1]
Study Guide Session 1. 
INTRODUCTION TO COURSE & DATA SCIENCE AND AI/ML  
 
INTRODUCTION TO THE DATA SCIENCE WORKFLOW 
 
The Data Science Workflow is a step-by-step process that guides practitioners from 
problem identification to model deployment and monitoring. It ensures structured, 
reproducible, and goal-oriented analysis. One well-known model is OSEMIN: 
Obtain, Scrub, Explore, Model, Interpret 
 
 
THE 9-STEP DATA SCIENCE PROCESS 
 
1. Problem Definition 
 
Theory 
Understand stakeholder goals 
Define measurable KPIs (e.g., accuracy, NPV uplift) 
Identify business constraints: data latency, computation limits, regulatory norms 
Practice 
Define a problem (e.g., predict churn in a telecom business) 
Template: Use Problem Framing Template 
 
2. Data Acquisition (Obtain) 
 
Theory 
Internal: CRM, POS 
External: APIs, web scraping 
Manual: Surveys, interviews 
Tools 
SQL (PostgreSQL, MySQL) 
Python: requests, pandas, BeautifulSoup 
APIs: Twitter API, OpenWeatherMap API 
Practice 
Scrape COVID-19 data from WHO or a news portal 
Save to CSV/JSON using pandas

[page 2]
3. Data Cleaning & Preprocessing (Scrub) 
 
Theory 
Handle missing values (mean/mode/ML-based imputation) 
Outliers: boxplots, z-score, IQR method 
Normalize: MinMaxScaler, 
StandardScaler 
Tools 
pandas, numpy, scikit-learn, matplotlib, 
seaborn 
Practice 
Dataset: Titanic Kaggle Dataset 
Code sample: Impute Age using 
median, detect outliers using IQR 
 
 
 
 
 
 
 
 
 
 
 
 
 
4. Exploratory Data Analysis (Explore) 
 
Theory 
Summary stats: describe() 
Visuals: Histograms, pairplots, heatmaps 
Feature distribution & correlation 
Tools 
matplotlib, seaborn, 
pandas_profiling, sweetviz 
Practice 
Use pandas_profiling to generate 
a full report of datasetVisualize 
outliers, missing values, 
correlation heatmaps

[page 3]
5. Feature Engineering 
 
Theory 
Encoding: LabelEncoder, OneHotEncoder 
Binning: Quantile, Decision Tree-based 
Dimensionality Reduction: PCA, t-SNE 
Tools 
sklearn.preprocessing, category_encoders, umap-learn 
Practice 
Dataset: Adult Income Dataset 
Convert education to ordinal scale; apply 
PCA 
 
 
 
6. Modeling 
 
Theory 
ML types: Regression, Classification, 
Clustering, RL 
Algorithms: Logistic Regression, Random 
Forest, XGBoost, KMeans 
Tools 
scikit-learn, XGBoost, LightGBM, 
TensorFlow, Keras 
Practice 
Dataset: Telco Churn 
Build & evaluate a Random Forest classifier 
 
 
 
7. Model Evaluation 
 
Theory 
Metrics: 
Classification: Accuracy, ROC-AUC, 
Confusion Matrix 
Regression: MSE, RMSE, R² 
Validation: K-fold, Grid Search 
Tools 
sklearn.metrics, yellowbrick 
Practice 
Plot ROC curves 
Calculate macro/micro precision, recall

[page 4]
8. Deployment 
 
 Theory 
Flask API: Create REST endpoint 
Cloud: AWS Lambda, GCP VertexAI 
CI/CD: Docker + GitHub Actions 
Tools 
Flask, FastAPI, Docker, Heroku, AWS 
SageMaker 
Practice 
Export model as joblib 
Serve via Flask and test using Postman 
 
 
 
 
 
9. Monitoring & Maintenance 
 
Theory 
Metrics Drift, Concept Drift 
MLOps Pipelines 
Retraining triggers 
 Tools 
mlflow, neptune.ai, evidently, prometheus, grafana 
 Practice 
Simulate drift by altering feature distributions 
Setup alert system for performance drop 
 
 
    Case Studies with Links 
1. Uber's Michelangelo MLOps Platform 
2. Zomato Price Optimization using XGBoost 
3. Flipkart ML Platform for Recommendations 
4. Google AI: ReCaptcha, Gmail, Translate 
5. LinkedIn AI for Skill Matching

[page 5]
Tools & Libraries 
• Languages: Python, R 
• Data Handling: pandas, numpy 
• Visualization: seaborn, matplotlib, plotly, tableau 
• ML Frameworks: scikit-learn, TensorFlow, Keras, PyCaret 
• Deployment: Docker, Flask, FastAPI, AWS/GCP 
• Monitoring: MLflow, Neptune.ai, Prometheus 
• Version Control: Git, GitHub, DVC 
 
              Practice Resources 
• DataCamp Skill Tracks 
• Kaggle Competitions 
• Practice Exams – 365 Data Science 
• Data Science Project Templates – GitHub 
 
Introduction to Artificial Intelligence (AI) 
• Definition & Scope: AI refers to the simulation of human intelligence 
processes by machines, especially computer systems. These processes 
include learning, reasoning, problem-solving, perception, and language 
understanding. 
• Historical Evolution: From the inception of AI in the 1950s to the current 
advancements in neural networks and deep learning. 
• Core Subfields: 
o Natural Language Processing (NLP): Enables machines to 
understand and interpret human language. Applications include 
chatbots, language translation, and sentiment analysis. 
o Computer Vision: Allows machines to interpret and process visual 
information from the world. Used in facial recognition, object detection, 
and medical imaging. 
o Expert Systems: Computer programs that mimic the decision-making 
abilities of human experts. Utilized in medical diagnosis and financial 
forecasting.

[page 6]
Machine Learning (ML): The Driving Force Behind AI 
• Definition: ML is a subset of AI that enables systems to learn and improve 
from experience without being explicitly programmed. 
• Types of ML: 
o Supervised Learning: Models are trained on labeled data. Examples 
include linear regression and support vector machines. 
o Unsupervised Learning: Models identify patterns in unlabeled data. 
Examples include k-means clustering and hierarchical clustering. 
o Reinforcement Learning: Models learn optimal actions through trial 
and error interactions with an environment. Used in robotics and game 
playing. 
• Algorithms & 
Applications: 
Detailed exploration 
of algorithms like 
decision trees, 
random forests, and 
their applications in 
sectors like finance, 
healthcare, and e-
commerce. 
 Deep Learning (DL): 
Mimicking the Human 
Brain 
• Overview: DL is a 
subset of ML that uses neural networks with multiple layers (deep neural 
networks) to model complex patterns in data. 
• Key Architectures: 
o Convolutional Neural Networks (CNNs): Primarily used for image 
and video recognition. 
o Recurrent Neural Networks (RNNs): Effective for sequential data like 
time series and natural language. 
o Transformers: Advanced models for NLP tasks, enabling parallel 
processing of data sequences. 
• Applications: From autonomous vehicles to voice assistants, DL has 
revolutionized numerous industries.

[page 7]
Generative AI (GenAI): 
Creating New Content 
• Definition: GenAI 
involves models that 
can generate new 
data instances 
resembling the 
training data. 
• Core Models: 
o Generative 
Adversarial 
Networks (GANs): Consist of a generator and a discriminator 
competing to produce realistic data. 
o Variational Autoencoders (VAEs): Encode input data into a latent 
space and decode it back to generate new data. 
o Transformer-based Models: Such as GPT-3 and BERT, capable of 
generating human-like text. 
• Use Cases: Content creation, image synthesis, music composition, and code 
generation. 
Agentic AI: Autonomous 
Decision-Making Systems 
• Definition: Agentic AI refers to 
systems capable of autonomous 
goal-setting, planning, decision-
making, and learning from 
outcomes without human 
intervention. 
• Characteristics: 
o Autonomy: Ability to operate 
without human guidance. 
o Adaptability: Learning from new data and experiences to improve 
performance. 
o Goal-Oriented Behavior: Pursuing objectives through planning and 
action. 
• Examples: AutoGPT, BabyAGI, and advanced robotics systems.

[page 8]
Interconnections Among AI Subfields 
• Hierarchical Structure: 
o AI: The overarching field encompassing all intelligent systems. 
▪ ML: A subset focusing on learning from data. 
▪ DL: A further subset 
utilizing deep neural networks. 
▪ GenAI: Specialized DL 
models generating new data. 
▪ Agentic AI: Advanced 
systems leveraging DL and 
GenAI for autonomous 
operations. 
• Integration in 
Applications: How these 
subfields collaborate in real-
world applications like 
autonomous vehicles and 
intelligent virtual assistants. 
 
 
Practical Case Studies 
1. Healthcare: Tumor Detection Using CNNs 
Problem: 
Detect tumors in MRI or X-ray scans automatically using image classification.  
Technique: 
Convolutional Neural Networks (CNNs) excel at pattern recognition in pixel data. 
 Tools: 
• Frameworks: TensorFlow, PyTorch, OpenCV 
• Dataset: Breast Histopathology Images – Kaggle 
Reference: 
• Stanford’s CheXNet Project

[page 9]
2. Finance: Fraud Detection 
Problem: 
Identify fraudulent transactions based on unusual patterns in customer data. 
Technique: 
Use supervised ML models like Random Forest, Isolation Forest, XGBoost. 
 Tools: 
• Frameworks: scikit-learn, LightGBM, XGBoost 
• Dataset: Credit Card Fraud Detection – Kaggle 
 Reference: 
• SAS on Fraud Detection 
 
3. Retail: Recommendation Systems 
Problem: 
Suggest personalized products to users. 
Technique: 
Collaborative Filtering (User-User / Item-Item), Matrix Factorization, Deep Learning 
Tools: 
• Libraries: Surprise, LightFM, TensorFlow Recommenders 
• Dataset: Movielens 100K Dataset 
 Reference: 
• Amazon’s Personalization at Scale 
 
4. Manufacturing: Predictive Maintenance 
 Problem: 
Predict machine failures using time series sensor data. 
 Technique: 
LSTM networks, ARIMA models, Time Series Forecasting 
 Tools: 
• Libraries: Prophet, tsfresh, sktime, TensorFlow

[page 10]
• Dataset: NASA Turbofan Engine Degradation Dataset 
   Reference: 
• Predictive Maintenance using ML 
 
5. Human Resources: Resume Screening with NLP & GenAI 
Problem: 
Automatically match resumes with job descriptions. 
 Technique: 
Text Embeddings (BERT, SBERT), Entity Extraction (Spacy), GPT-based matching 
Tools: 
• APIs: OpenAI GPT-3, HuggingFace Transformers 
• Libraries: Spacy, NLTK, TfidfVectorizer 
• Dataset: Job Descriptions and Resumes Dataset 
   Reference: 
• AI-Powered Resume Parser 
 
 
Hands-On Projects and Exercises 
Project 1: Supervised ML Model for Churn Prediction 
Goal: 
Predict whether a customer will churn using historical data. 
 Dataset: 
• Telco Customer Churn – Kaggle 
 Steps: 
1. Import and clean data 
2. Feature engineering 
3. Train model (e.g., Logistic Regression, XGBoost) 
4. Evaluate with ROC-AUC, Confusion Matrix 
Tutorial:

[page 11]
• Churn Prediction Notebook 
 
 Project 2: Image Classifier Using CNN in TensorFlow 
 Goal: 
Classify handwritten digits using CNNs. 
Dataset: 
• MNIST Dataset 
 Steps: 
1. Preprocess image data 
2. Define CNN architecture in Keras 
3. Train and validate model 
4. Test on new images 
 Tutorial: 
• TensorFlow CNN Image Classifier 
 
 Project 3: Text Generator with GPT-3 
 Goal: 
Generate text (e.g., email, paragraph, answer) from a prompt. 
Tool: 
• OpenAI GPT-3 via OpenAI API 
Steps: 
1. Register and get OpenAI API key 
2. Write Python script to call text-davinci-003 model 
3. Customize temperature, token length 
4. Generate multiple variations and fine-tune 
 Tutorial: 
• OpenAI Python Quickstart

[page 12]
Project 4: Autonomous Agent Using Reinforcement Learning 
 Goal: 
Create an AI agent that can make decisions in a dynamic environment. 
 Tools: 
• gym, stable-baselines3, LangChain, AutoGPT 
 Steps: 
1. Define environment (e.g., CartPole, custom task) 
2. Train using PPO/DQN algorithm 
3. Evaluate and visualize performance 
4. For advanced: Use AutoGPT for goal planning 
 References: 
• OpenAI Gym + Stable Baselines3 
• LangChain Official Docs 
• AutoGPT GitHub 
 
 Career Pathways in AI 
• Roles and Responsibilities: 
o AI/ML Engineer: Developing algorithms and models. 
o Data Scientist: Analyzing data to extract insights. 
o Research Scientist: Exploring new methodologies in AI. 
o AI Product Manager: Overseeing AI product development.