# 1.4 [Optional] Study Guide- Introduction to AI ML and Data Science
course: Module 1 — Foundations of AI & ML
module: Module-1-Foundations-AI-ML
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-1-Foundations-AI-ML/General/1.4_[Optional]_Study_Guide-_Introduction_to_AI_ML_and_Data_Science.pdf
pages: 12
---
[page 1]
Study Guide Session 1.
INTRODUCTION TO COURSE & DATA SCIENCE AND AI/ML
INTRODUCTION TO THE DATA SCIENCE WORKFLOW
The Data Science Workflow is a step-by-step process that guides practitioners from
problem identification to model deployment and monitoring. It ensures structured,
reproducible, and goal-oriented analysis. One well-known model is OSEMIN:
Obtain, Scrub, Explore, Model, Interpret
THE 9-STEP DATA SCIENCE PROCESS
1. Problem Definition
Theory
Understand stakeholder goals
Define measurable KPIs (e.g., accuracy, NPV uplift)
Identify business constraints: data latency, computation limits, regulatory norms
Practice
Define a problem (e.g., predict churn in a telecom business)
Template: Use Problem Framing Template
2. Data Acquisition (Obtain)
Theory
Internal: CRM, POS
External: APIs, web scraping
Manual: Surveys, interviews
Tools
SQL (PostgreSQL, MySQL)
Python: requests, pandas, BeautifulSoup
APIs: Twitter API, OpenWeatherMap API
Practice
Scrape COVID-19 data from WHO or a news portal
Save to CSV/JSON using pandas
[page 2]
3. Data Cleaning & Preprocessing (Scrub)
Theory
Handle missing values (mean/mode/ML-based imputation)
Outliers: boxplots, z-score, IQR method
Normalize: MinMaxScaler,
StandardScaler
Tools
pandas, numpy, scikit-learn, matplotlib,
seaborn
Practice
Dataset: Titanic Kaggle Dataset
Code sample: Impute Age using
median, detect outliers using IQR
4. Exploratory Data Analysis (Explore)
Theory
Summary stats: describe()
Visuals: Histograms, pairplots, heatmaps
Feature distribution & correlation
Tools
matplotlib, seaborn,
pandas_profiling, sweetviz
Practice
Use pandas_profiling to generate
a full report of datasetVisualize
outliers, missing values,
correlation heatmaps
[page 3]
5. Feature Engineering
Theory
Encoding: LabelEncoder, OneHotEncoder
Binning: Quantile, Decision Tree-based
Dimensionality Reduction: PCA, t-SNE
Tools
sklearn.preprocessing, category_encoders, umap-learn
Practice
Dataset: Adult Income Dataset
Convert education to ordinal scale; apply
PCA
6. Modeling
Theory
ML types: Regression, Classification,
Clustering, RL
Algorithms: Logistic Regression, Random
Forest, XGBoost, KMeans
Tools
scikit-learn, XGBoost, LightGBM,
TensorFlow, Keras
Practice
Dataset: Telco Churn
Build & evaluate a Random Forest classifier
7. Model Evaluation
Theory
Metrics:
Classification: Accuracy, ROC-AUC,
Confusion Matrix
Regression: MSE, RMSE, R²
Validation: K-fold, Grid Search
Tools
sklearn.metrics, yellowbrick
Practice
Plot ROC curves
Calculate macro/micro precision, recall
[page 4]
8. Deployment
Theory
Flask API: Create REST endpoint
Cloud: AWS Lambda, GCP VertexAI
CI/CD: Docker + GitHub Actions
Tools
Flask, FastAPI, Docker, Heroku, AWS
SageMaker
Practice
Export model as joblib
Serve via Flask and test using Postman
9. Monitoring & Maintenance
Theory
Metrics Drift, Concept Drift
MLOps Pipelines
Retraining triggers
Tools
mlflow, neptune.ai, evidently, prometheus, grafana
Practice
Simulate drift by altering feature distributions
Setup alert system for performance drop
Case Studies with Links
1. Uber's Michelangelo MLOps Platform
2. Zomato Price Optimization using XGBoost
3. Flipkart ML Platform for Recommendations
4. Google AI: ReCaptcha, Gmail, Translate
5. LinkedIn AI for Skill Matching
[page 5]
Tools & Libraries
• Languages: Python, R
• Data Handling: pandas, numpy
• Visualization: seaborn, matplotlib, plotly, tableau
• ML Frameworks: scikit-learn, TensorFlow, Keras, PyCaret
• Deployment: Docker, Flask, FastAPI, AWS/GCP
• Monitoring: MLflow, Neptune.ai, Prometheus
• Version Control: Git, GitHub, DVC
Practice Resources
• DataCamp Skill Tracks
• Kaggle Competitions
• Practice Exams – 365 Data Science
• Data Science Project Templates – GitHub
Introduction to Artificial Intelligence (AI)
• Definition & Scope: AI refers to the simulation of human intelligence
processes by machines, especially computer systems. These processes
include learning, reasoning, problem-solving, perception, and language
understanding.
• Historical Evolution: From the inception of AI in the 1950s to the current
advancements in neural networks and deep learning.
• Core Subfields:
o Natural Language Processing (NLP): Enables machines to
understand and interpret human language. Applications include
chatbots, language translation, and sentiment analysis.
o Computer Vision: Allows machines to interpret and process visual
information from the world. Used in facial recognition, object detection,
and medical imaging.
o Expert Systems: Computer programs that mimic the decision-making
abilities of human experts. Utilized in medical diagnosis and financial
forecasting.
[page 6]
Machine Learning (ML): The Driving Force Behind AI
• Definition: ML is a subset of AI that enables systems to learn and improve
from experience without being explicitly programmed.
• Types of ML:
o Supervised Learning: Models are trained on labeled data. Examples
include linear regression and support vector machines.
o Unsupervised Learning: Models identify patterns in unlabeled data.
Examples include k-means clustering and hierarchical clustering.
o Reinforcement Learning: Models learn optimal actions through trial
and error interactions with an environment. Used in robotics and game
playing.
• Algorithms &
Applications:
Detailed exploration
of algorithms like
decision trees,
random forests, and
their applications in
sectors like finance,
healthcare, and e-
commerce.
Deep Learning (DL):
Mimicking the Human
Brain
• Overview: DL is a
subset of ML that uses neural networks with multiple layers (deep neural
networks) to model complex patterns in data.
• Key Architectures:
o Convolutional Neural Networks (CNNs): Primarily used for image
and video recognition.
o Recurrent Neural Networks (RNNs): Effective for sequential data like
time series and natural language.
o Transformers: Advanced models for NLP tasks, enabling parallel
processing of data sequences.
• Applications: From autonomous vehicles to voice assistants, DL has
revolutionized numerous industries.
[page 7]
Generative AI (GenAI):
Creating New Content
• Definition: GenAI
involves models that
can generate new
data instances
resembling the
training data.
• Core Models:
o Generative
Adversarial
Networks (GANs): Consist of a generator and a discriminator
competing to produce realistic data.
o Variational Autoencoders (VAEs): Encode input data into a latent
space and decode it back to generate new data.
o Transformer-based Models: Such as GPT-3 and BERT, capable of
generating human-like text.
• Use Cases: Content creation, image synthesis, music composition, and code
generation.
Agentic AI: Autonomous
Decision-Making Systems
• Definition: Agentic AI refers to
systems capable of autonomous
goal-setting, planning, decision-
making, and learning from
outcomes without human
intervention.
• Characteristics:
o Autonomy: Ability to operate
without human guidance.
o Adaptability: Learning from new data and experiences to improve
performance.
o Goal-Oriented Behavior: Pursuing objectives through planning and
action.
• Examples: AutoGPT, BabyAGI, and advanced robotics systems.
[page 8]
Interconnections Among AI Subfields
• Hierarchical Structure:
o AI: The overarching field encompassing all intelligent systems.
▪ ML: A subset focusing on learning from data.
▪ DL: A further subset
utilizing deep neural networks.
▪ GenAI: Specialized DL
models generating new data.
▪ Agentic AI: Advanced
systems leveraging DL and
GenAI for autonomous
operations.
• Integration in
Applications: How these
subfields collaborate in real-
world applications like
autonomous vehicles and
intelligent virtual assistants.
Practical Case Studies
1. Healthcare: Tumor Detection Using CNNs
Problem:
Detect tumors in MRI or X-ray scans automatically using image classification.
Technique:
Convolutional Neural Networks (CNNs) excel at pattern recognition in pixel data.
Tools:
• Frameworks: TensorFlow, PyTorch, OpenCV
• Dataset: Breast Histopathology Images – Kaggle
Reference:
• Stanford’s CheXNet Project
[page 9]
2. Finance: Fraud Detection
Problem:
Identify fraudulent transactions based on unusual patterns in customer data.
Technique:
Use supervised ML models like Random Forest, Isolation Forest, XGBoost.
Tools:
• Frameworks: scikit-learn, LightGBM, XGBoost
• Dataset: Credit Card Fraud Detection – Kaggle
Reference:
• SAS on Fraud Detection
3. Retail: Recommendation Systems
Problem:
Suggest personalized products to users.
Technique:
Collaborative Filtering (User-User / Item-Item), Matrix Factorization, Deep Learning
Tools:
• Libraries: Surprise, LightFM, TensorFlow Recommenders
• Dataset: Movielens 100K Dataset
Reference:
• Amazon’s Personalization at Scale
4. Manufacturing: Predictive Maintenance
Problem:
Predict machine failures using time series sensor data.
Technique:
LSTM networks, ARIMA models, Time Series Forecasting
Tools:
• Libraries: Prophet, tsfresh, sktime, TensorFlow
[page 10]
• Dataset: NASA Turbofan Engine Degradation Dataset
Reference:
• Predictive Maintenance using ML
5. Human Resources: Resume Screening with NLP & GenAI
Problem:
Automatically match resumes with job descriptions.
Technique:
Text Embeddings (BERT, SBERT), Entity Extraction (Spacy), GPT-based matching
Tools:
• APIs: OpenAI GPT-3, HuggingFace Transformers
• Libraries: Spacy, NLTK, TfidfVectorizer
• Dataset: Job Descriptions and Resumes Dataset
Reference:
• AI-Powered Resume Parser
Hands-On Projects and Exercises
Project 1: Supervised ML Model for Churn Prediction
Goal:
Predict whether a customer will churn using historical data.
Dataset:
• Telco Customer Churn – Kaggle
Steps:
1. Import and clean data
2. Feature engineering
3. Train model (e.g., Logistic Regression, XGBoost)
4. Evaluate with ROC-AUC, Confusion Matrix
Tutorial:
[page 11]
• Churn Prediction Notebook
Project 2: Image Classifier Using CNN in TensorFlow
Goal:
Classify handwritten digits using CNNs.
Dataset:
• MNIST Dataset
Steps:
1. Preprocess image data
2. Define CNN architecture in Keras
3. Train and validate model
4. Test on new images
Tutorial:
• TensorFlow CNN Image Classifier
Project 3: Text Generator with GPT-3
Goal:
Generate text (e.g., email, paragraph, answer) from a prompt.
Tool:
• OpenAI GPT-3 via OpenAI API
Steps:
1. Register and get OpenAI API key
2. Write Python script to call text-davinci-003 model
3. Customize temperature, token length
4. Generate multiple variations and fine-tune
Tutorial:
• OpenAI Python Quickstart
[page 12]
Project 4: Autonomous Agent Using Reinforcement Learning
Goal:
Create an AI agent that can make decisions in a dynamic environment.
Tools:
• gym, stable-baselines3, LangChain, AutoGPT
Steps:
1. Define environment (e.g., CartPole, custom task)
2. Train using PPO/DQN algorithm
3. Evaluate and visualize performance
4. For advanced: Use AutoGPT for goal planning
References:
• OpenAI Gym + Stable Baselines3
• LangChain Official Docs
• AutoGPT GitHub
Career Pathways in AI
• Roles and Responsibilities:
o AI/ML Engineer: Developing algorithms and models.
o Data Scientist: Analyzing data to extract insights.
o Research Scientist: Exploring new methodologies in AI.
o AI Product Manager: Overseeing AI product development.