# [Optional] Study Guide - EDA
course: Module 1 β Foundations of AI & ML
module: Module-1-Foundations-AI-ML
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-1-Foundations-AI-ML/General/[Optional]_Study_Guide_-_EDA.pdf
pages: 22
---
[page 1]
STUDY
GUIDE
:
1.4
Exploratory
Data
Analysis
(EDA):
Mastering
the
Art
of
Data
Interrogation
Welcome
to
this
advanced
study
guide
on
Exploratory
Data
Analysis
(EDA).
For
professionals
like
you
in
the
PGC
GenAI
program
at
IIT
Roorkee,
powered
by
Futurense,
with
substantial
technical
experience,
EDA
transcends
a
mere
procedural
checklist.
It
becomes
a
strategic
imperativeβa
deeply
intellectual
process
of
inquiry,
skepticism,
and
discovery
that
forms
the
bedrock
of
impactful
AI
and
GenAI
solutions.
Think
of
yourselves
as
seasoned
investigators
arriving
at
a
complex
scene;
your
ability
to
meticulously
examine
evidence,
question
assumptions,
and
connect
disparate
clues
through
EDA
will
directly
determine
the
success
and
reliability
of
the
high-stakes
models
and
systems
you'll
build.
This
guide
is
designed
to
foster
that
investigative
mindset,
moving
beyond
basic
techniques
to
explore
the
"why,"
the
"what
if,"
and
the
"so
what"
of
your
data
explorations,
especially
pertinent
for
those
stepping
into
or
operating
in
roles
like
Chief
AI
Officer,
Principal
Data
Scientist,
AI
Research
Lead,
or
Senior
Manager
of
AI/ML
Solutions.
π―
Why
EDA
is
Your
Indispensable
Strategic
Compass
(Objectives)
For
senior
practitioners,
the
objectives
of
EDA
gain
additional
layers
of
strategic
importance:
β
Profound
Data
Acumen
&
Contextual
Intelligence
(Beyond
Surface-Level
Understanding):
β
Elaboration:
This
isn't
just
about
variable
types;
it's
about
understanding
data
provenance,
lineage,
and
potential
biases
embedded
during
collection.
It
involves
questioning
the
proxies
used
for
real-world
constructs
(e.g.,
is
'click-through
rate'
truly
the
best
proxy
for
'user
engagement'
in
this
specific
GenAI
application
?).
β
Advanced
Consideration
(AI
Research
Lead):
When
evaluating
datasets
for
training
foundational
models,
EDA
must
critically
assess
not
just
content,
but
also
representational
biases,
potential
copyright
[page 2]
issues
hinted
at
by
data
sources,
and
the
ethical
implications
of
the
data's
origin.
β
Pitfall
to
Avoid:
Assuming
data
is
"ground
truth"
without
critically
examining
its
collection
methodology
and
inherent
limitations.
β
Ensuring
Impeccable
Data
Integrity
&
Robustness
(The
Cornerstone
of
Trustworthy
AI):
β
Elaboration:
Beyond
simple
imputation,
this
involves
understanding
the
mechanism
of
missingness
(MCAR,
MAR,
MNAR)
to
choose
an
appropriate
advanced
imputation
strategy
(e.g.,
regression
imputation,
multiple
imputation).
For
outliers,
it's
about
distinguishing
between
data
errors,
genuine
extreme
values,
and
"black
swan"
events.
β
Advanced
Consideration
(Principal
Data
Scientist):
Undetected
data
quality
issues
(e.g.,
subtle
data
drift,
inconsistent
labeling
in
training
data
for
a
supervised
GenAI
model)
can
lead
to
models
that
perform
well
in
testing
but
fail
catastrophically
or
unfairly
in
production.
EDA
is
the
first
line
of
defense.
β
Visualization
Cue:
Imagine
using
missing
data
patterns
(e.g.,
from
msno.matrix
in
Python's
missingno
library)
not
just
to
see
where
data
is
missing,
but
to
hypothesize
why
,
discussing
this
with
data
owners.
(You
can
explore
the
missingno
library
visuals
by
searching
"missingno
python
library
examples"
on
the
web).
β
Illuminating
Latent
Structures,
Complex
Interactions
&
Emergent
Patterns
(The
Hunt
for
Alpha):
β
Elaboration:
This
involves
looking
for
non-linear
relationships,
interaction
effects
(where
the
effect
of
one
variable
depends
on
the
level
of
another),
and
clustering
tendencies
that
might
not
be
obvious
from
simple
bivariate
plots.
This
is
where
domain
expertise
combined
with
advanced
visualization
shines.
β
Advanced
Consideration
(Chief
AI
Officer):
EDA
insights
can
identify
opportunities
for
novel
feature
engineering
that
could
give
your
AI
models
a
competitive
edge.
For
instance,
discovering
a
non-linear
threshold
effect
in
sensor
data
could
lead
to
a
critical
feature
for
predictive
maintenance.
[page 3]
β
Analogy
for
Indian
Professionals:
Think
of
it
as
discerning
the
subtle,
underlying
Raga
(melodic
framework)
from
a
complex
musical
piece,
even
when
multiple
instruments
(variables)
are
playing
intricate
patterns.
β
Guiding
Sophisticated
Feature
Engineering
&
Architecting
Optimal
Models
(Strategic
Resource
Allocation):
β
Elaboration:
EDA
informs
decisions
like
whether
to
use
tree-based
models
(less
sensitive
to
outliers
and
scaling)
or
neural
networks
(which
might
require
careful
normalization
and
outlier
treatment).
It
helps
identify
when
to
create
interaction
terms,
polynomial
features,
or
embeddings
(especially
for
GenAI
dealing
with
text/image).
β
Advanced
Consideration:
The
"shape"
of
your
data
(distributions,
sparsity,
dimensionality)
heavily
influences
the
choice
of
algorithms
and
the
need
for
techniques
like
dimensionality
reduction
(e.g.,
PCA,
UMAP,
t-SNE
initially
explored
during
EDA).
For
GenAI,
understanding
token
distributions,
sequence
lengths,
etc.,
is
critical
for
EDA.
β
Rigorous
Hypothesis
Formulation,
Iterative
Testing
&
Assumption
Validation
(The
Scientific
Backbone):
β
Elaboration:
Encourage
formulating
S.M.A.R.T.
(Specific,
Measurable,
Achievable,
Relevant,
Time-bound)
hypotheses
during
EDA.
For
example,
instead
of
"Are
sales
related
to
marketing
spend?",
a
better
hypothesis
is
"Does
a
10%
increase
in
digital
marketing
spend
lead
to
a
>5%
increase
in
online
sales
within
the
same
quarter,
holding
other
factors
constant?"
EDA
provides
the
initial
evidence.
β
Pitfall
to
Avoid:
Confirmation
bias
β
only
looking
for
evidence
that
supports
pre-existing
beliefs.
EDA
should
be
an
honest
exploration.
β
Compelling
Narrative
Construction
&
Influential
Stakeholder
Communication
(Driving
Action):
β
Elaboration:
For
senior
roles,
EDA
outputs
often
feed
into
strategic
decisions.
The
ability
to
translate
complex
data
findings
into
clear,
[page 4]
concise,
and
actionable
narratives
for
diverse
audiences
(technical
peers,
executive
leadership,
regulatory
bodies)
is
paramount.
β
Advanced
Consideration:
Your
EDA
story
should
not
just
present
findings
but
also
articulate
the
uncertainty
and
limitations
of
the
data
and
the
analysis.
πΊ
The
EDA
Expedition:
A
Detailed
&
Tactical
Itinerary
This
expedition
requires
meticulous
planning
and
execution.
1.
Problem
Articulation,
Strategic
Alignment
&
Data
Ecosystem
Understanding:
β
Elaboration:
Start
by
co-defining
the
Key
Performance
Indicators
(KPIs)
for
the
project
and
for
the
EDA
phase
itself
(e.g.,
"Identify
top
3
data
quality
issues
impacting
X,"
"Characterize
user
segments
based
on
Y
behavior").
Map
the
data
sources
to
the
business
processes
they
represent.
Understand
data
governance
policies
and
access
constraints
early.
β
Job
Role
Context
(Senior
Manager,
AI/ML
Solutions):
This
involves
extensive
stakeholder
interviews
β
from
business
users
to
data
[page 5]
engineers
β
to
capture
requirements,
assumptions,
and
potential
data
"gotchas"
before
a
single
line
of
code
is
written.
β
Tools
for
Thought:
Use
mind
maps
or
influence
diagrams
to
map
out
variables,
expected
relationships,
and
business
impacts.
2.
Forensic
Data
Cleaning
&
Advanced
Preprocessing
β
Establishing
a
Gold
Standard:
β
Handling
Missing
Data
(
NaN
,
None
,
Null
):
β
Advanced
Imputation:
β
Regression
Imputation:
Predict
missing
values
using
other
variables.
Pro:
Can
be
accurate.
Con:
May
artificially
reduce
variance
and
strengthen
correlations.
β
Stochastic
Regression
Imputation:
Adds
a
random
error
term
to
regression
predictions
to
preserve
variance.
β
Multiple
Imputation
(e.g.,
MICE
-
Multivariate
Imputation
by
Chained
Equations):
Creates
multiple
complete
datasets,
runs
analysis
on
each,
then
pools
results.
More
robust
but
complex.
(Scikit-learn
offers
IterativeImputer
for
this:
https://scikit-learn.org/stable/modules/generated/sklearn.i
mpute.IterativeImputer.html
)
β
Consideration
for
GenAI:
For
sequential
data
(text,
time
series),
specialized
imputation
methods
that
respect
sequence
order
might
be
needed
(e.g.,
forward/backward
fill,
interpolation,
or
even
model-based
imputation
using
LSTMs/Transformers
if
data
is
rich
enough).
β
Outlier
Detection
&
Sophisticated
Treatment:
β
Robust
Methods:
Beyond
Z-score/IQR,
consider
methods
less
sensitive
to
extreme
outliers
themselves,
like
using
the
Median
Absolute
Deviation
(MAD).
Model-based
outlier
detection
(e.g.,
isolation
forests,
one-class
SVM)
can
be
part
of
advanced
EDA.
(Scikit-learn
examples:
Isolation
Forest
-
[page 6]
https://scikit-learn.org/stable/modules/generated/sklearn.ensem
ble.IsolationForest.html
,
One-Class
SVM
-
https://scikit-learn.org/stable/modules/generated/sklearn.svm.On
eClassSVM.html
)
β
Philosophy:
When
is
an
outlier
an
error
vs.
a
critical
insight?
If
it's
an
error,
what
process
generated
it?
Can
this
be
fixed
at
the
source?
If
it's
a
genuine
rare
event,
how
should
the
model
be
designed
to
handle
it
or
learn
from
it?
β
Data
Type
Integrity
&
Semantic
Consistency:
β
Elaboration:
Beyond
astype()
,
this
includes
checking
for
semantic
consistency
(e.g.,
a
'country'
column
having
"USA"
and
"United
States"
β
requiring
normalization).
For
numerical
data,
check
if
units
are
consistent
(e.g.,
kgs
vs.
lbs).
β
High
Cardinality
Categorical
Variables:
EDA
needs
to
identify
these
early.
Strategies
include
grouping
less
frequent
categories,
target
encoding
(with
care
to
avoid
leakage),
or
embedding
techniques.
3.
Deep
Univariate
Analysis
β
Profiling
Individual
Data
Personalities:
β
Numerical
Variables
-
Beyond
Basic
Shapes:
β
Binning
Strategies
for
Histograms:
Don't
rely
on
defaults.
Conceptually
understand
rules
like
Freedman-Diaconis
(robust
to
outliers)
or
Sturges'
Law.
The
goal
is
to
reveal,
not
obscure,
the
underlying
distribution.
(Pandas
hist()
or
Seaborn
histplot()
allow
bin
customization).
β
Interpreting
Skewness
&
Kurtosis
Together:
β
High
Skew
+
High
Kurtosis:
Very
asymmetric
with
many
extreme
outliers
on
one
side.
β
Low
Skew
+
High
Kurtosis:
Symmetric
but
with
heavy
tails
(more
outliers
than
normal).
β
High
Skew
+
Low
Kurtosis:
Asymmetric
but
with
fewer
extreme
outliers
than
expected.
(Pandas
Series
have
.skew()
and
.kurt()
methods).
[page 7]
β
Categorical
Variables
-
Imbalance
and
Rarity:
β
Impact
of
Imbalance:
Highly
imbalanced
categorical
features
(e.g.,
99%
Class
A,
1%
Class
B)
pose
challenges
for
models.
EDA
must
quantify
this.
β
Rare
Categories:
How
to
handle
categories
that
appear
only
a
few
times?
(Group,
remove,
treat
as
special
case).
4.
Nuanced
Bivariate/Multivariate
Analysis
β
Deciphering
Complex
Interplays:
β
Interaction
Effects:
Actively
look
for
these.
E.g.,
the
impact
of
'Ad
Spend'
on
'Sales'
might
differ
significantly
across
'Customer
Segments'.
Visualized
via
grouped
plots
where
the
trend
lines
for
different
groups
are
not
parallel.
(Seaborn's
lmplot
with
hue
and
col
/
row
arguments
is
excellent
for
this).
β
Advanced
Visualizations
(Conceptual
with
Search
Cues):
β
Parallel
Coordinate
Plots:
For
visualizing
many
variables
at
once.
(Search
"parallel
coordinate
plot
python
pandas"
or
"parallel
coordinate
plot
plotly"
for
examples.
Plotly
offers
good
interactive
versions:
https://plotly.com/python/parallel-coordinates-plot/
)
β
Andrews
Curves:
Represents
each
multivariate
observation
as
a
curve.
(Pandas
plotting
includes
Andrews
curves:
https://pandas.pydata.org/pandas-docs/stable/user_guide/visuali
zation.html#andrews-curves
)
β
Dimensionality
Reduction
as
EDA:
Techniques
like
Principal
Component
Analysis
(PCA),
t-SNE,
or
UMAP
can
be
used
in
an
exploratory
manner
to
project
high-dimensional
data
into
2D
or
3D
for
visualization.
β
Visualization
Cue
(t-SNE/UMAP):
Search
"t-SNE
visualization
of
MNIST
python"
or
"UMAP
for
clustering
python"
for
code
examples
and
visual
outputs.
(Scikit-learn:
TSNE
-
https://scikit-learn.org/stable/modules/generated/sklearn.manifol
d.TSNE.html
,
UMAP
library:
https://umap-learn.readthedocs.io/
)
β
Statistical
Tests
(Conceptual
Understanding):
[page 8]
β
MANOVA
(Multivariate
Analysis
of
Variance):
Conceptually,
an
extension
of
ANOVA
to
multiple
dependent
numerical
variables.
(Statsmodels
library
in
Python
offers
MANOVA:
https://www.statsmodels.org/stable/multivariate.html
)
β
Canonical
Correlation
Analysis
(CCA):
Explores
relationships
between
two
sets
of
variables.
(Scikit-learn:
CCA
-
https://scikit-learn.org/stable/modules/generated/sklearn.cross_d
ecomposition.CCA.html
)
5.
Insight
Synthesis
&
Causal
Hypothesis
Generation
(The
"Aha!"
Moments):
β
Elaboration:
Move
beyond
observation
to
interpretation
and
synthesis.
Create
an
"Insight
Log."
For
senior
roles,
this
includes
asking
"What
are
the
second-order
effects
of
this
finding?"
or
"How
does
this
challenge
our
current
business
assumptions?"
β
Correlation
vs.
Causation:
EDA
will
primarily
show
correlations.
Emphasize
that
correlation
does
not
imply
causation.
However,
EDA
can
help
formulate
hypotheses
for
further
causal
inference
studies
(e.g.,
A/B
testing,
quasi-experimental
designs).
β
Example:
EDA
shows
a
strong
positive
correlation
between
ice
cream
sales
and
crime
rates.
The
insight
isn't
that
ice
cream
causes
crime,
but
that
a
confounding
variable
(e.g.,
hot
weather)
likely
influences
both.
6.
Strategic
Reporting
&
Actionable
Recommendations
(Driving
Change):
β
Tailoring
Reports:
β
For
Technical
Leads/Peers:
Detailed
plots,
statistical
test
results,
discussion
of
data
structures,
and
code
snippets.
β
For
C-Suite/Business
Stakeholders:
High-level
summaries,
key
insights
visualized
simply
(e.g.,
a
single
impactful
chart),
business
implications,
and
strategic
recommendations.
Focus
on
the
"so
what."
[page 9]
β
Quantify
Uncertainty:
Where
possible,
express
the
confidence
in
findings.
Acknowledge
limitations
of
the
data
or
analysis.
Your
EDA
Arsenal:
Techniques
&
Visualizations
β
An
Expanded
View
Non-Graphical
EDA
(Precision
and
Robustness):
β
Descriptive
Statistics
-
Robust
Alternatives:
β
Trimmed
Mean/Winsorized
Mean:
Mitigate
outlier
impact
by
removing
or
adjusting
a
percentage
of
extreme
values
before
calculating
the
mean.
(SciPy
offers
trim_mean
:
https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.trim_m
ean.html
)
β
Median
Absolute
Deviation
(MAD):
Robust
measure
of
dispersion
(MAD=median(
β£
X_iβmedian(X)
β£
)).
(Statsmodels
has
robust.mad
:
https://www.statsmodels.org/dev/generated/statsmodels.robust.scale.m
ad.html
)
β
Correlation
Analysis
-
Beyond
Pearson/Spearman:
β
Distance
Correlation:
Can
detect
non-linear
relationships.
(The
dcor
library
in
Python
implements
this:
https://pypi.org/project/dcor/
)
β
Mutual
Information:
Measures
non-linear
dependencies.
(Scikit-learn:
mutual_info_regression
,
mutual_info_classif
-
https://scikit-learn.org/stable/modules/feature_selection.html#mutual-ior
mation
)
β
Hypothesis
Testing
during
EDA:
Simple
tests
(t-tests,
chi-squared)
can
quickly
validate
observations.
(SciPy.stats:
ttest_ind
,
chi2_contingency
-
https://docs.scipy.org/doc/scipy/reference/stats.html
)
Graphical
EDA
(Clarity,
Depth,
and
Nuance
in
Visuals):
β
Histograms
-
Advanced
Binning:
[page 10]
β
Rules
like
Freedman-Diaconis,
Sturges'
Law,
Scott's
Normal
Reference
Rule
guide
bin
selection.
β
Key:
Experiment
and
choose
what
best
reveals
structure.
β
Q-Q
Plots
-
Interpreting
Deviations:
β
Deviations
indicate
non-normality,
skewness,
heavy/light
tails.
(Statsmodels
qqplot
:
https://www.statsmodels.org/dev/generated/statsmodels.graphics.gofpl
ots.qqplot.html
)
β
Pair
Plots
-
Considerations
for
High
Dimensions:
β
Become
unwieldy
for
many
features.
(Seaborn
pairplot
:
https://seaborn.pydata.org/generated/seaborn.pairplot.html
)
β
Mosaic
Plots:
For
visualizing
relationships
between
multiple
categorical
variables.
(Statsmodels
mosaic
:
https://www.statsmodels.org/dev/generated/statsmodels.graphics.mosaicplot.
mosaic.html
)
β
Visualization
Cue:
Search
"mosaic
plot
example
titanic
dataset
python"
for
visual
examples.
β
Residual
Plots
(after
initial
simple
modeling):
Can
reveal
non-linearity,
heteroscedasticity.
(Seaborn
residplot
:
https://seaborn.pydata.org/generated/seaborn.residplot.html
)
[page 11]
Solved
Example:
[page 15]
Industry
case
study
for
EDA
:
https://www.linkedin.com/pulse/exploratory-data-analysis-eda-case-study-archana-kh
ewariya-jqquf/
Quizzes
(MCQs
to
Test
Your
Understanding)
1.
Which
of
the
following
is
NOT
a
primary
objective
of
EDA?
a)
Identifying
missing
values
and
outliers.
b)
Building
a
production-ready
machine
learning
model.
c)
Understanding
the
relationships
between
variables.
d)
Visualizing
the
distribution
of
key
variables.
2.
You
have
a
dataset
of
customer
transaction
amounts
and
want
to
quickly
see
its
distribution
and
identify
potential
extreme
values.
Which
plot
is
most
suitable?
a)
Scatter
Plot
[page 16]
b)
Bar
Chart
c)
Box
Plot
d)
Line
Plot
3.
If
two
numerical
variables
have
a
Pearson
correlation
coefficient
of
-0.85,
it
means:
a)
They
have
a
weak
negative
linear
relationship.
b)
They
have
a
strong
positive
linear
relationship.
c)
They
have
a
strong
negative
linear
relationship.
d)
There
is
no
linear
relationship
between
them.
4.
To
analyze
the
relationship
between
two
categorical
variables,
"Region"
(North,
South,
East,
West)
and
"Product
Preference"
(A,
B,
C),
which
technique
is
most
appropriate?
a)
Calculating
the
mean
for
each
region.
b)
Creating
a
scatter
plot.
c)
Generating
a
contingency
table
(crosstab).
d)
Plotting
a
histogram
for
product
preference.
5.
Which
Python
library
is
primarily
used
for
data
manipulation
and
cleaning
during
EDA?
a)
Matplotlib
b)
Seaborn
c)
Pandas
d)
SciPy
6.
"Skewness"
in
a
dataset
refers
to:
a)
The
peakedness
of
the
distribution.
b)
The
measure
of
how
spread
out
the
data
is.
c)
The
asymmetry
of
the
distribution.
d)
The
average
value
of
the
dataset.
7.
When
performing
EDA,
if
you
encounter
a
feature
with
70%
missing
values,
what
is
a
generally
reasonable
first
approach
for
an
experienced
analyst?
[page 17]
a)
Immediately
drop
the
feature.
b)
Impute
all
missing
values
with
the
mean.
c)
Investigate
why
the
data
is
missing
and
then
decide
to
drop
or
impute.
d)
Replace
missing
values
with
zeros.
8.
A
Data
Science
Architect
is
examining
sensor
data
from
manufacturing
equipment.
They
notice
that
temperature
readings
occasionally
spike
to
unrealistic
values
(e.g.,
1000Β°C
for
a
water
pipe).
These
are
likely:
a)
Seasonal
trends.
b)
Outliers.
c)
Normal
variations.
d)
Missing
data.
9.
What
is
the
primary
purpose
of
using
a
heatmap
in
EDA?
a)
To
display
the
distribution
of
a
single
numerical
variable.
b)
To
visualize
the
relationship
between
two
categorical
variables.
c)
To
show
trends
over
time
for
multiple
variables.
d)
To
graphically
represent
a
matrix
of
values,
like
a
correlation
matrix,
using
color
intensities.
10.
In
the
context
of
EDA,
"Univariate
Analysis"
means:
a)
Analyzing
the
relationship
between
two
variables.
b)
Analyzing
a
single
variable
in
isolation.
c)
Building
a
model
with
only
one
feature.
d)
Analyzing
data
from
a
single
source.
(Answers
at
the
end
of
the
guide)
Practice
Questions
1.
Scenario
(Short
Answer):
You
are
given
a
dataset
of
employee
performance
with
features
like
'Years_Experience',
'Last_Review_Score'
(1-5),
'Projects_Completed',
'Department',
and
'Salary'.
What
would
be
your
first
3
[page 18]
EDA
steps
using
Pandas,
and
what
specific
insights
would
you
be
looking
for
with
each?
2.
Application
(Outliers):
Describe
a
situation
in
a
domain
you're
familiar
with
(e.g.,
e-commerce,
finance,
manufacturing)
where
an
outlier
might
be
a
critical
piece
of
information
rather
than
an
error.
How
would
your
EDA
approach
help
you
distinguish
this?
3.
Scenario
(Visualization):
You
need
to
explain
the
relationship
between
'Customer_Age_Group'
(Categorical:
Young,
Middle-aged,
Senior)
and
'Average_Monthly_Spend'
(Numerical)
to
a
non-technical
stakeholder.
Which
visualization
would
you
choose
and
why?
Sketch
or
describe
what
it
would
look
like
and
what
story
it
would
tell.
4.
Data
Cleaning
(Missing
Data):
Imagine
a
'Customer_Feedback_Text'
column
in
your
dataset.
If
10%
of
these
text
entries
are
missing,
discuss
two
different
strategies
for
handling
these
missing
values
and
the
potential
pros
and
cons
of
each
in
the
context
of
preparing
data
for
a
sentiment
analysis
model.
5.
Multivariate
Analysis
(Job
Role
Focus
-
AI
Lead):
An
AI
Lead
is
investigating
data
for
a
model
to
predict
employee
attrition.
They
have
features
like
'Job_Satisfaction_Score',
'Last_Promotion_Date',
'Distance_From_Home',
'Monthly_Income'.
How
could
they
use
a
correlation
matrix
and
pair
plots
to
gain
initial
insights
into
potential
drivers
of
attrition?
What
specific
patterns
might
they
look
for?
6.
Transformation
(Short
Answer):
You
plot
a
histogram
for
'Household_Income'
and
find
it's
heavily
right-skewed.
Why
might
this
be
a
problem
for
some
analytical
models,
and
what
common
transformation
could
you
apply
during
EDA
to
address
this?
What
would
you
check
after
applying
the
transformation?
7.
Categorical
Data
(Application):
You
have
two
categorical
variables:
'User_Device_Type'
(Desktop,
Mobile,
Tablet)
and
'Subscription_Plan'
(Basic,
Premium,
Enterprise).
How
would
you
explore
if
there's
an
association
between
the
device
type
and
the
chosen
subscription
plan?
Mention
both
a
statistical
test
(optional,
if
known)
and
a
visualization.
[page 19]
8.
EDA
for
Time
Series
(Scenario):
You
have
daily
sales
data
for
a
retail
store
over
the
past
3
years.
What
specific
EDA
techniques
and
visualizations
would
you
use
to
understand
seasonality,
trends,
and
any
unusual
patterns?
9.
Feature
Importance
(Conceptual):
While
EDA
doesn't
perform
formal
feature
selection,
how
can
the
insights
gained
from
EDA
(e.g.,
from
correlation
analysis
or
observing
distributions
across
target
classes)
give
you
early
hints
about
which
features
might
be
more
important
for
a
predictive
model?
10.
EDA
Ethics
(Short
Answer):
How
can
EDA
potentially
uncover
biases
in
a
dataset
(e.g.,
gender
or
racial
bias
in
loan
application
data)?
What
responsibility
does
the
analyst
have
when
such
patterns
are
discovered
during
EDA?
(Hints
to
solve
at
the
end
of
this
guide)
FURTHER
LEARNING
Reading
Materials
&
Articles
1.
What
is
Exploratory
Data
Analysis?
(GeeksforGeeks):
β
Link:
https://www.geeksforgeeks.org/what-is-exploratory-data-analysis/
2.
Exploratory
Data
Analysis
(Chapter
from
a
Stat
Book
-
CMU):
β
Link:
https://www.stat.cmu.edu/~hseltman/309/Book/chapter4.pdf
3.
Towards
Data
Science
/
Medium
/
Analytics
Vidhya
/
KDnuggets:
β
Strategic
Consumption:
Search
these
platforms
for
specific
EDA
topics.
β
Towards
Data
Science:
https://towardsdatascience.com/
(Search
for
"EDA,"
"Exploratory
Data
Analysis
Python,"
etc.)
β
KDnuggets:
https://www.kdnuggets.com/
(Search
for
EDA
topics)
4.
Kaggle
Notebooks:
β
Strategic
Consumption:
Analyze
high-ranking
public
notebooks
for
EDA
structure
and
insights.
β
Link:
https://www.kaggle.com/notebooks
(Filter
by
votes
and
search
for
relevant
competitions/datasets)
[page 20]
5.
Academic
Papers
(Google
Scholar,
arXiv,
etc.):
β
Strategic
Consumption
for
AI/GenAI:
Search
for
EDA
applications
in
advanced
AI
topics.
β
Google
Scholar:
https://scholar.google.com/
β
arXiv:
https://arxiv.org/
Recommended
YouTube
Videos
1.
StatQuest
with
Josh
Starmer:
Highly
recommended
for
clear
explanations
of
EDA
and
related
statistical
concepts.
β
Channel
Link:
https://www.youtube.com/c/StatQuestwithJoshStarmer
(Search
within
channel
for
"EDA,"
"PCA,"
"t-SNE")
2.
Krish
Naik:
Practical,
code-along
tutorials
often
using
real-world
datasets.
β
Channel
Link:
https://www.youtube.com/user/krishnaik06
(Search
for
"EDA
playlist,"
"Feature
Engineering")
3.
Corey
Schafer
(Python
Pandas
Tutorial
Series):
Foundational
for
data
manipulation
skills
essential
for
EDA.
β
Channel
Link:
https://www.youtube.com/user/schafer5
(Look
for
his
Pandas
playlist)
Quiz
Answers:
1.
b)
Building
a
production-ready
machine
learning
model.
2.
c)
Box
Plot.
3.
c)
They
have
a
strong
negative
linear
relationship.
4.
c)
Generating
a
contingency
table
(crosstab).
5.
c)
Pandas.
6.
c)
The
asymmetry
of
the
distribution.
7.
c)
Investigate
why
the
data
is
missing
and
then
decide
to
drop
or
impute.
8.
b)
Outliers.
[page 21]
9.
d)
To
graphically
represent
a
matrix
of
values,
like
a
correlation
matrix,
using
color
intensities.
10.
b)
Analyzing
a
single
variable
in
isolation.
Hints
to
solve
Practice
problems:
1.
First
3
EDA
Steps
(Employee
Data)
β
Use
.info()
,
.describe()
,
.isnull()
to
understand
structure.
β
Check
value
distributions
and
missing
data.
β
Group
by
categories
like
'Department'
to
find
patterns.
2.
Outliers
as
Signals
(Domain
Example)
β
In
finance
or
e-commerce,
outliers
can
show
fraud
or
VIPs.
β
Donβt
remove
them
blindlyβcheck
with
boxplots
or
z-scores.
3.
Visualizing
Categorical
vs
Numerical
β
Use
a
boxplot
or
bar
chart
to
compare
age
groups
and
spending.
β
Make
it
clear
and
readable
for
non-technical
viewers.
4.
Handling
Missing
Text
Data
β
Option
1:
Drop
missing
rows
(easy,
but
risky
if
many).
β
Option
2:
Fill
with
placeholders
like
βNo
feedbackβ
for
text
models.
5.
Correlation
&
Pair
Plots
for
Attrition
β
Use
correlation
to
spot
strong
numeric
links.
β
Pair
plots
show
how
features
relate
visually
to
attrition.
6.
Skewed
Income
Distributions
β
Right-skewed
data
can
hurt
model
performance.
β
Try
a
log
transformation
and
recheck
the
histogram.
7.
Device
vs
Subscription
(Categorical
Association)
[page 22]
β
Use
crosstab
and
bar
plots
to
explore
relationships.
β
(Advanced)
Apply
a
Chi-Square
test
to
check
statistical
link.
8.
EDA
for
Time
Series
(Retail
Sales)
β
Use
line
plots
to
spot
trends
and
seasonality.
β
Look
out
for
spikes
during
holidays
or
special
events.
9.
Feature
Importance
Clues
from
EDA
β
Look
for
features
with
strong
target
separation
or
correlation.
β
Low-variance
or
flat
distributions
might
be
less
useful.
10.
EDA
Ethics
&
Bias
β
Compare
group-level
metrics
(e.g.,
approval
rates
by
gender).
β
Report
unfair
patternsβdon't
ignore
them.