# [Optional] Study Guide - EDA

course: Module 1 β€” Foundations of AI & ML
module: Module-1-Foundations-AI-ML
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-1-Foundations-AI-ML/General/[Optional]_Study_Guide_-_EDA.pdf
pages: 22

---
[page 1]
STUDY
 
GUIDE
 
:
 
1.4
 
Exploratory
 
Data
 
Analysis
 
(EDA):
 
Mastering
 
the
 
Art
 
of
 
Data
 
Interrogation
 
 
Welcome
 
to
 
this
 
advanced
 
study
 
guide
 
on
 
Exploratory
 
Data
 
Analysis
 
(EDA).
 
For
 
professionals
 
like
 
you
 
in
 
the
 
PGC
 
GenAI
 
program
 
at
 
IIT
 
Roorkee,
 
powered
 
by
 
Futurense,
 
with
 
substantial
 
technical
 
experience,
 
EDA
 
transcends
 
a
 
mere
 
procedural
 
checklist.
 
It
 
becomes
 
a
 
strategic
 
imperativeβ€”a
 
deeply
 
intellectual
 
process
 
of
 
inquiry,
 
skepticism,
 
and
 
discovery
 
that
 
forms
 
the
 
bedrock
 
of
 
impactful
 
AI
 
and
 
GenAI
 
solutions.
 
Think
 
of
 
yourselves
 
as
 
seasoned
 
investigators
 
arriving
 
at
 
a
 
complex
 
scene;
 
your
 
ability
 
to
 
meticulously
 
examine
 
evidence,
 
question
 
assumptions,
 
and
 
connect
 
disparate
 
clues
 
through
 
EDA
 
will
 
directly
 
determine
 
the
 
success
 
and
 
reliability
 
of
 
the
 
high-stakes
 
models
 
and
 
systems
 
you'll
 
build.
 
This
 
guide
 
is
 
designed
 
to
 
foster
 
that
 
investigative
 
mindset,
 
moving
 
beyond
 
basic
 
techniques
 
to
 
explore
 
the
 
"why,"
 
the
 
"what
 
if,"
 
and
 
the
 
"so
 
what"
 
of
 
your
 
data
 
explorations,
 
especially
 
pertinent
 
for
 
those
 
stepping
 
into
 
or
 
operating
 
in
 
roles
 
like
 
Chief
 
AI
 
Officer,
 
Principal
 
Data
 
Scientist,
 
AI
 
Research
 
Lead,
 
or
 
Senior
 
Manager
 
of
 
AI/ML
 
Solutions.
 
 
🎯
 
Why
 
EDA
 
is
 
Your
 
Indispensable
 
Strategic
 
Compass
 
(Objectives)
 
For
 
senior
 
practitioners,
 
the
 
objectives
 
of
 
EDA
 
gain
 
additional
 
layers
 
of
 
strategic
 
importance:
 
●
 
Profound
 
Data
 
Acumen
 
&
 
Contextual
 
Intelligence
 
(Beyond
 
Surface-Level
 
Understanding):
 
β—‹
 
Elaboration:
 
This
 
isn't
 
just
 
about
 
variable
 
types;
 
it's
 
about
 
understanding
 
data
 
provenance,
 
lineage,
 
and
 
potential
 
biases
 
embedded
 
during
 
collection.
 
It
 
involves
 
questioning
 
the
 
proxies
 
used
 
for
 
real-world
 
constructs
 
(e.g.,
 
is
 
'click-through
 
rate'
 
truly
 
the
 
best
 
proxy
 
for
 
'user
 
engagement'
 
in
 
this
 
specific
 
GenAI
 
application
?).
 
β—‹
 
Advanced
 
Consideration
 
(AI
 
Research
 
Lead):
 
When
 
evaluating
 
datasets
 
for
 
training
 
foundational
 
models,
 
EDA
 
must
 
critically
 
assess
 
not
 
just
 
content,
 
but
 
also
 
representational
 
biases,
 
potential
 
copyright

[page 2]
issues
 
hinted
 
at
 
by
 
data
 
sources,
 
and
 
the
 
ethical
 
implications
 
of
 
the
 
data's
 
origin.
 
β—‹
 
Pitfall
 
to
 
Avoid:
 
Assuming
 
data
 
is
 
"ground
 
truth"
 
without
 
critically
 
examining
 
its
 
collection
 
methodology
 
and
 
inherent
 
limitations.
 
 
●
 
Ensuring
 
Impeccable
 
Data
 
Integrity
 
&
 
Robustness
 
(The
 
Cornerstone
 
of
 
Trustworthy
 
AI):
 
β—‹
 
Elaboration:
 
Beyond
 
simple
 
imputation,
 
this
 
involves
 
understanding
 
the
 
mechanism
 
of
 
missingness
 
(MCAR,
 
MAR,
 
MNAR)
 
to
 
choose
 
an
 
appropriate
 
advanced
 
imputation
 
strategy
 
(e.g.,
 
regression
 
imputation,
 
multiple
 
imputation).
 
For
 
outliers,
 
it's
 
about
 
distinguishing
 
between
 
data
 
errors,
 
genuine
 
extreme
 
values,
 
and
 
"black
 
swan"
 
events.
 
β—‹
 
Advanced
 
Consideration
 
(Principal
 
Data
 
Scientist):
 
Undetected
 
data
 
quality
 
issues
 
(e.g.,
 
subtle
 
data
 
drift,
 
inconsistent
 
labeling
 
in
 
training
 
data
 
for
 
a
 
supervised
 
GenAI
 
model)
 
can
 
lead
 
to
 
models
 
that
 
perform
 
well
 
in
 
testing
 
but
 
fail
 
catastrophically
 
or
 
unfairly
 
in
 
production.
 
EDA
 
is
 
the
 
first
 
line
 
of
 
defense.
 
β—‹
 
Visualization
 
Cue:
 
Imagine
 
using
 
missing
 
data
 
patterns
 
(e.g.,
 
from
 
msno.matrix
 
in
 
Python's
 
missingno
 
library)
 
not
 
just
 
to
 
see
 
where
 
data
 
is
 
missing,
 
but
 
to
 
hypothesize
 
why
,
 
discussing
 
this
 
with
 
data
 
owners.
 
(You
 
can
 
explore
 
the
 
missingno
 
library
 
visuals
 
by
 
searching
 
"missingno
 
python
 
library
 
examples"
 
on
 
the
 
web).
 
●
 
Illuminating
 
Latent
 
Structures,
 
Complex
 
Interactions
 
&
 
Emergent
 
Patterns
 
(The
 
Hunt
 
for
 
Alpha):
 
β—‹
 
Elaboration:
 
This
 
involves
 
looking
 
for
 
non-linear
 
relationships,
 
interaction
 
effects
 
(where
 
the
 
effect
 
of
 
one
 
variable
 
depends
 
on
 
the
 
level
 
of
 
another),
 
and
 
clustering
 
tendencies
 
that
 
might
 
not
 
be
 
obvious
 
from
 
simple
 
bivariate
 
plots.
 
This
 
is
 
where
 
domain
 
expertise
 
combined
 
with
 
advanced
 
visualization
 
shines.
 
β—‹
 
Advanced
 
Consideration
 
(Chief
 
AI
 
Officer):
 
EDA
 
insights
 
can
 
identify
 
opportunities
 
for
 
novel
 
feature
 
engineering
 
that
 
could
 
give
 
your
 
AI
 
models
 
a
 
competitive
 
edge.
 
For
 
instance,
 
discovering
 
a
 
non-linear
 
threshold
 
effect
 
in
 
sensor
 
data
 
could
 
lead
 
to
 
a
 
critical
 
feature
 
for
 
predictive
 
maintenance.

[page 3]
β—‹
 
Analogy
 
for
 
Indian
 
Professionals:
 
Think
 
of
 
it
 
as
 
discerning
 
the
 
subtle,
 
underlying
 
Raga
 
(melodic
 
framework)
 
from
 
a
 
complex
 
musical
 
piece,
 
even
 
when
 
multiple
 
instruments
 
(variables)
 
are
 
playing
 
intricate
 
patterns.
 
 
●
 
Guiding
 
Sophisticated
 
Feature
 
Engineering
 
&
 
Architecting
 
Optimal
 
Models
 
(Strategic
 
Resource
 
Allocation):
 
β—‹
 
Elaboration:
 
EDA
 
informs
 
decisions
 
like
 
whether
 
to
 
use
 
tree-based
 
models
 
(less
 
sensitive
 
to
 
outliers
 
and
 
scaling)
 
or
 
neural
 
networks
 
(which
 
might
 
require
 
careful
 
normalization
 
and
 
outlier
 
treatment).
 
It
 
helps
 
identify
 
when
 
to
 
create
 
interaction
 
terms,
 
polynomial
 
features,
 
or
 
embeddings
 
(especially
 
for
 
GenAI
 
dealing
 
with
 
text/image).
 
β—‹
 
Advanced
 
Consideration:
 
The
 
"shape"
 
of
 
your
 
data
 
(distributions,
 
sparsity,
 
dimensionality)
 
heavily
 
influences
 
the
 
choice
 
of
 
algorithms
 
and
 
the
 
need
 
for
 
techniques
 
like
 
dimensionality
 
reduction
 
(e.g.,
 
PCA,
 
UMAP,
 
t-SNE
 
initially
 
explored
 
during
 
EDA).
 
For
 
GenAI,
 
understanding
 
token
 
distributions,
 
sequence
 
lengths,
 
etc.,
 
is
 
critical
 
for
 
EDA.
 
 
●
 
Rigorous
 
Hypothesis
 
Formulation,
 
Iterative
 
Testing
 
&
 
Assumption
 
Validation
 
(The
 
Scientific
 
Backbone):
 
β—‹
 
Elaboration:
 
Encourage
 
formulating
 
S.M.A.R.T.
 
(Specific,
 
Measurable,
 
Achievable,
 
Relevant,
 
Time-bound)
 
hypotheses
 
during
 
EDA.
 
For
 
example,
 
instead
 
of
 
"Are
 
sales
 
related
 
to
 
marketing
 
spend?",
 
a
 
better
 
hypothesis
 
is
 
"Does
 
a
 
10%
 
increase
 
in
 
digital
 
marketing
 
spend
 
lead
 
to
 
a
 
>5%
 
increase
 
in
 
online
 
sales
 
within
 
the
 
same
 
quarter,
 
holding
 
other
 
factors
 
constant?"
 
EDA
 
provides
 
the
 
initial
 
evidence.
 
β—‹
 
Pitfall
 
to
 
Avoid:
 
Confirmation
 
bias
 
–
 
only
 
looking
 
for
 
evidence
 
that
 
supports
 
pre-existing
 
beliefs.
 
EDA
 
should
 
be
 
an
 
honest
 
exploration.
 
 
●
 
Compelling
 
Narrative
 
Construction
 
&
 
Influential
 
Stakeholder
 
Communication
 
(Driving
 
Action):
 
 
β—‹
 
Elaboration:
 
For
 
senior
 
roles,
 
EDA
 
outputs
 
often
 
feed
 
into
 
strategic
 
decisions.
 
The
 
ability
 
to
 
translate
 
complex
 
data
 
findings
 
into
 
clear,

[page 4]
concise,
 
and
 
actionable
 
narratives
 
for
 
diverse
 
audiences
 
(technical
 
peers,
 
executive
 
leadership,
 
regulatory
 
bodies)
 
is
 
paramount.
 
β—‹
 
Advanced
 
Consideration:
 
Your
 
EDA
 
story
 
should
 
not
 
just
 
present
 
findings
 
but
 
also
 
articulate
 
the
 
uncertainty
 
and
 
limitations
 
of
 
the
 
data
 
and
 
the
 
analysis.
 
 
πŸ—Ί
 
The
 
EDA
 
Expedition:
 
A
 
Detailed
 
&
 
Tactical
 
Itinerary
 
 
 
This
 
expedition
 
requires
 
meticulous
 
planning
 
and
 
execution.
 
1.
 
Problem
 
Articulation,
 
Strategic
 
Alignment
 
&
 
Data
 
Ecosystem
 
Understanding:
 
 
β—‹
 
Elaboration:
 
Start
 
by
 
co-defining
 
the
 
Key
 
Performance
 
Indicators
 
(KPIs)
 
for
 
the
 
project
 
and
 
for
 
the
 
EDA
 
phase
 
itself
 
(e.g.,
 
"Identify
 
top
 
3
 
data
 
quality
 
issues
 
impacting
 
X,"
 
"Characterize
 
user
 
segments
 
based
 
on
 
Y
 
behavior").
 
Map
 
the
 
data
 
sources
 
to
 
the
 
business
 
processes
 
they
 
represent.
 
Understand
 
data
 
governance
 
policies
 
and
 
access
 
constraints
 
early.
 
β—‹
 
Job
 
Role
 
Context
 
(Senior
 
Manager,
 
AI/ML
 
Solutions):
 
This
 
involves
 
extensive
 
stakeholder
 
interviews
 
–
 
from
 
business
 
users
 
to
 
data

[page 5]
engineers
 
–
 
to
 
capture
 
requirements,
 
assumptions,
 
and
 
potential
 
data
 
"gotchas"
 
before
 
a
 
single
 
line
 
of
 
code
 
is
 
written.
 
β—‹
 
Tools
 
for
 
Thought:
 
Use
 
mind
 
maps
 
or
 
influence
 
diagrams
 
to
 
map
 
out
 
variables,
 
expected
 
relationships,
 
and
 
business
 
impacts.
 
 
2.
 
Forensic
 
Data
 
Cleaning
 
&
 
Advanced
 
Preprocessing
 
–
 
Establishing
 
a
 
Gold
 
Standard:
 
 
β—‹
 
Handling
 
Missing
 
Data
 
(
NaN
,
 
None
,
 
Null
):
 
β– 
 
Advanced
 
Imputation:
 
β– 
 
Regression
 
Imputation:
 
Predict
 
missing
 
values
 
using
 
other
 
variables.
 
Pro:
 
Can
 
be
 
accurate.
 
Con:
 
May
 
artificially
 
reduce
 
variance
 
and
 
strengthen
 
correlations.
 
β– 
 
Stochastic
 
Regression
 
Imputation:
 
Adds
 
a
 
random
 
error
 
term
 
to
 
regression
 
predictions
 
to
 
preserve
 
variance.
 
β– 
 
Multiple
 
Imputation
 
(e.g.,
 
MICE
 
-
 
Multivariate
 
Imputation
 
by
 
Chained
 
Equations):
 
Creates
 
multiple
 
complete
 
datasets,
 
runs
 
analysis
 
on
 
each,
 
then
 
pools
 
results.
 
More
 
robust
 
but
 
complex.
 
(Scikit-learn
 
offers
 
IterativeImputer
 
for
 
this:
 
https://scikit-learn.org/stable/modules/generated/sklearn.i
mpute.IterativeImputer.html
)
 
β– 
 
Consideration
 
for
 
GenAI:
 
For
 
sequential
 
data
 
(text,
 
time
 
series),
 
specialized
 
imputation
 
methods
 
that
 
respect
 
sequence
 
order
 
might
 
be
 
needed
 
(e.g.,
 
forward/backward
 
fill,
 
interpolation,
 
or
 
even
 
model-based
 
imputation
 
using
 
LSTMs/Transformers
 
if
 
data
 
is
 
rich
 
enough).
 
β—‹
 
Outlier
 
Detection
 
&
 
Sophisticated
 
Treatment:
 
β– 
 
Robust
 
Methods:
 
Beyond
 
Z-score/IQR,
 
consider
 
methods
 
less
 
sensitive
 
to
 
extreme
 
outliers
 
themselves,
 
like
 
using
 
the
 
Median
 
Absolute
 
Deviation
 
(MAD).
 
Model-based
 
outlier
 
detection
 
(e.g.,
 
isolation
 
forests,
 
one-class
 
SVM)
 
can
 
be
 
part
 
of
 
advanced
 
EDA.
 
(Scikit-learn
 
examples:
 
 
Isolation
 
Forest
 
-

[page 6]
https://scikit-learn.org/stable/modules/generated/sklearn.ensem
ble.IsolationForest.html
 
,
 
 
One-Class
 
SVM
 
-
 
https://scikit-learn.org/stable/modules/generated/sklearn.svm.On
eClassSVM.html
)
 
β– 
 
Philosophy:
 
When
 
is
 
an
 
outlier
 
an
 
error
 
vs.
 
a
 
critical
 
insight?
 
If
 
it's
 
an
 
error,
 
what
 
process
 
generated
 
it?
 
Can
 
this
 
be
 
fixed
 
at
 
the
 
source?
 
If
 
it's
 
a
 
genuine
 
rare
 
event,
 
how
 
should
 
the
 
model
 
be
 
designed
 
to
 
handle
 
it
 
or
 
learn
 
from
 
it?
 
β—‹
 
Data
 
Type
 
Integrity
 
&
 
Semantic
 
Consistency:
 
β– 
 
Elaboration:
 
Beyond
 
astype()
,
 
this
 
includes
 
checking
 
for
 
semantic
 
consistency
 
(e.g.,
 
a
 
'country'
 
column
 
having
 
"USA"
 
and
 
"United
 
States"
 
–
 
requiring
 
normalization).
 
For
 
numerical
 
data,
 
check
 
if
 
units
 
are
 
consistent
 
(e.g.,
 
kgs
 
vs.
 
lbs).
 
β—‹
 
High
 
Cardinality
 
Categorical
 
Variables:
 
EDA
 
needs
 
to
 
identify
 
these
 
early.
 
Strategies
 
include
 
grouping
 
less
 
frequent
 
categories,
 
target
 
encoding
 
(with
 
care
 
to
 
avoid
 
leakage),
 
or
 
embedding
 
techniques.
 
 
3.
 
Deep
 
Univariate
 
Analysis
 
–
 
Profiling
 
Individual
 
Data
 
Personalities:
 
 
β—‹
 
Numerical
 
Variables
 
-
 
Beyond
 
Basic
 
Shapes:
 
β– 
 
Binning
 
Strategies
 
for
 
Histograms:
 
Don't
 
rely
 
on
 
defaults.
 
Conceptually
 
understand
 
rules
 
like
 
Freedman-Diaconis
 
(robust
 
to
 
outliers)
 
or
 
Sturges'
 
Law.
 
The
 
goal
 
is
 
to
 
reveal,
 
not
 
obscure,
 
the
 
underlying
 
distribution.
 
(Pandas
 
hist()
 
or
 
Seaborn
 
histplot()
 
allow
 
bin
 
customization).
 
β– 
 
Interpreting
 
Skewness
 
&
 
Kurtosis
 
Together:
 
β– 
 
High
 
Skew
 
+
 
High
 
Kurtosis:
 
Very
 
asymmetric
 
with
 
many
 
extreme
 
outliers
 
on
 
one
 
side.
 
β– 
 
Low
 
Skew
 
+
 
High
 
Kurtosis:
 
Symmetric
 
but
 
with
 
heavy
 
tails
 
(more
 
outliers
 
than
 
normal).
 
β– 
 
High
 
Skew
 
+
 
Low
 
Kurtosis:
 
Asymmetric
 
but
 
with
 
fewer
 
extreme
 
outliers
 
than
 
expected.
 
(Pandas
 
Series
 
have
 
.skew()
 
and
 
.kurt()
 
methods).

[page 7]
β—‹
 
Categorical
 
Variables
 
-
 
Imbalance
 
and
 
Rarity:
 
β– 
 
Impact
 
of
 
Imbalance:
 
Highly
 
imbalanced
 
categorical
 
features
 
(e.g.,
 
99%
 
Class
 
A,
 
1%
 
Class
 
B)
 
pose
 
challenges
 
for
 
models.
 
EDA
 
must
 
quantify
 
this.
 
β– 
 
Rare
 
Categories:
 
How
 
to
 
handle
 
categories
 
that
 
appear
 
only
 
a
 
few
 
times?
 
(Group,
 
remove,
 
treat
 
as
 
special
 
case).
 
 
4.
 
Nuanced
 
Bivariate/Multivariate
 
Analysis
 
–
 
Deciphering
 
Complex
 
Interplays:
 
 
β—‹
 
Interaction
 
Effects:
 
Actively
 
look
 
for
 
these.
 
E.g.,
 
the
 
impact
 
of
 
'Ad
 
Spend'
 
on
 
'Sales'
 
might
 
differ
 
significantly
 
across
 
'Customer
 
Segments'.
 
Visualized
 
via
 
grouped
 
plots
 
where
 
the
 
trend
 
lines
 
for
 
different
 
groups
 
are
 
not
 
parallel.
 
(Seaborn's
 
lmplot
 
with
 
hue
 
and
 
col
/
row
 
arguments
 
is
 
excellent
 
for
 
this).
 
β—‹
 
Advanced
 
Visualizations
 
(Conceptual
 
with
 
Search
 
Cues):
 
β– 
 
Parallel
 
Coordinate
 
Plots:
 
For
 
visualizing
 
many
 
variables
 
at
 
once.
 
(Search
 
"parallel
 
coordinate
 
plot
 
python
 
pandas"
 
or
 
"parallel
 
coordinate
 
plot
 
plotly"
 
for
 
examples.
 
Plotly
 
offers
 
good
 
interactive
 
versions:
 
https://plotly.com/python/parallel-coordinates-plot/
)
 
β– 
 
Andrews
 
Curves:
 
Represents
 
each
 
multivariate
 
observation
 
as
 
a
 
curve.
 
(Pandas
 
plotting
 
includes
 
Andrews
 
curves:
 
https://pandas.pydata.org/pandas-docs/stable/user_guide/visuali
zation.html#andrews-curves
)
 
β—‹
 
Dimensionality
 
Reduction
 
as
 
EDA:
 
Techniques
 
like
 
Principal
 
Component
 
Analysis
 
(PCA),
 
t-SNE,
 
or
 
UMAP
 
can
 
be
 
used
 
in
 
an
 
exploratory
 
manner
 
to
 
project
 
high-dimensional
 
data
 
into
 
2D
 
or
 
3D
 
for
 
visualization.
 
β– 
 
Visualization
 
Cue
 
(t-SNE/UMAP):
 
Search
 
"t-SNE
 
visualization
 
of
 
MNIST
 
python"
 
or
 
"UMAP
 
for
 
clustering
 
python"
 
for
 
code
 
examples
 
and
 
visual
 
outputs.
 
(Scikit-learn:
 
TSNE
 
-
 
https://scikit-learn.org/stable/modules/generated/sklearn.manifol
d.TSNE.html
,
 
UMAP
 
library:
 
https://umap-learn.readthedocs.io/
)
 
β—‹
 
Statistical
 
Tests
 
(Conceptual
 
Understanding):

[page 8]
β– 
 
MANOVA
 
(Multivariate
 
Analysis
 
of
 
Variance):
 
Conceptually,
 
an
 
extension
 
of
 
ANOVA
 
to
 
multiple
 
dependent
 
numerical
 
variables.
 
(Statsmodels
 
library
 
in
 
Python
 
offers
 
MANOVA:
 
https://www.statsmodels.org/stable/multivariate.html
)
 
β– 
 
Canonical
 
Correlation
 
Analysis
 
(CCA):
 
Explores
 
relationships
 
between
 
two
 
sets
 
of
 
variables.
 
(Scikit-learn:
 
CCA
 
-
 
https://scikit-learn.org/stable/modules/generated/sklearn.cross_d
ecomposition.CCA.html
)
 
 
5.
 
Insight
 
Synthesis
 
&
 
Causal
 
Hypothesis
 
Generation
 
(The
 
"Aha!"
 
Moments):
 
 
β—‹
 
Elaboration:
 
Move
 
beyond
 
observation
 
to
 
interpretation
 
and
 
synthesis.
 
Create
 
an
 
"Insight
 
Log."
 
For
 
senior
 
roles,
 
this
 
includes
 
asking
 
"What
 
are
 
the
 
second-order
 
effects
 
of
 
this
 
finding?"
 
or
 
"How
 
does
 
this
 
challenge
 
our
 
current
 
business
 
assumptions?"
 
β—‹
 
Correlation
 
vs.
 
Causation:
 
EDA
 
will
 
primarily
 
show
 
correlations.
 
Emphasize
 
that
 
correlation
 
does
 
not
 
imply
 
causation.
 
However,
 
EDA
 
can
 
help
 
formulate
 
hypotheses
 
for
 
further
 
causal
 
inference
 
studies
 
(e.g.,
 
A/B
 
testing,
 
quasi-experimental
 
designs).
 
β—‹
 
Example:
 
EDA
 
shows
 
a
 
strong
 
positive
 
correlation
 
between
 
ice
 
cream
 
sales
 
and
 
crime
 
rates.
 
The
 
insight
 
isn't
 
that
 
ice
 
cream
 
causes
 
crime,
 
but
 
that
 
a
 
confounding
 
variable
 
(e.g.,
 
hot
 
weather)
 
likely
 
influences
 
both.
 
 
6.
 
Strategic
 
Reporting
 
&
 
Actionable
 
Recommendations
 
(Driving
 
Change):
 
 
β—‹
 
Tailoring
 
Reports:
 
β– 
 
For
 
Technical
 
Leads/Peers:
 
Detailed
 
plots,
 
statistical
 
test
 
results,
 
discussion
 
of
 
data
 
structures,
 
and
 
code
 
snippets.
 
β– 
 
For
 
C-Suite/Business
 
Stakeholders:
 
High-level
 
summaries,
 
key
 
insights
 
visualized
 
simply
 
(e.g.,
 
a
 
single
 
impactful
 
chart),
 
business
 
implications,
 
and
 
strategic
 
recommendations.
 
Focus
 
on
 
the
 
"so
 
what."

[page 9]
β—‹
 
Quantify
 
Uncertainty:
 
Where
 
possible,
 
express
 
the
 
confidence
 
in
 
findings.
 
Acknowledge
 
limitations
 
of
 
the
 
data
 
or
 
analysis.
 
 
Your
 
EDA
 
Arsenal:
 
Techniques
 
&
 
Visualizations
 
–
 
An
 
Expanded
 
View
 
Non-Graphical
 
EDA
 
(Precision
 
and
 
Robustness):
 
●
 
Descriptive
 
Statistics
 
-
 
Robust
 
Alternatives:
 
β—‹
 
Trimmed
 
Mean/Winsorized
 
Mean:
 
Mitigate
 
outlier
 
impact
 
by
 
removing
 
or
 
adjusting
 
a
 
percentage
 
of
 
extreme
 
values
 
before
 
calculating
 
the
 
mean.
 
(SciPy
 
offers
 
trim_mean
:
 
https://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.trim_m
ean.html
)
 
β—‹
 
Median
 
Absolute
 
Deviation
 
(MAD):
 
Robust
 
measure
 
of
 
dispersion
 
(MAD=median(
∣
X_iβˆ’median(X)
∣
)).
 
(Statsmodels
 
has
 
robust.mad
:
 
https://www.statsmodels.org/dev/generated/statsmodels.robust.scale.m
ad.html
)
 
 
●
 
Correlation
 
Analysis
 
-
 
Beyond
 
Pearson/Spearman:
 
β—‹
 
Distance
 
Correlation:
 
Can
 
detect
 
non-linear
 
relationships.
 
(The
 
dcor
 
library
 
in
 
Python
 
implements
 
this:
 
https://pypi.org/project/dcor/
)
 
β—‹
 
Mutual
 
Information:
 
Measures
 
non-linear
 
dependencies.
 
(Scikit-learn:
 
mutual_info_regression
,
 
mutual_info_classif
 
-
 
https://scikit-learn.org/stable/modules/feature_selection.html#mutual-ior
mation
)
 
 
●
 
Hypothesis
 
Testing
 
during
 
EDA:
 
Simple
 
tests
 
(t-tests,
 
chi-squared)
 
can
 
quickly
 
validate
 
observations.
 
(SciPy.stats:
 
ttest_ind
,
 
chi2_contingency
 
-
 
https://docs.scipy.org/doc/scipy/reference/stats.html
)
 
Graphical
 
EDA
 
(Clarity,
 
Depth,
 
and
 
Nuance
 
in
 
Visuals):
 
●
 
Histograms
 
-
 
Advanced
 
Binning:

[page 10]
β—‹
 
Rules
 
like
 
Freedman-Diaconis,
 
Sturges'
 
Law,
 
Scott's
 
Normal
 
Reference
 
Rule
 
guide
 
bin
 
selection.
 
β—‹
 
Key:
 
Experiment
 
and
 
choose
 
what
 
best
 
reveals
 
structure.
 
 
●
 
Q-Q
 
Plots
 
-
 
Interpreting
 
Deviations:
 
β—‹
 
Deviations
 
indicate
 
non-normality,
 
skewness,
 
heavy/light
 
tails.
 
(Statsmodels
 
qqplot
:
 
https://www.statsmodels.org/dev/generated/statsmodels.graphics.gofpl
ots.qqplot.html
)
 
 
●
 
Pair
 
Plots
 
-
 
Considerations
 
for
 
High
 
Dimensions:
 
β—‹
 
Become
 
unwieldy
 
for
 
many
 
features.
 
(Seaborn
 
pairplot
:
 
https://seaborn.pydata.org/generated/seaborn.pairplot.html
)
 
 
●
 
Mosaic
 
Plots:
 
For
 
visualizing
 
relationships
 
between
 
multiple
 
categorical
 
variables.
 
(Statsmodels
 
mosaic
:
 
https://www.statsmodels.org/dev/generated/statsmodels.graphics.mosaicplot.
mosaic.html
)
 
β—‹
 
Visualization
 
Cue:
 
Search
 
"mosaic
 
plot
 
example
 
titanic
 
dataset
 
python"
 
for
 
visual
 
examples.
 
 
●
 
Residual
 
Plots
 
(after
 
initial
 
simple
 
modeling):
 
Can
 
reveal
 
non-linearity,
 
heteroscedasticity.
 
(Seaborn
 
residplot
:
 
https://seaborn.pydata.org/generated/seaborn.residplot.html
)

[page 11]
Solved
 
Example:

[page 15]
Industry
 
case
 
study
 
for
 
EDA
 
:
   
                  
https://www.linkedin.com/pulse/exploratory-data-analysis-eda-case-study-archana-kh
ewariya-jqquf/
Quizzes
 
(MCQs
 
to
 
Test
 
Your
 
Understanding)
 
1.
 
Which
 
of
 
the
 
following
 
is
 
NOT
 
a
 
primary
 
objective
 
of
 
EDA?
 
a)
 
Identifying
 
missing
 
values
 
and
 
outliers.
 
b)
 
Building
 
a
 
production-ready
 
machine
 
learning
 
model.
 
c)
 
Understanding
 
the
 
relationships
 
between
 
variables.
 
d)
 
Visualizing
 
the
 
distribution
 
of
 
key
 
variables.
 
 
2.
 
You
 
have
 
a
 
dataset
 
of
 
customer
 
transaction
 
amounts
 
and
 
want
 
to
 
quickly
 
see
 
its
 
distribution
 
and
 
identify
 
potential
 
extreme
 
values.
 
Which
 
plot
 
is
 
most
 
suitable?
 
a)
 
Scatter
 
Plot

[page 16]
b)
 
Bar
 
Chart
 
c)
 
Box
 
Plot
 
d)
 
Line
 
Plot
 
 
3.
 
If
 
two
 
numerical
 
variables
 
have
 
a
 
Pearson
 
correlation
 
coefficient
 
of
 
-0.85,
 
it
 
means:
 
a)
 
They
 
have
 
a
 
weak
 
negative
 
linear
 
relationship.
 
b)
 
They
 
have
 
a
 
strong
 
positive
 
linear
 
relationship.
 
c)
 
They
 
have
 
a
 
strong
 
negative
 
linear
 
relationship.
 
d)
 
There
 
is
 
no
 
linear
 
relationship
 
between
 
them.
 
 
4.
 
To
 
analyze
 
the
 
relationship
 
between
 
two
 
categorical
 
variables,
 
"Region"
 
(North,
 
South,
 
East,
 
West)
 
and
 
"Product
 
Preference"
 
(A,
 
B,
 
C),
 
which
 
technique
 
is
 
most
 
appropriate?
 
a)
 
Calculating
 
the
 
mean
 
for
 
each
 
region.
 
b)
 
Creating
 
a
 
scatter
 
plot.
 
c)
 
Generating
 
a
 
contingency
 
table
 
(crosstab).
 
d)
 
Plotting
 
a
 
histogram
 
for
 
product
 
preference.
 
 
5.
 
Which
 
Python
 
library
 
is
 
primarily
 
used
 
for
 
data
 
manipulation
 
and
 
cleaning
 
during
 
EDA?
 
a)
 
Matplotlib
 
b)
 
Seaborn
 
c)
 
Pandas
 
d)
 
SciPy
 
 
6.
 
"Skewness"
 
in
 
a
 
dataset
 
refers
 
to:
 
a)
 
The
 
peakedness
 
of
 
the
 
distribution.
 
b)
 
The
 
measure
 
of
 
how
 
spread
 
out
 
the
 
data
 
is.
 
c)
 
The
 
asymmetry
 
of
 
the
 
distribution.
 
d)
 
The
 
average
 
value
 
of
 
the
 
dataset.
 
 
7.
 
When
 
performing
 
EDA,
 
if
 
you
 
encounter
 
a
 
feature
 
with
 
70%
 
missing
 
values,
 
what
 
is
 
a
 
generally
 
reasonable
 
first
 
approach
 
for
 
an
 
experienced
 
analyst?

[page 17]
a)
 
Immediately
 
drop
 
the
 
feature.
 
b)
 
Impute
 
all
 
missing
 
values
 
with
 
the
 
mean.
 
c)
 
Investigate
 
why
 
the
 
data
 
is
 
missing
 
and
 
then
 
decide
 
to
 
drop
 
or
 
impute.
 
d)
 
Replace
 
missing
 
values
 
with
 
zeros.
 
 
8.
 
A
 
Data
 
Science
 
Architect
 
is
 
examining
 
sensor
 
data
 
from
 
manufacturing
 
equipment.
 
They
 
notice
 
that
 
temperature
 
readings
 
occasionally
 
spike
 
to
 
unrealistic
 
values
 
(e.g.,
 
1000Β°C
 
for
 
a
 
water
 
pipe).
 
These
 
are
 
likely:
 
a)
 
Seasonal
 
trends.
 
b)
 
Outliers.
 
c)
 
Normal
 
variations.
 
d)
 
Missing
 
data.
 
 
9.
 
What
 
is
 
the
 
primary
 
purpose
 
of
 
using
 
a
 
heatmap
 
in
 
EDA?
 
a)
 
To
 
display
 
the
 
distribution
 
of
 
a
 
single
 
numerical
 
variable.
 
b)
 
To
 
visualize
 
the
 
relationship
 
between
 
two
 
categorical
 
variables.
 
c)
 
To
 
show
 
trends
 
over
 
time
 
for
 
multiple
 
variables.
 
d)
 
To
 
graphically
 
represent
 
a
 
matrix
 
of
 
values,
 
like
 
a
 
correlation
 
matrix,
 
using
 
color
 
intensities.
 
 
10.
 
In
 
the
 
context
 
of
 
EDA,
 
"Univariate
 
Analysis"
 
means:
 
a)
 
Analyzing
 
the
 
relationship
 
between
 
two
 
variables.
 
b)
 
Analyzing
 
a
 
single
 
variable
 
in
 
isolation.
 
c)
 
Building
 
a
 
model
 
with
 
only
 
one
 
feature.
 
d)
 
Analyzing
 
data
 
from
 
a
 
single
 
source.
 
(Answers
 
at
 
the
 
end
 
of
 
the
 
guide)
 
 
Practice
 
Questions
 
1.
 
Scenario
 
(Short
 
Answer):
 
You
 
are
 
given
 
a
 
dataset
 
of
 
employee
 
performance
 
with
 
features
 
like
 
'Years_Experience',
 
'Last_Review_Score'
 
(1-5),
 
'Projects_Completed',
 
'Department',
 
and
 
'Salary'.
 
What
 
would
 
be
 
your
 
first
 
3

[page 18]
EDA
 
steps
 
using
 
Pandas,
 
and
 
what
 
specific
 
insights
 
would
 
you
 
be
 
looking
 
for
 
with
 
each?
 
2.
 
Application
 
(Outliers):
 
Describe
 
a
 
situation
 
in
 
a
 
domain
 
you're
 
familiar
 
with
 
(e.g.,
 
e-commerce,
 
finance,
 
manufacturing)
 
where
 
an
 
outlier
 
might
 
be
 
a
 
critical
 
piece
 
of
 
information
 
rather
 
than
 
an
 
error.
 
How
 
would
 
your
 
EDA
 
approach
 
help
 
you
 
distinguish
 
this?
 
3.
 
Scenario
 
(Visualization):
 
You
 
need
 
to
 
explain
 
the
 
relationship
 
between
 
'Customer_Age_Group'
 
(Categorical:
 
Young,
 
Middle-aged,
 
Senior)
 
and
 
'Average_Monthly_Spend'
 
(Numerical)
 
to
 
a
 
non-technical
 
stakeholder.
 
Which
 
visualization
 
would
 
you
 
choose
 
and
 
why?
 
Sketch
 
or
 
describe
 
what
 
it
 
would
 
look
 
like
 
and
 
what
 
story
 
it
 
would
 
tell.
 
4.
 
Data
 
Cleaning
 
(Missing
 
Data):
 
Imagine
 
a
 
'Customer_Feedback_Text'
 
column
 
in
 
your
 
dataset.
 
If
 
10%
 
of
 
these
 
text
 
entries
 
are
 
missing,
 
discuss
 
two
 
different
 
strategies
 
for
 
handling
 
these
 
missing
 
values
 
and
 
the
 
potential
 
pros
 
and
 
cons
 
of
 
each
 
in
 
the
 
context
 
of
 
preparing
 
data
 
for
 
a
 
sentiment
 
analysis
 
model.
 
5.
 
Multivariate
 
Analysis
 
(Job
 
Role
 
Focus
 
-
 
AI
 
Lead):
 
An
 
AI
 
Lead
 
is
 
investigating
 
data
 
for
 
a
 
model
 
to
 
predict
 
employee
 
attrition.
 
They
 
have
 
features
 
like
 
'Job_Satisfaction_Score',
 
'Last_Promotion_Date',
 
'Distance_From_Home',
 
'Monthly_Income'.
 
How
 
could
 
they
 
use
 
a
 
correlation
 
matrix
 
and
 
pair
 
plots
 
to
 
gain
 
initial
 
insights
 
into
 
potential
 
drivers
 
of
 
attrition?
 
What
 
specific
 
patterns
 
might
 
they
 
look
 
for?
 
6.
 
Transformation
 
(Short
 
Answer):
 
You
 
plot
 
a
 
histogram
 
for
 
'Household_Income'
 
and
 
find
 
it's
 
heavily
 
right-skewed.
 
Why
 
might
 
this
 
be
 
a
 
problem
 
for
 
some
 
analytical
 
models,
 
and
 
what
 
common
 
transformation
 
could
 
you
 
apply
 
during
 
EDA
 
to
 
address
 
this?
 
What
 
would
 
you
 
check
 
after
 
applying
 
the
 
transformation?
 
 
7.
 
Categorical
 
Data
 
(Application):
 
You
 
have
 
two
 
categorical
 
variables:
 
'User_Device_Type'
 
(Desktop,
 
Mobile,
 
Tablet)
 
and
 
'Subscription_Plan'
 
(Basic,
 
Premium,
 
Enterprise).
 
How
 
would
 
you
 
explore
 
if
 
there's
 
an
 
association
 
between
 
the
 
device
 
type
 
and
 
the
 
chosen
 
subscription
 
plan?
 
Mention
 
both
 
a
 
statistical
 
test
 
(optional,
 
if
 
known)
 
and
 
a
 
visualization.

[page 19]
8.
 
EDA
 
for
 
Time
 
Series
 
(Scenario):
 
You
 
have
 
daily
 
sales
 
data
 
for
 
a
 
retail
 
store
 
over
 
the
 
past
 
3
 
years.
 
What
 
specific
 
EDA
 
techniques
 
and
 
visualizations
 
would
 
you
 
use
 
to
 
understand
 
seasonality,
 
trends,
 
and
 
any
 
unusual
 
patterns?
 
9.
 
Feature
 
Importance
 
(Conceptual):
 
While
 
EDA
 
doesn't
 
perform
 
formal
 
feature
 
selection,
 
how
 
can
 
the
 
insights
 
gained
 
from
 
EDA
 
(e.g.,
 
from
 
correlation
 
analysis
 
or
 
observing
 
distributions
 
across
 
target
 
classes)
 
give
 
you
 
early
 
hints
 
about
 
which
 
features
 
might
 
be
 
more
 
important
 
for
 
a
 
predictive
 
model?
 
10.
 
EDA
 
Ethics
 
(Short
 
Answer):
 
How
 
can
 
EDA
 
potentially
 
uncover
 
biases
 
in
 
a
 
dataset
 
(e.g.,
 
gender
 
or
 
racial
 
bias
 
in
 
loan
 
application
 
data)?
 
What
 
responsibility
 
does
 
the
 
analyst
 
have
 
when
 
such
 
patterns
 
are
 
discovered
 
during
 
EDA?
 
(Hints
 
to
 
solve
 
at
 
the
 
end
 
of
 
this
 
guide)
 
 
FURTHER
 
LEARNING
 
 
Reading
 
Materials
 
&
 
Articles
 
1.
 
What
 
is
 
Exploratory
 
Data
 
Analysis?
 
(GeeksforGeeks):
 
β—‹
 
Link:
 
https://www.geeksforgeeks.org/what-is-exploratory-data-analysis/
 
2.
 
Exploratory
 
Data
 
Analysis
 
(Chapter
 
from
 
a
 
Stat
 
Book
 
-
 
CMU):
 
β—‹
 
Link:
 
https://www.stat.cmu.edu/~hseltman/309/Book/chapter4.pdf
 
3.
 
Towards
 
Data
 
Science
 
/
 
Medium
 
/
 
Analytics
 
Vidhya
 
/
 
KDnuggets:
 
β—‹
 
Strategic
 
Consumption:
 
Search
 
these
 
platforms
 
for
 
specific
 
EDA
 
topics.
 
β– 
 
Towards
 
Data
 
Science:
 
https://towardsdatascience.com/
 
(Search
 
for
 
"EDA,"
 
"Exploratory
 
Data
 
Analysis
 
Python,"
 
etc.)
 
β– 
 
KDnuggets:
 
https://www.kdnuggets.com/
 
(Search
 
for
 
EDA
 
topics)
 
4.
 
Kaggle
 
Notebooks:
 
β—‹
 
Strategic
 
Consumption:
 
Analyze
 
high-ranking
 
public
 
notebooks
 
for
 
EDA
 
structure
 
and
 
insights.
 
β—‹
 
Link:
 
https://www.kaggle.com/notebooks
 
(Filter
 
by
 
votes
 
and
 
search
 
for
 
relevant
 
competitions/datasets)

[page 20]
5.
 
Academic
 
Papers
 
(Google
 
Scholar,
 
arXiv,
 
etc.):
 
β—‹
 
Strategic
 
Consumption
 
for
 
AI/GenAI:
 
Search
 
for
 
EDA
 
applications
 
in
 
advanced
 
AI
 
topics.
 
β– 
 
Google
 
Scholar:
 
https://scholar.google.com/
 
β– 
 
arXiv:
 
https://arxiv.org/
 
 
 
Recommended
 
YouTube
 
Videos
 
1.
 
StatQuest
 
with
 
Josh
 
Starmer:
 
Highly
 
recommended
 
for
 
clear
 
explanations
 
of
 
EDA
 
and
 
related
 
statistical
 
concepts.
 
β—‹
 
Channel
 
Link:
 
https://www.youtube.com/c/StatQuestwithJoshStarmer
 
(Search
 
within
 
channel
 
for
 
"EDA,"
 
"PCA,"
 
"t-SNE")
 
2.
 
Krish
 
Naik:
 
Practical,
 
code-along
 
tutorials
 
often
 
using
 
real-world
 
datasets.
 
β—‹
 
Channel
 
Link:
 
https://www.youtube.com/user/krishnaik06
 
(Search
 
for
 
"EDA
 
playlist,"
 
"Feature
 
Engineering")
 
3.
 
Corey
 
Schafer
 
(Python
 
Pandas
 
Tutorial
 
Series):
 
Foundational
 
for
 
data
 
manipulation
 
skills
 
essential
 
for
 
EDA.
 
β—‹
 
Channel
 
Link:
 
https://www.youtube.com/user/schafer5
 
(Look
 
for
 
his
 
Pandas
 
playlist)
 
 
Quiz
 
Answers:
 
1.
 
b)
 
Building
 
a
 
production-ready
 
machine
 
learning
 
model.
 
2.
 
c)
 
Box
 
Plot.
 
3.
 
c)
 
They
 
have
 
a
 
strong
 
negative
 
linear
 
relationship.
 
4.
 
c)
 
Generating
 
a
 
contingency
 
table
 
(crosstab).
 
5.
 
c)
 
Pandas.
 
6.
 
c)
 
The
 
asymmetry
 
of
 
the
 
distribution.
 
7.
 
c)
 
Investigate
 
why
 
the
 
data
 
is
 
missing
 
and
 
then
 
decide
 
to
 
drop
 
or
 
impute.
 
8.
 
b)
 
Outliers.

[page 21]
9.
 
d)
 
To
 
graphically
 
represent
 
a
 
matrix
 
of
 
values,
 
like
 
a
 
correlation
 
matrix,
 
using
 
color
 
intensities.
 
10.
 
b)
 
Analyzing
 
a
 
single
 
variable
 
in
 
isolation.
 
Hints
 
to
 
solve
 
Practice
 
problems:
 
1.
 
First
 
3
 
EDA
 
Steps
 
(Employee
 
Data)
 
●
 
Use
 
.info()
,
 
.describe()
,
 
.isnull()
 
to
 
understand
 
structure.
 
●
 
Check
 
value
 
distributions
 
and
 
missing
 
data.
 
●
 
Group
 
by
 
categories
 
like
 
'Department'
 
to
 
find
 
patterns.
 
2.
 
Outliers
 
as
 
Signals
 
(Domain
 
Example)
 
●
 
In
 
finance
 
or
 
e-commerce,
 
outliers
 
can
 
show
 
fraud
 
or
 
VIPs.
 
●
 
Don’t
 
remove
 
them
 
blindlyβ€”check
 
with
 
boxplots
 
or
 
z-scores.
 
3.
 
Visualizing
 
Categorical
 
vs
 
Numerical
 
●
 
Use
 
a
 
boxplot
 
or
 
bar
 
chart
 
to
 
compare
 
age
 
groups
 
and
 
spending.
 
●
 
Make
 
it
 
clear
 
and
 
readable
 
for
 
non-technical
 
viewers.
 
4.
 
Handling
 
Missing
 
Text
 
Data
 
●
 
Option
 
1:
 
Drop
 
missing
 
rows
 
(easy,
 
but
 
risky
 
if
 
many).
 
●
 
Option
 
2:
 
Fill
 
with
 
placeholders
 
like
 
β€œNo
 
feedback”
 
for
 
text
 
models.
 
5.
 
Correlation
 
&
 
Pair
 
Plots
 
for
 
Attrition
 
●
 
Use
 
correlation
 
to
 
spot
 
strong
 
numeric
 
links.
 
●
 
Pair
 
plots
 
show
 
how
 
features
 
relate
 
visually
 
to
 
attrition.
 
6.
 
Skewed
 
Income
 
Distributions
 
●
 
Right-skewed
 
data
 
can
 
hurt
 
model
 
performance.
 
●
 
Try
 
a
 
log
 
transformation
 
and
 
recheck
 
the
 
histogram.
 
7.
 
Device
 
vs
 
Subscription
 
(Categorical
 
Association)

[page 22]
●
 
Use
 
crosstab
 
and
 
bar
 
plots
 
to
 
explore
 
relationships.
 
●
 
(Advanced)
 
Apply
 
a
 
Chi-Square
 
test
 
to
 
check
 
statistical
 
link.
 
8.
 
EDA
 
for
 
Time
 
Series
 
(Retail
 
Sales)
 
●
 
Use
 
line
 
plots
 
to
 
spot
 
trends
 
and
 
seasonality.
 
●
 
Look
 
out
 
for
 
spikes
 
during
 
holidays
 
or
 
special
 
events.
 
9.
 
Feature
 
Importance
 
Clues
 
from
 
EDA
 
●
 
Look
 
for
 
features
 
with
 
strong
 
target
 
separation
 
or
 
correlation.
 
●
 
Low-variance
 
or
 
flat
 
distributions
 
might
 
be
 
less
 
useful.
 
10.
 
EDA
 
Ethics
 
&
 
Bias
 
●
 
Compare
 
group-level
 
metrics
 
(e.g.,
 
approval
 
rates
 
by
 
gender).
 
●
 
Report
 
unfair
 
patternsβ€”don't
 
ignore
 
them.