# [Optional] Study Guide - Data Pre Processing and Curation III

course: Module 1 — Foundations of AI & ML
module: Module-1-Foundations-AI-ML
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-1-Foundations-AI-ML/General/[Optional]_Study_Guide_-_Data_Pre_Processing_and_Curation_III.pdf
pages: 30

---
[page 1]
STUDY
 
GUIDE:
 
1.4
 
Data
 
Preprocessing
 
and
 
Curation
 
III:
 
Advanced
 
Techniques
 
for
 
Generative
 
AI
 
Modules
 
Covered:
 
Dimensionality
 
Reduction,
 
Feature
 
Selection,
 
Data
 
Compression
 
&
 
Reduction,
 
Similarity
 
Measures
 
Module
 
Introduction:
 
Mastering
 
Data
 
Preprocessing
 
and
 
Curation
 
for
 
Advanced
 
Generative
 
AI
 
The
 
Indispensable
 
Role
 
in
 
the
 
GenAI
 
Pipeline
 
1.
 
Clean
 
data
 
is
 
the
 
starting
 
point
 
Generative
 
AI
 
models
 
learn
 
directly
 
from
 
the
 
information
 
you
 
feed
 
them.
 
If
 
the
 
text,
 
images,
 
or
 
code
 
you
 
give
 
them
 
are
 
messy
 
or
 
inconsistent,
 
the
 
model
 
will
 
learn
 
those
 
mistakes.
 
2.
 
Big
 
datasets
 
make
 
problems
 
bigger
 
GenAI
 
often
 
works
 
with
 
thousands
 
of
 
features
 
at
 
once.
 
Extra
 
noise
 
or
 
repeated
 
information
 
slows
 
training
 
and
 
makes
 
it
 
hard
 
to
 
understand
 
why
 
the
 
model
 
behaves
 
a
 
certain
 
way.
 
In
 
tools
 
like
 
Retrieval-Augmented
 
Generation,
 
poor
 
data
 
makes
 
the
 
system
 
pull
 
back
 
the
 
wrong
 
facts
 
and
 
can
 
cause
 
“hallucinated”
 
answers.
 
3.
 
Bad
 
input
 
→
 
bad
 
(and
 
risky)
 
output
 
Biases,
 
errors,
 
or
 
gaps
 
in
 
the
 
data
 
don’t
 
disappear—they
 
get
 
copied
 
and
 
sometimes
 
exaggerated
 
in
 
the
 
model’s
 
responses.
 
That
 
can
 
lead
 
to
 
unfair,
 
confusing,
 
or
 
unreliable
 
answers,
 
which
 
is
 
a
 
serious
 
issue
 
in
 
areas
 
like
 
healthcare
 
or
 
finance.
 
Take-away
 
Spend
 
time
 
on
 
careful,
 
bias-aware
 
preprocessing.
 
It
 
keeps
 
GenAI
 
models
 
faster,
 
clearer,
 
and
 
safer
 
to
 
use.

[page 2]
Chapter
 
1:
 
Dimensionality
 
Reduction:
 
From
 
High-Dimensional
 
Chaos
 
to
 
Actionable
 
Insights
 
1.1.
 
Confronting
 
the
 
"Curse
 
of
 
Dimensionality"
 
in
 
GenAI
 
What
 
it
 
is:
 
When
 
your
 
data
 
has
 
loads
 
of
 
columns
 
(features),
 
everything
 
gets
 
harder—think
 
of
 
trying
 
to
 
spot
 
patterns
 
in
 
a
 
spreadsheet
 
that’s
 
miles
 
wide
.
 
Why
 
it
 
hurts:
 
●
 
Slow
 
&
 
expensive
:
 
More
 
features
 
mean
 
longer
 
training
 
times
 
and
 
bigger
 
computers.
 
●
 
Overfitting
 
risk
:
 
Models
 
latch
 
onto
 
random
 
noise
 
because
 
true
 
patterns
 
are
 
buried.
 
●
 
Hard
 
to
 
pi
cture:
 
You
 
can
 
plot
 
2-D
 
or
 
3-D
 
data,
 
but
 
not
 
10,000-D!
 
●
 
“Empty
 
space”
 
problem
:
 
Points
 
spread
 
so
 
thin
 
that
 
measuring
 
similarity
 
or
 
density
 
becomes
 
unreliable.
 
In
 
Generative
 
AI:
 
●
 
Word,
 
sentence,
 
image
 
embeddings
 
often
 
sit
 
in
 
thousands
 
of
 
dimensions;
 
multimodal
 
models
 
stack
 
even
 
more.
 
●
 
Trillions
 
of
 
parameters
 
amplify
 
the
 
problem—training
 
and
 
debugging
 
would
 
be
 
nearly
 
impossible
 
without
 
trimming
 
dimensions.
 
Fix:
 
Use
 
dimensionality-reduction
 
tools
 
(e.g.,
 
PCA,
 
t-SNE,
 
auto-encoders)
 
to
 
compress
 
data
 
into
 
fewer
 
but
 
still
 
informative
 
dimensions.
 
This
 
keeps
 
training
 
affordable,
 
reduces
 
overfitting,
 
and
 
makes
 
patterns
 
easier
 
to
 
see.
 
 
1.2.
 
Principal
 
Component
 
Analysis
 
(PCA):
 
Unveiling
 
Dominant
 
Patterns

[page 3]
Principal
 
Component
 
Analysis
 
(PCA)
 
is
 
a
 
cornerstone
 
of
 
unsupervised
 
dimensionality
 
reduction.
 
It
 
takes
 
a
 
dataset
 
whose
 
columns
 
are
 
often
 
tangled
 
together
 
and
 
mathematically
 
“rotates”
 
it
 
to
 
produce
 
new,
 
independent
 
axes
 
called
 
principal
 
components.
 
These
 
components
 
are
 
ranked
 
so
 
that
 
the
 
first
 
one
 
explains
 
the
 
largest
 
share
 
of
 
the
 
original
 
data’s
 
patterns,
 
the
 
second
 
explains
 
the
 
next
 
largest
 
share,
 
and
 
so
 
on,
 
letting
 
you
 
keep
 
just
 
the
 
leading
 
few
 
components
 
to
 
work
 
with
 
a
 
much
 
smaller
 
yet
 
still
 
informative
 
version
 
of
 
the
 
data.
 
Algorithmic
 
Steps,
 
Assumptions,
 
Pros
 
&
 
Cons
 
The
 
fundamental
 
goal
 
of
 
PCA
 
is
 
to
 
identify
 
new,
 
uncorrelated
 
features
 
(principal
 
components)
 
that
 
capture
 
the
 
maximum
 
possible
 
variance
 
from
 
the
 
original
 
data.
 
This
 
is
 
achieved
 
by
 
projecting
 
the
 
data
 
onto
 
a
 
lower-dimensional
 
subspace
 
defined
 
by
 
these
 
components.
 
The
 
core
 
Algorithmic
 
Steps
 
involved
 
are:
 
Algorithmic
 
Steps:
 
1.
 
Standardization/Centering:
 
It
 
is
 
crucial
 
to
 
standardize
 
the
 
data
 
(mean
 
of
 
0,
 
standard
 
deviation
 
of
 
1)
 
before
 
applying
 
PCA,
 
especially
 
if
 
features
 
are
 
on
 
different
 
scales.
 
PCA
 
is
 
sensitive
 
to
 
the
 
variance
 
of
 
the
 
initial
 
variables;
 
features
 
with
 
larger
 
ranges
 
would
 
otherwise
 
dominate
 
the
 
principal
 
components.
 
2.
 
Covariance
 
Matrix
 
Computation:
 
Calculate
 
the
 
covariance
 
matrix
 
from
 
the
 
standardized
 
data.
 
3.
 
Eigendecomposition:
 
Compute
 
the
 
eigenvectors
 
and
 
eigenvalues
 
of
 
the
 
covariance
 
matrix.
 
4.
 
Component
 
Selection:
 
Sort
 
the
 
eigenvectors
 
by
 
their
 
corresponding
 
eigenvalues
 
in
 
descending
 
order.
 
Select
 
the
 
top-k
 
eigenvectors
 
(where
 
k
 
is
 
the
 
desired
 
number
 
of
 
dimensions)
 
that
 
capture
 
a
 
significant
 
portion
 
of
 
the
 
total
 
variance
 
(e.g.,
 
95%
 
or
 
99%
 
).
 
5.
 
Data
 
Transformation:
 
Project
 
the
 
original
 
standardized
 
data
 
onto
 
the
 
selected
 
k
 
eigenvectors
 
to
 
obtain
 
the
 
new
 
k-dimensional
 
feature
 
space.

[page 4]
Benefits
 
●
 
Dimensionality
 
cut-down:
 
Fewer
 
features
 
speed
 
up
 
training
 
and
 
reduce
 
memory
 
use.
 
●
 
Noise
 
filtering:
 
Minor
 
PCs
 
often
 
capture
 
noise,
 
so
 
dropping
 
them
 
can
 
improve
 
model
 
generalization.
 
●
 
Visualization:
 
Plotting
 
the
 
first
 
two
 
or
 
three
 
PCs
 
offers
 
a
 
quick
 
glance
 
at
 
structure
 
and
 
outliers.
 
Caveats
 
●
 
Interpretability
 
loss:
 
PCs
 
are
 
linear
 
blends
 
of
 
original
 
features,
 
so
 
their
 
meaning
 
isn’t
 
always
 
intuitive.
 
●
 
Linear
 
assumption:
 
PCA
 
only
 
captures
 
linear
 
relationships;
 
complex
 
non-linear
 
patterns
 
may
 
remain
 
hidden.
 
●
 
Sensitivity
 
to
 
scale
 
and
 
outliers:
 
Always
 
standardize
 
data
 
and
 
inspect
 
for
 
extreme
 
values
 
before
 
applying
 
PCA.
 
PCA
 
Code
 
Example
 
Loads
 
and
 
standardizes
 
the
 
Iris
 
dataset,
 
reduces
 
it
 
to
 
two
 
dimensions
 
with
 
PCA,
 
then
 
plots
 
the
 
samples
 
so
 
you
 
can
 
visually
 
see
 
how
 
much
 
variance
 
the
 
first
 
two
 
components
 
capture
 
and
 
how
 
the
 
three
 
species
 
cluster.

[page 5]
GenAI
 
Applications:
 
Visualizing
 
LLM
 
Embeddings,
 
Feature
 
Engineering
 
for
 
GenAI
 
Turns
 
big
 
numbers
 
into
 
small
 
pictures
 
●
 
LLMs
 
give
 
every
 
sentence
 
thousands
 
of
 
numbers
 
(an
 
embedding).
 
●
 
PCA
 
squeezes
 
those
 
numbers
 
down
 
to
 
just
 
2
 
or
 
3
 
so
 
you
 
can
 
plot
 
them
 
like
 
dots
 
on
 
a
 
graph.
 
Helps
 
you
 
see
 
what
 
the
 
model
 
knows
 
●
 
Similar
 
sentences
 
land
 
close
 
together;
 
weird
 
or
 
wrong
 
ones
 
drift
 
away.\Easy
 
way
 
to
 
check
 
why
 
a
 
chatbot
 
answer
 
feels
 
“off.”
 
Gives
 
lighter
 
inputs
 
to
 
other
 
models
 
●
 
The
 
smaller
 
set
 
of
 
numbers
 
still
 
keeps
 
the
 
main
 
meaning
 
but
 
trains
 
faster
 
and
 
with
 
less
 
noise.
 
Shows
 
what
 
happens
 
inside
 
a
 
big
 
model
 
●
 
Apply
 
PCA
 
to
 
the
 
hidden
 
layers
 
to
 
watch
 
how
 
raw
 
words
 
slowly
 
turn
 
into
 
higher-level
 
ideas.
 
One
 
limitation
 
●
 
PCA
 
only
 
sees
 
straight-line
 
patterns.
 
If
 
you
 
need
 
curved
 
or
 
very
 
subtle
 
relationships,
 
add
 
tools
 
like
 
t-SNE
 
or
 
UMAP.
 
GenAI
 
Job
 
Scenario:
 
An
 
ML
 
Engineer
 
at
 
a
 
GenAI
 
startup
 
in
 
Bangalore
 
needs
 
to
 
debug
 
inconsistent
 
LLM
 
outputs
 
by
 
visualizing
 
text
 
embeddings.
 
●
 
Application:
 
The
 
engineer
 
extracts
 
text
 
embeddings
 
for
 
both
 
problematic
 
(inconsistent)
 
and
 
well-behaved
 
outputs
 
generated
 
by
 
their
 
LLM.
 
PCA
 
is
 
then
 
applied
 
to
 
reduce
 
these
 
high-dimensional
 
embeddings
 
to
 
a
 
2D
 
or
 
3D
 
space
 
for
 
plotting.
 
●
 
Insights
 
Sought:
 
The
 
visualization
 
aims
 
to
 
answer
 
questions
 
such
 
as:
 
Do
 
embeddings
 
corresponding
 
to
 
inconsistent
 
outputs
 
form
 
distinct
 
clusters
 
separate
 
from
 
the

[page 6]
well-behaved
 
ones?
 
Are
 
these
 
problematic
 
embeddings
 
located
 
far
 
from
 
expected
 
semantic
 
regions
 
in
 
the
 
embedding
 
space?
 
Do
 
they
 
exhibit
 
unusually
 
high
 
variance
 
along
 
certain
 
principal
 
components
 
that
 
might
 
not
 
align
 
with
 
semantic
 
meaning?
 
Discovering
 
such
 
patterns
 
can
 
provide
 
valuable
 
clues,
 
guiding
 
prompt
 
refinement
 
strategies,
 
identifying
 
issues
 
in
 
the
 
training
 
data,
 
or
 
indicating
 
the
 
need
 
for
 
further
 
model
 
fine-tuning
 
to
 
improve
 
consistency
 
and
 
reliability.
 
PCA
 
Use
 
case
 
code
-
 
https://github.com/aminfvthi/Principal-Component-Analysis
 
link
 
to
 
external
 
site
 
1.3.
 
t-Distributed
 
Stochastic
 
Neighbor
 
Embedding
 
(t-SNE):
 
Mapping
 
Local
 
Relationships
 
t-SNE
 
is
 
a
 
non-linear
 
dimensionality
 
reduction
 
technique
 
particularly
 
well-suited
 
for
 
visualizing
 
high-dimensional
 
datasets
 
in
 
low-dimensional
 
space
 
(typically
 
2D
 
or
 
3D).
 
It
 
is
 
renowned
 
for
 
its
 
ability
 
to
 
reveal
 
local
 
data
 
structure
 
and
 
create
 
well-separated
 
clusters
.
 
Core
 
Ideas
 
1.
 
Goal
 
—
 
keep
 
neighbors
 
together
 
 
t-SNE
 
rearranges
 
your
 
high-dimensional
 
data
 
into
 
2-D
 
or
 
3-D
 
so
 
that
 
items
 
that
 
were
 
close
 
before
 
stay
 
close
 
and
 
items
 
that
 
were
 
far
 
apart
 
stay
 
(mostly)
 
far
 
apart.
 
2.
 
Step
 
1:
 
measure
 
“closeness”
 
in
 
the
 
original
 
space
 
 
For
 
every
 
pair
 
of
 
points,
 
it
 
turns
 
their
 
distance
 
into
 
a
 
probability
 
of
 
being
 
neighbors
 
(using
 
a
 
bell-curve/Gaussian
 
formula).
 
3.
 
Step
 
2:
 
guess
 
closeness
 
in
 
the
 
low-dimensional
 
map
 
 
It
 
puts
 
the
 
same
 
points
 
on
 
a
 
small
 
canvas
 
and
 
calculates
 
new
 
neighbor
 
probabilities,
 
but
 
this
 
time
 
with
 
a
 
heavy-tailed
 
t-distribution.
 
Those
 
heavy
 
tails
 
give
 
extra
 
room
 
so
 
distant
 
points
 
don’t
 
get
 
squished
 
together.
 
4.
 
Step
 
3:
 
make
 
the
 
two
 
sets
 
of
 
probabilities
 
match
 
 
The
 
algorithm
 
nudges
 
points
 
around
 
to
 
shrink
 
the
 
difference
 
(called
 
KL
 
divergence)
 
between
 
“how
 
close
 
they
 
should
 
be”
 
(original
 
space)
 
and
 
“how
 
close
 
they
 
are
 
now”
 
(map).
 
When
 
that
 
difference
 
is
 
tiny,
 
the
 
map
 
is
 
ready.
 
Result:
 
you
 
get
 
an
 
easy-to-see
 
plot
 
where
 
nearby
 
dots
 
really
 
are
 
similar,
 
and
 
separate
 
clusters
 
stand
 
out
 
clearly.
 
Key
 
Parameters
 
Effective
 
t-SNE
 
visualization
 
heavily
 
depends
 
on
 
careful
 
tuning
 
of
 
its
 
parameters:
 
1.
 
Perplexity
 
(≈
 
neighborhood
 
size)
 
○
 
Pick
 
a
 
value
 
from
 
5
 
–
 
50.
 
○
 
Small
 
value
 
→
 
more
 
tiny
 
clusters,
 
risk
 
of
 
“speckled”
 
plot.
 
○
 
Large
 
value
 
→
 
smoother,
 
big-picture
 
view,
 
but
 
fine
 
detail
 
may
 
disappear.
 
○
 
Must
 
be
 
lower
 
than
 
your
 
total
 
number
 
of
 
points.
 
2.
 
Learning
 
rate
 
(how
 
far
 
points
 
move
 
each
 
step)

[page 7]
○
 
Try
 
10
 
–
 
1000.
 
○
 
Too
 
low
 
→
 
very
 
slow
 
convergence;
 
too
 
high
 
→
 
points
 
shoot
 
past
 
good
 
positions
 
and
 
the
 
plot
 
looks
 
random.
 
3.
 
Iterations
 
(how
 
many
 
optimization
 
steps)
 
○
 
A
 
few
 
hundred
 
to
 
a
 
few
 
thousand
 
usually
 
does
 
the
 
trick.
 
○
 
t-SNE
 
starts
 
with
 
an
 
“early
 
exaggeration”
 
phase
 
that
 
briefly
 
stretches
 
clusters
 
apart
 
so
 
they
 
settle
 
more
 
clearly
 
later.
 
Tweak
 
these
 
three
 
knobs
 
until
 
the
 
clusters
 
in
 
your
 
2-D/3-D
 
plot
 
look
 
stable
 
Algorithmic
 
Overview,
 
Strengths
 
&
 
Limitations
 
for
 
Visualization
 
1.
 
Compute
 
pairwise
 
affinities
 
(conditional
 
probabilities
 
pj
∣
i )
 
in
 
the
 
high-dimensional
 
space
 
using
 
a
 
Gaussian
 
kernel,
 
where
 
the
 
variance
 
for
 
each
 
point
 
is
 
determined
 
by
 
the
 
perplexity.
 
2.
 
Symmetrize
 
these
 
probabilities
 
to
 
get
 
joint
 
probabilities
 
pij .
 
3.
 
Initialize
 
low-dimensional
 
map
 
points
 
randomly
 
(or
 
sometimes
 
with
 
PCA).
 
4.
 
Compute
 
joint
 
probabilities
 
qij 
 
in
 
the
 
low-dimensional
 
space
 
using
 
a
 
Student's
 
t-distribution.
 
5.
 
Minimize
 
the
 
KL
 
divergence
 
between
 
P
 
and
 
Q
 
using
 
gradient
 
descent.
 
Code
 
Example

[page 8]
Loads
 
the
 
Iris
 
data,
 
scales
 
it,
 
compresses
 
it
 
to
 
two
 
dimensions
 
with
 
t-SNE
 
(preserving
 
local
 
neighbourhood
 
relationships),
 
and
 
plots
 
the
 
result
 
so
 
you
 
can
 
visually
 
inspect
 
how
 
the
 
three
 
species
 
separate.
 
 
Benefits
 
●
 
Shows
 
clusters
 
clearly
 
–
 
often
 
uncovers
 
groups
 
that
 
other
 
methods
 
miss.
 
●
 
Catches
 
non-linear
 
patterns
 
–
 
handles
 
curved
 
or
 
tangled
 
relationships
 
that
 
PCA
 
flattens.
 
●
 
Keeps
 
close
 
points
 
close
 
–
 
local
 
neighborhoods
 
are
 
faithfully
 
preserved.
 
Caveats
 
●
 
Slow
 
on
 
big
 
datasets
 
–
 
even
 
the
 
faster
 
versions
 
struggle
 
past
 
~100
 
k
 
points.
 
●
 
Global
 
layout
 
can
 
deceive
 
–
 
distances
 
between
 
clusters
 
and
 
their
 
sizes
 
are
 
not
 
reliable.
 
●
 
Random
 
each
 
run
 
–
 
results
 
change
 
slightly
 
unless
 
you
 
fix
 
the
 
seed.
 
●
 
No
 
easy
 
“add
 
new
 
point”
 
–
 
can’t
 
project
 
fresh
 
data
 
without
 
rerunning
 
the
 
algorithm.
 
●
 
Fussy
 
settings
 
–
 
needs
 
trial-and-error
 
tuning
 
of
 
perplexity
 
and
 
learning
 
rate.
 
GenAI
 
Job
 
Scenario:
 
A
 
Data
 
Scientist
 
at
 
an
 
Indian
 
e-commerce
 
company
 
(e.g.,
 
Flipkart,
 
Myntra)
 
uses
 
t-SNE
 
to
 
understand
 
user
 
clusters
 
from
 
interaction
 
embeddings
 
with
 
a
 
GenAI-powered
 
recommendation
 
system.
 
●
 
Application:
 
The
 
data
 
scientist
 
extracts
 
user
 
embeddings
 
based
 
on
 
a
 
variety
 
of
 
interactions,
 
such
 
as
 
products
 
clicked,
 
items
 
purchased,
 
and
 
queries
 
made
 
to
 
a
 
GenAI-powered
 
shopping
 
assistant.
 
t-SNE
 
is
 
then
 
employed
 
to
 
reduce
 
these
 
high-dimensional
 
user
 
embeddings
 
into
 
a
 
2D
 
or
 
3D
 
space
 
for
 
visualization.
 
●
 
Insights
 
Sought:
 
The
 
goal
 
is
 
to
 
identify
 
if
 
distinct
 
user
 
personas
 
or
 
segments
 
emerge
 
as
 
visually
 
separable
 
clusters.
 
For
 
example,
 
do
 
users
 
who
 
respond
 
positively
 
to

[page 9]
GenAI-driven
 
recommendations
 
cluster
 
differently
 
from
 
those
 
who
 
ignore
 
them?
 
Are
 
there
 
clusters
 
representing
 
bargain
 
hunters,
 
brand
 
loyalists,
 
or
 
users
 
with
 
specific
 
stylistic
 
preferences?
 
Such
 
visualizations
 
can
 
help
 
refine
 
the
 
recommendation
 
strategy,
 
personalize
 
the
 
GenAI
 
assistant's
 
interactions,
 
and
 
tailor
 
marketing
 
campaigns
 
more
 
effectively.
 
1.4.
 
Linear
 
Discriminant
 
Analysis
 
(LDA):
 
Maximizing
 
Class
 
Separability
 
Linear
 
Discriminant
 
Analysis
 
(LDA)
 
is
 
a
 
supervised
 
dimensionality
 
reduction
 
technique.
 
Unlike
 
PCA,
 
which
 
is
 
unsupervised
 
and
 
aims
 
to
 
capture
 
the
 
maximum
 
variance
 
in
 
the
 
data,
 
LDA
 
focuses
 
on
 
finding
 
a
 
feature
 
subspace
 
that
 
maximizes
 
the
 
separability
 
between
 
classes.
 
Foundations,
 
Assumptions,
 
and
 
Comparison
 
to
 
PCA
 
LDA
 
operates
 
by
 
projecting
 
the
 
feature
 
space
 
onto
 
a
 
lower-dimensional
 
space
 
(with
 
at
 
most
 
C−1
 
dimensions,
 
where
 
C
 
is
 
the
 
number
 
of
 
classes)
 
in
 
such
 
a
 
way
 
that
 
the
 
ratio
 
of
 
between-class
 
variance
 
to
 
within-class
 
variance
 
is
 
maximized.This
 
ensures
 
that
 
in
 
the
 
transformed
 
space,
 
data
 
points
 
belonging
 
to
 
different
 
classes
 
are
 
as
 
far
 
apart
 
as
 
possible,
 
while
 
data
 
points
 
within
 
the
 
same
 
class
 
are
 
as
 
close
 
as
 
possible.
 
Assumptions:
 
1.
 
Gaussian
 
Distribution:
 
The
 
data
 
within
 
each
 
class
 
is
 
assumed
 
to
 
be
 
normally
 
(Gaussian)
 
distributed.
 
2.
 
Homoscedasticity
 
(Equal
 
Covariance):
 
All
 
classes
 
are
 
assumed
 
to
 
share
 
the
 
same
 
covariance
 
matrix.
 
If
 
this
 
assumption
 
is
 
violated,
 
Quadratic
 
Discriminant
 
Analysis
 
(QDA)
 
might
 
be
 
more
 
appropriate.
 
3.
 
Feature
 
Independence
 
(often
 
assumed
 
for
 
simplicity,
 
though
 
not
 
strictly
 
required
 
by
 
the
 
math):
 
Features
 
are
 
statistically
 
independent
 
or
 
at
 
least
 
have
 
low
 
multicollinearity.
 
LDA
 
can
 
handle
 
some
 
multicollinearity
 
by
 
transforming
 
data.
 
4.
 
Linear
 
Separability:
 
The
 
technique
 
implicitly
 
assumes
 
that
 
classes
 
are
 
linearly
 
separable
 
to
 
some
 
extent.
 
Comparison
 
to
 
PCA:
 
●
 
Supervision:
 
LDA
 
is
 
a
 
supervised
 
algorithm
 
that
 
uses
 
class
 
labels,
 
whereas
 
PCA
 
is
 
unsupervised
 
and
 
does
 
not
 
use
 
class
 
labels.
 
●
 
Objective:
 
LDA
 
aims
 
to
 
find
 
a
 
subspace
 
that
 
maximizes
 
class
 
separability.
 
PCA
 
aims
 
to
 
find
 
a
 
subspace
 
that
 
maximizes
 
variance.
 
●
 
Application:
 
LDA
 
is
 
primarily
 
used
 
for
 
classification
 
or
 
as
 
a
 
dimensionality
 
reduction
 
step
 
before
 
classification.
 
PCA
 
is
 
used
 
for
 
general
 
dimensionality
 
reduction,
 
data
 
compression,
 
and
 
visualization.
 
●
 
Number
 
of
 
Components:
 
LDA
 
will
 
find
 
at
 
most
 
C−1
 
discriminant
 
dimensions.
 
PCA
 
can
 
find
 
up
 
to
 
d
 
principal
 
components
 
(where
 
d
 
is
 
the
 
original
 
number
 
of
 
features).

[page 10]
Implementation
 
Steps,
 
Advantages
 
&
 
Use
 
Cases
 
Implementation
 
Steps:
 
1.
 
Compute
 
Mean
 
Vectors:
 
Calculate
 
the
 
d-dimensional
 
mean
 
vector
 
μ
i
 
for
 
each
 
class
 
(m1 ,m2 ,...,mC )
 
and
 
the
 
overall
 
mean
 
vector
 
μ.
 
2.
 
Compute
 
Scatter
 
Matrices:
 
              
 
3.
 
Solve
 
Generalized
 
Eigenvalue
 
Problem:
 
Find
 
the
 
eigenvectors
 
and
 
eigenvalues
 
for
 
the
 
equation
 
below.
 
The
 
eigenvectors
 
corresponding
 
to
 
the
 
largest
 
eigenvalues
 
are
 
the
 
linear
 
discriminants
 
(directions
 
for
 
the
 
new
 
subspace).
 
           
 
4.
 
Select
 
Linear
 
Discriminants:
 
Choose
 
the
 
top
 
k
 
eigenvectors
 
(where
 
k≤C−1)
 
to
 
form
 
the
 
transformation
 
matrix
 
W.
 
5.
 
Transform
 
Data:
 
Project
 
the
 
original
 
data
 
onto
 
the
 
new
 
k-dimensional
 
subspace:
 
Y=XW.
 
LDA
 
Code
 
Example
 
Loads
 
and
 
scales
 
the
 
Iris
 
measurements,
 
uses
 
LDA
 
to
 
find
 
two
 
axes
 
that
 
maximise
 
between-species
 
separation,
 
then
 
plots
 
the
 
samples
 
so
 
you
 
can
 
see
 
how
 
cleanly
 
the
 
classes
 
split
 
(
 ∼
99
 
%
 
of
 
the
 
discriminative
 
power
 
lies
 
on
 
LD1).

[page 12]
Advantages:
 
●
 
Effective
 
Class
 
Separation:
 
By
 
design,
 
LDA
 
excels
 
at
 
finding
 
dimensions
 
that
 
maximize
 
discrimination
 
between
 
classes.
 
●
 
Computational
 
Efficiency:
 
It
 
is
 
relatively
 
computationally
 
efficient
 
compared
 
to
 
some
 
non-linear
 
DR
 
techniques.
 
●
 
Performance
 
with
 
Small
 
Samples:
 
LDA
 
can
 
perform
 
well
 
even
 
when
 
the
 
number
 
of
 
samples
 
per
 
class
 
is
 
limited,
 
making
 
it
 
useful
 
in
 
fields
 
like
 
bioinformatics.
 
Use
 
Cases:
 
●
 
Face
 
Recognition:
 
Identifying
 
individuals
 
based
 
on
 
facial
 
features
 
(Fisherfaces).
 
●
 
Medical
 
Diagnosis:
 
Classifying
 
diseases
 
or
 
patient
 
states
 
based
 
on
 
medical
 
data.
 
●
 
Customer
 
Segmentation/Marketing:
 
Identifying
 
distinct
 
customer
 
groups
 
for
 
targeted
 
marketing.
 
●
 
Text
 
Classification:
 
Topic
 
modeling,
 
sentiment
 
analysis,
 
spam
 
detection
 
by
 
projecting
 
text
 
features
 
into
 
a
 
space
 
that
 
separates
 
document
 
categories.
 
●
 
Credit
 
Risk
 
Assessment:
 
Identifying
 
applicants
 
likely
 
to
 
default
 
on
 
loans.
 
GenAI
 
Job
 
Scenario:
 
A
 
GenAI
 
Ethics
 
Officer
 
at
 
a
 
media
 
company
 
in
 
India
 
uses
 
LDA
 
to
 
improve
 
the
 
classification
 
of
 
potentially
 
harmful
 
or
 
biased
 
content
 
generated
 
by
 
an
 
LLM
 
used
 
for
 
news
 
summarization.
 
●
 
Application:
 
The
 
company
 
uses
 
an
 
LLM
 
to
 
generate
 
summaries
 
of
 
news
 
articles.
 
To
 
ensure
 
ethical
 
compliance,
 
these
 
summaries
 
need
 
to
 
be
 
screened
 
for
 
harmful
 
or
 
biased
 
content.
 
Features
 
(e.g.,
 
TF-IDF
 
scores,
 
sentiment
 
scores,
 
or
 
even
 
embeddings
 
from
 
a
 
separate
 
model)
 
are
 
extracted
 
from
 
the
 
generated
 
summaries.
 
A
 
dataset
 
of
 
summaries
 
has
 
been
 
manually
 
labeled
 
as
 
'harmful'
 
or
 
'non-harmful'.
 
The
 
GenAI
 
Ethics
 
Officer
 
applies
 
LDA
 
to
 
this
 
feature
 
set
 
to
 
reduce
 
its
 
dimensionality,
 
specifically
 
aiming
 
to
 
find
 
a
 
projection
 
that
 
best
 
separates
 
the
 
harmful
 
from
 
non-harmful
 
summaries.
 
This
 
reduced
 
feature
 
set
 
is
 
then
 
used
 
to
 
train
 
a
 
classifier
 
(e.g.,
 
SVM
 
or
 
Logistic
 
Regression).
 
●
 
Insights
 
Sought:
 
The
 
officer
 
would
 
analyze
 
which
 
linear
 
combinations
 
of
 
the
 
original
 
features
 
(the
 
discriminants
 
found
 
by
 
LDA)
 
are
 
most
 
effective
 
in
 
distinguishing
 
harmful
 
content.
 
This
 
could
 
reveal
 
key
 
linguistic
 
patterns
 
or
 
semantic
 
elements
 
associated
 
with
 
problematic
 
summaries.
 
Furthermore,
 
they
 
would
 
assess
 
if
 
using
 
LDA
 
as
 
a
 
preprocessing
 
step
 
improves
 
the
 
accuracy,
 
precision,
 
and
 
recall
 
of
 
the
 
downstream
 
harmful
 
content
 
classifier,
 
and
 
potentially
 
makes
 
the
 
classification
 
model
 
simpler
 
and
 
faster
 
to
 
train

[page 13]
Chapter
 
2:
 
Feature
 
Selection:
 
Isolating
 
Critical
 
Signals
 
for
 
Robust
 
GenAI
 
2.1.
 
Why
 
Feature
 
Selection
 
is
 
Paramount
 
in
 
GenAI
 
Systems
 
Think
 
of
 
it
 
like
 
packing
 
for
 
a
 
trip.
 
 
Your
 
suitcase
 
(the
 
model)
 
can’t
 
fit
 
everything
 
in
 
your
 
room
 
(the
 
dataset),
 
so
 
you
 
only
 
take
 
items
 
(features)
 
that
 
will
 
actually
 
be
 
useful
 
at
 
your
 
destination.
 
Leave
 
behind
 
duplicates
 
and
 
“just-in-case”
 
clutter,
 
and
 
the
 
trip
 
is
 
lighter,
 
cheaper,
 
and
 
less
 
stressful.
 
 
What
 
exactly
 
is
 
a
 
“feature”?
 
●
 
In
 
a
 
spreadsheet
 
of
 
house
 
prices,
 
each
 
column—square
 
footage,
 
number
 
of
 
bedrooms,
 
distance
 
to
 
school—is
 
a
 
feature.
 
●
 
For
 
an
 
image
 
model,
 
every
 
pixel
 
value
 
is
 
a
 
feature.
 
●
 
In
 
text,
 
hidden
 
“embedding”
 
numbers
 
that
 
represent
 
meaning
 
become
 
features.
 
A
 
model
 
learns
 
patterns
 
by
 
mixing
 
and
 
matching
 
these
 
columns.
 
Too
 
many
 
weak
 
or
 
irrelevant
 
columns
 
confuse
 
it.
 
Why
 
remove
 
features?
 
1.
 
Avoid
 
over-fitting
 
–
 
The
 
model
 
won’t
 
chase
 
tiny
 
quirks
 
that
 
only
 
appear
 
in
 
the
 
training
 
data.
 
2.
 
Train
 
faster
 
–
 
Fewer
 
numbers
 
=
 
fewer
 
calculations.
 
3.
 
Save
 
money
 
–
 
Less
 
GPU/CPU
 
time
 
and
 
memory.
 
4.
 
Explain
 
results
 
–
 
It
’s
 
eas
ier
 
to
 
show
 
stakeholders
 
which
 
columns
 
really
 
drove
 
a
 
decision.
 
2.2.
 
Filter
 
Methods:
 
Statistical
 
Sieving
 
for
 
Relevance
 
Filter
 
methods
 
for
 
feature
 
selection
 
operate
 
by
 
assessing
 
the
 
relevance
 
of
 
features
 
based
 
on
 
their
 
intrinsic
 
statistical
 
properties,
 
independent
 
of
 
any
 
specific
 
machine
 
learning
 
model.
 
These
 
techniques
 
are
 
typically
 
applied
 
as
 
a
 
preprocessing
 
step
 
before
 
model
 
training.
 
The
 
core
 
idea
 
is
 
to
 
"filter
 
out"
 
less
 
important
 
features
 
based
 
on
 
scores
 
derived
 
from
 
various
 
statistical
 
tests
 
that
 
measure
 
their
 
correlation
 
with
 
the
 
target
 
variable
 
or
 
their
 
individual
 
characteristics.
 
Techniques:
 
●
 
Information
 
Gain:
 
Measures
 
the
 
reduction
 
in
 
entropy
 
(uncertainty)
 
about
 
the
 
target
 
variable
 
when
 
the
 
value
 
of
 
a
 
feature
 
is
 
known.
 
Features
 
that
 
provide
 
more
 
information
 
(higher
 
entropy
 
reduction)
 
are
 
ranked
 
higher.
 
It
 
is
 
commonly
 
used
 
with
 
categorical
 
target
 
variables.
 
●
 
Chi-squared
 
Test
 
(X2):
 
Used
 
primarily
 
for
 
categorical
 
features
 
and
 
a
 
categorical
 
target.
 
It
 
tests
 
the
 
independence
 
between
 
a
 
feature
 
and
 
the
 
target
 
variable.
 
A
 
high
 
X2
 
statistic

[page 14]
(and
 
a
 
low
 
p-value)
 
suggests
 
that
 
the
 
feature
 
is
 
dependent
 
on
 
the
 
target
 
and
 
thus
 
relevant.
 
●
 
ANOVA
 
F-test:
 
Suitable
 
for
 
numerical
 
features
 
and
 
a
 
categorical
 
target.
 
It
 
compares
 
the
 
variance
 
between
 
the
 
means
 
of
 
the
 
feature
 
values
 
across
 
different
 
classes
 
to
 
the
 
variance
 
within
 
each
 
class.
 
A
 
high
 
F-statistic
 
indicates
 
that
 
the
 
feature
 
has
 
different
 
distributions
 
for
 
different
 
classes,
 
making
 
it
 
discriminative.
 
●
 
Pearson's
 
Correlation
 
Coefficient:
 
Measures
 
the
 
linear
 
relationship
 
between
 
two
 
numerical
 
variables
 
(a
 
feature
 
and
 
a
 
numerical
 
target).
 
Values
 
range
 
from
 
-1
 
(perfect
 
negative
 
correlation)
 
to
 
+1
 
(perfect
 
positive
 
correlation).
 
Features
 
with
 
high
 
absolute
 
correlation
 
values
 
are
 
considered
 
more
 
relevant.
 
●
 
Variance
 
Threshold:
 
A
 
simple
 
approach
 
that
 
removes
 
features
 
whose
 
variance
 
does
 
not
 
meet
 
a
 
certain
 
threshold.
 
The
 
assumption
 
is
 
that
 
features
 
with
 
very
 
low
 
variance
 
provide
 
little
 
information
 
for
 
discrimination.
 
By
 
default,
 
it
 
often
 
removes
 
features
 
with
 
zero
 
variance.
 
Pros:
 
●
 
Fast
 
and
 
Computationally
 
Inexpensive:
 
Filter
 
methods
 
do
 
not
 
involve
 
training
 
machine
 
learning
 
models,
 
making
 
them
 
very
 
quick
 
to
 
execute,
 
especially
 
on
 
high-dimensional
 
datasets.
 
●
 
Model-Agnostic:
 
The
 
feature
 
selection
 
is
 
independent
 
of
 
the
 
choice
 
of
 
the
 
learning
 
algorithm,
 
making
 
the
 
selected
 
subset
 
potentially
 
generalizable
 
across
 
different
 
models.
 
●
 
Good
 
for
 
Initial
 
Screening:
 
Effective
 
for
 
an
 
initial
 
culling
 
of
 
features,
 
especially
 
when
 
dealing
 
with
 
a
 
very
 
large
 
number
 
of
 
potential
 
features,
 
to
 
reduce
 
the
 
search
 
space
 
for
 
more
 
complex
 
methods.
 
Cons:
 
●
 
Ignores
 
Feature
 
Interactions:
 
Filter
 
methods
 
evaluate
 
each
 
feature
 
independently
 
and
 
do
 
not
 
consider
 
the
 
combined
 
effect
 
or
 
interaction
 
between
 
features.
 
A
 
feature
 
might
 
be
 
individually
 
weak
 
but
 
highly
 
informative
 
when
 
combined
 
with
 
others.
 
●
 
May
 
Select
 
Redundant
 
Features:
 
Since
 
feature
 
dependencies
 
are
 
often
 
ignored,
 
filter
 
methods
 
might
 
select
 
multiple
 
features
 
that
 
are
 
highly
 
correlated
 
with
 
each
 
other
 
and
 
thus
 
provide
 
redundant
 
information.
 
●
 
Suboptimal
 
for
 
Specific
 
Models:
 
The
 
selected
 
feature
 
subset
 
is
 
not
 
tailored
 
to
 
any
 
specific
 
machine
 
learning
 
model
 
and
 
thus
 
may
 
not
 
be
 
the
 
optimal
 
set
 
for
 
a
 
particular
 
algorithm.
 
GenAI
 
Job
 
Scenario:
 
A
 
Research
 
Scientist
 
at
 
an
 
Indian
 
AI
 
lab
 
(e.g.,
 
working
 
on
 
agricultural
 
GenAI
 
solutions
 
for
 
crop
 
advice
 
)
 
performs
 
initial
 
feature
 
screening
 
from
 
a
 
diverse
 
dataset
 
including
 
weather
 
patterns,
 
soil
 
parameters,
 
historical
 
yields,
 
and
 
market
 
prices.
 
This
 
data
 
will
 
inform
 
a
 
GenAI
 
model
 
designed
 
to
 
predict
 
crop
 
yield
 
and
 
generate
 
tailored
 
farming
 
advice.

[page 15]
●
 
Application:
 
The
 
scientist
 
employs
 
filter
 
methods
 
like
 
Pearson's
 
correlation
 
(for
 
numerical
 
yield
 
targets)
 
and
 
information
 
gain
 
(if
 
advice
 
categories
 
are
 
the
 
target)
 
to
 
quickly
 
assess
 
the
 
individual
 
relevance
 
of
 
numerous
 
environmental,
 
agricultural,
 
and
 
economic
 
factors.
 
●
 
Insights
 
Sought:
 
The
 
primary
 
goal
 
is
 
to
 
identify
 
which
 
factors
 
demonstrate
 
the
 
strongest
 
standalone
 
statistical
 
relationships
 
with
 
crop
 
yield
 
or
 
advice
 
categories.
 
This
 
initial
 
screening
 
helps
 
to
 
narrow
 
down
 
the
 
vast
 
feature
 
set,
 
creating
 
a
 
more
 
manageable
 
input
 
for
 
subsequent,
 
more
 
computationally
 
intensive
 
wrapper
 
or
 
embedded
 
feature
 
selection
 
methods,
 
or
 
directly
 
for
 
the
 
GenAI
 
model
 
if
 
it
 
can
 
handle
 
structured
 
inputs.
 
This
 
ensures
 
that
 
the
 
most
 
promising
 
features
 
are
 
prioritized
 
early
 
in
 
the
 
development
 
pipeline.
 
2.3.
 
Wrapper
 
Methods:
 
Model-Centric
 
Feature
 
Optimization
 
Idea
 
Test
 
different
 
sets
 
of
 
columns
 
by
 
actually
 
training
 
the
 
model
 
each
 
time
 
and
 
keep
 
the
 
set
 
that
 
scores
 
highest.
 
Common
 
strategies
 
●
 
Forward
 
Selection
 
–
 
start
 
empty,
 
add
 
one
 
column
 
at
 
a
 
time
 
if
 
it
 
boosts
 
accuracy.
 
●
 
Backward
 
Elimination
 
–
 
start
 
with
 
all
 
columns,
 
drop
 
the
 
one
 
that
 
hurts
 
accuracy
 
least.
 
●
 
Recursive
 
Feature
 
Elimination
 
(RFE)
 
–
 
train,
 
rank
 
columns
 
by
 
importance,
 
drop
 
the
 
weakest,
 
repeat.
 
Workflow
 
1.
 
Pick
 
(or
 
change)
 
a
 
feature
 
set.
 
2.
 
Train
 
the
 
model,
 
check
 
its
 
score.
 
3.
 
Add
 
or
 
drop
 
a
 
column
 
using
 
one
 
of
 
the
 
strategies.
 
4.
 
Stop
 
when
 
more
 
changes
 
no
 
longer
 
help—or
 
when
 
you
 
reach
 
the
 
desired
 
number
 
of
 
features.
 
Pros
 
●
 
Finds
 
column
 
combinations
 
that
 
really
 
matter
 
to
 
this
 
model.
 
●
 
Often
 
beats
 
filter
 
methods
 
on
 
accuracy.
 
Cons
 
●
 
Slow
 
and
 
compute-hungry
 
(model
 
retrains
 
many
 
times).
 
●
 
Can
 
overfit
 
if
 
the
 
dataset
 
is
 
small
 
or
 
validation
 
is
 
weak.
 
●
 
Best
 
feature
 
set
 
may
 
change
 
if
 
you
 
switch
 
to
 
a
 
different
 
model
 
type.

[page 16]
2.4.
 
Embedded
 
Methods:
 
Integrated
 
Feature
 
Selection
 
What
 
it
 
means
 
The
 
model
 
picks
 
its
 
own
 
important
 
columns
 
while
 
it
 
trains,
 
thanks
 
to
 
built-in
 
penalties
 
or
 
split
 
rules.
 
Popular
 
built-in
 
pickers
 
●
 
LASSO
 
/
 
L1
 
regularization
 
–
 
linear
 
model
 
shrinks
 
some
 
weights
 
to
 
exactly
 
0,
 
so
 
those
 
columns
 
disappear.
 
●
 
Ridge
 
/
 
L2
 
regularization
 
–
 
shrinks
 
weights
 
toward
 
0
 
(helps
 
stability)
 
but
 
rarely
 
kills
 
them
 
outright.
 
●
 
Elastic
 
Net
 
–
 
mixes
 
L1
 
and
 
L2,
 
balancing
 
true
 
selection
 
and
 
stability.
 
●
 
Tree-based
 
models
 
(Random
 
Forest,
 
Gradient
 
Boosting)
 
–
 
at
 
each
 
split
 
the
 
tree
 
chooses
 
the
 
best
 
column;
 
columns
 
used
 
near
 
the
 
top
 
are
 
usually
 
most
 
important.
 
Pros
 
●
 
Fast
 
–
 
one
 
training
 
run;
 
no
 
looping
 
through
 
feature
 
subsets.
 
●
 
Built-in
 
over-fit
 
guard
 
–
 
regularization
 
keeps
 
the
 
model
 
from
 
memorizing
 
noise.
 
●
 
Captures
 
interactions
 
–
 
model
 
decides
 
which
 
combos
 
matter.
 
Cons
 
●
 
Model-dependent
 
–
 
features
 
chosen
 
for
 
a
 
LASSO
 
model
 
might
 
not
 
suit,
 
say,
 
an
 
SVM.
 
●
 
Transparency
 
varies
 
–
 
zeroed
 
LASSO
 
weights
 
are
 
obvious;
 
tree
 
importance
 
scores
 
or
 
deep-net
 
pruning
 
can
 
be
 
harder
 
to
 
explain.
 
2.5.
 
Advanced
 
Applications
 
in
 
GenAI
 
Feature
 
Selection
 
for
 
Multimodal
 
Data
 
(Image,
 
Text,
 
etc.)
 
GenAI
 
models
 
increasingly
 
operate
 
on
 
multimodal
 
data,
 
combining
 
inputs
 
like
 
images,
 
text,
 
audio,
 
and
 
structured
 
data.
 
Feature
 
selection
 
in
 
this
 
context
 
presents
 
unique
 
challenges
 
and
 
opportunities.
 
●
 
Challenges:
 
○
 
Heterogeneity:
 
Different
 
modalities
 
have
 
distinct
 
data
 
types
 
and
 
structures
 
(e.g.,
 
pixel
 
arrays
 
for
 
images,
 
token
 
sequences
 
for
 
text).
 
○
 
Dimensionality
 
Mismatch:
 
Features
 
extracted
 
from
 
different
 
modalities
 
can
 
have
 
vastly
 
different
 
dimensionalities.
 
○
 
Finding
 
Joint
 
Relevance:
 
Identifying
 
features
 
or
 
combinations
 
of
 
features
 
across
 
modalities
 
that
 
are
 
jointly
 
relevant
 
to
 
the
 
GenAI
 
task
 
is
 
complex.
 
●
 
Approaches:
 
○
 
Early
 
Fusion:
 
Features
 
from
 
different
 
modalities
 
are
 
concatenated
 
into
 
a
 
single
 
feature
 
vector
 
early
 
in
 
the
 
pipeline.
 
Standard
 
feature
 
selection
 
techniques
 
can
 
then
 
be
 
applied
 
to
 
this
 
combined
 
vector.
 
However,
 
this
 
approach
 
might
 
not

[page 17]
effectively
 
capture
 
inter-modal
 
relationships
 
and
 
can
 
be
 
dominated
 
by
 
higher-dimensional
 
modalities.
 
○
 
Late
 
Fusion:
 
Separate
 
models
 
are
 
trained
 
for
 
each
 
modality,
 
potentially
 
with
 
modality-specific
 
feature
 
selection.
 
The
 
outputs
 
or
 
decisions
 
from
 
these
 
models
 
are
 
then
 
combined
 
at
 
a
 
later
 
stage.
 
This
 
allows
 
for
 
specialized
 
processing
 
but
 
might
 
miss
 
subtle
 
cross-modal
 
interactions.
 
○
 
Intermediate/Hybrid
 
Fusion:
 
More
 
sophisticated
 
methods
 
aim
 
to
 
learn
 
joint
 
representations
 
or
 
allow
 
interactions
 
between
 
modalities
 
at
 
intermediate
 
layers
 
of
 
a
 
deep
 
learning
 
model.
 
○
 
Deep
 
Learning
 
for
 
Feature
 
Extraction:
 
For
 
modalities
 
like
 
images
 
and
 
text,
 
deep
 
learning
 
models
 
such
 
as
 
Convolutional
 
Neural
 
Networks
 
(CNNs)
 
and
 
Transformers/RNNs
 
are
 
powerful
 
feature
 
extractors.
 
Feature
 
selection
 
can
 
then
 
be
 
applied
 
to
 
these
 
learned
 
high-level
 
features.
 
○
 
Reinforcement
 
Learning
 
(RL)
 
for
 
Dynamic
 
Selection:
 
RL
 
agents
 
can
 
be
 
trained
 
to
 
dynamically
 
select
 
the
 
optimal
 
subset
 
of
 
features
 
from
 
multimodal
 
data
 
by
 
interacting
 
with
 
the
 
environment
 
(the
 
model
 
and
 
data)
 
and
 
receiving
 
rewards
 
based
 
on
 
performance.
 
○
 
Optimization
 
Algorithms:
 
Novel
 
optimization
 
algorithms
 
are
 
being
 
developed
 
specifically
 
for
 
multimodal
 
feature
 
selection,
 
aiming
 
to
 
find
 
the
 
most
 
relevant
 
features
 
across
 
modalities
 
to
 
improve
 
classification
 
or
 
generation
 
performance.
 
Feature
 
Engineering
 
&
 
Selection
 
in
 
LLM
 
Fine-tuning
 
When
 
fine-tuning
 
Large
 
Language
 
Models
 
(LLMs)
 
for
 
tasks
 
that
 
involve
 
structured
 
data
 
in
 
addition
 
to
 
text,
 
the
 
selection
 
and
 
representation
 
of
 
these
 
structured
 
features
 
become
 
critical.
 
●
 
Context:
 
For
 
example,
 
fine-tuning
 
an
 
LLM
 
to
 
predict
 
customer
 
churn
 
might
 
involve
 
using
 
historical
 
transaction
 
data
 
(structured)
 
and
 
customer
 
review
 
texts
 
(unstructured).
 
The
 
LLM
 
needs
 
to
 
effectively
 
integrate
 
information
 
from
 
both
 
sources.
 
●
 
Challenges:
 
Deciding
 
which
 
structured
 
features
 
are
 
relevant,
 
how
 
to
 
encode
 
them
 
(e.g.,
 
numerical
 
scaling,
 
categorical
 
embedding),
 
and
 
how
 
to
 
present
 
them
 
to
 
the
 
LLM
 
alongside
 
the
 
text
 
(e.g.,
 
as
 
part
 
of
 
the
 
prompt,
 
or
 
through
 
dedicated
 
input
 
channels
 
in
 
multimodal
 
architectures)
 
are
 
key
 
challenges.
 
●
 
LLMs
 
for
 
Feature
 
Understanding:
 
An
 
emerging
 
approach
 
involves
 
using
 
LLMs
 
themselves
 
to
 
assess
 
the
 
importance
 
of
 
features,
 
especially
 
when
 
features
 
have
 
descriptive
 
names
 
or
 
textual
 
context.
 
The
 
LLM
 
can
 
be
 
prompted
 
with
 
feature
 
descriptions
 
and
 
task
 
context
 
to
 
provide
 
relevance
 
scores
 
or
 
select
 
features.

[page 18]
Chapter
 
3:
 
Data
 
Compression
 
&
 
Reduction:
 
Making
 
GenAI
 
Models
 
Leaner
 
and
 
Faster
 
Why
 
Compress
 
GenAI
 
Models?
 
Generative
 
AI
 
models,
 
especially
 
Large
 
Language
 
Models
 
(LLMs)
 
like
 
GPT-3,
 
are
 
often
 
huge.
 
GPT-3
 
has
 
175
 
billion
 
parts
 
called
 
parameters!
 
This
 
large
 
size
 
causes
 
problems:
 
●
 
They
 
need
 
a
 
lot
 
of
 
computer
 
memory
 
to
 
store
 
and
 
run.
 
●
 
They
 
require
 
powerful
 
computers,
 
which
 
use
 
a
 
lot
 
of
 
energy
 
and
 
can
 
be
 
expensive.
 
These
 
issues
 
can
 
make
 
it
 
hard
 
to
 
use
 
advanced
 
GenAI
 
everywhere,
 
especially
 
on
 
smaller
 
devices
 
like
 
smartphones
 
or
 
in
 
places
 
with
 
slow
 
internet.
 
Model
 
and
 
data
 
compression
 
techniques
 
help
 
by:
 
●
 
Making
 
models
 
and
 
data
 
smaller.
 
●
 
Speeding
 
up
 
how
 
fast
 
they
 
work.
 
●
 
Using
 
less
 
power.
 
This
 
allows
 
these
 
smart
 
AI
 
systems
 
to
 
be
 
used
 
in
 
more
 
places
 
and
 
on
 
more
 
devices.
 
For
 
LLMs,
 
compression
 
isn't
 
just
 
a
 
nice-to-have;
 
it's
 
often
 
essential
 
to
 
use
 
them
 
outside
 
of
 
big
 
data
 
centers.
 
Making
 
these
 
models
 
smaller
 
and
 
cheaper
 
to
 
run
 
also
 
means
 
more
 
people
 
and
 
organizations
 
can
 
use
 
them,
 
leading
 
to
 
new
 
ideas
 
and
 
applications,
 
like
 
on-device
 
personal
 
assistants.
 
Basic
 
Choices:
 
Lossless
 
vs.
 
Lossy
 
Compression
 
When
 
we
 
compress
 
data,
 
we
 
have
 
two
 
main
 
choices:
 
Lossless
 
Compression:
 
●
 
What
 
it
 
is:
 
This
 
method
 
makes
 
files
 
smaller
 
by
 
finding
 
and
 
removing
 
repetitive
 
information.
 
The
 
best
 
part?
 
You
 
can
 
get
 
the
 
original
 
data
 
back
 
perfectly,
 
with
 
no
 
information
 
lost.
 
●
 
When
 
to
 
use
 
it:
 
Ideal
 
when
 
you
 
can't
 
afford
 
to
 
lose
 
any
 
detail.
 
Think
 
of
 
text
 
documents,
 
computer
 
code,
 
or
 
spreadsheets.
 
●
 
Examples:
 
ZIP
 
files,
 
PNG
 
images,
 
FLAC
 
audio
 
files.
 
●
 
For
 
GenAI:
 
Good
 
for
 
storing
 
training
 
data
 
(especially
 
text),
 
model
 
settings,
 
or
 
the
 
exact
 
text
 
an
 
LLM
 
produces
.
 
Lossy
 
Compression:
 
●
 
What
 
it
 
is:
 
This
 
method
 
makes
 
files
 
much
 
smaller
 
by
 
throwing
 
away
 
some
 
data
 
that's
 
considered
 
less
 
important
 
or
 
hard
 
to
 
notice.
 
You
 
can't
 
get
 
the
 
original
 
data
 
back
 
perfectly;
 
there's
 
some
 
quality
 
loss,
 
but
 
often
 
it's
 
barely
 
noticeable.

[page 19]
●
 
When
 
to
 
use
 
it:
 
Mostly
 
for
 
media
 
like
 
images,
 
audio,
 
and
 
video,
 
where
 
a
 
tiny
 
drop
 
in
 
quality
 
is
 
okay
 
for
 
a
 
big
 
drop
 
in
 
file
 
size.
 
●
 
Examples:
 
JPEG
 
images,
 
MP3
 
audio,
 
MP4
 
videos.
 
●
 
For
 
GenAI:
 
The
 
idea
 
behind
 
lossy
 
compression
 
is
 
very
 
similar
 
to
 
how
 
we
 
compress
 
AI
 
models.
 
Techniques
 
like
 
quantization
 
(using
 
simpler
 
numbers
 
for
 
model
 
parts)
 
or
 
pruning
 
(removing
 
less
 
important
 
model
 
parts)
 
mean
 
some
 
model
 
information
 
is
 
lost
 
to
 
make
 
the
 
model
 
smaller,
 
hoping
 
it
 
still
 
works
 
well.
 
It's
 
also
 
used
 
for
 
compressing
 
large
 
image
 
or
 
video
 
datasets
 
for
 
training
 
GenAI.
 
The
 
main
 
idea
 
in
 
lossy
 
compression
 
–
 
giving
 
up
 
a
 
little
 
bit
 
of
 
quality
 
for
 
a
 
much
 
smaller
 
size
 
–
 
is
 
the
 
same
 
for
 
shrinking
 
AI
 
models.
 
We
 
try
 
to
 
remove
 
the
 
least
 
important
 
bits,
 
whether
 
it's
 
tiny
 
details
 
in
 
a
 
photo
 
or
 
less
 
critical
 
connections
 
in
 
a
 
neural
 
network,
 
to
 
make
 
things
 
smaller
 
and
 
faster
 
without
 
a
 
big
 
drop
 
in
 
how
 
well
 
they
 
work.
 
Shrinking
 
Big
 
AI
 
Models
 
(LLMs
 
and
 
Deep
 
Learning)
 
Modern
 
AI
 
models,
 
especially
 
LLMs,
 
are
 
so
 
big
 
and
 
power-hungry
 
that
 
we
 
need
 
special
 
ways
 
to
 
shrink
 
them.
 
The
 
goal
 
is
 
smaller,
 
faster,
 
more
 
energy-efficient
 
models
 
that
 
still
 
perform
 
well.
 
The
 
main
 
methods
 
are
 
pruning,
 
quantization,
 
and
 
knowledge
 
distillation.
 
Pruning:
 
Trimming
 
the
 
Fat
 
●
 
Concept:
 
Pruning
 
is
 
like
 
trimming
 
off
 
unnecessary
 
branches
 
from
 
a
 
tree.
 
It
 
removes
 
unneeded
 
parts
 
(weights,
 
neurons,
 
or
 
even
 
whole
 
layers)
 
from
 
a
 
trained
 
AI
 
model
 
to
 
make
 
it
 
smaller
 
and
 
faster.
 
The
 
idea
 
is
 
that
 
big
 
models
 
often
 
have
 
many
 
redundant
 
parts.
 
●
 
Types:
 
1.
 
Unstructured
 
Pruning:
 
Removes
 
individual
 
weights.
 
Can
 
make
 
the
 
model
 
very
 
small,
 
but
 
the
 
resulting
 
"sparse"
 
model
 
can
 
be
 
hard
 
for
 
regular
 
computers
 
to
 
speed
 
up.
 
2.
 
Structured
 
Pruning:
 
Removes
 
whole
 
chunks
 
like
 
neurons
 
or
 
layers.
 
This
 
keeps
 
the
 
model's
 
structure
 
neat,
 
making
 
it
 
easier
 
for
 
hardware
 
to
 
run
 
faster.
 
●
 
How
 
it's
 
done:
 
1.
 
Train
 
the
 
full
 
model.
 
2.
 
Figure
 
out
 
which
 
parts
 
are
 
least
 
important
 
(e.g.,
 
weights
 
with
 
small
 
values).
 
3.
 
Remove
 
those
 
parts.
 
4.
 
Fine-tune
 
(retrain
 
a
 
bit)
 
the
 
smaller
 
model
 
to
 
get
 
back
 
any
 
lost
 
performance.
 
●
 
For
 
GenAI:
 
Pruning
 
is
 
useful
 
for
 
LLMs.
 
SparseGPT
 
is
 
one
 
such
 
method.
 
However,
 
it
 
can
 
be
 
tricky,
 
and
 
sometimes
 
accuracy
 
drops
 
more
 
than
 
desired,
 
especially
 
for
 
certain
 
LLM
 
types.
 
Quantization:
 
Using
 
Simpler
 
Numbers
 
●
 
Concept:
 
Instead
 
of
 
using
 
very
 
precise
 
numbers
 
(like
 
32-bit
 
numbers)
 
for
 
a
 
model's
 
weights
 
and
 
calculations,
 
quantization
 
uses
 
simpler,
 
less
 
precise
 
numbers
 
(like
 
8-bit

[page 20]
integers).
 
This
 
is
 
like
 
rounding
 
numbers
 
to
 
make
 
them
 
take
 
up
 
less
 
space.
 
It
 
makes
 
the
 
model
 
smaller,
 
uses
 
less
 
memory,
 
and
 
can
 
speed
 
up
 
calculations
 
on
 
hardware
 
that
 
supports
 
these
 
simpler
 
numbers.
 
●
 
Types:
 
○
 
Quantization-Aware
 
Training
 
(QAT):
 
The
 
model
 
is
 
trained
 
knowing
 
it
 
will
 
use
 
simpler
 
numbers.
 
This
 
often
 
gives
 
better
 
accuracy
 
but
 
needs
 
retraining.
 
○
 
Post-Training
 
Quantization
 
(PTQ):
 
Simpler
 
numbers
 
are
 
applied
 
after
 
the
 
model
 
is
 
already
 
trained.
 
It's
 
faster
 
and
 
easier,
 
very
 
popular
 
for
 
LLMs
 
where
 
retraining
 
is
 
too
 
expensive.
 
●
 
For
 
GenAI:
 
Quantization
 
is
 
vital
 
for
 
LLM
 
compression.
 
Methods
 
like
 
GPTQ
 
can
 
shrink
 
LLMs
 
to
 
use
 
very
 
few
 
bits
 
(e.g.,
 
3
 
or
 
4)
 
with
 
acceptable
 
performance.
 
A
 
challenge
 
is
 
that
 
LLMs
 
sometimes
 
have
 
a
 
few
 
very
 
large,
 
important
 
weights
 
that
 
need
 
careful
 
handling
 
during
 
quantization.
 
Knowledge
 
Distillation:
 
Learning
 
from
 
a
 
"Teacher"
 
●
 
Concept:
 
A
 
smaller
 
"student"
 
model
 
learns
 
from
 
a
 
larger,
 
more
 
capable
 
"teacher"
 
model.
 
The
 
goal
 
is
 
for
 
the
 
student
 
to
 
be
 
almost
 
as
 
smart
 
as
 
the
 
teacher
 
but
 
much
 
smaller
 
and
 
faster.
 
●
 
How
 
it's
 
done:
 
1.
 
Train
 
a
 
big,
 
powerful
 
teacher
 
model.
 
2.
 
The
 
student
 
model
 
is
 
then
 
trained
 
to
 
copy
 
the
 
teacher's
 
outputs.
 
Often,
 
the
 
student
 
learns
 
from
 
the
 
teacher's
 
"soft
 
targets"
 
(the
 
probabilities
 
the
 
teacher
 
assigns
 
before
 
making
 
a
 
final
 
decision),
 
which
 
gives
 
more
 
information
 
than
 
just
 
the
 
final
 
answer.
 
Often,
 
the
 
best
 
results
 
come
 
from
 
using
 
these
 
techniques
 
together,
 
like
 
pruning
 
a
 
model
 
and
 
then
 
quantizing
 
it.
 
The
 
order
 
in
 
which
 
you
 
do
 
them
 
can
 
also
 
make
 
a
 
difference.
 
GenAI
 
Job
 
Scenario:
 
Deploying
 
LLMs
 
on
 
Edge
 
Devices
 
A
 
GenAI
 
Deployment
 
Engineer
 
in
 
India
 
needs
 
to
 
make
 
a
 
sophisticated
 
multilingual
 
LLM
 
run
 
on
 
a
 
small
 
edge
 
device
 
(like
 
a
 
Jetson
 
Nano
 
)
 
for
 
an
 
on-device
 
personal
 
assistant
 
that
 
can
 
summarize
 
and
 
answer
 
questions
 
in
 
various
 
Indian
 
languages.
 
●
 
How
 
they
 
might
 
do
 
it:
 
1.
 
Knowledge
 
Distillation
 
(maybe
 
first):
 
Start
 
with
 
a
 
smaller
 
"student"
 
LLM
 
that
 
learned
 
from
 
a
 
bigger
 
one.
 
2.
 
Pruning:
 
Use
 
structured
 
pruning
 
to
 
make
 
the
 
student
 
LLM
 
even
 
smaller
 
and
 
faster
 
for
 
the
 
hardware.
 
3.
 
Quantization:
 
Apply
 
Post-Training
 
Quantization
 
(PTQ),
 
maybe
 
to
 
8-bit
 
integers,
 
using
 
tools
 
like
 
NVIDIA's
 
TensorRT-LLM
 
or
 
ONNX
 
Runtime
 
to
 
optimize
 
it
 
for
 
the
 
Jetson
 
Nano.

[page 21]
●
 
What
 
they
 
consider:
 
They
 
need
 
to
 
balance
 
model
 
size
 
(to
 
fit
 
on
 
the
 
device),
 
speed
 
(for
 
quick
 
responses),
 
power
 
use
 
(for
 
battery
 
life),
 
and
 
accuracy
 
(good
 
quality
 
answers
 
in
 
all
 
languages).
 
●
 
Goal:
 
Find
 
the
 
best
 
mix
 
of
 
compression
 
that
 
fits
 
the
 
device's
 
limits
 
without
 
making
 
the
 
LLM
 
perform
 
poorly
 
on
 
its
 
main
 
tasks.
 
This
 
means
 
lots
 
of
 
testing
 
on
 
the
 
actual
 
device.
 
Ensuring
 
Compressed
 
Models
 
Are
 
Still
 
Smart:
 
Metrics
 
like
 
SrCr
 
Just
 
making
 
a
 
model
 
smaller
 
isn't
 
enough.
 
We
 
need
 
to
 
ensure
 
it
 
still
 
works
 
well
 
and
 
understands
 
things
 
correctly.
 
Old
 
metrics
 
like
 
"perplexity"
 
(how
 
fluent
 
the
 
text
 
is)
 
don't
 
always
 
catch
 
if
 
a
 
compressed
 
LLM
 
has
 
lost
 
its
 
reasoning
 
ability
 
or
 
factual
 
accuracy.
 
The
 
Semantic
 
Retention
 
Compression
 
Rate
 
(SrCr)
 
is
 
a
 
newer
 
idea
 
to
 
measure
 
this
 
better.
 
●
 
SrCr
 
looks
 
at
 
two
 
things:
 
○
 
How
 
much
 
the
 
model
 
was
 
compressed
 
(Theoretical
 
Compression
 
Rate
 
-
 
TCr).
 
○
 
How
 
well
 
the
 
compressed
 
model
 
still
 
performs
 
on
 
actual
 
tasks
 
compared
 
to
 
the
 
original
 
(Semantic
 
Retention
 
-
 
Sr).
 
●
 
Why
 
it's
 
important:
 
Metrics
 
like
 
SrCr
 
help
 
us
 
focus
 
on
 
making
 
models
 
that
 
are
 
both
 
small
 
and
 
still
 
useful
 
for
 
what
 
they
 
were
 
designed
 
to
 
do.
 
It
 
reminds
 
us
 
to
 
test
 
compressed
 
models
 
on
 
real
 
tasks.
 
For
 
GenAI,
 
especially
 
LLMs,
 
we
 
must
 
check
 
if
 
the
 
compressed
 
version
 
can
 
still
 
reason,
 
summarize,
 
and
 
answer
 
questions
 
accurately.
 
The
 
goal
 
is
 
to
 
keep
 
the
 
"semantic
 
understanding"
 
–
 
the
 
model's
 
ability
 
to
 
grasp
 
meaning
 
and
 
context
 
–
 
even
 
after
 
making
 
it
 
smaller.

[page 22]
Chapter
 
4:
 
Similarity
 
Measures
 
What
 
is
 
Similarity?
 
Often,
 
we
 
represent
 
data
 
(like
 
text,
 
images,
 
or
 
user
 
preferences)
 
as
 
lists
 
of
 
numbers
 
called
 
vectors
 
or
 
embeddings.
 
These
 
vectors
 
live
 
in
 
a
 
multi-dimensional
 
space.
 
If
 
two
 
items
 
are
 
similar
 
in
 
meaning
 
or
 
characteristics,
 
their
 
vectors
 
should
 
be
 
close
 
together
 
in
 
this
 
space.
 
Similarity
 
measures
 
are
 
mathematical
 
tools
 
that
 
tell
 
us
 
exactly
 
how
 
"close"
 
or
 
"similar"
 
these
 
vectors
 
are
 
to
 
each
 
other.
  
 
Being
 
able
 
to
 
measure
 
similarity
 
is
 
key
 
to
 
many
 
data
 
tasks:
 
●
 
Information
 
Retrieval/Search:
 
Finding
 
documents
 
or
 
items
 
that
 
are
 
most
 
similar
 
to
 
a
 
user's
 
query.
  
 
●
 
Recommendation
 
Systems:
 
Suggesting
 
items
 
a
 
user
 
might
 
like
 
based
 
on
 
similarity
 
to
 
what
 
they've
 
liked
 
before,
 
or
 
finding
 
users
 
with
 
similar
 
tastes.
  
 
●
 
Clustering:
 
Grouping
 
similar
 
data
 
points
 
together.
 
●
 
Anomaly
 
Detection:
 
Identifying
 
data
 
points
 
that
 
are
 
very
 
different
 
from
 
others.
 
●
 
In
 
essence,
 
if
 
vectors
 
are
 
the
 
language
 
of
 
our
 
data,
 
similarity
 
measures
 
are
 
the
 
grammar
 
that
 
helps
 
us
 
understand
 
the
 
relationships
 
between
 
them.
 
Euclidean
 
Distance:
 
The
 
Straight
 
Line
 
This
 
is
 
the
 
most
 
straightforward
 
way
 
to
 
measure
 
distance:
 
the
 
shortest,
 
straight-line
 
distance
 
between
 
two
 
points.
  
 
The
 
Formula
 
&
 
Idea
 
Imagine
 
two
 
points,
 
P
 
and
 
Q.
 
If
 
P
 
has
 
coordinates
 
(p1 ,p2 ,...,pn )
 
and
 
Q
 
has
 
(q1 ,q2 ,...,qn )
 
in
 
an
 
n-dimensional
 
space,
 
the
 
Euclidean
 
distance
 
is:

[page 23]
This
 
formula
 
comes
 
from
 
the
 
Pythagorean
 
theorem
 
(like
 
finding
 
the
 
hypotenuse
 
of
 
a
 
triangle).
  
 
Pros
:
 
●
 
Very
 
intuitive
 
–
 
it's
 
how
 
we
 
naturally
 
think
 
about
 
distance.
  
 
●
 
Easy
 
to
 
calculate
.
 
 
 
Cons:
 
●
 
"Curse
 
of
 
Dimensionality":
 
In
 
very
 
high-dimensional
 
spaces
 
(data
 
with
 
many
 
features),
 
Euclidean
 
distance
 
can
 
become
 
less
 
useful.
 
All
 
points
 
might
 
start
 
to
 
look
 
equally
 
far
 
apart.
  
 
●
 
Sensitive
 
to
 
Scale:
 
If
 
your
 
features
 
are
 
on
 
different
 
scales
 
(e.g.,
 
one
 
feature
 
is
 
0-1,
 
another
 
is
 
0-1000),
 
the
 
larger-scale
 
feature
 
will
 
dominate
 
the
 
distance.
 
Always
 
standardize
 
your
 
data
 
(make
 
features
 
have
 
a
 
similar
 
scale,
 
like
 
mean
 
0
 
and
 
standard
 
deviation
 
1)
 
before
 
using
 
Euclidean
 
distance.
  
 
●
 
Magnitude
 
Matters:
 
It
 
considers
 
the
 
"length"
 
or
 
magnitude
 
of
 
the
 
vectors.
 
This
 
might
 
not
 
be
 
what
 
you
 
want
 
if,
 
for
 
example,
 
you're
 
comparing
 
two
 
documents
 
on
 
the
 
same
 
topic
 
but
 
one
 
is
 
much
 
longer
 
than
 
the
 
other.
 
Cosine
 
Similarity:
 
Measuring
 
the
 
Angle
 
Cosine
 
similarity
 
measures
 
how
 
similar
 
two
 
vectors
 
are
 
by
 
looking
 
at
 
the
 
angle
 
between
 
them.
 
It
 
cares
 
about
 
the
 
direction
 
of
 
the
 
vectors,
 
not
 
their
 
length
 
or
 
magnitude.
 
 
 
The
 
Formula
 
&
 
Idea
 
For
 
two
 
vectors,
 
A
 
and
 
B,
 
the
 
cosine
 
similarity
 
is:

[page 24]
Interpretation:
 
●
 
+1:
 
The
 
vectors
 
point
 
in
 
the
 
exact
 
same
 
direction
 
(angle
 
is
 
0°).
 
They
 
are
 
perfectly
 
similar
 
in
 
orientation.
 
●
 
0:
 
The
 
vectors
 
are
 
at
 
a
 
90°
 
angle
 
(orthogonal).
 
They
 
have
 
no
 
similarity
 
in
 
orientation.
 
●
 
-1:
 
The
 
vectors
 
point
 
in
 
opposite
 
directions
 
(angle
 
is
 
180°).
 
They
 
are
 
perfectly
 
dissimilar
 
in
 
orientation.
 
Because
 
it
 
only
 
looks
 
at
 
the
 
angle,
 
it's
 
not
 
affected
 
by
 
how
 
long
 
the
 
vectors
 
are.
  
 
Why
 
It's
 
Good
 
for
 
Certain
 
Data
 
●
 
Magnitude
 
Invariance:
 
This
 
is
 
great
 
when
 
the
 
length
 
of
 
vectors
 
doesn't
 
reflect
 
similarity.
 
For
 
example,
 
in
 
text
 
analysis,
 
a
 
short
 
article
 
and
 
a
 
long
 
essay
 
on
 
the
 
same
 
topic
 
should
 
be
 
considered
 
similar.
 
Cosine
 
similarity
 
can
 
capture
 
this
 
because
 
their
 
content
 
(and
 
thus
 
vector
 
direction)
 
is
 
similar,
 
even
 
if
 
their
 
lengths
 
differ.
  
 
●
 
High
 
Dimensions:
 
It
 
often
 
works
 
better
 
than
 
Euclidean
 
distance
 
in
 
high-dimensional
 
spaces,
 
especially
 
with
 
sparse
 
data
 
(data
 
with
 
many
 
zeros,
 
common
 
in
 
text).
  
 
Jaccard
 
Index:
 
Overlap
 
Between
 
Sets
 
The
 
Jaccard
 
Index
 
(or
 
Jaccard
 
similarity
 
coefficient)
 
measures
 
how
 
similar
 
two
 
sets
 
are
 
by
 
looking
 
at
 
how
 
many
 
elements
 
they
 
share.
  
 
The
 
Formula
 
&
 
Idea
 
For
 
two
 
sets,
 
A
 
and
 
B:

[page 25]
●
 
Range:
 
The
 
Jaccard
 
Index
 
is
 
always
 
between
 
0
 
and
 
1.
  
 
○
 
0:
 
The
 
sets
 
have
 
no
 
elements
 
in
 
common.
 
○
 
1:
 
The
 
sets
 
are
 
exactly
 
the
 
same.
 
Use
 
with
 
Categorical/Binary
 
Data
 
&
 
Text
 
The
 
Jaccard
 
Index
 
is
 
very
 
useful
 
when
 
you're
 
comparing
 
items
 
based
 
on
 
shared
 
characteristics
 
that
 
are
 
either
 
present
 
or
 
absent
 
(binary
 
data),
 
or
 
fall
 
into
 
categories.
 
For
 
example,
 
if
 
A
 
and
 
B
 
are
 
binary
 
vectors
 
(lists
 
of
 
0s
 
and
 
1s):
 
J(A,B)=M01 +M10 +M11 M11  
 
Where:
  
 
●
 
M11 :
 
Number
 
of
 
positions
 
where
 
both
 
A
 
and
 
B
 
have
 
a
 
1.
 
●
 
M01 :
 
Number
 
of
 
positions
 
where
 
A
 
has
 
0
 
and
 
B
 
has
 
1.
 
●
 
M10 :
 
Number
 
of
 
positions
 
where
 
A
 
has
 
1
 
and
 
B
 
has
 
0.
 
 
Choosing
 
the
 
Right
 
Similarity
 
Measure
 
The
 
best
 
similarity
 
measure
 
depends
 
on
 
your
 
data
 
and
 
what
 
you
 
mean
 
by
 
"similar":
 
●
 
Euclidean
 
Distance:
 
Use
 
when
 
the
 
actual
 
values
 
and
 
magnitudes
 
(lengths)
 
of
 
your
 
vectors
 
are
 
important,
 
and
 
your
 
features
 
are
 
on
 
a
 
similar
 
scale
 
(or
 
standardized).
 
Good
 
for
 
dense
 
numerical
 
data
 
where
 
absolute
 
differences
 
matter.
  
 
●
 
Cosine
 
Similarity:
 
Best
 
when
 
the
 
direction
 
(orientation)
 
of
 
vectors
 
is
 
more
 
important
 
than
 
their
 
magnitude.
 
This
 
is
 
often
 
the
 
case
 
for
 
high-dimensional
 
data
 
like
 
text,
 
where
 
you
 
care
 
about
 
meaning
 
(direction)
 
more
 
than
 
document
 
length
 
(magnitude).
  
 
●
 
Jaccard
 
Index:
 
Use
 
when
 
you
 
are
 
comparing
 
sets
 
of
 
items,
 
or
 
features
 
that
 
are
 
binary
 
(yes/no,
 
present/absent).
 
It's
 
about
 
shared
 
presence
 
of
 
elements,
 
not
 
their
 
order
 
or
 
frequency
 
beyond
 
presence.
  
 
Consider
 
if
 
your
 
vectors
 
are
 
normalized
 
(all
 
scaled
 
to
 
have
 
a
 
length
 
of
 
1).
 
If
 
they
 
are,
 
Euclidean
 
distance
 
and
 
cosine
 
similarity
 
become
 
mathematically
 
related,
 
though
 
cosine
 
similarity
 
is
 
often
 
still
 
preferred
 
for
 
its
 
interpretability
 
in
 
terms
 
of
 
angles.
 
For
 
very
 
large
 
datasets,
 
how
 
quickly
 
the
 
measure
 
can
 
be
 
calculated
 
is
 
also
 
important.
 
Code  example

[page 26]
Cosine*  gauges  how  closely  the  two  numeric  vectors  point  in  the  same  direction,  Euclidean  gives  their  raw  
geometric
 
distance,
 
and
 
Jaccard
 
measures
 
overlap
 
between
 
two
 
binary
 
sets
 
(here,
 
word-presence
 
vectors).

[page 27]
Multiple
 
Choice
 
Questions
 
(MCQs)
 
1.
 
The
 
"curse
 
of
 
dimensionality"
 
in
 
the
 
context
 
of
 
GenAI
 
primarily
 
means
 
that:
 
A.
 
Models
 
become
 
too
 
simple
 
and
 
underfit
 
the
 
data.
 
 
B.
 
It's
 
difficult
 
to
 
find
 
enough
 
diverse
 
data
 
for
 
training.
 
 
C.
 
In
 
high-dimensional
 
spaces,
 
data
 
points
 
become
 
sparse,
 
distances
 
less
 
meaningful,
 
and
 
models
 
may
 
overfit
 
noise.
 
 
D.
 
The
 
ethical
 
implications
 
of
 
using
 
many
 
features
 
become
 
unmanageable.
 
 
(Correct
 
Answer:
 
C.
 
Rationale:
 
Section
 
1.1
 
explains
 
that
 
the
 
"curse
 
of
 
dimensionality"
 
leads
 
to
 
slower
 
models,
 
overfitting
 
risk
 
due
 
to
 
sparsity,
 
and
 
difficulty
 
in
 
visualizing
 
patterns
 
because
 
points
 
spread
 
out
 
and
 
distances
 
become
 
less
 
reliable.)
 
 
2.
 
A
 
key
 
difference
 
between
 
Principal
 
Component
 
Analysis
 
(PCA)
 
and
 
Linear
 
Discriminant
 
Analysis
 
(LDA)
 
is:
 
 
A.
 
PCA
 
is
 
supervised
 
and
 
aims
 
to
 
maximize
 
variance,
 
while
 
LDA
 
is
 
unsupervised
 
and
 
aims
 
for
 
class
 
separability.
 
 
B.
 
LDA
 
is
 
supervised
 
and
 
aims
 
to
 
maximize
 
class
 
separability,
 
while
 
PCA
 
is
 
unsupervised
 
and
 
aims
 
to
 
maximize
 
variance.
 
 
C.
 
Both
 
are
 
supervised,
 
but
 
PCA
 
focuses
 
on
 
variance
 
and
 
LDA
 
on
 
feature
 
correlation.
 
 
D.
 
Both
 
are
 
unsupervised,
 
but
 
LDA
 
is
 
better
 
for
 
non-linear
 
data.
 
 
(Correct
 
Answer:
 
B.
 
Rationale:
 
Section
 
1.4
 
states,
 
"Unlike
 
PCA,
 
which
 
is
 
unsupervised
 
and
 
aims
 
to
 
capture
 
the
 
maximum
 
variance
 
in
 
the
 
data,
 
LDA
 
focuses
 
on
 
finding
 
a
 
feature
 
subspace
 
that
 
maximizes
 
the
 
separability
 
between
 
classes."
 
and
 
"LDA
 
is
 
a
 
supervised
 
algorithm
 
that
 
uses
 
class
 
labels,
 
whereas
 
PCA
 
is
 
unsupervised.")
 
 
3.
 
When
 
visualizing
 
high-dimensional
 
data
 
with
 
t-SNE,
 
what
 
is
 
a
 
critical
 
point
 
to
 
remember
 
about
 
interpreting
 
the
 
resulting
 
plot?
 
 
A.
 
The
 
exact
 
distances
 
between
 
clusters
 
and
 
their
 
relative
 
sizes
 
always
 
accurately
 
reflect
 
their
 
true
 
separation
 
in
 
the
 
original
 
high-dimensional
 
space.
 
 
B.
 
t-SNE
 
is
 
primarily
 
useful
 
for
 
linear
 
data
 
structures,
 
similar
 
to
 
PCA.
 
 
C.
 
The
 
global
 
layout,
 
such
 
as
 
distances
 
between
 
clusters
 
and
 
cluster
 
sizes,
 
can
 
be
 
misleading
 
and
 
should
 
be
 
interpreted
 
with
 
caution.
 
 
D.
 
t-SNE
 
is
 
computationally
 
less
 
expensive
 
than
 
PCA
 
for
 
large
 
datasets.
 
 
(Correct
 
Answer:
 
C.
 
Rationale:
 
Section
 
1.3
 
(Caveats)
 
mentions,
 
"Global
 
layout
 
can
 
deceive
 
–
 
distances
 
between
 
clusters
 
and
 
their
 
sizes
 
are
 
not
 
reliable.")
 
 
4.
 
What
 
is
 
the
 
primary
 
benefit
 
of
 
using
 
feature
 
selection
 
techniques
 
in
 
building
 
GenAI
 
models?
 
 
A.
 
To
 
artificially
 
increase
 
the
 
dataset
 
size
 
for
 
better
 
training.
 
 
B.
 
To
 
select
 
a
 
subset
 
of
 
relevant
 
features,
 
which
 
can
 
lead
 
to
 
simpler
 
models,
 
faster
 
training,
 
reduced
 
overfitting,
 
and
 
potentially
 
better
 
accuracy.
 
 
C.
 
To
 
ensure
 
all
 
features
 
are
 
converted
 
to
 
a
 
numerical
 
format.
 
 
D.
 
To
 
always
 
increase
 
the
 
number
 
of
 
dimensions
 
for
 
more
 
complex
 
pattern
 
recognition.
 
 
(Correct
 
Answer:
 
B.
 
Rationale:
 
Section
 
2.1
 
lists
 
benefits
 
like
 
avoiding
 
overfitting,
 
faster

[page 28]
training,
 
saving
 
money
 
(less
 
compute),
 
and
 
easier
 
result
 
explanation.)
 
 
5.
 
Which
 
category
 
of
 
feature
 
selection
 
methods
 
involves
 
training
 
a
 
model
 
multiple
 
times
 
to
 
evaluate
 
different
 
subsets
 
of
 
features
 
based
 
on
 
the
 
model's
 
performance?
 
 
A.
 
Filter
 
methods
 
 
B.
 
Embedded
 
methods
 
 
C.
 
Wrapper
 
methods
 
 
D.
 
Intrinsic
 
methods
 
 
(Correct
 
Answer:
 
C.
 
Rationale:
 
Section
 
2.3
 
(Wrapper
 
Methods
 
Idea)
 
states
 
they
 
"Test
 
different
 
sets
 
of
 
columns
 
by
 
actually
 
training
 
the
 
model
 
each
 
time
 
and
 
keep
 
the
 
set
 
that
 
scores
 
highest.")
 
 
6.
 
LASSO
 
(L1
 
Regularization)
 
is
 
an
 
embedded
 
feature
 
selection
 
method
 
that
 
works
 
by:
 
 
A.
 
Boosting
 
the
 
coefficients
 
of
 
important
 
features.
 
 
B.
 
Assigning
 
scores
 
to
 
features
 
based
 
on
 
statistical
 
tests
 
before
 
model
 
training.
 
 
C.
 
Shrinking
 
the
 
coefficients
 
of
 
some
 
features
 
exactly
 
to
 
zero,
 
effectively
 
removing
 
them
 
from
 
the
 
model.
 
 
D.
 
Creating
 
new
 
features
 
that
 
are
 
linear
 
combinations
 
of
 
the
 
original
 
ones.
 
 
(Correct
 
Answer:
 
C.
 
Rationale:
 
Section
 
2.4
 
(Embedded
 
Methods
 
-
 
LASSO)
 
explains
 
that
 
LASSO
 
"shrinks
 
some
 
weights
 
to
 
exactly
 
0,
 
so
 
those
 
columns
 
disappear.")
 
 
7.
 
In
 
data
 
compression,
 
what
 
is
 
the
 
fundamental
 
difference
 
between
 
lossless
 
and
 
lossy
 
compression?
 
 
A.
 
Lossless
 
compression
 
achieves
 
higher
 
compression
 
ratios
 
than
 
lossy
 
compression.
 
 
B.
 
Lossy
 
compression
 
allows
 
perfect
 
reconstruction
 
of
 
the
 
original
 
data,
 
while
 
lossless
 
does
 
not.
 
 
C.
 
Lossless
 
compression
 
allows
 
perfect
 
reconstruction
 
of
 
the
 
original
 
data,
 
while
 
lossy
 
compression
 
discards
 
some
 
information,
 
meaning
 
the
 
original
 
cannot
 
be
 
perfectly
 
rebuilt.
 
 
D.
 
Lossy
 
compression
 
is
 
only
 
used
 
for
 
text
 
data,
 
and
 
lossless
 
for
 
images.
 
 
(Correct
 
Answer:
 
C.
 
Rationale:
 
Under
 
Chapter
 
3,
 
"Basic
 
Choices:
 
Lossless
 
vs.
 
Lossy
 
Compression"
 
explains
 
lossless
 
allows
 
perfect
 
reconstruction
 
and
 
lossy
 
discards
 
some
 
data.)
 
 
8.
 
Which
 
model
 
compression
 
technique
 
involves
 
training
 
a
 
smaller
 
"student"
 
model
 
to
 
replicate
 
the
 
behavior
 
and
 
knowledge
 
of
 
a
 
larger,
 
more
 
capable
 
"teacher"
 
model?
 
 
A.
 
Pruning
 
 
B.
 
Quantization
 
 
C.
 
Knowledge
 
Distillation
 
 
D.
 
Parameter
 
Sharing
 
 
(Correct
 
Answer:
 
C.
 
Rationale:
 
Under
 
Chapter
 
3,
 
"Knowledge
 
Distillation:
 
Learning
 
from
 
a
 
'Teacher'"
 
describes
 
this
 
process.)
 
 
9.
 
Cosine
 
similarity
 
is
 
often
 
preferred
 
over
 
Euclidean
 
distance
 
for
 
comparing
 
text
 
embeddings
 
primarily
 
because:

[page 29]
A.
 
It
 
is
 
less
 
computationally
 
intensive
 
for
 
all
 
types
 
of
 
data.
 
 
B.
 
It
 
focuses
 
on
 
the
 
magnitude
 
(length)
 
of
 
the
 
vectors,
 
which
 
is
 
crucial
 
for
 
text.
 
 
C.
 
It
 
measures
 
the
 
angle
 
(orientation)
 
between
 
vectors,
 
making
 
it
 
robust
 
to
 
differences
 
in
 
document
 
length
 
while
 
capturing
 
semantic
 
similarity.
 
 
D.
 
Euclidean
 
distance
 
cannot
 
handle
 
sparse
 
data
 
typically
 
found
 
in
 
text
 
embeddings.
 
 
(Correct
 
Answer:
 
C.
 
Rationale:
 
Section
 
4.3
 
(Why
 
It's
 
Good
 
for
 
Text)
 
highlights
 
magnitude
 
invariance
 
and
 
focus
 
on
 
orientation
 
for
 
semantic
 
similarity.)
 
 
10.
 
The
 
Jaccard
 
Index
 
is
 
most
 
appropriately
 
used
 
for
 
measuring
 
similarity
 
between:
 
 
A.
 
Two
 
continuous
 
time-series
 
datasets.
 
 
B.
 
Two
 
sets
 
of
 
items,
 
such
 
as
 
the
 
unique
 
words
 
in
 
two
 
different
 
documents.
 
 
C.
 
The
 
numerical
 
outputs
 
of
 
two
 
different
 
regression
 
models.
 
 
D.
 
Two
 
high-dimensional,
 
dense
 
numerical
 
embeddings
 
where
 
magnitude
 
is
 
important.
 
 
(Correct
 
Answer:
 
B.
 
Rationale:
 
Section
 
4.4
 
(Use
 
with
 
Categorical/Binary
 
Data
 
&
 
Text)
 
states
 
it's
 
useful
 
for
 
comparing
 
items
 
described
 
by
 
categorical/binary
 
features
 
or
 
sets
 
of
 
words.)
 
 
Practice
 
Questions
 
1.
 
Scenario:
 
Interpreting
 
Dimensionality
 
Reduction
 
for
 
LLM
 
Embeddings
 
You
 
are
 
an
 
ML
 
Engineer
 
working
 
on
 
a
 
GenAI
 
application.
 
You've
 
extracted
 
768-dimensional
 
text
 
embeddings
 
from
 
user
 
queries
 
to
 
your
 
chatbot.
 
○
 
(A)
 
If
 
you
 
use
 
PCA
 
to
 
reduce
 
these
 
embeddings
 
to
 
2D
 
for
 
visualization
 
and
 
find
 
that
 
queries
 
with
 
very
 
different
 
intents
 
(e.g.,
 
"book
 
a
 
flight"
 
vs.
 
"tell
 
me
 
a
 
joke")
 
are
 
plotted
 
close
 
together,
 
what
 
are
 
two
 
potential
 
reasons
 
related
 
to
 
PCA's
 
characteristics
 
that
 
could
 
explain
 
this?
 
○
 
(B)
 
What
 
alternative
 
dimensionality
 
reduction
 
technique
 
is
 
often
 
better
 
for
 
visualizing
 
local
 
clusters
 
in
 
such
 
embedding
 
spaces,
 
and
 
what
 
key
 
parameter
 
of
 
this
 
alternative
 
technique
 
would
 
you
 
need
 
to
 
tune
 
carefully?
 
Hint
 
for
 
answering:
 
○
 
For
 
(A),
 
consider
 
PCA's
 
linearity
 
and
 
its
 
objective
 
of
 
maximizing
 
variance.
 
○
 
For
 
(B),
 
think
 
about
 
non-linear
 
techniques
 
known
 
for
 
preserving
 
local
 
structure
 
and
 
their
 
common
 
hyperparameters.
 
2.
 
Application:
 
Feature
 
Selection
 
and
 
Model
 
Compression
 
Strategy
 
Imagine
 
you
 
are
 
tasked
 
with
 
building
 
a
 
GenAI
 
model
 
for
 
a
 
fintech
 
company
 
in
 
India
 
to
 
detect
 
potentially
 
fraudulent
 
transaction
 
descriptions
 
(text
 
data)
 
augmented
 
with
 
structured
 
transaction
 
metadata
 
(e.g.,
 
transaction
 
amount,
 
time
 
of
 
day,
 
user
 
location
 
category).
 
○
 
(A)
 
From
 
the
 
structured
 
metadata,
 
how
 
would
 
you
 
select
 
the
 
most
 
relevant
 
features
 
to
 
combine
 
with
 
the
 
text
 
analysis?
 
Describe
 
one
 
type
 
of
 
feature
 
selection

[page 30]
method
 
(Filter,
 
Wrapper,
 
or
 
Embedded)
 
you
 
might
 
use
 
and
 
why
 
it
 
would
 
be
 
appropriate.
 
○
 
(B)
 
If
 
the
 
final
 
GenAI
 
model
 
(after
 
incorporating
 
text
 
and
 
selected
 
structured
 
features)
 
is
 
too
 
large
 
and
 
slow
 
for
 
real-time
 
fraud
 
alerts,
 
name
 
two
 
distinct
 
model
 
compression
 
techniques
 
you
 
would
 
consider
 
applying
 
and
 
briefly
 
state
 
the
 
core
 
idea
 
of
 
each.
 
Hint
 
for
 
answering:
 
○
 
For
 
(A),
 
consider
 
the
 
trade-offs
 
(speed,
 
interaction
 
capture,
 
model
 
dependency)
 
of
 
different
 
feature
 
selection
 
categories.
 
○
 
For
 
(B),
 
recall
 
the
 
main
 
approaches
 
to
 
making
 
models
 
smaller
 
and
 
faster
 
(e.g.,
 
removing
 
parts,
 
simplifying
 
numbers,
 
learning
 
from
 
a
 
bigger
 
model).
 
3.
 
Conceptual:
 
Choosing
 
Similarity
 
Measures
 
Explain
 
a
 
specific
 
scenario
 
or
 
type
 
of
 
data
 
where
 
Euclidean
 
distance
 
would
 
be
 
a
 
more
 
suitable
 
similarity
 
measure
 
than
 
cosine
 
similarity,
 
and
 
vice-versa.
 
Justify
 
your
 
choices
 
by
 
highlighting
 
the
 
key
 
properties
 
of
 
each
 
measure.
 
 
Hint
 
for
 
answering:
 
 
○
 
Think
 
about
 
when
 
the
 
magnitude/length
 
of
 
vectors
 
is
 
important
 
versus
 
when
 
only
 
the
 
direction/orientation
 
matters.
 
Further
 
Learning
 
1.
 
Reference
 
Youtube
 
Videos
 
a.
 
 
Dimensionality Reduction | ML-005 Lecture 14 | Stanford University | Andre…
link
 
to
 
external
 
site
 
b.
 
 
link
 
to
 
Feature Selection Techniques Easily Explained | Machine Learning
external
 
site
 
2.
 
Reference
 
Websites
 
a.
 
https://www.geeksforgeeks.org/dimensionality-reduction/
  
link
 
to
 
external
 
site
 
b.
 
https://www.ibm.com/think/topics/feature-selection
 
link
 
to
 
external
 
site
 
c.
 
https://www.geeksforgeeks.org/data-reduction-in-data-mining/
 
link
 
to
 
external
 
site