# [Optional] Study Guide - Data Pre Processing and Curation III
course: Module 1 — Foundations of AI & ML
module: Module-1-Foundations-AI-ML
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-1-Foundations-AI-ML/General/[Optional]_Study_Guide_-_Data_Pre_Processing_and_Curation_III.pdf
pages: 30
---
[page 1]
STUDY
GUIDE:
1.4
Data
Preprocessing
and
Curation
III:
Advanced
Techniques
for
Generative
AI
Modules
Covered:
Dimensionality
Reduction,
Feature
Selection,
Data
Compression
&
Reduction,
Similarity
Measures
Module
Introduction:
Mastering
Data
Preprocessing
and
Curation
for
Advanced
Generative
AI
The
Indispensable
Role
in
the
GenAI
Pipeline
1.
Clean
data
is
the
starting
point
Generative
AI
models
learn
directly
from
the
information
you
feed
them.
If
the
text,
images,
or
code
you
give
them
are
messy
or
inconsistent,
the
model
will
learn
those
mistakes.
2.
Big
datasets
make
problems
bigger
GenAI
often
works
with
thousands
of
features
at
once.
Extra
noise
or
repeated
information
slows
training
and
makes
it
hard
to
understand
why
the
model
behaves
a
certain
way.
In
tools
like
Retrieval-Augmented
Generation,
poor
data
makes
the
system
pull
back
the
wrong
facts
and
can
cause
“hallucinated”
answers.
3.
Bad
input
→
bad
(and
risky)
output
Biases,
errors,
or
gaps
in
the
data
don’t
disappear—they
get
copied
and
sometimes
exaggerated
in
the
model’s
responses.
That
can
lead
to
unfair,
confusing,
or
unreliable
answers,
which
is
a
serious
issue
in
areas
like
healthcare
or
finance.
Take-away
Spend
time
on
careful,
bias-aware
preprocessing.
It
keeps
GenAI
models
faster,
clearer,
and
safer
to
use.
[page 2]
Chapter
1:
Dimensionality
Reduction:
From
High-Dimensional
Chaos
to
Actionable
Insights
1.1.
Confronting
the
"Curse
of
Dimensionality"
in
GenAI
What
it
is:
When
your
data
has
loads
of
columns
(features),
everything
gets
harder—think
of
trying
to
spot
patterns
in
a
spreadsheet
that’s
miles
wide
.
Why
it
hurts:
●
Slow
&
expensive
:
More
features
mean
longer
training
times
and
bigger
computers.
●
Overfitting
risk
:
Models
latch
onto
random
noise
because
true
patterns
are
buried.
●
Hard
to
pi
cture:
You
can
plot
2-D
or
3-D
data,
but
not
10,000-D!
●
“Empty
space”
problem
:
Points
spread
so
thin
that
measuring
similarity
or
density
becomes
unreliable.
In
Generative
AI:
●
Word,
sentence,
image
embeddings
often
sit
in
thousands
of
dimensions;
multimodal
models
stack
even
more.
●
Trillions
of
parameters
amplify
the
problem—training
and
debugging
would
be
nearly
impossible
without
trimming
dimensions.
Fix:
Use
dimensionality-reduction
tools
(e.g.,
PCA,
t-SNE,
auto-encoders)
to
compress
data
into
fewer
but
still
informative
dimensions.
This
keeps
training
affordable,
reduces
overfitting,
and
makes
patterns
easier
to
see.
1.2.
Principal
Component
Analysis
(PCA):
Unveiling
Dominant
Patterns
[page 3]
Principal
Component
Analysis
(PCA)
is
a
cornerstone
of
unsupervised
dimensionality
reduction.
It
takes
a
dataset
whose
columns
are
often
tangled
together
and
mathematically
“rotates”
it
to
produce
new,
independent
axes
called
principal
components.
These
components
are
ranked
so
that
the
first
one
explains
the
largest
share
of
the
original
data’s
patterns,
the
second
explains
the
next
largest
share,
and
so
on,
letting
you
keep
just
the
leading
few
components
to
work
with
a
much
smaller
yet
still
informative
version
of
the
data.
Algorithmic
Steps,
Assumptions,
Pros
&
Cons
The
fundamental
goal
of
PCA
is
to
identify
new,
uncorrelated
features
(principal
components)
that
capture
the
maximum
possible
variance
from
the
original
data.
This
is
achieved
by
projecting
the
data
onto
a
lower-dimensional
subspace
defined
by
these
components.
The
core
Algorithmic
Steps
involved
are:
Algorithmic
Steps:
1.
Standardization/Centering:
It
is
crucial
to
standardize
the
data
(mean
of
0,
standard
deviation
of
1)
before
applying
PCA,
especially
if
features
are
on
different
scales.
PCA
is
sensitive
to
the
variance
of
the
initial
variables;
features
with
larger
ranges
would
otherwise
dominate
the
principal
components.
2.
Covariance
Matrix
Computation:
Calculate
the
covariance
matrix
from
the
standardized
data.
3.
Eigendecomposition:
Compute
the
eigenvectors
and
eigenvalues
of
the
covariance
matrix.
4.
Component
Selection:
Sort
the
eigenvectors
by
their
corresponding
eigenvalues
in
descending
order.
Select
the
top-k
eigenvectors
(where
k
is
the
desired
number
of
dimensions)
that
capture
a
significant
portion
of
the
total
variance
(e.g.,
95%
or
99%
).
5.
Data
Transformation:
Project
the
original
standardized
data
onto
the
selected
k
eigenvectors
to
obtain
the
new
k-dimensional
feature
space.
[page 4]
Benefits
●
Dimensionality
cut-down:
Fewer
features
speed
up
training
and
reduce
memory
use.
●
Noise
filtering:
Minor
PCs
often
capture
noise,
so
dropping
them
can
improve
model
generalization.
●
Visualization:
Plotting
the
first
two
or
three
PCs
offers
a
quick
glance
at
structure
and
outliers.
Caveats
●
Interpretability
loss:
PCs
are
linear
blends
of
original
features,
so
their
meaning
isn’t
always
intuitive.
●
Linear
assumption:
PCA
only
captures
linear
relationships;
complex
non-linear
patterns
may
remain
hidden.
●
Sensitivity
to
scale
and
outliers:
Always
standardize
data
and
inspect
for
extreme
values
before
applying
PCA.
PCA
Code
Example
Loads
and
standardizes
the
Iris
dataset,
reduces
it
to
two
dimensions
with
PCA,
then
plots
the
samples
so
you
can
visually
see
how
much
variance
the
first
two
components
capture
and
how
the
three
species
cluster.
[page 5]
GenAI
Applications:
Visualizing
LLM
Embeddings,
Feature
Engineering
for
GenAI
Turns
big
numbers
into
small
pictures
●
LLMs
give
every
sentence
thousands
of
numbers
(an
embedding).
●
PCA
squeezes
those
numbers
down
to
just
2
or
3
so
you
can
plot
them
like
dots
on
a
graph.
Helps
you
see
what
the
model
knows
●
Similar
sentences
land
close
together;
weird
or
wrong
ones
drift
away.\Easy
way
to
check
why
a
chatbot
answer
feels
“off.”
Gives
lighter
inputs
to
other
models
●
The
smaller
set
of
numbers
still
keeps
the
main
meaning
but
trains
faster
and
with
less
noise.
Shows
what
happens
inside
a
big
model
●
Apply
PCA
to
the
hidden
layers
to
watch
how
raw
words
slowly
turn
into
higher-level
ideas.
One
limitation
●
PCA
only
sees
straight-line
patterns.
If
you
need
curved
or
very
subtle
relationships,
add
tools
like
t-SNE
or
UMAP.
GenAI
Job
Scenario:
An
ML
Engineer
at
a
GenAI
startup
in
Bangalore
needs
to
debug
inconsistent
LLM
outputs
by
visualizing
text
embeddings.
●
Application:
The
engineer
extracts
text
embeddings
for
both
problematic
(inconsistent)
and
well-behaved
outputs
generated
by
their
LLM.
PCA
is
then
applied
to
reduce
these
high-dimensional
embeddings
to
a
2D
or
3D
space
for
plotting.
●
Insights
Sought:
The
visualization
aims
to
answer
questions
such
as:
Do
embeddings
corresponding
to
inconsistent
outputs
form
distinct
clusters
separate
from
the
[page 6]
well-behaved
ones?
Are
these
problematic
embeddings
located
far
from
expected
semantic
regions
in
the
embedding
space?
Do
they
exhibit
unusually
high
variance
along
certain
principal
components
that
might
not
align
with
semantic
meaning?
Discovering
such
patterns
can
provide
valuable
clues,
guiding
prompt
refinement
strategies,
identifying
issues
in
the
training
data,
or
indicating
the
need
for
further
model
fine-tuning
to
improve
consistency
and
reliability.
PCA
Use
case
code
-
https://github.com/aminfvthi/Principal-Component-Analysis
link
to
external
site
1.3.
t-Distributed
Stochastic
Neighbor
Embedding
(t-SNE):
Mapping
Local
Relationships
t-SNE
is
a
non-linear
dimensionality
reduction
technique
particularly
well-suited
for
visualizing
high-dimensional
datasets
in
low-dimensional
space
(typically
2D
or
3D).
It
is
renowned
for
its
ability
to
reveal
local
data
structure
and
create
well-separated
clusters
.
Core
Ideas
1.
Goal
—
keep
neighbors
together
t-SNE
rearranges
your
high-dimensional
data
into
2-D
or
3-D
so
that
items
that
were
close
before
stay
close
and
items
that
were
far
apart
stay
(mostly)
far
apart.
2.
Step
1:
measure
“closeness”
in
the
original
space
For
every
pair
of
points,
it
turns
their
distance
into
a
probability
of
being
neighbors
(using
a
bell-curve/Gaussian
formula).
3.
Step
2:
guess
closeness
in
the
low-dimensional
map
It
puts
the
same
points
on
a
small
canvas
and
calculates
new
neighbor
probabilities,
but
this
time
with
a
heavy-tailed
t-distribution.
Those
heavy
tails
give
extra
room
so
distant
points
don’t
get
squished
together.
4.
Step
3:
make
the
two
sets
of
probabilities
match
The
algorithm
nudges
points
around
to
shrink
the
difference
(called
KL
divergence)
between
“how
close
they
should
be”
(original
space)
and
“how
close
they
are
now”
(map).
When
that
difference
is
tiny,
the
map
is
ready.
Result:
you
get
an
easy-to-see
plot
where
nearby
dots
really
are
similar,
and
separate
clusters
stand
out
clearly.
Key
Parameters
Effective
t-SNE
visualization
heavily
depends
on
careful
tuning
of
its
parameters:
1.
Perplexity
(≈
neighborhood
size)
○
Pick
a
value
from
5
–
50.
○
Small
value
→
more
tiny
clusters,
risk
of
“speckled”
plot.
○
Large
value
→
smoother,
big-picture
view,
but
fine
detail
may
disappear.
○
Must
be
lower
than
your
total
number
of
points.
2.
Learning
rate
(how
far
points
move
each
step)
[page 7]
○
Try
10
–
1000.
○
Too
low
→
very
slow
convergence;
too
high
→
points
shoot
past
good
positions
and
the
plot
looks
random.
3.
Iterations
(how
many
optimization
steps)
○
A
few
hundred
to
a
few
thousand
usually
does
the
trick.
○
t-SNE
starts
with
an
“early
exaggeration”
phase
that
briefly
stretches
clusters
apart
so
they
settle
more
clearly
later.
Tweak
these
three
knobs
until
the
clusters
in
your
2-D/3-D
plot
look
stable
Algorithmic
Overview,
Strengths
&
Limitations
for
Visualization
1.
Compute
pairwise
affinities
(conditional
probabilities
pj
∣
i )
in
the
high-dimensional
space
using
a
Gaussian
kernel,
where
the
variance
for
each
point
is
determined
by
the
perplexity.
2.
Symmetrize
these
probabilities
to
get
joint
probabilities
pij .
3.
Initialize
low-dimensional
map
points
randomly
(or
sometimes
with
PCA).
4.
Compute
joint
probabilities
qij
in
the
low-dimensional
space
using
a
Student's
t-distribution.
5.
Minimize
the
KL
divergence
between
P
and
Q
using
gradient
descent.
Code
Example
[page 8]
Loads
the
Iris
data,
scales
it,
compresses
it
to
two
dimensions
with
t-SNE
(preserving
local
neighbourhood
relationships),
and
plots
the
result
so
you
can
visually
inspect
how
the
three
species
separate.
Benefits
●
Shows
clusters
clearly
–
often
uncovers
groups
that
other
methods
miss.
●
Catches
non-linear
patterns
–
handles
curved
or
tangled
relationships
that
PCA
flattens.
●
Keeps
close
points
close
–
local
neighborhoods
are
faithfully
preserved.
Caveats
●
Slow
on
big
datasets
–
even
the
faster
versions
struggle
past
~100
k
points.
●
Global
layout
can
deceive
–
distances
between
clusters
and
their
sizes
are
not
reliable.
●
Random
each
run
–
results
change
slightly
unless
you
fix
the
seed.
●
No
easy
“add
new
point”
–
can’t
project
fresh
data
without
rerunning
the
algorithm.
●
Fussy
settings
–
needs
trial-and-error
tuning
of
perplexity
and
learning
rate.
GenAI
Job
Scenario:
A
Data
Scientist
at
an
Indian
e-commerce
company
(e.g.,
Flipkart,
Myntra)
uses
t-SNE
to
understand
user
clusters
from
interaction
embeddings
with
a
GenAI-powered
recommendation
system.
●
Application:
The
data
scientist
extracts
user
embeddings
based
on
a
variety
of
interactions,
such
as
products
clicked,
items
purchased,
and
queries
made
to
a
GenAI-powered
shopping
assistant.
t-SNE
is
then
employed
to
reduce
these
high-dimensional
user
embeddings
into
a
2D
or
3D
space
for
visualization.
●
Insights
Sought:
The
goal
is
to
identify
if
distinct
user
personas
or
segments
emerge
as
visually
separable
clusters.
For
example,
do
users
who
respond
positively
to
[page 9]
GenAI-driven
recommendations
cluster
differently
from
those
who
ignore
them?
Are
there
clusters
representing
bargain
hunters,
brand
loyalists,
or
users
with
specific
stylistic
preferences?
Such
visualizations
can
help
refine
the
recommendation
strategy,
personalize
the
GenAI
assistant's
interactions,
and
tailor
marketing
campaigns
more
effectively.
1.4.
Linear
Discriminant
Analysis
(LDA):
Maximizing
Class
Separability
Linear
Discriminant
Analysis
(LDA)
is
a
supervised
dimensionality
reduction
technique.
Unlike
PCA,
which
is
unsupervised
and
aims
to
capture
the
maximum
variance
in
the
data,
LDA
focuses
on
finding
a
feature
subspace
that
maximizes
the
separability
between
classes.
Foundations,
Assumptions,
and
Comparison
to
PCA
LDA
operates
by
projecting
the
feature
space
onto
a
lower-dimensional
space
(with
at
most
C−1
dimensions,
where
C
is
the
number
of
classes)
in
such
a
way
that
the
ratio
of
between-class
variance
to
within-class
variance
is
maximized.This
ensures
that
in
the
transformed
space,
data
points
belonging
to
different
classes
are
as
far
apart
as
possible,
while
data
points
within
the
same
class
are
as
close
as
possible.
Assumptions:
1.
Gaussian
Distribution:
The
data
within
each
class
is
assumed
to
be
normally
(Gaussian)
distributed.
2.
Homoscedasticity
(Equal
Covariance):
All
classes
are
assumed
to
share
the
same
covariance
matrix.
If
this
assumption
is
violated,
Quadratic
Discriminant
Analysis
(QDA)
might
be
more
appropriate.
3.
Feature
Independence
(often
assumed
for
simplicity,
though
not
strictly
required
by
the
math):
Features
are
statistically
independent
or
at
least
have
low
multicollinearity.
LDA
can
handle
some
multicollinearity
by
transforming
data.
4.
Linear
Separability:
The
technique
implicitly
assumes
that
classes
are
linearly
separable
to
some
extent.
Comparison
to
PCA:
●
Supervision:
LDA
is
a
supervised
algorithm
that
uses
class
labels,
whereas
PCA
is
unsupervised
and
does
not
use
class
labels.
●
Objective:
LDA
aims
to
find
a
subspace
that
maximizes
class
separability.
PCA
aims
to
find
a
subspace
that
maximizes
variance.
●
Application:
LDA
is
primarily
used
for
classification
or
as
a
dimensionality
reduction
step
before
classification.
PCA
is
used
for
general
dimensionality
reduction,
data
compression,
and
visualization.
●
Number
of
Components:
LDA
will
find
at
most
C−1
discriminant
dimensions.
PCA
can
find
up
to
d
principal
components
(where
d
is
the
original
number
of
features).
[page 10]
Implementation
Steps,
Advantages
&
Use
Cases
Implementation
Steps:
1.
Compute
Mean
Vectors:
Calculate
the
d-dimensional
mean
vector
μ
i
for
each
class
(m1 ,m2 ,...,mC )
and
the
overall
mean
vector
μ.
2.
Compute
Scatter
Matrices:
3.
Solve
Generalized
Eigenvalue
Problem:
Find
the
eigenvectors
and
eigenvalues
for
the
equation
below.
The
eigenvectors
corresponding
to
the
largest
eigenvalues
are
the
linear
discriminants
(directions
for
the
new
subspace).
4.
Select
Linear
Discriminants:
Choose
the
top
k
eigenvectors
(where
k≤C−1)
to
form
the
transformation
matrix
W.
5.
Transform
Data:
Project
the
original
data
onto
the
new
k-dimensional
subspace:
Y=XW.
LDA
Code
Example
Loads
and
scales
the
Iris
measurements,
uses
LDA
to
find
two
axes
that
maximise
between-species
separation,
then
plots
the
samples
so
you
can
see
how
cleanly
the
classes
split
(
∼
99
%
of
the
discriminative
power
lies
on
LD1).
[page 12]
Advantages:
●
Effective
Class
Separation:
By
design,
LDA
excels
at
finding
dimensions
that
maximize
discrimination
between
classes.
●
Computational
Efficiency:
It
is
relatively
computationally
efficient
compared
to
some
non-linear
DR
techniques.
●
Performance
with
Small
Samples:
LDA
can
perform
well
even
when
the
number
of
samples
per
class
is
limited,
making
it
useful
in
fields
like
bioinformatics.
Use
Cases:
●
Face
Recognition:
Identifying
individuals
based
on
facial
features
(Fisherfaces).
●
Medical
Diagnosis:
Classifying
diseases
or
patient
states
based
on
medical
data.
●
Customer
Segmentation/Marketing:
Identifying
distinct
customer
groups
for
targeted
marketing.
●
Text
Classification:
Topic
modeling,
sentiment
analysis,
spam
detection
by
projecting
text
features
into
a
space
that
separates
document
categories.
●
Credit
Risk
Assessment:
Identifying
applicants
likely
to
default
on
loans.
GenAI
Job
Scenario:
A
GenAI
Ethics
Officer
at
a
media
company
in
India
uses
LDA
to
improve
the
classification
of
potentially
harmful
or
biased
content
generated
by
an
LLM
used
for
news
summarization.
●
Application:
The
company
uses
an
LLM
to
generate
summaries
of
news
articles.
To
ensure
ethical
compliance,
these
summaries
need
to
be
screened
for
harmful
or
biased
content.
Features
(e.g.,
TF-IDF
scores,
sentiment
scores,
or
even
embeddings
from
a
separate
model)
are
extracted
from
the
generated
summaries.
A
dataset
of
summaries
has
been
manually
labeled
as
'harmful'
or
'non-harmful'.
The
GenAI
Ethics
Officer
applies
LDA
to
this
feature
set
to
reduce
its
dimensionality,
specifically
aiming
to
find
a
projection
that
best
separates
the
harmful
from
non-harmful
summaries.
This
reduced
feature
set
is
then
used
to
train
a
classifier
(e.g.,
SVM
or
Logistic
Regression).
●
Insights
Sought:
The
officer
would
analyze
which
linear
combinations
of
the
original
features
(the
discriminants
found
by
LDA)
are
most
effective
in
distinguishing
harmful
content.
This
could
reveal
key
linguistic
patterns
or
semantic
elements
associated
with
problematic
summaries.
Furthermore,
they
would
assess
if
using
LDA
as
a
preprocessing
step
improves
the
accuracy,
precision,
and
recall
of
the
downstream
harmful
content
classifier,
and
potentially
makes
the
classification
model
simpler
and
faster
to
train
[page 13]
Chapter
2:
Feature
Selection:
Isolating
Critical
Signals
for
Robust
GenAI
2.1.
Why
Feature
Selection
is
Paramount
in
GenAI
Systems
Think
of
it
like
packing
for
a
trip.
Your
suitcase
(the
model)
can’t
fit
everything
in
your
room
(the
dataset),
so
you
only
take
items
(features)
that
will
actually
be
useful
at
your
destination.
Leave
behind
duplicates
and
“just-in-case”
clutter,
and
the
trip
is
lighter,
cheaper,
and
less
stressful.
What
exactly
is
a
“feature”?
●
In
a
spreadsheet
of
house
prices,
each
column—square
footage,
number
of
bedrooms,
distance
to
school—is
a
feature.
●
For
an
image
model,
every
pixel
value
is
a
feature.
●
In
text,
hidden
“embedding”
numbers
that
represent
meaning
become
features.
A
model
learns
patterns
by
mixing
and
matching
these
columns.
Too
many
weak
or
irrelevant
columns
confuse
it.
Why
remove
features?
1.
Avoid
over-fitting
–
The
model
won’t
chase
tiny
quirks
that
only
appear
in
the
training
data.
2.
Train
faster
–
Fewer
numbers
=
fewer
calculations.
3.
Save
money
–
Less
GPU/CPU
time
and
memory.
4.
Explain
results
–
It
’s
eas
ier
to
show
stakeholders
which
columns
really
drove
a
decision.
2.2.
Filter
Methods:
Statistical
Sieving
for
Relevance
Filter
methods
for
feature
selection
operate
by
assessing
the
relevance
of
features
based
on
their
intrinsic
statistical
properties,
independent
of
any
specific
machine
learning
model.
These
techniques
are
typically
applied
as
a
preprocessing
step
before
model
training.
The
core
idea
is
to
"filter
out"
less
important
features
based
on
scores
derived
from
various
statistical
tests
that
measure
their
correlation
with
the
target
variable
or
their
individual
characteristics.
Techniques:
●
Information
Gain:
Measures
the
reduction
in
entropy
(uncertainty)
about
the
target
variable
when
the
value
of
a
feature
is
known.
Features
that
provide
more
information
(higher
entropy
reduction)
are
ranked
higher.
It
is
commonly
used
with
categorical
target
variables.
●
Chi-squared
Test
(X2):
Used
primarily
for
categorical
features
and
a
categorical
target.
It
tests
the
independence
between
a
feature
and
the
target
variable.
A
high
X2
statistic
[page 14]
(and
a
low
p-value)
suggests
that
the
feature
is
dependent
on
the
target
and
thus
relevant.
●
ANOVA
F-test:
Suitable
for
numerical
features
and
a
categorical
target.
It
compares
the
variance
between
the
means
of
the
feature
values
across
different
classes
to
the
variance
within
each
class.
A
high
F-statistic
indicates
that
the
feature
has
different
distributions
for
different
classes,
making
it
discriminative.
●
Pearson's
Correlation
Coefficient:
Measures
the
linear
relationship
between
two
numerical
variables
(a
feature
and
a
numerical
target).
Values
range
from
-1
(perfect
negative
correlation)
to
+1
(perfect
positive
correlation).
Features
with
high
absolute
correlation
values
are
considered
more
relevant.
●
Variance
Threshold:
A
simple
approach
that
removes
features
whose
variance
does
not
meet
a
certain
threshold.
The
assumption
is
that
features
with
very
low
variance
provide
little
information
for
discrimination.
By
default,
it
often
removes
features
with
zero
variance.
Pros:
●
Fast
and
Computationally
Inexpensive:
Filter
methods
do
not
involve
training
machine
learning
models,
making
them
very
quick
to
execute,
especially
on
high-dimensional
datasets.
●
Model-Agnostic:
The
feature
selection
is
independent
of
the
choice
of
the
learning
algorithm,
making
the
selected
subset
potentially
generalizable
across
different
models.
●
Good
for
Initial
Screening:
Effective
for
an
initial
culling
of
features,
especially
when
dealing
with
a
very
large
number
of
potential
features,
to
reduce
the
search
space
for
more
complex
methods.
Cons:
●
Ignores
Feature
Interactions:
Filter
methods
evaluate
each
feature
independently
and
do
not
consider
the
combined
effect
or
interaction
between
features.
A
feature
might
be
individually
weak
but
highly
informative
when
combined
with
others.
●
May
Select
Redundant
Features:
Since
feature
dependencies
are
often
ignored,
filter
methods
might
select
multiple
features
that
are
highly
correlated
with
each
other
and
thus
provide
redundant
information.
●
Suboptimal
for
Specific
Models:
The
selected
feature
subset
is
not
tailored
to
any
specific
machine
learning
model
and
thus
may
not
be
the
optimal
set
for
a
particular
algorithm.
GenAI
Job
Scenario:
A
Research
Scientist
at
an
Indian
AI
lab
(e.g.,
working
on
agricultural
GenAI
solutions
for
crop
advice
)
performs
initial
feature
screening
from
a
diverse
dataset
including
weather
patterns,
soil
parameters,
historical
yields,
and
market
prices.
This
data
will
inform
a
GenAI
model
designed
to
predict
crop
yield
and
generate
tailored
farming
advice.
[page 15]
●
Application:
The
scientist
employs
filter
methods
like
Pearson's
correlation
(for
numerical
yield
targets)
and
information
gain
(if
advice
categories
are
the
target)
to
quickly
assess
the
individual
relevance
of
numerous
environmental,
agricultural,
and
economic
factors.
●
Insights
Sought:
The
primary
goal
is
to
identify
which
factors
demonstrate
the
strongest
standalone
statistical
relationships
with
crop
yield
or
advice
categories.
This
initial
screening
helps
to
narrow
down
the
vast
feature
set,
creating
a
more
manageable
input
for
subsequent,
more
computationally
intensive
wrapper
or
embedded
feature
selection
methods,
or
directly
for
the
GenAI
model
if
it
can
handle
structured
inputs.
This
ensures
that
the
most
promising
features
are
prioritized
early
in
the
development
pipeline.
2.3.
Wrapper
Methods:
Model-Centric
Feature
Optimization
Idea
Test
different
sets
of
columns
by
actually
training
the
model
each
time
and
keep
the
set
that
scores
highest.
Common
strategies
●
Forward
Selection
–
start
empty,
add
one
column
at
a
time
if
it
boosts
accuracy.
●
Backward
Elimination
–
start
with
all
columns,
drop
the
one
that
hurts
accuracy
least.
●
Recursive
Feature
Elimination
(RFE)
–
train,
rank
columns
by
importance,
drop
the
weakest,
repeat.
Workflow
1.
Pick
(or
change)
a
feature
set.
2.
Train
the
model,
check
its
score.
3.
Add
or
drop
a
column
using
one
of
the
strategies.
4.
Stop
when
more
changes
no
longer
help—or
when
you
reach
the
desired
number
of
features.
Pros
●
Finds
column
combinations
that
really
matter
to
this
model.
●
Often
beats
filter
methods
on
accuracy.
Cons
●
Slow
and
compute-hungry
(model
retrains
many
times).
●
Can
overfit
if
the
dataset
is
small
or
validation
is
weak.
●
Best
feature
set
may
change
if
you
switch
to
a
different
model
type.
[page 16]
2.4.
Embedded
Methods:
Integrated
Feature
Selection
What
it
means
The
model
picks
its
own
important
columns
while
it
trains,
thanks
to
built-in
penalties
or
split
rules.
Popular
built-in
pickers
●
LASSO
/
L1
regularization
–
linear
model
shrinks
some
weights
to
exactly
0,
so
those
columns
disappear.
●
Ridge
/
L2
regularization
–
shrinks
weights
toward
0
(helps
stability)
but
rarely
kills
them
outright.
●
Elastic
Net
–
mixes
L1
and
L2,
balancing
true
selection
and
stability.
●
Tree-based
models
(Random
Forest,
Gradient
Boosting)
–
at
each
split
the
tree
chooses
the
best
column;
columns
used
near
the
top
are
usually
most
important.
Pros
●
Fast
–
one
training
run;
no
looping
through
feature
subsets.
●
Built-in
over-fit
guard
–
regularization
keeps
the
model
from
memorizing
noise.
●
Captures
interactions
–
model
decides
which
combos
matter.
Cons
●
Model-dependent
–
features
chosen
for
a
LASSO
model
might
not
suit,
say,
an
SVM.
●
Transparency
varies
–
zeroed
LASSO
weights
are
obvious;
tree
importance
scores
or
deep-net
pruning
can
be
harder
to
explain.
2.5.
Advanced
Applications
in
GenAI
Feature
Selection
for
Multimodal
Data
(Image,
Text,
etc.)
GenAI
models
increasingly
operate
on
multimodal
data,
combining
inputs
like
images,
text,
audio,
and
structured
data.
Feature
selection
in
this
context
presents
unique
challenges
and
opportunities.
●
Challenges:
○
Heterogeneity:
Different
modalities
have
distinct
data
types
and
structures
(e.g.,
pixel
arrays
for
images,
token
sequences
for
text).
○
Dimensionality
Mismatch:
Features
extracted
from
different
modalities
can
have
vastly
different
dimensionalities.
○
Finding
Joint
Relevance:
Identifying
features
or
combinations
of
features
across
modalities
that
are
jointly
relevant
to
the
GenAI
task
is
complex.
●
Approaches:
○
Early
Fusion:
Features
from
different
modalities
are
concatenated
into
a
single
feature
vector
early
in
the
pipeline.
Standard
feature
selection
techniques
can
then
be
applied
to
this
combined
vector.
However,
this
approach
might
not
[page 17]
effectively
capture
inter-modal
relationships
and
can
be
dominated
by
higher-dimensional
modalities.
○
Late
Fusion:
Separate
models
are
trained
for
each
modality,
potentially
with
modality-specific
feature
selection.
The
outputs
or
decisions
from
these
models
are
then
combined
at
a
later
stage.
This
allows
for
specialized
processing
but
might
miss
subtle
cross-modal
interactions.
○
Intermediate/Hybrid
Fusion:
More
sophisticated
methods
aim
to
learn
joint
representations
or
allow
interactions
between
modalities
at
intermediate
layers
of
a
deep
learning
model.
○
Deep
Learning
for
Feature
Extraction:
For
modalities
like
images
and
text,
deep
learning
models
such
as
Convolutional
Neural
Networks
(CNNs)
and
Transformers/RNNs
are
powerful
feature
extractors.
Feature
selection
can
then
be
applied
to
these
learned
high-level
features.
○
Reinforcement
Learning
(RL)
for
Dynamic
Selection:
RL
agents
can
be
trained
to
dynamically
select
the
optimal
subset
of
features
from
multimodal
data
by
interacting
with
the
environment
(the
model
and
data)
and
receiving
rewards
based
on
performance.
○
Optimization
Algorithms:
Novel
optimization
algorithms
are
being
developed
specifically
for
multimodal
feature
selection,
aiming
to
find
the
most
relevant
features
across
modalities
to
improve
classification
or
generation
performance.
Feature
Engineering
&
Selection
in
LLM
Fine-tuning
When
fine-tuning
Large
Language
Models
(LLMs)
for
tasks
that
involve
structured
data
in
addition
to
text,
the
selection
and
representation
of
these
structured
features
become
critical.
●
Context:
For
example,
fine-tuning
an
LLM
to
predict
customer
churn
might
involve
using
historical
transaction
data
(structured)
and
customer
review
texts
(unstructured).
The
LLM
needs
to
effectively
integrate
information
from
both
sources.
●
Challenges:
Deciding
which
structured
features
are
relevant,
how
to
encode
them
(e.g.,
numerical
scaling,
categorical
embedding),
and
how
to
present
them
to
the
LLM
alongside
the
text
(e.g.,
as
part
of
the
prompt,
or
through
dedicated
input
channels
in
multimodal
architectures)
are
key
challenges.
●
LLMs
for
Feature
Understanding:
An
emerging
approach
involves
using
LLMs
themselves
to
assess
the
importance
of
features,
especially
when
features
have
descriptive
names
or
textual
context.
The
LLM
can
be
prompted
with
feature
descriptions
and
task
context
to
provide
relevance
scores
or
select
features.
[page 18]
Chapter
3:
Data
Compression
&
Reduction:
Making
GenAI
Models
Leaner
and
Faster
Why
Compress
GenAI
Models?
Generative
AI
models,
especially
Large
Language
Models
(LLMs)
like
GPT-3,
are
often
huge.
GPT-3
has
175
billion
parts
called
parameters!
This
large
size
causes
problems:
●
They
need
a
lot
of
computer
memory
to
store
and
run.
●
They
require
powerful
computers,
which
use
a
lot
of
energy
and
can
be
expensive.
These
issues
can
make
it
hard
to
use
advanced
GenAI
everywhere,
especially
on
smaller
devices
like
smartphones
or
in
places
with
slow
internet.
Model
and
data
compression
techniques
help
by:
●
Making
models
and
data
smaller.
●
Speeding
up
how
fast
they
work.
●
Using
less
power.
This
allows
these
smart
AI
systems
to
be
used
in
more
places
and
on
more
devices.
For
LLMs,
compression
isn't
just
a
nice-to-have;
it's
often
essential
to
use
them
outside
of
big
data
centers.
Making
these
models
smaller
and
cheaper
to
run
also
means
more
people
and
organizations
can
use
them,
leading
to
new
ideas
and
applications,
like
on-device
personal
assistants.
Basic
Choices:
Lossless
vs.
Lossy
Compression
When
we
compress
data,
we
have
two
main
choices:
Lossless
Compression:
●
What
it
is:
This
method
makes
files
smaller
by
finding
and
removing
repetitive
information.
The
best
part?
You
can
get
the
original
data
back
perfectly,
with
no
information
lost.
●
When
to
use
it:
Ideal
when
you
can't
afford
to
lose
any
detail.
Think
of
text
documents,
computer
code,
or
spreadsheets.
●
Examples:
ZIP
files,
PNG
images,
FLAC
audio
files.
●
For
GenAI:
Good
for
storing
training
data
(especially
text),
model
settings,
or
the
exact
text
an
LLM
produces
.
Lossy
Compression:
●
What
it
is:
This
method
makes
files
much
smaller
by
throwing
away
some
data
that's
considered
less
important
or
hard
to
notice.
You
can't
get
the
original
data
back
perfectly;
there's
some
quality
loss,
but
often
it's
barely
noticeable.
[page 19]
●
When
to
use
it:
Mostly
for
media
like
images,
audio,
and
video,
where
a
tiny
drop
in
quality
is
okay
for
a
big
drop
in
file
size.
●
Examples:
JPEG
images,
MP3
audio,
MP4
videos.
●
For
GenAI:
The
idea
behind
lossy
compression
is
very
similar
to
how
we
compress
AI
models.
Techniques
like
quantization
(using
simpler
numbers
for
model
parts)
or
pruning
(removing
less
important
model
parts)
mean
some
model
information
is
lost
to
make
the
model
smaller,
hoping
it
still
works
well.
It's
also
used
for
compressing
large
image
or
video
datasets
for
training
GenAI.
The
main
idea
in
lossy
compression
–
giving
up
a
little
bit
of
quality
for
a
much
smaller
size
–
is
the
same
for
shrinking
AI
models.
We
try
to
remove
the
least
important
bits,
whether
it's
tiny
details
in
a
photo
or
less
critical
connections
in
a
neural
network,
to
make
things
smaller
and
faster
without
a
big
drop
in
how
well
they
work.
Shrinking
Big
AI
Models
(LLMs
and
Deep
Learning)
Modern
AI
models,
especially
LLMs,
are
so
big
and
power-hungry
that
we
need
special
ways
to
shrink
them.
The
goal
is
smaller,
faster,
more
energy-efficient
models
that
still
perform
well.
The
main
methods
are
pruning,
quantization,
and
knowledge
distillation.
Pruning:
Trimming
the
Fat
●
Concept:
Pruning
is
like
trimming
off
unnecessary
branches
from
a
tree.
It
removes
unneeded
parts
(weights,
neurons,
or
even
whole
layers)
from
a
trained
AI
model
to
make
it
smaller
and
faster.
The
idea
is
that
big
models
often
have
many
redundant
parts.
●
Types:
1.
Unstructured
Pruning:
Removes
individual
weights.
Can
make
the
model
very
small,
but
the
resulting
"sparse"
model
can
be
hard
for
regular
computers
to
speed
up.
2.
Structured
Pruning:
Removes
whole
chunks
like
neurons
or
layers.
This
keeps
the
model's
structure
neat,
making
it
easier
for
hardware
to
run
faster.
●
How
it's
done:
1.
Train
the
full
model.
2.
Figure
out
which
parts
are
least
important
(e.g.,
weights
with
small
values).
3.
Remove
those
parts.
4.
Fine-tune
(retrain
a
bit)
the
smaller
model
to
get
back
any
lost
performance.
●
For
GenAI:
Pruning
is
useful
for
LLMs.
SparseGPT
is
one
such
method.
However,
it
can
be
tricky,
and
sometimes
accuracy
drops
more
than
desired,
especially
for
certain
LLM
types.
Quantization:
Using
Simpler
Numbers
●
Concept:
Instead
of
using
very
precise
numbers
(like
32-bit
numbers)
for
a
model's
weights
and
calculations,
quantization
uses
simpler,
less
precise
numbers
(like
8-bit
[page 20]
integers).
This
is
like
rounding
numbers
to
make
them
take
up
less
space.
It
makes
the
model
smaller,
uses
less
memory,
and
can
speed
up
calculations
on
hardware
that
supports
these
simpler
numbers.
●
Types:
○
Quantization-Aware
Training
(QAT):
The
model
is
trained
knowing
it
will
use
simpler
numbers.
This
often
gives
better
accuracy
but
needs
retraining.
○
Post-Training
Quantization
(PTQ):
Simpler
numbers
are
applied
after
the
model
is
already
trained.
It's
faster
and
easier,
very
popular
for
LLMs
where
retraining
is
too
expensive.
●
For
GenAI:
Quantization
is
vital
for
LLM
compression.
Methods
like
GPTQ
can
shrink
LLMs
to
use
very
few
bits
(e.g.,
3
or
4)
with
acceptable
performance.
A
challenge
is
that
LLMs
sometimes
have
a
few
very
large,
important
weights
that
need
careful
handling
during
quantization.
Knowledge
Distillation:
Learning
from
a
"Teacher"
●
Concept:
A
smaller
"student"
model
learns
from
a
larger,
more
capable
"teacher"
model.
The
goal
is
for
the
student
to
be
almost
as
smart
as
the
teacher
but
much
smaller
and
faster.
●
How
it's
done:
1.
Train
a
big,
powerful
teacher
model.
2.
The
student
model
is
then
trained
to
copy
the
teacher's
outputs.
Often,
the
student
learns
from
the
teacher's
"soft
targets"
(the
probabilities
the
teacher
assigns
before
making
a
final
decision),
which
gives
more
information
than
just
the
final
answer.
Often,
the
best
results
come
from
using
these
techniques
together,
like
pruning
a
model
and
then
quantizing
it.
The
order
in
which
you
do
them
can
also
make
a
difference.
GenAI
Job
Scenario:
Deploying
LLMs
on
Edge
Devices
A
GenAI
Deployment
Engineer
in
India
needs
to
make
a
sophisticated
multilingual
LLM
run
on
a
small
edge
device
(like
a
Jetson
Nano
)
for
an
on-device
personal
assistant
that
can
summarize
and
answer
questions
in
various
Indian
languages.
●
How
they
might
do
it:
1.
Knowledge
Distillation
(maybe
first):
Start
with
a
smaller
"student"
LLM
that
learned
from
a
bigger
one.
2.
Pruning:
Use
structured
pruning
to
make
the
student
LLM
even
smaller
and
faster
for
the
hardware.
3.
Quantization:
Apply
Post-Training
Quantization
(PTQ),
maybe
to
8-bit
integers,
using
tools
like
NVIDIA's
TensorRT-LLM
or
ONNX
Runtime
to
optimize
it
for
the
Jetson
Nano.
[page 21]
●
What
they
consider:
They
need
to
balance
model
size
(to
fit
on
the
device),
speed
(for
quick
responses),
power
use
(for
battery
life),
and
accuracy
(good
quality
answers
in
all
languages).
●
Goal:
Find
the
best
mix
of
compression
that
fits
the
device's
limits
without
making
the
LLM
perform
poorly
on
its
main
tasks.
This
means
lots
of
testing
on
the
actual
device.
Ensuring
Compressed
Models
Are
Still
Smart:
Metrics
like
SrCr
Just
making
a
model
smaller
isn't
enough.
We
need
to
ensure
it
still
works
well
and
understands
things
correctly.
Old
metrics
like
"perplexity"
(how
fluent
the
text
is)
don't
always
catch
if
a
compressed
LLM
has
lost
its
reasoning
ability
or
factual
accuracy.
The
Semantic
Retention
Compression
Rate
(SrCr)
is
a
newer
idea
to
measure
this
better.
●
SrCr
looks
at
two
things:
○
How
much
the
model
was
compressed
(Theoretical
Compression
Rate
-
TCr).
○
How
well
the
compressed
model
still
performs
on
actual
tasks
compared
to
the
original
(Semantic
Retention
-
Sr).
●
Why
it's
important:
Metrics
like
SrCr
help
us
focus
on
making
models
that
are
both
small
and
still
useful
for
what
they
were
designed
to
do.
It
reminds
us
to
test
compressed
models
on
real
tasks.
For
GenAI,
especially
LLMs,
we
must
check
if
the
compressed
version
can
still
reason,
summarize,
and
answer
questions
accurately.
The
goal
is
to
keep
the
"semantic
understanding"
–
the
model's
ability
to
grasp
meaning
and
context
–
even
after
making
it
smaller.
[page 22]
Chapter
4:
Similarity
Measures
What
is
Similarity?
Often,
we
represent
data
(like
text,
images,
or
user
preferences)
as
lists
of
numbers
called
vectors
or
embeddings.
These
vectors
live
in
a
multi-dimensional
space.
If
two
items
are
similar
in
meaning
or
characteristics,
their
vectors
should
be
close
together
in
this
space.
Similarity
measures
are
mathematical
tools
that
tell
us
exactly
how
"close"
or
"similar"
these
vectors
are
to
each
other.
Being
able
to
measure
similarity
is
key
to
many
data
tasks:
●
Information
Retrieval/Search:
Finding
documents
or
items
that
are
most
similar
to
a
user's
query.
●
Recommendation
Systems:
Suggesting
items
a
user
might
like
based
on
similarity
to
what
they've
liked
before,
or
finding
users
with
similar
tastes.
●
Clustering:
Grouping
similar
data
points
together.
●
Anomaly
Detection:
Identifying
data
points
that
are
very
different
from
others.
●
In
essence,
if
vectors
are
the
language
of
our
data,
similarity
measures
are
the
grammar
that
helps
us
understand
the
relationships
between
them.
Euclidean
Distance:
The
Straight
Line
This
is
the
most
straightforward
way
to
measure
distance:
the
shortest,
straight-line
distance
between
two
points.
The
Formula
&
Idea
Imagine
two
points,
P
and
Q.
If
P
has
coordinates
(p1 ,p2 ,...,pn )
and
Q
has
(q1 ,q2 ,...,qn )
in
an
n-dimensional
space,
the
Euclidean
distance
is:
[page 23]
This
formula
comes
from
the
Pythagorean
theorem
(like
finding
the
hypotenuse
of
a
triangle).
Pros
:
●
Very
intuitive
–
it's
how
we
naturally
think
about
distance.
●
Easy
to
calculate
.
Cons:
●
"Curse
of
Dimensionality":
In
very
high-dimensional
spaces
(data
with
many
features),
Euclidean
distance
can
become
less
useful.
All
points
might
start
to
look
equally
far
apart.
●
Sensitive
to
Scale:
If
your
features
are
on
different
scales
(e.g.,
one
feature
is
0-1,
another
is
0-1000),
the
larger-scale
feature
will
dominate
the
distance.
Always
standardize
your
data
(make
features
have
a
similar
scale,
like
mean
0
and
standard
deviation
1)
before
using
Euclidean
distance.
●
Magnitude
Matters:
It
considers
the
"length"
or
magnitude
of
the
vectors.
This
might
not
be
what
you
want
if,
for
example,
you're
comparing
two
documents
on
the
same
topic
but
one
is
much
longer
than
the
other.
Cosine
Similarity:
Measuring
the
Angle
Cosine
similarity
measures
how
similar
two
vectors
are
by
looking
at
the
angle
between
them.
It
cares
about
the
direction
of
the
vectors,
not
their
length
or
magnitude.
The
Formula
&
Idea
For
two
vectors,
A
and
B,
the
cosine
similarity
is:
[page 24]
Interpretation:
●
+1:
The
vectors
point
in
the
exact
same
direction
(angle
is
0°).
They
are
perfectly
similar
in
orientation.
●
0:
The
vectors
are
at
a
90°
angle
(orthogonal).
They
have
no
similarity
in
orientation.
●
-1:
The
vectors
point
in
opposite
directions
(angle
is
180°).
They
are
perfectly
dissimilar
in
orientation.
Because
it
only
looks
at
the
angle,
it's
not
affected
by
how
long
the
vectors
are.
Why
It's
Good
for
Certain
Data
●
Magnitude
Invariance:
This
is
great
when
the
length
of
vectors
doesn't
reflect
similarity.
For
example,
in
text
analysis,
a
short
article
and
a
long
essay
on
the
same
topic
should
be
considered
similar.
Cosine
similarity
can
capture
this
because
their
content
(and
thus
vector
direction)
is
similar,
even
if
their
lengths
differ.
●
High
Dimensions:
It
often
works
better
than
Euclidean
distance
in
high-dimensional
spaces,
especially
with
sparse
data
(data
with
many
zeros,
common
in
text).
Jaccard
Index:
Overlap
Between
Sets
The
Jaccard
Index
(or
Jaccard
similarity
coefficient)
measures
how
similar
two
sets
are
by
looking
at
how
many
elements
they
share.
The
Formula
&
Idea
For
two
sets,
A
and
B:
[page 25]
●
Range:
The
Jaccard
Index
is
always
between
0
and
1.
○
0:
The
sets
have
no
elements
in
common.
○
1:
The
sets
are
exactly
the
same.
Use
with
Categorical/Binary
Data
&
Text
The
Jaccard
Index
is
very
useful
when
you're
comparing
items
based
on
shared
characteristics
that
are
either
present
or
absent
(binary
data),
or
fall
into
categories.
For
example,
if
A
and
B
are
binary
vectors
(lists
of
0s
and
1s):
J(A,B)=M01 +M10 +M11 M11
Where:
●
M11 :
Number
of
positions
where
both
A
and
B
have
a
1.
●
M01 :
Number
of
positions
where
A
has
0
and
B
has
1.
●
M10 :
Number
of
positions
where
A
has
1
and
B
has
0.
Choosing
the
Right
Similarity
Measure
The
best
similarity
measure
depends
on
your
data
and
what
you
mean
by
"similar":
●
Euclidean
Distance:
Use
when
the
actual
values
and
magnitudes
(lengths)
of
your
vectors
are
important,
and
your
features
are
on
a
similar
scale
(or
standardized).
Good
for
dense
numerical
data
where
absolute
differences
matter.
●
Cosine
Similarity:
Best
when
the
direction
(orientation)
of
vectors
is
more
important
than
their
magnitude.
This
is
often
the
case
for
high-dimensional
data
like
text,
where
you
care
about
meaning
(direction)
more
than
document
length
(magnitude).
●
Jaccard
Index:
Use
when
you
are
comparing
sets
of
items,
or
features
that
are
binary
(yes/no,
present/absent).
It's
about
shared
presence
of
elements,
not
their
order
or
frequency
beyond
presence.
Consider
if
your
vectors
are
normalized
(all
scaled
to
have
a
length
of
1).
If
they
are,
Euclidean
distance
and
cosine
similarity
become
mathematically
related,
though
cosine
similarity
is
often
still
preferred
for
its
interpretability
in
terms
of
angles.
For
very
large
datasets,
how
quickly
the
measure
can
be
calculated
is
also
important.
Code example
[page 26]
Cosine* gauges how closely the two numeric vectors point in the same direction, Euclidean gives their raw
geometric
distance,
and
Jaccard
measures
overlap
between
two
binary
sets
(here,
word-presence
vectors).
[page 27]
Multiple
Choice
Questions
(MCQs)
1.
The
"curse
of
dimensionality"
in
the
context
of
GenAI
primarily
means
that:
A.
Models
become
too
simple
and
underfit
the
data.
B.
It's
difficult
to
find
enough
diverse
data
for
training.
C.
In
high-dimensional
spaces,
data
points
become
sparse,
distances
less
meaningful,
and
models
may
overfit
noise.
D.
The
ethical
implications
of
using
many
features
become
unmanageable.
(Correct
Answer:
C.
Rationale:
Section
1.1
explains
that
the
"curse
of
dimensionality"
leads
to
slower
models,
overfitting
risk
due
to
sparsity,
and
difficulty
in
visualizing
patterns
because
points
spread
out
and
distances
become
less
reliable.)
2.
A
key
difference
between
Principal
Component
Analysis
(PCA)
and
Linear
Discriminant
Analysis
(LDA)
is:
A.
PCA
is
supervised
and
aims
to
maximize
variance,
while
LDA
is
unsupervised
and
aims
for
class
separability.
B.
LDA
is
supervised
and
aims
to
maximize
class
separability,
while
PCA
is
unsupervised
and
aims
to
maximize
variance.
C.
Both
are
supervised,
but
PCA
focuses
on
variance
and
LDA
on
feature
correlation.
D.
Both
are
unsupervised,
but
LDA
is
better
for
non-linear
data.
(Correct
Answer:
B.
Rationale:
Section
1.4
states,
"Unlike
PCA,
which
is
unsupervised
and
aims
to
capture
the
maximum
variance
in
the
data,
LDA
focuses
on
finding
a
feature
subspace
that
maximizes
the
separability
between
classes."
and
"LDA
is
a
supervised
algorithm
that
uses
class
labels,
whereas
PCA
is
unsupervised.")
3.
When
visualizing
high-dimensional
data
with
t-SNE,
what
is
a
critical
point
to
remember
about
interpreting
the
resulting
plot?
A.
The
exact
distances
between
clusters
and
their
relative
sizes
always
accurately
reflect
their
true
separation
in
the
original
high-dimensional
space.
B.
t-SNE
is
primarily
useful
for
linear
data
structures,
similar
to
PCA.
C.
The
global
layout,
such
as
distances
between
clusters
and
cluster
sizes,
can
be
misleading
and
should
be
interpreted
with
caution.
D.
t-SNE
is
computationally
less
expensive
than
PCA
for
large
datasets.
(Correct
Answer:
C.
Rationale:
Section
1.3
(Caveats)
mentions,
"Global
layout
can
deceive
–
distances
between
clusters
and
their
sizes
are
not
reliable.")
4.
What
is
the
primary
benefit
of
using
feature
selection
techniques
in
building
GenAI
models?
A.
To
artificially
increase
the
dataset
size
for
better
training.
B.
To
select
a
subset
of
relevant
features,
which
can
lead
to
simpler
models,
faster
training,
reduced
overfitting,
and
potentially
better
accuracy.
C.
To
ensure
all
features
are
converted
to
a
numerical
format.
D.
To
always
increase
the
number
of
dimensions
for
more
complex
pattern
recognition.
(Correct
Answer:
B.
Rationale:
Section
2.1
lists
benefits
like
avoiding
overfitting,
faster
[page 28]
training,
saving
money
(less
compute),
and
easier
result
explanation.)
5.
Which
category
of
feature
selection
methods
involves
training
a
model
multiple
times
to
evaluate
different
subsets
of
features
based
on
the
model's
performance?
A.
Filter
methods
B.
Embedded
methods
C.
Wrapper
methods
D.
Intrinsic
methods
(Correct
Answer:
C.
Rationale:
Section
2.3
(Wrapper
Methods
Idea)
states
they
"Test
different
sets
of
columns
by
actually
training
the
model
each
time
and
keep
the
set
that
scores
highest.")
6.
LASSO
(L1
Regularization)
is
an
embedded
feature
selection
method
that
works
by:
A.
Boosting
the
coefficients
of
important
features.
B.
Assigning
scores
to
features
based
on
statistical
tests
before
model
training.
C.
Shrinking
the
coefficients
of
some
features
exactly
to
zero,
effectively
removing
them
from
the
model.
D.
Creating
new
features
that
are
linear
combinations
of
the
original
ones.
(Correct
Answer:
C.
Rationale:
Section
2.4
(Embedded
Methods
-
LASSO)
explains
that
LASSO
"shrinks
some
weights
to
exactly
0,
so
those
columns
disappear.")
7.
In
data
compression,
what
is
the
fundamental
difference
between
lossless
and
lossy
compression?
A.
Lossless
compression
achieves
higher
compression
ratios
than
lossy
compression.
B.
Lossy
compression
allows
perfect
reconstruction
of
the
original
data,
while
lossless
does
not.
C.
Lossless
compression
allows
perfect
reconstruction
of
the
original
data,
while
lossy
compression
discards
some
information,
meaning
the
original
cannot
be
perfectly
rebuilt.
D.
Lossy
compression
is
only
used
for
text
data,
and
lossless
for
images.
(Correct
Answer:
C.
Rationale:
Under
Chapter
3,
"Basic
Choices:
Lossless
vs.
Lossy
Compression"
explains
lossless
allows
perfect
reconstruction
and
lossy
discards
some
data.)
8.
Which
model
compression
technique
involves
training
a
smaller
"student"
model
to
replicate
the
behavior
and
knowledge
of
a
larger,
more
capable
"teacher"
model?
A.
Pruning
B.
Quantization
C.
Knowledge
Distillation
D.
Parameter
Sharing
(Correct
Answer:
C.
Rationale:
Under
Chapter
3,
"Knowledge
Distillation:
Learning
from
a
'Teacher'"
describes
this
process.)
9.
Cosine
similarity
is
often
preferred
over
Euclidean
distance
for
comparing
text
embeddings
primarily
because:
[page 29]
A.
It
is
less
computationally
intensive
for
all
types
of
data.
B.
It
focuses
on
the
magnitude
(length)
of
the
vectors,
which
is
crucial
for
text.
C.
It
measures
the
angle
(orientation)
between
vectors,
making
it
robust
to
differences
in
document
length
while
capturing
semantic
similarity.
D.
Euclidean
distance
cannot
handle
sparse
data
typically
found
in
text
embeddings.
(Correct
Answer:
C.
Rationale:
Section
4.3
(Why
It's
Good
for
Text)
highlights
magnitude
invariance
and
focus
on
orientation
for
semantic
similarity.)
10.
The
Jaccard
Index
is
most
appropriately
used
for
measuring
similarity
between:
A.
Two
continuous
time-series
datasets.
B.
Two
sets
of
items,
such
as
the
unique
words
in
two
different
documents.
C.
The
numerical
outputs
of
two
different
regression
models.
D.
Two
high-dimensional,
dense
numerical
embeddings
where
magnitude
is
important.
(Correct
Answer:
B.
Rationale:
Section
4.4
(Use
with
Categorical/Binary
Data
&
Text)
states
it's
useful
for
comparing
items
described
by
categorical/binary
features
or
sets
of
words.)
Practice
Questions
1.
Scenario:
Interpreting
Dimensionality
Reduction
for
LLM
Embeddings
You
are
an
ML
Engineer
working
on
a
GenAI
application.
You've
extracted
768-dimensional
text
embeddings
from
user
queries
to
your
chatbot.
○
(A)
If
you
use
PCA
to
reduce
these
embeddings
to
2D
for
visualization
and
find
that
queries
with
very
different
intents
(e.g.,
"book
a
flight"
vs.
"tell
me
a
joke")
are
plotted
close
together,
what
are
two
potential
reasons
related
to
PCA's
characteristics
that
could
explain
this?
○
(B)
What
alternative
dimensionality
reduction
technique
is
often
better
for
visualizing
local
clusters
in
such
embedding
spaces,
and
what
key
parameter
of
this
alternative
technique
would
you
need
to
tune
carefully?
Hint
for
answering:
○
For
(A),
consider
PCA's
linearity
and
its
objective
of
maximizing
variance.
○
For
(B),
think
about
non-linear
techniques
known
for
preserving
local
structure
and
their
common
hyperparameters.
2.
Application:
Feature
Selection
and
Model
Compression
Strategy
Imagine
you
are
tasked
with
building
a
GenAI
model
for
a
fintech
company
in
India
to
detect
potentially
fraudulent
transaction
descriptions
(text
data)
augmented
with
structured
transaction
metadata
(e.g.,
transaction
amount,
time
of
day,
user
location
category).
○
(A)
From
the
structured
metadata,
how
would
you
select
the
most
relevant
features
to
combine
with
the
text
analysis?
Describe
one
type
of
feature
selection
[page 30]
method
(Filter,
Wrapper,
or
Embedded)
you
might
use
and
why
it
would
be
appropriate.
○
(B)
If
the
final
GenAI
model
(after
incorporating
text
and
selected
structured
features)
is
too
large
and
slow
for
real-time
fraud
alerts,
name
two
distinct
model
compression
techniques
you
would
consider
applying
and
briefly
state
the
core
idea
of
each.
Hint
for
answering:
○
For
(A),
consider
the
trade-offs
(speed,
interaction
capture,
model
dependency)
of
different
feature
selection
categories.
○
For
(B),
recall
the
main
approaches
to
making
models
smaller
and
faster
(e.g.,
removing
parts,
simplifying
numbers,
learning
from
a
bigger
model).
3.
Conceptual:
Choosing
Similarity
Measures
Explain
a
specific
scenario
or
type
of
data
where
Euclidean
distance
would
be
a
more
suitable
similarity
measure
than
cosine
similarity,
and
vice-versa.
Justify
your
choices
by
highlighting
the
key
properties
of
each
measure.
Hint
for
answering:
○
Think
about
when
the
magnitude/length
of
vectors
is
important
versus
when
only
the
direction/orientation
matters.
Further
Learning
1.
Reference
Youtube
Videos
a.
Dimensionality Reduction | ML-005 Lecture 14 | Stanford University | Andre…
link
to
external
site
b.
link
to
Feature Selection Techniques Easily Explained | Machine Learning
external
site
2.
Reference
Websites
a.
https://www.geeksforgeeks.org/dimensionality-reduction/
link
to
external
site
b.
https://www.ibm.com/think/topics/feature-selection
link
to
external
site
c.
https://www.geeksforgeeks.org/data-reduction-in-data-mining/
link
to
external
site