# Data Pre-processing and Curation -I - Sunday (09.11.2025)
course: Module 1 — Foundations of AI & ML
module: Module-1-Foundations-AI-ML
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-1-Foundations-AI-ML/General/Data_Pre-processing_and_Curation_-I_-_Sunday_(09.11.2025).pdf
pages: 41
---
[page 1]
IIT Roorkee – Futurense
PG Certification
GenAI and Agentic AI for Engineers
Data Pre-Processing
and Curation - I
Dr Durga Toshniwal
Professor (HAG)
Department of Computer Science & Engg.
Indian Institute of Technology Roorkee
durga.toshniwal@cs.iitr.ac.in, durgatoshniwal@gmail.com
www.durgatoshniwal.in
[page 2]
Data Pre-processing
and
Curation - I
GenAI and Agentic AI - Prof Durga Toshniwal
[page 3]
Major Tasks in Data Pre-processing
• Data cleaning
– Fill in missing values, smooth noisy data, and resolve
inconsistencies
• Data integration
– Integration of multiple databases, or files
• Data transformation
– Normalization and aggregation
• Data reduction
– Obtain reduced representation in volume with same or similar
analytical results
• Data discretization
– Part of data reduced by binning
GenAI and Agentic AI - Prof Durga Toshniwal
[page 4]
Data Cleaning - Handling
Missing Data
GenAI and Agentic AI - Prof Durga Toshniwal
• How can we handle missing value (5%)?
• C1: 3, C2: 6 (Majority Class)
Rowid A1 A2 A3 A4 A5 X
R1 1 2 3 1 5 C1
R2 1 4 3 2 6 C2
R3 3 2 1 2 5 C2
R4 4 2 2 1 5 C1
R5 1 1 1 3 4 C2
R6 2 2 1 1 C2
R7 1 3 6 3 9 C1
R8 1 1 1 1 10
R9 3 6 7 1 C2
R10 2 2 5 9 8 C2
[page 5]
Data Cleaning - Handling
Missing Data
GenAI and Agentic AI - Prof Durga Toshniwal
• How can we handle missing value ?
• Most Simple, remove tuples
Rowid A1 A2 A3 A4 A5 X
R1 1 2 3 1 5 C1
R2 1 4 3 2 6 C2
R3 3 2 1 2 5 C2
R4 4 2 2 1 5 C1
R5 1 1 1 3 4 C2
R7 1 3 6 3 9 C1
R10 2 2 5 9 8 C2
[page 6]
Data Cleaning - Handling
Missing Data
GenAI and Agentic AI - Prof Durga Toshniwal
• An incorrect handling of missing value can lead
to loss in data and complete Class itself !
[page 7]
Data Cleaning - Handling
Missing Data
GenAI and Agentic AI - Prof Durga Toshniwal
• How can we handle missing value ? (10%)
Some other simple Solution ?
Fill in a default value… ?
Rowid A1 A2 A3 A4 A5
R1 1 2 100 1 5
R2 1 4 3 2 6
R3 3 1 100 2 5
R4 4 2 2 1 5
R5 1 1 100 3 4
R6 2 2 100 1 1
R7 1 3 6 3 9
R8 1 1 100 1 10
R9 3 100 6 7 1
R10 2 2 5 9 8
[page 8]
GenAI and Agentic AI - Prof Durga Toshniwal
[page 9]
Imputing Missing Values with
Variable Mean
GenAI and Agentic AI - Prof Durga Toshniwal
[page 10]
Imputing Missing Values with Variable Median
GenAI and Agentic AI - Prof Durga Toshniwal
• The medians used for imputing the missing values are:
• Attribute1 (X): 12.0, Attribute2 (Y): 12.0
[page 11]
Imputing Missing Values with Variable Mode
GenAI and Agentic AI - Prof Durga Toshniwal
• The Mode Values Used:
• Attribute1 (X): 9.0, Attribute2 (Y): 20.0
[page 12]
Imputing Missing Values with Variable Mean Per Class
GenAI and Agentic AI - Prof Durga Toshniwal
• Class-wise Means Used for Imputation:
• Attribute1 (X) Class 0 Mean: 9.19, Class 1 Mean: 19.61
• Attribute2 (Y) Class 0 Mean: 9.64 Class 1 Mean: 19.65
[page 13]
Imputing Missing Values with Most Probable Value Per Class
GenAI and Agentic AI - Prof Durga Toshniwal
The missing values have been filled using the most probable (mode) value per class
for each attribute, Yellow stars represent values imputed using class-wise modes.
[page 14]
Data Cleaning - Handling Noisy Data
• Binning method:
– first sort data and partition into bins
– can smooth by bin means, smooth by bin median, smooth
by bin boundaries, etc.
• Clustering
– detect and remove outliers
• Regression
– smooth by fitting the data into regression functions
GenAI and Agentic AI - Prof Durga Toshniwal
[page 15]
Binning Methods for Data
Smoothing
• Equal-width (distance) partitioning:
– It divides the range into N intervals of equal size: uniform
grid
– if A and B are the lowest and highest values of the
attribute, the width of intervals will be: W = (B-A)/N.
GenAI and Agentic AI - Prof Durga Toshniwal
[page 16]
Example 1 of Equi-width Binning
GenAI and Agentic AI - Prof Durga Toshniwal
Age
Medical_
Expenses Bin_ID
20 571.5 0
24 474 0
25 1367.5 0
25.5 614.5 0
25.5 940 0
26 1132 0
28 1630 0
28 1083.5 0
28 949 0
30 1018.5 0
32 1054.5 1
32 932 1
34.5 1277.5 1
37 1277 1
38.5 971.5 1
Example data:
It has 50 rows, only few shown
here for illustration
Number of bins: 5 (equi-width)
• Minimum age: 20.0 years
• Maximum age: 70.0 years
Bin boundaries (age ranges):
• Bin 0: 20.0 to 30.0 years
• Bin 1: 30.0 to 40.0 years
• Bin 2: 40.0 to 50.0 years
• Bin 3: 50.0 to 60.0 years
• Bin 4: 60.0 to 70.0 years
[page 17]
Example 2 of Equi-width Binning
GenAI and Agentic AI - Prof Durga Toshniwal
Dataset Overview:
• Independent Variable: Daily Screen
Time (in hours)
• Dependent Variable: Sleep
Duration (in hours)
• Total Points: 52
• Screen Time Range: 1.4 to 7.5
hours
• Number of Bins: 6
• Bin Boundaries:
• Bin 0: 1.0 – 2.0
• Bin 1: 2.0 – 3.0
• Bin 2: 3.0 – 4.0
• Bin 3: 4.0 – 5.0
• Bin 4: 5.0 – 6.0
• Bin 5: 6.0 – 7.5
Screen
_Time
Sleep
Duration
Bin_ID
1.4 4.6 0
1.8 3.8 0
2 2.7 0
2.8 3.9 1
3.5 5.2 2
3.5 4 2
5 5.2 3
5 2.9 3
5.1 3.2 4
5.1 6.1 4
5.1 2.6 4
5.2 3.2 4
5.2 4.5 4
5.3 1.9 4
[page 18]
Binning Methods for Data
Smoothing
• Equal-width (distance) partitioning:
– It divides the range into N intervals of equal size: uniform
grid
– if A and B are the lowest and highest values of the
attribute, the width of intervals will be: W = (B-A)/N.
– In the data as per Example 2, Equi-width binning is not
suitable as all points tend to fall in just 2 bins and the rest
of the bins are almost empty
GenAI and Agentic AI - Prof Durga Toshniwal
[page 19]
Binning Methods for Data
Smoothing
• Equal-width (distance) partitioning:
– It divides the range into N intervals of equal size: uniform
grid
– if A and B are the lowest and highest values of the
attribute, the width of intervals will be: W = (B-A)/N.
– Disadvantage : If data is skewed, some bins become
crowded, some sparse
• Equal-depth (frequency) partitioning:
– It divides the range into N intervals, each containing
approximately same number of samples
GenAI and Agentic AI - Prof Durga Toshniwal
[page 20]
Example 2 Using Equi-Depth Binning
• Equal-depth (frequency) partitioning:
– It divides the range into N intervals, each containing approximately same
number of samples
GenAI and Agentic AI - Prof Durga Toshniwal
Screen_
Time
Sleep
Duration
Earlier
Bin_ID
EqFreq
Bin_ID
7 0.6 5 0
6.6 1.5 5 0
5.9 1.9 4 0
1.8 2 0 0
7.1 2.2 5 0
5.8 2.4 4 0
6.5 2.6 5 0
6.2 2.6 5 0
7 2.8 5 0
5.1 2.8 4 0
5.9 2.9 4 1
6.1 3 5 1
6.6 3 5 1
6.3 3.3 5 1
[page 21]
Example 2 Using
Equi-Depth Binning
GenAI and Agentic AI - Prof Durga Toshniwal
Bin ID Coun
t
Sleep Duration
Range (Hours)
0 10 0.6 – 2.8
1 10 2.9 – 3.7
2 10 3.8 – 4.4
3 10 4.7 – 6.6
4 12 6.7 – 10.2
[page 22]
Binning Methods for Data
Smoothing
* Sorted data for price (in INR): 4, 8, 9, 15, 21, 21, 24, 25, 26, 28,
29, 34
* Partition into (equi-depth) bins:
- Bin 1: 4, 8, 9, 15
- Bin 2: 21, 21, 24, 25
- Bin 3: 26, 28, 29, 34
* Smoothing by bin means:
- Bin 1: 9, 9, 9, 9
- Bin 2: 23, 23, 23, 23
- Bin 3: 29, 29, 29, 29
* Smoothing by bin boundaries:
- Bin 1: 4, 4, 4, 15
- Bin 2: 21, 21, 25, 25
- Bin 3: 26, 26, 26, 34
GenAI and Agentic AI - Prof Durga Toshniwal
[page 23]
Clustering for Noise Removal
GenAI and Agentic AI - Prof Durga Toshniwal
[page 24]
Regression for Noise Removal
x
y
y = x + 1
X1
Y1
Y1’
GenAI and Agentic AI - Prof Durga Toshniwal
[page 25]
Regression for Noise Removal
• Linear regression finds a model that minimizes the
distance between the fitted line and all of the data
points.
• Ordinary least squares (OLS) regression usually
minimizes the sum of the squared residuals.
• In general, if there are differences between the
observed values and the model's predicted values,
then those points away from the fitted curve, will
be removed as noise
GenAI and Agentic AI - Prof Durga Toshniwal
[page 26]
Data Transformation
• Aggregation: Data summarization
• Normalization: scaled to fall within a small, specified
range
– min-max normalization
– z-score normalization
• Attribute/feature construction:
– New attributes constructed from the given ones
GenAI and Agentic AI - Prof Durga Toshniwal
[page 27]
Data Transformation:
Normalization
• min-max normalization
• z-score normalization
AAA
AA
A
minnewminnewmaxnewminmax
minvv _)__(' +−−
−=
A
A
devstand
meanvv _' −=
GenAI and Agentic AI - Prof Durga Toshniwal
[page 28]
Data Transformation: Normalization
min-max normalization
• Enables us to convert from one range to some
other range of values
z-score normalization
• Enables us to compare two scores that are
from different normal distributions.
• The standard score does this by converting
scores in a normal distribution to z-scores in a
standard normal distribution
GenAI and Agentic AI - Prof Durga Toshniwal
[page 29]
Data Transformation: Normalization
X Y
1 1
2 3
3 3
4 3
4 6
5 7
6 4
7 9
8 11
9 9
10 10
11 7
12 15
13 17
14 14
15 15
16 21
17 17
18 15
19 24
20 25
• min-max normalization
• z-score normalization
AAA
AA
A
minnewminnewmaxnewminmax
minvv _)__(' +−−
−=
A
A
devstand
meanvv _' −=
GenAI and Agentic AI - Prof Durga Toshniwal
Z-score normalization centers the data around 0 and
scales it based on standard deviation, making it useful
for comparing variables with different units or scales.
[page 30]
Min-Max Normalization
GenAI and Agentic AI - Prof Durga Toshniwal
Minimum X: 1, Maximum X: 20, New Min X is 5, Max X is 25
Minimum Y: 1, Maximum Y: 25, New Min Y is 10, Max Y is 40
[page 31]
Z Score Normalization X Y
1 1
2 3
3 3
4 3
4 6
5 7
6 4
7 9
8 11
9 9
10 10
11 7
12 15
13 17
14 14
15 15
16 21
17 17
18 15
19 24
20 25GenAI and Agentic AI - Prof Durga Toshniwal
•X (Z-score normalized) : Minimum: -1.548, Maximum: 1.652
•Y (Z-score normalized) : Minimum: -1.448, Maximum: 1.946
[page 32]
Dimensionality Reduction
• Feature selection (i.e., attribute subset selection):
– Select a minimum set of features such that the probability
distribution of different classes given the values for those
features is as close as possible to the original distribution
given the values of all features
• Feature construction / extraction
– Derive alternate features for any given data with the aim
to reduce the data
GenAI and Agentic AI - Prof Durga Toshniwal
[page 33]
Feature Selection - Example of Decision
Tree Induction
Initial attribute set:
{A1, A2, A3, A4, A5, A6}
A4 ?
A1? A6?
Class 1 Class 2 Class 1 Class 2
Reduced attribute set: {A1, A4, A6}
GenAI and Agentic AI - Prof Durga
Toshniwal
[page 34]
Feature Selection - Decision Tree
• All internal nodes denote test on attributes, branch
corresponds to the outcome
• External nodes correspond to class prediction
• The “best” attribute is chosen to partition the data
• All attributes that do not appear on the tree are
assumed to be irrelevant
• The set of attributes appearing on the tree form
the subset
GenAI and Agentic AI - Prof Durga Toshniwal
[page 35]
Data Reduction by Numerosity
Reduction
• Parametric methods
– Assume the data fits some model, estimate model
parameters, store only the parameters, and discard the
data (except possible outliers)
• Non-parametric methods
– Do not assume models
– Major families: histograms, clustering, sampling
GenAI and Agentic AI - Prof Durga Toshniwal
[page 36]
• A popular data reduction technique
• To construct a histogram, the first step is to "bin" the
range of values or divide the entire range of values into
a series of intervals
• Then count how many values fall into each interval.
• The bins are usually specified as consecutive, non-
overlapping intervals of a variable. The bins (intervals)
must be adjacent, and are usually equal size.
Data Reduction by Numerosity
Reduction - Histograms
GenAI and Agentic AI - Prof Durga Toshniwal
[page 37]
GenAI and Agentic AI - Prof Durga Toshniwal
Example Histogram
• Histogram of the frequency of occurrence of
alphabets in first name in the Group formed by all
of you
[page 38]
Data Reduction by Sampling
• Choose a representative subset of the data
– Simple random sampling may have very poor performance
in the presence of skew
• Develop adaptive sampling methods
– Stratified sampling:
• Approximate the percentage of each class (or
subpopulation of interest) in the overall database
• Used in conjunction with skewed data
GenAI and Agentic AI - Prof Durga Toshniwal
[page 39]
Data Reduction by Sampling
Raw Data Cluster/Stratified Sample
GenAI and Agentic AI - Prof Durga
Toshniwal
[page 40]
Data Reduction
• Discretization
– reduce the number of values for a given continuous
attribute by dividing the range of the attribute into
intervals. Interval labels can then be used to replace
actual data values.
– Methods – Binning, Histogram, Clustering
• Concept hierarchies
– reduce the data by collecting and replacing low level
concepts (such as numeric values for the attribute age)
by higher level concepts (such as young, middle-aged,
or senior).
GenAI and Agentic AI - Prof Durga Toshniwal
[page 41]
Concept Hierarchy Generation
Concept hierarchy can be automatically generated
based on the number of distinct values per
attribute in the given attribute set. The attribute
with the most distinct values is placed at the
lowest level of the hierarchy.
country
province_or_ state
city
street
15 distinct values
65 distinct values
3567 distinct values
674,339 distinct values
GenAI and Agentic AI - Prof Durga Toshniwal