# Data Pre-processing and Curation -I - Sunday (09.11.2025)

course: Module 1 — Foundations of AI & ML
module: Module-1-Foundations-AI-ML
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-1-Foundations-AI-ML/General/Data_Pre-processing_and_Curation_-I_-_Sunday_(09.11.2025).pdf
pages: 41

---
[page 1]
IIT Roorkee – Futurense
PG Certification 
GenAI and Agentic AI for Engineers
Data Pre-Processing 
and Curation - I
Dr Durga Toshniwal
Professor (HAG)
Department of Computer Science & Engg.
Indian Institute of Technology Roorkee
durga.toshniwal@cs.iitr.ac.in, durgatoshniwal@gmail.com
www.durgatoshniwal.in

[page 2]
Data Pre-processing 
and 
Curation - I
GenAI and Agentic AI - Prof Durga Toshniwal

[page 3]
Major Tasks in Data Pre-processing
• Data cleaning
– Fill in missing values, smooth noisy data, and resolve 
inconsistencies
• Data integration
– Integration of multiple databases, or files
• Data transformation
– Normalization and aggregation
• Data reduction
– Obtain reduced representation in volume with same or similar 
analytical results
• Data discretization
– Part of data reduced by binning
GenAI and Agentic AI - Prof Durga Toshniwal

[page 4]
Data Cleaning  -  Handling 
Missing Data
GenAI and Agentic AI - Prof Durga Toshniwal
• How can we handle missing value (5%)?
• C1: 3, C2: 6 (Majority Class)
Rowid A1 A2 A3 A4 A5 X
R1 1 2 3 1 5 C1
R2 1 4 3 2 6 C2
R3 3 2 1 2 5 C2
R4 4 2 2 1 5 C1
R5 1 1 1 3 4 C2
R6 2 2 1 1 C2
R7 1 3 6 3 9 C1
R8 1 1 1 1 10
R9 3 6 7 1 C2
R10 2 2 5 9 8 C2

[page 5]
Data Cleaning  -  Handling 
Missing Data
GenAI and Agentic AI - Prof Durga Toshniwal
• How can we handle missing value ?
• Most Simple, remove tuples
Rowid A1 A2 A3 A4 A5 X
R1 1 2 3 1 5 C1
R2 1 4 3 2 6 C2
R3 3 2 1 2 5 C2
R4 4 2 2 1 5 C1
R5 1 1 1 3 4 C2
R7 1 3 6 3 9 C1
R10 2 2 5 9 8 C2

[page 6]
Data Cleaning  -  Handling 
Missing Data
GenAI and Agentic AI - Prof Durga Toshniwal
• An incorrect handling of missing value can lead 
to loss in data and complete Class itself !

[page 7]
Data Cleaning  -  Handling 
Missing Data
GenAI and Agentic AI - Prof Durga Toshniwal
• How can we handle missing value ? (10%)
    Some other simple Solution ?
     Fill in a default value… ?
    
Rowid A1 A2 A3 A4 A5
R1 1 2 100 1 5
R2 1 4 3 2 6
R3 3 1 100 2 5
R4 4 2 2 1 5
R5 1 1 100 3 4
R6 2 2 100 1 1
R7 1 3 6 3 9
R8 1 1 100 1 10
R9 3 100 6 7 1
R10 2 2 5 9 8

[page 8]
GenAI and Agentic AI - Prof Durga Toshniwal

[page 9]
Imputing Missing Values with 
Variable Mean
GenAI and Agentic AI - Prof Durga Toshniwal

[page 10]
Imputing Missing Values with Variable Median
GenAI and Agentic AI - Prof Durga Toshniwal
• The medians used for imputing the missing values are:
• Attribute1 (X): 12.0,       Attribute2 (Y): 12.0

[page 11]
Imputing Missing Values with Variable Mode
GenAI and Agentic AI - Prof Durga Toshniwal
• The Mode Values Used:
• Attribute1 (X): 9.0,     Attribute2 (Y): 20.0

[page 12]
Imputing Missing Values with Variable Mean Per Class
GenAI and Agentic AI - Prof Durga Toshniwal
• Class-wise Means Used for Imputation:
• Attribute1 (X)   Class 0 Mean: 9.19, Class 1 Mean: 19.61
• Attribute2 (Y)   Class 0 Mean: 9.64 Class 1 Mean: 19.65

[page 13]
Imputing Missing Values with Most Probable Value Per Class
GenAI and Agentic AI - Prof Durga Toshniwal
The missing values have been filled using the most probable (mode) value per class 
for each attribute, Yellow stars represent values imputed using class-wise modes.

[page 14]
Data Cleaning  -  Handling Noisy Data
• Binning method:
– first sort data and partition into bins
– can smooth by bin means,  smooth by bin median, smooth 
by bin boundaries, etc.
• Clustering
– detect and remove outliers
• Regression
– smooth by fitting the data into regression functions
GenAI and Agentic AI - Prof Durga Toshniwal

[page 15]
Binning Methods for Data 
Smoothing
• Equal-width (distance) partitioning:
– It divides the range into N intervals of equal size: uniform 
grid
– if A and B are the lowest and highest values of the 
attribute, the width of intervals will be: W = (B-A)/N.
GenAI and Agentic AI - Prof Durga Toshniwal

[page 16]
Example 1 of Equi-width Binning 
GenAI and Agentic AI - Prof Durga Toshniwal
Age
Medical_
Expenses Bin_ID
20 571.5 0
24 474 0
25 1367.5 0
25.5 614.5 0
25.5 940 0
26 1132 0
28 1630 0
28 1083.5 0
28 949 0
30 1018.5 0
32 1054.5 1
32 932 1
34.5 1277.5 1
37 1277 1
38.5 971.5 1
Example data:
It has 50 rows, only few shown 
here for illustration
    Number of bins: 5 (equi-width)
• Minimum age: 20.0 years
• Maximum age: 70.0 years
    Bin boundaries (age ranges):
• Bin 0: 20.0 to 30.0 years
• Bin 1: 30.0 to 40.0 years
• Bin 2: 40.0 to 50.0 years
• Bin 3: 50.0 to 60.0 years
• Bin 4: 60.0 to 70.0 years

[page 17]
Example 2 of Equi-width Binning 
GenAI and Agentic AI - Prof Durga Toshniwal
Dataset Overview:
• Independent Variable: Daily Screen 
Time (in hours)
• Dependent Variable: Sleep 
Duration (in hours)
• Total Points: 52
• Screen Time Range: 1.4 to 7.5 
hours
• Number of Bins: 6
• Bin Boundaries:
• Bin 0: 1.0 – 2.0 
• Bin 1: 2.0 – 3.0 
• Bin 2: 3.0 – 4.0 
• Bin 3: 4.0 – 5.0 
• Bin 4: 5.0 – 6.0 
• Bin 5: 6.0 – 7.5 
Screen
_Time
Sleep 
Duration
Bin_ID
1.4 4.6 0
1.8 3.8 0
2 2.7 0
2.8 3.9 1
3.5 5.2 2
3.5 4 2
5 5.2 3
5 2.9 3
5.1 3.2 4
5.1 6.1 4
5.1 2.6 4
5.2 3.2 4
5.2 4.5 4
5.3 1.9 4

[page 18]
Binning Methods for Data 
Smoothing
• Equal-width (distance) partitioning:
– It divides the range into N intervals of equal size: uniform 
grid
– if A and B are the lowest and highest values of the 
attribute, the width of intervals will be: W = (B-A)/N.
– In the data as per Example 2, Equi-width binning is not 
suitable as all points tend to fall in just 2 bins and the rest 
of the bins are almost empty
GenAI and Agentic AI - Prof Durga Toshniwal

[page 19]
Binning Methods for Data 
Smoothing
• Equal-width (distance) partitioning:
– It divides the range into N intervals of equal size: uniform 
grid
– if A and B are the lowest and highest values of the 
attribute, the width of intervals will be: W = (B-A)/N.
– Disadvantage : If data is skewed, some bins become 
crowded, some sparse
• Equal-depth (frequency) partitioning:
– It divides the range into N intervals, each containing 
approximately same number of samples
GenAI and Agentic AI - Prof Durga Toshniwal

[page 20]
Example 2 Using Equi-Depth Binning
• Equal-depth (frequency) partitioning:
– It divides the range into N intervals, each containing approximately same 
number of samples
GenAI and Agentic AI - Prof Durga Toshniwal
Screen_
Time
Sleep 
Duration
Earlier 
Bin_ID
EqFreq 
Bin_ID
7 0.6 5 0
6.6 1.5 5 0
5.9 1.9 4 0
1.8 2 0 0
7.1 2.2 5 0
5.8 2.4 4 0
6.5 2.6 5 0
6.2 2.6 5 0
7 2.8 5 0
5.1 2.8 4 0
5.9 2.9 4 1
6.1 3 5 1
6.6 3 5 1
6.3 3.3 5 1

[page 21]
Example 2 Using 
Equi-Depth Binning
GenAI and Agentic AI - Prof Durga Toshniwal
Bin ID Coun
t
Sleep Duration 
Range (Hours)
0 10 0.6 – 2.8
1 10 2.9 – 3.7
2 10 3.8 – 4.4
3 10 4.7 – 6.6
4 12 6.7 – 10.2

[page 22]
Binning Methods for Data 
Smoothing
*  Sorted data for price (in INR): 4, 8, 9, 15, 21, 21, 24, 25, 26, 28, 
29, 34
*  Partition into (equi-depth) bins:
      - Bin 1: 4, 8, 9, 15
      - Bin 2: 21, 21, 24, 25
      - Bin 3: 26, 28, 29, 34
*  Smoothing by bin means:
      - Bin 1: 9, 9, 9, 9
      - Bin 2: 23, 23, 23, 23
      - Bin 3: 29, 29, 29, 29
*  Smoothing by bin boundaries:
      - Bin 1: 4, 4, 4, 15
      - Bin 2: 21, 21, 25, 25
      - Bin 3: 26, 26, 26, 34
GenAI and Agentic AI - Prof Durga Toshniwal

[page 23]
Clustering for Noise Removal
GenAI and Agentic AI - Prof Durga Toshniwal

[page 24]
Regression for Noise Removal
x
y
y = x + 1
X1
Y1
Y1’
GenAI and Agentic AI - Prof Durga Toshniwal

[page 25]
Regression for Noise Removal
• Linear regression finds a model that minimizes the 
distance between the fitted line and all of the data 
points. 
• Ordinary least squares (OLS) regression usually 
minimizes the sum of the squared residuals.
• In general, if there are differences between the 
observed values and the model's predicted values, 
then those points away from the fitted curve, will 
be removed as noise
GenAI and Agentic AI - Prof Durga Toshniwal

[page 26]
Data Transformation
• Aggregation: Data summarization
• Normalization: scaled to fall within a small, specified 
range
– min-max normalization
– z-score normalization
• Attribute/feature construction:
– New attributes constructed from the given ones
GenAI and Agentic AI - Prof Durga Toshniwal

[page 27]
Data Transformation: 
Normalization
• min-max normalization
• z-score normalization
AAA
AA
A
minnewminnewmaxnewminmax
minvv _)__(' +−−
−=
A
A
devstand
meanvv _' −=
GenAI and Agentic AI - Prof Durga Toshniwal

[page 28]
Data Transformation: Normalization
min-max normalization
• Enables us to convert from one range to some 
other range of values
z-score normalization
• Enables us to compare two scores that are 
from different normal distributions. 
• The standard score does this by converting 
scores in a normal distribution to z-scores in a 
standard normal distribution
GenAI and Agentic AI - Prof Durga Toshniwal

[page 29]
Data Transformation: Normalization
X Y
1 1
2 3
3 3
4 3
4 6
5 7
6 4
7 9
8 11
9 9
10 10
11 7
12 15
13 17
14 14
15 15
16 21
17 17
18 15
19 24
20 25
• min-max normalization
• z-score normalization
AAA
AA
A
minnewminnewmaxnewminmax
minvv _)__(' +−−
−=
A
A
devstand
meanvv _' −=
GenAI and Agentic AI - Prof Durga Toshniwal
Z-score normalization centers the data around 0 and 
scales it based on standard deviation, making it useful 
for comparing variables with different units or scales.

[page 30]
Min-Max Normalization
GenAI and Agentic AI - Prof Durga Toshniwal
Minimum X: 1, Maximum X: 20, New Min X is 5, Max X is 25
Minimum Y: 1, Maximum Y: 25, New Min Y is 10, Max Y is 40

[page 31]
Z Score Normalization X Y
1 1
2 3
3 3
4 3
4 6
5 7
6 4
7 9
8 11
9 9
10 10
11 7
12 15
13 17
14 14
15 15
16 21
17 17
18 15
19 24
20 25GenAI and Agentic AI - Prof Durga Toshniwal
•X (Z-score normalized) : Minimum: -1.548, Maximum: 1.652
•Y (Z-score normalized) : Minimum: -1.448, Maximum: 1.946

[page 32]
Dimensionality Reduction
•     Feature selection (i.e., attribute subset selection):
– Select a minimum set of features such that the probability 
distribution of different classes given the values for those 
features is as close as possible to the original distribution 
given the values of all features
• Feature construction / extraction
– Derive alternate features for any given data with the aim 
to reduce the data
GenAI and Agentic AI - Prof Durga Toshniwal

[page 33]
Feature Selection - Example of Decision 
Tree Induction
Initial attribute set:
{A1, A2, A3, A4, A5, A6}
A4 ?
A1? A6?
Class 1 Class 2 Class 1 Class 2
Reduced attribute set:  {A1, A4, A6}
GenAI and Agentic AI - Prof Durga 
Toshniwal

[page 34]
Feature Selection - Decision Tree
• All internal nodes denote test on attributes, branch 
corresponds to the outcome
• External nodes correspond to class prediction
• The “best” attribute is chosen to partition the data
• All attributes that do not appear on the tree are 
assumed to be irrelevant
• The set of attributes appearing on the tree form 
the subset
GenAI and Agentic AI - Prof Durga Toshniwal

[page 35]
Data Reduction by Numerosity 
Reduction
• Parametric methods
– Assume the data fits some model, estimate model 
parameters, store only the parameters, and discard the 
data (except possible outliers)
• Non-parametric methods 
– Do not assume models
– Major families: histograms, clustering, sampling 
GenAI and Agentic AI - Prof Durga Toshniwal

[page 36]
• A popular data reduction technique
• To construct a histogram, the first step is to "bin" the 
range of values or divide the entire range of values into 
a series of intervals
• Then count how many values fall into each interval. 
• The bins are usually specified as consecutive, non-
overlapping intervals of a variable. The bins (intervals) 
must be adjacent, and are usually equal size.
Data Reduction by Numerosity 
Reduction - Histograms
GenAI and Agentic AI - Prof Durga Toshniwal

[page 37]
GenAI and Agentic AI - Prof Durga Toshniwal
Example Histogram
• Histogram of the frequency of occurrence of   
alphabets in first name in the Group formed by all 
of you

[page 38]
Data Reduction by Sampling
• Choose a representative subset of the data
– Simple random sampling may have very poor performance 
in the presence of skew
• Develop adaptive sampling methods
– Stratified sampling: 
• Approximate the percentage of each class (or 
subpopulation of interest) in the overall database 
• Used in conjunction with skewed data
GenAI and Agentic AI - Prof Durga Toshniwal

[page 39]
Data Reduction by Sampling
Raw Data Cluster/Stratified Sample
GenAI and Agentic AI - Prof Durga 
Toshniwal

[page 40]
Data Reduction
• Discretization 
– reduce the number of values for a given continuous 
attribute by dividing the range of the attribute into 
intervals. Interval labels can then be used to replace 
actual data values.
– Methods – Binning, Histogram, Clustering
• Concept hierarchies 
– reduce the data by collecting and replacing low level 
concepts (such as numeric values for the attribute age) 
by higher level concepts (such as young, middle-aged, 
or senior).
GenAI and Agentic AI - Prof Durga Toshniwal

[page 41]
Concept Hierarchy Generation
Concept hierarchy can be automatically generated 
based on the number of distinct values per 
attribute in the given attribute set. The attribute 
with the most distinct values is placed at the 
lowest level of the hierarchy.
country
province_or_ state
city
street
15 distinct values
65 distinct values
3567 distinct values
674,339 distinct values
GenAI and Agentic AI - Prof Durga Toshniwal