# 02 2025-11-09 Data Pre-Processing and Curation

course: Module 1 — Foundations of AI & ML
module: Module-1-Foundations-AI-ML
date: 2025-11-09
type: transcript
video_url: https://personal-learn.armco.dev/files/_Recordings/Module-1-Foundations-AI-ML/02_2025-11-09_Data_Pre-Processing_and_Curation.mp4

---
[00:19:34] There are immense applications of classification.
[00:19:38] Such as segmentation, frog detection, etc.
[00:19:43] Then, uh, I introduced to you the concept of discriminative models and the classification, uh,
[00:19:50] is a discriminative model.
[00:19:53] Which helps us to distinguish between the different categories that are available.
[00:19:58] In the data.
[00:20:00] And discriminative models usually
[00:20:03] help to classify, and then use this classification model.
[00:20:07] To predict the class labels from the input data.
[00:20:11] And, uh, the idea is to find out a function.
[00:20:15] Uh, that helps to, uh, make a boundary, uh, that separates the different classes that are there in the given data.
[00:20:23] And the output here would be the class label.
[00:20:27] Now, talking about unsupervised learning…
[00:20:30] That, in term, unsupervised learning is coined
[00:20:34] Because, uh, here in this case, we do not have any class labels given to us, we just have that data.
[00:20:41] And we try to group that data, or find out the natural groupings that is available in the data.
[00:20:47] With the help of, uh, similarity measures.
[00:20:52] So, similarity measures are those neat tricks or measures that help us to
[00:20:56] assess or quantify the similarity between data points.
[00:21:01] And based on that singularity, the points are clustered together, or grouped together.
[00:21:06] In such a way that points which are…
[00:21:09] Uh, very close together.
[00:21:11] output in the same cluster, thereby…
[00:21:14] Minimizing the intra-cluster distances and maximizing the inter-cluster distances.
[00:21:20] So, then we talked about, uh…
[00:21:24] clustering applications,
[00:21:29] Uh, then we discuss regression and forecasting.
[00:21:32] So, I, uh, had introduced to you, uh, what is predictive analytics and how
[00:21:38] Classification plays a role in predictive analytics.
[00:21:43] Uh, when we talk about predicting value, then it's called forecasting, and regression can be used for forecasting.
[00:21:51] We also talked about anomaly detection, where we try to identify abnormal behavior.
[00:21:57] from the given data, it's also called outlier.
[00:22:01] identification, and different types of supervised and unsupervised learning methods can help, too.
[00:22:08] perform anomaly detection.
[00:22:11] Then, I introduce to you deep learning techniques, where, uh, neural networks are used.
[00:22:17] Which are inspired by the human brain and human neurons.
[00:22:21] Which are there in human brain.
[00:22:23] And whenever the data that we are having is very complex, and the patterns are also very complicated,
[00:22:30] Uh, then deep learning methods are used.
[00:22:33] Such as on data like text, images, audio, etc.
[00:22:38] And there are many deep learning techniques, like convolutional neural networks, recurrent neural networks, long-short-term memory,
[00:22:46] LSTM. Uh, and so on and so forth, which we'll be covering in the subsequent, uh, sections.
[00:22:54] Then, deep learning has immense applications in various domains, like computer vision, NLP, speech recognition, healthcare, finance, and…
[00:23:03] So, so for…
[00:23:06] And finally, we talked about generative AI.
[00:23:10] Uh, well, um…
[00:23:12] There are… we talk about AI models that are able not just to
[00:23:18] Analyze the data, but to create new content.
[00:23:21] Uh, that new content could be text, audio, video, images, code,
[00:23:25] Or, uh, any other, uh, data types.
[00:23:29] And they are capable of understanding the context, the style of writing, the structure,
[00:23:35] That should be available. Uh…
[00:23:37] In a way that is similar to human-like content generation.
[00:23:42] And, uh, generative AI, uh, models work on huge quantities of data, or corpus.
[00:23:49] And they use this to learn the patterns, and then…
[00:23:52] After learning, they generate new patterns, which may be similar
[00:23:57] to the patterns, they learn, or they may be of different types.
[00:24:02] Then, uh…
[00:24:04] So, uh, unlike the discriminative models,
[00:24:09] Generative AI talks about generative models.
[00:24:14] So, the purpose of discriminative model is just to classify and predict class labels.
[00:24:18] However, the purpose of generative AI is to generate new samples.
[00:24:23] And uh… there are different, uh…
[00:24:27] types of algorithms that fall in each of these categories. For example, discriminative model could be logistic regression.
[00:24:34] SVM or other classifiers, whereas generative model.
[00:24:38] could be GPT, but…
[00:24:40] guns, and so on and so forth.
[00:24:44] And then, um, the…
[00:24:47] Generative AI models progress to agentic AI.
[00:24:52] Where there are agents, uh…
[00:24:54] one or multiple agents. Usually, there are multiple agents.
[00:25:00] start to work in collaboration with each other, and…
[00:25:03] They are, uh, autonomous in nature, and they are capable of…
[00:25:07] learning and objective.
[00:25:09] or goal, then working towards…
[00:25:12] Completion of the goal.
[00:25:14] planning, coordinating…
[00:25:17] Uh, dynamically changing, taking feedback.
[00:25:20] improving, and so on and so forth. They have huge, uh, capabilities. And they act like assistants.
[00:25:28] Collaborative assistance, uh, which help humans.
[00:25:31] to perform the task with minimal human intervention.
[00:25:36] So, there are, again, a lot of applications of agent-type AI.
[00:25:41] So this is what we had covered.
[00:25:43] In the, uh, initial part.
[00:25:47] Then, uh…
[00:25:50] Going to the next portion.
[00:25:54] I will just share the slides.
[00:25:59] Uh… so I think I could see a query on the slides on LMS portal.
[00:26:06] Uh, so I have shared the slides, uh, today morning, uh, yesterday it was a little late.
[00:26:12] And I'm sure that they will be uploaded very soon on the LMS, so you will find, uh…
[00:26:17] the slides for yesterday on LMS, and also for today. So, I will be…
[00:26:22] Uh, sharing, uh, to the concerned Futurance people in the afternoon today. Today's slides.
[00:26:29] And they'll also be available on LMS, so you may… please don't worry, you will find them.
[00:26:36] Okay, so let me now…
[00:26:42] Share the other slides.
[00:27:01] Okay, so I hope you all can…
[00:27:04] see my screen.
[00:27:08] And let me maximize it.
[00:27:13] Okay, so yesterday we started with the second portion, which was on data preprocessing and curation.
[00:27:20] So, uh, I discussed that there is a requirement of the data to be, uh, quality data, so that
[00:27:27] the analysis on such a data.
[00:27:30] could yield, uh, good results.
[00:27:33] Uh, and there are different tasks that can be performed as part of data preprocessing.
[00:27:39] Such as data cleaning,
[00:27:42] data integration, data transformation, reduction, and discretization.
[00:27:47] We started out discussing some of these yesterday.
[00:27:51] So, we discussed, uh, in-data cleaning, handling of missing data.
[00:27:56] And, uh, we discussed how, uh, different, uh, percentages of missing
[00:28:01] values in the given data could impact
[00:28:06] Uh, the missing value imputation.
[00:28:09] So, we concluded that if the percentage
[00:28:12] of missing values are very less, then we could actually go with removing
[00:28:17] those records. However, removing of records is not a very good solution in case
[00:28:23] Uh, the missing values are…
[00:28:26] higher… higher in percentage.
[00:28:30] Uh, so, for example, uh, if the…
[00:28:34] We discussed that if around 30% of the data
[00:28:37] has missing values, and uh… in this example, if we remove such rows,
[00:28:43] then it is possible that the majority, and if there is a majority of one particular class, like Class C2,
[00:28:50] It is possible that it can become a uniform or almost uniform distribution.
[00:28:55] Uh, and it can also be possible
[00:28:59] that one of the classes can be completely lost, and the minority class could become a majority class.
[00:29:05] Uh, if the percentage of…
[00:29:08] Missing values is high.
[00:29:15] then we, uh, talked about other solutions.
[00:29:20] So, uh…
[00:29:24] So these are just the examples that I had shown to you with different percentages of missing values.
[00:29:31] Also, uh…
[00:29:35] There, uh, we talked about different solutions, uh, for this problem.
[00:29:40] So, another solution could be to fill in with a default value.
[00:29:44] say, any value, say 100, 9, 10, or whatever.
[00:29:48] And, uh, we also, uh, discussed at length that feeling of missing values could actually lead to the creation of a new class.
[00:29:57] In the data, if the percentage of the missing value is high.
[00:30:03] Then, uh, the next solution was filling, um, a mean value.
[00:30:07] So, when, uh, mean value is filled,
[00:30:10] Again, there is a possibility, and the mean could be a classified mean, or whatever way.
[00:30:17] Uh, there's again a possibility.
[00:30:19] that, uh, uh…
[00:30:22] The points that are, uh…
[00:30:25] The points that get those values that are being filled on the basis of mean,
[00:30:29] are leading to the creation of points, uh, which are very different from the distribution.
[00:30:35] The regular distribution of the given set of points.
[00:30:39] So, feeling of, uh…
[00:30:41] Missing values with mean is also not a very good solution.
[00:30:46] Uh, then, uh, we, uh, also discussed filling off, uh, missing values using median.
[00:30:54] So, when we use median,
[00:30:57] Then, filling of missing values.
[00:30:58] uh… happens in a better fashion, but…
[00:31:02] It's still not a very good, uh, example, because again, we are creating points.
[00:31:07] Which are actually very different from the distribution of the points given to us.
[00:31:14] Then we also tried to fill with mode, and we found out, again, that, uh, in some cases,
[00:31:22] It is a good solution, whereas in the other cases,
[00:31:24] It is not a very good solution.
[00:31:27] Uh, to do that.
[00:31:31] So, uh, then we also talked about…
[00:31:37] About filling of, uh, missing values using other methods, so mean, median, and mode.
[00:31:45] are somewhat better, but they are not that much better.
[00:31:52] So, we, uh, actually were in a need to have other methods also.
[00:31:58] Sorry.
[00:32:01] And, uh, we talked about feeling of mean based on, uh, classes, uh, class labels.
[00:32:08] For the corresponding attribute, and this…
[00:32:11] came out to be a much better solution.
[00:32:14] And even, uh, more beta solution was to use probabilistic
[00:32:19] methods to fill the missing values.
[00:32:23] And this turned out to be probably the best solution.
[00:32:28] Uh, that we discussed. So, I think we had finished this much yesterday. Now I'm going to continue with the…
[00:32:35] New portion.
[00:32:36] Okay.
[00:32:39] So… so these are the different solutions that we discussed.
[00:32:44] Uh, ignoring of the tuple or removing it.
[00:32:47] Then filling, uh, with missing values, some missing values, uh, manually.
[00:32:52] Filling a global constant.
[00:32:56] Or, uh, using the mean media norm mode, using attribute mean.
[00:33:00] On using the most probable value to fill the missing value, which turned out to be the best solution.
[00:33:07] Now, in data cleaning, uh, in addition to, uh,
[00:33:12] handling missing value.
[00:33:14] Uh, the next thing that we wanted to do was handling noisy data.
[00:33:20] So, we are going to discuss methods where we are going to handle noisy data.
[00:33:26] Just a second… yeah.
[00:33:29] So, there are… there can be different ways. Uh, one, um, common way of
[00:33:34] handling noisy data is to perform binning.
[00:33:37] So, in beginning, as the word bin implies, VIN is a container.
[00:33:42] or a bucket, where we first sort the data points.
[00:33:46] And then partition them into groups.
[00:33:49] On bins? Are we…
[00:33:51] Uh, assume that we are partitioning and putting the points in bins.
[00:33:55] And then we can apply different types of smoothing.
[00:33:59] Uh, on the bins.
[00:34:01] Which I will discuss as we go along.
[00:34:03] So, binning is one popular method of handling noisy data.
[00:34:08] Where we just retain the VIN information and do away with the noise.
[00:34:13] Then, clustering also can be used to remove noise or outliers.
[00:34:18] Regression also can be used to fit a curve, and all the points that do not lie on the fitted
[00:34:25] Uh, regression function.
[00:34:28] are actually removed as noise points.
[00:34:32] So, let's take each of these methods in more details as we go along.
[00:34:36] So, we have binning, uh, for our data smoothing, or removing noise.
[00:34:42] Again, in binning, there are different methods to perform binning.
[00:34:45] So, the first method to perform binning
[00:34:48] uh… is…
[00:34:50] equivalent partitioning, or equivalent bins.
[00:34:56] So, in this particular case, what we do is that we take the range of values for that… for the…
[00:35:01] a variable that we have to bid. For example, if we have height,
[00:35:05] And we want to bin the heights. Then we take the range of values,
[00:35:10] Uh, within which the height is falling in the given set of data.
[00:35:15] And then, uh, we divide it by the total number of bins that we want to form.
[00:35:22] For example, here, we are interested to form capital N number of bins.
[00:35:27] This doesn't hang up. So, these are the number of VINs that we want to form. So, we will divide
[00:35:33] The given range of values, let's say the highest value was B,
[00:35:38] And the lowest value was A for that particular attribute that we want to build.
[00:35:44] So, uh, the width of the VIN will be given by B minus A divided by N.
[00:35:49] So, uh, this is how, uh…
[00:35:52] VINs are formed, just a second.
[00:36:02] So, in this particular case, uh, all the VINs are having equal, uh, width.
[00:36:07] So, what are we trying to do? Let's say, uh…
[00:36:11] This is the VIN number.
[00:36:14] And, uh…
[00:36:16] What we are having here…
[00:36:19] is some variable.
[00:36:22] And what we try to do is that we try to divide it into equal width
[00:36:28] partitions, like this.
[00:36:30] So we are actually kind of dividing it into strips.
[00:36:35] So here's an example. Uh, let's assume we have a data in which there are 50 rows. Not all 50 are shown here. Some of them are shown here only.
[00:36:44] For illustration purpose.
[00:36:46] So, equivalent winning is to be done on the data, assuming the total number of bins that we want to firm up is 5.
[00:36:53] The minimum age in this data is 20, and the maximum age is 70.
[00:37:00] Then, uh, what we will do is, uh, we'll determine the bin size as…
[00:37:06] Uh, uh, you know, given by…
[00:37:08] 50, 70 minus 20.
[00:37:10] So it will be 70 minus 20.
[00:37:13] divided by 5, which will be 10.
[00:37:16] So, the width of the VIN will be 10.
[00:37:19] And the different VINs that can be formed will be VIN0, VIN 1, VIN 2, VIN 3, VIN 4.
[00:37:24] Which will be of size 10 each, like 20 to 30, 30 to 40, and so on and so forth, till 70.
[00:37:31] Now, all the points are assigned to these pins.
[00:37:35] So, for example, the age 20.
[00:37:37] So, age 20 is going to belong to bin 0.
[00:37:41] So, it is assigned bin ID 0.
[00:37:43] So, any age in 20s is assigned to bin 0, so we can see here.
[00:37:48] Because the data is sorted, so the initial values
[00:37:52] will be assigned to, till here, will be assigned to.
[00:37:55] bin 0, and uh…
[00:37:57] So, we can assume that, uh…
[00:38:01] Any particular VIN is inclusive of the
[00:38:03] upper bound. So, this means that this pin…
[00:38:06] includes values from 20 up to…
[00:38:10] 30. And any value greater than 30 will come in bin 1.
[00:38:14] So, up to 30, we are having bin 0.
[00:38:17] Similarly, up to 40 we are having been…
[00:38:20] 2, and so on and so forth. VIN 1 and so and so forth.
[00:38:24] So now, if we visualize this data…
[00:38:28] What we can see here is these are the VIN IDs.
[00:38:34] Justice…
[00:38:36] Yeah, so we have the different VIN IDs, as you can see here.
[00:38:40] And, uh, let's say we are binning…
[00:38:43] Medical expenses in rupees.
[00:38:46] Uh, so the medical expenses, when they are binned,
[00:38:49] Then what we find out, uh…
[00:38:52] On the basis of VIN information.
[00:38:56] that, uh, these are the 5 bins, and the crosses show the points that have been assigned to these VINs. So, these many points are assigned.
[00:39:03] 2.0, bin 1, bin 2, bin 3, and bin 4.
[00:39:08] So, uh…
[00:39:10] So, what are your comments on this kind of assignment?
[00:39:20] any comments or anything that, uh…
[00:39:22] to note, here.
[00:39:24] To form a equivalent cluster, like, because the data is going to be in a specific range, so it is, like, a more likely
[00:39:33] Okay.
[00:39:35] And the distance is less, uh, compared to…
[00:39:39] those, like, the 4 has a distance, uh, lesser than…
[00:39:44] Uh, the two are…
[00:39:48] 3…
[00:39:50] Uh, you're talking about the…
[00:39:52] distance the points.
[00:39:53] The distance between the, uh, each other.
[00:39:56] Oh. Oh.
[00:39:57] Good point.
[00:39:58] intercluster, yeah.
[00:40:01] It is.
[00:40:02] As it seems like we have customized, uh, this.
[00:40:03] Just one question. So, is winning software.
[00:40:08] Can you come back again, Neeraj?
[00:40:09] data… Yeah, it seems like that we have categorized the data.
[00:40:17] Yeah, so…
[00:40:18] Into… Into 5 correctly.
[00:40:19] Yeah, so actually, bins are kind of, uh, creating categories only.
[00:40:25] So, we have now created categories out of our data. One category is people in their 20s up to 30.
[00:40:33] people in their 30s up to 40 like that.
[00:40:36] And, uh, definitely,
[00:40:38] Uh, we can see the distribution of the points, uh, more or less.
[00:40:42] Uh, average number of points in the VIN might be…
[00:40:46] not exactly equal, but…
[00:40:49] somewhat, uh…
[00:40:51] uniform. Of course, there have been 3 and 4 are more populated than VIN 0, 1, and 2.
[00:40:58] Uh, but all the VINs do have points in them.
[00:41:02] So, Deepak, do you have a question?
[00:41:05] Yeah, I'm just trying to get the relation between binning and the clustering, so is it… Somewhere related because I think we are doing the same thing here as well.
[00:41:16] You want to understand, uh, the relationship between bins and clusters?
[00:41:22] Yeah, so is it the part of the process, uh, to create a cluster?
[00:41:27] Uh, you can think of them as clusters, but these are kind of static clusters.
[00:41:33] So, when we perform clustering of the data,
[00:41:37] We find out the similarity between the points.
[00:41:40] And on the basis of the similarity, we cluster them or group them.
[00:41:45] Here, we are not finding out
[00:41:47] Any similarity across the points. All we are doing is, based on the value,
[00:41:52] of the attribute that we want to bin, we are just assigning the VIN ID, some bin ID,
[00:41:58] to the point. So, there is no point-to-point similarity that is getting calculated here.
[00:42:04] So, binning is not exactly clustering.
[00:42:07] However, it does form groups in some way.
[00:42:09] But these are not exactly clusters.
[00:42:12] Is it okay?
[00:42:15] Yeah. Uh…
[00:42:17] Then, uh, of course, uh…
[00:42:20] While my circles were there, so…
[00:42:22] There can be points, I mean,
[00:42:25] So these points, for example,
[00:42:27] Either they could be considered… so here, since we are not talking about, uh…
[00:42:33] Anomaly detection, so we are assuming all the points to be regular or normal points.
[00:42:38] So, uh, we may not say that they are outliers, but if required, they can be treated as outliers if they are really different from the rest of the points.
[00:42:48] As of now, we are assuming that there are no outliers in the data. So, actually, all of them will be form of
[00:42:54] will be part of the bin. All the points. So, even though my circles…
[00:42:58] got a little smaller.
[00:43:01] Okay, so Neil, what question do you have?
[00:43:03] Yeah, hello, ma'am. So, I'm just trying to understand, so…
[00:43:07] What is the noise in this data? What we are trying to do here?
[00:43:12] When we are dividing it into the bins.
[00:43:16] So, uh, so actually, uh,
[00:43:19] Noise points are removed. Yeah, so if we are doing denoising using binning, which is what I started out with,
[00:43:29] So, uh, actually, uh…
[00:43:32] What is happening here is that there are, let's say, if we were to plot these points,
[00:43:36] Like this, let's say, then they would be, like, jittery, right? There will be points here in all the range, like this.
[00:43:44] So, what are we doing? We are smoothing it out by replacing it with some value.
[00:43:49] And that value would be the VIN ID. Say, I call this…
[00:43:54] You know, this as all the values inside this, I'm replacing it with some range, R1.
[00:43:59] And I call it VIN0. All these values, I'm replacing with some other
[00:44:04] So, uh, like this, we are actually removing all the jitter and just replacing them by these VIN IDs.
[00:44:12] So, we are kind of smoothing this data, replacing all these jitters with these simple numbers.
[00:44:16] This is how we are denoising the data here.
[00:44:18] Okay, so I think it is somewhat related to the classes or clusters.
[00:44:26] So, once again, I'll say that in clustering, we try to look at the similarity of the points.
[00:44:32] Here, we are not looking at similarity, we are just dividing the given range, which I already showed you.
[00:44:38] If this is the highest value, this is the lowest value.
[00:44:41] We just subtract them and divide them by the total number of bins to be formed.
[00:44:46] So this gives us the width of the bin.
[00:44:48] And then, based on the width, we just keep on assigning points. You can visualize it like this. Let's say if I have 5 wins, you think of 5 buckets.
[00:44:57] And let's say if you have some colored balls, then each bucket is going to hold one particular color.
[00:45:02] So, as you have the balls in your hand, you keep tossing it.
[00:45:05] In the respective bucket, if there's a bucket for white balls, you toss it in the…
[00:45:09] And you have a white ball with you, then you toss it in the respective bucket and like that.
[00:45:14] So instead of having all the balls separately with different colors, now what we have are buckets.
[00:45:20] And we just handle the buckets rather than the individual balls. So, similarly here,
[00:45:26] We are having bins or buckets. We, uh, just place the points in each of these bins, and then forget about the points.
[00:45:33] And we are only interested to retain the VIN information.
[00:45:37] And we are going to use that only. So, in a way, we are kind of categorizing the data.
[00:45:41] from, uh, continuous valued, uh, you know, range to a range, uh, which is, uh,
[00:45:50] finite ranges. Like, for example,
[00:45:53] If we talk about…
[00:45:55] Uh, the age from 20 up to 30.
[00:45:57] Then there could be infinite values. Instead of that, I'm replacing it by…
[00:46:03] one value, zero.
[00:46:04] Like that.
[00:46:05] Go to the area, thanks.
[00:46:13] So then, uh…
[00:46:14] We have another, uh, example.
[00:46:18] Where we are talking about, uh, daily, uh, screen time and sleep duration.
[00:46:24] And, uh, the dependent variable here is leave duration.
[00:46:29] Uh, so, uh…
[00:46:31] So, sleep duration is what is going to be, uh, the dependent variable
[00:46:38] And the rest of the screen time is independent variable. The total number of points that
[00:46:43] are given a 52, and the screen, uh, time range is given.
[00:46:48] And then, uh…
[00:46:50] We want to form 6 VINs, and according to that, uh…
[00:46:54] According to the…
[00:46:57] Equiva at the beginning, uh, we divide the given set of points from the highest to the lowest value.
[00:47:03] into 6. So, the ranges that we get are shown here.
[00:47:08] And, uh, now, when we do equi-width binning,
[00:47:13] Then, uh, we have these ranges.
[00:47:16] Right? You have these ranges with us.
[00:47:19] And when we try to accommodate the points in these ranges, then this is the distribution that you are seeing here.
[00:47:26] So, uh, this graph shows the distribution.
[00:47:31] So, before I proceed, I could see a hand raised. Uh, is there any question?
[00:47:39] Okay, G2 do, what do you want to know?
[00:47:40] Oh, yeah.
[00:47:41] Uh, about that binning already, uh, like, uh, when the data distribution is large, we are binning the…
[00:47:49] data to small sets, and we'll be analyzing only that particular range of values, right?
[00:47:55] For that, only we are using Building, right?
[00:47:56] Yes, so we are, uh, as of now,
[00:48:01] Binning is used in a lot of…
[00:48:03] things. As of now, I've talked about denoising the data.
[00:48:08] By ignoring the individual values, and only retaining the VIN IDs.
[00:48:13] And these bill IDs will be used for further analysis, okay?
[00:48:21] Okay, so now I'm just showing you another example of EQ with winning.
[00:48:27] So, any comments on this kind of winning?
[00:48:30] On equivalent winning on this particular data.
[00:48:35] Uh, yes, so… Unlike the earlier example here, it seems all the points are constituted on bin 4 and 5.
[00:48:46] And… so here, binning is not a good representation of the.
[00:48:52] Uh, screen time. In this case. So if we replace… The actual screen times with these beans.
[00:49:01] It would appear all of them are only dependent on 4 and 5, which kind of makes the other bins.
[00:49:07] Yeah, so what you're saying is correct, Pallavi?
[00:49:08] Uh, not useful.
[00:49:11] So, if we are using EQI with binning,
[00:49:14] In this particular case, then what are we having? We are having most of the points are crowded,
[00:49:20] In these two VINs, whereas the other VINs are practically empty.
[00:49:26] So remember why we want to form the bins. We want to form the bins so that we uniformly distribute the
[00:49:32] So, we should not have bins just for the sake of having them.
[00:49:36] Here, we have formed bins like bin 0, 1, 2, 3, and 4.
[00:49:40] And these VINs are kind of just formally there.
[00:49:43] But they don't really have much points in them.
[00:49:46] 1-1 point, or there might be a… it is… it is also possible that there might be a VIN having no points at all.
[00:49:54] in it. And therefore, equie with winning in this particular case, where the data distribution is such,
[00:50:01] is not very useful.
[00:50:03] Because some bins are practically empty, whereas others are crowded.
[00:50:07] So, this defies the…
[00:50:11] logic-affirming bins.
[00:50:12] So that's why, in this case, equi with winning is not a good solution.
[00:50:20] Okay, Deepak, what question do you have?
[00:50:23] Yeah, I think on the similar line, uh… That… uh, so basically, the purpose of billing is to categorize data or make it more simplistic, so that we can do the next.
[00:50:35] set of process, like clustering or whatever it is. But in this scenario, uh, probably the data distribution is.
[00:50:44] Like, there is huge difference, where, you know, we see few points on one line.
[00:50:47] And 4 and 5 years used. So we would ignore this step. Am I right on saying this?
[00:50:53] Uh, I couldn't understand what step you want to ignore here.
[00:50:57] Can you…?
[00:50:58] Uh, the billing… the process of billing. Or categorizing the data with the help of bidding.
[00:51:01] In this data set, you are saying?
[00:51:04] Okay, so the… we want to bin the data.
[00:51:05] Yes, yes.
[00:51:08] Right? We have to form the bins.
[00:51:11] But when we are using equivalent winning, definitely the meaning of, uh…
[00:51:16] Binning is getting defied because, uh, there are practically these two wins that are having points.
[00:51:22] And most of the points are coming in these two bins, and the rest of the bins are
[00:51:27] Uh, you know, just practically empty.
[00:51:30] You can also take another example. Suppose, uh, we want to bin.
[00:51:35] The students in a… in a particular course. Let's not talk about PG certification, because there are a variety of
[00:51:43] people from different backgrounds, but let's say we have students in
[00:51:47] We take fourth year, and we want to win them.
[00:51:50] On the basis of age.
[00:51:54] And we use equie with binning.
[00:51:56] So, if you are using EQ with binning on the students of BTEC 4th year,
[00:52:02] Uh, on the basis of age, then what would happen?
[00:52:06] we may form…
[00:52:08] N number of bins, but mostly all of them, because they'll have similar ages.
[00:52:14] All the points will, uh, you know, form part of a single VIN only.
[00:52:18] mostly single VIN, not even two.
[00:52:21] So, therefore, the purpose of forming bins is lost, because
[00:52:25] VINs are there, but they are not having any points, so…
[00:52:28] This is a useless scenario.
[00:52:31] So, but we have to perform binning, uh, so, uh…
[00:52:35] So, we have to look at some other way of forming bins, and equal with binning is not the solution.
[00:52:40] That is code here.
[00:52:43] So then, uh…
[00:52:46] So, we already discussed, uh…
[00:52:48] the impact of equivirth winning.
[00:52:52] Uh, in certain datasets, in which case, the…
[00:52:58] points are in, uh, are given in such a way that all of them form part of one single VIN, or
[00:53:05] some humans, and the rest of the bins are empty.
[00:53:10] So then, uh, since equally with binning is not suitable in many cases, we have to go for another
[00:53:18] Some other form of winning, and what we have is equally depth binning.
[00:53:21] If we have the winning is also called equip frequency binning, where
[00:53:25] We do not, uh, divide
[00:53:28] Uh, the given range of values into equal…
[00:53:32] size to bins. Rather,
[00:53:35] we divide the given ranges in such a way that each range of values
[00:53:42] Forms one bin, and…
[00:53:44] Approximately, the total number of samples
[00:53:48] per VIN are equal.
[00:53:50] So, here we are talking about
[00:53:52] frequency, uh, of the points in the VINs.
[00:53:57] And we are not talking about the width of the bin, so we don't look at the range of values.
[00:54:02] Rather, we look at the total number of points in the VIN.
[00:54:16] Sorry, I got muted here.
[00:54:18] So, uh, in this particular case, we are interested to form the VINs but equal
[00:54:25] Uh, worth winning is not going to serve the purpose. So, we have equal depth or equal frequency binning.
[00:54:32] In which case, uh, the number of samples
[00:54:35] per VIN are going to be approximately equal, and when such a number, uh…
[00:54:41] is achieved in any VIN.
[00:54:43] There, we cut the VIN boundary.
[00:54:47] So, we don't have regular VIN boundaries of equal width, but they are based on frequency.
[00:54:56] So, for example, uh…
[00:54:58] Let's say that we have certain, uh, the same set of points that we discussed earlier.
[00:55:05] And we have the screen time, then we have the sleep duration, and the points are actually sorted on sleep duration.
[00:55:12] And then we have the VIN IDs that were assigned to them based on equivalent binning, and then we have equally frequency binning.
[00:55:20] So, in case of frequency binning, we are going to associate
[00:55:26] Uh, we first sort the data and associate equal number of points per VIN.
[00:55:32] In this particular case, we have 52 points, and we wanted to form 5 bins, so approximately 10 points are accommodated.
[00:55:39] carbon. So, uh… and accordingly, when 10 points are accommodated, then the VIN boundary is
[00:55:47] found out. So once there are 10 points, then we look at the… what are the range of values of, uh,
[00:55:53] of the points,
[00:55:56] For that particular attribute, because they are all sorted.
[00:56:00] And use that build boundary as…
[00:56:02] The range of values for that particular VIN. Similarly,
[00:56:06] The next pin, next spin, like that.
[00:56:08] So here, we are focusing on the frequency of points, so you can see…
[00:56:13] That, more or less, all the VINs are uniformly packed.
[00:56:16] Uh, for example, in this VIN, you're not able to see many points, because they are lying one on top of the other.
[00:56:23] So, like this. So, these are the winds formed. All of them are equally populated.
[00:56:28] And there are no VINs that are empty, or no VINs that are…
[00:56:33] sparse or overly crowded.
[00:56:36] They are all equal.
[00:56:37] So, equidepth binning is used when we have the likelihood of having
[00:56:42] All points, uh, falling in the same range.
[00:56:45] For that, uh, for some particular attribute that is used for binning.
[00:56:50] If that is the case, then don't use equal with billing, use equil depth winning or some other form of winning.
[00:57:01] So, uh, I think this is the example. I already showed it to you.
[00:57:09] Now, uh, there are other methods also to form, uh, VINs.
[00:57:14] For example, uh…
[00:57:16] We partition the given data, we have the…
[00:57:20] price of some item in INR. So, first step would be to sort them in
[00:57:24] some order, let's say in the increasing order.
[00:57:28] And then we partition into equidepth, uh, bins. So, uh, when we, uh, talk about
[00:57:35] Equidact means equal number of points are going to be put per bin.
[00:57:39] So here, we are going to put 4-4 points per VIN, because the total number of points are…
[00:57:44] 12, so first 4 points, then next 4 points, like that.
[00:57:48] We will keep on doing, just a second.
[00:57:52] Let me choose mine.
[00:58:02] I'm just choosing my pen.
[00:58:10] So, this is about equal depth binning, using these points. Then we have other methods, like
[00:58:15] Uh, smoothing by VIN means. So, first, we actually partition the given data
[00:58:21] into equi-depth bins, and then we denoise the bins so that
[00:58:26] We are doing smoothing on the points by VIN means. So, we look at all these points, find out the bin mean,
[00:58:33] Uh, let's say 9 is the VIN mean, then we replace all the points by 9.
[00:58:37] So, it looks something like this.
[00:58:39] Let's say these are the points.
[00:58:42] You know, these are the values that we are having.
[00:58:45] So, what do we do? We have some number of bins that we want to form. Let's say 3 number of bins are to be formed.
[00:58:51] So these are the three wins.
[00:58:53] And we replace each bin by the mean.
[00:58:57] So, what will we have? These are the points. So, this is the VIN mean.
[00:59:00] This may be the VIN mean for the next VIN.
[00:59:04] This has been one.
[00:59:06] This has been 2, this is bin 3.
[00:59:09] And then this is also replaced by the…
[00:59:12] Meanwhile. So now, these are the wins that we are having here.
[00:59:18] Right? So, like this, we do smoothing by VIN means, and we, uh, lose the jitters in the data.
[00:59:25] Uh… here.
[00:59:27] So now, uh, the next thing, uh, is, uh, winning by VIN boundaries.
[00:59:32] So, we can do a binning on VIN boundaries.
[00:59:36] So, whatever the value, uh,
[00:59:40] Let's say the VIN boundary for this VIN is either 4 or 15.
[00:59:44] So, whatever the other values are,
[00:59:46] Uh, if 8 is more closer to 4 than to 15, 8 will be replaced by 4.
[00:59:53] Similarly, 9 if it is…
[00:59:56] Now, nearer to 4 than to 15, then it will be replaced by 4.
[01:00:01] Because we can see that 9 minus 4 is 5, and 15 minus 9 is 6.
[01:00:05] So 9 is closer to 4 than to 5, so we replace it by 4.
[01:00:10] Like this, we do smoothing using VIN boundaries. So, these are some different ways. There are so many other ways also possible.
[01:00:17] I'm just showing some of them.
[01:00:20] Then, the other method to do noise removal is to perform clustering on the given data.
[01:00:26] When we cluster the given data, then we see
[01:00:30] that the point, that is a noise point, will not be part of any cluster, and it will be excluded.
[01:00:36] And we can remove this as noise.
[01:00:41] Then, using regression for removing noise, so we have these points, which are shown here.
[01:00:46] We fit our model, or a curved line,
[01:00:48] on it. It could be linear, non-linear. Here it is a linear…
[01:00:53] line. And, uh, now this…
[01:00:56] 9 is retained, which is given by Y is equal to X plus 1.
[01:01:01] And all those points which are not lying on this line, are actually ignored or removed as noise points.
[01:01:11] So, uh, in regression, we try to minimize the distance
[01:01:18] So, here, we want to minimize the distance between the model that we have found
[01:01:22] And the points. So we want to minimize these.
[01:01:25] this distance is some of such distances.
[01:01:30] Uh, this is the idea to fit our line, uh, on most of the points.
[01:01:36] Okay.
[01:01:38] So, here's a diagrammatic representation. Let's say we have points which are shown there.
[01:01:44] And these points are actually shown in this purple graph.
[01:01:48] Then, uh, we could use different form of…
[01:01:51] regression, like logarithmic, exponential, and so on and so forth, and accordingly, different lines will be there.
[01:02:00] which represent the, uh, regression models that have been fitted on these points.
[01:02:05] And the model is written, and the points are ignored from further analysis.
[01:02:11] Okay. So that was about, uh…
[01:02:14] denoising of the data, any questions, anyone?
[01:02:18] About denoising.
[01:02:24] While learning Python, we understood the. DF command to, uh, use the value from the previous, so the next record.
[01:02:35] So, when that will be used in general?
[01:02:39] Uh, are you talking about when… what Python command to use, or your… I couldn't understand what you wanted to ask.
[01:02:49] Okay, for example, uh, we have on record for which one field is missing.
[01:02:54] So, we can use the value from the previous record from the same field.
[01:03:00] Or we can use the value from the next record to fill that value, like that.
[01:03:04] Okay.
[01:03:06] So, command, so any, any application we have here about…
[01:03:10] Uh, so let's say you have, uh…
[01:03:14] You know, uh, let's say you have some data here.
[01:03:17] And these are the different values, and this is attribute A1.
[01:03:22] And this is the missing value, and you are having value 1 here, and maybe value 2 here.
[01:03:28] So what you are saying is to fill this value, you could use this value.
[01:03:31] Or you could use this value, something like that.
[01:03:35] So, yeah. So, um…
[01:03:36] Yes, yes.
[01:03:37] This, uh, can… such kind of feeling of missing values can only be used
[01:03:43] When you have dependency in the attribute, let's say, uh, I have an attribute,
[01:03:49] Uh, which represents the sales.
[01:03:52] And that sales is for different countries.
[01:03:56] Right? So, let's say, uh…
[01:04:01] I don't know why it's not coming. So, I have got 3 countries.
[01:04:06] Let's say one is India.
[01:04:08] One is USA.
[01:04:10] And one is UK.
[01:04:13] And I have a missing value in USA.
[01:04:15] Whereas there are values for India and UK.
[01:04:18] Now, in the scenario that you mentioned,
[01:04:21] Uh, would it be logical to use the sales of, uh, of, uh…
[01:04:26] over India for that product to be… to fill
[01:04:29] The sales in USA, or the sales of UK to fill.
[01:04:33] Sales in USA.
[01:04:35] Uh, will it be logical?
[01:04:37] Let's say we talk.
[01:04:39] Yeah, yeah. So…
[01:04:40] No, certainly not, yes.
[01:04:43] Because sales in different countries are, uh, uncorrelated,
[01:04:47] They don't have any relationship, even if the product is same.
[01:04:50] Even if we talk about sales of, let's say,
[01:04:53] Any food item, say chips. So, the sales in India versus sales in UK or USA will be different.
[01:05:01] Even if the brand is the same, the chips type is the same, and all that stuff.
[01:05:07] So, in not all cases, you can use, but you can use this kind of a…
[01:05:12] thing, when I'm having a time series with me.
[01:05:15] Let's say, here I'm having sales.
[01:05:18] For a product.
[01:05:20] I don't know why I'm facing problems with my digital pen, sorry for that.
[01:05:25] So these are sales of one particular product, and this is time, T1.
[01:05:29] time T2, uh…
[01:05:32] time T3, and so on and so forth. This is time 0.
[01:05:38] So, uh, if we talk about sales of a product over different instances of time, and we have a missing value here.
[01:05:44] at time T2. Then we can definitely use
[01:05:48] Value from T1 or T3 to decide what needs to be filled here.
[01:05:53] In that case, you can do this, because the data is having a dependency.
[01:05:57] Uh, in sales over time.
[01:05:59] So, if such kind of dependencies exist, then you can use this method. Otherwise, you can't use it.
[01:06:06] Okay?
[01:06:08] Yep. So let's talk about data transformation.
[01:06:13] In data transformation, we apply techniques
[01:06:15] I'm sorry to interrupt. Just in case we don't have data, uh, like, time dependency.
[01:06:22] Mm-hmm.
[01:06:23] So, what we'll do in that scenario?
[01:06:26] Yeah, so if you don't have time dependency, then I already discussed so many cases, how to fill missing values.
[01:06:33] For example, remove the…
[01:06:35] Record containing missing values, fill with a default value, fill with a mean, median, mode.
[01:06:41] Okay.
[01:06:42] filled with mean per attribute, I mean, per class, and all that. That is for uncorrelated data only.
[01:06:48] Okay? Yeah.
[01:06:49] Okay, got it.
[01:06:52] Uh, okay. So, uh, let's talk about data transformation. Here, we try to…
[01:06:58] transform or, uh, change… not exactly change the data distribution, but
[01:07:04] modify the data in such a way.
[01:07:07] Uh, that it is more, uh…
[01:07:10] more applicable, or it is better suited for analysis.
[01:07:14] So, in this, there are many techniques, like data aggregation, normalization of the data,
[01:07:20] Attribute construction, and so on and so forth.
[01:07:23] And we'll talk about each of these as we go along.
[01:07:27] So, first, um, transformation on the data is data normalization.
[01:07:32] So, data normalization means…
[01:07:36] that we are actually, uh, mapping
[01:07:39] One range of the data to some other range.
[01:07:42] For example, initially, our values were, uh, belonging to
[01:07:47] Uh, the initial value was… min value A.
[01:07:50] And max value, A, for the attribute A.
[01:07:54] No, I want to map it to…
[01:07:57] Newman A.
[01:08:00] And new max A.
[01:08:04] I want to do this. So, how can I do it? I can do it with min-max normalizing.
[01:08:09] Where any value mu prime is given as mu minus min A,
[01:08:14] divided by max minus min into…
[01:08:18] into new max, A minus new min A plus new max, so…
[01:08:22] Actually, here, what are we doing? Mapping one range of values to the other. For example, if I've got sales,
[01:08:30] In INR.
[01:08:32] And I've got sales in USD.
[01:08:35] And I want to map them both to the same units.
[01:08:39] Then, either I want to map INR to UST or UST to INR, then I can use min-max normalization.
[01:08:46] To bring the values to the same range.
[01:08:49] Where many Min, and max say designate the minimum and maximum values.
[01:08:55] For the attribute A,
[01:08:57] uh… initially,
[01:09:00] And then Newman A and new max A.
[01:09:02] are the values that are the new values.
[01:09:05] Uh, that are to be used.
[01:09:08] Uh, where we want to map them.
[01:09:16] So, what is he, ma'am?
[01:09:18] Uh, sorry, the new, you're saying?
[01:09:20] No, what is V?
[01:09:23] So, that is any value.
[01:09:24] Okay.
[01:09:25] So, the new value is given by nu prime, or V prime, whatever. This is the initial value.
[01:09:32] So, we'll subtract the minimum value, then divide it by the original range,
[01:09:37] Multiplied by the new range, and add the minimum value of the new range. So, in a way, what are we doing?
[01:09:43] Where actually, uh, first removing the offset
[01:09:46] of the original range.
[01:09:48] From the given value, and then…
[01:09:51] Mapping it to a new range?
[01:09:53] And offsetting it with a new minimum.
[01:09:55] This is what we are doing in min-max novelization.
[01:09:59] Then we also have z-score normalization.
[01:10:01] Where we have, uh, the value.
[01:10:04] V, given us V', which is…
[01:10:08] Uh, the value minus mean over the attribute, say, A, divided by standard deviation.
[01:10:13] So, this is also another method.
[01:10:15] Many times, marks are normalized using z-score normalization.
[01:10:20] And, uh, this is not the entire list. We have many other methods also.
[01:10:25] However, I'm just including some important ones.
[01:10:31] So then, uh…
[01:10:32] Ma'am, can we have one example for normalization?
[01:10:37] Yeah, I think I… I do have some graphs. I'll show you.
[01:10:42] Just, uh, if you'll wait for a little while, I'll be able to show you examples.
[01:10:44] Okay.
[01:10:46] on graphs, you know.
[01:10:49] So then, uh, the application of minima access to convert one range to a new range.
[01:10:57] And the application of z-score normalization is to compare
[01:11:01] Uh, two scores.
[01:11:04] Uh, that are coming from different normal distributions.
[01:11:07] And, uh, we use z-score normalization in that case.
[01:11:12] So, uh…
[01:11:14] So, here's an example. Let's say we have the data as shown here, where we have X and Y as the two variables.
[01:11:21] And we want to do min-max normalization as well as z-score normalization.
[01:11:26] They already mentioned that G-score normalization is going to
[01:11:30] Center the data, uh, around the mean.
[01:11:34] or around zero. So, we are subtracting the…
[01:11:37] original mean value from it.
[01:11:39] So that what we are having is only the fluctuations around the mean.
[01:11:44] So, uh, and then we scale it on the basis of the standard deviation.
[01:11:50] Uh, this is how we get z-score normalization. So, let's, uh, see.
[01:11:55] Uh, this is the data, original data.
[01:11:58] With minimum and maximum values given as 1 and 20,
[01:12:02] For X and Y, and for Y… for X, and accordingly for Y also.
[01:12:08] Just a second, I don't know what's happening.
[01:12:20] Now, uh, we, uh, perform min-max normalization.
[01:12:24] And, uh…
[01:12:35] So, the initial range of values is 1 to 20 and 1 to 25.
[01:12:39] And we want to map these values to a new range.
[01:12:43] Uh, which is from 5 to 25 for the X coordinates, and 10 to 40 for the Y coordinates.
[01:12:50] So, the minimum was 1, we want to map it to 5, this one.
[01:12:55] And maximum is 20, we want to map it to…
[01:12:59] I don't know what's going wrong with my…
[01:13:02] spend today.
[01:13:07] Yeah, and we want to map it to 25, so 20, we want to map to 25.
[01:13:11] Similarly, Y is 1, we want to map it to 10, 25, we want to map to 40.
[01:13:17] This is how we want to perform the…
[01:13:19] mapping. So, you can see here, this is the original data shown in the yellow color.
[01:13:24] And it's starting with 1, and ending with 20.
[01:13:28] For the x-axis, and for the y-axis, it starts with 1.
[01:13:33] And ends with 25 here.
[01:13:36] This is 25.
[01:13:39] Now, we want the minimum value to be 5, so we are going to perform min-max normalization.
[01:13:44] And this point is actually going to be mapped to…
[01:13:47] The new minimum, which is 5.
[01:13:50] So this is actually mapped to 5.
[01:13:52] 5 on the x-axis, and 10 on the y-axis. So, this is the 10 value.
[01:13:58] This is the phi value, this is the point.
[01:14:00] So we start from here, and we end the data with
[01:14:05] Uh, the new minimum value, uh, the maximum value is 40.
[01:14:11] So, it was 25 initially.
[01:14:12] Now, this has got mapped to 40 out here.
[01:14:15] So this is the new, uh, mapped values we get.
[01:14:19] Uh, from the given data. The thing to notice is that initially the range was 25 minus 1.
[01:14:26] And now the range is 40 minus 10.
[01:14:29] Okay, so this is how we perform min-max normalization.
[01:14:33] If we have to perform z-score normalization on the same data, then
[01:14:37] Oh, sorry, ma'am, to, uh, disturb you. Can you please go back to the previous slide?
[01:14:42] Mm-hmm.
[01:14:43] Can we see how this new mean X is getting calculated?
[01:14:48] Yeah, so you just have to use this formula.
[01:14:50] So you have, uh…
[01:14:53] the value minus the minimum value divided by the original range,
[01:14:57] into new range plus the new minimum.
[01:15:00] For example, you have a point.
[01:15:03] Uh, the initial point is X is equal to 1, Y is equal to 1.
[01:15:07] Right? So, the new value will be…
[01:15:11] The old value, you can see this formula, which is actually…
[01:15:15] Uh, the minimum value. So, this new value will be…
[01:15:19] Uh, if we talk about one… the X value, that will be one
[01:15:24] Minus. So here, it is a…
[01:15:28] given value, let's say I… you can just use this formula. So, let's say I talk about the point 2 and 3.
[01:15:34] Okay? Then it will be the original value.
[01:15:38] Minus the minimum value.
[01:15:40] Which will be, uh, for this particular case, the original minimum value is 1, so we'll subtract 1.
[01:15:48] divided by…
[01:15:51] Just a second now, facing some issues with this.
[01:15:54] divided by the original range.
[01:15:57] What is the original range?
[01:15:59] For X, it is 20 minus 1, so it is going to be 19.
[01:16:06] I'm actually trying to write a 9 here.
[01:16:10] Yeah. This multiplied by the new range,
[01:16:13] The new range is going from…
[01:16:16] We are having the…
[01:16:20] the new range going from…
[01:16:23] 5 to 25, so…
[01:16:25] This will be the new range, which will be…
[01:16:28] 25 minus 5.
[01:16:30] plus the minimum value of the new range. So what we'll have here, 1 upon 19,
[01:16:35] In 220,
[01:16:37] plus 5, so it will be 5…
[01:16:39] Five point, I mean, 5 plus 1 point something.
[01:16:42] So, it will be around 6.
[01:16:44] That is the point 2, 3.
[01:16:47] So, this is the point 2, 3.
[01:16:50] in the original data, here.
[01:16:52] This is a point in correlational data.
[01:16:53] But how should we decide the new Minix and new meanwhile?
[01:16:59] Sorry, how do you decide what?
[01:17:01] How should we decide new minimum X and. new maximum X.
[01:17:04] So that is the range that you want to map the data into, right?
[01:17:09] Suppose, uh, I have, uh, you know, the value…
[01:17:13] Let's say I've got values for different variables, like, for example, I've got height.
[01:17:18] Then I've got weight.
[01:17:20] And maybe I've got age.
[01:17:22] All of them will have different ranges.
[01:17:24] Right? And, uh, I want to compare them.
[01:17:28] So then, in that case, it is possible that I may want to bring them to some similar ranges, so that
[01:17:34] Uh, because age may range, let's say, from 1 to 100.
[01:17:39] And weight may range, again,
[01:17:42] from, let's say, 10-200, or maybe 200.
[01:17:46] And height may range from…
[01:17:48] in feet from something, let's say, 2 feet to 7 feet, or 6.5.
[01:17:55] So, what can happen is with these different ranges, some particular attribute
[01:18:00] will become dominant. For example, weight or age.
[01:18:04] So, to prevent that, we might want them to…
[01:18:06] scale them to the same range. So that new range will be decided by you only.
[01:18:13] Right? Or the user.
[01:18:15] Whosoever wants to use, whatever, but what I'm showing you is how to do…
[01:18:20] How to calculate when you have the new range with you.
[01:18:25] Okay. So, I request all of you to please, uh…
[01:18:26] Yeah, got it, ma'am. Thank you.
[01:18:29] right on the chat, and I will, uh, so that it doesn't disrupt the flow of explanation, and I'll take questions in between.
[01:18:36] I'll stop and take in case the chat is disabled, uh…
[01:18:41] A request of insurance co-host to please enable.
[01:18:43] the chart, so that people can write.
[01:18:46] Uh, on the chat for any doubts.
[01:18:51] Okay, so this is z-score normalization. Here, what we are doing is…
[01:18:54] That we are bringing, uh, all the data into the range of 0 to
[01:18:59] one, something like that. So…
[01:19:01] So, you can see the data was spread over this much range, now it has…
[01:19:06] being brought to this range.
[01:19:08] So, this is how z-score normalization works.
[01:19:12] Then, uh, the next thing that we will talk about is dimensionality reduction. It's a very, very important topic.
[01:19:19] And there are many methods to do it. One important way is doing… using feature selection.
[01:19:25] Another is to use feature construction.
[01:19:28] So, what is feature selection?
[01:19:30] Feature selection means we have certain attributes that are given in the data.
[01:19:35] And we select some important attributes.
[01:19:38] And you retain those, and just ignore the rest of them.
[01:19:42] So, we are actually finding out, uh, important subset.
[01:19:47] Uh, of attributes from the given set.
[01:19:50] And using that only to represent the data, and ignoring the others.
[01:19:55] This is feature selection.
[01:19:57] So, we are doing a select operation, or we are doing a selection.
[01:20:02] from the entire set.
[01:20:04] Then, uh, comes feature construction or feature extraction.
[01:20:08] So, in feature construction, what we do is that we…
[01:20:12] Actually, form new features out of the given data.
[01:20:16] So, how do we form new features that can be different methods? For example, uh, we might map the given points
[01:20:24] into, uh, you know, some other…
[01:20:27] coordinate plane completely.
[01:20:29] So, we may form new features. For example, if we have text data,
[01:20:33] then we can represent this text.
[01:20:36] Uh, in some way. For example, by sets of keywords or something like that. So…
[01:20:41] The original data was the text document.
[01:20:44] And it got mapped to new features.
[01:20:47] By some method. Okay, so feature construction and extraction,
[01:20:51] means deriving completely new set of features from the Q1 data.
[01:20:55] Now, what is dimensionality reduction?
[01:20:58] Dimensionality reduction means we are reducing the number of attributes or the dimensions in the data.
[01:21:06] So that the data becomes more easy to handle, to…
[01:21:09] archive to access and all that.
[01:21:12] So, when we are doing feature selection or feature construction,
[01:21:17] We may be also doing dimensionality reduction. For example, if I've got some
[01:21:22] 30 attributes with me. Out of these 30, I'm retaining only 5.
[01:21:26] Uh, and doing a feature selection in which 5 gets selected and rest of the…
[01:21:31] Uh, 25 are not so important, and they are ignored.
[01:21:36] Then I'm reducing the dimensions from 30 to 5.
[01:21:39] This is what is dimension reduction, and it's very important.
[01:21:43] Because, uh, higher the dimensionality, it's more difficult to handle the data.
[01:21:49] And it's also very difficult to visualize such high-dimensional data.
[01:21:53] Um, and so on and so forth.
[01:21:59] So, here's an example of feature selection.
[01:22:01] Uh, we are assuming that, uh, we are building a decision tree.
[01:22:06] How we'll actually build our decision tree, I will be showing you.
[01:22:10] In the chapter on supervised learning. However, I've just, for illustration,
[01:22:15] shown one here. So, the initial attribute sets are A1, A2, A3, A4, A5, A6. So, there are 6 attributes.
[01:22:23] given to us, and we want to do a feature selection.
[01:22:26] There can be many methods to perform feature selection.
[01:22:30] And one of them is this, where we try to build a decision tree.
[01:22:34] And all the attributes that are represented
[01:22:38] On the tree, or are utilized in building the tree, are retained
[01:22:42] And the rest of them are ignored.
[01:22:45] Because when we build a decision tree, we…
[01:22:49] actually use this to classify the data.
[01:22:50] And if three attributes are enough to form… to arrive at the class labels,
[01:22:55] Then the others are just redundant, and we can ignore.
[01:22:59] So, this means that we can ignore A2.
[01:23:02] Then we can ignore A3.
[01:23:04] then we can ignore A5. So, we are left with…
[01:23:08] A1, then A3… A4, and A6. So, we are left
[01:23:15] Uh, with 50% of the attributes, and we have…
[01:23:17] done attribute selection out here.
[01:23:21] Uh, while the sanctity of the data is still getting maintained because the rest of the attributes are kind of redundant for us.
[01:23:28] So we ignored them. There are many other methods also we'll be discussing them as we go along.
[01:23:37] Then, uh, uh…
[01:23:40] Then there are methods for data reduction.
[01:23:43] Uh, so data reduction is another thing that is required if the data is very huge.
[01:23:48] We somehow want to reduce it so that it becomes…
[01:23:52] Uh, smaller, and it becomes easier to handle.
[01:23:55] a huge amount of data could also be compute-intensive.
[01:23:59] And if there can be some methods by which we can reduce the data, it will definitely help.
[01:24:04] And one such method is numerosity reduction.
[01:24:08] So, numerosity reduction. So, here we… numerosity means numbers. If we can somehow reduce
[01:24:14] The numbers in the data, then we are talking about numerosity reduction.
[01:24:20] So, how can we do that?
[01:24:22] There are two methods. One is parametric methods, the others are non-parametric methods.
[01:24:28] What are parametric methods?
[01:24:31] Here, what we do is we take the entire data, that may be some…
[01:24:34] Uh, let's say we have got some 1 lakh records with us.
[01:24:37] And, uh, these are a huge number, and therefore, we…
[01:24:42] somehow want to reduce the data, so what we do is we fit a model on that data,
[01:24:46] And after fitting the model, we just, uh… and to fit the model, we do, uh…
[01:24:51] parameter estimation that may be hyperparameters or other parameters that may require to model the data.
[01:24:58] We find out those parameters, model the data,
[01:25:01] And retain the model.
[01:25:03] And discard.
[01:25:05] Uh, the individual, the actual points.
[01:25:09] So, we don't really keep the points, we just keep the model with us, because that's a good representation.
[01:25:15] Assuming that's a good representation of the data,
[01:25:18] And then just, uh…
[01:25:21] user model and forget about the points. So, these are parametric methods, because here we have to assess the parameters
[01:25:30] that are involved, uh, in making up a model.
[01:25:35] Then we also have, uh…
[01:25:38] Non-parametric methods. So, non-parametric methods are those where
[01:25:42] We do not build any models and do not need to estimate any parameters of the models.
[01:25:48] For example, clustering.
[01:25:50] So, in clustering, we just form the clusters or groups, out of the given data.
[01:25:55] So there are no parameters or hyperparameters that need to be estimated.
[01:25:59] In such methods, all we need to do is just cluster the data,
[01:26:04] Or if you are doing some histograms out of that data, we just form the histograms, but we are not…
[01:26:09] Using any other hyperparameters or any other parameter estimation here.
[01:26:14] Therefore, such methods are called
[01:26:17] non-parametric methods, and
[01:26:19] All these methods help to reduce the volume of the data.
[01:26:22] For example, in the parametric methods, the volume gets represented by our single model, or a set of models.
[01:26:29] In the non-parametric methods, the entire set of or volume of data gets
[01:26:35] represented by some few clusters.
[01:26:37] So, in both these cases, we are actually reducing the volume of the data.
[01:26:43] So, we can, uh, also perform numerosity reduction, as I mentioned here.
[01:26:49] Not just by clustering, but also forming histograms. Now, what are histograms?
[01:26:54] Histograms are nothing but, again, they are similar to bins only.
[01:26:59] And, uh, the method to form histogram is again similar to the methods for binning, which I've already told you.
[01:27:04] That we could do binning by equivalent binning, or by equip frequency binning, and so on and so forth.
[01:27:10] And similar methods can be used.
[01:27:14] to form histograms. Often histograms are done on counts, or frequency-based histograms are very, very popular.
[01:27:21] methods of data reduction, where we retain the histogram, and we just do away with the actual
[01:27:29] volume of the data.
[01:27:34] So this is a histogram that I had shown you already, where I'm…
[01:27:39] I've done the histogram based on the
[01:27:41] First alphabet of, uh, the… of your first name.
[01:27:45] Uh, for all this group of students, and we find out that
[01:27:49] This program for the alphabet S is having the maximum count in it, or frequency of values, followed by a
[01:27:56] RM, and so on and so forth. We…
[01:28:02] Okay, so before I, uh, proceed, any questions on what I just discussed?
[01:28:08] Maybe you can write on the chat.
[01:28:20] Okay, so what I'll do is I'll proceed and then invite questions once again.
[01:28:25] Um, as we progress.
[01:28:27] So now, another form of performing data reduction is by sampling.
[01:28:33] So, what is sampling? Sampling of the data is
[01:28:36] So, let's say we have got some…
[01:28:38] millions of points with us. It's really very difficult.
[01:28:43] to handle a huge amount of data. So, instead of having the entire set of millions of points,
[01:28:47] We could have a good representative subset of this data,
[01:28:52] And retain the subset rather than retaining all those millions of points.
[01:28:56] Now, how can the subset be formed?
[01:28:59] One way of forming the subset is doing simple random sampling.
[01:29:04] So what we do, out of the given set of points, whatever the number is, let's say 1 million or 1 lakh,
[01:29:09] We just randomly select some points.
[01:29:13] And, uh, retain it.
[01:29:15] As a representative subset of the given data.
[01:29:19] Uh, so the…
[01:29:23] Method looks very simple and nice. However, the primary problem with this simple random sampling is
[01:29:29] That, let's say if the data distribution is skewed.
[01:29:34] Let's say that in certain ranges of values are very less,
[01:29:39] In some other ranges, the values are very high in number.
[01:29:44] And when we do a random sampling, it is possible
[01:29:47] that some of the lower-range values get lost because they never were selected.
[01:29:52] in the random sample.
[01:29:54] Right, so then in that case, the distribution of the data will…
[01:29:58] get modified, and we don't want to modify the data distribution.
[01:30:03] What we want to do is get a representative subset of the given data.
[01:30:07] And reduce the volume of the data by doing sampling.
[01:30:11] So then, uh, random, uh, sampling, simple random sampling, is a good, uh, and, uh,
[01:30:18] a good method, which is intuitive,
[01:30:21] Uh, and understandable. However, it doesn't work when the…
[01:30:25] Data distribution is skewed.
[01:30:27] So, in such cases, we can use something called stratified sampling.
[01:30:33] So, stratified sampling means that, let's say, we have got
[01:30:37] Uh, two classes in the data.
[01:30:39] One class is majority class, where 80% of the samples belong to Class C1.
[01:30:44] Let's say we have 80% of the samples.
[01:30:49] Belonging to Class C1.
[01:30:52] And the rest of the 20% samples belong to class C2.
[01:30:56] Now, uh, when we are doing a stratified sample, assuming, let's say, we have got a total of 100 samples,
[01:31:03] Then, the ratio in which the samples are 80s to 20.
[01:31:08] Now, suppose I want to draw a stratified sampling.
[01:31:13] And I want to reduce the data to 50%.
[01:31:15] Let's say, now, I want only 50 samples. So, what I'll do…
[01:31:20] I will pick samples from each of this group in the ratio.
[01:31:24] of the size of that group.
[01:31:26] For example, the size of this group is 0.8, and this is 0.2.
[01:31:31] So, 40 per… uh, so 80% of the samples should come from this group, and 20% should come
[01:31:37] from this group, so that the distribution, overall distribution,
[01:31:40] of the class labels remains same.
[01:31:43] So this means if I want to make 50 samples, I will draw 40 from here.
[01:31:47] And 10 from here. So, in this case, the distribution remains the same because the
[01:31:54] Uh, the…
[01:31:56] Uh, the amount of Class C1 and C2 samples.
[01:32:00] are similar to the original distribution, which is 80 to 20%.
[01:32:06] So, uh, stratified sampling is used when we have skewed data distribution. Most of the real-world data is skewed in nature, as you all already know.
[01:32:18] Then another method of, uh, doing, uh, data reduction.
[01:32:23] is by, uh, doing, uh…
[01:32:26] So, if we have, uh, here, the raw data, we are again illustrating, sorry, this is not…
[01:32:32] So, clustering is also another method of doing, uh…
[01:32:36] sampling. So this is the raw data, where we have got 3
[01:32:40] Uh, classes or clusters available in the data.
[01:32:44] So, if we do a stratified sample,
[01:32:46] Then, uh…
[01:32:49] samples, uh, will be drawn from each of these classes in the same ratio as they are available in the original data.
[01:32:55] And therefore, you will see, let's say if this was class C1, this is C2,
[01:33:00] And this is C3. Then, in the final, uh, data also, after doing the random sampling,
[01:33:07] We still have all these classes in the data.
[01:33:10] And none of the classes will go missing.
[01:33:12] Uh, because in random sampling, we just randomly pick, but in stratified sampling,
[01:33:18] We pick the points in the ratio in which they are available in the original data.
[01:33:27] Okay, any questions, anyone, uh, on whatever we have discussed?
[01:33:35] Yeah.
[01:33:36] Oh, ma'am, I have a question here. So, like, we are doing this data reduction, right? So, what are the activities where data reduction.
[01:33:42] will come handy, because I was under the assumption the more the later, better the machine learning model.
[01:33:47] But, uh, if we are reducing the data, then we are reducing the number of samples on which.
[01:33:53] We are training the model as well, right, for decision tree that you mentioned, I understand that some of the attributes are not having any.
[01:34:01] details, or any… Implication on the machine learning model.
[01:34:06] But how this will help?
[01:34:08] Okay. Uh, definitely, we generally say that, uh, the more the data, the better will be the learning.
[01:34:17] That is true. Why? Because we want to have
[01:34:20] data samples that represent each and every use case.
[01:34:25] Uh, in the given data.
[01:34:27] Right? So that's what, uh, what I'm saying is, let's say that…
[01:34:31] We are having some, you know, the Aadhaar, uh, or let's not talk about Aadhar. We are having some data
[01:34:39] Say tweets, and there are millions of tweets from the given dataset.
[01:34:43] So it's, uh… so if we have to analyze the tweets to find out what the topic of discussion.
[01:34:49] It's going to take a huge amount of compute.
[01:34:53] Or probably the algorithm will keep running for days, and you won't be able to get your result in time.
[01:34:58] So, in all such cases,
[01:35:00] We can, uh, perform data reduction.
[01:35:03] In such a way that the overall distribution of the data in the classes that is there
[01:35:10] doesn't change. So, uh, here in this case, we are just making a smaller sample.
[01:35:17] To make the data more easily handy, uh, you know, easy for handling.
[01:35:23] So, if we have a sample of the data, then we can use it for different purposes, say clustering, classification,
[01:35:30] or any other purpose. So, this is where data reduction comes into.
[01:35:35] picture. It is not just this. Suppose we want to do exploratory data analysis, or EDA, on the data.
[01:35:43] So, it will be very difficult to do EDA, let's say, on millions of points. So, we may want to reduce the data and get a set which is a good representative of the original set.
[01:35:54] So, we can easily perform EDA on it. So, here also, data reduction is handy.
[01:35:58] So, there are many such cases where data reduction is important.
[01:36:04] And it is handy for us. So, in all such cases, we perform.
[01:36:07] Uh, data reduction by different methods. Sampling is one of them.
[01:36:12] Is that okay, Ankles?
[01:36:13] So, yeah, ma'am, so we are saying that the time to result is what we are prioritizing over accuracy.
[01:36:18] Uh, so that's why we are going with data reduction. Second thing is the exploratory data analysis would be easier to do.
[01:36:25] Uh, over the smaller data set. And that's why we are using it, right?
[01:36:31] Yeah, these are two reasons. There are many other reasons also, and uh…
[01:36:36] Uh, definitely compute time will be less if…
[01:36:39] Uh, the amount of data will be the same. However, the data distribution is still the same. The most important thing is
[01:36:48] that when we are drawing a sample of the
[01:36:50] given data, it should not happen that the sample's distribution is different from the original data distribution.
[01:36:57] So that we must not change. That is why we go for stratified sample, where
[01:37:01] Uh, all the classes are retained in the data that is there in the original data.
[01:37:07] Yeah.
[01:37:10] Uh, so I can see, uh…
[01:37:11] Understood.
[01:37:13] that, uh, Sivanch, you have some question, can you please ask what you want to know?
[01:37:18] Yes, sure. So, I was just thinking about the data sets of form and max normalization. Let's say if I have data, like, 10 to.
[01:37:29] From 2 to 10, uh, 20. Uh, and uh… Uh, let's say my training set is having this done as a max value and 2 as a main value.
[01:37:41] Uh, and the test value has just provided to be as, uh, 15.
[01:37:45] By the model. And… and when we are trying to calculate the… The min-max normalization value from that formula which you have showed.
[01:37:57] Yeah.
[01:37:58] It gives them a fractional number. So, uh, it's just acceptable, uh, for… to apply this approach, or it will not work in such type of data, certainly.
[01:38:06] Uh, you're talking about datasets where you get fractions and not numbers. Is that what you're saying?
[01:38:14] Yeah, with the formula, you get fractions.
[01:38:15] Yeah, with the formula, I mean, uh… The new value will… yes.
[01:38:18] Yeah, see, if you are having one range,
[01:38:20] Let's say your initial range is, uh, let's say 10 to…
[01:38:25] say, 50, and you want to map these values to a range like
[01:38:30] Say, 1, 2, say, 10.
[01:38:32] Let's say. Then, definitely,
[01:38:35] Something that was 10-year is coming to 1.
[01:38:38] Then, it will be a fraction only, no?
[01:38:40] So, having fractions is not a problem.
[01:38:43] Uh, as long as we are allowing any kind of values. Suppose we…
[01:38:49] don't want, uh, values in points.
[01:38:52] Uh, as per our algorithm, then it's different. Otherwise, it's okay to do that, because…
[01:38:58] We ourselves decide what is the new range. If the new range is compressed,
[01:39:05] Then some of the larger values which are there in the original range will come out in points only.
[01:39:10] So that's fine.
[01:39:14] Okay.
[01:39:15] Okay. And one question I have related to binning.
[01:39:16] Mm-hmm.
[01:39:17] Uh, what I understood is we are… today we have covered about numerical.
[01:39:23] techniques of those spinning equally width and equal frequency.
[01:39:29] Mm-hmm.
[01:39:30] So, in the… EQ frequency case of what happened if the dataset contains many repeated value?
[01:39:38] At the VIN boundaries.
[01:39:41] Uh, okay, so you're talking about equal frequency binning, right?
[01:39:45] And… and we are talking about values that are repeatedly coming.
[01:39:46] Great.
[01:39:50] So that's okay. Accordingly, see, in case of equip frequency binning,
[01:39:51] Hello.
[01:39:55] The bins may not be of equal width.
[01:39:57] Only the frequency remains wants… is to be remaining the same.
[01:39:58] Mm-hmm.
[01:40:02] So, let's say in the range of…
[01:40:04] Uh, let's say 1 to 10, we are having lots of values, say 999 coming again and again.
[01:40:10] Then, uh, what are we going to do?
[01:40:13] So, the frequency, or the count that we want to keep on the VIN,
[01:40:18] will depend on the average values we are expecting per bin.
[01:40:23] Let's say we already know, we have the data with us, and we want to bin it, and we know that
[01:40:28] a lot of 9 values are coming in here. So, accordingly, we'll bin the data.
[01:40:34] So that the rest of the bins will have also a similar
[01:40:38] uh… distribution. However, uh, we have to keep the count. I mean, if we are having 10 values,
[01:40:45] Uh, which are, uh, 10 numbers, which are, you know, having the same value. It has to be put in the same bin only.
[01:40:52] We cannot ignore them.
[01:40:55] They will account for the frequency.
[01:40:56] Um…
[01:40:59] Understood, yeah. Oh, one last question, uh… Just, uh, having thought is, can we consider the negative values in this binning method, uh, changes?
[01:41:10] Or it's for… Positive natural numbers.
[01:41:15] You're gonna have any kind of range.
[01:41:17] The range will accordingly work
[01:41:18] Okay.
[01:41:19] Because what we are doing, we are subtracting the maximum minus the minimum value.
[01:41:24] If the minimum is negative, it will get added to the range, so that's no problem.
[01:41:31] So, you could have values going from minus 1 to 10 also in the new range.
[01:41:32] Okay, thank you.
[01:41:35] So, because minus 1 to 10 means it is a set of 11 values, so it will be 10 minus minus 1, which will be 11.
[01:41:42] So that's okay.
[01:41:44] Okay, thank you.
[01:41:48] So then, uh, we talked about, so far, different methods for data reduction, like sampling, like numerosity reduction,
[01:41:57] Uh, using, um…
[01:41:59] parametric methods, non-parametric methods, and so on and so forth.
[01:42:04] So, another way to do data reduction is via data discretization. What is data discretization? Again, it's similar to binning only.
[01:42:12] I can explain it with the help of binarization. Let's say that
[01:42:16] Uh, you know, I have got, uh…
[01:42:18] I've got to represent the data in form of two numbers.
[01:42:22] Let's say I have a range of values from 0 to 10.
[01:42:25] But I want to represent this using binary values.
[01:42:29] So, I can do this in a couple of ways. One way could be
[01:42:33] Anything from 0 up to 5.
[01:42:37] could be 0, and anything, uh, greater than 5, up to
[01:42:42] 10 could be 1.
[01:42:44] So here, we are actually binarizing the data from 0 to 10 to 0 and 1, 2 values.
[01:42:50] So, we all know that computers work on binary data, and whatever is the actual
[01:42:55] value, it is represented in form of binary numbers. So, binarization is a simple example.
[01:43:01] Or a special case of discretization.
[01:43:03] However, discretization may not just be binarization. It could be anarization.
[01:43:09] Say, for example, I have the range 0 to 15,
[01:43:13] And I want to represent it in 3 sets, in 3 discrete value sets. So what I could do, I could have 0 up to 5,
[01:43:21] As one said, then greater than 5 up to
[01:43:26] 10. As another set.
[01:43:30] greater than… and then greater than 10, up to 15 in another set. This could be…
[01:43:36] Uh, that discrete value represented by 0, 1, and 2.
[01:43:41] So, here, binning is being used, or we also call it discretization.
[01:43:46] Where we are representing this, uh, infinite number of values in each of these range,
[01:43:52] by a finite value, or a discrete value.
[01:43:56] Which is, like, 0, 1, 2, something.
[01:43:58] And this can also be said to be a label that we are assigning to all the points in this range. So this is called discretization.
[01:44:06] of the given data, and we use it to convert a set of continuous valued, uh…
[01:44:12] attribute.
[01:44:15] to a categorical attribute. Let's say I've got, uh, say, high, uh, age.
[01:44:25] Right? So, let's say age ranges from 1 to 100.
[01:44:28] So, there could be infinite points in this range.
[01:44:31] But I want to represent a data by a finite set so that I can do some calculations over it.
[01:44:37] Having an infinite value set is very difficult to handle,
[01:44:41] So what I can do, I can discretize it into…
[01:44:44] 10, uh, bins, or 10 ranges,
[01:44:47] say, 1 to 10, then greater than 10, up to 20, greater than 20, up to 30, like that, so I'll have
[01:44:54] 10 bins with me, like, 1, 2, 3, till 10, or 0, up to 9.
[01:45:00] So, this is nothing but data discretization.
[01:45:03] Where we are dividing the data.
[01:45:06] Uh, in such a way that certain set of data falls in some bin, and we call it
[01:45:12] some… and assign it some discrete value.
[01:45:14] And we call this data discretization.
[01:45:18] Then there's something called concept hierarchy.
[01:45:22] So, what is a concept hierarchy? In case of a concept hierarchy, we are having…
[01:45:27] Uh, some implicit hierarchy in the data. For example, if we talk about location,
[01:45:33] So, location is an attribute.
[01:45:36] Now, the location could be over a country, it could be over state, it could be over city, it could be our street.
[01:45:41] And so on and so forth. So there's an inherent…
[01:45:44] hierarchy involved
[01:45:46] In the attribute-like location, and we can use this
[01:45:51] a hierarchy to represent, uh, this particular value called location.
[01:45:57] Suppose I want a very fine-grained representation. I can go up to the street level,
[01:46:03] If I have, uh…
[01:46:05] More courses are on a broader representation, I can go up to city,
[01:46:08] or up to state, up to country, and so on and so forth.
[01:46:11] So, we can have different levels in the hierarchy.
[01:46:15] And we can utilize those values from corresponding level in the hierarchy that we want.
[01:46:22] If you want lots of values, we go to the…
[01:46:25] Street level, if we got, uh, somewhat less number of values, we go to city level, still lesser state, still lesser country, like that.
[01:46:34] So, uh, this is how we utilize the concept hierarchy.
[01:46:40] So…
[01:46:42] Okay. Just a second…
[01:46:48] So I think there's some question that I could see here on…
[01:46:52] Um, okay.
[01:46:55] So, I think I could see a question that, uh…
[01:46:58] Uh, about parametric reduction.
[01:47:01] Okay, so Srinam, parametric reduction means…
[01:47:05] that we are using model.
[01:47:07] We are using a model to represent the data. Obviously, when we build a model, we will need to represent the model with the help of some parameters.
[01:47:17] So that is parametric representation of the data.
[01:47:19] Okay? Say a model will need some… for example, if I have a model like Y is equal to
[01:47:25] MX plus C.
[01:47:27] Suppose our model of this kind.
[01:47:30] Where this is a linear model.
[01:47:32] So then, the parameters that I'll need to estimate are the slope M and the constant C.
[01:47:38] So, no matter what we use,
[01:47:41] to, uh, designate
[01:47:43] Uh, you know, a given set of points,
[01:47:46] What we will use are some parameters. So, this is the parametric method.
[01:47:52] of data representation. Is it okay?
[01:47:59] Yeah, any other questions, anyone?
[01:48:00] Ah yes, ma'am, got it. Thank you.
[01:48:11] So, uh…
[01:48:13] Uh, if there are any questions, then I'll proceed further on one or two topics before we wrap up for today.
[01:48:21] Okay, I can see…
[01:48:28] Okay.
[01:48:35] So, the next topic will be on similarity and dissimilarity, and since it's a very different topic,
[01:48:44] I like to take it in a new, uh, class. However, I'll just introduce what we mean by similarity or dissimilarity.
[01:48:51] So the amount of, uh, uh, alikeness
[01:48:55] Across points is said to be the similarity in the points.
[01:49:00] Those points, which are very similar,
[01:49:02] or, uh, are said to be very alike, or those points which are…
[01:49:06] Uh, like each other are said to be similar.
[01:49:10] And usually, similarity would be represented in a range, say, 0 to 1, where 0 means no similarity.
[01:49:16] And 1 means maximum similarity, where…
[01:49:20] Uh, the one data element is equal to the other.
[01:49:23] The converse of similarities, dissimilarity,
[01:49:28] Where we try to measure how different the data elements are with respect to each other.
[01:49:33] So, uh, if we have a high amount of dissimilarity,
[01:49:37] Then, uh, this means that…
[01:49:39] The data elements, uh, being, uh, discussed are very different.
[01:49:45] If you have a low amount of dissimilarity, means they are similar to each other.
[01:49:49] So, minimum value of this similarity could be 0. When dissimilarity, 0 means the data elements are exactly equal.
[01:49:58] Or, uh, they are exactly alike each other.
[01:50:00] And if we have a dissimilarity value equal to 1,
[01:50:03] means that the data elements are very different, almost reverse of each other.
[01:50:09] So, uh, this is what similarity and dissimilarity mean.
[01:50:13] Sometimes, we also represent similarity by proximity.
[01:50:17] So, proximity is a generic term that could be used to represent similarity or dissimilarity.
[01:50:24] So, when we say proximity of a point, we actually mean the distance
[01:50:28] or the closeness of a point with respect to some other point.
[01:50:32] So, uh, we could represent the similarity using proximity.
[01:50:37] Um, uh, like that.
[01:50:39] So, uh… so then we'll… in the next turn, we'll be talking about different measures.
[01:50:45] about similarity and how to find out the similarity between points, or how to find out the dissimilarity between points.
[01:50:53] So, uh, as of now, I'll just, uh…
[01:50:57] stop here and invite any questions, if there are.
[01:51:03] Ma'am, can we say that similarity, dissimilarity are same as relation and correlation, which we do in heat map?
[01:51:13] Um, correlation often means co-occurrence.
[01:51:17] or correlation can mean, uh, some kind of relationship
[01:51:22] Uh, between the two, similarity may not be correlation always. Similarity means the…
[01:51:28] Okay.
[01:51:29] Um, the closeness between the two points. For example, I can say that
[01:51:37] Uh, this cloth material is similar to the other one.
[01:51:41] So, the two are alike.
[01:51:43] each other. Uh, right? But if I say co-occurrence, or if I talk about correlation,
[01:51:44] Mm-hmm.
[01:51:50] Then I can say that
[01:51:52] Uh, material A may be correlated
[01:51:55] with Material B. For example, Material A, short, may be correlated with Material B, trouser.
[01:52:02] So, that is correlation.
[01:52:03] Okay. Got it.
[01:52:08] Any other questions, anyone?
[01:52:16] Okay, I think there's a question on whether we'll be having any hands-on, uh, covered for whatever we have covered today, yes?
[01:52:22] You'll have hands-on on whatever I teach, not just today, always throughout the course.
[01:52:28] Whatever I teach, the theory will be followed by a hands-on session.
[01:52:32] So, we will be having hands-on, either starting next week or next to next week.
[01:52:37] So, either I'll cover some more theory next week, and then start with the hands-on, or whatever I'll, uh…
[01:52:43] have it announced on the LMS.
[01:52:45] Uh, so far, uh, we have used Colab.
[01:52:48] Uh, I think it will be…
[01:52:51] a link for collab will be available on your LMS, uh, and I'll also ask, uh,
[01:52:57] The FutureN's team to provide one if it is not already provided.
[01:53:00] So, you can use Colabs separately also.
[01:53:03] Uh, or whatever way you want to do the coding. But for us, we'll be using collab so that it's a single
[01:53:10] environment, uh, used by everyone, and you can…
[01:53:14] run and see the results, similar to what we are doing here.
[01:53:19] So, we will definitely have hands-on.
[01:53:21] Okay.
[01:53:25] So, uh…
[01:53:28] Any other questions?
[01:53:30] Uh, anyone?
[01:53:34] Uh, definitely, I'll suggest those who are familiar and those who are new.
[01:53:37] All of you to, uh, just have a recap on whatever I've thought.
[01:53:42] Uh, if you will, uh, keep reading whatever I've taught,
[01:53:46] Or it will be good because you'll remember.
[01:53:48] And everything will… is actually sequenced out in the sense that what you're learning today will be used tomorrow.
[01:53:54] And so on and so forth. So right now, I'm…
[01:53:57] Talking about data preprocessing, it will be used throughout the…
[01:54:01] course, because whenever we have data, it will require pre-processing. So, you need to know these concepts.
[01:54:08] If you are knowing already, it's good. If you are not knowing, and if you are new to the field,
[01:54:13] Uh, then you should, uh, actually go through the material that is provided by
[01:54:18] future rents are by me, the study material.
[01:54:21] You can also consult a net and find out some free…
[01:54:25] 30 material for any particular topic.
[01:54:28] And familiarize yourself with the topic.
[01:54:34] Anyone, anything else?
[01:54:42] Yeah, so I think there is some question about the detailed syllabus that, uh…
[01:54:47] So I had shown you the syllabus. You will be provided with that same, uh, syllabus.
[01:54:53] Uh, what I had shown you, only thing is that, uh…
[01:54:56] It is just, uh, under preparation. That is, I've shown you the topics, but uh…
[01:55:01] We are just associating the weeks and all, so you should… hopefully by the…
[01:55:06] Next week, you'll get the entire schedule. So, don't worry, we will not give you a partial one.
[01:55:12] We'll give you the entire set of topics, week by week, how they'll be followed.
[01:55:18] Okay?
[01:55:20] Okay, Aditya, you have a question?
[01:55:22] Uh, yes, ma'am. It's not a question, it's a concern. It's not regarding this class.
[01:55:28] Mm-hmm.
[01:55:29] It's for the batch manager. Uh, I've sent a couple of emails more than a week ago, and I've got no reply on that. Batch Manager, can you please look into it?
[01:55:37] I'll send you a DM also in the chat, but I don't know if you've read it, I've not gotten any reply as of now.
[01:55:44] Uh, so, uh, I think the future is co-host, if you can please take a note of the concern of Aditya.
[01:55:52] Uh, I hope, Atul, you are there.
[01:55:58] So, actually, uh…
[01:56:02] Um, so usually we have Simran, who takes care of the sessions.
[01:56:07] So, today, she's not here.
[01:56:10] And we were supposed to have Atul in her place. I'm sure Atul must be there. Atul, can you please, uh, respond?
[01:56:22] Also, I'm not sure about, uh…
[01:56:28] what I thought was… was Atul would be there. Anyhow, uh…
[01:56:37] GenAI Batch 2 Manager, could you please respond?
[01:56:47] Yeah, I've tried pinging, uh, him or her in the chat, but I've not gotten any response from him.
[01:56:53] So, yeah, so actually today, Simran is away. Probably you can write to Simran.
[01:56:59] Uh, about the concerns, uh, on the email ID provided.
[01:57:03] Uh, and I'll also, uh,
[01:57:07] So I'll request, uh…
[01:57:09] Yeah, we have only one email ID, ma'am, uh, as far as I know, and uh…
[01:57:14] They're very unresponsive on that email.
[01:57:20] Okay.
[01:57:21] No, no, no. Actually, actually, they had given one more email for the Simbrown one, but the Simbran email address. I'm always getting mail delivery only.
[01:57:32] Okay.
[01:57:33] I think I've saved with me as well, Aditya. So, I think I…
[01:57:37] So, I'll just, uh…
[01:57:40] take a note of, uh…
[01:57:43] Uh, this thing, and I'll try. I hope the co-host here is also…
[01:57:49] hearing.
[01:57:52] So, I'll request Simran to…
[01:57:55] uh… respond.
[01:57:57] under concern that you have raised about the management.
[01:58:02] Uh, okay.
[01:58:03] Yeah, if you can, uh, put in a word, ma'am, I think that would be great help for us.
[01:58:08] Uh, yeah, I'm doing it right away, I'm just texting Simran.
[01:58:12] Uh, she's away today, otherwise she would be here.
[01:58:17] Uh, okay, so I'm just doing that. I've done it already.
[01:58:22] So, yeah.
[01:58:24] So, anything else before we wrap up for today?
[01:58:32] Any other concerns, questions?
[01:58:37] Ma'am, again, there's a question regarding attenders on LMS.
[01:58:41] There is a tenant section, so again, this question will be for Simbran.
[01:58:45] Once you join next time.
[01:58:47] Yeah, so you can, uh, have, uh, take a note, and then next time she'll be there.
[01:58:53] So, you can just, uh, please find out from her. So, she'll be the best person to answer.
[01:58:59] Thanks.
[01:59:00] Sure. Thank you, ma'am.
[01:59:02] Gitu, do you have a question?
[01:59:03] Yeah, ma'am. Uh, from, uh, dimensional reproduction, you have explained one feature selection and feature discussion, right?
[01:59:13] Yes.
[01:59:14] Can you explain a brief that which is a selection part?
[01:59:17] Just show. So, uh, let's say we have a data in which there are some 15 attributes given to us.
[01:59:27] But sometimes it may happen that not all 15 may be useful for our analysis purpose.
[01:59:33] And we may want to retain
[01:59:35] say, out of 15, some…
[01:59:37] 10 of them, or 9 of them.
[01:59:40] So, that is what is feature selection about, that we try to choose the relevant features.
[01:59:45] And retain those, and uh…
[01:59:48] The rest of them we ignore for our further analysis. So, we are actually selecting a subset
[01:59:54] of the features from the given full set.
[01:59:57] This is what is feature selection.
[01:59:59] And then there are different ways to perform feature selection.
[02:00:02] For example, forming a decision tree is one of them.
[02:00:07] Is it okay, cheetu?
[02:00:14] Okay, there are no further questions or comments, then we can break for today. I will be passing the slides to…
[02:00:23] Same done, and hopefully she'll upload on the LMS. She's away today, uh, so I'm not sure.
[02:00:28] But I've also shared, uh, yesterday's slides with her.
[02:00:33] So, hopefully, you'll get it today, but…
[02:00:36] If she's away, probably by tomorrow, so maybe I'll request you to bear for this time.
[02:00:40] However, she's quite prompt, and immediately we upload the materials. I mean, I just pass it on to her, and…
[02:00:47] She uploads on the items.
[02:00:51] I'll also try to share in future classes a set of some
[02:00:56] questions, maybe some, uh, few multiple-choice questions, which are not graded.
[02:01:02] But those are exercise questions. If you wish, you can just attempt them and see what you have learned.
[02:01:07] Okay?
[02:01:10] So, uh, with that, I'll just wrap up today's class. Thank you all, have a great Sunday. Thank you.
[02:01:16] Bye-bye.
[02:01:17] Yeah, it's okay.
[02:01:18] Thanks so… I'm spelled.
[02:01:19] Thank you.
[02:01:20] Thank you.
[02:01:21] Thanks, man, bye.
[02:01:22] Thank you.