# 02 2025-11-09 Data Pre-Processing and Curation course: Module 1 — Foundations of AI & ML module: Module-1-Foundations-AI-ML date: 2025-11-09 type: transcript video_url: https://personal-learn.armco.dev/files/_Recordings/Module-1-Foundations-AI-ML/02_2025-11-09_Data_Pre-Processing_and_Curation.mp4 --- [00:19:34] There are immense applications of classification. [00:19:38] Such as segmentation, frog detection, etc. [00:19:43] Then, uh, I introduced to you the concept of discriminative models and the classification, uh, [00:19:50] is a discriminative model. [00:19:53] Which helps us to distinguish between the different categories that are available. [00:19:58] In the data. [00:20:00] And discriminative models usually [00:20:03] help to classify, and then use this classification model. [00:20:07] To predict the class labels from the input data. [00:20:11] And, uh, the idea is to find out a function. [00:20:15] Uh, that helps to, uh, make a boundary, uh, that separates the different classes that are there in the given data. [00:20:23] And the output here would be the class label. [00:20:27] Now, talking about unsupervised learning… [00:20:30] That, in term, unsupervised learning is coined [00:20:34] Because, uh, here in this case, we do not have any class labels given to us, we just have that data. [00:20:41] And we try to group that data, or find out the natural groupings that is available in the data. [00:20:47] With the help of, uh, similarity measures. [00:20:52] So, similarity measures are those neat tricks or measures that help us to [00:20:56] assess or quantify the similarity between data points. [00:21:01] And based on that singularity, the points are clustered together, or grouped together. [00:21:06] In such a way that points which are… [00:21:09] Uh, very close together. [00:21:11] output in the same cluster, thereby… [00:21:14] Minimizing the intra-cluster distances and maximizing the inter-cluster distances. [00:21:20] So, then we talked about, uh… [00:21:24] clustering applications, [00:21:29] Uh, then we discuss regression and forecasting. [00:21:32] So, I, uh, had introduced to you, uh, what is predictive analytics and how [00:21:38] Classification plays a role in predictive analytics. [00:21:43] Uh, when we talk about predicting value, then it's called forecasting, and regression can be used for forecasting. [00:21:51] We also talked about anomaly detection, where we try to identify abnormal behavior. [00:21:57] from the given data, it's also called outlier. [00:22:01] identification, and different types of supervised and unsupervised learning methods can help, too. [00:22:08] perform anomaly detection. [00:22:11] Then, I introduce to you deep learning techniques, where, uh, neural networks are used. [00:22:17] Which are inspired by the human brain and human neurons. [00:22:21] Which are there in human brain. [00:22:23] And whenever the data that we are having is very complex, and the patterns are also very complicated, [00:22:30] Uh, then deep learning methods are used. [00:22:33] Such as on data like text, images, audio, etc. [00:22:38] And there are many deep learning techniques, like convolutional neural networks, recurrent neural networks, long-short-term memory, [00:22:46] LSTM. Uh, and so on and so forth, which we'll be covering in the subsequent, uh, sections. [00:22:54] Then, deep learning has immense applications in various domains, like computer vision, NLP, speech recognition, healthcare, finance, and… [00:23:03] So, so for… [00:23:06] And finally, we talked about generative AI. [00:23:10] Uh, well, um… [00:23:12] There are… we talk about AI models that are able not just to [00:23:18] Analyze the data, but to create new content. [00:23:21] Uh, that new content could be text, audio, video, images, code, [00:23:25] Or, uh, any other, uh, data types. [00:23:29] And they are capable of understanding the context, the style of writing, the structure, [00:23:35] That should be available. Uh… [00:23:37] In a way that is similar to human-like content generation. [00:23:42] And, uh, generative AI, uh, models work on huge quantities of data, or corpus. [00:23:49] And they use this to learn the patterns, and then… [00:23:52] After learning, they generate new patterns, which may be similar [00:23:57] to the patterns, they learn, or they may be of different types. [00:24:02] Then, uh… [00:24:04] So, uh, unlike the discriminative models, [00:24:09] Generative AI talks about generative models. [00:24:14] So, the purpose of discriminative model is just to classify and predict class labels. [00:24:18] However, the purpose of generative AI is to generate new samples. [00:24:23] And uh… there are different, uh… [00:24:27] types of algorithms that fall in each of these categories. For example, discriminative model could be logistic regression. [00:24:34] SVM or other classifiers, whereas generative model. [00:24:38] could be GPT, but… [00:24:40] guns, and so on and so forth. [00:24:44] And then, um, the… [00:24:47] Generative AI models progress to agentic AI. [00:24:52] Where there are agents, uh… [00:24:54] one or multiple agents. Usually, there are multiple agents. [00:25:00] start to work in collaboration with each other, and… [00:25:03] They are, uh, autonomous in nature, and they are capable of… [00:25:07] learning and objective. [00:25:09] or goal, then working towards… [00:25:12] Completion of the goal. [00:25:14] planning, coordinating… [00:25:17] Uh, dynamically changing, taking feedback. [00:25:20] improving, and so on and so forth. They have huge, uh, capabilities. And they act like assistants. [00:25:28] Collaborative assistance, uh, which help humans. [00:25:31] to perform the task with minimal human intervention. [00:25:36] So, there are, again, a lot of applications of agent-type AI. [00:25:41] So this is what we had covered. [00:25:43] In the, uh, initial part. [00:25:47] Then, uh… [00:25:50] Going to the next portion. [00:25:54] I will just share the slides. [00:25:59] Uh… so I think I could see a query on the slides on LMS portal. [00:26:06] Uh, so I have shared the slides, uh, today morning, uh, yesterday it was a little late. [00:26:12] And I'm sure that they will be uploaded very soon on the LMS, so you will find, uh… [00:26:17] the slides for yesterday on LMS, and also for today. So, I will be… [00:26:22] Uh, sharing, uh, to the concerned Futurance people in the afternoon today. Today's slides. [00:26:29] And they'll also be available on LMS, so you may… please don't worry, you will find them. [00:26:36] Okay, so let me now… [00:26:42] Share the other slides. [00:27:01] Okay, so I hope you all can… [00:27:04] see my screen. [00:27:08] And let me maximize it. [00:27:13] Okay, so yesterday we started with the second portion, which was on data preprocessing and curation. [00:27:20] So, uh, I discussed that there is a requirement of the data to be, uh, quality data, so that [00:27:27] the analysis on such a data. [00:27:30] could yield, uh, good results. [00:27:33] Uh, and there are different tasks that can be performed as part of data preprocessing. [00:27:39] Such as data cleaning, [00:27:42] data integration, data transformation, reduction, and discretization. [00:27:47] We started out discussing some of these yesterday. [00:27:51] So, we discussed, uh, in-data cleaning, handling of missing data. [00:27:56] And, uh, we discussed how, uh, different, uh, percentages of missing [00:28:01] values in the given data could impact [00:28:06] Uh, the missing value imputation. [00:28:09] So, we concluded that if the percentage [00:28:12] of missing values are very less, then we could actually go with removing [00:28:17] those records. However, removing of records is not a very good solution in case [00:28:23] Uh, the missing values are… [00:28:26] higher… higher in percentage. [00:28:30] Uh, so, for example, uh, if the… [00:28:34] We discussed that if around 30% of the data [00:28:37] has missing values, and uh… in this example, if we remove such rows, [00:28:43] then it is possible that the majority, and if there is a majority of one particular class, like Class C2, [00:28:50] It is possible that it can become a uniform or almost uniform distribution. [00:28:55] Uh, and it can also be possible [00:28:59] that one of the classes can be completely lost, and the minority class could become a majority class. [00:29:05] Uh, if the percentage of… [00:29:08] Missing values is high. [00:29:15] then we, uh, talked about other solutions. [00:29:20] So, uh… [00:29:24] So these are just the examples that I had shown to you with different percentages of missing values. [00:29:31] Also, uh… [00:29:35] There, uh, we talked about different solutions, uh, for this problem. [00:29:40] So, another solution could be to fill in with a default value. [00:29:44] say, any value, say 100, 9, 10, or whatever. [00:29:48] And, uh, we also, uh, discussed at length that feeling of missing values could actually lead to the creation of a new class. [00:29:57] In the data, if the percentage of the missing value is high. [00:30:03] Then, uh, the next solution was filling, um, a mean value. [00:30:07] So, when, uh, mean value is filled, [00:30:10] Again, there is a possibility, and the mean could be a classified mean, or whatever way. [00:30:17] Uh, there's again a possibility. [00:30:19] that, uh, uh… [00:30:22] The points that are, uh… [00:30:25] The points that get those values that are being filled on the basis of mean, [00:30:29] are leading to the creation of points, uh, which are very different from the distribution. [00:30:35] The regular distribution of the given set of points. [00:30:39] So, feeling of, uh… [00:30:41] Missing values with mean is also not a very good solution. [00:30:46] Uh, then, uh, we, uh, also discussed filling off, uh, missing values using median. [00:30:54] So, when we use median, [00:30:57] Then, filling of missing values. [00:30:58] uh… happens in a better fashion, but… [00:31:02] It's still not a very good, uh, example, because again, we are creating points. [00:31:07] Which are actually very different from the distribution of the points given to us. [00:31:14] Then we also tried to fill with mode, and we found out, again, that, uh, in some cases, [00:31:22] It is a good solution, whereas in the other cases, [00:31:24] It is not a very good solution. [00:31:27] Uh, to do that. [00:31:31] So, uh, then we also talked about… [00:31:37] About filling of, uh, missing values using other methods, so mean, median, and mode. [00:31:45] are somewhat better, but they are not that much better. [00:31:52] So, we, uh, actually were in a need to have other methods also. [00:31:58] Sorry. [00:32:01] And, uh, we talked about feeling of mean based on, uh, classes, uh, class labels. [00:32:08] For the corresponding attribute, and this… [00:32:11] came out to be a much better solution. [00:32:14] And even, uh, more beta solution was to use probabilistic [00:32:19] methods to fill the missing values. [00:32:23] And this turned out to be probably the best solution. [00:32:28] Uh, that we discussed. So, I think we had finished this much yesterday. Now I'm going to continue with the… [00:32:35] New portion. [00:32:36] Okay. [00:32:39] So… so these are the different solutions that we discussed. [00:32:44] Uh, ignoring of the tuple or removing it. [00:32:47] Then filling, uh, with missing values, some missing values, uh, manually. [00:32:52] Filling a global constant. [00:32:56] Or, uh, using the mean media norm mode, using attribute mean. [00:33:00] On using the most probable value to fill the missing value, which turned out to be the best solution. [00:33:07] Now, in data cleaning, uh, in addition to, uh, [00:33:12] handling missing value. [00:33:14] Uh, the next thing that we wanted to do was handling noisy data. [00:33:20] So, we are going to discuss methods where we are going to handle noisy data. [00:33:26] Just a second… yeah. [00:33:29] So, there are… there can be different ways. Uh, one, um, common way of [00:33:34] handling noisy data is to perform binning. [00:33:37] So, in beginning, as the word bin implies, VIN is a container. [00:33:42] or a bucket, where we first sort the data points. [00:33:46] And then partition them into groups. [00:33:49] On bins? Are we… [00:33:51] Uh, assume that we are partitioning and putting the points in bins. [00:33:55] And then we can apply different types of smoothing. [00:33:59] Uh, on the bins. [00:34:01] Which I will discuss as we go along. [00:34:03] So, binning is one popular method of handling noisy data. [00:34:08] Where we just retain the VIN information and do away with the noise. [00:34:13] Then, clustering also can be used to remove noise or outliers. [00:34:18] Regression also can be used to fit a curve, and all the points that do not lie on the fitted [00:34:25] Uh, regression function. [00:34:28] are actually removed as noise points. [00:34:32] So, let's take each of these methods in more details as we go along. [00:34:36] So, we have binning, uh, for our data smoothing, or removing noise. [00:34:42] Again, in binning, there are different methods to perform binning. [00:34:45] So, the first method to perform binning [00:34:48] uh… is… [00:34:50] equivalent partitioning, or equivalent bins. [00:34:56] So, in this particular case, what we do is that we take the range of values for that… for the… [00:35:01] a variable that we have to bid. For example, if we have height, [00:35:05] And we want to bin the heights. Then we take the range of values, [00:35:10] Uh, within which the height is falling in the given set of data. [00:35:15] And then, uh, we divide it by the total number of bins that we want to form. [00:35:22] For example, here, we are interested to form capital N number of bins. [00:35:27] This doesn't hang up. So, these are the number of VINs that we want to form. So, we will divide [00:35:33] The given range of values, let's say the highest value was B, [00:35:38] And the lowest value was A for that particular attribute that we want to build. [00:35:44] So, uh, the width of the VIN will be given by B minus A divided by N. [00:35:49] So, uh, this is how, uh… [00:35:52] VINs are formed, just a second. [00:36:02] So, in this particular case, uh, all the VINs are having equal, uh, width. [00:36:07] So, what are we trying to do? Let's say, uh… [00:36:11] This is the VIN number. [00:36:14] And, uh… [00:36:16] What we are having here… [00:36:19] is some variable. [00:36:22] And what we try to do is that we try to divide it into equal width [00:36:28] partitions, like this. [00:36:30] So we are actually kind of dividing it into strips. [00:36:35] So here's an example. Uh, let's assume we have a data in which there are 50 rows. Not all 50 are shown here. Some of them are shown here only. [00:36:44] For illustration purpose. [00:36:46] So, equivalent winning is to be done on the data, assuming the total number of bins that we want to firm up is 5. [00:36:53] The minimum age in this data is 20, and the maximum age is 70. [00:37:00] Then, uh, what we will do is, uh, we'll determine the bin size as… [00:37:06] Uh, uh, you know, given by… [00:37:08] 50, 70 minus 20. [00:37:10] So it will be 70 minus 20. [00:37:13] divided by 5, which will be 10. [00:37:16] So, the width of the VIN will be 10. [00:37:19] And the different VINs that can be formed will be VIN0, VIN 1, VIN 2, VIN 3, VIN 4. [00:37:24] Which will be of size 10 each, like 20 to 30, 30 to 40, and so on and so forth, till 70. [00:37:31] Now, all the points are assigned to these pins. [00:37:35] So, for example, the age 20. [00:37:37] So, age 20 is going to belong to bin 0. [00:37:41] So, it is assigned bin ID 0. [00:37:43] So, any age in 20s is assigned to bin 0, so we can see here. [00:37:48] Because the data is sorted, so the initial values [00:37:52] will be assigned to, till here, will be assigned to. [00:37:55] bin 0, and uh… [00:37:57] So, we can assume that, uh… [00:38:01] Any particular VIN is inclusive of the [00:38:03] upper bound. So, this means that this pin… [00:38:06] includes values from 20 up to… [00:38:10] 30. And any value greater than 30 will come in bin 1. [00:38:14] So, up to 30, we are having bin 0. [00:38:17] Similarly, up to 40 we are having been… [00:38:20] 2, and so on and so forth. VIN 1 and so and so forth. [00:38:24] So now, if we visualize this data… [00:38:28] What we can see here is these are the VIN IDs. [00:38:34] Justice… [00:38:36] Yeah, so we have the different VIN IDs, as you can see here. [00:38:40] And, uh, let's say we are binning… [00:38:43] Medical expenses in rupees. [00:38:46] Uh, so the medical expenses, when they are binned, [00:38:49] Then what we find out, uh… [00:38:52] On the basis of VIN information. [00:38:56] that, uh, these are the 5 bins, and the crosses show the points that have been assigned to these VINs. So, these many points are assigned. [00:39:03] 2.0, bin 1, bin 2, bin 3, and bin 4. [00:39:08] So, uh… [00:39:10] So, what are your comments on this kind of assignment? [00:39:20] any comments or anything that, uh… [00:39:22] to note, here. [00:39:24] To form a equivalent cluster, like, because the data is going to be in a specific range, so it is, like, a more likely [00:39:33] Okay. [00:39:35] And the distance is less, uh, compared to… [00:39:39] those, like, the 4 has a distance, uh, lesser than… [00:39:44] Uh, the two are… [00:39:48] 3… [00:39:50] Uh, you're talking about the… [00:39:52] distance the points. [00:39:53] The distance between the, uh, each other. [00:39:56] Oh. Oh. [00:39:57] Good point. [00:39:58] intercluster, yeah. [00:40:01] It is. [00:40:02] As it seems like we have customized, uh, this. [00:40:03] Just one question. So, is winning software. [00:40:08] Can you come back again, Neeraj? [00:40:09] data… Yeah, it seems like that we have categorized the data. [00:40:17] Yeah, so… [00:40:18] Into… Into 5 correctly. [00:40:19] Yeah, so actually, bins are kind of, uh, creating categories only. [00:40:25] So, we have now created categories out of our data. One category is people in their 20s up to 30. [00:40:33] people in their 30s up to 40 like that. [00:40:36] And, uh, definitely, [00:40:38] Uh, we can see the distribution of the points, uh, more or less. [00:40:42] Uh, average number of points in the VIN might be… [00:40:46] not exactly equal, but… [00:40:49] somewhat, uh… [00:40:51] uniform. Of course, there have been 3 and 4 are more populated than VIN 0, 1, and 2. [00:40:58] Uh, but all the VINs do have points in them. [00:41:02] So, Deepak, do you have a question? [00:41:05] Yeah, I'm just trying to get the relation between binning and the clustering, so is it… Somewhere related because I think we are doing the same thing here as well. [00:41:16] You want to understand, uh, the relationship between bins and clusters? [00:41:22] Yeah, so is it the part of the process, uh, to create a cluster? [00:41:27] Uh, you can think of them as clusters, but these are kind of static clusters. [00:41:33] So, when we perform clustering of the data, [00:41:37] We find out the similarity between the points. [00:41:40] And on the basis of the similarity, we cluster them or group them. [00:41:45] Here, we are not finding out [00:41:47] Any similarity across the points. All we are doing is, based on the value, [00:41:52] of the attribute that we want to bin, we are just assigning the VIN ID, some bin ID, [00:41:58] to the point. So, there is no point-to-point similarity that is getting calculated here. [00:42:04] So, binning is not exactly clustering. [00:42:07] However, it does form groups in some way. [00:42:09] But these are not exactly clusters. [00:42:12] Is it okay? [00:42:15] Yeah. Uh… [00:42:17] Then, uh, of course, uh… [00:42:20] While my circles were there, so… [00:42:22] There can be points, I mean, [00:42:25] So these points, for example, [00:42:27] Either they could be considered… so here, since we are not talking about, uh… [00:42:33] Anomaly detection, so we are assuming all the points to be regular or normal points. [00:42:38] So, uh, we may not say that they are outliers, but if required, they can be treated as outliers if they are really different from the rest of the points. [00:42:48] As of now, we are assuming that there are no outliers in the data. So, actually, all of them will be form of [00:42:54] will be part of the bin. All the points. So, even though my circles… [00:42:58] got a little smaller. [00:43:01] Okay, so Neil, what question do you have? [00:43:03] Yeah, hello, ma'am. So, I'm just trying to understand, so… [00:43:07] What is the noise in this data? What we are trying to do here? [00:43:12] When we are dividing it into the bins. [00:43:16] So, uh, so actually, uh, [00:43:19] Noise points are removed. Yeah, so if we are doing denoising using binning, which is what I started out with, [00:43:29] So, uh, actually, uh… [00:43:32] What is happening here is that there are, let's say, if we were to plot these points, [00:43:36] Like this, let's say, then they would be, like, jittery, right? There will be points here in all the range, like this. [00:43:44] So, what are we doing? We are smoothing it out by replacing it with some value. [00:43:49] And that value would be the VIN ID. Say, I call this… [00:43:54] You know, this as all the values inside this, I'm replacing it with some range, R1. [00:43:59] And I call it VIN0. All these values, I'm replacing with some other [00:44:04] So, uh, like this, we are actually removing all the jitter and just replacing them by these VIN IDs. [00:44:12] So, we are kind of smoothing this data, replacing all these jitters with these simple numbers. [00:44:16] This is how we are denoising the data here. [00:44:18] Okay, so I think it is somewhat related to the classes or clusters. [00:44:26] So, once again, I'll say that in clustering, we try to look at the similarity of the points. [00:44:32] Here, we are not looking at similarity, we are just dividing the given range, which I already showed you. [00:44:38] If this is the highest value, this is the lowest value. [00:44:41] We just subtract them and divide them by the total number of bins to be formed. [00:44:46] So this gives us the width of the bin. [00:44:48] And then, based on the width, we just keep on assigning points. You can visualize it like this. Let's say if I have 5 wins, you think of 5 buckets. [00:44:57] And let's say if you have some colored balls, then each bucket is going to hold one particular color. [00:45:02] So, as you have the balls in your hand, you keep tossing it. [00:45:05] In the respective bucket, if there's a bucket for white balls, you toss it in the… [00:45:09] And you have a white ball with you, then you toss it in the respective bucket and like that. [00:45:14] So instead of having all the balls separately with different colors, now what we have are buckets. [00:45:20] And we just handle the buckets rather than the individual balls. So, similarly here, [00:45:26] We are having bins or buckets. We, uh, just place the points in each of these bins, and then forget about the points. [00:45:33] And we are only interested to retain the VIN information. [00:45:37] And we are going to use that only. So, in a way, we are kind of categorizing the data. [00:45:41] from, uh, continuous valued, uh, you know, range to a range, uh, which is, uh, [00:45:50] finite ranges. Like, for example, [00:45:53] If we talk about… [00:45:55] Uh, the age from 20 up to 30. [00:45:57] Then there could be infinite values. Instead of that, I'm replacing it by… [00:46:03] one value, zero. [00:46:04] Like that. [00:46:05] Go to the area, thanks. [00:46:13] So then, uh… [00:46:14] We have another, uh, example. [00:46:18] Where we are talking about, uh, daily, uh, screen time and sleep duration. [00:46:24] And, uh, the dependent variable here is leave duration. [00:46:29] Uh, so, uh… [00:46:31] So, sleep duration is what is going to be, uh, the dependent variable [00:46:38] And the rest of the screen time is independent variable. The total number of points that [00:46:43] are given a 52, and the screen, uh, time range is given. [00:46:48] And then, uh… [00:46:50] We want to form 6 VINs, and according to that, uh… [00:46:54] According to the… [00:46:57] Equiva at the beginning, uh, we divide the given set of points from the highest to the lowest value. [00:47:03] into 6. So, the ranges that we get are shown here. [00:47:08] And, uh, now, when we do equi-width binning, [00:47:13] Then, uh, we have these ranges. [00:47:16] Right? You have these ranges with us. [00:47:19] And when we try to accommodate the points in these ranges, then this is the distribution that you are seeing here. [00:47:26] So, uh, this graph shows the distribution. [00:47:31] So, before I proceed, I could see a hand raised. Uh, is there any question? [00:47:39] Okay, G2 do, what do you want to know? [00:47:40] Oh, yeah. [00:47:41] Uh, about that binning already, uh, like, uh, when the data distribution is large, we are binning the… [00:47:49] data to small sets, and we'll be analyzing only that particular range of values, right? [00:47:55] For that, only we are using Building, right? [00:47:56] Yes, so we are, uh, as of now, [00:48:01] Binning is used in a lot of… [00:48:03] things. As of now, I've talked about denoising the data. [00:48:08] By ignoring the individual values, and only retaining the VIN IDs. [00:48:13] And these bill IDs will be used for further analysis, okay? [00:48:21] Okay, so now I'm just showing you another example of EQ with winning. [00:48:27] So, any comments on this kind of winning? [00:48:30] On equivalent winning on this particular data. [00:48:35] Uh, yes, so… Unlike the earlier example here, it seems all the points are constituted on bin 4 and 5. [00:48:46] And… so here, binning is not a good representation of the. [00:48:52] Uh, screen time. In this case. So if we replace… The actual screen times with these beans. [00:49:01] It would appear all of them are only dependent on 4 and 5, which kind of makes the other bins. [00:49:07] Yeah, so what you're saying is correct, Pallavi? [00:49:08] Uh, not useful. [00:49:11] So, if we are using EQI with binning, [00:49:14] In this particular case, then what are we having? We are having most of the points are crowded, [00:49:20] In these two VINs, whereas the other VINs are practically empty. [00:49:26] So remember why we want to form the bins. We want to form the bins so that we uniformly distribute the [00:49:32] So, we should not have bins just for the sake of having them. [00:49:36] Here, we have formed bins like bin 0, 1, 2, 3, and 4. [00:49:40] And these VINs are kind of just formally there. [00:49:43] But they don't really have much points in them. [00:49:46] 1-1 point, or there might be a… it is… it is also possible that there might be a VIN having no points at all. [00:49:54] in it. And therefore, equie with winning in this particular case, where the data distribution is such, [00:50:01] is not very useful. [00:50:03] Because some bins are practically empty, whereas others are crowded. [00:50:07] So, this defies the… [00:50:11] logic-affirming bins. [00:50:12] So that's why, in this case, equi with winning is not a good solution. [00:50:20] Okay, Deepak, what question do you have? [00:50:23] Yeah, I think on the similar line, uh… That… uh, so basically, the purpose of billing is to categorize data or make it more simplistic, so that we can do the next. [00:50:35] set of process, like clustering or whatever it is. But in this scenario, uh, probably the data distribution is. [00:50:44] Like, there is huge difference, where, you know, we see few points on one line. [00:50:47] And 4 and 5 years used. So we would ignore this step. Am I right on saying this? [00:50:53] Uh, I couldn't understand what step you want to ignore here. [00:50:57] Can you…? [00:50:58] Uh, the billing… the process of billing. Or categorizing the data with the help of bidding. [00:51:01] In this data set, you are saying? [00:51:04] Okay, so the… we want to bin the data. [00:51:05] Yes, yes. [00:51:08] Right? We have to form the bins. [00:51:11] But when we are using equivalent winning, definitely the meaning of, uh… [00:51:16] Binning is getting defied because, uh, there are practically these two wins that are having points. [00:51:22] And most of the points are coming in these two bins, and the rest of the bins are [00:51:27] Uh, you know, just practically empty. [00:51:30] You can also take another example. Suppose, uh, we want to bin. [00:51:35] The students in a… in a particular course. Let's not talk about PG certification, because there are a variety of [00:51:43] people from different backgrounds, but let's say we have students in [00:51:47] We take fourth year, and we want to win them. [00:51:50] On the basis of age. [00:51:54] And we use equie with binning. [00:51:56] So, if you are using EQ with binning on the students of BTEC 4th year, [00:52:02] Uh, on the basis of age, then what would happen? [00:52:06] we may form… [00:52:08] N number of bins, but mostly all of them, because they'll have similar ages. [00:52:14] All the points will, uh, you know, form part of a single VIN only. [00:52:18] mostly single VIN, not even two. [00:52:21] So, therefore, the purpose of forming bins is lost, because [00:52:25] VINs are there, but they are not having any points, so… [00:52:28] This is a useless scenario. [00:52:31] So, but we have to perform binning, uh, so, uh… [00:52:35] So, we have to look at some other way of forming bins, and equal with binning is not the solution. [00:52:40] That is code here. [00:52:43] So then, uh… [00:52:46] So, we already discussed, uh… [00:52:48] the impact of equivirth winning. [00:52:52] Uh, in certain datasets, in which case, the… [00:52:58] points are in, uh, are given in such a way that all of them form part of one single VIN, or [00:53:05] some humans, and the rest of the bins are empty. [00:53:10] So then, uh, since equally with binning is not suitable in many cases, we have to go for another [00:53:18] Some other form of winning, and what we have is equally depth binning. [00:53:21] If we have the winning is also called equip frequency binning, where [00:53:25] We do not, uh, divide [00:53:28] Uh, the given range of values into equal… [00:53:32] size to bins. Rather, [00:53:35] we divide the given ranges in such a way that each range of values [00:53:42] Forms one bin, and… [00:53:44] Approximately, the total number of samples [00:53:48] per VIN are equal. [00:53:50] So, here we are talking about [00:53:52] frequency, uh, of the points in the VINs. [00:53:57] And we are not talking about the width of the bin, so we don't look at the range of values. [00:54:02] Rather, we look at the total number of points in the VIN. [00:54:16] Sorry, I got muted here. [00:54:18] So, uh, in this particular case, we are interested to form the VINs but equal [00:54:25] Uh, worth winning is not going to serve the purpose. So, we have equal depth or equal frequency binning. [00:54:32] In which case, uh, the number of samples [00:54:35] per VIN are going to be approximately equal, and when such a number, uh… [00:54:41] is achieved in any VIN. [00:54:43] There, we cut the VIN boundary. [00:54:47] So, we don't have regular VIN boundaries of equal width, but they are based on frequency. [00:54:56] So, for example, uh… [00:54:58] Let's say that we have certain, uh, the same set of points that we discussed earlier. [00:55:05] And we have the screen time, then we have the sleep duration, and the points are actually sorted on sleep duration. [00:55:12] And then we have the VIN IDs that were assigned to them based on equivalent binning, and then we have equally frequency binning. [00:55:20] So, in case of frequency binning, we are going to associate [00:55:26] Uh, we first sort the data and associate equal number of points per VIN. [00:55:32] In this particular case, we have 52 points, and we wanted to form 5 bins, so approximately 10 points are accommodated. [00:55:39] carbon. So, uh… and accordingly, when 10 points are accommodated, then the VIN boundary is [00:55:47] found out. So once there are 10 points, then we look at the… what are the range of values of, uh, [00:55:53] of the points, [00:55:56] For that particular attribute, because they are all sorted. [00:56:00] And use that build boundary as… [00:56:02] The range of values for that particular VIN. Similarly, [00:56:06] The next pin, next spin, like that. [00:56:08] So here, we are focusing on the frequency of points, so you can see… [00:56:13] That, more or less, all the VINs are uniformly packed. [00:56:16] Uh, for example, in this VIN, you're not able to see many points, because they are lying one on top of the other. [00:56:23] So, like this. So, these are the winds formed. All of them are equally populated. [00:56:28] And there are no VINs that are empty, or no VINs that are… [00:56:33] sparse or overly crowded. [00:56:36] They are all equal. [00:56:37] So, equidepth binning is used when we have the likelihood of having [00:56:42] All points, uh, falling in the same range. [00:56:45] For that, uh, for some particular attribute that is used for binning. [00:56:50] If that is the case, then don't use equal with billing, use equil depth winning or some other form of winning. [00:57:01] So, uh, I think this is the example. I already showed it to you. [00:57:09] Now, uh, there are other methods also to form, uh, VINs. [00:57:14] For example, uh… [00:57:16] We partition the given data, we have the… [00:57:20] price of some item in INR. So, first step would be to sort them in [00:57:24] some order, let's say in the increasing order. [00:57:28] And then we partition into equidepth, uh, bins. So, uh, when we, uh, talk about [00:57:35] Equidact means equal number of points are going to be put per bin. [00:57:39] So here, we are going to put 4-4 points per VIN, because the total number of points are… [00:57:44] 12, so first 4 points, then next 4 points, like that. [00:57:48] We will keep on doing, just a second. [00:57:52] Let me choose mine. [00:58:02] I'm just choosing my pen. [00:58:10] So, this is about equal depth binning, using these points. Then we have other methods, like [00:58:15] Uh, smoothing by VIN means. So, first, we actually partition the given data [00:58:21] into equi-depth bins, and then we denoise the bins so that [00:58:26] We are doing smoothing on the points by VIN means. So, we look at all these points, find out the bin mean, [00:58:33] Uh, let's say 9 is the VIN mean, then we replace all the points by 9. [00:58:37] So, it looks something like this. [00:58:39] Let's say these are the points. [00:58:42] You know, these are the values that we are having. [00:58:45] So, what do we do? We have some number of bins that we want to form. Let's say 3 number of bins are to be formed. [00:58:51] So these are the three wins. [00:58:53] And we replace each bin by the mean. [00:58:57] So, what will we have? These are the points. So, this is the VIN mean. [00:59:00] This may be the VIN mean for the next VIN. [00:59:04] This has been one. [00:59:06] This has been 2, this is bin 3. [00:59:09] And then this is also replaced by the… [00:59:12] Meanwhile. So now, these are the wins that we are having here. [00:59:18] Right? So, like this, we do smoothing by VIN means, and we, uh, lose the jitters in the data. [00:59:25] Uh… here. [00:59:27] So now, uh, the next thing, uh, is, uh, winning by VIN boundaries. [00:59:32] So, we can do a binning on VIN boundaries. [00:59:36] So, whatever the value, uh, [00:59:40] Let's say the VIN boundary for this VIN is either 4 or 15. [00:59:44] So, whatever the other values are, [00:59:46] Uh, if 8 is more closer to 4 than to 15, 8 will be replaced by 4. [00:59:53] Similarly, 9 if it is… [00:59:56] Now, nearer to 4 than to 15, then it will be replaced by 4. [01:00:01] Because we can see that 9 minus 4 is 5, and 15 minus 9 is 6. [01:00:05] So 9 is closer to 4 than to 5, so we replace it by 4. [01:00:10] Like this, we do smoothing using VIN boundaries. So, these are some different ways. There are so many other ways also possible. [01:00:17] I'm just showing some of them. [01:00:20] Then, the other method to do noise removal is to perform clustering on the given data. [01:00:26] When we cluster the given data, then we see [01:00:30] that the point, that is a noise point, will not be part of any cluster, and it will be excluded. [01:00:36] And we can remove this as noise. [01:00:41] Then, using regression for removing noise, so we have these points, which are shown here. [01:00:46] We fit our model, or a curved line, [01:00:48] on it. It could be linear, non-linear. Here it is a linear… [01:00:53] line. And, uh, now this… [01:00:56] 9 is retained, which is given by Y is equal to X plus 1. [01:01:01] And all those points which are not lying on this line, are actually ignored or removed as noise points. [01:01:11] So, uh, in regression, we try to minimize the distance [01:01:18] So, here, we want to minimize the distance between the model that we have found [01:01:22] And the points. So we want to minimize these. [01:01:25] this distance is some of such distances. [01:01:30] Uh, this is the idea to fit our line, uh, on most of the points. [01:01:36] Okay. [01:01:38] So, here's a diagrammatic representation. Let's say we have points which are shown there. [01:01:44] And these points are actually shown in this purple graph. [01:01:48] Then, uh, we could use different form of… [01:01:51] regression, like logarithmic, exponential, and so on and so forth, and accordingly, different lines will be there. [01:02:00] which represent the, uh, regression models that have been fitted on these points. [01:02:05] And the model is written, and the points are ignored from further analysis. [01:02:11] Okay. So that was about, uh… [01:02:14] denoising of the data, any questions, anyone? [01:02:18] About denoising. [01:02:24] While learning Python, we understood the. DF command to, uh, use the value from the previous, so the next record. [01:02:35] So, when that will be used in general? [01:02:39] Uh, are you talking about when… what Python command to use, or your… I couldn't understand what you wanted to ask. [01:02:49] Okay, for example, uh, we have on record for which one field is missing. [01:02:54] So, we can use the value from the previous record from the same field. [01:03:00] Or we can use the value from the next record to fill that value, like that. [01:03:04] Okay. [01:03:06] So, command, so any, any application we have here about… [01:03:10] Uh, so let's say you have, uh… [01:03:14] You know, uh, let's say you have some data here. [01:03:17] And these are the different values, and this is attribute A1. [01:03:22] And this is the missing value, and you are having value 1 here, and maybe value 2 here. [01:03:28] So what you are saying is to fill this value, you could use this value. [01:03:31] Or you could use this value, something like that. [01:03:35] So, yeah. So, um… [01:03:36] Yes, yes. [01:03:37] This, uh, can… such kind of feeling of missing values can only be used [01:03:43] When you have dependency in the attribute, let's say, uh, I have an attribute, [01:03:49] Uh, which represents the sales. [01:03:52] And that sales is for different countries. [01:03:56] Right? So, let's say, uh… [01:04:01] I don't know why it's not coming. So, I have got 3 countries. [01:04:06] Let's say one is India. [01:04:08] One is USA. [01:04:10] And one is UK. [01:04:13] And I have a missing value in USA. [01:04:15] Whereas there are values for India and UK. [01:04:18] Now, in the scenario that you mentioned, [01:04:21] Uh, would it be logical to use the sales of, uh, of, uh… [01:04:26] over India for that product to be… to fill [01:04:29] The sales in USA, or the sales of UK to fill. [01:04:33] Sales in USA. [01:04:35] Uh, will it be logical? [01:04:37] Let's say we talk. [01:04:39] Yeah, yeah. So… [01:04:40] No, certainly not, yes. [01:04:43] Because sales in different countries are, uh, uncorrelated, [01:04:47] They don't have any relationship, even if the product is same. [01:04:50] Even if we talk about sales of, let's say, [01:04:53] Any food item, say chips. So, the sales in India versus sales in UK or USA will be different. [01:05:01] Even if the brand is the same, the chips type is the same, and all that stuff. [01:05:07] So, in not all cases, you can use, but you can use this kind of a… [01:05:12] thing, when I'm having a time series with me. [01:05:15] Let's say, here I'm having sales. [01:05:18] For a product. [01:05:20] I don't know why I'm facing problems with my digital pen, sorry for that. [01:05:25] So these are sales of one particular product, and this is time, T1. [01:05:29] time T2, uh… [01:05:32] time T3, and so on and so forth. This is time 0. [01:05:38] So, uh, if we talk about sales of a product over different instances of time, and we have a missing value here. [01:05:44] at time T2. Then we can definitely use [01:05:48] Value from T1 or T3 to decide what needs to be filled here. [01:05:53] In that case, you can do this, because the data is having a dependency. [01:05:57] Uh, in sales over time. [01:05:59] So, if such kind of dependencies exist, then you can use this method. Otherwise, you can't use it. [01:06:06] Okay? [01:06:08] Yep. So let's talk about data transformation. [01:06:13] In data transformation, we apply techniques [01:06:15] I'm sorry to interrupt. Just in case we don't have data, uh, like, time dependency. [01:06:22] Mm-hmm. [01:06:23] So, what we'll do in that scenario? [01:06:26] Yeah, so if you don't have time dependency, then I already discussed so many cases, how to fill missing values. [01:06:33] For example, remove the… [01:06:35] Record containing missing values, fill with a default value, fill with a mean, median, mode. [01:06:41] Okay. [01:06:42] filled with mean per attribute, I mean, per class, and all that. That is for uncorrelated data only. [01:06:48] Okay? Yeah. [01:06:49] Okay, got it. [01:06:52] Uh, okay. So, uh, let's talk about data transformation. Here, we try to… [01:06:58] transform or, uh, change… not exactly change the data distribution, but [01:07:04] modify the data in such a way. [01:07:07] Uh, that it is more, uh… [01:07:10] more applicable, or it is better suited for analysis. [01:07:14] So, in this, there are many techniques, like data aggregation, normalization of the data, [01:07:20] Attribute construction, and so on and so forth. [01:07:23] And we'll talk about each of these as we go along. [01:07:27] So, first, um, transformation on the data is data normalization. [01:07:32] So, data normalization means… [01:07:36] that we are actually, uh, mapping [01:07:39] One range of the data to some other range. [01:07:42] For example, initially, our values were, uh, belonging to [01:07:47] Uh, the initial value was… min value A. [01:07:50] And max value, A, for the attribute A. [01:07:54] No, I want to map it to… [01:07:57] Newman A. [01:08:00] And new max A. [01:08:04] I want to do this. So, how can I do it? I can do it with min-max normalizing. [01:08:09] Where any value mu prime is given as mu minus min A, [01:08:14] divided by max minus min into… [01:08:18] into new max, A minus new min A plus new max, so… [01:08:22] Actually, here, what are we doing? Mapping one range of values to the other. For example, if I've got sales, [01:08:30] In INR. [01:08:32] And I've got sales in USD. [01:08:35] And I want to map them both to the same units. [01:08:39] Then, either I want to map INR to UST or UST to INR, then I can use min-max normalization. [01:08:46] To bring the values to the same range. [01:08:49] Where many Min, and max say designate the minimum and maximum values. [01:08:55] For the attribute A, [01:08:57] uh… initially, [01:09:00] And then Newman A and new max A. [01:09:02] are the values that are the new values. [01:09:05] Uh, that are to be used. [01:09:08] Uh, where we want to map them. [01:09:16] So, what is he, ma'am? [01:09:18] Uh, sorry, the new, you're saying? [01:09:20] No, what is V? [01:09:23] So, that is any value. [01:09:24] Okay. [01:09:25] So, the new value is given by nu prime, or V prime, whatever. This is the initial value. [01:09:32] So, we'll subtract the minimum value, then divide it by the original range, [01:09:37] Multiplied by the new range, and add the minimum value of the new range. So, in a way, what are we doing? [01:09:43] Where actually, uh, first removing the offset [01:09:46] of the original range. [01:09:48] From the given value, and then… [01:09:51] Mapping it to a new range? [01:09:53] And offsetting it with a new minimum. [01:09:55] This is what we are doing in min-max novelization. [01:09:59] Then we also have z-score normalization. [01:10:01] Where we have, uh, the value. [01:10:04] V, given us V', which is… [01:10:08] Uh, the value minus mean over the attribute, say, A, divided by standard deviation. [01:10:13] So, this is also another method. [01:10:15] Many times, marks are normalized using z-score normalization. [01:10:20] And, uh, this is not the entire list. We have many other methods also. [01:10:25] However, I'm just including some important ones. [01:10:31] So then, uh… [01:10:32] Ma'am, can we have one example for normalization? [01:10:37] Yeah, I think I… I do have some graphs. I'll show you. [01:10:42] Just, uh, if you'll wait for a little while, I'll be able to show you examples. [01:10:44] Okay. [01:10:46] on graphs, you know. [01:10:49] So then, uh, the application of minima access to convert one range to a new range. [01:10:57] And the application of z-score normalization is to compare [01:11:01] Uh, two scores. [01:11:04] Uh, that are coming from different normal distributions. [01:11:07] And, uh, we use z-score normalization in that case. [01:11:12] So, uh… [01:11:14] So, here's an example. Let's say we have the data as shown here, where we have X and Y as the two variables. [01:11:21] And we want to do min-max normalization as well as z-score normalization. [01:11:26] They already mentioned that G-score normalization is going to [01:11:30] Center the data, uh, around the mean. [01:11:34] or around zero. So, we are subtracting the… [01:11:37] original mean value from it. [01:11:39] So that what we are having is only the fluctuations around the mean. [01:11:44] So, uh, and then we scale it on the basis of the standard deviation. [01:11:50] Uh, this is how we get z-score normalization. So, let's, uh, see. [01:11:55] Uh, this is the data, original data. [01:11:58] With minimum and maximum values given as 1 and 20, [01:12:02] For X and Y, and for Y… for X, and accordingly for Y also. [01:12:08] Just a second, I don't know what's happening. [01:12:20] Now, uh, we, uh, perform min-max normalization. [01:12:24] And, uh… [01:12:35] So, the initial range of values is 1 to 20 and 1 to 25. [01:12:39] And we want to map these values to a new range. [01:12:43] Uh, which is from 5 to 25 for the X coordinates, and 10 to 40 for the Y coordinates. [01:12:50] So, the minimum was 1, we want to map it to 5, this one. [01:12:55] And maximum is 20, we want to map it to… [01:12:59] I don't know what's going wrong with my… [01:13:02] spend today. [01:13:07] Yeah, and we want to map it to 25, so 20, we want to map to 25. [01:13:11] Similarly, Y is 1, we want to map it to 10, 25, we want to map to 40. [01:13:17] This is how we want to perform the… [01:13:19] mapping. So, you can see here, this is the original data shown in the yellow color. [01:13:24] And it's starting with 1, and ending with 20. [01:13:28] For the x-axis, and for the y-axis, it starts with 1. [01:13:33] And ends with 25 here. [01:13:36] This is 25. [01:13:39] Now, we want the minimum value to be 5, so we are going to perform min-max normalization. [01:13:44] And this point is actually going to be mapped to… [01:13:47] The new minimum, which is 5. [01:13:50] So this is actually mapped to 5. [01:13:52] 5 on the x-axis, and 10 on the y-axis. So, this is the 10 value. [01:13:58] This is the phi value, this is the point. [01:14:00] So we start from here, and we end the data with [01:14:05] Uh, the new minimum value, uh, the maximum value is 40. [01:14:11] So, it was 25 initially. [01:14:12] Now, this has got mapped to 40 out here. [01:14:15] So this is the new, uh, mapped values we get. [01:14:19] Uh, from the given data. The thing to notice is that initially the range was 25 minus 1. [01:14:26] And now the range is 40 minus 10. [01:14:29] Okay, so this is how we perform min-max normalization. [01:14:33] If we have to perform z-score normalization on the same data, then [01:14:37] Oh, sorry, ma'am, to, uh, disturb you. Can you please go back to the previous slide? [01:14:42] Mm-hmm. [01:14:43] Can we see how this new mean X is getting calculated? [01:14:48] Yeah, so you just have to use this formula. [01:14:50] So you have, uh… [01:14:53] the value minus the minimum value divided by the original range, [01:14:57] into new range plus the new minimum. [01:15:00] For example, you have a point. [01:15:03] Uh, the initial point is X is equal to 1, Y is equal to 1. [01:15:07] Right? So, the new value will be… [01:15:11] The old value, you can see this formula, which is actually… [01:15:15] Uh, the minimum value. So, this new value will be… [01:15:19] Uh, if we talk about one… the X value, that will be one [01:15:24] Minus. So here, it is a… [01:15:28] given value, let's say I… you can just use this formula. So, let's say I talk about the point 2 and 3. [01:15:34] Okay? Then it will be the original value. [01:15:38] Minus the minimum value. [01:15:40] Which will be, uh, for this particular case, the original minimum value is 1, so we'll subtract 1. [01:15:48] divided by… [01:15:51] Just a second now, facing some issues with this. [01:15:54] divided by the original range. [01:15:57] What is the original range? [01:15:59] For X, it is 20 minus 1, so it is going to be 19. [01:16:06] I'm actually trying to write a 9 here. [01:16:10] Yeah. This multiplied by the new range, [01:16:13] The new range is going from… [01:16:16] We are having the… [01:16:20] the new range going from… [01:16:23] 5 to 25, so… [01:16:25] This will be the new range, which will be… [01:16:28] 25 minus 5. [01:16:30] plus the minimum value of the new range. So what we'll have here, 1 upon 19, [01:16:35] In 220, [01:16:37] plus 5, so it will be 5… [01:16:39] Five point, I mean, 5 plus 1 point something. [01:16:42] So, it will be around 6. [01:16:44] That is the point 2, 3. [01:16:47] So, this is the point 2, 3. [01:16:50] in the original data, here. [01:16:52] This is a point in correlational data. [01:16:53] But how should we decide the new Minix and new meanwhile? [01:16:59] Sorry, how do you decide what? [01:17:01] How should we decide new minimum X and. new maximum X. [01:17:04] So that is the range that you want to map the data into, right? [01:17:09] Suppose, uh, I have, uh, you know, the value… [01:17:13] Let's say I've got values for different variables, like, for example, I've got height. [01:17:18] Then I've got weight. [01:17:20] And maybe I've got age. [01:17:22] All of them will have different ranges. [01:17:24] Right? And, uh, I want to compare them. [01:17:28] So then, in that case, it is possible that I may want to bring them to some similar ranges, so that [01:17:34] Uh, because age may range, let's say, from 1 to 100. [01:17:39] And weight may range, again, [01:17:42] from, let's say, 10-200, or maybe 200. [01:17:46] And height may range from… [01:17:48] in feet from something, let's say, 2 feet to 7 feet, or 6.5. [01:17:55] So, what can happen is with these different ranges, some particular attribute [01:18:00] will become dominant. For example, weight or age. [01:18:04] So, to prevent that, we might want them to… [01:18:06] scale them to the same range. So that new range will be decided by you only. [01:18:13] Right? Or the user. [01:18:15] Whosoever wants to use, whatever, but what I'm showing you is how to do… [01:18:20] How to calculate when you have the new range with you. [01:18:25] Okay. So, I request all of you to please, uh… [01:18:26] Yeah, got it, ma'am. Thank you. [01:18:29] right on the chat, and I will, uh, so that it doesn't disrupt the flow of explanation, and I'll take questions in between. [01:18:36] I'll stop and take in case the chat is disabled, uh… [01:18:41] A request of insurance co-host to please enable. [01:18:43] the chart, so that people can write. [01:18:46] Uh, on the chat for any doubts. [01:18:51] Okay, so this is z-score normalization. Here, what we are doing is… [01:18:54] That we are bringing, uh, all the data into the range of 0 to [01:18:59] one, something like that. So… [01:19:01] So, you can see the data was spread over this much range, now it has… [01:19:06] being brought to this range. [01:19:08] So, this is how z-score normalization works. [01:19:12] Then, uh, the next thing that we will talk about is dimensionality reduction. It's a very, very important topic. [01:19:19] And there are many methods to do it. One important way is doing… using feature selection. [01:19:25] Another is to use feature construction. [01:19:28] So, what is feature selection? [01:19:30] Feature selection means we have certain attributes that are given in the data. [01:19:35] And we select some important attributes. [01:19:38] And you retain those, and just ignore the rest of them. [01:19:42] So, we are actually finding out, uh, important subset. [01:19:47] Uh, of attributes from the given set. [01:19:50] And using that only to represent the data, and ignoring the others. [01:19:55] This is feature selection. [01:19:57] So, we are doing a select operation, or we are doing a selection. [01:20:02] from the entire set. [01:20:04] Then, uh, comes feature construction or feature extraction. [01:20:08] So, in feature construction, what we do is that we… [01:20:12] Actually, form new features out of the given data. [01:20:16] So, how do we form new features that can be different methods? For example, uh, we might map the given points [01:20:24] into, uh, you know, some other… [01:20:27] coordinate plane completely. [01:20:29] So, we may form new features. For example, if we have text data, [01:20:33] then we can represent this text. [01:20:36] Uh, in some way. For example, by sets of keywords or something like that. So… [01:20:41] The original data was the text document. [01:20:44] And it got mapped to new features. [01:20:47] By some method. Okay, so feature construction and extraction, [01:20:51] means deriving completely new set of features from the Q1 data. [01:20:55] Now, what is dimensionality reduction? [01:20:58] Dimensionality reduction means we are reducing the number of attributes or the dimensions in the data. [01:21:06] So that the data becomes more easy to handle, to… [01:21:09] archive to access and all that. [01:21:12] So, when we are doing feature selection or feature construction, [01:21:17] We may be also doing dimensionality reduction. For example, if I've got some [01:21:22] 30 attributes with me. Out of these 30, I'm retaining only 5. [01:21:26] Uh, and doing a feature selection in which 5 gets selected and rest of the… [01:21:31] Uh, 25 are not so important, and they are ignored. [01:21:36] Then I'm reducing the dimensions from 30 to 5. [01:21:39] This is what is dimension reduction, and it's very important. [01:21:43] Because, uh, higher the dimensionality, it's more difficult to handle the data. [01:21:49] And it's also very difficult to visualize such high-dimensional data. [01:21:53] Um, and so on and so forth. [01:21:59] So, here's an example of feature selection. [01:22:01] Uh, we are assuming that, uh, we are building a decision tree. [01:22:06] How we'll actually build our decision tree, I will be showing you. [01:22:10] In the chapter on supervised learning. However, I've just, for illustration, [01:22:15] shown one here. So, the initial attribute sets are A1, A2, A3, A4, A5, A6. So, there are 6 attributes. [01:22:23] given to us, and we want to do a feature selection. [01:22:26] There can be many methods to perform feature selection. [01:22:30] And one of them is this, where we try to build a decision tree. [01:22:34] And all the attributes that are represented [01:22:38] On the tree, or are utilized in building the tree, are retained [01:22:42] And the rest of them are ignored. [01:22:45] Because when we build a decision tree, we… [01:22:49] actually use this to classify the data. [01:22:50] And if three attributes are enough to form… to arrive at the class labels, [01:22:55] Then the others are just redundant, and we can ignore. [01:22:59] So, this means that we can ignore A2. [01:23:02] Then we can ignore A3. [01:23:04] then we can ignore A5. So, we are left with… [01:23:08] A1, then A3… A4, and A6. So, we are left [01:23:15] Uh, with 50% of the attributes, and we have… [01:23:17] done attribute selection out here. [01:23:21] Uh, while the sanctity of the data is still getting maintained because the rest of the attributes are kind of redundant for us. [01:23:28] So we ignored them. There are many other methods also we'll be discussing them as we go along. [01:23:37] Then, uh, uh… [01:23:40] Then there are methods for data reduction. [01:23:43] Uh, so data reduction is another thing that is required if the data is very huge. [01:23:48] We somehow want to reduce it so that it becomes… [01:23:52] Uh, smaller, and it becomes easier to handle. [01:23:55] a huge amount of data could also be compute-intensive. [01:23:59] And if there can be some methods by which we can reduce the data, it will definitely help. [01:24:04] And one such method is numerosity reduction. [01:24:08] So, numerosity reduction. So, here we… numerosity means numbers. If we can somehow reduce [01:24:14] The numbers in the data, then we are talking about numerosity reduction. [01:24:20] So, how can we do that? [01:24:22] There are two methods. One is parametric methods, the others are non-parametric methods. [01:24:28] What are parametric methods? [01:24:31] Here, what we do is we take the entire data, that may be some… [01:24:34] Uh, let's say we have got some 1 lakh records with us. [01:24:37] And, uh, these are a huge number, and therefore, we… [01:24:42] somehow want to reduce the data, so what we do is we fit a model on that data, [01:24:46] And after fitting the model, we just, uh… and to fit the model, we do, uh… [01:24:51] parameter estimation that may be hyperparameters or other parameters that may require to model the data. [01:24:58] We find out those parameters, model the data, [01:25:01] And retain the model. [01:25:03] And discard. [01:25:05] Uh, the individual, the actual points. [01:25:09] So, we don't really keep the points, we just keep the model with us, because that's a good representation. [01:25:15] Assuming that's a good representation of the data, [01:25:18] And then just, uh… [01:25:21] user model and forget about the points. So, these are parametric methods, because here we have to assess the parameters [01:25:30] that are involved, uh, in making up a model. [01:25:35] Then we also have, uh… [01:25:38] Non-parametric methods. So, non-parametric methods are those where [01:25:42] We do not build any models and do not need to estimate any parameters of the models. [01:25:48] For example, clustering. [01:25:50] So, in clustering, we just form the clusters or groups, out of the given data. [01:25:55] So there are no parameters or hyperparameters that need to be estimated. [01:25:59] In such methods, all we need to do is just cluster the data, [01:26:04] Or if you are doing some histograms out of that data, we just form the histograms, but we are not… [01:26:09] Using any other hyperparameters or any other parameter estimation here. [01:26:14] Therefore, such methods are called [01:26:17] non-parametric methods, and [01:26:19] All these methods help to reduce the volume of the data. [01:26:22] For example, in the parametric methods, the volume gets represented by our single model, or a set of models. [01:26:29] In the non-parametric methods, the entire set of or volume of data gets [01:26:35] represented by some few clusters. [01:26:37] So, in both these cases, we are actually reducing the volume of the data. [01:26:43] So, we can, uh, also perform numerosity reduction, as I mentioned here. [01:26:49] Not just by clustering, but also forming histograms. Now, what are histograms? [01:26:54] Histograms are nothing but, again, they are similar to bins only. [01:26:59] And, uh, the method to form histogram is again similar to the methods for binning, which I've already told you. [01:27:04] That we could do binning by equivalent binning, or by equip frequency binning, and so on and so forth. [01:27:10] And similar methods can be used. [01:27:14] to form histograms. Often histograms are done on counts, or frequency-based histograms are very, very popular. [01:27:21] methods of data reduction, where we retain the histogram, and we just do away with the actual [01:27:29] volume of the data. [01:27:34] So this is a histogram that I had shown you already, where I'm… [01:27:39] I've done the histogram based on the [01:27:41] First alphabet of, uh, the… of your first name. [01:27:45] Uh, for all this group of students, and we find out that [01:27:49] This program for the alphabet S is having the maximum count in it, or frequency of values, followed by a [01:27:56] RM, and so on and so forth. We… [01:28:02] Okay, so before I, uh, proceed, any questions on what I just discussed? [01:28:08] Maybe you can write on the chat. [01:28:20] Okay, so what I'll do is I'll proceed and then invite questions once again. [01:28:25] Um, as we progress. [01:28:27] So now, another form of performing data reduction is by sampling. [01:28:33] So, what is sampling? Sampling of the data is [01:28:36] So, let's say we have got some… [01:28:38] millions of points with us. It's really very difficult. [01:28:43] to handle a huge amount of data. So, instead of having the entire set of millions of points, [01:28:47] We could have a good representative subset of this data, [01:28:52] And retain the subset rather than retaining all those millions of points. [01:28:56] Now, how can the subset be formed? [01:28:59] One way of forming the subset is doing simple random sampling. [01:29:04] So what we do, out of the given set of points, whatever the number is, let's say 1 million or 1 lakh, [01:29:09] We just randomly select some points. [01:29:13] And, uh, retain it. [01:29:15] As a representative subset of the given data. [01:29:19] Uh, so the… [01:29:23] Method looks very simple and nice. However, the primary problem with this simple random sampling is [01:29:29] That, let's say if the data distribution is skewed. [01:29:34] Let's say that in certain ranges of values are very less, [01:29:39] In some other ranges, the values are very high in number. [01:29:44] And when we do a random sampling, it is possible [01:29:47] that some of the lower-range values get lost because they never were selected. [01:29:52] in the random sample. [01:29:54] Right, so then in that case, the distribution of the data will… [01:29:58] get modified, and we don't want to modify the data distribution. [01:30:03] What we want to do is get a representative subset of the given data. [01:30:07] And reduce the volume of the data by doing sampling. [01:30:11] So then, uh, random, uh, sampling, simple random sampling, is a good, uh, and, uh, [01:30:18] a good method, which is intuitive, [01:30:21] Uh, and understandable. However, it doesn't work when the… [01:30:25] Data distribution is skewed. [01:30:27] So, in such cases, we can use something called stratified sampling. [01:30:33] So, stratified sampling means that, let's say, we have got [01:30:37] Uh, two classes in the data. [01:30:39] One class is majority class, where 80% of the samples belong to Class C1. [01:30:44] Let's say we have 80% of the samples. [01:30:49] Belonging to Class C1. [01:30:52] And the rest of the 20% samples belong to class C2. [01:30:56] Now, uh, when we are doing a stratified sample, assuming, let's say, we have got a total of 100 samples, [01:31:03] Then, the ratio in which the samples are 80s to 20. [01:31:08] Now, suppose I want to draw a stratified sampling. [01:31:13] And I want to reduce the data to 50%. [01:31:15] Let's say, now, I want only 50 samples. So, what I'll do… [01:31:20] I will pick samples from each of this group in the ratio. [01:31:24] of the size of that group. [01:31:26] For example, the size of this group is 0.8, and this is 0.2. [01:31:31] So, 40 per… uh, so 80% of the samples should come from this group, and 20% should come [01:31:37] from this group, so that the distribution, overall distribution, [01:31:40] of the class labels remains same. [01:31:43] So this means if I want to make 50 samples, I will draw 40 from here. [01:31:47] And 10 from here. So, in this case, the distribution remains the same because the [01:31:54] Uh, the… [01:31:56] Uh, the amount of Class C1 and C2 samples. [01:32:00] are similar to the original distribution, which is 80 to 20%. [01:32:06] So, uh, stratified sampling is used when we have skewed data distribution. Most of the real-world data is skewed in nature, as you all already know. [01:32:18] Then another method of, uh, doing, uh, data reduction. [01:32:23] is by, uh, doing, uh… [01:32:26] So, if we have, uh, here, the raw data, we are again illustrating, sorry, this is not… [01:32:32] So, clustering is also another method of doing, uh… [01:32:36] sampling. So this is the raw data, where we have got 3 [01:32:40] Uh, classes or clusters available in the data. [01:32:44] So, if we do a stratified sample, [01:32:46] Then, uh… [01:32:49] samples, uh, will be drawn from each of these classes in the same ratio as they are available in the original data. [01:32:55] And therefore, you will see, let's say if this was class C1, this is C2, [01:33:00] And this is C3. Then, in the final, uh, data also, after doing the random sampling, [01:33:07] We still have all these classes in the data. [01:33:10] And none of the classes will go missing. [01:33:12] Uh, because in random sampling, we just randomly pick, but in stratified sampling, [01:33:18] We pick the points in the ratio in which they are available in the original data. [01:33:27] Okay, any questions, anyone, uh, on whatever we have discussed? [01:33:35] Yeah. [01:33:36] Oh, ma'am, I have a question here. So, like, we are doing this data reduction, right? So, what are the activities where data reduction. [01:33:42] will come handy, because I was under the assumption the more the later, better the machine learning model. [01:33:47] But, uh, if we are reducing the data, then we are reducing the number of samples on which. [01:33:53] We are training the model as well, right, for decision tree that you mentioned, I understand that some of the attributes are not having any. [01:34:01] details, or any… Implication on the machine learning model. [01:34:06] But how this will help? [01:34:08] Okay. Uh, definitely, we generally say that, uh, the more the data, the better will be the learning. [01:34:17] That is true. Why? Because we want to have [01:34:20] data samples that represent each and every use case. [01:34:25] Uh, in the given data. [01:34:27] Right? So that's what, uh, what I'm saying is, let's say that… [01:34:31] We are having some, you know, the Aadhaar, uh, or let's not talk about Aadhar. We are having some data [01:34:39] Say tweets, and there are millions of tweets from the given dataset. [01:34:43] So it's, uh… so if we have to analyze the tweets to find out what the topic of discussion. [01:34:49] It's going to take a huge amount of compute. [01:34:53] Or probably the algorithm will keep running for days, and you won't be able to get your result in time. [01:34:58] So, in all such cases, [01:35:00] We can, uh, perform data reduction. [01:35:03] In such a way that the overall distribution of the data in the classes that is there [01:35:10] doesn't change. So, uh, here in this case, we are just making a smaller sample. [01:35:17] To make the data more easily handy, uh, you know, easy for handling. [01:35:23] So, if we have a sample of the data, then we can use it for different purposes, say clustering, classification, [01:35:30] or any other purpose. So, this is where data reduction comes into. [01:35:35] picture. It is not just this. Suppose we want to do exploratory data analysis, or EDA, on the data. [01:35:43] So, it will be very difficult to do EDA, let's say, on millions of points. So, we may want to reduce the data and get a set which is a good representative of the original set. [01:35:54] So, we can easily perform EDA on it. So, here also, data reduction is handy. [01:35:58] So, there are many such cases where data reduction is important. [01:36:04] And it is handy for us. So, in all such cases, we perform. [01:36:07] Uh, data reduction by different methods. Sampling is one of them. [01:36:12] Is that okay, Ankles? [01:36:13] So, yeah, ma'am, so we are saying that the time to result is what we are prioritizing over accuracy. [01:36:18] Uh, so that's why we are going with data reduction. Second thing is the exploratory data analysis would be easier to do. [01:36:25] Uh, over the smaller data set. And that's why we are using it, right? [01:36:31] Yeah, these are two reasons. There are many other reasons also, and uh… [01:36:36] Uh, definitely compute time will be less if… [01:36:39] Uh, the amount of data will be the same. However, the data distribution is still the same. The most important thing is [01:36:48] that when we are drawing a sample of the [01:36:50] given data, it should not happen that the sample's distribution is different from the original data distribution. [01:36:57] So that we must not change. That is why we go for stratified sample, where [01:37:01] Uh, all the classes are retained in the data that is there in the original data. [01:37:07] Yeah. [01:37:10] Uh, so I can see, uh… [01:37:11] Understood. [01:37:13] that, uh, Sivanch, you have some question, can you please ask what you want to know? [01:37:18] Yes, sure. So, I was just thinking about the data sets of form and max normalization. Let's say if I have data, like, 10 to. [01:37:29] From 2 to 10, uh, 20. Uh, and uh… Uh, let's say my training set is having this done as a max value and 2 as a main value. [01:37:41] Uh, and the test value has just provided to be as, uh, 15. [01:37:45] By the model. And… and when we are trying to calculate the… The min-max normalization value from that formula which you have showed. [01:37:57] Yeah. [01:37:58] It gives them a fractional number. So, uh, it's just acceptable, uh, for… to apply this approach, or it will not work in such type of data, certainly. [01:38:06] Uh, you're talking about datasets where you get fractions and not numbers. Is that what you're saying? [01:38:14] Yeah, with the formula, you get fractions. [01:38:15] Yeah, with the formula, I mean, uh… The new value will… yes. [01:38:18] Yeah, see, if you are having one range, [01:38:20] Let's say your initial range is, uh, let's say 10 to… [01:38:25] say, 50, and you want to map these values to a range like [01:38:30] Say, 1, 2, say, 10. [01:38:32] Let's say. Then, definitely, [01:38:35] Something that was 10-year is coming to 1. [01:38:38] Then, it will be a fraction only, no? [01:38:40] So, having fractions is not a problem. [01:38:43] Uh, as long as we are allowing any kind of values. Suppose we… [01:38:49] don't want, uh, values in points. [01:38:52] Uh, as per our algorithm, then it's different. Otherwise, it's okay to do that, because… [01:38:58] We ourselves decide what is the new range. If the new range is compressed, [01:39:05] Then some of the larger values which are there in the original range will come out in points only. [01:39:10] So that's fine. [01:39:14] Okay. [01:39:15] Okay. And one question I have related to binning. [01:39:16] Mm-hmm. [01:39:17] Uh, what I understood is we are… today we have covered about numerical. [01:39:23] techniques of those spinning equally width and equal frequency. [01:39:29] Mm-hmm. [01:39:30] So, in the… EQ frequency case of what happened if the dataset contains many repeated value? [01:39:38] At the VIN boundaries. [01:39:41] Uh, okay, so you're talking about equal frequency binning, right? [01:39:45] And… and we are talking about values that are repeatedly coming. [01:39:46] Great. [01:39:50] So that's okay. Accordingly, see, in case of equip frequency binning, [01:39:51] Hello. [01:39:55] The bins may not be of equal width. [01:39:57] Only the frequency remains wants… is to be remaining the same. [01:39:58] Mm-hmm. [01:40:02] So, let's say in the range of… [01:40:04] Uh, let's say 1 to 10, we are having lots of values, say 999 coming again and again. [01:40:10] Then, uh, what are we going to do? [01:40:13] So, the frequency, or the count that we want to keep on the VIN, [01:40:18] will depend on the average values we are expecting per bin. [01:40:23] Let's say we already know, we have the data with us, and we want to bin it, and we know that [01:40:28] a lot of 9 values are coming in here. So, accordingly, we'll bin the data. [01:40:34] So that the rest of the bins will have also a similar [01:40:38] uh… distribution. However, uh, we have to keep the count. I mean, if we are having 10 values, [01:40:45] Uh, which are, uh, 10 numbers, which are, you know, having the same value. It has to be put in the same bin only. [01:40:52] We cannot ignore them. [01:40:55] They will account for the frequency. [01:40:56] Um… [01:40:59] Understood, yeah. Oh, one last question, uh… Just, uh, having thought is, can we consider the negative values in this binning method, uh, changes? [01:41:10] Or it's for… Positive natural numbers. [01:41:15] You're gonna have any kind of range. [01:41:17] The range will accordingly work [01:41:18] Okay. [01:41:19] Because what we are doing, we are subtracting the maximum minus the minimum value. [01:41:24] If the minimum is negative, it will get added to the range, so that's no problem. [01:41:31] So, you could have values going from minus 1 to 10 also in the new range. [01:41:32] Okay, thank you. [01:41:35] So, because minus 1 to 10 means it is a set of 11 values, so it will be 10 minus minus 1, which will be 11. [01:41:42] So that's okay. [01:41:44] Okay, thank you. [01:41:48] So then, uh, we talked about, so far, different methods for data reduction, like sampling, like numerosity reduction, [01:41:57] Uh, using, um… [01:41:59] parametric methods, non-parametric methods, and so on and so forth. [01:42:04] So, another way to do data reduction is via data discretization. What is data discretization? Again, it's similar to binning only. [01:42:12] I can explain it with the help of binarization. Let's say that [01:42:16] Uh, you know, I have got, uh… [01:42:18] I've got to represent the data in form of two numbers. [01:42:22] Let's say I have a range of values from 0 to 10. [01:42:25] But I want to represent this using binary values. [01:42:29] So, I can do this in a couple of ways. One way could be [01:42:33] Anything from 0 up to 5. [01:42:37] could be 0, and anything, uh, greater than 5, up to [01:42:42] 10 could be 1. [01:42:44] So here, we are actually binarizing the data from 0 to 10 to 0 and 1, 2 values. [01:42:50] So, we all know that computers work on binary data, and whatever is the actual [01:42:55] value, it is represented in form of binary numbers. So, binarization is a simple example. [01:43:01] Or a special case of discretization. [01:43:03] However, discretization may not just be binarization. It could be anarization. [01:43:09] Say, for example, I have the range 0 to 15, [01:43:13] And I want to represent it in 3 sets, in 3 discrete value sets. So what I could do, I could have 0 up to 5, [01:43:21] As one said, then greater than 5 up to [01:43:26] 10. As another set. [01:43:30] greater than… and then greater than 10, up to 15 in another set. This could be… [01:43:36] Uh, that discrete value represented by 0, 1, and 2. [01:43:41] So, here, binning is being used, or we also call it discretization. [01:43:46] Where we are representing this, uh, infinite number of values in each of these range, [01:43:52] by a finite value, or a discrete value. [01:43:56] Which is, like, 0, 1, 2, something. [01:43:58] And this can also be said to be a label that we are assigning to all the points in this range. So this is called discretization. [01:44:06] of the given data, and we use it to convert a set of continuous valued, uh… [01:44:12] attribute. [01:44:15] to a categorical attribute. Let's say I've got, uh, say, high, uh, age. [01:44:25] Right? So, let's say age ranges from 1 to 100. [01:44:28] So, there could be infinite points in this range. [01:44:31] But I want to represent a data by a finite set so that I can do some calculations over it. [01:44:37] Having an infinite value set is very difficult to handle, [01:44:41] So what I can do, I can discretize it into… [01:44:44] 10, uh, bins, or 10 ranges, [01:44:47] say, 1 to 10, then greater than 10, up to 20, greater than 20, up to 30, like that, so I'll have [01:44:54] 10 bins with me, like, 1, 2, 3, till 10, or 0, up to 9. [01:45:00] So, this is nothing but data discretization. [01:45:03] Where we are dividing the data. [01:45:06] Uh, in such a way that certain set of data falls in some bin, and we call it [01:45:12] some… and assign it some discrete value. [01:45:14] And we call this data discretization. [01:45:18] Then there's something called concept hierarchy. [01:45:22] So, what is a concept hierarchy? In case of a concept hierarchy, we are having… [01:45:27] Uh, some implicit hierarchy in the data. For example, if we talk about location, [01:45:33] So, location is an attribute. [01:45:36] Now, the location could be over a country, it could be over state, it could be over city, it could be our street. [01:45:41] And so on and so forth. So there's an inherent… [01:45:44] hierarchy involved [01:45:46] In the attribute-like location, and we can use this [01:45:51] a hierarchy to represent, uh, this particular value called location. [01:45:57] Suppose I want a very fine-grained representation. I can go up to the street level, [01:46:03] If I have, uh… [01:46:05] More courses are on a broader representation, I can go up to city, [01:46:08] or up to state, up to country, and so on and so forth. [01:46:11] So, we can have different levels in the hierarchy. [01:46:15] And we can utilize those values from corresponding level in the hierarchy that we want. [01:46:22] If you want lots of values, we go to the… [01:46:25] Street level, if we got, uh, somewhat less number of values, we go to city level, still lesser state, still lesser country, like that. [01:46:34] So, uh, this is how we utilize the concept hierarchy. [01:46:40] So… [01:46:42] Okay. Just a second… [01:46:48] So I think there's some question that I could see here on… [01:46:52] Um, okay. [01:46:55] So, I think I could see a question that, uh… [01:46:58] Uh, about parametric reduction. [01:47:01] Okay, so Srinam, parametric reduction means… [01:47:05] that we are using model. [01:47:07] We are using a model to represent the data. Obviously, when we build a model, we will need to represent the model with the help of some parameters. [01:47:17] So that is parametric representation of the data. [01:47:19] Okay? Say a model will need some… for example, if I have a model like Y is equal to [01:47:25] MX plus C. [01:47:27] Suppose our model of this kind. [01:47:30] Where this is a linear model. [01:47:32] So then, the parameters that I'll need to estimate are the slope M and the constant C. [01:47:38] So, no matter what we use, [01:47:41] to, uh, designate [01:47:43] Uh, you know, a given set of points, [01:47:46] What we will use are some parameters. So, this is the parametric method. [01:47:52] of data representation. Is it okay? [01:47:59] Yeah, any other questions, anyone? [01:48:00] Ah yes, ma'am, got it. Thank you. [01:48:11] So, uh… [01:48:13] Uh, if there are any questions, then I'll proceed further on one or two topics before we wrap up for today. [01:48:21] Okay, I can see… [01:48:28] Okay. [01:48:35] So, the next topic will be on similarity and dissimilarity, and since it's a very different topic, [01:48:44] I like to take it in a new, uh, class. However, I'll just introduce what we mean by similarity or dissimilarity. [01:48:51] So the amount of, uh, uh, alikeness [01:48:55] Across points is said to be the similarity in the points. [01:49:00] Those points, which are very similar, [01:49:02] or, uh, are said to be very alike, or those points which are… [01:49:06] Uh, like each other are said to be similar. [01:49:10] And usually, similarity would be represented in a range, say, 0 to 1, where 0 means no similarity. [01:49:16] And 1 means maximum similarity, where… [01:49:20] Uh, the one data element is equal to the other. [01:49:23] The converse of similarities, dissimilarity, [01:49:28] Where we try to measure how different the data elements are with respect to each other. [01:49:33] So, uh, if we have a high amount of dissimilarity, [01:49:37] Then, uh, this means that… [01:49:39] The data elements, uh, being, uh, discussed are very different. [01:49:45] If you have a low amount of dissimilarity, means they are similar to each other. [01:49:49] So, minimum value of this similarity could be 0. When dissimilarity, 0 means the data elements are exactly equal. [01:49:58] Or, uh, they are exactly alike each other. [01:50:00] And if we have a dissimilarity value equal to 1, [01:50:03] means that the data elements are very different, almost reverse of each other. [01:50:09] So, uh, this is what similarity and dissimilarity mean. [01:50:13] Sometimes, we also represent similarity by proximity. [01:50:17] So, proximity is a generic term that could be used to represent similarity or dissimilarity. [01:50:24] So, when we say proximity of a point, we actually mean the distance [01:50:28] or the closeness of a point with respect to some other point. [01:50:32] So, uh, we could represent the similarity using proximity. [01:50:37] Um, uh, like that. [01:50:39] So, uh… so then we'll… in the next turn, we'll be talking about different measures. [01:50:45] about similarity and how to find out the similarity between points, or how to find out the dissimilarity between points. [01:50:53] So, uh, as of now, I'll just, uh… [01:50:57] stop here and invite any questions, if there are. [01:51:03] Ma'am, can we say that similarity, dissimilarity are same as relation and correlation, which we do in heat map? [01:51:13] Um, correlation often means co-occurrence. [01:51:17] or correlation can mean, uh, some kind of relationship [01:51:22] Uh, between the two, similarity may not be correlation always. Similarity means the… [01:51:28] Okay. [01:51:29] Um, the closeness between the two points. For example, I can say that [01:51:37] Uh, this cloth material is similar to the other one. [01:51:41] So, the two are alike. [01:51:43] each other. Uh, right? But if I say co-occurrence, or if I talk about correlation, [01:51:44] Mm-hmm. [01:51:50] Then I can say that [01:51:52] Uh, material A may be correlated [01:51:55] with Material B. For example, Material A, short, may be correlated with Material B, trouser. [01:52:02] So, that is correlation. [01:52:03] Okay. Got it. [01:52:08] Any other questions, anyone? [01:52:16] Okay, I think there's a question on whether we'll be having any hands-on, uh, covered for whatever we have covered today, yes? [01:52:22] You'll have hands-on on whatever I teach, not just today, always throughout the course. [01:52:28] Whatever I teach, the theory will be followed by a hands-on session. [01:52:32] So, we will be having hands-on, either starting next week or next to next week. [01:52:37] So, either I'll cover some more theory next week, and then start with the hands-on, or whatever I'll, uh… [01:52:43] have it announced on the LMS. [01:52:45] Uh, so far, uh, we have used Colab. [01:52:48] Uh, I think it will be… [01:52:51] a link for collab will be available on your LMS, uh, and I'll also ask, uh, [01:52:57] The FutureN's team to provide one if it is not already provided. [01:53:00] So, you can use Colabs separately also. [01:53:03] Uh, or whatever way you want to do the coding. But for us, we'll be using collab so that it's a single [01:53:10] environment, uh, used by everyone, and you can… [01:53:14] run and see the results, similar to what we are doing here. [01:53:19] So, we will definitely have hands-on. [01:53:21] Okay. [01:53:25] So, uh… [01:53:28] Any other questions? [01:53:30] Uh, anyone? [01:53:34] Uh, definitely, I'll suggest those who are familiar and those who are new. [01:53:37] All of you to, uh, just have a recap on whatever I've thought. [01:53:42] Uh, if you will, uh, keep reading whatever I've taught, [01:53:46] Or it will be good because you'll remember. [01:53:48] And everything will… is actually sequenced out in the sense that what you're learning today will be used tomorrow. [01:53:54] And so on and so forth. So right now, I'm… [01:53:57] Talking about data preprocessing, it will be used throughout the… [01:54:01] course, because whenever we have data, it will require pre-processing. So, you need to know these concepts. [01:54:08] If you are knowing already, it's good. If you are not knowing, and if you are new to the field, [01:54:13] Uh, then you should, uh, actually go through the material that is provided by [01:54:18] future rents are by me, the study material. [01:54:21] You can also consult a net and find out some free… [01:54:25] 30 material for any particular topic. [01:54:28] And familiarize yourself with the topic. [01:54:34] Anyone, anything else? [01:54:42] Yeah, so I think there is some question about the detailed syllabus that, uh… [01:54:47] So I had shown you the syllabus. You will be provided with that same, uh, syllabus. [01:54:53] Uh, what I had shown you, only thing is that, uh… [01:54:56] It is just, uh, under preparation. That is, I've shown you the topics, but uh… [01:55:01] We are just associating the weeks and all, so you should… hopefully by the… [01:55:06] Next week, you'll get the entire schedule. So, don't worry, we will not give you a partial one. [01:55:12] We'll give you the entire set of topics, week by week, how they'll be followed. [01:55:18] Okay? [01:55:20] Okay, Aditya, you have a question? [01:55:22] Uh, yes, ma'am. It's not a question, it's a concern. It's not regarding this class. [01:55:28] Mm-hmm. [01:55:29] It's for the batch manager. Uh, I've sent a couple of emails more than a week ago, and I've got no reply on that. Batch Manager, can you please look into it? [01:55:37] I'll send you a DM also in the chat, but I don't know if you've read it, I've not gotten any reply as of now. [01:55:44] Uh, so, uh, I think the future is co-host, if you can please take a note of the concern of Aditya. [01:55:52] Uh, I hope, Atul, you are there. [01:55:58] So, actually, uh… [01:56:02] Um, so usually we have Simran, who takes care of the sessions. [01:56:07] So, today, she's not here. [01:56:10] And we were supposed to have Atul in her place. I'm sure Atul must be there. Atul, can you please, uh, respond? [01:56:22] Also, I'm not sure about, uh… [01:56:28] what I thought was… was Atul would be there. Anyhow, uh… [01:56:37] GenAI Batch 2 Manager, could you please respond? [01:56:47] Yeah, I've tried pinging, uh, him or her in the chat, but I've not gotten any response from him. [01:56:53] So, yeah, so actually today, Simran is away. Probably you can write to Simran. [01:56:59] Uh, about the concerns, uh, on the email ID provided. [01:57:03] Uh, and I'll also, uh, [01:57:07] So I'll request, uh… [01:57:09] Yeah, we have only one email ID, ma'am, uh, as far as I know, and uh… [01:57:14] They're very unresponsive on that email. [01:57:20] Okay. [01:57:21] No, no, no. Actually, actually, they had given one more email for the Simbrown one, but the Simbran email address. I'm always getting mail delivery only. [01:57:32] Okay. [01:57:33] I think I've saved with me as well, Aditya. So, I think I… [01:57:37] So, I'll just, uh… [01:57:40] take a note of, uh… [01:57:43] Uh, this thing, and I'll try. I hope the co-host here is also… [01:57:49] hearing. [01:57:52] So, I'll request Simran to… [01:57:55] uh… respond. [01:57:57] under concern that you have raised about the management. [01:58:02] Uh, okay. [01:58:03] Yeah, if you can, uh, put in a word, ma'am, I think that would be great help for us. [01:58:08] Uh, yeah, I'm doing it right away, I'm just texting Simran. [01:58:12] Uh, she's away today, otherwise she would be here. [01:58:17] Uh, okay, so I'm just doing that. I've done it already. [01:58:22] So, yeah. [01:58:24] So, anything else before we wrap up for today? [01:58:32] Any other concerns, questions? [01:58:37] Ma'am, again, there's a question regarding attenders on LMS. [01:58:41] There is a tenant section, so again, this question will be for Simbran. [01:58:45] Once you join next time. [01:58:47] Yeah, so you can, uh, have, uh, take a note, and then next time she'll be there. [01:58:53] So, you can just, uh, please find out from her. So, she'll be the best person to answer. [01:58:59] Thanks. [01:59:00] Sure. Thank you, ma'am. [01:59:02] Gitu, do you have a question? [01:59:03] Yeah, ma'am. Uh, from, uh, dimensional reproduction, you have explained one feature selection and feature discussion, right? [01:59:13] Yes. [01:59:14] Can you explain a brief that which is a selection part? [01:59:17] Just show. So, uh, let's say we have a data in which there are some 15 attributes given to us. [01:59:27] But sometimes it may happen that not all 15 may be useful for our analysis purpose. [01:59:33] And we may want to retain [01:59:35] say, out of 15, some… [01:59:37] 10 of them, or 9 of them. [01:59:40] So, that is what is feature selection about, that we try to choose the relevant features. [01:59:45] And retain those, and uh… [01:59:48] The rest of them we ignore for our further analysis. So, we are actually selecting a subset [01:59:54] of the features from the given full set. [01:59:57] This is what is feature selection. [01:59:59] And then there are different ways to perform feature selection. [02:00:02] For example, forming a decision tree is one of them. [02:00:07] Is it okay, cheetu? [02:00:14] Okay, there are no further questions or comments, then we can break for today. I will be passing the slides to… [02:00:23] Same done, and hopefully she'll upload on the LMS. She's away today, uh, so I'm not sure. [02:00:28] But I've also shared, uh, yesterday's slides with her. [02:00:33] So, hopefully, you'll get it today, but… [02:00:36] If she's away, probably by tomorrow, so maybe I'll request you to bear for this time. [02:00:40] However, she's quite prompt, and immediately we upload the materials. I mean, I just pass it on to her, and… [02:00:47] She uploads on the items. [02:00:51] I'll also try to share in future classes a set of some [02:00:56] questions, maybe some, uh, few multiple-choice questions, which are not graded. [02:01:02] But those are exercise questions. If you wish, you can just attempt them and see what you have learned. [02:01:07] Okay? [02:01:10] So, uh, with that, I'll just wrap up today's class. Thank you all, have a great Sunday. Thank you. [02:01:16] Bye-bye. [02:01:17] Yeah, it's okay. [02:01:18] Thanks so… I'm spelled. [02:01:19] Thank you. [02:01:20] Thank you. [02:01:21] Thanks, man, bye. [02:01:22] Thank you.