# 02 2025-11-30 Hands On Supervised Learning course: Module 2 — Machine Learning Algorithms module: Module-2-Machine-Learning-Algorithms date: 2025-11-30 type: transcript video_url: https://personal-learn.armco.dev/files/_Recordings/Module-2-Machine-Learning-Algorithms/02_2025-11-30_Hands_On_Supervised_Learning.mp4 --- [00:14:22] aditya shrivastava: Good morning, all. [00:14:41] Durga Toshniwal: A very good morning to all of you, and welcome to today's session. So, we'll continue with the… what we had been discussing yesterday, that is decision tree induction. [00:14:51] Durga Toshniwal: So let me share my slides, and then we'll go on from there. Just allow me a second. [00:15:19] Durga Toshniwal: I'll just a moment. [00:15:32] Durga Toshniwal: Okay, so I hope my screen is visible, and I'm audible. [00:15:37] Durga Toshniwal: So, we'll just start with a quick recap of what we did yesterday, and then I'll start with the [00:15:46] Durga Toshniwal: New bullshit. [00:15:49] Durga Toshniwal: So, yesterday we discussed about what classification is. So, classification primarily is a two-step process. So, one involves building up the model, and then the other step is, [00:16:01] Durga Toshniwal: Testing the performance using the test data. [00:16:04] Durga Toshniwal: And then, once the model is built, then it is used for making predictions for future unseen records, where unseen record is a record whose class label is not known, and we want to predict it using the classifier model that we have built. [00:16:20] Durga Toshniwal: So, in the process of classification, we always are given a special data, which is called the training data. [00:16:29] Durga Toshniwal: Training data is special because, it has the class labels given to, along with the attribute values. We also have the class labels given to us as part of the data. [00:16:40] Durga Toshniwal: So we can think of this data as having a set of attributes X. The attributes are also called independent variables, inputs or predictors, whereas the output is the class label, and it is the dependent variable. [00:16:56] Durga Toshniwal: And, the process of building a classifier is actually mapping the input to the output. [00:17:04] Durga Toshniwal: In the best possible fashion, so that the output is… [00:17:08] Durga Toshniwal: Almost, equal to the input, where input is the given class labels, and the attributes, of course. [00:17:17] Durga Toshniwal: So, the process of classification is finding out the function F, just a second. [00:17:28] Durga Toshniwal: Yeah. [00:17:29] Durga Toshniwal: So the process of classification involves finding out this function f. [00:17:35] Durga Toshniwal: That we had, discussed yesterday. [00:17:38] Durga Toshniwal: Where, input X is the set of attributes that we are having. [00:17:43] Durga Toshniwal: Or the predictors or independent variables that we are having. And… [00:17:47] Durga Toshniwal: Y is the class label that we want to predict, and to predict this, we need to find out the function f, which is nothing but learning of the classifier. It could be linear, nonlinear function, and all that. [00:18:00] Durga Toshniwal: So, then I explained to you how the classifier works, and so, once the classifier model is learned. [00:18:10] Durga Toshniwal: Then, whenever an unseen data is presented to it, then we can find out the class lady. [00:18:18] Durga Toshniwal: Then classification could be a binary classification problem, or an NRA classification problem, where binary means two class labels, and an NRA means multiple class labels. [00:18:30] Durga Toshniwal: Then we also discussed that we require ample amount of training data to be able to have the classifier learn all the cases that it is supposed to predict in future. So, if our training data is such that some of the use cases that the model is supposed to predict is missing. [00:18:50] Durga Toshniwal: from the data, then those cases will not be learned by the classifier. As a result, if on the unseen data, such kind of cases are encountered, then the classifier will not know what to do, and it will fail. [00:19:05] Durga Toshniwal: So to explain this, I've given an example of a very simple, linear function, Y is equal to mx plus c, which is a straight line having slope n and intercept C. [00:19:18] Durga Toshniwal: And, if we have just 1.X1Y1 given to us, then we will not be able to find out MNC. [00:19:25] Durga Toshniwal: So, even when we have just one independent variable, X, [00:19:32] Durga Toshniwal: Then also, we at least need 2 points to be able to find out, correctly the model. [00:19:38] Durga Toshniwal: That is, the values of M and C. So this means that some minimum number of points are essential to finding out the parameters of the model. [00:19:48] Durga Toshniwal: And therefore, the training data must be adequate, and it must also contain all use cases which the model is supposed to see. [00:19:56] Durga Toshniwal: Let me give you an example, which I had not given yesterday. Let's say that our model is supposed to predict some values which are anomalous, or it is supposed to predict some, unusual, you know, patterns or unusual events. [00:20:14] Durga Toshniwal: If our model is supposed to do that. [00:20:19] Durga Toshniwal: Let's say if the model is being built to predict credit card fraud. [00:20:25] Durga Toshniwal: But if the training data does not contain any use case representing fraudulent activity, then the model will never be able to learn the patterns [00:20:35] Durga Toshniwal: pertaining to fraudulent activity. As a result, it will only be able to predict the legitimate transactions, and it will not learn how to predict fraudulent ones, and hence it will fail whenever fraudulent activity comes into picture. [00:20:54] Durga Toshniwal: Then, we started out discussing how a decision tree is built, which is one of the most popular methods. [00:21:02] Durga Toshniwal: So, in a decision tree, we have something called a root node. So, the root is very similar to the real-world root of a tree. However, the difference is, in the real world, the root is inside the soil. [00:21:16] Durga Toshniwal: That is, it is the lowest level, entity in a tree. [00:21:21] Durga Toshniwal: However, in a decision tree, the root is the topmost node. [00:21:25] Durga Toshniwal: To start with. [00:21:27] Durga Toshniwal: Then, we have internal nodes. Internal nodes are those that have one input edge going to it, but the output edges could be two or multiple. [00:21:40] Durga Toshniwal: And then we have the leaf nodes. Leaf nodes are special notes because they have, input, they have edges which are going. [00:21:49] Durga Toshniwal: inside the node. However, there are no outgoing edges from the leaf node, and they are used to signify the class label. That's why they are called the leaf nodes. So, they are similar to the leaves in a real tree. [00:22:08] Durga Toshniwal: Because the tree, the boundary of the tree is nothing but its leaves. Similarly, in this tree, decision tree also, which is an upside-down kind of a real tree, with root at the top and leaves at the bottom. [00:22:22] Durga Toshniwal: Beyond the leaves, the tree doesn't grow here also. [00:22:26] Durga Toshniwal: And the trees designate the class labels. So, I'd explained, to you building of a tree for which, we will need always categorical attributes. So, if our attributes are continuous, like this taxable income. [00:22:40] Durga Toshniwal: Then we'll need to actually discretize it. I already discussed the methods of discretization earlier. [00:22:46] Durga Toshniwal: And then, we choose attribute by attribute to do the split, and the objective is that when we are doing a split, what we obtain is a subset of the training samples, which have a clearer majority of the class labels. [00:23:08] Durga Toshniwal: So, we also discussed that, given a set of attributes and a training data, there could be multiple trees that could be made based on which attributes we choose. [00:23:21] Durga Toshniwal: Of course, there are methods to decide which attribute would be the best to choose at any point in time, which I'll be discussing today. [00:23:29] Durga Toshniwal: So, I already discussed this. [00:23:33] Durga Toshniwal: then, we also discussed, [00:23:38] Durga Toshniwal: how we actually stop growing our tree. So, the tree… we, stopped growing the tree when the leaf node shows a clear majority. [00:23:50] Durga Toshniwal: Of the class labels of samples that are going to the leaf node, or there are no samples left to classify. [00:23:59] Durga Toshniwal: So now, there are many algorithms that actually, many, real-world, algorithms. [00:24:07] Durga Toshniwal: That involved decision pre-induction. [00:24:11] Durga Toshniwal: Such as the Hunts algorithm, which we discussed yesterday. Then there are some other commercial packages containing CART, ID3, C4.5, Slick, Sprint, all these are nothing but variants of decision-free. [00:24:24] Durga Toshniwal: So, let's start with the general structure of Hunt's algorithm. So, yesterday, I discussed that already. [00:24:32] Durga Toshniwal: So, we start out with, all the training samples incident at the root node. [00:24:39] Durga Toshniwal: Right, so we start out like this. [00:24:41] Durga Toshniwal: So we start out with the root node, this is the root, and all 10 samples will be incident here, because all 10 of them do not have the same class label, so we'll go ahead and split the data using some attribute. [00:24:56] Durga Toshniwal: So, for example, we chose refund. So, if we choose refund as the first one, then we'll go ahead and use the values in the refund, like refund yes or no. [00:25:08] Durga Toshniwal: And based on that, we will do the split. And we had identified that by using refund, 3 of the samples would go towards the left side, and 7 of them would go towards the right side, and so on and so forth, so we'll go on using the attributes till they are available, or till the tree has been built. [00:25:28] Durga Toshniwal: So, these were the different steps that we will… that I had shown to you to build a tree. [00:25:36] Durga Toshniwal: Then there were certain constraints that we need to satisfy if we are to build a tree on the given data. [00:25:44] Durga Toshniwal: And the most important one is that a tree can only be built when the records have unique value pair combinations for the attributes. So, if there are two set of records that have the same values that I had shown yesterday. [00:26:00] Durga Toshniwal: Sorry, these were the two. [00:26:05] Durga Toshniwal: These were the two. So… [00:26:07] Durga Toshniwal: The values of the attributes are just the same. However, the class labels are different. So, what does it mean? [00:26:14] Durga Toshniwal: It means that probably there was some other attribute, either there was some other attribute that should have been measured. [00:26:22] Durga Toshniwal: And it's not available in the given data. [00:26:25] Durga Toshniwal: Or, or we need some other, so, because the creating the data may not always be in our hand, many times we are just given the data, so the other option is… [00:26:37] Durga Toshniwal: That, we use some other classifi- classification algorithm, and we don't use decision tree, because a decision tree cannot [00:26:47] Durga Toshniwal: We built, in such cases where the records may have the same values for the attributes, yet have different class labels. [00:26:59] Durga Toshniwal: Then there were other conditions, for example, [00:27:04] Durga Toshniwal: If we are building a tree, it is possible that some particular leaf node may be empty at that point in time. [00:27:13] Durga Toshniwal: Still, it's a good idea, and that may be because none of the training, records satisfy those conditions. However, it's a good idea to keep that leaf note. [00:27:27] Durga Toshniwal: So that, so that if such, value pair combinations occur in, the real-world unseen data, then the path can be traversed correctly. [00:27:41] Durga Toshniwal: That is one thing. [00:27:42] Durga Toshniwal: And the class label of the… [00:27:46] Durga Toshniwal: The glass label of such a node… [00:27:49] Durga Toshniwal: Could be the class label of the parent. [00:27:54] Durga Toshniwal: Or otherwise, some default class label could be used. [00:27:59] Durga Toshniwal: So, like this, we have… Then, maybe you're wrong. [00:28:07] Durga Toshniwal: Oh… [00:28:13] Durga Toshniwal: Then sometimes what can happen is that, [00:28:19] Durga Toshniwal: It, so I already discussed this condition. [00:28:26] Durga Toshniwal: Also, if, all of the… [00:28:29] Durga Toshniwal: Records follow the same tra… follow the same, similar value pair combinations. [00:28:37] Durga Toshniwal: Then also, it's no use drawing the tree, because the tree will have only a set of edges and nodes, and all the records will follow the same path. [00:28:46] Durga Toshniwal: So, such, condition is also, not good to grow the tree. [00:28:55] Durga Toshniwal: Then we discussed, the condition where we actually, you know, do that split. That split could be a binary split or a multi-way split. [00:29:06] Durga Toshniwal: Binary split means that we will have two edges, outgoing edges, and there'll be two categories in the data. So in such cases, even if there are three categories, for example, if we have car type. [00:29:21] Durga Toshniwal: Satisfying family, sports, luxury cars. Then, also, we'll… we can combine two of these. [00:29:28] Durga Toshniwal: together to form two categories, or two class labels, like this. So we could combine sports and luxury, or we could combine… combine family and luxury, like that. [00:29:40] Durga Toshniwal: So this is a binary split. [00:29:43] Durga Toshniwal: There could be multi-way split also, where we could have car type, and we could have all three categories. So, it's okay to have any of these kinds of nodes in a decision tree. [00:29:55] Durga Toshniwal: Then, to… if you are having continuous valued attributes, which I already discussed to you, we'll actually need to discretize them. [00:30:03] Durga Toshniwal: And to discretize, we will need some decision boundary. Say, for example, if the value of the attribute is less than V, then it will [00:30:12] Durga Toshniwal: be, you know, designated by some category, say, 0, and if it is greater than or equal to V, it could be designated by one category 1. It is very similar to the binary numbers. [00:30:24] Durga Toshniwal: That we have in computer systems, right? [00:30:28] Durga Toshniwal: So… Oh… [00:30:31] Durga Toshniwal: So this is like a binary split, where we are discretizing the taxable income on the condition greater than or equal to 80. [00:30:40] Durga Toshniwal: And, [00:30:42] Durga Toshniwal: This is, another condition, MV split, where we are having multiple bins, like less than 10K, 10 to 25, 25 to 50, 50 to 80, and greater than 80K. Both of them are, 5. [00:30:58] Durga Toshniwal: And then we stopped at this point, where we discussed which split out of these is best. So, let's say we had attributes like owning of a car, car type, and student ID, and we want to decide which particular one to use. [00:31:13] Durga Toshniwal: To do the split. [00:31:15] Durga Toshniwal: The owning of a car leads to the formation of two subsets, and let's assume that the total number of samples of car C0 type and C1 type are both 10 is to 10. [00:31:26] Durga Toshniwal: So both are equally… Prevalent in the given data. [00:31:31] Durga Toshniwal: Then, by doing, let's say, owning of a car, the subsets that we obtained are having a distribution 6s to 4 and 4 is to 6, for classes C0 and C1, respectively. [00:31:43] Durga Toshniwal: Then if we use car type. [00:31:47] Durga Toshniwal: Then, and the three categories are family, sports, and luxury cars, and the subsets that we obtain are 1 is 2, 3, 8 is 2-0, 1 is 27 distribution. [00:31:57] Durga Toshniwal: And then we are having student ID also available with us. [00:32:01] Durga Toshniwal: And obviously, for each student, there'll be one single ID. So, based on that, as many number of students are there, that many number of nodes we will have. [00:32:11] Durga Toshniwal: And accordingly, it will have one class label, either C0 or C1. [00:32:17] Durga Toshniwal: So, we had to decide which one is the best split. So, if we consider them one by one, owning of a car results in two child nodes. [00:32:25] Durga Toshniwal: But both of these are having almost similar, class distribution, that is C0 to C1, 6 is to 4, C0 to C1, 4 is to 6. Now, by looking at these two child nodes, it is really difficult to infer [00:32:40] Durga Toshniwal: the class label, because C0 and C1 are almost 50-50 in both of these child nodes. [00:32:48] Durga Toshniwal: So we cannot, you know, make a… Final decision. [00:32:53] Durga Toshniwal: that what will be the class label representing this particular… these two… each of these two child nodes. So then, such an attribute is not good, we will not choose it. If we look at student ID, definitely. [00:33:06] Durga Toshniwal: Each of the child nodes gives a clear majority of the class label. [00:33:11] Durga Toshniwal: But… We can see that each node is having one sample only. [00:33:16] Durga Toshniwal: So the whole idea of making groups, because classification is nothing but [00:33:21] Durga Toshniwal: Forming the groups based on the class information that is available in the data. [00:33:27] Durga Toshniwal: So here, there are no… there is no grouping. [00:33:30] Durga Toshniwal: Each student is in a group on his own, on his or her own. [00:33:35] Durga Toshniwal: And, the resultant tree also can be very bushy, assuming there may be number of students. Therefore, we'll not go for this split also. If we look at car type. [00:33:46] Durga Toshniwal: Then the child nodes that we get are having distribution of C0 to C1 as 1 is to 3, 8 is to 0, and 1 is to 7, respectively. [00:33:56] Durga Toshniwal: So each of these child nodes, have a distribution which can clearly give us the class label. For example, the first one will be C1, second one will be C0, and the third one will be C1. [00:34:08] Durga Toshniwal: And, additionally, the trees are not that very bushy also. So, therefore, we will go with car type as the node of our, I mean, as the attribute of our choice to do the split. [00:34:23] Durga Toshniwal: So, the test condition, that will be the most optimally, the best one is car type at this point in time. [00:34:31] Durga Toshniwal: So, this is what we had covered yesterday. [00:34:34] Durga Toshniwal: Saibe Hardwaj, you have a question. So, do you want to, ask it right now. [00:34:42] Saibharadwaj C: Yes, ma'am. Morning, we discussed a record where the salary… only the salary was empty, and [00:34:50] Saibharadwaj C: We discussed that we can delete the record. [00:34:53] Saibharadwaj C: Can we… Do, any data filling methods to add… [00:35:02] Saibharadwaj C: Data based on median or mean, and include that record in the dataset. [00:35:08] Durga Toshniwal: Oh, dude. [00:35:08] Saibharadwaj C: Will removing… will removing, shift the data? [00:35:12] Saibharadwaj C: Shift the model leftwards, or will adding the data shift it rightwards? [00:35:18] Saibharadwaj C: What would be a right approach for classification? [00:35:22] Durga Toshniwal: Just a second, yeah. [00:35:24] Durga Toshniwal: So, the thing is that if you are having some missing values, okay. [00:35:30] Durga Toshniwal: As a pre-processing step, remember, pre-processing is always done prior to doing anything else on the data. [00:35:39] Durga Toshniwal: So, if you are having a missing value, and if it is possible to resolve the missing value. [00:35:45] Durga Toshniwal: Then it's always a good idea to do that. [00:35:49] Durga Toshniwal: the method could be anything. It could be filling a media… of a median mode. [00:35:54] Durga Toshniwal: median based on class information, or whatever the method is that we have studied. So, it's a good idea to fill [00:36:02] Durga Toshniwal: And definitely, it may not be as good to actually remove the sample. However, if it is absolutely essential, then only we go ahead, and if those kinds of samples are very few, only then we go for that option, okay? [00:36:27] Durga Toshniwal: Does that answer your question, Saig Vaj? [00:36:33] Saibharadwaj C: Oh, yes, ma'am, that answers. Thank you. [00:36:35] Durga Toshniwal: Oh. [00:36:36] Durga Toshniwal: So… Anyone else before we now start with the new portion? [00:36:45] Durga Toshniwal: Okay, so let me proceed. [00:36:51] Durga Toshniwal: So now that we, saw… [00:36:55] Durga Toshniwal: That what would be the best split? Just visually, we saw it. [00:37:01] Durga Toshniwal: So then, whenever we have to decide what will be the best split, [00:37:08] Durga Toshniwal: So, a best split will be one. That would result, [00:37:13] Durga Toshniwal: in, you know, a node which is less impure. So, we will use the node impurity [00:37:21] Durga Toshniwal: To find out whether the split is a good split or not. [00:37:26] Durga Toshniwal: So, let's say we have, after the split, we have this kind of a note. [00:37:30] Durga Toshniwal: with the distribution of classes C0 to C1, as 5 is to 5, that is 50-50. [00:37:35] Durga Toshniwal: And then we have another node, which is C0 to C1, 9 is to 1. [00:37:41] Durga Toshniwal: So, we can see that, if we talk about a class distribution here, so class distribution of C0 to C1 is 50-50. [00:37:51] Durga Toshniwal: Right? So it's a completely uniform class distribution. [00:37:56] Durga Toshniwal: Whereas, in this case, it's a skewed class distribution, which is, 90s to 1, and 90s to 10%. [00:38:04] Durga Toshniwal: So, [00:38:06] Durga Toshniwal: By having a uniform class distribution of 50s to 50, it will be very difficult to infer the class label, right? [00:38:14] Durga Toshniwal: By looking at this kind of distribution, we cannot really say [00:38:17] Durga Toshniwal: If this was a leaf node, then what would be the class label being designated by such a leaf node? But if we had a distribution like this at the leaf node, with C090% and C1 10%, [00:38:31] Durga Toshniwal: Then, in that particular case, we can clearly say that this particular deep node designates class C0. [00:38:41] Durga Toshniwal: So, if we talk about the class distribution, though usually we would say that [00:38:48] Durga Toshniwal: a uniform class distribution is what we would want, but in this case, we don't really want a uniform class distribution. What we want is a skewed class distribution, right? So, this… this particular, [00:39:04] Durga Toshniwal: This particular notion is a little different in our case. We don't want uniformity, we want skewness. [00:39:14] Durga Toshniwal: That is one thing. Now, if we look at the degree of impurity. [00:39:20] Durga Toshniwal: The degree of impurity at this particular node is high. Why it is high? Because… [00:39:27] Durga Toshniwal: There are 50-50% samples, so it's complete mixture. [00:39:31] Durga Toshniwal: Whereas if we look at this particular note, the degree of impurity is low, because, most of the records are having similar class label. [00:39:43] Durga Toshniwal: Then we have, if we talk about the class labels themselves. [00:39:48] Durga Toshniwal: So, the class label distribution, or the class labels are heterogeneously placed. [00:39:55] Durga Toshniwal: Whereas, if we talk about a class distribution here. [00:39:59] Durga Toshniwal: So, the classes are homogeneous in nature here, because [00:40:04] Durga Toshniwal: Most of them are having the same class label, which is C0. [00:40:09] Durga Toshniwal: So, whenever the child node that, or whenever a split is such that it is resulting into a node which is having low degree of impurity in terms of class label, homogeneous class distribution. [00:40:23] Durga Toshniwal: I should not say class distribution, I should say homogeneous class labels. It should be labels here, sorry for that. [00:40:32] Durga Toshniwal: Yep. [00:40:33] Durga Toshniwal: So this one is correct in terms of class label. And skewed class distribution, then, then such a… [00:40:42] Durga Toshniwal: Such a split is a good split, and such a node that we get is a good, child node. [00:40:49] Durga Toshniwal: Then… [00:40:52] Durga Toshniwal: Now, the thing is that this we saw visually, that we are having this kind of a split and all that, and we want a node which is having a low degree of impurity, skewed class distribution, but how do we actually, [00:41:08] Durga Toshniwal: determine this. So, we will need some measure to quantitatively assess [00:41:13] Durga Toshniwal: whether an attribute is going to result into a good split or a bad split. And for that, we have certain measures. I think somebody had asked about… I think a couple of you had asked how we can determine which attribute to use. [00:41:28] Durga Toshniwal: So, in order to determine which attribute to use, we will need to measure the node impurity. [00:41:35] Durga Toshniwal: And that we can measure with the help of something called, entropy. [00:41:41] Durga Toshniwal: There are other measures also, but this is one of the most popular ones, so I'm just discussing entropy. [00:41:48] Durga Toshniwal: So, what is entropy? [00:41:53] Durga Toshniwal: I'm sure all of you would have studied entropy in class 8th, or something like that, 10th. [00:42:01] Durga Toshniwal: In chemistry. And what is entropy meaning? Entropy means the amount of randomness or disorder. [00:42:08] Durga Toshniwal: That is what we had studied in chemistry. [00:42:11] Durga Toshniwal: And this is exactly the same definition that we will make use of entropy here. So, entropy, measures the amount of randomness, or the amount of [00:42:23] Durga Toshniwal: Impurity, or the, amount of, you know, mixedness that is available. [00:42:33] Durga Toshniwal: In the given data. So how are we going to use entropy? [00:42:37] Durga Toshniwal: to find out which particular attribute to choose to result into the best split. So, let's say we have a given data with two classes, C0 and C1. [00:42:48] Durga Toshniwal: And, class C0 has N00 number of samples, and C1 has N01 number of samples. [00:42:55] Durga Toshniwal: And we are having two attributes. One is attribute A, and the other is attribute B. And we have to decide which one to use. [00:43:02] Durga Toshniwal: To do a split. [00:43:05] Durga Toshniwal: And we have to do the split, which is the better one out of these two. [00:43:10] Durga Toshniwal: Let's say A results, into… after… if we do a split on A, it results into node N1 and N2. [00:43:18] Durga Toshniwal: And B results into N3 and N4. [00:43:21] Durga Toshniwal: And, at N1, we are having the class distribution, like, N10, N11, N20, N21. [00:43:27] Durga Toshniwal: And at N3 and N4, we have N3, NC1, N40, N41, respectively. [00:43:34] Durga Toshniwal: So now, what do we want? Let's say that, [00:43:40] Durga Toshniwal: This is the… So let's assume that M0 is the entropy before doing the split, and M1 is the one… M1 and M2 are the entropies after doing the split, and M3 and M4 are the ones. [00:43:55] Durga Toshniwal: After the… doing the split on A and B, respectively. [00:43:59] Durga Toshniwal: And let's say M1 and M2 put together becomes M12, and M3 and M4 put together becomes M34. [00:44:07] Durga Toshniwal: Then, we want to go for a split. [00:44:10] Durga Toshniwal: Where the gain after doing the split is maximum. What is gain? [00:44:16] Durga Toshniwal: Gain means the entropy value before doing the split, minus the one that is after doing the split. [00:44:27] Durga Toshniwal: So if we want to do a split on attribute A, then it is M12, so the gain will be M0 minus M12. [00:44:35] Durga Toshniwal: And if you want to do the split on B, then the gain will be M0 minus M34. [00:44:43] Durga Toshniwal: So then, [00:44:46] Durga Toshniwal: Now, the thing is that we want to maximize the gain, and accordingly, we will decide which split to use. [00:44:55] Durga Toshniwal: So now, obviously we'll need to measure this gain, and to measure the gain. [00:45:01] Durga Toshniwal: We will have a measure like, let's say, entropy. [00:45:06] Durga Toshniwal: So, entropy at node T So, entropy at node T is given as entropy T. [00:45:14] Durga Toshniwal: Where T may be some child node. [00:45:17] Durga Toshniwal: That is, given by this formula, which is… [00:45:21] Durga Toshniwal: the relative frequency of the, the GS class [00:45:28] Durga Toshniwal: at node T. So, let me just explain to you once again. [00:45:33] Durga Toshniwal: So, it means that, and it is logged to the base 2, it is the probability of having the jth class at the node T. So. [00:45:42] Durga Toshniwal: whenever we do the split, the samples may have the same class label, or it may have different class label. So, there may be some 1, 2, 3, till J number of classes. So, the probability of having the Jth class at the node T [00:45:56] Durga Toshniwal: log of that into that probability. And this is done, the summation is done over all the classes, and it is minus. So this is how we measure the entropy. [00:46:07] Durga Toshniwal: At any particular node T. So whenever we do the split, Then the resultant node T, [00:46:14] Durga Toshniwal: for that node, we can find out the entropy at T, and then use this. Now, the question is that when we get the entropy, would we prefer a low value of entropy, or would we prefer a high one? Maybe someone, [00:46:30] Durga Toshniwal: some… somebody can comment on this. Which entropy should be good for us, based on the discussion that we are doing, whether it will be high entropy that will be good, or a low entropy. [00:46:45] Deepak Bobade: Entropy should be low. [00:46:48] Durga Toshniwal: Tropie should be low. [00:46:49] Deepak Katara: Low entropy, yeah. [00:46:51] Durga Toshniwal: Okay. [00:46:52] Durga Toshniwal: Anyone for a high-end trophy? [00:46:57] Durga Toshniwal: So, the answer is correct. We definitely want a low value of entropy, because entropy, designates randomness. [00:47:07] Durga Toshniwal: So, if you remember, in the previous slide, I showed to you that, here, the randomness is very high because the class distribution is uniform. We don't want that. We want [00:47:19] Durga Toshniwal: something which is queued, so that there's a clear majority of the class label. So, therefore, we want the entropy to be low. We don't want randomness, because high value of entropy means high amount of randomness. [00:47:36] Durga Toshniwal: Okay, so then, [00:47:39] Durga Toshniwal: If the randomness is minimum, then it will indicate that all the classes will belong to one class, and this will be when all the classes will belong to one class, then the gain will be maximized. [00:47:53] Durga Toshniwal: So, and if, the number of records that belong to, different classes [00:48:01] Durga Toshniwal: Or they are equally distributed across the classes, then in that particular case, the, [00:48:08] Durga Toshniwal: The homogeneity in terms of the class label will be maximized, and this is what we don't want. [00:48:16] Durga Toshniwal: So, in this formula, if, if all the records The total number of records. [00:48:24] Durga Toshniwal: at… at node T. [00:48:27] Durga Toshniwal: It's deep. [00:48:31] Durga Toshniwal: At node P is T. [00:48:35] Durga Toshniwal: Let's say. So then, if… If we have J equal to T, [00:48:41] Durga Toshniwal: Then it becomes log, then, this becomes equal to what? [00:48:47] Durga Toshniwal: It becomes equal to 0, because, this expression becomes 1. [00:48:54] Durga Toshniwal: And, so that will be 2 to the power 0. [00:49:00] Durga Toshniwal: So the log of this 2 to the power 0 will be 0, and this expression will become 0. [00:49:06] Durga Toshniwal: So, when all the records belong to the same class T, then this expression, that is, entropy becomes what? Entropy becomes zero. And when entropy becomes zero, then what happens? If this was our initial entropy. [00:49:21] Durga Toshniwal: Then what we'll have, the gain will be maximized. [00:49:26] Durga Toshniwal: So, this is just an example where you can look at the… how the calculations vary. Let's say, initially, our… this was our child node. [00:49:38] Durga Toshniwal: And, the distribution of C1 is to 0 is 0 is 0 is to 6, 1 is to 5, 2 is to 4. So, you can see that the skewness is decreasing as we go from here. [00:49:50] Durga Toshniwal: So, if we try to find out the entropy, then the probability of class C1 at this node is 0 by 6, and C2 is… [00:49:57] Durga Toshniwal: 6 by 6. So this is 0, this is 1. If we use this formula to calculate the entropy, then what we'll have is 0. [00:50:08] Durga Toshniwal: Whereas if we try to use this kind of a distribution at the child node, then probability for Class C1 will be 1x6, and this will be 5x6. And then entropy, if we calculate, it will become higher, you can see, because the… [00:50:23] Durga Toshniwal: The mixedness, or the randomness in terms of the class label is increasing. [00:50:28] Durga Toshniwal: If it increases still further, then you can see entropy has gone up. [00:50:33] Durga Toshniwal: If it becomes completely equal 3s to 3, then it might… the entropy might go up to 1, like that. [00:50:41] Durga Toshniwal: So, as we increase the, [00:50:46] Durga Toshniwal: Randomness, in terms of the class label, the entropy keeps on increasing, or the disorder is increasing, which is what we don't want. [00:50:56] Durga Toshniwal: So now we have, the formula that we had discussed earlier, with a little bit of a modification. So we call this information gate. [00:51:04] Durga Toshniwal: So, gain by doing a split. [00:51:07] Durga Toshniwal: It's defined as the entropy at the parent, This is the parent. [00:51:13] Durga Toshniwal: Right? [00:51:14] Durga Toshniwal: Minus the entropy, summation of, you know, if there are K number of child nodes. [00:51:23] Durga Toshniwal: Then, the summation of the entropy of all the child nodes. [00:51:28] Durga Toshniwal: Where i is equal to 1 2K. [00:51:31] Durga Toshniwal: And total KHIL nodes are there, but there's an additional term [00:51:36] Durga Toshniwal: Which is, introduced here, which is NI by N. What is Ni by N? [00:51:42] Durga Toshniwal: So, the entropy at any particular node [00:51:46] Durga Toshniwal: Is given a weight, that depends. [00:51:50] Durga Toshniwal: On how many samples out of the training data are going to that particular node. [00:51:55] Durga Toshniwal: So, for example, here. [00:52:04] Durga Toshniwal: So, for example, let's say that we are having a parent node P, [00:52:11] Durga Toshniwal: This is our parent node P, and let's say it's resulting into a split. [00:52:16] Durga Toshniwal: Like this. [00:52:18] Durga Toshniwal: And let's say here we are having… and total samples here are 100. Then 50 of them are going here. [00:52:25] Durga Toshniwal: Maybe 30 are going here, and 20 are going here. [00:52:28] Durga Toshniwal: And let's say the entropy at this child node is E1, this is E2, and then this is E3. [00:52:35] Durga Toshniwal: Then, as per this, the weight that will be given to entropy E1 will be 50 upon 100. [00:52:43] Durga Toshniwal: Then, it will be 30 upon 100. [00:52:47] Durga Toshniwal: And then it will be 20 upon 100. [00:52:51] Durga Toshniwal: So, this is a penalty factor, or this is the weight that is given, because any node which is having more number of samples, accordingly, the weight… more weight should be given to it. So, the weight here will be 1 by 2, here it will be 3.3, [00:53:08] Durga Toshniwal: So it will be 0.5, 0.3, and 0.2, based on the size of these nodes. [00:53:13] Durga Toshniwal: So this is how we calculate the node, the, information gain, and we want to maximize the gain, or we want to go for a gain, which is the highest one, after doing the split. [00:53:27] Durga Toshniwal: So then, now, whenever we are given multiple attributes, like A, B, C, let's say we are given 3 attributes. [00:53:36] Durga Toshniwal: And we want to choose which one to do the split at any point in time. We'll have to find out the information gain. [00:53:43] Durga Toshniwal: By doing the split on each of these, like A, B, and C, respectively, and then go for the one that gives the highest information gain. [00:53:52] Durga Toshniwal: then let's say if A gave the highest gain, we chose it. Then we are left with B and C, so next time when we do the split, we'll have to find out the gain again. [00:54:01] Durga Toshniwal: and see which one is giving better gain. Let's say B gave, so we chose this. And then, if required, then we'll choose C. So, like this, we will decide which particular attribute to use. [00:54:13] Durga Toshniwal: At any given point in time. [00:54:18] Durga Toshniwal: The stopping criteria for growing a tree. So, there can be situations when there are lots of attributes, say 50 of them, or 100 of them are available. And if we keep on using all of the attributes, then in that particular case, the tree could be overly deep, or it could be very bushy. [00:54:37] Durga Toshniwal: So, therefore, it's not necessary that we will use all the attributes, and we will stop expanding the node. [00:54:47] Durga Toshniwal: or the tree. Either when all the records belong to the same class level, then that particular node becomes the leaf node. We stop growing beyond it. [00:54:58] Durga Toshniwal: Or, when all the attributes [00:55:02] Durga Toshniwal: that are, you know, of all the training records that are falling at any particular… that are incident at any particular node, they have the same values, then we cannot split it further, because the values are all uniform. [00:55:19] Durga Toshniwal: Or else, we will go for early termination, that is, we decide the max depth to which we are going to do the tree. So, early termination is one of the most popular methods that is used. [00:55:31] Durga Toshniwal: And in the Python function, for a decision tree, usually you define the max depth. [00:55:38] Durga Toshniwal: And the tree is built up to that depth only. [00:55:42] Durga Toshniwal: So now, there are a lot of advantages of using Decision Tree. First of all. [00:55:47] Durga Toshniwal: I think there was a comment that it's very compute-intensive. So, out of the available methods, decision tree is relatively inexpensive to construct. It is very fast because, again, it is requiring less, resources. [00:56:11] Durga Toshniwal: Sorry, I got muted. [00:56:13] Durga Toshniwal: So, So, since it is inexpensive to construct, therefore it is also quite fast. [00:56:20] Durga Toshniwal: At, performing classification of the unseen record. [00:56:25] Durga Toshniwal: It is quite, intuitive, interpretable, because the tree, once built, can be easily understood, it's understandable, and the accuracy of a tree is quite high. [00:56:42] Durga Toshniwal: The decision tree, though it looks very simple, but it results into very complex decision boundaries based on the decision rules. [00:56:51] Durga Toshniwal: that, you know, that the tree follows. [00:56:56] Durga Toshniwal: And therefore, decision tree is a… results into a non-linear classification method. So, decision tree-based method is a nonlinear classification method, because the decision boundaries that we [00:57:10] Durga Toshniwal: obtain are nonlinear. Now, what are nonlinear decision boundaries, and what are linear decision boundaries? [00:57:17] Durga Toshniwal: Just a second. [00:57:23] Durga Toshniwal: So, for example, if we had a data, something like this, there are two classes in the data. [00:57:32] Durga Toshniwal: Let's say these are the two classes. [00:57:35] Durga Toshniwal: In the given data. [00:57:37] Durga Toshniwal: And we want to classify it. [00:57:39] Durga Toshniwal: So, a decision boundary, which is something like this. So, here, the hyperplane is linear. [00:57:45] Durga Toshniwal: So, any method that results into a linear hyperplane is said to be a linear classification method. However, if our data was such. [00:57:54] Durga Toshniwal: Let us say this is the class distribution. [00:58:01] Durga Toshniwal: And, [00:58:15] Durga Toshniwal: Suppose this was a kind of class distribution, then a decision boundary, would be something, it won't be a straight line like this, because we are having a mixture. [00:58:27] Durga Toshniwal: So then, in that particular case, probably the class [00:58:31] Durga Toshniwal: Distribution would be something like this, where, both of the similar, the groups having similar classes are clubbed together. [00:58:39] Durga Toshniwal: And these two are clubbed together. So this is a nonlinear decision boundary. [00:58:44] Durga Toshniwal: With the help of a decision tree, we are able to actually, [00:58:49] Durga Toshniwal: Have rules. Say, this is our decision tree. [00:58:52] Durga Toshniwal: And we are doing splits on the tree, on the attributes. [00:58:57] Durga Toshniwal: And like that. So, based on these, split, attribute values, we are doing the split. And finally, the decision boundary. [00:59:06] Durga Toshniwal: would be defined by these values, thresholds that we are using, or the values. So actually, these decision boundaries will not be linear, they'll be complex and non-linear. So even though the method looks very simple, however, it's able to classify, [00:59:23] Durga Toshniwal: In such a way that if we are having nonlinear decision boundaries, then this will work. [00:59:30] Durga Toshniwal: And that's why, generally, decision trees have a very good accuracy. [00:59:36] Durga Toshniwal: Even on complex data. [00:59:39] Durga Toshniwal: Now, we have something called Random Forest. So, random forest is a forest… it's called a forest because it is made up of a tree, so the concept is drawn. [00:59:49] Durga Toshniwal: from the real-world analogy of a forest, which is made up of trees. So here, multiple decision trees, when they are put together, they form a random forest. [01:00:01] Durga Toshniwal: So I'll explain to you why the word random is used here, but before that, it's important to know that it's an ensemble machine learning method. [01:00:11] Durga Toshniwal: Because we are having a combination of multiple trees here, and together they are used to make the final prediction. [01:00:20] Durga Toshniwal: So, each of the decision tree in a random forest actually, is, trained [01:00:28] Durga Toshniwal: on a random subset of the data samples. So, we have a random subset of the data samples. [01:00:35] Durga Toshniwal: And, it also considers a random subset of the features. So, both are randomly chosen. So, a subset of the samples are chosen, and a subset of the features are also chosen. [01:00:46] Durga Toshniwal: And, based on this subset having a lesser number of samples, or a random subset of the samples and the features, this puts together the training data on which [01:00:59] Durga Toshniwal: One particular decision tree in the random forest will be trained, and there will be multiple decision trees. [01:01:06] Durga Toshniwal: Similar decision trees drawn. And therefore, the forest is said to be random forest. Because of the randomness. [01:01:15] Durga Toshniwal: That, you know, on which the tree is derived. [01:01:26] Durga Toshniwal: Yeah, sorry for that, I got muted. [01:01:29] Durga Toshniwal: So then I explained to you what the relevance of the word forest is, because it's made up of multiple decision trees. Random is because of the randomness of the attributes and the data samples. [01:01:42] Durga Toshniwal: And, all these trees, make the prediction… predictions. And, [01:01:49] Durga Toshniwal: So, to avoid bias of having a single tree do the prediction, we have this random subset of samples and attributes per tree, and so many trees doing the prediction together. [01:02:02] Durga Toshniwal: And then, the majority vote of the class labels that are predicted by all these trees is used to make the final prediction. So, therefore, when we are using a random forest, we are actually combining lots of decision trees. [01:02:20] Durga Toshniwal: To make our model more strong and more accurate. [01:02:25] Durga Toshniwal: Now, since I told you that it's an ensemble that we are using, so ensemble can be done in a lot of ways, in actually two ways, primarily. One is said to be bagging, another is said to be boosting. It's called bagging because here, the different models are combined in parallel. [01:02:45] Durga Toshniwal: And this combination in parallel gives the final prediction model. Whereas boosting means the result of one is input into the other, then the other. So here, the outcome [01:03:00] Durga Toshniwal: is actually boosted. [01:03:09] Durga Toshniwal: Why not? [01:03:13] Durga Toshniwal: So, then we have here the result of one going as an input to the next, and so on and so forth. So, this is called boosting because the goodness gets boosted because of the sequence. [01:03:25] Anurag Krishnam: Sorry, ma'am, we lost you in between. Can you please repeat bagging and boosting again? [01:03:32] Durga Toshniwal: Okay, okay, so what I said was that bag… in case of bagging. [01:03:39] Durga Toshniwal: What we do is that, whatever the models are. [01:03:42] Durga Toshniwal: In this kind of ensemble, we parallelly combine the models. So all the models actually choose random subset of the data and random set of attributes, and then, let's say if you're talking about a decision tree, then each of the tree is combined in parallel. [01:04:01] Durga Toshniwal: And then, the ensemble is actually a parallel combination of all of these, giving the final outcome. [01:04:09] Durga Toshniwal: The other option is to make and ensemble is boosting. So what is boosting? Boosting means the outcome of one is improved. [01:04:18] Durga Toshniwal: By the next, decision tree, or next classifier, by giving the output as an input to the next one. Then the outcome of the next one is given, and the input to the next one, like that. So the, [01:04:33] Durga Toshniwal: This works in a sequential fashion, like this. [01:04:36] Durga Toshniwal: And then the… So the outcome of this will be something [01:04:42] Durga Toshniwal: And then the outcome of this will actually depend on the outcome of the previous unit. And the outcome of this will be governed by the output of the previous unit like this, and the outcomes of them, all of them are combined together. [01:04:55] Durga Toshniwal: Whereas, the combination of the, of the classifiers are sequential. Here, the combination of the classifiers are in parallel. [01:05:05] Durga Toshniwal: Accordingly, we have, their names coined as bagging and boosting, respectively. Bagging is parallel combination. [01:05:13] Durga Toshniwal: Boosting is sequential combination, and then the final outcome is obtained by taking the decisions of all the units that make up the ensemble. [01:05:27] Durga Toshniwal: So then, in random forests, the concept of bagging is used. [01:05:32] Durga Toshniwal: As I already showed you, what is baggy? Baggy is a parallel combination of the units. In this case, it is the decision tree. [01:05:41] Durga Toshniwal: So, bagging is also called bootstrap aggregation. [01:05:45] Durga Toshniwal: Okay, so what is, so the bees obtained here, and aggregation. [01:05:52] Durga Toshniwal: From here. So, bagging, is actually also called bootstrap aggregation. [01:05:58] Durga Toshniwal: What is bootstrap aggregation? Bootstrap aggregation is a mechanism where multiple decision trees [01:06:05] Durga Toshniwal: Which are trained independently on different subsets of the given training data are combined together. [01:06:13] Durga Toshniwal: Right? In an ensemble. So, so the idea here is that each tree will have a different view of the given training data. How it will have a given, different view? It will have a subset of the training samples, and it will also have a subset of the [01:06:31] Durga Toshniwal: Attributes that are available in the training samples. [01:06:35] Durga Toshniwal: So, the randomness in this random forest, which is made up of trees, comes from two things, that is the random… randomly selected samples and the random subset of the attributes. [01:06:48] Durga Toshniwal: Which I already discussed. [01:06:52] Durga Toshniwal: And then, [01:06:55] Durga Toshniwal: what is bootstrapping? Bootstrapping actually means that to create a sample dataset that will be used by the different trees that make up the forest. [01:07:09] Durga Toshniwal: Can be obtained by repeatedly drawing, samples with replacement. [01:07:16] Durga Toshniwal: So, what is drawing samples with replacement? So, let's say I'll show you with an example. Let's say I have this original box. This box contains so many balls, and each ball is having a unique number on it, going from 1, 2, 3, whatever, like this. [01:07:35] Durga Toshniwal: So, we are having some 8 volts, each having a unique number. [01:07:39] Durga Toshniwal: on it Now, [01:07:43] Durga Toshniwal: Drawing a random sample with replacement means that, let's say, there are 3 samples, set of samples that we are drawing. [01:07:53] Durga Toshniwal: And in these three set of samples, what is happening is that [01:07:58] Durga Toshniwal: First of all, let's say we are drawing… we are randomly drawing some particular ball, let's say this box is having balls, and we randomly put our hand in the box, and we just blindly took out a ball. Let's say it was the ball having number 1. [01:08:15] Durga Toshniwal: So, we… what we did was, we took it out and put it in the… another box where we are going to put the sample. [01:08:23] Durga Toshniwal: Okay? [01:08:25] Durga Toshniwal: So now, what is going to happen is that [01:08:29] Durga Toshniwal: Because we are following the policy of replacement, therefore, what will happen? The ball with number 1 will be replaced again in the original box. [01:08:41] Durga Toshniwal: So, once again, when we put our hands… so, now, in each of the, resulting samples, we'll have… because there are 8 number of samples here, so we will have a total of 8 samples only in each of these subsets also. [01:08:55] Durga Toshniwal: So, again, when we put our hand, because the ball with number 1 has been replaced, therefore it can be drawn again. [01:09:02] Durga Toshniwal: Once it is drawn again, then because of the policy of replacement, this will be replenished back in the original box. And once again, it can just randomly happen by chance that we are once again drawing number 1 again. So, as many times one is drawn, or whichever ball is drawn. [01:09:19] Durga Toshniwal: it will be replaced back in the original box. A new one with the same number. [01:09:24] Durga Toshniwal: Will be replaced back inside the, inside the box. [01:09:29] Durga Toshniwal: Okay, so, here what we are doing, we are actually, following the principle of replacement. [01:09:37] Durga Toshniwal: So, the total number of samples in the original box and the new boxes should be the same. [01:09:45] Durga Toshniwal: That is it. [01:09:47] Durga Toshniwal: So, whenever we are drawing any particular ball. [01:09:50] Durga Toshniwal: As per the strategy of replacement, a similar… a ball with the same number will be replenished back in the original box. Say, for example, in the second subset, so in the first subset, we are having [01:10:03] Durga Toshniwal: Out of the 8 balls, 6 of them having number 1 and 2 of them having number 5, because when we were drawing the random sample. [01:10:11] Durga Toshniwal: It just so happened that we, the ball that we had taken out had number 1, but it had got replaced again. Again, when we put our hand… again, we got a number 1 like that, by chance. [01:10:23] Durga Toshniwal: And similarly, two walls of number 5. [01:10:26] Durga Toshniwal: Came in our hand. [01:10:28] Durga Toshniwal: The second random sample, then, when it was drawn, then let's say the first ball that got drawn by a blind, you know, pick of the ball was number 5. [01:10:40] Durga Toshniwal: And then it was number 4. When the ball number 4 was picked up, then a new ball with the same number was replenished back in the original box. [01:10:49] Durga Toshniwal: And when we did a blind pick again, it so happened that we just got the number 4 again, 3 got picked 3 times, 6 got picked twice, and so on and so forth, and it could be unique numbers also. [01:11:01] Durga Toshniwal: Right? Whereas there are certain ball IDs that never got picked up, say, for example, number 2, number 7, none of these samples have. [01:11:10] Durga Toshniwal: Right, so this is the strategy, [01:11:13] Durga Toshniwal: that I wanted to explain to you, that is drawing random sample with replacement. So, this is exactly what is used in a random forest. So, as you hear. [01:11:27] Durga Toshniwal: That this original box was nothing but your training data. [01:11:31] Durga Toshniwal: Okay. [01:11:32] Durga Toshniwal: This is your training data, and number of samples in it are 8. [01:11:36] Durga Toshniwal: And whenever a particular sample is drawn, you assume that a similar sample is replenished back in the original training data. [01:11:45] Durga Toshniwal: And it is ready to be picked again. [01:11:48] Durga Toshniwal: So, [01:11:49] Durga Toshniwal: So this is the analogy that is used in random forest in the method of bootstrap sampling or bagging, that we use a replacement policy, where the original sample, when it is drawn, a similar sample is replaced back. [01:12:07] Durga Toshniwal: Like that, okay? And this can be done multiple times. [01:12:11] Durga Toshniwal: So then, once we do this, we can see that, now we have these, views on the training data. Say we call it V1, version V1, V2 version, V3 version of the same training data. [01:12:26] Durga Toshniwal: Now, different trees are built on each of these, you know, random samples, drawn with replacement of the training data. [01:12:37] Durga Toshniwal: And, each of these trees are built, and these decision trees give some outcome. [01:12:43] Durga Toshniwal: And all of the outcomes are combined together using majority voting. So, in this case, none of the single tree has, can have any bias on the result. [01:12:56] Durga Toshniwal: And, because the number of trees that are drawn are very large. [01:13:01] Durga Toshniwal: In the random forest. Therefore, there'll be no skewness, nothing, no bias, nothing would be there, and the result will be a very fair result. [01:13:13] Durga Toshniwal: So then, when we talk about bootstrapping, the important thing is that, we use, resampling with replacement, which I already told you that each data point, in the original training data [01:13:28] Durga Toshniwal: Has the equal chance of being selected. [01:13:31] Durga Toshniwal: And after one particular data point gets selected in the subset, it is replenished back. A similar record is replenished back, which is called replacement. And therefore, any particular data point may get selected multiple times across multiple subsets, or even within a subset. [01:13:49] Durga Toshniwal: Then… The sample size will be typically the same as the original training data. [01:13:55] Durga Toshniwal: Only, the versions of the original data will be different, or the views may be different. [01:14:02] Durga Toshniwal: And the number of times such samples are drawn and the total number of trees that are built are huge. Say, for example, 1,000, 100, something like that. [01:14:14] Durga Toshniwal: And the purpose is to actually build a much better or a stronger tree. [01:14:19] Durga Toshniwal: So that the overall performance of, the classifier gets improved. [01:14:26] Durga Toshniwal: Byrd this. [01:14:29] Durga Toshniwal: So now, since we are talking about decision tree, there are, certain issues, also. We talked about the advantages of a decision tree, like, it's very interpretable, it's quite, simple to build. [01:14:46] Durga Toshniwal: It's quite understandable and all that. But some important issue is that of scalability. That is that [01:14:53] Durga Toshniwal: Whenever we are talking about training data. [01:14:57] Durga Toshniwal: If the size of the training data is high, let's say if the training data is some thousands or millions of records, to build a tree, all of these records should be, available in the main memory, inside the data. [01:15:11] Durga Toshniwal: Why it should be available in the main memory? Because we are going to build the tree using the training data. So, if the training data is available in some secondary or tertiary memory, such as maybe on a hard disk or pen drive, or whatever. [01:15:28] Durga Toshniwal: Then, every time we want to obtain samples from the training data, we'll actually have to access the secondary memory. And secondary memory is always very slow as compared to the main memory. As a result, the time to build the tree will be huge. [01:15:44] Durga Toshniwal: And therefore, the prediction time will also increase a lot. So, whenever we are building a tree, we must have all the training samples inside the main memory. So, if [01:15:56] Durga Toshniwal: The number of samples which are used to build the tree are very large, then definitely the tree will be less scalable. [01:16:05] Durga Toshniwal: Also, [01:16:07] Durga Toshniwal: We know that in the real world, training data can be very huge, and it may be very difficult to fit all of that training data into the main memory, because main memory is a very costly resource. [01:16:23] Durga Toshniwal: And in such cases, building of the tree will be very inefficient due to swapping in and out of the training samples from the main memory into the second memory, and then bringing the next subset, then swapping it out, then bringing the next subset, and all that. So this is one important issue. [01:16:41] Durga Toshniwal: That can be there. [01:16:44] Durga Toshniwal: So for… in order to solve this problem, we usually restrict the depth of the tree. [01:16:50] Durga Toshniwal: And, we, [01:16:52] Durga Toshniwal: We might also go for some kind of stratified sampling if we are not able to fit the training data inside the main memory. [01:17:00] Durga Toshniwal: So, this is one particular solution. [01:17:05] Durga Toshniwal: Now, once the tree is built, how do we assess it? [01:17:10] Durga Toshniwal: So, there are two important kinds, and generally for any classification method, not just for a tree, we have two kinds of errors, called the training error and the generalization error. [01:17:22] Durga Toshniwal: So, what are training errors and what are generalization error? Training error is the error that is encountered on the training data itself. [01:17:32] Durga Toshniwal: That is the number of misclassifications on the training data. [01:17:37] Durga Toshniwal: Itself, while we are building the classification model on the training data. [01:17:43] Durga Toshniwal: And the other error is said to be generalization error, which is the error that is expected on some generic data or unseen data. [01:17:53] Durga Toshniwal: So, we have these two kinds of error, training error, generalization error. Now, my question to you is, that training error I told you is the error on the training data. [01:18:03] Durga Toshniwal: Now, remember that we are building the classifier on the training data itself. [01:18:08] Durga Toshniwal: Since we are building the classifier on the training data, therefore, probably the training errors may be zero, because all the training samples may be getting very nicely classified, because the classifier is built on this data itself. [01:18:23] Durga Toshniwal: So, when the training errors are zero. [01:18:26] Durga Toshniwal: Then is, then, whether a classifier is going to be a very good classifier, or it's not going to be a good classifier. [01:18:34] Durga Toshniwal: Any comments on that? [01:18:39] Durga Toshniwal: I'm saying that… [01:18:40] Deepak Bobade: repeat that? [01:18:41] Durga Toshniwal: Yeah, so what I'm saying is, we have two kinds of errors, training error and generalization error. [01:18:47] Durga Toshniwal: So let's talk about training error. Training error means that whenever we are building a classifier. [01:18:53] Durga Toshniwal: Then, the number of training records that are misclassified by that model is the training error. [01:19:01] Durga Toshniwal: But since we are building our classifier on the training data itself, therefore, is it okay to have zero training error? Because if our model is such that all the training samples are correctly classified, then the training error becomes zero, right? So, is this a good scenario or a bad? [01:19:24] Hemanth Gunturu: And this will lead to water treatment of the model. [01:19:27] Durga Toshniwal: This will lead to overfitting of the model. [01:19:30] Durga Toshniwal: Okay. [01:19:32] Anurag Krishnam: This would be a bad classifier. [01:19:34] Anurag Krishnam: Because, if training errors are zero, it definitely means it's, training the examples too perfectly. [01:19:42] Durga Toshniwal: So, perfection is not good? Is it bad? [01:19:46] Deepak Bobade: I think it is the goal, right? Accuracy is the goal, right? [01:19:50] Durga Toshniwal: Yeah, accuracy is the goal, so perfection should be good. [01:19:54] Deepak Bobade: Yes, it should be quit, yeah. [01:19:57] Durga Toshniwal: Anyone else? [01:20:03] Durga Toshniwal: So, we have some answers saying… [01:20:05] Durga Toshniwal: Having training error 0 is bad. Maybe equal number of votes for training error to be 0 is perfect and good. [01:20:16] Durga Toshniwal: What is the answer? [01:20:17] Gunjan Bhaiya: Typically, error is… can't be zero, I mean, so that is okay to… it's good to have that, some errors, because that shows that actual real-time data, not, manipulated data. [01:20:29] Anurag Krishnam: And I definitely think that it does not generalize well to unseen or test data, actually. [01:20:36] Midhun VM: Absolutely, it should… it should have some ironic. [01:20:39] Durga Toshniwal: It should have some error, but what I'm saying is, what may be the reason? See, we are building a classifier on the training data. [01:20:46] Durga Toshniwal: So this means that if you are building a… you are having a perfect classifier, then it should be able to correctly classify all the samples on which it is built, right? The model itself is built on a certain set of data. So all that should be correctly classified. So, what's the harm in it? [01:21:04] Durga Toshniwal: I mean, just saying that, it is going to be bad or good… [01:21:10] Midhun VM: In real time, it will be the training data that we have used, right? The data that we're testing on might be different from what we are training on. [01:21:18] Midhun VM: So, some errors should be, there so that we can understand what, what, where we are going from. [01:21:24] Deepak Katara: What I'm saying is, because we have already processed the data, and if it is too perfect, that means we don't have different attributes, meaning that we sort of either misinterpreted our outliers. [01:21:37] Deepak Katara: That could be one thing. And the second is, like, we have perfectly processed the data, there is no noise, there is no anomaly, which is sort of not a scenario in real world case, I think. [01:21:51] Durga Toshniwal: Okay. [01:21:53] Durga Toshniwal: Okay, of course, we assume that the data is pre-processed in a nice fashion. We don't have any anomalies, and we don't have any noise, or we don't have anything missing value also in the data. Still, if the training errors [01:22:07] Durga Toshniwal: R0 or not is a good scenario or a bad scenario. [01:22:12] Durga Toshniwal: So, let me give you an example. Of course, it's not directly related, but maybe it will help to understand… gain understanding. So, you know, there's a student who says that he has, you know, he has stopped the class. [01:22:28] Durga Toshniwal: Okay. So, whether that thing is… and all his answers are correct in the question paper that was given to him for the exam. [01:22:37] Durga Toshniwal: Now, the question is that definitely it's good to have a student who gets the highest score in the class, and who answers all the questions correctly. [01:22:48] Durga Toshniwal: But if we think about the background, which the student did not tell us. [01:22:54] Durga Toshniwal: Number one, he was the only student in the class. [01:22:58] Durga Toshniwal: Then, so, whatever marks he got will definitely be the highest. [01:23:03] Durga Toshniwal: Secondly, the questions were designed by the student, and so were the solution. So whatever was written by the student was marked as correct only. Whether actually it was correct or not correct is a different question. So then, a scenario [01:23:18] Durga Toshniwal: Where a student says he taught the class where… which had only one student, and the paper attempted was made by the student, and so are the solutions. And therefore, everything was correct. [01:23:29] Durga Toshniwal: So, such a scenario where even the student tops or gets full marks is actually… looks ideal, but it's not so ideal, actually. It's… it's not a good scenario. [01:23:41] Durga Toshniwal: Similar is this case when we are building a classifier on a training data. [01:23:46] Durga Toshniwal: And all the samples in the training data are very nicely or very correctly mapped. [01:23:53] Durga Toshniwal: to the, you know, to the classifier, and we get a zero training error. So, just crudely thinking, we always think of error, something which is not required, and therefore, zero error is something that looks good. [01:24:08] Durga Toshniwal: But actually, it is deceptive. [01:24:10] Durga Toshniwal: So, in this particular case, zero error is actually deceptive. It's not good. [01:24:16] Durga Toshniwal: I'll show you within… with the help of an example. Let me see if I have space on the next slide. Oh, so I'll do it here itself. [01:24:25] Anurag Krishnam: So, ma'am, you're saying that a model actually remembers everything, but it doesn't understand the pattern, which you were talking about the example previously about the student. [01:24:35] Durga Toshniwal: Yes. What you're saying is correct, and I'll substantiate with the help of this diagram. [01:24:42] Durga Toshniwal: Let's say that, okay, so let's say once again, I'll… draw something. [01:24:52] Durga Toshniwal: Okay. [01:24:54] Durga Toshniwal: I'm good. [01:24:57] Durga Toshniwal: We have a binary class problem. We have two classes shown by the pluses and the minuses. [01:25:04] Durga Toshniwal: And then… [01:25:13] Durga Toshniwal: This is the distribution of the points that are there in the given data space, and we want to draw a decision boundary. [01:25:23] Durga Toshniwal: So, a decision boundary. [01:25:26] Durga Toshniwal: you know, that gives zero, misclassification, and this is the training data. On this training data would look something like this. [01:25:45] Durga Toshniwal: And let's say that… I also add some more points, you know. [01:25:51] Durga Toshniwal: So then… It passes like this. [01:25:55] Durga Toshniwal: And it passes like this. So this is our decision boundary. [01:25:58] Durga Toshniwal: On this given data. [01:26:01] Durga Toshniwal: Now, the question that I want to ask you is. [01:26:05] Durga Toshniwal: That this is our training data. [01:26:09] Durga Toshniwal: And, here, the training error, or the error, or the misclassification. [01:26:15] Durga Toshniwal: on the training data is zero, because you can see that all the points having same class label as plus or the circle are falling on either one side of the decision boundary. Now, is this decision boundary a good boundary or not? [01:26:32] Deepak Bobade: It's a good boundary. Yeah, it's a good boundary. It is able to separate, right? [01:26:38] Durga Toshniwal: Yeah, it's a good boundary because it's actually able to separate the points as it should have done. [01:26:44] Durga Toshniwal: Okay, anyone else? [01:26:51] Durga Toshniwal: Is it a good one? [01:26:52] Durga Toshniwal: a bird? [01:26:53] Saibharadwaj C: It might be overfitting the data. [01:26:56] Durga Toshniwal: Why do you say it's overfitting? What do you mean by overfitting? [01:27:01] Saibharadwaj C: The data is… More finely classified between categories, and the division, it's done… [01:27:12] Saibharadwaj C: The, the model is trying to [01:27:14] Saibharadwaj C: Achieve a higher level of classification for the data. [01:27:18] Durga Toshniwal: Yeah, whatever you're saying, how do you map it to this diagram that I'm drawing? [01:27:23] Saibharadwaj C: So, all the X and all the zeros, they are classified perfectly. [01:27:28] Saibharadwaj C: On the either side into two categories. [01:27:32] Saibharadwaj C: So, my understanding would be a model, if it's trying to show that it's 100% correct, then… [01:27:39] Saibharadwaj C: I would believe it's overfitting the data into Classification. [01:27:45] Durga Toshniwal: So what you are saying is correct, but I will, show here… [01:27:53] Durga Toshniwal: So, actually, the decision boundary would have been something like this. [01:27:58] Durga Toshniwal: Which is a very simple and a linear decision boundary. [01:28:02] Durga Toshniwal: Right? But we made it overcomplicate to have all the training samples… I mean, to have the misclassification on the training samples to be zero. [01:28:12] Durga Toshniwal: Actually, in real-world data, or whatever data we are having, there might be some noise. So here, only one sample is getting… on the training data is getting misclassified because it's falling on this plus side. And one sample is getting… [01:28:26] Durga Toshniwal: misclassified because it's going on the other side. But generally, if you see most of the data, that is plus side goes on this side, and the other one goes on this side, right? So there might be, training samples. [01:28:40] Durga Toshniwal: That, you know, that are noise points. But if we use them to define our decision boundary, it might end up getting overly complex. [01:28:51] Durga Toshniwal: And it's not just about complexity, it's also incorrect, because the correct decision boundary would be like this. Now, suppose if in… if we were using this boundary, and we had a new point. [01:29:03] Durga Toshniwal: Let's say we had a point coming in here, which is 0. [01:29:07] Durga Toshniwal: I mean, the circle. Then it would be misclassified as… [01:29:11] Durga Toshniwal: plus point, right? It is actually a circle, but it is getting misclassified as plus because of this kind of a nonlinear nature, overly complex nature of the decision boundary. Actually, if you look at this, this portion is like this. [01:29:26] Durga Toshniwal: So, if there is a point which is a circle, should broadly fall on the circle side, not on the plus side. So, this point is getting misclassified because of the over nonlinear nature, or over, you know, fitment. [01:29:42] Durga Toshniwal: of the training samples, of the boundary, decision boundary on the training samples. It shouldn't be like this. The… so, in this case, however, if we have a circle here. [01:29:55] Durga Toshniwal: It will be classified as circle only, because the decision boundary is not based on certain noise points. [01:30:02] Durga Toshniwal: Similarly, if we have a plus point out here. [01:30:06] Durga Toshniwal: Like, if we had a plus here, it would get misclassified into a circle. Here, it won't happen. [01:30:12] Durga Toshniwal: Or in other words, having training error to be zero is, for real-world data is actually not good, because we are unnecessary… we might be unnecessarily complicating the decision boundary. [01:30:25] Durga Toshniwal: Right? And this is what we will show you. I'll show you in the hands-on that we'll be doing. [01:30:30] Durga Toshniwal: How we'll use the training error and the test error to find out [01:30:35] Durga Toshniwal: Whether our model is having training error zero, and so on and so forth. So, error on the training data is called training error. We don't want it to be equal to 0. [01:30:47] Durga Toshniwal: Though, deceptively, error should be zero, but training error should not be zero. Because means… it means that our model is unnecessarily complex. [01:30:58] Durga Toshniwal: Because it's overfitting, it's too much tailor-made on the data, right? [01:31:03] Durga Toshniwal: Now, the question to you is, it's too much tailor-made, then is it good? So, again, now the question that I'll ask you is with another example. Let's say you go to a shop. [01:31:14] Durga Toshniwal: Let's say it's keeping some, wardrobe, or it's keeping like that. Let's say it's… it keeps shirts. [01:31:21] Durga Toshniwal: One scenario is that when you go to the shop. [01:31:25] Durga Toshniwal: Then, you have shirt fitting exactly custom-made for your size. [01:31:31] Durga Toshniwal: Right? You may be a regular customer, and they have short fitting to your size. The other scenario could be the shop could be having some sizes that may be generic. [01:31:44] Durga Toshniwal: So, which one will be better? [01:31:46] Durga Toshniwal: Let's say that there is a generic will be better. So, let's say that if the size is generic, then anybody having a petite [01:31:55] Durga Toshniwal: Size, or a small, or a medium, or a large, or a extra large. [01:32:01] Durga Toshniwal: All will be able to wear it. [01:32:03] Durga Toshniwal: Okay, because it's one generic size. [01:32:06] Durga Toshniwal: So, that's going to be good. [01:32:09] Deepak Bobade: Yes. [01:32:11] Anurag Krishnam: Yes. [01:32:12] Durga Toshniwal: Okay, so that's going to be good. Now, imagine a person who's having petite size, wearing a shirt, let's say the max size or the most generic form is 3XL. [01:32:25] Durga Toshniwal: So, of course, the person, because we had one generic size, it fitted everyone. So, it has to fit 3XL also, person, which is the max size, or the max person size. So, petite wearing a 3XL shirt, will it be good for him? [01:32:41] Deepak Bobade: No. Don't. [01:32:43] Durga Toshniwal: It won't be good for him. [01:32:45] Durga Toshniwal: And similarly, if there's someone who's XL and is wearing a… let's say the size was petite, it wouldn't be good, it will be too very tight. [01:32:54] Durga Toshniwal: Then, this means that having generic size, actually. [01:32:59] Durga Toshniwal: you know, is not good. [01:33:02] Durga Toshniwal: Because the size may be overly generic. [01:33:05] Durga Toshniwal: Now, the other scenario is that the person walks into the store and finds a shirt exactly custom-tailored on his or her size. Now, is it going to be good? [01:33:19] Deepak Bobade: It's good, yeah. [01:33:21] Durga Toshniwal: It is good. [01:33:22] Anurag Krishnam: No, so, ma'am, you're saying that it's tailored to his size exactly. But anyone new comes in, it won't be fitting his… that particular size, right? So, any new person comes in, it will be difficult for him using the same size. [01:33:36] Anurag Krishnam: So, might not be good. [01:33:39] Durga Toshniwal: Okay. [01:33:40] Durga Toshniwal: So, assume that, you know, they are able to create [01:33:44] Durga Toshniwal: Shots, in a finite… [01:33:48] Durga Toshniwal: amount of time, which is bearable for any customer, so you assume that the custom-made size is available for any customer, whether he's a repeating customer or a new customer, then will it be okay? [01:34:02] aditya shrivastava: Yes, yes. [01:34:04] Durga Toshniwal: Okay. [01:34:05] Durga Toshniwal: Okay, now, imagine the scenario that you walked into a store, and you saw a shirt. [01:34:12] Durga Toshniwal: or a set of shirts that exactly fit to whatever your size is. Of course you like that, but you didn't like the color, or you didn't like the pattern. Then will you buy it? [01:34:26] Deepak Bobade: No. [01:34:27] Durga Toshniwal: So, having the same size may be good, but then the… obviously, the storekeeper may not be, having all colors made on your size, or all patterns made on your size. Therefore, having a good fit, or exact fit. [01:34:46] Durga Toshniwal: But, may also not be very good. [01:34:48] Durga Toshniwal: Not for you, and not for the shopkeeper, because for you, the pattern and color variance will be lacking to some extent, and for the shopkeeper, it's a waste of money if you don't buy it. [01:35:00] Durga Toshniwal: Therefore, if it's an exact fit, it is not good. But if it is generic also, it's not good. Then what is good? [01:35:10] Durga Toshniwal: So… [01:35:11] Durga Toshniwal: Who is a little generic and a little specific? So, what do you mean by little generic and a little specific? So, instead of having generic size, fit for all kind of a thing. [01:35:22] Durga Toshniwal: we have certain segregation, like, we have groups of sizes, like medium, large, very large, maybe XL, XXL, like that. [01:35:31] Durga Toshniwal: So, we are having generic size, so we are generic to some extent that certain group of people will be able to fit to a certain size. Obviously, it may not be an exact fit, but more or less, it will be a fit. [01:35:43] Durga Toshniwal: But we are not going for a, you know, zero, looseness or, exact fit also. We are not going for that. [01:35:54] Durga Toshniwal: So, we are, still generic, so that a group of people can fit. Now, from the shopkeeper's perspective, somebody out of that group might like the color pattern and everything and will buy. [01:36:05] Durga Toshniwal: And there'll be no loss, in terms of the money, invested. [01:36:10] Durga Toshniwal: Similarly, from the buyer's point of view also, having this kind of a size is better, because the buyer gets more variety, more number of patterns, colors, and all that, because the sizes are a little generic. [01:36:23] Durga Toshniwal: So, this means having exact fit or having over-generic size, both are not good. What is good is in between, a little generic and a little specific. [01:36:34] Durga Toshniwal: Now, I gave you this example to relate to this training error. If we have exact fitment and zero misclassification on the training data, it's not going to be good, as I have shown you in this example. This is not good, this is good. So what are we allowing? We are allowing a little generality. [01:36:52] Durga Toshniwal: So, therefore, we want generalization errors also Not to be zero. [01:36:58] Durga Toshniwal: We want them to be finite, but we don't… [01:37:01] Durga Toshniwal: Don't want them to be very high. [01:37:03] Durga Toshniwal: they should be very high also. We want them… we want generalization error. [01:37:08] Durga Toshniwal: Like, the errors of this kind, this one, this one. We want these, but… [01:37:14] Durga Toshniwal: At the same time, we don't want them to be very high. So, we want training errors. [01:37:19] Durga Toshniwal: to be non-zero, but not to be very high. We want generalization errors to be non-zero, but not to be very high. So, we want little of both of these. So, we want a model that we build to be a little generic and a little specific, both of these. [01:37:34] Durga Toshniwal: So that's exactly what we want. And in such cases, the model will be not overfitting, this is an overfitting model. [01:37:42] Durga Toshniwal: Similarly, the model will also not be underfitting, which is a case of generalization, where it's overly generic, that everything fits [01:37:50] Durga Toshniwal: So, it should not happen like that also. So, we don't want overfitting, we don't want underfitting, we want just the right. [01:37:57] Durga Toshniwal: Or the most optimal combination, right? So, I've already told you, when the model overly fits the training data. [01:38:04] Durga Toshniwal: then it's said to be overfitting situation. We don't want that. [01:38:11] Durga Toshniwal: So, this is a graph in which we can illustrate. So, we have two kinds of errors, the training error, which is shown in the red color, and the test error. [01:38:20] Durga Toshniwal: So, as I told you, that once the model is built, it is tested on another data, which is similar to the training data, but not exactly the same, which is the training test data. [01:38:30] Durga Toshniwal: So, the model is built on the training data, it is verified on the test data. [01:38:36] Durga Toshniwal: So, this is the training error, and this is the test error. Test error is also same as generalization error, because it's an error on unseen data. [01:38:45] Durga Toshniwal: Data is unseen by the model. [01:38:47] Durga Toshniwal: So then, what does that mean? [01:38:51] Durga Toshniwal: So, here we can see, initially, when the model is built, both the training error and the test error are high, because the model is not yet stable, it hasn't learned the patterns from the data. [01:39:01] Durga Toshniwal: Now, slowly, when we increase the number of nodes, we are talking about a decision tree. So, when we increase the number of nodes, slowly, the model is learning the pattern in the data, and the error on the training sample is [01:39:15] Durga Toshniwal: Becoming lesser and lesser. [01:39:18] Durga Toshniwal: Whereas, if we look at the test error, initially it is high. When more amount of, you know, nodes are used to represent the decision tree. [01:39:29] Durga Toshniwal: Initially, the test error will go down. As you can see here, it is going down, right? And then, beyond a point, increasing, [01:39:40] Durga Toshniwal: Increasing the number of nodes will have no impact [01:39:43] Durga Toshniwal: Right? In this situation, it has no impact. [01:39:47] Durga Toshniwal: On the test error, because addition is not, improving, and substantially the error. [01:39:56] Durga Toshniwal: And beyond a point, again, the error will start increasing. [01:40:01] Durga Toshniwal: So, when the number of nodes in a decision tree are optimally high, they're not too high, then the model is not overfitted, and in that particular case, the training error and the test error both are optimally low. Not exactly low at the same time, but they'll be optimally low. [01:40:21] Durga Toshniwal: But if we are having both the training and the test error to be high means we are having underfitting, because neither the model has learned the pattern, nor it is able to predict properly. [01:40:32] Durga Toshniwal: And in case of overfitting, what will happen? Only the training error will be low, because misclassification on the training data [01:40:40] Durga Toshniwal: is going to be, there, not there. So that is going to be zero. Whereas, if we talk about, [01:40:48] Durga Toshniwal: If we talk about the test data. [01:40:52] Durga Toshniwal: the misclassification on the test data is going to be high, right? And therefore. [01:41:01] Deepak Bobade: Event thumb. [01:41:03] Durga Toshniwal: Yeah, sorry. [01:41:06] Durga Toshniwal: So then, in such cases, though the error on the, on the training data is very low, but the generalization error, so model is overfitted, and the error on the unseen data is going to be high. [01:41:21] Durga Toshniwal: So, therefore, we don't want overfitting, we don't want this, we don't want underfitting. We want this… this kind of a division, where both of them are optimally low. [01:41:32] Durga Toshniwal: Hoping so, [01:41:35] Durga Toshniwal: Now, what could be the reasons of overfitting? I already showed you with the help of the example. In this example, you have two class labels, one shown in the red, one shown in the blue, and our decision boundary actually should have been like this. [01:41:48] Durga Toshniwal: But to accommodate this noise point, it actually had… has been made like this. [01:41:53] Durga Toshniwal: So it's unnecessarily been, overfitted on the model, on the training data, right? And, it has… the training data has, you know, it has just become [01:42:06] Durga Toshniwal: Just a second, let me connect my charger. [01:42:14] Durga Toshniwal: So it has unnecessarily become more complex. [01:42:18] Durga Toshniwal: Right? [01:42:20] Durga Toshniwal: Then, overfitting… the reasons for overfitting may be many. One, the overfitting may be due to noise, as you are seeing here. The overfitting may be also due to insufficient, examples in the training data. [01:42:32] Durga Toshniwal: to… for the use cases, as I already showed you. Sometimes the use cases which the model is supposed to see are missing, and therefore it just unnecessarily becomes overfitted. [01:42:44] Durga Toshniwal: So, therefore, we should have adequate and all kinds of use cases in the training data for the classifier to work properly. [01:42:53] Durga Toshniwal: So then, overfitting actually results into a tree that is overly complex, than it should be. [01:43:00] Durga Toshniwal: And in this case, the training error provides… is no longer providing a good mechanism of estimating the performance on the tree. Because we… actually, by looking at the training error, we really cannot say how good or bad, you know, the performance on the unseen data would be. [01:43:25] Durga Toshniwal: Right, so we want the training error to help us assess the performance of the tree. [01:43:30] Durga Toshniwal: Or the classifier. But training data, in case of overfitting. [01:43:35] Durga Toshniwal: Provides us no such estimate, because it's just overly fitting on the classifier. [01:43:42] Durga Toshniwal: So I think with this, I will just wrap up here, and I'm open to any questions that you may be having. [01:43:49] Deepak Bobade: Yes, I got two questions. [01:43:52] Deepak Bobade: First one is, like, if we talk about boundaries, right, we got two, linear and non-linear. So, is it, like, nonlinear is a bad boundary or something? [01:44:03] Durga Toshniwal: No, no. Having a nonlinear boundary is not bad. [01:44:07] Durga Toshniwal: But nonlinear boundary should be made only on truly nonlinear data. [01:44:12] Durga Toshniwal: See, for example, I showed you this thing, right? Let me go back. [01:44:18] Durga Toshniwal: Just a second. [01:44:24] Durga Toshniwal: So, if you look at this example, here the decision boundary is actually what? It is drawn like this. [01:44:30] Durga Toshniwal: But if you look at the data. [01:44:33] Durga Toshniwal: It is actually linearly separable only, because these are the regions where one class is there, this is the region where another class is there. [01:44:40] Durga Toshniwal: But unnecessarily to fit this point. [01:44:44] Durga Toshniwal: the decision boundary has been changed. If there are 100 points to fit 1 or 2 points, the decision boundary is changed. [01:44:50] Durga Toshniwal: So, a nonlinear decision boundary is not bad, but it should be used only when the data is truly nonlinear in nature. Say, for example, I'll give you an example where you really [01:45:04] Durga Toshniwal: should be having a nonlinear decision boundary. Like, for example. [01:45:10] Durga Toshniwal: You have one class like this. [01:45:30] Durga Toshniwal: So, these are the two classes that you are having. One shown in the pluses, and one shown in the circles in this particular data space, which is shown in this two-dimensional box. Now, if we were to draw a decision boundary between this, it has to be like this. [01:45:46] Durga Toshniwal: It cannot be a linear line at all, because here, definitely, the two classes are nonlinearly separable. They cannot be, separated by a linear boundary. [01:45:57] Durga Toshniwal: So, having a nonlinear boundary is not bad, but it is bad when the data is linearly separable. Then you are unnecessarily complicating the decision boundary like this, and this is the overfitting case. However, this is not overfitting. It is just correct. [01:46:12] Durga Toshniwal: It's okay. [01:46:13] Deepak Bobade: Just correct. Yeah. [01:46:15] Deepak Bobade: So, is it like we'll have to draw a scatterplot out of the data samples, and then we'll have to validate it manually? Is it what we're gonna do? [01:46:25] Deepak Bobade: So, it's always a good idea to visualize the data, try to visualize it, and for that, if your data is having multiple dimensions. [01:46:33] Durga Toshniwal: then you reduce it using PCA, and there's another thing similar to PCA, which is called TCNE plot, which I'll show you in the hands-on. So, you can use some of these mechanisms to [01:46:46] Durga Toshniwal: do a dimensionality reduction on the data, and then visualize it in 2D or 3D, and you see whether the data looks… what kind of separation is there on the data. [01:46:57] Durga Toshniwal: So that gives you an idea. [01:46:59] Durga Toshniwal: Other than that, even if you don't do that, you can find out the training error, test error, and all that, and that can also help you to assess whether your model is overfitting or underfitting, or it's right. [01:47:11] Durga Toshniwal: Which also you will see in the hands-on. [01:47:14] Deepak Bobade: Okay. And then, second question, we'll have to go to one slide back. [01:47:21] Deepak Bobade: The previous slide. [01:47:23] Deepak Bobade: previous. [01:47:24] Deepak Bobade: That, right. [01:47:27] Deepak Bobade: But I'd say… The graph was there, right? [01:47:31] Durga Toshniwal: Oh, you want to go this one? [01:47:34] Durga Toshniwal: This one. [01:47:36] Durga Toshniwal: Which graph you're talking about? [01:47:37] Deepak Bobade: A graph before this. It was about the overfitting and underfitting. [01:47:43] Durga Toshniwal: Oh, this one. [01:47:44] Deepak Bobade: This graph, yeah. [01:47:46] Deepak Bobade: So, the overfitting is happening because we have, more data points, right? If we talk about number of nodes at, between 150 to 200, that was perfect, right? But now we have extra node, which is, going towards 250, right? So that is the reason it is being overfitted. So if we just wrap up with 0 to… [01:48:08] Deepak Bobade: Whatever range in between 100 and 200… 150 and 200, right? That should be good, right? That… [01:48:15] Deepak Bobade: If we have that number of nodes, the model would be perfect. Is that what I can implement? [01:48:22] Durga Toshniwal: So, this doesn't have anything to do with the data points. [01:48:27] Durga Toshniwal: It is the depth of the tree. Let's say I have a tree, depth of the tree. [01:48:32] Durga Toshniwal: And based on the… let's say I had some, you know, 300 attributes with me, I go on using them and keep on increasing the depth of the tree, right? It will increase with the number of attributes. [01:48:46] Durga Toshniwal: I keep on increasing with the aim to keep on minimizing the training error. [01:48:51] Durga Toshniwal: So, when I keep on doing that, the tree becomes overly deep, and because it is overly deep, therefore, it is no longer generic. [01:48:59] Durga Toshniwal: Right? It will give zero error or lower error only on that kind of data. But if some generic data will come, it will give higher error, right? It's too much tailor-made on the training data only. [01:49:12] Durga Toshniwal: So, we don't want that. We want it to be… [01:49:15] Durga Toshniwal: You know, we don't want the training error to be overly low. It should be generalizable to some limit. So, in this case, the test error is very high. We don't want that, because that's the error on the unseen data. We want both of them to be optimally low. So, here it is not the lowest, but it is low. [01:49:34] Durga Toshniwal: And at the same, point, where the number of nodes is maybe 180 or something, maybe something like this, 180, 190, something like this. [01:49:43] Deepak Bobade: 90, yeah. Yeah, so then… [01:49:45] Durga Toshniwal: In this case, the test error is not the highest, it's not the lowest, but it is optimally low. [01:49:52] Durga Toshniwal: So that's what we look forward to. [01:49:55] Deepak Bobade: Okay, so whatever the number we get, we try to cap the depth of the tree. [01:50:02] Durga Toshniwal: Exactly, exactly. [01:50:03] Deepak Bobade: Okay. [01:50:04] Deepak Bobade: Okay, thank you, that was all. [01:50:08] Durga Toshniwal: Any other question, anyone? [01:50:15] Durga Toshniwal: bubble. [01:50:17] Durga Toshniwal: So, probably from the next week onwards, you might be having a… either next week or next to next week. [01:50:25] Durga Toshniwal: As I said, the class would be 10 to 12. [01:50:28] Durga Toshniwal: probably it might take one more week, and then from, in the… on Saturdays, you'll still have 6 to 8, and on Sundays, you might have 10 to 12, starting the week after the next week. [01:50:41] Durga Toshniwal: So, I think with this, I'd like to wrap up for today. If there are any further questions, you can let me know. [01:50:48] Durga Toshniwal: And, there's a possibility that I might get delayed in sharing the slides with you. Probably I'll share them tomorrow, because I have some… I have to… I have, to travel somewhere, so… [01:51:02] Durga Toshniwal: That I'm already telling you. I'll try to share them today, but if that may not be possible, then I'll do it tomorrow. [01:51:09] laxmi sahu: Oh, mom, I have a question regarding the curriculum. [01:51:14] Durga Toshniwal: Yeah. [01:51:14] laxmi sahu: So that… that can be shared, then it gives us a prior idea, then that would be helpful for us to organize other things. [01:51:24] GenAI Batch-2 Manager: Lakshmi, just adding here, in your Module 1, under the General section, the curriculum has already been added on the LMS. [01:51:32] laxmi sahu: Okay, thank you. [01:51:35] Chandrasekhar Sahu: Ma'am, can we have class Saturday morning as well? Like… [01:51:41] Chandrasekhar Sahu: Instead of emailing, is it possible? [01:51:44] Durga Toshniwal: I don't think that may be possible because of the fact that Saturday mornings, many a times at my end. [01:51:52] Durga Toshniwal: I have classes, examinations, and other things, so for that [01:51:57] Durga Toshniwal: Becomes a, you know, that may become a little difficult. [01:52:02] Chandrasekhar Sahu: Okay. [01:52:04] Durga Toshniwal: And mostly, some… also, some of you might be having offices, because not everyone would be having 5 days and a week. [01:52:13] Durga Toshniwal: So also, it's good to have it in the evening for those who are having offices. So, both ways, it works better. That's the thing. [01:52:22] Chandrasekhar Sahu: Okay, okay, okay. [01:52:24] Durga Toshniwal: How can I get the previous class, slides? [01:52:27] SHUBHAM GOSWAMI: So, I need to download it, using LMS, or the future provide the slides? [01:52:36] GenAI Batch-2 Manager: all of it is uploaded on the LMS already. [01:52:42] SHUBHAM GOSWAMI: Okay, okay, thanks. [01:52:44] Durga Toshniwal: Yeah, and Lokesh, I think your question is already answered, so I'm not taking it separately. [01:52:52] Durga Toshniwal: Okay, then I think if there are no further questions, then we can break for today. Thank you all, have a great day. [01:52:58] Nirav Mehta: Thank you. [01:53:01] Krishnakumar MS: Thank you. [01:53:02] Durga Toshniwal: Thank you, bye-bye. [01:53:07] Durga Toshniwal: Thank you. Bye-bye.