# 01 2025-11-29 Supervised Learning course: Module 2 — Machine Learning Algorithms module: Module-2-Machine-Learning-Algorithms date: 2025-11-29 type: transcript video_url: https://personal-learn.armco.dev/files/_Recordings/Module-2-Machine-Learning-Algorithms/01_2025-11-29_Supervised_Learning.mp4 --- [00:06:56] aditya shrivastava: Hello, everyone! [00:11:22] Neeraj Kumar: Hi, good afternoon, all of you. [00:11:38] Abhishek Mhatre: Everybody join you up? [00:14:05] Durga Toshniwal: A very good evening to all of you, and welcome to today's session. [00:14:10] Durga Toshniwal: So, I'll just share my slides, and we'll go on from there. [00:14:16] Durga Toshniwal: Allow me a minute, please. [00:14:25] Durga Toshniwal: Just a second. [00:15:30] Durga Toshniwal: Okay, so I hope my screen is visible and I'm audible. [00:15:35] Durga Toshniwal: So… In the last, we had discussed, different similarity measures. [00:15:43] Durga Toshniwal: We also discussed what are the different kinds of attributes, like discrete and continuous attributes. Then we talked about similarity measures. [00:15:53] Durga Toshniwal: We discussed the properties of a distance measure, such as symmetry, self-similarity, positivity, and triangular inequality. [00:16:04] Durga Toshniwal: Which any metric should satisfy for it to be called a measure. [00:16:10] Durga Toshniwal: So, self-similarity means that [00:16:13] Durga Toshniwal: The distance of any point with respect to itself should be greater… should be equal to zero. Positivity means distance will always be positive or zero. [00:16:24] Durga Toshniwal: Triangular inequality means that distance of [00:16:27] Durga Toshniwal: A to B and A to C will be greater than or equal to distance between B and C. [00:16:36] Durga Toshniwal: And so and so forth. We talked about the measured Euclidean distance. [00:16:41] Durga Toshniwal: And we took some examples. We talk about the generalization of Euclidean distance. [00:16:47] Durga Toshniwal: Which is, Minskow Whiskey Distance. [00:16:50] Durga Toshniwal: We also talked about Hamming on Manhattan distance, which is the L1 norm. Euclidean distance is called the L2 norm. [00:16:57] Durga Toshniwal: And these two are specific cases of the generic distance measure called Minskovic key distance. [00:17:04] Durga Toshniwal: Then we talked about similarity between binary vectors, using simple matching coefficient, as well as checkout. [00:17:12] Durga Toshniwal: We talked about cosine similarity. [00:17:15] Durga Toshniwal: Then we discussed what is confusion matrix, which consisted of, [00:17:22] Durga Toshniwal: the actual versus the predicted class labels, in which there are four values, true positives, two negatives, false positives, and false negatives. These are the four entries inside the confusion matrix. [00:17:37] Durga Toshniwal: And based on that, we have certain measures, like precision, recall, F1 score, and, [00:17:47] Durga Toshniwal: False positive rate, false negative rate, and so on and so forth. And we can find this out with the help of confusion matrix. [00:17:57] Durga Toshniwal: So, for example, precision here talks about the, [00:18:02] Durga Toshniwal: Actual true positives that have been predicted. [00:18:06] Durga Toshniwal: Divided by… the total number of predicted positives. So, in this particular case, the denominator actually includes [00:18:15] Durga Toshniwal: The, predicted class labels. [00:18:20] Durga Toshniwal: So, the total predicted positive class labels, which is true positive plus false positive. [00:18:27] Durga Toshniwal: Then, sensitivities also, same as precision. [00:18:39] Durga Toshniwal: Then we have… [00:18:44] Durga Toshniwal: We have recall. In recall, the focus was on the actual positives. [00:18:50] Durga Toshniwal: Not just on the predicted positive, it was on the actual positives. [00:18:54] Durga Toshniwal: So here, we had the actual two positives, that are there, that are predicted, divided by the actual positives that are [00:19:06] Durga Toshniwal: Available in the given data. [00:19:08] Durga Toshniwal: Actual positives means true positives plus positives that are predicted as negatives, which means false negatives. [00:19:15] Durga Toshniwal: So, so, this kind of measure, which is recall, is very important. [00:19:22] Durga Toshniwal: Because, [00:19:25] Durga Toshniwal: In this particular case, we want the two positives to be high, or in other words, in the denominator, false negatives should be low. [00:19:34] Durga Toshniwal: If false negative is 0, then recall is 100%, or equal to 1. And what is false negative? False negative means, for example, a person who's having a disease, but is declared healthy. [00:19:47] Durga Toshniwal: So, example that I had given was a person having a corona, what is, tested as negative and declared healthy, and then he or she mixes with other healthy people and infects them. [00:20:04] Durga Toshniwal: Then we have, So, I already explained to you precision and recall. [00:20:14] Durga Toshniwal: So I give you an example where recall is important. Precision is also very important, and we want to have high precision. [00:20:23] Durga Toshniwal: For which false positives should be equal to zero. [00:20:26] Durga Toshniwal: False positive is equal to zero means. [00:20:29] Durga Toshniwal: Precision will be equal to 100%, or will have a value 1. [00:20:34] Durga Toshniwal: And false positive means something which is not positive, but is declared as positive. [00:20:41] Durga Toshniwal: So, for example, a healthy person is declared as diseased. [00:20:46] Durga Toshniwal: So, in this particular case, it can happen that a healthy person is declared a disease, and he or she may have to take a lot of medicines unnecessarily. So, it's also important to have false positives as zero. [00:21:01] Durga Toshniwal: Then we have, the other things, such as… Oh… [00:21:10] Durga Toshniwal: Yep. [00:21:11] Durga Toshniwal: Then we had discussed the trade-off between precision and recall, and we had discussed that [00:21:17] Durga Toshniwal: Precision and recall have an inverse relationship and cannot be high at the same time. If precision is high, then recall is going to be low. If recall is high, then precision is going to be low. [00:21:35] Durga Toshniwal: So, [00:21:42] Durga Toshniwal: Then, [00:21:47] Durga Toshniwal: This is all about the trade-off that is there. [00:21:51] Durga Toshniwal: Then I also told, discussed that precision and recall, because both of them are very important. [00:21:58] Durga Toshniwal: Therefore, we want both of them to be higher, but because they are in an inverse relationship. [00:22:04] Durga Toshniwal: with each other. Therefore, both cannot be high at the same time, and we definitely want, [00:22:10] Durga Toshniwal: a balanced score to measure them, which is F1 score, which is the harmonic mean between precision and recall, and it is equal to 2 times precision recall divided by precision plus recall. [00:22:24] Durga Toshniwal: Then, we have accuracy. Accuracy is something. [00:22:28] Durga Toshniwal: That is how much correctness is there in the model. So, correctness refers to the correct positives [00:22:36] Durga Toshniwal: and the negatives both. So, in this particular case, we have True positives plus two negatives. [00:22:44] Durga Toshniwal: plus, divided by some total of all, which is TP plus, that is true positive plus false positive plus true negative plus false negatives. [00:22:57] Durga Toshniwal: Then, [00:23:02] Durga Toshniwal: So if we want the accuracy to be high, then in that particular case, both false positives and false negatives should be zero. [00:23:09] Durga Toshniwal: If both are zero, then definitely accuracy would be high. [00:23:20] Durga Toshniwal: Then we had discussed sensitivity and recall. [00:23:25] Durga Toshniwal: So, sensitivity and recall, as I already told you, are the same things. I already discussed recall, which focuses on true positives divided by true positives plus false negatives. [00:23:41] Durga Toshniwal: Then, we have something called specificity. [00:23:46] Durga Toshniwal: Specificity means two negatives. [00:23:49] Durga Toshniwal: divided by 2 negative plus false positives. So here, the focus is on how… how many true negatives are actually announced as true negatives, and again, false positives should be zero. If they are zero, then specificity will be 1. [00:24:10] Durga Toshniwal: Then we have false positive rate and false negative rate. [00:24:13] Durga Toshniwal: I discussed false positive rate is the number of false positives. [00:24:18] Durga Toshniwal: Divided by 2 negative plus false positives. [00:24:23] Durga Toshniwal: So, here, when we are talking about false positives, we are actually talking about [00:24:29] Durga Toshniwal: the actual, negatives that are there. So, we have false positives divided by 2 negatives, plus false positives. [00:24:36] Durga Toshniwal: That is the false positive rate. [00:24:38] Durga Toshniwal: False negative rate will be false negatives. [00:24:41] Durga Toshniwal: divided by… [00:24:43] Durga Toshniwal: True positive plus false negative. So this is the rate at which false negatives are present in the data, so it will be false negative. [00:24:51] Durga Toshniwal: Divided by 2 positive plus false negative. [00:24:57] Durga Toshniwal: Then I discussed how sensitivity and specificity are related. [00:25:01] Durga Toshniwal: And, suppose this is the threshold, or the decision boundary. [00:25:06] Durga Toshniwal: And anything to the right of it is true positive, and left to… it will be true negative. So as we vary the decision boundary, we found out [00:25:15] Durga Toshniwal: How the relationship between two positive and two negative changes. If we move it completely towards the left side, everything becomes too positive, and such a test is useless, because everyone is positive, no one is negative. [00:25:30] Durga Toshniwal: Similarly, if it is moved towards the right side, then everything becomes too negative. [00:25:37] Durga Toshniwal: So if everything is too negative, means, again, the test is useless, because no one is positive, everyone is negative only. [00:25:47] Durga Toshniwal: So, 100% sensitivity and specificity. [00:25:52] Durga Toshniwal: Are not practically possible, and when one increases, the other decreases, so they have a… [00:25:59] Durga Toshniwal: Inverse relationship amongst them. [00:26:03] Durga Toshniwal: And then we also talked about ROC curve, and I showed to you [00:26:09] Durga Toshniwal: the situation where the ROC curve, is ideal. And ROC curve, is the curve between the two positive rate drawn [00:26:19] Durga Toshniwal: Versus false positive rate. [00:26:21] Durga Toshniwal: So, when we talk about the ROC curve, we are actually… ROC is an acronym standing for Deceiver Operating Characteristic. [00:26:29] Durga Toshniwal: However, this full form doesn't carry much meaning or value now. We… it's generally just used as a notation. [00:26:39] Durga Toshniwal: And then we have area under curve. The higher… the better the, the prediction is, the higher will be the area under the curve. Ideally, it will be 0.921. So, this is what we had discussed. [00:26:54] Durga Toshniwal: In the last turn. [00:26:56] Durga Toshniwal: Now, we'll start with today's portion. [00:27:00] Durga Toshniwal: And before that, just allow me a second, I'm just trying to… I just… [00:27:08] Durga Toshniwal: Just give me a minute, please. [00:27:28] Durga Toshniwal: So now, I'll be sharing today's, slides. [00:27:46] Durga Toshniwal: Just allow me mute. [00:28:13] Durga Toshniwal: Just a second, please. [00:28:29] Durga Toshniwal: Oh. [00:28:30] Durga Toshniwal: So today, we'll be starting a very interesting and extremely important topic, which is, unsupervised learning. [00:28:39] Durga Toshniwal: So, by the way, anyone, any questions on the last one? [00:28:45] Durga Toshniwal: part. [00:28:52] Durga Toshniwal: Any questions, anyone? [00:28:59] Durga Toshniwal: So, if there are no questions, we'll start today's portion, which is on supervised learn. [00:29:04] Durga Toshniwal: So, [00:29:06] Durga Toshniwal: So today, we'll be first discussing what is supervised learning, then we'll be discussing the decision tree algorithm, which is a supervised learning, very popular supervised learning algorithm. [00:29:18] Durga Toshniwal: So, what is supervised learning? So, supervised learning, is called supervised learning. [00:29:25] Durga Toshniwal: Because in this particular case, we are given the class labels. We know the… what are the class labels available in the given data, and because we have this information, therefore, this kind of an algorithm [00:29:39] Durga Toshniwal: Falls under the category of supervised learning. [00:29:42] Durga Toshniwal: Often, supervised learning, [00:29:45] Durga Toshniwal: is used along with classification and prediction. So, we develop a classification model and use it for making a prediction. So, usually, it is used in predictive analytics, supervised learning, and classification and prediction. So, we'll be discussing a couple of, [00:30:04] Durga Toshniwal: Algorithms for, supervised learning. [00:30:08] Durga Toshniwal: So, before we go into the details of how we can do supervised learning, first of all. [00:30:14] Durga Toshniwal: We should understand what is classification, and so what is the basic idea of classification? So, just a second, let me just grab my pen. Allow me a moment to… [00:30:31] Durga Toshniwal: Just a second, please. [00:30:45] Durga Toshniwal: Just a sec. [00:30:47] Durga Toshniwal: Sticky. [00:30:48] Durga Toshniwal: Different time. [00:31:10] Durga Toshniwal: So, what is the task of classification? [00:31:13] Durga Toshniwal: So, first of all, in classification, we have a special data called the training data. [00:31:19] Durga Toshniwal: This training data is special because, in addition to the different attribute values, or the field values, it also contains a column called the class label column. [00:31:34] Durga Toshniwal: And in that, the value of the class labels is given. So, we have the different attributes, which are also called independent variables, or inputs, or the predictors. [00:31:47] Durga Toshniwal: So, these are all meaning the same thing. So, we have a set of attributes. [00:31:51] Durga Toshniwal: These are used for making the predictions, so they are called predictors. They are also called the input variables, independent variables and all. And we have Y, which is called the class label, and this is the dependent variable because it is, [00:32:07] Durga Toshniwal: Derived from the independent variable or the input, and this is nothing but the output. [00:32:13] Durga Toshniwal: So, now, what is the task of classification? [00:32:17] Durga Toshniwal: The task of classification is actually, by using the training data, or the attributes in the training data, we are supposed to obtain a model [00:32:31] Durga Toshniwal: In such a way that this model predicts the class label, [00:32:36] Durga Toshniwal: Correctly for the, given data. [00:32:40] Durga Toshniwal: So, so the whole process of classification is nothing but mapping a set of attributes to one of the, you know, values of the class label. So, very generally, we can talk about [00:32:56] Durga Toshniwal: Classification, as finding out this function X, where, function F, [00:33:06] Durga Toshniwal: So, X is the set of, inputs. [00:33:10] Durga Toshniwal: or the attributes? [00:33:12] Durga Toshniwal: That are there. [00:33:14] Durga Toshniwal: Y is the class label? [00:33:16] Durga Toshniwal: That we will predict. [00:33:19] Durga Toshniwal: And F is nothing but the classification function. [00:33:26] Durga Toshniwal: That we are supposed to find out. So, if the function is properly found out, then whenever we give the input. [00:33:34] Durga Toshniwal: this input, which is the set of attributes given to us. So, and this function, when it receives the input values, that is the attribute or the predictor values, then it will find out the value of Y, [00:33:51] Durga Toshniwal: And this is nothing but the class label that we require. So the whole task of classification is nothing but learning the given data set of attributes X to obtain the correct class label Y, and in that process, finding out the function F. [00:34:07] Durga Toshniwal: So, the function f could actually be a linear function or a nonlinear function. It could be any kind of function, right? [00:34:15] Durga Toshniwal: Now the process of classification involves building the model, and then using it for making the prediction. Say, for example, if we are using a decision tree. [00:34:30] Durga Toshniwal: Just take a second. [00:34:37] Durga Toshniwal: Missing some issues with my… [00:34:43] Durga Toshniwal: So, for example, if we are having a decision tree. [00:34:47] Durga Toshniwal: It could be anything. I'm just drawing some… one tree. [00:34:53] Durga Toshniwal: Just for the purpose of… [00:34:56] Durga Toshniwal: an example. And let's say these are the leaf nodes, and this is one particular class type. I'm denoting this. Let's say there are two classes. [00:35:07] Durga Toshniwal: Class 1 and Class… Zero. [00:35:10] Durga Toshniwal: So, this kind of a shading indicates Class 1. [00:35:14] Durga Toshniwal: And, empty, box or a node represents class 0. So we have class 1 here, class 0 here, and maybe… [00:35:24] Durga Toshniwal: Class 1 here and class 0, like that. [00:35:27] Durga Toshniwal: And these… these are the edges. [00:35:30] Durga Toshniwal: Like this. [00:35:32] Durga Toshniwal: Now, this tree is actually built with the help of [00:35:38] Durga Toshniwal: The training data, which is provided to us. [00:35:41] Durga Toshniwal: And this data is special because it does not just have the attribute values, it also has the class label given to us. Now, once this model is built, let's say we have an unseen record. [00:35:54] Durga Toshniwal: So, we have this unseen record. [00:35:58] Durga Toshniwal: So this record has the values of the attributes, that is, it has the values for the settings, and it wants to find out the class label. So, what we'll do is that, with the help of these values, we are going to map this set of values onto the tree. Let's say, for example, we start from here. [00:36:18] Durga Toshniwal: Right? And then, let's say it traverses this path. After that, let's say it traverses this path, and it arrives at the leaf node. [00:36:26] Durga Toshniwal: Then the class label that we are predicting, Is this… [00:36:32] Durga Toshniwal: Which is nothing but your Class 1. [00:36:35] Durga Toshniwal: So then, for the unseen record. [00:36:38] Durga Toshniwal: The value of the class label will be 1. [00:36:40] Durga Toshniwal: Similarly, it could be zero, or it could be something. I'm just very simply showing you how a classification model can be used for the purpose of prediction. [00:36:53] Durga Toshniwal: So now, so, first of all, in order to use a model, what we will need is [00:37:03] Durga Toshniwal: That we'll need to build a model, right? [00:37:05] Durga Toshniwal: And to build a model, we will need, first of all, The… Predefined class labels. [00:37:14] Durga Toshniwal: So the class label should always be given to us, or somehow known to us when we start with this process. [00:37:20] Durga Toshniwal: And we have this set of samples, which is our training data. [00:37:25] Durga Toshniwal: And this training data is actually learned by the classifier, and once the data is learned, then it can be used to make prediction. How it will make the prediction? [00:37:40] Durga Toshniwal: So, the input data, or the training data, must be… have… it must have all kinds of use cases which, which the classifier is required to predict in future. [00:37:56] Durga Toshniwal: Right? If there are certain use cases that are not contained in the training data, and so the classifier will not learn them, and as a result, if the unseen data is similar to some of these use cases which have not been learned. [00:38:13] Durga Toshniwal: By the classifier, and let's say the unseen, data [00:38:19] Durga Toshniwal: Falls is similar to one such kind of [00:38:25] Durga Toshniwal: You know, use case that is not learned, then what will happen is that when it encounters such a data, it fails, because it doesn't know what to do. [00:38:35] Durga Toshniwal: Because the process of classification involves learning from the data. The data that it learns is the training data. It's also sometimes called the historical data. [00:38:47] Durga Toshniwal: Because it contains the attribute values, and it also contains the class labels. And with the help of that, we learn a model. [00:38:55] Durga Toshniwal: And this learning actually involves finding out a correct function f that maps the input, which is the attributes to the output, which is the class label, in the most… in the best possible fashion. [00:39:11] Durga Toshniwal: So, very crudely, when I try to explain, a very crude example of classifier is that of a talking parrot. [00:39:17] Durga Toshniwal: So, a talking parrot is, so what do you do when talking parrots usually memorizes or learns certain set of questions and their answers? [00:39:30] Durga Toshniwal: So now, if you keep asking the same similar question. [00:39:35] Durga Toshniwal: Then the parrot will just respond with the answer that it has learned. [00:39:39] Durga Toshniwal: But… In case… [00:39:41] Durga Toshniwal: You ask a new question, then the parrot doesn't know what to do, because the parrot hasn't learned that question at all. [00:39:49] Durga Toshniwal: And, because the response of the parrot purely depends on what it has learned, as a result, a new question, when is encountered by the parrot, the parrot fails. It doesn't know what to do. It just responds in an ambiguous fashion, or doesn't respond at all. [00:40:05] Durga Toshniwal: So, very similar to this is the classifier. You make it learn all the use cases that it is supposed to predict in future. [00:40:15] Durga Toshniwal: And if it learns all these kind of use cases, then it will be able to predict them whenever similar unseen record is encountered. An unseen record is one whose class label is not known to us. [00:40:29] Durga Toshniwal: But if some of the case… use cases are missing, then such a prediction cannot be made by the classifier. [00:40:39] Durga Toshniwal: So it sounds a little weird that, you know, as humans, we always have cognitive powers, we visualize, we understand, and even if we haven't learned something, we can still [00:40:51] Durga Toshniwal: Info. [00:40:53] Durga Toshniwal: Or we can ponder and infer. However, computer systems and algorithms, they do not have brains, so to say. [00:41:00] Durga Toshniwal: And they cannot… Think on their own. [00:41:04] Durga Toshniwal: They are made to think something which they have learned, and that's all they can do. [00:41:09] Durga Toshniwal: Right, so this is about the classification process. Let me elaborate some more details about the function f. So I told you. [00:41:18] Durga Toshniwal: Making a classifier involves giving input X, and getting output Y, and learning the function F. [00:41:26] Durga Toshniwal: And this function f is nothing but a classifier. [00:41:29] Durga Toshniwal: So, this function f could be a linear or a nonlinear, or a simple or a difficult function. Let's say, for example, we have a very simple function. I'm sure all of you read it years back, like, Y is equal to MX plus C. [00:41:45] Durga Toshniwal: So Y is equal to MX plus C is nothing but a equation of a straight line. [00:41:51] Durga Toshniwal: having an offset C, and having a slope M. [00:41:55] Durga Toshniwal: So, [00:41:57] Durga Toshniwal: The whole idea is that if we want to predict the value of Y properly, then we should be able to find the value of MNC, [00:42:05] Durga Toshniwal: Correctly. [00:42:07] Durga Toshniwal: And, the whole process of building a classifier is actually finding these values correctly, so that the correct value of Y could be found out, given value of X. So this is a very simple classifier, where we have X can be a single attribute. [00:42:24] Durga Toshniwal: Now, suppose I told you that, [00:42:28] Durga Toshniwal: The learning must be complete, and if it is not complete, then the classifier will not be able to predict what all it is supposed to predict. [00:42:38] Durga Toshniwal: So, let's say if we want to find out MNC [00:42:42] Durga Toshniwal: and I have a point, and we are talking about two-dimensional data, then I have a point like X1Y1. Now, with a single point X1Y1, [00:42:52] Durga Toshniwal: Will we be able to find out the values of MNC? [00:42:56] vinit shah: No. [00:42:58] Durga Toshniwal: Yeah, so this is correct, that if we just have a single point, X1Y1, then we cannot find MNC. [00:43:06] Durga Toshniwal: Because a single value is insufficient. So, we should have at least two values. [00:43:12] Durga Toshniwal: like X1Y1 and X2Y2 to be able to infer the values of M and C, and then find out, this function MX plus C. [00:43:22] Durga Toshniwal: So, this means that there are some minimum number of points, data points, that are required for us to be able to [00:43:31] Durga Toshniwal: Obtain the correct function. [00:43:33] Durga Toshniwal: Or map the given… so this whole process involves mapping the input to the correct output. [00:43:41] Durga Toshniwal: We can also call it mapping, where input is nothing but the set of input attributes, and output is nothing but the class label. And we need some, so training data should be enough. [00:43:54] Durga Toshniwal: It should not be very less, because if it is very less, then the correct function cannot be found, and therefore, [00:44:02] Durga Toshniwal: The classifier is not going to work properly. [00:44:07] Durga Toshniwal: So now, [00:44:09] Durga Toshniwal: Once the model is built, that is, the function f has been found out, then it will be used to predict the class label. [00:44:18] Durga Toshniwal: In future, for those records whose class labels are not [00:44:22] Durga Toshniwal: known. They are called unseen records. [00:44:24] Durga Toshniwal: And then, [00:44:26] Durga Toshniwal: Once we, you know… so while we are using the training data to find out the function f and build the classifier model. [00:44:34] Durga Toshniwal: Once the classifier model is actually, [00:44:39] Durga Toshniwal: ready, then what we will do is we'll try to see, so you know, I already told you that the training data has class labels given in it. So what we'll do is that, we will, [00:44:56] Durga Toshniwal: I have another data, which is called the test data. [00:45:00] Durga Toshniwal: So, test data is also a special data, because it not just has the attribute values, it also has the class labels given to us. So, once we find out a model, or learn the model. [00:45:14] Durga Toshniwal: For the classifier, using the training data, then we are going to find out the performance of this model that we have built using the test data. [00:45:25] Durga Toshniwal: So, as I told you, that in the test data, we have the class labels given to us, and what we will do is that we will give the attribute values, or the input in form of the rest of the attributes to the classifier model that we have built. [00:45:41] Durga Toshniwal: And then, what we'll do is that we'll do the prediction. [00:45:45] Durga Toshniwal: And then compare the predicted, values [00:45:49] Durga Toshniwal: and the given values of the class label in the test data, and see how many match. The amount of match, refers to the accuracy of the model. So, if the predicted class label and the given class label match, then the model is working correctly, and therefore. [00:46:09] Durga Toshniwal: We define accuracy, which I have already defined earlier for you, as the total number of class labels [00:46:18] Durga Toshniwal: That are predicted, so the match between the predicted class labels and the given class labels, divided by some total of all, class labels. That is what is accuracy. [00:46:34] Durga Toshniwal: So, while we are building a classifier model. [00:46:38] Durga Toshniwal: We try to assess, its performance with the help of the test data. We find out its accuracy, and if the accuracy is within the bound or within the limit that we want it to be. [00:46:50] Durga Toshniwal: Then, if it is acceptable to our model, then we can, you know, start using the classifier for making predictions in future. [00:47:04] Durga Toshniwal: So, now we are going to talk about, Building a classifier [00:47:10] Durga Toshniwal: using decision tree induction. So, decision tree is one particular algorithm, which is one of the most popular algorithms. [00:47:18] Durga Toshniwal: To, build a classifier. [00:47:22] Durga Toshniwal: It's a very, very, very popular method, and it involves learning a decision tree from the given class labels. [00:47:32] Durga Toshniwal: From the given training data that contains the class labels. [00:47:36] Durga Toshniwal: So, what is a decision tree like? A decision tree… [00:47:39] Durga Toshniwal: It's actually, because it's a tree. [00:47:44] Durga Toshniwal: So, it will definitely have a root. However, because it's not the real tree. [00:47:49] Durga Toshniwal: So, the root of a decision tree will be its topmost node. [00:47:56] Durga Toshniwal: This will be the route. [00:47:58] Durga Toshniwal: And then, from the root, will come out other notes. [00:48:04] Durga Toshniwal: These nodes we call the internal nodes. [00:48:07] Durga Toshniwal: So these are internal notes. [00:48:11] Durga Toshniwal: And then… [00:48:12] Durga Toshniwal: there might be, bifurcation, trifurcation, or whatever. So, this split may be binary, as I'm showing you. Binary because two edges are coming out of it. But it could be an array also. [00:48:26] Durga Toshniwal: Means it could have multiple edges coming out of it. [00:48:31] Durga Toshniwal: And then, like this, we build the tree. So this is the root node. Any node which is not at the, [00:48:39] Durga Toshniwal: at the last level is said to be internal node. The nodes that are at the end of the decision tree are called the leaf nodes, and these are the ones that are used to predict the class level. [00:48:53] Durga Toshniwal: So then, we have internal loads. Internal nodes are just used to test the attribute onto them. [00:49:00] Durga Toshniwal: Then, we have… The leaf node that finds out the outcome of the test. [00:49:07] Durga Toshniwal: And, the leaf node is also going to contain a class label. [00:49:14] Durga Toshniwal: Say, for example, what you are seeing here is a dataset. [00:49:18] Durga Toshniwal: This can be considered as a training data. This training data has attributes like transaction ID, refund, marital status, taxable income, and cheating. [00:49:29] Durga Toshniwal: Okay, this cheating is cheating or, filing of income tax. [00:49:34] Durga Toshniwal: Now, if we are to build a tree, let's say, for example, we are starting with this particular attribute, a refund. [00:49:43] Durga Toshniwal: Let's say, for example, we are starting with, refund. [00:49:47] Durga Toshniwal: So then, [00:49:50] Durga Toshniwal: This, when we start the decision tree with one particular attribute, we can choose the, from the rest of the attribute, some other attribute to continue drawing the tree. For example, after refund, we could do the, continue the tree on marital status. [00:50:09] Durga Toshniwal: So, refund could be yes or no. [00:50:11] Durga Toshniwal: Then followed by marital status, which could be single, divorced, or married. [00:50:17] Durga Toshniwal: For the purpose of, this particular problem, we have clubbed together single and divorced, and put them together. [00:50:24] Durga Toshniwal: In this, category. [00:50:27] Durga Toshniwal: And then we have taxable income, again, [00:50:31] Durga Toshniwal: the taxable income is actually, a continuous valued field, so… [00:50:38] Durga Toshniwal: Here, you can see refund is a categorical field. A categorical field is one that is having some [00:50:46] Durga Toshniwal: finite number of values in it, which are actually not numbers. They are values like yes, no. [00:50:52] Durga Toshniwal: And like that. So, categorical data is one which has certain categories being defined by it. I already had explained to you what is categorical data. So, refund. [00:51:03] Durga Toshniwal: Yes or no. It has two values. It is categorical in nature. Marital status could have single, married, and divorced. It is, again, categorical data. Then we have, [00:51:16] Durga Toshniwal: another attribute, which is called the taxable income. So, taxable income could take up any value, let's say, from $60,000 [00:51:27] Durga Toshniwal: to something like 220,000. So, this is called continuous valued attribute because in this particular case, from 60 up to 220,000, there could be infinite number of values, and it could be really very difficult [00:51:44] Durga Toshniwal: If we have such a big range of values, and so, therefore, what we will do is, we'll actually, discretize this particular column. I've already explained to you what is discretization, etc. [00:51:59] Durga Toshniwal: And so we'll actually, take the container's valued attributes and divide it into discrete categories, and then use those categories only. Then we have the final thing, which is the class table. [00:52:17] Durga Toshniwal: Right? And this is the one that is predicted. [00:52:20] Durga Toshniwal: So, here, what you are seeing in the figure is one particular decision tree. So, we have the refund. [00:52:27] Durga Toshniwal: Refund is, you know, so we start out with the root node. [00:52:33] Durga Toshniwal: So the root node will have, all the records incident onto it. [00:52:40] Durga Toshniwal: So here, in this particular case, we are having 10 number of records. You can see here, 1 up to 10. [00:52:46] Durga Toshniwal: So, in the starting, [00:52:49] Durga Toshniwal: Okay, so I'll discuss the building of the tree slightly later. What I want to presently show to you is one example of a decision tree where we have chosen refund as the first attribute. [00:53:01] Durga Toshniwal: Which is followed by marital status, taxable income, and all that. If refund is yes, we traverse the left side. If refund is no, we traverse the right side. [00:53:11] Durga Toshniwal: Then we use marital status, then we use taxable income, and finally, we arrive at the leave notes. [00:53:18] Durga Toshniwal: Which are indicating the class labels. The blue color nodes are indicating that a pure, that class labels can be obtained at this particular point. [00:53:31] Durga Toshniwal: So, I will explain to you why refund has been chosen, or how do we choose an attribute slightly later. As of now, you just assume that we have chosen to use refund followed by marital status, followed by taxable income. [00:53:46] Durga Toshniwal: To build the decision tree. [00:53:48] Durga Toshniwal: Now, the question is that since we have built the decision tree, is the tree going to be a unique tree over the training data? [00:53:58] azad choubey: No, ma'am, that can be changed. Actually, we can change the route as well, right? [00:54:03] Durga Toshniwal: Yes, yes. What you're saying is correct, because in the previous case, we actually chose a refund as the first particular attribute, at the root node. [00:54:13] Durga Toshniwal: However, we could have chosen some other attribute. For example, we could have chosen marital status first. This could be followed by refund and taxable income. We could also start at taxable income, then marital status refund, and so on and so forth. [00:54:27] Durga Toshniwal: So, there could be many number of trees that could be obtained from the single, training data. So, decision trees are not unique. [00:54:38] Durga Toshniwal: There could be multiple decision trees that could be built from the given data. [00:54:44] Durga Toshniwal: But… [00:54:45] Durga Toshniwal: The thing is that, though there can be multiple decision trees, however, the most optimal tree will only be one. There will not be multiple, you know, most optimal trees that could be drawn from the given date. [00:55:01] Durga Toshniwal: Sorry. [00:55:02] Durga Toshniwal: So, this is what I wanted to explain to you so far. [00:55:07] Durga Toshniwal: So, there can be multiple trees. Here, you're seeing another example. Here, it is marital status, followed by refund, followed by taxable income. [00:55:16] Durga Toshniwal: And then the prediction of the class label. Earlier, we had refund, followed by, marital status and taxable income, like that. [00:55:27] Durga Toshniwal: So I'm facing some bandwidth issues, that's why I've turned off my video, okay? [00:55:32] Durga Toshniwal: So, if it gets resolved, then I'll turn it on again. [00:55:36] Durga Toshniwal: I hope that's fine. [00:55:39] Durga Toshniwal: So then, in a decision tree, we have a root node, we have internal nodes, and we have leaf nodes. [00:55:47] Durga Toshniwal: A root node is 1, that has no incoming edges, but only outgoing edges. [00:55:53] Durga Toshniwal: So, a node that has only outgoing edges from it is said to be a root node, and this is the starting point for building a tree. [00:56:03] Durga Toshniwal: Just a second, yeah. [00:56:05] Durga Toshniwal: Then we have internal nodes. Internal nodes, [00:56:09] Durga Toshniwal: have one incoming edge, but outgoing edges could be multiple. It could be binary, or it could be an array, depending on what kind of tree we are choosing. [00:56:19] Durga Toshniwal: If you are choosing a binary tree, Then there'll be two, [00:56:25] Durga Toshniwal: Edges coming out of any particular internal node. [00:56:29] Durga Toshniwal: Whereas, the, incoming edge will still be 1. [00:56:34] Durga Toshniwal: Then, the third thing are the leaf nodes, or the terminal nodes. [00:56:38] Durga Toshniwal: So, leaf nodes are those that do not have any, [00:56:42] Durga Toshniwal: Outgoing edges, all they have are incoming edge only. [00:56:48] Durga Toshniwal: Okay. [00:56:49] Durga Toshniwal: So these are the different kinds of notes that you're, that you're having in a decision tree. [00:56:55] Durga Toshniwal: So now, I think, this will come slightly later. [00:57:01] Durga Toshniwal: Oh. [00:57:03] Durga Toshniwal: So, any questions, by the way? [00:57:06] vinit shah: Ma'am, one… one lame question I have. [00:57:09] vinit shah: Sure. So, I'm assuming that there will be no tree which has two root nodes, right? Or is there a, I don't know, something there at a later point of time where we can have two root nodes? [00:57:18] Durga Toshniwal: So, definitely, there cannot be a tree with two root nodes. Root node will only be one. [00:57:24] Durga Toshniwal: Just like a real-world tree always will have a single root only. [00:57:29] vinit shah: Cutting. [00:57:29] vinit shah: Okay. [00:57:30] Durga Toshniwal: Yeah, so the concept is very similar to that only. [00:57:34] vinit shah: Okay, thank you. [00:57:35] Durga Toshniwal: Yeah. Any other questions, anyone? [00:57:39] Durga Toshniwal: Before we continue… [00:57:43] azad choubey: Ma'am, like, how we can decide, like, which is the, like, the most optimal way, like. [00:57:49] azad choubey: The most optimal tree that you can, yeah, put it as a node, like a root node. [00:57:55] Durga Toshniwal: Yep. [00:57:55] Durga Toshniwal: So that I will be explaining to you shortly from now. [00:57:59] Durga Toshniwal: While we are trying to understand the algorithms, okay? So till then, I'll request you to just bear with me, and I will explain to you how we choose the attributes. [00:58:10] azad choubey: Sure. [00:58:14] Durga Toshniwal: Anyone? Any other question? [00:58:19] Durga Toshniwal: If there are no further questions, then let me just… share my slides. [00:58:27] Durga Toshniwal: Just a second. [00:58:41] Durga Toshniwal: Just a moment. I had done a stop shares. [00:58:54] Durga Toshniwal: Just maximizing it. [00:59:09] Durga Toshniwal: So now, how do we build a decision tree? There are definitely lots of algorithms that are available that can help us to build the decision tree. However, one of the most basic algorithms [00:59:21] Durga Toshniwal: that, has been used to build that decision trees, the Hunt's algorithm. [00:59:26] Durga Toshniwal: Okay, how… what is the Hansa algorithm? [00:59:29] Durga Toshniwal: So, in Hunt's algorithm, and, let's say we have, [00:59:34] Durga Toshniwal: Just a second. So we have, [00:59:37] Durga Toshniwal: Paining data that you are seeing. [00:59:40] Durga Toshniwal: On the upper right corner, it has got 10 transactions in it. [00:59:45] Durga Toshniwal: And this is similar to the dataset that we just had seen. [00:59:49] Durga Toshniwal: In a different perspective. So, we have this refund here, we have marital status, we have taxable income, and we are cheating on filing of income tax, yes or no? [01:00:00] Durga Toshniwal: Okay. [01:00:01] Durga Toshniwal: So now, the whole idea of one's algorithm is to keep partitioning the given set of training records into more and more pure subsets, and this is done in a recursive fashion. [01:00:15] Durga Toshniwal: So, how do we, obtain the tree? First of all, At the starting. [01:00:23] Durga Toshniwal: what we do is that we try to find out what is the majority class label. Here, if we look at these records, we can see no is there in 1, 2, 3, 4, 5, 6, 7. [01:00:36] Durga Toshniwal: So, 7 records out of 10 are having, no as the majority class label, the others are having, yes, and they are minority class label. So, we start out with the majority class label. [01:00:51] Durga Toshniwal: At the root node, okay? [01:00:54] Durga Toshniwal: Now, if, [01:00:57] Durga Toshniwal: So we start out with all the records at the root node, no matter whether they are having a yes as the class label, or no as the class label, okay? [01:01:06] Durga Toshniwal: As per the Hunt's algorithm. [01:01:08] Durga Toshniwal: And, then we try to see whether the class labels are all same or they are different. So if we look at all these 10 records, we can see that the class labels are not same. They are different because [01:01:22] Durga Toshniwal: For the first four, records, we are having class label no, but for this one, the fifth one, we are having class label yes. Then again, for the, [01:01:35] Durga Toshniwal: This one, the 8th one, then the 10th one, we are having class label, yes. So… [01:01:40] Durga Toshniwal: Because the class label, all the class labels in all the training samples is not the same. [01:01:47] Durga Toshniwal: Therefore, and it is a mixture of yeses and nos, therefore, what we are going to do, we are going to actually, from the root node, where all these records of our incident, we are going to perform splits, okay? [01:02:03] Durga Toshniwal: And, the idea of doing a split is to obtain a purer subset of the given data. So I'll explain to you more as we go along. [01:02:13] Durga Toshniwal: Okay? So… Boom. [01:02:22] Durga Toshniwal: Okay, so we had started out with refund here. [01:02:28] Durga Toshniwal: Refund, actually, [01:02:30] Durga Toshniwal: So, this is the first attribute that we have chosen. Why we have chosen, or how to choose, I will explain to you as we go along. Right now, you just assume that you have chosen, refund as the first attribute to split the… to do the split. [01:02:46] Durga Toshniwal: So, when we use refund, then what happens? [01:02:49] Durga Toshniwal: Because refund has got two values, yes and no. So, if refund is yes, we have transaction ID 1. [01:02:57] Durga Toshniwal: Otherwise, we have transaction ID 4. [01:03:00] Durga Toshniwal: Then we have got transaction ID. [01:03:03] Durga Toshniwal: 7. [01:03:04] Durga Toshniwal: So, these are the transactions. [01:03:07] Durga Toshniwal: For which… [01:03:09] Durga Toshniwal: When we are, having the value, when we are trying to see the value of refund, it is yes. [01:03:15] Durga Toshniwal: So, there are just 3 records that are having value, yes, for the refund. [01:03:22] Durga Toshniwal: Whereas all others are having value no. This means that the records at the root node are not having the same class labels. Therefore, we will need to perform a split on the given data, okay? [01:03:37] Durga Toshniwal: That is one thing, that we'll have to do a split on the… [01:03:41] Durga Toshniwal: given data, with the objective that after the split. [01:03:47] Durga Toshniwal: Whatever we obtain will be a clearer, I mean, a purer subset of What we started out with. [01:03:54] Durga Toshniwal: So now, because, we started out with this data that you are seeing here, and we tried to verify with refund yes or no. [01:04:04] Durga Toshniwal: So, if we see here, what will happen? [01:04:07] Durga Toshniwal: So, 3 records will go this side, those which are having yes, yes, and yes, whereas all of them, one, record number 2, 3, 4, sorry, 5, 6, [01:04:20] Durga Toshniwal: And then 8, 9, and 10. [01:04:23] Durga Toshniwal: So, there are, 7 records that are having a no there. [01:04:29] Durga Toshniwal: So… 3 records. [01:04:32] Durga Toshniwal: We'll traverse this path of a refund equal to yes, and 7 of them will follow [01:04:38] Durga Toshniwal: Refund equal to no. [01:04:40] Durga Toshniwal: Okay. [01:04:41] Durga Toshniwal: So now, if we look at all the records with refund equal to yes, we can see here [01:04:48] Durga Toshniwal: that, these are the records, and we don't need to do anything else. With all the yeses, there's the same class label, which is when refund is yes, cheating is no, refund is yes, cheating is no. Refund is yes. [01:05:03] Durga Toshniwal: Then cheating is no, which means that whenever the value of refund is yes, as per this particular training data set. [01:05:10] Durga Toshniwal: the value of… Cheating will always be equal to what? When it is yes, cheating is no. [01:05:19] Durga Toshniwal: Okay. [01:05:20] Durga Toshniwal: So… so, this is… don't cheat out here. [01:05:24] Durga Toshniwal: Then, we'll consider the rest of the samples. So, refund equal to yes is done. Refund equal to no. [01:05:32] Durga Toshniwal: So, refund equal to no is 2, 3. [01:05:37] Durga Toshniwal: 5, 6… 8, and 9. These are the records that are having A value of, No. [01:05:48] Durga Toshniwal: for the refund. [01:05:50] Durga Toshniwal: Now, if you look at these, some of them… [01:05:54] Durga Toshniwal: Those which are having default equal to no. [01:05:57] Durga Toshniwal: Then, out of these, not all records are still having the same class label. As you can see here. [01:06:04] Durga Toshniwal: Let me just, C. [01:06:08] Durga Toshniwal: Say, for example, this record number 8. [01:06:12] Durga Toshniwal: You can see here that for record number 8, [01:06:18] Durga Toshniwal: What is happening? Here, when we are having refund equal to no, the class label is yes. [01:06:25] Durga Toshniwal: Right? Whereas, and similarly, for record number 10, when we are saying refund equal to no, class label is yes. [01:06:34] Durga Toshniwal: Whereas for all others, like record number 2, 3, Then, 4… And then… Like that, if we see… [01:06:46] Durga Toshniwal: So, what… what I want to say here is that with refund equal to no. [01:06:52] Durga Toshniwal: We are still having a mix of class labels. [01:06:55] Durga Toshniwal: So, we are still having the 7 records, some records are having yes. [01:07:00] Durga Toshniwal: Cheating equal to yes, some are having cheating equal to no. [01:07:04] Durga Toshniwal: With, refund equal to yes. [01:07:07] Durga Toshniwal: all of them are having cheating equal to no. So, so therefore, this part is fine, but when we talk about refund equal to no, then 7 samples go towards this side, and out of these 7 samples, all of them don't have the same class label, it's still a mix. [01:07:27] Durga Toshniwal: So, since it is still a mix, then we'll need to use another attribute that is given to us. [01:07:34] Durga Toshniwal: To do another split, so as to obtain a purer subset. [01:07:39] Durga Toshniwal: So now, let's choose the next, possible attribute. [01:07:44] Durga Toshniwal: And that is, let's say, marital status. [01:07:47] Durga Toshniwal: So, now, after refund equal to… [01:07:51] Durga Toshniwal: No, we have marital status. The marital status in this particular case, you can see here how many values it's having. Single, married. [01:08:02] Durga Toshniwal: divorced. [01:08:04] Durga Toshniwal: So these are the three values that it is having. [01:08:07] Durga Toshniwal: Let's say that we are clubbing single and divorced, and [01:08:13] Durga Toshniwal: So, let's say that we are clubbing, single and divorced together, and we are keeping marriage separate. [01:08:19] Durga Toshniwal: So now, when we use the next attribute, marital status, so we have refund equal to yes or no. [01:08:25] Durga Toshniwal: With refund equal to no, we use the next possible attribute, let's assume it's marital status. [01:08:32] Durga Toshniwal: And then, based on that. [01:08:35] Durga Toshniwal: we again do a binary split. So, marital status, you can see here, has actually got 3 values, single, married, and divorced. We don't take them as 3 values. We clog two of them together and, keep the third one, as a separate one. [01:08:53] Durga Toshniwal: So, single, divorced are clubbed together, married is kept separately. And when we do this, then, [01:09:01] Durga Toshniwal: When marital status is equal to single or divorced, then cheating will always be equal to yes. [01:09:08] Durga Toshniwal: Now, let's look at the other side, which is marital status equal to married. [01:09:14] Durga Toshniwal: Okay, so now what we are going to do is, again, look at all the records that are going to go in this direction. So let me just write here, we had, 10 records here. Out of 10 records, 3 went this side, and then with 7 went this side. Out of these 7, [01:09:32] Durga Toshniwal: We check the marital status, single or divorced. How many are there? [01:09:38] Durga Toshniwal: With refund, no, refund, no, refund, no. Refund, no. [01:09:46] Durga Toshniwal: Okay. [01:09:47] Durga Toshniwal: So now, out of these, we have to check how many are having marital status, single or divorced. So this one is having single. [01:09:55] Durga Toshniwal: Then, this one is having single. [01:09:59] Durga Toshniwal: then… Single, single. [01:10:05] Durga Toshniwal: Then… We have divorced here. [01:10:11] Durga Toshniwal: Then we have single, then we have married, like that. [01:10:15] Durga Toshniwal: So, so, what we can observe is that when we use the marital status. [01:10:22] Durga Toshniwal: And, we do the spread based on the marital status, like loving single and divorced together and married separately. [01:10:30] Durga Toshniwal: We had already obtained, The class label here. [01:10:37] Durga Toshniwal: With marital status single or divorced, as cheating equal to yes, whereas with marital status equal to married. [01:10:47] Durga Toshniwal: you know, what we'll have is all the values are having the same class label. So our objective is that whenever we are doing the split, all the instances that are going to, you know. [01:11:01] Durga Toshniwal: That are getting split will have the same class label. [01:11:05] Durga Toshniwal: So, what you can see here, for marital status equal to married, when refund was no. So, you look at refund no, this is refund no, marital status married. [01:11:17] Durga Toshniwal: And what is happening? Cheating is no. [01:11:20] Durga Toshniwal: So, for one record, it is no. [01:11:24] Durga Toshniwal: Then, again, we'll have cheating equal to no, this is not married. [01:11:29] Durga Toshniwal: Cheating equal to no. Here, this is another record. [01:11:33] Durga Toshniwal: And marital status is married, and cheating is again no. [01:11:38] Durga Toshniwal: And then, let's look at another one. This is no. This is married, this is no. [01:11:44] Durga Toshniwal: So, pre-records are there that are having marital status as married. [01:11:51] Durga Toshniwal: And for all these three, the class label is known only. This means that whenever [01:11:57] Durga Toshniwal: As per this particular training data, whenever the marital status is married, class label is going to be known. [01:12:04] Durga Toshniwal: So, therefore, we have all the 3 records with no class label, or don't cheat. Out of these 7, 3 are going here, and 4 are going on this side. [01:12:14] Durga Toshniwal: And because all these three are actually having the class label no. [01:12:20] Durga Toshniwal: They are having the same class level, no, so we don't need to do anything else. [01:12:25] Durga Toshniwal: We don't need to do anything else, and we will, not split this further, because all the class labels have the same, all the records have the same class label. Now, let's look at these four records, which are having marital status, single or divorced. [01:12:39] Durga Toshniwal: So… Here, we have refund no, marital status, this one. [01:12:45] Durga Toshniwal: Marital status single. This is the record. [01:12:49] Durga Toshniwal: It is having a class label? No. [01:12:51] Durga Toshniwal: This is record number 3. [01:12:54] Durga Toshniwal: Then, marital status… This one? [01:12:58] Durga Toshniwal: No, divorced. [01:13:00] Durga Toshniwal: It is having a class label, yes. It is record number 5. [01:13:06] Durga Toshniwal: Then we have, [01:13:08] Durga Toshniwal: Record number 7, sorry, not 7, it is having refund, yes, so we want to have refund known only. So, this is the one, 8. [01:13:18] Durga Toshniwal: So, this is it. It is having what? [01:13:20] Durga Toshniwal: Task label, yes. [01:13:22] Durga Toshniwal: And then… We have the record number 10. [01:13:26] Durga Toshniwal: Which is heavy. [01:13:27] Durga Toshniwal: Class label? Yes. So, out of these 4 records that are going here, we still don't have the same class label. [01:13:35] Durga Toshniwal: So this means that we still need to split this set further, and to do that, we are going to utilize the next available, [01:13:45] Durga Toshniwal: attribute. [01:13:46] Durga Toshniwal: So, the next available attribute is taxable income. [01:13:49] Durga Toshniwal: And this taxable income is a continuous valued attribute. Because it is continuous valued attribute, we cannot use it directly, we need to discretize it. [01:13:59] Durga Toshniwal: And the values that we are going to use. [01:14:01] Durga Toshniwal: to discretize this is going to be, the range 80 to 100K. [01:14:09] Durga Toshniwal: And then, less than 80, or greater than 100. So, this particular category is less than 80. [01:14:17] Durga Toshniwal: Or greater than 100, like that. [01:14:20] Durga Toshniwal: Of course, equal to sign will come. [01:14:23] Durga Toshniwal: So, this will be greater than only equal to 8200 is here. Now, if we look at these records, record number 3, [01:14:30] Durga Toshniwal: Record number 3. [01:14:32] Durga Toshniwal: Had refund equal to no, so we came here. [01:14:36] Durga Toshniwal: Then had marital status as single, so we came here. [01:14:41] Durga Toshniwal: Then, taxable income is $70. [01:14:44] Durga Toshniwal: So, 70 is not in the range of 80 to 100, so 70 means here. [01:14:49] Durga Toshniwal: And then, it has cheating equal to no. So, here, the cheating is no for this particular, one, which is the third one. [01:14:59] Durga Toshniwal: Now, let's look at the next one. [01:15:02] Durga Toshniwal: Here. This is record number 5. [01:15:08] Durga Toshniwal: So, we have a no, so we come here. Then we have marital status divorce, so we go here. [01:15:14] Durga Toshniwal: And then, cheating is yes. [01:15:17] Durga Toshniwal: So, when cheating is yes, we see the taxable income, it is $95K, so this particular record will traverse this path, and cheating here is yes. And this is which one? This is record number 5. [01:15:30] Durga Toshniwal: Then we go to record number 8. [01:15:33] Durga Toshniwal: Again, refund is no. [01:15:35] Durga Toshniwal: Marital status is single. [01:15:38] Durga Toshniwal: And the value of the income is $85K, so it will go from here, it will traverse this path, and cheating is yes. [01:15:46] Durga Toshniwal: And then we look at the 10th record, so refund is no. [01:15:50] Durga Toshniwal: Marital status is single, and then we have value of taxable income equal to 90K, which is this side, and cheating is yes. So, all three records have the same class label, and this record, the record going this side also has only no here. So, now our task is done, because [01:16:10] Durga Toshniwal: All the records that are going on this side, refund equal to yes, they are having marital status no only. [01:16:17] Durga Toshniwal: Sorry, cheating, no, that is don't cheat. And how many, are there? [01:16:22] Durga Toshniwal: Refund yes, refund yes, refund yes, all of them have no. So there are 3 records going here, 7 records going here. Out of these 7 records, I told you that, 3 of them have merit, so 3 are here and 4 are here. Out of these 4, 1 goes here, and 3 go here. [01:16:42] Durga Toshniwal: So, like this, we have built a tree. Now, this decision tree correctly, identifies the class label. [01:16:50] Durga Toshniwal: And now, if we have an unseen record, Then what will we do? [01:16:56] Durga Toshniwal: Unseen record, will be made incident on this tree. [01:17:00] Durga Toshniwal: So, let's say… Oh… [01:17:04] Durga Toshniwal: Let's say we have some particular record, so what we'll do, its class label is not known. [01:17:10] Durga Toshniwal: So, what we'll do, we'll make it incident here, and allow it to follow the path based on these attribute values, refund, marital status, taxable income, and all. [01:17:20] Durga Toshniwal: And then wherever, whatever path it follows, it is going to arrive at a leaf node, and the leaf node will give the class label. So this is how a decision tree is used to predict the class label. [01:17:33] Durga Toshniwal: So I'll stop at this point and, invite questions, if any. [01:17:38] Durga Toshniwal: Any questions, anyone? [01:17:41] vinit shah: Ma'am, on the taxable income, right, where we categorized 60 to 80 as 1, 80-200K as one, and above 100 to 50K as one. [01:17:49] Durga Toshniwal: So, was this done to… [01:17:51] vinit shah: make… [01:17:52] vinit shah: classification easier, or is there some sort of an idea behind it, or is it something at our discretion? [01:17:59] Durga Toshniwal: You're talking about the range that has been chosen? [01:18:02] vinit shah: Yes, yes. The last node where we are trying to, you know, again, have all the three, yes, yes. [01:18:08] vinit shah: Like, example, if I made this, instead of 80 or 100K, if I change this to, say, 90 to 100K, then I would have one record 85K, which would, again, not fall under that category. [01:18:18] Durga Toshniwal: Yo. [01:18:20] Durga Toshniwal: So, remember, there are two, three things you should know. First of all, discretization is a step that is involved in pre-processing. Remember, I explained to you. So, once done, we cannot change it. So, at this point, you cannot change it, because this discretization will be… would have been done at the start. [01:18:39] Durga Toshniwal: Okay, and what you will be having here are the discrete… in this taxable income, you would only be having the discrete values, rather than… I did not show them, but it will be done at the pre-processing step, number one. Number two, the discretization is usually done keeping the objective. Here, in this case. [01:18:59] Durga Toshniwal: Oh. [01:19:00] Durga Toshniwal: So here, in this case, we want to perform classification. [01:19:04] Durga Toshniwal: And, accordingly, we have done the discretization. But there are many ways of discretization. Sometimes we just divide the range. We discussed all those cases, right? Sometimes we divide the range into equal width drains, sometimes we do equal frequency discretization and all that. [01:19:23] Durga Toshniwal: So, yeah. [01:19:25] vinit shah: So the doubt around this, okay, maybe I'll put it in a different way. Then, say, for example, is the discretization done [01:19:32] vinit shah: kind of looking at the data and then trying to see that, you know, this discretization might fit it, or in this case, it's a coincidence, because if I… if I ask you that, you know, you introduce another case here instead of 80K to 100K and, you know, make it as a 90K to 100K, then will my decision tree again go into another loop where you have to do something else, or… Yes. [01:19:51] Durga Toshniwal: Yes, yes. Then you'll have to do something else if… if the discretization was done, like, 90 to 100 and rest of the values. [01:19:59] vinit shah: Okay. [01:19:59] Durga Toshniwal: If it was done, then we'll have to do something more to complete this tree. That is one thing. [01:20:05] Durga Toshniwal: And then, secondly, definitely, we want a most optimal tree, and accordingly, we do the discretization. One thing very important we should remember is that as data scientists, or as people dealing with data, definitely nothing can be done without having a first look at the data. [01:20:25] Durga Toshniwal: So, my suggestion is, anyone who's using our data, first have a look at what it is like. Only then you can decide what you need to do with it, right? So, definitely, that is required. [01:20:37] Durga Toshniwal: Looking at the data to decide what task labels will look like and all that, that we'll definitely need to see. [01:20:44] Durga Toshniwal: Oh my god. [01:20:46] vinit shah: Sorry, one last question on this, ma'am. Sorry, Trento. So, can I then think it in this way, right? I use the Hunch algorithm. I say, for example, my first discretization had 9200K, okay? I did all this thing, and then, you know, I realized that maybe if I would have changed my discretization. [01:21:01] vinit shah: I could have had a much better, decision tree or something. So, then can I again go change my data, and then again start the process? [01:21:08] vinit shah: Is that a good practice? [01:21:11] Durga Toshniwal: So, number one, you can definitely do that, but you have to go back and start from scratch, because preprocessing… the data is the first step. Correct. So, and if the data is having lots of records, say, thousands or millions of records, pre-processing itself can be costly. It could. [01:21:30] vinit shah: Okay. [01:21:30] Durga Toshniwal: compute-intensive. So once you do it, then it's not a good practice to do it again. [01:21:36] Durga Toshniwal: Okay. Because it requires a lot of compute power, and if you have, you can afford it, then you can do it. [01:21:44] vinit shah: Okay. [01:21:45] Durga Toshniwal: Most of the times, we don't have that much of time, neither resources, to repeat this whole process. Therefore, it's advisable to, you know, judiciously decide [01:21:56] Durga Toshniwal: What, you know, how we are discretizing a given attribute. [01:22:02] vinit shah: Okay. [01:22:03] vinit shah: Thanks, thanks, Mo. [01:22:05] Gunjan Bhaiya: Ma'am, a follow-up question, what Vinit told. Even to do, like, manually, right, we have to do some hit and trial first, right? To make that balanced approach, right? [01:22:14] Gunjan Bhaiya: Then only we have to go with the actual… because in the real world, the data size is going to be huge. [01:22:19] Durga Toshniwal: Yes, yes. [01:22:20] Gunjan Bhaiya: Right? So, we have to do… even if we say we have to do, like, a data scientist, we have to look at the data first, right? [01:22:25] Gunjan Bhaiya: We might need to do some hit-and trial to arrive at the stage. We can say, yes, this, on the paper, this seems a balanced or a efficient way of creating a tree, right? [01:22:36] Durga Toshniwal: Yeah, so it is always a good idea, you can take a sample, or you can take a stratified sample, as I already told you, and then… [01:22:45] Durga Toshniwal: You know, try to play around with that stratified sample and see what kind of discretization looks optimally good, and then you can decide to use it. [01:22:57] Durga Toshniwal: You can do that. Or else, the other way is that you think of a discretization, do it, then look at the performance of the model, then go back, redo it. Of course, that will be compute intensive. So, it's a good thing what you are saying, is to first do some hidden trial and see [01:23:16] Durga Toshniwal: What discretization would work? [01:23:19] Durga Toshniwal: Better. And then choose that. [01:23:23] Gunjan Bhaiya: Thank you. Thanks. [01:23:24] Deepan Kanagaraj: Ma'am, one question. So, every time, do we need to have only, like, two leaf nodes, or that… can that be more? So that's one… one question. [01:23:34] Durga Toshniwal: No, it can be anything. I just showed you a binary tree. A binary is one where two edges come out, but it could be an array also. So you can have a tree like this, like this. It could be a combination of two. [01:23:48] Durga Toshniwal: Or multiple. It could be anything. Everything is fine. [01:23:52] Deepan Kanagaraj: Okay. Another question is, do we, completely has to expect all the, labels to match properly before even we get started, or can we just assume that, okay, maybe, this is giving me 90% of accuracy, and then can I move forward with the real… I'm talking about in real cases, ma'am. Yes, yes. How do we normally do? [01:24:11] Durga Toshniwal: Yeah, so actually, in the real world, the 100% match will definitely not be possible at the least low, and so you have to actually take majority voting. [01:24:21] Durga Toshniwal: I will be discussing that also. However, I'm just telling you. [01:24:25] Deepan Kanagaraj: Pokemon. [01:24:26] Durga Toshniwal: is correct. That 100%, [01:24:29] Durga Toshniwal: Saying class labels may not be possible in the real world. Therefore, we look at majority, and if that value of the majority is acceptable to us, we accept the tree. If it is not, then we redo it. [01:24:45] Deepan Kanagaraj: Okay, ma'am, got it, thank you. [01:24:46] Durga Toshniwal: Yeah, Deepak, I think you have a question. [01:24:49] Deepak Katara: I mean, what all scenario, we do use this hunt algorithm, because it seems like it's very costly, and also, right now, we already consider binary data. If it is beyond that, then it looks pretty complex in that way. So, in what all cases, we use this. [01:25:06] Durga Toshniwal: Yeah, so actually, Hunt's algorithm, I'm, telling you, is a algorithm for building the decision tree. [01:25:13] Durga Toshniwal: And the decision tree, definitely, at this point of time, you might feel that it is very compute-intensive, it looks very complex, and all that. But believe me, as you progress, you'll realize that it is one of the simplest methods. [01:25:29] Durga Toshniwal: And the other methods are much more costlier, they are much more complex, and that's why our decision-free-based classification algorithm is one of the most popular classification algorithms. [01:25:46] Durga Toshniwal: Okay, so… [01:25:48] Durga Toshniwal: It's looking that it is difficult, it is, but as you go along, you'll have… have other methods which are much more complex. So this will now… then it will seem to you much simpler. [01:26:01] Deepak Katara: So, is there a possibility that we go with, like, go-to approach in this hunt algorithm to visualize our classification and everything, and then we can probably look forward to optimize it with other methods? [01:26:13] Durga Toshniwal: Other classification methods, you are saying? [01:26:16] Deepak Katara: Yeah, other decision tree, as you said, like, there are a few other algorithms that we are going to read in future. [01:26:23] Deepak Katara: Definitely, you can use one method in conjunction with the other. [01:26:27] Durga Toshniwal: You can do that. It purely depends on what you want to do, how you want to use it. You can do that. That's okay to do. [01:26:36] Deepak Katara: Okay. That makes sense as of now, yeah? Thank you. [01:26:43] Durga Toshniwal: Sivanch, you have placed a hand, do you have a question? [01:26:47] Shivansh Sharma: Yes, ma'am. [01:26:48] Shivansh Sharma: I just wanted to understand any… in practicality, do we need to any measures with respect to this Decision 3 algorithm usage? [01:26:59] Shivansh Sharma: Like, the limit of the depth. [01:27:02] Shivansh Sharma: Or the maximum sample splits, where we get them, [01:27:08] Shivansh Sharma: optimized result, whether we should focus on the breadth, or… [01:27:14] Shivansh Sharma: Or depth, which will give the better results to us. [01:27:17] Durga Toshniwal: Okay, so that's an interesting question. I will, I will, answer that. [01:27:23] Durga Toshniwal: But definitely, I just started explaining to you what Hunt's algorithm broadly does, but how we build a decision tree? [01:27:34] Durga Toshniwal: Like, here we are just manually looking at the records and all. Definitely, an algorithm will do it automatically, so I'll explain to you the measures that will be used. So you'll have to wait for some more time. [01:27:47] Durga Toshniwal: And you'll come to know how we decide, how we choose attributes, and all that stuff, how we do it. [01:27:53] Durga Toshniwal: And what are the other parameters that are required? So, all that, what we are saying, like max depth and all that, may be required. [01:28:01] Durga Toshniwal: Okay, we'll also shoot in the hands-on also. [01:28:04] Durga Toshniwal: Now, the question to… another question that you had asked was that [01:28:09] Durga Toshniwal: Do we go for deep trees? I think that was the question, right? [01:28:13] Shivansh Sharma: Yeah. [01:28:14] Durga Toshniwal: So, whether we should prefer a deep tree or we should not prefer a deep tree. [01:28:19] Durga Toshniwal: Whether we should prefer a short trigger and all that. [01:28:24] Shivansh Sharma: Miss. [01:28:25] Durga Toshniwal: Yeah, so definitely overly deep trees are not preferred, but optimally deep trees are are required. [01:28:33] Durga Toshniwal: But you will learn more as you go along. I'll explain to you what this means and what are the after-effects of having a very deep tree. So, you'll have to wait for some more time. I'll explain to you. [01:28:46] Durga Toshniwal: Okay. [01:28:47] Shivansh Sharma: Okay, man. [01:28:48] Durga Toshniwal: Thank you. Anyone else? Any other questions? [01:28:57] Durga Toshniwal: So, there are no further questions. [01:29:00] Durga Toshniwal: Oh. [01:29:01] Durga Toshniwal: I feel good. [01:29:02] Sonam Manwal: Hmm. [01:29:04] Durga Toshniwal: Yes, go ahead. [01:29:04] Sonam Manwal: How are we deciding this? Like, first we took a refund, then we are taking marital status. [01:29:14] Durga Toshniwal: So, I told you that I will be explaining to you how we choose the attribute, so I… right now, I'm just broadly showing you how it works. I mean. [01:29:22] Durga Toshniwal: manually how we are doing it. But as I already pointed out to you, that no algorithm will work on any manual calculation, so I will explain to you how you will choose the attribute, okay? [01:29:35] Sonam Manwal: Okay. [01:29:36] Durga Toshniwal: Just wait for some more time. [01:29:39] Sonam Manwal: Thank you. [01:29:39] Durga Toshniwal: So, I think there's a question, Lokesh, how does the decision tree decide which attribute gives the highest information gain and all that? [01:29:48] Durga Toshniwal: So, I will explain to you. Please wait for some time. [01:29:54] Lokesh R: Yes. [01:30:14] Durga Toshniwal: Sorry, I got painted. [01:30:15] Durga Toshniwal: So now, let's take another example. [01:30:21] Durga Toshniwal: Remember that once the pre-processing is done on the, [01:30:26] Durga Toshniwal: On the… these continuous valued attributes, actually. [01:30:31] Durga Toshniwal: You cannot change that, right? So, whatever range you had started out, that remains fixed. So what was that? Taxable income is broken up into 80 to 100K. Less than 80 and greater than 100 will be the other category. So, these are the two categories. Marital status has been broken down into marriage. [01:30:50] Durga Toshniwal: And then single divorced as another one, refund is yes and no. So we have all these 6 categories available for refund, two, marital status, two categories, and taxable income also, two categories. [01:31:02] Durga Toshniwal: Now, we have this particular data that you can see here. [01:31:06] Durga Toshniwal: Okay. [01:31:07] Durga Toshniwal: Can you build a tree using this? [01:31:15] Durga Toshniwal: You can have a look for a minute, and then, let me know. [01:31:20] Durga Toshniwal: Whether you can build a tree. [01:31:23] Durga Toshniwal: Or you can't build a tree. [01:31:26] Durga Toshniwal: I mean, whether you can build a tree. [01:31:35] Durga Toshniwal: So, in the meantime, while you're thinking, there is a question. I think Shikar said that, is it necessary to consider all attributes? [01:31:43] Durga Toshniwal: In a decision tree, and can we make multiple trees to come to a decision? So, the answer to the first question, it's not necessary to choose all the attributes. [01:31:54] Durga Toshniwal: No, it's not necessary. [01:31:56] Durga Toshniwal: We only choose the ones that are enough for us to be able to make a good decision. [01:32:01] Durga Toshniwal: About the class label. [01:32:04] Durga Toshniwal: So, we don't necessarily need to use all the attributes, that's number one. Number two, can we use multiple trees? Yes. And there are algorithms like Random Forest, which I'll discuss as we go along. [01:32:16] Durga Toshniwal: That make use of multiple trees, okay? [01:32:19] Durga Toshniwal: So you'll have to wait for some time. [01:32:23] Durga Toshniwal: Yeah. [01:32:24] vinit shah: We can build a tree, madam. We can build a tree. [01:32:27] Durga Toshniwal: We can't build that rule. [01:32:28] Deepak Bobade: we can build a tree, but, definitely, I think refund can't be the first route. [01:32:35] vinit shah: It can be, I guess it's against him, refund is just cheating. [01:32:40] vinit shah: Yeah, correct, sorry, my bad, yeah. [01:32:42] Durga Toshniwal: Let me see… let us see. So, we have refund, yes. [01:32:46] Durga Toshniwal: Then we have refund, yes. [01:32:48] Durga Toshniwal: And both of them have a known. So, if we started out with a tree. [01:32:53] Durga Toshniwal: And we started out with refund. [01:32:57] Durga Toshniwal: We could have a yes year. [01:33:00] Durga Toshniwal: And we could have a leaf node, which says a class labeled no. [01:33:04] Durga Toshniwal: Okay. And how many records are there? [01:33:07] Durga Toshniwal: Yvonne? [01:33:08] Durga Toshniwal: And 7, so there are 2 reports going here. Now, the rest of the 8 go there. [01:33:15] Durga Toshniwal: So, Saihardvaja, I will take your query, just wait for some time. [01:33:20] Durga Toshniwal: So, 8 of them go here. [01:33:23] Durga Toshniwal: And let's say in the previous case, we had used marital status, so let's say we are using the same one again. [01:33:30] Durga Toshniwal: So, we are testing on, marital status. [01:33:34] vinit shah: Wherever it is married to a cheating is known. [01:33:37] Durga Toshniwal: Yeah, so now it is married. [01:33:40] Durga Toshniwal: The cheating is no. Let me put this one as… Marital status, married. [01:33:47] Durga Toshniwal: And… Let's say cheating is no… [01:33:51] Durga Toshniwal: So, this is no. Then, we have marital status. [01:33:56] Durga Toshniwal: Patrick? [01:33:58] Durga Toshniwal: And then we have… Heating, no? [01:34:01] Durga Toshniwal: And marital status marriage eating. [01:34:04] Durga Toshniwal: We have 3 reports. [01:34:06] Durga Toshniwal: Cool. [01:34:07] Durga Toshniwal: It is no, and we are having 3 records going. [01:34:10] Durga Toshniwal: Then… [01:34:14] Durga Toshniwal: Out of 8, 5 of them go here, and that is single and divorced. So, we have the third one. [01:34:20] Durga Toshniwal: Which says, no, single, and [01:34:24] Durga Toshniwal: And let's say that out of these five, because all of them are not the same, because there is a yes here, then there's a no here, so it's a mix. So we are again going to use the taxable income here. [01:34:37] Durga Toshniwal: And the split that we already defined. [01:34:41] Durga Toshniwal: And then we will build the tree. So, we are having which one? We are done with this. [01:34:48] Durga Toshniwal: Then we are done with this. So, we are having refund no, then we are having marital status, single. [01:34:56] Durga Toshniwal: So, it will go this side. [01:34:59] Durga Toshniwal: And then, the… Sorry, third one. [01:35:03] Durga Toshniwal: then it has a value 70K. [01:35:06] Durga Toshniwal: So, let's say I define greater than 80, sorry, less than 80 and greater than 100 this side, and 80 to 100 this side. So, it will follow this path. [01:35:18] Durga Toshniwal: Okay, and it will follow this path, and it will, let's say, have a value, yes, here. [01:35:27] Durga Toshniwal: Now we have the fourth one. It goes on no. [01:35:31] Durga Toshniwal: After a new marital status is single, it comes this side. [01:35:35] Durga Toshniwal: The value of the income is $70K, so it again comes this side. [01:35:41] Durga Toshniwal: And the value of the class label is no here. [01:35:47] Durga Toshniwal: Then, if we consider the rest of them, this is divorced. It has, [01:35:52] Durga Toshniwal: 95k? Values, yes. The 95 will fall here, and values, yes. [01:35:59] Durga Toshniwal: Then… We have, diverse 220K value is no. [01:36:07] vinit shah: That's completed, ma'am. It's under Yes Refund, so we put it in the starting. [01:36:11] Durga Toshniwal: Yeah, yeah, sorry. [01:36:13] Durga Toshniwal: So, this is single, it is 85, it is IES. [01:36:17] Durga Toshniwal: So, another yes here. [01:36:19] Durga Toshniwal: And then, this is singles, this is 90, another year. [01:36:23] Durga Toshniwal: So, this is final. This will be yes. [01:36:27] Durga Toshniwal: But now this one… is comprising of two records. One is yes, one is no. [01:36:33] Durga Toshniwal: Then, how are we building the tree? [01:36:35] Gunjan Bhaiya: We can take, again, like, a single and divorce. [01:36:39] Gunjan Bhaiya: as a… Different node. [01:36:43] Gunjan Bhaiya: Or pop. [01:36:44] Durga Toshniwal: But we have, already… we are finished with the attributes. We just had how many? 1, 2, 3. We are done with 1, 2, 3, and how we'll choose? We can't keep choosing them again and again. [01:36:57] Durga Toshniwal: Because that will make the tree unnecessarily deep. [01:37:02] Gunjan Bhaiya: At the metal restrictors, can we go. [01:37:04] Nirav Mehta: Don't believe it. [01:37:05] Nirav Mehta: We'll go by majority of the status since 8 of… For it, we have identified. [01:37:13] Durga Toshniwal: Yeah, but in here, there's no majority. One is yes, one is no. [01:37:18] Durga Toshniwal: It's 50-50. [01:37:19] Deepak Bobade: Yeah, it's 50-50, yeah. [01:37:23] vinit shah: I don't know, we leave, we're telling 95% it's correct, so… [01:37:27] vinit shah: I'll go with it or something. [01:37:29] Durga Toshniwal: So you also… I mean, some of you said that it's possible to build a tree. Now, here we are at a point that we have built a tree, but we have not converged to the solution. [01:37:40] Durga Toshniwal: Then what should we do? [01:37:42] Deepak Bobade: Change the mood. [01:37:44] vinit shah: Maybe change the route? [01:37:46] Durga Toshniwal: Change the route. [01:37:48] Durga Toshniwal: Even if you change the route, if there's this thing, it will come. [01:37:51] vinit shah: This will, yeah, yeah, this record will stay the same, correct? [01:37:54] Durga Toshniwal: Yep. [01:37:56] Gunjan Bhaiya: Ma'am, at the multi-status, can we start at 3 nodes instead of single, married, and diverse 3 instead of 2 only? [01:38:02] Durga Toshniwal: Yeah, but remember, Gunjan, that we have done that deep free processing again, we don't have enough compute to do that all over again. [01:38:10] Durga Toshniwal: And that's why I said these are fixed once done. [01:38:15] Durga Toshniwal: This free processing is over. [01:38:18] Deepak Katara: Can we change the range of taxable income? [01:38:22] Deepak Katara: So the method of discretization of taxable income. If we change it, then it probably will fall into the right place. [01:38:28] Durga Toshniwal: So, yeah, so what I'm saying is, once the preprocessing is done, then it's done. Then you cannot change it. [01:38:35] Durga Toshniwal: If you have to change it, then you have to start all over again. In this data, there are 10 samples. In real data, there might be millions of samples. Imagine pre-processing them again. That will be too much of a, you know, resource requirement. [01:38:51] Durga Toshniwal: So, once pre-processed, it cannot be done again. [01:38:54] Sushree Dash: So you have to… Both are 70K also. [01:38:58] Durga Toshniwal: Both are 70K, yeah. [01:39:01] Durga Toshniwal: That's okay, I mean, two people can have 70K. [01:39:05] Durga Toshniwal: As the income, so… Both are 70, that's okay. [01:39:10] Durga Toshniwal: That's how the data is. [01:39:15] Durga Toshniwal: Okay, in the meantime, Saif Advaj, I think you raised a hand a long time back. What is your question? [01:39:22] Saibharadwaj C: So ma'am, you have started [01:39:24] Saibharadwaj C: The classification based on the refund as first. [01:39:28] Saibharadwaj C: a column. [01:39:30] Saibharadwaj C: How can we decide in the dataset which to pick, and what order should we follow? [01:39:36] Durga Toshniwal: Yeah, so I told you earlier also, I will be telling you. [01:39:39] Durga Toshniwal: How to choose the attribute to do the split. [01:39:43] Durga Toshniwal: I'll be telling you, you'll have to wait for a little more while, okay? [01:39:47] Saibharadwaj C: Okay. [01:39:48] Durga Toshniwal: They need to. [01:39:51] Durga Toshniwal: Okay, in the meantime, we are still standing at this problem, and we don't know how to resolve it. [01:39:58] Durga Toshniwal: Okay, so… So the tree is actually… [01:40:06] Durga Toshniwal: That was our original one with that problem we are still having. [01:40:13] Durga Toshniwal: And… [01:40:19] Durga Toshniwal: So this is the tree we are at right now, and we don't know what to do here. [01:40:25] Durga Toshniwal: And the problem is arising because, you see here, refund is no, refund is no. [01:40:31] Durga Toshniwal: Marital status, single and single. Taxable income, 70, or anything between. I mean, if it… if it had [01:40:40] Durga Toshniwal: Let's say instead of 70, if it had, say. [01:40:45] Durga Toshniwal: 240, then also it would be the same thing only. [01:40:49] Durga Toshniwal: But… The problem is, when all the attributes are same, one class table is yes, one is no. [01:40:57] Durga Toshniwal: Because the two class labels for the same set of attribute values [01:41:01] Durga Toshniwal: are different. Therefore, we are arriving at a situation that is not converging. [01:41:09] Durga Toshniwal: So… The… remember that the decision tree will only work [01:41:16] Durga Toshniwal: When you have unique value pair combinations in the attributes. [01:41:23] Durga Toshniwal: If your attributes are such that unique value pair combinations are not possible, then don't use a decision tree. [01:41:33] Durga Toshniwal: What I mean by unique value pair combinations? It means that if I had a combination like no, single, and 70, [01:41:41] Durga Toshniwal: then… It will have some particular class label, and if it occurs again. [01:41:47] Durga Toshniwal: That kind of combination, that it should have the same class label as earlier. [01:41:52] Durga Toshniwal: So that a yes is linked. [01:41:55] Durga Toshniwal: With some particular combination of values. [01:41:58] Durga Toshniwal: But it should not happen that for the same combination, we are having multiple class labels. So this is the first condition of using Hunt's algorithm that [01:42:11] Durga Toshniwal: We should have unique value pair combinations. If we don't have unique value pair combinations, we can't build a decision tree on it. [01:42:19] Durga Toshniwal: Okay. [01:42:21] Durga Toshniwal: Okay, Deepak, what question do you have? [01:42:23] Deepak Katara: One other thing that we discussed, that we can drop the data, as well. It's not necessary to… [01:42:29] Deepak Katara: pick everything, so would that… [01:42:32] Durga Toshniwal: Okay, so dropping and all is done as part of preprocessing. [01:42:36] Durga Toshniwal: After you have pre-processed the data, it is ready to be classified, and there's nothing more you do on it. [01:42:42] Durga Toshniwal: Except for classification, because preprocessing part is done. [01:42:46] Durga Toshniwal: While you were preprocessing, you could have dropped the row, if at all. [01:42:50] Deepak Katara: You could have afforded to drop it, but assuming… [01:42:54] Durga Toshniwal: That dropping the row was possible, then you would have done it already. [01:42:57] Durga Toshniwal: Now, at this point, you cannot drop. [01:43:02] Deepak Katara: Yeah, maintenance, so… [01:43:04] Durga Toshniwal: just think of it that you have made a data frame, let's say in Python. Then, on the data frame, you did some transformations, like… [01:43:12] Durga Toshniwal: You know, so… Discretizing it, making bins, all that, something, something you did. [01:43:19] Durga Toshniwal: And now, that part is ready, now you can't do anything on it. [01:43:26] Durga Toshniwal: Any other question, anyone? [01:43:36] Durga Toshniwal: So let me proceed from here. [01:43:44] Ankit Sood: Ma'am, one question here. So, what are we supposed to do here? Are you going to go through that solution now? [01:43:52] Durga Toshniwal: So, at this point in time, where we are having a situation. [01:43:58] Durga Toshniwal: this situation, where we have equal number of yeses and nos at a node, and we don't have more attributes to use to do the split. We can't do anything. This means our decision tree algorithm has failed, because it cannot predict the class label here. [01:44:16] Durga Toshniwal: And so we'll need to do… use some other, classification algorithm. That's the only solution. [01:44:24] Ankit Sood: But we can't actually drop the row, so as somebody mentioned in the call, right, like, we can actually check if there are certain rows, or a certain combination of columns for which we are getting more than one. [01:44:37] Ankit Sood: values for cheats, so we can drop one of them, and we can still use this algorithm, right? Only thing is that the accuracy might be lesser, but still it's usable, correct? [01:44:49] Durga Toshniwal: So, I'm asking you, let's say we have these two records, okay? [01:44:52] Ankit Sood: new. [01:44:53] Durga Toshniwal: Actually, this one was standing for Person A, this one was standing for Person B. Let's assume. [01:44:59] Durga Toshniwal: Okay? Though the names are not here. [01:45:02] Durga Toshniwal: Now, you're saying that if we can drop at this point, if we had the freedom, though of course we don't have it, let's assume you had the freedom to drop. [01:45:11] Durga Toshniwal: How will you decide to draw… whether to draw A's record or B's record? Let's assume that after doing all this, that is the building of the decision tree. [01:45:22] Durga Toshniwal: Oh… [01:45:24] Durga Toshniwal: This decision tree is going to be used to decide upon some kind of award or something that will be given to someone. Now we have A and B two people, two people record. [01:45:38] Durga Toshniwal: And, the outcome of the decision is going to… [01:45:42] Durga Toshniwal: you know, decide whether the person may get a price or not get a price, then how will you decide to drop A or B? [01:45:51] Ankit Sood: So, ma'am, it's, like, not that binary, right? In practical situation, it's not going to be… it may not be a binary decision that we are making, right? [01:46:01] Ankit Sood: So, even if we actually try to drop either of those rows, let's say I drop A, which means that half of the time I'm still going to be correct. [01:46:12] Durga Toshniwal: Yeah, but half of the times, if you are correct, you don't need to make any, classification algorithm and do any kind of hocus-pocus, you know? That you can say without anything also. [01:46:24] Ankit Sood: But if it's an approximation problem, where I have to make a decision, and I have to give an answer? [01:46:30] Durga Toshniwal: Right, but what I'm saying is that having a probability of 0.5, you don't need any algorithm, you can say yourself also. [01:46:39] Ankit Sood: Okay. [01:46:39] Durga Toshniwal: I can say, anyone can say. Then why do all so much of our work? [01:46:44] Durga Toshniwal: If you need 0.5 as, the value. [01:46:48] Ankit Sood: 0.5 only for certain cases, right, ma'am? It, like, we are coming through 3 hops already. It's not like on the first hop we are saying. [01:46:57] Durga Toshniwal: No, that's okay, so this is the final rewrite. [01:47:01] Durga Toshniwal: Now. [01:47:01] Ankit Sood: Correct. So, which means that we came till 3… after 3 ops only, we are making a decision, not at the refund we are saying. That's where the probability will be 0.5, right? [01:47:10] Ankit Sood: That half of the time I'm going to make correct decision, and half of the time I'm going to make wrong decision. [01:47:15] Durga Toshniwal: You're talking about refund equal to yes and no? [01:47:17] Ankit Sood: Yes. [01:47:18] Ankit Sood: As we are… as we are moving down the tree, the probability point is not 0.5 anymore, right? [01:47:26] Durga Toshniwal: Yeah, but you can see here, if you talk about probability yes or no at this point, which is the first one, you have how many refund equal to yes one. [01:47:36] Durga Toshniwal: 2, and I think the… [01:47:38] Ankit Sood: No, ma'am, in generalization sense, right, I'm not talking just from this dataset perspective. [01:47:42] Durga Toshniwal: Yeah, yeah, not just this data set. I'm trying to explain the same thing using this data, though it will be true. [01:47:49] Ankit Sood: Okay. [01:47:49] Durga Toshniwal: data. [01:47:50] Ankit Sood: Okay, sure. [01:47:50] Durga Toshniwal: The thing is that the probability will depend on the number of samples going this side. [01:47:56] Durga Toshniwal: Yes, there are 2 on 10, and it is 8 on 10. [01:47:59] Ankit Sood: Okay. [01:48:00] Durga Toshniwal: Probability is just 1 upon 5, and this one is higher anyway. [01:48:04] Ankit Sood: Okay. [01:48:05] Ankit Sood: But as we move down, then this probability of… It's actually become lesser, right? [01:48:12] Durga Toshniwal: Yeah, but then you have to consider the joint probability of no, then this, and this. [01:48:19] Ankit Sood: Yes, which means that we can't… it's going to be an approximation in that sense. We can't really say that half of the time we are making wrong decisions. [01:48:29] Durga Toshniwal: Yeah, you are not making wrong decisions half of the time, but you are even lower than 0.5. It is how much? You are… it is 0.2. [01:48:38] Durga Toshniwal: You are saying 0.5 times you are correct, or you are incorrect. Now you're… [01:48:43] Durga Toshniwal: Correct, only 0.2 times. That's what I want to say. [01:48:47] Durga Toshniwal: It's going worse, not becoming better. [01:48:50] Ankit Sood: No, ma'am. On the sec- right-hand side, it is 8x10, correct? And then, again, you're going to calculate at each node what is the probability. So, by the time you actually come at the last leaf node, at that time, you will see the difference there, right? [01:49:04] Durga Toshniwal: No, but it will all be a joint probability. How can you say this will be… [01:49:08] Ankit Sood: a better one only. You need to calculate and see, definitely. [01:49:13] Durga Toshniwal: It all depends on how many records go this side, how many go this side, and all that stuff. [01:49:18] Durga Toshniwal: So, definitely, we are not making use of any probability here. [01:49:23] Ankit Sood: Okay. [01:49:24] Durga Toshniwal: Because what we are making use of is just the split information itself, and that's all. [01:49:34] Durga Toshniwal: So I think I already told this. [01:49:41] Durga Toshniwal: So now, let's have, this data that you are seeing here. [01:49:47] Durga Toshniwal: and the tree. Now, what are your observations or comments on the tree? [01:50:07] Durga Toshniwal: Anyone, any comments? [01:50:11] vinit shah: Ma'am, can you come again? Sorry, I didn't get the question. [01:50:14] Durga Toshniwal: Yeah, so what I'm seeing is, now we have this data. Forget about the earlier one, we have this data. [01:50:22] Durga Toshniwal: Now, using this data, What observations do you have? [01:50:27] Durga Toshniwal: You know, while you are building a decision tree on this data. [01:50:34] vinit shah: So I guess that same record is kind of gone away. [01:50:38] Durga Toshniwal: Pardi? [01:50:39] vinit shah: Sorry to go ahead. [01:50:43] Abhilash Daniel: Will it be fair to say… [01:50:45] Abhilash Daniel: Single and divorced people are more prone to cheating than married people. [01:50:50] Durga Toshniwal: That probably you may say, but right now I'm saying that if you have this kind of data. [01:50:57] Durga Toshniwal: Then, what happens to the tree? [01:51:00] vinit shah: We can, I guess, create it because that same record kind of a thing has gone away right now, right? Like, same record repeating twice or something. [01:51:08] Durga Toshniwal: Yeah, that's true, that that problem is, burn. [01:51:13] Durga Toshniwal: Buzz… You have another issue. [01:51:17] Durga Toshniwal: If you build a tree, you will notice that… I'd… [01:51:22] Anurag Krishnam: Ma'am, taxable income is the problem, I think. [01:51:26] Guru Raghavendran: Range is, still the same. [01:51:29] Durga Toshniwal: See, income is always a problem, right? You have low income, you have problem. If you have high income, you have problem. So, it's creating problem here also. [01:51:39] Anurag Krishnam: Yeah, absolutely. [01:51:41] Durga Toshniwal: Anyway, yes, you are… what you're saying is right, that this attribute taxable income is causing a lot of problem. [01:51:49] Durga Toshniwal: But in any case. [01:51:51] Durga Toshniwal: If we use this data, then we will observe that while we are traversing this tree. [01:51:58] Durga Toshniwal: There will be no record incident, or there'll be no record coming this side at all. [01:52:05] Durga Toshniwal: See? [01:52:06] Durga Toshniwal: If you have yes, then you have yes, yes, and yes. All 3 go here, so 3 are gone. [01:52:12] Durga Toshniwal: Then for this no, it is married, so it is no, then it is married, so it goes this side. [01:52:19] Durga Toshniwal: Then we have next as, no. [01:52:22] Durga Toshniwal: And this is divorced, so it is no… And then it goes here. [01:52:28] Durga Toshniwal: And then, it comes to taxable income 95. [01:52:34] Durga Toshniwal: So 95 means, it will go… probably go this side, right? [01:52:40] Durga Toshniwal: And then, we have, this one, no. [01:52:48] Durga Toshniwal: Which is single, 85 years. [01:52:51] Durga Toshniwal: So, no. [01:52:53] Durga Toshniwal: Single, then 85. [01:52:56] Durga Toshniwal: 85 will be here. [01:52:58] Durga Toshniwal: And then, cheating is yes here, okay? [01:53:04] Durga Toshniwal: So, actually, cheating is yes means what? Sorry, 85 will go this side. I'm so sorry. [01:53:11] Durga Toshniwal: So, for this one… No. [01:53:15] Durga Toshniwal: Then, singan. [01:53:17] Durga Toshniwal: And then 85 means it will go here. [01:53:21] vinit shah: There's no record directly on the left-hand side, because we don't have anything that is 60 to 80K or 100 to 50K under the… [01:53:29] vinit shah: Single or divorced category. [01:53:31] Durga Toshniwal: Yes, exactly. So what I'm saying is, we don't have any record here. [01:53:36] Durga Toshniwal: So then, what do we do? [01:53:39] Durga Toshniwal: We have records, at all other, [01:53:43] Durga Toshniwal: Points, or leaf nodes, but we don't have anything here. [01:53:49] Durga Toshniwal: So, the answer to this question is… that… While building the tree. [01:53:55] Durga Toshniwal: like, for example, here, it looks that it is unnecessary to have this kind of a leaf node, because it's actually not having anything incident onto it. But in the real world, what would happen is that [01:54:09] Durga Toshniwal: There could be samples which are having value combinations of this kind. [01:54:15] Durga Toshniwal: And therefore, they will… Become incident on it. [01:54:19] Durga Toshniwal: Right now, we are building this tree only with the training data, and in training data, there's no such sample. But in the real-world data, there could be samples. [01:54:29] Durga Toshniwal: And therefore, it is always good to have [01:54:33] Durga Toshniwal: This kind of a node, even if it is empty at that point. [01:54:38] Durga Toshniwal: Okay. [01:54:40] Durga Toshniwal: So, this is one thing. [01:54:43] Durga Toshniwal: Then, no… [01:54:49] Durga Toshniwal: About the rest of it, I think I've told you on this. [01:54:53] Durga Toshniwal: Another thing is that, [01:54:56] Durga Toshniwal: How do we resolve this? Because here, we have kept it, suppose, as I said, that we need to keep it. [01:55:03] Durga Toshniwal: In the decision tree. But when we keep it in the decision tree, then what is a label we assign to it? [01:55:11] Durga Toshniwal: So, the label is assigned. [01:55:13] Durga Toshniwal: Either what we do is that we assign a default class label here. [01:55:19] Durga Toshniwal: some default class label, whatever we decide. Or else, what we do is that we merge this with the parent node. [01:55:27] Durga Toshniwal: And give the class label of the parent node only. [01:55:31] Durga Toshniwal: Okay, so this is how we resolve this. I've already told you the comments. [01:55:36] Durga Toshniwal: Okay. [01:55:38] Durga Toshniwal: Now, if the training data was something like what you're seeing here, So, what do you observe? [01:55:48] Durga Toshniwal: on this data. [01:55:50] aditya shrivastava: All are refund, all are no for refund. [01:55:54] Durga Toshniwal: Oh, we have… All going northside. And then after that, You have refund here. [01:56:03] Durga Toshniwal: You have marital status here? [01:56:08] Durga Toshniwal: Then, what will happen? [01:56:10] Shikhar Gupta: No case for merit. [01:56:12] Durga Toshniwal: So, all will go single drivers. [01:56:15] Durga Toshniwal: Okay. [01:56:16] Durga Toshniwal: And then… If we talk about taxable income. [01:56:22] Durga Toshniwal: then they are all satisfying one particular criteria. That is, either less than 80 or greater than unread. [01:56:31] vinit shah: 192. [01:56:32] Durga Toshniwal: So, all of them will go on that particular criteria, this one. [01:56:37] Durga Toshniwal: Whatever that is. [01:56:38] Durga Toshniwal: Right? So now, it means that if this is the kind of training data that is available, where the value pairs [01:56:47] Durga Toshniwal: are common, then what happens? All of the, transactions [01:56:54] Durga Toshniwal: Get mapped to a single subtree. [01:56:57] Durga Toshniwal: Or a single path. [01:57:00] Durga Toshniwal: If they get mapped to a single path, then there's no point doing a decision tree here, right? The goodness of the decision tree relies on the fact [01:57:09] Durga Toshniwal: or lies on the fact that it is bushy, it has, you know, it has two or multiple branches coming out of it, and all that stuff. Here, that facility is not there, because it's just a single [01:57:22] Durga Toshniwal: Structure with a single path and all. [01:57:24] Durga Toshniwal: So, it doesn't make any… it doesn't add any value for us. [01:57:29] Durga Toshniwal: It's only just making a single path, no matter what happens. So therefore, this kind of a tree is also not very useful, right? Because [01:57:39] Durga Toshniwal: It has just one path. [01:57:44] Durga Toshniwal: Then, I think we, I told you that how should we do the split? [01:57:50] Durga Toshniwal: We have to decide. [01:57:51] Durga Toshniwal: So, to do the split. [01:57:53] Durga Toshniwal: We can actually use two ways. One is MV split, one is binary split. MV means we could have, let's say, weight here. [01:58:04] Durga Toshniwal: And we are having 3 situations. [01:58:07] Durga Toshniwal: like, lightweight. [01:58:09] Durga Toshniwal: And then, you know. [01:58:13] Durga Toshniwal: Let's say medium weight, so this is light, this is medium, this is heavyweight, like that. [01:58:20] Durga Toshniwal: So… MV splits are also possible. There's nothing wrong in having it. [01:58:26] Durga Toshniwal: And, binary sprits are those where we are having [01:58:30] Durga Toshniwal: The records being splitted into two groups. [01:58:34] Durga Toshniwal: Okay. [01:58:35] Durga Toshniwal: And accordingly, the… The… The binary values, or the binarization of car type is also done. [01:58:46] Durga Toshniwal: That is one example. [01:58:51] Durga Toshniwal: So, now, the thing is that we had continuous valued attribute. For example, This one. [01:59:03] Durga Toshniwal: Any one for that matter. We can have any particular example. We could have a binary split, we could have an split. [01:59:11] Durga Toshniwal: And, whenever we are doing the discretization, there'll be a decision boundary. And based on the decision boundary, only we can do the split. For example, here, if the value of the attribute is less than, V, then it will be one particular, [01:59:28] Durga Toshniwal: discrete value, and if it is greater than or equal to V, there'll be another set of values. So, there's a decision boundary. [01:59:38] Durga Toshniwal: And based on the decision boundary, the values will get divided. [01:59:43] Durga Toshniwal: Okay. [01:59:47] Durga Toshniwal: So then we have, [01:59:50] Durga Toshniwal: Again, this I've already showed you, that on the basis of continuous valued attributes, we could have a binary split, we could have a [01:59:59] Durga Toshniwal: investment. [02:00:01] Durga Toshniwal: And both are correct. [02:00:04] Durga Toshniwal: Okay, now is… I'm going to answer the question that many of you had asked. How do we decide which split to use? [02:00:12] Durga Toshniwal: So, how do we decide that? Let's say that we have a class distribution. [02:00:19] Durga Toshniwal: Okay? And, let's say, the class distribution as fraction of records belonging to class I at any given node T. [02:00:30] Durga Toshniwal: And the class distribution, so there are two classes, C0, C1, and both have 1010 records each. [02:00:38] Durga Toshniwal: Okay. [02:00:40] Durga Toshniwal: So this is what we are having with us, 20 records with 10 belonging to C0 class category, and 10 belonging to C1 category. [02:00:50] Durga Toshniwal: Okay. [02:00:51] Durga Toshniwal: Now, Even, you know, without doing any split, we can see that [02:01:00] Durga Toshniwal: The class distribution is, like, 50-50 here, 10 records of C0 type, 10 records of C1 type. [02:01:06] Durga Toshniwal: Now, and let's assume we have these attributes, owning a car, car type, and student ID. These are the ones that are available to do the split. [02:01:16] Durga Toshniwal: Okay, so let's do the split on owning of the car. So, out of those 20 samples, we got 2 groups, 10 and 10. [02:01:26] Durga Toshniwal: Where owning a car, yes or no, with value, yes, it is C0C1, 6 is to 4, and with no, it is C0C1, 4 is to 6. Okay. [02:01:37] Durga Toshniwal: So, we use this, we divided it. Then we use the student ID. Accordingly, we divided the [02:01:44] Durga Toshniwal: It into these subsets, as you are seeing here. [02:01:48] Durga Toshniwal: And then we have the third category, where, based on the car type, we did a MV split, like family, sports, and luxury. [02:01:57] Durga Toshniwal: And then, we, you know, continue to do the tree. Now, the question is, which split is the best? You can see the outcome of the split, like C0 to C1, 6 is to 4, 4 is to 6. [02:02:10] Durga Toshniwal: Then we have 1, 3, 1, 2, something, something, and then at the end, we have, like, 8017 and all for this particular split. And for this split, we are having 10101, like that. So now, my question to you is which one is the best split? [02:02:29] Durga Toshniwal: Out of these three. [02:02:30] Deepak Bobade: C. C. [02:02:32] Durga Toshniwal: This is the best bit. [02:02:35] Deepak Bobade: Yes. [02:02:36] Gunjan Bhaiya: It will be A, because it is A to equal. [02:02:40] Nirav Mehta: Better. [02:02:40] Durga Toshniwal: Why? A will be better, you said? [02:02:43] Gunjan Bhaiya: Because the. [02:02:43] Nirav Mehta: It is equal distribution. [02:02:45] Gunjan Bhaiya: Near to equal distribution. [02:02:47] Durga Toshniwal: Okay. [02:02:48] Durga Toshniwal: Other than that, anyone else? Yes? [02:02:52] Sushree Dash: I feel it is B… [02:02:55] aditya shrivastava: And C is also even, because 1010 BO in all that are the same, right? Because even though it is multiple, but that distribution is also equal. [02:03:10] Durga Toshniwal: Okay, then what do you want to say? C is good or sees bad? [02:03:15] Ankit Sood: I think 3 is bad. [02:03:17] Durga Toshniwal: he's back. [02:03:18] Ankit Sood: Bad, yeah. [02:03:19] Gunjan Bhaiya: Yes, he is back. [02:03:20] aditya shrivastava: I think he is good, because multiple… [02:03:23] Gunjan Bhaiya: That is going to be increased a lot. [02:03:25] Ankit Sood: Correct. [02:03:26] Ankit Sood: Just based on… you are making. [02:03:29] laxmi sahu: The tree would be more deeper. [02:03:33] Durga Toshniwal: 3 should be more deeper, you are saying? [02:03:36] laxmi sahu: Yeah, with C, T will go more deeper. [02:03:41] Durga Toshniwal: With C, it is not going deep, it is going very deep. [02:03:44] Shikhar Gupta: Blessy. [02:03:45] Gunjan Bhaiya: But the number of students, like, 20 or 50 inches in number of breath is going to be increased a lot, which is going to be a good problem. [02:03:52] Gunjan Bhaiya: So, C is not a good… [02:03:56] Durga Toshniwal: Solid. [02:03:56] Ankit Sood: Yeah, I feel like with C, we are only considering one attribute, and we are actually making a decision. Whereas if we are actually splitting it, in the equal proportions, then we would try to consider maximum possible attributes. [02:04:11] Sushree Dash: Shouldn't it depend… [02:04:13] Durga Toshniwal: What we are actually deciding here. [02:04:16] Sushree Dash: Because with student ID, I don't know what we will be deciding. But if we are going by this B1, where we're deciding the car type, maybe we can go further, like, what is the family income, and why they opted for family car or sports car, like that. [02:04:34] Durga Toshniwal: What do you… what do you want to say? This is good, this is good, or this is good? [02:04:37] Sushree Dash: the B one is good, because we can decide, things here in B. [02:04:43] Sushree Dash: Cool. [02:04:43] Durga Toshniwal: Okay. [02:04:44] vinit shah: I might go with A, because it has, less number of, DC, like. [02:04:50] Sathish Prabu Chelliah Krishnasamy: Yeah, I also… I also… [02:04:54] vinit shah: liquid. [02:04:55] Durga Toshniwal: Want to go with the… [02:04:57] Ankit Sood: Yeah. [02:04:58] Dwarakesh T P: Yes, yes. [02:04:59] Sathish Prabu Chelliah Krishnasamy: I also want to go with you, because the number of decisions and, like, the deep also, it won't go much, actually. It will be, yeah, very quick, we can make a… decision, actually. [02:05:15] Ankit Sood: Also, it is actually… [02:05:16] Anurag Krishnam: Mom, I should be B. [02:05:20] Anurag Krishnam: The reason being. [02:05:22] Durga Toshniwal: green… [02:05:23] Anurag Krishnam: It's best for, testing the conditions, so I think it should be B. B. [02:05:29] Ravindra Singh: I think B would be better. [02:05:32] Durga Toshniwal: I'll just share my screen again. [02:05:37] Deepak Katara: I think either A or B, based on what we want to see as outcome. [02:05:41] Durga Toshniwal: Okay, so I think a couple of you said A, many of you said C, some of you said B. [02:05:48] Durga Toshniwal: And all that. So the answer to your question is, let's take it one by one. Let's take owning of the car. [02:05:56] Durga Toshniwal: So, when we take owning of the car, it is splitting on owning car, yes or no. And based on that, we get two, subsets, and the subset contains C0 is to C1, 6 is to 4, C0 to C1, 4 is to 6. [02:06:12] Durga Toshniwal: So remember, what was our objective? Our objective is that we want to do a split to obtain a, you know, clear consensus or a clarity about what is going to be the class label. But if we go for this particular split. [02:06:28] Durga Toshniwal: then what happens? We, obtain, you know, subsets. [02:06:33] Durga Toshniwal: But both these subsets, like, for example, the first one contains C026, C1 as 4. The other one contains C0 as 4 and C1 as 6. [02:06:45] Durga Toshniwal: So, they are almost 50-50 or equal, proportion. [02:06:50] Durga Toshniwal: And so, if we do a split of this kind, that is, owning of a car, yes or no, if we use this. [02:06:57] Durga Toshniwal: Then, this split that we are finally having is of no good, because it is containing a complete, [02:07:06] Durga Toshniwal: equal or a uniform mixture of the class labels. So, owning of a car is not good. We'll not go with this. [02:07:15] Durga Toshniwal: Because the split, definitely will give us, two subsets, but these subsets should be formed in such a way that they help us to form a clear majority. [02:07:29] Durga Toshniwal: Right? [02:07:30] Durga Toshniwal: So, definitely… Option 1 or A is not the best. Let's go for C. [02:07:38] Durga Toshniwal: So, if you look at C, remember what is the idea? [02:07:42] Durga Toshniwal: of classification. [02:07:45] Durga Toshniwal: So, what is the idea of classification? [02:07:48] Durga Toshniwal: So the idea of classification is to group the data in such a way that [02:07:54] Durga Toshniwal: We obtain, the groups that are purer, or that have, you know, clear majority about the class label. [02:08:02] Durga Toshniwal: But if you see here, what is happening, there is one record placed in one particular, you know, node, then in the next node, again, one sample, in the next node, another one sample. So, per node, there's going to be one one sample only. [02:08:19] Durga Toshniwal: And if they have… if we have, you know, let's say 100 students, then we'll have so many such [02:08:26] Durga Toshniwal: nodes, with each having, any one of those, like 1001, or, like that. So the purpose of classification is completely, you know, just destroyed here. We are not doing any classification, we are just putting one record as a separate group. [02:08:45] Durga Toshniwal: So, then, owning off a car using this is not a good option, right? Because it's leading to subsets which are having 50-50 of Class C0 and C1. Having a split-like student ID, [02:09:01] Durga Toshniwal: is not 100% good, but it is better than the others. Sorry, it is worse than the others, because the idea of doing classification is to group the data. Here, there's no grouping. [02:09:15] Durga Toshniwal: All of them are separately placed. [02:09:18] Durga Toshniwal: And suppose we have new, more number of students, this keeps on becoming very bushy. [02:09:24] Durga Toshniwal: So, this is not a good option, this is also not a good option. [02:09:29] Durga Toshniwal: Then, the best option will be this, why it will be because the subsets that we are obtaining are having a distribution of 1 is to 3, 8 is to 0, or 1 is to 7. [02:09:40] Durga Toshniwal: which are quite… which are simply showing a very nice majority. Here it is C, C1 as the majority, here again C0 as the majority, C1 as the majority, and these three are giving clear majority. Therefore. [02:09:55] Durga Toshniwal: car type. [02:09:57] Durga Toshniwal: Will be the one that will be… that is most preferred. [02:10:01] Durga Toshniwal: Because owning of a car is resulting to subsets, that is okay, but they don't have any majority, they're having 50-50 mix-up. [02:10:09] Durga Toshniwal: Student ID is giving us a pure, distribution of the subsets, that is true. So all of them are, you know. [02:10:18] Durga Toshniwal: Plays separately, but if you look at it. [02:10:22] Durga Toshniwal: We are actually not doing any grouping at all, we are just separating each point [02:10:28] Durga Toshniwal: In a group on its own. [02:10:30] Durga Toshniwal: So this is also not good. The best is this one. [02:10:34] Durga Toshniwal: That is the splitting done on car time. That will be the best. [02:10:39] Durga Toshniwal: And it will be most optimal. Here, it's resulting into 3 subcategories, with each of one having [02:10:46] Durga Toshniwal: you know, a clear majority of the class label, like C0 to C1, 1 is to 3. [02:10:54] Durga Toshniwal: So, it is going to be C1, then C0 to C1 is… 8 is to 0, so it's going to be C0. And then C0 to C1, it's going to be 1 is to 7, so again, this is also clear. [02:11:08] Durga Toshniwal: So, this is the best one. [02:11:10] Durga Toshniwal: Is it okay, everyone? Any questions? [02:11:13] vinit shah: So is it because the pure sets are comparatively more over here with Class B that it is better, ma'am? Like… [02:11:22] Durga Toshniwal: I couldn't understand. Can you please come back again? [02:11:25] vinit shah: So, the car type B, the classification B, we just have 3 branches, but it has comparatively higher pure subclasses, right? So, is that the reason that we're calling it out as better, or… [02:11:38] vinit shah: If I want to summarize it. [02:11:39] Durga Toshniwal: Yeah, so the objective… what is the objective? The objective… why are we doing the split? We are doing the split so that the subsets are having a clear majority of the class label. [02:11:52] vinit shah: Got it. [02:11:53] Durga Toshniwal: So, this one is satisfying this purpose. [02:11:56] Durga Toshniwal: This is not satisfying. Again, this is not leading to formation of any groups. This is all, you know, I mean, I should not say for… yeah, so these are no groups. [02:12:09] Durga Toshniwal: These are all individually placed only, so the idea of grouping is defied here. [02:12:14] Durga Toshniwal: So this is also not good. [02:12:17] Durga Toshniwal: This is West. [02:12:19] vinit shah: Got it, man. [02:12:20] Durga Toshniwal: Yeah, so I think we can wrap up here, if there are any questions. [02:12:24] Durga Toshniwal: You can ask me. [02:12:28] Durga Toshniwal: Right now. [02:12:32] Durga Toshniwal: Any questions, anyone? [02:12:38] Gunjan Bhaiya: And this is currently, we have done manually, right, with a small set of data and everything, but if you look at the larger part, right? So, any algorithm we have to use, right? The hands-on, where we can supply a bulk number of records. [02:12:51] Gunjan Bhaiya: Where we have, tens of attributes and thousands of rows, right? [02:12:56] Durga Toshniwal: Yes, so definitely, you'll be doing it algorithmically also, and this Hunts algorithm [02:13:03] Durga Toshniwal: for decision tree induction, it's a very popular one, and it's actually pre-coded and available to you in Python. [02:13:14] Durga Toshniwal: Okay. [02:13:15] Gunjan Bhaiya: Oh, so that… then it automatically, the entire tree structure, based on the looking data pattern and everything, or we have to… [02:13:21] Gunjan Bhaiya: So, configure the nodes and everything from our side. [02:13:25] Durga Toshniwal: So, definitely, you will have to give inputs about the parameters, like what will be the max depth, what will be the number of… there are various things that we'll… I'll show in the hands-on that… [02:13:37] Gunjan Bhaiya: Okay. [02:13:38] Durga Toshniwal: Oh, sorry, I think I'm loading the right one. So, you'll have to define the… some parameters. Definitely, yes, you will need to do that. [02:13:47] Durga Toshniwal: Oh… But anyway, the basic algorithm [02:13:51] Durga Toshniwal: That is working, will be pre-coded, and it will be available to you. [02:13:56] Durga Toshniwal: You don't really have to do it from scratch until unless you are asked to do it. [02:14:01] Gunjan Bhaiya: Okay. [02:14:03] Durga Toshniwal: Okay? Say, for example, you're, you know, you're asked to design a decision tree. [02:14:11] Durga Toshniwal: So if you're asked to design a decision tree, then you have to start from scratch and show [02:14:18] Durga Toshniwal: how each and every step works and all that, right? Otherwise… [02:14:22] Gunjan Bhaiya: Good. [02:14:23] Durga Toshniwal: Oh, it's available. [02:14:27] Gunjan Bhaiya: Okay. [02:14:28] Durga Toshniwal: Yeah. Any other questions, anyone, before we take part away? [02:14:32] Shikhar Gupta: And we need to check, like, before making a decision tree, that… and no two rows have same attributes and different results in the end. [02:14:40] Shikhar Gupta: like… [02:14:41] Durga Toshniwal: Yeah. [02:14:41] Shikhar Gupta: Yeah, that case. [02:14:43] Durga Toshniwal: Yeah, so if you have such cases, definitely you should not use this particular method. Otherwise, what you can do is, if there are such cases, then you can handle them at the time of pre-processing. [02:14:57] Durga Toshniwal: So, for such cases, what you could do is that when they are having similar, all similar values, then you could define some kind of a default class label, or [02:15:07] Durga Toshniwal: You know, you can do… you can have some ways to handle it. [02:15:11] Durga Toshniwal: Beforehand. [02:15:13] Shikhar Gupta: But, like, we cannot remove our data, because it might affect… [02:15:18] Durga Toshniwal: Yeah, so I'm not talking about removing of the data. [02:15:21] Durga Toshniwal: What I'm saying is, at the point where preprocessing is being done, you can… [02:15:26] Durga Toshniwal: You can devise some algorithm to pre-process, I mean, not algorithm, just some method to… [02:15:33] Durga Toshniwal: correctly pre-process the data. Or else, let's say that if you're having, the cases where attributes are same and class labels are different. [02:15:45] Durga Toshniwal: Then in that case, don't use decision tree, use something else. So, you can check it. Or otherwise, the second option, which is more common, is that if there are such [02:15:56] Durga Toshniwal: Value pair combinations, which are not unique. [02:16:00] Durga Toshniwal: And they are impacting your classification. [02:16:03] Durga Toshniwal: In such cases, both of them could be assigned some default class, like, there could be different ways of [02:16:10] Durga Toshniwal: Oh… dissolving this problem. That's what I'm saying. Instead of removing the rows. [02:16:18] Sushree Dash: Ma'am, does the library, whatever Python library we will use for this, will it, will it, I mean, tell us that whether decision tree will be best for this or not? [02:16:30] Durga Toshniwal: No, that… it won't be. That you have to decide. [02:16:33] Durga Toshniwal: What it will do is that if you invoke that library, it will build the decision tree, whatever it is like, if it is good or if it is bad, but it will build it. [02:16:45] Sushree Dash: But we can decide the root nodes for this, right? The attributes which will fall on the root note, or the labels. [02:16:52] Durga Toshniwal: So, yeah, you can decide some of the things. Some things, to some extent, can be decided by you. The rest will be decided by the algorithm. [02:17:03] Sushree Dash: Okay. Thank you. [02:17:07] Durga Toshniwal: Any other questions, anyone? [02:17:14] Durga Toshniwal: Okay, there are no further questions than, Weekend break for today? [02:17:20] Deepak Katara: Ma'am, just one question last, to wrap up. [02:17:22] Durga Toshniwal: Yeah, sure. [02:17:23] Deepak Katara: So, when we have some duplicate data, as we have seen, like, with the different classification with the same data, so, mathematically, what all, things or parameters do we have, which will indicate that, okay, we have such data? [02:17:37] Deepak Katara: Because if we have occurrence of such data, then probably we might need to take a step back and do the preprocessing again, right? So mathematically, what are the identifiers to understand your data set? [02:17:51] Durga Toshniwal: So the best thing to do is, as I've mentioned earlier, is that before doing anything on the data, I have a [02:17:58] Durga Toshniwal: Have some kind of, you know, information about the data, look through it. [02:18:04] Durga Toshniwal: Or, you know, just have a feel of the data, just have a look. [02:18:08] Durga Toshniwal: So that is important, or you identify you know… Oh. [02:18:14] Durga Toshniwal: whatever. Suppose there are… how many unique values are there? Or whatever is the point of contention, it's always good to [02:18:22] Durga Toshniwal: You know, understand the data before applying any preprocessing to it. [02:18:29] Durga Toshniwal: That is the thing. [02:18:33] Durga Toshniwal: Does that answer your question? [02:18:35] Deepak Katara: Okay, so we need to manually see through and understand the data. [02:18:39] Deepak Katara: Rather than relying on some equation. [02:18:42] Durga Toshniwal: Yeah, so a very good thing is to manually, in the sense that I'm not saying literally manually. For example, you can, you know, have a quick… you can find out what are the number of unique values [02:18:56] Durga Toshniwal: Then, how many records are having the same set of attributes? Like that, some simple things you can see on the data that you feel will create some kind of contention. [02:19:09] Durga Toshniwal: And then, once you see all that, and you, you, you know, you see there are not many or very few are there, then you can use it as such. [02:19:19] Durga Toshniwal: Otherwise, if… Multiple of them are there, then you need to probably choose something else. [02:19:26] Durga Toshniwal: And so on and so forth, like that. [02:19:31] Deepak Katara: Sure. [02:19:32] Durga Toshniwal: Cool. [02:19:34] Durga Toshniwal: Any other questions, anyone, before we wrap up for today? [02:19:41] Durga Toshniwal: Okay, so I think then, if there are no further questions, then when… then we can wrap it up for today. Tomorrow, I plan to have hands-on, though I'd expected that I will complete decision trees and also random forest. However, I have not been able to complete decision trees almost towards the end, but [02:20:01] Durga Toshniwal: Random Forest I'm not covered. So, would you like to take the hands-on on Random Forest, or not? [02:20:08] Durga Toshniwal: For tomorrow, I'm asking. [02:20:13] Durga Toshniwal: Because I haven't told you what is Random Forest. [02:20:16] Durga Toshniwal: So, would you still want to do the hands-on, and then, parallelie, I can explain to you [02:20:22] Durga Toshniwal: In the next turn, or something like that. Will it be okay? [02:20:26] Durga Toshniwal: Or we just, you know, how do you think we should do in the hands-on? [02:20:33] Durga Toshniwal: Should we do the random forest, or… [02:20:37] Durga Toshniwal: How should we do it? [02:20:39] Durga Toshniwal: So the hands-on that you have… Sorry, you want to power? [02:20:43] Swagat Pattnaik: I prefer to cover it before, and then, you know. [02:20:46] Swagat Pattnaik: Under the hands on. [02:20:48] Durga Toshniwal: Okay, that's why I wanted to ask. [02:20:51] Durga Toshniwal: That, because I plan to have hands-on, tomorrow. [02:20:55] Durga Toshniwal: But you haven't studied what is random forest and all that? [02:21:00] Deepak Katara: If it's a continuation in the topic of decision tree, so if it can give us the better picture, we can probably have some introduction tomorrow, and then later you can cover. Depends on how related it is. [02:21:14] Durga Toshniwal: The random forest definitely is related to decision tree. [02:21:18] Durga Toshniwal: Oh… [02:21:21] Durga Toshniwal: It is definitely related to decision tree, because a random forest is made up of multiple decision trees. The only question is that in the hands-on, because I haven't told you what is a random forest. [02:21:33] Durga Toshniwal: I mean, I'm telling you now, but I haven't told you a lot of other things about it, so do you want that to be covered in the hands-on? [02:21:42] Durga Toshniwal: Or you don't want to be covered in the hands-on tomorrow. [02:21:49] Sunil Saini: Momino. [02:21:51] Sunil Saini: In the real world, which one be used more? [02:21:56] Durga Toshniwal: Decision trees are also used, and random forests also used. Both are used equally. [02:22:02] Durga Toshniwal: So, it's very difficult to say which one will be used first. [02:22:05] Sunil Saini: But there might be some segregation, right, depending on the type of the data and the combination of it. [02:22:12] Durga Toshniwal: There might be segregation on what side? [02:22:15] Sunil Saini: So, which precision tree we should use? The random forest and… or maybe this one? [02:22:21] Sunil Saini: Bye! [02:22:22] Durga Toshniwal: I don't think there will be any such kind of decision-making thing possible. [02:22:27] Durga Toshniwal: It all depends, like… [02:22:30] Sunil Saini: It's all depend, then, how… [02:22:33] Sunil Saini: Using which method, it will provide us the right set of the answer, or right set of the tree. [02:22:42] Sunil Saini: As depend… I mean, there's no such clear-cut… Line, bottom line, that… [02:22:48] Sunil Saini: We should go ahead for this one and that one. [02:22:50] Sunil Saini: It all depends on the output, outcome. [02:22:53] Sunil Saini: computation tool. [02:22:56] Durga Toshniwal: You know, what do you want to say? Whether to use decision tree or not? Is that what you're asking? [02:23:02] Sunil Saini: No, no, I'm saying the random forest, or maybe this one. So, there's no other way to decide that which one we should go. [02:23:10] Durga Toshniwal: So that, actually, the only best way to decide is to look at the performance. [02:23:16] Durga Toshniwal: If the performance of a decision tree is good, you use it. If it isn't, then go for random forest. [02:23:22] Sunil Saini: That's all we can try. [02:23:25] Durga Toshniwal: Okay. [02:23:26] Sunil Saini: So anyway, right now, I… the question that I wanted to ask you, we haven't done random forest, do you want to do the hands-on? [02:23:34] Durga Toshniwal: Okay, let'. [02:23:34] Sonam Manwal: Yeah, we'll check. Can we go Theory first? [02:23:38] Gautam Sharma: Can we cover that first, and then we can do the address? [02:23:41] Durga Toshniwal: Yeah, okay. So let me see, once again, what all we planned in the hands-on. [02:23:47] Durga Toshniwal: So, we'll try to cover… [02:23:49] Durga Toshniwal: decision tree in the hands-on only, if we are having the hands-on. Otherwise, I will take the theory. [02:23:56] Durga Toshniwal: And then we'll take the entire thing, okay? So I haven't yet decided, let me have a look, and then, accordingly, we'll act. [02:24:04] Durga Toshniwal: Bye. [02:24:07] Deepak Bobade: Okay. [02:24:08] Durga Toshniwal: So… so I think with this, we'll wrap up here then. [02:24:12] Durga Toshniwal: Thank you so much, and have a great evening. [02:24:18] Durga Toshniwal: That's it, we'll meet tomorrow. Thank you, bye-bye. [02:24:21] Shivansh Sharma: Thank you, Lord. [02:24:22] Swagat Pattnaik: Thank you, ma'am. [02:24:24] Durga Toshniwal: Thank you. [02:24:25] Surya Bobbala: Thank you. [02:24:26] Neeraj Kumar: Thank you, bye-bye.