# 04 2025-12-07 Machine Learning Classifiers

course: Module 2 — Machine Learning Algorithms
module: Module-2-Machine-Learning-Algorithms
date: 2025-12-07
type: transcript
video_url: https://personal-learn.armco.dev/files/_Recordings/Module-2-Machine-Learning-Algorithms/04_2025-12-07_Machine_Learning_Classifiers.mp4

---
[00:10:31] Deepan Kanagaraj: Why?
[00:10:37] Deepan Kanagaraj: You know what?
[00:13:36] Deepan Kanagaraj: Very good morning.
[00:13:37] Durga Toshniwal: To all of you, and welcome to today's session.
[00:13:41] Durga Toshniwal: So… We'll be, starting.
[00:13:45] Durga Toshniwal: a new topic, so I'll just share my slides, and we'll go on from there.
[00:13:49] Durga Toshniwal: Just allow me a sec.
[00:13:56] Durga Toshniwal: That's mine.
[00:14:22] Durga Toshniwal: So now, in the last one, we had, discussed, about,
[00:14:28] Durga Toshniwal: classifier, what are classifiers? Then we started out with decision tree induction. I explained to you how to make a decision, tree.
[00:14:38] Durga Toshniwal: And in that, we discussed underfitting, overfitting, how to grow a tree, how to choose an attribute to grow at a time.
[00:14:47] Durga Toshniwal: To add in the tree.
[00:14:49] Durga Toshniwal: And that was based on entropy-based information gain.
[00:14:54] Durga Toshniwal: And then, we also discussed random forest, which was nothing but a set of trees.
[00:15:01] Durga Toshniwal: So, now, today, we'll be discussing about another very popular classifier, which is the…
[00:15:07] Durga Toshniwal: K nearest neighbor classifier, just allow me a second, I'll just… That might be.
[00:15:14] Durga Toshniwal: Let's done.
[00:15:24] Durga Toshniwal: So, in the K, your neighbor base, and so, first of all, our decision tree…
[00:15:30] Durga Toshniwal: Classifier involved building of a tree.
[00:15:36] Durga Toshniwal: As soon as the training data was made available, so when the training data was made available, the,
[00:15:44] Durga Toshniwal: Attributes were chosen one by one to grow the tree, and the decision tree was built.
[00:15:50] Durga Toshniwal: So, this kind of classifiers actually are called eager learners.
[00:15:55] Durga Toshniwal: And they are called eager learners because of the fact that, When, whenever the…
[00:16:04] Durga Toshniwal: Training data is made available.
[00:16:07] Durga Toshniwal: Such kind of classifiers use threat training data to build the classification model.
[00:16:14] Durga Toshniwal: And the classification model is nothing but the mapping of the input attributes to the output variable, which is the dependent variable, or the class variable.
[00:16:26] Durga Toshniwal: So, a decision tree classifier is called a eager learner, because it falls in the same category, that whenever the training data is made available, then it's going to build a classifier and keep it ready for making predictions in the future.
[00:16:43] Durga Toshniwal: So now, having talked about eager learners, definitely when we have eager learners, then there would be an opposite strategy also.
[00:16:53] Durga Toshniwal: So, the opposite strategy of eager learners Would just be what?
[00:16:59] Durga Toshniwal: The opposite of eager, that is, lazy learners.
[00:17:05] Durga Toshniwal: So, lazy learners are those, or lazy classifiers are those, that delay the process of learning of the training data.
[00:17:13] Durga Toshniwal: Until the time of making the classification. So, eager learners, when the training data is available, they learn the classifier and keep it ready to make predictions in future, whenever an unseen data is made available to them.
[00:17:29] Durga Toshniwal: However, lazy learners They just,
[00:17:34] Durga Toshniwal: They do not make any classification model ready.
[00:17:40] Durga Toshniwal: And they try to consult the training data at the time of making predictions.
[00:17:46] Durga Toshniwal: Also, a very good example of this, or rather a…
[00:17:52] Durga Toshniwal: A simple example of this I give.
[00:17:56] Durga Toshniwal: Many a times.
[00:17:57] Durga Toshniwal: In reference to students.
[00:18:01] Durga Toshniwal: So… There can be two categories of students.
[00:18:05] Durga Toshniwal: In fact, more than two, but however, we'll talk about these two categories. One, that as soon as the reading material is made available.
[00:18:15] Durga Toshniwal: The student will go through, And,
[00:18:19] Durga Toshniwal: I'll try to understand the concepts and keep the learning ready.
[00:18:23] Durga Toshniwal: To be, utilized later, whenever there's examination or whenever there's requirement.
[00:18:30] Durga Toshniwal: So, such kind of students may be what eager learners. And then, we have just the opposite.
[00:18:37] Durga Toshniwal: Which, actually, just sit with the data, sit with the documents, or sit with the material, and, just wait. They wait for the exam to come, and just before the exam, time is there, and when they are required.
[00:18:54] Durga Toshniwal: To, understand the material, then they start.
[00:18:58] Durga Toshniwal: Utilizing it, or then they start going through it.
[00:19:02] Durga Toshniwal: So, so, these two categories are very much analogous to eager learners and lazy learners.
[00:19:11] Durga Toshniwal: Eagle learners, those, which,
[00:19:13] Durga Toshniwal: Utilize the material, understand it, keep it ready. And lazy learners are those which keep the material with them and just sit with it until the time of exam.
[00:19:25] Durga Toshniwal: When they start actually, thinking of learning it.
[00:19:29] Durga Toshniwal: So then, in classifiers also, we have these two examples.
[00:19:34] Durga Toshniwal: So, in the category of eager learners, we had decision trees. Now we'll talk about lazy learners, in which we have instance-based classifiers.
[00:19:44] Durga Toshniwal: So, instance-based classifiers are those that just store the training data, or training records, and use them at the time of making prediction of an unseen data.
[00:19:58] Durga Toshniwal: So, for example, we have a set of, training records, as shown in the
[00:20:04] Durga Toshniwal: Diagram in the table on the slide, and there are some attribute values.
[00:20:09] Durga Toshniwal: And the instance-based classifiers just sit with the data and do nothing on it. And the moment they have an unseen data, they… what they try to do is… just give me a minute, I'll just…
[00:20:25] Durga Toshniwal: So, the moment they have some unseen data, and suppose these are some values, like.
[00:20:33] Durga Toshniwal: Just a second.
[00:20:35] Durga Toshniwal: I guess my…
[00:20:42] Durga Toshniwal: Say, for example, this contains values like X1, X2, something, something.
[00:20:47] Durga Toshniwal: Then they try to search all these, training sets and find the same combination of values.
[00:20:55] Durga Toshniwal: In the records, so they try to search.
[00:20:59] Durga Toshniwal: For example, some values are there, and maybe this is having XN. Then try to look for same combination of values.
[00:21:08] Durga Toshniwal: Something, something like this.
[00:21:10] Durga Toshniwal: And then whenever, let's say they found a match here, say, for example, this is a…
[00:21:17] Durga Toshniwal: record that is completely matching. Then.
[00:21:21] Durga Toshniwal: What they will do, look at this, look at these values, there is an exact…
[00:21:27] Durga Toshniwal: Match, then they will use this class label to make the inference.
[00:21:31] Durga Toshniwal: This is how.
[00:21:33] Durga Toshniwal: The instance which classifiers work.
[00:21:37] Durga Toshniwal: Now, and they are also called rote learners, so an example of instance-based classifiers is root learner.
[00:21:45] Durga Toshniwal: And a road learner will perform classification if and only if all the attribute values, like you saw here.
[00:21:53] Durga Toshniwal: Suppose there are 3 attributes, each of them will match. If there's an exact match between the attribute values in the training record and the unseen record, then they are going to use that exact match
[00:22:09] Durga Toshniwal: To infer the class label of the given record.
[00:22:14] Durga Toshniwal: So, any comments on this? Is it good, bad? How is it?
[00:22:18] Durga Toshniwal: This process where we try to
[00:22:22] Durga Toshniwal: Just match the attribute values of the,
[00:22:27] Durga Toshniwal: unseen record versus those of the given samples, which is the training data. And then try to see where there's a match, and whenever there's a match, that particular
[00:22:39] Durga Toshniwal: Sample from the training data is utilized to give the class limit.
[00:22:44] Durga Toshniwal: So, so, does this look nice and good, or there can be some issues in this also? Anyone?
[00:23:02] Durga Toshniwal: Yes?
[00:23:04] Neeraj Kumar: Yeah, actually, I think that it looks good, because it creates an instance of this, and I think that the process will be fast, as per my understanding, ma'am.
[00:23:17] Durga Toshniwal: Anyone else? So, it looks good.
[00:23:25] Sonam Manwal: Brookly
[00:23:27] Sonam Manwal: Predictions will be very few, like, if it is not getting exact match, it might not be able to predict.
[00:23:35] Durga Toshniwal: Okay, if there's no exact match, then…
[00:23:38] Durga Toshniwal: It may not be able to predict.
[00:23:42] Abhishek Diwan: Here, also, we'll have to store data in the primary memory. That issue will still exist.
[00:23:48] Durga Toshniwal: Okay, data storage may still be an issue.
[00:23:56] Durga Toshniwal: So I think let's go over these, inputs.
[00:24:01] Durga Toshniwal: So, first talking about the data storage.
[00:24:04] Durga Toshniwal: Definitely.
[00:24:06] Durga Toshniwal: The training data needs to be in the main memory.
[00:24:10] Durga Toshniwal: All the time.
[00:24:13] Durga Toshniwal: And if it is not in the main memory, if the training data is too big, then some input-output, so it needs to be swapped in and out of the main memory, and that might take extra time. And that is true for any classifier, not just instance-based classifiers.
[00:24:29] Durga Toshniwal: So, that is one thing. Now, the second question is that whether there is going to be an issue in this
[00:24:35] Durga Toshniwal: A rote learner in the basic form, or it looks nice and good.
[00:24:38] Durga Toshniwal: So I think already, as discussed by someone.
[00:24:44] Durga Toshniwal: that, -Oh.
[00:24:49] Durga Toshniwal: So,
[00:24:57] Durga Toshniwal: So, as discussed by a few of you, Definitely, rote learner.
[00:25:02] Durga Toshniwal: Has the limitation that we try to find. It tries to find an exact match between
[00:25:11] Durga Toshniwal: the training samples attributes, and those in the unseen data. But there could be possibilities, let's say that there are a number of attributes, and therefore having an exact match may always not be possible.
[00:25:25] Durga Toshniwal: And what if there's no exact match?
[00:25:28] Durga Toshniwal: To the unseen data in the training record.
[00:25:31] Durga Toshniwal: So then, in that particular case, the classifier will fail, because when it doesn't find an exact match, it doesn't really know what to do.
[00:25:40] Durga Toshniwal: So then This leads to… A modified version.
[00:25:45] Durga Toshniwal: off-road learner.
[00:25:47] Durga Toshniwal: And therefore, the most obvious thing is, instead of Exact match.
[00:25:55] Durga Toshniwal: We just go for an approximate match. So, this is what is the solution.
[00:26:01] Durga Toshniwal: Because approximate match definitely is a much better solution, because it is a match, but it doesn't rely on exact match.
[00:26:09] Durga Toshniwal: And so, to have an approximate match.
[00:26:12] Durga Toshniwal: What we do is that instead of looking for exactly similar, training samples, we just look for relatively similar training samples. And relatively similar means the attributes are having
[00:26:27] Durga Toshniwal: Similar values, but not exactly same. So, the…
[00:26:31] Durga Toshniwal: There will be some non-zero… so, in the first case, in the case of RoadLearner.
[00:26:37] Durga Toshniwal: the dissimilarity between the unseen record and the training data was, dissimilarity was zero. In this case, there will be some positive dissimilarity, that is, the records will not match exactly, but the
[00:26:53] Durga Toshniwal: They'll be similar to, to an extent which is tolerable.
[00:27:00] Durga Toshniwal: So then, based on this particular notion, what we have is… Call the nearest neighbor.
[00:27:07] Durga Toshniwal: So, now, to satisfy the approximate match constraints.
[00:27:12] Durga Toshniwal: What we do is that we look at the k nearest neighbors, or k closest points.
[00:27:18] Durga Toshniwal: To the unseen sample, and then use those to find out the class label.
[00:27:24] Durga Toshniwal: So, the basic idea here is that you know, any person If, the person
[00:27:33] Durga Toshniwal: You know, moves out, along with a set of
[00:27:36] Durga Toshniwal: People who are his close friends, then the nature of the person or his…
[00:27:42] Durga Toshniwal: Attributes must be similar, definitely, to the people in that group.
[00:27:47] Durga Toshniwal: Or the basic idea is that if there's an animal that looks like a duck.
[00:27:52] Durga Toshniwal: And moves, in the flock of duck.
[00:27:55] Durga Toshniwal: and coax like a duck, then probably it is a duck for sure. This is the same intuition that we use.
[00:28:03] Durga Toshniwal: Here.
[00:28:04] Durga Toshniwal: So…
[00:28:08] Durga Toshniwal: So, we have a very important and very popular nearest neighbor-based algorithm, classification algorithm, which is called the K-Nearest neighbor, or the KNN classifier.
[00:28:22] Durga Toshniwal: And in the KNN classifier, what it does is that, each and every point can be thought of as a data point in a D-dimensional space, where D will be the total number of attributes.
[00:28:34] Durga Toshniwal: Right? And we need to, identify an appropriate distance measure.
[00:28:40] Durga Toshniwal: This distance measure is going to compute the similarity between the unseen sample and the training samples in the D-dimensional space. And after measuring the similarity between the unseen record and the training sample.
[00:28:56] Durga Toshniwal: K number of points from the training samples, which are nearest neighbors, or they are closest to the unseen sample, are identified, and their class labels are used to infer the class label of the unseen record.
[00:29:13] Durga Toshniwal: Say, for example… sorry for this scribbling, I don't know. Anyhow, for example, we require, in this particular algorithm.
[00:29:23] Durga Toshniwal: Three things. One, the set of stored training data. So, here we have the training data.
[00:29:33] Durga Toshniwal: And then, we want an appropriate distance measure. It's very, very important, because if we use a long, wrong distance measure, then the notion of distance or similarity will be incorrect.
[00:29:46] Durga Toshniwal: And then we also need an appropriate value of K, which will help us to identify the… those many number of nearest neighbors to use.
[00:29:56] Durga Toshniwal: So now, so, when we are doing the classification, first of all, we'll have to compute the distance
[00:30:05] Durga Toshniwal: of the unseen record versus the other training records, and then identify, out of those distances the k-nearest neighbor, and then use the class labels of these k-nearest neighbor to find out the class label of the unseen record. For example, this circle here shows the unseen record.
[00:30:23] Durga Toshniwal: And we can use any distance measure, let's say we use Euclidean distance.
[00:30:29] Durga Toshniwal: to calculate.
[00:30:31] Durga Toshniwal: And then, suppose we use,
[00:30:36] Durga Toshniwal: k equal to 3. Then we'll consult the three nearest neighbors, so we'll find out the distance of this point with respect to all other points, right?
[00:30:46] Durga Toshniwal: All other points, because…
[00:30:48] Durga Toshniwal: A computer system doesn't have eyes to visualize or just see, so it will need to compute all the pairwise distances and sort them in a
[00:30:58] Durga Toshniwal: Increasing order, and then use the top 3.
[00:31:02] Durga Toshniwal: For example, here, these are the three neighbors which are closest to this unseen data.
[00:31:08] Durga Toshniwal: And then use their class labels to infer the class label. In this case, it will be a plus, to identify the class label of the unseen record.
[00:31:17] Durga Toshniwal: Let me just, stop here, I don't know what's happening, because… Oh…
[00:31:22] Durga Toshniwal: Give me a… give me a second.
[00:31:26] Durga Toshniwal: Sorry for that, as well.
[00:31:33] Durga Toshniwal: Just a second.
[00:31:43] Durga Toshniwal: Just a second, please, sorry for that.
[00:32:07] Durga Toshniwal: Just a moment.
[00:32:27] Durga Toshniwal: It's fine.
[00:32:51] Durga Toshniwal: So, we were here.
[00:32:56] Durga Toshniwal: Oh, I'll just choose my page.
[00:33:05] Durga Toshniwal: So then we have, in this case, say, for example, K is equal to 3.
[00:33:10] Durga Toshniwal: And then we'll use this value of K to infer the class label of the unseen record.
[00:33:16] Durga Toshniwal: And for that, we'll actually have to… to find out the… Closest point…
[00:33:22] Durga Toshniwal: We'll have to actually find out the distance of…
[00:33:25] Durga Toshniwal: this unseen record with respect to all other records, because actually we really don't know which one is going to be the closest one. So we'll have to actually find out the distance of this point with respect to all other points in the data space, like this.
[00:33:40] Durga Toshniwal: Because the computer system otherwise cannot understand, visualize, or interpret. And like that, it will do for all the points, and then arrange them in an increasing fashion. And then, out of these, identify the three pairs which are the lowest.
[00:33:58] Durga Toshniwal: So, this is how we identify 3 neighbors if we are talking about 3NN.
[00:34:05] Durga Toshniwal: So, the distance measure that could be used, as I mentioned, would be Euclidean, or it could be anything else. I just mentioned the…
[00:34:14] Durga Toshniwal: formula for Euclidean, for the sake of,
[00:34:18] Durga Toshniwal: recollection. So, Euclidean means that if we are talking about two data points, P and Q, then we find out the difference between each of the IATH attributes.
[00:34:29] Durga Toshniwal: square it, sum it, and then take out the root. For example, let's say if i is equal to 3,
[00:34:37] Durga Toshniwal: Then, this means that there are 3 attributes.
[00:34:40] Durga Toshniwal: So, let's assume… We have…
[00:34:44] Durga Toshniwal: to use it here. P… we have P equal to X1 by 1.
[00:34:48] Durga Toshniwal: and Z1, and P2 is equal to… X2, Y2, and C2.
[00:34:55] Durga Toshniwal: Then, the distance between P and Q. Using Euclidean distance will be equal to X2 minus X1 squared.
[00:35:05] Durga Toshniwal: plus Y2 minus Y1 squared.
[00:35:09] Durga Toshniwal: Plus… Z2 minus C1 squared, and whole of this rooted.
[00:35:16] Durga Toshniwal: Because here, I is equal to 3.
[00:35:19] Durga Toshniwal: And we are going to find out the distance between, the IF value, So, this is P1P… P1.
[00:35:28] Durga Toshniwal: And this is Q1.
[00:35:30] Durga Toshniwal: This is P2.
[00:35:33] Durga Toshniwal: This is Q2. F, P1, P2, P3 are the three attributes.
[00:35:38] Durga Toshniwal: And then we find out the distance between them.
[00:35:41] Durga Toshniwal: And then add them, square them, add them, and root them.
[00:35:45] Durga Toshniwal: So, we can use Euclidean distance, or we could use any other distance measure also.
[00:35:52] Durga Toshniwal: Now, the question is that we decided on using your nearest neighbor.
[00:35:56] Durga Toshniwal: And, we know that we are going to use K number of neighbors, or closest point, to the unseen data to decide the class name.
[00:36:05] Durga Toshniwal: However.
[00:36:07] Durga Toshniwal: K could actually take up multiple values. For example, in this case, we are having K equal to 1, if this point is the unseen data.
[00:36:15] Durga Toshniwal: Then, within the closest distance, if the circle designates the distance within which we are going to choose the sample, and here, only one record is coming, so K will be equal to 1.
[00:36:29] Durga Toshniwal: If, K has slightly higher value, in this case, K is equal to 2, this is the unseen record, and we are having two samples.
[00:36:38] Durga Toshniwal: Here, K is equal to 3. Like that.
[00:36:41] Durga Toshniwal: So, now the question is, what value of K shall we use?
[00:36:46] Durga Toshniwal: Should K be a small value, or what should it be?
[00:36:50] Durga Toshniwal: So, let's say K is a very small value, and let's say K is equal to 1.
[00:36:57] Durga Toshniwal: So what will… what will be your comments on it?
[00:37:00] Durga Toshniwal: Should K be equal to 1?
[00:37:03] Durga Toshniwal: Or should it be… Higher, or what?
[00:37:09] Durga Toshniwal: And by that time you all respond, I can see here that,
[00:37:14] Durga Toshniwal: Question that how do we take k equal to 3, whether it is primary 1 to check the closest point?
[00:37:21] Durga Toshniwal: I think the question, Lokesh, that you have is, how do we decide the value of K, if I understood it? And that's what we are discussing.
[00:37:29] Lokesh R: Yes, ma'am.
[00:37:31] Durga Toshniwal: We are discussing that only. Let us assume K to take up a very small value, and let it be equal to 1.
[00:37:38] Durga Toshniwal: Then, is it nice and good? Because we'll consult the nearest point.
[00:37:43] Durga Toshniwal: And then use it as the…
[00:37:45] Durga Toshniwal: As the closest neighbor to infer the unseen record. Will that be fine?
[00:37:52] Durga Toshniwal: Anyone, any comments?
[00:37:55] Aditya Banda: If he's…
[00:37:58] Durga Toshniwal: One by one, Pallavi, and then followed by who was speaking. Yes, Pallavi?
[00:38:03] Pallavi Chakravarty: So, let's say, if k is equal to 1, we are choosing the point with the closest distance.
[00:38:09] Pallavi Chakravarty: But let's say the second closest point is very close to the…
[00:38:14] Pallavi Chakravarty: Closest point, like, there's not much difference.
[00:38:17] Durga Toshniwal: Then we could not say that…
[00:38:19] Pallavi Chakravarty: Like, ideally, it won't be good. Say, one is positive and one is negative.
[00:38:25] Durga Toshniwal: Okay.
[00:38:26] Pallavi Chakravarty: Then it might be wrong to infer based on just the closest spot.
[00:38:33] Durga Toshniwal: Who was next?
[00:38:35] Aditya Banda: Yeah, just in addition to that, let's say we take K is equal to 3,
[00:38:39] Aditya Banda: The first is very close, the next two are a little bit far away, but not that far away, and because you're taking majority class.
[00:38:48] Aditya Banda: What if the other two are, belong to same class, and the first belongs to a different class?
[00:38:55] Durga Toshniwal: I'll discuss that.
[00:38:58] Durga Toshniwal: I'll definitely discuss it.
[00:39:00] Durga Toshniwal: Okay, any other comments, anyone?
[00:39:04] Ravindra Singh: maybe I like, ma'am, as you suggest, K, maybe we can take it, maybe 2 or either 3, and then from out of that, we can take top 1, or…
[00:39:14] Ravindra Singh: Topto, basically.
[00:39:16] Durga Toshniwal: Okay.
[00:39:21] Durga Toshniwal: So first, consider a case when K is equal to 1 or k is very small. Let's say it is as small as 1.
[00:39:29] Durga Toshniwal: The intuition could be it's nice and good that we take the closest neighbor and finish it.
[00:39:35] Durga Toshniwal: So, what will be the problem?
[00:39:39] Anurag Krishnam: Ma'am, in that case, I think data might even change the prediction if the data is wrong.
[00:39:45] Durga Toshniwal: Okay, data is strong means what in this case?
[00:39:49] Abhishek Diwan: So anyways.
[00:39:51] Abhishek Diwan: Nearby, and that may have… Maybe.
[00:39:53] Ravindra Singh: Yeah, we can miss the nearest one.
[00:39:58] Durga Toshniwal: Okay, fine. So, I think a couple of you, whatever you said.
[00:40:03] Durga Toshniwal: are correct. Let's say this is our unseen record.
[00:40:07] Durga Toshniwal: And then, we want to look at the neighbors.
[00:40:10] Durga Toshniwal: And let's say there are… it's a binary class problem, so there are two classes.
[00:40:16] Durga Toshniwal: And actually…
[00:40:22] Durga Toshniwal: So there's no point here, I…
[00:40:24] Durga Toshniwal: Don't want to use the eraser.
[00:40:27] Durga Toshniwal: Or just assume this is nothing.
[00:40:35] Durga Toshniwal: Okay, sorry for this, this is nothing.
[00:40:38] Durga Toshniwal: Now, if k is equal to 1,
[00:40:41] Durga Toshniwal: What will it do? It will find out the pairwise distance of all the points with respect to the unseen record, like this, and out of this, choose the one that is nearest.
[00:40:54] Durga Toshniwal: And let's say this is the point that is nearest.
[00:40:58] Durga Toshniwal: And it has a negative, class label. So the class label that would be inferred will be negative.
[00:41:05] Durga Toshniwal: However.
[00:41:06] Durga Toshniwal: Most of the nearest neighbors are having positive class label. So, what does it mean? It doesn't mean that this
[00:41:15] Durga Toshniwal: This point should have been… shouldn't have been there, but it does mean that there could be noise in the data.
[00:41:22] Durga Toshniwal: And this noise can have a very adverse impact on the class label if we are using just very small value of K, or maybe equal to 1.
[00:41:31] Durga Toshniwal: So, erroneously, the nearest point was a noise point, and its class label was chosen to infer
[00:41:40] Durga Toshniwal: The class label of the unseen record.
[00:41:43] Durga Toshniwal: So, therefore, it's definitely not a very good idea to have a, you know, a very small
[00:41:51] Durga Toshniwal: Value of K, because then it makes our classifier sensitive to noise.
[00:41:56] Durga Toshniwal: Now, let's come to the case of K to be very large.
[00:42:00] Durga Toshniwal: After,
[00:42:02] Durga Toshniwal: Very small, which makes it sensitive to noise points. Let's talk about very large value of K, or…
[00:42:08] Durga Toshniwal: Say, K is equal to something like 30, or 50, something like that.
[00:42:13] Durga Toshniwal: And let that… let this be the situation, which is shown by k equal to 30, where we assume that approximate number of samples included in the neighborhood are 30. So, if we count these points, all these points within this circle, which is the neighborhood, then we assume that there are 30 points.
[00:42:33] Durga Toshniwal: No.
[00:42:34] Durga Toshniwal: If we include, so K, very small, is not good.
[00:42:40] Durga Toshniwal: That we, that we just discussed, because it makes the classifier
[00:42:45] Durga Toshniwal: Was it, possibly sensitive to noise?
[00:42:49] Durga Toshniwal: Now, if we take K, larger, then will it be okay?
[00:42:56] Durga Toshniwal: Any idea, anyone?
[00:42:59] Anurag Krishnam: I think, again, it will become too generalized, and it might ignore the important details.
[00:43:05] Durga Toshniwal: Okay, it will…
[00:43:06] Ravindra Singh: It won't be good, I think. Unnecessary, you know, like, we have to calculate the distance, I mean.
[00:43:12] Ravindra Singh: So, it should not be…
[00:43:16] Abhishek Diwan: I mean, the points near the decision boundary, especially, will be, like, we won't be able to calculate them properly.
[00:43:26] Durga Toshniwal: So, if the value of K is very large, such as what you're seeing here, then what is happening is.
[00:43:33] Durga Toshniwal: That?
[00:43:34] Durga Toshniwal: actually case so large that the real neighborhood is getting lost. If we look at this point, which is the unseen record, then probably its real neighborhood is this.
[00:43:45] Durga Toshniwal: But we are extending it, so probably K should have been 5 in this case.
[00:43:52] Durga Toshniwal: Ideally.
[00:43:53] Durga Toshniwal: Right, but we are extending it to such an extent
[00:43:57] Durga Toshniwal: That lot of other samples which are not relevant to the unseen record are getting included.
[00:44:05] Durga Toshniwal: And as a result, the class label might get dominated by these samples, which are unnecessarily get included, getting included, and then what happens is that instead of inferring it as positive, this positive should have been there, but this won't be inferred. The negative class label will be inferred.
[00:44:25] Durga Toshniwal: And this is actually wrong.
[00:44:27] Durga Toshniwal: So, therefore, Here, what is happening? Unnecessarily, too much large value is getting considered.
[00:44:36] Durga Toshniwal: Now, coming to the question that one of you said, that we will be requiring too many computations here.
[00:44:42] Durga Toshniwal: So, so, once again, I'd like to know… Will, if…
[00:44:49] Durga Toshniwal: Sorry, K is equal to 1.
[00:44:52] Durga Toshniwal: be the case, or k is equal to 30 be the case.
[00:44:55] Durga Toshniwal: Which one will require more computations?
[00:44:59] Durga Toshniwal: Any idea, anyone, which one will require more computations?
[00:45:03] Deepak Bobade: A equal to 1.
[00:45:12] Durga Toshniwal: Anyone else?
[00:45:13] Krishnakumar MS: Both will be the same.
[00:45:14] Pallavi Chakravarty: It requires the same computation.
[00:45:16] Nirav Mehta: Because we are calculating distance from all the points.
[00:45:19] Pallavi Chakravarty: To get the nearest ones.
[00:45:22] Durga Toshniwal: Yeah, so actually, both will require the same amount of computations.
[00:45:27] Durga Toshniwal: Nothing will require less. Why? Because we… the algorithm won't know, or doesn't have this kind of visualization technique, so…
[00:45:37] Durga Toshniwal: He'll actually be required to calculate the distance of all the points with respect to the unseen record.
[00:45:45] Durga Toshniwal: Right.
[00:45:46] Durga Toshniwal: So, let's say if we are having some 100 points, then we'll have to calculate the pairwise distance between if there are P1, P2, P3, till P100, and this is the unseen record, say, X.
[00:46:00] Durga Toshniwal: So, we'll have to calculate the pairwise distance of all the points with respect to X.
[00:46:06] Durga Toshniwal: Right? And then, we'll have to identify the nearest neighbors. If k is equal to 1, we identify one of it. If k is equal to 30, we identify the nearest 30 neighbors, and use them to infer the class table. So, computation-wise, it's going to be the same.
[00:46:26] Durga Toshniwal: So then… Now, coming to the fact…
[00:46:30] Durga Toshniwal: That these are the two cases which we discussed, and this was the unseen record. And let's say
[00:46:38] Durga Toshniwal: These are some of the glass labels around it.
[00:46:48] Durga Toshniwal: This was unseen record.
[00:46:53] Durga Toshniwal: And then, Marine, one more diagram I'm drawing for the same purpose.
[00:47:11] Durga Toshniwal: Now, this was the case when K is equal to 1. When k is equal to 1, and we find out these distances, so… so this point comes nearest, and the negative class label is inferred. So, what is this case, actually? This case is showing overfitting.
[00:47:31] Durga Toshniwal: Why? Because…
[00:47:33] Durga Toshniwal: the class label is overly specific to this training data. Why? Because it is also considering noise point, which may be in the vicinity of this unseen record, as a correct point. So, this is the case of overfitting.
[00:47:53] Durga Toshniwal: Here, our number of neighbors is too small. It's k equal to 1 only.
[00:47:58] Durga Toshniwal: So, it's very, very specific, only just to one neighbor. Now, if we consider the case of k equal to 30, and we are having this kind of a decision boundary, in which we are considering all the neighbors, then instead of considering maybe 4 or 5 neighbors, it's considering 30 of them.
[00:48:15] Durga Toshniwal: And giving a class label, which is again wrong. Here, in this case, it was negative. And here, also, it will negative… it will be negative class label, because probably the number of
[00:48:26] Durga Toshniwal: negatively labeled samples were much higher in comparison to the positive one, because the value of K was too large. So, this is the case of underfitting.
[00:48:36] Durga Toshniwal: So remember, as we discussed earlier, we don't want underfitting.
[00:48:40] Durga Toshniwal: And we don't want an overfitting. We just want the correct,
[00:48:45] Durga Toshniwal: correct combination. So, how do we actually get The correct,
[00:48:51] Durga Toshniwal: value of K. So, the correct value of K may not be 1, it may not be very large.
[00:48:57] Durga Toshniwal: And let's say that we need to decide whether it's going to be 4, 5, 3,
[00:49:03] Durga Toshniwal: 10 what it should be. So, for that, what we will do is, we will draw a graph, which you'll see in the hands-on.
[00:49:11] Durga Toshniwal: Of the performance metric.
[00:49:15] Durga Toshniwal: Whatever that… that is. Say, for example, accuracy.
[00:49:19] Durga Toshniwal: Versus the value of K. So, what will happen is that…
[00:49:24] Durga Toshniwal: The perform… and… or we could have precision…
[00:49:28] Durga Toshniwal: recall, all such kind of things being drawn. So, what will happen is that…
[00:49:34] Durga Toshniwal: For… initially, the accuracy may be very low.
[00:49:39] Durga Toshniwal: And then it keeps on increasing with the value of k, and at some time it becomes plateaued out.
[00:49:45] Durga Toshniwal: You might have seen this yesterday in the hands-on also. So, the value of the K for beyond which
[00:49:51] Durga Toshniwal: If we increase the number of neighbors, and it doesn't increase the performance too much of the model.
[00:49:57] Durga Toshniwal: then we stop at… we stop at that value of K, or we just consider that value of K.
[00:50:05] Durga Toshniwal: So, this is how we decide the correct value of K. I think there's a question by one of you that how do we decide the correct value of K?
[00:50:12] Durga Toshniwal: We just draw a graph between the performance measures.
[00:50:17] Durga Toshniwal: Like, accuracy, precision, recall, and all.
[00:50:20] Durga Toshniwal: And,
[00:50:21] Durga Toshniwal: ROC and other things, and then see which value of K is giving us the best performance, and we choose that value of K. Usually, that value of K will not be too large, and it will not be
[00:50:33] Durga Toshniwal: Too small.
[00:50:35] Durga Toshniwal: So, this is how, the value of K is decided.
[00:50:40] Durga Toshniwal: Now, an important feature about, or important characteristic of k-nearest neighbor is that it does not maintain any, first of all.
[00:50:50] Durga Toshniwal: No abstraction or model on the training data is built.
[00:50:54] Durga Toshniwal: So, this method does not rely on any model finding. Whenever there's an unseen record, it's going to consult the training samples and use the training samples to make the inference. So, that is one thing. Secondly, the predictions are not based on any global,
[00:51:13] Durga Toshniwal: Global type of thing.
[00:51:15] Durga Toshniwal: But, so there are no global models that are used. For example.
[00:51:21] Durga Toshniwal: In a… in a classifier like a decision tree, the entire data is used
[00:51:27] Durga Toshniwal: to make the decision model, for example, the entire training data will be used to grow the tree, and once the trees
[00:51:37] Durga Toshniwal: built, then it is used to make the prediction. So, that is kind of a global model, because the entire training data is used.
[00:51:44] Durga Toshniwal: In this case, the entire training data is not used at all to make the prediction. Rather, only a local data is used, that is, whatever is closest to the given point. Only that particular training samples are used.
[00:51:58] Durga Toshniwal: So, so, because it relies only on the local information, therefore k nearest neighbor has a limitation that they may be quite…
[00:52:07] Durga Toshniwal: susceptible to noise. If our data is noisy, we need to be careful.
[00:52:13] Durga Toshniwal: Whether to use,
[00:52:16] Durga Toshniwal: The nearest neighbor, or not to use, or… or else we'll need to preprocess that data.
[00:52:22] Durga Toshniwal: Before actually making use of it.
[00:52:24] Durga Toshniwal: So, we discussed, two things, eager learners and lazy learners. Eager learners, for example, the decision tree where the entire model is built.
[00:52:34] Durga Toshniwal: Using the training data as soon as it is available, and then that model is used to make prediction. The other example of a learner was lazy learner, like the nearest neighbor.
[00:52:47] Durga Toshniwal: And in the k-nearest neighbor.
[00:52:49] Durga Toshniwal: There was no modeling that was happening. Whenever the unseen record was made available, it used the training samples to find out its neighbor and use those neighbors to infer the class name. So my question to you is which one is better? Is it eager learners, or is it lazy learning? Which one do you think is more useful or is more better?
[00:53:10] Durga Toshniwal: Anybody, any comments?
[00:53:14] Deepak Bobade: lazy learner, I think it will, reduce… Computing. Please?
[00:53:24] Deepak Bobade: lazy learning, because it will reduce the computing, right? Because it is computing, whenever we are trying to predict, right? It's not, like, training everything.
[00:53:35] Deepak Bobade: In the beginning.
[00:53:37] Durga Toshniwal: Okay, so you feel that the lazy learners are better?
[00:53:41] Deepak Bobade: Because…
[00:53:42] Durga Toshniwal: Computations are less.
[00:53:44] Deepak Bobade: Yeah. Yeah.
[00:53:45] Neeraj Kumar: Ma'am, actually… hello?
[00:53:48] Durga Toshniwal: Yes, yes, go ahead.
[00:53:50] Neeraj Kumar: Yeah, as per my understanding that, ma'am, whenever the large amount of data is there, then it is better to use a learner.
[00:53:57] Neeraj Kumar: And, as compared to whenever the small data means
[00:54:03] Neeraj Kumar: then it is better to use LazyTurner.
[00:54:05] Durga Toshniwal: Yeah, but assume that the training data wouldn't be small.
[00:54:10] Durga Toshniwal: Usually, the data may be large enough, or at least
[00:54:14] Durga Toshniwal: comparably large, because we need to learn all the use cases which are supposed to be predicted, right? So, usually it wouldn't be too small.
[00:54:24] Durga Toshniwal: Okay, so, so, I can see two people say Pavan and Sachin say lazy learners.
[00:54:32] Durga Toshniwal: Anyone?
[00:54:32] Nirav Mehta: Ma'am, won't it depend on the type of the data and the use case that we are trying to solve?
[00:54:40] Durga Toshniwal: Just allow me a second, just a second.
[00:55:00] Durga Toshniwal: Yeah, sorry.
[00:55:01] Durga Toshniwal: So, what was the comment? .
[00:55:05] Nirav Mehta: Yeah, won't it depend on the type of the data that we have, the training data, and the actual use case that we are trying to solve?
[00:55:14] Durga Toshniwal: Use case, so what do you mean by use case here?
[00:55:18] Nirav Mehta: Yeah, let's say if the, unseen data, can be easily classified based on the kind of, similarity, then we might go ahead with an easy diagnose.
[00:55:30] Nirav Mehta: But if it's, if it requires modeling.
[00:55:33] Nirav Mehta: Then we might go out with the ego learner.
[00:55:38] Durga Toshniwal: Okay.
[00:55:40] Guru Raghavendran: When the data set is small, we can go for laziness, is what I think.
[00:55:45] Durga Toshniwal: Okay, so as I already mentioned, you may actually not expect to have a small training data.
[00:55:52] Durga Toshniwal: Because usually the small training data might be missing… missing on the…
[00:55:57] Durga Toshniwal: all the use cases that this classifier is supposed to predict. Therefore, you can assume that the training data may be moderately, or it may be high or moderately high.
[00:56:11] Guru Raghavendran: Okay, then… then eager learners are the big term.
[00:56:16] Banu Prakash Reddy Mareddy: Ma'am, I feel the eagerlands are better.
[00:56:19] Banu Prakash Reddy Mareddy: Because, lazy learners seems like an algorithm where every, prediction is, going through the entire training data and, is trying to find the better
[00:56:32] Banu Prakash Reddy Mareddy: Least distance, already existing data.
[00:56:37] Banu Prakash Reddy Mareddy: So, if you develop a model with eager learners, that would be better.
[00:56:44] Durga Toshniwal: Okay, so let's discuss it one by one. So, in the Eager Learner, we,
[00:56:54] Durga Toshniwal: Just give me a second.
[00:57:01] Durga Toshniwal: What's going on?
[00:57:06] Durga Toshniwal: I'm trying to choose my pen.
[00:57:08] Durga Toshniwal: But… Just doing something, something.
[00:57:21] Durga Toshniwal: So, in case of eager learners, we have this,
[00:57:25] Durga Toshniwal: Model, which is built on the training data.
[00:57:29] Durga Toshniwal: We have model building, or in other words, the training time
[00:57:34] Durga Toshniwal: There will be some finite training time.
[00:57:38] Durga Toshniwal: training time that will be required to build this kind of a model, right? So, to use this kind of a learner, which is a train… easy learner, and training time will be high, because usually,
[00:57:51] Durga Toshniwal: The data may not be that very small. But once the model is built and it is kept ready, then the prediction time
[00:58:02] Durga Toshniwal: What's happening with my pen that I'm not able to write?
[00:58:13] Durga Toshniwal: Somehow, there's some issue.
[00:58:17] Durga Toshniwal: Continuously going here and… So I'll be prepared.
[00:58:24] Durga Toshniwal: So now, the prediction time here… Once the model is ready.
[00:58:29] Durga Toshniwal: is going to be very low, because prediction will be very, very fast, because the model will be ready in the memory, and as soon as unseen record is presented to it, it will quickly make the prediction. And in this, lazy learner case, training time
[00:58:48] Durga Toshniwal: Will be very low.
[00:58:50] Durga Toshniwal: Why? Because it doesn't make any model. And what about the prediction time?
[00:58:57] Durga Toshniwal: The prediction time will be high based on the amount of the
[00:59:02] Durga Toshniwal: Training data, because there's no model that has been built, and it's actually going to
[00:59:07] Durga Toshniwal: look at the training samples. At the time, it's,
[00:59:11] Durga Toshniwal: At the time, it's making the model, right?
[00:59:14] Durga Toshniwal: Now, the thing is that, one might feel that what good it is going to be when the prediction time is going to be very high, then
[00:59:24] Durga Toshniwal: What will be, the requirement? So, definitely.
[00:59:29] Durga Toshniwal: there is… so this is, the eager learner, because it relies on building the model, therefore prediction time is very less. So this means that it is very, very useful for cases where our prediction time is very, very, you know, requires to be very fast.
[00:59:47] Durga Toshniwal: So, suppose we are, you know, we are utilizing some safety-critical systems. For example, it is hitting off a target, and a missile is being launched. Definitely, in that particular case, we cannot afford to
[01:00:03] Durga Toshniwal: have the training done in real time, and then using the samples to make the prediction. At that time, our model should be ready, so that the prediction time is very small, and the correct, you know, the correct prediction can be done.
[01:00:20] Durga Toshniwal: So, for example, if we are trying to… if there's an image and we try to identify that this is the object that needs to be hit by the missile, then it needs to be very fast before the aircraft even moves past this. So, in such cases, eager learners are better.
[01:00:37] Durga Toshniwal: But that doesn't really mean that lazy learners are of no good, because
[01:00:41] Durga Toshniwal: There can be many cases where
[01:00:44] Durga Toshniwal: We may afford to have prediction time high, but training time may be low, or…
[01:00:51] Durga Toshniwal: for… or another case where actually training may not be possible. For example, we may be having a data
[01:00:57] Durga Toshniwal: which is a data stream, where data is coming in continuously, you know, continuously the data is coming in. For example, there might be a sensor that is deployed, and the sensor is measuring the temperature
[01:01:12] Durga Toshniwal: off.
[01:01:13] Durga Toshniwal: You know, some particular region or location. And we want to predict what will be the temperature in,
[01:01:21] Durga Toshniwal: You know, some next few hours.
[01:01:23] Durga Toshniwal: In such cases, what is happening is that, because the data is not finite at all.
[01:01:31] Durga Toshniwal: Data is just coming in.
[01:01:38] Durga Toshniwal: So, in this particular case, what is happening is, because the training data, and this training data is just continuously coming in, because it is being measured by the sensor.
[01:01:48] Durga Toshniwal: And it is never becoming complete.
[01:01:51] Durga Toshniwal: And since the training data is continuously coming in, and it is not finite, and it is not…
[01:01:57] Durga Toshniwal: Ending at any time. Therefore, we cannot build a model on a data when, on a training data which is not complete and which is having samples continuously coming to it.
[01:02:07] Durga Toshniwal: So, in such cases, it is not possible, actually, to build a model on the data, and therefore, in such cases, we need to utilize lazy learners only.
[01:02:18] Durga Toshniwal: And that will actually consult the training data at the time of making prediction.
[01:02:24] Durga Toshniwal: So therefore, we can really say that eager learners rely on global models, training time is high, although, but prediction time is very low.
[01:02:35] Durga Toshniwal: So, therefore, it is useful for such applications. Whereas, in lazy learners, the training time is really very
[01:02:44] Durga Toshniwal: low, because…
[01:02:45] Durga Toshniwal: hardly any minimal training is done, and it is just a local training. So, in cases where local information is of utmost importance, and prediction time can be anything but training time needs to be low, then we use lazy learn.
[01:03:00] Durga Toshniwal: So this is how we decide, whether to go for lazy learners or to go for eagerland.
[01:03:09] Anurag Krishnam: Any comments on the quality or accuracy of these learners?
[01:03:16] Durga Toshniwal: Okay, so the, quality or accuracy, actually, we cannot comment on it, because the two, work in a completely different fashion.
[01:03:28] Durga Toshniwal: One rely on a global model. Therefore, it looks at the global intricacies of the data.
[01:03:34] Durga Toshniwal: The other relies on local model, and therefore, it is good for those applications where there might be local, local variations, or local fluctuations in the data.
[01:03:48] Durga Toshniwal: So, we cannot say that this one is going to give a better performance, or the other one is going to give better. It purely depends on the data on the application.
[01:03:57] Durga Toshniwal: and what we want to do, right? For example, as I said that
[01:04:03] Durga Toshniwal: If we want the prediction time to be low, then we use Eagle Learner. But at the same time, if we want some local, information to be given more importance as compared to the entire data. For example, there might be some region of interest.
[01:04:19] Durga Toshniwal: Say, for example, if we have the map of the estate, and we are not interested in all the districts, we are interested in a specific district.
[01:04:29] Durga Toshniwal: then it might be lazy learners that may be better, because we are having a region of interest, and that vicinity we want to see. So, it's difficult to say, in terms of performance, which one will be better, but it… the use case varies.
[01:04:45] Anurag Krishnam: Thank you.
[01:04:46] Durga Toshniwal: Yeah. Any other questions, anyone?
[01:04:54] Durga Toshniwal: So, because if there are no further questions… Then,
[01:05:01] Durga Toshniwal: We'll break for the theory, and because yesterday in the hands-on.
[01:05:06] Durga Toshniwal: What was happening was that random forest was left.
[01:05:09] Durga Toshniwal: So then we'll continue with Random Forest Hands-On that was left yesterday.
[01:05:14] Durga Toshniwal: So, I'd like you to continue there. So, any other questions, or shall we break and continue with the hands-on of the part that was left yesterday?
[01:05:27] Deepak Bobade: Yeah, I have one question. I got confused with the, K. K is our hyperparameter, right?
[01:05:34] Durga Toshniwal: Yes.
[01:05:35] Deepak Bobade: So, like, is, okay, the, you could say cap or boundary, right? For, like, we should look for a number of, you know, neighbors. Is it what it's supposed to be?
[01:05:49] Durga Toshniwal: Yes, so…
[01:05:49] Deepak Bobade: Make the…
[01:05:50] Durga Toshniwal: for K number of neighbors.
[01:05:53] Deepak Bobade: Whatever, maybe.
[01:05:54] Durga Toshniwal: Yeah, so when we look for K number of neighbors, in a way, we are actually looking at if the unseen record is the radius.
[01:06:02] Durga Toshniwal: and there are, let's say, K number of neighbors, then in a way, we are actually looking at a circular boundary only, because those k-nearest neighbors will fall in a certain radius around the unseen record, and that's the reason in the diagram that I showed.
[01:06:16] Durga Toshniwal: I've drawn those boundaries, just to indicate that though here it is based on the nearest neighbor, but we are actually looking at some epsilon neighborhood.
[01:06:25] Durga Toshniwal: Of that unseen record.
[01:06:31] Deepak Bobade: Okay, okay.
[01:06:34] Durga Toshniwal: Any other questions? Anyone?
[01:06:41] Durga Toshniwal: Okay, so I'll just, have the… Hands-on started.
[01:06:47] Durga Toshniwal: So, I hope you all enjoyed the hands-on yesterday, where we started out with the decision tree.
[01:06:55] Durga Toshniwal: On a dataset, and then, we illustrated how you could actually
[01:07:00] Durga Toshniwal: Choose the different parameters associated with the decision tree, like depth and all the number of
[01:07:06] Durga Toshniwal: So today, you'll be, going ahead and doing random forest. In a random forest, you'll be having number of trees, the depth of the tree, and all that stuff as the hyperparameters.
[01:07:18] Durga Toshniwal: That you can use. Of course, in the default function also, there are values that you can use directly.
[01:07:25] Durga Toshniwal: So,
[01:07:27] Durga Toshniwal: Let me see if we have Ashish here, who had taken you through the rest of the product.
[01:07:32] Durga Toshniwal: Yesterday.
[01:07:35] Durga Toshniwal: Oh.
[01:07:37] Durga Toshniwal: So, once Ashish joins in, then, he'll start with the hands-on.
[01:07:44] Durga Toshniwal: Any, any comments, anything regarding the hands-on, any questions regarding what you did yesterday?
[01:07:52] Durga Toshniwal: In the handsome.
[01:07:54] Durga Toshniwal: I hope you all followed the hands-on.
[01:08:01] Durga Toshniwal: Yeah. Good morning, Ashish. Welcome to the session.
[01:08:05] Ashish Kumar: Good morning, everyone.
[01:08:08] Durga Toshniwal: So, I'll request Ashish to share his screen.
[01:08:12] Durga Toshniwal: And carry on from where, he had left yesterday.
[01:08:16] Durga Toshniwal: I believe decision trees, was covered.
[01:08:20] Durga Toshniwal: Random forest is still there.
[01:08:23] Durga Toshniwal: So, Ashish, over to you to carry on with the hands-on session.
[01:08:30] Ashish Kumar: Good morning, everyone.
[01:08:32] Ashish Kumar: So, we were, doing…
[01:08:34] Ashish Kumar: Yesterday, just a quick recap, we were doing… we were working on telco customer churn dataset, and there's a classification problem where the target variable is churn, and we have to, like, predict whether the customer will churn or not. It's a binary classification problem.
[01:08:53] Ashish Kumar: And,
[01:08:54] Ashish Kumar: This is the first part, where we were loading the data frame. So, we'll pick up from where we left yesterday.
[01:09:03] Ashish Kumar: And, so if I go to, like, table of contents.
[01:09:07] Ashish Kumar: We have finished A and B.
[01:09:10] Ashish Kumar: Of the decision tree.
[01:09:12] Ashish Kumar: One more way of optimizing hyperparameter is left footage entry, and then random forest. So…
[01:09:18] Ashish Kumar: I am at the heading, if you can find in your notebook, another way to find optimal depth.
[01:09:24] Ashish Kumar: I am at this place.
[01:09:27] Ashish Kumar: So, can everyone, like, in the notebook follow me? I'm at this place.
[01:09:31] Durga Toshniwal: Yeah, please let us know if you all are ready, and then we can start with the hands-on.
[01:09:38] Durga Toshniwal: Or you would like us to wait for a minute to load the… Sport and everything.
[01:09:45] Durga Toshniwal: Is it okay, everyone? Chill?
[01:09:47] Durga Toshniwal: Shall we begin?
[01:09:51] Swagat kumar pattnaik: Yes.
[01:09:52] Lokesh R: kiss.
[01:09:54] Ashish Kumar: Okay.
[01:09:57] Ashish Kumar: Okay, so… okay, so, before this, we were… till now, we were doing decision tree, and we were finding ways how to find the optimal depth, and how to find the optimal train and test split size.
[01:10:11] Ashish Kumar: So, till now, in decision tree, what we were doing is, as we discussed yesterday, like, we splitted the dataset in 90 to 10 train and test size ratio, and based on that, we selected a depth
[01:10:25] Ashish Kumar: And that depth came out to be depth of 6, and then use it throughout our code to decide what will be the optimal train-test ratio. But this can be, like, reversed also, that you fix a train-test ratio.
[01:10:39] Ashish Kumar: And then, on that trend test ratio, you find the optimal depth.
[01:10:44] Ashish Kumar: So, this section is just reverse of the previous section. So, basically, like, people were asking, like, how to find optimal depth, like, various scenarios. So, one scenario was that, the earlier one, where we were using… we first split using 90-10 ratio, and then we were finding optimal depth.
[01:11:03] Ashish Kumar: Now we are fixing the ratio first.
[01:11:06] Ashish Kumar: And then, over that split ratio, we will find the optimal depth and see the performance.
[01:11:11] Ashish Kumar: So, let's begin with the code, so…
[01:11:15] Dwarakesh T P: Hey, Arashesh, could you share me the notebook link once? I lost it.
[01:11:20] Dwarakesh T P: Honestly.
[01:11:23] Ashish Kumar: Notebook link, maybe…
[01:11:33] Ashish Kumar: Maybe manager can just…
[01:11:38] azad choubey: I pinged in the chat, you can take it from there. That zip file is there, okay? You can take it.
[01:11:44] Dwarakesh T P: Great, thank you.
[01:11:45] Durga Toshniwal: Simran, are you there? Can you please share the link for these, hands on file.
[01:11:52] Durga Toshniwal: If you had the… I think it is in LMS Portal also.
[01:11:57] Durga Toshniwal: Yeah, so, you can actually log into the LMS, and you'll find the link to all the materials there itself.
[01:12:09] Durga Toshniwal: Okay, I think, Ashish, you can carry on. Dwarkish has found… Okay.
[01:12:15] Ashish Kumar: So let's begin, like, the… the way we are doing.
[01:12:19] Ashish Kumar: The final optimal depth in this part.
[01:12:22] Ashish Kumar: And so, again, some of the, like, module… import statements is there, so…
[01:12:30] Ashish Kumar: Yeah, so this is SKLN.model selection. We are importing the train test split, which will do… which will randomly select a subset of the dataset as train, and subset of a dataset as test, and bring for us, so we don't have to write the code. Code is inside this thing, so…
[01:12:48] Ashish Kumar: How to use it? So, this is the function to interpret. You give it your X, you give it your Y. So, yesterday we saw that
[01:12:59] Ashish Kumar: X and Y were, somewhat like…
[01:13:05] Ashish Kumar: Yeah, so this is, like, original total data. Say, we have 7032 samples, and X is, like, 29 features are there. And for those 7,000 samples, Y is a single list of entry of 1 or 0 for each sample. So, this is the original data. But we have to split in train and test, so we give this original data here.
[01:13:27] Ashish Kumar: and then tell it that this is my… will be my train size, 60%, because I've set this 0.6, so 60% of it will be train size. Then, there's a random state, to fix the randomness that, for each time. And then there'll… there's a stratify equal to Y, which… which tells me to keep the same
[01:13:47] Ashish Kumar: distribution of, classes of the, of, that were in the original dataset for the Y label. So, a lot of people were asking yesterday for this, like, what, what is this? So, more illustration is, like, suppose this was my original dataset, Y,
[01:14:06] Ashish Kumar: And in that proportion was 73% and 26%.
[01:14:10] Ashish Kumar: Now, when I, split, when I do trained split, using stratify equal to Y,
[01:14:17] Ashish Kumar: You can see that for train also, and test also, the proportion of 0 and 1 labels, they are, like, same.
[01:14:26] Ashish Kumar: So, if we did… had not used stratify equal to Y, they may or may not be same because of randomness, but since we are specifying that we have to group them according to the Y class.
[01:14:39] Ashish Kumar: the distribution of the Y class, so that's why they are exactly mimicking the distribution which was originally in Y.
[01:14:48] Ashish Kumar: So, that's what, like, the stratify equal to Y do in print a split.
[01:14:54] Ashish Kumar: So, once we call this function, we have, like.
[01:14:59] Ashish Kumar: We get… in this sequence, we get four variables. So, first is, like, your X train, you will get your 60% of data of the X. Then your remaining 40% of data of the test X. So, first two are, Xs, train X and TestX.
[01:15:17] Ashish Kumar: And then, similarly, train Y and TestY.
[01:15:21] Ashish Kumar: Now, we have got this, and we will now, like, so first size at 60%, so on this, we will find… on this 60% of data, we will train a decision tree for various values of depth, and see
[01:15:39] Ashish Kumar: on which we are having the highest performance on test set. So, let's begin. So, here is, some variables.
[01:15:49] Ashish Kumar: To collect the results. So, this is a list, which is right now empty, train accuracy. This is another list, test accuracy, right now empty. So, as I… as we see, there are 29 features.
[01:16:01] Ashish Kumar: So, in order to loop through them, I have… there's a list of… in the range 1 to 30, and 30 is out of, like, inclusive number is 29, so it will look to 29, and 30, it will not go to 30 value. So…
[01:16:18] Ashish Kumar: Here, now I am looping through this range, 1 to 29.
[01:16:23] Ashish Kumar: So, for each value of depth.
[01:16:26] Ashish Kumar: I'm calling my decision tree classifier.
[01:16:29] Ashish Kumar: criteria is… can be, like, entropy or Guinea, like, I have used here entropy.
[01:16:35] Ashish Kumar: Then, max depth is whatever you will get from this range, so initially it will be, like, 1, then 2, like that.
[01:16:41] Ashish Kumar: And once you initialize your classifier.
[01:16:44] Ashish Kumar: with the parameters, hyperparameters. Then you call the fit function, which will
[01:16:52] Ashish Kumar: train, on which you pass your training data. So my training data is X train and Y train.
[01:16:58] Ashish Kumar: I pass it through this… Classified here.
[01:17:02] Ashish Kumar: And at this point, as we discussed yesterday, the training finishes, so training happens at this point. Now, once it is trained.
[01:17:11] Ashish Kumar: We will collect two values, to, like, see how our model is performing during training and test time. So, for that, we call the predict function.
[01:17:22] Ashish Kumar: So, there is clf.predict.
[01:17:25] Ashish Kumar: And what this predict function does is, you just give it the X part, and it will give you the prediction, the Y part of itself. And that too also in 1 and 0 form, because right now it's binary classification, so it will give you a value of either 1 or 0.
[01:17:44] Ashish Kumar: So, I, I pass it X trained, so, because I want to see how much is my, like, predictions, even on, like, the data on it was trained, how much, how it is, how it is performing. And also, more important is.
[01:17:58] Ashish Kumar: when I do CLF.predict on test set, because this is the dataset which was not used for training, so we want to see the performance on this.
[01:18:07] Ashish Kumar: So, predictions of train, they are, like, stored in this variable, train underscore print.
[01:18:12] Ashish Kumar: Predictions of, test set, they are scored in this variable.
[01:18:16] Ashish Kumar: And, then… Like, this is a, so basically, Just to show…
[01:18:29] Ashish Kumar: So, yeah, so basically, it's just a train prediction will, finally be a list.
[01:18:37] Ashish Kumar: Where each element is either 0 or 1.
[01:18:42] Ashish Kumar: Like this. So, yeah. So, like this. And we also have a…
[01:18:48] Ashish Kumar: Like, the Y train is my ground truth variables.
[01:19:00] Ashish Kumar: So, white tooth is my ground-tooth variable, and this is what my model predicted. Assume it like that. So.
[01:19:08] Ashish Kumar: Now, based on these two values, we can, like, compute accuracy. We can either write a function which looks at each entry, and if they're equal, it can compute accuracy and everything.
[01:19:19] Ashish Kumar: And so, basically, for this… all of these metrics in SQLN, there is already functions defined, so we use them. So, accuracy score is a function to which you pass this predicted list, and your… to which you pass your ground truth list and your predicted list, and it will…
[01:19:39] Ashish Kumar: give you a single scalar accuracy value for this particular depth, and which we are, like, adding in this list. So, it is a list.
[01:19:49] Ashish Kumar: And at each value of the loop, we will add one accuracy score. Similarly, for the test data, this Y test, Y underscore test is the ground truth.
[01:20:01] Ashish Kumar: And, well underscore, is that what we would got after prediction. So, we obtain accuracy, for this part also, and, whatever is the accuracy for this particular depth, this will be stored in this list.
[01:20:16] Ashish Kumar: And finally, train ACC and test ACC. They will be like…
[01:20:22] Ashish Kumar: They are like a list, if you see.
[01:20:29] Ashish Kumar: Yeah, so this is basically a list of 1 to 29 depth, accuracy at each value of the depth 1 to 29, and similarly, train test ACC. So, yeah, so this is, like, something.
[01:20:43] Ashish Kumar: So, since my depth was 29,
[01:20:47] Ashish Kumar: Since my depth was 29, so this is a type… this is a value for accuracy at each of the depths. And similarly, test accuracy.
[01:20:54] Ashish Kumar: Now, like, this is a 1D range of value of 1 to 29. For trade and test. We can directly plot them. So that's why, after that, there's a plt.figure.
[01:21:06] Ashish Kumar: this basically matplotlib function, PLT, which is telling… first thing, I am, like, specifying how big is my figure size to be. It's just for that. Now the actual plotting happens here, plt.plot.
[01:21:21] Ashish Kumar: On X, I want to specify the… on X axis, I want to plot the range of the depth, which is 1 to 30.
[01:21:29] Ashish Kumar: On Y axis, the value, which is the accuracy value, that, and I'm specif- the label for the first plot is… because both of these will be plotted on the same graph. That's why there is a, like, I'm specifying a label. That first plot, there is a label train accuracy.
[01:21:49] Ashish Kumar: And for… and there's some… a marker to denote, like, to differentiate between… so marker is basically… you can see the dots in the blue one, they are, like.
[01:22:00] Ashish Kumar: O, O, kind of O, and in other, there are, like, the shape is different, so marker, means that.
[01:22:07] Ashish Kumar: So, where we… yeah. So, markers for that. And then…
[01:22:13] Ashish Kumar: These are just formatting functions, so basically on X, X is I want to, put a label, put a label with the name tree depth. On Y axis, I want to put a, information that this is accuracy, so these are, like, that information. This is a title for overall graph.
[01:22:30] Ashish Kumar: And, yeah, and then I want to, like, add a legend to tell that, like, what is trade and test differentiator of the color.
[01:22:43] Ashish Kumar: And, this is also, like, optional if you want to grid in the background or not. And then, finally, plT.show, it will display this plot.
[01:22:53] Ashish Kumar: So, yeah. So, after this part.
[01:22:56] Ashish Kumar: So basically, this, graph gets displayed, after this part.
[01:23:03] Ashish Kumar: Till… till here. Yeah, so till here, this graph got displayed.
[01:23:07] Ashish Kumar: And again, as we see earlier, so there's always this scenario of overfitting and underfitting, so as we go on increasing the depth.
[01:23:17] Ashish Kumar: The train accuracy increases.
[01:23:20] Ashish Kumar: And it reaches, like, more than 95% when the depth is higher.
[01:23:26] Ashish Kumar: But we see that after depth of file, there is a sharp reduction in test accuracy.
[01:23:34] Ashish Kumar: So, so basically, this is the point where the model starts to overfit.
[01:23:41] Ashish Kumar: So, this part is… the first part of the code is this graph here.
[01:23:48] Ashish Kumar: Yeah, we are here. So, once we get, once I get this.
[01:23:53] Ashish Kumar: I'm using the… I'm using the test ACC, which is a list of the accuracy for all the depth sizes. I am printing first few of them, first 15 of them.
[01:24:06] Ashish Kumar: And,
[01:24:08] Ashish Kumar: which is, like, displayed here. So, I'm printing few of them, and within this thing, either I can manually… I can manually see that highest accuracy is on 5, or I can put a max over this list.
[01:24:23] Ashish Kumar: and obtain the maximum, like, the depth value for which the accuracy is highest. So, here, instead of doing manual inspection, I've just added… so here, yeah. So, the printing happens here for the first 15 depth sizes, and then I apply a max.
[01:24:42] Ashish Kumar: over the, depth value, depth value, and it will give me the best depth, which is here 5, and that gets, like, the print statement is also there to denote what is the best size. Now, I get the, best depth size. Now, this depth size, I will use.
[01:24:59] Ashish Kumar: To train our final model. So, this is, like, the whole pipeline. So, again, the split is same. So, basically, split is same, 60-40, but we were not sure which depth size was optimal. So, we did this plotting, and…
[01:25:17] Ashish Kumar: and store train and test accuracy to see which is the depth size at which our model… after which our model reduces its accuracy. Then we obtain 5. So, again, this is like the same train-test split, the same code.
[01:25:34] Ashish Kumar: Then I, use the decision tree classifier again, and this time, I'm passing it to the best depth, which is, like, coming as 5.
[01:25:42] Ashish Kumar: And again, we are doing model training.
[01:25:46] Ashish Kumar: So, this is, like…
[01:25:48] Ashish Kumar: CLF.fit, again, and, then, what we do is, the same, the pipeline is the same, just now, right now, I'm not looping through, just on… working on the single depth.
[01:26:02] Ashish Kumar: So, then I use the predict function to get my… to get predictions for train and test data. Test data, we… train data, we usually don't have to, like… we are using only for comparison purpose, but what matters more is our performance on test data.
[01:26:19] Ashish Kumar: And then we obtain probabilities for train and test data.
[01:26:24] Ashish Kumar: the predicted probabilities, because when we plot a UC score, we need probabilities.
[01:26:29] Ashish Kumar: So, we obtain all of this, and then there is the… there is, like, from SQL.metric, there are functions, already there for calculating accuracy, for calculating precision, for calculating recall, F1.
[01:26:42] Ashish Kumar: ROC, so we pass all these values of the predictions we obtain, and also the, like, the ground truth. So, ground truth for train, if we want to obtain metrics for train data, so ground truth is by the name Y underscore train.
[01:26:57] Ashish Kumar: Every time.
[01:26:58] Ashish Kumar: And, predicted value is by the name white underscore trend pred, for predictions, and for probability, it is by name white underscore trend underscore prop.
[01:27:09] Ashish Kumar: And similarly for test metrics, we will, like, obtain
[01:27:14] Ashish Kumar: like, ground truth is by the name Y underscore test, and the prediction is by this name, Y underscore test underscore plate. So, this we obtained, and we can, like, we discussed, we can obtain a classification report also at the end. So, yeah, so after we obtained accuracy of 6, this is the final result.
[01:27:33] Ashish Kumar: we obtained, on depth accuracy of 5, and where, split ratio is 60-40. So, this is also, like, one way we can do.
[01:27:43] Ashish Kumar: And I repeated, the same code is repeated for, 7030.
[01:27:50] Ashish Kumar: 80-20, and then, 90-10 sprint test split. So, yeah, so this will finish our, total decision tree, and then we can, like, move to random forest. So, any doubts on this part?
[01:28:17] Ashish Kumar: Okay, so… I see, there are no doubts. So, shall I proceed to the next section?
[01:28:29] Deepan Kanagaraj: Yep, yep.
[01:28:30] Ashish Kumar: Okay.
[01:28:31] Ashish Kumar: Okay, so, yeah.
[01:28:33] Ashish Kumar: Just before proceeding, just to… just, we can, like, look at the value of accuracy, how it is coming at various interspit. So, for 6040, it was coming 5.
[01:28:46] Ashish Kumar: And, this is, like, 70-30. It is also… in 70-30 also, it's coming 5 only.
[01:28:53] Ashish Kumar: And,
[01:28:55] Ashish Kumar: Like, there is very much less difference between 5 and 6, difference in, in decimal places only.
[01:29:02] Ashish Kumar: And,
[01:29:03] Ashish Kumar: For 80-20, it is coming, against 6, but 5 and 6 are, like, same only. So…
[01:29:12] Ashish Kumar: Yeah, yeah, no, it's coming 7… it's coming to be 7 here, 7875. And, again, 9010, this is coming to be 6. As we already saw yesterday, this is coming to be 6.
[01:29:26] Ashish Kumar: So, yeah. So, let's… after this, let's now move to Random Forest.
[01:29:32] Ashish Kumar: So, as we saw in this entry classifier, like, the classification code is minimal, but in order to do classification, we have to first, like, obtain these optimal values of the hyperparameters.
[01:29:46] Ashish Kumar: So, again, for random forest, we will follow somewhat same procedure, to obtain the optimal values, and then once we obtain the optimal values, then we can train a final model on our dataset.
[01:30:01] Ashish Kumar: So, in…
[01:30:03] Ashish Kumar: Random forest, two important hyperparameters there. One is called… in SKLN, it is called by the name N underscore estimators. Basically, it is the number of trees you will consider in your forest.
[01:30:14] Ashish Kumar: And,
[01:30:16] Ashish Kumar: Second is the max features. It is the number of features which has to be considered at each split.
[01:30:22] Ashish Kumar: So, you don't consider all the 29 features at each split, only a subset of the features. So, this… these are the two major hyperparameters which we… which will optimize for random forest.
[01:30:35] Ashish Kumar: So, like, how to get the number of trees? So, we manually, like, are select… searching in the range, this. So, I will make a list, like, with the range, where… with the values, like.
[01:30:53] Ashish Kumar: 50, 7500, 150. Of course, we can add over it, we can make many and more entries, 200 to 50. Like, it just increases the computation, the code will take some more time to run.
[01:31:05] Ashish Kumar: And, nothing, like, we can increase this list also, like, how many trees to add. But it, just to give a, like, small illustration, we are using, like, these, four values of the number of trees.
[01:31:18] Ashish Kumar: So, we begin.
[01:31:21] Ashish Kumar: So, this time, the difference is, we are importing, so decision tree is, like, is imported, like, from skln.trees, but randomforest is, like, in skln.Ensemble, and we import randomforest classifier from there.
[01:31:37] Ashish Kumar: And this thing is same. In skln.metrics, we have the metric function written for all the standard classification metrics, accuracy, precision, recall, F1, and AUC score, which is called here ROCUC score.
[01:31:52] Ashish Kumar: Again, trend, test bit, the same thing. For plotting, we are using matplotlib, and importing it as plot, matplotlib.pyplot, and numpy, as and when required.
[01:32:03] Ashish Kumar: So, yeah. So,
[01:32:06] Ashish Kumar: So, here, like we have seen, like, in our illustration again and again, so, we could have experimented with the, again, with the trend split size also.
[01:32:17] Ashish Kumar: But then it would have, like, again been redundant, seeing it again and again, so I fixed it as 80-20. I'm using 80-20 as my train test speed ratio for doing these experiments on random forest.
[01:32:30] Ashish Kumar: So, based on that.
[01:32:32] Ashish Kumar: I obtain my train and test data by, like, I'm giving the original data here again, and train size is 80%.
[01:32:40] Ashish Kumar: And based on that, I will get my train and test data at this line.
[01:32:46] Ashish Kumar: Then, again, as I told, there is a… this is the, like, the, values of the estimator, the number of trees I will be, searching for, that, among them, which is the, like, the… which is, like, a good value for the number of trees.
[01:33:02] Ashish Kumar: So, like, these are,
[01:33:04] Ashish Kumar: Like, these could be, like, we could have them in a, like, a continuous range of values also. These are, like, kind of arbitrary only, because I wanted to only show only 4 of them, because it takes time, because random forest
[01:33:19] Ashish Kumar: makes, for, like, for this one, it will make 50 trees, it will make 75 trees, so it becomes very… computation takes, longer to run. So that's why we are experimenting only on, like, 4 of them, 4 values.
[01:33:32] Ashish Kumar: And then, to store these metrics, again, there is a dictionary, and inside dictionary, there's, like, a list for train accuracy, test accuracy, similarly, train precision, test precision, train recall, test recall, and everything, which will be stored here.
[01:33:49] Ashish Kumar: So, so, again, the code is, like, we will loop through this… four values in the list.
[01:33:57] Ashish Kumar: And this is my random forest classifier. The, the class, the class, for the…
[01:34:05] Ashish Kumar: The constructor for it to call this class.
[01:34:08] Ashish Kumar: And here, we pass on the number of estimators, as in this list, so first time it will be, like, 50, then 75, so on. And here also, because ultimately, inside it's a tree, so we have the option to, like.
[01:34:23] Ashish Kumar: specify the, like, criteria, either it could be in Tropi or Guinea, and then,
[01:34:29] Ashish Kumar: Then again, because it's still… we can… there's also, like, criteria to specify the max depth. So, we have, like, arbitrarily, like, we have fixed, because we have so interestingly, that max depth was 6 was coming, okay?
[01:34:43] Ashish Kumar: 605, so we have fixed a 6 here. And again, a random state, because, it will, because, yeah, so random, for instance, the name is also random, so it picks up random subset of the features, and there are many, like, these all random steps are there, so we have… so in order to, like, reproduce the result.
[01:35:01] Ashish Kumar: to fix the randomness that only… at each time, select from the same set of… same split, we are using this random state. So, if I use a different number, the results might occur different when you run them, and just the defense is only that.
[01:35:19] Ashish Kumar: It is like to fix the randomness, that every time we run the code, the same results are reproduced.
[01:35:27] Ashish Kumar: So, yeah, so… At this stage, I'm getting a, like, a…
[01:35:33] Ashish Kumar: like, I am making a class object of random forest classifier, where setting various hyperparameters. So, you see, there are many hyperparameters available, but I am, like, finding optimal value for only one of them. Otherwise, it becomes very hard to manage for everyone.
[01:35:51] Ashish Kumar: So, once I make a class object, then again, as decision tree. So, basically, the training code is minimal. The effort lies in to find the optimal value. So, again, for training, there's just this line, clf.fit.
[01:36:06] Ashish Kumar: You give it your training data, the X's and Y's.
[01:36:09] Ashish Kumar: And at this point, it will, like, finished it, training, and obtained, like, finished all its… taken all its decision that which… where to split, which will be the first, at, node to consider
[01:36:24] Ashish Kumar: First attribute to consider for the root node and all. So, after that.
[01:36:28] Ashish Kumar: Once training is finished, like, as before.
[01:36:31] Ashish Kumar: We will do a predict over train. We will do a predict over test.
[01:36:37] Ashish Kumar: And this gives me hard-coded values, as we saw in the list also. You get a 1 or 0, a list of… a binary list of either 1 or 0 for the
[01:36:46] Ashish Kumar: Predictions on train and test data.
[01:36:50] Ashish Kumar: And, because AUC score, as I mentioned, it requires probabilities, to, for its metric. So, in order to calculate probability, we have this predict underscore prob A.
[01:37:02] Ashish Kumar: It will give me, instead of hard-coded predictions, it will, like, hard, hard binary values, it will give me a
[01:37:10] Ashish Kumar: probabilities in the range 0 to 1, soft values of my predictions. So, for training, it will be stored in this variable, and for test, it will be stored in this variable.
[01:37:23] Ashish Kumar: And here, then, once I get it, I will, like, I have initialized this huge, big dictionary in which there are several keys. So, first, for train accuracy.
[01:37:35] Ashish Kumar: I will, for this particular range of estimator, which is, like, starting at this is 50, when number of trees is 50, I will, I will first calculate the accuracy on the trained… on the train dataset. So, this is accuracy on train dataset.
[01:37:52] Ashish Kumar: Similarly, I'm calculating accuracy here on the test dataset, and storing it in the list corresponding to test accuracy. And in the previous line, it was, like, stored in the list corresponding to train accuracy.
[01:38:06] Ashish Kumar: Similarly, here, inside the code itself, I am storing the precision also, for the trained dataset.
[01:38:14] Ashish Kumar: The precision value for the print dataset, and the, test dataset at, at, for, for number of trees as 50. For the number of trees, 50, we're storing all of this.
[01:38:24] Ashish Kumar: Similarly, recall is stored.
[01:38:27] Ashish Kumar: And then we store the F1, and at last, we store the AUC values.
[01:38:33] Ashish Kumar: So, this particular code here, it is just, like, we are looping through all these four values and storing these metrics, accuracy, precision, recall, and everything.
[01:38:44] Ashish Kumar: And, then, what I'm doing is.
[01:38:48] Ashish Kumar: I am plotting all of them, accuracy, precision recall, F1 AUC, so these are, like, five,
[01:38:54] Ashish Kumar: These are, like, 5 metrics. So, I will plot a 2 across 3 subgraph, like, the 2 rows and 3 columns, and the last entry will be empty, because, like, these are only 5, so it can have up to 6 subplots. So, I will use that.
[01:39:12] Ashish Kumar: And, just after that, I'm just… I've made this big list here, metrics underscore RF. So, basically, what I'm doing is, I will just
[01:39:23] Ashish Kumar: Yeah, so I'm just… I will… I will run through it, I will loop through it.
[01:39:29] Ashish Kumar: And, based on, whatever is the, like, the value of the metric, I'm doing this plot. The main plotting is happening here. For the train value, for the train metric, and for the test metrics.
[01:39:40] Ashish Kumar: The plotting is happening here. The X's is the number of… X is the number of trees, which is, like, 50, 7,500, yeah, it's coming also. The list is 50, 7500, 150.
[01:39:53] Ashish Kumar: And the y-axis is the, metric value. So, it can be, like, test accuracy, train, test precision, test, AUC, everything. So, yeah.
[01:40:05] Ashish Kumar: So, these are plotted here, and then for formatting purpose, these are again the same, that you set the title, you set the labels for X, you set the label on the Y axis, you set the label on X axis, whether you want a background, the greatest to be required.
[01:40:19] Ashish Kumar: And… and everything. So…
[01:40:23] Ashish Kumar: After that, yeah, so, all, everything is, like, this is to, like, this code is there to just to hide the,
[01:40:31] Ashish Kumar: like, the hide the sixth cell, which is, like, because that will be empty, so we are telling it to hide that sixth cell.
[01:40:39] Ashish Kumar: Just a minute.
[01:40:41] Ashish Kumar: Yeah, so,
[01:40:43] Ashish Kumar: Yeah, so once we, now, I have created a function for this, we have created a function for this, and we'll pass or obtain this matrixRF to it.
[01:40:54] Ashish Kumar: And it will, then plot this kind of graph.
[01:40:58] Ashish Kumar: So, we will see, we are seeing… what we are seeing is, performance of a random forest,
[01:41:05] Ashish Kumar: Across the different number of estimators, 50, 75, 100, and…
[01:41:12] Ashish Kumar: I think, 50, 70, 100, and 150. We are on the x-axis, these are there. And we are saying, train and test accuracy, train and test precision, train and test recall, train and test F1, and train and test AUC. So.
[01:41:26] Ashish Kumar: based on that, we usually take a, we take a decision that, which of, is, like, performing well. So, here, if we, like, see.
[01:41:40] Ashish Kumar: So basically, accuracy, accuracy, it is like, we… if you see accuracy, so it is coming highest on 75.
[01:41:49] Ashish Kumar: And, the second highest is on 150, and, no, no, so second highest on 50, then 75, and third highest on, 150. So, for 100 estimator, it is the lowest, the accuracy among all these four.
[01:42:03] Ashish Kumar: Similarly, like, the precision,
[01:42:07] Ashish Kumar: Is varying. And, similarly, we can… so basically, the same problem arised as yesterday, that there is… there will be no single.
[01:42:16] Ashish Kumar: like, value of estimator, which is, like, will be good for all of them, so we will have to, like.
[01:42:21] Ashish Kumar: A sort of, like.
[01:42:23] Ashish Kumar: like, compromise, or, like, see, like, which works… which works as a balance for all of them. So, in terms of recall, I'm seeing that, the test performance is, like.
[01:42:37] Ashish Kumar: test and train both performances best on 150. And similarly, F1 also, the train and test is based on 150.
[01:42:46] Ashish Kumar: And,
[01:42:48] Ashish Kumar: AOC, the best is, like, on 100, but on 150 also, like, there is not too much dip.
[01:42:55] Ashish Kumar: So, here, like,
[01:42:59] Ashish Kumar: I selected, like, 150 as the number of estimators, like, because, like, 50 estimators, like, 75 estimators seems to be too low here.
[01:43:12] Ashish Kumar: And, like, if you see after 100, the performance again increases. Even in accuracy, if you are seeing that, the performance… if you are seeing that, if it might look that at 150 the performance is decreasing, it's because we have, like, maybe not, used lesser number of trees.
[01:43:30] Ashish Kumar: But, you see, after 100, the performance again starts increasing.
[01:43:33] Ashish Kumar: So, that's why I like to keep a fair number of trees, like, 50-75 becomes far less for the, like, random forest, and I think the default is also, like, 100 or so, the default value of estimators, if you don't give anything. So, I selected 150 for them.
[01:43:49] Ashish Kumar: So, but accordingly, like, how we interpret this graph, and, we can, like, check for other values also, like, what is best value of estimators coming.
[01:44:01] Ashish Kumar: So, this is how we will, like, give, like, an estimate of 150 is best. So, basically, if my… so, basically, what we will see is, if my criteria is, like, I want to have a good recall, so I'm saying that best recall is coming at 150, so, we will pick up.
[01:44:20] Ashish Kumar: and estimate is at 150. So, this is also a criteria, like, for your classification system, either you want to have a high precision or a high recall. So, if I want a system to have a high recall, so I will select as 150.
[01:44:32] Ashish Kumar: So, it, this, decision, also, like, these, requirements of our system also will also influence our decision.
[01:44:41] Ashish Kumar: So, with this, we, like, finish the first hyperparameter.
[01:44:45] Ashish Kumar: And, then we'll, we'll begin with the second hyperparameter. Any, like, questions on this one?
[01:45:00] Ashish Kumar: Okay, so I think let's just finish it, and then we can discuss. This will finish our total end of our section.
[01:45:09] Ashish Kumar: So, again,
[01:45:11] Ashish Kumar: Now, now the second paragraph… so basically, now I've obtained, and a number of estimators I have obtained, so I will use that only in render first classifier, and I will vary the maximum number of features.
[01:45:27] Ashish Kumar: So, so basically, so yeah, so we… I will vary the number of features. So, earlier, in the earlier code.
[01:45:35] Ashish Kumar: This was, like, not shown explicitly inside the loop. So, it was taking the default value.
[01:45:42] Ashish Kumar: of it, whatever it is in the, like, the sklearn libraries, it has a default, way to handle max features. It was taking that. So this also, like, changes the thing. If we first start with maximum features, and then we obtain a maximum feature value, and then we will go to the number of estimators.
[01:46:01] Ashish Kumar: So, the sequence also matters. So, the problem is that because of the computation, we are not able to do
[01:46:08] Ashish Kumar: all of them simultaneously, otherwise we could have… make, like, a 2D grid of the values of the number of max features to consider and number of estimators to consider.
[01:46:17] Ashish Kumar: So, we could have… make a 2D grid type of, and search for each combination. But since it takes… right now, the competition takes too much time in collapse, so that's why we are doing one pyramid estimator at a time.
[01:46:31] Ashish Kumar: And even if you, like, see on this dataset, like, this data set is slightly challenging.
[01:46:38] Ashish Kumar: even if you try a various combination of them, the performance, like, there is… there would be, like, very less chance that performance increases. So, it's very hard to obtain, like, a good accuracy on it, like, beyond 80% or something. So, this is also, like, a challenge with this dataset.
[01:46:58] Ashish Kumar: So, here, first step. So, basically, total number of features, so basically, if I do…
[01:47:05] Ashish Kumar: So, X.shape is like this.
[01:47:09] Ashish Kumar: So, this is 0th position, this is 1ATH position.
[01:47:14] Ashish Kumar: Okay, so,
[01:47:16] Ashish Kumar: So, yeah, so X, so basically, it's slightly, wrong, it should be, like, I have to… it should be encoded.
[01:47:25] Ashish Kumar: So because, I have to take from encoded one, yeah. So, it should be, like this.
[01:47:34] Ashish Kumar: Okay, okay.
[01:47:37] Ashish Kumar: Yeah, so it should be from here that we have, like, taken our parameters.
[01:47:43] Ashish Kumar: So, yeah. So, then, but the flow stays more or less the same. So, number of estimators, I fixed the 150 obtained from previously.
[01:47:53] Ashish Kumar: Depth of each tree, I'm fixing a 6. So, another search of, could be the… what will be the optimal depth of all these trees in the, random forest.
[01:48:04] Ashish Kumar: This could be, like, another search. So, but because depth was previously covered in vision tree in this, so we have not, like, separately done for this, but if this module would have been just for random forest, so one search would be, like, one part of the search would be dedicated to find the max depth. Right now, we are not doing that.
[01:48:23] Ashish Kumar: So, again, now next we proceed to obtain the train and test data. So, here is my, like.
[01:48:30] Ashish Kumar: X is in, like, this is my original, data, X and Y, the X encoded in Y, and when I pass through this function, I obtain train and test, and, train and test, dataset for both X and, Y.
[01:48:45] Ashish Kumar: And the number of features… so, yeah, yeah, so basically, I'm not even sure, like, maybe I'm not even using this, so… because the range here, I am fixing like this. So, my features are… there are total 29 features, I am going through like this.
[01:49:02] Ashish Kumar: 5, 10, 15, 20, 25, and then the maximum 29 features.
[01:49:08] Ashish Kumar: So, is my, screen, visible to everyone? The tab, second tab?
[01:49:18] Swagat kumar pattnaik: Yes.
[01:49:20] Ashish Kumar: Okay, so what I wanted to show is that if you go through a random forest classifier in SKLearn, if you start over this, this SKLearn link will come.
[01:49:32] Ashish Kumar: And, you can, like, see what is the default value, like, there are so many parameters to consider. Like, we are… we are hyper-optimizing for only two.
[01:49:42] Ashish Kumar: So, you see, there are so many parameters to consider. The default value of a number increases 100, the default criteria is GINI, if you don't specify. If you don't specify a depth value, it considers none. So, it will go to the max depth. So, basically, it is saying, if max depth is none.
[01:50:00] Ashish Kumar: then nodes are expanded until all leaves are pure, or until all leaves are less than number of… minimum number of samples to consider for a split. So, yeah, so there are various set of parameters which I've considered. So, basically.
[01:50:15] Ashish Kumar: By default, value of max feature is square root. So, whatever is your, suppose you have n features, it will consider root n number of features as max features by default.
[01:50:26] Ashish Kumar: So, like, we can have a look, for all, similarly for decision tree, the default values of each of them, so for… for our better understanding, like, like, what are these, all these hyperparameters out there.
[01:50:39] Ashish Kumar: And what are their, like, default values? So, yeah, so…
[01:50:44] Ashish Kumar: Let's come… I'm coming again to code. So, this is my feature range I will loop through, and again, this is, like, the same dictionary structure, like, as before. I'm having the train, list… I'm storing the list for both train data, train metrics, train accuracy, train precision, train recall.
[01:51:03] Ashish Kumar: And similarly, test accuracy, vision, recall.
[01:51:06] Ashish Kumar: And like that.
[01:51:07] Ashish Kumar: So, then… I loop through my, this list.
[01:51:13] Ashish Kumar: Max feature range, this time. So, number of estimators, I have hardcore fixed as 150, so here it is, like, 150, it is fixed.
[01:51:22] Ashish Kumar: And, then I will, max depth is also fixed, and for max features, I will loop.
[01:51:28] Ashish Kumar: And criteria is, again, entropy and random statistics 52. And then, as we all do every time, we do a fit, and this training finishes over there. And then for predictions, we are calling predict function. It will give us the predictions, and for obtaining prediction probabilities, we'll call predict underscore probate.
[01:51:47] Ashish Kumar: And then, like as before, we are storing all of these in a, like a…
[01:51:52] Ashish Kumar: in… inside this… inside this dictionary, which contains lists for storing train accuracy, test accuracy and everything. So, here we are, like, storing accuracy. This is white underscore train is, again, the ground truth. This is my prediction, predicted values, which I obtain at this step, CLF.predict.
[01:52:12] Ashish Kumar: Similarly, for test 1, I am storing here Y underscore test ground truth, and this is the printed value obtained, like, at this step.
[01:52:22] Ashish Kumar: And, likewise, since we have both of them, and even these, so we can, like, calculate precision. Precision score is the function to, like, call and pass on these parameters, and recall score is to pass on here. Pass on these parameters, and F1 score.
[01:52:42] Ashish Kumar: And finally, an AUC score, we pass on the probabilities. Y underscore 10 is the ground truth.
[01:52:49] Ashish Kumar: And this is the predicted probabilities. So, these are stored in dictionary.
[01:52:54] Ashish Kumar: And again, like, as previously, like, this graph is, like, plotted for each of them, like, how is the number of features.
[01:53:02] Ashish Kumar: So, here again, like,
[01:53:06] Ashish Kumar: There's no single value of max features which will be, like, which is performing good for all metrics, but with a focus on the recall, max… maximizing the recall, I can, like, look that 415 is the max features.
[01:53:22] Ashish Kumar: Recall is, like, highest, highest, and F1 is also highest, here, on that, on the test set.
[01:53:31] Ashish Kumar: And, also, like, if, like, if we, so here, what if we are seeing is.
[01:53:40] Ashish Kumar: That somewhat it feels like that,
[01:53:42] Ashish Kumar: test accuracy is coming highest on, like, taking all the features. But it is like, if you take all the features, then it tends to, like, it tends to…
[01:53:56] Ashish Kumar: defeat the purpose of random forest itself. It is coming on this particular data set we are training. It is showing that, like, accuracy is highest on, like.
[01:54:06] Ashish Kumar: like, all… if you consider all 29, it is showing highest, what… what trade and test. But yeah. But, we use… because random words usually take a subset of the features, so, and by default, also, it takes root n. So, we use, 15 also. 15 as our number of features.
[01:54:26] Ashish Kumar: So once we, so…
[01:54:30] Ashish Kumar: So, this is, like, I've also mentioned, like, observation that, test equities highest as this.
[01:54:36] Ashish Kumar: So, it is indicating that model performs best when all features are considered.
[01:54:41] Ashish Kumar: But, the trade-off is that,
[01:54:45] Ashish Kumar: recall, like, if we see… recall, it is very low for… if we consider all the features.
[01:54:53] Ashish Kumar: So, so accuracy itself is not a, like, a, like, a good criteria, good criteria, if you want to, like, if a system either want to have a good precision or recall, so that's why, like, we switched to, like.
[01:55:08] Ashish Kumar: 15 here.
[01:55:10] Ashish Kumar: So, yeah. So finally, with the values of estimators, number of trees and the number of max features, we'll train one final model.
[01:55:23] Ashish Kumar: So, again, this is my train test bit.
[01:55:26] Ashish Kumar: obtained, and I have, like, in Random Forest, I am uniformly using 80% of the data is trained, all throughout.
[01:55:33] Ashish Kumar: So, then my number estimate is 50, my max feature is 15, and I've used criteria entropy and read this. Again, I train, and I obtain the predictions.
[01:55:47] Ashish Kumar: And, after obtaining predictions, I will… once I have the predictions, I can call metric function to calculate the train accuracy, train, precision, train recall, and similarly, test accuracy, test, precision, test recall. And, we get something like, we get these values.
[01:56:05] Ashish Kumar: So, if you see, if you will compare, there's a final comparison table at the end. I think there is, like, less than one, or somewhat like that improvement is there. In decision tree, it was coming 78, here it is coming 79.88. Like, not much improvement.
[01:56:23] Ashish Kumar: So, this is what we see here. Like, only a 1% improvement, again, is there. So, but yeah, so if we were to, like.
[01:56:33] Ashish Kumar: more optimized over these values, so maybe, like, we can squeeze more performance out of it, because right now we are, like, fixing that max step to be 6, we are not optimizing over this.
[01:56:44] Ashish Kumar: So… At last, the last, so basically.
[01:56:49] Ashish Kumar: In table content, there's, like, the labels written A, B,
[01:56:55] Ashish Kumar: C and D. These are compared at the end. These are, like, 4 final configurations, 2A and B are of configurations of Digentry, and C and D are configuration of random forest. So, just the last one. The same code is repeated.
[01:57:10] Ashish Kumar: just, I am using the criteria Ginny to see, like, if there is, like, any improvement can happen. So, if you see.
[01:57:18] Ashish Kumar: Yeah, some improvement is there, although not much. If you see here, it was 79.88.
[01:57:25] Ashish Kumar: And here, it is, like, now 80. It has crossed 80. So, like, so therefore, we see that,
[01:57:32] Ashish Kumar: the criteria also, in decision tree, changing the criteria was not changing the performance. It was more or less… it was similar. It was, like, 78 in both the cases, but here, slightly, somewhat increase.
[01:57:47] Ashish Kumar: Like, we don't know yet whether it is a statistically significant increase or not, but yeah. But changing the criteria also, like, can help sometimes to, like, get more performance out of it.
[01:58:01] Ashish Kumar: And, yeah, so after that, like, so this is my, like.
[01:58:06] Ashish Kumar: final, comparison on, like, the test set. So, so for the markings I've marked over the table of content, A, B, and C, D, I have, like, written here separately. So, you see, that, in decision tree.
[01:58:23] Ashish Kumar: When we are using criteria either entropy or GINI, the accuracy tend to stay the same.
[01:58:29] Ashish Kumar: No matter what. And, the other important criteria could be, like, F1 score, we can see that in Gini, the F1 score was decreasing.
[01:58:42] Ashish Kumar: So, basically, in addition 3, Ginny was not, giving, for one single decision tree, Ginny criteria was not giving, like, a good performance. But for random forest.
[01:58:54] Ashish Kumar: We can see that Gini, so this F1 score is also not decreasing. It is, like, more or less same like this, the F1 score of random forest of entropy, it's same to that, but accuracy is, but, what improved, even the, like.
[01:59:11] Ashish Kumar: like, there is… improvement is still in terms of decimal only. If you see, it's 79.88, and now it is 80.09. So, like, if you, like, say it in that way, that…
[01:59:23] Ashish Kumar: there is not that much increase, but yeah. But what you wanted to focus is that this criteria also can play a role sometime to, like, obtain, like, a good performance.
[01:59:37] Ashish Kumar: So, with that, we, like, finish, this, notebook.
[01:59:43] Ashish Kumar: And, we can, like, there are, like, still 4 minutes left. Any questions are there?
[01:59:58] Ashish Kumar: Okay, so… Okay, just a little,
[02:00:04] Ashish Kumar: For the, like, for the sake of it, just, one last, if, are there any questions, or we should, like, we can stop then?
[02:00:18] Ashish Kumar: So, shall we, like, wrap up for today?
[02:00:31] Sonam Manwal: Ashish, just one question. Yeah. Yesterday, we did a rep, plotting for… we did plot 3 also, right?
[02:00:40] Sonam Manwal: So, is it possible to plot the random forest to see, or we have to plot trees only, one by one?
[02:00:47] Ashish Kumar: I think, because the number of trees are too much, so, they have not… there's not a library function, so this is the problem. Because I… just yesterday, I showed you the, like, the tree which I showed you. It was from the library function.
[02:01:02] Ashish Kumar: The visualization was from this, there's a… in SKLearn, there is a function called plot tree.
[02:01:08] Ashish Kumar: Which, works for this. Now, because it's a single tree, even… this looks cluttered, because at the end of it, it becomes very cluttered.
[02:01:17] Ashish Kumar: So, now, because there are, like, 50, like, minimum 50 trees, and then they are increasing, so it becomes hard to, like, visualize each of them. So, I don't think there is any, as such, liability function, although, like, we can search if there is any.
[02:01:34] SHUBHAM GOSWAMI: Huh?
[02:01:36] Ashish Kumar: I'm saying that there is… as such, as I'm aware, there is no library function to, like, plot this random forest.
[02:01:42] Sonam Manwal: Yeah, sure, it was someone else, I think.
[02:01:45] Ashish Kumar: Okay, okay.
[02:01:47] SHUBHAM GOSWAMI: She got nigga.
[02:01:50] SHUBHAM GOSWAMI: As you build them.
[02:01:53] Ashish Kumar: Okay, so, any more… any questions? Anyone else?
[02:01:57] Jithu Tagore: Yeah, am I audible?
[02:01:59] Ashish Kumar: Yeah.
[02:02:01] Jithu Tagore: I would like to know for, like, for random forest, if the estimator value is very high, like, more than 200 or 300-something, will the model accuracy and positional increase?
[02:02:15] Ashish Kumar: Like, if we can… Just, here.
[02:02:20] Ashish Kumar: Like, what do you… what should I put? Should I put 300 or 200?
[02:02:24] Jithu Tagore: Yeah, 300.
[02:02:26] Ashish Kumar: Okay.
[02:02:29] Ashish Kumar: The only pronun might take a…
[02:02:32] Ashish Kumar: Bit of time? Yeah, you're saying that it… here it stays the same, right? You're doing 300? So…
[02:02:39] Ashish Kumar: Not… not helping here, please.
[02:02:43] Ashish Kumar: So, basically, it depends. Increasing the tree also is, like, more number of trees is also, like, it will not do much good. It's, like, an optimum value only till which it gains.
[02:02:55] Jithu Tagore: Okay, thank you. Sorry, why I asked? Because, when I saw the graphs, I felt like when I… when the…
[02:03:04] Jithu Tagore: value we're increasing.
[02:03:05] Ashish Kumar: So maybe you're more willing to.
[02:03:08] Jithu Tagore: Yeah, I thought like that. That's why I cast that question.
[02:03:13] Ashish Kumar: Okay, so maybe… 200.
[02:03:17] Ashish Kumar: But yeah, yeah, overall, the values are not. So basically, when… the thing is, these graphs are, like, also, I would say, not that accurate in the sense, because we are fixing one thing and then looking at another.
[02:03:32] Ashish Kumar: And, afterwards, like, so basically, if you see this max number of features, we have fixed that number of estimators will be 150, and then we are looking at it.
[02:03:43] Ashish Kumar: But in the final model, whatever, we are, like, separately using them. So… and if you see in the first plot, this one, where we were, we were varying the number of estimators, so max features is fixed as root 10, which is the default value of the random forest.
[02:04:00] Ashish Kumar: So, so these, I, were for, these are for illustration that this is how we do, but, because, the, like, the… the ideal way would have been that, would have been, like, as I, like, told yesterday, that suppose you have your
[02:04:18] Ashish Kumar: estimators is 100. Then, for this 100, you iterate through all the features. Like, if my features are 15, 20, 25, you iterate for this
[02:04:29] Ashish Kumar: And then, similarly, for number of estimators as 200, you iterate for your all, these features again. So that would have been given the highest, like, reliance that, like, what you are pointing out, that it looks like that it might have increased.
[02:04:47] Ashish Kumar: So… but, we did not do that, just because… just because of the limitation of computation here. It would… to run the code, it would have taken more time.
[02:04:57] Ashish Kumar: Otherwise, the best way is to do, like, a 2D thing, a 2D loop, like, two loops through this.
[02:05:05] Jithu Tagore: Okay, we need to check all these parameters for analyzing which model parameters are best for best, right?
[02:05:13] Ashish Kumar: Yeah, yeah, yeah. That would be the most ideal approach, but the thing is, you will see that, like, here, the comp… so, because the dataset is very smaller.
[02:05:26] Ashish Kumar: So, that thing comes in play. There are only 7,000 samples, so maybe if you will do two for loops, you will have, that will give result within, like, maybe even 4 to 5 minutes, you will get the result. But suppose we have million of samples, even tabular data has million of samples or more sometimes.
[02:05:42] Ashish Kumar: So then we can't afford to do 2D looping. So then we will have to, like, do… we can estimate by taking one parameter at a time.
[02:05:53] Ashish Kumar: So, this is the thing.
[02:05:56] Jithu Tagore: Yeah, I got it. Like you have done now.
[02:05:59] Ashish Kumar: Yeah, yeah, yeah.
[02:06:05] Ashish Kumar: Okay, so, any more questions?
[02:06:12] Ashish Kumar: Okay, so I think, we can wrap up for today. It's, past 11th.
[02:06:19] Swagat kumar pattnaik: Sure.
[02:06:20] Ashish Kumar: Okay. Thank you.
[02:06:21] Ashish Kumar: Okay, so, thank you, everyone.
[02:06:24] Ashish Kumar: Nice talking to all of you.
[02:06:27] Swagat kumar pattnaik: Thank goodness.
[02:06:29] aditya shrivastava: Thank you.
[02:06:30] Ashish Kumar: Bye-bye.