# 10 2025-12-28 Hands On Logistic Regression course: Module 2 — Machine Learning Algorithms module: Module-2-Machine-Learning-Algorithms date: 2025-12-28 type: transcript video_url: https://personal-learn.armco.dev/files/_Recordings/Module-2-Machine-Learning-Algorithms/10_2025-12-28_Hands_On_Logistic_Regression.mp4 --- [01:01:05] Uh, what we can… what we are doing is, uh, to this object, we pass on our total full X and Y. [01:01:12] And over that X and Y, it will, uh, when we do kf.split, it will start giving me folds. [01:01:18] So, this train index is basically… suppose my fold is 3, so basically there will be… my fold value is 3, so there will be 2 folds for training, and one fold will be for testing. [01:01:30] So, for these two folds, [01:01:33] What particular… what is returning is that index… [01:01:35] Corresponding the original data, which particular index [01:01:39] come to the two-folds which are in training, and which particular index [01:01:42] come in the, uh, that test data for the one-fold. Uh, so this is giving me index. [01:01:51] So once I have indexed, so I go in my original data, [01:01:53] And I locate, and I obtain the subset of the data, [01:01:57] Uh, based on the indexes, which correspond to my, uh, uh, like, the folds which are there in training. So, here, uh, and similarly, for the test, for the… I have the [01:02:08] indexes which will be part of the test fold. So, from my original data. [01:02:13] I obtained that subset, and that's why it is like, uh… and finally, [01:02:18] This will be, like, if I'm doing… [01:02:20] like, three-fold, so two-fold data will be there in extra in row, and last one-fold will be therein. [01:02:26] X test row. And similarly, Y will also be divided. So, Y is also divided in the same manner, that [01:02:33] The train indexes are passed on to obtain a subset of the train fold and test index passed down to obtained subset of the test pool. [01:02:43] So, once you get this, [01:02:46] So, this… now this full, uh… [01:02:49] Now, this full data is my train, and uh… and this is my test one. [01:02:54] So, yeah. So now, uh… [01:02:56] Inside of it, because we have to deal with [01:02:59] Uh, because if you remember, there are, like, categorical features which are in form of strings here in this data. So, what we are doing is, [01:03:08] integer encoding, uh, so in integer encoding in SQL learn is by name ordinary encoder. [01:03:13] And this is, like, two parameters are there to tell us what to do. [01:03:19] If you encounter an unknown category during testing, [01:03:25] Uh, because you have… you will fit the data on your… [01:03:28] training parts, so you will fit the data on this extend row. So, whatever categories [01:03:34] are coming here, uh… [01:03:36] It will encode them, but suppose in test data there is a new category for that particular column. What to do then? Either, uh… [01:03:44] If you don't use this, it will throw an error, so we are saying it not to throw an error, just, uh… [01:03:49] silently insert a minus 1 over there. [01:03:51] So, that's why it's written like this. [01:03:53] And similarly, we will want to normalize our numeric features so that all of them are in the same scale, and all of them have zero [01:04:01] mean and unit variants of 1. [01:04:03] So, we are calling standard scalar also. [01:04:06] Now, this is just, uh… like, I'm copying this because [01:04:10] the usual way we understand the convention we understand is by these names. [01:04:14] The train split and test fit, we understand by these names, so that's why I've written a copy. Otherwise, uh, this… [01:04:20] thing is the same. So, yeah. So, once we have extended next test, [01:04:24] So, uh, what I'm doing is, again, you can see that this is encoder, auto-encoder. I'm doing a fit and transform. [01:04:33] over my categorical columns, uh, for the training set. [01:04:36] And, uh, I'm doing only transform. [01:04:40] over my test set. So… [01:04:42] So, to learn, like, which category map to which integer, [01:04:46] We are using a train set, and once we have, uh, and then doing its transformation also to convert it into integers. [01:04:53] But for test set, uh, we are relying on what mapping we learned, and that particular mapping is used to transform the [01:05:01] data in the test set. [01:05:03] So, you can see that what I've used is, like, [01:05:06] Categorical, uh, like, this exchange the entire dataset. [01:05:10] Only for the categorial columns, I am doing it integer encoding here. [01:05:14] This is for the test part, and this is for the… so, this is for the train part, and this is for the test part. [01:05:21] Now, this finish… this finishes the categorical part, but we have our numerical columns also, and which we want to normalize. So, again, uh, for our numerical columns, this corresponds to our numerical, uh, data of the numerical columns. [01:05:34] I'm, uh… I'm learning mean and variance for each column, and based on that, transforming the original features. [01:05:42] So, for train part, I'm, again, doing fit underscore transform, fit and transform, both are performed in one step. [01:05:47] For the test data, only transformation is done. So, basically, the mean, the variants are learned from the trained data, and based on that, that… [01:05:56] Test data is just transformed. So, yeah. So, after that, uh, you done, uh, integer encoding and, uh, tra… uh, like, uh, like, normalizing of the numeric features. [01:06:07] Uh, then, uh, like, simply, you call, uh, we are calling logistic regression model here. [01:06:12] And, uh, passing in this data, that train data, X train and Y train. [01:06:17] So, once we have that, [01:06:20] Once we have this data, uh, the training finishes here. So, all the effort was to do stratified K-folding, you see? The training function is, like, it sounds, uh, seems, like, almost trivial. [01:06:33] But to obtain the whole struggle is to how to obtain this in as good manner as we want. Otherwise, training is just, like, one line here. [01:06:41] Okay, so once model has been trained, uh, we can then predict from that model, do predictions. So here, I'm doing both… saving both the predictions on train and test data. [01:06:54] So, you can see model.predict, I'm using the X train. [01:06:59] And I'm doing an X test. Of course, uh… [01:07:02] Performance on test is much more, uh, like, importance to us, because it was not shown during testing. [01:07:07] Uh, yeah. So, and model.predict gives me a hard-coded, uh, ones and zeros. [01:07:12] So, you all… so these lists, this will be a list of 1 and 0s. [01:07:18] And uh… and if I want product probability values, so this is like model.predictPropA, [01:07:24] If you pass on your… whatever your X over here, what you will want, uh, obtain is, uh, probabilistic values of [01:07:33] for each sample. These will not be hard-coded. [01:07:35] So, we obtain the probabilistic values also, because one of our matrix, uh, AUC score, it expects that we give it probabilities, not the hardcoded values. [01:07:43] So, these are our predictions stored here. [01:07:46] Once we have our prediction, [01:07:48] Uh, we obtain accuracy, and, uh, so basically, because, uh, uh, [01:07:54] like, this model, uh, this model will be trained 3 times because, uh, we are having three-fold. So, uh, we are saving the performance each time, so that's why I have, like, initialized that list. [01:08:06] Initially, in the outer… [01:08:08] In the outer loop. [01:08:11] So, for this particular inner loop, [01:08:14] whatever you will get the accuracy, it will be stored here. [01:08:19] So, similarly, the train accuracy will be stored. Similarly, the test accuracy will be stored. [01:08:24] Again, precision. [01:08:26] And, uh, in precision, uh, we are passing the ground truth. This is ground truth, and this is the predictions obtained from the, uh, from here at this step. [01:08:34] You are seeing this zero division equal to zero, that means sometimes, if you see the precision formula, there is a, like, a denominator over there, uh, TP upon FP, if I'm not wrong. TP upon F is there, so sometimes it can lead to zero value. [01:08:48] like, true positive, false positive is equal to 0, and then it will be undefined number, so that's why it's saying that if division is by zero, [01:08:57] How to treat it. If division is by 0, give that particular instance a zero value. [01:09:03] Otherwise, it will throw an error. So that's why, in precision and recall, if you see, and in F1 also, we are using zero division, we are using this particular [01:09:13] like, argument in the function. [01:09:16] So yeah, so… ground truth, and this is why trained by ground truth? This is my prediction values I am giving to the precision score. [01:09:23] Similarly, I'm giving to the… [01:09:25] Recall score, uh, then I'm giving to the F1 score, and lastly, [01:09:32] for calculating RUCUC, as I said, we will give a ground truth. [01:09:36] And for the predictions, we are giving the predicted probabilities, not the hard predictions of 1 and 0. So, this will give… [01:09:42] And you see, this is an inner loop, so, uh… [01:09:45] If we are having 3-fold, this will run, uh, 3 times. If you are in 5-fold, it will run 5 times, so you will have, uh… so if you are running it 3 times for the three-fold, [01:09:54] You will have a list of 3 accuracies. Similarly, for 5, you will have a list of 5 accuracy, and for each metric, there will be 5 of them. [01:10:03] So, at the end, when we come out of it, what we will, uh, do is… [01:10:09] We will mean over our list. So, you see, this was my… [01:10:14] like, in the outer loop, I had defined these lists. So, uh, this is now again part of the outer loop. So we, you see, I'm doing a mean over all my, uh, accuracy, train accuracy, test accuracy, similarly train precision, test precision. I'm doing a mean over them. [01:10:29] And all this, like, I have hard-codedly defined a dictionary previously. [01:10:35] store of all of this. That's why I'm able to, like, now… [01:10:41] and just append. So, I'm also appending the fold, uh, number, uh, like, how much is the fold value. [01:10:46] And then I'm appending that particular train test accuracy, everything. So, this all, like, is stored in a… [01:10:53] dictionary here, in this manner, and the reason I'm doing this is because then I will print a data frame, and it is very easier to print a data frame if the, like, information is stored in a dictionary, sort of. [01:11:06] So, you can see, at the last, when all that has been, uh, we are out of the loop now. So, when we do, uh, pd.dataframe, we want a data frame. [01:11:16] And if you just pass it a dictionary, so based on the dictionary, it itself, uh… [01:11:20] prints you a nice data frame. So, that's why we are adding everything in the dictionary. [01:11:25] So, yeah, so for readability, because these decimal values can extend to many digits, so I am restricting to first 4 digits. First four values of the decimal. [01:11:35] And, uh, yeah, this I have, like, printed all the… [01:11:40] test metresol only, although… [01:11:43] in metrics DF, like, there are… [01:11:48] many columns, basically, there are, uh, I have, like, [01:11:50] For printing, I'm only showing you the test ones, although if you do a MetalsDF, we can see all the columns. [01:11:58] So, but, uh, that makes it very, like, congested, so only for the test ones. [01:12:03] So you can see here, in this particular dataset, [01:12:06] Uh, uh, Lake… [01:12:09] changing the number of folds is not like… the performance is almost stable, and number of folds changing is not helping that much. If you see, uh, in the [01:12:18] In the fold 5, the performance highest. [01:12:21] But the gains are, like, marginal, like, in fourth fold, like, uh, in fourth fold, it was 6411, now it is 6425. [01:12:31] And 4 and 5 are only, like, better. The 3 and 6, 7, again, like, reduces to 63%. [01:12:39] Uh, yeah. So, this is this data frame, and similarly, because this data frame is there, so I, at the end, [01:12:48] Uh, like, plotted a graph over it, so we are plotting a 2, 3, like, 6 subplots are there, and the last one is, like, empty, so overall, 5 are there. [01:12:57] And again, this information, which is in data frame to you, uh, this is presented here in form of a graph. So, here, both train and test things are printed. [01:13:08] Uh, but decision is usually made on how the performance is on test set, so you see [01:13:12] like, for the first four graphs, [01:13:15] The highest is on, uh… [01:13:17] five-fold. When the fold is five, the accuracy, precision, recall, and F1 are highest. [01:13:22] Uh, so, like, for… there are clear, like, four metrics on which one thing is highest, so we decide for 5. [01:13:29] Uh, but you see, the gay… but it looks higher, but you see the gay is not that much, so that's what they're saying, that… [01:13:37] The lowest value in this graph, in the first graph in accuracy, is 0.6385. [01:13:42] And the highest accuracy, then, is going to 6425. So, gain is, uh, like, not that much. So, yeah. Uh, what we can infer is that, uh… [01:13:50] The results are almost, like, stable across different folds, uh, with minor increments. [01:13:57] So, yeah. So, this finishes the first way how stratified K-Fold. Any questions on this? [01:14:12] Like, now we got this, uh, let's say we assume this third [01:14:17] My fault was giving us the… [01:14:19] most relevant results, or most accurate result. [01:14:22] So now, how… like, where are we using that fold to create the final model? Because… [01:14:29] We haven't saved any model in our previous code. [01:14:30] Yeah, you know, at the end, after this, there's a code where you will now… [01:14:36] Uh, because you have obtained this information that you will do, uh, that 5 is working better. So, you will make your final model here with this. [01:14:45] Okay, okay. [01:14:49] So yeah, but this initial code was just to decide the number of folds which is working the best. So, when you, like, uh… [01:14:56] When you come… when you have that decision, then you, uh, make one single model. [01:15:05] Yeah, yeah, okay. So, just let me, uh, just finish this second section also, then it will be, like, fully completed. So, uh, the second section is, uh, uh… [01:15:14] Yeah, just a minute. [01:15:17] The second section is everything is same as the previous block, just the categorical features, they are one-hot encoded this time, instead of integer encoded. [01:15:26] So, if you see, uh, let's say everything is repeated, uh, the features X and Y obtained in the same manner as before. [01:15:32] Categorical and numerical columns are separated. There are fold values, again. [01:15:37] And this dictionary is also the same. We are looping through fold values, this is also the same. [01:15:41] Uh, just the differences, uh, in earlier one, it had written one audio encoder, now it has written 100 encoder. [01:15:49] So, this is the difference. So, uh, how this different correspondence is, uh, the difference just come in this part. [01:15:54] Now, uh, all this code below it is also the same, the modeling part is same. [01:15:59] The difference comes here. I am using one-out encoder for categorical columns. So, for my categorical columns, when I do a one-off encoding, [01:16:08] Uh, so as we know that if there are k features in one… if there are K categories in one column, so you will have new K columns, uh, uh… [01:16:17] in ordinary code or indigent encoder, for each column, you only obtain one single column only, so… [01:16:23] Here, you will opt in more columns. So, when I do this step, [01:16:27] I obtain all those columns here, like, the newer columns, because for each feature, for each category, it's making a separate column. [01:16:35] For each value of category. So, that thing is stored here. [01:16:40] Similarly, for test set also, it, uh, based on, uh… [01:16:45] One-hot encoding, it is making these particular columns. [01:16:50] And, uh… [01:16:52] the numeric features, the transformation of numeric features is same that you do a pit transform over train and only transform over test. [01:16:59] But at the end, uh, what, uh, we are, uh, we are doing is, like, we are doing in this part is that, uh, [01:17:07] We have just not… we, uh… we have separately, uh, kept separate the numeric columns information, and, uh, [01:17:16] test column information, and the one hot columns information, one hot coded. And then, again, like, concatenated them. [01:17:22] So, suppose… so, uh, assume there are 10 features of numeric. So, when I do this… [01:17:29] I will opt, uh, I will obtain a 10-column data set in which all the numeric features are normalized. [01:17:36] And assume that, suppose there are, like, 4 categorical features initially, but when I do this, uh… [01:17:43] based on, like, number of categories I have, I obtain, like, 14 or 20 different, uh, like, data set of these 100 gold features. [01:17:53] So, assume, uh, like, so, 10 numeric features here, 14 features here, both of them concatenated with 24 features. That is what is done in this stacking over here. [01:18:04] And once you obtain this, this is your new full dataset. [01:18:07] And once… this is your one… again, this is your full dataset, you can do a similar thing, train and test over it. [01:18:14] and obtain. So, just the differences of the encoding of the categorical features, rest thing is same. [01:18:22] And if you see the results, also, uh, so here, if you see, the results are more stable. [01:18:27] But, uh, of course, in this, you can see that in one hot… [01:18:31] integer encoding, it was coming 64, here, it is not touching 64 anywhere, but uh… [01:18:37] across the fold, results are more stable, but overall, we can see… see… say that, like, [01:18:42] NDR encoding, uh, is, like, working best here. Usually, it does not, but in this case, [01:18:48] However, the improvement is not that much, but yeah, integer encoding worked better here, as compared to one-hot encoding. [01:18:55] And, uh, similarly, [01:18:57] Like, these are their… [01:18:58] These are our, uh, like, the similar graph as previous one on train and test folds. [01:19:07] Uh, yeah. So, uh, so after… once we do this, we have, uh, yeah, so one more, like, information. So here, if you see the highest performances of 4, uh, [01:19:18] I have taken… chosen folds, uh, like, four, because it is, like, performing, uh, if you see the graph also. [01:19:27] The peak is at 4 in, uh… [01:19:30] In almost all of them. So I've chosen 4. So, you can… so that this decision can also change based on if you're doing one-hot encoding or categorical, uh, integer encoding, the number of folds, they might change. [01:19:42] So, this thing is there. So… [01:19:44] So, yeah, so after that, we will, based on the selection of fold, we will do the final model training. So, uh, yeah, so this finishes the selecting of the number of folds. So, this thing is okay. [01:20:02] Okay, thank you. So, yeah. [01:20:04] So, uh, so final model can also be, like, made, uh, like, in two ways. [01:20:12] Yeah, Ashish, you can share this code in the chart to us, in the link. [01:20:15] Uh, the code. [01:20:19] Uh, the code… [01:20:20] Yeah. Accordingly, since the whole link. [01:20:24] Okay, so, uh… [01:20:25] What happened with the code. [01:20:27] The manager can share the lab. [01:20:30] files link. Actually, I don't have access to. [01:20:48] And does anyone else have access to a link? They can share, because actually I don't have access to the link of the lab files and everything. [01:21:02] I have shared it in the fold. [01:21:05] In the chat. [01:21:07] Okay, thank you so much. Bye-bye. [01:21:15] So, uh, Prakash is okay now? [01:21:24] Okay, so let me proceed further. [01:21:27] Yeah, so, uh… so there are two decisions we came up with, above all, uh, is that, uh, [01:21:35] K equal to 5 worked best when features were in design encoded, and k equal for whoever worked best when… [01:21:41] features work, one hot encoded. [01:21:44] So, based on that, uh, we do a final model, uh, creation. And also, uh, [01:21:50] I'm also, like, comparing with decision tree and random forest classifier, because we already know of these classifiers, and usually, uh, like, you don't… like, [01:22:00] have this, uh, target that you will only use large distribution. You will use the classifier which works best to you. So, on the same dataset, you train multiple classifier and see which one is working best on your performance. [01:22:12] So, yeah. So, for that… [01:22:15] Uh, once, uh, this time, the outer loop is not there, only the… [01:22:21] only one, uh, like, uh, there are only two loops, otherwise, uh, there would, uh, so basic, yeah, so basically, uh, [01:22:28] This time, there would have been a single loop, but since there are… but you are seeing two loops, because… [01:22:34] There is another loop of which model I am using for fitting. So, that's why it's coming. So, let's start. [01:22:41] So, this thing is same, that to obtain X and Y, uh… [01:22:45] How to obtain X and Y, and then this thing is also the same that to… [01:22:49] like, separately obtain the categorical columns and numerical columns, and then here, [01:22:55] I am using a fold of 5. [01:22:58] And, uh, this… and these are my three models on which I will show the results. [01:23:04] logistic decision and, uh, like, random, uh, forest. [01:23:09] One. And, uh, yeah, so… [01:23:13] Yeah, so these are the models I will show the results, and then I will store the results in this, like, list. [01:23:22] And, uh, so first, uh, I'm iterating over this dictionary, which are, like, 3 models, so it will first fit for logistic regression, then decision and random forest. This I'm able to do because, as I said, that the training is very… the training code is very abstract that you just have to do a model.fit, [01:23:37] And all of these classifiers have that function, model.predict and model.fit, everyone has them. So, that's why we can loop through everyone, because the function of training and predictions are same for everyone. [01:23:49] Just the name of the classifier is different, either it is logistic, either it's addition 3 or random forest. [01:23:57] So, yeah, so, uh… [01:24:00] Yeah, so, uh, for each one, I am, again, using a list of everything. [01:24:05] And uh… yeah, so this thing is, uh, now the, uh… [01:24:10] the inner loop is same that you obtain, you do a kf.split based on your fold value, and you obtain indexes for the train fold, you obtain the index for the test fold. [01:24:20] And separate out your data. [01:24:23] And, uh, once your data is separated, you do ordinal encoding over it. [01:24:27] Uh, because in this, uh, this time we are doing, uh, in this code, we are doing ordinal or integer encoding of the categorical features. [01:24:35] For numeric features, we are just normalizing them. [01:24:38] And, uh, yeah, so this thing is just repeated from the previous blog, everything is almost same. [01:24:44] Uh, here we are, uh, uh, this is the place where integer encoding all the categorical features. [01:24:50] This thing corresponds to the train part. We are doing fit underscore transform. [01:24:54] And in test particular, only doing transform. [01:24:56] And similarly for numerical columns, uh, we are doing the normalization, so we are using the scalar, which is the standard scalar. [01:25:03] variable, and doing this. [01:25:10] Uh, is there a question? [01:25:15] On a practical sense, that's what we need to know. [01:25:19] We need to learn and pick the best one. [01:25:22] Okay, yeah. Uh… [01:25:24] Krishna has asked this question, Aya. [01:25:26] Uh, so yeah, uh, uh… [01:25:29] Uh, I think the question, Phil, was slightly long. We need to run and pick up the best one based reason. [01:25:36] or judgmentally decide models. [01:25:38] Yeah, so, uh, normally, uh, the EDA part, so if a newer dataset come to you, you always do the EDA thing, you always do, because… [01:25:49] the more time is spent on preprocessing and exploring the data, hitting the… like, exploring the aspects of it, which we… as much as we can from the, like, explored data analysis, because that thing can be visualized. [01:26:04] So, as much possible, we first do the EDA, the preprocessing, like, cleaning of the dataset, everything. [01:26:11] And, uh, like, we… like, normally we don't take, uh, we, we are not judgmental about it. We will only want to fit a particular classifier. [01:26:20] We have a range of clusters available to us. So, usually, we will, uh, from the same dataset, we will, uh, like, [01:26:28] obtain the performance of all of the classifiers, so at least more than one classifier, or multiple classifier, 3, 4, 5, we will obtain. Because, uh, logistical relation, as you see, it can only uncover the linear relationship from the, uh, data. [01:26:44] And these classified enough for us, uh, they can model nonlinear relationships also. [01:26:49] So, whatever works best on my data, maybe my data has nonlinear relationships, [01:26:55] Maybe my data is only linear relationship of the input and target variable. So, we always, like, do… [01:27:01] on a multiple, uh, multiple classifier to see the performance difference. [01:27:06] So, this thing is, like, always there. We, uh, don't just usually, uh, like, fix with one… just one class. At least for machine learning models, this is the case, that because that, uh… [01:27:17] the training time is, like, almost same for everyone. It's just that the scale of data, that is different, and the coding, like, the… [01:27:27] training code and prediction code is almost same for everyone. So, yeah, so we train multiple classifiers in order to, like, see what works best for us. [01:27:38] So, but obviously, I think in machine learning models, these newer [01:27:43] boosting trees and everything, they work far best, like XGBoost and everything. [01:27:49] So, that's why, like, in many people in production, they only use them. But when we are having a newer dataset, we will first see its performance on multiple classifiers to see, like… like, suppose if our model is performing [01:28:00] good using our simple logistic regression, then why to go for a complex model? So, that's like a… [01:28:06] That's why we explore multiple. [01:28:08] things. So, I think this question is addressed. It's okay, Krishna? [01:28:16] Uh, so just I want to add here… [01:28:17] Uh, the question was on… [01:28:20] Whether we need to do exploratory data analysis prior to doing any coding, [01:28:27] And, uh, this will be followed by… [01:28:30] you know, developing the different models. [01:28:32] And then picking the best one. So, I think answer already given by Ashish, but the question… [01:28:39] was this, and definitely… [01:28:41] EDA is a good idea, and then we iterate on the models and see which one is the best one. [01:28:46] Uh, so… and it will be difficult to make a decision, uh, [01:28:52] Or judge which model will be best without actually looking at the performance of the model. [01:28:57] Yeah, Ashish, over to you, please carry on. [01:28:59] Okay. [01:29:01] Okay, so, uh… [01:29:03] Yeah, so, uh, so once we have, uh, integer encoded our data and normalized the numeric features, [01:29:09] Uh, then we do, again, the similarly model.fit, and model.predict over everything. [01:29:16] And obtain the, like, the, uh, like, the accuracy score, the precision, uh, precision score, recall score, reference score. [01:29:23] Uh, so this is the inner loop, because, uh, remember, uh, like, if there are three folds, the model will run 3 times, so each time, it will compute an accuracy, and we will store. So, basically, there will be 3 values of accuracy. [01:29:34] three-value precision, 3 value of recall, F1. And when we come outside of it, uh, so the outer loop is the, uh, model name, because we are looping through the… here, the outer loop corresponds to which model we are doing. So, we are doing three models here. [01:29:47] logistic decision tree random for us, so… [01:29:49] We are storing that name here, and because there are 3 values of accuracy and everything, [01:29:56] After, uh, each fold. So basically, there are five folds here. We have… so basically, there will be 5 values of accuracy, so we'll… mean over them to obtain one single value. [01:30:06] Minor them to precision, Minor them to obtain recall and everything. And, uh, this will, uh, this is, like, at the end presented in this form. So, uh… [01:30:16] We see that, uh… [01:30:18] Although the gain is not that much, [01:30:21] Uh, but, uh, like, we can see that, uh, random forest is working better. [01:30:26] Even, uh, although the decision to use 5Volts came from logistic radiation, [01:30:31] Uh, but, uh, random forest is working better, and right now, we have not, like, optimized any parameters for random forest, just… [01:30:38] Use the default parameters, but still, it turned out better. [01:30:41] as compared to last iteration. So, there are maybe some nonlinear patental data. [01:30:47] that it is able to work better here. [01:30:51] So, similarly, uh… [01:30:54] This is, like, again, the final model comparison, but this time, [01:30:59] Uh, features are, uh, the category features are one-hot encoded, and if you, uh… [01:31:05] Good question. [01:31:10] Can we use GINI or Entropy? [01:31:13] Again, to see the results as one more. [01:31:17] Yeah, we can use linear entropy to see the difference, like, if you are, uh, if you first, uh, like, from this basic analysis, so, uh, the default is, like, the entropy, uh, here, uh, [01:31:29] the default one is use entropy. So, we, uh, in addition to random for the default entropy, so we came up with this value. [01:31:36] Now, if you are inter… now, if you double down that we, uh, we are seeing that random forest is working better. So then, we can experiment with either GINI as the splitting criteria, [01:31:47] that if there are more gains over that. But the thing is that, very carefully, we have to tune, because, uh… [01:31:55] Because there are so many parameters to be tuned here. So, first, number of holds was one thing, so I, like, showed that we are looping through number of folds for logic regression. [01:32:05] Then, we settled on five-fold, uh, which, uh, here. Now, if we want to, uh, like, uh, we are choosing that, uh, random forest will work well for us. [01:32:17] Then we can separately tune the hyperparameter random forest, which are working well on our dataset. So, yeah. [01:32:24] that thing a little bit. [01:32:31] sigmoid graph of the logistic regression model. Uh… [01:32:37] Okay, so, uh… [01:32:40] like, uh… [01:32:44] I don't… yeah. [01:32:45] Let me take this, uh, sheet. So, actually, the sigmoid graph that I had [01:32:51] shown you was only to represent how… [01:32:53] Uh, the decision boundary is there when we are talking about nonlinear data. [01:32:59] However, because what we are using logistic regression is for classification. [01:33:04] So, what you will see here as an outcome of the code will be the classification which is being done by the logistic regression. [01:33:12] You will not be able to view a sigmoid graph. [01:33:15] Okay. So that was only a… [01:33:18] theoretical part that I took, that I showed you, that how… [01:33:23] Uh, the sigmoid curve… [01:33:24] maps onto that nonlinear data, and then how the logouts of it is used. [01:33:31] to make the classification. [01:33:33] However, you don't really require to see all that. All you need is a classification outcome, which is what you're seeing here. [01:33:41] Okay? [01:33:45] Yeah, okay. So, just to add, if you… here, we are using logical relation from SKLAN, so that's why we are not… we, uh, don't have control over its sigma derivative, everything. But if you write your own logical relation function, then maybe you can… we can see the sigmoid graph also. [01:33:59] So, yeah. So, let's finish with this section, uh, then, uh… [01:34:04] Okay, so let's finish with just this section. So, here, uh, in the last, uh, again, [01:34:12] I'm, uh, this time integers, the category of features are one-hot encoded, and as from the above observation, we saw that when features were hot encoded in logistics, [01:34:22] Now, k equal to 4 fold was coming best. So, again, uh, the code is, uh, the same as previous block, just for the difference of the 100 good features. [01:34:31] So, X and Y, we obtained are same as above. [01:34:35] And, uh, again, uh, we are selecting the categorical columns separately, and selecting the numerical columns separately. [01:34:43] And then, this time, I'm using a split of 4, 4-fold. And again, these are the, like, 3 models on which I'm [01:34:49] Doing the, uh, like, comparing the benchmark results. [01:34:54] So, I will… I will loop through each of these models first. [01:34:59] And then, within each of these models, I will make a 4-fold… [01:35:05] Like, I will do a four-fold cross-validation. [01:35:09] So, within each model, the four-fold validation code is inside this inner loop. [01:35:13] So, again, the same thing, uh, when you have made, uh, like, full… this is an object of the split, uh, of the split. [01:35:22] And when you do a kf.split x comma Y here, so this time there is four-fold. [01:35:28] So, the indices which belong to the training folds, we come in this train index, and indices which belong to [01:35:35] the test fold will come in this test index. So, based on that, it returns me indexes, so I can obtain the subset of the data, which is trained, based on their indexes here. [01:35:49] from the X, and similarly, I can obtain that subset of the data which belong to test, uh, based, uh, from here. [01:35:55] So this gives me a subset of the trained data and subset of the test data based on the fold indexes. [01:36:03] And similarly, for a test for the Y values are also… we are doing the same. And as we see… yeah, this thing is same as previous one. The difference is just that [01:36:10] We are doing a one-hot encoder this time to encode the category features, not the integer encoder, uh, as used in previous one. [01:36:19] So yeah, so, uh, this thing is a… so, these are the parameter, uh… [01:36:24] We are using, and this is the standard scalar thing. [01:36:28] So, again, uh, this first parameters handle unknown equal to ignore, so basically, we are using, uh, we are encoding, we are using the trained data, and encoding all the categories that we see in train data. So, sometime it might happen that, uh, in our test data, there is a new category. [01:36:43] Which was not seen in trains. So, what to do for that? So, that's why it is saying that ignore that particular, uh… [01:36:49] sample together, uh, that… together. If you see an unknown value. [01:36:56] And, uh, yeah, so… and sparse output equal to false is that we want the entire 2D matrix to return. Otherwise, it returns a compressed version, because suppose there are… if there are 10 categories for one particular column, it will return you 10 columns. [01:37:10] And many of the values will be 0 in it. So, it returns, I think, uh, it returns a very compressed version of the, uh, of it, so, but we want a full 2D matrix, uh, which is consisting of 1s and zeros, so that's why… [01:37:22] We did a sparse output equal to false. [01:37:25] And then the same thing, you do a fit transform, [01:37:27] over your, uh, categorical columns, and obtain the new, uh, uh, [01:37:33] like, the one-hot encoded features. [01:37:35] for the train and test. [01:37:37] And the difference is, again, that fit and transform, both are done on, uh, train, and only transform is done on test. [01:37:43] Similarly, for the, uh, uh, like, the numeric features, fit and transform is done on the [01:37:49] train, and only transform is done on the test. [01:37:51] Then we combine them together, like we discussed in the previous above block. [01:37:56] At this point, uh, X train is having [01:38:01] integer, uh, categorical features are 100 encoded, and integer features… and numeric features are, uh, like, normalized. [01:38:08] At this point. So, once we have all of them, uh, we will do a X train comma, uh, Y-train, uh, do the model training. [01:38:17] And again, uh, uh, what we can do is, like, we can use our X test set and see what are the predictions over it. [01:38:26] And, uh, to get the probabilities, uh, we can do… [01:38:30] predict underscore prop A, which will give me probabilities, uh, so, uh, as I said earlier, this will give me hard-coded values, 1 and 0s. [01:38:38] predictions, model.predict, and model.predict underscore probabay will give me probabilities. [01:38:43] probabilistic values between 0 and 1. So, uh… so then we… I'm storing my accuracy. [01:38:51] Y underscore test my ground truth, true values. This one is the prediction. [01:38:55] And again, I can obtain the accuracy. Accuracy score is, like, the function from the scale on. [01:39:00] And similarly, precision. [01:39:03] ground truth, the true values and the predictor values, similarly recall, similarly F1. [01:39:09] And uh… and in the AUC, we give the ground truth, but we don't give the, uh, [01:39:14] predicted values of 0 and 1, we give the probabilities. [01:39:17] to compute the EUC score, so that's why [01:39:19] Y underscore prop, this… this… what we obtained from this step, we pass on here. [01:39:25] And finally, uh… [01:39:26] So, uh, so basically, this inner loop, so this, uh, this is part of inner loop, and this is part of outer loop. [01:39:32] So, inner loop, so since there are four folds, so for each [01:39:36] model, 4 times training will happen. [01:39:38] Since four-time training will, uh, 4 time training will happen for one particular model, so there will be four values of each metric, accuracy, precision, recall, everything. [01:39:47] So, at the end, we will do a mean over that to obtain a single value per model of accuracy position, recall, and everything. So, that's why it is, like, mean over them. [01:39:58] And, uh, finally, we see the… we pin the values. [01:40:01] So, uh, here also the same conclusion is there, that in this particular dataset, if you see, this is, like, 65, [01:40:09] And this was 66. So, in this particular dataset, integer encoding is working better than one-hot encoding. [01:40:15] Uh, you know, so, uh, like, man, I also said in last class that, uh, we can just, uh, like, decide that which one better, we have to take all, both of them. So, in this particular dataset. [01:40:25] 100 encoding, uh, is performing slightly inferior, even for, like, uh, even if you change the model and everything. [01:40:31] So yeah, so, uh… [01:40:34] like, this finishes, like… [01:40:36] One particular part of code, uh, of, uh, using K for validation, so, yeah. [01:40:42] So, any questions till this part? [01:40:55] Okay, so let's just finish with the code notebook then, uh, there are no questions. [01:40:58] So, uh, so the main part is over, uh, just added this section to show that, uh, we can use kfold and logistic regression for multi-class classification also. So here, uh, in this dataset, [01:41:11] There are, like, 4 classes to classify from. Here, we are using two. And this is added to, like, to show, uh, that… [01:41:19] In the particular… in the above dataset, we were seeing that integer encoding was working best. [01:41:26] But here, we will see that the reverse is happening, the 100 encoding will work best. So, otherwise, most of them is same. [01:41:32] So, let's see the other dataset which we have shared with you by this name. [01:41:37] This will run in this particular section, car underscore V1.data. [01:41:41] So, in this one, there is just, like, information of the, uh, cars, and there are, uh, uh, like, and, uh, [01:41:50] All of these data types are categorical, so this thing is interesting that… [01:41:54] All of these data types are categorical, and the target variable, the predictor variable, it is like, there are four categories, whether the car is in an unacceptable condition, acceptable, [01:42:04] good condition or very good. So… [01:42:07] based on, like, these particular information, buying price level, maintenance level, number of doors, capacity, luggage boot size, safety rating, everything. [01:42:16] They are predicting the acceptability class for the car. [01:42:19] And they have, like, defined for acceptability card. [01:42:23] classes. And here, when we load the dataset, there are only [01:42:29] a thousand, uh, like, 1.7K, uh, 1700 rows are there. So, that means… [01:42:35] When we do K4 validation, now we can try the bigger values also. So, we can try 10, 250, like, 10 till… even 15 folds we can try, because the samples are very less. Earlier, we were restricted in the above code, we were having only 11,000 entries. [01:42:50] And it was taking slight time, like, more than 30 seconds it was taking for, uh, if you are looping through all the 3-4… all the folds, 3 to 7. So here, it will not take that much time, so I have, like, experimented with 10 folds here. [01:43:03] Because the samples are very less. [01:43:05] And, uh, moving on, uh, if we just again see that the thing is same that we do df.info. [01:43:12] Uh, let's start from here. [01:43:15] So, again, like, as the information is same as in the data card, that all the [01:43:21] features are categorical, and the target variable is also, like, in string here, so… which is, like, the categories, 4 categories here. [01:43:29] the class of the car. And again, uh, like previously, these things are added similarly, that we check for missing values. [01:43:36] So, there are no such missing values, then we can see if there are any empty strings, because in category, there are strings. [01:43:43] the categories are in the form of strings, so no, there are no empty strings here. And again, the same thing if any other [01:43:49] value like question mark, NA, null, none are there in any column? The same thing. [01:43:54] And, uh, then, uh, in the previous code, I was, like, separately obtaining all the categorical columns, uh, because the data was a mixture of both, uh, [01:44:05] the data type was a mixture of both object and integer, but here, only object categories are there, so that's why… [01:44:12] Here, I am… what I'm doing is looping through all my columns, because everything is categorical. [01:44:18] And then I am printing the unique values corresponding to each column. So, the last column is class. Apart from that, all other, these columns, these are input features. [01:44:28] So, and even the values are not that much, so you can see that, uh, for this, there are 4 unique values. [01:44:34] 4, 4, and then 3 unique values for each of these columns. [01:44:38] And, uh, similarly, as we proceed, that we do before model building, we do a [01:44:43] distribution of the class proportion of the target variable. The target variable here is by the name class. [01:44:49] And which is, uh, so… [01:44:51] Well, this is the target variable class. When I do value counts over it, and when I set normalize equal to true, I obtain [01:44:58] percentage proportion of each class, so… [01:45:00] When I printed this, [01:45:03] I obtained this particular block of code, so there is 70% is unacceptable cars. [01:45:08] 22% acceptable, and then there is… [01:45:11] like, 3%, 3% good and very good. So, it's very highly imbalanced data, and it's, like, totally tilted towards most of the entries are unacceptable cars. [01:45:21] So, like, very challenging to, uh, like, build a… [01:45:25] Uh, like, even a random model can give me a 70% if I just print unacceptable everywhere, it can give me 70%. So, yeah. [01:45:33] So, it's very challenging, and even its very smaller data, only 1700 samples are there. [01:45:38] So, uh, that, uh, that… [01:45:42] Approach is same for obtaining the optimal number of K4s. That approach remains same. [01:45:49] Uh, that you… so here, there is no ID as such. I think there is… yeah, so… [01:45:56] So, in… since there is no ID column, there is nothing… there is nothing to be dropped, just you separate, like, you just wanted these four, all columns will be your input, X, [01:46:05] And this will be my input, uh, this will be my output Y target variable. So, this is what we are, uh, doing here. [01:46:12] Uh, the target column class, we are dropping from our [01:46:16] or list, and uh… [01:46:18] Let's just show again that… [01:46:20] DF.shape… [01:46:25] X.shape. [01:46:34] Yeah, so it will run a bit, because I have… [01:46:37] Uh, I'm running it for 10 epochs, so… so 10 for loops only. [01:46:42] And just taking a bit of time. [01:46:45] So, uh, yeah, it will… [01:46:47] Yeah, so you can see, uh, there were 7… in the dataset, there were 7 columns altogether. The last column is my… what I have to predict, so that's why, when I'm doing df.drop over it, I'm obtaining, uh… [01:47:01] the sixth column, which are the input features on which I want to predict. So that's why I access this. [01:47:06] The shape of X is this, and similarly, the last column, the one single particular column, is my Y. [01:47:12] And, uh, yeah, so here, categorical columns are obtained separately, uh, because this is being done… although all columns are categorical only, so there was no point of opting them separately, I just wanted [01:47:26] the name of all the columns, so that I don't have to manually write the name. So that's why I'm doing that, select data types. [01:47:32] Which are object in data type, all columns, and put them in a list. So, I have a name of all the categorical columns with me. [01:47:38] So, yeah, so this time, as you can see, I have tried till 10. I have tried 10 folds. [01:47:45] And, uh, again, these are dictionary to store all the matrices, so… [01:47:50] The code is almost the same that I'm, uh, doing, uh, like, I'm looping through, uh, all these values, these values of the list. [01:48:00] I'm looping through all of them at one at a time, and then I'm doing a stratified care fold. So, stratified is more important here, because the class proportion is very imbalanced, 70% of one class. [01:48:10] Then, uh… [01:48:13] 22% next class. So, we want all our folds to also have the same proportion. That's why stratified K-Fold. [01:48:19] And again, the code is similar, that you do a split over the entire dataset, kf.split, x comma Y. [01:48:26] And you obtain the indices of the train fold, you obtain the indices of the test fold, and then obtain the subset of trade and test [01:48:35] data from that, and uh… [01:48:36] This time, there is no numeric feature, so no point of doing standard scalar, so that thing has been commented out here. [01:48:43] You only have categorical features, so that's why… [01:48:48] First, we are doing ordinal encoding of these categorical features. [01:48:52] So, encoder, order encoder, and uh… [01:48:55] And then, as earlier, this thing is, like, copied from the above block. [01:49:00] So, just for naming convention, uh, [01:49:03] distinguish called earlier raw, I have now… and I'm calling train and test, because to keep, uh, so that it is easier to read the code. [01:49:11] So yeah, so… [01:49:14] for integer features, for doing the integer encoding. On trained data, we are doing fit underscore transform. On test data, we are only doing transform. So, here, since there is no numeric feature, at this end, [01:49:24] Data will be integer encoded, and it is ready for fitting at this… at this point. [01:49:29] after this block. So, once it is done, [01:49:33] I'm calling my logistic regression, doing a fitting over it, learning… training is happening in this part, and then [01:49:40] Predictions are in, uh, this part. [01:49:43] uh, like, each of them are, uh, train and test, both are, like, [01:49:49] train predictions and test predictions, both are being done here to store, and this all, like, is like the above previous blocks. These things are written, like, repeated, that you store the accuracy of the train. [01:50:01] You score the accuracy of the test. Similarly, precision recall and everything. And at the end, uh… [01:50:06] you obtain, uh, like a, uh, [01:50:09] performance of the, like, you obtain a performance of different folds. [01:50:15] ranging from 3 to 10, their test precision recall, everything is obtained. [01:50:20] So, uh, what I want to point out is that, uh, although we are, like, use 3 to 10 all folds, but if you see the highest proportion of class was 70%, and the model is, when we are doing integer encoding, [01:50:34] The model is not even able to learn that. You see, the accuracy is hovering around 69%. [01:50:40] So, this is the problem here with the, like, one integer encoding, but as soon as I change in the last… this is the last block, then we can take the question. [01:50:52] So, in the last block, as soon as I change, I have the same code is there, just that I am doing 100 encoding this time. [01:50:58] So, you see, there is a good gain in all the metrics. [01:51:04] And if we see both 3 and 4 folds are working very, uh, like, best, then other one. [01:51:11] So, uh, but yeah, so… the encoding of the categorical features, like, change the performance much here in this one. [01:51:22] So, yeah, so there were questions… [01:51:29] Uh, Sushil, how do we define the threshold for the classification? [01:51:32] So, right now, uh, we are using the default threshold is 0.5, but in scikit-LAN, they have this option that you can, uh, so basically, you have a pre-trade probabilities, but if you want that… [01:51:44] for your particular dataset, uh, like, you want a 0.6 threshold of points and threshold. [01:51:49] You can… so basically, we can experiment with different threshold values also. And if we see that our dataset 0.6 is giving better gains, then we may shift to 0.6, [01:51:59] Or maybe 0.4, like that. So yeah, this is also a parameter that can be tuned. [01:52:05] That… for whichreshold we are coming… we are coming better. But obviously, this is not concerned with training. Once our model is trained, then we can experiment with whatever has been trained, because in training, there is no such, uh… [01:52:18] Yeah. [01:52:21] Ma'am, you want to, like, add on this threshold? Like, I'm a bit confused. [01:52:28] Sure. So, actually, if we talk about, uh… Deciding the threshold for classification, I believe this was the question that was asked. [01:52:38] So, that purely depends on, uh… Uh, actually, it might depend on a couple of things. If we have some idea about the data, or if we have some domain expertise, or we iterate and see. [01:52:51] What kind of threshold gives the best performance and least amount of error? [01:52:56] That way, we can decide the threshold, but there is no clear-cut rule. [01:53:00] So, to say that this will be the threshold or this won't be the threshold, okay? [01:53:04] So, threshold needs to be decided, either it needs to be defined looking at the domain or at the data. [01:53:11] Or we need to take a variety of… a range of thresholds and see which threshold. [01:53:16] gives the best performance of classification accordingly, we decide that. [01:53:22] Okay. [01:53:30] Yeah, Sushri has asked the question. Sushri, is that okay? [01:53:31] Is it okay, whosoever asked the question? Okay. [01:53:37] Uh, yes, ma'am, but I was asking, where do we define? I mean, while calling the procedure or function. [01:53:43] First of scikit-learn, where do we define that threshold? [01:53:47] Uh, so, Ashish, you can just come to it. [01:53:48] Yeah, yeah. [01:53:51] Yeah, so I think, uh… [01:53:56] Metal observation is killed on. [01:54:05] Just give me a minute, uh… [01:54:18] Actually, yeah, so basically, uh, is my… this screen visible? The browser, SKL and I'm open. [01:54:29] Yes. [01:54:30] Yeah, so actually, what they are doing is they… like, they have given the source code and all. If you direct, directly do model.predict, [01:54:37] There, uh, whatever the internally they have used threshold, and most probably it's 0.5, you are getting hard decisions. They have given the other option that you get the [01:54:47] probabilities instead of the hard predictions. [01:54:50] So, we can then manually define our own threshold. If we get probability, then we can manually decide threshold. Otherwise, we will have to see, like, inside the source code, uh, how… what the actual values, like, how they are, like, using threshold value. [01:55:06] But what we can do is, we can use this particular log, like, I have also used my function. [01:55:12] Uh, okay, so yeah, this one. This one worked best, prob A. So, you have probability estimate between 0 and 1. [01:55:17] Then you can manually just write a single block of line that, uh, greater than this, convert them to 1, less than that, less than this value, convert them to 0. [01:55:31] Okay. Yeah, so there was, like, one more question. [01:55:38] As part of EDA, if the correlation [01:55:42] We're doing input is low, can I go with logistic? [01:55:45] Elibration helps with other classification model. [01:55:50] Yeah, yeah, this thing is, like, uh… like, this thing is established in theory, when… [01:55:56] Input features themselves have very low correlation, and they are highly correlated with the output feature. People usually… and there's a classification problem, people usually go with logistic regression. [01:56:08] But again, uh, uh, like, uh… [01:56:12] These things are very data-dependent, and it does not harm to train more models, because, as you can see in SKLearn, these are good to use modeling productions. They have been very much optimized, and they don't take much time to train. So, [01:56:25] either you have a very fixed computation budget that, uh, training many classular modulas, like, very resource intensive, [01:56:33] Then, uh, based on your EDA analysis, you shift to just one logistic regression. [01:56:39] Otherwise, the normal advice which people give… get from expert machine learning experts is that you train multiple models, like, and see the performance, at least train more than one model and see the performance. [01:56:58] Okay, so, uh… [01:57:01] I hope, uh, Krishna, uh, I've answered the question. [01:57:05] So yeah, so code is, uh, like, finished, uh, with this, uh, this last one was just added to, like, because… [01:57:13] In the previous one, there was not much gain, so here we can see, uh, like, the difference in encoding can sometimes [01:57:19] change the performance very much so. That's why I was added. Otherwise, yeah, so… [01:57:30] Any questions regarding this last part, this last quote? [01:57:38] Okay, so, uh, before we finish, like, just me, just, um, I think regarding code, I, um, will add some more information. [01:57:47] So, one thing, uh, different in this one is, in the second blog, is that there are four categories to be classified. [01:57:53] So, this is not a binary classification problem. [01:57:55] So, although Leilistic regression is able to handle it, I'm just giving the same function. [01:58:01] So, but the difference, if you see, is in the… [01:58:03] metrics of the, uh… [01:58:05] Uh, like, the metrics which you are getting after, uh, training and testing. [01:58:09] Uh, in the earlier one, there was no such value as average equal to max. If you go to the previous binary trash part, [01:58:14] your code… I had till this line only. I don't have this particular line. So, it was because, uh… [01:58:22] It was a binary classification, there were only two values, 0 and 1, but now there, we have four categories here, acceptable and acceptable if good, very good. [01:58:31] So, what will… what is to be done is, for each class, we will learn separately what is their, uh, for each class, what is their, uh, [01:58:37] For each class of separate configuration metrics will be built, and based on the accuracy for each class, F1, sorry, not accuracy, F1 score, precision and requirement of each class, [01:58:46] We will obtain a value, and then we will average over it. So, that's why it is written average equal to macro, that because for each class, there will be separate confusion metrics, and then you will, uh, whatever precision recall you will obtain, you will average them. [01:59:00] That's why, when you are doing multiple classes classification, you are… have to, like, specify the averaging criteria which you are taking. We are taking here micro. Sorry, macro criteria. [01:59:12] And similarly, the ROC also, for when you are doing multiple classes, uh, because ROC score is defined for binary classification problem, but now you are doing multiple classes, [01:59:23] So, what you will do is, uh, when you do ROC, so suppose you have 4 class, [01:59:28] So, we have written multi-class OVR, it is one verse. It means one versus rest. So, if classes are acceptable, [01:59:34] So, all acceptable will be 1, and other all three classes will make 0. Uh, will make 0, and then you will calculate the AUC value, considering it a binary problem. And similarly, this will be repeated for all the four classes, and then averaged out value will be my [01:59:47] our OCSU score. So, yeah, so this thing is… gets changed when there are… [01:59:52] multiple classes. [01:59:54] Other than that, the code is almost same, totally same from the previous one. Only difference in… [02:00:00] the metrics computation. [02:00:07] So, any questions, uh, regarding this second part? [02:00:26] Any questions, anyone, before we wrap up for today? [02:00:32] Yeah, you can view this link here, Copilot, uh, this related to code, we'll go through it again. [02:00:39] The link… the code is uploaded on the LMS. [02:00:42] Already, you can just have a look. Okay, it's available for you to download and to study, see, do whatever. [02:00:44] Okay. Sure, sure. [02:00:53] Any questions? [02:00:58] So I think then we can wrap up for today. I'll request all of you to stay for a minute, whereas I'll thank Ashe. [02:01:07] For taking us through ashes, you may feel free to leave the session, whereas the others may just stay for a small announcement. [02:01:14] Okay, everyone, okay, uh, thank you, everyone. It was nice talking to you all. [02:01:20] Thanks, Ashish. Okay, so for all of you, uh, the announcement is which I actually had made last time also. [02:01:29] That next weekend will be off on account of New Year. [02:01:34] Uh, so we will not be having any sessions on both Saturday and Sunday the next week. [02:01:39] So that was our announcement. Already, I made it, but I thought I'll just re-announce for those who are not there. [02:01:44] So, then we can break for today. Wish you all a very happy and prosperous New Year. [02:01:51] Thank you. Bye-bye. So, meet. Uh, on the second week. [02:01:55] In January. Thank you, thank you everyone. Thank you so much. Thank you. [02:01:57] Happy New Year, madam. [02:02:04] Thank youI've been doing, man