# 05 2026-02-07 Hands On DL Contd course: Module 3 — Deep Learning & NLP module: Module-3-Deep-Learning-NLP date: 2026-02-07 type: transcript video_url: https://personal-learn.armco.dev/files/_Recordings/Module-3-Deep-Learning-NLP/05_2026-02-07_Hands_On_DL_Contd.mp4 --- [00:06:34] Durga Toshniwal: Good evening to all of you. [00:06:36] Durga Toshniwal: And so, welcome to today's session. So, we'll be continuing what we had done. [00:06:43] Durga Toshniwal: What we had been doing in the last turn. [00:06:47] Durga Toshniwal: And, so, we were, discussing about the hands-on. [00:06:54] Durga Toshniwal: In which we were discussing, the, deep learning part. [00:06:59] Durga Toshniwal: Using MLP, multi-layer Perceptron, and we were doing hyperparameter tuning. [00:07:05] Durga Toshniwal: Where different kinds of hyperparam… different hyperparameters we were trying to tune. [00:07:10] Durga Toshniwal: Including the number of layers, the number of neurons per layer, the activation functions. [00:07:18] Durga Toshniwal: The epochs, and so on and so forth. [00:07:22] Durga Toshniwal: So I think, because that was a rather lengthy quote, so last time we had done it partly. [00:07:31] Durga Toshniwal: We'll continue it today. [00:07:33] Durga Toshniwal: And for that, I'd like to invite Asmita. [00:07:36] Durga Toshniwal: To continue with where we had left. [00:07:40] Durga Toshniwal: Asmit, I'll request you to please enable your video. [00:07:44] Asmita Mahajan: You may please, share your screen. [00:07:49] Durga Toshniwal: And we can start, [00:07:56] Durga Toshniwal: So, I think you can take it from the beginning, and we can come up to that point. [00:08:03] Durga Toshniwal: Very broadly, we can just take the… [00:08:07] Durga Toshniwal: broadheads that we had done. So, it was the EEG eye state, dataset that was being discussed, where the state of the eye was the classification, or the class label. [00:08:21] Durga Toshniwal: Then, So, we tried to convert this ARFF file into the standard Pandas data-free. [00:08:34] Durga Toshniwal: We tried to do, you know, just see what are the columns like, so we looked at the data frame, did some… [00:08:43] Durga Toshniwal: Initial data preprocessing, so we had a description of the data frame, along with [00:08:49] Durga Toshniwal: The standard, descriptors, like, how many null values, and it was, like, 4,980. [00:08:57] Durga Toshniwal: Rows in that data frame, okay? [00:09:01] Durga Toshniwal: Then we checked for missing values, so there were no missing values, and… [00:09:07] Durga Toshniwal: And what you are seeing are the different probes of the EEG, and they were accordingly named. [00:09:15] Durga Toshniwal: Like, you can see here the… FD… [00:09:19] Durga Toshniwal: F3, something, whatever, TT, PT, all these are some standard medical terminologies. [00:09:25] Durga Toshniwal: Which are the EEG signals that are measured from the scalp to identify the state of the eye. [00:09:35] Durga Toshniwal: Okay, carry on. [00:09:37] Durga Toshniwal: Then, the standard, data frame describe was used, function2C. [00:09:44] Durga Toshniwal: What was the statistical characteristics of the data? Some plots. [00:09:50] Durga Toshniwal: for the target variable revealed that the data is more or less balanced. It's not a very unbalanced data. [00:09:58] Durga Toshniwal: Because 0 and 1 are almost equal. [00:10:01] Durga Toshniwal: Then there were other, standard data preprocessing, things that we did to prepare the data, like normalization. [00:10:12] Durga Toshniwal: And then, using, features like zero mean, variance, and all that, you can just [00:10:20] Durga Toshniwal: Then we started with the thing in focus, that is the multi-layer perceptron, which is what we studied in the deep learning part. [00:10:28] Durga Toshniwal: And for that, PyTorch tensors are being utilized. So the first hyperparameter that we tried to use was, the optimizer. Of course, there are a… [00:10:40] Durga Toshniwal: large number of, large types of optimizers that are available. We tried to go with some of the [00:10:48] Durga Toshniwal: more popular ones, like ADAM SGD, and with learning rates of 0.001 and 0.01. [00:10:55] Durga Toshniwal: And we looked at the different combinations, and for that, we, studied the performance. [00:11:02] Durga Toshniwal: The specific code was discussed last time, so I'm not just mentioning the code, I'm just showing the broad flow. [00:11:11] Durga Toshniwal: And with this, it was very much clear, as you can see here, that the train and the test losses are very different. [00:11:19] Durga Toshniwal: We are quite far apart, where the train loss is much less than the test loss for the first [00:11:25] Durga Toshniwal: two cases, which is SJD with 0.001 and 0.01 learning rate, which means that, the model is very, very [00:11:36] Durga Toshniwal: overfitted in these two cases, wherein the other two cases, it… the performance of the model is much better, which… with the ADAM optimizer, and with .001 and 0.01 learning rate. And finally, we chose 0.01 because of the convergence being at a [00:11:55] Durga Toshniwal: Better rate, so we chose to use that. [00:12:00] Durga Toshniwal: Okay, so… -Oh. [00:12:04] Durga Toshniwal: And then, after that, so, I think we stopped at this point, or… [00:12:10] Asmita Mahajan: Yes, ma'am, we have to, like, continue from here now. [00:12:15] Durga Toshniwal: Okay, so I think this was the… [00:12:17] Durga Toshniwal: Broad flow up to this point, where we had only, discussed about [00:12:25] Durga Toshniwal: The optimizer and the learning grid. [00:12:28] Durga Toshniwal: Any questions? Anyone who had gone through and would want to ask anything? [00:12:33] Durga Toshniwal: So far, before we start. I hope all of you have loaded the code. Are there people who haven't loaded? [00:12:42] Durga Toshniwal: Is everyone ready with the code? [00:12:48] Durga Toshniwal: Shall we start the discussion on the core? [00:12:54] Durga Toshniwal: Yes, ma'am, Minister. Yeah. [00:12:56] Deepan Kanagaraj: One question is with respect to Adam. So why the second one, seems to be better? Because I see there is a divergence, but what basis we will figure it out, which is the best, best one? [00:13:10] Durga Toshniwal: Yeah, so, what you can see here, that in the last curve, which is Adam with 0.01 versus Adam with 0.01, you can see that [00:13:21] Durga Toshniwal: The convergence is much faster, which is at around 75 epochs. [00:13:27] Durga Toshniwal: Whereas in the, in the… [00:13:30] Durga Toshniwal: So you can see here that the… [00:13:33] Durga Toshniwal: Ideally, that, if our model is, working fine, then in that… or our model is properly done. [00:13:42] Durga Toshniwal: There are, train losses. [00:13:45] Durga Toshniwal: The test losses and the train losses should be approximately equal, which is… [00:13:51] Durga Toshniwal: Coming out to be our same are at around 75. That is a thing. [00:13:56] Durga Toshniwal: So, in the fourth one, the convergence is faster, hence we chose to use that. [00:14:04] Durga Toshniwal: Otherwise, [00:14:06] Durga Toshniwal: Otherwise, there's quite less difference between the third and the fourth cases, which is the Adam with 0.001 versus Adam with 0.01. [00:14:16] Durga Toshniwal: So that's the only major difference. You can just scroll down a little, the epochs are getting cut, yeah. [00:14:24] Durga Toshniwal: So… That is a thing. [00:14:28] Deepan Kanagaraj: Okay, ma'am. [00:14:29] Durga Toshniwal: Bye. [00:14:31] Durga Toshniwal: Okay, so I think, you can continue, Asmita, from here. [00:14:38] Asmita Mahajan: Okay, ma'am. Thank you. [00:14:42] Asmita Mahajan: So, in the previous plots, we have plotted, the train and test loss, using the history dictionary that we, that was returned from the train and evaluation function. [00:14:57] Asmita Mahajan: Now we'll be plotting the matrices, like the accuracy score, the, precision, F1 score. So, to plot that, we'll be using the, [00:15:08] Asmita Mahajan: So, in the previous, code, we have also created a data frame, from the, metrics and history, [00:15:19] Asmita Mahajan: from the results that we captured, in the results dictionary, and in the results dictionary, we had the matrices. All the… from… for every, experiment, we had a corresponding matrix, dictionary, which has stored the evaluated scores. [00:15:38] Asmita Mahajan: So, using this data frame, we'll be now plotting our bar plots to compare the errors, or the metrics measures for all the experiments. [00:15:52] Asmita Mahajan: So, what we are doing is, we are just, firstly, we are creating a function to plot the bar plots, in which we are passing the data frame, the rows, how many rows we want, columns, and the title of the overall plot. [00:16:08] Asmita Mahajan: Now, we'll be creating a list metrics, which will store the columns. So, in the columns of the data frame, we had the, metrics, names. So, in metrics, we'll be storing the names of all the metrics, which is accuracy, precision, recall, F1 score, and the AUC score. [00:16:28] Asmita Mahajan: And the index, the data frame index had the, experiment name, which we stored in the form of SGD at the rate 0.001. So we had the optimizer name, at the rate, the learning rate. [00:16:43] Asmita Mahajan: So, we are extracting the index and converting, and before passing it to the labels, these index labels, we are converting it into a string, and then passing it to the index labels. [00:16:57] Asmita Mahajan: Now, what we are doing is, we are, calling the subplot function. [00:17:02] Asmita Mahajan: Passing the rows and columns, the size of the figure… the size of each subplot, and we are creating the overall figure and the subplots. So axis will have the location of the subplots. [00:17:15] Asmita Mahajan: Here, we are flattening the axis, because axis will be a 2D NP array, so we want… so for easy iteration, we want it in a 1D array, so we are just flattening the axis location, and we are just creating a list out of it. [00:17:31] Asmita Mahajan: And iterating over the metrics, list. [00:17:35] Asmita Mahajan: We'll capture the index and the metric name. [00:17:40] Asmita Mahajan: And, from… from the list that we had just created, we are passing the index that the… that we want the location… so this is the location of this first subplot. So, first subplot will be of accuracy. So, we are, just extracting [00:17:57] Asmita Mahajan: The accuracy from the… the value of the accuracies from the data frame, and passing it to values. [00:18:05] Asmita Mahajan: now calling the bar function, which will create the bar plots. So… [00:18:11] Asmita Mahajan: For the bar plot, at the x-axis, we want the index labels, which will be the experiment name. [00:18:19] Asmita Mahajan: On the y-axis, we want the values for each. [00:18:23] Asmita Mahajan: experiment. So, the first plot will be of accuracy. So, we'll compare [00:18:29] Asmita Mahajan: accuracy for each experiment, so that's why we have the index, like the experiment name at the x-axis, and the values for each experiment, so value will be the accuracy value for each experiment. [00:18:44] Asmita Mahajan: And each color is black. So, as you see here, for each, bar. [00:18:50] Asmita Mahajan: we have an outline of black color, so if you will… you can also change it. So, if we'll just try it red here, so… [00:19:00] Asmita Mahajan: Yeah, so you can see the difference. [00:19:03] Asmita Mahajan: So you can just, experiment with this, either whatever you like, you can put that, value here. [00:19:13] Asmita Mahajan: So, this will create the bar for each value. [00:19:17] Asmita Mahajan: Now, we are just setting the title. So, the title of each subplot will be the metric. So, the first plot is of accuracy, so the title is accuracy, Precision, Recall, F1 score, and AUC score. [00:19:32] Asmita Mahajan: We'll also set the X label, which is the optimizer, at the rate LR. [00:19:37] Asmita Mahajan: the Y label will be… will, have the metrics, so accuracy, so we have set the Y label. [00:19:45] Asmita Mahajan: And the while limit is, we are also setting the Y limit. So, here, 0 to 1, so you can see the limit. If we'll not [00:19:52] Asmita Mahajan: If we don't include this, you can see [00:19:58] Asmita Mahajan: That every subplot will have, like, [00:20:02] Asmita Mahajan: it will have a different label, and also, you can see that now my values here, the text is getting, like, it is not included in the plot. So to, like, make our plot's presentable. [00:20:14] Asmita Mahajan: What we will do, we will just set the limit so that, it doesn't, it, the one is also included in our limit. [00:20:23] Asmita Mahajan: So, to make our plots more presentable. [00:20:31] Asmita Mahajan: Also, you, here, the, labels are, or the X sticks, so we have, every experiment name. [00:20:40] Asmita Mahajan: So, what we are doing is, this, X tick para, tick parameters axis X. So, it is telling that the ticks at the X axis, we want them, at a rotation of 30. [00:20:53] Asmita Mahajan: So, you can also experiment with this. [00:20:56] Asmita Mahajan: If I'll rotate it to 90, it will, [00:21:02] Asmita Mahajan: it will be presented, or it will be displayed like this. So you can, it's… it is just to make the plots, readable and, visibly presentable. So we are just, doing… experimenting with every, feature. [00:21:21] Asmita Mahajan: So what it does is, it just, [00:21:24] Asmita Mahajan: rotate the label at the X sticks to 45 degree. [00:21:31] Asmita Mahajan: Now, as you can see here, we… for each bar, we have the value also associated with it, and displayed here. So this text, how you can, display it, so what you can do is, you can enumerate over the values that we extracted here from the data frame. [00:21:48] Asmita Mahajan: So, you can, so what this for loop is doing, it is iterating over each value, so it is iterating over the values in the values list, extracting it in V, [00:22:00] Asmita Mahajan: and the index. So, we are telling users… so, at each subplot, the text… so, using text, we can display the [00:22:09] Asmita Mahajan: values associated with each plot. So what we are doing is, we are telling where the text should be there. So X and here, these two values tell the X and Y [00:22:18] Asmita Mahajan: location. So, this location, how we got it, we have told it that I want… at x-axis, I want to go to X, and on Y-axis, I want [00:22:30] Asmita Mahajan: Oh… [00:22:31] Asmita Mahajan: So we will have the value, the associated value, which is 0.4 something, okay? And we are adding 0.01 so that [00:22:41] Asmita Mahajan: it… the text doesn't appear on the line, so it appears just above the line. So, in this manner, we are, like, giving the location for the text where we want to display the text. [00:22:53] Asmita Mahajan: So this, is telling that I want the value still two decimal places. [00:23:00] Asmita Mahajan: And also, we are, like, telling what should be the horizontal alignment of the text. It should be at center, and the vertical alignment is at bottom, and the font size is 9. [00:23:11] Asmita Mahajan: So, in this manner, you can just, write the text or the values associated with each, [00:23:18] Asmita Mahajan: a barge. [00:23:21] Asmita Mahajan: Now, in the previous… last week, I, explained this, like, we are creating [00:23:30] Asmita Mahajan: we have told, so, what we are doing is we are creating a grid of 3 columns and 2 rows, but here. [00:23:40] Asmita Mahajan: we only have 5 subplots. So, if we're not off the… off this… the location of this grid, so if… [00:23:48] Asmita Mahajan: It is… Yeah, so you can see, so… [00:23:53] Asmita Mahajan: We have 6, we have the, position for 6 subplots, but, in our case, we only have 5 matrices, so only 5 [00:24:04] Asmita Mahajan: subplots will be displayed. But at the sixth position, as it… as, it doesn't have any value, it will be displayed, but we don't want it to be displayed. So that's why what we are doing is, at that location, we are just telling that we… I don't want any subplot to be, displayed there, because I don't have any value for that subplot. So I just… what I will tell, that, I don't want that location, so I'm [00:24:28] Asmita Mahajan: blocking that location. So that's why this line of code will tell our subplot that we don't want any subplot to be displayed at that location. So, this, empty subplot will be… will not be there, it will off… [00:24:44] Asmita Mahajan: that value. [00:24:47] Asmita Mahajan: So it is gone from here. [00:24:50] Asmita Mahajan: So in this manner, and now we are just, telling that, the overall [00:24:59] Asmita Mahajan: Title of the plot? [00:25:01] Asmita Mahajan: And we are just displaying this. [00:25:04] Asmita Mahajan: plot, using plot.show. And also, at the end, we are also displaying the data frame of the, for each [00:25:12] Asmita Mahajan: experiment, and the accuracy, precision, matrices associated with each experiment. And as, here you can see that Adam with learning rate 0.01 has performed well, and has classified well. [00:25:29] Asmita Mahajan: But SGD, with the same learning rate, is not able… was not able to, you know, correctly classify the dataset. [00:25:38] Asmita Mahajan: So here, precision, recall, and F1 score are zero, but Adam is performing well. So in this manner, we will choose… so in our further experiments, we will fix the optimizer to Atom, and then tune other, hyperparameters. [00:25:56] Asmita Mahajan: Any doubt till now? [00:26:02] Asmita Mahajan: Okay, so let's move ahead. [00:26:05] Asmita Mahajan: So now, as we have, fixed our optimizer, we have tuned our optimizer and found out that Adam is performing well, we will fix Adam optimizer, and now we will try to tune our learning rate. So for the learning rate, we are taking 5 values, starting from 0.001 to 0.01. [00:26:24] Asmita Mahajan: And for each learning rate, we will see, which, like, which… at which learning rate our model performs better. [00:26:35] Asmita Mahajan: So, we'll create a MLP class to… [00:26:38] Asmita Mahajan: train, so, why we are doing this? We are tr… so that we can train for each learning rate a new model. [00:26:48] Asmita Mahajan: So what we are doing here is we are passing… we are creating an initialization constructor, and passing input dimensions, and the… [00:26:56] Asmita Mahajan: hidden layer… the size of the hidden layer, and the size of the input layer to the constructor. We are creating a network using sequential [00:27:06] Asmita Mahajan: function, so in a sequence, so we want our input layer, then our hidden layer, and then our output layer. So, in a sequence, we are creating it. [00:27:15] Asmita Mahajan: So, this linear will create our hidden… so this linear function will create the input and hidden layer, so it will, creating a fully connected layer from, input to hidden layer. [00:27:28] Asmita Mahajan: Then, the outputs from the input layer are passed to the hidden layer, but before passing, we are using ReLU activation function for that, and the outputs… so this is… this is a fully connected layer from the hidden to output layer, and output layer has only one neuron, and we are passing [00:27:47] Asmita Mahajan: And in this manner, we are creating a fully connected layer from hidden to output layer. [00:27:52] Durga Toshniwal: Just, I wanted to come in. If you'll just go up in the previous portion, where we were… yeah, so, in these graphs, as well as in the table. [00:28:04] Durga Toshniwal: As you can see here, as per the table also. [00:28:07] Durga Toshniwal: Even if you are able to, sorry to interrupt the flow, but I just forgot to mention, so I thought I will do that. [00:28:15] Durga Toshniwal: So you can see here that if we look at the accuracy, precision, recall, F1 score, and AUC, [00:28:22] Durga Toshniwal: First of all, in the table here, which is the output of all these metrics. [00:28:28] Durga Toshniwal: We can see that, with the atom, and at the rate 0.01, [00:28:33] Durga Toshniwal: All of these metrics are showing the best performance. [00:28:39] Durga Toshniwal: So that is also an indication. I think there was some doubt that, which combination should we use, and from the graph, it wasn't very clear. So if it is also not very clear from the graph, you can actually look at all the metrics, performance metrics. [00:28:57] Durga Toshniwal: And then C. [00:28:59] Durga Toshniwal: where the performance is coming best, it's very clear that with the ADAM and learning rate of 0.01, the model is performing the best. Now, if you look at the bar plots also, you can just scroll up. [00:29:13] Durga Toshniwal: So you can look at, accuracy. Accuracy is, again, highest. As you can see, with the last combination, Adam plus 0.01 learning rate. [00:29:25] Durga Toshniwal: Then you look at precision. So, so for the, accuracy, precision, recall, F1 score and AUC, ROC, AUC. [00:29:35] Durga Toshniwal: Definitely, the last two cases are giving better results as compared to the first two. That is the first observation. Out of the last two. [00:29:45] Durga Toshniwal: Though they are close, but still, Adam with 0.01 learning rate is giving the best, so you can see here with accuracy. [00:29:53] Durga Toshniwal: It's coming to be 0.84, then if you look at the, precision, then you look at recall, you look at F1 score, and then you look at AUC. [00:30:05] Durga Toshniwal: So, it is very clear, there's no doubt that this is a combination that must be our choice out of those that we have experimented with. [00:30:15] Durga Toshniwal: Probably, if we have more combinations and all, then we can think of other things also. Right now, we just tried this, so this is what I wanted to show. [00:30:25] Durga Toshniwal: Just for those. [00:30:27] Durga Toshniwal: For which it wasn't very clear what combination to use by looking only at the graph. [00:30:33] Durga Toshniwal: Okay, Asmita, now you can carry on. [00:30:37] Asmita Mahajan: Thank you, ma'am. [00:30:40] Asmita Mahajan: So… [00:30:42] Asmita Mahajan: Yeah, so, in this manner, we have created the full network, and we are just collecting or returning the raw outputs from this, network, where we are calling it. [00:30:55] Asmita Mahajan: So, the forward function will, pass the inputs from, through, like, it will propagate the inputs through the network. Now, we have train and evaluation function, where we are passing the learning rate. [00:31:12] Asmita Mahajan: and the epochs, and we have set epochs to be 50 now, and have different learning rates, this function will be called. So we have 5 learning rates, for the experiments, and 5 cases, which we'll compare for the learning rates. [00:31:27] Asmita Mahajan: So, this model will be creating the object from the MLP class, and to the MLP class, we are passing the size of the input [00:31:37] Asmita Mahajan: layer, which we are extracting from the Xtrain tensor, and using shape, and from that shape, we are just extracting the columns, because that is what the number of features we have, and the input layer will contain the [00:31:55] Asmita Mahajan: So then, the size of the input layer will be equal to the number of features we have in the dataset. And the size of the heater layer, we are fixing it to 64. [00:32:05] Asmita Mahajan: So, we'll be using, so to calculate the loss, we'll be using the binary cross entropy loss, and in this, the sigmoid function is, implemented implicitly. [00:32:17] Asmita Mahajan: So for Optimizer, here we are just, directly using Adam Optimizer, because we have fixed it, and we have tuned it to Atom Optimizer, and we'll be passing the learning rates, different learning rates, to this Atom Optimizer. [00:32:34] Asmita Mahajan: Now, we'll be creating training and test lost list, so that for each iteration, we can capture the training and test losses, to further plot the curves for the training and test losses for each experiment. [00:32:50] Asmita Mahajan: iterating over the epox, so we have 50 epochs, we have, fixed it to 50, and we'll train our model [00:33:00] Asmita Mahajan: Or we'll, we'll train for 50 times, so for every training process will be done, for 50 times. [00:33:10] Asmita Mahajan: So here, we are opening our model, in the training mode. [00:33:15] Asmita Mahajan: Using this train function. [00:33:18] Asmita Mahajan: Then we, we'll be initializing the gradients, the weights, for the model, and also to clear the gradients from the previous iteration, we'll be using this zero gradients, so that, the gradients are… doesn't leak [00:33:34] Asmita Mahajan: from the previous iteration to the new iteration. So that's why we are initializing the gradients to 0. [00:33:40] Asmita Mahajan: Now, we are capturing the output from the model. [00:33:45] Asmita Mahajan: into lodges, so we are calling a model using, extrane Tensor as an input. So, extrane Tensor contains the training dataset, and we are passing it to the model function, so the dataset, or the data points. [00:34:01] Asmita Mahajan: propagate through the network, and the output is calculated or evaluat- calculated, and that output is now captured in Lodges. [00:34:13] Asmita Mahajan: Binary cross entropy loss is used to calculate the loss. So, to that function, we are passing the lodges that we captured, and the true labels of the training dataset. [00:34:26] Asmita Mahajan: And we are evaluating the loss. This loss that we calculated in the previous step is now propagated backward to the network using backpropagation, and that will be, [00:34:39] Asmita Mahajan: Implemented using this backward function. [00:34:42] Asmita Mahajan: So in, so we'll be, [00:34:45] Asmita Mahajan: optimizing our weights using the Atom optimizer, and here, it, this, optimizer is called so that the weights are updated. [00:34:56] Asmita Mahajan: And after that, we'll be appending the train laws. [00:35:02] Asmita Mahajan: To the, to the train loss list. [00:35:05] Asmita Mahajan: Using append function, and the loss that we calculated in the previous step, we are converting it into float, and then appending to the train loss for this particular iteration. [00:35:16] Asmita Mahajan: Similarly, for this particular iteration, we'll also be calculating our test loss, so that we can see how our model is converging at each iteration. [00:35:26] Asmita Mahajan: So for that, we are opening our model in evaluation mode. Now, we don't want any gradients. We don't want to update gradients, or we don't want to update the weights, because we'll be using the, [00:35:38] Asmita Mahajan: Learned weights from the training. [00:35:41] Asmita Mahajan: And using the learn weights, we will be test… using those weights to test the dataset. [00:35:47] Asmita Mahajan: So that's why, torch.noGradient is used, that we don't want any gradients, or we don't want to update the gradients, when we are passing the test dataset to the model. [00:35:59] Asmita Mahajan: So with this, we'll be, propagating the test dataset in the network, and then capturing the outputs in the test logist. [00:36:11] Asmita Mahajan: And in a similar manner, we'll be calculating the test loss by giving test logist and the true labels from the test dataset to the criterion function, and converting into new float, and then capturing it in test loss variable. And that loss is then appended to the list. [00:36:31] Asmita Mahajan: For that particular iteration. [00:36:34] Asmita Mahajan: So, in this manner, our training will be done. [00:36:37] Asmita Mahajan: Using 50 epochs. So after the training has been done, now we will, use the network and use the, updated weights, the trained weights from the network, to, [00:36:51] Asmita Mahajan: Calculate the predictions. [00:36:54] Asmita Mahajan: So, we'll be again, opening our model in evaluation mode with, no, weights, and now we'll be calculating, we'll be, passing the X tensor to the model, capturing the outputs, raw outputs. [00:37:10] Asmita Mahajan: Now we'll… we are using a sigmoid activation function, and [00:37:15] Asmita Mahajan: Converting the raw outputs into probabilities, and using a threshold, we are converting the probabilities into the hard predictions. [00:37:25] Asmita Mahajan: So in this manner, we have calculated our predictions. [00:37:30] Asmita Mahajan: Now, to use or to evaluate the scores, or the matrices, the performance matrices, we will convert [00:37:38] Asmita Mahajan: all our tensors into NumPy array. [00:37:42] Asmita Mahajan: So, as, in the previous week, I told you, when we are using GPU, so to bring back the dataset from GPU to CPU, we are… we use this CPU command, or the CPU function. So, it just brings back our data, all the data from the GPU to the CPU. [00:37:58] Asmita Mahajan: After bringing back the data to the CPU, we are converting it into NumPy array, and then we are flattening the dataset from 2D array, or ND array to 1D array. [00:38:09] Asmita Mahajan: And in this manner, the tensors are converted into NumPy array. So we'll have Y2 labels, Y predicted labels, and Y probabilities. [00:38:21] Asmita Mahajan: So this matrix will store all the performance matrices, the accuracy, precision score, recall, F1 score, and ROC curve. History dictionary will store the training and test losses from every iteration, and for each, [00:38:36] Asmita Mahajan: Experiment, and this function will return the matrix and history dictionaries. [00:38:42] Asmita Mahajan: So… [00:38:44] Asmita Mahajan: Here, we have created the learning rate list, which will contain… which will have 5 elements, and 5 learning rate values. So, we'll be calling the train and evaluation function 5 times. [00:38:56] Asmita Mahajan: And result dictionary will store. [00:38:59] Asmita Mahajan: For every… so, this… this value will be the key to the result dictionary. [00:39:06] Asmita Mahajan: And for every learning rate, we'll be storing the metrics. [00:39:11] Asmita Mahajan: And history will store for every learning grade the train and testing laws. [00:39:16] Asmita Mahajan: So, this history dictionary will store that, that. [00:39:20] Asmita Mahajan: So now we are iterating over the learning rates from this list, and we are calling the train and evaluation function on that learning rate with 300 epochs. [00:39:32] Asmita Mahajan: And, just capturing the return matrices and history in, metrics and history variable. [00:39:39] Asmita Mahajan: And, in results for the learning rate. [00:39:44] Asmita Mahajan: we are, storing the matrix for that learning rate, for that experiment, which is iterated over here. And histories, dictionary will store the history, which is the train and test loss. So now, we'll, using histories, we'll plot the curves for the train and test loss for each learning rate. [00:40:03] Asmita Mahajan: So, we are here, defining a function. [00:40:09] Asmita Mahajan: We'll pass the histories dictionary to that. The maximum columns is 3, and the subtitle to that whole plot would be Train Test Loss Per Learning Rate. [00:40:22] Asmita Mahajan: So, histories.keys will have… [00:40:25] Asmita Mahajan: these keys, which are the learning grades. So, these keys are extracted from history's dictionary and are stored in, LRS. So, this, LRS is a list of history's, keys. [00:40:44] Asmita Mahajan: N will be the value of the… the length of the keys, so we have 5 learning rates, so N will contain 5. [00:40:55] Asmita Mahajan: Now, we are calculating the number of rows and columns for the grid, for the subplots, and how we want to display the subplots. So, that grid, will have [00:41:06] Asmita Mahajan: columns and rows, so we are calculating it using, this minimum of the max columns, and mini… so minimum of max columns and, this N would be 3, so columns will be 3. [00:41:19] Asmita Mahajan: Ish? [00:41:20] Asmita Mahajan: And rose is the mat ceiling, and 5x3, so it is ceiling, so it will store 2. So, we have… [00:41:29] Asmita Mahajan: 2 in the rows. [00:41:31] Asmita Mahajan: So, there will be 2 rows and 3 columns. [00:41:34] Asmita Mahajan: Now, using subplots function, we'll be, creating our subplots, and axis will contain the location for each subplot, and figure will be the overall figure. Rows will be 2, columns are 3, this is the figure size. [00:41:49] Asmita Mahajan: So, while iterating over this, list, which contains the learning rates, the name of the learning rates, so we will, we will get the index and the learning rate name. [00:42:05] Asmita Mahajan: And we'll be calculating the rows and columns. So, this RC. So, in the previous… [00:42:11] Asmita Mahajan: code. So, these are the different methods that you can use to flatten your axis. So, previously, what we did, we flattened it using the RABL function. But here, what we are doing is, we are just directly extracting, [00:42:25] Asmita Mahajan: we are just directly, like, calling RNC on the NDA and directly getting the value or the location from [00:42:34] Asmita Mahajan: the NDA into AX. So, using this RNC, we are just, we… [00:42:40] Asmita Mahajan: like, directly getting the location. But, previously, what we did was we flattened the… we flattened our NDARA into 1D array and iterated over the 1D array. So these are different methods to get the location for the subplots. [00:43:00] Asmita Mahajan: And history. So here, what we are, getting in HIST is the, dictionary for that particular learning rate. And that dictionary will contain the training losses and testing losses for, that particular learning rate. So history… HIST is a dictionary here. [00:43:18] Asmita Mahajan: In that dictionary, what we want to plot, we want to plot the train loss curve and the test loss curve. So, what we are doing is, in the first plot, we are telling that from this hist… [00:43:30] Asmita Mahajan: dictionary, I just want the train loss. So, I just want to plot the train loss curve, and I'm labeling that, train loss curve as [00:43:39] Asmita Mahajan: train loss, and in the second plot, I am plotting the test loss. So, from the history dictionary, I am extracting test loss values, and I am plotting them and labeling that as test loss. [00:43:51] Asmita Mahajan: the title will be the, value of the learning rate. So here, LR is equals to 0.001. This is the value of our first learning rate, and in the same manner, there will be 5 plots. So, 0.003, [00:44:06] Asmita Mahajan: We have… and we have a subplot for 0.001, we have a subplot for 0.003, and a subplot for 0.01 for different values of learning rate. [00:44:18] Asmita Mahajan: So, title… so we have set the title for each subplot. [00:44:22] Asmita Mahajan: This will set the X label, so X label will contain the epoch, and Y label will contain the loss. [00:44:30] Asmita Mahajan: and legend to, differentiate between the train loss and test loss. So again, as we have 3 columns and 2 rows, we'll be… so there will be position for 6, subplots, but we don't want 6 subplots. [00:44:46] Asmita Mahajan: We want… we have the values for only 5, so we'll just… we don't want to display the 6 subplots, that's why we are telling it to be off. [00:44:56] Asmita Mahajan: We'll, this, code will set the, title for the whole figure. [00:45:03] Asmita Mahajan: And… The plots will be displayed. [00:45:09] Durga Toshniwal: So, I think, once again, you can, have a look at the plots, and then, to see, [00:45:16] Durga Toshniwal: That which learning rate, is most suited, and yeah, I think so. [00:45:23] Durga Toshniwal: There… there must be a tabular, if you just scroll down a little bit… [00:45:28] Durga Toshniwal: It will also show the… Values. [00:45:33] Asmita Mahajan: Ma'am, it is shown in the other plot, our plot. [00:45:35] Durga Toshniwal: in the next block. So, once again, you can have a look at the different, [00:45:42] Durga Toshniwal: Values, you can just, make it a little smaller, because the ticks are getting cut a little on the lower, [00:45:50] Asmita Mahajan: Okay, ma'am. Lower figures. [00:45:53] Durga Toshniwal: Let's scroll up a little, please. [00:45:56] Asmita Mahajan: Is it… is the booth. [00:45:57] Durga Toshniwal: Yeah, yeah, I think it's okay, no. [00:46:00] Durga Toshniwal: So now here, when you look at, the… so in the previous code, we did choose two learning, rates, that is 0.01 and 0.001. [00:46:12] Durga Toshniwal: But at that time, the objective was… [00:46:15] Durga Toshniwal: Primarily, to decide the optimizer, that was ADAM and STD, and we decided on ADAM. Obviously, we took some combinations of learning rate also. [00:46:25] Durga Toshniwal: Now we are finalizing on… learning rate. So, actually, [00:46:30] Durga Toshniwal: The thing is that one might, just wonder that why we are picking up one hyperparameter at a time. [00:46:38] Durga Toshniwal: So, ideally, all the hyperparameter [00:46:43] Durga Toshniwal: that we want to, iterate on actually should be chosen all together, and the different combinations should be tried. That is the best way to do. However, in collab, it is not possible, and we will need GPU support. [00:46:59] Durga Toshniwal: To do all that. [00:47:00] Durga Toshniwal: Right, so therefore, what we are trying to show here is that at a time, we just try to choose one while keeping others fixed. [00:47:09] Durga Toshniwal: To some extent. For example, in the earlier case, we took two options for the optimizers, like Adam and SGD, and obviously with that, we had to choose some learning rate, so we chose that. [00:47:25] Durga Toshniwal: So, at that time, we… the learning rate was just taken based on, you know, some standard choices in the previous code. However, here, we are finalizing on the learning rate, so now we have fixed on Adam. [00:47:41] Durga Toshniwal: So, ideally, we could have taken the optimizer set, the different learning rates, like what you are seeing here, 0.0013, like that, all that. [00:47:51] Durga Toshniwal: then you might see the number of layers to be taken, number of neurons per layer, then activation for all these combinations together, and then, you know, experiment with them. However, you can note that the combinations will be huge, and the search space will be very large. [00:48:11] Durga Toshniwal: For the most optimal combinations. [00:48:14] Durga Toshniwal: And therefore, it will be very, very compute-intensive, and it won't be possible to show it here. [00:48:20] Durga Toshniwal: And so take… so to make things more… [00:48:23] Durga Toshniwal: Simpler, and a little less compute-intensive, that's how we are showing it. [00:48:28] Durga Toshniwal: Okay, so now let's look at the different learning rates, and Adam is fixed so far. [00:48:35] Durga Toshniwal: as the optimizer. Again, you can see here. [00:48:39] Durga Toshniwal: That, the train loss, with learning weight of 0.0001 and 0003, [00:48:47] Durga Toshniwal: In both these cases, the… [00:48:51] Durga Toshniwal: The train lost is actually losses [00:48:54] Durga Toshniwal: Quite low, and the test, loss is high. [00:48:59] Durga Toshniwal: Especially with .003, it's very, very noticeable. [00:49:03] Durga Toshniwal: Which, definitely means that the model is overfitted to the training data. [00:49:08] Durga Toshniwal: And so we will not choose these learning rates. [00:49:12] Durga Toshniwal: Now, looking at .001.003, and then, 0.01. [00:49:21] Durga Toshniwal: So, the first three combinations at .001… I mean, 00010003. [00:49:29] Durga Toshniwal: You can see that these learning rates are really very low also. [00:49:33] Durga Toshniwal: Now, you look at .001 and 0.003. Again, you see here, definitely, the model is not that very overfitted in these two cases. [00:49:43] Durga Toshniwal: Now the point to see is where the model… how fast the model is converging. [00:49:49] Durga Toshniwal: so… So what we want is that the model should also converge quite fast. [00:49:56] Durga Toshniwal: So, although 0.001, 0.003, and 0.01 [00:50:01] Durga Toshniwal: All of them do have, [00:50:05] Durga Toshniwal: They are… they are not overfitted, in the sense that the… [00:50:09] Durga Toshniwal: The training loss is not much lower than the test loss. [00:50:15] Durga Toshniwal: So, they are not very, very overfitted, definitely. However. [00:50:19] Durga Toshniwal: You can see that the convergence in the case of 0.001 [00:50:25] Durga Toshniwal: It comes at around 200 in case of 0.003. It's coming at around, 175. [00:50:36] Durga Toshniwal: And if you look at 0.01, [00:50:39] Durga Toshniwal: It's at around 75 or something like that, where you see that the two curves start separating out. [00:50:46] Durga Toshniwal: So, again, we can conclude that learning rates [00:50:51] Durga Toshniwal: Of 0.01 is still coming out to be the most optimal out of the 5 combinations we tried. [00:50:59] Durga Toshniwal: And this further can be substantiated, as we did earlier. Here, also, we can do that we can actually look at the different performance metrics, like accuracy, precision, recall, etc. [00:51:11] Durga Toshniwal: With all these learning rates, and then see. [00:51:14] Durga Toshniwal: which learning rate gives us the best result, which is what you are going to see in the next section. And that will help us to [00:51:22] Durga Toshniwal: To actually verify the results further. [00:51:26] Durga Toshniwal: So I think, Asmita, you can continue. [00:51:31] Asmita Mahajan: Yes, ma'am. [00:51:33] Asmita Mahajan: So, in these plots, we, looked at the train and test losses, and compared the train and test loss curve for each learning rate. [00:51:44] Asmita Mahajan: Now, we will, [00:51:49] Asmita Mahajan: Yeah, so now, we will plot the bar plots of the different matrices that we evaluated, the accuracy precision for each learning rate, and we'll plot that. So, here, what we are doing is, we are just extracting, [00:52:09] Asmita Mahajan: So, we haven't converted our, results. So, this result dictionary, we have… we haven't converted it to a data frame. So, what we are doing is, we are just, converting it, this dictionary. [00:52:24] Asmita Mahajan: to the… to a data frame. So, this is, another way of converting a dictionary into a data frame. So what we are doing is, from pandas, we are calling the function data frame, which will be… which will convert the data. It has been passed to a data frame. [00:52:40] Asmita Mahajan: So, what we are doing is, we are just, creating one key and telling that the LR would… will be the, value. [00:52:49] Asmita Mahajan: at this LR key, and this LR is, how we are getting the value at LR is, we are, using this for loop, iterating over the learning rates. [00:53:00] Vidushi Sidana: So, LR is a variable which will have the names and… [00:53:05] Asmita Mahajan: Of all the learning grades. [00:53:06] Vidushi Sidana: Deluxe. [00:53:07] Asmita Mahajan: And that will be passed here. [00:53:09] Asmita Mahajan: and all the other values in this dictionary will be extracted from the result dictionary at that LR. [00:53:21] Asmita Mahajan: And, we are setting… so the index will be, after the data frame has been created, we are setting the index. [00:53:30] Asmita Mahajan: to be the LR column. So, the index will, will have the values for each LR. [00:53:37] Asmita Mahajan: And, the columns will be now the matrix, the accuracy, precision, recall, and F1 score. [00:53:45] Asmita Mahajan: So, we… in this manner, we have created the data frame. [00:53:49] Asmita Mahajan: Now, we… using this data frame, we'll be plotting our, bars. We'll be plotting, the bar plots. So, similarly, we have, 2 number of rows, 3 number of columns, and the title, the main title of the figure, will be matrix versus learning rate. [00:54:10] Asmita Mahajan: So this matrix list will have the name of all the matrices, performance matrices, And, LR… [00:54:19] Asmita Mahajan: Will… is a list of all the learning grades, the 5 learning grades. [00:54:25] Asmita Mahajan: So, how we are getting that? From index in the data frame. So, this dotindex will, convert the indexes [00:54:36] Asmita Mahajan: And, from, we'll extract the indexes from the data frame and convert it into a list. [00:54:42] Asmita Mahajan: And before passing it to LR, or before passing it to the list, we are converting the values to string data type first. [00:54:51] Asmita Mahajan: Now, using subplots… sorry. Using subplots function, we'll be creating the figure and subplots. [00:54:59] Asmita Mahajan: Our subplots. So the grid will be, will have 2 rows and 3 columns. [00:55:06] Asmita Mahajan: This, this will tell the figure size. [00:55:12] Asmita Mahajan: We'll flatter. [00:55:17] Asmita Mahajan: With… yeah. [00:55:18] Asmita Mahajan: So, we'll flatten the access, array, ND array, into 1D array, and, we'll convert it using RABL, and store it in access list. [00:55:31] Asmita Mahajan: iterating over the matrices. So I will have the index, and metric will have the metric name. [00:55:38] Asmita Mahajan: And, for each location, so this will extract the first location. So at first location, we have the accuracy subplot. [00:55:45] Asmita Mahajan: So, this will extract the first location from the access list. [00:55:52] Asmita Mahajan: So, in values, we'll get the values from the, data frame, and what values we want for this first subplot, the accuracy values. So, the metric will have accuracy name. [00:56:04] Asmita Mahajan: And, from that data frame, we'll, we'll extract the values, and it will be stored in, values column, here, values variable. [00:56:13] Asmita Mahajan: So, to the bar plot, we are, so on the x-axis, we want the different learning rates, and on the y-axis, we want the values for those learning rates, the accuracy values for those learning rates. So, we are passing these values to the bar function, and it will create the bar plots. [00:56:33] Asmita Mahajan: And the title will be set, so, the accuracy, precision, and recall. So, for each subplot, the title is the metric, the X label is the learning rate, Y label is, again, the metric, and [00:56:46] Asmita Mahajan: Similar, like, previously, as we set the limit to be from 0 to 1, we'll again set it so that the plot is presentable, and all the bars are inside the plot. [00:57:00] Asmita Mahajan: So that's why we are setting the limit from 0 to 1. [00:57:05] Asmita Mahajan: And also, these, tick labels, to display the tick labels at a 45 degree angle, so we are using these tick parameters, and we are telling it… [00:57:18] Asmita Mahajan: that. [00:57:19] Vidushi Sidana: dividend debit? [00:57:20] Asmita Mahajan: X axis, we want it to be at a rotation of 45. [00:57:29] Asmita Mahajan: So, again, we want, for each bar, we want the value associated with each bar to also be displayed, so we'll be using it using the text function, so we are enumerating over the values. [00:57:45] Asmita Mahajan: Variable, which is the list, and we are just getting the index and the value at that index, passing it to the… [00:57:54] Vidushi Sidana: Oh, weird. [00:57:55] Asmita Mahajan: Oh. [00:57:56] Vidushi Sidana: Thanks for. [00:57:56] Asmita Mahajan: function. [00:57:57] Asmita Mahajan: Where, these two values will tell. [00:58:00] Laxmi SAHU: Sorry, Asmita, to interrupt. I'm so sorry. Vidhushi, can you please, mute yourself? We are getting disturbed, we're not able to understand what Asmita is saying. [00:58:11] Vidushi Sidana: Listen. [00:58:13] Laxmi SAHU: Vidushi. [00:58:15] Vidushi Sidana: I don't… [00:58:16] Laxmi SAHU: Thank you. [00:58:19] Laxmi SAHU: Oh, please go and ask me the… [00:58:25] Vidushi Sidana: Actually, it's a billionaire. [00:58:27] Vidushi Sidana: It bubbled last week. [00:58:28] Deepak Bobade: I think it's accident. [00:58:31] Deepak Bobade: on and off, probably. You might not even know. Yeah. [00:58:36] Shikhar Gupta: She should leave and join me instead. [00:58:39] Deepak Bobade: Yeah, probably. [00:58:40] Vidushi Sidana: Found that photo. [00:58:42] Deepak Bobade: Vedushi, hello? [00:58:47] Deepak Bobade: She doesn't know the way the mic is turning on and off. [00:58:52] Vidushi Sidana: Aye. [00:58:54] Durga Toshniwal: I think the admin can mute her. [00:58:59] Deepak Bobade: You can remove, actually. Probably muting is not the solution, because it's turning on and off. [00:59:05] Durga Toshniwal: No, I think it's turning on and off from her end, probably if the admin, I mean… [00:59:12] Durga Toshniwal: If it is muted from here… Sindran, are you there? [00:59:19] Vidushi Sidana: Oh, but I'll pass. [00:59:22] Durga Toshniwal: Either she can… she can be made to exit and join again, or something like that, or… [00:59:29] Durga Toshniwal: She can be muted. I wonder… [00:59:31] Vidushi Sidana: Asmita, can you please hold on? I think we need to sort this out. [00:59:36] Asmita Mahajan: M. [00:59:39] Vidushi Sidana: Can you speak to that moment? [00:59:43] Deepak Bobade: Ma'am, maybe you have the privileges. [00:59:46] Durga Toshniwal: Yeah, I'm just going to try that only. I'm actually not the host, but [00:59:53] Durga Toshniwal: Where shall I see this? [00:59:55] Vidushi Sidana: Don't tell. [00:59:56] Durga Toshniwal: Probably, I may be having it. [01:00:05] Durga Toshniwal: I think… let me see. The participants, probably, I can mute the participants from my end. [01:00:16] Durga Toshniwal: No, I cannot mute, I just saw… [01:00:18] Vidushi Sidana: It's a photo! [01:00:21] Deepak Bobade: Could you remove the person? [01:00:24] Durga Toshniwal: That will be a little difficult for me to do, right? I… I shouldn't be removing, I don't know. [01:00:31] Deepak Bobade: Anyway, she is not… she's not, I think, having the attention over here. [01:00:36] Shikhar Gupta: Cool. [01:00:37] Durga Toshniwal: that's. [01:00:37] Deepak Bobade: Otherwise, she would have responded, right? [01:00:39] Durga Toshniwal: Yeah. [01:00:41] Durga Toshniwal: Where has she gone? [01:00:46] Durga Toshniwal: I can't see the gnome. [01:00:50] Durga Toshniwal: Petushi. [01:00:52] Laxmi SAHU: Oh, she's there, ma'am. [01:00:55] Durga Toshniwal: Yeah. [01:00:55] Laxmi SAHU: muted now. [01:00:59] Durga Toshniwal: Actually, she got muted, and then she automatically gets unmuted, that's the problem. If she's muted, then it's okay, I think we can continue. Let's see if she gets unmuted again, then we'll see. [01:01:10] Durga Toshniwal: So… Asmita, you can carry on now, let's see. [01:01:15] Asmita Mahajan: Okay, ma'am. [01:01:17] Asmita Mahajan: So, [01:01:18] Asmita Mahajan: I think I was here, yeah. So, we are just, we just want that values, to be displayed associated with each bar. So, using text function, we'll be doing that. The first two values will be telling the location, of that text, where that text should be. So, X will be the index. [01:01:38] Asmita Mahajan: So X would be, like, 0, 1, 2, so 0 at 0… [01:01:43] Asmita Mahajan: X, so at 0th location, and Y location will tell us, so we'll be getting the value from V, and just adding a 0.01 to that, so that the values or the text doesn't appear on the line of the bar. [01:02:00] Asmita Mahajan: So, for that, we are adding 0.01. [01:02:02] Asmita Mahajan: And, this will just tell us that the values we want are… [01:02:08] Asmita Mahajan: We want them to be displayed till, two decimal places. [01:02:13] Asmita Mahajan: The horizontal alignment, the vertical alignment will be set to bottom, the horizontal alignment is set to center, and the font size is 9 for the text. [01:02:23] Asmita Mahajan: So in this manner, we'll be displaying the text associated with, or sorry, the values associated with each bar. [01:02:32] Asmita Mahajan: So, to just not… so that, this sixth subplot, the empty subplot is not displayed, we are using this, [01:02:45] Asmita Mahajan: sorry, function. So what it is doing is, it is iterating over the range, like, so length will tell, length of the matrix will tell that our fifth, at fifth location. [01:03:00] Asmita Mahajan: Right? And at… Row 2, and column, 3. [01:03:06] Asmita Mahajan: So, this will tell us the location. So, for that, J position, I want my axis to be off. [01:03:17] Asmita Mahajan: And now we are labeling the title, the main title of the figure, and displaying the plot, and also displaying the data frame after that. [01:03:30] Asmita Mahajan: So these are the plots in which we can see that, learning rate 0.01 [01:03:42] Asmita Mahajan: Yes, so this learning rate in each subplot [01:03:45] Asmita Mahajan: Is, is having the best matrices, or is performing… the model is performing well at, the learning rate, 0.01. [01:03:55] Durga Toshniwal: I think you can scroll down to the, yeah, these results. So, it becomes very much evident and much more clear, as you can see the, results in a tabular fashion. [01:04:08] Durga Toshniwal: And it is, you can see that with 0.01, accuracy is 0. [01:04:15] Durga Toshniwal: 0.859, precision, again, higher, 0.87. [01:04:20] Durga Toshniwal: Recall 0.80, F1.83, AUC 0.94. So the model is quite good at .01. [01:04:29] Durga Toshniwal: So that will be our choice. [01:04:32] Durga Toshniwal: That's what, I think it has shown. [01:04:35] Durga Toshniwal: So, we are done with two things. One is the optimizer, one is with the learning rate. Now, we'll take on the third one, which is the… [01:04:44] Durga Toshniwal: Size of the first redundant. [01:04:47] Durga Toshniwal: Yeah, please continue. [01:04:49] Asmita Mahajan: So, now we'll be optimizing the first hidden layer, the size of the first hidden layer, so we'll be optimizing how many neurons will be there in this hidden layer. [01:04:59] Asmita Mahajan: And we'll be fixing our optimizer and the learning rate, which we have… which we tuned, like, in the previous code. So Adam will be our optimizer. [01:05:11] Asmita Mahajan: Will be the choice of optimizer we'll be using, and 0.01 will be the learning rate that we'll be using to update the weights. [01:05:20] Asmita Mahajan: So, for the sizes, we'll be, taking 32, 64, 128, and 256. These values, on these values, we'll be evaluating our model. [01:05:32] Asmita Mahajan: So, [01:05:34] Asmita Mahajan: Now, as here we have passed it as a parameter, but when we are calling this function, or this… when we are creating the instance of this class, at that time, the… [01:05:46] Asmita Mahajan: a number which will be passed will be used. So this number will not be used. This number will only… this number will only tell us if the hidden… if for this parameter, the number is not passed. [01:05:58] Asmita Mahajan: So, what default value should be taken? So. [01:06:01] Asmita Mahajan: Whenever we are, defining our function, so at that time. [01:06:06] Asmita Mahajan: Sometimes we give the value, we give the default value for the parameter, so it will only be used if we have forgotten, or if we have not given any value at the time of, calling the function. So this 64 hidden… 64 number will only be used when, at the time of calling, we haven't defined, or we haven't passed. [01:06:27] Asmita Mahajan: the value. [01:06:28] Asmita Mahajan: Otherwise, the value that has been passed at the time of calling will be used. So this 64 will be, will be set as a default value. [01:06:37] Asmita Mahajan: So, we are creating the MLP class, and in the MLP class, we are creating an initialization instructor, passing the input dimension, or the input size, and the hidden layer size. [01:06:51] Asmita Mahajan: So, again, [01:06:52] Asmita Mahajan: very similarly… it's very similar to the previous code that we did. We are creating the network in a sequential manner, creating the first hidden layer using the radioactivation function. [01:07:04] Asmita Mahajan: Sorry, creating the first layer, and then the input layer, the hidden layer, and then from the hidden layer to output layer, we are creating this. [01:07:13] Asmita Mahajan: fully connected layers. [01:07:17] Asmita Mahajan: So, this function will be used to propagate the inputs in the forward pass through the network. So, this forward function will pass the inputs X [01:07:26] Asmita Mahajan: to the network, which is using the net function, and in the net function, we are passing the X. [01:07:36] Asmita Mahajan: So, the learning rate, we, as it is fixed, so we are just, fixing it here using best LR. So, in the variable best LR, we are passing 0.01 as the learning rate, because we had already, optimized it. [01:07:52] Asmita Mahajan: Now we are creating the train and evaluation function. So here, in this, in this code, we'll be optimizing, we'll… our focus will be to optimize the hidden layer. So, hidden will tell us the size of the hidden layer, and it will be passed at the time of calling of this train and evaluation function. [01:08:10] Asmita Mahajan: EPOCs are 200, so epochs are defined using the default value 200, and LR is fixed, and it is the, it is 0.01 best LR. [01:08:22] Asmita Mahajan: So now we are just creating an object of the class MLP, passing the shape of the, extrane tensor, or passing the input size, which is 14. [01:08:34] Asmita Mahajan: And the hidden layer size, which is hidden. So hidden will have the value, which will be passed, using the calling, which is passed at the time of calling of this function. [01:08:48] Asmita Mahajan: So we're using, binary cross entropy, loss function as a criteria to evaluate losses. [01:08:57] Asmita Mahajan: Atom Optimizer is used to optimize the learning rates. [01:09:03] Asmita Mahajan: Using the LR, which is the best LR, which is 0.01. [01:09:09] Asmita Mahajan: Train and test loss list will be… are initialized here, which will store train and test losses for each epoch, after every. [01:09:18] Asmita Mahajan: Train, process… after, after training the model. [01:09:23] Asmita Mahajan: So, we're iterating over the epoch, which is 200. [01:09:27] Asmita Mahajan: We'll open the model in training mode using train command or train function, initializing the gradients to zero using zero gradient function. [01:09:38] Asmita Mahajan: So that, the gradients from the previous iteration are not passed on to the next iteration. [01:09:46] Asmita Mahajan: Now, calculating the outputs using the model, so, here we are passing the training dataset to the model, and after propagating the dataset. [01:09:59] Asmita Mahajan: throughout, through the whole network, after the forward passed, the outputs will be calculated, and to capture the outputs, we are capturing it in logist variable. [01:10:12] Asmita Mahajan: We are calculating the training loss for this iteration using binary cross entropy loss function, and passing the raw outputs and the [01:10:24] Asmita Mahajan: Y train labels, the actual training labels to the function, and calculating the loss, propagating the loss backward using a backward function, optimizing the… or updating the weights using optimizer.step function, and then the calculated loss [01:10:43] Asmita Mahajan: Is now appended to the train laws. [01:10:46] Asmita Mahajan: list for this particular iteration. And in a similar manner, we will calculate the test loss for this particular iteration by opening our model in evaluation mode. We will not use any gradients because the trained gradients or the trained weights will be used to evaluate the test loss. [01:11:06] Asmita Mahajan: So, the model which has already been trained will be used to, evaluate the output, and we'll be passing the test, X test dataset. [01:11:18] Asmita Mahajan: Capturing the outputs for the test dataset in test logist, calculating the loss using the BCE binary cross entropy loss, and then appending the particular loss for this particular iteration to the test losses list. [01:11:39] Asmita Mahajan: Now, we have trained our model, for, this particular, hidden, value of hidden layer. So, we have, trained it, and now this trained model will be used to, predict, make the hard predictions for the model. [01:11:56] Asmita Mahajan: So, opening our model in evaluation mode, with no gradients, because we want the train model to be used on the test set, passing the test set, and capturing the outputs and logist. [01:12:08] Asmita Mahajan: Using a sigmoid activation function, we are converting the raw outputs into probabilities, and using the threshold as 0.05, we are converting our… or we are just [01:12:20] Asmita Mahajan: Pass, like, using this threshold, we are converting the probabilities, into hard predictions. [01:12:30] Asmita Mahajan: now we need NumPy arrays, so the tensors that we have created, the PyTorch tensors, are now… have to be converted to NumPy arrays so that we can use the evaluation, matrices, these, [01:12:45] Asmita Mahajan: functions. So, these are sklearn functions, and sklearn function uses NumPyArray, so that's why we have to convert the tensor into NumPyArray. So, I'm bringing back the data set from GPU to CPU, converting the numpy… converting the tensor to NumPyArray, and flattening the [01:13:05] Asmita Mahajan: tensors, or the NumPy array into 1D array, and now storing it in by true, list. [01:13:14] Asmita Mahajan: So in this manner, we will create Y true labels, Y predicted labels, and Y probabilities, which were calculated here. [01:13:26] Asmita Mahajan: Now we'll create a matrix, dictionary, which will store all the performance matrices, which is accuracy score, precision, recall F1 score, and AUC score. History dictionary will, [01:13:40] Asmita Mahajan: Just store the training and test losses for each iteration, and this function will return the matrix dictionary and the history dictionary. [01:13:51] Asmita Mahajan: Now, here, we have defined the hidden sizes, so we'll be experimenting, with 4, hidden layer sizes. [01:14:01] Asmita Mahajan: So, firstly, we'll use 32 neurons in the hidden layer, then we'll use 64 neurons, then 128, and then 256. And for each size, we will see how the model will perform, with Adam as an optimizer and learning rate as 0.01, which is fixed. [01:14:20] Asmita Mahajan: So, results and histories, these are the overall dictionaries, which will store the values of the… results will store the values of the matrices for each size of the hidden layer, and histories will also store the values of the train and test slots for each iteration. [01:14:39] Asmita Mahajan: For this particular value of the hidden layer. [01:14:45] Asmita Mahajan: So, we are calling, we are iterating over these values, and we are just passing it to the train and, evaluation function. [01:14:57] Asmita Mahajan: And here, we'll pass the size of the hidden layer. Hidden is equals to H, and H will be, the variable which will store these values. [01:15:08] Asmita Mahajan: So, iterating over these values, we'll be, training the model, and then, this, function will return the matrix and history for that particular hero layer. Using, using these dictionaries, we'll create an overall dictionary, which will have [01:15:27] Asmita Mahajan: key as the size of the hidden layer, and the metrics for the calculated performance matrices for that particular size of the hidden layer. And histories dictionary will store the training and test losses for that particular hidden layer. [01:15:45] Asmita Mahajan: Now, again, as we… [01:15:48] Asmita Mahajan: previously converted our dictionary into a data frame. So, in similar manner, we are converting the results, dictionary into a data frame using this, [01:16:02] Asmita Mahajan: fun, using this line of code. So what we are doing is, we are iterating over the result dictionary. [01:16:08] Asmita Mahajan: For each hidden size, we'll be storing the results dictionary there, and also creating one key-value pair, which will tell us, what, like, what is the particular value of the hidden. [01:16:25] Asmita Mahajan: size, or, sorry, I will reframe it. So, this, [01:16:33] Asmita Mahajan: A new key-value pair will tell us that for this particular hidden layer, the results, are the… [01:16:42] Asmita Mahajan: accuracy, precision, F1 score, and, recall values are for that particular hidden size. And we are, setting the index as the size, like, this, [01:16:57] Asmita Mahajan: So, hidden 1 will have value 32, or it will have 64, 128, 256, so this will be set as the index, and the columns will be the matrices names. Accuracy, precision, recall F1 score, and AUC curve… AUC score. [01:17:15] Asmita Mahajan: Now, we will plot the test and training laws. So, we are defining this function. [01:17:24] Asmita Mahajan: So, histories dictionary will be used to plot the test and training laws, because history dictionary stores the test and training laws for each [01:17:34] Asmita Mahajan: Okay, so I have not run this code, I guess. [01:17:38] Asmita Mahajan: I've run it, okay. So, history will store, the, test and training, training loss for every iteration, and for each, hidden layer size. Maximum column is 3. The overall title will be train and test loss versus first hidden layer size. [01:17:57] Asmita Mahajan: Keys will store the… Keys from the histories dictionary, which are the hidden sizes, or the hidden layer sizes. [01:18:06] Asmita Mahajan: N is the length of the keys, so we have 4… we have defined 4 hidden layer size, 32, 64, 128, and 256, so N will store 4. [01:18:15] Asmita Mahajan: And we'll have the value 4. Columns is calculated, using this minimum function, so columns will be 3. [01:18:24] Asmita Mahajan: And rows will be? [01:18:27] Asmita Mahajan: Two rows will be there. [01:18:37] Asmita Mahajan: Okay, so now, using subplot function, we'll be plotting, the, we'll be plotting for each hidden [01:18:46] Asmita Mahajan: layer size. We'll be plotting test and training laws for each hidden layer size. So, rows will have value 2, column will have value 3. [01:19:00] Asmita Mahajan: And figure will store the overall figure, and access will store the location of each subplot. Enumerating over the keys list, we are getting the index and key from the keys, so keys will have value from as 32, [01:19:17] Asmita Mahajan: 64. [01:19:20] Asmita Mahajan: 128 and 256. [01:19:23] Asmita Mahajan: So, these are the key values, and index will be 0, 1, 2, 3, 4. So… [01:19:29] Asmita Mahajan: From the axis, we are, getting the location for each subplot. [01:19:36] Asmita Mahajan: his history is K. So, histories for 32, hidden layer will, so from here, we'll be getting the training and test losses for this particular hidden layer size, and HIST will, store that. So, from this HIST dictionary, we are extracting the train loss first, and plotting it. [01:19:56] Asmita Mahajan: And labeling that curve as strain loss. And from the same dictionary, we'll be extracting test loss values, plotting the curve, and labeling that curve as test loss. [01:20:08] Asmita Mahajan: Setting the title to be hidden is equals to 32, so K will store the values of the keys. [01:20:17] Asmita Mahajan: We are setting the X label as epoch, and setting the Y label as loss, displaying the legend, because we want to distinguish between the train and test loss curve. [01:20:27] Asmita Mahajan: Sim… [01:20:29] Asmita Mahajan: We don't want, any subplots, on the empty locations, or the, any empty subplots, so that's why we are just, telling that that axis should be off. [01:20:42] Asmita Mahajan: Displaying or labeling the overall title, and showing the subplots. [01:20:51] Asmita Mahajan: So these are the subplots. [01:20:53] Durga Toshniwal: So once again, yeah, you can, yeah, zoom down a little to show. [01:21:02] Durga Toshniwal: So, you can see here that, in, [01:21:07] Durga Toshniwal: Basically, we have tried different combinations of the hidden layers. [01:21:13] Durga Toshniwal: And, as we are increasing the first, hidden layer size, what is happening, generally is that, [01:21:21] Durga Toshniwal: First of all, the brain and the tossed… sorry. [01:21:26] Durga Toshniwal: the… [01:21:28] Durga Toshniwal: as we are increasing the size of the hidden layer, we can see that the, the convergence, I mean, the… [01:21:37] Durga Toshniwal: Train and test losses. [01:21:39] Durga Toshniwal: our… I mean, there is the overfitting is not happening much. [01:21:45] Durga Toshniwal: Mostly, in most of the cases, it is not happening much. [01:21:49] Durga Toshniwal: However, we will then, if the train and test losses are not [01:21:55] Durga Toshniwal: Very different. They are quite similar, means overfitting is not there. [01:22:01] Durga Toshniwal: So then we'll go for the convergence. [01:22:04] Durga Toshniwal: So, once again, if we have a look at the convergence, then we can see here, probably, [01:22:11] Durga Toshniwal: The convergence, maybe you can show with the… Arrow also, [01:22:17] Durga Toshniwal: So, as you can see here, the convergence for 32, we can start with 32, is somewhere here. [01:22:23] Durga Toshniwal: Which is, like, around 125 or 100 plus something. [01:22:29] Durga Toshniwal: And if we go with 64, then the convergence is… [01:22:35] Durga Toshniwal: Slightly more than 50, probably. Yeah, must be around 60 or something like that. [01:22:41] Durga Toshniwal: Then, with 128, if we see, of course, we can see here that may be due to data bias at around 250, there are some test losses that are getting very high, but the convergence is happening around somewhere similar to 64, which is, like, 70-something. [01:23:00] Durga Toshniwal: And, then if we look at 256, [01:23:04] Durga Toshniwal: The convergence is happening at around 50 here. [01:23:09] Durga Toshniwal: So, this is the case, probably, which shows quite good con… fast convergence. [01:23:16] Durga Toshniwal: Whereas the… if we look at the losses, that is the train and test losses, almost all four cases, show similar behavior only, in terms of the highest, lowest loss, and the, and the difference between the train and the test losses are also parallel only. [01:23:36] Durga Toshniwal: Only, thing is that you can see here, another thing to notice. [01:23:42] Durga Toshniwal: That if you look at the shape of the curve itself. [01:23:46] Durga Toshniwal: When we take 256 number of, [01:23:50] Durga Toshniwal: The layer size. You can see that it is more curvical in shape. [01:23:57] Durga Toshniwal: Right, so the… you can see that the decline is happening quite fast. [01:24:02] Durga Toshniwal: If you look at 32, it's just a small little dip. [01:24:07] Durga Toshniwal: Probably a very small dip as compared to if it was a 45 degree line, slightly, slightly curved version of a 45 degree line. If you look at 64, [01:24:20] Durga Toshniwal: then the angle or the convergence is increasing. If you look at 128, it is still further increasing. Look at 256. [01:24:30] Durga Toshniwal: Now you can see that the dip in the line, or the curve-ness in the line is quite high, which is standing for the convergence only. So you can see that the rate of convergence is quite high, and we [01:24:45] Durga Toshniwal: I should probably choose this. [01:24:48] Durga Toshniwal: And now, this will be more clear as we actually look at the numeric figures of the performance. [01:24:59] Durga Toshniwal: Probably, we can go down to that, and you can explain the code of it later, if you'll just scroll down. [01:25:06] Durga Toshniwal: And let's have a look at the tabular format. [01:25:09] Durga Toshniwal: So here, what you can see here, now in this case, you please enlarge slightly, yeah. [01:25:16] Durga Toshniwal: Mute. [01:25:17] Durga Toshniwal: So, you can see here that, the best figures, although the performance is quite good with 32, 64, 128, and 256, which is what we saw earlier also, however, the best results are with 256. [01:25:32] Durga Toshniwal: here. [01:25:35] Durga Toshniwal: And if you look at, yeah, so you can see it with the cursor out there. [01:25:41] Durga Toshniwal: If you look at accuracy, it's 0.895. You look at precision, it is 0.9. Decall is 0.86. F1 score is 0.88. [01:25:52] Durga Toshniwal: And AUC is 0.96, which is very close to 1. Now, if you look at the bar plots themselves. [01:26:00] Durga Toshniwal: You can see that the performance is quite high in all four cases. However, the best performance is definitely with [01:26:07] Durga Toshniwal: 256. So, we'll go with a hidden layer size of 256. [01:26:12] Durga Toshniwal: So, this is, the analysis of the experiments. I think the code, [01:26:18] Durga Toshniwal: Asmita, you can take over to the… [01:26:20] Durga Toshniwal: For the code of the block. [01:26:29] Asmita Mahajan: Yes, ma'am. [01:26:30] Asmita Mahajan: So, [01:26:33] Asmita Mahajan: like, we previously plotted the bar plots using the data frame, results. So, in a similar manner, we'll be doing here also. So, we have converted the result dictionary into a data frame. [01:26:48] Asmita Mahajan: And now, using that data frame, we'll be plotting the bar plots. So, this function that we have defined here will be plotting the bar subplots. So, we are passing the data frame to the function. [01:27:01] Asmita Mahajan: Rows will be 3, columns are 2. Here we… these are the default values, but, we can calc… like, these can be calculated or passed while calling the function also. [01:27:14] Asmita Mahajan: And the overall, label, or the overall title of the figure would… is metrics versus setting. [01:27:21] Asmita Mahajan: Setting means the hidden layer, size setting. [01:27:25] Asmita Mahajan: So, we are extracting the columns from the data frame, so columns will have the names of the, performance matrices, like accuracy, precision, recall, F1 score, and AUC score. And we are storing it in a list named Metrics. [01:27:42] Asmita Mahajan: Now, we are extracting the index, and index will have the value, of each setting, which is 32, 64, 128, and 256. So, index list will have these values. [01:27:56] Asmita Mahajan: So index labels will have these values, but in the form of a string, because we are converting it into strings, so it… these are not integers, but it will be a list of strings. [01:28:09] Asmita Mahajan: So in this manner, the index labels are stored. [01:28:12] Asmita Mahajan: Now, we are plotting the… we are calling the subplots on rows and columns, on the number of rows and columns. The figure size is 5 in 2 columns and 3.82 rows, and, we are just capturing the overall figure and the axis, or the subplot location in the axis variable. [01:28:31] Asmita Mahajan: Now, this travel is used to flatten the 2D array to a list, to a 1D array, iterating over the matrix list, which [01:28:41] Asmita Mahajan: Has the value, like, accuracy, precision, and all the other performance, matrix names. [01:28:50] Asmita Mahajan: So we are enumerating or iterating over, this metrics, list. [01:28:56] Asmita Mahajan: Extracting the first matrix and the first index. [01:29:00] Asmita Mahajan: And using the index, we are extracting from the axis location, the first location of the subplot. [01:29:07] Asmita Mahajan: from the DF result data frame, and for that particular metric, for the first metric, which is the accuracy, we are extracting the values and passing those values to the valves variable. [01:29:19] Asmita Mahajan: And this FALS variable will be used as the y-axis for the bar plot, and the index label, which is the first index, this label is used as the… will be used as the x-axis. [01:29:33] Asmita Mahajan: And each color is black. [01:29:36] Asmita Mahajan: So, in this manner, the bars will be created in the subplot. The title will be the name of the metric, which is accuracy for the first metric. [01:29:45] Asmita Mahajan: X label, so we are setting the X label to be the number of neurons, so this will tell us the number of hidden layer, hidden, neurons in the hidden layer. [01:29:58] Asmita Mahajan: We are setting the Y label as the name of the metric, which is accuracy, limiting it to 0 to… limiting the y-axis values from 0 to 1. [01:30:09] Asmita Mahajan: And also the labels, like these, 32, 64, 120, and 256, we want them to be displayed at a 45 degree angle, so we are using that. [01:30:23] Asmita Mahajan: So, if we haven't used it, if we not… to this. [01:30:29] Asmita Mahajan: Actually, I have to run the previous code also. So, if we'll not do this, they will not be displayed at the 45 degree angle, they'll be horizontal. You can do that. You can just comment it out and see what happens. [01:30:43] Asmita Mahajan: Now, we want, with each bar, we want the values associated with each bar to also be displayed. So, enumerating over the values. [01:30:54] Asmita Mahajan: List? [01:30:56] Asmita Mahajan: getting the values, in V variable, and the indexes in X variable, passing that, [01:31:04] Asmita Mahajan: Both the values to find out the location where the text will be displayed, and passing them in the text function. [01:31:12] Asmita Mahajan: We only want the values to be displayed till 2 decimal places, the horizontal alignment is center, the vertical alignment is bottom, and the font size is 9. And in this manner, the values associated with each part, are being displayed. [01:31:29] Asmita Mahajan: We don't want the empty plot to be displayed, so that's why we are setting it to off. [01:31:36] Asmita Mahajan: The overall title will be displayed using this function, using… the overall figure title will be displayed, and will show the bar plots and the data frame also. [01:31:51] Asmita Mahajan: Of the performance matrices. [01:31:54] Asmita Mahajan: So this, we have already discussed how, the… [01:32:00] Asmita Mahajan: Accuracy, or the performance of the model using the [01:32:05] Asmita Mahajan: using 256 as the size of the hidden layer is performing best. So, we'll set the size of the first hidden layer to be 256, so the first hidden layer will have 256 neurons. Now. [01:32:21] Durga Toshniwal: Just sorry to interrupt. Any questions, anyone? [01:32:25] Durga Toshniwal: On whatever has been discussed. [01:32:32] Durga Toshniwal: So, just one thing I wanted to point out here for all of you, that, because this is the first code on deep learning. [01:32:41] Durga Toshniwal: So, we are going a little slow on explaining each and every, statement. [01:32:47] Durga Toshniwal: And as you can see here, there are a large number of fiber parameters to be tuned. [01:32:53] Durga Toshniwal: So, would you want us to go faster and, skip the individual instruction explanation? [01:33:01] Durga Toshniwal: And cover the entire food, or would you want… [01:33:04] Muni Prakash Ganji: Training, man. [01:33:06] Durga Toshniwal: No, no password, this space is fine. [01:33:09] Durga Toshniwal: Okay, and yeah, we also thought that a little slightly slower space would be better, so that you can understand what's happening. And once this is clear, because every time in future code, we may not take in so much detail. [01:33:27] Durga Toshniwal: So it's good for now. So why I'm asking is that, it may be possible that a part of it might, spill over tomorrow's portion also. [01:33:40] Durga Toshniwal: That's why I was asking. So, if y'all are fine, then we are also fine. In fact, I wanted it to be slow, so that… a little slow, so that you all can understand nicely what is happening. [01:33:52] Durga Toshniwal: Okay, so Asmita, you can continue with the same pace. [01:33:56] Durga Toshniwal: No problem. [01:33:58] Asmita Mahajan: Okay, ma'am. [01:34:00] Asmita Mahajan: So now, as we have tuned our, three hyperparameters till now, the optimizer, the learning rate, and the size of the first hidden layer, so we will fix this, and we will move on to the, onto tuning the other hyperparameters. [01:34:20] Asmita Mahajan: So now, we will see how to tune the, size of the second header layer. So what we are doing is, we are just taking, [01:34:30] Asmita Mahajan: Now we are taking the values as 16, 32, 64, and 128 as the size of the second handle layer. We are setting learning rate to 0.01. [01:34:41] Asmita Mahajan: So… [01:34:44] Asmita Mahajan: Yes. So, these are the, libraries that will be imported. So, you have already imported that. So… [01:34:54] Asmita Mahajan: It need not to be here. [01:34:56] Asmita Mahajan: Okay. [01:34:58] Asmita Mahajan: So, as we have already imported those libraries in the previous code, so we'll just move on to creating the class. So, what we are doing is, we are creating the MLP class to tune the hyperparameter, the size of the second header layer. [01:35:15] Asmita Mahajan: So, what we are doing is, [01:35:19] Asmita Mahajan: we are initialize… we are creating a constructor initialize, passing it the input dimension, and the size of the second header layer. So here, we are not [01:35:28] Asmita Mahajan: passing it the default parameter. It will… this, parameter will… we will get it when we are calling the function. [01:35:37] Asmita Mahajan: So, we are creating the network using… in a sequential manner. So, here you can see a difference, in how we are creating the network. Previously, we are using only one hidden layer, so there was only one linear unit, and one linear unit from [01:35:52] Asmita Mahajan: hidden to output. One linear unit from input to hidden layer, and one linear unit for hidden to output. But now, as you can see, there is one linear unit from the first hidden layer to the second hidden layer also. [01:36:06] Asmita Mahajan: So, the code, is slightly different here. So, this linear unit is creating a fully connected layer from the input layer to the first, hidden layer, and the first hidden layer, for the first hidden layer, the size is 256, that's why we are passing directly 256 here. [01:36:25] Asmita Mahajan: So the activation function, which we are using to pass the outputs from the first input layer to the first hidden layer, is the ReLU activation function. [01:36:37] Asmita Mahajan: Now, in the second linear unit, we are creating a fully connected layer from the first hidden layer to the second hidden layer. And hidden 2 is, not defined yet, because we are optimizing it, so it will be passed at the calling of the function. [01:36:53] Asmita Mahajan: So, 256 is the size of the first hidden layer, and hidden 2 will be the size of the second hidden layer. And it is a fully connected network, and that's why a linear unit is used here. [01:37:05] Asmita Mahajan: And similarly, a ReLU activation function is used to pass the output… outputs from the first hidden layer to the… as the inputs to the second hidden layer. [01:37:17] Asmita Mahajan: And the last fully connected layer is created from the, second hidden layer to the output. [01:37:26] Asmita Mahajan: And we have only one neuron in the output, because we are doing binary classification. [01:37:31] Durga Toshniwal: So, just one thing I want to point out here. See, remember that [01:37:36] Durga Toshniwal: as we explained that, as I had explained earlier, so whenever we are making a… [01:37:43] Durga Toshniwal: Neural network, in addition to the standard input and output, we have these hidden layers. [01:37:50] Durga Toshniwal: So, we, we have a fully connected layer, means, each neuron of any preceding layer is connected to all other new… to all neurons in the second successive… in the next successive layer. [01:38:08] Durga Toshniwal: So, that is a fully connected layer. Now, this, when we are combining, so we are taking a, [01:38:17] Durga Toshniwal: Actually, each neuron is feeding, or is giving weighted input to all other neurons in the next layer. Then the second neuron in the previous layer is giving input to all neurons in the next layer, like this. And they're all getting added together. [01:38:37] Durga Toshniwal: At the successive layer. So, what is happening? This is just a linear combination, as discussed. And, so what you're seeing is this, linear unit, nn.linear.input, whatever input size, which is 256. [01:38:53] Durga Toshniwal: And the nonlinearities coming from the activation function, which is the relu. So you can… when you're seeing this code, you can visualize it also. [01:39:03] Durga Toshniwal: In the visualization, you can think of it as the input layer, then you have the linear combination of the inputs going to the first hidden layer, which is of size 256, means 256 neurons are there. [01:39:17] Durga Toshniwal: You can just visualize while I'm telling you this. Then the activation at this hidden layer will be RELU. [01:39:25] Durga Toshniwal: Then you are having the second hidden layer, and again, it is fully connected. [01:39:30] Durga Toshniwal: And the size of it will be decided, obviously, and so we are having a kind of a runtime variable, which is hidden to… defined hidden to, with the list of the values in it. [01:39:44] Durga Toshniwal: And then that, again, will be a linear combination, and then the, the nonlinearity will be brought in by the relu function, so you can just visualize this. These are the inputs. [01:39:57] Durga Toshniwal: From the input, you have 256 neurons in the first hidden layer, then you are going to the second hidden layer, which you are designing right now, and then you'll have the output layer. And at the second first hidden layer, and the second hidden layer. [01:40:13] Durga Toshniwal: the activation is really. So I was just trying to help you to visualize what is getting modeled out here. That's why, I just added. [01:40:23] Durga Toshniwal: So you can just continue, Asmita, with the code. [01:40:27] Asmita Mahajan: Hey, Mom. [01:40:28] Asmita Mahajan: So, in this manner, we will create, the network with two hidden layers, one input layer and one output layer. [01:40:36] Asmita Mahajan: Now, we are creating this forward function for forward pass, and the input X is propagated through the network using this forward function. [01:40:50] Asmita Mahajan: Now, the train and evaluation function is defined here, where hidden will tell us the size of the second hidden layer, the number of epochs for which, like, the number of times the training will happen. [01:41:06] Asmita Mahajan: And the best LR, which is 0.01, is the learning rate that we will be using. [01:41:16] Asmita Mahajan: In this experiment. [01:41:18] Asmita Mahajan: So, we are creating the object model. [01:41:21] Asmita Mahajan: from the MLP class, which we have just created, passing the size of the input [01:41:28] Asmita Mahajan: Layer, which is 14, which is the number of columns. [01:41:32] Asmita Mahajan: in the extreme tensor, so how we… this is the function, or this is the code that we'll be using to extract that. So, shape 1. So, shape will have two parameters, the number of rows in the, X-Train dataset, and the number of columns. And we want the number of columns to be passed here, that's why we are using 1. [01:41:51] Asmita Mahajan: And hidden 2 will, have the value, or the size of the second hidden layer. [01:41:58] Asmita Mahajan: So the criteria to calculate the loss is binary cross entropy, which we have used earlier also. So we will just calling this function and storing it in criterion variable. [01:42:10] Asmita Mahajan: Atom optimizer will be used, because we, as we have seen, that the optimizer, the optimal optimizer of choice is the ADAM, so that's what we are using, and LR is the learning rate, which is 0.01. [01:42:24] Asmita Mahajan: Train and test losses, [01:42:27] Asmita Mahajan: Sorry, train and test losses lists are being initialized here to capture the losses, after the training. [01:42:36] Durga Toshniwal: just a second. I think there was a question if we can have different activation layers for different, layers. Yes, you can have different activation functions for different layers. So, activation function will… though it works at the neuron level, but it is per layer. [01:42:56] Durga Toshniwal: So, for… in this case, we have chosen both Releu, but it would be different for different layers as well. There's no… nothing necessary that, it needs to be the same. [01:43:08] Durga Toshniwal: So, I hope that answers your question, Sushri. [01:43:14] Durga Toshniwal: Okay, Asmita, you can carry. [01:43:17] Asmita Mahajan: So, we have initialized, training and test losses list to capture the train and test losses, after each training. [01:43:26] Asmita Mahajan: And for each iteration. [01:43:28] Asmita Mahajan: So we are just iterating over the number of epochs, and opening the model in training mode. [01:43:35] Asmita Mahajan: using… [01:43:37] Asmita Mahajan: So we are initializing the gradients to zero, so that we don't have the gradients from the previous iteration in the successive iteration, so there should be no leakage. So that's why we are initializing. At each iteration, we are initializing the weights to be 0. [01:43:56] Asmita Mahajan: now propagating the, extrane dataset through the model, and then calculating the outputs and capturing them in the largest variable. [01:44:06] Asmita Mahajan: using criterion BCE, that is the boundary cross entropy loss, we are calculating, or, yeah, we are calculating the loss, for that iteration, and, these raw outputs and the [01:44:20] Asmita Mahajan: true Y train labels are passed to this criterion function, and using that, we are calculating the loss. [01:44:28] Asmita Mahajan: Propagating that loss backward through the network, using that propagation, and using this backward, function. [01:44:36] Asmita Mahajan: optimizing, or updating the weights at each iteration using optimizer.step function, and [01:44:44] Asmita Mahajan: Appending the previously computed, computed loss to the train loss, list using the append function. [01:44:52] Asmita Mahajan: And item is used to convert the loss into float. [01:44:57] Asmita Mahajan: And in a similar manner, for this particular iteration, we'll be calculating the testing loss also, the test loss, and appending that loss to the test loss list. [01:45:09] Asmita Mahajan: So, for that, we are opening the model in evaluation mode. We don't want any weights, because the learned weights will be used to propagate or to pass the X test dataset, and the learned weights will be used for that. [01:45:26] Asmita Mahajan: We are capturing the outputs from the model after passing the X test dataset and capturing them in test logist, calculating the loss using the binary cross entropy loss, and appending that loss to the [01:45:42] Asmita Mahajan: Test losses, list. [01:45:48] Asmita Mahajan: So, for each iteration, in this manner, we will be… our test loss list and train loss list will be filled. [01:45:56] Asmita Mahajan: For each iteration. And now, after the model has been trained for 300 or 200 epochs, now we will be evaluating the model on the test set using the learned weights. [01:46:11] Asmita Mahajan: So we are opening the model in evaluation mode. We don't want any weights, we are… we want the long waits to be used, capturing the raw outputs in logist while passing the X test dataset to the model. [01:46:25] Asmita Mahajan: Using the sigmoid activation function, we are converting our raw logic to probabilities, and using 0.5 as the threshold, we are converting the probabilities to hard predictions. [01:46:38] Asmita Mahajan: So in the… in this manner, we have now, made predictions from the X test dataset. [01:46:45] Asmita Mahajan: So, again, YTRUE, YPREDICT, and Y probabilities, so we have to convert it into numpy array, the tensors. So, using, these commands, or these functions, or this code, we are converting our X test tensor into YTRUE, labels, array. [01:47:05] Asmita Mahajan: The predictions are being converted, so predictions also were in the form of a tensor. As I am hovering over this variable, you can see Preds tensor. So Preds is a tensor, variable. [01:47:18] Asmita Mahajan: So to convert it into a NumPy array, now, you can see YPRED has a ND array, written over it. [01:47:24] Asmita Mahajan: So… [01:47:26] Asmita Mahajan: YPRED will be a… has been converted into NumPy array. Probabilities… probability tensor has been converted into a NumPy array using, this… [01:47:37] Asmita Mahajan: Good. [01:47:39] Asmita Mahajan: So, metrics dictionary is created to capture the performance matrices, accuracy, precision, recall, F1, and AUC score. History dictionary is created to capture the train and test laws for each iteration, and this function returns the metrics and history dictionary. [01:47:56] Asmita Mahajan: So now, we are creating the list, where the list has the elements, and the elements will tell us the size of the [01:48:04] Asmita Mahajan: second, hidden layer. So here, we are defining it to be 16, 32, 64, and 128. Result in history dictionary will be storing the, performance matrices for each [01:48:21] Asmita Mahajan: Hidden layer size. [01:48:23] Asmita Mahajan: And history will store the test and training laws for each hidden layer size. So we are iterating over each size. [01:48:31] Asmita Mahajan: And, calling the train and evaluation function for each hidden layer size, for 300 epochs and learning rate as 0.01, and then, storing the metrics and history dictionary. [01:48:44] Asmita Mahajan: And creating the result dictionary for that particular hidden layer size, and storing the performance metrics in the result dictionary, and training and test laws for that particular hidden layer size in the history dictionary. [01:48:59] Asmita Mahajan: Now, using the similar command as we have seen previously, we are creating the data frame, so… [01:49:07] Asmita Mahajan: iterating over the, hidden layer size, we'll be, cap… we'll be capturing the dictionaries at the result, because result is a… is a dictionary of dictionary. So, results will contain [01:49:24] Asmita Mahajan: So, how results is stored is that, is, like, it has a key of, like, 16, and for that particular key, the value will be, the metrics. So, metrics dictionary is there. [01:49:40] Asmita Mahajan: And this matrix will have this. So this dictionary will be there for each [01:49:47] Asmita Mahajan: key. And in a similar manner, it will have 32 as a key, and metrics as value. And this matrix is, again, a dictionary. So it's a dictionary of dictionaries, so nested dictionary is there. [01:50:00] Asmita Mahajan: So, in this manner, results, from this, results, we'll… we are creating the data frame, [01:50:08] Asmita Mahajan: And setting the… Val- these values as the index. [01:50:15] Asmita Mahajan: For the data frame. And the columns will be the accuracy, precision, Recall F1 score, and AUC score. [01:50:22] Asmita Mahajan: So now we'll be plotting the train and test losses for each, hidden layer size. So how we are doing that? Histories will have the values for the train and test loss. So we'll be passing, so we are creating this function to plot the train and test curves. [01:50:42] Asmita Mahajan: train loss and test loss curves, passing histories, dictionary to the function, maximum column S3, and the overall title of the figure. [01:50:54] Asmita Mahajan: So, keys will, keys, list will, have… this. [01:51:06] Asmita Mahajan: Not written it, okay. So, keys, a list will be like this. It will have 16. [01:51:15] Asmita Mahajan: So, these are the elements of… the keys list. [01:51:29] Asmita Mahajan: So keys will, will contain these values. N will contain the length of keys, which is 4. [01:51:38] Asmita Mahajan: columns will be, the minimum of these two, the maximum columns defined here, and the keys, so it will contain. So, actually. [01:51:49] Asmita Mahajan: Yeah, so maximum column is 3, and N is 4, so minimum of the two will be stored in columns, so it will have 3, and rows will be… [01:52:00] Asmita Mahajan: One, two. [01:52:04] Asmita Mahajan: So, using subplot function, we will be creating the subplots and the grid for the subplots. So, we will have two rows, three columns, and the overall figure and axis location of each subplot. [01:52:19] Asmita Mahajan: enumerating over the keys, so enumerating over these values. So each key will be stored in K, and the index, will be stored here, 01234. [01:52:31] Asmita Mahajan: Now, this code will, extract the location of the first subplot, which is for the first hidden layer. [01:52:39] Asmita Mahajan: For the first, the size of the hidden layer. [01:52:44] Asmita Mahajan: So, we are extracting for 16, [01:52:47] Asmita Mahajan: So, K will have, value 16 in the first iteration. [01:52:54] Asmita Mahajan: So the dictionary for that will be stored in HIST. So we are extracting the train loss first, and then plotting it, and, labeling as train… it as train loss. [01:53:04] Asmita Mahajan: So in the second… so the second line of test loss is extracted using HIST test loss, and the values for this are plotted on the subplot, and we are labeling it as test loss. [01:53:16] Asmita Mahajan: Again, the title is set, the X label is set as epoch, and the Y label is set as loss. We'll be displaying the legend to distinguish between the train and test losses, and so that the empty subplots are not displayed here. That's why the axis is off here. [01:53:33] Asmita Mahajan: The final figure subtitle is displayed using this command, and the plot is shown. So, in this manner, we are creating, or we are displaying different subplots. [01:53:47] Asmita Mahajan: for each. [01:53:49] Asmita Mahajan: Hidden layer, size. [01:53:52] Durga Toshniwal: So again, you can notice, if you look at, hit number of hidden layers, 16. [01:53:58] Durga Toshniwal: Versus, 32 versus 64. [01:54:02] Durga Toshniwal: Versus 128. [01:54:05] Durga Toshniwal: So, here, what is happening is that, [01:54:10] Durga Toshniwal: Actually, the graphs are looking a little jittery, definitely. [01:54:16] Durga Toshniwal: But, we can see that the test loss, if we look at hidden, layer size, I mean, hidden layer 2 size equal to 16, [01:54:26] Durga Toshniwal: So, the test loss is quite, coming out to be high as compared to the train loss. [01:54:33] Durga Toshniwal: So, definitely, we may not go for this option. If we look at, hidden, size of 32, [01:54:43] Durga Toshniwal: Which is the second, case. [01:54:46] Durga Toshniwal: So here, the train loss and the test loss are showing differences, all through. [01:54:53] Durga Toshniwal: May not be a good option again. [01:54:55] Durga Toshniwal: With hidden size equal to 64. [01:54:59] Durga Toshniwal: What is happening is that the train and test losses match. However, there are some jitters that can be seen for the test loss over a range of [01:55:11] Durga Toshniwal: somewhere between 150, I mean, between, yeah, 150 and 200 at this point in time. [01:55:19] Durga Toshniwal: And if you could scroll a little down, so what we can see here… [01:55:23] Durga Toshniwal: Is, that in this particular case, the last case with number of neurons to be 128, again, the train loss looks to be much lower than the test loss. [01:55:40] Durga Toshniwal: So, so, from these figures. [01:55:45] Durga Toshniwal: Probably 128 may not be a very good option. 16 also doesn't look to be a very good option. [01:55:52] Durga Toshniwal: 32 also doesn't look to be a good option. 64 might be a better option to go for. However, things will be more clear with the help of the figures that will be shown in the next section. So, what we'll do is that we'll not do the code for that section, I will just, [01:56:11] Durga Toshniwal: I request Asmita to scroll down. [01:56:15] Durga Toshniwal: to that part. [01:56:17] Durga Toshniwal: Please, scroll down to the table. [01:56:20] Durga Toshniwal: Yeah, so now if you'll just, yeah, enlarge slightly… [01:56:25] Durga Toshniwal: So, from here, what we can see is that, [01:56:31] Durga Toshniwal: At 128, it is coming out to be 0.92, the accuracy. [01:56:39] Durga Toshniwal: And the precision also is coming out higher. [01:56:45] Durga Toshniwal: And recall, an F1 score, and AUC. [01:56:49] Durga Toshniwal: So, AUC is comparable to 64. [01:56:53] Durga Toshniwal: However, the others, that is accuracy, precision, recall, and everything, are coming out, higher, [01:57:03] Durga Toshniwal: Even though there was a lot of jitter and all that. [01:57:07] Durga Toshniwal: So, this means that, probably 128 is still a better option. [01:57:14] Durga Toshniwal: Now, even in the second layer. [01:57:17] Durga Toshniwal: Even though, earlier it was looking… it was a little confusing when we looked only at the performance graph because of the, [01:57:27] Durga Toshniwal: Amount of, jitter that was coming in. [01:57:32] Durga Toshniwal: If you look at the bar plots also. [01:57:36] Durga Toshniwal: You will see a better performance. Actually, 64 and 128 are almost comparable. [01:57:43] Durga Toshniwal: However, 128 is slightly higher. [01:57:48] Durga Toshniwal: one as compared to 64. So, we may go for one of these options, so 128. Because it is slightly higher, we may go for that. [01:57:59] Durga Toshniwal: Okay, now you can go up back again, Asmita, to those plots. [01:58:04] Durga Toshniwal: So, from the plots, however, So, actually, if you look at the figures, definitely. [01:58:12] Durga Toshniwal: 128 is looking the best. Probably the convergence is very fast here, because at 50, it has converged. [01:58:21] Durga Toshniwal: Whereas in the other cases, the convergence is much slower, that is the case. However, I would say that [01:58:29] Durga Toshniwal: Even though in this code we have chosen 128 as the best choice because of the figures on the numbers that were coming out to be lower. [01:58:39] Durga Toshniwal: But 64 also may turn out to be a good choice, so 64 and 128 are comparable. We can choose one of these two. [01:58:47] Durga Toshniwal: If we want a less complex neural network, we can go for 64. Otherwise, we can go for 128. [01:58:55] Durga Toshniwal: So, okay, Pallavi, I can see a hand raised. Do you have a question? [01:59:01] Pallavi Chakravarty: Yes, ma'am. So, before this, second layer, when we start, before we added second layer. [01:59:07] Pallavi Chakravarty: In the first layer, we could see that gradually, when we were increasing the neuron size, in 256, it was giving, like. [01:59:15] Pallavi Chakravarty: Aoc was about 96%, right? [01:59:18] Pallavi Chakravarty: So… If we had increased the neuron size to, say, 512, [01:59:24] Pallavi Chakravarty: We probably could have, gotten to 97%, or it might have stayed stable. [01:59:30] Pallavi Chakravarty: I think… Only after that, we can see for second layer, because as we can see. [01:59:37] Pallavi Chakravarty: Here it's… the test loss is not as consistent as it was earlier, when we had only one hidden layer. [01:59:46] Durga Toshniwal: Yeah, so… So, what you're saying… yeah, please continue. [01:59:53] Durga Toshniwal: Yeah, ma'am, please continue. Yeah, so what you're saying is actually very right. [01:59:58] Durga Toshniwal: that, if you'll go up, Asmita, to the previous layers, graphs, or figures. [02:00:06] Durga Toshniwal: So, what you're seeing is very right, that we are thinking that we are achieving a good performance at 256. [02:00:16] Durga Toshniwal: However, if we would have increased it to 512, [02:00:19] Durga Toshniwal: Both of the cases, 50-50 chance. It might have stabilized at 512, or it might have even gone up. [02:00:26] Durga Toshniwal: So, as I mentioned earlier to you, that because of the constraints on compute and all that, and we are doing it only on collab, we are taking a li- [02:00:38] Durga Toshniwal: A limited number of choices, right? But it is true that, wherever you feel that it is stabilizing, it is a good idea to go a one step beyond. [02:00:49] Durga Toshniwal: Say, in this case, let's say if we feel that 256 is best, then it's a good idea to go up to 512 and see whether it is improving further or it is stabilizing. [02:01:00] Durga Toshniwal: If it is stabilizing, then 256 is the best. If it is improving, then you might as well look at the next higher size and see. So that is the correct way of doing it, but because of the constraint on the compute power and all, we are just showing the possibilities. [02:01:16] Durga Toshniwal: Which is opening up, opening you all up with how to actually look at these possibilities, but how many possibilities you can see [02:01:26] Durga Toshniwal: or how many combinations you can go over, actually depends on, how much compute you have at hand. If you have a GPU, you can actually try a lot of combinations and all that. So what you're saying is correct, you should do that. [02:01:39] Durga Toshniwal: However, this is only exemplary. [02:01:42] Durga Toshniwal: I pointed out right in the beginning also that we are just choosing some limited choices to illustrate how we go about this. [02:01:53] Durga Toshniwal: But, yes, we should look at where it is stabilizing. Okay, Falavi? [02:01:59] Pallavi Chakravarty: Yes, ma'am, I was just asking for, like, in general, if that should be the approach. [02:02:04] Pallavi Chakravarty: On that same line, I had another question. Is it for the same, like, compute size or exemplary purpose that we are selecting the, second layer sizes to be 128, since we have already chosen 256 to be, [02:02:22] Pallavi Chakravarty: the size of the neuron for the first layer? Like, is it just for example purpose, or is there any… [02:02:28] Pallavi Chakravarty: particular reason. Yeah. [02:02:29] Durga Toshniwal: Again, I will say, I pointed out in the beginning also, it's for… not just for Pallavi, for everyone, we have taken some number of choices based on some, previous domain, you know, knowledge, or some common numbers that are often chosen. [02:02:48] Durga Toshniwal: But in actual cases, when you'll be doing some, you know, the industry-wide project, or when you are actually doing some live projects. [02:03:00] Durga Toshniwal: It's always a good idea to go for different sizes, more or less. I mean, it could be 256, 512 and all. [02:03:09] Durga Toshniwal: So the… so the choice of the size for the second… second hidden layer are also just exemplary to illustrate this. However, a bigger range will always be helpful. [02:03:22] Durga Toshniwal: We couldn't have done it on Polar. [02:03:25] Durga Toshniwal: And with the limited time and resources we have, that's the only reason why we didn't choose. [02:03:32] Durga Toshniwal: Okay. [02:03:36] Pallavi Chakravarty: Yes, Pram, just one last question. [02:03:40] Durga Toshniwal: kill it. [02:03:41] Pallavi Chakravarty: On… if we compare this first hidden layer plot and the second hidden layer plot, here we can see that, like, with every epoch, test sizes, like, test loss is almost following the same pattern as strain loss. It's not that much jittery or fluctuating. [02:03:59] Pallavi Chakravarty: As we saw in the second hidden layer series. So… [02:04:05] Pallavi Chakravarty: Let's… if we keep this… keep to this example, would we… can we say that it's better that we use, the first, keep only first, like, one hidden layer? [02:04:17] Pallavi Chakravarty: Or we can't see that. [02:04:19] Durga Toshniwal: No, not by looking at these plots. So, actually, whether we should finally keep one layer, two layer, or we should have another one, that depends on the final, you know, the final prediction of the target variables. So, right now, you're just building a model, okay? [02:04:38] Durga Toshniwal: And based on whatever we decide right now, we will be building the final model. Your final model will contain all the values that we are finalizing. Say we finalize Adam, it will have Adam. Then we finalize learning rate of 0.01. [02:04:53] Durga Toshniwal: So it will have a learning rate, 0.01. Then we decided a size of whatever, 256 for hidden layer 1, hidden layer 2, whatever we are deciding. [02:05:03] Durga Toshniwal: We will go with all these, and then build the final model. Now, what you can do to see whether you should be having a second hidden layer or not, you can make the model with whatever parameters. If you could go scroll down at the end, Asmita, I think the model will be there. [02:05:20] Durga Toshniwal: At the end. So, you can look at the final performance of the model, okay? And then what you can do is build another model having only one layer. [02:05:30] Durga Toshniwal: And then look at the, overall performance of the full model. [02:05:35] Durga Toshniwal: So, this is the final model, performance values. If you can enlarge slightly, Asmita. [02:05:41] Durga Toshniwal: So, what you can see here, with whatever parameters, hyperparameters we are using, we have a model of MLP with accuracy 0.936, precision 0.936, recall 0.919, [02:05:56] Durga Toshniwal: F1.927 and AUC.98, which are pretty good. [02:06:01] Durga Toshniwal: Now, if you want to see whether we should be going with one layer or two, this is with two layers. [02:06:08] Durga Toshniwal: But, to see if that works better, you remove the second layer and build a final MLP model, and see what is the performance coming up to, and then compare. If that is coming better, then don't go for a second layer, only use the first layer. [02:06:25] Durga Toshniwal: So, this is the way you see. Just by looking at the intermediate results, you cannot really conclude whether you should include a second layer or you should. Finally, it is the final model that will tell you. [02:06:42] Pallavi Chakravarty: Welcome, thank you. [02:06:45] Durga Toshniwal: Okay. Any other questions, anyone? [02:06:49] Nirav Mehta: Yes, ma'am. So, in this learning rate, it will, impact, it will impact the weights, or… or loss, or it will impact both, as in, coming up with the learning rate and plugging it into the model and running the model. [02:07:05] Durga Toshniwal: Yeah, so as I told you in the last turn, The learning rate… [02:07:10] Durga Toshniwal: actually decides the step size, right? So, the step size is the size. [02:07:17] Durga Toshniwal: or the delta, which you use to change or update the weight. Let's say, as I told you initially, there'll be some initial weights that will be randomly decided. [02:07:29] Durga Toshniwal: Now, with those initial choice, the model will do a forward pass. It will calculate the target, variable, which is a… [02:07:41] Durga Toshniwal: Your, the, the class labels. [02:07:47] Durga Toshniwal: right, it will predict the class labels, and then what will happen is that there'll be some loss, or there'll be an error. So, that error will be back propagated, and this back-propagation will help to update the weights and the bias. [02:08:03] Durga Toshniwal: So, the update of the weight and bias will be done with the help of the learning rate in terms of the step size. So, how much [02:08:12] Durga Toshniwal: So, the update will be done to what extent that will be the step size. So, that step size is actually decided by learning rate times the gradient. So, learning rate decides how fast the weights are updated. [02:08:29] Durga Toshniwal: This is what is the role of learning rate. [02:08:31] Durga Toshniwal: And, based on that, your gradient descent will happen, which I had explained, that the convergence towards the [02:08:41] Durga Toshniwal: minimization of the loss will happen. So, learning rate will decide how much of the update on the wait will happen at a time. [02:08:49] Durga Toshniwal: That, because learning weight times the gradient is the step size. So, step size, you can think of step size as a delta. So, whenever you have a weight, let's say some weight values there, say W1, [02:09:03] Durga Toshniwal: So, whenever the error is back propagated, and the weight is updated, how much will be updated will be that step size. So, you will update the weight with W1 plus minus delta, that delta will be the step size. [02:09:17] Durga Toshniwal: And how… how big the step size will be. [02:09:20] Durga Toshniwal: So, accordingly, the weight will be decreased or increased based on the size of the delta, and the delta is dependent on the learning rate, because gradient into learning rate is the step size. So, if the learning rate is high. [02:09:35] Durga Toshniwal: then the convergence will be faster, as we are seeing here also. But if it is overly high, then it will never converge. [02:09:44] Durga Toshniwal: And if it is overly slow, then the rate of convergence will be very, very slow, which is also what you saw here. [02:09:50] Durga Toshniwal: Is it okay? [02:09:52] Nirav Mehta: Yes, ma'am. Thank you. [02:09:54] Durga Toshniwal: Yeah. Any other questions? [02:10:00] Durga Toshniwal: Okay, there are no further questions. We'll conclude for today, and the rest of the code will take it tomorrow. [02:10:07] Durga Toshniwal: Okay, and hopefully we'll finish it. And then, if time permits, I'll also start the next part of the theory, which I had scheduled for tomorrow. [02:10:18] Durga Toshniwal: So let's see how we proceed, tomorrow. Okay, so with that, we'll just end the session here. Thank you all. Thank you very much. [02:10:27] Durga Toshniwal: Okay, so we'll just wrap it up here.