# 06 2026-02-08 Theory DL Multilayer Perceptron MLP

course: Module 3 — Deep Learning & NLP
module: Module-3-Deep-Learning-NLP
date: 2026-02-08
type: transcript
video_url: https://personal-learn.armco.dev/files/_Recordings/Module-3-Deep-Learning-NLP/06_2026-02-08_Theory_DL_Multilayer_Perceptron_MLP.mp4

---
[00:13:59] Durga Toshniwal: A very good morning to all of you. Welcome to today's session. We'll just shortly begin in 2 minutes.
[00:14:05] Durga Toshniwal: We'll just wait for,
[00:14:08] Durga Toshniwal: Some more people to join, and we'll just start in another 2 minutes.
[00:16:41] Durga Toshniwal: A very good morning to all of you, we'll just begin.
[00:16:45] Durga Toshniwal: So we'll continue with, what we had, left in between yesterday. That was the code.
[00:16:53] Durga Toshniwal: So I'll request a smithal.
[00:16:56] Durga Toshniwal: to enable her video.
[00:16:58] Durga Toshniwal: And, to start from where we left.
[00:17:01] Durga Toshniwal: I think the code for this block was left. That is the… we just did a discussion on this.
[00:17:10] Durga Toshniwal: the… the…
[00:17:12] Durga Toshniwal: Yeah, this block, did we complete the discussion on the code for this one? I think we didn't. We just…
[00:17:20] Durga Toshniwal: Did an analysis of the results, so you can start with the code for this block.
[00:17:25] Durga Toshniwal: And, yeah, this one.
[00:17:28] Durga Toshniwal: Okay, over to you, Osmith.
[00:17:31] Asmita Mahajan: Thank you, ma'am.
[00:17:32] Asmita Mahajan: So, now we will… Try to plot the, bar graph, of the matrices.
[00:17:41] Asmita Mahajan: To compare the performance, for different, values of the… Neurons in the hidden layer.
[00:17:49] Asmita Mahajan: So we have chosen 32, 64, 128, 256. So these were the four values, that were there for the… sorry, 16, 32, 64, and 128 for the second hidden layer.
[00:18:06] Asmita Mahajan: So now, using the, results data frame, we'll be plotting the bar plots. So what we are doing is, we are, just, extracting the
[00:18:17] Asmita Mahajan: Column names in the list matrices, and also the index of the data frame in the list, index labels, and the index are, first converted into string data type.
[00:18:31] Asmita Mahajan: So, using the subplot function, we are calling the subplot function on the respective number of rows and columns, and then just extracting the overall figure location and the axis location.
[00:18:45] Asmita Mahajan: And then, flattening the axis from 2D array to 1D array and converting it into a list.
[00:18:52] Asmita Mahajan: Iterating over the matrices, and matrices contain the names of the performance matrices, we'll be extracting the location of the first subplot, where we are plotting the accuracy plot.
[00:19:05] Asmita Mahajan: Then, calling out, like, extracting the values of the first metric from the data frame, and, capturing it in the container values, or the variable value… valves.
[00:19:20] Asmita Mahajan: We are calling bar function to plot the bars, and on the x-axis, we want the labels. The labels would be, like, 16, 32, 64, 128, and on the y-axis, the values are there, and edge color is black.
[00:19:35] Asmita Mahajan: We're setting the, title for each subplot, using the metric name, which is Accuracy, Precision, Recall, F1 score, and AUC score.
[00:19:45] Asmita Mahajan: Setting the X label as the number of neurons, the Y label, as the metric name, and also setting the limit, to 1.2.
[00:19:57] Asmita Mahajan: 0 to 1.2.
[00:20:00] Asmita Mahajan: the Y limit. And the X stick, as we have already discussed, the name, the label of the X stick would be, rotated to 45 degree for, presentation purposes.
[00:20:13] Asmita Mahajan: Now, we also want the values associated with each graph, on the bar of each plot. So, for that, we'll be enumerating over the valves, list, and extracting each value from there, and the index.
[00:20:28] Asmita Mahajan: And using text function, we'll be passing the location of the text, how many decimal points we want, for that value, and the alignment.
[00:20:41] Asmita Mahajan: of the text. What would be the horizontal alignment, and what would be the vertical alignment for the text? So, in this manner, and the font size, lastly.
[00:20:49] Asmita Mahajan: In this manner, we'll be displaying the values associated with each graph.
[00:20:56] Asmita Mahajan: And the extra, subplots, we don't want them, so we are just, telling the code that we don't want that to be displayed.
[00:21:07] Asmita Mahajan: Now, setting the overall, like, figure's, title, which is metric versus second hidden layer size, and then displaying the plot, and also displaying the data frame, which stores the, performance values.
[00:21:20] Asmita Mahajan: So, in this manner, The graphs are displayed.
[00:21:27] Asmita Mahajan: So, here we can see that 64 and 128 have performed nearly, equally, but,
[00:21:35] Asmita Mahajan: We will be choosing 128 as the number of HDLE, because 128 is performing slightly better than the 64 size.
[00:21:45] Asmita Mahajan: So now, so, to, tune the second hidden layer, we used,
[00:21:53] Asmita Mahajan: the learning rate of 0.01, because we had already optimized or tuned the learning rate to 0.01. But now we will also check if we choose a slightly smaller learning rate, what will happen, what will be the performance, and how the performance changes, in the model. So,
[00:22:13] Asmita Mahajan: Now, we will be fixing our learning rate to 0.001, and the sizes will be the same, and we will see how the performance changes and how the model converges with this learning rate.
[00:22:28] Asmita Mahajan: So, firstly, we will set the, like, we will define the class, which is the MLP class. We will have the initialization function.
[00:22:37] Asmita Mahajan: And, the network is defined in a sequential manner using linear unit, which will connect the input layer to the hidden… first hidden layer, and the dimension will be the input dimension, which is 14.
[00:22:53] Asmita Mahajan: And the hidden layer size, which is 256. This is the size of the first hidden layer, using the ReLU activation function.
[00:22:59] Asmita Mahajan: Then we'll have a second linear unit, which will connect the first hidden layer to the second hidden layer. The size of the first hidden layer is 256, and the size of the second hidden layer is, will be passing it. So it will be 16, 32, 64, and 128.
[00:23:15] Asmita Mahajan: with Raylu as the activation function. Then we will have the last linear unit, which will connect hidden layer 2 with the output, with one neuron at the output. And the raw outputs will be… will be returned.
[00:23:33] Asmita Mahajan: This forward function will propagate the input through the network using the forward pass.
[00:23:41] Asmita Mahajan: So, this completes our MLP class. Now we will be defining our train and evaluation function to actually train the overall model. So, hidden will contain the size of the hidden layer.
[00:23:54] Asmita Mahajan: We have 200 epochs, and now the learning rate would be 0.001, because we are now checking at 0.001 learning rate.
[00:24:04] Asmita Mahajan: So, we are creating an object of the MLP class, passing the number of features to the input dimension, and also passing the hidden size.
[00:24:17] Asmita Mahajan: BC will be used, which is binary cross entropy, loss function is used to calculate or evaluate the loss of the model, and atom optimizer is used to optimize the weights with 0.001 learning rate.
[00:24:35] Asmita Mahajan: So, these are the containers, test laws and training loss containers, or the list, to capture or to store the loss associated with each iteration, with each epoch.
[00:24:46] Asmita Mahajan: So we are now iterating over EGPOC, training, opening the model in the training mode, initializing or clearing the gradients from each iteration, and then, propagating or, performing the forward pass on the training dataset, and, capturing the raw outputs in Logist.
[00:25:06] Asmita Mahajan: Now, using criterion, or BC, binary cross entropy laws, we are calculating the loss using the raw outputs and the Y true labels.
[00:25:16] Asmita Mahajan: Now, this loss is propagated backwards in the network using backpropagation and using the backward function, and optimizer.step will update the weights.
[00:25:25] Asmita Mahajan: Now, the calculated loss will be appended to the train loss container using append function.
[00:25:33] Asmita Mahajan: For this particular iteration.
[00:25:37] Durga Toshniwal: I'm just… I'm interrupting you.
[00:25:39] Durga Toshniwal: Are there any questions, anyone, on what was done yesterday, or what we are doing presently?
[00:25:54] Durga Toshniwal: Okay, if there are no further questions, you may proceed, Asmita.
[00:25:58] Muni Prakash Ganji: running rate, we are doing, change today, but the same, the code is same, and
[00:26:05] Muni Prakash Ganji: It is… I think it is, for… for us, to get code priority, so you guys are repeating
[00:26:12] Muni Prakash Ganji: with the…
[00:26:15] Durga Toshniwal: Yeah, so here.
[00:26:16] Muni Prakash Ganji: What is happening?
[00:26:17] Durga Toshniwal: is that in the previous part, if you can just go up, Asmita, to those plots.
[00:26:24] Durga Toshniwal: So, you can see that, the upper plots, yeah. So, you can see here.
[00:26:31] Durga Toshniwal: That, with the learning rate of 0.01 that we had chosen.
[00:26:35] Durga Toshniwal: The train and test losses are not very much stabilized, there are a lot of jitters. So, although we have selected the hidden layer size to be 128, we just wanted to explore
[00:26:52] Durga Toshniwal: Whether some other learning rate could be beneficial in place of 0.01.
[00:26:59] Durga Toshniwal: So, in that regard, the optimizer and all remain the same. However, now, if you'll go down to the current block as well…
[00:27:08] Durga Toshniwal: So, what is happening here? Again, the sizes are… that are being, explored are still… the set is same, 16, 32, 64, 128, which we took earlier.
[00:27:20] Durga Toshniwal: But because of those jitters, we thought that probably if we can go with a lesser learning rate, and see if that makes things better.
[00:27:30] Durga Toshniwal: So, the rest of the code will be just the same.
[00:27:33] Durga Toshniwal: But only learning rate is going to be different. And now we'll see what is the impact of reducing the learning rate by this much, by point… around 10%, so it reduces from 0.01 to 0.001.
[00:27:50] Durga Toshniwal: And then, we'll see the impact and accordingly decide.
[00:27:54] Durga Toshniwal: What, shall be our final learning rate?
[00:27:58] Durga Toshniwal: So, this is what we are trying to do.
[00:28:00] Durga Toshniwal: Okay, so the code here is, just the same as the previous block, but as requested by all of you yesterday.
[00:28:09] Durga Toshniwal: That, we… you all would like you to go, and would like to go slowly over the code.
[00:28:16] Durga Toshniwal: So we are doing that today.
[00:28:19] Muni Prakash Ganji: It's okay.
[00:28:20] Muni Prakash Ganji: Thanks, and we need… because for me, especially, it's new, so it will be better if it is repeating.
[00:28:28] Durga Toshniwal: Yeah, yeah, not just you, I think the majority, confirmed, that's what I had asked, because I wanted to start a new portion today, that you wanted to go over the code slightly slowly.
[00:28:40] Durga Toshniwal: State written by statement, and I think right now it's reasonable to go like that, so that you get a good foundation.
[00:28:49] Durga Toshniwal: Of how the code is written, because later on, when we have further, more complex deep learning models, this is just the starting. It may not be possible also to go that very…
[00:29:00] Durga Toshniwal: I mean, at this pace, then we'll definitely need to go more fast-paced. So, right now, that's why we can afford to do that.
[00:29:10] Durga Toshniwal: So, okay, so Asmita, you can continue and go over, in a slow pace only, so that people can understand, even if it is repeating, it will be for the benefit of everyone, okay?
[00:29:24] Asmita Mahajan: Yes, ma'am.
[00:29:28] Asmita Mahajan: So, after evaluating the training laws for this particular… for a particular iteration, we will be evaluating the test laws similarly, to, for that particular iteration. So now, we are opening the model in evaluation mode using eval function.
[00:29:45] Asmita Mahajan: And now we want to use, because as we are, we will be passing the test data to the network, so we need not to have the, updates in the weights. We'll be using the, trained, weights to test the model.
[00:30:02] Asmita Mahajan: So that's why, with torch.noGrad, which means we don't want any gradients, or we don't want to update the gradients or the weights. So now, the forward pass is performed on the model using the test tensor dataset, and test lodges will capture the raw outputs,
[00:30:21] Asmita Mahajan: From the model.
[00:30:23] Asmita Mahajan: Now, we'll calculate the loss using criterion function, and, we'll append that calculated loss to the test losses, list.
[00:30:34] Asmita Mahajan: Using a pen function. So in this manner, we will train it, for each, iteration for each epoch, which is 200, which is set to 200 pair, or 300 using, at the time of calling.
[00:30:47] Asmita Mahajan: So now, the mod… after the model is trained, the overall model is used to test the test dataset. So we'll be, again, opening the model in evaluation mode for the final,
[00:31:01] Asmita Mahajan: outputs using torch.lograd, because we don't want any gradients to be updated. We'll use the, learned, weights.
[00:31:10] Asmita Mahajan: So now the model is, you, model, like, we'll, perform the forward pass using the X test datasets, and, capture the raw outputs using, in the largest variable.
[00:31:23] Asmita Mahajan: Using the sigmoid activation function, we are now converting our raw outputs into probabilities, and using 0.5 as the threshold, we are converting the probabilities and getting our predictions.
[00:31:39] Asmita Mahajan: So now we have to calculate the performance matrices. So for that, our data or our true labels should be in a NumPy format. So, previously they were, like, in tensor, so now we have to convert them to NumPy. So, before doing that, we are, bringing back our data from the GPU to the CPU, converting the data
[00:32:03] Asmita Mahajan: or the Y tensor… Y test tensor into a NumPy array, and then flattening the array.
[00:32:11] Asmita Mahajan: into 1D array, and getting our byTrue list.
[00:32:15] Asmita Mahajan: So, Y2List will contain the two labels, Y predict will contain the predicted labels which we… which the model predicted, and Y probability will have or will contain the probabilities associated with the predictions.
[00:32:33] Asmita Mahajan: This matrix, dictionary will store the performance… the evaluated performance matrices, the values.
[00:32:41] Asmita Mahajan: of all this accuracy, precision, recall, F1 score, and AUC score, and history will store the training loss and test loss for each epoch. And we will return these, both these, dictionaries.
[00:32:56] Asmita Mahajan: Now, we are defining a second layer size list, which will contain the size of… size that we will be passing at the time of calling of the function. Results and history will… dictionaries will,
[00:33:08] Asmita Mahajan: store the matrices associated, or the performance matrices associated with each size. And history will also con… history will contain the train and test laws for each iteration, associated with each size of the hidden layer.
[00:33:27] Asmita Mahajan: Now, iterating over the, hidden layer size, will pass, one by one, each size to the train and evaluation function.
[00:33:35] Asmita Mahajan: Then, in return, we'll get the matrices, performance matrices, and train and test losses for each, hidden layer size, and that will be stored in result and history.
[00:33:48] Asmita Mahajan: dictionaries.
[00:33:50] Asmita Mahajan: Using result dictionary, we will be converting, the dictionary into a data frame, so…
[00:33:59] Asmita Mahajan: this function is used, and we are setting the index to the name of the hidden… to the size of the hidden layer, and the columns will be the performance metrice name, which is accuracy, precision, recall, L1 score, and AUC score.
[00:34:15] Asmita Mahajan: Now we will be plotting, we will see how the train and test curves, differ,
[00:34:23] Asmita Mahajan: From each other for these sizes. And also, we have to make… we have to see
[00:34:29] Asmita Mahajan: that the learning rate that has been passed here in the train and evaluation function is 0.001, which is slightly, less than the previous learning rate, which was 0.01. And we will see, like, using this learning rate, how our curves have converged.
[00:34:47] Asmita Mahajan: So, to test that, or to see that, we will be plotting the curves using histories, because history contains our trade and test loss values for each iteration and for each, value of the, size of the hidden… second hidden layer.
[00:35:03] Asmita Mahajan: And the title is Train Test Law vs. Second hidden layer size.
[00:35:08] Asmita Mahajan: Keys will contain the history keys, which will be… so, keys will have the value as 16.
[00:35:21] Asmita Mahajan: 32.
[00:35:24] Asmita Mahajan: 120… sorry, 64.
[00:35:31] Asmita Mahajan: 128. So, this is what the keys list will look like. So, N is the length of the keys, which is 4. Column will be the minimum of max column, and N, which is, 3.
[00:35:45] Asmita Mahajan: And, rows will contain 2.
[00:35:48] Asmita Mahajan: So now, using subplots, we are plot… we are… we will be plotting, our subplots with the two rows and three columns.
[00:35:57] Asmita Mahajan: And the figure and axis is the overall figure and the location of the subplots they will contain.
[00:36:05] Asmita Mahajan: enumerating over the keys, these keys, we'll be capturing each key and plotting it in the subplot.
[00:36:12] Asmita Mahajan: And using, this axis, so here we are not raveling or flattening the axis, we are just, extracting from each row and column value, we are extracting that location value.
[00:36:25] Asmita Mahajan: And passing it to acts, func, variable.
[00:36:29] Asmita Mahajan: And, we are also extracting the train and test laws for that particular hidden layer size, and storing it in HIST. So HIST will be a dictionary.
[00:36:42] Asmita Mahajan: containing the train and test loss for that particular, hidden layer size. And then plotting the first curve, which is the training loss curve, using the his train loss, and labeling it at train loss. Now plotting the second curve, which is a test loss curve, using his test loss, and labeling it as test loss.
[00:37:02] Asmita Mahajan: Setting the title to the hidden is equals to K, where K is the value of the hidden layer size.
[00:37:10] Asmita Mahajan: setting the X label as epoch and Y label as loss, and also displaying the legend to distinguish between train and test losses. Now, we don't want, the extra subplots, the empty subplots, so we will be, like,
[00:37:25] Asmita Mahajan: We'll… using this code, we are telling that we don't want that empty subplots. Setting the subtitle of the overall figure, and then displaying the plots.
[00:37:41] Durga Toshniwal: So I think we have… oh.
[00:37:44] Durga Toshniwal: I just, seen the plots.
[00:37:47] Durga Toshniwal: I mean, made the shots. Now, what you can see here…
[00:37:51] Durga Toshniwal: That, in contrast to what we had drawn earlier.
[00:37:55] Durga Toshniwal: These plots are much more smooth. That is the…
[00:37:58] Durga Toshniwal: The noise in the learning rate and all that is…
[00:38:03] Durga Toshniwal: Has now gone. It has become quite smooth.
[00:38:06] Durga Toshniwal: However, the important thing to notice is, here, you can see, each of these,
[00:38:12] Durga Toshniwal: Draft, starting with hidden layer 18, then 32, 64, and 128.
[00:38:18] Durga Toshniwal: The highest value.
[00:38:21] Durga Toshniwal: Of the loss, it starts with something around…
[00:38:29] Durga Toshniwal: Slightly, I mean, 0.7, and it goes up to, you can see 0.3, somewhere around 0.3.
[00:38:38] Durga Toshniwal: So, out of these four, Again, we can see that…
[00:38:43] Durga Toshniwal: The hidden layer size of 128.
[00:38:48] Durga Toshniwal: Shows the maximum and the fastest convergence.
[00:38:52] Durga Toshniwal: Which is… which is at around 75, or 70, like that. So…
[00:39:03] Durga Toshniwal: Sorry, I got muted.
[00:39:05] Durga Toshniwal: So, these, so the convergence here is faster.
[00:39:12] Durga Toshniwal: So, I'd like to show you the results, so you can go to the results of the next block, Asmita.
[00:39:19] Durga Toshniwal: So, if you look at this tabular form now, I can just enlarge slightly.
[00:39:26] Durga Toshniwal: So, definitely, out of the four of these.
[00:39:31] Durga Toshniwal: With hidden layer size, I mean, hidden layer… second hidden layer size of 16,
[00:39:38] Durga Toshniwal: Then 32, 64, and 128, we can see that
[00:39:42] Durga Toshniwal: Definitely 128 is giving… 64 is giving better results here, not 128.
[00:39:49] Durga Toshniwal: So, the thing is that we might feel that,
[00:39:55] Durga Toshniwal: you know, the size of 60 or 4 may be better. Now, you can notice that at 64, the accuracy, precision, recall, F1 score, and AUC that are achieved are, like, 0.88.
[00:40:10] Durga Toshniwal: I mean, 0.89, 0.89, 0.86.87, and 0.96, something like that.
[00:40:19] Durga Toshniwal: Now, let's go to the previous case with the learning rate of 0.01. Can you please go to that table?
[00:40:27] Muni Prakash Ganji: Ma'am, both 128 also similar, correct? 64 and 128.
[00:40:31] Durga Toshniwal: Similar, but slightly… yeah, it's very similar.
[00:40:35] Durga Toshniwal: And if you recollect in the last turn, I mean, yesterday also, I said that 64 and 128 are very similar, with very minor difference only. However, at that time, because 128 was coming out the best, we had chosen 128.
[00:40:49] Durga Toshniwal: With learning rate of 0.01, that is what we had chosen last time. Now we are trying to see, with 0.01, what is the impact of this learning rate on the performance of this classifier, that is MLP.
[00:41:05] Durga Toshniwal: And those were the results that I was showing, and with a learning rate of 0.001,
[00:41:11] Durga Toshniwal: the… and the size of 64 for the hidden layer 2, the results were coming slightly better. However, now, if we look in totality, those were the results that we saw with a learning rate of 0.001.
[00:41:28] Durga Toshniwal: And size was 64 was giving slightly better results. Now, if you look at this, if we look at the results with learning rate 0.01 and size 64, because there, 64 was giving better with 0.001. So, these are the results with 0.01, and
[00:41:47] Durga Toshniwal: You can just highlight, here, 64.
[00:41:51] Durga Toshniwal: first 64, because there, 64 was coming better. But you see, the 64 out there was, giving results like what?
[00:42:00] Durga Toshniwal: The results were something like, for 64, it was, like, 0.889 for accuracy, for 0.893 for, precision, like that.
[00:42:12] Durga Toshniwal: So now, here you can see, even if we chose, 64, the results are better.
[00:42:19] Durga Toshniwal: And they are still better with 128, that is 0.92, 0.92, and on. So, in totality, with the learning rate of 0.001, if we choose 64 or we choose 128, the results are still not as good.
[00:42:35] Durga Toshniwal: As the overall performance with a learning rate of 0.01,
[00:42:40] Durga Toshniwal: And the size of 64 or 128.
[00:42:43] Durga Toshniwal: Here, we have chosen 128, so we'll finally go with 128.
[00:42:48] Durga Toshniwal: Okay, is it okay with everyone? Any questions on the interpretation part?
[00:42:56] Durga Toshniwal: So, here we have just tried, because earlier we had chosen, 0.01 as the learning rate, so initially we experimented with 0.01, but just for the sake of completeness, we also tried to see, because of this jitter, if
[00:43:10] Durga Toshniwal: 0.001 would help. It wasn't helping because the overall performance was still coming out to be lesser at 64 and 128 both.
[00:43:19] Durga Toshniwal: So, we continue to… we will continue to use 0.01 only, and not 0.001.
[00:43:26] Durga Toshniwal: Because, whether there were jitters or not, or the overall performance of the MLP classifier is better with a size of 128 and a learning rate of 0.01.
[00:43:38] Durga Toshniwal: This is a point that we wanted to make by doing the next block that you were doing.
[00:43:43] Durga Toshniwal: Is it okay, everyone? Any questions?
[00:43:49] Durga Toshniwal: Okay, Asmita, you can take the…
[00:43:52] Durga Toshniwal: The next section of that block in which we are currently there.
[00:43:58] Durga Toshniwal: Yeah, this one, I think.
[00:44:02] Durga Toshniwal: Yeah, thanks.
[00:44:08] Durga Toshniwal: Yeah, I think we were in the second part of, learning rate .01. Yeah, this code you have to explain. I just took the results.
[00:44:15] Asmita Mahajan: Beautiful.
[00:44:19] Asmita Mahajan: So now, to plot the bar plot.
[00:44:22] Asmita Mahajan: For the performance matrices, to compare the performance matrices for each, hidden layer size, we'll be using the data frame that we, that we converted, from the dictionary results. So,
[00:44:39] Asmita Mahajan: We will be defining the function, passing the data frame, and passing the row and columns and the overall subtitle of the figure.
[00:44:49] Asmita Mahajan: Now, extracting the matrices from… matrices name from the data frame using, dfresults.columns and converting it into list.
[00:44:59] Asmita Mahajan: Also extracting the index of the data frame, which will be the name, or the values of the hidden layer, and converting it into a string, and then storing it in labels, index labels.
[00:45:12] Asmita Mahajan: Now, calling out the subplots. So, function to, plot the subplots, in a, in a grid of, rows and columns. So, rows and columns are defined here as 3, and column will be 2.
[00:45:27] Asmita Mahajan: So, at the time of calling, we have defined the rows and columns as 2 and columns are 3.
[00:45:33] Asmita Mahajan: So we'll be, plotting the overall figure and the subplots.
[00:45:38] Asmita Mahajan: So, here, we have flattened the access, ND array into, 1D array and stored it in a list.
[00:45:45] Asmita Mahajan: Now, enumerating or, iterating over the matrices list, which contains the names of the performance matrices, we are getting it… so the first performance metric would be accuracy, so we'll be plotting the bar plots for the accuracy values for different, size of the hidden layer.
[00:46:08] Asmita Mahajan: Getting the location of the first subplot, and then, extracting the values of the first performance matrix, which is, accuracy, and storing the values in val.
[00:46:19] Asmita Mahajan: And calling the bar function, and giving… on the x-axis, we want the labels, and the labels would be, like, 16, 32, 64, and 128, and the values, the valves contain the value of the performance metrics accuracy. Edge color is black.
[00:46:38] Asmita Mahajan: Setting the title to metric, which is the accuracy, precision, recall, F1 score, and AOC score. X label is set to, number of neurons. Y label is set to metric, which is accuracy.
[00:46:52] Asmita Mahajan: For the first subplot, and setting the value limit from 0 to 1 for all the subplots.
[00:46:59] Asmita Mahajan: Now, for display purposes, we are just, rotating the X stick
[00:47:05] Asmita Mahajan: at 45 degree angle. And then, for each bar, we want the values associated with it. So, we will be iterating over the valves list, and getting the index and value from them. And then, using text function, we'll be passing the location where we want the text.
[00:47:25] Asmita Mahajan: how many decimal places should be displayed, and the horizontal and vertical alignment, and what would be the font size. And we don't want any extra, empty subplots, so we are just, taking them off.
[00:47:39] Asmita Mahajan: Setting the overall title of the figure, and then displaying the plots, and also displaying the data frame.
[00:47:49] Asmita Mahajan: So in this manner, the subplots will be displayed and plotted.
[00:47:55] Asmita Mahajan: So, we will choose, the 128 as the hidden layer size, and learning rate as 0.001, because the performance, was better, at learning rate 0.01.
[00:48:13] Asmita Mahajan: So, till now, we have,
[00:48:16] Asmita Mahajan: tuned our optimization hyperparameter, the learning rate, the size of the first hidden layer, and the size of the second hidden layer. Now, we will be tuning
[00:48:27] Asmita Mahajan: the activation function for, that we will be using, in, while passing the outputs from the input layer to the hidden layer, and from the hidden layer to the second hidden layer. So, we… we will be seeing which,
[00:48:44] Asmita Mahajan: Activation function would perform better.
[00:48:47] Asmita Mahajan: And for that, we are choosing TANH and RELU activation function, so… sorry.
[00:48:55] Asmita Mahajan: Yeah, so we'll compare the performance for both the activation function, and we will then choose which performed better.
[00:49:04] Asmita Mahajan: So for that, we'll be defining our class, MLP, and in a very similar manner, initialization constructor would be there, and we will define it with the parameters as the input dimension. And now, it will have
[00:49:20] Asmita Mahajan: Activation as a parameter, which will be passed at the calling time.
[00:49:24] Asmita Mahajan: And which will contain the value either Tanetch or ReLU.
[00:49:28] Asmita Mahajan: So…
[00:49:30] Asmita Mahajan: We'll be defining our network in a sequential manner now, using… so, the first linear, unit will have the
[00:49:43] Asmita Mahajan: the size of the input, it will contain the parameter, which would be the size of the input layer, the size of the activation layer, which is set to 256. Then.
[00:49:53] Asmita Mahajan: Here, previously, as you may have noticed, that, like, we would have written directly the function here, the activation function here, but now, as we have to tune the activation function, so here, the activation function would be either Tanach or Reelu, so it will depend at the time of calling.
[00:50:11] Asmita Mahajan: So that's why we have written activation here. Now the second, linear unit, which will connect the first hidden layer to the second hidden layer, and the size is 256 and 128, these are tuned, and we are using directly.
[00:50:28] Asmita Mahajan: Here, then comes the second activation function, so it will be decided at the time of calling either Tanach or ReLU. Now, the last linear unit will connect the second hidden layer to the output layer, and the, outputs… raw outputs would be returned from here.
[00:50:46] Asmita Mahajan: So this forward function will pass, will propagate the inputs through the network using the forward pass, so…
[00:50:55] Asmita Mahajan: This function will do that. Now our train and evaluation function will come. We'll be defining that to evaluate our model and to train the model.
[00:51:05] Asmita Mahajan: So, we are passing activation function as the parameter, which will, contain values either Tanetch or ReLU, the hidden, layer,
[00:51:16] Asmita Mahajan: We have defined hidden as 128, because we have tuned it. Epochs is 300, and learning rate is 0.01, because the learning rate is also tuned.
[00:51:27] Asmita Mahajan: So now, creating an object of the class MLP, passing the size of the input layer, which is the number of features, and the activation function, which will be defined at the time of calling.
[00:51:41] Asmita Mahajan: So, binary cross and property loss will be used as the criteria to evaluate the loss. So, here, criterion will contain that function. The optimizer is ADAM, which we have already tuned using the learning rate, 0.01.
[00:51:56] Asmita Mahajan: Train and test losses, list will contain the train and test losses from each epoch. So, we have initialized the containers here.
[00:52:07] Asmita Mahajan: Now, iterating over the epochs and training the model. So, we will be opening our model in train mode with zero grads, which is, like, initializing the gradients to zero, or clearing the gradients from the previous iteration, so that there is no leakage of the gradients in the next iteration.
[00:52:25] Asmita Mahajan: Now, propagating the training dataset, using the forward pass through the model, through the network, and capturing the raw outputs and logists.
[00:52:37] Asmita Mahajan: Calculating or evaluating the loss for this particular iteration, the training loss, using the raw outputs and the vtrue labels of the train dataset, and using binary cross entropy, we'll be calculating the loss.
[00:52:50] Asmita Mahajan: Now, this calculated loss is, propagated backward, using the backward function, and then the weights would be updated using Atom Optimizer.
[00:53:00] Asmita Mahajan: And the loss is also appended to the train loss list using append function.
[00:53:05] Asmita Mahajan: For this particular iteration.
[00:53:08] Asmita Mahajan: And in a similar manner, we will calculate the test loss also, and we are now opening our model in evaluation mode, and telling the model that we don't want any gradients for this forward pass propagation, because,
[00:53:24] Asmita Mahajan: We will be using the, trained, weights.
[00:53:29] Asmita Mahajan: From this epoch.
[00:53:31] Asmita Mahajan: And we don't want to update the weights for the test set. So, we'll be passing the X test dataset in the model, and capturing the raw outputs in test logist.
[00:53:45] Asmita Mahajan: Calculating the test laws using BC entropy loss, using the function criterion, and then appending the test laws into the test loss list.
[00:53:57] Asmita Mahajan: So, in this manner, for each epoch, we'll be training our model, and the final model will be trained after 300 epochs.
[00:54:07] Asmita Mahajan: So, using that final model now, we will be making the predictions, and for that, we are opening the model in valuation mode. We don't want any gradients, or we don't want to update the weights, because we will be using the train weights to make the predictions.
[00:54:23] Asmita Mahajan: passing the X test dataset to the model and capturing the raw largest, or the raw outputs in largest. Using the sigmoid function, we are converting these raw outputs into probabilities, and using 0.5 threshold, we are converting the probabilities into hard predictions.
[00:54:44] Asmita Mahajan: So, the… this is the code to convert the tensor dataset into NumPy dataset, so we'll be getting Y true, labels, as a NumPy array, Y predicted labels as a NumPy array, and Y probabilities as NumPy array, to… so that we can, calculate the
[00:55:02] Asmita Mahajan: performance matrices.
[00:55:05] Asmita Mahajan: So, matrix will store the accuracy, precision, recall F1 score and ROC, for this training, for this particular training at that particular, activation function, and history will store the train and test loss for each epoch.
[00:55:23] Asmita Mahajan: And matrix and history dictionary are returned through this, function.
[00:55:29] Asmita Mahajan: So now in activations, we are defining the ReLU activation function. So we are defining a dictionary. So, we're telling if the
[00:55:39] Asmita Mahajan: if we'll be using… if we are passing railu as the activation function, so from the neural network, railu function will be used as the activation function. And if the name of my activation function is TanH, so from the neural network,
[00:55:55] Asmita Mahajan: library, TANH activation function should be used, at the activation, which is defined here.
[00:56:03] Asmita Mahajan: in the MLP class.
[00:56:06] Asmita Mahajan: So here, either Tanetch activation function will be called, or ReLU will be called, depending upon the name that we have given to… at the time of calling.
[00:56:16] Asmita Mahajan: And results in history's dictionary will store the performance metrics result for the particular activation function, and the train and test losses for each iteration for that particular activation function.
[00:56:30] Asmita Mahajan: So now we are iterating over the dictionary using activation.items. So, items will give us the…
[00:56:38] Asmita Mahajan: and the values, separately. So, name will contain the key.
[00:56:44] Asmita Mahajan: like this ReLU, and ACT will contain the function.
[00:56:50] Asmita Mahajan: When we are… so for the first iteration, the name will have… Name will be?
[00:57:00] Asmita Mahajan: Relu?
[00:57:02] Asmita Mahajan: And in a similar manner, ACT will be… Nn dot.
[00:57:09] Asmita Mahajan: Hello.
[00:57:11] Asmita Mahajan: So, this would be the value of name and act for the first iteration.
[00:57:16] Asmita Mahajan: So, now what we are doing is, we are calling the train and, evaluation function. We are telling that ACT function is equals to ACT, epox will be 300, and, learning rate is 0.01. So here, ACT is passed, which is the function itself.
[00:57:35] Asmita Mahajan: And results,
[00:57:37] Asmita Mahajan: at the results dictionary using the name as key. So, using railu as key, at this key, matrix dictionary will be stored as the value.
[00:57:48] Asmita Mahajan: And similarly, in the same manner, in the histories dictionary, using railua as key, history dictionary, which is the… which contains the train and test losses, will be stored in the value.
[00:58:01] Asmita Mahajan: So in this manner, we will, pass,
[00:58:05] Asmita Mahajan: The parameters to the train and evaluation function for both the
[00:58:10] Asmita Mahajan: activation functions, and we will then get…
[00:58:14] Asmita Mahajan: the results and history dictionary. And, here we are converting the result dictionary into, the data frame.
[00:58:24] Asmita Mahajan: Which is DFACT.
[00:58:26] Asmita Mahajan: So now… so that this data frame will be used to plot the bar plots.
[00:58:33] Asmita Mahajan: Now, using the history dictionary, we will be plotting the train and test losses for both the activation function, and the subtitle, the overall subtitle would be train and test loss versus activation function. Keys, so here, keys will have
[00:58:51] Asmita Mahajan: Value as… Raylu?
[00:58:58] Asmita Mahajan: Hun?
[00:59:00] Asmita Mahajan: Garn it.
[00:59:06] Asmita Mahajan: And n length of the keys would be 2 here.
[00:59:10] Asmita Mahajan: The column is, 3.
[00:59:13] Asmita Mahajan: We have 3 columns.
[00:59:17] Asmita Mahajan: Minimum of max columns, and… oh, sorry, we have… we will have… we will be having 2 columns, and rows will be 1, because the minimum of 3 and 2 is 2, and then we are just dividing N by columns, which is 2, which… and it will give us 1.
[00:59:33] Asmita Mahajan: Now, calling the subplots function to plot the subplots with the one row and two columns.
[00:59:40] Asmita Mahajan: iterating over the Keys, which is Railu and Tanish. We are getting, so here, we are just getting the location of the first subplot, and for that, in that subplot, a Relu would be, plotted, the, train and test loss of Reilu would be plotted.
[00:59:56] Asmita Mahajan: Using history dictionary, we are just extracting the dictionary from the histories and storing it in HIST, plotting the first curve, which is the train loss curve, and using this, we are plotting the second curve, which is the test loss curve.
[01:00:14] Asmita Mahajan: Setting the title to activation function equals to ReLU, and activation function it is equals to TanH, setting the X label to epoch and Y label to loss, and also displaying the legend. And we don't want any extra, empty subplots.
[01:00:29] Asmita Mahajan: This will set the overall title of the figure, and we'll be displaying the plots.
[01:00:37] Durga Toshniwal: So now, coming to the plots themselves, you can see that
[01:00:41] Durga Toshniwal: We tried to explore the use of different activation functions. Obviously, there could be many more in the choice also.
[01:00:49] Durga Toshniwal: So, as I mentioned yesterday, this is just exemplary. So, in your code, when you're trying to actually
[01:00:57] Durga Toshniwal: write similar code, then you can include more number of activation functions. For example, here we have written RELU and tanh. However, there could be others that you want to include and explore for your experiment.
[01:01:14] Durga Toshniwal: So you can see here that with the ReLU versus tanage, so in relu, you can see that the train loss is, much lesser than the test loss, and the test loss is very
[01:01:27] Durga Toshniwal: jittery out here, whereas if you look at the results of the activation function tanh, then you can observe
[01:01:37] Durga Toshniwal: Here, that, the losses are, quite smooth out, and the convergence is also quite fast, as you can see here, that at around 50 or 75, the convergence would have been achieved.
[01:01:54] Durga Toshniwal: And, not much of differences there between the train and the Teslas. So, as the graph for the 10H looks much smoother and better, we will choose, at this level, 10H as the
[01:02:08] Durga Toshniwal: Activation function.
[01:02:11] Durga Toshniwal: And then we will just also, have a look at
[01:02:15] Durga Toshniwal: What is the performance of the overall model? So, if you can go to that block… yeah. So, if you look at the final results of the activation function, tanh versus delu.
[01:02:28] Durga Toshniwal: What you can see here, although ReLU is also giving very good results, but there's a minor difference between ReLU and Tanet that you can see here.
[01:02:37] Durga Toshniwal: So, tanage is, like, 0.92, ReLU is 0.91.
[01:02:41] Durga Toshniwal: Then tanage for prediction, is 0.92, and, value is 0.91, like that. So they are very close. Actually.
[01:02:51] Durga Toshniwal: given any freedom, you could have chosen ReLU also. However, since we are just trying to illustrate, the better, model performance, so we will choose Tanish, because it's slightly better than ReLU, okay, to include in our model.
[01:03:10] Durga Toshniwal: Now you can go to the upper block and explain the code.
[01:03:16] Asmita Mahajan: Yes, ma'am.
[01:03:17] Asmita Mahajan: So, now we'll be plotting the, performance matrices, to compare the ReLU and Tanish, activation function. So, using DF results, which we have created, using the result dictionary, we are passing it to this, function, and then,
[01:03:37] Asmita Mahajan: The number of rules, the number of columns, and what will be the overall title of the figure.
[01:03:43] Asmita Mahajan: We are extracting the, name of the columns from the DF function, which is the matrices, itself, for example, accuracy.
[01:03:55] Asmita Mahajan: Precision Recall F1 score and AUC score. And the index labels, so in our case, that would be relevant and edge, so that would be extracted from the data frame and passed to the index labels.
[01:04:10] Asmita Mahajan: Calling the subplot function, and passing the rows and columns to the function, so we will be plotting the overall figure and the subplots.
[01:04:19] Asmita Mahajan: And then flattening the axis 2D array to convert it into a 1D list.
[01:04:28] Asmita Mahajan: And so, here we are iterating over the matrices and capturing the metric name and the index at which the metric is present. So, for accuracy, for the first metric, which is the accuracy, so a metric will contain accuracy, and index would be 0.
[01:04:46] Asmita Mahajan: And then, getting the, location of the first subplot, and also extracting the values, from the data frame, and storing that in Val's list.
[01:04:58] Asmita Mahajan: Now, calling the bar, and the buy,
[01:05:02] Asmita Mahajan: sorry, the X axis will contain the name of the labels, which is Reelu and Tanage. The Y-axis will contain the…
[01:05:10] Asmita Mahajan: We'll have the… so, at the y-axis, we'll have the values, and the edge color is black.
[01:05:17] Asmita Mahajan: So, this will set the subplot title, which is the name of the metric.
[01:05:24] Asmita Mahajan: The X label is activation function, Y label would be metric, and setting the limit to 0 to 1 for all the subplots. Also rotating the X text to 45 degree for presentation purposes, and also displaying the
[01:05:38] Asmita Mahajan: Now, using this code, we will be displaying, or yes, we'll be displaying the values associated with each bar.
[01:05:47] Asmita Mahajan: So, X and, so the first two values in this text function will tell us where the value should be presented, or where the value should be displayed, the location, how many decimal points we want, the horizontal alignment and the vertical alignment, and the font size.
[01:06:05] Asmita Mahajan: And we don't want any extra empty subplots in our overall figure, so that this will off that.
[01:06:13] Asmita Mahajan: And the overall figure title will be displayed using this code subtitle, and passing the subtitle to it, which we have defined here at the time of defining the function, and we are displaying the plot, and also displaying the data frame.
[01:06:32] Asmita Mahajan: So these are the plots.
[01:06:37] Asmita Mahajan: So now that we have tuned the, activation function, so, use… so now we will be using optimizer, Adam Optimizer, the learning rate, which is already tuned, the hidden layer sizes for both the layers.
[01:06:54] Asmita Mahajan: and the activation function. These are all fixed. Now we will see or tune the epochs, like how, at the time of training, the training function, how many times the training function should be called, so that, our model is trained. So for that, we will be now tuning the epochs, and the value of epochs will go
[01:07:16] Asmita Mahajan: like 50, 100, 200, 300, or 500. So, for these 5 values, we'll be testing and evaluating the performance of the MLP model.
[01:07:27] Asmita Mahajan: So for that, we are defining our class MLP, and, passing the parameters, the input dimension and the hidden layer dimensions, calling, or defining our network in a sequential manner with the first linear unit, which will connect the input
[01:07:44] Asmita Mahajan: Layer to the first hidden layer, then we'll have the activation function, then the second linear unit, which will connect the first hidden layer to the second hidden layer.
[01:07:55] Asmita Mahajan: And 256 and hidden 2 will have value as 128, because the size, are already tuned for these hidden layers. Then comes the second activation function.
[01:08:07] Asmita Mahajan: And then the third linear unit, which will connect this second header layer to the output. And we will be returning the raw outputs.
[01:08:15] Asmita Mahajan: hair.
[01:08:18] Asmita Mahajan: With one neuron.
[01:08:19] Asmita Mahajan: Because we have… we are just doing, binary classification.
[01:08:24] Asmita Mahajan: Now, the inputs, the input to the network is passed or propagated, through the network using the forward pass, and this function forward will be doing that, which we have defined here. So this completes the, class MLP. Now, we'll come to the train and evaluation.
[01:08:43] Asmita Mahajan: because we now have to check the model performance for the number of epochs, which we will be defining at the time of calling the function. So, epochs will tell us how many epochs, which is 50, 100, 200, 300, or 500, so that we will be passing it at the time of calling.
[01:09:02] Asmita Mahajan: And learning rate would be 0.01.
[01:09:06] Asmita Mahajan: Now, we are calling, the… we are, creating an object of the MLP class, and passing the parameters, which is the size of the input layer, which is the number of features, and the size of the hidden layer, which is 128.
[01:09:20] Asmita Mahajan: The cross entropy loss function would be used to evaluate the losses, so criterion will contain the object for that.
[01:09:31] Asmita Mahajan: Atom Optimizer is used using LR.
[01:09:35] Asmita Mahajan: activation function, sorry, LR learning rate, which is 0.01. Train and test losses list, are initialized here to store the train and test loss for each epoch.
[01:09:47] Asmita Mahajan: And iterating over the number of epochs, which will be… we will be defining at the time of calling. We are opening our model, in train mode, clearing the gradients, or initializing the gradients to 0.
[01:10:01] Asmita Mahajan: Also, we are performing now the forward pass using the X-Train dataset, and, capturing the raw outputs in logists.
[01:10:12] Asmita Mahajan: Now, evaluating the loss using binary loss entropy loss, and propagating the loss backward in the network using backward function, and also updating the
[01:10:21] Asmita Mahajan: weights using Atom Optimizer.
[01:10:24] Asmita Mahajan: So the evaluated loss, or the calculated loss, is now added to the train loss.
[01:10:31] Asmita Mahajan: list using append function, and in a similar manner for this particular iteration, we'll be, evaluating, or we'll be also calculating the test loss. So, we are opening the model in evaluation mode with no gradients, because we want to use learned, weights.
[01:10:51] Asmita Mahajan: So we are now propagating the test dataset in the model, and capturing the raw outputs and test logist. Using those test logists and the vitrue test labels, we'll be calculating the loss.
[01:11:05] Asmita Mahajan: And that loss is now appended to the test loss dictionary.
[01:11:09] Asmita Mahajan: For this particular epoch. And after we have trained the model for a number of epochs, or we have run the model for this particular number of epochs, we will now be evaluating the model
[01:11:26] Asmita Mahajan: On the test set data set to finally, evaluate or calculate the matrices.
[01:11:32] Asmita Mahajan: So what we are doing is, we are opening the model in evaluation mode with no gradients, and then passing the X test dataset to the model.
[01:11:42] Asmita Mahajan: Capturing the raw outputs in logists, using the sigmoid activation function, converting the logists into probabilities, and using 0.5 as threshold, we are converting those probabilities into predictions.
[01:11:56] Asmita Mahajan: And, this code will just, convert our tensor dataset into NumPy array, so why true will contain the true predictions in a NumPy 1D array, because we are also flattening it after converting it into NumPy array, so it will be a 1D array.
[01:12:15] Asmita Mahajan: We have the Y predictions, which we just got, here, in Ypredict, 1DRA and Y probabilities associated with the predictions.
[01:12:26] Asmita Mahajan: Now, we will be calculating the matrices, which is the accuracy, precision, recall, F1 score, and ROC score, and storing them in a dictionary, matrices, metrics.
[01:12:37] Asmita Mahajan: History will contain the train and test laws for each epoch, and we'll be returning the matrix and history dictionary, after, from this function.
[01:12:52] Asmita Mahajan: So, we are now defining epoch count list, which will contain the number of epochs for which we have to train the model.
[01:13:01] Asmita Mahajan: History… result and history dictionary will, store the performance metrics and the train and test laws for each particular epoch.
[01:13:11] Asmita Mahajan: So, iterating over the epochs, we are now calling the train and eval function for that particular epoch, and the learning rate is 0.01. So…
[01:13:22] Asmita Mahajan: in our first iteration, 50 would be passed as the number of epochs, so the model would be trained 50 times. In the second iteration, the model would be trained 100 times. In the third iteration, the model would be trained, 200 times, and so on, so forth.
[01:13:37] Asmita Mahajan: And, the performance matrices for, the first iteration, like, we have passed 50, and the model is trained 50 times, so for that, the performance… the performance of the model is stored in the result.
[01:13:52] Asmita Mahajan: dictionary, and also the train and test laws associated with that, is stored in history's dictionary.
[01:13:59] Asmita Mahajan: Now, using the result dictionary, we are converting it into a data frame, in this code.
[01:14:07] Asmita Mahajan: Now, we will be plotting the train curve and test curve to see, the, convergence, how the model has converged.
[01:14:17] Asmita Mahajan: To see that and to evaluate that, we'll be using histories, because it contains the train and test losses for each epoch that we have.
[01:14:31] Asmita Mahajan: Keys will contain… so here, keys would be, like, 50… 100.
[01:14:41] Asmita Mahajan: 100.
[01:14:44] Asmita Mahajan: 300 and 500. So, this will be the value of the keys list. So, N is 5 here, because we have 5… we have defined 5 different, epoch sizes, or a number of epochs.
[01:15:00] Asmita Mahajan: So, columns is 3.
[01:15:03] Asmita Mahajan: And number of rows would be 2.
[01:15:07] Asmita Mahajan: So, using the subplot, we are now plotting the train and test loss curve, loss curves for each epoch. Enumerating over the keys, we are taking the first key, which is 50, and then plotting the train and test loss for that particular epoch.
[01:15:24] Asmita Mahajan: value. So here, the first plot will plot the train curve, the second plot will plot the test curve, the title would be set as epox is equals to 50, epox is equals to 100.
[01:15:37] Asmita Mahajan: X label is epoch, Y label is loss, and legend would be displayed, and we don't want any extra or empty plots.
[01:15:46] Asmita Mahajan: The overall figure is, figure title is set, and the plots are displayed using the plot.show function.
[01:15:55] Durga Toshniwal: Okay, and you can see here, I think there was a question, first of all, that we have finalized on 10H, then why are we using drilling?
[01:16:05] Durga Toshniwal: I think value has been, used here.
[01:16:08] Durga Toshniwal: So we could do that, we could… actually, if 10H is finalized, 10H should be.
[01:16:14] Durga Toshniwal: used here.
[01:16:16] Durga Toshniwal: So, you're using ReLU at the first layer, and ReLU also at the second layer.
[01:16:21] Asmita Mahajan: So probably at the second layer, we can go with Tanach.
[01:16:25] Durga Toshniwal: And rebuild the model and see the results that we can do shortly from now. Let me first…
[01:16:31] Durga Toshniwal: Okay, you can change it to Tanish and rerun it. I think that will be better.
[01:16:36] Asmita Mahajan: Oh, yes, ma'am, I have to run it from the start.
[01:16:39] Durga Toshniwal: Yeah, I think it wouldn't take too much of a time.
[01:16:43] Asmita Mahajan: Ms.
[01:16:47] Durga Toshniwal: So, we'll just rerun.
[01:16:49] Durga Toshniwal: So, the first activation function, I mean, the one at the first layer is still relevant.
[01:16:55] Durga Toshniwal: And at the second layer, we had iterated and found out that 10H is giving better results. So, in the model, I think,
[01:17:04] Durga Toshniwal: Tanet should be there, which is what is included now.
[01:17:07] Durga Toshniwal: And, the code is just being redone.
[01:17:13] Durga Toshniwal: For the first layer, we had found out ReLU to be the better one, and we included ReLU only in the model. So, you can visualize the model.
[01:17:20] Durga Toshniwal: Having the input layer, then you are having, The first hidden layer…
[01:17:26] Durga Toshniwal: first hidden layer with the activation function value. Then you imagine the second subsequent layer.
[01:17:34] Durga Toshniwal: And the subsequent layer is having, 128 or whatever number of neurons, and the activation function is damage. And then after that, you will have the output, which is a… because it's a binary class problem, so your output will be one neuron only.
[01:17:52] Durga Toshniwal: Which will show you what is the final outcome.
[01:17:56] Durga Toshniwal: So I think with just a…
[01:17:59] Durga Toshniwal: Few more seconds, and, we'll be done with the…
[01:18:03] Durga Toshniwal: Tuning, with the… this model generation and the results, and then we'll see.
[01:18:09] Durga Toshniwal: How the outcome is coming with… with the second activation function as 10H.
[01:18:18] Durga Toshniwal: In the meantime, any other questions, anyone?
[01:18:42] Durga Toshniwal: Okay, it's still running.
[01:19:30] Sonam Manwal: Hello, ma'am?
[01:19:32] Durga Toshniwal: Yes.
[01:19:33] Sonam Manwal: Ma'am, this code is completely running on CPU,
[01:19:36] Sonam Manwal: If we have to run it on GP also, then, like, code-wise, we have to change something, or it might be my environment issue.
[01:19:44] Durga Toshniwal: I think the… wherever required, wherever you could have run it on GPU, it's already in PyTorch. So PyTorch, by default, will support GPU.
[01:19:54] Durga Toshniwal: And I think then, we are bringing from GPU to CPU also. So it shouldn't give any issue, as far as I understand.
[01:20:03] Durga Toshniwal: You do have a GPU at your end?
[01:20:06] Sonam Manwal: Yes, yes ma'am.
[01:20:08] Durga Toshniwal: Okay, see, what I suggest is you please try running on the GPU. I think it should work, and that's the reason why it's done in PyTorch. However, if you face any issues, please come back to us and let us see. I think this shouldn't be giving any problem on GPU.
[01:20:25] Sonam Manwal: it is not giving error, but, like, if I will go to this, task Explorer.
[01:20:33] Sonam Manwal: task manager, so there I can see a complete load is on CPU, not… GPU is not getting used when I was exiting this complete file.
[01:20:43] Durga Toshniwal: Okay.
[01:20:44] Durga Toshniwal: Okay, let me look into it, and I'll come back to you, on that.
[01:20:49] Sonam Manwal: General.
[01:20:51] Laxmi SAHU: Oh, ma'am, would this file, work on Jupyter notebook?
[01:20:57] Laxmi SAHU: Like, if I run on the.
[01:20:58] Durga Toshniwal: or not.
[01:20:59] Laxmi SAHU: Jupiter.
[01:21:00] Durga Toshniwal: It should run. I think, Jupyter and Colab are quite,
[01:21:05] Durga Toshniwal: Similar only, I mean, the… just the environment that has been provided, you can try, it should run.
[01:21:11] Durga Toshniwal: On Jupiter.
[01:21:13] Durga Toshniwal: Okay, I think we are done with,
[01:21:16] Durga Toshniwal: The losses, okay, we are at the epochs.
[01:21:20] Durga Toshniwal: Yeah, fine. So now we are having, Tanish as the… As a choice, so…
[01:21:28] Durga Toshniwal: We had chosen, for our second activation function.
[01:21:32] Durga Toshniwal: And, now we can see here,
[01:21:36] Durga Toshniwal: that with epoch equal to 50. So now, remember what is epoch? Epoch is the number of times we are showing the entire data.
[01:21:46] Durga Toshniwal: To our, to our model.
[01:21:50] Durga Toshniwal: So we can see here that, with epoch equal to 50, the convergence, the train and test losses are quite stable, however, the convergence is…
[01:22:04] Durga Toshniwal: happening, hasn't actually happened also. They are… they are not, converged. Probably. It will take more number of epochs, so 50 may not be the best solution.
[01:22:16] Durga Toshniwal: With 100 also, probably, again, I think the losses are not stabilized.
[01:22:25] Durga Toshniwal: But 200, probably, around maybe 75 or 70.
[01:22:34] Durga Toshniwal: The losses seem to be…
[01:22:37] Durga Toshniwal: Stabilized, though again, they show jitter.
[01:22:40] Durga Toshniwal: With epochs equal to 300 or 500, the results are just very similar only.
[01:22:47] Durga Toshniwal: Only thing to note here is…
[01:22:49] Durga Toshniwal: I think, everything is just the same.
[01:22:53] Durga Toshniwal: So if we actually look at epoch equal to 500 and take a portion of it, up to 300, what you'll get is the previous plot. So the behavior is very much similar, only thing is that I still cannot see
[01:23:08] Durga Toshniwal: The final convergence at 500 also. Let's have a look at the tabular results.
[01:23:16] Durga Toshniwal: They will… they might help us to decide what is the best out of what we are having.
[01:23:23] Durga Toshniwal: So, if we look at the performance of the model.
[01:23:27] Durga Toshniwal: I'll just scroll up a little bit, yeah.
[01:23:30] Durga Toshniwal: So, if you look at the performance of the model, at 500 epochs, like URC is coming out to be.
[01:23:38] Durga Toshniwal: 0.93. Precision is 0.94. Recall is 0.91.
[01:23:45] Durga Toshniwal: And then F1 score 0.92, and AUC 0.98.
[01:23:51] Durga Toshniwal: And again, there's very little difference between 300 and 500 epochs.
[01:23:56] Durga Toshniwal: So, if you'll go back to that, those graphs…
[01:24:01] Durga Toshniwal: So, the best performance so far is coming with epochs equal to 500.
[01:24:07] Durga Toshniwal: However, there's very minute difference between 300 and 500.
[01:24:12] Durga Toshniwal: So if we do have compute support, we can go with EPOC 500, and I would rather suggest that you… you could even raise the number of epochs beyond 500 to see where the results are stabilized, because I think the difference between 300 and 500 is not much. So it's kind of on the verge of
[01:24:32] Durga Toshniwal: stabilizing. Probably, if we could increase the number of epochs.
[01:24:37] Durga Toshniwal: I suggest, Asmita, can you please increase the number of epochs?
[01:24:41] Durga Toshniwal: say to, 600, and then to maybe 800 in the list, and just redrun this part, and let's see how the results vary. I think that with 500, around 500,
[01:24:55] Durga Toshniwal: It might have got stabilized.
[01:24:58] Durga Toshniwal: We'll just rerun this portion for all of you, and…
[01:25:01] Durga Toshniwal: then we can see what is the impact. You all can also do… you all can take smaller steps, like you could have 50, 100, 150, all that.
[01:25:11] Durga Toshniwal: And then see the… the… The specific, outcomes.
[01:25:17] Durga Toshniwal: Here, we are having some exemplary numbers that we have chosen. I'm still adding a few more to see whether the results have got stabilized or not.
[01:25:28] Durga Toshniwal: Though, so far, epochs equal to 500 is giving the best results. As I already pointed out to you.
[01:25:35] Durga Toshniwal: That it's not that we go on increasing the number of epochs, that the results will keep on improving. At some point in time, they'll become saturated, the results will become saturated, and there'll be no game, which is the point beyond which we don't increase the number of epochs.
[01:25:52] Durga Toshniwal: So, we'll allow this to run.
[01:25:55] Durga Toshniwal: I think it will take a little while, because epochs have increased accordingly.
[01:26:01] Durga Toshniwal: Some compute time will be required.
[01:26:09] Durga Toshniwal: So, in the meantime, I can see, I have noted…
[01:26:18] Durga Toshniwal: That, there are a total of…
[01:26:21] Durga Toshniwal: 121 participants in this batch.
[01:26:26] Durga Toshniwal: So, accordingly, we can decide the team size for the number of groups.
[01:26:32] Durga Toshniwal: That is the final number, 120mm, so I'll be shortly sharing with you the team size also.
[01:26:42] Durga Toshniwal: I think it's still running, so we'll have to wait slightly for more amount of time.
[01:27:19] Durga Toshniwal: In the meantime, the team size that we can have for this, for all of you could be around 8.
[01:27:26] Durga Toshniwal: She could have around 8 people per team.
[01:27:29] Durga Toshniwal: Now you can start to…
[01:27:31] Durga Toshniwal: Making your teams, if you haven't started yet, and maybe in a week or so, we'll discuss the teams also.
[01:27:42] Durga Toshniwal: Okay, please note the team size, which is 8.
[01:27:46] Durga Toshniwal: Okay, it's still running. Maybe we shouldn't have started that.
[01:27:51] Durga Toshniwal: I was a little curious to show the…
[01:27:56] Durga Toshniwal: To show the saturation, that's why…
[01:28:00] Durga Toshniwal: We started, but it's going on and on.
[01:28:09] Asmita Mahajan: Also, ma'am, with the Realu as the activation function, as we have changed it to tanage.
[01:28:14] Durga Toshniwal: Yeah.
[01:28:15] Asmita Mahajan: the convergence was a little smoother using Raylu. So, the jitter happened because we have used Tanish at the second hidden layer.
[01:28:28] Durga Toshniwal: Yeah, that's true, and the tanage also provides a bigger range of values from minus 1 to plus 1.
[01:28:37] Durga Toshniwal: Okay.
[01:28:39] Durga Toshniwal: I think we are true.
[01:28:42] Durga Toshniwal: And can we have a look at… I think, probably.
[01:28:49] Durga Toshniwal: 300 would be the stability point. That's what I'm seeing, because the difference Can we come to the…
[01:28:58] Durga Toshniwal: Has the other next section also run already?
[01:29:08] Durga Toshniwal: Okay, you can enlarge it slightly.
[01:29:14] Durga Toshniwal: Yeah, now this will be interesting to see.
[01:29:18] Durga Toshniwal: So, now we were earlier deciding for 500. You can see very nicely, that if we look at 600,
[01:29:28] Durga Toshniwal: The accuracy is almost the same.
[01:29:31] Durga Toshniwal: The precision is also similar only.
[01:29:34] Durga Toshniwal: Recall also similar, because if you see, number of epochs as 500, accuracy is 0.92, precision is 0.918, which is 0.92 only.
[01:29:46] Durga Toshniwal: Recall is 0.907, which is 0.91. F1 score is 0.91.
[01:29:52] Durga Toshniwal: 2, so it is 0.91, and AUC is 0.978, which is 0.98.
[01:29:58] Durga Toshniwal: Now, if you look at epochs equal to 600,
[01:30:02] Durga Toshniwal: Then you see that the accuracy is 0.92 again.
[01:30:06] Durga Toshniwal: Precision is 0.92 again.
[01:30:08] Durga Toshniwal: Recall is 0.91 again, F1 score is 0.91 again, and AUC is 0.98 again. At 800 also, it is 0.92, 0.91, 0.91, 0.91.
[01:30:22] Durga Toshniwal: And 0.98. So, you can assume that the model has saturated with epochs equal to 500. And I think increasing beyond 500 is not going to do any benefit, and we can choose 500 as the final number of epochs. Okay, this is what I wanted to illustrate. Luckily, we did
[01:30:42] Durga Toshniwal: we did have convergence with 500, so we'll proceed with…
[01:30:46] Durga Toshniwal: Number of epochs as 500. Okay, so now you can proceed to the next section.
[01:30:58] Durga Toshniwal: Yeah, you can proceed. Sivanch, I will come back to you towards the end. Let us just finish this part.
[01:31:05] Shivansh Sharma: Okay.
[01:31:06] Asmita Mahajan: So, we'll be tuning the batch size now. So, previously, what we were doing is for… while training the model at each epoch, at every epoch, we are, for… we are, like, propagating the whole training dataset through the network. So,
[01:31:27] Asmita Mahajan: like, because we didn't have the data set as large, like, in the real-world scenarios, we will see that the data, would have many number of rows. It couldn't be in thousands, but it could be in millions also. So to overcome that, to overcome the training of the model in a better manner.
[01:31:51] Asmita Mahajan: We pass batches of the dataset to the model, and then see how the model performs.
[01:31:58] Asmita Mahajan: So, for that, we will be tuning the batch size, and here we are just… we are, like, keeping the batch size to 16, 32, and 64. So, for these 3 batch sizes, we will be seeing how the model performs, the MLP model performs.
[01:32:16] Asmita Mahajan: So, what we are doing is, we are creating the class.
[01:32:20] Asmita Mahajan: MLP class with input dimension and the hidden layer dimension. We are creating the net
[01:32:28] Asmita Mahajan: work in a sequential manner, and using linear unit, connecting the input to the first hidden layer, then the activation function, ReLU, then the second, linear unit, which will connect the first hidden layer to the second hidden layer.
[01:32:43] Asmita Mahajan: defining the activation function, and the third, linear unit, which will connect the second hidden layer to the, outputs.
[01:32:53] Asmita Mahajan: Now, using the forward pass function, or, to propagate the inputs, for, through the network, we are using forward function, and defining it here.
[01:33:05] Asmita Mahajan: The best learning rate is 0.01, which we have already tuned, so we are defining it here.
[01:33:12] Asmita Mahajan: Now, the train and evaluation function will be defined, where we are passing, batch size as the parameter, the number of epochs. So here, 200, it is written at the time of defining, but, as we have already set it to 500, at the time of calling, we are, so for,
[01:33:30] Asmita Mahajan: computation purposes, we are defining our epochs to be 20, because at each epoch, for
[01:33:37] Asmita Mahajan: 3 batch sizes, our model will be trained.
[01:33:41] Asmita Mahajan: So, that's why, here we have defined our epochs to be 20, so that, it doesn't take much time to train the model. So, for exemplary purposes, we have done that. Otherwise, to train the full model, we will be using 500 as the number of epochs.
[01:33:58] Asmita Mahajan: Now, weight decay is also, used here, so,
[01:34:02] Asmita Mahajan: So, weight decay will just add a regularization parameter at the time of optimizing or updating the weights. So, to overcome the problem of overfitting the model, we, add a regularization parameter, in the loss function. So, firstly, we were reducing the loss, the overall loss, that has been,
[01:34:26] Asmita Mahajan: evaluated at the time of training. But now, using weight decay, we are just adding one more parameter to the loss, and now we are,
[01:34:35] Asmita Mahajan: reducing the… this loss. The loss, which is the actual loss, plus the regularization loss. But here, as weight decay is zero, so we are not adding any regularization parameter to our loss function. So, we are only reducing the overall loss.
[01:34:53] Asmita Mahajan: Of the model.
[01:34:54] Asmita Mahajan: and the, LR function is used, LR function… sorry, not LR function, the LR value, as 0.01 is used, is passed to this train and evaluation function.
[01:35:06] Asmita Mahajan: Now, we are creating the object of the MLP class, passing the, number of input,
[01:35:15] Asmita Mahajan: number of neurons in the input layer, which will be equal to the number of features, the hidden layer size, which is 128. Bandary cross entropy loss will be used to calculate or evaluate the loss, and passing it to criterion.
[01:35:30] Asmita Mahajan: variable, atom optimization, optimizer is used with LR as the learning rate, and weight decay is zero, so there will be no regularization parameter here.
[01:35:41] Asmita Mahajan: Now, using, data loader, we will be, dividing our Full data set into batches.
[01:35:50] Asmita Mahajan: So, what we are doing is, we are firstly, combining, because we have separated the features and the labels into two different, datasets, or into two different variables using the X-Train tensor and Y-Train tensor. So, X-Train tensor will contain all the features and the values for that features, and Y-Train tensor
[01:36:13] Asmita Mahajan: contains the labels, or the target variable. But, what we will be doing here is, using Tensor dataset, we are combining it into one dataset and, like, storing it in TrainDS, and using Data Loader.
[01:36:31] Asmita Mahajan: We are, Dividing this whole dataset into
[01:36:35] Asmita Mahajan: batches using the batch size. So, batch size will just tell us how many samples will be there in one batch. So, 16 will tell that there will be 16 samples in one batch, and the whole dataset is divided into batches of size 16.
[01:36:54] Asmita Mahajan: Shuffle is true, which means that for each batch, we need
[01:36:58] Asmita Mahajan: Like, we can randomly choose from the dataset, from the whole dataset. It doesn't have to be in a sequential manner, which means the first 16 samples are going into the first batch, and the next 16 samples are going into the next patch. No, we don't want it like that. We can have the 16 samples from the whole dataset chosen randomly.
[01:37:19] Asmita Mahajan: Drop last is fast. So, it could be a case that my, whole data set is not perfectly divided. Like, it could happen that the last portion of the dataset can have less number of samples compared to the batch size. So we don't want to… so by default, the data loader drops that batch if we have… if we have an incomplete batch, but we don't want that.
[01:37:44] Asmita Mahajan: We want those samples to be there, and we want those samples, we want the model to be trained on those samples also from the training set, so we are not dropping that. That's why we have set it to false.
[01:37:57] Asmita Mahajan: So, the data loader will just take the whole training data set and divide it into the batches of size, batch size. So, 16 for the first batch size, which is 16, every batch will have 16 samples from the dataset.
[01:38:15] Asmita Mahajan: And trainloader is an iteratable object which will be iterated over in the training, process.
[01:38:25] Asmita Mahajan: So now, these, train and test loss containers will contain the train and test losses from each epoch. So, we are iterating over the epochs,
[01:38:37] Asmita Mahajan: Opening our model in, train mode.
[01:38:41] Asmita Mahajan: Now, what we are doing is, we are, for each epoch, we are also calculating a total loss. So, total loss will be calculated, so now what is happening is, for every batch, there will be, we'll be calculating a loss. And, in a similar manner, if we have, if the dataset is divided into 100 batches, so for every batch, for the 100 batches.
[01:39:04] Asmita Mahajan: We will have… Loss associated with each batch. So, total loss will, we will be calculating
[01:39:10] Asmita Mahajan: the total loss of the training data set, which is, which was, of 14,000 rows. So, for 14,000 rows, we want to calculate the loss. So, this total loss will contain that, and total N is, the number of samples in the training, set. So, we are setting it to zero first, because we'll be calculating it, in this,
[01:39:35] Asmita Mahajan: The further for loop.
[01:39:39] Asmita Mahajan: So what we are doing is, we are, from the train loader, we are extracting the X batch and the Y batch. So the X batch will contain the features, the 14 features, the values of the 14 features, and Y batch… YB, which is, is a… which is represented as Y batch.
[01:39:57] Asmita Mahajan: It will contain the labels corresponding to those features.
[01:40:00] Asmita Mahajan: So, because trainloader will, it will be a… so, here you can see, it is a data…
[01:40:08] Asmita Mahajan: tuple. It is a list of lists.
[01:40:10] Asmita Mahajan: trainloader, where each tuple will have, the features and also the labels associated with the features. So we are extracting them in separate, variables here.
[01:40:24] Asmita Mahajan: Now, we are just, initializing our gradients to 0.
[01:40:30] Asmita Mahajan: Calling the model, or propagating the particular batch into the model, and capturing the largest, which is the raw output from the model.
[01:40:43] Asmita Mahajan: evaluating the loss for that particular batch. So, here you can see we have, propagated X batch to the model, which is the 16 samples, the first 16 samples to the model, and corresponding to those 16 samples, we'll be having the Y labels.
[01:40:58] Asmita Mahajan: And that's why we are using YB in the criterion function now, instead of Y train dataset, or Y-train tensor.
[01:41:08] Asmita Mahajan: Now, propagating this calculated loss backward, using the back propagation, we are use… we are doing it here.
[01:41:15] Asmita Mahajan: Updating the, weights for this particular batch.
[01:41:20] Asmita Mahajan: Now, total loss is calculated using, this, so,
[01:41:25] Asmita Mahajan: how it is calculated, it will be, like, total loss plus… so the calc… the previous total loss. So previously, we have initialized it to 0 for this particular batch, so it will be 0 here, plus the loss that we have calculated for this particular batch into the size of the batch.
[01:41:45] Asmita Mahajan: Which is, 16 here.
[01:41:47] Asmita Mahajan: So, for the first patch, the size would be 16, and it could be possible that for the last batch, the size is not 16, it can be less than 16, so that's why we are calculating the size for each batch, and we are not hard-coding it here as 16, because the last size can have a different, batch size.
[01:42:06] Asmita Mahajan: So, the size of the batch, which is the batch size. So, in this manner, we are calculating the total loss, and also total number of,
[01:42:16] Asmita Mahajan: samples. So, total N will be calculated as the previous total N plus the size of that particular batch.
[01:42:25] Asmita Mahajan: And, after this, the overall training loss is calculated using total loss, which is calculated here, divided by the total number of samples in that particular batch.
[01:42:39] Asmita Mahajan: So, in this manner, the train loss is, calculated, and this value is then appended to the train loss list.
[01:42:50] Asmita Mahajan: Now, in a similar manner, we'll be calculating the test loss for the full batch. So, training… for the training, we will be dividing our data set into batches, but for testing, we are not dividing it into batches, we are taking the full batch, and then, you know, evaluating our test loss.
[01:43:09] Asmita Mahajan: So we are opening the model in evaluation mode. We don't want any gradients to be updated, so that's why no grad is there.
[01:43:17] Asmita Mahajan: Now, passing the whole test dataset into the model, and capturing the test logists.
[01:43:24] Asmita Mahajan: Evaluating the test laws, and then appending the test laws to the test laws, List.
[01:43:33] Asmita Mahajan: So, after the training, so after we have, trained the model.
[01:43:37] Asmita Mahajan: for a particular number of epochs, and for each batch, now we will be making the final predictions using the evaluation function. So we are opening the model in evaluation mode. We don't want any gradients, we want the learned gradients to be used to test the dataset.
[01:43:55] Asmita Mahajan: Passing the test dataset to the model and capturing the raw outputs, converting the raw outputs to probabilities using the sigmoid activation function, and then converting the probabilities to predictions using the 0.5 as the threshold.
[01:44:15] Asmita Mahajan: now, we are now converting the tensor, datasets into NumPy array. So, Y2 will be the… will contain the two labels of the dataset in a NumPy array. Y predictions, the predictions that we have made in this step, so the predictions will be converted from tensor to the NumPy.
[01:44:34] Asmita Mahajan: Array, and probabilities also.
[01:44:37] Asmita Mahajan: We'll be evaluating the matrices, which is the accuracy, precision, recall, F1, and AOC score, and also capturing the history in history dictionary, which will contain the train loss and test loss.
[01:44:51] Asmita Mahajan: And we are returning these, both the dictionaries, to the… After the function is returned.
[01:44:58] Asmita Mahajan: So, we are defining here the batch sizes, so we'll be using 16, 32, and 64 as different batch sizes, and results and history dictionary will contain the performance matrices and train and test loss for each batch size.
[01:45:12] Asmita Mahajan: So, iterating over the batch sizes and calling the train and evaluation function on the batch sizes with 20 perks, because otherwise the model will run indefinitely, and we don't have that computation time.
[01:45:25] Asmita Mahajan: And weight decay is 0, so it is telling that we are not using any regularization parameter at the time of optimizing the weights.
[01:45:35] Asmita Mahajan: Also, the hair results for that particular, batch is stored.
[01:45:41] Asmita Mahajan: And the history, the train and test losses are also stored. Using the result dictionary, we are now converting it to a data frame. And now we'll be plotting the histories, which is the train and test losses for each batch size, with maximum column history, and the subtitle as train test loss per batch size.
[01:46:01] Asmita Mahajan: Keys will contain, keys will have 3 values, 16, 32, and 64. Length of the keys would be… n will contain 3, so the columns will be 3, and the rows will be,
[01:46:14] Asmita Mahajan: So we'll have one row for that.
[01:46:18] Asmita Mahajan: Now, we are calling the subplot function on the number of, on, one number of row for, like, we are creating a grid, of one row and three columns, and the overall figure and axis will be stored here.
[01:46:31] Asmita Mahajan: Enumerating over the keys, or iterating over the keys, which is 16, 32, and 64.
[01:46:36] Asmita Mahajan: We'll be plotting the subplots. So, our first subplot would be of 16 batch size, and train and test loss are extracted for that particular batch size, and are plotted in two different, like, at two different curves.
[01:46:52] Asmita Mahajan: We are setting the title using this setTitle function, batch size 16, batch size 32, batch size 64, setting the label to epoch, Y label to loss, displaying the legend, and we want… we don't want any empty subplot, and now setting the overall figure title, and displaying the plots.
[01:47:19] Durga Toshniwal: Okay, so we can see here.
[01:47:22] Durga Toshniwal: That, the results with batch size 16, 32, and 64 are in front of us.
[01:47:29] Durga Toshniwal: You can see that, with a batch size of 32,
[01:47:35] Durga Toshniwal: The training and test losses. So, test, test loss doesn't see, smoothness, it is quite jittery, it is varying.
[01:47:47] Durga Toshniwal: Similarly, with batch size 64 also, it's not stabilized, the test loss, so probably the most smooth results are coming with a bad size of 16.
[01:47:59] Durga Toshniwal: If we look at, the results, in form of numbers, if you could just scroll down…
[01:48:06] Durga Toshniwal: So, here you see that the results for…
[01:48:10] Durga Toshniwal: The bed size of 16 versus 32 versus 64.
[01:48:15] Durga Toshniwal: Now, it's interesting to note that if you look at accuracy, then it's 0.92 for both 16 and 64, for 32, it is 0.91. I'm rounding to the second digit, decimal.
[01:48:29] Durga Toshniwal: If you look at, precision, it is 0.90 for 16, 0.88 for 32, and 0.8, again, 0.90 for, 64.
[01:48:44] Durga Toshniwal: Then, if you look at, recall, it is 0.92 for, size of… batch size of 16. For 32, it is 0.92.
[01:48:55] Durga Toshniwal: And for 94 also 0.92, so all 3 are equal.
[01:48:59] Durga Toshniwal: Then, F1 score for size 16 is 0.91.
[01:49:05] Durga Toshniwal: For 32.90, and for 64, it is 0.91, so it is equal in 16 and 64.
[01:49:14] Durga Toshniwal: And then AUC is 0.97 for 16, 0.97 for 32, 0.98 if you round up, for 64.
[01:49:25] Durga Toshniwal: So, actually, more or less, the results for 16 and 64, although all three are almost equal. However, for 16 and 64, they are comparable, so we could go with 16 or 64.
[01:49:39] Durga Toshniwal: In this piece of code, we have selected 16.
[01:49:42] Durga Toshniwal: However, if the data size is very large, in our case, the data size is around, like, 14,000-something records. However, for larger data sets, larger batch sizes will be preferred, because batch size is the amount of data
[01:49:58] Durga Toshniwal: We, process at one go.
[01:50:01] Durga Toshniwal: So, if the batch size is very slow, the compute time will be very high.
[01:50:06] Durga Toshniwal: However, here, 16 has been chosen, and we'll proceed with those, with that number for our model.
[01:50:15] Durga Toshniwal: So I think you can take the code now, Asmita.
[01:50:19] Asmita Mahajan: Yes, ma'am.
[01:50:20] Asmita Mahajan: So, to plot these bar plots, we'll be using the data frame, results.
[01:50:27] Asmita Mahajan: And,
[01:50:28] Asmita Mahajan: From that, we are extracting the columns, which contains the metric names, and also we are extracting the index, which will contain the value for the batch size, which is 16, 32, and 64, and converting it to string before storing it in the list label.
[01:50:48] Asmita Mahajan: Calling the subplots on the number of rows and columns, which would be 1 and 3 in our case… sorry, which will be, 3 and 2, because, we have 5 matrices to be plotted.
[01:51:02] Asmita Mahajan: And then flattening the location of the axis from 2D to 1D array, enumerating over the name of the matrices, which is accuracy, precision, recall, F1 score, and AUC score, and extracting the location of the first subplot, which is… we are now plotting accuracy at the first location.
[01:51:21] Asmita Mahajan: So, from that, we are,
[01:51:23] Asmita Mahajan: Extracting the values of the accuracy metric and storing it in val' list, calling the bar function on the index labels, which is 16, 32, and 64,
[01:51:35] Asmita Mahajan: And, on the y-axis, we want the values for the accuracy scores, and each color is black. Setting the title of the subplot to metric, setting the X label to batch size and Y label to metric, and setting the limit of the Y axis to… from 0 to 1.
[01:51:54] Asmita Mahajan: Also, for display purposes, we are also rotating the X stick labels at 45 degree, and also we want the values associated with each bar. For that, we are using text, and giving the location for that value, how much decimal places we want, like how many decimal places we want for that value, horizontal alignment, vertical alignment, and the form size.
[01:52:19] Asmita Mahajan: And we don't want any extra empty plots, so we are offing the access for that. We don't want the access, so we are saying that access should be off.
[01:52:30] Asmita Mahajan: Displaying the overall, subtitle, and plotting the subplots, and also displaying the
[01:52:38] Asmita Mahajan: data frame. So we… this… we have already discussed, we are choosing 16 as the batch size. Now, we have trained, the overall. We have already tuned all the parameters of the model.
[01:52:51] Durga Toshniwal: So, Atom Optimizer will be used with learning rate of 0.05. Just, I… sorry to interrupt, I just want to add here that after, tuning all the hyperparameters, we…
[01:53:03] Durga Toshniwal: we must use those that we have finalized. However, you will notice here that Adam Optimizer has been used, learning rate is what we decided, 0.01, bad size 16.
[01:53:14] Durga Toshniwal: And the hidden layer sizes are 256 and 128, respectively. However, the activation functions are not relevant TANH, and as already pointed out, 10H is more compute-intensive, because of various reasons. So, just for the sake of,
[01:53:32] Durga Toshniwal: Example here, we have chosen both as ReLU.
[01:53:36] Durga Toshniwal: And then the epochs have also been chosen as 50 only, because the batch size is 16… the chosen batch size is very small, 16. So if we have epochs equal to 15, 500, with a batch size of 16, it will take a huge amount of time to run.
[01:53:52] Durga Toshniwal: Because batch size is very small, so you can just imagine that 14,000 records will be taken 16-16 at a time.
[01:54:00] Durga Toshniwal: And this will be repeated 500 times, so that is going to be taking a huge amount of time. So that's the reason, because the batch size has come out to be quite small, so the number of epochs chosen here for just the purpose of showing or illustrating is 50. However.
[01:54:18] Durga Toshniwal: Actually, if you really build a model in… the correct model with all the chosen hyperparameters, the activation function should be relevant tanage, and the epoch should be 500. So I'm telling you this in advance for those who would be wondering why this is coming.
[01:54:34] Durga Toshniwal: Okay, now you may proceed, Rasmita.
[01:54:37] Asmita Mahajan: Yeah, so,
[01:54:38] Asmita Mahajan: with these hyperparameters that we have fixed or tuned, we'll be training the overall MLP model. So, what we are doing is, we are,
[01:54:50] Asmita Mahajan: like, fixing the things that we have tuned. So, batch size is 16.
[01:54:56] Asmita Mahajan: So, here also, like we did, when we were tuning pad size, we will,
[01:55:04] Asmita Mahajan: combine the dataset, the X train and Y train dataset, and then pass it to the data loader to separate it into different batch sizes, different batches of size 16.
[01:55:17] Asmita Mahajan: So, trainloader will now be an iteratable object, which will, contain the, X,
[01:55:26] Asmita Mahajan: The features for that batch, and the labels for that… for a particular batch.
[01:55:32] Asmita Mahajan: we are now, defining our MLP class in a very similar manner, so what we are doing is, we are just building the network in a sequential manner with a linear unit, which will connect the input layer to the
[01:55:48] Asmita Mahajan: First hidden layer, which has the 256 as the size of the first hidden layer. Railu activation function is used. Now, the second linear unit will connect the first hidden layer to the second hidden layer.
[01:56:00] Asmita Mahajan: with ReLU activation function, and the last linear unit, connects the second hidden layer to the output layer.
[01:56:08] Asmita Mahajan: And this forward function will just pass or propagate the inputs through the network use, for the forward pass, which is defined here. So this completes the definition of the class MLP.
[01:56:21] Asmita Mahajan: Now, we are,
[01:56:23] Asmita Mahajan: creating the object of the MLP class in model, and passing it the input dimension, or the input layer size, which is the number of features from the X-Train dataset.
[01:56:39] Asmita Mahajan: Using criteria as the binary cross entropy loss, we are storing that in criterion. Atom Optimizer is used to update the weights with learning rate of 0.01.
[01:56:51] Asmita Mahajan: Now, epochs are defined here. We have initialized epochs to 50, because we are training our model 50 times. So, for that number of epochs, we are iterating over the train.
[01:57:05] Asmita Mahajan: We are training the model, yeah, and opening the model in training mode, train mode. So, for the batch size, for each batch, we are now, we will be now, training the model. So…
[01:57:18] Asmita Mahajan: We are initializing the gradients to zero, so that the gradients from the previous iteration should not be, used or leaked into the next iteration.
[01:57:29] Asmita Mahajan: Passing the batch into the model, or the network, using the forward pass here, and, capturing the outputs in lodges.
[01:57:41] Asmita Mahajan: We are now, evaluating the loss using the binary cross entropy loss function, and calculating the loss here. So this loss is now back-propagated through the network using the backward function, and the,
[01:57:59] Asmita Mahajan: The weights are now updated using the optimizer function.
[01:58:05] Asmita Mahajan: So, we have, with this…
[01:58:08] Asmita Mahajan: for loop, we have trained our model, for the number of, for the particular batches also. So we have trained the full model.
[01:58:16] Asmita Mahajan: Now, we will be evaluating the trained model on the test set. So what we are doing is we are opening our model in evaluation mode with no gradients, because we want to use the learned weights to test the model. We are using the whole batch, the whole test batch, to propagate through the network.
[01:58:33] Asmita Mahajan: And then, capturing the outputs in the lodges. Using the sigmoid activation function, we are converting the raw outputs into probabilities, and making the hard predictions using 0.5 as the threshold.
[01:58:48] Asmita Mahajan: Now, simply, we are now converting the tensors into NumPy array using NumPy function to evaluate… to further evaluate the accuracy, precision, recall F1 score, and AOC score.
[01:59:01] Asmita Mahajan: So, the matrix will store the performance matrices, and then we are just converting the matrix into a data frame with index as MLP.
[01:59:12] Asmita Mahajan: So, this is the… Oh, result, these are the results.
[01:59:19] Asmita Mahajan: So now, as we were discussing…
[01:59:22] Durga Toshniwal: So now, just, I wanted to add here, so the final accuracy, precision recall, F1 score, and AUC,
[01:59:29] Durga Toshniwal: That we are getting are quite good, as you can see here, 0.93, 0.94, 0.91, 0.9, F1 is 0.93, and AUC is 0.98. Pretty good.
[01:59:43] Durga Toshniwal: Yeah.
[01:59:44] Durga Toshniwal: And the addition that you won't be having in your code is having the model with just one hidden layer, because there was a question yesterday.
[01:59:53] Durga Toshniwal: that what if we try the model with one hidden layer, and not with two? So we just added it just for illustration purpose, so please note that you won't be having it.
[02:00:05] Durga Toshniwal: So you can just, I think, run it and show the results, because the code is just the same.
[02:00:09] Asmita Mahajan: Similar.
[02:00:10] Asmita Mahajan: Yes, ma'am. Just we don't have that linear unit here, otherwise the code is similar, and with that, we are getting slightly,
[02:00:20] Asmita Mahajan: Reduced results.
[02:00:22] Durga Toshniwal: Yeah, so… so you can see here that when we are having one layer, the accuracy is 0.90, 0.91, precision is 0.
[02:00:34] Durga Toshniwal: 8, 9, recall is point…
[02:00:37] Durga Toshniwal: 90, F1 score is 0.90, AUC is 0.
[02:00:42] Durga Toshniwal: 9, 8, or 99, something like that. So, though the results are not very different, however, definitely by increasing the layers to 2, hidden layers to 2, the results are better. You can just go over to those results once again.
[02:00:58] Durga Toshniwal: Just… yeah. So you can see there is a difference.
[02:01:02] Durga Toshniwal: In all the, these scores of accuracy, precision, recall, and front score, AUC,
[02:01:10] Durga Toshniwal: So, definitely, we will choose two hidden layers.
[02:01:14] Durga Toshniwal: For our model, and not one. So, this was just us yesterday, so we just included. So, I think with this, we have come to the end of the present code.
[02:01:23] Durga Toshniwal: We have discussed at length,
[02:01:26] Durga Toshniwal: The different models that we wanted to do.
[02:01:30] Durga Toshniwal: One final thing that we included here is the comparison with other classifiers, which are nonlinear models. So I think you can just show the results. The models have been discussed
[02:01:43] Durga Toshniwal: Earlier at length.
[02:01:44] Durga Toshniwal: So if you recollect, we already discussed, random forest, decision tree and k nearest neighbor. All of them are nonlinear classifiers, so what we have done is that we have compared our results with the other, nonlinear classifiers.
[02:02:01] Durga Toshniwal: And, just noted which ones are the best. Out of all these that you are seeing here, definitely for this dataset.
[02:02:09] Durga Toshniwal: the MLP model that we have built with two hidden layers, 256, 128 neurons per layer, relu, ReLU,
[02:02:17] Durga Toshniwal: and batch size of 16, and epochs equal to 50. That was our final model. With that, we are getting better results for accuracy, which is 0.93, precision 0.94, recall 0.91, F1 score 0.93,
[02:02:35] Durga Toshniwal: and AUC.98. So these are better than all other nonlinear classifiers, like
[02:02:41] Durga Toshniwal: Random forest, decision tree, k nearest neighbor, and all.
[02:02:45] Durga Toshniwal: However, these results cannot be generalized over any data. So, for any given data, we'll actually have to find out the results and see, how they benchmark with the different types of nonlinear classifiers. MLP is one of them.
[02:03:01] Durga Toshniwal: Okay, so here MLP is outperforming the, other important
[02:03:06] Durga Toshniwal: nonlinear classifiers. So I think with this, we come to an end for this part, where we illustrated how to build, how to do hyperparameter tuning, then having done the hyperparameter tuning, how to build the final model, and then to compare it with some other nonlinear models.
[02:03:25] Durga Toshniwal: Any questions? Anything? Any comments on the code? Any questions on the code?
[02:03:34] Durga Toshniwal: I think we took it at a lot of details. However, as we will now proceed, it may not be possible to go to all such details, and I suggest that you please increase your proficiency in Python.
[02:03:48] Durga Toshniwal: and PyTorch, so that there are no issues in the further sets of codes.
[02:03:56] Durga Toshniwal: Any questions, anyone?
[02:04:02] Durga Toshniwal: Okay, if there are no further questions, then we can break for today. I'll request Asmita to… she can feel free to leave the meeting. Thank you, Asmita, for taking us through the code.
[02:04:14] Durga Toshniwal: And thank you all. We can now break for today. We'll meet next week. Okay, thank you. Bye-bye.
[02:04:21] Nirav Mehta: Thanks.