# 04 2026-02-01 Hands On Deep Learning

course: Module 3 — Deep Learning & NLP
module: Module-3-Deep-Learning-NLP
date: 2026-02-01
type: transcript
video_url: https://personal-learn.armco.dev/files/_Recordings/Module-3-Deep-Learning-NLP/04_2026-02-01_Hands_On_Deep_Learning.mp4

---
[00:19:49] Durga Toshniwal: A very good morning to all of you.
[00:19:52] Durga Toshniwal: And, welcome to today's session.
[00:19:55] Durga Toshniwal: So, I'm facing some technical issues, so I'm not turning on my video. Today, we'll be actually, taking the hands-on for deep learning.
[00:20:05] Durga Toshniwal: In which we'll be using a multi-layer perceptron, or MLP,
[00:20:09] Durga Toshniwal: So, the whole idea will be to do hyperparameter tuning. We discussed at length the different hyperparameters. So, we are going to do hyperparameter tuning.
[00:20:20] Durga Toshniwal: On our data, and then use it for purpose of classification, and then we'll compare the result of the classification with other classifiers. We'll choose the best set of hyperparameters to build a final model.
[00:20:33] Durga Toshniwal: And for taking us through the hands-on, we have a senior research scholar from my group, Ms. Asmita, here. Ms. Asmita, can I request you to please enable your video?
[00:20:47] Asmita Mahajan: Yes, ma'am. Good morning, ma'am.
[00:20:50] Durga Toshniwal: Good morning, good morning, Asmita. Welcome to the session. Thank you for joining.
[00:20:55] Durga Toshniwal: So, I request you to please share your screen.
[00:20:59] Durga Toshniwal: And,
[00:21:00] Durga Toshniwal: take over from here. I request all of the learners to please, load the files on the LMS. Are you all ready, or shall we allow 2 minutes time to load the files?
[00:21:15] Sonam Manwal: Please give us 2 minutes, ma'am.
[00:21:18] Durga Toshniwal: Okay, okay, no issues.
[00:21:22] Durga Toshniwal: Yeah, that's funny.
[00:21:51] Durga Toshniwal: So, the link for the files was shared yesterday. Also, on the chat, I request the host to please share the link in case some people may not have seen,
[00:22:02] Durga Toshniwal: the LMS, for the files.
[00:22:08] Durga Toshniwal: So, in the meantime, please do let us know. Smita is going to take over. In case you find, you, find the pace, slow or fast, you can please let us know. And, maybe every… in between, say, 15-20 minutes, we can take questions.
[00:22:28] Durga Toshniwal: So, Asmita, I'll request you to please start with the explanation of the code.
[00:22:34] Durga Toshniwal: And, I'll be there anyway. Please take over from here. Thank you.
[00:22:39] Asmita Mahajan: Okay, ma'am, thank you.
[00:22:40] Asmita Mahajan: So, today we'll be learning how we'll tune the hyperparameters of a simple MLP model, multi-layer perceptron.
[00:22:50] Asmita Mahajan: So for that, we'll be using the dataset of EEG iState dataset, which has 14 feature attributes and 1 target variable, which is,
[00:23:02] Asmita Mahajan: a binary variable which has a 0 and 1 value. 0 means the eyes are open, and 1 means the eyes are closed.
[00:23:11] Asmita Mahajan: So, we have these 14 features, and using these 14 features, we'll be predicting.
[00:23:18] Asmita Mahajan: the target variable, and then we'll be checking, like, which model performed the best. And to tune the MLP, we will have many hyperparameters, for example, the activation function between the layers, the hidden layers, how many hidden layers should be there, or what would be the size of the hidden layer.
[00:23:37] Asmita Mahajan: what would be the optimizer that we'll be using? So, batch size, how many iterations we have to run, to train our model. So, there are different hyperparameters which will be eventually,
[00:23:54] Asmita Mahajan: learning.
[00:23:56] Asmita Mahajan: So, this data that we have extracted from the UCI machine learning, repository, it's a ARFF file. So last time, when we did k-means clustering, we used the,
[00:24:11] Asmita Mahajan: CSV file, or the Excel file. So, we extracted the Excel file into a data frame, and then we used that data frame for further tasks. But today, we have the ARFF file, which is the attribute relation file format, which is a text-based dataset format. It contains data in the form of text.
[00:24:34] Asmita Mahajan: And also, it has some information about the data, the relation of different… of the data, or what is the data about. So, it has that also. So, let me first… so for that, you have to firstly, install the library using PIP install.
[00:24:53] Asmita Mahajan: And you, using this command, and the library will be installed, because this is a different file format, so for that, we need a different library to be imported into our system, so that we can extract it,
[00:25:07] Asmita Mahajan: Correctly. So, I hope everybody has installed this library.
[00:25:18] Asmita Mahajan: Okay, I'll move ahead now.
[00:25:20] Asmita Mahajan: So, I'll be moving the ARFF, importing ARFF, file format and Pandas as PD, because, with this, library, we'll be, performing tasks for our dataset.
[00:25:35] Asmita Mahajan: So, this line with open, EEGISTATE.ARF, this is our file name.
[00:25:43] Asmita Mahajan: So, what we are doing is we are opening this file, on read, mo… in read mode, so as F. So, F is the variable where we will be storing the instance of the file, okay?
[00:25:56] Asmita Mahajan: So, this command, open EEG.ARF, ARFF, it opens our file in read mode, and then we will be, using arFF.load
[00:26:09] Asmita Mahajan: Function to load our dataset… to load our data into a dataset, variable.
[00:26:17] Asmita Mahajan: And F is where we have stored the instance of the file, okay?
[00:26:24] Asmita Mahajan: And then we'll… and then, like last time we did it, to convert anything into a data frame, we use data frame function. So, from the pandas library, we… we are extracting the data frame, function, and then we are passing… okay. So, let me first run this.
[00:26:44] Asmita Mahajan: Okay, so I have to install this.
[00:27:10] Asmita Mahajan: So, if I'll try to… See what dataset…
[00:27:16] Asmita Mahajan: contains. So, dataset is the raw, file. Dataset is the variable where the raw file is stored. So, if I'll print it.
[00:27:28] Asmita Mahajan: If you try to print this dataset, You will see… How it is stored.
[00:27:38] Asmita Mahajan: Okay, so here, as you can see, we have… it is a dictionary, so it is extracted as a dictionary, where a description is a key, relation is a key, then EEG data is the value, attributes is a key.
[00:27:53] Asmita Mahajan: And inside attributes, we have all the names of the attributes that we have in the dataset, and also what is their data type. So, AF3 is our first column.
[00:28:05] Asmita Mahajan: And it is of numeric data type. So, in this manner, we have all the metadata, all the information about the data.
[00:28:12] Asmita Mahajan: And after attributes, we have the key data, where the actual data is stored, in a row-wise manner. So, it is a 2D array where the first row contains the record of the first
[00:28:28] Asmita Mahajan: a row in the dataset. So, here you can see. So, in a row-wise manner, the data has been stored. But it is a text-based data, so to extract the text-based data, we need the ARFF file, okay?
[00:28:39] Asmita Mahajan: So, in this manner, the dataset has been… the dataset is actually present in this format, but then we converted it into a data frame using dataset
[00:28:51] Asmita Mahajan: and key data. So, this, line of code, or this, portion of the code, extract this.
[00:29:00] Asmita Mahajan: so, it will match data with the key, and then it will take all the values.
[00:29:07] Asmita Mahajan: the data, actual data, of… which is stored in the row-wise form, it will extract this, and it will give it to the data frame, function. Then, for columns, what we are doing is.
[00:29:20] Asmita Mahajan: We are calling the, we are calling it on dataset Attributes.
[00:29:28] Asmita Mahajan: So here you… here you will see the key attribute, so it will match the key, and it will then extract this list of tuples. So, we have a… so the first element gives us the name of the…
[00:29:42] Asmita Mahajan: attribute, and then tells us that the data type is numeric. But we only need the name of the attribute, because we are passing it to the columns, variable, parameter here. So, we only need the, first,
[00:29:57] Asmita Mahajan: element of the tuple. So that's why we… what we are doing is we are iterating over the dataset attributes, we are iterating over this list, and from this list.
[00:30:08] Asmita Mahajan: the first element is this full tuple, and from that full tuple, we only want the first element of the tuple. That's why we have used attribute 0.
[00:30:19] Asmita Mahajan: Okay?
[00:30:22] Asmita Mahajan: Because, here, we will, in the first iteration, we will get this full tuple, but we only want the name of the attribute, so that's why we are using 0 here.
[00:30:34] Asmita Mahajan: Okay?
[00:30:35] Asmita Mahajan: I think… I hope this is understood.
[00:30:39] Asmita Mahajan: So, when we run this, we'll get our data set in the form of a data frame. So, we have 14,980 rows and 15 columns in total, where one is the target variable, and the 14 others are the features of the… which we'll use to predict the target variable.
[00:30:58] Asmita Mahajan: Now we'll run the info function, and we will see that we don't have any null values in the dataset, and what are the data types of the attributes.
[00:31:09] Asmita Mahajan: So, my target variable is of object data type, and all the other variables are numeric.
[00:31:16] Asmita Mahajan: And they have float64 data type.
[00:31:22] Asmita Mahajan: So, we'll explicitly check if there is any null value in the dataset, and as you can see, there is zero null value in the dataset. So, for all the attributes, we don't have any null value.
[00:31:36] Asmita Mahajan: Describe, will tell us,
[00:31:40] Asmita Mahajan: the count, mean, standard deviation, and minimum, and the interquartile range of the dataset. So, using describe, you can check the descriptive statistics of the numerical variable.
[00:31:55] Asmita Mahajan: So, how the dataset is distributed, what is the variance of the data, Each value in the attribute.
[00:32:05] Durga Toshniwal: So, Neerat, I see you have a question. What question do you have?
[00:32:09] Neeraj Kumar: Yeah, hi, sorry to disturb you, ma'am. Actually, can you…
[00:32:13] Neeraj Kumar: Describe a bit about the data. What does it mean that F7, FC5, what does it mean? Actually, I'm not able to understand this data.
[00:32:23] Asmita Mahajan: Okay, so at the very start, we have, the,
[00:32:28] Asmita Mahajan: this feature, like, this description of the dataset, where AF3 is the EEG signal from the frontal electrode.
[00:32:36] Asmita Mahajan: And it, it has given the name AF3. So, the EEG dataset contains the, signals from the brain. So, it has been collected for, one patient, and,
[00:32:50] Asmita Mahajan: like,
[00:32:51] Asmita Mahajan: the description is given here. It is a reading from multiple scalp electrodes recorded over time from a single patient. So,
[00:33:01] Asmita Mahajan: F7 is the signal from the frontal electrode, the electrode which was connected at the frontal lobe, and left side. F3 is the,
[00:33:12] Asmita Mahajan: Again, from the frontal electrode, but it is from the left hemisphere of the brain.
[00:33:18] Durga Toshniwal: Just to come in, sorry to interrupt, Smita. So, these are the signals that have been taken from the scalp, or the human brain, using different probes.
[00:33:29] Durga Toshniwal: So, the names of the probes are, like, F1, F4, and all that. And of course, they are covering the different regions of the brain to get a full picture of the EEG signals from the brain. Yeah, please carry on.
[00:33:45] Asmita Mahajan: Yes. So,
[00:33:48] Asmita Mahajan: So, F3 is the electrode which is connected to the left hemisphere. Then again, F… FC5 is the central electrode, which is connected to the frontal central electrode. So, these are all the, electrode names. So, the description is given here for all the electrodes, where the electrode is connected. And using this, like, whenever we
[00:34:11] Asmita Mahajan: open our eyes, or we close our eyes, what is actually going on in the brain? So, these electrodes will capture that. These electrodes will capture the data, and then it is converted into the numeric data. So.
[00:34:24] Asmita Mahajan: we are getting, like, so this data contains the values, like, what is actually happening inside the brain when we open the eyes or when we close the eyes. And,
[00:34:37] Asmita Mahajan: The eye detection target variable captures that, or tells us that.
[00:34:45] Asmita Mahajan: Is it clear?
[00:34:50] Neeraj Kumar: Yes, ma'am. Bit.
[00:34:52] Neeraj Kumar: Thank you, man.
[00:35:00] Asmita Mahajan: Okay. So, we have,
[00:35:03] Asmita Mahajan: displayed, or we have seen, the descriptive statistics of the numerical attributes in the data which we had. So, as we… our target variable was object, so target variable will not be here, but other than that, all the numeric attributes are displayed here.
[00:35:22] Asmita Mahajan: So now, let us see what is the, like, distribution of the target variable inside our data. So for that, we'll use the valueCount function, which will count the number of zeros and ones in the target variable.
[00:35:38] Asmita Mahajan: So… and now, this time we are normalizing it, because we want to see the proportion of the data, not the actual values. We don't want to see the actual values, but we want to see the proportion. Like, how much proportion, is, in…
[00:35:52] Asmita Mahajan: in the data is 0, and how much proportion, is 1. So, that's why we are using normalize true, because what it will do is, it will, just normalize our data according to the total number of, you know.
[00:36:08] Asmita Mahajan: records in the data. So,
[00:36:12] Asmita Mahajan: And after that, what we are doing is, we are also plotting it, using a bar plot. So, here you can see, we extracted the counts, we normalized it, then we multiplied it by 100. Why? Because, normalized true will give us the proportion. But to get the percentage, we multiplied it by 100, and then we plotted the data using the kind bar.
[00:36:35] Asmita Mahajan: So, bar will tell us that we want bars for the data.
[00:36:40] Asmita Mahajan: Here, you can also write, line, so for, so it will plot the line plot of the data.
[00:36:46] Asmita Mahajan: So, if we'll run this…
[00:36:50] Asmita Mahajan: you see 55… 0.55, proportion of the data, contains 0, as in the target variable, and 0.44 proportion of the data contains 1 in the target variable. And, again, it is plotted using the bar plot.
[00:37:18] Aditya Banda: Just one question. So, are we doing this to make sure the class distribution is not uneven? Is that the purpose why we are checking this?
[00:37:26] Asmita Mahajan: Yes.
[00:37:27] Aditya Banda: Okay.
[00:37:28] Asmita Mahajan: Yes, so, for checking the class imbalance, if the classes are imbalanced in the dataset, that's why we are doing this, to see, how much proportion is, 0 and how much proportion is 1.
[00:37:41] Aditya Banda: Okay, here it looks like it's not imbalanced, but what if it's imbalanced? What do we do then?
[00:37:46] Asmita Mahajan: Then we have, further methods to, you know,
[00:37:52] Asmita Mahajan: correct it. Like, if, I have…
[00:37:55] Asmita Mahajan: For example, if, here, 1 would be, like, if it comes out to be only 30%, then we have to increase the instances of 1, or we have to augment the data in such a manner so that my classes are nearly balanced, because otherwise, the influence of 0 would be more, in our… in training our model.
[00:38:17] Asmita Mahajan: So, for that, we have to make the glasses balance.
[00:38:21] Asmita Mahajan: So, for that, there are various methods which you will learn in upcoming classes.
[00:38:27] Aditya Banda: Okay.
[00:38:34] Asmita Mahajan: Okay, so, after this, we'll be, doing our preprocessing steps and train and split,
[00:38:45] Asmita Mahajan: we will be splitting our data set into train set and test set. So, to do that, so why we are doing this, we are, like, we'll be training our model using the train set, and then we'll be testing the test instances, or the test set, which is unseen to the model, so that we can, have a…
[00:39:04] Asmita Mahajan: like,
[00:39:07] Asmita Mahajan: we can tell that, our model is not overfitted on the training set, because it will give the full model to the data. The model will just, it will not learn from the data, it will not learn the underlying, you know, behavior of the data, but instead, it will just mimic the data. So, to,
[00:39:27] Asmita Mahajan: to,
[00:39:29] Asmita Mahajan: you know, correct this, mistake that the model should not, mimic the data. What we'll be doing is we'll be, splitting our data set into train, set and the test set, so using only…
[00:39:41] Asmita Mahajan: Only the train set will be training the model, and using the test set, we'll be seeing if the model has learned anything or not.
[00:39:51] Asmita Mahajan: So, firstly, we'll be, dividing the dataset into feature and target variables. So, X will, contain the feature variables, where we are dropping, the eye detection column from the dataset, and X is 1, because, we are telling,
[00:40:08] Asmita Mahajan: the drop function that, this, eye detection is a column.
[00:40:13] Asmita Mahajan: variable.
[00:40:15] Asmita Mahajan: So, by default, axis is 0 here in this method, so that's why we have to explicitly mention that axis is 1. So, 1 tells us that the eye detection is a column.
[00:40:25] Asmita Mahajan: And then we are, passing the values, to the X variable, and in the Y variable, we are just, passing this iDetection column, which is the target variable directly, and we are converting into the type int, because,
[00:40:43] Asmita Mahajan: in the… using the info command, we saw that, ID reduction is… target variable is of type object, so that's why we are converting it to int, and then, we are passing the values from that column to the by variable, which is our target variable.
[00:41:00] Asmita Mahajan: This print command is, showing the shape of the dataset. What is the shape, of X?
[00:41:08] Asmita Mahajan: variable, and what is the shape of the Y variable. So X will have values like the total number of rows.
[00:41:15] Asmita Mahajan: comma, the columns. So, there were 14 columns. And Y will have only one column, and all the rows. So, using shape, we can see if both the variables are compatible, or both the variables have equal shape.
[00:41:35] Asmita Mahajan: So, using train-test-split method, we'll be splitting our dataset. So, what we'll… we are going to do, we are passing the
[00:41:44] Asmita Mahajan: features, features variable. We are passing the target variable.
[00:41:50] Asmita Mahajan: then test sites will tell us that how much data we have to keep for the test set. So here we are splitting the dataset into 80-20 rule, like, 80% of the data we'll be using for training, and 20% of the data we'll be using for
[00:42:05] Asmita Mahajan: testing the model. So that's why we have given it the value as 0.2, which will tell the train-test split to,
[00:42:14] Asmita Mahajan: Keep 20% of the data for the test set, and the remaining data set should be given to the training set.
[00:42:21] Asmita Mahajan: Random state is just, too, you know, for reproducibility. So, if we'll use 42, every time in the random state, we'll get the same split.
[00:42:34] Asmita Mahajan: So…
[00:42:36] Asmita Mahajan: So that we get, we always get the same records in the training set, and same records in the test set, whenever we are,
[00:42:47] Asmita Mahajan: You know, running the… This, line of code.
[00:42:52] Asmita Mahajan: Stratify by means… so, either it can be true or false. So, stratification we do is, like, we… we are telling that you… you have… you don't have to take the records sequentially.
[00:43:04] Asmita Mahajan: Or in sequence, in the records, in which the records are present in the dataset, in the actual dataset. But, otherwise, what you are going to do is.
[00:43:13] Asmita Mahajan: Either it will be true, so it tells that you can shuffle the data and then take it, so in classification, it is a good practice if you stratify the data and then split it. Otherwise, it could be false, so for…
[00:43:28] Asmita Mahajan: Time series forecasting, we usually keep it false, because we need the data in sequence in time series forecasting.
[00:43:34] Asmita Mahajan: And Y tells us that I want my data to be strat… like, shuffled according to the Y variable, according to the target variable, so that my classes,
[00:43:45] Asmita Mahajan: So… so there's a representation of my classes in the…
[00:43:49] Asmita Mahajan: Test set, and in the training set doesn't change.
[00:43:53] Asmita Mahajan: So if, as we have seen here, we have 55% of the data as 0, and 44% as 1, so in the test set and training set, the proportion of the classes of these 0 and 1 shouldn't change, so that the represent… the test set is a pure representation of the training set.
[00:44:14] Asmita Mahajan: That's why we are using, stratify as Y here.
[00:44:21] Asmita Mahajan: So, after this split, after we have splitted, like, separated our dataset into train and test set, now we'll be, scaling it, or we'll be normalizing our dataset, because, as we can see.
[00:44:35] Asmita Mahajan: Here, the standard deviation for all the val- of all the variables or the attribute is very different. So, we'll be scaling the dataset, or we'll be normalizing the data, so that one variable doesn't influence the training of the model.
[00:44:54] Asmita Mahajan: So for that, we'll be using a standard scaler.
[00:44:57] Asmita Mahajan: library from the SK… module from the sklearn preprocessing, library. So, we have imported it here. And also, train, test, split, we have imported from the model selection library.
[00:45:14] Asmita Mahajan: So, what we are doing is, we are, creating an instance of standard scalar, in the variable scalar, and using that scalar… So, standard scalar is the z-score transformation of the data. So, we are transforming our data into,
[00:45:30] Asmita Mahajan: Such like… so it has, the…
[00:45:34] Asmita Mahajan: The formula for this is, we have X, where X is the data point, from the data, minus the mean of the data, divided by the standard deviation of the data.
[00:45:46] Asmita Mahajan: So, using this transformation, we are transforming our dataset into… or we are normalizing our data set. So, we have standard scalar, we also have min-max scaler, where the minimum and maximum values are used to scale the data. So, there, the formula would be X minus the minimum value divided by maximum minus minimum value. But here, we are using the z-score transformation, so that the
[00:46:09] Asmita Mahajan: Whole data has a zero mean and one standard deviation.
[00:46:19] Asmita Mahajan: Okay, so we have created the scalar instance, from the standard scalar module. Then, using that scalar instance, what we are doing is, we are fitting and transforming the, XStrain dataset.
[00:46:33] Asmita Mahajan: Which we, got here from the train-test split.
[00:46:40] Asmita Mahajan: So, we are passing Xtrain to the fitTransform function, and then, we are storing it in XtrainScaled.
[00:46:48] Asmita Mahajan: And, as you can see here, we have just transformed the X test. So, why we have done that? So, to prevent data leakage.
[00:47:02] Asmita Mahajan: Because…
[00:47:03] Asmita Mahajan: We only want our test set… we want our test set to be unseen by any model, here.
[00:47:10] Asmita Mahajan: So, to prevent that, to prevent data leakage, so that my test set is not seen earlier, so what I am doing is I'm not fitting it. So, whatever the model has been fitted here.
[00:47:25] Asmita Mahajan: for the X train, I'm using that model, and I'm using that model to transform my X test.
[00:47:32] Asmita Mahajan: And to scale. So the mean and standard deviation, which we have evaluated or calculated here for the X-trained dataset, the same mean and standard deviation I'm using to transform my X test dataset, so that
[00:47:45] Asmita Mahajan: I can prevent data leakage problem. So, data leakage is where, you know, your…
[00:47:52] Asmita Mahajan: You are showing the model, the…
[00:47:55] Asmita Mahajan: test instances also. So, to prevent that, what we are doing is, we are not fitting it here. As you can see, we have fitTransform function, on the training set, but only… we have only transform function on the test set. So, whatever my mean and standard deviation, I have calculated,
[00:48:11] Asmita Mahajan: In this step, I'm using the same mean and standard deviation to transform my X test.
[00:48:19] Asmita Mahajan: So, this will, transform or normalize my, X test data, and it will be stored in test scaled.
[00:48:26] Asmita Mahajan: And then I'm just printing the shape of both the variables that I've just created.
[00:48:34] Asmita Mahajan: So when you run this, You will see that, before, standardizing.
[00:48:41] Asmita Mahajan: Sorry, before splitting the data, I had 14,980 rows and 14 columns in my X feature variable.
[00:48:51] Asmita Mahajan: And only 14,980 in the target variable. But after the split, you can see that my test set contains approximately 3,000 records, and the remaining records are there in the train set.
[00:49:12] Asmita Mahajan: Any question till now?
[00:49:19] Asmita Mahajan: Okay, we'll move forward.
[00:49:21] Deepan Kanagaraj: So, with respect to scalar transformation, right, when we will go with min-max type of scaling and the other way, is there any constraint or conditions for that?
[00:49:34] Asmita Mahajan: There are no constraints for that. It is just a best practice that you use a standard scalar as your transformation function. So you can use either, either min-max or standard scalar, so there are no constraints.
[00:49:52] Asmita Mahajan: Okay.
[00:49:53] Asmita Mahajan: For using that.
[00:49:57] Ankit Sood: Ma'am, can you please explain the data leakage part again?
[00:50:01] Asmita Mahajan: Okay, so…
[00:50:02] Asmita Mahajan: why I have splitted my data into train and test set is I don't want my model to learn from the full data, because, as I said.
[00:50:11] Asmita Mahajan: I don't want my model to mimic the behavior of the data. I want my model to learn from the data. So, I am doing… I am separating my data set into two parts. One is used for training, and one is,
[00:50:24] Asmita Mahajan: explicitly used for only testing. So, this test data
[00:50:27] Asmita Mahajan: it is unseen to the model. Whenever I'm training my model, I'm only passing X train, the training set, to the model. I'm not passing testing set. So, for that, we have separated our data set into two sets, X train and xtest. So, data leakage is where your model
[00:50:48] Asmita Mahajan: Where, you, you performs,
[00:50:52] Asmita Mahajan: Practices where the model has seen the test set earlier.
[00:50:56] Asmita Mahajan: or the dataset, or the test, instances earlier. So… To prevent that.
[00:51:06] Asmita Mahajan: Here, we have fitted our standard scaler on the training set only, and we have used that model, or that fit, to transform our text, X test data.
[00:51:21] Asmita Mahajan: So, to prevent data leakage, so that my model doesn't see the test data earlier or previous, previous to the training. So, for that, I have just used the fitted model on the X train.
[00:51:36] Asmita Mahajan: to transform the X test data.
[00:51:40] Ankit Sood: Hmm, okay.
[00:51:41] Asmita Mahajan: Is it clear?
[00:51:43] Ankit Sood: Ew.
[00:51:51] Asmita Mahajan: Okay, so, till… this part, we have pre-processed the data, so now we are ready
[00:51:58] Asmita Mahajan: to, you know, pass our data, this preprocessed data, normalized data, to train our model. But before that, as we are using, deep learning, and for that, we are… we will be using PyTorch, library, PyTorch module. So, PyTorch,
[00:52:15] Asmita Mahajan: So the neural network, methods in the PyTorch module, con- takes the data in the form of tensors. So, tensors are nothing, but it is just a form that PyTorch uses. So, as, here, if I'll display white train.
[00:52:32] Asmita Mahajan: You can see that, Y train is an array, or if I'll…
[00:52:42] Asmita Mahajan: Extreme scary.
[00:52:47] Asmita Mahajan: So you can see that Xtrain is a NumPy 2D array. So, where, every element is the record, or the scaled record from the data, okay? So,
[00:53:00] Asmita Mahajan: So, this array…
[00:53:02] Asmita Mahajan: we want to convert this NumPy array into Tensor, so Tensor is also a form of array only. So, we'll be using… we'll be importing our torch module from,
[00:53:14] Asmita Mahajan: sorry, we'll be importing a torch module to perform, to, you know, extract the neural network layers and the optimizers from the torch module, that's why we are importing this. Then we'll be importing the neural network. NN stands for the neural network, so to,
[00:53:33] Asmita Mahajan: To form the neural network layers, we will be needing this NN module. We'll be importing Optim. Optim stands for the optimizer, so the ADAM, SGD, so all these optimizers will be, coming, we'll be using from the Optim module, and then NumPy as, NP.
[00:53:52] Asmita Mahajan: This we are, using to reproduce the same,
[00:53:59] Asmita Mahajan: Values, or the same weights, further in the… when we are performing this.
[00:54:06] Asmita Mahajan: Afterwards, also.
[00:54:08] Asmita Mahajan: So, when I run this.
[00:54:16] Asmita Mahajan: Okay, so here you can see I have used torch.tensor to convert my dataset, to convert this extreme scale data into a tensor data, because it, previously it was in a NumPy array, as you can see here also, if I hover over this, name. So, I can see it is a ND array with shape 11984.
[00:54:39] Asmita Mahajan: Rows and 14 columns. So… Here, what I'm doing is I'm converting it to a tensor, and if…
[00:54:46] Asmita Mahajan: After this, I print this, it will be… The seam?
[00:54:53] Asmita Mahajan: Array, but only written in a different form.
[00:54:56] Asmita Mahajan: So here, it was written as array, but the only difference is that it is now written as a tensor, because the modules that I'll be using, or the functions that I'll be using further to train my neural network model, they…
[00:55:11] Asmita Mahajan: take tensor as an input. So that's why we have to convert it into Tensor, because we are using PyTorch.
[00:55:17] Asmita Mahajan: Also, we have to explicitly mention that the data type of this train scaled, variable value is off-float data type. It is numeric data type. And in a similar manner, we have to convert the xtest scaled data type, Y train.
[00:55:37] Asmita Mahajan: And also, so, here I have converted Y-Train into Tensor, and also you can see here there is a view
[00:55:45] Asmita Mahajan: function called. So what view does is, So, let me…
[00:55:57] Asmita Mahajan: So, here you can see, my array is a…
[00:56:03] Asmita Mahajan: row vector is in the form of a row vector. Okay, so… to convert… so,
[00:56:11] Asmita Mahajan: When we'll be performing
[00:56:14] Asmita Mahajan: training on the model, when we'll be running, we'll be, we'll be doing the model training. The predictions that I am getting, or the output that we will be getting, is in the form of a column vector.
[00:56:27] Asmita Mahajan: So, to make both the, true labels and the predicted labels compatible with each other, we are, reshaping our, this, tensor
[00:56:39] Asmita Mahajan: into a column vector. So, If I don't do this… And I run this.
[00:56:50] Asmita Mahajan: So, you will see…
[00:56:52] Asmita Mahajan: the NumPy array, which was… the Y train, which was a NumPy array before, it has been converted into a tensor, but it is in the same shape, like the… in a row vector form. But I want it to be converted into a column vector, so that it would be compatible with the outputs of the model. So, for that.
[00:57:12] Asmita Mahajan: what I'm doing is, I'm changing, or I'm reshaping it using view command.
[00:57:17] Asmita Mahajan: So, view does… view, what it does is, it will tell that I want to convert it into a column vector, and the records will be similar. Minus one is that, the records that were previously present in the data, they are similar, but I want it into… I want it to be converted into a column vector.
[00:57:34] Asmita Mahajan: So, after this, if…
[00:57:37] Asmita Mahajan: So you can see my dataset is now converted into a column vector, which was a row vector before. So it has been converted to tensor also, and the shape is also changed.
[00:57:49] Asmita Mahajan: That's why we are using view here.
[00:57:53] Asmita Mahajan: And further, you will see why we have converted it, because our predictions will be, or the output from the model, will be in the form of a column vector, so to make it compatible with the predictions, so that's why we are changing our true labels also into a column vector.
[00:58:22] Asmita Mahajan: Okay.
[00:58:23] Asmita Mahajan: So…
[00:58:25] Asmita Mahajan: now my dataset, is preprocessed. It has been in proper format. Now we will… we are ready to pass it into the model.
[00:58:35] Asmita Mahajan: from scratch, and we will see, how our predictions work, or how they… how we will tune the, hyperparameters. So the first thing that we are tuning is, the first hyperparameter that we have chosen is the optimizer.
[00:58:52] Asmita Mahajan: like, how we will optimize the weights, in the back propagation. So, to tune that, we have taken two optimizers, ADAM and SGD.
[00:59:02] Asmita Mahajan: We will check for both, and with the two learning rates, the first one is 0.001, and the second learning rate would be 0.01. So we will see how the model converges, and how fast the model converges using these two learning rates, and using the ADAM optimizer.
[00:59:22] Asmita Mahajan: So, at the very first step.
[00:59:26] Asmita Mahajan: We'll be exporting, or importing, sorry, we'll be importing these matrices to…
[00:59:33] Durga Toshniwal: Just, I wanted to come in, sorry to interrupt, Asmita. So, actually, there are a variety of optimizers, as we discussed yesterday.
[00:59:42] Durga Toshniwal: However…
[00:59:43] Durga Toshniwal: For sake of making the code computationally slightly lighter, because we can take a variety of other optimizers also.
[00:59:52] Durga Toshniwal: And try, similarly, the learning rate or other parameters that we'll consider. We have considered some range of values or certain set of values, however.
[01:00:04] Durga Toshniwal: That could be much larger.
[01:00:06] Durga Toshniwal: Provided we have the support of compute power behind us, because Colab does not allow a very heavy code, and that would become… make it very, very slow.
[01:00:19] Durga Toshniwal: So, we, we have chosen, comparatively lesser number of…
[01:00:28] Durga Toshniwal: for each of the hyperparameters. So, we have tuned, we have illustrated the tuning of each and every hyperparameter. However, if you have good compute resource, such as if you have a GPU or something available at your end, then you could actually take a bigger variety of the values and try it.
[01:00:47] Durga Toshniwal: So, because we wanted you all to learn how to do it, so we have, included some set.
[01:00:55] Durga Toshniwal: So, that's all I wanted to add here. For those who might be thinking that why two optimizers have been used, why not 3 or 4? So, some common ones we have in… which are very popular, we have included.
[01:01:08] Durga Toshniwal: And, some range of values later on for learning rate and other things we have included. However, if you have better compute, please go ahead and add more values and other optimizers and all that stuff and try it.
[01:01:23] Durga Toshniwal: Okay, that's it from my side, Asmita. You can take over now.
[01:01:28] Asmita Mahajan: Thank you, ma'am.
[01:01:30] Asmita Mahajan: So, we are importing these, matrices so that we can compare the performance of both the optimizers and at different… at these two learning rates.
[01:01:39] Asmita Mahajan: So we have accuracy score, which will evaluate the accuracy of the model, the precision, recall, F1 score, and the AUC score. So, ROC, AUC score will evaluate the AUC of the model.
[01:01:57] Asmita Mahajan: Then, we'll be importing Math Library to…
[01:02:01] Asmita Mahajan: perform the ceiling and floor functions, or to perform the division and power functions. And then, this is the matplotlib to visualize the plots.
[01:02:15] Asmita Mahajan: Again, we are using our seed as 42, so that all the computations that I'm doing, it is reproducible.
[01:02:24] Asmita Mahajan: So…
[01:02:26] Asmita Mahajan: Firstly, we'll be creating an MLP class so that we can easily use this class to prepare our model, or to
[01:02:36] Asmita Mahajan: You know, at each iteration, we create a new model for different hyperparameters.
[01:02:44] Asmita Mahajan: So, this is the initialization function, you know, in the…
[01:02:50] Asmita Mahajan: In this class. So, we will have the input dimension.
[01:02:53] Asmita Mahajan: So, input dimension will tell us, how… what is the size of the input layer.
[01:02:59] Asmita Mahajan: So, we are… using this, we are creating a simple, neural network, which will have only one hidden layer, one input layer, and one output layer. So, input dimension will tell us the size of the input layer. Here, hidden will tell us the size of the hidden layer, how many neurons will be present in the hidden layer.
[01:03:17] Asmita Mahajan: And then, what we are doing is, we are initializing our, this initialization function. Then, we are telling that I want a network.
[01:03:28] Asmita Mahajan: And in that network, I want a neural network in a sequential manner. So we have a sequence of layers. We have input layer, we have hidden layer, and then we have output layer. So I'm telling that I want a sequential network.
[01:03:40] Asmita Mahajan: Oh… And the first…
[01:03:43] Asmita Mahajan: layer, which is the input layer, I want it as a linear layer. Linear layer means we have a, we have, like, the input dimension will,
[01:03:54] Asmita Mahajan: will be used to, define the size of the input layer, and then, linearly, I want my inputs to be inputted into the network, and then, using hidden, so this is a fully connected layer. So, linear will tell us that the first layer
[01:04:14] Asmita Mahajan: That is the input layer, is a fully connected layer, so I will have 14, input, 14 input neurons in the input layer, and then, from those 14 neurons, it is connected to all the neurons in the hidden layer.
[01:04:30] Asmita Mahajan: Which is 64. So these, the 14 neurons in the input layer are fully connected to the 64 neurons in the hidden layer. So linearly, they are connected.
[01:04:41] Asmita Mahajan: So, again, I'm, telling what activation function will be used before passing it to the hidden layer, before the outputs from the input layer are passed to the… as inputs to the hidden layer. So, before that, I'm using the ReLU activation function for that.
[01:04:58] Asmita Mahajan: In between.
[01:04:59] Asmita Mahajan: And then, I'm passing, the raw outputs from the, hidden layer to the, output layer, and here, output layer has only 1. The size is 1.
[01:05:10] Asmita Mahajan: Because we are doing the binary classification, our target class has only two values, or two classes, which is 0 and 1. So, one output neuron is enough to predict my binary classification. That's why
[01:05:25] Asmita Mahajan: In the output layer, we have… it is a fully connected layer. All the, hidden layer
[01:05:32] Asmita Mahajan: Neurons are connected to the output layer, which has only one neuron.
[01:05:38] Asmita Mahajan: Yes, Ankit?
[01:05:40] Ankit Sood: Ma'am, where are we specifying number of hidden layers? Because you mentioned hidden is the number of neurons, right?
[01:05:46] Asmita Mahajan: Number of new releases?
[01:05:49] Ankit Sood: But where are we specifying how many layers are we going to have? How many hidden layers?
[01:05:53] Asmita Mahajan: Here, as I, you can see that, I defined my network to be sequential, so the first
[01:06:02] Asmita Mahajan: linear…
[01:06:04] Asmita Mahajan: layer is the input layer, okay? And as I have passed, the input dimension and the hidden dimension, it automatically tells us that if I have one more linear layer here, and again, one more hidden hair, like, if I'll do like this.
[01:06:26] Asmita Mahajan: Okay? So, I have created two hidden layers.
[01:06:32] Asmita Mahajan: Okay.
[01:06:33] Ankit Sood: portrait.
[01:06:35] Ankit Sood: So we have to write all the layers in this way?
[01:06:39] Asmita Mahajan: We are sequentially, defining our network. So, I have used sequential here, so that I can sequentially define my network. So, the first linear is the, input, layer. The second linear would be the
[01:06:57] Asmita Mahajan: So, first linear will tell us that I have a fully connected layer from the input to the first hidden layer. The second linear will tell me that I have a fully connected neurons from the first hidden layer to the second hidden layer.
[01:07:12] Asmita Mahajan: And the last linear will tell me that, from the second hidden layer to the output layer, I have a fully connected network.
[01:07:21] Ankit Sood: Okay, and ma'am, we have to write, like, if we have 64 hidden layers, we have to write it 64 times?
[01:07:26] Asmita Mahajan: No, if you have… if you're then using 64, then we… we will not define our network sequentially.
[01:07:32] Asmita Mahajan: Then we will use something else, or we will use, where we can define, how many hidden layers we want. Then we will not use sequential.
[01:07:41] Asmita Mahajan: Because my hidden layer were very less here, because I'm using only one hidden layer here. That's why I've, defined my network in the form of sequential… in sequential manner. Like, I will manually, write.
[01:07:53] Asmita Mahajan: But if I have many hidden layers, then I will not use sequential hair.
[01:07:58] Muni Prakash Ganji: But we mentioned hidden as equal to 64, correct?
[01:08:02] Asmita Mahajan: Yes, hidden is 64. It means that the neurons in my hidden layer are 64.
[01:08:09] Asmita Mahajan: It is not that I have 64 hidden layers. It is not that.
[01:08:15] Asmita Mahajan: So, I have only one hidden layer, and one hidden layer, the size of one hidden layer is 64, which means that, the number of neurons in my hidden layer are 64.
[01:08:31] Asmita Mahajan: Is it clear?
[01:08:34] Ankit Sood: Yes, ma'am.
[01:08:37] Asmita Mahajan: Okay.
[01:08:41] Asmita Mahajan: So, I will repeat myself again. So, I'm creating a network where I am sequentially defining my network. So, the first layer will be my input layer. So, input dimension will tell me the size of my input layer.
[01:08:59] Asmita Mahajan: And hidden will tell me the size of my hidden layer. So, linear is, like, I'm telling the network that my input layer and the hidden layer is a fully connected network, so all the
[01:09:14] Asmita Mahajan: outputs from… the input layer are connected to the… all the, neurons in the hidden layer, or all the outputs from the input layer will go as inputs to the hidden layer. So, before
[01:09:29] Asmita Mahajan: passing them as inputs to the hidden layer, I'll be using the ReLU activation function.
[01:09:34] Asmita Mahajan: So here, as we are only attuning the optimizer parameter, all of the other hyperparameters, I am hard-coding it, or I am explicitly mentioning what will be the value. So as here, for the hidden layer, I have chosen 64,
[01:09:50] Asmita Mahajan: 64 as the size of the hidden layer, but we will tune it.
[01:09:55] Asmita Mahajan: We will tune… we will also tune this size in, like, in our further, code.
[01:10:02] Asmita Mahajan: But here, as I am only tuning the optimizer as the hyperparameter, I will not, I am explicitly mentioning the size of the hidden layer. Also, I am explicitly mentioning the activation function that I am using. So.
[01:10:16] Asmita Mahajan: so that my processing is easier. So, I'm using Railu as the activation function in the hidden layer, and in the output layer, all the outputs from the hidden layer are passed as input to the,
[01:10:31] Asmita Mahajan: output layer, and I have only one neuron in the output layer, because I'm doing binary classification.
[01:10:39] Asmita Mahajan: Okay?
[01:10:42] Asmita Mahajan: this forward function will tell us that, how the input will be forwarded to the network. So, it will, the input will, pass on
[01:10:53] Asmita Mahajan: through the whole, network. So, net… so it came from here. We have defined, our network here, so my input X, will be forwarded to the whole network.
[01:11:06] Asmita Mahajan: In the forward pass.
[01:11:08] Dwarakesh T P: I'm sorry to interrupt, just one question. So, you have chosen the number of neurons in the hidden layer of 64, right? So, is there any best practice around it, or how did we start with 64?
[01:11:20] Asmita Mahajan: No, we can choose, any number from, which is, like, we can choose any, number, but it should be in the form of 2 to the power something. So, 64 is, 2 to the power 6, so that's why we have chosen 64 here.
[01:11:38] Asmita Mahajan: So, it is not the hidden layer, it is not the number of hidden layers, it is the size of one hidden layer.
[01:11:44] Asmita Mahajan: Okay.
[01:11:45] Dwarakesh T P: Okay.
[01:11:46] Asmita Mahajan: Yeah, so, we are also… we are also going to optimize the size of the hidden layer,
[01:11:53] Asmita Mahajan: as we move forward. But here, as we are only optimizing the optimizer as the hyperparameter, that's why I've explicitly mentioned that I'm using 64 as my size of the hidden layer.
[01:12:07] Dwarakesh T P: Okay, and .
[01:12:08] Asmita Mahajan: Any number, but it could be in the 2, in the power of 2.
[01:12:13] Dwarakesh T P: Got it, okay. And then can you also exp… how the LR values are taken here? Like, you have mentioned 0.001 and 0.6.
[01:12:21] Asmita Mahajan: So, let me go through the whole code. I'll be, like, explaining that also, how we are passing LR as the hyperparameter.
[01:12:33] Dwarakesh T P: Okay, and one last question.
[01:12:35] Asmita Mahajan: Absolutely.
[01:12:36] Dwarakesh T P: How do we, decide on the number of layers? So, here you are just, taking as one hidden layer, right? So, is that also something that we are going to, in the iterative process, we'll keep on changing it?
[01:12:49] Asmita Mahajan: Yes, yes.
[01:12:50] Asmita Mahajan: Yes.
[01:12:52] Asmita Mahajan: So, the number of hidden layers is also a hyperparameter that we will see further, that we will also optimize. So, it depends on how my model is performing.
[01:13:03] Dwarakesh T P: Okay, so to start with, we can always start with one, or is there any criteria around it? Like, what do we start with when we start building the model?
[01:13:12] Asmita Mahajan: So, no, we don't always start with one, because, it will be computationally very expensive, or it will not… it will take
[01:13:21] Asmita Mahajan: time to run my model if I have chosen a hidden… a large number of hidden layers. Like, if I've chosen 3 as my number of hidden layers, so the time it will take to run this code, it will be much higher. So,
[01:13:39] Asmita Mahajan: like, we are… I'm, running this code side by side, so it will be… it will not be possible to, you know,
[01:13:47] Asmita Mahajan: Complete the whole, processing in one class only.
[01:13:52] Asmita Mahajan: So, to make it simple, or to make it computationally less expensive, or to make it computationally feasible, we are choosing one hidden layer here.
[01:14:02] Asmita Mahajan: But it is not always the case that we choose… we start with one hidden layer.
[01:14:08] Dwarakesh T P: Okay, got it. Thank you.
[01:14:11] Ankit Sood: Ma'am, quick question on the comment that you made, right? Like, we always have to go with 2 to the power something.
[01:14:18] Ankit Sood: Like, is it, like, something which is identified by experimenting with neural networks or deep learning? Because odd numbers can also be there, right? So how have we came to that conclusion that it always have to be 2 raised to the power something?
[01:14:34] Asmita Mahajan: Okay, so it comes from the very basic when we learn about, like, computers and CPUs. So, the computer works in a binary form.
[01:14:46] Asmita Mahajan: So, yeah, so whenever we are, you know, doing anything, we do it in… we try to do it in… or we practice to do it in a 2 to the power form, so that it is easy for, my, you know.
[01:15:04] Asmita Mahajan: The low-level language to convert it into,
[01:15:10] Asmita Mahajan: like, what I'm doing here is I am writing a high-level language code. So to… but my computer doesn't understand this language. It will always convert it into zeros and ones.
[01:15:21] Asmita Mahajan: So, when.
[01:15:24] Ankit Sood: I understand that part, ma'am, but how is that related to the number of neurons that you're keeping?
[01:15:30] Ankit Sood: Here, right? Like, you're keeping 64, and you mentioned that
[01:15:33] Ankit Sood: it is… it always has to be two raised to the power of something. I was also reading it yesterday about deep learning, and they also mentioned that
[01:15:40] Ankit Sood: it has to be 2… or, like, it's beneficial if you keep it as 2 raised to the power something, but I couldn't relate, like, why it has to be 2 raised to power something. Why not 63? Why 64?
[01:15:52] Ankit Sood: And why not 62?
[01:15:54] Ankit Sood: is what I was trying to grasp, that why I have to keep it always in 2H to the power something.
[01:15:59] Asmita Mahajan: Because whenever we are performing… so here I'm, we are not, talking about the mathematics behind the, our model.
[01:16:07] Asmita Mahajan: So,
[01:16:09] Asmita Mahajan: when we… when… if, you will learn about the, mathematics behind deep learning, you will always see that we are using gradients, or we are using, partial derivatives. So there, or we are using logs. So there, we always use, logs to the base 2.
[01:16:28] Asmita Mahajan: So, when we are using logs to the base 2, so, here, if I convert it, it will be 2 to the power 6. So, when… if you are using logs, and if you are using gradients, partial derivatives, so it is a best practice that we always use it in the form of 2 to the power, so that the conversion is easier.
[01:16:48] Asmita Mahajan: So to make the conversion easier.
[01:16:50] Asmita Mahajan: For the mathematics behind deep learning, we always, it is a best practice to always choose numbers which are in the power of twos.
[01:16:59] Ankit Sood: Okay, and it will not have any impact on my accuracy, like, if I go with 62 versus 64, though 64 is computationally less intensive, but on the accuracy, it will not make much of a difference, is what we are trying to do.
[01:17:12] Asmita Mahajan: like, 62 and 64 will not make much of a difference, but if you will say, 128 and 64, that will make a difference.
[01:17:22] Ankit Sood: Okay.
[01:17:25] Ankit Sood: Good. Okay, thank you.
[01:17:27] Asmita Mahajan: Thank you.
[01:17:28] Dwarakesh T P: One follow-up question. So, when you have, like, multiple hidden layers, is there any criteria, like, the first hidden layer and second hidden layer should have the same number of neurons, whether it should increase, decrease, or how is it?
[01:17:40] Asmita Mahajan: No, there is no criteria. It is also, like, it is also tunable. Like, how many, neurons should be there in the first hidden layer, and how many neurons should be there in the second hidden layer. You have to tune that.
[01:17:53] Asmita Mahajan: Also. So there is no criteria that the first hidden layer should have, or the number of neurons in the hidden layers should be same.
[01:18:02] Asmita Mahajan: So, it is not, you have to tune that.
[01:18:06] Asmita Mahajan: You have to see, like, at what number your model is performing well.
[01:18:14] Dwarakesh T P: Okay, good. Thank you.
[01:18:16] Asmita Mahajan: Thank you.
[01:18:18] Asmita Mahajan: Okay, so we have, created our network, we have created our MLP network, which is a very basic network with one handle layer. Now, we are defining the makeup optimizer, function.
[01:18:31] Asmita Mahajan: which takes the model as the parameter. The model is, like, when we are using the training. So.
[01:18:40] Asmita Mahajan: We are creating Model for each.
[01:18:44] Asmita Mahajan: parameter, like, a hyperparameter. The name, tells us what is the name of the optimizer that I am passing it on, and the LR, what is the value of the learning rate. So, these three parameters will be inputted to this function here.
[01:19:02] Asmita Mahajan: So, if my name is Adam.
[01:19:06] Asmita Mahajan: If, I have inputted, like, if I've given the name parameter here as Adam, so what I am doing is, I will…
[01:19:15] Asmita Mahajan: pass on, or I will call the function atom from the Optim module.
[01:19:20] Asmita Mahajan: And, I will pass the LR, the value of LR, which I have defined here, to the optimizer atom. Else, if my name is SGD, if I've passed SGD in the name parameter, so I'll be calling the SGD
[01:19:37] Asmita Mahajan: function, or the SGD function from the Optim module, which I have loaded here.
[01:19:44] Asmita Mahajan: So here we imported the optimizer module. So from that optimizer module, I'm calling the function SGD, and I'm passing the LR as the hyperparameter, the value of LR, which I have defined here, okay?
[01:19:59] Asmita Mahajan: So… This function will only call my,
[01:20:05] Asmita Mahajan: atom optimizer function or the SGDOptimizer function. And else, if I have passed any other name, so I will just raise the error that, the name that I have mentioned is, wrong.
[01:20:20] Asmita Mahajan: Okay?
[01:20:21] Asmita Mahajan: Now, the main code comes here, which is the train and evaluation function.
[01:20:28] Asmita Mahajan: So, it takes optimizer as the parameter, the learning rate as the parameter, and epochs I have already defined, which is 200. I want to run my… I want to train my model
[01:20:40] Asmita Mahajan: 200 times, so, it will iterate 200 times, it will optimize, and then my final model will be created. So, Epox tells us that.
[01:20:51] Asmita Mahajan: So, this is predefined, because I'm not optimizing epochs right now, I'm only optimizing my optimizer, or I am… sorry, I'm not tuning epochs right now, I'm only tuning optimizer, so that's why I have explicitly mentioned, or the epochs are predefined.
[01:21:09] Asmita Mahajan: Okay, so I'm, storing, I am creating an instance of the class MLP, and I'm storing it, in model. So, what I'm doing is, I'm telling
[01:21:19] Asmita Mahajan: it, firstly… so, MLP takes input dimension as a parameter and hidden as a parameter. So, hidden, I have already mentioned that, my hidden layer will be 60… will have 64 neurons, but what will be the input dimension?
[01:21:36] Asmita Mahajan: So, I have 14 features.
[01:21:38] Asmita Mahajan: So, 14 features from where it came. The dataset has 14 attributes, so those 14 attributes will be passed… the values of these 14 attributes will be passed to the input layer. So that's why my input layer should have 14 neurons. And to… how I'm going to extract that.
[01:21:55] Asmita Mahajan: So, when, we saw here the shape.
[01:22:01] Asmita Mahajan: Like here. As you can see, when we printed this X.shape, we got what?
[01:22:08] Asmita Mahajan: the number of records, the number of rows, and the columns. So, from this shape.
[01:22:13] Asmita Mahajan: the, so shape is returning a tuple, and that tuple contains the number of rows and the number of columns. So I want to extract these, the number of columns, and it is at the position 1. So what I'm doing is, I'm passing it.
[01:22:27] Asmita Mahajan: In my input dimension. So I'm calling X, it on X train tensor, and I'm calling the…
[01:22:33] Asmita Mahajan: function shape. So, shape will give me a return… me a tuple, which has my number of rows and my number of columns. The number of rows will be at index 0, but my number of columns is at index 1, that's why I'm, extracting my number of columns using 1.
[01:22:49] Asmita Mahajan: Okay? So, this whole portion of the code will give me, the input dimension.
[01:23:01] Asmita Mahajan: Okay, so, how I'm, optimizing, or how I'm, like,
[01:23:08] Asmita Mahajan: passing my inputs, into the… in the forward pass, so I'm using BCE, which is binary cross entropy, with logits. So, logits, here also, as you can see, in, in the, when I was passing my,
[01:23:25] Asmita Mahajan: outputs from my hidden layer to the output layer, I haven't used any activation function. So my raw outputs are passed as inputs to the output layer, and my raw outputs are,
[01:23:36] Asmita Mahajan: I'll be returning the raw outputs from the output layer. So, I have not converted the raw outputs into any probability, or into hard classes. Hard classes means that I have not converted the output to 0 and 1, because here in the output layer, some value would be there.
[01:23:56] Asmita Mahajan: Okay, but that value will not be 0 or 1. We have to convert it into 0 or 1.
[01:24:01] Asmita Mahajan: So that I can predict my target class. But here, I'm getting raw outputs from the output layer. So, that's why I have used the criterion here as binary.
[01:24:14] Asmita Mahajan: cross entropy with logit's loss. Logit's loss means that I am using raw outputs to, you know, evaluate my losses.
[01:24:24] Asmita Mahajan: So that criteria I have used.
[01:24:28] Asmita Mahajan: then I'm, creating an instance of the optimizer, like, here.
[01:24:36] Asmita Mahajan: whatever…
[01:24:38] Asmita Mahajan: is returned from here, like, this atom has been called, and it has been returned from my class. So, I'm storing it in the optimizer variable, and I'm calling, makeOptimizer, function with the model.
[01:24:53] Asmita Mahajan: which… with the MLP class instance, optimizer name, and the LR. So…
[01:25:01] Asmita Mahajan: We'll also be, defining these further.
[01:25:06] Asmita Mahajan: Then, I'm, creating train loss and test losses, array to store the train losses and the test losses for each epoch.
[01:25:15] Asmita Mahajan: For each epoch means for each iteration. So, we'll be having 200 iterations of the model, and for each iteration, there will be a train loss, there will be a test loss, and that we'll be storing in an
[01:25:26] Asmita Mahajan: in a list.
[01:25:29] Asmita Mahajan: So that's why I've created this. So, this will be for… this will be further used to plot, the train, loss curve and the test loss curve.
[01:25:44] Asmita Mahajan: So for each iteration, or for each epoch, what I'm doing is, I'm training my model, so model.train will,
[01:25:53] Asmita Mahajan: open this MLP, it will, convert this MLP into a train, training mode. So, this will train my model. So,
[01:26:04] Asmita Mahajan: my MLP, net… so my network has gone into the training mode using this, line of code.
[01:26:12] Asmita Mahajan: optimizer.zeroGrad will just clear the, previous gradients, so in each epoch, we'll be… we will be, calculating gradients, or we'll be calculating the partial derivatives in the, backward pass. So, to clear that, so that in each iteration, I have,
[01:26:32] Asmita Mahajan: We'll initialize the gradients as 0, so that's why we are doing this, so that the previous… the gradients from the previous part doesn't leak into the next iteration.
[01:26:44] Asmita Mahajan: So, this line of code is just clearing my gradients and initializing them, freshly.
[01:26:51] Asmita Mahajan: Okay, so model extraint tensor, so I'm passing.
[01:26:55] Asmita Mahajan: So, we have, opened the network in the training mode, then I'm passing the, input
[01:27:03] Asmita Mahajan: the training set to the model, and I'm… Logis will have my raw outputs from the model. The output layer, the… whatever the output layer has predicted, it will be passed on to Logis.
[01:27:17] Asmita Mahajan: Now, I'm using those outputs and the true labels to, calculate my loss. So here, we have used the binary cross entropy loss, so that loss will be stored in this loss variable.
[01:27:32] Asmita Mahajan: So, what I'm doing is lodges. Lodges we got from the output layer, and by-train tensor, we already had it. These are the true labels, or the true classes of the dataset, and using this, we are, calculating the cross entropy loss, and we are, storing it in the loss variable.
[01:27:52] Asmita Mahajan: Then, this loss is being passed in the back propagation process, so loss.backward, so I'm calling this loss, and I'm calling the backward propagation, or the backpropagation function, to propagate my gradients to the network.
[01:28:11] Asmita Mahajan: Also, I'm using optimizer.step, so what it is doing is, it is telling that, I need to optimize the weights, or update the weights using the atom optimizer, or either the SGD optimizer, whatever we have passed in the name, variable.
[01:28:30] Asmita Mahajan: So, optimizer step will, basically just call out that I have to update the weights using the specified, optimizer.
[01:28:43] Asmita Mahajan: Then, here I'm simply appending, to my list the loss that, I have computed. So here, whatever the loss I have computed in this iteration, I'm appending it to the train loss, list.
[01:29:01] Asmita Mahajan: Now, to evaluate the model on the test set, I am,
[01:29:07] Asmita Mahajan: converting my model in the evaluate mode, and with…
[01:29:11] Asmita Mahajan: torch.noGradient. So here, I don't want to calculate the gradient, because what I'm doing is I'm just, whatever the model has learned, using the training set, I just want to use those weights, and I'm passing my test set as the parameter now.
[01:29:28] Asmita Mahajan: So here you can see, to the model, I have passed the test set.
[01:29:33] Asmita Mahajan: And I don't want any gradients, like, I don't want any…
[01:29:36] Asmita Mahajan: updates to the weights, so that's why no grad is used. So I, I don't want my model to learn. What it has already learned, I want, those weights to be used, to, you know, in the forward pass.
[01:29:51] Asmita Mahajan: when I'm passing the test set. So here, when I'm, like, calling the model on the test set, and my output will be stored in the test logist.
[01:30:02] Asmita Mahajan: Similarly, that, like, what we have done here, when I was training my model. So, on the test set, also, I will do that only. I will, calculate the loss using the cross entropy loss, bypassing the test lodges, the output from the model, and the true label.
[01:30:19] Asmita Mahajan: In the test set, and…
[01:30:22] Asmita Mahajan: So, item is, like, we have a tensor. So, as we are working, in PyTorch, we are working on tensors, so to convert that to a float.
[01:30:34] Asmita Mahajan: what I'm doing is, I'm, writing it in, like, I'm converting… dot item will convert the tensor into float, and it will be stored in the test loss function, test loss variable.
[01:30:46] Asmita Mahajan: And this test loss, I'm appending it to the list, test losses.
[01:30:51] Asmita Mahajan: For this iteration.
[01:30:55] Asmita Mahajan: Okay?
[01:30:57] Asmita Mahajan: So this was basically… what I have done is, I have just run one iteration. So, it will, so this for loop will run for 200 iterations, and after that, my model will be, finalized.
[01:31:12] Asmita Mahajan: So, after every iteration, the optimizer step will be called to update the weights, and it will run for 200 times, and after the 200 iteration, my… the final weights of the model, are my fixed rates, and that I will be using on the full test set.
[01:31:31] Asmita Mahajan: So, again, I'm, converting my model into evaluate mode with no gradients, because I'm, like, the final model, this final model now, I want to use it on the test set, so…
[01:31:44] Asmita Mahajan: I'm calling my model on the test set, and storing the output in lodges.
[01:31:50] Asmita Mahajan: And no!
[01:31:52] Asmita Mahajan: So, I want my output in the form of 0 and 1.
[01:31:57] Asmita Mahajan: So, to do that, what I'm… as I have not used any,
[01:32:01] Asmita Mahajan: activation function here, when I was defining my model. So, to convert the raw outputs into probabilities, what I will do is, I will call the sigmoid activation function.
[01:32:13] Asmita Mahajan: So here, as you can see, torch.sigmoid, so the sigmoid activation function is being called, and lodges is given as the… the raw outputs from the model is given as the parameter to this sigmoid activation function. So what it will do is, it will convert my raw outputs in the range 0 to 1, so that, they are converted into probabilities. So these probabilities I'll be storing in probs,
[01:32:35] Asmita Mahajan: variable, and using this prop variable now, I will be, predicting, or I will be getting the predictions, in the form of 0 and 1. So for that, what we are using, we are using a threshold. If these probabilities are greater than 0.5, then 1 will be stored in the
[01:32:54] Asmita Mahajan: prediction, in the predictions, but if these probabilities, the probability that I'm getting, if they're less than 0.5, 0 will be stored. And, .float is just converting the, predictions to float.
[01:33:10] Asmita Mahajan: So it will ensure that my prediction variable is a float variable. Yes, Ange?
[01:33:18] Ankit Sood: Ma'am, this prediction I thought we'll do in each iteration. This is outside of the for loop, right? So, we'll do it only once?
[01:33:25] Asmita Mahajan: Yes, we'll do it only once, because we are… we are using the whole model, the whole learned model, to, test our training, sorry, the whole learned model on our testing dataset.
[01:33:38] Asmita Mahajan: So, it will be done only once.
[01:33:40] Ankit Sood: So, before…
[01:33:42] Asmita Mahajan: Yes, yes.
[01:33:43] Ankit Sood: Sorry, please, please complete your sentence right.
[01:33:45] Asmita Mahajan: So before, when we were, using this, law… so to…
[01:33:50] Asmita Mahajan: like, we haven't done it here. In every iteration, we haven't done it, because we don't need to do it.
[01:33:57] Asmita Mahajan: Why? Because, the… we are just learning. Here, we are… in this, for loop, we are making our model learn from the training set.
[01:34:07] Asmita Mahajan: So here, we need not to convert, or we need not to get the predictions from the model. We only had… we are only working with the losses.
[01:34:18] Asmita Mahajan: Okay? Which is the cross entropy loss.
[01:34:25] Ankit Sood: So we'll not ask our model to predict, to understand whether the prediction was correct or wrong in this iteration, so that… because we are adjusting the weights, right, in this, for loop.
[01:34:37] Ankit Sood: So, without even, calculating the probability, we would know that it was wrong or right, the prediction?
[01:34:45] Asmita Mahajan: Yes, because, my cross-entropy laws, it takes raw, raw outputs from the model, and it uses that raw… those raw outputs to, you know,
[01:35:01] Asmita Mahajan: compute the laws. So this, BCE function.
[01:35:05] Asmita Mahajan: Why we haven't, converted our,
[01:35:10] Asmita Mahajan: raw outputs into probabilities, because BC function internally does that. It internally uses a sigmoid function to convert my raw outputs into probabilities, and then using those probabilities, it converts into predictions internally, and compute the loss.
[01:35:28] Ankit Sood: Okay.
[01:35:30] Asmita Mahajan: So, if, see, this loss combines… if you will, learn about this BC function, so I hover over this BC function, and, you can see, the… what this does is, this loss combines the sigmoid layer and the BC loss in one sig… single class. So, internally, it is doing that.
[01:35:50] Asmita Mahajan: So, I need not to do it explicitly.
[01:35:54] Ankit Sood: So then why we are doing it at the end, if it is already doing it?
[01:35:57] Asmita Mahajan: that, that time I'm not computing laws, that time I have to make predictions. So I have to explicitly make the predictions.
[01:36:05] Ankit Sood: Oh, okay.
[01:36:06] Asmita Mahajan: This function is computing loss. It is, computing loss, okay?
[01:36:11] Ankit Sood: Got it, man. Yeah.
[01:36:18] Asmita Mahajan: So, after training my model for 200 iterations, I'm using the same, weights which were finalized in these iterations to run my test set on the model, and then,
[01:36:34] Asmita Mahajan: like, extracting or capturing the raw outputs in the logist variable. Using this logist variable, I'm inputting it to the sigmoid function to convert my raw output into probabilities, and using those probabilities and this 0.5 as a threshold, I'm converting… I am getting my predictions.
[01:36:55] Asmita Mahajan: Okay?
[01:36:56] Asmita Mahajan: No.
[01:36:58] Asmita Mahajan: Oh… Okay, so, here I have got my predictions, I also have my two, true labels, but
[01:37:07] Asmita Mahajan: as, we have already converted our data into a tensor, now, to run this, accuracy, so this we have imported from the sklearn mat, sklearn library. So, the sklearn library works with NumPyArray. So,
[01:37:27] Asmita Mahajan: like, it's a to-and-flow process. If we are using PyTorch, it works with tensors, so that's why we have to convert the NumPy array into tensors. But when I have to use the functions from the sklearn library, so sklearn works with NumPyArray, and now the converted tensors
[01:37:45] Asmita Mahajan: have to be reconverted into a NumPy array, so that's why this line of code is important. These lines of code is important. So what I'm doing is, I'm,
[01:37:56] Asmita Mahajan: converting my true labels into the NumPy array. So, this CPU, what it does is, it brings back my data set on the CPU. So PyTorch is basically… was basically introduced because,
[01:38:10] Asmita Mahajan: as we will learn, my deep learning models would be very, like, they will have a vast architecture where we have many number, like, a huge number of hidden layers, we have huge number of neurons. So, to run my,
[01:38:29] Asmita Mahajan: full architecture, parallelly, I'll… we will be… we will use the GPUs.
[01:38:35] Asmita Mahajan: further. So,
[01:38:37] Asmita Mahajan: PyTorch gives us the capability to run my… each proportion of my code on different GPUs parallelly, and then, you know, combine the data,
[01:38:49] Asmita Mahajan: After the, processing. So, to,
[01:38:55] Asmita Mahajan: to bring my data back on the CPU, I'm using… we will be using this CPU command, but as we haven't used any GPU, so this command is, like, of no use here, but
[01:39:08] Asmita Mahajan: If we have used GPUs, so this command would be of use, because I want to bring my data back on the CPU, and then I will convert my tensor into a NumPy array, and Ravel is used to flatten.
[01:39:21] Asmita Mahajan: So as, before, what we did was we reshaped the tensor into a column vector, but now I want it, back into my row vector, so RAVL is used to flatten my
[01:39:34] Asmita Mahajan: Column vector into a row vector.
[01:39:38] Asmita Mahajan: So, in this manner, my Y2 labels, are stored, in Y2, which is a numpy array and a row vector.
[01:39:49] Asmita Mahajan: Similarly, with the predictions, so, these predictions I have captured here, so what I am doing is I'm bringing back the predictions to the CPU, I'm converting into a NumPy array, and I'm flattening into a… it into a row vector, and then I'm storing it in by predict variable.
[01:40:07] Asmita Mahajan: And also, these probabilities, which I have calculated here.
[01:40:11] Asmita Mahajan: I'm also storing these probabilities, bringing back it on the CPU, converting into NumPy, and then flattening it to rhoVector.
[01:40:21] Asmita Mahajan: And then, storing the buy probabilities. Yes, Ankit?
[01:40:26] Ankit Sood: Ma'am, will it matter if you are using Ravel here or not? Because it has to be only one element inside this matrix, right? Because it's 0 or 1.
[01:40:36] Asmita Mahajan: No, it will be 0 or 1, but it will be for the whole record.
[01:40:41] Asmita Mahajan: Whiteest has 3,000 rows.
[01:40:45] Ankit Sood: Right, okay, understood. Okay.
[01:40:49] Ankit Sood: So it, like, 1 cross 3000 versus 3,000 cross 1, is what we're trying.
[01:40:54] Asmita Mahajan: Yeah, yes, yes, that's what we are doing.
[01:40:56] Ankit Sood: Okay.
[01:41:01] Asmita Mahajan: So,
[01:41:03] Asmita Mahajan: Okay, so we have these, Y2 labels, Y predictions, and Y probabilities. So now, what we are doing is, we are just calculating the accuracy, precision, recall, F1 score, and AUC score for the model that we have just built.
[01:41:21] Asmita Mahajan: Yes, do you have any questions?
[01:41:25] Jithu Tagore: We already, convert… use that probability to get the prediction, right? Like.
[01:41:32] Jithu Tagore: Produce equal to prop greater than or equal to 0.5.4.
[01:41:37] Jithu Tagore: So, since we got the prediction, why are you… why are you converting that prob… probably… using the probability again, like, using… converting to Numpe and Rava?
[01:41:49] Asmita Mahajan: Yes. So, to calc… to evaluate my AUC score, my AUC score doesn't take hard prob… hard predictions as input. It takes probabilities as inputs, as you can see here.
[01:42:02] Asmita Mahajan: I have passed the probabilities to the AUC score function. So that's why I need to store these probabilities, so that I can use it here.
[01:42:13] Asmita Mahajan: It doesn't take hard predictions as the input.
[01:42:16] Asmita Mahajan: So, all the other matrices, it takes hard predictions, like the Y prediction that we have calculated as input, and the true labels as input, and we'll be passing it to the functions, their designated functions, their accuracy score, precision score, recall score, and F1 score, and using this, we'll be evaluating the matrices. For the AUC score, we'll be taking the true labels.
[01:42:41] Asmita Mahajan: And the probabilities that we got.
[01:42:43] Asmita Mahajan: From the output layer.
[01:42:45] Asmita Mahajan: So, I'm passing, so we'll be passing this, and we'll be evaluating the AUC score. And, so metrics, we are… we are creating a dictionary, where each key is the name of the metric… metrics that we have evaluated, and the value would be the
[01:43:03] Asmita Mahajan: a value which has been evaluated. So, this matrix is a dictionary.
[01:43:07] Asmita Mahajan: In this form.
[01:43:10] Asmita Mahajan: Now, history… history will contain the training loss and the testing loss that we have, evaluated and stored during the… during each iteration, during the 200 epochs. So, here, it is also a dictionary, which, history is a dictionary, which will, store our training losses and test losses, and we are returning
[01:43:30] Asmita Mahajan: So, this train and evaluation, function, We'll return the metrics.
[01:43:35] Asmita Mahajan: and the history. So using these now, we'll be seeing how our model has performed on different optimizers and different learning rates.
[01:43:47] Asmita Mahajan: So, config is a list of tuples. So, we have created a list of tuples where we are defining what would be the optimizer that I am using, and what will be the value of the learning rate. So.
[01:44:00] Asmita Mahajan: So, I have 2 optimizers and 2 learning rates, so 4 different combinations will be there, and 4 different experiments will be performing. So, for the first experiment, SGT will be used as the optimizer, and 0.001 will be used as the learning rate.
[01:44:14] Asmita Mahajan: For the second, similar, SGD is the optimizer, but 0.01 is the learning rate now.
[01:44:20] Asmita Mahajan: For the third experiment, Adam is the optimizer, and 0.001 is the learning rate. And for the fourth experiment, Adam is the optimizer, and 0.01 is the learning rate. So these four experiments will be performing.
[01:44:32] Asmita Mahajan: Results and histories will store, the result. The result means the, values of the metrics, from each experiment.
[01:44:45] Asmita Mahajan: And histories will store the train and test losses from each experiment now.
[01:44:50] Asmita Mahajan: Okay?
[01:44:52] Asmita Mahajan: And it will, they are dictionaries, so we'll be storing them in a dictionary.
[01:44:58] Asmita Mahajan: So, now I'm calling… I'm iterating over each element in my config list, and each element in the config list is a tuple, so I'm extracting the optimizer and the learning rate from each element in the config.
[01:45:15] Asmita Mahajan: storing it in OPT and LR, and iterating over, them these four ex… these four values, or elements. So, key… so, what I'm doing here is, I'm defining a key which will be used as a
[01:45:30] Asmita Mahajan: key to the result and history dictionaries, which we have initialized here. So, the key to these results and histories will be in the form
[01:45:41] Asmita Mahajan: So, this is defining the form that… how I want my keys in this result and history dictionaries. So, how they will be? So, OPT contains what? In the first experiment, OPT will be SGD.
[01:45:54] Asmita Mahajan: At the rate, LR will be 0.001. So my key will be in the form this.
[01:46:00] Asmita Mahajan: So, SGT at the rate 0.001 will be my first key, and the value corresponding to that, what will… what will be the value corresponding to that? It will be the metrics.
[01:46:10] Asmita Mahajan: and the history, corresponding to this key. And it will be stored in result and history dictionary, okay? So here, I'm just defining how I want my key to be.
[01:46:21] Asmita Mahajan: So, it will be, SGD,
[01:46:24] Asmita Mahajan: Then, for the second experiment, it will be SGD at the rate 0.01.
[01:46:31] Asmita Mahajan: For the third, it will be Adam.
[01:46:34] Asmita Mahajan: Sorry. At the rate 0.001. And for the fourth experiment, my key will be atom, at the rate 0.01.
[01:46:43] Asmita Mahajan: Okay?
[01:46:45] Asmita Mahajan: So these are, my keys will look like this.
[01:46:50] Asmita Mahajan: Then, we are calling the train an evaluation function, where we are passing the optimizer, the LR, and this optimizer for the first experiment will have the value SGD, and LR for the first experiment will have the value 0.001, and epochs we have already defined, which is 200.
[01:47:10] Asmita Mahajan: And, after calling this function, what will I… what will my function return? My function will return metrics, which
[01:47:19] Asmita Mahajan: are evaluated here, and the history, the trading loss and test losses. So, we are capturing them here, using metrics and history variables, and then we are storing the metrics
[01:47:34] Asmita Mahajan: for the first experiment, and the key would be this, SGD at the rate 0.001. So, at this key, the value would be metrics. Metrics is the whole dictionary.
[01:47:45] Asmita Mahajan: This whole dictionary will be stored at key SGD, at the rate 0.001. And, in history's, dictionary, at key the, this.
[01:47:56] Asmita Mahajan: My full… this dictionary will be stored.
[01:48:01] Asmita Mahajan: the loss dictionary.
[01:48:04] Asmita Mahajan: Okay?
[01:48:06] Asmita Mahajan: So, in this manner, we will perform the four experiments? Yes?
[01:48:13] Dwarakesh T P: Yeah, so ma'am.
[01:48:14] Dwarakesh T P: Why are we coming up with, these LR values? Could you just explain, like, what is the significance of each of it?
[01:48:21] Asmita Mahajan: So, in my optimizer, when I'm calling my optimizer class, here.
[01:48:28] Asmita Mahajan: So, this LR function, when you have learned about the gradient descent algorithm, so there will… there was a learning rate, like, how slowly my model will learn, or how fast my model… so these are the steps. So, at every iteration, my weights will be updated. So, how my weights will be updated, so there is a formula, like, old weight.
[01:48:51] Asmita Mahajan: Plus?
[01:48:52] Asmita Mahajan: the… this LR into the,
[01:48:57] Asmita Mahajan: So, this LR is the step.
[01:49:00] Asmita Mahajan: Like, how I'm going to update my weight. And if, this LR is very small.
[01:49:06] Asmita Mahajan: my updations will be small, okay? But if I have a higher LR, my updations will be fast. But this also, we have to see, if my LR is very small, the model will converge very slowly, because at every iteration, it will take a small step to the optimum value.
[01:49:25] Asmita Mahajan: in the gradient descent algorithm. But if I have a large LR… if I have a very large LR, my model will oscillate between
[01:49:35] Asmita Mahajan: and it will never reach the optimal value. So that's why we have… we are using here two, LR values to see, at which LR value my model will perform better. If we, if we have a smaller LR, or if we have a slightly greater LR.
[01:49:54] Asmita Mahajan: So these are the steps.
[01:49:56] Dwarakesh T P: That'd be…
[01:49:57] Dwarakesh T P: Which, or best practices around this, LR values, like, any predefined range, or best practices, okay, your LR can start… it's a good practice to start with this LR, and you can try up to this one, so any such range defined?
[01:50:13] Asmita Mahajan: So this is also a tunable parameter, but here, as we are only tuning optimizer, that's why we are predefining my LR values.
[01:50:22] Asmita Mahajan: Okay, but we will also optimize this, or we will also tune this in the further code.
[01:50:29] Asmita Mahajan: But here, as we are only tuning my optimizer, that's why we have taken a predefined number here.
[01:50:36] Asmita Mahajan: But it is also a tunable parameter. That's why I said if my LR would be very small, my model will converge slowly. But if it is large, it will oscillate between the values, but it will not converge, or it will not reach the local minimum, or the optimum value for the weights.
[01:50:54] Asmita Mahajan: So, it is also tunable, which we will tune, further in the code, but here, as we are only tuning the optimizer, that's why we have chosen it.
[01:51:04] Asmita Mahajan: Explicitly.
[01:51:07] Dwarakesh T P: Okay.
[01:51:08] Dwarakesh T P: Thank you.
[01:51:10] Asmita Mahajan: Okay? So, this for loop will perform my experiment 4 times, because I have 4 different combinations here.
[01:51:21] Asmita Mahajan: After this.
[01:51:24] Asmita Mahajan: what I'm doing is I'm creating a data frame using the data frame function, and I'm passing results to it.
[01:51:34] Asmita Mahajan: So, results will have…
[01:51:38] Asmita Mahajan: these… these as keys, the optim… the name of the optimizer, and the value of the LR, the learning rate. So, these are the keys, these are the four keys, and corresponding these four keys, I will have this dictionary.
[01:51:57] Asmita Mahajan: full dictionary, okay? So, this full data will be… so if, let me show you if I'll run this…
[01:52:11] Asmita Mahajan: It will take 10 seconds.
[01:52:20] Asmita Mahajan: Okay, so in my results, if I'll print my result…
[01:52:29] Asmita Mahajan: As you can see, I have this first key, and corresponding to that key, I have the whole matrix for that key.
[01:52:37] Asmita Mahajan: Then, I have my second key, as I have shown you here, that this will… my keys will be in this form.
[01:52:44] Asmita Mahajan: So, you can see in the result, when we are printing result also, this is my key, and corresponding to that key, I have my full metrics for that key, for that experiment. So, in this manner, I will have 4 keys and 4 values. So, now I'm using this data.
[01:53:01] Asmita Mahajan: to, build a data frame. So how I'm doing that, I'm taking
[01:53:07] Asmita Mahajan: results, and I'm passing it to the data frame function, but Okay.
[01:53:15] Asmita Mahajan: I'm also transposing it. So, T is, like, we are transposing, we are converting the rows into columns and columns into rows. So, we'll also see what it does is… so PD dot…
[01:53:28] Asmita Mahajan: data frame results. If I'll…
[01:53:31] Asmita Mahajan: converted, without transposing, so I will get, the keys on the… as my columns, so my key, the keys will be converted as columns, and the matrices, are being given to the index. But I don't want it in this form.
[01:53:48] Asmita Mahajan: So, because I want to display it in another form, so that's why I just want my keys to be the indexes of my data frame, and these, metric values, or the metrics name, I want this to be my columns. So that's why I'm transposing it.
[01:54:08] Asmita Mahajan: So, when I transposed it, you can see that my keys are being the indexed of the data frame, and the name of the matrix are the columns now.
[01:54:21] Asmita Mahajan: So, it's just how I want to display my data frame, so that's why I want my data frame in this form, that's why I've transposed it, okay?
[01:54:37] Asmita Mahajan: So, after converting my results data frame, now I will plot it. So,
[01:54:44] Asmita Mahajan: we'll be plotting the loss grid, the training and test loss that we have captured for every iteration, so we will be plotting that, using this function, so I'm making a function for the plot.
[01:54:59] Asmita Mahajan: Histories is a… here, we… we have made a dictionary.
[01:55:04] Asmita Mahajan: history.
[01:55:05] Asmita Mahajan: Where, for each experiment, I have my training loss and test loss.
[01:55:12] Asmita Mahajan: So, using that history dictionary, maximum column, I have defined that, so we'll be plotting subplots. Like before, we did in k-means, we plotted subplots. So, for those, subplots, my, I want my maximum columns to be 2.
[01:55:29] Asmita Mahajan: And the title for the whole plot is, being given here. So I have defined the, title. I have written the title. So, this title, which you… is displayed for the whole plot, is given hair.
[01:55:43] Asmita Mahajan: So, now I'm converting the keys from the history dictionary into a list, and I'm storing it in keys variable.
[01:55:52] Asmita Mahajan: So length will just calculate the length of this, list, so it… n will have the value 4, okay?
[01:56:01] Asmita Mahajan: I'm also, getting the columns and rows,
[01:56:07] Asmita Mahajan: I'm calculating them explicitly. So, what I'm doing is, whatever the maximum value I have defined here, and whatever the value I get for, here, after the… after calculating the length of the keys.
[01:56:21] Asmita Mahajan: So, the minimum value I'm taking as columns, and rows I'm calculating, like, how… the number of keys I have, divided by the columns, because as I have 4 keys, but I only want 2 columns, so how many rows will be there? 4 divided by 2, 2 rows will be there.
[01:56:38] Asmita Mahajan: Okay, so in this manner, so seal function has been used, so math.seal. So that's why we have imported a math, library before.
[01:56:48] Asmita Mahajan: to use the ceiling function. Ceiling function is just, if I have 1.6, so it will convert my number to 2. If I have 1. Like, greater than 0.5, the decimal… if the decimal value is greater than 0.5,
[01:57:02] Asmita Mahajan: It will take the higher number, and if it is less than the 0.5, it will take the lower number.
[01:57:10] Asmita Mahajan: Now I'm creating the instances, like the figure and the subplots. This is… this will be used for the whole figure, and the axis will be used for the subplots. So we are calling the subplot functions from the matplotlib library. We are giving the row… how many rows I want for the subplot, and how many columns I want.
[01:57:28] Asmita Mahajan: what will be the figure size, and we are telling that I want to squeeze. Like, if there are more, columns or more rows, I don't want them, I want to squeeze my plot.
[01:57:41] Asmita Mahajan: So, for each index and, K value in the key, so key is a list, so I'm enumerating, or I'm iterating over the keys, and I'm extracting the index and the key value. So k will, get the values, this.
[01:57:55] Asmita Mahajan: So, K will store these values.
[01:57:58] Asmita Mahajan: these key values, and index will be 0, 1, 2, 3, okay?
[01:58:03] Asmita Mahajan: Now I'm calling my axis, this subplot function, and I'm giving the row and columns to the subplot, and extract. So.
[01:58:13] Asmita Mahajan: in k-means, what we did was we reshaped, my axis
[01:58:19] Asmita Mahajan: variable. But here, what, to iterate over it.
[01:58:25] Asmita Mahajan: I'm not, reshaping it, because it is a 2D array, we converted it into a 1D array, when we did k-means. But here, I'm not converting it into a 1D array, I'm just taking the values from the 2D array, passing it to the
[01:58:41] Asmita Mahajan: acts, so I'm iterating over 2D array, taking the location, extracting the location, and then I'm just plotting it at that location, the subplot at that location.
[01:58:51] Asmita Mahajan: So, from histories, okay, so, this is used to extract the training laws and the testing laws for the particular key, and hist will store the training laws and the testing laws for that particular key. So, history, this hist will store the dictionary.
[01:59:10] Asmita Mahajan: So, this dictionary For my first experiment will be, extracted here.
[01:59:17] Asmita Mahajan: is extracted here. So, using this now, I will be plotting it. So what I'm doing is I'm creating… I got the location of the subplot, so I'm calling the function plot. I'm passing the training laws, because I have a dictionary here. Hist is a dictionary, so I'm telling it
[01:59:36] Asmita Mahajan: that I want the training laws from that dictionary, because train laws is a key here, so that's why I'm passing train laws as a key to this his dictionary, and I'm, taking this training loss data.
[01:59:49] Asmita Mahajan: And plotting it. Again, in the same plot, I'm taking the test loss data, and I'm plotting it, and labeling it as test loss, loss.
[01:59:59] Asmita Mahajan: For this particular subplot, I want the experiment name to be the key. So here, as you can see, this SGD at the rate 0.001 has been displayed, so this is the title of the subplot, which we are setting it here.
[02:00:16] Asmita Mahajan: Using this line of code.
[02:00:18] Asmita Mahajan: Okay?
[02:00:19] Asmita Mahajan: Then, set X label. So, X label, I want my X label to be epoch, and my Y label to be as lost. So, how I'm doing is, I'm using, these, this
[02:00:32] Asmita Mahajan: this line of codes. So, set XLabel will, label my X axis, and Y label, setY label will label my y-axis. And I also want the legend, so that I will explicitly know that which line is my, which curve,
[02:00:48] Asmita Mahajan: is my training loss curve, and which curve is my test loss curve? So, Legend will help me with that.
[02:00:55] Asmita Mahajan: So,
[02:00:57] Asmita Mahajan: this, what we are doing is, as we have mentioned, as we are not explicitly mentioning our rows and columns, and we are calculating it. So, it could be possible that my,
[02:01:10] Asmita Mahajan: grid is of, the form 3 across 3, but we are only plotting 4 plots, okay? So, what will happen is, the…
[02:01:20] Asmita Mahajan: Plots, the subplots which are empty will also be displayed here, but they will have no data.
[02:01:25] Asmita Mahajan: Okay, so further, we will see that when my, my grid is, my grid has empty spaces, so to not display those empty spaces, we are just
[02:01:38] Asmita Mahajan: switching it off, using this line of code, which has no purpose here, because here, my grid was not… my grid was off 4 cross 4 only, and I have only, sorry, 2 cross 2, and I have only 4 subplots. So, here it will not matter, but in the further code, we will… I will show you where, there will be a empty subplot here. I don't want to display that empty subplot, that's why I'm switching it off.
[02:02:03] Asmita Mahajan: So, this line of code will switch off that empty subplot.
[02:02:08] Asmita Mahajan: And, now I'm just setting my full title, or the whole title of the figure, which is this, which I have passed in the subtitle variable here.
[02:02:20] Asmita Mahajan: So using that, I'm, mentioning my…
[02:02:24] Asmita Mahajan: Or labeling my figure with the title, and plot.show will show the plot.
[02:02:34] Asmita Mahajan: So here, as you can see.
[02:02:37] Asmita Mahajan: with the first experiment, with SGD as my training, as my optimizer, and 0.001 as my learning rate. So, you can see the
[02:02:47] Asmita Mahajan: train loss and the test loss curves, they are almost parallel, so that… so it tells us that my model hasn't been converged. So even with 200 iterations, my model is not performing well on the test.
[02:03:02] Durga Toshniwal: Just, I wanted to come in here, Smitha, sorry to interrupt.
[02:03:06] Durga Toshniwal: So the point which is, very significantly noticeable here is that the… you can see that the test loss is much higher than the training loss.
[02:03:17] Durga Toshniwal: Which means that, in this particular case, since the training loss is very less and test loss is high, so the model is overfitted, actually, and we don't really want an overfitted model.
[02:03:28] Durga Toshniwal: Similarly, if we now look at the next plot, which is the SGT with 0.01 as the learning rate, again, you can see that there's a lot of difference, and the test loss is much higher than the train loss.
[02:03:43] Durga Toshniwal: And as a result, obviously the model is overfitted, and so the test loss is still high, and the model is not generalizable. Now, when we try to use the ADAM optimizer.
[02:03:57] Durga Toshniwal: we can see that things are much more better. The learning rate,
[02:04:02] Durga Toshniwal: even if the learning rate is 0.001 or 0.01, the model is much less overfitted. In fact, it is not overfitted.
[02:04:15] Durga Toshniwal: And, we can see that the best performance is coming with, .01 and Atom. So, what we can see here is that, and, you can see that in this particular case.
[02:04:29] Durga Toshniwal: With a, with a…
[02:04:33] Durga Toshniwal: 0.01 learning rate, and Adam Optimizer. What is happening is that…
[02:04:38] Durga Toshniwal: Finally, the train loss is, the test loss is just almost close to the train loss. Therefore, we will go ahead and choose
[02:04:49] Durga Toshniwal: Add them along with 0.01 as the final, hyperparameters, for… for the model.
[02:04:57] Durga Toshniwal: So, this was just some analysis part, so how you decide which optimizer and what learning rate to use out of the combinations that are… that are used here. However, you can use many other combinations, provided you have the compute to support that.
[02:05:14] Durga Toshniwal: So, I think, Smita, you can carry on, and since we are almost, converging to the end of the time, so what I will suggest is that we can stop here for today, and the rest of the code will continue in the next turn, which will be next week.
[02:05:31] Durga Toshniwal: However, if there are any questions so far on whatever has been covered in the code or otherwise, you may please feel free to ask.
[02:05:47] Durga Toshniwal: So, any questions anyone, regarding whatever has been covered in the code, or analysis, or anything, please feel free to ask. Otherwise, also, you can go through the code, because we'll be continuing… it's a big code, so we cannot, it will be difficult to cover.
[02:06:06] Durga Toshniwal: In one term, so we will be covering the rest of it in the next term, and you can still go through the code, and if you have any problems.
[02:06:15] Durga Toshniwal: Then you can actually ask.
[02:06:18] Durga Toshniwal: Yes, Deepak, what, doubt do you have?
[02:06:22] Deepak Bobade: Yeah, ma'am. So, the last graph, right, where you said, we should go with .01, right?
[02:06:29] Deepak Bobade: the right one.
[02:06:31] Deepak Bobade: So, I can see that the, test loss, right, is, as we go ahead, on the epoch side, right, it is increasing. So, why we chose the left, sorry, why we chose the right one instead of the left one?
[02:06:48] Asmita Mahajan: Oh, yeah.
[02:06:48] Deepak Bobade: Adam.
[02:06:49] Asmita Mahajan: for Adam, because you can see, we are not seeing where my… so, as you can see here, my test laws and, train laws are…
[02:07:01] Asmita Mahajan: like, overlapping, the curves are overlapping. So I can say, like, at around iteration 50, my model converges. But for this, my model converges around… above 75 iterations after.
[02:07:16] Asmita Mahajan: So, the conversions also, so the speed also we have to see.
[02:07:20] Asmita Mahajan: Okay, so at 200 iteration, obviously, my training laws and test laws are diverging apart, but where it is converging… so here, it is converging around 50.
[02:07:32] Deepak Bobade: 75, yeah. Oh, 50, yeah, correct. Oh, so we can reduce the number of apps, if we choose that one. Okay, okay, that is what it is. Okay, got it, got it, ma'am. Thank you.
[02:07:44] Durga Toshniwal: Any other questions, anyone?
[02:07:57] Durga Toshniwal: Any questions?
[02:08:05] Durga Toshniwal: Okay, if there are no further questions, then probably we can break for today.
[02:08:11] Durga Toshniwal: I think we can thank Asmita for taking us through the hands-on,
[02:08:17] Durga Toshniwal: special. Thank you all, have a great day, thank you, bye-bye. Then we'll meet the next week.
[02:08:23] Durga Toshniwal: to complete the rest of the quote. Thank you all.
[02:08:25] Pawan Misra: We'll wait.
[02:08:26] Akshat Pandya: Thank you.
[02:08:27] Krishnakumar MS: Cool.