# 08 2025-12-21 Hands On Naive Bayes and Linear Regression

course: Module 2 — Machine Learning Algorithms
module: Module-2-Machine-Learning-Algorithms
date: 2025-12-21
type: transcript
video_url: https://personal-learn.armco.dev/files/_Recordings/Module-2-Machine-Learning-Algorithms/08_2025-12-21_Hands_On_Naive_Bayes_and_Linear_Regression.mp4

---
[00:14:06] Then we'll split the train, uh, the data set into train and test sets.
[00:14:13] And then use a linear regression model after doing one hot encoding.
[00:14:19] Um, on the, uh, data.
[00:14:22] And then some standard other things, like normalization and all that will be performed.
[00:14:29] And, uh, then, um…
[00:14:32] Uh, the target variable, the actual values of the target variable versus those which are predicted will also be seen… shown as just an illustration of how
[00:14:43] Uh, the prediction model is…
[00:14:46] performed, uh, is performing.
[00:14:49] And then the log transform categorical variables.
[00:14:53] Uh, that also will be shown.
[00:14:56] using integer encoding and one-hot encoding. These are the two different…
[00:14:59] ways in which encoding is performed.
[00:15:03] Uh, Part B of the code, uh, comprises of a section on knife-based classification.
[00:15:10] In which another dataset, which is the adult census income dataset, is included.
[00:15:15] It shows the income of people.
[00:15:17] And the other attributes, like education level,
[00:15:21] Marital status, occupation, and other fields.
[00:15:25] are shown here.
[00:15:27] Again, the standard steps would be loading of the data in the data frame.
[00:15:31] followed by looking at the information, like,
[00:15:34] What are the attribute types, whether they contain
[00:15:38] Uh, not… I mean, null values.
[00:15:41] And then looking at missing values, if there are any, and there are certain…
[00:15:47] Um, attributes, uh, that have missing values.
[00:15:52] And then treating the data by some method. In this case, dropping of the rows that are having missing values.
[00:15:59] And, uh, then, uh…
[00:16:02] Uh, so this was some preliminary data preprocessing.
[00:16:06] That was performed, that is performed on the data. Then, looking at the unique values,
[00:16:12] per categorical attribute.
[00:16:15] And then, uh…
[00:16:17] doing some EDA, like, plotting of the distribution of values in the target variable.
[00:16:23] Then, uh, this will be followed by using the night-based classifier.
[00:16:29] And then, in this case, we are talking about the multinomial, um…
[00:16:33] classifier, which works on categorical values.
[00:16:36] And then, um…
[00:16:40] the usual steps would be…
[00:16:42] splitting in the training and test sets, and then…
[00:16:48] There's another form, which is a Gaussian knife-based classifier.
[00:16:52] Uh, that also has been used, which is primarily, it works on numeric data.
[00:16:58] Uh, and the numeric data is encoded.
[00:17:01] Using integer encoding.
[00:17:04] And uh… then, finally, uh…
[00:17:08] We look at the performance of a multinomial versus Gaussian knife base.
[00:17:13] And, uh, see which one performs better in each of the encoding methods, that is.
[00:17:19] Uh, one hot encoding versus the…
[00:17:22] integer encoding.
[00:17:24] And uh… and I think pretty much that's it for today.
[00:17:29] So, this is the overall flow of what we'll be following.
[00:17:34] Uh, so I think today, what is not included here is K-Fold cross-validation.
[00:17:39] Which I'll probably take up in the next turn.
[00:17:41] And also, uh, use of R-squared.
[00:17:45] Uh, that is not there.
[00:17:47] But we'll cover it, don't worry.
[00:17:49] So I think, uh…
[00:17:51] Over to you, Ashish, to take the code in more details. I just took over…
[00:18:04] Uh, good morning, everyone. So, let's start with the part A of the notebook, which is linear regression.
[00:18:12] And for linear regression, we are using a bike-sharing dataset.
[00:18:15] Whose original link is provided here, the original source.
[00:18:19] And, uh, the goal is here to predict the number of bike rentals based on various conditions like time, weather, and there are various user behavior variables here.
[00:18:30] So, it's a… as you can see, it's a data collected over a 2-year period.
[00:18:33] And, uh, yeah, so…
[00:18:37] This is our… in this dataset, CNT by the target variable is by the name CNT.
[00:18:41] Which is the total number of bikes which are rented in an hour.
[00:18:46] So, uh, so basically, uh, there are two, uh, separate features also, causal plus registered, whose values sum to the variable, uh, count.
[00:18:56] So, in our analysis,
[00:18:58] Causal and registered columns, uh, we will not be using, because these two columns together sum to count variable.
[00:19:05] And, uh, yeah. So, uh, let's look at the features.
[00:19:10] So, there are various features like instant, which is a record index, which is just an identifier, so it can also be dropped, it is a kind of unique identifier. And then there is season, which is categorical, the year, it's a binary variable, because two years are… data is of 2 years.
[00:19:26] month, hour, uh, holiday, uh, weekday, and so on.
[00:19:32] And, uh, like, as I said, that, uh, we will not be using the features causal and registered, because when predicting the count variable, because the sum of the values of both of these columns is actually equal to the count.
[00:19:46] So, uh, we will not be including them, uh, otherwise model will get…
[00:19:51] model will get extra knowledge before, uh, instead of doing prediction.
[00:19:57] So, uh, there are two data files shared today. So, the first file is by the name BikeSharing.csv. So, if you load it in the collab, and then run this command, it will, uh, then run this block.
[00:20:11] a period of read CSV, the file name, it will show something like this.
[00:20:17] So, there are 17,000…
[00:20:19] rose in this CSV file, and there are 16 columns.
[00:20:25] 16 columns, and uh…
[00:20:28] Yeah, so this is a data frame which is showing
[00:20:32] Top 5, uh, values head, and last 5 values, the tail of the data frame.
[00:20:38] And this is how it…
[00:20:40] will look, uh, when we load it into Panda's data frame.
[00:20:45] Next, like, as we have seen, first, we usually do our df.info, uh, to see the data types of the
[00:20:53] columns which are there in this, uh…
[00:20:56] Uh, in this, uh, dataset. Uh, how pandas perceive, uh, how it, uh, how, what it assigns data type to each of the columns.
[00:21:04] which are there in this CSV file. So, you can see that, uh, as such, there are no…
[00:21:12] there seems to be, again, there seems to be no categorical variable on the face of it. It's because it either treats like an integer or in float.
[00:21:24] So, basically, if unless, uh, if unless there is a string in the columns,
[00:21:29] Even for categorical variable, if categorical variables are categories in numbers, it assigns them integer.
[00:21:36] As its data type. So, yeah, so although the information in the dataset was given, that there are categorical and numeric variables, but initially,
[00:21:45] Pandas is treating everything as integer or a float. So, we will have to make this distinction before modeling that whatever are the categorical variables, we will have to assign them as, like, object columns, or treat them as categorical.
[00:22:00] So, that's why we always, like, do a div.info to see how natively the Pandas is treating each of these
[00:22:07] columns.
[00:22:10] So, then, uh, we can see from here also that, uh,
[00:22:13] There are no missing values. So, this step becomes redundant. If… but if there had been any missing value,
[00:22:22] If we did a df.isNull.sum, for each of the columns,
[00:22:28] It would tell me the number of missing values in each of the columns.
[00:22:31] So, this step is basically for completeness. In the second dataset, there are missing values.
[00:22:37] Where, uh, we will see that when we do, uh, this line, it will tell me the number of rows per each column where there are missing values.
[00:22:47] Okay, so, uh, first things, uh, as we, uh, as I told, that causal plus registered,
[00:22:54] So, even in this, like, let's see in this data frame, what I was saying, so…
[00:22:59] This is the column causal, this is the column registered.
[00:23:02] And we can see that some of causal and registered is equal to the count variable.
[00:23:07] Everywhere. So, that's why.
[00:23:10] We will be dropping this causal and registered for further, uh, for further model building.
[00:23:17] These are not required. And similarly, instant.
[00:23:20] Instant is also like a unique identifier at which instant of the time the…
[00:23:26] A rental was taken, and you can see that it is, like, similar to the index.
[00:23:32] What part does assign, everything is just same, just at the pandas assign index starting from 0, and it is starting from instant 1.
[00:23:41] So, yeah, so this is also not, like, usually not required for training linear regression, so…
[00:23:47] Here, uh, in this line, in this block of code, what I'm first doing is, I'm dropping some columns, which will not be required. So, that's why I've written, uh, df.drop columns, I've mentioned in a list.
[00:24:01] There are 3 columns, instant,
[00:24:03] Causal and registered.
[00:24:06] So, uh, after that,
[00:24:08] Uh, because from prior information of the dataset, I know that, uh,
[00:24:12] These are the specific columns, uh, which are having category… which are categorical. So, what I do is, I change their data type to object.
[00:24:21] So, from the data card information, I knew which one of them are categorical variables.
[00:24:27] So, I… so I picked them up,
[00:24:30] And, uh, like, hard-coded, uh, like, explicitly, uh,
[00:24:34] explicitly change their data type with object.
[00:24:39] So, one way is, like, we can do it like this, or otherwise, what we…
[00:24:44] could have. What we can also do is, for each of the variables,
[00:24:48] We can loop through the unique values of each of the variables.
[00:24:52] So, for categorical columns, uh, here we, uh, here I have explicitly first defined categorical columns, and then I'm showing their unique values, but it can be done for all the columns in the data frame, in a data set, and we can see that if unique values are fewer, that is an indication that these are categorical variables.
[00:25:13] So, for the categorical variables, because we, in this dataset, we already know which are categorical variables, so that's why we, like, explicitly defined here that these will be our categorical variables.
[00:25:25] And then, after…
[00:25:28] We use them and change their data type to object.
[00:25:32] And just for, like, looking, like, how many unique values are there, uh,
[00:25:38] Uh, there, we then looped over these categorical columns, and for each of the, uh,
[00:25:43] like, categorical columns is a list,
[00:25:45] For each of the columns, uh, we store there, uh, we store the unique values by doing div call.unique.
[00:25:52] After that, uh, like, this is printed, that for, uh, that there are, for this value, for season, there are 4 unique, for year, there are 2. And, uh, like, this, uh, for R, there are 24 unique, like this. So there are, uh, these are categorical columns, but, uh, you can see that, uh,
[00:26:07] Unique values are very limited, uh, like…
[00:26:10] There are… the highest is for R, which is 24.
[00:26:13] Otherwise, uh…
[00:26:15] almost all of them are less than 10, and, like, month is, uh…
[00:26:19] 12 categories. Yeah. So, after that,
[00:26:23] We always, like, even if it's… the target variable is a continuous value called count,
[00:26:32] So, in regression, it's a continuous value. Even for regression, even for all for classification, we always plot the distribution of the target variable.
[00:26:40] So here, uh, since it's a continuous variable, what we plot is a histogram of the target variable. So,
[00:26:48] Target variable is the variable count.
[00:26:52] So, in Pandas, you can directly plot histogram. So, DF underscore clean, and then we're doing plot over it, the plot function, and tell it which kind of histogram I want.
[00:27:02] the kind is here, the histogram. So, we can plot the bar graph also, uh, with… by specifying here in count.
[00:27:09] So, if it would have been a categorical variable. But right now, it's a continuous variable, so we do hist, uh, equal to hist.
[00:27:18] And then, uh, histogram-specific information that how many bins are to be there in each, uh, there, and everything. So…
[00:27:25] And title, etc. has also been specified. So, uh, just by looking at this distribution, we can see that there's right skewness in the data set, there is a long right tail.
[00:27:39] Uh, which is extending. So, uh, much of the data is concentrated in the range
[00:27:45] Uh, like, the counts are higher in the range 0 to 200, but again,
[00:27:50] there are fewer values, but there is number of values in the right, but the tail extends as long as, like, thousand.
[00:28:00] So, the observation is, like, the data distribution is skewed, and there are many low rental counts.
[00:28:08] Majority of the counts are very lower-end. The counts are very less for the rental majority of them.
[00:28:15] So, this is the first thing.
[00:28:17] And now, uh, like, if we want to, uh, like, do a…
[00:28:23] We want to build a regression model over this.
[00:28:26] So, the first step we do is, like,
[00:28:30] Uh, we will split the dataset into train and test.
[00:28:34] So, uh, in order to split the dataset, so… let me… see, so this is the dataset we are… data frame we are working on, DF underscore clean.
[00:28:43] So, uh… so there are, right now, 13 columns in it.
[00:28:47] And this is our target variable, CNT.
[00:28:50] So, we will remove this target variable. All variables excluding this target variable will be our input variables.
[00:28:57] So, for doing that,
[00:29:00] First line is this.
[00:29:02] I'm, uh, dropping the count variable.
[00:29:07] Uh, from that, uh, from this, uh, DF clean, uh, variable.
[00:29:12] I'm dropping the count column.
[00:29:13] And after… and whatever is left, I will assign it as my X. So, this is my X, which will be my, like, the input features, which are required to predict the output feature, which is
[00:29:24] which is the number of columns here.
[00:29:26] And individually alone, the target variable, uh, the, uh, is the, like, the column count, we will treat it as Y.
[00:29:34] So, now, once we have an X and Y,
[00:29:37] Then, we can, uh, like, use it for, uh, when we can split over it and, uh, make trade and test split.
[00:29:44] And we can also, like,
[00:29:47] use them for, uh, like, uh, like, we can use them later for model building, because, uh, like, these kind of algorithms, they expect, uh, you to provide them with both X and Y.
[00:30:00] So, you can see that, as usual, I am using, uh, sklun.modelselection, the function train test it from there.
[00:30:09] So, this is trend test split.
[00:30:12] And here, I gave it my total, uh, X, so my total X is like…
[00:30:17] 17,379 rows.
[00:30:20] Uh, like, all the data I've given it, and this is my, uh, Y, which is, like,
[00:30:27] the target values for all these input data. So, here, like, in last class, we saw that we can, like, perform, uh, train test split.
[00:30:38] Like, even this, uh, the value of the trend split can be a hyperparameter, but here, uh,
[00:30:43] like, we are, like, fixing it
[00:30:46] 80-20 as a trend test split.
[00:30:49] Uh, just a minute here, because it…
[00:30:52] Then earlier, so it was loading from…
[00:31:03] Yeah. So, uh…
[00:31:06] So yeah. So, uh, finally, as you can see, I initially, I showed the DF underscore clean. So, there were 13 columns in the original data frame after the minimal
[00:31:18] cleaning, and uh… or we did. So, uh, the input features are 12.
[00:31:23] So, after doing this trend test plating and what we saw is that our… there are input features 12, and these are the number of samples on which we will be doing, training our model.
[00:31:35] And 3000, uh, the 20% of data, uh, so 20% of data is, like, uh, R.
[00:31:42] test data. 3,000 samples are in our test data.
[00:31:46] And we can also see that, uh, so that Y… Y is a single column, because it's a single value, which has to be predicted, and train…
[00:31:54] It's a 2D type of matrix.
[00:31:57] There are 12 columns per sep… there are 12 entries, there are 12 variables per… and, uh…
[00:32:02] row, we can say. So yeah, so this is the trend test split.
[00:32:07] Uh, which we'll be using.
[00:32:09] So now, uh, as we can see, that some variables are numeric.
[00:32:15] Some are categorical. So, uh… so categorical variables have to be, like, treated, uh, like, they have to be treated in a manner, uh,
[00:32:26] like, how do we provide that category of variables? So in… there are two ways. How do we deal with categorical variables?
[00:32:33] One is integer encoding.
[00:32:35] Uh, the other is one-hot encoding.
[00:32:38] So, uh, before giving to model, we will either integer encode the, uh,
[00:32:44] categorical variables, or do their one-hot encoding. Uh, for the remaining numeric features,
[00:32:49] Uh, just for, like, stable training, what we will do is, we will, uh, normalize each feature. So, uh, that we… it will have, uh, the features will have zero mean and standard deviation ones. So, uh, like,
[00:33:02] Uh, if our variable is, suppose, it's a continuous variable, uh, X,
[00:33:06] We will divide it by its, uh… by its mean.
[00:33:11] And, uh…
[00:33:13] We will divide by its mean value, and uh…
[00:33:16] And its standard deviation, so that, uh, overall,
[00:33:20] for that feature, among all the values, the mean then becomes 0, and it has standard deviation 1.
[00:33:27] This we'll do for continuous, uh, features, uh, but for categorical variables,
[00:33:33] Either we… we will either handle them by integer encoding or one-hot encoding.
[00:33:38] So, uh, if we see, uh, in this dataset,
[00:33:41] Already, the continuous variables, uh, they are, like, they are already numbers.
[00:33:42] in the flow of what is the overall flow for today.
[00:33:48] So, uh, so that means they are kind of, like, integer-encoded already, uh, but if they were in, like, in form of string and all, so integer encoding, uh, would, uh, integer encoding would have been necessary, but, uh, here, uh, like, since they are already numbers, that means they are already integer encoded. So, uh…
[00:34:08] So, these things to be noted.
[00:34:09] So, just one second, sorry to interrupt.
[00:34:13] So, uh, I want to add here that encoding here is of two kinds. One is one for encoding.
[00:34:19] So, in one hot encoding, what we do is that…
[00:34:23] Whatever, uh…
[00:34:25] Uh, let's say, uh, you know, whatever cat, uh, attribute is in forecast, or whatever value we are talking about, we assign it 1, and for all others, we assign 0.
[00:34:36] So, uh, the encoding, one-hot encoding is form of zeros and ones.
[00:34:42] And if we talk about integer encoding, in case of integer encoding, what we do is that each categorical value
[00:34:51] These are assigned an integer number.
[00:34:54] For example, if we had values like small, medium, and large,
[00:34:59] So, integer encoding would assign values like 1, 2, 3, or 0, 1, 2,
[00:35:04] Something like that, to all those categories. So, as mentioned by Ashish, that in some of the attributes which are categorical,
[00:35:13] already integer encoding is available, because the values are available in form of integers. So, if you could go up, Ashish, to…
[00:35:21] illustrate that.
[00:35:23] So, you can see here, for example, when we talk about R, or…
[00:35:28] uh… month and all, you can see that it's already available in the form of 0 to 23.
[00:35:35] So, 24 integers are assigned from 0 up to 23.
[00:35:40] And for a month also, 1 to 12, 12 integers are assigned.
[00:35:44] So, these are already integer encoded.
[00:35:47] They look like numbers, but these are nothing but integers. There's nothing like…
[00:35:52] 1.5 or something like that.
[00:35:54] So, this is already available. The other form is one hot encoding.
[00:35:58] In one not encoding, we have something like 0001, something like that. So we just have ones and zeros, we don't have any numbers out there.
[00:36:07] So just I wanted to add it.
[00:36:10] No, before we proceeded, so…
[00:36:13] I think Ashish, you can take on from here.
[00:36:15] And we have one small example with the code, if that will be there, ma'am, uh, how that can be converted to…
[00:36:21] One more encoding order.
[00:36:24] data reporting.
[00:36:27] Okay. Uh, you want to see what one hot encoding looks like, right?
[00:36:32] And, uh,
[00:36:33] Yes, yes, ma'am.
[00:36:35] Yeah, so I think right now we are doing code only.
[00:36:38] Okay.
[00:36:39] So, what I'll do is, I'll take a note of it.
[00:36:42] And I'll show you example in my next theory class next week.
[00:36:46] Well, we will do one hot encoding. I'll show you some example.
[00:36:51] and integer encoding also, I'll show you example, okay?
[00:36:54] Okay, then.
[00:36:55] So right now, I've broadly told you what it means, but I'll also include examples.
[00:37:03] Yeah. So, Ashish, you can carry on from here.
[00:37:06] Okay, so now, uh…
[00:37:09] So, as we said that, since there are categorical variables, so we have to take a decision whether
[00:37:14] to encode them in, like, integer encode them, or why not encode them. Uh, and for numeric features, uh, like, we… they are all… we use them as it is.
[00:37:23] Uh, so this is all done because we want them all to be in numeric form.
[00:37:28] to give to the model. So, uh, let's start with, like, how we will be modeling this using linear regression.
[00:37:36] So, uh, first, like, as I, uh, like, as we have done is, like,
[00:37:40] We have, uh… like, if we see our X…
[00:37:45] If we see CRX here…
[00:37:50] So X has, uh, uh, like, these are the input variables, so, uh, object, all those which are object,
[00:37:57] They are categorical, and all those that are float, they are numeric.
[00:38:01] So, these are the data types. So, first line, uh, what is written here is that x.selectDataTypes
[00:38:08] and include npit.number, that means it is basically indicating to include, uh, that data types where there is a number. So, right now,
[00:38:20] There is only a distinction that either they are categorical or object, and remaining will be numeric, which can be integer or float. Like, we have made it in this form, that
[00:38:33] numeric columns will be indesigned and flowed, and all the categorical variables, they will either… even if they are originally stored as numbers,
[00:38:42] We will treat them as object data type. So, that's why…
[00:38:45] So, all the data types which have a number will be selected, uh, and will be stored in numeric columns, which is like a kind of list. So, you can see that if I do this separately…
[00:39:01] Okay, so…
[00:39:07] So, yeah. So, uh, basically, the last four values,
[00:39:12] They are float, so that they were included here, and the numeric columns.
[00:39:17] When we, like, run this, uh, line of code. Similarly,
[00:39:22] When I say that select data types and include the data type as object,
[00:39:28] So, what is, uh, it will do is, uh, it will pick all these data types which have object,
[00:39:34] And we'll, uh, like, store, uh, in this, uh, in this column, uh, like, categorical columns. So, if I run this…
[00:39:48] So, basically, in the first… in these lines, I'm just, like, in a variable, I'm storing what are my categorical variables.
[00:39:56] And what are my numeric variables?
[00:39:59] Now, uh, this is being done because I will treat them separately.
[00:40:03] for variables which are categorical, I will either do one hot, or I will either do integer or one, uh, integer encoding or one not encoding.
[00:40:12] For numeric variables, I will, like, normalize them so that for each variable, for each column, the overall mean will be 0, and standard deviation will be 1.
[00:40:24] Uh, for that, I am, like, first…
[00:40:25] like, taking, filtering out, uh, based on whether the columns or variables or features, they are numeric or categorical.
[00:40:34] So, once I get this list,
[00:40:36] I will, uh, split my train data set and my test dataset, uh, like, for numeric… so…
[00:40:43] this line will give me a subset of the columns which are only numeric.
[00:40:48] When I, uh, like, index using this list, uh, numeric columns is nothing but a list containing all the names or numeric columns, so this will give me only all my, like, numeric columns, so…
[00:40:59] If we see here…
[00:41:07] Yeah. So, the same 13,000 entries, but only for the
[00:41:13] four numeric columns.
[00:41:15] Similarly, and this is done for the test set also.
[00:41:21] And also, uh, the same thing is repeated for categorical columns also. So, all the categorical columns, like, if we run this…
[00:41:34] So, let's directly run this diagonal.
[00:41:41] Yeah, so the same 13,000 entries, but with the 8 columns, which correspond to categories. So, these are basically categories.
[00:41:50] Okay, so, uh, first, uh, what we are, uh, so first we are, uh, like, normalizing these numeric features. So, what we are using it, either we could hardcore, uh, like,
[00:42:00] We could write 2-3… few lines of code ourselves that, uh, calculate the mean… loop over each cat… uh, loop over each numeric feature.
[00:42:07] calculate the mean of all the values, calculate the mean for each column separately,
[00:42:13] Calculated standard deviation, and then divide each value, then subtract each value by its mean and divide by standard deviation. So, uh…
[00:42:20] instead of doing this, uh, in… because sklearn, there are, uh, like, all these functions are already available. So, what we are using is, uh,
[00:42:29] There is a pre-processing module in SkLearn.
[00:42:31] From there, we are using Standard Scaler, and we'll use ordinal encoder also, but right now, we are using standard scalar.
[00:42:38] So, uh…
[00:42:40] Uh, just a minute.
[00:42:43] Okay. Uh…
[00:42:48] Yeah, so, uh…
[00:42:53] Okay, so we'll be using Standard Scaler. So, first I call the standard scalar, uh, first I initialized it, uh, the… an object of the scalar, a standard scalar, and we are calling it variable scalar.
[00:43:05] And then, uh, what I do is…
[00:43:08] Uh, I, uh, from a trained dataset, so the numeric columns were in this, for the trainer set, numeric columns were in this variable. I do a fit and transform. So, basically, fitting is that for each variable, it will find whatever… for each variable, it will loop, it will find its mean, it will find its standard deviation.
[00:43:25] Uh, this is called, like, fitting, this thing. And then,
[00:43:29] based on whatever is the mean and standard deviation for each column, uh, you are… you will transform your initial value by subtracting by mean and dividing by standard division. That is called transform.
[00:43:39] So, fit underscore transform.
[00:43:42] will perform both of these steps together on the train dataset.
[00:43:47] So, uh, and after doing this,
[00:43:50] whatever we get is our, like, transformed, uh, uh…
[00:43:56] transformed, uh, numeric columns.
[00:44:01] by doing, like, normalization. So we can see if, uh…
[00:44:10] Yeah, so, uh, because, like, the column and information has been, like, removed because it does not return the column, columns thing. So, if we see here, uh, X, the originally,
[00:44:24] Originally, they were like, uh…
[00:44:28] Or… originally, they were, like, these values. So, basically, this could happen that all of them are in separate scales. So, what we do is,
[00:44:37] For each of these columns, when we normalize them, each… now, for each of these columns, we'll have a zero mean, and unit variance of 1. So that we do this for all the numeric columns. So, to bring them in a similar scale.
[00:44:50] So, this is done, and now, for, uh, so what I showed you was for the test train dataset.
[00:44:57] Now, for the test dataset, uh, we will not be doing, uh, fit, we will just transform.
[00:45:02] Because the learning of the mean and variance for each of the columns was done from the train part of the dataset.
[00:45:10] For the test dataset, uh, because test is, like, how will we be seeing in production, when a new… after training one model, whenever a new sample comes. So, we will not be, like, learning a mean for adding it in the dataset.
[00:45:25] whatever was our, uh, mean and variance for the training data.
[00:45:29] We will use that, and then just transform.
[00:45:32] Here, uh, for the test dataset.
[00:45:35] So, the same thing, but just applied the transform thing, and we will get the normalized values for the numeric columns.
[00:45:41] Now, uh, for the audi- now for the categorical features, as we said, there are, uh, two ways we can handle them.
[00:45:48] So, uh, uh, yeah, so one is that you either, uh, either do an ordinal encoder. So, uh, this has been added.
[00:45:57] to, like, illustrate, uh, how it is done, here, it is kind of redundant because the numbers are already in integers, so that means they are already orderly encoded, or integer encoded. So, basically, it is integer encoded, but, like, in SKLearn.
[00:46:13] They call it ordinal encoder.
[00:46:15] So, uh, it is already done, so this step becomes redundant, but in the next step, in the next dataset, we'll see that when categories are in the, like, in the form of string, so, uh, this is required to convert them into integers.
[00:46:29] So here, uh, again, uh, like, uh, ordinal encoder we call.
[00:46:34] And use it as a, like, use it using a variable, encoder variable is by the name, uh, we are using it. And then again,
[00:46:43] Uh, fit underscore transform, uh, we'll do on the train dataset.
[00:46:46] And for the test dataset, we'll already… we will only be using
[00:46:51] transform. So, basically, the thing is that…
[00:46:54] If you, uh… so suppose in the trained dataset, uh, you saw, uh,
[00:46:59] in the tens dataset, suppose that there are, like, 6, and if you see only 5 of them,
[00:47:05] And now in that test dataset, uh, the sixth category will come, so it will throw an error. So, that's why, like, we are using only for the train dataset.
[00:47:18] Uh, because we are expecting that, uh, whatever the categories are there in the, uh, for each column that are in the train dataset, that only accounts.
[00:47:23] So, this is, uh, as I said, uh, fit underscore transform.
[00:47:28] So, I get this variable after doing ordinary encoding of my categorical features.
[00:47:37] And this, after doing ordinal encoding of my categorical feature for the test dataset.
[00:47:42] Now, uh, in order to give to the model,
[00:47:45] So now, these are in, like, two separate forms now. So, this is my train dataset.
[00:47:51] But only for the numeric features. This X train num scale.
[00:47:56] This is my train dataset, but only the part of categorical variables. So, these two have to be, like, now, again, concatenated together, because they together represent the data.
[00:48:08] This represents only the numeric features of the data.
[00:48:12] Input data, and this represents only the, uh, like, the categorical input of the data.
[00:48:18] So, now when we come… now what will be… now what we are doing is…
[00:48:21] Uh, we are doing, uh, this function is called np.edhtech. It is in the NumPy library, horizontal stacking.
[00:48:29] We are, uh, like, uh, this basically means that, uh, we have two, uh, we basically support, uh, okay, so let's see what, by code, it will be more clear.
[00:48:42] So, what it is being done here.
[00:48:52] So,
[00:48:55] Okay. So, basically, uh, if I… if we see…
[00:49:18] So, we can see that, uh, since we do… did a separate normalization of the numerical features, there were four of them.
[00:49:26] We did separate encoding of the categorical features, which were 8 of them. Then, in order to give to the, like, the linear regression model, we first concatenate them again together, because all of them together represent data. So that's why, after I did a horizontal stacking here,
[00:49:41] It… the number of rows are same, just the columns are added. So, 4 plus 8, now there are 12 columns.
[00:49:48] So, this is done for the…
[00:49:50] trained dataset also, and this is done for the test dataset also.
[00:49:55] Uh, just I wanted to come in, Ashish.
[00:49:59] So, I think it will be helpful if I'll just, uh…
[00:50:02] I'll quickly give an example as requested of integer encoding and, uh,
[00:50:07] One hot encoding, so addition of…
[00:50:10] more number of columns and all could be more clear.
[00:50:12] So, if you could just do a stop share for just 2 minutes, I'll share my screen.
[00:50:18] And, uh, then, uh, we can… I'll just give an example.
[00:50:25] Yes, screen is stopped.
[00:50:26] I just, I don't know. Yeah, yeah.
[00:50:27] I just saw it.
[00:50:31] Yeah, so I hope that my screen is visible.
[00:50:35] To all of you, I'm just going to give you an example.
[00:50:41] Off, of…
[00:50:44] of what is integer encoding and what would… could be one-hot encoding. Just allow me a second.
[00:50:53] So, let's say that, uh…
[00:50:55] We have an, uh… we have an attribute.
[00:51:00] And that attributes, uh, is a shirt color.
[00:51:07] And the shirt color could be either red,
[00:51:11] Or it could be blue.
[00:51:13] Or it could be green, as per our data.
[00:51:16] And, uh, let's say that, uh…
[00:51:19] Uh, we are having some 5 persons in the data, so this is our data.
[00:51:25] And we are having some attributes, and we are having 5 persons, P1, P2, P3.
[00:51:31] P4 and P5 are 5 persons.
[00:51:34] And one of the attributes is this shirt color.
[00:51:38] Which is the color of shirt of each person, P1, P2, then maybe other attributes also, which I'm not mentioning here.
[00:51:45] And for person P1,
[00:51:49] Uh, the color may be something, say red.
[00:51:53] Followed by another red.
[00:51:55] Followed by green.
[00:51:57] followed WebDo, then green.
[00:52:00] Let's say this is the attribute, and these are the values.
[00:52:03] Now, what do we do in integer encoding?
[00:52:06] If we want to do integer encoding,
[00:52:10] Then, I should be able to assign some integer for each of these values.
[00:52:17] What are the values that the attribute shirt color could take? These are red, blue, and green.
[00:52:21] So, I'll assign the values…
[00:52:24] 0, 1, and 2.
[00:52:26] in integer encoding to these colors.
[00:52:29] So now, this dataset for person 1, 2, 3, 4, and 5,
[00:52:33] And the attribute shirt color would…
[00:52:35] After doing an integer encoding,
[00:52:38] would have something like this. So, here are the person's P1, P2, P3.
[00:52:44] before… I'm just rewriting it for sake of clarity.
[00:52:48] And this is the shirt color after…
[00:52:50] doing, uh, one-hot encoding.
[00:52:55] So then, what…
[00:52:58] Uh, green… red would be replaced by its integer value. Then again, we have red, which is 0.
[00:53:03] Then we have green, which is 2, then we have blue, which is 1, followed by green again.
[00:53:09] So now, this column is mapped to integer-encoded column like this.
[00:53:13] Now, suppose we are doing one hot encoding.
[00:53:16] So, important hot encoding, actually,
[00:53:19] We need to add columns.
[00:53:22] one column each for the value of that particular attribute.
[00:53:27] So now, if we still talk about the same data, P1, P2, P3, P4, and P5,
[00:53:34] So, and we are doing, uh, one hot encoding.
[00:53:38] then how many columns will I have?
[00:53:40] I will have as many columns as the values.
[00:53:44] So, there are 3 columns.
[00:53:46] And 3 values of the color of the shirt, that is possible.
[00:53:49] red, blue, and green. So, instead of one attribute, which is short color,
[00:53:55] Uh, this shirt color is actually replaced by these three columns.
[00:54:00] Person 1 is having red shirt, so red will be 1.
[00:54:03] Blue and green will be zero.
[00:54:05] Person 2 is again having red shirt, so this will be 1, these two will be 0. Person 3 is wearing a green shirt, so green will be 1, others will be 0.
[00:54:14] Person 4 is wearing a blue one, so it will be 010, like that.
[00:54:19] Person 5, uh, maybe is wearing green, so it will be 001.
[00:54:25] So, one hot encoding, actually.
[00:54:27] Uh, here, increases the total number of columns to as many
[00:54:32] Columns, as there are values. In this case, there are 3 values in the shortcut.
[00:54:37] So, instead of one attribute shirt color, we got three of them, red, blue, and green.
[00:54:42] Okay.
[00:54:45] But in integer encoding, what we do is we replace any of these values via the corresponding integer value.
[00:54:52] So, this is about integer encoding and one-hot encoding. Any questions on this?
[00:54:58] Anyone?
[00:55:00] integer.
[00:55:01] Well, my mother, like, is there any best practice or something, like, how can we choose between these two, like, which one to be used under which scenario?
[00:55:09] If both are equally popular. Say, for example, uh, let's say if the number of values
[00:55:16] Did I do a stop share, or what did I do? I don't…
[00:55:19] Let's say if the number of values
[00:55:23] In the attribute are very high.
[00:55:24] Say, for example, let's say if I've got 24 hours.
[00:55:29] Right, so the… so if I have to do one-hot encoding, so in place of the hour,
[00:55:35] I have to replace the hour attribute with 24-hour attributes.
[00:55:40] So that will increase the dimensionality of the data to a large extent.
[00:55:45] So, if the number of values in the attribute are very high,
[00:55:50] Then we go for integer encoding.
[00:55:54] If the number of values in the
[00:55:55] attribute is not that very high, we can go for integer, or…
[00:56:00] One hot encoding, any of these.
[00:56:02] Uh, similarly, if we are having 12 months in a year.
[00:56:06] If I do integer encoding, then I'll have just one column, and we'll have 12 values in that one column.
[00:56:14] One for each month, January one, then like that, up to December 12.
[00:56:17] But if we have to do one-hot encoding, then I have to replace the month attribute with 12 other attributes, starting from 1, 2, 3 till 12.
[00:56:26] So that will make the data much more sparse and highly dimensional.
[00:56:30] So, accordingly, you decide if number of values in the attribute are too high, don't go for…
[00:56:36] One hot encoding, if they are limited, you can go for any of these. Is it okay?
[00:56:43] So, in that case, can I assume that I can always go for integer encoding, and not… I mean, like… Is there a reason to prefer one-hot encoding in that case? Why can't we default integer encoding by default?
[00:56:54] So, there's no specific reason, so to say.
[00:56:57] Uh, to prefer one hot encoding, uh, over integer encoding.
[00:57:02] But many a times, one-hot encoding, uh, you might…
[00:57:06] You know, you… let's say you want to do classification.
[00:57:10] And, uh, sometimes, uh, doing integer encoding
[00:57:14] might actually, uh…
[00:57:17] not give very good performance.
[00:57:20] For example, let's say if I'm having the numbers 1 to 12 for month.
[00:57:25] And I'm doing classification in which Euclidean distance is used to find out the similarity between the month values.
[00:57:32] Now, if the… if the computer system thinks these are numbers,
[00:57:37] And finds out the distance between them as numbers, so that won't actually come out correct, because we are having 1 to 12, 12 numbers.
[00:57:46] And if we do a Euclidean distance,
[00:57:48] Then, the value of the month 2 and month 1 will also match.
[00:57:54] One-in-one, like, let's say there are two…
[00:57:56] two particular records in which…
[00:57:59] One particular record has the value for the month 1, which is January.
[00:58:04] And the other one is having the value 2, which is February.
[00:58:07] So, January and February are two different months, which should come out different.
[00:58:12] But if we use some simple distance measures, like you'd… Euclidean,
[00:58:17] Then the values 1 and 2 will also come out to be similar only.
[00:58:22] Or in other words, if we use integer encoding and use Euclidean distance,
[00:58:27] And then, uh, the months February and January will come out to be similar.
[00:58:33] But in applications where we are very, very, uh, you know, particular that we don't want this kind of
[00:58:39] problems. Then we use one-hot encoding, and accordingly, we use distance measures that work on Boolean data.
[00:58:47] Because 1 and 0 is nothing but Boolean data, something like simple matching coefficient,
[00:58:52] or Jacquard and all that, which will give you better results in terms of similarity.
[00:58:57] Okay.
[00:58:59] Got it.
[00:59:00] Okay. So, any other questions? Anyone on this before we proceed?
[00:59:09] One follow-up. So, uh, now, if you have, like, multiple categorical variables, like, do we have to… Choose the encoding depending on the specific categorical feature, whether we go with one-hode encoding or integer encoding.
[00:59:22] Yeah, so as I already mentioned, if you have multiple categorical values,
[00:59:27] And you have to choose the encoding, then…
[00:59:30] Uh, you have a look at what are the number of values. If generally the number of values are not very large, you can go for any of those integer or one-hot.
[00:59:39] And preferably integer, but if the number of values
[00:59:44] Uh, sorry, are not very large, then you can go for any of those.
[00:59:47] One hot encoding would be a good solution, but if the number of values in the attribute field are very large, then don't go for one-hot encoding.
[00:59:56] That's the way we usually do.
[01:00:01] But, like, for example, if there are, like, multiple categorical variables, we can take some… some variables can be treated with 1Hz, and some can be done with integers, depending on the requirement.
[01:00:10] It can be done, but better to have a similar encoding only.
[01:00:15] That can be done also, but uh…
[01:00:16] Okay.
[01:00:17] Uh, internally, they would be treated as Boolean data or whatever, but it's better to go for a single
[01:00:24] scheme rather than using two of them.
[01:00:27] For the same, uh, data frame, or same data set.
[01:00:32] Okay? Yes.
[01:00:35] So, over to you, Ashish. Please continue from there.
[01:00:40] Okay, so, uh, there is a…
[01:00:43] doubt regarding the code.
[01:00:46] Uh, okay, Lokesh. So, uh…
[01:00:49] So, basically, the doubt is regarding the scale… the standard scalar thing. So, uh, so I have written something more to add to it. So, basically, we are doing, uh, we are doing fit underscore transform in the train data.
[01:01:03] And on the test dataset, we are only doing a transform. So, this is the, uh, like, the doubt, uh, location, is this correct?
[01:01:11] Yeah, that associates.
[01:01:14] Yeah. So, basically, what does, like, fitting means here? So, I have written here that scale… scalar equals standard scalar, and then we are doing fitroscope underscore transforms.
[01:01:23] After this line, uh, what is basically learned is, in fitting, so there are four, as we saw, there are four input features. So, it learns mean for all of these features,
[01:01:34] And it learns variance for each… all of these features, so I have written here mean and variance. So, uh, and what will… it will do is, for each of the entries, it will go
[01:01:43] and subtract the value of that column entry with its mean, and divide by its variance, so that overall, now when we will do a mean, it will give mean as 0 and variance as 1. So, what we want is…
[01:01:56] We want here only to learn this mean and variance only from our training data.
[01:02:00] Because, uh… and during testing, whatever the mean and variance we learn, we will directly use it and only apply transform.
[01:02:08] So, that's why we are using only fit underscore transform we are using for train, and this we are using test, because during runtime,
[01:02:16] VA… when a single sample will come,
[01:02:21] We will, uh, we already have, from our training set,
[01:02:24] fitness transform, we will only apply transform over there. Yeah, uh, so here, it seems like some kind of odd. So, basically, this is like for model development. After…
[01:02:34] You have, like, come up with a model, you will use all of your data and learn fit under a transform, and now, during runtime, only a single sample will come each time.
[01:02:43] to predict. That time, you will only use its transform. So, since we have, uh, we are giving some kind of, like, sense how it will work on unseen dataset, so that's why, since test is unseen dataset, we are only doing transform over it.
[01:02:59] Okay, uh, thank you. So, yeah, so…
[01:03:03] Okay, so, uh, we were at this place, so finally, at this place, we have, like, combined…
[01:03:11] train and test, and as we see here, so this was the shape.
[01:03:15] This was a shape. This will overall extract final. This is my shape, 1390312, uh…
[01:03:23] 13,000 rows and 12 columns.
[01:03:26] So, uh, fire, uh, so, when this has been combined, now, uh, uh,
[01:03:32] We will call the linear regression, uh, like the class.
[01:03:37] From a skill run, so here it is, SK linear model import linear regression.
[01:03:44] So, we call this… and as we saw in the legendary class,
[01:03:47] the commands are standard here.
[01:03:49] So, there is a fit function over here, uh, where the actual training happens. It learns the slope, the coefficients for each of the variables, whatever the coefficients, everything. It is learned here.
[01:04:01] based on the trend data.
[01:04:03] So, this is being learned here, and once, uh, so basically, training happens here. Now, uh, once it has been learned,
[01:04:12] Uh, we can, like, use it to predict, and here, we could also… so, basically, now we are using… so basically, we could also predict the trend dataset, but obviously, on which… on that, the training had happened. So, like, there is, like, so we just… so that would be obviously of good value. So we are only showing here that performance on test.
[01:04:34] So, once training has happened,
[01:04:36] We'll use the same classifier, so REG is the, uh, sorry, regressor here, the same regressor, linear regressor.
[01:04:43] And now, there's an internal function called predict, to which you only pass the
[01:04:49] like, the test samples, uh, or the samples on which you want to predict. So, basically, these are the samples on which model was not trained. So, here, these correspond to the 20% test.
[01:04:58] split we took. So, they are in here. And, uh, remember, we, uh…
[01:05:04] like, uh, here, they started from here.
[01:05:07] Uh, from this place, we started them, uh, like, from here. So, we…
[01:05:13] applied the same thing, the same things are being applied to convert the test sample also, so the test samples were transformed. Remember, not fit transform, only transformed when we are doing their… when we are normalizing the numeric features.
[01:05:27] And similarly, ordinal encoding, where when the ordinal encoding, there also we applied only transform.
[01:05:33] So, overall, uh, this is also converted into the same form of test dataset, and, uh, but… so the difference is, we are not using this for training, but the form which we will give to the model, because it has only seen that data in the form,
[01:05:50] that we give the train data. So, the exact form, we have to, like, make our test data, and then only we can give it for prediction. So here, it will, uh, predict, uh, continuous value, which is our target, uh, of which, what is considered the target variable, given this input variables.
[01:06:06] So, here, at this step, we get the predictions. So, uh, once we have the predictions,
[01:06:11] Uh, so, uh, and we, uh, because we have ourselves created 80 and 20% split, so for these
[01:06:18] This data, we also have the ground truth value, which are what was the actual count of the bicycles.
[01:06:22] So, we can use that, and in regression, like in class regression, there were accuracy, precision metrics, so similarly, there are metrics in regression.
[01:06:31] So, uh, just to show you the library, so…
[01:06:35] I don't have to write, like, separate function for everything, because SKL and metrics, as we… as I told earlier, also, that it has all these things already defined here, because it is a, like, a library for machine learning. So, we are using
[01:06:47] Regression metrics from the same SKLN.metrics, it contains metrics for both classification and regression, so there is
[01:06:54] mean absolute error, mean squared error, R2 score, uh, we are importing from there.
[01:07:00] So, basically, fastest mean absolute error, it will directly… we give it the ground truth, and the prediction, it will calculate the mean absolute error over it.
[01:07:08] Uh, similarly, if we give… then we call mean squared error over it.
[01:07:14] We give it the ground truth and the prediction, and then if we do a route over the mean squared error, we can get the root mean square value.
[01:07:22] And, uh, and the R2 score?
[01:07:25] Again, we pass on the ground truth and the predicted value.
[01:07:29] And then, uh, like, this code is just prints those values, so if you see the first part here,
[01:07:34] Uh, just sorry to interrupt.
[01:07:37] Any questions, anyone, so far, or any part of the code?
[01:07:43] whatever we have done, any questions?
[01:07:47] One question, ma'am.
[01:07:49] Please go ahead, yeah.
[01:07:50] So, uh, here we have used, uh, linear regressions, so…
[01:07:55] Uh, we have, uh, this number.
[01:07:59] accuracy. So, we may use some other model also, right?
[01:08:04] to check this accuracy. So, it is, like, already, it is decided, or we will check each and every, uh,
[01:08:09] algorithm, and we will check the accuracies of this model.
[01:08:14] Yeah, you'll have to actually, because for, uh, let's say if we have a new dataset,
[01:08:19] And we don't know much about the dataset, then…
[01:08:23] And you don't know whether it's linear, non-linearly separable, and what it is like.
[01:08:28] So, actually, you'll have to go ahead and use different classification models on it.
[01:08:33] And see which one is giving the best results.
[01:08:37] Uh, so here, because we want to focus on linear regression only, so that is the one that, uh, you are seeing here.
[01:08:43] However, uh, you'll also see at the end, uh,
[01:08:47] Uh, that are summary of different, uh, classifiers and options are also shown.
[01:08:52] To see which one is going to give the best performance. So, always, whenever, usually whenever you have a data,
[01:08:59] Then you try different classification models on it.
[01:09:02] And then conclude that this one is giving the best performance, and you finally decide to use that for further analysis.
[01:09:09] That's the usual way.
[01:09:11] to do.
[01:09:12] Okay, okay, ma'am.
[01:09:15] Another thing I wanted to add here is that as we had discussed R-squared yesterday, you can see here the value of R squared is, like, 0.388.
[01:09:25] Which means that this particular model captures 38.8% of the variance
[01:09:31] In the target variable.
[01:09:33] Uh, right, and similarly, if you see… so this is the linear regression with integer encoding.
[01:09:40] Now, the other option is linear regression with one hot encoding.
[01:09:44] And when you use one-hot encoding, what you can see here, that the value of R squared is 60…
[01:09:50] 8.14, so it's almost kind of double.
[01:09:53] So, if one hot encoding is used, then the model
[01:09:57] captures 68.14% of the variance in the target variable.
[01:10:01] Which is, uh, quite a good…
[01:10:03] thing, or it's much, much better than what, uh, you know, linear regression with.
[01:10:10] integer encoding could do.
[01:10:12] So, this is also one way to decide
[01:10:14] which particular encoding to use, whether integer encoding is fine, or whether we should go for one-hot encoding.
[01:10:22] Because you can try, in addition to the other, um…
[01:10:26] Uh, to the other performance metrics, like MAE, MSC, RMSC, which also you are seeing here.
[01:10:33] For example, mean absolute error is…
[01:10:35] much higher in case of…
[01:10:38] Integer encoding as compared to one-hot encoding.
[01:10:41] Uh, MSC is also much higher. Our MSC is also quite high.
[01:10:46] And R-squared value…
[01:10:48] Uh, should be high, um, and it is lower in case of the first model with integer encoding.
[01:10:56] So, all in all, the second model, uh, using linear regression with one-hot encoding,
[01:11:00] is giving much better performance in terms of the error.
[01:11:04] As well as in terms of capturing the variance in the dependent variable.
[01:11:09] Okay, so this is what I wanted to add. Yes?
[01:11:10] I have, uh, one more question. Like, how we can interpret MAE, MSE, the Ranger?
[01:11:17] So, lower the range is good for MAE, like…
[01:11:22] So, yeah, so errors, see, or MAE, MSC, or MSC.
[01:11:27] All of them represent error.
[01:11:29] So, definitely, we want the error… what is error? Error is the difference between the actual and the predicted value.
[01:11:36] We want the predicted value to be as close as possible to the actual value.
[01:11:42] Or, in other words, because the difference between the predicted and the actual value is the error,
[01:11:48] We want the error to be very low.
[01:11:50] Ideally, it should be 0. In real-world applications, it will not be zero, but it will be something which is non-zero.
[01:11:57] So, we want the error to be as low as possible.
[01:12:00] So, you can see here, mean absolute error, MSC, RMSE.
[01:12:04] So, they don't have a standard range, but in any case, if we are having different models,
[01:12:09] To look at the values and see which model is having lower.
[01:12:13] So, in this case, the linear regression with one-hot encoding is having lower values for
[01:12:18] Error when we measure mean absolute error, which is MA, or mean squared error, which is MSC.
[01:12:24] or root-mean-square, which is RMSE.
[01:12:27] So, all three error metrics are giving lower values when we talk about one-hot encoding.
[01:12:32] So, this means the model which uses one hot encoding is giving lower error, means it is
[01:12:40] Predicting values,
[01:12:42] Uh, which are much closer to the actual given values.
[01:12:46] So, this is a better model.
[01:12:49] It's okay.
[01:12:51] Yeah. Deepak, what question do you have?
[01:12:56] Yeah, ma'am. So, what is our take-up in terms of why one hot encoding is performing better than the integer one?
[01:13:04] So, uh, probably… see, the thing is that…
[01:13:06] Uh, when we talk about one-hot encoding,
[01:13:09] We are actually replacing each and every value
[01:13:13] with an integer. So, the advantage of this is that we are not increasing the number of attributes, the number of attributes is still 1.
[01:13:22] And we are replacing that attribute with integer values, right?
[01:13:26] But the downside of it is that
[01:13:28] These values actually are not numbers.
[01:13:31] For example, if we talk about month 1 to 12, these are January, February, and all.
[01:13:36] They're not numbers, so to say. They are represented as numbers, but they are not numbers.
[01:13:42] Now, what happens is that suppose some distance measure is being used,
[01:13:46] like Euclidean or some other distance measure,
[01:13:49] These numbers are treated as numbers, actual peer-not numbers, they are just categories.
[01:13:54] Representing months.
[01:13:56] However, if we talk about one-hot encoding, this problem will not come in one-hot encoding, because
[01:14:02] We are adding as many columns as the number of values in that particular field.
[01:14:07] If we are having 12 months, then we'll be adding 12 value… 12 attributes.
[01:14:13] Uh, one each for one particular value. For example, for an ear, we'll be adding 12 months, January, February,
[01:14:19] March, April, so on and so forth, till December.
[01:14:23] And based on whatever month,
[01:14:25] is mentioned in the record.
[01:14:27] Corresponding month will be assigned 1, all other months will be assigned 0.
[01:14:32] So, the, uh, downside of this is we are increasing the dimensionality
[01:14:37] To a large extent, but the advantage is that the…
[01:14:42] Data, you know, the…
[01:14:43] Uh, here, there is no such confusion that numbers will be treated as
[01:14:49] Uh, you know, I mean, categories will be treated as numbers. That will not happen.
[01:14:55] And that is probably the reason why one hot encoding is performing better.
[01:15:01] As compared to integer encoding, because in integer encoding, encoding is representing months,
[01:15:07] But they are being inferred as numbers.
[01:15:12] Whereas in one-hot encoding, these will be…
[01:15:14] Uh, vectors having zero ones like that. So, that problem will never arise.
[01:15:20] Is it okay, Deepak?
[01:15:23] Any other question, anyone?
[01:15:24] Shalom, thanks.
[01:15:25] Okay. Yeah. Yes, ma'am, yes.
[01:15:28] Okay, over to you, Ashish.
[01:15:29] Yes. Ah, sis…
[01:15:30] Okay, you have a question?
[01:15:31] Yeah, uh, since, uh, uh, like, uh, in the core,
[01:15:35] Uh, we are… we are calculating, uh, ordinal encoding and one-hot encoding, right? Both we are calculating, but, uh, in terms of inputs,
[01:15:44] Uh, we are getting the same input from the same dataset, uh, prepare feature.
[01:15:52] just slightly go up.
[01:15:57] Select… yeah, yeah, yes, yes, yes, uh…
[01:16:00] there itself, huh?
[01:16:03] Yeah, which…
[01:16:04] No, please go up, up, up, up, up. No, means I just want to know, like, how this, uh, ordinal encoding and one-hot encoding
[01:16:12] Having that different… using the different function to getting the different data.
[01:16:16] Like, uh, I was comparing the score almost the same, I could not find the difference.
[01:16:21] Like, yeah.
[01:16:22] Okay, so till now, I covered only the coding part, I covered this till this part, which is one-out encoding.
[01:16:27] So the difference I will show now, like, at which place that difference was there, you know, why not only going to be.
[01:16:28] Yeah, yeah, okay, okay, yeah.
[01:16:33] Okay, so till now, uh, like, the first… so this…
[01:16:37] uh, like, the print statement is having two blocks, so only the code is covered till this part, the first print statement. So there are two print statements in this block.
[01:16:45] So, at this time apart, the integer encoding is finished.
[01:16:49] So now, uh, we, uh, so as there are two options to the category variables, to either do integer encoding or one-hot, we will see the one-hot encoding.
[01:16:59] So, almost everything's our same,
[01:17:02] Uh, like, as point about Chanchar, that, uh, like, almost… so, just a difference of the one-out inputting will be there, almost all other things are same. So, for that,
[01:17:11] In SKL under preprocessing, there is a, like, here it was…
[01:17:16] Uh, ordinal encoder, which we imported. So here, we are importing one-off encoder, which is also there in SKL node preprocessing.
[01:17:22] And we will split in the same manner. We'll obtain the numeric columns separately, because we will do their normalization.
[01:17:30] And we will, uh, like, split the subset of the category columns separately, because we will do their one-hot encoding.
[01:17:37] So this part, standard scalar, where we are normalizing the numeric features to have the 0, median unit variance, this is exactly same as discussed. The difference come in this part.
[01:17:49] Here, in the… as… in the earlier part, as you saw, there was ordinal encoder, which was called.
[01:17:54] So here, for categorical features, the… now we are calling One Hot Encoder.
[01:17:59] So, this is the, uh…
[01:18:01] like, the major difference, uh, which has come, one-hot encoder. So, uh…
[01:18:06] And basically, like, what is the usual behavior of this function is that when you use it, so because, uh, as ma'am said that we will have one column for each of its value. So, one… every column will be binary, so…
[01:18:20] When there are too many columns, so matrix… the number of columns become very large, so it's…
[01:18:26] Default behavior is to give a very compressed type of metrics, only metrics where… not to give the full 2D metrics, because there are too many zeros.
[01:18:35] So, in order to not do that, I have mentioned here that sparse output equal to false, otherwise it will give only a sparse kind of matrix.
[01:18:44] So, we want the full matrix where, uh,
[01:18:46] no matter how long it is, so…
[01:18:49] We want a full 2D metrics, that's why it is in sparse equal to false.
[01:18:53] And then, uh, handle unknown equal to ignore. So basically, it is like… it is for handling the cases, uh, because we will doing, uh, because one-hot encoding will… one-hot encode will encode all the categories which it sees in the trend data.
[01:19:07] So now one, uh, one new category, if it comes during test data, it does not know to which
[01:19:14] of the, uh, one hot-encoded values, it should map to. So, we will ignore that particular
[01:19:18] row all together, because we don't know, uh, which it will map to. So, so that's why handle unknown equality ignore is written here.
[01:19:26] So, uh, so again, the other thing is, like, the same. I pass it the categorical data,
[01:19:32] And I do a fit underscore transform only over the, uh, like, the, uh, over the training data.
[01:19:38] Of the categorical features. So, this is the part here. So, when I do this, I will obtain
[01:19:46] like, the one hot encoded, uh, training data.
[01:19:49] And similarly, I will do only transform. Fitting, I will, uh, fit underscore transform is for trained data. Only transform is for the test data.
[01:19:59] And then I will get 100 encoded test data.
[01:20:01] Now, again, as we saw, that these are now two different views. One is only numeric features, and one is only categorical. We combine them together before giving to the model.
[01:20:11] So, at this place, horizontal stacking, we are doing that thing, we are combining both of these views together. So, you can see, like, I have separately put, uh, printed this code. The one hot encoding part. You can see that numeric features were only 4,
[01:20:26] And originally, so there were 8.
[01:20:31] Uh, columns for the categorical, and it was integer encoded, so for each column, there was only one. For each categorical column, there was one column.
[01:20:37] So, now, when we do one-hot encoding,
[01:20:40] We see that categorical columns have risen up till 57.
[01:20:44] So, this is, uh, like I mentioned, sometimes it can be problem. Here, only there were, uh, like, uh, for… there was one column where there were 24 categories, and one column where there were 24, otherwise there were very small categories. So, you see…
[01:20:57] Uh, just to go upwards, so yeah.
[01:21:00] So, as many countries are there, it creates those many columns. So,
[01:21:06] If all of my columns would have a greater number of categories, so it would have risen over 200, 200…
[01:21:13] more like that. So, right now, since number of categories are limited, so what we are seeing is,
[01:21:19] that it has given us 57 different columns. So, when we add them together… So, this is the data, 61 columns will go.
[01:21:28] inside the linear regression model. This is the input data, which will pass to the linear regression in this case. Remember, in case of integer encoding, only these 12 columns
[01:21:38] per sample, they were being passed, so 13,000, 12. These 2D metric was passed during integer encoding.
[01:21:46] Right now, this 2D matrix is being passed. So, this is the, like, the difference at the data level, how… which, uh, of using the different encoding strategies.
[01:21:56] So, here, uh, once we have this, uh, again, the same linear regression.
[01:22:02] We do a fit over it.
[01:22:03] And then, at this place, model training finishes.
[01:22:08] And then, since your model has been trained, we can use our test data and see how the predictions for our
[01:22:15] test data. So, see there, here, we are only passing the trained, uh, we are passing X, we are not passing Y.
[01:22:21] And during training, both X and Y were being given to the model for training. So once we have the predictions, again, we… since we have the ground truth also, and predicted values also, we can again do these metrics of error, mean absolute error, mean squared error, uh, RMSC, root mean squared error, and R2 score.
[01:22:39] And as we discussed, for one hot encoding, our results came… are coming better here.
[01:22:46] So, why not encoding is having benefit here.
[01:22:49] So, and sim… and in this block, we just, like, sometimes these values become, uh, like…
[01:22:55] Uh, like, these four… we are doing so much, and just looking at these four values, it does not give that much, like, perspective what we are doing, so…
[01:23:03] At the data level, we… what… at this code, what we have done is, like, we have taken the test dataset,
[01:23:09] And printed the ground truth and the predicted values, the original values and the ground truth value, predicted values, and the original values simultaneously together.
[01:23:19] Uh, for just 10 samples, we have taken it, so there is, uh, 10. So, we can, uh, so this has been done for both ordinal and integer encoding, so you can see at the end, so this is a data sample, and…
[01:23:32] At the end, you can see that this was the actual value given all these input variables.
[01:23:38] Uh, so this is integer encoding, so that's why you are seeing, like, only 12, 13 of them. And…
[01:23:44] This is the actual value, this was what predicted by, for this sample, what was predicted by the linear regression when integer encoded.
[01:23:51] So, these are for few of the samples, which we are showing here.
[01:23:56] And similarly, like, this can be shown for the, uh,
[01:24:02] one-hot encoding. So, remember in one-hot encoding, there were too many columns, uh, but what I'm doing is I'm showing, uh…
[01:24:09] the columns, the original columns of how the original columns look. The initial columns which were there in the dataset, because showing 57 columns, these were, like, our transformation steps.
[01:24:22] But originally, what were the initial column? We are showing for that. So, you can see now, these are the predicted values when one hot encoding is used.
[01:24:29] So, at individual sample level, we can also see, like, how much difference is there, and the error and R2 score, of course, gives us an overall performance on our test data.
[01:24:39] So, this thing, like, uh…
[01:24:42] is finishes, like, the first block of our… first part of our linear regression. Any, like, doubt over it?
[01:25:05] Okay, so… I think we can proceed and finish this linear regression then. There is no question.
[01:25:10] Okay. Yeah. So, uh, just, uh, like, linearization is almost finished there, but, uh, we can…
[01:25:16] further do some improvement over here, uh, so that's why, like, we predict, uh, print, uh, like,
[01:25:23] plotted the target variable Y.
[01:25:25] And we saw that it has the right skewness over it. So sometimes, when there is a, like, a long tail over this, what we can do is, if we lock transform our…
[01:25:33] Uh, target variable.
[01:25:35] It becomes slightly normal in distribution, as you can see here. Whenever, uh, so basically, whenever there is a skewness, right tail, uh, there is a tail, long tail, uh, in the data,
[01:25:46] in the right skewed data. Log transformation sometimes helps.
[01:25:50] So, what we'll do is, the same setup which we used, everything will be same, just the…
[01:25:57] like, the target variable,
[01:25:58] will be, uh, we will, uh, we will do a log over it. Instead of using the original target variable, we will use the log version of it, and then, uh, plot… then do the same linear regression.
[01:26:10] So, if you see the code here,
[01:26:12] The code is almost same that we use the numeric column, we use the categorical column.
[01:26:19] And the difference now is that the training will happen on the log transform values of the Ys instead of the original Y's. So, you can see that what I am doing is, from NumPy, there is a… this NP denotes numpy,
[01:26:33] So, yeah, so they're importing NumPy as NP.
[01:26:37] I'm using log over this, and what it will do is, it will do a, uh, like a log transformation.
[01:26:47] of these values of the wide train and wide test. So now, uh, if you see the rest of the code is pretty much same, that…
[01:26:53] You obtain the numeric columns,
[01:26:54] You obtain the categorical columns.
[01:26:57] You do normalization of the numeric features using standard scalar, and then, uh, for the categorical features, you either do this integer encoding,
[01:27:06] Using this ordinary encoder,
[01:27:08] And then, uh, combine them together.
[01:27:11] And then, here, instead of just doing Y train, I am using the log version of it.
[01:27:16] Uh, that's the difference.
[01:27:18] And, uh, yeah, so… but now, my original values were in, uh, like, my ground truth was in non-lock form. It was… so in order to compare with my ground tooth,
[01:27:28] whatever is my predictions will come, finally. I will apply, like, a exponential over it, uh, so in order to reduce log. So, exponential over log will, uh, doing exponential over a log value.
[01:27:40] will again give me the original value back.
[01:27:43] So that's why my predictions are again in the same scale as my ground tooth was. So, I obtained them here.
[01:27:49] And once they are in the same scale, I use the mean absolute error, uh, mean squared error, and everything. So, this thing is again repeated for integer encoded.
[01:27:50] Uh, you can, I think, proceed.
[01:27:58] And a one-hot encoding.
[01:28:02] Uh, the same thing is repeated.
[01:28:04] Just, uh, wanted to add here, sorry to interrupt Rashid, that, uh…
[01:28:09] This log transformation is done usually when the data is skewed.
[01:28:12] As shown in the distribution, you can see that the data is right-tailed, you can see it is, uh, distribution is very much skewed.
[01:28:22] Uh, towards the lower values. So, wherever the data is, uh, skewed in this way,
[01:28:28] specifically write skewed.
[01:28:30] Then, uh, we prefer to use, um,
[01:28:35] Log transformation, so that the model, uh, that is created on that data is more
[01:28:41] Uh, specifically when we are talking about, uh, linear regression.
[01:28:45] So that's the reason for using log transformation.
[01:28:49] So you can keep that in mind, that whenever your data is…
[01:28:52] Uh, ride scoot, like what you are seeing, you can think of using log transformation.
[01:28:58] Because it will give you a better model with lower errors and all that.
[01:29:02] Yeah, that's it. Please continue.
[01:29:04] Yeah. So, everything, almost the code is same, like we discussed, uh, and again, we use these metrics, uh, so, uh, we can see, uh, like, this is the…
[01:29:16] metrics for the integer encoded, when that target was log transform during learning.
[01:29:21] For, uh, both one-hot encoding integer. So, uh, like, we can, like, compare all four cases together. So, to have a better look, uh, and we can see that…
[01:29:32] Log-transformed help held more in integer encoding.
[01:29:36] For, uh, did not help in integer encoding, the values are almost, like, same.
[01:29:42] Even R2 score has decreased, but for one-out encoding, uh, like log transform has even improved the result. So, from 68 of R2 score, it is now 71.
[01:29:53] So, for one-off encoding, lock transformed has helped. So, overall, we can conclude that integer encoded, uh, here, uh,
[01:30:01] overall is not helped, so be it the lock transform target or the normal target variable. So, the performance is, uh, same, similar, similarly inferior. But for, uh,
[01:30:10] normal one-hot encoding, when we now did its log transform, that performance increased more.
[01:30:16] So, this finishes the, like, the linear regression part.
[01:30:19] So, any doubts till this part? Then we'll move to… then we'll move to naive base after that.
[01:30:31] I think you can carry on.
[01:30:32] Okay. Okay. So, uh, 9 ways, uh, so basically, uh, the thing is, uh, knives works only for classification problems, so that's why we have to, like, uh, introduce a
[01:30:43] Second dataset also in this class.
[01:30:45] So… so basically, this dataset is, uh, called Adult Sensors Income Dataset.
[01:30:50] And it contained demographic data from the Census of USA, of the year 1994.
[01:30:57] And it's available, uh, the links are given, the original sources from which dataset his, uh…
[01:31:02] has been fetched. And here, uh, what they are doing is, they are…
[01:31:07] Uh, the target variable is a binary variable.
[01:31:10] which is having value greater than 50K or less than 50K. So, basically, it's a 0-1 kind of problem.
[01:31:15] Whether the person income is
[01:31:17] less than $50,000, less than equal to $15,000, or greater than $50,000?
[01:31:22] So, uh, so this makes it a binary classification problem, because target variable has only two values, and the earlier we saw regression where target variable is continuous value.
[01:31:33] So, again, uh, if, uh, like, we have shown the various features of this dataset, various feature…
[01:31:41] So, some are categorical, some are numeric, and…
[01:31:48] Okay, some are categorical.
[01:31:49] Come on!
[01:31:55] Okay, some are, uh…
[01:31:56] Categorical, some are numeric features. And overall, all one information, like, initially, we get from the data is that, uh, this column
[01:32:07] final sample weight is for… it was more used for population… basically, it was a variable,
[01:32:15] That represents, uh, uh, like, uh, how many people in a record represents in a population. So, basically, it's because it's a sample.
[01:32:21] So, it does not help directly in predicting it's an individual income, so we'll be dropping it. It was basically more for, uh, how the data was fetched.
[01:32:29] It was for that, uh, it was taken, but for a classification,
[01:32:36] This final sample weight, because it basically represents how many… each sample represents how many people in a population, so it's kind of a weighting value, so we will not include it, uh, like, classifier, the input features for classifier will not include this.
[01:32:51] variable. So…
[01:32:54] Again, uh, like, uh…
[01:32:56] So, this data…
[01:32:58] If you load it yourself, if you open it, it's… it is by the name Adel.data. If you'll open it in a text reader, uh, you will see that the column names, I think, are not there.
[01:33:10] Uh, so that's why, like, uh, so I… we have… header is none is written, because if header name… none I will not give. So,
[01:33:19] The first row, it will treat as header. So, pandas will treat the first row as header, that's why I've specified header, name is none.
[01:33:24] And I am manually supplying the column names.
[01:33:30] So, explicitly, I've written this column names by the knowledge of this… the data guard of this model.
[01:33:36] What are the column names? So here, names I am passing as the column names. And if you look at the raw version of the data in the text editor or something, you will see there are some question marks in the dataset.
[01:33:46] So, what we want is, uh, we want pa- we- we want Pandas to treat this question mark as NN values.
[01:33:53] So, that's why we have mentioned here,
[01:33:55] that any underscore values, that wherever there will be a question mark, and sometimes there is the exact question mark, sometimes there is a space followed by a question mark.
[01:34:05] So you… whenever you encounter this value,
[01:34:08] type them as NEN, enter them as NAN in your, uh, when you read them.
[01:34:14] So, we pass on this thing, and finally,
[01:34:17] Like, we print the data frame here.
[01:34:19] So, this is that, uh, data frame which we will get. So, uh…
[01:34:25] Initially, also, we know that there are, uh…
[01:34:27] like, missing values in a dataset, and this can be…
[01:34:30] like, confirm if we do a df.info. Uh, so there are, I think there are just two columns where there are missing values.
[01:34:38] In all other columns, uh, there are 32,000, uh, entries are there. So, yeah, in one more column in third, there are.
[01:34:46] So, basically, there are 3 columns where there are missing values.
[01:34:49] So, uh…
[01:34:51] This can be, like, uh… so, uh, we can, like, also print, uh, and see…
[01:34:57] So, yeah. So…
[01:35:00] Uh, so basically, uh…
[01:35:03] You can see that since I was told that question mark should be NAN, so it… wherever there are missing values, it has shown, like,
[01:35:09] NAN values are there, uh, here. So, we can see that some values are treated as NAN. So, uh…
[01:35:18] So, in order to get how many are missing, what we can do is df.isNum,
[01:35:25] And sum over each of the columns.
[01:35:27] And we can see that, uh, for 3 specific columns, for work class,
[01:35:31] There are 1800 missing for occupation, similarly, 1800 missing.
[01:35:36] And native country, uh, 583 missing.
[01:35:40] So, uh, what, uh, further we have done is, like, uh, we, uh, since we could use that missing value strategy, imputation strategies,
[01:35:48] But since we are doing multi-novel, and we want to do knive ways, so we stick to a simpler way of, like, dropping these missing values.
[01:35:55] So, uh, first, in order to, uh, like, what we are doing is, we are…
[01:36:01] Just printing how much of the total samples are the missing values. So, basically,
[01:36:08] For work class and occupation, it is 5% of the total SAM values are missing, and native country, only 1% are missing.
[01:36:16] Uh, so, uh, uh, for simplicity and not complicating anything, we just dropped
[01:36:21] All the rows where there are missing values. So, overall,
[01:36:24] After original rows were 3 to 561.
[01:36:27] And for this line, df.dropNA,
[01:36:32] we… you were… after this line,
[01:36:34] we get these many values.
[01:36:37] So, this is now my, uh…
[01:36:39] data without any missing values, and once I… again, I do a df.cleanInfo, this is called df underscore clean now, this variable, uh, the…
[01:36:49] data set… data frame without any missing values. And you can see all of them are, like, uniformly
[01:36:54] 30,000 values are there, without any missing one, and again, for just a sanity check,
[01:37:02] It is the same number that for any column.
[01:37:04] all the missing values are zero.
[01:37:06] So, yeah. So, once we have, like,
[01:37:09] Sorted with missing value. Again, uh…
[01:37:12] based on, uh, like, the…
[01:37:16] Uh, first, based on the data type. So here, uh, if we see, uh, in this dataset, uh,
[01:37:22] It has, uh… the problem we faced in the earlier one was that even for the categorical variable, it was treating them as object ones. So, but in this dataset,
[01:37:31] Uh, the categorical variables.
[01:37:33] are more or less, uh, in the object data frame earlier, because the categorical variables here are generally in the string form.
[01:37:40] And the, uh, whatever is in the integer, uh, variables, it is more like a numeric feature right now.
[01:37:47] So, yeah, so…
[01:37:50] Yeah, even that target variable, it is calling its object, because it's a binary classification, and they have used string for it right now, so it is calling them object.
[01:37:59] So, that's why I have directly done is, uh, selected, uh, data types wherein object is included. That is our categorical right now.
[01:38:08] And selected data types, where there is integer and float, that will be a numerical columns for… because in this dataset, uh, the information which was shown in data card that these are the categorical and these are numeric,
[01:38:19] the pandas loaded in that format only, so that's why.
[01:38:22] We directly use this. So, these are our, uh… so basically, 9 are categorical columns.
[01:38:29] And, uh, 6 are numerical columns.
[01:38:32] So, after that, uh, we can, uh, also see…
[01:38:37] The unique values in each of the categorical columns, because in order to see, like, how many… if you do a 100 encoding, how many unique values will, uh, like, it will explode to. So, we can see that, uh,
[01:38:49] for these 9 columns,
[01:38:51] like, there are 7 unique, uh, for another column, there's 16 unique.
[01:38:55] And so on, and uh… like, there is a column native country where there are 41 unique values. So, this is the, like, the highest.
[01:39:03] Uh, rest of them are, like, slightly lower.
[01:39:07] And again, like before proceeding to modeling, uh, what we will, uh, do is, like, we will do a, uh, the distribution of the target variable. The target variable here is income.
[01:39:17] And we do its value count of both the values, so there are only two values greater than 50K, less than equal to 50K. And you can see the proportion here also.
[01:39:31] So, the first line, when I print the first line, it will give me this.
[01:39:35] thing. So, you can see that 75% are, like, less than, and 24% are greater than. So, again, it's, like, an imbalanced distribution. And same thing has been plotted in the…
[01:39:47] Like, the graph here. So, uh, so it, uh, is, uh, we can see here, uh, from the face of it, that it imbalanced distribution on the target variable.
[01:39:57] So, uh, now, for proceeding to multi-novel knife base,
[01:40:03] Uh, we have, uh… so, basically, we are showing two versions of NiveWiz. So, the first is, like, multinomial knife-based.
[01:40:11] So, uh…
[01:40:12] Uh, multi-novel knife base, uh, the problem is that it expects everything to be in categorical form.
[01:40:18] But right now, our data… in our data, some variables are numeric, some are categorical.
[01:40:23] So, what we will have to do is, even for numeric features, we will have to discretize them.
[01:40:29] Only then we can use multinomial knive ways, because it cannot understand these numeric features, these continuous values.
[01:40:37] So, that's why.
[01:40:39] Uh, uh, again, uh, like, this code is redundant, the same thing, the categorical and numerical features.
[01:40:45] I've printed out, uh, by using that same select data types. So, as you remember in vision tree, we made the numeric features into
[01:40:56] categorical one by doing a discretizer.
[01:40:59] So this is the same cable-ins discretizer.
[01:41:03] we are calling, and I've specified the binches 5,
[01:41:08] and encode an ordinal 5 and strategies quantile. So, basically, what is… we'll be doing is, it will sort for my numeric column, it will sort the data in increasing order.
[01:41:20] And, uh, so, uh, and a number of wins is 5, so first 25% of data will be assigned group 0. Second 25% of data will be assigned Group 1, and so on, and the last
[01:41:33] 20% of your data will be assigned, uh, Group 4, 0 to 4. So…
[01:41:38] like this, each of my continuous values will be, like,
[01:41:43] encoded into form of a number, into a form of group, either group 0, group 1, group 2, group 3, group 4, like this.
[01:41:52] So, in that way, all of my numeric features
[01:41:56] will now be into a categorical form.
[01:41:58] So, this thing is, uh…
[01:42:01] being done, so that's why, first, we call this K-bins discretizer.
[01:42:06] And then, uh, what we are doing is…
[01:42:08] Taking all my numeric data,
[01:42:12] So, because here, in this fit transcript, everything is performed in the same step. So, we will, uh, take a numerical columns.
[01:42:23] And do a fit underscore transform over it to give me the…
[01:42:27] that discretized data, discretized version of my numeric features.
[01:42:33] So, after that, I obtain it, the discretized version of numeric feature,
[01:42:38] Uh, uh, this… I'm creating a new data, uh, like a…
[01:42:43] new data frame, just to first show my discretized version of the discretized values of the dataset, of the…
[01:42:50] numeric columns. And, uh, yeah, so…
[01:42:53] This is pd.dataframe, this is passed here, uh, this information is passed there, this variable, and we are…
[01:43:01] where this column is what we want the column names to be for each of them. So we are calling whatever is the original name,
[01:43:09] We'll just add bin after that. So, whatever the original name of the numeric column, you will just add bin over it. So, let's see, like, how it looks.
[01:43:16] So, just wait a minute.
[01:43:21] Okay, just a minute.
[01:43:40] Okay, so if I just print…
[01:43:47] So, yeah, so if you see, uh…
[01:43:50] Uh, if I just print, uh…
[01:43:59] In order to compare. So, these were the original, uh…
[01:44:05] Uh, yeah, just, um, just give me one more minute.
[01:44:13] So, uh, just to compare…
[01:44:17] Uh, uh, yeah, uh, just available, yeah. So, just to compare, uh, these were my original, uh, values.
[01:44:27] These were my original values in the numeric form. So, when I discretize them, now, as I told you, that all of them are converted into bins.
[01:44:35] 0 to 4. So, uh, each of these columns…
[01:44:38] was first sorted in an increasing order, and…
[01:44:44] And since we are using quantile strategy, so… and breaking into 5 groups. So, first 25… 20% of the data will be, like, passed to the bin 0, then
[01:44:53] been 1, and so on. So, that's why all of these datas right now, as you can see, they are in, like, bin form.
[01:44:59] So, overall, this is like, uh… and you can see the same number of column length is there, so the number of columns which are here.
[01:45:08] So, if we can… if we do a shape over it. So, 6 columns are there, same 6 columns are over here, right now.
[01:45:14] Just, uh, is that for… from their numeric features, we have
[01:45:18] transform them into, uh, like, they are integer encoded or discretized version right now.
[01:45:24] So, with this, all of the data right now is in, uh…
[01:45:28] After that, all of the data right now is in… is categorical in nature. So, this is required to do multi-level NIFPS. So, that's why we are doing this, otherwise we would not have, like, done in this manner. So, yeah. So, once we have this discretized data,
[01:45:45] Uh, and you remember this, this variable only shows me the discretized numeric features, but there are categorical features also there. So, finally, what we'll be doing is,
[01:45:56] I will… from DF underscore clean, which contained my all the features, categorical numerical, I will drop the numerical columns, because, uh, I have the discretized version right now. And once, after dropping them,
[01:46:09] I'm making a new data frame where
[01:46:12] Uh, uh, basically, uh, the categorical features are there, and the discretized version of numeric features are there all together. So, now, if you will see…
[01:46:24] if I run this…
[01:46:26] So, now, if you will see, so these are… till here, these are the categorical features, which were originally categorical here, because all of these are in form of strings, so pandas, uh, consider them as categorical. And for our numeric features.
[01:46:41] I can… I used their discretized version. I am using their discretized version right now.
[01:46:47] So, this thing…
[01:46:49] So,
[01:46:50] Please, uh, one question here.
[01:46:52] Yeah.
[01:46:53] So, if you discretize the data,
[01:46:54] Uh, 201.
[01:46:57] Some category, 5 category.
[01:47:00] Uh, you have put right. So data, indirectly, it's lost, right? Like, uh…
[01:47:04] The previously bought data, it was having earlier, 30, 40, and we are assigning that one to 0.
[01:47:09] So, how this model will calculate internally, and how it will keep the data.
[01:47:16] how the model will?
[01:47:17] like, uh, we are, uh, categorized all those data into 5, right? 0, 1, 2, 3, 4, 5.
[01:47:23] Yeah, yeah, yeah. Hmm.
[01:47:25] So, in first, uh…
[01:47:27] 20 set, okay, so first half, where we have assigned those data to zero.
[01:47:33] Hmm.
[01:47:34] So, how this model will interface, like, we have put that one to zero.
[01:47:39] And, uh, if new data will come into that field,
[01:47:43] Then, uh…
[01:47:45] So, if new data will come, we will just do a transform over it. So, here we are doing Physical transform. If new data will come,
[01:47:53] You will just do a transform over it, and it will also try to fit in one of the five columns, in one of the five groups.
[01:47:56] Okay, okay.
[01:47:59] So that's… this is how.
[01:48:00] Okay, bye.
[01:48:05] Okay, so… so, basically, all this is being done because, uh, uh, multi-nominal live base, it can't work on numeric features, because it expects that all the input features have to be categorical, because the formula of it is like that it has to treat everything as categorical.
[01:48:19] Uh, the next version is Gaussian Knife. In there, it… same as we saw in linear regression, it treats numeric feature as it is. So, there it is not a problem.
[01:48:29] But in multinomial libraries, this is a problem, so, uh, that's why we are… have to convert everything into categorical.
[01:48:38] So, okay, so…
[01:48:40] So once everything at this step, uh,
[01:48:43] Everything is now in the form of categorical DF underscore discretize, everything in the form of
[01:48:48] categories. So, now, again, everything is a form of category, so, uh, in category, there are, again, two ways.
[01:48:56] Either you do ordinal encoding, or you do integer encoding. So, uh, so here, encode… ordinal encoding is required, uh, is required, because you can see that category features, right now, they are in the form of strings. These are often in the form of strings.
[01:49:14] Uh, had it been in the form of numbers, that means they are already integer-encoded. But you can see, only the, uh…
[01:49:20] whatever we have discretized these numeric features, these are in the form of numbers, but our categorical features are in the form of string right now. So, we have to assign integers to them also before passing to the model.
[01:49:34] 0, 1, 2, 3, like that. So, that's why…
[01:49:38] there are two options. Now, either do integer encoding for all features, or either do one-hot encoding.
[01:49:44] So, first we begin with integer encoding. So, for integer encoding,
[01:49:52] So, right now, uh, we can see, uh…
[01:49:58] So, uh, so basically, here, what I'm doing is select data types which are object. So, uh…
[01:50:04] object, so it will select object all these data types right now, which are in string. So, uh…
[01:50:10] Although these, uh, so basically, since I have already integer-encoded them, I don't want these values to be again either encoded, these are already done. So that's why, uh, just integer encode these, uh, values which are form in string, so that's why I've used object.
[01:50:24] So, it will select, uh, the original values which were categorical.
[01:50:29] So, using this line. Once you have them selected, these columns,
[01:50:35] Again, I'm using the same ordinal encoder.
[01:50:38] And, uh, like…
[01:50:41] If, uh… and then I've given special parameters, like, uh, if some unknown, uh, if some…
[01:50:48] If something maps to an unknown value, like right now, so everything will get mapped, because it is showing all categories. Some new data comes, and it is not mapping to the existing categories, so you give its value as minus 1.
[01:51:02] So, otherwise, if you… if you don't fit a… if you don't use handleUnknown, so what happened is, if you don't give this parameter, so if some, uh, during test time, some, uh,
[01:51:11] some value comes with a new category, it will throw an error.
[01:51:16] So, because it has not seen that category doing, uh, training.
[01:51:19] So, that's why, uh, it is used. So, encoder is, uh, just same, you do fit transform.
[01:51:24] Over all your categorical columns.
[01:51:27] And this is the, like, the final data frame we obtain.
[01:51:31] So, this will now go as input in the, uh…
[01:51:36] The multi-novel knife base. This is my overall dataset, and in order to, like, go input over it, just, I will just remove this income column, which is my thing which I have to predict.
[01:51:48] So, I will remove this one. Otherwise, everything else will, uh, go inside my…
[01:51:56] Like, as an input feature from multinomial knife base. So, for doing that,
[01:52:01] Uh, I just, uh, for starting it, I just, uh, dropped the income column.
[01:52:07] And whatever is left, I treat it as my X.
[01:52:10] And uh… whatever is rema… and the single income column.
[01:52:15] right now is my Y.
[01:52:19] So, this is my X and Y here. I do a train test
[01:52:24] split over, uh, this right now.
[01:52:26] And, uh, so 80-20, the similar 80-20 split. And see, since it is a classification problem, so the… there is an added parameter I am using here, as used in an earlier class, stratify. So, basically, it is
[01:52:40] telling that to spread the dataset, to have the same distribution,
[01:52:45] as the original one. So, if we see
[01:52:49] in original dataset, uh, the distribution of Y was that there was 75%
[01:52:55] less than 50K, and…
[01:52:57] 24% greater than 50K. So, uh, the train and test split should also have the same kind of distribution.
[01:53:05] So that's why we are using
[01:53:07] stratify L. Stratify with using Y.
[01:53:11] So, this thing over here, and after that, you obtain your
[01:53:17] X train and, uh, like, Y train.
[01:53:20] you obtain extra and vitamin X test, and then everything is, like, normal, like, you… hair, just… the thing is that we are calling multi-normal knife base from Skln.knivebase.
[01:53:33] So, as we call, similarly, the multi-level database, then we do a fit over it, over the training set. Once we do a fit over it, so here, uh…
[01:53:42] Uh, then we are, uh, model training finishes at this step, and then we can use it, like, similarly use it for… to calculate predictions. So, uh, we, uh, use… here we are showing for both train and test set the predictions.
[01:53:56] So, passing on the train dataset, X only, passing the text at Y only, we will get the prediction. So these are, like, binary predictions, uh, uh, based on the input data. So…
[01:54:08] Yeah, and since, uh, for this, I will show… we'll show AUC score also. So, these are binary predictions, hard-coded. We also want their probabilities. So, to obtain the probabilities, we have this predict underscore prob A. So, for each production, it will give a probabilistic value, uh, how much it is confident of.
[01:54:26] So, for that, we do this, and these values are obtained here. So, once we have… now we have the predictions, the original values. So, as usual, we can, uh…
[01:54:36] So, you see the same SKL and not
[01:54:39] matrices there. So, right now, I'm importing all the classification-related metrics, accuracy, score, precision, recall, F1, ROC, AUC, classification report. All these are, uh…
[01:54:49] classification-related metrics.
[01:54:52] So, once you call them,
[01:54:55] You can get your accuracy score, you can get precision. Recall F1.
[01:54:59] And, uh, and see, for ROC, AUC, you have to pass on probability values, not the hard-coded predictions like binary predictions 0 and 1, so that's why we are passing the
[01:55:10] probabilities for the trained dataset, and probabilities for the test dataset, which we predicted. So, once we…
[01:55:18] use this, then we can, like, do a, uh…
[01:55:23] After that, we are just printing all these values, what we got.
[01:55:27] So, this is, uh, here.
[01:55:30] Uh, that, uh, this is… this thing is here. And, uh, similarly, in the next column,
[01:55:37] What we did was, uh, everything is same, just we are now using, uh, we are one hot encoding all the features. Instead of just doing a, uh, integer encoding. So, when we one-hot encode
[01:55:50] everything. Uh, so, uh, basically…
[01:55:55] So, uh, yeah, so basically, when you… one-hot encode everything. So, you can see that for each
[01:56:02] column, there is a bi… for each column, if there are multiple categories, so one column, uh…
[01:56:08] created. So this, basically, work class.
[01:56:11] Uh, one column is for federal government category, one column is for local government category, one column is for private. So,
[01:56:18] basically, if I just see…
[01:56:21] DF underscore clean.
[01:56:33] Okay, so, some issue.
[01:56:36] Okay.
[01:56:40] Yeah. Okay…
[01:56:48] Yeah, so there are 7 categories in work class column. So, there will be 7 different binary columns now, uh, for this particular column bar class.
[01:56:57] So, as you can see, so, uh, similarly for all other columns, so you can see now there are total 118 columns which have increased.
[01:57:05] when I'm doing one-hot encoding. So, the overall columns have increased, uh, because…
[01:57:10] one column for each category will be introduced. So, this is how they look. So, this is a binary column 0 and 1.
[01:57:18] For federal government category, similarly for local government category and everything. And at the last also.
[01:57:23] So, when… for hours per week, when VIN value was 0, there is one column. When VIN value was 1, there was one column.
[01:57:31] VIN values 2, there is one column. VIN value is 3, one column.
[01:57:35] And so, like this, uh, yeah, so… so there will be categories like this. So…
[01:57:42] Yeah, so, uh, that's why these have increased.
[01:57:46] So, uh, rest of it, the code is, like, same, new, uh, take the entire dataset.
[01:57:53] So, basically, you take the entire dataset that will be…
[01:57:57] Uh, my…
[01:57:59] So, here, uh, already, like, income column is, uh, like, not there.
[01:58:04] So, while creating this dataset,
[01:58:07] Uh, what we, uh, what we did was that
[01:58:10] We remove the income columns, and then created one hot encoding. So, income column is already not there. So, that's why
[01:58:17] or I directly used this one-hot encoded discretize as X.
[01:58:23] income column, uh, like, Y was, like, already here. I had created this variable.
[01:58:30] which is having Y. So, these are treated as my X and Y, and then I do a split, the same train test split, and use multinomial knife base over this.
[01:58:40] Uh, thing, and you can see that, uh, like, using multi-novel libraries, there was 78% accuracy, now accuracy is improved. Other metrics have also been improved.
[01:58:51] Uh, when I'm using, uh, one-hot encoding with multinomial live base, as compared to
[01:58:56] integer inputting. So, at the end, we have all the comparisons.
[01:59:00] Now, similarly, we can also use Gaussian knife waves. So, Gaussian knife waves,
[01:59:07] you don't have to discretize your numeric features, so there was this question that information is lost when we are discretizing them and everything.
[01:59:14] So, uh, yeah, that thing is true. Uh, so, uh, sometimes, uh, you just, uh, so basically, the assumption is that, uh, each fall… column, it follows a normal distribution. It follows this, like, assumption.
[01:59:31] So, now the problem is that you have to, uh…
[01:59:34] integer encode your categorical variables, or one-hot encode your categorical variable. That thing is also always there.
[01:59:41] But apart from that, for numeric features, you don't do any, like,
[01:59:46] like anything else, you use them as it is.
[01:59:49] So, here you can see…
[01:59:52] That, again, uh…
[01:59:54] Like, for creating… so here, first I have… what we have done is, we have integer-encoded categorical features.
[02:00:01] So, how do we have integrated encoded is that, uh, first you have
[02:00:06] obtain your categorical columns.
[02:00:09] all your category columns are obtained.
[02:00:11] And numerical columns are obtained in the same manner as above.
[02:00:14] And, uh, basically, this special code is returned that the target column is income.
[02:00:20] So we don't want that to be integer encoded. We remove it from the category columns list here.
[02:00:27] And these are my categorical features, these are my numerical features, and for the
[02:00:32] categorical features, I will do ordinal encoding, and once I do ordinal encoding over it,
[02:00:38] Uh, using FitTransform, I obtain all of them in a data frame, and the numerical column… so basically, this is my categorical features, integer encoded.
[02:00:48] This is my numerical features in their true form.
[02:00:51] And this is my target variable. Over all of this, I am just putting them… this is just for display purpose.
[02:00:57] I'm just putting them in one single data frame.
[02:00:59] And I've shown ahead of this. So, this is my overall data will look before I will pass it to Gojian Naive, so you can see.
[02:01:07] The category variables, right? Right now, our integer encoded.
[02:01:10] So, there is one column for each one of them.
[02:01:13] And that numeric columns, education, capital gain, these are, like, in their original form, as they were.
[02:01:20] And this is my target variable, income. So, before passing to Gojan Naibes, I will just have to remove
[02:01:26] this column, income.
[02:01:29] So, yeah, so…
[02:01:31] So, if you see in Gojian LiveWest, the code is almost same, just that, uh…
[02:01:37] First, I, uh, drop this income column.
[02:01:40] And, uh, yeah, so if I drop this income column, and if you see this income column is kind of string, so I will first convert into 0 and 1.
[02:01:50] So, for that, I'm having this mapping, that less than 50K is 0, greater than 50K is 1. So, this will convert Y into 0 and 1.
[02:01:57] And then the same, uh, train-test split is happening. And here, instead of
[02:02:04] multi-novel live base, you just have to replace it with Gaussian envy, and we are just importing from same askln.nive base. Instead of multinomial library, I'm importing Gaussian live base.
[02:02:15] So, like, the flow of the… that is, uh, like, fairly similar. Uh, you are fit training over extended and wide train.
[02:02:22] And then you are predicting, uh, for both trade and test data, all the metrics.
[02:02:27] And this is, like, how they come for the Gojian knife base. When they are integer encoded.
[02:02:34] And similarly, the categorical features can also be one-hot encoded.
[02:02:38] Uh, so if we do a one-hot encoding of them, so this time,
[02:02:43] Numeric features are not touched.
[02:02:45] They remain same, so we get slightly fewer columns.
[02:02:49] And, uh, uh, rest of the thing is, like, same.
[02:02:54] Here, and overall, there's a final comparison over it, so…
[02:02:59] You can see that Gaussian knife base, be it whatever encoding, is not performing that great.
[02:03:06] And, uh, like, the greatest improvement is multi-nomial LiveBase with one-out encoding.
[02:03:11] So, yeah. So…
[02:03:13] This is the overall conclude. Yeah, so this is the overall conclusion we draw from the second part.
[02:03:19] that. So, in our case, multi… one-hot encoding with multinomial ways perform very well. Gaussian knife ways, be it ordinal or one-off encoding, performed, like, sim… similar.
[02:03:34] So, like, any questions on the second part?
[02:04:02] Uh, yeah, uh, so basically the…
[02:04:05] Code is finished, uh, like, any doubts or any questions, then we can, like, conclude if…
[02:04:12] Just one question. So, in any of these, uh… You know, patterns, uh… I think we've used 3. So, in terms of data processing, what are the things we have to do?
[02:04:24] I… in last two, I think we have used discretization, uh, right?
[02:04:28] But, in general, what I'll preprocessing of the data that we have to perform.
[02:04:35] I think the first one, uh, you said, the missing value.
[02:04:37] that will drop them. Like, we shouldn't, uh, go and impute those values.
[02:04:44] So, yeah, I just want to understand these three.
[02:04:46] Yeah.
[02:04:48] And what is our ways of, you know, handling the data?
[02:04:50] So, uh, basically, missing values are generally not dropped. We are dropping this because, like, our concentration is on a particular topic, like knife-based or decision tree.
[02:05:00] You know, every class. So, uh, that thing, uh, so, the first class, I think it was just to show the various important strategies. So, generally, we use them.
[02:05:09] before modeling. Uh, but if missing values are… and other thing is, sometimes missing values are very low. Say, in our case, like, only 5% of them were missing values. So, uh, if…
[02:05:19] it's something that's not hampering much, so we drop them. Otherwise, yes, first we will have to think of the imputation strategies, which will help us impute those missing values.
[02:05:31] So, first thing begins from there.
[02:05:41] Okay, and uh… missing value is 1, and I think, uh, we also need to, uh, look, uh, for dropping the columns, like, you know.
[02:05:46] Okay, so, uh, yeah, so, sorry, so sorry for not taking that second part. So, second part is, like, once you have, uh, imported all missing values, then you will, uh, the general pipeline is, you see if you have categorical columns or numeric columns, a mixture of them, or just numeric or categorical columns.
[02:06:02] So, if you have a… normally you have a mixture of both of them. So, when you have a mixture of both of them, you will have to decide
[02:06:09] that how you will treat your categorical columns, how will you encode them? Because categorical columns cannot be understood directly, so either you integer encode them, or one-hot encode them.
[02:06:18] Numeric columns, you can treat them as it is and do normal scaling of them, normalization of them. So, the second decision has to be taken, how to treat these categorical columns.
[02:06:27] So yeah. Like, these are the two things, and apart from this, as you can see, the modeling part is basically the same. You take the train test split, you call that model, uh, so here, in SKLN, everything is, like, similar, same, just the model name is different.
[02:06:42] the fitting function, the predict function, everything is same. The pipeline is same. The major decisions is regarding this, uh, how to treat this categorical columns.
[02:07:02] Any other questions, anyone?
[02:07:05] Before we break for today?
[02:07:16] So, if there are no further questions, then I think then, uh, we can, uh…
[02:07:21] Just complete the session, and then we can break for today.
[02:07:26] Yeah, so, like, we finished, uh…
[02:07:31] We have finished the Navis part, like, the codebook is finished.
[02:07:35] Yeah.
[02:07:37] Okay, then I think we are done with what we wanted to cover, so then we'll, uh…
[02:07:43] We have come to an end of the session, so…
[02:07:47] Thank you all, have a great day. Thank you. Thanks, Ashish. We'll break for now, and we'll meet next week.
[02:07:56] Thank you, everyone.
[02:07:57] Thank you. Bye-bye.