# 03 2025-12-06 Data Preprocessing and Classification course: Module 2 — Machine Learning Algorithms module: Module-2-Machine-Learning-Algorithms date: 2025-12-06 type: transcript video_url: https://personal-learn.armco.dev/files/_Recordings/Module-2-Machine-Learning-Algorithms/03_2025-12-06_Data_Preprocessing_and_Classification.mp4 --- [00:18:20] we… there is always a, like, uh, we don't… we keep some, uh, held-out dataset on which we don't train, and then we, uh, the train model is then given that data, and we see that, uh, how the model is performing. [00:18:32] Because it is the dataset on which we were not training. So, like, a simple function could be written for it, but yeah, but there's a handy function for it inside the skill learn. [00:18:43] And then, since today we are, like, doing decision tree, so there's SKL1.tree module, where decision tree classifier is [00:18:51] present. So, we are using Legendary for classification here. There is also a decision tree regressor. So, if we wanted to do regression, uh, like, if the target variable would have been continuous, so that time-digit regressor would have been used. [00:19:05] And similarly, we will also, at the end, see random forest classifier, so it comes in this ensemble module. [00:19:13] And likewise. So, even the metrics, like the… for classification, there are metrics like accuracy, precision, recall, F1, and everything. [00:19:21] So, either you write your own functions for every one of them, but these things have been, like, included inside the scalar. So that's why Sklearn encompasses everything which is related to machine learning. It is a good package for everything inside, so that's why we use it. [00:19:37] And then there is NumPy. It's of not so much usage right now, but for complex mathematical operation which we want to do on our dataset, it's use. Here, it is mainly used for, like, some calculations are there in between the code while plotting and everything, so that's why NumPy is used. [00:19:57] not that much of use right now. So yeah, so this was some introduction of the packages. [00:20:03] And the dataset, and then the first thing, import pandas SPD. So, yeah. So, this line, uh, I'm calling the Pandas library. [00:20:14] And then I will call this read CSV function, and I will give it my path of my file. [00:20:23] file right now is in my, like, in the current working directory, therefore, I have only to give the name. [00:20:30] Otherwise, if it's located somewhere else, you give the entire path along with the name. So, yeah, if I run it, and I just do a [00:20:37] Uh, like, printed DF, I will get, like, a nice data frame here, so… [00:20:43] This shows me first 5 entries, [00:20:46] And last 5 entries of this, like, the head and tail of the dataset. [00:20:53] So, uh, yeah, so… [00:20:54] As such, it, uh, like, if, uh, so basically, there are, like, you can see there are, uh, it has written, there are 7,000, uh, [00:21:03] rows and 21 columns. So, basically, there are 7,000 samples, [00:21:07] And 21 features are there. And since there are too many features, so as such, if we just look at the dataset, uh, we may not get any good idea what it is. So, yeah, so there is customer ID, and these are, like, some random identifiers which have they made by combination of digits and text. [00:21:25] And then there are these other features, like gender, senior citizen, partners, and I will have to, like, scroll, because there are so many of them features, so you scroll, and then at the last, you have this variable, which is to be predicted. [00:21:38] So, when they are saying that they have 21 columns, [00:21:42] Uh, the input columns are 20. [00:21:47] And the one target column is there, which is the last one, which is the churn, which we have to, like, predict, like, given these input features. [00:21:54] Whether the person will leave the service or not. [00:21:58] Okay, so… [00:22:01] Then we, like, uh… like, this is just loading of the dataset, so then we do a, like, the last class also in EDA, one, when we are doing exploratory data analysis. So, we'll do some exploratory data analysis to see the data set, how it is, and then we'll be moving to [00:22:16] the classification part, where we do classification. So, whenever we have a tabular dataset, the first thing is generally, in Pandas, what we do is df.info we run. [00:22:28] And then it tells us, [00:22:29] That, uh… what pandas understand of each column. [00:22:34] in its internal understanding, what it is understanding. So it is saying that, yeah, there are 7043 entries, [00:22:39] Yeah, the 7,000 rows are there. [00:22:43] And, uh, it is also giving me a nominal count for each of my columns. So, basically, it is telling me that, as such, there are no panda… according to Panda, there is no such missing values. [00:22:54] in each column, and everyone is, like, filled completely, so there are no missing values according to Pandas. [00:22:59] But the other thing, uh, it is also telling me the data types of each of these. [00:23:04] Uh, like columns. So, uh, if you see here… [00:23:08] The first, uh, uh, difference which… what comes is, [00:23:13] that if I… [00:23:14] map it to my original data information. [00:23:17] There, I was told that there are only, like, 3 categories, uh, like, 3 numeric features, and everything other is categorical. So I… if we just fix our… [00:23:28] point to this senior citizen, they are saying that it is having only two values, 0 and 1, and it should be categorical. [00:23:35] But according to how Pandas has treated this, it is saying senior citizen is an integer, because there are two values, 0 and 1, so it is stating as an integer dead type. [00:23:45] So, that's why we have a look at df.info, and what we will do is… [00:23:51] that all such [00:23:53] All such columns which the fund does is right now treating as numeric features, but they are indeed categorical, because there are only two type of values for senior citizen, 0 and 1. So, we know it's categorical. [00:24:07] So, it is better to convert all of them to objective data type, because it is the one for categorical, it is the data for categorical. [00:24:17] And to keep integer or flow data type only for the actual [00:24:22] data types, where the type is numeric. [00:24:25] So, if you see here, uh, so the first observation I have written is, like, uh… [00:24:31] So, yeah, so, uh, okay, so… [00:24:34] Yeah, so this is the second observation, sorry, the senior citizen one, that it is stored as India 64, but it's a binary categorical feature. So, yeah. So, what is to be done is that if this is this feature, [00:24:46] if we just run this comma, like a line over it, that type of it, you change its type as object. So, [00:24:55] Immediately, like, its type would be changed to categorical, and only, like, this is one of the anomaly. And the other anomaly is it's slightly hard to decipher, uh… no, it's not even try to add. So, basically, we were saying that there are 3… [00:25:11] We had 3 numeric columns. [00:25:13] But, uh, like it was written here, that it needs cleaning, so total charges [00:25:19] has to be numeric, it has to be… it is, like, an amount, so it should be integer, uh, or float. [00:25:25] But how Pandas is treating it? [00:25:27] If you see, total charges is object. [00:25:30] So why… so the question arises why it is treating as an object, like, it is treating these numbers as string. Why? [00:25:37] So, that's why it's, like, the first observation. So, uh, some sequence got altered, so, uh, I first told of second observation first. So, yeah, so first observation. [00:25:46] So, it is treating it as an object instead of numeric. [00:25:49] So, uh, like, it begs to the question, like, uh, maybe, uh, like, some entries are missing, but they are, like, empty string. [00:25:58] So, that's why, when we run over this, and to confirm our doubt, we see that [00:26:03] Uh, like, for the total charges, uh, what happens is, like, for the total charges, uh, if… I have printed 3 columns here, customer ID, tenure, and total charges. [00:26:14] So, you see, these are empty strings. [00:26:18] But these, uh… but because these are string, these are not null values like NAN, so it is saying that there are no null values, but they are actually [00:26:28] missing values, or null values, but since there are, like, blank strings, so it is not treating them as such, and that's why it has stored it like object. [00:26:37] And not as integer. So, this line here, [00:26:41] Just, like, give a filter. [00:26:45] That, uh, for this column, total charges, [00:26:48] Tell me the entries where there's an empty string. [00:26:52] And, uh, it gave me the, like, there are 11 such entries, and then I, uh, printed them, that if the… [00:26:59] entries for this column, total charges, empty string, then print me just these three, uh, like, column, any 3, because there are so many 29 columns, so I did not want it to print all of them, so just 3 of them. So, to see, yeah. So, see that these are, like, empty columns. [00:27:17] So, what we can do is, uh… [00:27:19] We can, uh, like, uh, so… [00:27:21] What we'll do is, we can just drop these 11, like, these are only 11 samples, so it's better to, like, drop them in order to… the simplest way to handle, uh, like, them is to, uh, [00:27:34] get them removed from a dataset. So, originally, there were 7043 entries. [00:27:41] And what we did was, what we are doing now is that Pandas is telling, uh, to this column to convert this to numeric. [00:27:50] We are just trying that, ordinary, its data type is object. We are telling it to convert, it is numeric. Now, [00:27:57] Uh, when we are converting into numeric, because some values are non-numeric, because these are empty strings, so otherwise error will arise, so that's why we have written error equal to course, [00:28:08] coercion, so whatever it will do is, it will remove those rows where there is empty string. That's why [00:28:15] Finally, what we have is, like, this 7032. [00:28:19] entries. Uh, like, 11 have been removed. So, yeah. So, these are, uh, the first two, uh, observations we see from the data. [00:28:29] Now, another, uh, observ… uh, yeah, so these are the first two observations. [00:28:34] And then what we are doing is, at this… now we are at this block, after observation 2, this block… this block was left. [00:28:42] So, what we are doing next is, we are separating for our purpose for later purpose, we are separating all the categorical columns and the numerical column. So, just… this is my data from DF. I am telling it to select all the data frames which are having type of objects, because these will be categorical. [00:28:59] And I'm telling it to select all the variables which are having type integer or float, because these will be numeric values. [00:29:07] So, in total, uh, you see, numeric columns are only 3. [00:29:12] And 18 columns are… [00:29:15] of the categorical type. [00:29:18] And, uh, yeah, so, uh, this thing here. [00:29:22] And other thing, uh, which is to be seen, which is to be seen in the data set, is that [00:29:27] When I… when I try on some, like, some columns, like, there is column called multiple lines, there is column called Online Security, Online Backup, [00:29:37] All of these columns, if I just print the unique values of them, like, uh, like how many… these are all categorical variables, so, but how many unique values are there in them? So, that's why, uh, for… I looped through all of them, [00:29:51] And I just print their value counts means that you print how many values are there for each of the category values that are there. So, you see for multiple line. [00:29:59] There is value no, there is value yes, and there is value no phone service. [00:30:02] For online security, no, yes, and another, no internet service. So, [00:30:07] for all of these, uh, there is a common pattern that there are two types of no's. There is a single, simple no, and then there is a no internet service also no. [00:30:16] So, what we will do is, we will simplify this. [00:30:19] That, uh, because two redundant nodes are there, so… [00:30:23] We will just add that this one, this no and this no will be, uh, the values will be added. So, in total, there will be only two categorical values for all of these variables. [00:30:34] Because, uh, like, for our purpose, like, these are redundant debt. No and no internet service. [00:30:39] So, yeah, so… [00:30:41] Uh, so, in, uh, so here… [00:30:44] So, first, in this block of code, columns to check, I just printed, I did nothing, just printed the unique value counts for… [00:30:51] these particular columns. So… [00:30:54] Then, uh, in the next [00:30:57] block here. Uh, this is, uh, the, like, the phrase that… these are the columns which you'll be seeing. [00:31:04] And, uh, so, uh, basically there are two things that, for this particular column, the phrase which is there is no internet service, and we will replace it by the phrase, no. [00:31:14] So, everything in these columns will be yes or no, but in the column multiple line, it is called no phone service. It, uh, for all these columns, the second entry of no was called no internet service. So, all these things have to be, like, looked in the dataset, and that's why… then only we can, like, do… [00:31:30] This kind of preprocessing. [00:31:32] So, yeah, so, uh, so basic, for this one, multiple line, we are replacing it with no, and for this one, no internet service, which is replacing with no. [00:31:40] And then we merge them together, and then we see that for all these columns, there is, like, only no and yes. [00:31:46] Uh, like, the categories are only no and yes. [00:31:50] So, yeah, so… and after that, [00:31:53] When I… [00:31:55] After that, when I now see my, like, I do this initial pre-processing, [00:32:01] And after, uh… then, when I see, uh, my data, uh, so… [00:32:06] Just for, like, clarity purpose, I will again see my data and will see that, yeah, so this is, like, [00:32:14] Uh, now there are 7, uh… [00:32:17] Yeah, so now there are 703 at 2 inches. [00:32:22] Does that, uh… yes, because… yeah, so yeah, so from 7043 entries, now it is reduced to 7032. [00:32:28] And, yeah, and this month total charges thing has also been sorted out. It is now float. [00:32:34] Uh, this thing, and a categorical thing, we have, uh, like, taken care of. [00:32:40] So yeah, so this is before any visualization also. These are some of the preprocessing that is to be done. [00:32:47] So, any questions still, uh, here? [00:32:50] Till this part. [00:32:57] So, to this step, when you're saying, we are trying to make a data uniform, right? The data types, like, yes and no. [00:33:03] Yeah, yeah, this part is just added to show that, uh, like, there are very much… the data set is always noisy, so we'll have to see, like, inside, like… [00:33:12] doing various techniques, that what all type of things are, like, missing inside the data. [00:33:17] So what will happen if, let's say, in the case when the value is null or blank, [00:33:23] Uh, value is… [00:33:24] So, what are we gonna do? Because as of now, as per this data, it's completely there, right? So, which is the best case. [00:33:28] But there might be possible always the data might be missing, it's empty data. [00:33:34] So, uh, when there is values missing on null or blank during modeling, it will create a problem, like, uh, then that error will come when it will appear in training or sim… training or testing time. [00:33:44] Yeah? [00:33:45] Even if that sample come in training time, error will come. And other… and, like, so basically, this is all is, like, the data is sourced from some distribution. [00:33:55] So, if some value is null or blank, so that means when we are modeling, we are missing some aspect of the data. [00:34:03] So, uh, then our model may not give such good response. Sometimes what happens, these are like, if missing values are very few, then always, uh, the effect is not such. [00:34:13] But if there are, like, missing values or too much, like in some features, uh, like 20-30% values are missing, so what do we do then? So, then it begs to question that either we drop that column altogether, everything, [00:34:26] So yeah, so many… many modeling time, the problem will occur, uh, when we have missing values. [00:34:32] Okay, so are we also, in certain cases, when value is null, replacing with some other variable? [00:34:35] Okay, so I don't… [00:34:38] a value? [00:34:43] Yeah? [00:34:44] Yeah, in the last class, we saw those methods, I guess, for, like, we replaced by the mean, median mode, like, the basic techniques. [00:34:48] And, uh, yeah, we'll have to, like, decide a strategy, like, how do we… what do… do we do for our missing values. [00:34:54] So, like here, our strategy was just simple, that since there are only 11 missing values, we totally dropped them. [00:35:00] Okay. Got it. [00:35:05] Okay, so… [00:35:07] Yeah. [00:35:08] Uh, assist. Uh… So, one, uh… [00:35:16] Yeah, go on. [00:35:37] Mm-hmm. [00:35:38] Probably our last discussion. So, we have one observation where we merge, like, a lot of columns, no internet related, no phone service, to yes and no. [00:35:40] So, how manual this process is, if the data is, uh, like, used. [00:35:54] Yeah, yeah, this… [00:35:55] Any Python function to check the data in such a way, or… Because in JavaScript or something, we do have some rejection expression, right? Where we pick… Sweet. [00:35:56] This step is totally, like, you have to manually do it if the data is too much. Also, like, there is… like, you have to find some automated way, like you told regex expression are there. [00:36:09] But we'll have to manually look inside. So basically, that, uh, the… [00:36:12] a loop is that, uh, usually, uh, people, whatever data they have, they look minimal thing, they look that there are missing values or not, and they build a model over it. If the results are not that good, then they again iterate and see, uh, is our data, uh, like, how much noise is in our data. [00:36:29] So, usually, it is like you have… you do iteration multiple times over data. Like, if you're not getting good results for modeling, then you, again, see, like, [00:36:38] how much noise did I, like, left out in the data? So, it's like that. [00:36:44] I mean, we don't over-engineering the data until we model it and figure out it's not. [00:36:49] being accurate, then we go back and then. [00:36:51] No, but, uh, yeah, so, uh, I'm just telling you the easiest path, like, the easiest part is that you… [00:36:57] do the minimal preprocessing and proceed. Uh, but yeah, but if… but plenty of time is first spent on this part. [00:37:05] If, uh, like, uh, like a, uh, like, uh… [00:37:08] a good, like, a good industrial party for a good industrial skill, you have to make a machine learning framework. So, plenty of time will be spent here. [00:37:16] to, like, uh… [00:37:18] see… manually see each of the features. If the noise is present or not. [00:37:22] So, uh, that thing is always there. [00:37:24] So, I was just telling you, uh, like, an ad hoc-based pipeline that if somebody wants to, like, just see [00:37:30] If it is… if our data is there, and we just want to, like, run a model over it and see, like, are we able to get something or not? [00:37:38] But this thing is always there, that much of the time is spent here. [00:37:42] Instead of modeling. [00:37:47] Yeah, thank you. [00:37:49] Okay, so let us proceed, because, like, the file is slightly bigger, so… [00:37:56] So, we first will have some minimal ED again, and then we'll go to the modeling part. [00:38:02] So, uh, there's, uh, some… this section is on heat maps. [00:38:07] So, uh… maybe this part is not in, uh, like, in theory part, but, uh, just to, uh, it, the… [00:38:15] The formula is not very much required, we just want to see, like, how much the, like, for… like, there is correlation for numeric features, so there is something called Grammar V, which measures association between categorical features. So, I'm not, right now, like, don't want to emphasize on how, like, this formula works and everything, but yeah, it's a… it's a recommended metric which is used to measure categorical feature. [00:38:38] And I wanted to see, like, are the… how are my categorical attributes associated or not, and that, and that for that thing? [00:38:45] I used directly its, uh, like, function. So… [00:38:50] Uh, yeah, uh… [00:38:53] So, yeah, it is here in Maya. [00:38:55] Yeah, so this is, uh, the heat map I'm talking about. [00:38:58] And this is the, uh, like, the last target variable, which, uh, I'm just looking at the last row right now. It's the target variable churn. [00:39:07] And these are all the input variable here. So, customer ID is not, uh, like, if you see customer ID, the first row, it is having association 1, uh, everywhere. So, values are between 0, 2, and 1, it is having high resolution. [00:39:20] But it is like a misnomer, because [00:39:23] each customer ID is a unique value, so that's why it's telling you, uh, like, it's having a higher association. So, but it has to be, uh, it is there in the notebook, but it has to be, like, dropped, it is not required, because customer ID is just an identifier, not much use during modeling. [00:39:39] So, apart from that, we can see only the contract type is having some slightly, uh, like, uh, high association of 0.4. [00:39:48] Uh, with the target variable churn. All other are, like, not, uh, that much, uh, like, [00:39:55] having a great association with that of the target variable. [00:39:58] And, uh, if we… and the other observation is that, uh, [00:40:04] If we see this internet service, [00:40:06] It is like having some good association of, like, 0.45 [00:40:11] 0.44 with other values, which are, like, [00:40:15] 0.45 is, like, phones, like other related services. So, the second point here is writtenness, the internet service. [00:40:22] And other related services, like online security, tech support, these. These are, uh, these inter… they have some intercorrelations with each other, but with target variable churn, [00:40:32] Uh, there is not, uh, like, all other categorical features, they are not that much associated. [00:40:38] So, this is the, uh, overall… [00:40:40] code for it, how it is obtained, this grammar thing. [00:40:45] So, basically, uh, let's… [00:40:47] have some abstraction, like, uh, first, that this is the ultimate… [00:40:53] function which we recall between two variables, if we have two variables, it will compute the association. [00:40:59] For those two variables. Uh, yeah, but to go over them, [00:41:03] First, I, uh, come… [00:41:05] obtain all my categorical columns. [00:41:08] So, first I obtained all my categorical columns, and my make a data frame, [00:41:14] like this, for all my… only my categorical columns, and it would look like something like this initially. [00:41:20] So, yeah, these are only categorical columns, and grammar matrix is this kind of a matrix. [00:41:25] And then I will iterate over [00:41:28] Each pair of column in this matrix. [00:41:31] And then, uh, what I will do is, uh, if it is equal, then the value is 1. Otherwise, I will call this function. [00:41:39] to calculate this, uh… [00:41:41] association value, and it will populate over this matrix, and then I'm printing this matrix. So, I'm telling… I will not go inside the, like, the code for this. [00:41:52] Uh, how it is computed, but I'm, like, telling how this, like, this matrix came and how this is plotted, so that, uh, thing… [00:42:01] I've, like, told her, yeah. So… [00:42:04] So, ultimately, this matrix is filled, and then, uh, if we pass on this matrix to this, uh, this is, like, import C1 as SNS, this is another library. [00:42:14] Which can, uh, which has a function called heat map. If you provided any heat map… if you provided any 2D matrix, [00:42:22] It can generate a heat map of sorts from that. [00:42:24] So, first we got this matrix, [00:42:28] We passed over this heat map over here, [00:42:29] And yeah, then we have this categorical, uh… [00:42:32] variables here. So, one thing more, uh, this warnings thing is written. So, there are basically in calculations, if I remove this, there are various, uh, uh, some warning statements which come. [00:42:43] So, uh, that's why, like, this thing is written, otherwise it's not needed. [00:42:51] But yeah, there are many warnings which are coming during [00:42:54] computation of this, so that's why it was, uh, like, introduced, these two lines. [00:42:59] So, this is how… like, this was given, like, I think in last class, where we did heat map, we only showed for numeric features, so that's why it was included. [00:43:12] That, uh, for categorical features also, there is, like, a way if we can show that how much they are associated with one another, and the… [00:43:19] target variable. [00:43:21] And similarly, [00:43:23] Uh, we can… it is a standard way of, uh, like, for the numerical features, the Pearson Core Richard heat map, which we do. [00:43:31] And in order to compute this, this is my target variable, and because it is in, like, yes or no, so I first map it to, like, yes has been mapped to 1, and no has mapped to 0. [00:43:42] And then, all, uh, like, for numerical columns, uh, this is a list which contains my, uh, 3 numerical columns. I added to this, uh, target variable churn. [00:43:52] And then I do, uh, this, uh, for these, only these four columns, I do this .corrr. So, this correlation, this is, like, [00:44:02] given in pandas only, so much, uh, so, uh, so much less calculation is here. And, and like for categorical variable, the correlation thing, that was not given in Pandas, so… [00:44:13] For each column, whatever is the correlation value has to become. We had to make a separate function. [00:44:18] And this is, like, I found from, like, the internet, how to do this, like, compute this grammar function. [00:44:26] grammar for two categorical variables. So this is already given here, so that's why I just had to just do CORR over it. [00:44:33] could do a correlation over it, and it computes a matrix for all the numerical features. And similarly, once I have a matrix, [00:44:42] correlation matrix for all my numerical features. [00:44:45] Then I pass it to my heat map. [00:44:48] And it computes me some heatmap like this. [00:44:51] And from this, like, we can again see that [00:44:55] Uh, tenure is, like, if you see tenure, it is like… [00:44:59] positively correlated with total charges. So, there are input features also, which are correlated with each other. [00:45:07] And uh… but with that output variable, we see that [00:45:12] the strongest relationship is, again, like, with tenure, but they are negatively correlated. It's… it's less than 0.5, minus 0.35 it is there, negative correlation. [00:45:22] So yeah. So, these two parts is where, like, a minimal idea. We could have plotted, uh, we could have done more EDA here, but yeah, so… [00:45:32] There's a basic idea we have done for categorical and numeric features. [00:45:37] So, okay, so after that, uh, before modeling, one more thing we are doing is that we have these three. [00:45:45] Uh, numeric features with us. [00:45:46] And what we'll be doing is, uh, uh, in order to do learning for our [00:45:52] vision tree and random forest. We will… what we will want is, we will want that all of our features [00:45:58] All our features are categorical, because decision tree and everything, like, work best when everything is categorical. [00:46:03] So, but these are, like, continuous numbers, so we'll convert them into categories. So, how we… so basically, that's why the title is, that we will want to discretize [00:46:14] Numeric attributes. So, these are 3 numeric attributes. [00:46:18] Which have to be discretized. And by discretized, I mean that we will split the dataset into some bins. So, basically, if I want that dataset to be split into [00:46:29] like 3 bins, so each value will, uh, like, each value will go either in bin 0, 1, and 2. [00:46:36] So, the strategy of how it will be due is, like, you can… we can… we are using a quantile strategy. [00:46:43] So, uh, the data will be sorted. [00:46:46] will be sorted, and if there are four quantile, so each quantile, each bin or quantile will have 25% of the data. [00:46:55] So, if your data is in the first 25% of the, uh, like, the sequence, so it will go to the bin 0. If it is in the next 25, it will go to the bin 1. That's… that's how it is being, uh, will be mapped. [00:47:10] So, first, let's, uh… [00:47:13] Let's start this, uh… [00:47:14] So, okay, so we'll start this. So, uh, I hope the correlation thing is okay with everyone, the code thing. [00:47:20] The code, uh, for coordination is okay? [00:47:24] I have one question, like… [00:47:25] Yeah, yeah, Deepa. [00:47:26] Uh, so what's… why, uh, what's the difference between the previous heat map and this heat map, and why we are not able to generate that heat map via Python itself? Why we have to use a different function altogether? [00:47:37] Okay, so, uh, for categorical features, [00:47:41] Uh, this Pearson correlation for function, the mathematical way it is defined, it does not work for… it works for only for numeric features. [00:47:49] So, that's why, for categorical features, where there are, like, more than two categories and everything, so that's why we have to have a separate function. [00:47:59] So, it's basically the thing which we are doing, the Pearson function which we are using for obtaining correlation. [00:48:06] It works only for numeric features, that's why. [00:48:13] Okay, uh, for discretizing numeric attributes, we are using this quantile strategy. Uh, can you just elaborate on that? [00:48:19] So, basically, I'm saying this, suppose I have, like, 0 to 100 is my numbers, and these are… let's support sort them out. So, 0, 1, 2, 0, 1, 2, 3, like, there are still 100 numbers. [00:48:30] So, first, 0 to 25… 0 to 24 numbers will go to the bin 0. [00:48:35] Next 25 will go to the bin 1. [00:48:37] And next, uh, we'll go to the bin 3. [00:48:41] This is how we are discretizing them. [00:48:45] Is it clear? [00:48:51] Oh. [00:48:52] Yeah, yeah, yeah, yeah, same thing. Just a name thing, uh, how they are naming it, they are calling it… [00:48:57] EQ frequencing, meaning… [00:49:01] Okay, binge. [00:49:03] Okay, so let's start with, uh, like, the discretizing thing. [00:49:07] So, yeah, so, again, this thing is, like, done multiple times, uh, but, uh, yeah. So, uh, be… [00:49:16] the churn variable, uh, it is… [00:49:19] Like, again, what I'm doing is, like, [00:49:24] I am resetting it for, like, and also because everything needs to be converted into a category, but for now and also, I'm just, like, again reversing it for now, because for correlation, I did it 1 and 0. [00:49:36] So, I'm reversing into yes and no. [00:49:37] So, map 1 to yes and 0 to no. [00:49:40] So, and apart from that, and here I'm just printing these three values, that these are right now in numeric features. [00:49:46] So, yeah, so now I will introduce the discretization of them. [00:49:52] So, this function is there in SQL and preprocessing, k-bin's discitizer, so it will create bins, uh, k-bins, how many bins you will, uh, specify that many bins. [00:50:03] So, first, we… when we call it, we have to create an object of it. [00:50:09] So, cabin… so this is the object name, disk tenure. So, I'm… [00:50:14] telling it cables discretizer, calling it, and in constructor, I'm giving it that I want 4 bins. [00:50:20] And encode is a… in which way I want to encode, in ordinal, basically categorical way I want to encode. [00:50:27] Uh, in categories. So, with this, they'd know that it has to store it in an object, like, type, data type. [00:50:36] And strategy is quantile. So, quantile strategy is there, so you'll have 4 quantiles of your data. [00:50:42] And uh… and this object here… [00:50:45] I'm, uh, like, then what I'm doing is, I created this object, [00:50:50] And then I'm calling fit underscore transform over it. [00:50:54] So, what it will do is… [00:50:56] Uh, it will take this column, I'm giving it 10-year column, [00:51:01] It will, uh, first, like, do whatever, it will sort, uh, it will sort it out in, uh, like, in its ascending order, divided into four… find four quantiles. [00:51:10] And, uh, this is the new column, and it will… for each value, it will tell whether it has been 0, 1, 2, and 3. Like this. [00:51:21] And, uh, we can… could use the same variable again, but just for, like, not any confusion and everything. [00:51:28] Uh, because, uh, yeah, so basically, no, we could not use same variable, because we have defined that how it, uh, its fit and transform will work. [00:51:39] So, for the next column, I'm again defining a new discretizer. [00:51:43] So, uh, and it will do a, uh, and it will do a fitting and transform over this monthly charges, the next column. So, why we are not using the same column is, [00:51:51] Because this one, this variable, it is learning, uh… [00:51:56] that what will be the quantiles based on this column. So, it has learned this. Now, again, we can't call it for the next column. [00:52:03] So, uh, that's why we are doing this. [00:52:07] And, uh, uh, yeah, so… [00:52:10] Here, I've initialized it again, [00:52:13] And uh… and doing fit transform over monthly charges, and similarly, [00:52:17] I'm doing fit transform over total charges. And whatever will be the output, I am creating a new column. [00:52:23] a name which will store them. [00:52:25] So, that's why, when I… [00:52:27] Finally, run them. So, these are the new columns. [00:52:31] Uh, like, uh, earlier they were numbers. [00:52:35] 10 or 134M anything. Now, there are, like, either in some bins, and bins are, like, 0 to 3. [00:52:43] bins are there. So, yeah. So, this is, uh, how we're doing it. But, uh, right now, if you now see… [00:52:52] If the data… in the data set. [00:52:56] a guy print. So it was… I think initially it was calling 21 columns, now it is called 24 column, because [00:53:02] The initial columns are also there, the numeric ones, and the discretized versions are also there. [00:53:07] So, if you see here, uh, the tenure VIN, monthly charges, these are also there. [00:53:13] And normal total charges are also there. [00:53:15] So, because we have made everything categorical, and we'll be working with categorical variables only, [00:53:21] So, we are planning to dropping the numeric attributes altogether before giving into the, like, training everything. [00:53:28] So, that's why, to do this, I, uh, just create a copy of my… [00:53:32] Original dataset, because, like, there are too many people happening here, so I created a new data frame. [00:53:39] And, uh, yeah, and now, before work… by proceeding, because we will proceed to modeling and everything. So, modeling… [00:53:48] could, uh, like, the immediate thing which we do is, [00:53:53] We drop this customer ID. So, see, in modeling, customer ID is of no use, it's just a unique identifier. So, in order to… let's not have more confusion, because maybe some wrong associations will be learned during modeling because of this cut feature. [00:54:05] So, it's best to drop it at the earliest, so… [00:54:09] Uh, could have dropped it, uh, so we are dropping it here. [00:54:12] So, we drop the customer ID column. [00:54:16] And also, we drop these numeric features. [00:54:18] Which are, like, because we have all… already they have discretized version now, so we drop them. [00:54:24] And finally, for this… [00:54:28] 10-year mean monthly charges and total charges bin, uh… [00:54:31] attributes. I am explicitly, like, specifying that make their data type as object, because I want to treat them as categorical. [00:54:40] So, yeah, so after that… [00:54:42] It, uh, looks like this. [00:54:45] Yeah, because of, uh, I will just do a full run of everything, because, like, I have, uh, the sequence I have, uh, slightly… [00:54:54] modified, so some NAN is coming right now, but this is, like, if I run it over sequentially again, so this will be… this is not a problem. This will, again, see yes and no only. [00:55:04] But you see, everything else is, like, now categorical. [00:55:07] And uh… like, let me, like, run… [00:55:13] do a run from initially, because… [00:55:16] Uh, otherwise, it will be confusing to everyone. Why NEN is coming there. [00:55:22] So… [00:55:31] Okay, so, uh… [00:55:35] Hmm. Okay, so, uh, [00:55:39] this thing, maybe it is, like, it has been passed over here. Just see. [00:55:44] Yeah, so you see, now, if I… because some… the way I was, like, up and down, I was… I was running individual code, so it ran into some bug over there. [00:55:55] But now, you see that only Noah and yes are there, and everything is, like, this… I was at this part. [00:56:01] Replacing numeric attributes. So, at this part, everything is now categorical, and this can be confirmed from here also. [00:56:09] In the columns, everything is being treated an object data type, and these are 7032 values. [00:56:16] And last, just the last thing before moving forward to the modeling part, [00:56:20] Uh, because it took… it's taking… So, the last thing is just we want to plot and see how… some basic plot to see the distribution of the categorical, because everything is categorical. Let's see, like, the bar plot distribution of everything. [00:56:33] So, this is the churn target variable. [00:56:35] And, uh, I'm just plotting, uh, here, I'm just doing a value, uh, so basically, this, again, the value counts. [00:56:42] So, value counts will give me these numbers. [00:56:45] And then I'm also, like, plotting their bar graph here. [00:56:49] So, this is value count. This first line. [00:56:52] When I run this, this gives me this thing here, and then I, uh, in pandas only, there is, like, a way in the function, they have specified. [00:57:02] That if you do value counts, and after that, uh, this is for multiplying by 100, because these were, like, in decimal, and after that, if you do dot plot, [00:57:11] And if you mention kind is bar, uh, bar, so it will give a vertical bar plot. If I do bar H, [00:57:17] It will… I think it will give me a horizontal bar plot. So yeah, these things are mentioned in the library in more detail. But yeah, so it tells us [00:57:25] that the target variable is balanced. So, 73% of the data. [00:57:32] So, like, zero, housing value zero, and… [00:57:38] No, the user has not churned, and only 26% of the time that it has turned, so yeah. [00:57:43] So… so that thing, we have to take care of. So, when we make a model, we'll have to see that, uh… [00:57:50] are more… that, uh… [00:57:52] since it is an imbalanced data, [00:57:55] So, if our model is, like, predicting only 50% accuracy and everything, it is not good, because if we take any dumb data also, [00:58:04] And ask it to only give answer as 0, and do nothing, and give answer as 0, then also it can achieve 73%. [00:58:11] of accuracy, because how the data is imbalanced. [00:58:15] So, yeah, and it also gives us, uh, insight that [00:58:19] Since the data is imbalanced, so there will be more challenging to train it. [00:58:25] Because, uh, samples for one are very less. [00:58:30] Very less samples are there in order to learn when the user will churn, but when the user will not churn, there are so many samples. [00:58:37] So, that thing is, like, VC. [00:58:40] So yeah, so now we move to… [00:58:43] preparation of the actual features, X and Y, because in order to give for this classification problem, we have to create an X, the input, and the Y, which is the expected output. [00:58:53] So, that thing we will do here. [00:58:56] And, uh, any questions till now? The ADA part is totally completed. We'll… [00:59:02] move to the modeling part. [00:59:03] This is, uh, please go slow, we are going so fast. Please go slow, yes? [00:59:04] Okay, okay, okay, sorry, sorry, sorry, sorry, sorry. [00:59:09] Actually, just one question. [00:59:12] Kate, in turn, if we had 50-50% data, [00:59:15] Hmm. [00:59:16] what impact it had. [00:59:18] Yeah, [00:59:19] That is also a good thing. No, or no? [00:59:20] If we had 50-50% data, [00:59:22] And, uh, so basically, we have an equal distribution. We would have an equal opportunity to learn when the model will return and when the model will not churn. [00:59:32] Correct. [00:59:33] Now, the problem is that we have, like, so much information when the model will churn. [00:59:37] But very much less information when it will not. So, the… we can't blame the classifier itself, that it is giving us less bad performance. That's, uh, that we will have to… that we… information we obtain from in, like, this label distribution. [00:59:52] So, what I was saying was that, uh, since [00:59:55] If it was, like, 50-50% was there, and even if I had a 60%, uh, like, performance on my test set, I would say that model has learned something. [01:00:06] But in this case, suppose my accuracy is only, uh, like, my accuracy is, uh… [01:00:12] 73%. So, I am saying it, that why to even develop a model? Just say, uh, just types 0, like, no for all the samples you see. So, 73% time, you will be correct. [01:00:26] Okay. [01:00:27] So, that thing is to be, like, a scene after seeing that there is distribution in the… [01:00:32] There is imbalance in the distribution. [01:00:34] So yeah, so… uh, yeah, so, uh, apologies for being fast, slightly fast, but, uh, yeah, so now I'll be more slow, uh, but all this part was EDA, and it was, uh, so that's why it's fluff. Now I'll be very slow. [01:00:47] Because the… this thing is, like, new, which is not covered in previous classes. [01:00:52] So yeah, okay. Anything else before we start? [01:00:56] Any more doubts? [01:01:02] So that, uh, binning you explained, right? So, they are, like, uh, the method, uh, itself will do it intelligently based on the. [01:01:11] Yeah, yeah. [01:01:12] Whatever the parameters given, uh… [01:01:13] Yeah, yeah, yeah. Yeah, so that's starting, I said that, uh, otherwise, we will have to write, like, function for each of the things. So, that's why sklearn [01:01:23] For all the things related to ML, the pre-processing, the training time, the modeling time, every thing, like, every imaginable thing, every possible thing which is there in the, like, what we do, [01:01:35] They have functions for that. So that's why we don't have to… we just… that's why we are calling directly the function. [01:01:45] Okay, so let's start with how, uh, first we'll create, uh, the input and target for the model. [01:01:52] And, uh, so to do that… [01:01:55] Uh, so we know that our, uh, target column is churned. [01:01:59] Uh, which is yes or no, and we are, again, first we are converting into [01:02:04] So, this is my target column. [01:02:06] And uh… so basically, initially it is string. So, basically, this strip function is to, like, [01:02:13] strip away any spaces are there, because in string, sometimes, there is a yes. [01:02:17] But there is a space after years. But when you see print it, the space is not directly visible, so that's why. [01:02:24] We are saying that in this column, [01:02:26] strip all the space, left, right, whatever is their space, and just that string. [01:02:31] And then, uh, map yes to 1. [01:02:35] And, uh, map node 0 to 0. [01:02:39] Yeah, so first we'll, again, like, we are doing 1 and 0 mapping. [01:02:42] And then, what I'm doing is, I'm saying that everything except the column churn [01:02:48] is my X right now. So, everything… because I've already dropped the ID column, the customer ID column, I've dropped previously. [01:02:56] So, this is… [01:02:58] DF2 right now. [01:03:01] So, right now, this is DF2, 20 columns, uh, ID is dropped. [01:03:05] So, basically, we will remove the churn column. [01:03:10] So, 19 columns will take particip- will become our X. [01:03:12] And one final column, churn, will be our Y. [01:03:15] So, that's why, like, it's written DF2.drop. [01:03:20] drop this column, uh, that's why X is equal to 1, uh, on column side. [01:03:24] And just take this column, that will be a Y. [01:03:28] So, this thing is done in this part. [01:03:32] And, uh, then what we do is, because all of these are, uh, like, categorical, [01:03:38] You must have, like, uh, like, heard the one-hot encoding. So, what we'll do for all these [01:03:44] input features. We are, uh, doing a one-hot encoding of them. [01:03:50] For that, we have used here, uh, like, in… [01:03:53] Pandas only, the function is getDummies, and uh… [01:03:59] And you give it your, uh, data frame, which is X, the input features. [01:04:02] And there is one thing mentioned here, drop first, uh, I will show you it. I will show it to you. [01:04:08] Uh, but yeah, but this is 2. [01:04:11] one-hot encode all the… [01:04:13] categorical features. So… [01:04:18] Again, maybe the problem is by running it twice, it sometimes goes into error. [01:04:26] Boom. [01:04:27] Okay, so it will take 1-2 minute just to do it. [01:04:53] Okay, so, uh… [01:04:57] Okay, so, uh, yeah, I think that's why I was saying that after that, X and Y, [01:05:03] I will do a… I will one-hot encode over my category. Basically, every variable of mine is categorical right now, so… [01:05:08] I will do one-hot encoding over them, and I will go at 1x encode it. [01:05:14] So, basically… so, again, if we see X… [01:05:18] Just first, we'll see X, so there are, like I said, [01:05:21] In DF2, there were 20 columns. [01:05:23] In order to make X, I just dropped one churn column, so that's why I have 19 columns. [01:05:30] So, there is no… the churned variable has been, like, is not there. [01:05:35] So, yeah. So… [01:05:37] Yeah, so in DF2, there is churn variable, and it is not present here, because now the count is 19. [01:05:46] And then, I am saying that what I will do is one hot encode. So, basically, uh… [01:05:53] Right now, uh, everything is, like, for a… [01:05:57] all the categorical features. [01:05:58] There is only one single column. [01:06:01] I hope you are all aware of, like, one hot encoding, so… [01:06:05] So, basically, right now, for all categorical features, there is only one column. [01:06:10] So, if a categorical feature has three categories, [01:06:15] We will make E1 column for… we will make one binary column for each of its category. [01:06:21] Uh, that's how we do one-hot encoding. [01:06:23] So, if we see now… [01:06:26] X encoded. [01:06:31] Okay, so you see, there's X encoded. [01:06:34] And it is written, uh, gender… so, uh, let me… [01:06:40] Just see for, like, some value we'll see, then it will be more clear. So, in the original X, [01:06:45] You see, in internet, there were 3 categories. [01:06:48] So, what… one hot encoding does is, it makes one separate column for [01:06:53] it makes one separate binary column for each of these categories, that internet services, fiber optic. [01:06:58] And one on all the places where there is fiber optic is one, and rest place is 0. [01:07:04] then it will do similarly, internet service equal to DSL. [01:07:06] one for all the places where service is DSL, and one, and similarly, one for all the places where there's no. [01:07:12] So, basically, 3. [01:07:15] Columns are created, and everyone, every column is binary column. [01:07:21] But what we do is, uh, like the third call, if [01:07:24] Both of these columns are false. [01:07:27] So, obviously, this will be true. So that's why, uh, for… [01:07:31] encoding and everything, what we do is, we see drop first equal to true. [01:07:36] Because if a category has n columns, n unique values are there in a category, so we don't want n features out of it. N-1 will do, because if n-1 features are false, it automatically means n-1 values are false, it automatically means [01:07:52] The init value will be true. [01:07:54] So, that's why it dropped first equal to true was written. [01:07:58] Otherwise, what it will do is… [01:08:01] It will… even for, like, a gender can be male and female, it will make a column for gender male, where [01:08:09] on… everything will be true when gender is male, and false, and similarly, it will make a column for genderic female. [01:08:15] But right now, it will make, uh… [01:08:19] one column less. So, if there are two values, it will just make one column. If there are three values, it will make two columns. [01:08:25] So, is this thing clear? Like, we can… [01:08:30] If any confusion on this thing, we can discuss more on this thing. [01:08:37] When there are multiple columns, and it is, you know, generating this n-1. So, let's say in case of, you know, the fiber optic, uh, the internet thing. [01:08:45] Yeah. Hmm. [01:08:46] We have 3 there, and for, you know, the, um… drop first equal to true, it will just have the two columns, right? [01:08:53] Yeah, yeah. [01:08:55] How are we going to infer what is available? Let's say DSL and, you know, no. So if DSL is 1 and no is, you know. [01:09:01] So, if a… so… [01:09:06] Okay. So basically, what, uh, after one-out encoding, what you're getting is, [01:09:07] Because DSL is already 1. [01:09:11] Every column is a binary column. It will be the true or false. [01:09:16] Okay. [01:09:17] So, if both fiber optic and DSL are false, so it is automatically inferred that this thing will be true. Like, we don't need to have a separate column for it. [01:09:26] That's what we are saying. [01:09:28] Like, if, uh, if it's a visible layer, I will show… [01:09:33] So… [01:09:36] monthly charges been. So… [01:09:38] what is the column name? Internet service. [01:09:41] So, basically, yeah. So, it is… [01:09:44] internet service, no. [01:09:48] Internet service fiber optic. [01:09:50] And I think DSL, it has dropped the DSL thing. [01:09:54] So yeah. So, there are two columns, internet service, fiber optic, and internet service, no. [01:10:00] Mm-hmm. Okay. [01:10:01] When both of them are false, that automatically means that the thing will be… that thing will be true. [01:10:05] DSL is true. [01:10:06] So, uh, that's why we don't need to separately encode it. [01:10:10] Like, basically, the more… actually, the model don't have the knowledge that these are the column names. It will have numbers only. [01:10:19] Okay. [01:10:20] So, it will itself infer that if these things are zero, so these, to be not considered. So that… that's what we are saying. Basically, [01:10:27] Having an explicit new column, yeah, okay. [01:10:28] Okay, okay, makes sense. Okay, we don't have to, you know, loop all the columns, inferring, you know, n-1 column is enough. [01:10:36] Yeah, yeah, yeah. [01:10:39] Yeah, yeah, yeah. [01:10:40] importing from n-1 column is, you know. Okay, yeah. Thank you. [01:10:43] Yeah, one question. So, what is the need of a changing or converting categorical data into binary? [01:10:48] Categorical data into binary, uh, uh, [01:10:49] Like… [01:10:52] how… so basically, [01:10:55] These machine learning models, like, they all, uh, they work only on, like, numbers. [01:11:02] numbers. So, when I convert it into binary, I will, like, convert them to 0 and 1, so… [01:11:07] That's why I've converted everything into num… like, the other way was to… that, uh, in order to convert them to numbers, [01:11:15] If I had 3 categories, I will mark everyone as, uh, like, if I have 3 categories, mark them as 0, 1, 2, 3, like I was doing earlier. [01:11:22] But it gives a false kind of, like, indexing, like a false order, that DSL is one, [01:11:29] fiber optic, uh, like, if I had these three categories, if I had given them 0, 1, and 2, so model might interpret that there is some kind of [01:11:38] like, sequence to them. So that's why the… [01:11:41] we follow the one-hot encoding. Everything becomes a separate column, so that's why this false sense does not come. [01:11:51] Is that the mandatory step in all the modeling? [01:11:52] Yeah, yeah, in every modeling for, uh, you have to… in order to deal with categorical features, because numeric features are numbers, model can directly understand. [01:12:01] Now, categorical features are categories. How to, uh… even if the categorical features are… initially, it is seeming to you like number, suppose there are four categories, [01:12:10] Suppose there are three, uh, like, two categories, 0 and 1. 0 means male, 1 means female. But, uh, these are categories just being seen as numbers. [01:12:22] So, that's why we have to, like, tell the more… like, encode them separately. [01:12:27] Like, otherwise we can't do. [01:12:33] So, uh, we don't, uh, uh, take the discreteized columns, right, over here. [01:12:39] discretized, I… I first discretized, I had done first two, like, bring everything into categorical features. [01:12:48] Mm-hmm. [01:12:49] everything in category features. But ultimately, when it will be passed to the, uh, like, the model, [01:12:54] It will pass it into… in some form of numbers. Now, the problem is… [01:12:59] Like, we are… in order to pass categorical features into number, there are very, uh, like, there are two main steps. Either do you do ordinal encoding, or either you do one-hot encoding. [01:13:11] Yes. [01:13:12] So, I have shown one hot encoding, how to do one hot encoding. [01:13:15] Yeah, yeah, the question was, like, one hot encoding, right? It has to be done only on the categorical features, right? [01:13:23] The discretized features which we have, right, we don't apply one-hot encoding on them. Is that correct to say? [01:13:32] No, no, even the, uh, discretized features. [01:13:34] Uh, uh, where were the discretized features? So… [01:13:39] Just a minute. So, okay, so you're saying they were already in, like, 0 and 1 state? [01:13:44] Yes, yes. [01:13:45] Yeah, yeah, yeah, yeah, yeah. It is redundant for them right now, yeah, yeah. [01:13:53] No, no, uh, okay, so, uh, Deepak lets me just come again. [01:13:54] Okay, got it too. [01:13:58] So, basically, initially, there was only… even in the discretized column, there was only one column. [01:14:03] been this column is tenure bin. So, it is having value 0, 1, 2, and 3. These are categories. [01:14:10] But when we will do its one-hot encoding for, uh, if there are three categories, I will make separate column for each category, tenure [01:14:16] been equal to zero, one column, tenure been equal to 1, uh, one column. [01:14:21] And entries will be binary. So, yeah, so for them also, you will have to… [01:14:25] like, do one-hot encoding. [01:14:27] Oh, okay. [01:14:30] Okay, but I want that, uh, increase the overall number of features, uh… [01:14:34] Yeah, it increases the overall number of features. That thing is to be, like, right now it is not giving that too much kind of, uh, too much features. [01:14:42] But that's why, uh, like, there are two ways. Either you directly work on, like, these numbers. These are also numbers. It is kind of ordinal encoding. [01:14:51] that these are converted into encode, like, this is one category, 0, this is one category 2. [01:14:56] But, uh, in a sense, it is, uh, like, model can wrongly assume that, uh, [01:15:02] Uh, this is kind of some weight. [01:15:03] type of category. That's why we converted to 100 encoding. But yeah, it leads to a, like, a great increase in features. [01:15:11] Uh, that thing is always there. So… [01:15:18] Okay, so… [01:15:21] Yeah, so, you know, we have, like, one hot encoded every… [01:15:25] features, and uh… [01:15:28] Yeah, if we, uh, like, uh, what… one more thing could have been done is, like, uh, like, I am modeling for all the features all together. [01:15:36] So, if, like, categorical features you don't want to do, like, encode, because everything, uh, in… while giving to the model, everything will be into numbers in the form. [01:15:45] So, just take the numeric features, [01:15:47] and make the model on that. So, on that, on that time, because these are numeric features, you will not have to do any, like, encoding of the one-hot encoding kind of thing. [01:15:56] So, as rightly pointed out by, like, Yudipag, so in column have increased. Now, there are 29 columns. [01:16:04] So, because it has created multiple columns for same feature. [01:16:09] So, yeah, so this thing is there. [01:16:11] So, at this point, X encoded. [01:16:16] I have my, uh, like, the call… like, these are the total columns I will be working on. [01:16:21] As you see here, uh, initially, I think there were 19 columns, not much columns have been increased, only 10 columns have increased. [01:16:29] But this is the… will be my… [01:16:31] total input dataset, which I will give for modeling. [01:16:36] So, uh, this is the step. Like, I will be finished with all the data preprocessing. I will proceed to splitting of them. [01:16:43] And, uh, yeah, so at this step… [01:16:46] Everything is, like, suitable for, uh, now I will just have to do a train-test split, but yeah, till here, I was… [01:16:52] making it in a form that will be suitable for model. [01:16:58] So, yeah, so let's begin with… [01:17:00] splitting the dataset. [01:17:03] So, if, uh, so basically, this 7,000 samples, uh, there was no separate test set. [01:17:09] or any set of… usually, if you… usually that is the case, we have only one set. [01:17:15] But in order to see how much the model learned, we usually give a separate test set and all. [01:17:21] And in order to obtain it, [01:17:23] Uh, like, we can write a separate function to, like, take 10% or 20% of our dataset. [01:17:28] So, similarly, SKLN.modelselection, there is a, like, this function called trend test split. [01:17:35] In which, uh, you provide your… [01:17:38] Uh, basically, X, [01:17:40] This is my X, and this is my, like, the Y. [01:17:42] So, basically, let me just… [01:17:45] print it. [01:17:47] Like, what is the shape of… [01:17:49] Like, how much is there? [01:17:55] Okay, so… [01:17:57] So basically, if you see, this is the input dataset, 7,000 samples. [01:18:01] and 29 features, and why, uh, why is the churn value for 0 or 1, like, for those 7,000 samples? [01:18:09] So, I will create a held-out… I will create a train set on which I will train the model. [01:18:16] And then, in order to see how good the performance is, I will keep a separate test set [01:18:23] and see, uh, like, how much it is performing in that, because it is the dataset on which it was not trained. [01:18:28] So, that's why I used my original dataset. [01:18:31] I told it that my… I will initially, for starting, I will take only 10% of the dataset as test [01:18:37] dataset. And this is just a, like, a random state. It could be different. So, basically, [01:18:44] from my 7,000 samples, [01:18:45] It is deciding that, uh, based on our, uh, randomly it is deciding that some part of it will be [01:18:52] 10% of it will be tested. [01:18:54] So, that randomness will differ if I change this. [01:18:58] So, since I've used randomState 42, [01:19:00] So, uh, if I now… every time when I use the same number, even you will use the same number. So… [01:19:07] the samples which will be in our trained and tested will be same. [01:19:12] So, that's why this random state is there, to reproduce the randomness. [01:19:16] And then, there is a… this value called stratify. [01:19:21] And, uh, it is because we have seen that, uh, that train and test were, like, 70… [01:19:25] 3% was one value, and… [01:19:28] And the remaining were the value, so there was imbalance. I was… I'm talking about this imbalance. [01:19:34] So, is 25 equal to Y tail? [01:19:37] that when you are splitting into train and test, [01:19:40] maintain this proportion. [01:19:42] in the train and test. It should not be like in your… [01:19:45] train set, or your test set, this distribution differs, because then it will, like, change. [01:19:51] So, that's why… [01:19:53] This, uh… [01:19:56] Okay, okay, so this stratify equal to Y is, uh, there. [01:20:02] Wait a minute. [01:20:06] Yeah, so that's why this try equal to Y is there. [01:20:10] And, uh, yeah, we can see now… [01:20:13] Uh, because 10% of the dataset is there, so 10% has gone to train… will form the train set, and [01:20:19] At, uh, 10% will form the test set, and remaining is the training set. [01:20:25] So, this is how you are splitting. Any question on this thing to interrupt? [01:20:33] Yeah, got it. [01:20:37] Initially, we told that, you know, the 70-30… 73 and 26% split is not ideal, right? [01:20:44] Um, then, I mean, did we do anything, you know, to equalize that sample? Like, 50%? [01:20:47] No, no, we will not equalize it. Like, then we are, like, interfering with the distribution of that. [01:20:54] It is just to, like, it is just to give an observation for us. We will not… we don't change that thing right now. [01:21:02] There are… [01:21:03] But it has, you know, the… sorry, go ahead. [01:21:05] Huh, yeah, yeah, there are… there are ways that you can do oversampling and… [01:21:06] Good, good. [01:21:10] increase more of the data. [01:21:11] But right now, we don't change it. [01:21:14] So, that thing is separate, uh, separate, uh, thing that… [01:21:20] But how are we, you know, mitigating, you know, the risk of. [01:21:22] model, uh… if the model is, you know, not correct, how will we, you know, judge that out of. [01:21:29] You know, the 70 26, you know, sample and all. [01:21:30] Yeah, so… mhm. [01:21:31] Because, I mean, as you told. There is a chance that if the model is not performing at all, then also, you know, it will… it will be treated as correct. [01:21:39] So, yeah, so that… [01:21:40] Right? Because we have more number of. Mm-hmm, go ahead. [01:21:41] Hmm, you are correct, that thing comes separate, that if you want to, before modeling also, if you want to, [01:21:47] But, see, sometimes you can mitigate it, but sometimes that… there is no chance that you get more data. [01:21:53] The dataset itself is balanced. So, like, in churn type of problem, [01:21:57] Like, in industrial setting. So, obviously, one value will be higher than others. So, this thing is like, uh, like, like, natural. [01:22:05] Then, either your model has to, uh, is capable of dealing with the imbalance distribution, [01:22:11] Or you, like, during sample, increase the samples of, uh, like, you create, like, [01:22:17] oversample the one which is less, you make multiple copies of it, like, that thing would happen. But right now, we are not, uh… like, we don't… basically, we want to see… the target was to [01:22:30] tell the class… how classification happens. [01:22:32] I'd like to improve the performance, that thing is, like, will be, like, separately done. [01:22:38] Okay. Okay, makes sense. Thank you. [01:22:41] Yeah, so this is train test bit. [01:22:45] And, uh, after that, we obtained this. So now, we're coming to the first part, modeling the decision tree. [01:22:51] And uh… and if, uh, like, decision tree, the main, uh, so basically, now, there are some hyperparameters which are to be obtained, like, how do we know, uh, like, [01:23:02] we call them hyperparameter because we have to tell… we have to set them. [01:23:08] Like, how much their value will be, so that we get a better performance. So, in… [01:23:12] Uh, like, in decision tree, the main, uh, hyperparameter is, like, uh, [01:23:17] How do you select the optimal depth? [01:23:19] Like, you want to, uh, like, you want to… how much depth should be there of the tree, like 5, 10, or, like, there are 29 features, or do you want to go till 29 depth? [01:23:30] So, uh, sort of… so, basically, if I see the model, uh, is only this part. [01:23:37] But in order to get that how much is the depth of it, [01:23:41] So, that's why we have to do, like, experiment, because we have to first find out what will be the optimal depth. [01:23:48] So, what we are doing is, uh, I have splitted my data into 90 and 10. [01:23:52] And I'm using the same data. [01:23:55] Uh, so let's see the code of here. So, this is, uh, matplotlib for plotting, and Skln.tree for decision tree classifier. [01:24:04] And accuracy also, there is a direct function to calculate accuracy and all. So, SKL and dot matrix, it is there. [01:24:10] So, as you see, there are 29 features, so I looped through [01:24:14] through all the depth, 1 to 20, uh, 1 to 29. [01:24:18] So, 30 here means, like, it is not to… it included. It will go till 29. [01:24:24] And… [01:24:26] Then I'm looping through all my depth, [01:24:28] And I'm creating a decision tree. [01:24:30] Nah, this is, like, the decision tree classifier, which is mentioned in the skill-on. [01:24:35] So, uh, so criteria is entropy, I'm using criteria entropy. The other thing is you use the criteria Guinea. [01:24:42] And then, I am using the depth. [01:24:45] Uh, like, first it will go for Depth 1, then depth 2, like this, it will loop, and then, again, a similar thing, random state. [01:24:52] So, if this will be different, uh, results will come slightly different, and it will not reproduce with [01:24:58] what the results I'm seeing here, your results will come slightly different with different random state. [01:25:03] So yeah, so this is the place where model is defined. [01:25:08] And after that, uh, I have, like, instantiated the model. [01:25:13] then what I will do is, I will call a, uh… then I will call a fit over it. [01:25:18] Fit basically here means that I am… [01:25:20] training the legitree, and by decision tree training, mean that it will learn that [01:25:26] At which value to split, and what, like, what will be my first root node, and then what will be my second node? [01:25:32] Uh, on which I will split in everything. [01:25:33] So, this is the training part which will happen. So, you see that we don't have to write the separate loop for training and everything, because this library handles it. [01:25:43] So, I'm just calling the object of the decision tree. [01:25:46] I am passing my train and test data to it, sorry, train data only, X and Y, I'm passing it. [01:25:53] It will train here, and at that… and at that instant only, the model is trained. [01:25:58] So, yeah, so after this, the model is trained. So, then what I'm doing is, [01:26:04] Uh, I'm calling predict function. It will… [01:26:08] give me, like, uh… [01:26:10] a prediction in the form of binary, 1 and 0 for my training data. [01:26:15] And similarly, I'm calling predict for my [01:26:18] test data here. And basically, the model, uh, training has happened at this place. These I am calling because I want to see, like, how much [01:26:27] performance is there. So, yeah, so this is the predictions of the train set. [01:26:33] And this is a prediction of the test set. [01:26:36] So, you see, during training, I was giving both X and Y, but when I want to predict, [01:26:41] Uh, so I'm only giving X. [01:26:43] And it will tell me its prediction, 1 or 0, whatever, for the… based on this input data. [01:26:50] After the model is trained. [01:26:52] So, then, uh, I'm calling accuracy function, [01:26:55] over my, uh, like, uh… [01:26:58] Since this is the prediction value, and Y train is my ground-tooth value. [01:27:01] Uh, which I already know from the data, so… [01:27:04] This is the ground truth, this is a prediction. [01:27:06] I'm calling accuracy over this, and similarly for train and test. [01:27:12] And then we plot this. [01:27:14] Uh, the rest of the code is, like, we plot this thing. [01:27:19] So, basically, it, uh… so, yeah, so… [01:27:23] So, uh, blue here denotes that training accuracy, and we see that as the depth increases, [01:27:30] Training, accuracy, [01:27:32] goes on increasing. So… [01:27:35] And uh… but the test accuracy, it increases to certain depth, and then it is, like, decreasing. [01:27:42] So, this is… this plotting, why we do is… [01:27:48] To show where the model is overfitting, and at the place where it is underfitting. [01:27:53] So, we see that… [01:27:55] the point, maybe a point of 5 or 6, a point where [01:28:01] Uh, like, that… after that, the train accuracy. [01:28:03] It goes on increasing, but the test one goes on decreasing. So, we can say that after 5 or 6 depth, [01:28:10] The model starts to overfit. [01:28:12] It, uh, performs very well on train data, but, uh, it performs poorly on test data. [01:28:18] As its performance is dipping after that. [01:28:21] And initially, for, uh, like, lower depth, like 1 or 2, the model is, like, performing lesser on train, uh… [01:28:29] Both train and test data. So that's why we, like… [01:28:33] plot this graph, and uh… [01:28:37] do, uh, uh, do modeling over entire range of depth to see, like, [01:28:42] which is the best depth which we can use? [01:28:45] So, uh, so yeah. So, at the end, we… I've, like, also… [01:28:50] plotted them separately, uh, like, what is the test accuracy coming? [01:28:54] So… so we see that, uh, first it is increasing the test accuracy. [01:28:59] Uh, till 6th point, it is, like, 7840 is there. [01:29:03] Uh, remember, test set is one on which model is not trained. We are just, uh… like, it was not shown during training. [01:29:11] And after that, it is decreasing. [01:29:13] So, based on that, [01:29:16] We take our decision that what depth will we take, uh, [01:29:21] for the final model, uh, for training and everything. [01:29:24] So, we see that depth 6 work. [01:29:27] okay for us. So, that's why, like, we can have a decision that, okay, we will train model till only depth 6, or, like, we can… [01:29:35] Or, yeah, so… [01:29:38] This is the first, uh, like, a param… hyperparameter, which we… [01:29:42] which we want to set. [01:29:45] That… for which depth we will train the model. [01:29:48] So, yeah, so… is this thing clear to everyone? [01:29:49] Okay, thank you. [01:29:53] So, as is a depth, 6 means, uh, how we can interpret, uh, internally, like, uh, [01:30:02] Okay, okay, uh, it's later, but it will… I'll show you. So, basically, so this is, like, the level 0. [01:30:09] some attribute will be at level 0. [01:30:12] a root node. Then, at the first depth, based on, like, this attribute is having value true or false, [01:30:17] So, we will go to either the left, or we will go either to right. [01:30:23] Then we can go to left or right to light. This thing I'm saying that will only go till 6th level. We'll not… [01:30:30] Okay, okay. [01:30:31] Got it, thanks. [01:30:33] Okay, so first thing was depth. [01:30:36] And, uh, [01:30:38] Yeah. [01:30:39] does, uh, this depth 6 will cover all the attributes that we are having. [01:30:44] Like, we are having around 29 attributes. [01:30:45] Yeah, this thing is there, like, it will… obviously, it will not cover all the attributes, but the problem is that you have 29 attributes. [01:30:54] So, if you're going… if you're using all the 29 attributes, we see that the performance is very good during training. [01:31:01] But test performance is very bad. [01:31:02] So, basically, we are restricting that to use only a subset of the attributes. [01:31:09] whatever is giving good performance. [01:31:12] Got it, and this function internally decides which attributes we need to [01:31:17] have at the first node. [01:31:19] And then, what would be the child nodes? [01:31:21] Yeah, so basically, the logic is a similar, that the first attribute will be one whatever that has reducing the highest gain in impurity is, uh, entropy is giving. [01:31:32] Like, basically, if you use all the attributes, at the end, you will get true purity. But, like, the data will be totally pure separately. But if you will reach to that end, [01:31:45] Uh, then we are seeing that our trade's performance is not good. [01:31:49] So, that is the issue. [01:32:02] Okay, okay, so… [01:32:03] Yeah, can you repeat why we are choosing depth as 6 only? Why not, uh, 4? Because in the 4, we see train and test accuracy, or… Like, bar near… [01:32:04] So, basically, I choose, uh, like, I… that's why I plotted, uh, like, just printed them here. So, uh, [01:32:11] you could choose 4 also, but we saw that, uh, like, the increment is very much less, but we see that after 6th, [01:32:18] Uh, it is a kind of decreasing. [01:32:22] Okay. [01:32:23] Till 6, it is increasing. That's why I fixed at 6, but yeah, even 5th, uh, 5 or 6 could have been experimented with, but yeah. [01:32:30] Okay, instead of float, we are checking the values. Okay, good. [01:32:32] Yeah, yeah. [01:32:36] Yeah, actually, sir, just a follow-up question. So, um. [01:32:40] like, how does a model, like, exactly choose us, like, which is a good attribute to choose at the root node? [01:32:47] So, does it, like, iterately, uh, do it and just find the best one, or how does it work? [01:32:51] So, it's simple that, uh, like, uh, like, uh, have you guys, like, done that, uh, the, uh, entropy thing, like, how Digientree works? [01:33:01] So, it decides that, like, at the first run, if there are 29 features, so it will decide out of them, [01:33:07] At the root node, which one to split, that it will give the highest priority. [01:33:13] So, whichever attribute is leading to the highest reduction in entropy, it will use that at the first level. Then, out of remaining 28, it will select the next one. [01:33:22] Okay, so at every level, it decides that which is… which one to choose? [01:33:25] Yeah, yeah, the algorithm for it, the standard algorithm, which is, it does it there. It does it itself. [01:33:31] Got it. Got it, yes, thank you. [01:33:36] Okay, so, uh, yeah, sure, only halfa is there. So, yeah, so, uh, the second thing is, uh, the other thing which could be, uh, like, [01:33:46] Uh, like, I selected here depth of 6, [01:33:50] And before that, if you remember, I had used a train and test test, 90% data was train, and only 10% was [01:33:57] test. So, other, uh, thing which can impact our decision-free performance is, like, how much ratio of the data we are considering in train. [01:34:06] And how much we are considering in test. [01:34:08] So, for that to find… so next parameter to find, what we are doing is, [01:34:14] Yet, fixing the depth of 6, [01:34:16] And now, iterating over various split ratio. So, we will, uh, fix the depth as 6, uh… [01:34:24] And we will use the split, like, we will use 60% of our data strain. Next, we will use 70% of our data strain, then we'll use 80% of our data strain, and then we'll use 90% of the data strain. [01:34:34] And then we will plot four legendary, and see how much is the performance of this, so… [01:34:41] So, if you fix best depth is 6, this is a list of split ratio. [01:34:45] This is the dictionary of matrix which I am tracking. [01:34:50] train accuracy, test accuracy, train precision, traced precision, everything. [01:34:54] And right now, these are all our, uh, lists which are empty. So, I am looping through my… [01:35:00] And these ratios, split ratio. [01:35:02] So… so first, again, this train test fit function will come. [01:35:07] This time, ratio is from this list. [01:35:09] So, it will take… [01:35:11] 60% of the thing here, first time, it will take 60% of your data strain, and 40% as test. [01:35:17] And on that 60% of data, [01:35:20] So, we will learn additional classifier. So, this thing is there, and depth is, uh, we have used fix. So, the problem… so, yeah, so the problem is I'm taking the same depth, uh, like, from previous depth, you know, uh, uh, so that thing, you can say that I'm, like, the… [01:35:38] That thing is sliding red, but the problem is, uh, this is giving you flexibility to… [01:35:43] to run. Otherwise, what will happen is, if both thing is variable, so there will be, like, a 2D matrix on which I have to search. Right now, I'm searching only linearly on a list. [01:35:54] Otherwise, if I have to, uh, like, I will have to create a 2D matrix, so suppose… [01:35:59] Uh, basically, if I… I will have, uh, basically, if I have to search, so basically, then it will happen, like… [01:36:07] For depth 4, I will search for all the combinations. For depth 5, I will have to search, like, all the combinations. This can also be done. It will take some time to run. [01:36:17] That's… that's it. Uh, that's why I'm… because in order to… we… I just wanted to show, so that's why I've used that 6 otherwise. [01:36:23] Uh, like, uh, to get a full sense of it, we would have to do, like, this. [01:36:29] kind of grid search. That for depth 4, go through all the split sizes, [01:36:34] For depth 5, go through all the sizes, like that. [01:36:39] So right now, I'm using the same depth as 6. [01:36:41] And then, I'm just iterating over different train sizes. So, right now, first is 60% [01:36:46] train size here. So, you see, again, I'm calling this fit function. [01:36:51] The model is getting trained, it is learning those… all the splits of the tree, whichever will come at the nodes and everything. [01:36:59] And the attributes, it is being learned here. [01:37:00] Once learning has happened, [01:37:03] We are using it to, uh, predict our train set. We are using it to predict the test set, only the input features. [01:37:10] So, you see, I have only given X here, not Y. [01:37:13] So, I will get some predictions. [01:37:16] So, these predictions will be there. These will be binary predictions. [01:37:20] And there is one metric I am plotting here, which is called area under curve. [01:37:25] Uh, it requires probabilities predicted, not the binary prediction. [01:37:30] So, for that, for that, [01:37:33] In SKLAN, there is all, uh, this is called predict underscore prob A. [01:37:38] So, when we use it, it gives us, uh, like, probabilities, uh, instead of the… [01:37:44] Uh, like, hard-coded 1 and 0 values. [01:37:47] So, uh, that thing is there, like, stored here. [01:37:51] And then what I'm doing is, uh, I have… remember, I've introduced this matrix here. [01:37:57] So, I, uh, for train accuracy, I will pass my Y train and the predicted value I got. [01:38:03] I will pass through this value here. [01:38:07] It will append for 60%, then it will append for 70%, like that. [01:38:10] And similarly for Test 1, test accuracy. [01:38:15] And now, uh, here I am… [01:38:17] Calculating, like, more things, like precision also I'm storing. So, uh, this function is already there, I don't have to make a separate function for precision. [01:38:26] Similarly, recall, there is function. [01:38:28] So, I recall also I'm calling on both train and test data. [01:38:33] And uh… and similarly, F1 score is there, so I'm calling F1 score on train and test data. [01:38:40] And uh… so, uh, and this is the, uh, like, the area undercraft score ROC value, which I will get. So for that, I need to have probabilities instead of hard-coded 0 and 1 predictions. [01:38:54] Uh, so that's why I pass on this prop values, and so that's why the special thing which I had done. [01:39:00] So, this is different from other ones. So, this is being done here. [01:39:05] So yeah, so when this is run… [01:39:09] Yeah. [01:39:10] Uh, one thing, when you're creating the train test split. [01:39:13] Uh, there's a parameter, stratify equal to Y. Uh, why are we using this? [01:39:18] Okay, so we are using it, that we… that whatever is the dist… okay, so… [01:39:24] like, I could show it, but it will… I don't remember the command. But yeah, so basically, this is the original distribution of Y's. [01:39:33] 73% is no, and 26% is yes. [01:39:34] Yes. Yes. [01:39:36] So, when we train and test pit, we want that in our train data also, in our test data also, this proportion should be maintained. [01:39:42] That's why we are using it. [01:39:45] Okay, so no matter the proportion of trend test split we are taking. [01:39:50] Approximately this type of ratio should be there of yes and no. [01:39:51] Both of them should have similar kind of… [01:39:58] Okay, so… [01:39:59] Yeah, so now, uh, this second one was to, like, see that impact of trend as a ratio. [01:40:06] And this thing is, like, the code to, like, store everything. [01:40:11] And the next is this graph, which will, uh, since I've stored everything in this dictionary, [01:40:16] I have, uh, like, four, uh, the accuracy value, uh, precision value for all the split ratios. [01:40:24] And this I'm plotting here. [01:40:26] And it's like nothing more, just plotting L4, like, all these graphs are there, and you can see. [01:40:31] That, uh, this is… blue one is again the train accuracy. [01:40:36] And orange one is the test, uh… [01:40:37] accuracy. So, you see that in the first graph, in accuracy, [01:40:42] So, 80% of the dataset, [01:40:45] Both train and test are… although what matters the test accuracy, but yeah, yeah, we can say that 80% of the… that 80% of the split [01:40:53] Accuracy is highest. [01:40:55] similar in precision, uh, [01:40:58] Uh, it, it is… [01:41:00] 90% of the split when we use, the precision is highest. [01:41:04] And in recall, when we use 80% of the dataset, the recall is highest. [01:41:10] And F1, again, uh, I think in F1, in 60% of the dataset, when we use that time, it is… [01:41:18] Like that. Uh, like, 60%, 80% are the same value, according to F1. [01:41:23] When we use. And, uh, AUC value, area under curve, [01:41:28] Uh, it is a value between 0 and 1, and it… it… [01:41:31] it all… with the increase and split, it is decreasing only, but you see, like, uh… [01:41:36] Yeah, it is kind of decreasing with the engagement screen. So, we will have to decide based on all of this, that what will be the optimal. [01:41:43] split. Not everything is coming good for all the values. [01:41:46] But 80-20 is kind of a compromise thing, that it is best for some metrics. For accuracy, it is best. [01:41:57] For, uh, like, recall, uh, for accuracy and F1 score at its best. For others, it is, like, second best. [01:42:02] So, based on that, uh, for further thing, I will be using 80-20 split. [01:42:08] And, uh, yeah. So… [01:42:10] This is how we obtain the best split ratio. [01:42:14] So, like, any question on this? [01:42:23] I can, like, go through this code, but it's just, uh, plotting code. The main thing, uh, the… [01:42:30] how I do this particular hair, this thing, and this code is mainly for plotting… I am plotting… [01:42:38] two, three legs, like, there are, uh, six subgraphs. Last one is empty. [01:42:42] So, this code is for plotting purpose only. [01:42:47] Um, so, just one thing, can you just go over once more, uh, how are you deciding on the 8021? [01:42:54] Um… [01:42:55] Okay, so basically, uh, for deciding on ED21, it's just, like, from the visualization, I saw that [01:43:00] For accuracy, like, we'll more focus on training, sorry, on the test thing. So, for we see that accuracy is highest on… [01:43:08] And, like, the 81. 81. [01:43:10] 80%. Similarly, the F1 is also, like, the highest on the, uh, highest or same as 60% here for the, uh, [01:43:20] 80% one split, 80% split. [01:43:22] And, uh, for precision, it is, like, second highest, it can be set to second highest. [01:43:28] And recall also, it is, like, second highest. So, basically, uh, what I was saying is that, uh, [01:43:37] not… there is no single metric which is, like, coming best for all the splits. [01:43:40] So, but we can see that 80% or 90%, [01:43:44] They are either second or first, uh, they are either, like, second or first best or second best. [01:43:49] So, that's why we are settling on one of them. The other reason is that, like, [01:43:57] using only 10% for evaluation is also, like, some… is very less of the data we are using for evaluation. Uh, so that's why we, like, used, uh, we switched for 80-20. [01:44:10] Yeah, look at it. [01:44:13] Yeah, yes, go on. [01:44:14] Yep, yeah, I was just gonna… am I audible? [01:44:16] Yeah, you're audible. [01:44:17] like, this is the color version, like, [01:44:18] like, it is good to give, uh, the dataset for 80% of training, [01:44:24] And 20% of testing, right? That's what I was saying. [01:44:25] Yeah, generally, generally it is. [01:44:30] Um, just to one point, uh, this is for the, uh, Depth6 [01:44:34] Right? [01:44:35] Yeah, yeah, and yeah, it is… I have fixed… [01:44:37] Can you make it, uh, once four? [01:44:41] Yeah, okay. [01:44:42] the message before, and run it again. [01:44:43] Okay, I will try to run it and see. [01:44:56] Yeah, so as I change the depth, [01:44:57] So, yeah, so how you want to read this determinants, like, because now it's overlapping. [01:45:02] No, no, [01:45:03] Right? It's a crossing to each other, right? So, then how are you going to read this? [01:45:06] like, what is over? Like, it is not available, it's… the blue thing is train thing. [01:45:14] Okay. [01:45:15] No, no, if you look at the point 80, right, so it is crossing, okay, so what is this? So just want to know, understand that how it is… how to interpret this kind of a data, because… [01:45:17] The previous NDI is running like a… [01:45:18] No, it's there. Okay, okay. [01:45:19] not… so I just want to understand how are we going to read this data from that perspective? [01:45:28] Mm-hmm. [01:45:29] Okay, so it's not like… like, it seems like overlapping, but train and testing are different. So, I just… I just plotted the train thing, but the train accuracy is not that much to our concern. We want to see, like, how we are performing on the unseen dataset. [01:45:38] So, that's why we'll… are more focused on… will only… always be on the test set, like performance on test set. [01:45:45] Correct. So, if you go with the support, right, so we found that [01:45:48] If we train with the 4-level depth, [01:45:52] Hmm. [01:45:53] Our outcome is more accurate, right? As per the train data, even… because it is, like, on the 80. [01:45:59] Yeah, yeah, yeah. [01:46:00] Sprint, right? We are getting better accuracy than the train. [01:46:01] better accuracy at train or… and test in, uh, like, most of the cases, 80% is coming better. [01:46:09] Yeah. [01:46:10] Right? So, uh, in that case, example, if we look for the depth level 6, [01:46:14] And four, then by looking at both the graphs, right, [01:46:18] What do you think which is the… now, because earlier you showed that, okay, uh, that near that level, levels are near to 77% and 78%, right? [01:46:26] But now, looking at this, both the graphs, [01:46:29] So, what, what do you… [01:46:30] Right? Uh, as outcome, which is the better one? The fourth one, should we go with the fourth one, or the sixth level? [01:46:35] So, uh, like, these decisions have to be taken, like, these are, like, what machine learning engineer takes a decision for, for, like, either to go for 4.1 or 6.1. So, what we'll do is, like, because, uh, what, what is generally done is, because, uh, here, uh, the computation thing is there, [01:46:51] So, we'll get a great grid and calculate for everyone, for… calculate for 4, 5, and 6, and everything at once. [01:46:57] And then decide accordingly. So, like, there is no fixed rule that, whatever, what to go with. Like, it is based on, like, this, uh, values. [01:47:04] Mm-hmm. [01:47:05] that we just have to take some decision that this is, like, this could be… this could work. [01:47:09] Yeah, so my only point is that, uh, 80, let's say 80-20 split is the best one, that is clear now, both the graph. [01:47:15] Hmm. [01:47:16] Now, for the level 4 depth or 6th depth. If I look at engineer, right? ML engineer, [01:47:23] Mm-hmm. [01:47:24] if I… let's say you have to choose, right? Out of these two looking at this, which one is the port level seems better one? [01:47:26] Promising all the 6-level make it more promising. [01:47:39] Better? [01:47:40] So, what you can do is that, uh, if there's more conversion like that, so, uh, like, if you have fixed that edit, uh, like, you will have to, like, because there are so many hyperparameters, you will have to, like, fix at least one. So, you are saying that 80-20 is coming best. [01:47:42] It is correct, yes, that is true, yeah. [01:47:44] Okay, so when during final time, I will give us 80, [01:47:48] Mm-hmm. [01:47:49] And this is the model I've trained, train and test metrics, everything is there. On 80 and depth 6. We can train a similar one for 6 also, and see, like, if it is, like, how much improvement is there, or decreases there. [01:48:02] Okay? [01:48:03] like, if it is, like, that much, like, if you are, uh, like, not able to decide, then we will, uh, do a separate… we can train our full model for a depth 6, we can train a full model for depth 4, both on 80% of the split. [01:48:15] Not you? [01:48:16] And then decide. But yeah, but these decisions are challenging, because there are so many hyperparameters, and these are all interconnected, like, these are in, uh, like, grid 2D, 3D, there are so many grids where, like, there is combination for each of them. So, yeah, so it is slightly challenging to decide which one of them will work. [01:48:33] Okay. [01:48:37] Okay, so… [01:48:39] Yeah, so, uh… [01:48:42] Okay, so based on, like, all of this, I selected, like, 80-20 as a split, [01:48:47] And, uh, as I told, like, I… in this code, I've used a depth of 6 only, [01:48:52] And, uh, based on that, this is, like, the final decision tree which will be used. [01:48:57] code before this was only for… to get what will be the optimum values. [01:49:01] So, once we get an optimum value, [01:49:04] Then again, we are, like, splitting, so here, I've used train size 80. [01:49:09] And random set of same, 35 equal to is same, we'll get an 80-20% split of print test. [01:49:16] And this is my decision declassifier again, giving the depth as 6. [01:49:19] criteria is entropy. I'm using entropy as the… for… [01:49:26] selecting the split, and… [01:49:27] this function here, dot fit, it will train my model. [01:49:32] And then, as usual, uh, for getting the hard-coded predictions 1 and 0, [01:49:35] I will do .predict. [01:49:37] And over my train set, and this over my test set. [01:49:41] And, uh, because I am plotting these, our AUC values also, so I want to get probabilities values also, probabilistic value. [01:49:49] So for that, I'm doing this predict underscore prob A. [01:49:52] this thing here. So, uh… [01:49:55] Yeah, so ultimately here… so basically, this is the final code which is to be run, that, uh, the code there, all the looping and everything, that was to decide which value will come. Otherwise, if we want, uh, like, if you want to start and we don't want to, uh, like, go through all the looping and stuff, [01:50:11] We can manually put some value here, maybe you directly run this block of code, you start with some value here, and start with some depth here, and see, like, what your train and test metrics are coming. So we see with, like, I'm seeing there is not too much [01:50:24] dropped in my performance, so… [01:50:28] train and test… it was 99%. [01:50:30] in test metric, it's 78%. So, like, modelist [01:50:35] somewhat, like, okay-ish, because there is not too much drop here. [01:50:40] So, yeah, so similarly, there is values oppression, recall, and AUC, and at the end, [01:50:47] In SKLAN, there is a method of, like, classification report. It will print every… like a… [01:50:54] This kind of matrix for everything, so… [01:50:58] You just give it your ground truth Y test, and this is my predicted values for Y test. [01:51:03] And I got a separate precision for each of the class. [01:51:07] overall accuracy and, uh, like, and all further metrics, like, it's called macro and weighted average, yeah. So… [01:51:15] So, this is, like, uh, we finish with the, like, how we will finally arrive with our model, which is to be trained, and this can be run, like, separately, once we have the… once we get the notion that this will be my best depth. [01:51:27] This will be my best train and split ratio. Then we can, like, run this. [01:51:33] And just to finally end this section, we can also visualize, like, how our tree, like, [01:51:39] Uh, until depth 6, how it is there. So… [01:51:44] In SQLN3, there is a method called plot underscore tree. [01:51:47] What it does it, it… you give it your classifiers, which you've trained, CLF is the one we trained. [01:51:53] You tell it all your feature names, [01:51:56] So, extend.columns, uh, basically is a… [01:52:01] list of all the feature names. [01:52:04] Well, yeah, so it is a list of all the feature names. [01:52:08] And, uh, it will, uh, like, uh… [01:52:10] use them, or… and the class, there's a variable called, uh… [01:52:16] class name that will, like, go through. [01:52:19] It will go through this. [01:52:21] And, uh, yeah, finally, you, uh, we get, uh, like, a… it's a very, like, uh… [01:52:28] I collected one graph, like, if you maybe click, double-click over it, it will… [01:52:32] zoom in, it gets zoomed better. So, basically, it is showing me that what is my first, uh, [01:52:39] Basically, the attribute… the feature on which it is splitting. [01:52:43] And based on it, the second feature, it is going to split, and… [01:52:47] we get some visualization, like, how model has decided on [01:52:50] which feature it has split. And it is showing the, like, the… [01:52:55] entropy value also, and, uh, like, I can see the… [01:53:00] Uh, samples also consider everything. So yeah, so this is the visualization things. Okay. [01:53:06] And lastly, in this code, like, just to finish up the first section, in this code, [01:53:14] this code, what I'm showing is that sometimes people want that they want to, like, tabularly look [01:53:21] like, uh… like, basically, these, uh, if you get these numbers, so it becomes, like, uh, like, these are just four numbers, like, what did we do, actually? [01:53:31] So, what we did was that if this is our input data, [01:53:35] This is our original input data, like, if you consider this one, this sample. [01:53:39] This is our original input data. [01:53:41] We obtained, uh, like, if this is a ground truth, this is the ground truth value 0, and our model predicted 1 for it. [01:53:47] So, for each individual sample also, we can, like, analyze for which sample it is giving wrong predictions, for which sample it is [01:53:55] Giving correct predictions, uh… [01:53:57] like this. [01:53:59] Okay, so, yeah. [01:54:02] So, this is, uh… anyone want to, uh… [01:54:04] Have any? [01:54:07] Any doubts regarding this? [01:54:09] Uh, we have one doubt, like, currently our model accuracy is around 78%. [01:54:14] Uh, ideally, it should be above 90, right? [01:54:18] In most of the cases. [01:54:19] No, no, no, no, no, I… it is not like that, that it should be. Like, uh… [01:54:25] It is good to have 90, but there are, like, challenging datasets where you can't cross. So here, it is very difficult to cross because dataset is such imbalanced. [01:54:33] Got it. Because, uh, in, uh, earlier we have discussed, right, 73% our model will work only when… [01:54:40] we are, uh, marking at 0000 for other cases, and we are having 78. [01:54:44] Yeah, only 5% gain, you're saying, no? [01:54:45] So, yes, yes, yes. [01:54:48] Yeah, because, uh, yeah, that is always a challenge in, like, imbalanced dataset to how to get more increase. [01:54:55] But this thing is there that it's not always that you need to have 90% accuracy, because the inherent challenges are dataset is such so much that… so if you see, [01:55:05] In this dataset, even the train accuracy is also not 90%. [01:55:09] So, if the model on which it's trained, on that also it is not giving 90%, so we can't expect it to, like, [01:55:18] give good performance on test set. [01:55:19] So that is auto. [01:55:20] Yes. [01:55:22] Okay, uh, yeah. Also, uh, what's the difference between entropy and Genie, these two criterias? [01:55:29] Okay, so, uh, like, uh, like, these are theoretical, like, maybe ma'am will, uh, ma'am will cover it. [01:55:36] So, basically, these are, like, different mathematical formulas which decide what will be the first root node. The difference is just that. [01:55:44] And both of them, like, works similarly, that, uh, that… [01:55:50] pure the dataset that you have to, like, you partition the dataset based on, like, [01:55:57] You have one partition where all the, uh, value is 0, and on the other side, there's all the values 1. So… [01:56:04] Yeah, maybe ma'am will cover it better, so yeah. [01:56:08] Sure, thanks. [01:56:12] Yes, Lukash. [01:56:13] Yeah, as you said that if we can give manually, [01:56:15] about the test sites and training dataset, like 80-20 and the maximum breadth is… [01:56:20] like, um, how can we fix the criteria for every data subject? [01:56:25] These things have to be decided. That's why I've included them, that these things have to be decided on. [01:56:32] Like, for… [01:56:33] Okay, without looping, we can design, like, we can give the resting size of 2020 and writing. [01:56:37] It is possible. [01:56:38] Yeah, yeah, yeah, yeah. If… generally, it is, like, good that you… the more data you give to training, it is better. [01:56:44] It is always the case, so you will have only to decide between maybe 70, 80, 90. Like, these are the speeds you can take. [01:56:51] Yeah, yeah, it's just… [01:56:55] Yeah, Shivanj. [01:57:01] Uh, hi, Ashes. I want a small check, uh, thanks for today's session, and where we emphasize about decision-free classifier library. [01:57:10] Uh, just to check in the… Upcoming session, are we going to touch upon. [01:57:15] Other libraries like gradient, boosting, decision tree. Or… actual… [01:57:21] I think this, like, regarding course content, ma'am, we can tell you better? [01:57:24] Like, I have… there is random forest and vision tree here in this… in this house. [01:57:31] But not about the concepts or algorithm. Decision 3 can be implemented by multiple libraries, like today we have. [01:57:35] No, no, okay, we will be… [01:57:38] Okay. [01:57:39] Focus on decision tree classifier. So, like, decision tree classifier is provided by SKKit. [01:57:41] Yes, Kellan. Mm-hmm. [01:57:45] Yeah, yeah, yeah. [01:57:46] Yeah, so, uh, there are some of our other libraries are available, like extreme gradient boosting, so are we going to touch upon, uh, in the upcoming sessions? [01:57:50] No, no, uh, that thing we usually don't… we, uh, we'll be using only SQL for all our machine learning stuff. [01:57:57] like, their syntax is similar, like, you people can try after, like, if you want to see the improvement in how they are implementing, if there is implemented performance, you can do. [01:58:08] Okay. Thank you. [01:58:21] Yeah, actually, it… time is already over. Random Forest is there inside it. [01:58:25] Uh, but yeah, but we were not able to cover it. [01:58:26] Ashish, uh, so whatever we did today, it was about a single decision tree, right? It wasn't about the random forest, right? [01:58:29] Okay, okay, got it. So, it's gonna be there. [01:58:33] Yeah, it's there in the notebook, uh… [01:58:35] let's see, like, I will ask, ma'am, how it will, like, whether… when to continue maybe next class, or what to… how to… [01:58:42] But, like, presented. [01:58:46] Okay, yeah, but there is one more question, which is kind of out of the scope. [01:58:50] Uh, so you said, like, uh, once we get the classifier, right, we can use it to predict. [01:58:56] Hmm. [01:58:57] Correct. So, we can predict individual value with it. [01:59:00] So, to reach to the classifier, uh, if you talk about this collab, we'll have to go through all the, uh, you know, blocks, we'll have to execute them one by one, and then reach there, right? [01:59:11] So, what are the ways, right, in which we could, like, retain that model, and then we could directly use it? [01:59:19] Without going through the processing step. [01:59:22] Okay, so, uh… [01:59:24] Yeah, yeah, so… [01:59:26] The major block, like, which is whatever written A here, [01:59:29] It is the, like, the, uh, like, one single block. Somewhere where values are there inside it, which is to be, like, coded. [01:59:37] So, you have to, like, go to the block where we are getting X encoded and Y. [01:59:43] Uh, that block, you have to run. So basically, what you can do is, like, it does not take much time. If you do a run all over everything. [01:59:50] Hmm. [01:59:51] So, it runs everything. So after that, you can, like, go to one, uh, like, insert your own text blocks, own code blocks, [01:59:59] And do, like, analysis more. [02:00:03] Like, if it's confusing to isolate everything. Otherwise, the main one block is this. [02:00:09] And whatever else is required to, like, implement this one block. So, what else is required is, uh, you have to include the loading code, the loading of the CSV file. [02:00:18] Mm-hmm. [02:00:19] And you have to include the code where we were obtaining this X encoded and Y. Only those initial blocks are required. [02:00:27] Okay, got it, got it. So, uh, my question was, uh, more on the, uh, practical part. [02:00:38] Hmm. [02:00:39] Or maybe we could say application part, right? So, you know the concept of, uh, server being listening to the requests, right, and being ready to reply. [02:00:46] Uh, whenever we send a request, right? [02:00:49] So, uh, I was asking about that thing, like, I wanna keep this model ready, uh, listening, right? [02:00:56] And if somebody sends a request to, uh, you know, get a protection, right? So… [02:01:03] Uh, at that moment, right? I will have to, uh, go through this process, right, again and again. [02:01:08] Uh, the, uh, the training part, no. [02:01:11] No, no, at that time, you will use that, uh, whatever, once you have offline, you are trained your model. You have done this clf.fit. [02:01:17] Mm-hmm. [02:01:20] Okay. [02:01:21] So, then you will all… you will use your trained model, and whatever will your request will come, you will do CLO.predictonly, per sample-wise. [02:01:28] Okay. [02:01:29] Yeah, now the thing which will differ is that there is this concept of data drift, because we are assuming that train and test comes from same distribution. So, uh, eventually, your… because multiple users are giving everyday requests, so data will… the distribution will differ. [02:01:44] So, once you have, like, after some days, you will again train on whatever the new data you have, you will accumulate it with your old data. [02:01:52] And train… train it again, offline. [02:01:54] Okay. Okay, so somewhere will store that object CLF, right? [02:01:59] Yeah, yeah, yeah. [02:02:00] Got it, got it, got it, yep. [02:02:05] Okay, so, uh, yeah, so no. [02:02:06] On the same lines which we were discussing, [02:02:10] Let's say we have trained our model, CLF, and we can predict the future data, like, row-wise, or record-wise. [02:02:19] But the new record will not be one code hard-coded. [02:02:23] So, because we know, because when you obtain your data in the JSON structure and everything, [02:02:32] Mm-hmm. [02:02:33] You will pass that through the same way we will pass in… have… we have passed through training. The same preprocessing will happen to it. [02:02:40] That thing has to be made, sure. You can't give it in raw form. [02:02:44] So for that, do we have to save something, like, encoding also, like, the way which we… [02:02:49] Yeah, yeah, yeah. [02:02:50] Like, we can save CLM. [02:02:52] we can save the preprocessed, uh, the step, where the preprocessor is there. We can save its state also, that it will, like, it will… if you get a raw sequence, it will structure in the form that the model expects. [02:02:53] Okay, got it. [02:03:07] Okay, okay. [02:03:08] Okay, so basically, it will convert the row in the format it needs, uh, for protection, right? [02:03:13] Yes, yes, like one-hot encoding and everything. [02:03:15] Okay, got it, got it. So, that would basically be a pipeline, right? [02:03:19] Yeah, yeah, yeah. [02:03:24] So, uh, okay, so, uh, time is also over. So, what I want to, like, [02:03:30] show his, like, uh, like, most of the part of this notebook is… maximum thing is redundant. [02:03:36] So, what that I mean is that in the Part B, [02:03:41] I'm using the same code, doing everything again, but just the difference is I'm showing the criteria is GINI. [02:03:45] So, I have used, uh… I think the same 80-20 split. [02:03:50] And I've used the max depth of 6, and criteria I've used Genie, and to see whether, uh, using… choosing this criteria, do I get something different in values and… [02:04:01] So, the only difference I see is [02:04:02] The train accuracy has increased a bit. [02:04:05] Uh, but the test accuracy is slightly, somewhat similar. So, and then, the whole code is same as before, just this criteria is changed. [02:04:16] And then I also plot this tree again to see, like, the features which were used in, uh, like, if you… if we just go… [02:04:24] Like, yeah, we'll be finishing, just give me a moment. So, if you see the first feature in Entropy, it was like contact 2 here. [02:04:30] the root node feature. When I used Ginny, [02:04:34] When I use Ginny, the feature was something else, internet service fiber. So, yeah. So, different criteria. [02:04:40] different, uh, features are selected for [02:04:43] being at each level. [02:04:45] Yeah. So, uh, so in this code, like, APAR, this is A and B, so similarly, there is, like, two more parts, C and D. [02:04:56] And, uh, that thing are left right now, and it's already to us. [02:05:00] So, we'll finish them, like, I will ask them to schedule a class for that. [02:05:05] And, uh, most of the things… but even if you go through itself, like, most of the things are now similar. Even for random forest, like, there are two hyperparameters, uh, for the random forest classifier, on which we'd first decide that what is the optimal number of trees which we'll use, because in random forest, many… there are multiple trees, like… [02:05:24] should we take 50 trees, 100 trees? And similarly, there is one more hyperparameter right now for us. First, we decide on them, and then simulate the code pipeline is same, then once we get a value of the… [02:05:36] that this is the value of the random forest. [02:05:38] Then we… this is the hyperparameter, then we do the same accuracy thing. This accuracy precision thing. [02:05:44] So, yeah, so without taking not much more time, yeah. So, this was… [02:05:51] So, shall we wrap up? [02:05:54] Or anyone want to, like, uh… [02:05:55] Okay, okay, got it. [02:05:56] Yeah, so the… [02:05:59] Okay, okay. [02:06:00] Yeah, yeah, please go ahead, yeah. [02:06:01] Oh, good, good. [02:06:02] So, one quick… [02:06:03] And I was talking about the, uh, hyperparameters, right? So, when we talk about, uh, definitely we will be having multiple trees in random forest, right? [02:06:11] So, first of all, we'll have to find the depth for the tree, and then we'll have to find the number of, uh, trees, right? Is that correct? [02:06:21] Okay, yeah, so things are different. If you are just using regent tree, then depth is, uh, then depth will concentrate on depth. [02:06:28] For random forest, the major criteria is this number of trees, because it runs on… it aggregates on many… multiple trees. [02:06:35] So, number of trees does not come in decision tree classified, it will come in random forest thing. [02:06:40] Okay, so, but we… we, uh, plotted few graphs, right, to identify what should be the depth and everything, right? So, do we have any… [02:06:47] Yeah. [02:06:49] plotting or any criteria, which… [02:06:50] Yeah, if… yeah, if you go inside a part… part… Part 2 random, plotted the same thing, like, [02:06:57] This is an estimator, tell me number of trees. So, I have, like, shown for 50, like, the values. I think I… [02:07:04] taken the value 50, 75, 100, 150, these are the number of trees I've taken. So, for that also, like, values are [02:07:11] plotted. So… [02:07:12] Okay. [02:07:13] That's how we decide. [02:07:14] So, one question? [02:07:16] In the random forest, right? [02:07:20] Yeah. [02:07:21] So, the depth of all tree must remain uniform, or it can change. [02:07:25] According to this library, once you… the depth is same, for that, you'll have to, like, implement our own random forest if we want it for… but, theoretically, it is like that depth remains same, like, how random 4 is defined. [02:07:38] Okay? [02:07:39] Like, there are multiple trees, but each will have the same depth. [02:07:41] Okay, and this print also the same. [02:07:45] Uh, split, uh… no, no, for, uh, basically by the name random forest is there, so basically it means that if there are 29 features, you randomly select a subset of features, [02:07:55] Yes, Mr. [02:07:56] It will be… so… so each one will get… maybe each one will get a separate, uh, set of features, each one… each one will be shown. [02:08:02] No, no, that is the number of records, uh, like, my point is that, uh, [02:08:09] But there are each and every tree, like the 50 trees selecting, right? [02:08:13] You can make clearly that, in theory, [02:08:14] It should be the depth remains the same. [02:08:16] Right? Secondly, the split, which we're doing for training, right? That will also remain the same, like, 50, 70, like, if it says 80-20, it will remain 80-20 for all, or it can be sometime for 80-20, some 30… [02:08:31] Got it? [02:08:32] No, no, no, okay, okay, okay, yeah, the split thing is different, that what you pass to that .fit function, that thing remains same, that it… you are passing 80% of the data as training. [02:08:36] So that will remain the same uniform. [02:08:37] Inside, yes, you are taking 80% of data, now you are in residential, you are training one single tree on 80% of data, now you are training a random forest for your 80% data. The difference is like this. [02:08:50] Okay, got it. [02:08:54] Okay, so, yeah, it's already 8.5. [02:08:58] So, let's… [02:09:00] Yeah. [02:09:01] Uh, one question to JNI batch to manager, maybe Simran, if you'll be available. [02:09:08] Yeah, so regarding this feedback, uh, the form which you were sending to Phil, so we are filling it, but internally, how you guys are taking any actions, if some suggestions we are giving. So, just want to know on that part. [02:09:09] Yeah? [02:09:23] Yes, yes, we will be taking action based on your feedback with Professor Toshnimal. [02:09:29] Okay, okay. [02:09:33] Okay, fine, fine, let me see. [02:09:34] And also, is there any sessions, like, uh… [02:09:37] Uh, there'll be a get-together or not get-togethers kind of orientations, like, when we will, uh… [02:09:46] or gather together, and we can discuss about the improvements and all. Maybe casual one. [02:09:54] Tier 1, [02:09:55] So, to make this class more intact. [02:09:56] Yeah, yeah, if you want such a session, y'all can write over email, and I'll see what we can do about that. [02:10:02] Okay. Right. [02:10:06] Thank you. [02:10:12] Okay, everyone, then have a good night. [02:10:15] I think, uh, all doubts covered, I guess? Okay. [02:10:18] Thank you, Jesus. [02:10:21] Thank you, Levon. [02:10:22] Okay, it gives… okay, go ahead. [02:10:24] Thank you. [02:10:31] Thank you