# 07 2026-02-14 Master Class Supervised ML

course: Module 3 — Deep Learning & NLP
module: Module-3-Deep-Learning-NLP
date: 2026-02-14
type: transcript
video_url: https://personal-learn.armco.dev/files/_Recordings/Module-3-Deep-Learning-NLP/07_2026-02-14_Master_Class_Supervised_ML.mp4

---
[00:05:58] Mithun DJ: Hello, guys.
[00:12:01] Mithun DJ: Hello, guys.
[00:12:04] Mithun DJ: Can you hear me?
[00:12:10] Mithun DJ: Hello.
[00:12:12] aditya shrivastava: Yes.
[00:12:13] Mithun DJ: Okay, okay.
[00:12:15] Mithun DJ: Should we get started, guys, everybody?
[00:12:21] Aditya Banda: Yeah, sure.
[00:12:24] Mithun DJ: Okay, just give me a moment, please.
[00:12:41] Mithun DJ: This, do let me know in case you're able to see my screen.
[00:12:47] azad choubey: Yes, we're gonna see.
[00:12:48] aditya shrivastava: Yes, we can.
[00:12:49] Mithun DJ: Okay, okay, thank you so much, guys, thank you. Okay. Guys,
[00:12:58] Mithun DJ: We'll be speaking about, you know,
[00:13:02] Mithun DJ: Supervised machine learning, and within supervised machine learning, there are a few challenges.
[00:13:06] Mithun DJ: We will be looking at some of these challenges, in today's session, okay?
[00:13:20] Mithun DJ: So, as far as the agenda for today's presentation is concerned, you know, I will just give you an introduction into a subject which is called as regularization.
[00:13:30] Mithun DJ: Some definition of regularization. There are mostly three different types of regularization techniques that you are expected to know. We'll look at that. The differences among these techniques.
[00:13:42] Mithun DJ: And we can look at some examples for each of these things, and I've got a couple of case studies for each of these things.
[00:13:54] Mithun DJ: Now, Grace, what do you understand when we say supervised machine learning techniques, Grace?
[00:14:01] Mithun DJ: supervised machine learning, when you say this, what comes to your mind?
[00:14:04] Aditya Banda: When there are labels to the data?
[00:14:06] Mithun DJ: Okay, when there are labels, so you have worked on any of these supervised machine learning techniques?
[00:14:12] Aditya Banda: Yes, we did.
[00:14:13] Mithun DJ: Okay, you have, right? So, when you are working, typically, which algorithms have you used so far?
[00:14:19] Aditya Banda: Linear regression.
[00:14:21] Mithun DJ: Okay, you've used linear regression, okay. Anything else that you've used?
[00:14:25] Muni Prakash Ganji: Classification.
[00:14:27] Mithun DJ: Okay, you have used classification. What exactly have you used within classification?
[00:14:33] Mithun DJ: Any specific technique?
[00:14:34] Mithun DJ: Any algorithm?
[00:14:41] Mithun DJ: Yes.
[00:14:43] Aditya Banda: I think we have used Random Forest for this.
[00:14:45] Mithun DJ: okay, you'll use a random forest that's slightly more advanced. At a very basic level, you can maybe look at a close cousin of linear regression, which is your logistic regression.
[00:14:57] Mithun DJ: Okay, so typically linear and logistic regression. So I'm just trying to tell you that these are some of the… at a very high level, you may have these machine learning types, like supervised, unsupervised.
[00:15:07] Mithun DJ: and reinforcement learning. I will not be focusing too much on unsupervised and reinforcement learning, because that's outside the scope, but when you're building a typical
[00:15:18] Mithun DJ: when you're building a typical supervised machine learning technique, you'll encounter certain challenges and systems. And these challenges is what we are trying to address, basically.
[00:15:29] Mithun DJ: Oh… Okay, now, supervised machine learning, we have already discussed this. See.
[00:15:35] Mithun DJ: Under supervised machine learning techniques, you have both regression as well as classification.
[00:15:40] Mithun DJ: So, under regression, you may have linear addition tree, neural network, ARIMA, REMAX. It is not that you'll find the need to regularize under all circumstances. You'll need to… you'll find the need to regularize only under certain circumstances, not under all circumstances.
[00:15:57] Mithun DJ: Typically, when you look at models like linear regression, or logistic regression, or your decision tree and such things.
[00:16:04] Mithun DJ: they, there you will encounter a problem. What is that problem? That problem is a problem of overfitting. Have you heard of this expression, overfitting is?
[00:16:17] Aditya Banda: Yeah.
[00:16:18] Mithun DJ: What do you understand when you say over-roffitting?
[00:16:25] Muni Prakash Ganji: I need data with, predefined values, like labels we given predefined, and everything we given.
[00:16:35] Muni Prakash Ganji: So, there is a no scope of thinking. It is all hard-coded.
[00:16:39] Mithun DJ: Okay.
[00:16:42] Aditya Banda: The model is fitting test data too well, and the scope of generalization is not great.
[00:16:48] Mithun DJ: Yes, you're absolutely right. Typically, the model, fits the training data very well.
[00:16:53] Mithun DJ: And it does a very poor job as far as the test data is concerned. So, when you're building a model and you see that in the training data we have an accuracy level of, let's say.
[00:17:03] Mithun DJ: 80% or 85%. At the moment you apply the same model, you're very happy. Firstly, that in the training data set, my model is doing very well, 90-95% accuracy is doing.
[00:17:14] Mithun DJ: The moment you apply it on the test dataset, your accuracy comes down drastically.
[00:17:18] Mithun DJ: by, 60%. So this is a problem that you see mostly in, some of the, regression techniques and such things, some of the supervised machine learning algorithms and things.
[00:17:29] Mithun DJ: So, under regression, I'm just mentioning that you have all of these algorithms, linear regression, decision tree, neural network, Arima, and Arimax. So, for neural network, ARIMA and Arimax and such things, you'll not encounter this problem, but for linear logistic, decision tree is a classic candidate for
[00:17:46] Mithun DJ: your regularization, or it suffers from the problem of overfitting, right? And we'll see this. Again, under classification, when you see logistic regression, you can have the problem of overfitting. I have, in a project, basically, I've spent a lot of time, you know, correcting for the problem of overfitting long back around 10 or 12 years back. Tishantry, again, guilty of overfitting.
[00:18:05] Mithun DJ: Neural network, we typically do not see the problem of forfeiting in such things, okay?
[00:18:10] Mithun DJ: So…
[00:18:12] Mithun DJ: some examples I've given, I'll just skip this, and now I will look at the core of regularization, right?
[00:18:19] Mithun DJ: Under regularization, there are some examples, correct? Where can you see this regularization? Where should you use? That's what I'm saying. Typically, when you're using linear regression, let's say you're trying to predict sales or price.
[00:18:33] Mithun DJ: Typically, you can expect a problem of multicollinearity. This happens because there's a problem of multicollinearity. I'd like to take a step back and explain, have you come across some of the assumptions of linear regression, guys?
[00:18:54] Mithun DJ: Anybody knows what are the assumptions behind regression?
[00:19:02] Mithun DJ: Yes.
[00:19:02] Muni Prakash Ganji: There will be… every output will follow the linear relationship between the x-axis and y-axis.
[00:19:10] Mithun DJ: Okay, that could be one of them. Is that the only one?
[00:19:18] Mithun DJ: Yes?
[00:19:28] Mithun DJ: Anybody else, guys?
[00:19:30] Mithun DJ: Do you know that there are assumptions, and why do we need these assumptions? Are you clear? What are these assumptions?
[00:19:36] Mithun DJ: What happens when these… when there is a violation of these assumptions? Are you familiar with those things?
[00:19:40] Mithun DJ: If not, I'll just quickly take some time to explain it, or else I'll just make a move on.
[00:19:48] Mithun DJ: Yes, no.
[00:19:52] Aditya Banda: Yeah, can you describe them?
[00:19:55] Mithun DJ: See, firstly, why do we need assumptions?
[00:19:59] Mithun DJ: That's the question.
[00:20:06] Mithun DJ: If you can understand that part, basically, the rest of the things will be very, very clear, right?
[00:20:12] Mithun DJ: why do we need these assumptions? Have you, you know, have you ever,
[00:20:18] Mithun DJ: spent time with, you know, removing these problems and such things, and why do you… why do you feel that these problems come up?
[00:20:26] Mithun DJ: See, understand something. When you talk about any of these fancy techniques, linear regression and logistic regression and such things, let's take a step back. What are these? These are all statistical frameworks.
[00:20:39] Mithun DJ: So, statistical framework. So, statistics as a subject itself, if you trace back the history, it came into existence in 1750s, right? Today, we're in 2026. So, what…
[00:20:51] Mithun DJ: So, when the early statisticians basically developed many of these frameworks and such things, there are certain assumptions behind the data. They had to assume
[00:21:00] Mithun DJ: the nature of the data and such things. Only under those controlled conditions, only when your data follows those assumptions will your statistical techniques give good result. I'm just giving you a very simple example. Let's take the example of mean. What is mean?
[00:21:13] Mithun DJ: A lot of people say mean is nothing but total divided by number of values. Is that it? No. There are certain barriers, assumptions, caveats, only then you can use mean. What are the barriers? If there are too many outliers and such things, you can't use mean, because mean will be inflated.
[00:21:28] Mithun DJ: Correct? One more, one more… in the case of, let's say, correlation, right? When you talk about correlation again,
[00:21:35] Mithun DJ: there are assumptions. The data should be linearly related. Second, basically, the dependent variable… sorry, both the variables should be normally distributed. If the data is not normally distributed, if you have too many outliers and such things, if there is… if there are… if there is a nonlinear relationship, you can't use correlation.
[00:21:53] Mithun DJ: So, there are assumptions, barriers, and caveats, and if you respect those assumptions and such things, your statistical techniques hold good. Otherwise, basically, your statistical techniques do not hold good. And I would like to, in this particular slide, I will specifically highlight
[00:22:08] Mithun DJ: the usual suspects here, correct? So, when we talk about any of these techniques.
[00:22:14] Mithun DJ: which I'm just sort of highlighting for you, linear regression, logistic regression, and such things, one of the most important assumptions. There are 4 or 5 different assumptions.
[00:22:24] Mithun DJ: One is that linear regression should be normally distributed. If it is normally distributed, it'll give you better results. R-square and all will be high. Second important assumption is that the independent variables must be uncorrelated.
[00:22:36] Mithun DJ: The independent variables among themselves should be uncorrelated, but the independent variable should show correlation with the dependent variable, correct? Third one is basically that there is a linear relationship. Fourth one is that the errors or the residual terms basically should follow normal distribution with a mean zero and a constant sigma square.
[00:22:54] Mithun DJ: If there is a non-constant variance, right, it leads to a problem called as heteroscedasticity, alright? Next, basically, you have a case wherein the
[00:23:03] Mithun DJ: If you specifically talk about linear regression and such things, there is a problem of autocorrelation, which means that the error terms are correlated within themselves, right? If the error terms are correlated with themselves, basically it leads to a problem which is called as autocorrelation.
[00:23:19] Mithun DJ: So there's a problem of, there's a problem of multicollinearity, autocorrelation, there's a problem of heteroschoasticity. Forget the last two, basically, because it'll take our discussion to another level. Let us mainly focus on the problem of multicollinearity. Do you understand the term multicollinearity, guys?
[00:23:38] Mithun DJ: Do you understand the term multicollinearity?
[00:23:43] Mithun DJ: Everybody?
[00:23:47] Mithun DJ: Yes or no?
[00:23:48] Muni Prakash Ganji: No.
[00:23:49] Mithun DJ: No, yes. See, when you, when you see a situation when there is high degree of correlation between the independent variables, that's a classic case of multicollinearity. Low degree of correlation is still fine. Here and there, one or two variables are showing correlation and such things, it is still fine.
[00:24:06] Mithun DJ: But if there is a high degree of correlation among the independent variables, then it leads to a problem of multicollinearity. You can ask me, so what? Let there be a problem of multicolinity, why are you worried? See, the moment you have a problem of multicollinearity, then there… it'll lead to another problem, which is called as overfitting.
[00:24:25] Mithun DJ: To correct for the problem of overfitting, we are doing regularization. That is the story behind it. I'm just talking about regularization, but the story behind it is, you have multicollinearity, the independent variables are usually correlated with themselves.
[00:24:38] Mithun DJ: while the demand of regression is that they be uncorrelated. What happens if there's multicollinearity? Too many, too many predictors correlated amongst themselves?
[00:24:46] Mithun DJ: that leads to a problem which is called as, overfitting. What happens if there is overfitting? If there is a problem of overfitting, the unstandardized quotients will be, will look up, it'll be inflated, the standard errors and such things will be very, very high. The model may learn, out of 10 or 15 independent variables, the model may actually learn from only top two variables, the remaining, the balanced 13 variables may not
[00:25:11] Mithun DJ: at all.
[00:25:12] Mithun DJ: So all of… it's not a balanced model at all. So your standard errors will be high, your beta coefficients and such things will be unreliable, the width of the confidence interval will be wider. So, all of these things, basically, are called as multicollinearity. Now, if multicollinearity is a disease, how do you attack this disease? The answer is regression… sorry, the answer is regularization, which we'll be learning today.
[00:25:35] Mithun DJ: Is this part clear to you, why we are learning regularization and systems?
[00:25:41] aditya shrivastava: Yes.
[00:25:42] Mithun DJ: Right? Is it clear to everybody? I hope I'm not confusing people.
[00:25:46] Mithun DJ: Is this clear?
[00:25:48] Sushree Dash: I think it will be more clear if we have example.
[00:25:53] Sushree Dash: I'll show you the examples as well, but as of now, I'm just telling you, why do we need regularization. Okay. Regularization is used, to…
[00:26:00] Mithun DJ: correct the problem of overfitting. Why do you get overfitting? Because you're violating the assumption of regression. What is the assumption of regression? The independent variable has too much of collinearity.
[00:26:10] Mithun DJ: correlation Is that clear?
[00:26:17] Sushree Dash: Yeah.
[00:26:21] Mithun DJ: See, I want you to just focus on three different types of, regularization techniques. Now, the first one, the second line, if you see, a rigid regression, then there is what is called as lasso regression and elastic net. As of now, if you want to regularize
[00:26:37] Mithun DJ: And you get into regularization when there is overfitting problem and such things. There are 3 medicines that you have. There are 3 weapons that you have.
[00:26:45] Mithun DJ: Sorry. The first weapon that you have is known as a ridge regression.
[00:26:49] Mithun DJ: Second weapon that you have is lasso regression, and the third weapon that you have is elastic net. I will basically give you greater details of all of these things, but, understand as of now, what is water, correct?
[00:27:03] Mithun DJ: address the problem of linear regression, we need 3 other… three other weapons. Rich, Lasso, and ElasticNet. What are all of these things? What is the difference? When do we use this? When do we not use it? I'll elaborate now all of these things, but as of now, just remember that it's a
[00:27:20] Mithun DJ: It's a regularization technique, or it's a way in which we can go ahead and drastically reduce the problem of overfitting at such things.
[00:27:29] Mithun DJ: Now, classification algorithms, as of now, basically, you can have classification algorithm. Here, again, basically, right at the top, when we speak about logistic regression, this also, basically, suffers from the problem of overfitting, especially when there is heavy problem of multicollinearity and such things.
[00:27:46] Mithun DJ: Now, to address it, basically, we use random forest classifier. This is used a lot, basically. So, random forest as an algorithm came into existence to correct the problem of overfitting.
[00:27:57] Mithun DJ: That is posed by distance, correct? So, whenever there's an algorithm which has a problem, to correct that, we create another algorithm, and here, random forest is used. It's a great technique because it removes the problem of overfitting and such things to a large extent.
[00:28:13] Mithun DJ: I'll skip this, skip this, yeah.
[00:28:16] Mithun DJ: Right.
[00:28:17] Mithun DJ: Okay, now I've spoken a lot of things. Now, if you want a formal definition of regularization.
[00:28:23] Mithun DJ: What exactly is regularization? Typically, when we begin with regression and such things, there are certain classic questions that were being asked. One is that regularization is a main technique in machine learning. Why do we use this? It is used to prevent
[00:28:38] Mithun DJ: the problem of overfitting. Question is, how do you prevent the problem of overfitting? The answer is very simple. We are adding a penalty term to the loss function of the model. Typically, under linear regression, we use a loss function. That is, loss function is used,
[00:28:58] Mithun DJ: it's nothing but your OLS model itself, ordinary least squares. We are getting into minimization of error sum of squares.
[00:29:04] Mithun DJ: Now, to that loss function, if I add a penalty term, basically, I can basically, to a large extent, solve the problem of overfitting. That's what I'm trying to tell you. Now, when does overfitting occur? That's the question. Overfitting typically occurs when a model fits the training data too well and performs poorly on the test data. This is what we have discussed. Last point is, basically.
[00:29:28] Mithun DJ: Your regularization keeps the model weights small. This is a very important point.
[00:29:34] Mithun DJ: instead of having inflated numbers for your model weights, your regularization basically keeps it very, very small. Instead of having large values like 100, 200, 2000, 250, and such things, it'll basically trim it off and reduce it to a small number.
[00:29:50] Mithun DJ: Second benefit of using regularization is that it reduces the problem of model complexity. Do you understand what is the meaning of model complexity here, guys?
[00:29:59] Mithun DJ: What does the word complex mean here?
[00:30:06] vinit shah: If you can give some context to it, it'll be easier to understand.
[00:30:09] Mithun DJ: The context is this itself, multicollinearity, assumptions, those sort of things.
[00:30:13] Aditya Banda: Maybe it'll… maybe it'll reduce the number of independent variables.
[00:30:18] Mithun DJ: Very good, very good. That is the answer, that's the right answer that I was looking at. Typically, you know, if you have too many variables, independent variables, you have, you know, 50, 60, or maybe you can also have, you know, unlimited number of variables, let's say 200 or 250 independent variables. That is called as a complex model.
[00:30:37] Mithun DJ: Please understand something, look at the contradiction here. When we kickstart the process of data collection for a model.
[00:30:43] Mithun DJ: we would like to have a complex model. I would like to include as many variables as possible, because when I kickstart the data collection procedure, I do not know which variable will turn out to be important, which variable will be unimportant. So, since I'm greedy.
[00:30:57] Mithun DJ: Any data scientist is greedy. I will take 100, 150, 200, you give me as many variables as possible. I'll take macroeconomic factors, microeconomic factors.
[00:31:07] Mithun DJ: If possible, I'll take, you know, product information of the customer, account information of the customer, demographic information of the customer, psychographic information of the customer, socioeconomic factors, as many as possible, I'll take, so that I will get 200, 300, as many variables. Why? Because before building the model, how do I know whether age is important, or whether gender is important.
[00:31:30] Mithun DJ: Microeconomic factors will be important or not.
[00:31:33] Mithun DJ: But… When you propose the final model.
[00:31:37] Mithun DJ: It's exactly the opposite. I want to propose the final model with as few variables as possible. Why is it?
[00:31:44] Mithun DJ: This is because you cannot build… you cannot propose the final model with 60, 70, or 80 variables. The final model, when you're proposing to the management or deploying in a production environment, it should be… it should be crisp.
[00:31:58] Mithun DJ: 5 or 6 variables, not more than that. Here and there, if you want, you can stretch it to 7, but beyond 7, it becomes a very complex model.
[00:32:06] Mithun DJ: The opposite of complexity is parsimonious. Parsimonious means somebody who's very poor, somebody who's basically very, very, very opposite of rich it is, right? Very few, right? So, parsimonious. We would like to build a parsimonious model, not a very, very complex model. Why?
[00:32:26] Mithun DJ: Typically, because if you have a parsimonious model that improves generalization. So all of these things basically are tied. If you can basically… if you have a problem of overfitting, typically you have too many variables. Too many variables means what? The model is complex. So what would I like to do? I like to reduce it, but can I reduce it manually? Not possible, right? Because, how will I know which variable is important, unimportant, and such things, correct?
[00:32:50] Mithun DJ: So, we have to reduce or shrink the model weight.
[00:32:53] Mithun DJ: We have to reduce the problem of complexity, and once you are basically able to reduce the problem of complexity, you'll automatically be able to improve the generalization. So this is the story of regularization that I wanted to convey to you. So far, are you clear, guys?
[00:33:10] Mithun DJ: So far, a repair.
[00:33:11] Aditya Banda: Yeah.
[00:33:12] azad choubey: Yes.
[00:33:13] Mithun DJ: Understand the concepts as much as possible, guys, because the code is there. I have… anybody should give the code, I'll also run the model. Code, anybody can run the code, guys. It's not a big deal. You use ChatGPT or any AI tool for everything, there's a code. But more than the code, understand the concept.
[00:33:27] Mithun DJ: Why regularization? What is the problem that it is trying to solve? It is trying to solve the problem of overfitting. So it's a beautiful concept, right, to understand this.
[00:33:36] Mithun DJ: Right? Now, regularization types. This is the definition of regularization, right? Now, the types of regularization. The most important
[00:33:44] Mithun DJ: And the most popular type of regularization is called as L1 regularization. I repeat, this is called as L1 regularization. Now, L1 regularization is also called as lasso regularization. I repeat, it is called as lasso regularization.
[00:34:05] Mithun DJ: Correct? Is this clear to you guys?
[00:34:10] Mithun DJ: Have you heard of Lasso Anybody?
[00:34:14] Mithun DJ: Have you heard of Lasso?
[00:34:18] Mithun DJ: Yes, sir.
[00:34:19] vinit shah: It's like a rope or something, right?
[00:34:21] Mithun DJ: Yeah, okay. Yeah. Here, when you say lasso, L-A-S-S-O,
[00:34:27] Mithun DJ: It is just an acronym. It actually stands for L stands for least, A for absolute, S for shrinkage, and then selection operator. Least absolute shrinkage selection parameter, that they have basically condensed and compressed it, and they've just given… given you an acronym called as LASSO regression.
[00:34:46] Mithun DJ: Now, what is the beauty of your lasso regression, or L1 regularization? We are adding the penalty term. How will the penalty term look like?
[00:34:55] Mithun DJ: The penalty term is nothing but the sum of the absolute values of the quotients. So, let's say you've run a regression, you have 10 variables, so you have approximately 10 quotient values. So typically, your loss function will have the RSS, or the residual sum of squares.
[00:35:09] Mithun DJ: plus some lambda times summation of modulus beta js, which means if you have 10 variables, you'll have 10 values for beta. I'm aggregating all of these things, and I'm multiplying by a factor called as lambda. So this is the penalty term
[00:35:24] Mithun DJ: that I'm using. It is the sum of the absolute value. So, let's say, for example, for price, you may get a negative value. Don't worry about it. When you're putting this in the loss function, you will always be using the modular sign, which means you are trying to minimize this.
[00:35:40] Mithun DJ: you're trying to minimize this modulus side, so it's only the modulus. We are considering the sum, and this lambda, I'll just… sometimes, when you are using this basically, we can use alpha, right? Or lambda, that is just a symbol. Now, what is the effect?
[00:35:55] Mithun DJ: So, the main thing is, the second term is the penalty term. So, that is what, if you're using this penalty term.
[00:36:02] Mithun DJ: So what you're effectively doing is, basically, you're building a lasso regression model. I repeat, if you're using this penalty term, the second term, it's a lasso regression. See, what is the effect of putting this penalty, lambda times this? So, it can shrink some quotients exactly to zero.
[00:36:20] Mithun DJ: So, if some of the coefficients can become exactly to zero, which means what? 0 multiplied by any of those variables will be 0, which means that variable itself is flushed out from the system.
[00:36:31] Mithun DJ: It is flushed out. 0. For example, if you forcefully put age to zero, what does this mean, basically? Age has a coefficient of zero. If you multiply zero by anything, you'll get zero. So, zero will be, age will become a useless variable. Gender may become useless. If you forcefully, basically, if it becomes zero. So.
[00:36:50] Mithun DJ: This technique is a
[00:36:52] Mithun DJ: is a great technique because it performs what is called as feature selection, right? Out of 15 variables, or 50 variables, you can reduce the complexity. So, if complexity is your headache.
[00:37:04] Mithun DJ: So, if you have 50, 60, 70 variables, which I was talking about, a very, very complex model, you can reduce this. How? You can reduce it using the first type of regularization, L1 regularization, how you're putting a penalty term, which is nothing but absolute of the sum… some of the absolute values of the coefficient.
[00:37:23] Mithun DJ: And if it shrinks to zero, if we're using the word if.
[00:37:26] Mithun DJ: If it shrinks to zero, it means that the effect of that variable becomes zero, and if it becomes zero, it's a feature selection. You are basically dropping a useless variable.
[00:37:37] Mithun DJ: That is what it is important. Sir, what is the great benefit of using a lasso regression? Please understand something.
[00:37:44] Mithun DJ: When we are using, using, lasso regression, it simplifies a very, very complex variable… model, sorry. It simplifies, the model complexity by removing unimportant and useless features.
[00:37:59] Mithun DJ: So, if you want to clean up your dirty data set, fill the data set with too many variables, use L1 regularization or LASSO regularization. So, long story short.
[00:38:10] Mithun DJ: These are the 3 or 4 important points that you need to keep in mind. LASSO is also called as L1 regularization. What is the biggest advantage? It encourages building sparse models.
[00:38:23] Mithun DJ: Sparse model means what? A model with too many variables becoming zero is called as a sparse model.
[00:38:29] Mithun DJ: Second, L1 is a fantastic technique
[00:38:33] Mithun DJ: Because it is robust to the presence of outliers. You may have outliers, L1 says, I don't care a damn, I can handle outliers. So that's the plus point. Third one is, it's a beautiful technique if you have a high-dimensional data, because you can do feature selection. How are you doing feature selection? By eliminating irrelevant features.
[00:38:53] Mithun DJ: Now, with any technique, you have both advantages as well as disadvantages. You can't have one technique which will solve all of your problems, all of the time, under all of your circumstances. L1 also has its disadvantages. First.
[00:39:05] Mithun DJ: It can be unstable.
[00:39:08] Mithun DJ: This is the most important thing. It can be unstable. That is, the weights can change significantly with small data changes. If you change the data even small, right, your weights can drastically vary. That is one of the biggest headaches of
[00:39:23] Mithun DJ: your L1 regularization. Second one is, it is not suitable for those features which have a high correlation, right? It selects one of the features arbitrarily. There is no science. One feature, it'll select arbitrarily. If X1 and X2 are correlated.
[00:39:37] Mithun DJ: This fellow does what? This L1 regularization? It may choose X2 or X1, depending upon its whims and fancy, and therefore, basically, there is an element of arbitrariness in its selection.
[00:39:48] Mithun DJ: So, these are the advantages and disadvantages of your L1 regularization, correct? Is this clear, guys, so far?
[00:39:56] Mithun DJ: Aditya…
[00:39:58] Aditya Banda: Yeah, can you, once explain how it's shrinking some coefficients to zero?
[00:40:05] Mithun DJ: See, what happens is that… see, this loss function is what? Error, correct?
[00:40:09] Aditya Banda: Right.
[00:40:10] Mithun DJ: loss function is errors. Now, RSS… I'm just giving a… the weight of RSS is basically, let's say it comes to some… around 50 units, I'm just giving a numerical example.
[00:40:20] Aditya Banda: Okay.
[00:40:21] Mithun DJ: Right? Now, if I'm basically saying that plus penalty means 50 plus 30 is 80, so the loss function is basically 80, which means
[00:40:30] Mithun DJ: if… Loss means what? It's basically the error sum of squares.
[00:40:34] Mithun DJ: Right? Correct. The unexplained variation part. So you are putting an additional penalty on that, which means any variable which has a small weightage, its weightage is basically changing to some extent, not a big, big, advantage, basically. It is pushing it closer and closer to zero.
[00:40:55] Mithun DJ: It's coefficient eventually, when you differentiate it basically, it becomes so close to zero that it'll… its effect is becoming useless.
[00:41:10] Aditya Banda: But the loss is increasing, right? When you put a penalty, loss is increasing.
[00:41:13] Mithun DJ: Overall value is increasing, but this term, if you look at the second term, basically, it's a combination of multiple variables.
[00:41:21] Mithun DJ: This beta J, if you say, there may be 10 variables. Some of them may be pushing it up, some of them may be pushing it down, depending on its magnitude.
[00:41:29] Aditya Banda: Okay.
[00:41:29] Mithun DJ: So this, fellow, this L1 regularization, basically, what happens is, basically, it depends upon… see, this is a hyperparameter. See, this lambda, it is upon you, basically. If you want a higher value, if you choose a higher value of lambda, bigger penalty.
[00:41:45] Mithun DJ: So, you are basically, those variables, basically, whose contribution is even slightly, if it is dropping, you are putting a bigger penalty. If you basically choose a smaller value of lambda, it means that there's a smaller weightage of penalty, so you are
[00:41:59] Mithun DJ: being more liberal to it. So it depends upon what is the choice that you are coming up with. It is not that for all values of lambda, basically, you'll get it as zero. So what we do is, basically, we try to incrementally increase lambda, so that, basically, some of these useless features, basically, are pushed to zero.
[00:42:18] Aditya Banda: Okay.
[00:42:23] Mithun DJ: Okay.
[00:42:25] Mithun DJ: Now, second one is basically what is called as a ridge regression base.
[00:42:30] Mithun DJ: So, this is what is called as lasso L1. Second one is what? Second one is a technique which is called as ridge regression. Now, what is the meaning of your ridge regression? Now, ridge regression is a different technique.
[00:42:43] Mithun DJ: Here, also, we have a penalty term. We are attacking the penalty term itself, we are putting the penalty term to RSS. How?
[00:42:49] Mithun DJ: Unlike in the first case, basically, wherein you use the sum of the
[00:42:55] Mithun DJ: some of the quotients. Here, what am I doing? I am taking the sum of the squares of the quotient, so it is beta j squared, summation of beta j square, which means if you add this, basically, if you take the square of it.
[00:43:08] Mithun DJ: Correct, if the coefficient is 2,
[00:43:11] Mithun DJ: Right? If the coefficient is 2, and by the time you put it as a penalty, it'll become square of that. 2 square is 4.
[00:43:18] Mithun DJ: And the penalty is, let's say, even a small penalty, 4 times you're putting a penalty on a particular variable.
[00:43:24] Mithun DJ: Four times, because you're squaring it up, so the penalty becomes, bigger and bigger.
[00:43:31] Mithun DJ: That's the thing. It's like Dohuna lagan, char guna lagan, satra guna lagan, if you just think of it, right? So you are putting a bigger and bigger penalty for a misbehavior of the model. So, that is what is happening, and also you have to choose the L here.
[00:43:46] Mithun DJ: It shrinks the coefficient towards zero, but never exactly zero. So this is one important point. In the case of, basically, your,
[00:43:56] Mithun DJ: In the case of, basically, your, your L2 regularization, this fellow here, it is basically going close to zero, but it never touches zero.
[00:44:18] Mithun DJ: Is this understood, guys?
[00:44:24] Ankit Sood: I'm not able to actually relate to it, because we are saying we are adding something to something, and then it is tending to zero.
[00:44:30] Mithun DJ: It is not something to something. This is RSS's residual sum of squares. So, in a typical regression model, you're trying to minimize the error sum of squares. You're adding a penalty to the sum of the squared quotients.
[00:44:45] Ankit Sood: But that's still an addition, right? Like, I'm still doing addition on two positive numbers here.
[00:44:49] Mithun DJ: Yes. So the penalty is becoming more, right?
[00:44:53] Ankit Sood: But how is it tending to zero, is what I'm not able to relate. After addition, how is it becoming zero?
[00:44:58] Mithun DJ: No, it is… this is the loss function. This loss function and all, you have to differentiate it. Just because you're adding it, it will not become zero.
[00:45:05] Mithun DJ: This is the loss function, and this loss function further your differentiating, and then, basically, you're arriving at the numbers. So, in the second iteration, basically, when you do, basically, the overall values decrease, that's what we are seeing.
[00:45:20] Ankit Sood: Okay, is there an example with which we can understand? Because I was thinking loss function is nothing but nothing but the line of slope Y equal to MXC plus C, right? This is what we are trying to minimize in case of regression.
[00:45:33] Mithun DJ: Yeah, Y minus Y hat.
[00:45:36] Mithun DJ: Actual minus predicted, the whole.
[00:45:38] Ankit Sood: Correct.
[00:45:39] Mithun DJ: That is what we are trying to minimize.
[00:45:40] Ankit Sood: Minimize, yeah. Correct. So, where Y is your actual value, and… sorry, Y might be your predicted value, and Y hat is… sorry, Y is your theoretical value, and Y hat is your practical value, correct?
[00:45:52] Mithun DJ: Correct, correct, correct.
[00:45:56] Ankit Sood: But still, I think with an example, I would be able to relate more to it. Currently, I'm not sure.
[00:46:01] Mithun DJ: I have few examples, I have few examples.
[00:46:07] Aditya Banda: Yeah, just one question. So, when you're trying to minimize… now you're adding penalty, and the model is trying to minimize the overall loss, which is including the new penalty, and when it tries to minimize, that is when you're saying differentiation comes into picture.
[00:46:23] Mithun DJ: Yes.
[00:46:23] Aditya Banda: And the model is trying to, basically minimize as much as possible the penalty part of it. When it does, some of the coefficients are tending towards zero.
[00:46:34] Mithun DJ: But not exactly zero.
[00:46:35] Aditya Banda: Not exactly zero. And it is random. It decides what coefficients to shrink and what not to shrink.
[00:46:42] Mithun DJ: It is not actually random. It is going to put a penalty to those variables which are causing the maximum variation.
[00:46:50] Mithun DJ: So, for example, if I have got 10 variables, and let us say my model is learning only from age and gender, the rest of the variables, it is not learning at all.
[00:47:00] Aditya Banda: Okay, so it will… it will check the collinearity, basically.
[00:47:03] Mithun DJ: Yes, yes, it'll check all of those things, and those variables… you cannot build a model with only two variables, correct? Because those two… you're heavily… you can't put all your eggs in one basket. How can you have a model which is just learning only from two independent variables, not from all the variables?
[00:47:21] Mithun DJ: So, other variables, their contribution may be, you know, very tiny, 6, 7, 8, 9.
[00:47:27] Aditya Banda: Right.
[00:47:27] Mithun DJ: Which is…
[00:47:29] Aditya Banda: Got it.
[00:47:29] Mithun DJ: So it'll try to shrink those variables. Yes, it'll try to shrink those variables, but it'll preserve it. Earlier method basically did not preserve it, because it pushes the coefficient to zero. But here, it tends closer to zero, but never actually becomes zero.
[00:47:45] Mithun DJ: So, it'll be there, but its effect is reduced. So here, we are saying, basically, the advantage is this fellow can handle multicollinearity well. It keeps all the features, but reduces its impact. You're not losing all the features.
[00:47:59] Mithun DJ: Before doing, if you have 10 variables, after doing also, you'll have 10 variables.
[00:48:05] Mithun DJ: the effect of the variables, basically, will be reduced to a large extent. So, before, age may be showing 16, 17. Now, after you take it through L2 regularization, the value from 16 may drop to 7 or 6.
[00:48:20] Mithun DJ: Right? So its impact has… so the over-prediction part will reduce to a large extent.
[00:48:25] Aditya Banda: Sure.
[00:48:27] Mithun DJ: So the earlier was a feature selection technique. This is not a feature selection technique. It is a technique which is used to handle multicollinearity, which we were talking about earlier.
[00:48:41] Mithun DJ: Now, the third one is there, guys. This is called as ElasticNet.
[00:48:45] Mithun DJ: This algorithm is called as ElasticNet. So, what do you have here? You have L1 regularization, L2 regularization, and ElasticNet.
[00:48:54] Mithun DJ: Now, what is this L… what is this ElasticNet? It is nothing but a combination of the earlier two techniques, wherein you have both L1 regularization as well as L2 regularization.
[00:49:06] Mithun DJ: So… You have to your inbuilt equation, to the loss function, basically, you have lambda 1 plus summation of
[00:49:13] Mithun DJ: modulus of beta j, so it's a combination of all the model weights plus lambda 2 into some of the sum of the squares of all of the coefficients.
[00:49:25] Mithun DJ: So the second term is the penalty parameter which we are picking up from your Elver regularization. The third term here, lambda 2 into beta j squared, basically, this is basically nothing but the penalty term that you're picking up from L2 regularization. So, when you're using a combination of both of these things.
[00:49:45] Mithun DJ: you use loss function. You use a different loss function here in the case of ElasticNet. When do you use it? Let us say you have too many features. Also, your data is suffering from multicollinearity. You are having two problems. It's a complex model.
[00:50:00] Mithun DJ: Plus, there is a lot of correlation systems, so use both of these penalty parameters.
[00:50:05] Mithun DJ: Here, you can use, rigid regression when you have… when you may not have too many variables, you may barely have around 10 or 12 variables, and you see that there's a lot of correlation and statistics within those 10 or 12 variables. Then, to address the problem of multicollinearity, you can use L2 regularization.
[00:50:24] Mithun DJ: Now, here, you have 50-60 variables. L1 regularization can be used when, by definition, you have a high-dimensional data set, 50, 60, 70 variables.
[00:50:33] Mithun DJ: You are not so concerned about the correlation and systems, but you are deeply concerned about to having too many features.
[00:50:41] Mithun DJ: When you have too many features and you want to do some variable selection and systems, Lasso comes into picture. So, if I have to just, if you have to take away one small message, forget the mathematics and systems,
[00:50:55] Mithun DJ: As of now, if you have to just look at the application point of view, L1 is used for performing feature selection.
[00:51:02] Mithun DJ: L2 is used to address the problem of multicollinearity. L3 is used to attack both of these problems, and when you use L3, basically, it is nothing but a case of ElasticNet.
[00:51:24] Mithun DJ: Is it clear?
[00:51:28] Aditya Banda: Yeah.
[00:51:33] Sacheen Adavinavar: Is…
[00:51:35] Mithun DJ: Okay, this is… don't worry about the mathematics, think of the usage here. Now, what is the difference among all of these things? Right? There are… I've just shown you a very simple way in which you can comprehend this.
[00:51:49] Mithun DJ: See, what is the penalty for each of these things, right? For L1, you have a penalty, for L2, you have a penalty, and for, elasticsNet also, you have penalty. The coefficients, some of them become zero here.
[00:52:01] Mithun DJ: Here, none of them become zero. Here, under elastic nets, some of them may become zero. That's what we're seeing. Feature selection. Feature selection, the best technique which handles feature selection for high-dimensional data is your L1 regularization.
[00:52:16] Mithun DJ: L2, its job is not to do feature selection.
[00:52:20] Mithun DJ: ElasticNet partially addresses the problem of feature selection, because you have that penalty parameter, which is able to take care of the feature selection. Now, in terms of model stability and such things, which one is the most stable model? L1 is less stable.
[00:52:37] Mithun DJ: L2 is more stable, and your elastic net is more of a balanced problem. And lastly, I'm summarizing by saying, when do you use what techniques?
[00:52:48] Mithun DJ: If you have sparse data, you… and you want to basically do some feature reduction systems, use multicollinearity. If you have sparse data, and you want to do some feature reduction systems, use the… if you want to do feature reduction.
[00:53:02] Mithun DJ: use L1. If you have multicollinearity, that is a lot of 0.8, 0.9 kind of correlation and such things for a lot of variables, then you can use Ridge regression. But if you have both of these problems, you have to use elastic fit.
[00:53:30] Mithun DJ: Correct? If you want, explanation, guys, I'll just share with you a couple of Excel sheets that will also help you.
[00:53:36] Mithun DJ: We have covered this, we have covered this, so, the mathematics behind some of these things, guys, which we have seen.
[00:53:43] Mithun DJ: For Rich and Lasso, what are the parameters?
[00:53:46] Mithun DJ: Right?
[00:53:49] Mithun DJ: just understand this concept, guys. Typically, you need to understand one, one concept, and this is called as bias and
[00:53:58] Mithun DJ: Model bias, as well as, error.
[00:54:05] Mithun DJ: you know, fluctuation. How does it fluctuate? You can see here, just think of it. The horizontal axis reflects the model complexity, which means you have a more and more complex model, from 20 to 30 to 40 to 50 to 60 to 80. What happens as you have a complex
[00:54:25] Mithun DJ: Model is, when you… for a very, very complex model, typically, what happens is, basically, you may have, error and such things, because you're using so many variables, your… have a look at this, this blue term, your error goes on reducing here.
[00:54:43] Mithun DJ: Correct? As you make the model complex, look at the blue line, error goes down, which means accuracy increases.
[00:54:50] Mithun DJ: But one more problem is there, right?
[00:54:52] Mithun DJ: Now, for a less complex model, which means if you take very, very few features, basically, your error is at the highest level.
[00:54:59] Mithun DJ: Correct? Now, observe what happens to the green line. Now, as you take more and more features, the variance becomes very, very high, which means the quotients basically become very, very
[00:55:10] Mithun DJ: high. And here, as you take fewer, variables, basically, maybe only 2 or 3 independent variables, your variance is very, very low. So this graph is called as variance bias relationship. So it's like a trade-off that we have to think of.
[00:55:24] Mithun DJ: Because nobody can give you a silver bullet saying that you use two, three, four variables. What is the optimal number? It just becomes a case wherein this green line and blue line at some point meet each other. Green line and blue line meet at each other. You can see here, let's say it is meeting at this point.
[00:55:42] Mithun DJ: This dotted line here indicates the optimal model complexity. Optimal model complexity means I'm not working with 100 variables, I may be happy with just, you know, 20, 30 variables here. Why? At this stage, my… look at this, my error has dropped
[00:55:59] Mithun DJ: to an appreciable level, and at this stage, my variance also has not increased out of proportion. So, I want a model without too high a variance. I want a model which is less biased. What is that optimal number? That is basically
[00:56:17] Mithun DJ: it could be any number. It depends on… it depends from one data set to another. So, this is that point, basically, wherein you have a less complex model and a more stable model also.
[00:56:28] Mithun DJ: It determines the optimal model complexity.
[00:56:48] Mithun DJ: Is this clear, this picture?
[00:56:50] vinit shah: Yes.
[00:56:52] Mithun DJ: Correct?
[00:56:53] Mithun DJ: So, just take a look at this. There are three scenarios when you speak about it. One is a scenario of underfitting, we'll say, alright? Can you say this? First case is a scenario of underfitting. Second one is a case of appropriate fitting, and this is how, basically, an overfitting case will look like.
[00:57:13] Mithun DJ: Underfitting, or fitting. I hope you're able to understand the difference.
[00:57:25] Mithun DJ: Are you able to understand the difference?
[00:57:27] vinit shah: Yes.
[00:57:29] Mithun DJ: See, this is a case wherein it's called as appropriate fitting, because you have something like a parabolic curve, and you may have a few misclassified cases.
[00:57:36] Mithun DJ: But in a very, very overfitting case, you can see here, it is trying to go to every corner here and somehow, basically, you know, try to sup… it is trying to separate the plus side.
[00:57:49] Mithun DJ: from the circles.
[00:57:51] Mithun DJ: Correct? You may have, basically, frauds from non-fraud cases, pluses and the circles it's just trying to… look at the kind of effort that it is putting. This kind of a model, you can't replicate it because we want the model to be generalizable. We don't want a model to be sample-specific.
[00:58:09] Mithun DJ: This is underfitting because there are so many cases, basically, when you put a line like this, we expect a majority of the cases, the pluses to be below, and majority of the circles to be above, which is not happening. There are so many circles, basically, which are below, and there are so many pluses which are above. So this is an underfitting case, the accuracy will be less.
[00:58:29] Mithun DJ: Right? We are trying to find out that optimal mix. So, this picture is usually important from an examination point of view, as well as from an interview point of view, because this is called as the bias and error trade-off. If they ask you in an interview or any time, do you know what is the meaning of bias and error trade-off?
[00:58:46] Mithun DJ: you should be in a position to at least draw this and explain what is happening with basically a highly biased model and a highly veering model, okay? So here, you have four diagrams.
[00:58:59] Mithun DJ: Please do let me know if you're able to understand this. These four scenarios are very important, because to address these four scenarios itself, we are talking about regularization and sustenance.
[00:59:25] Mithun DJ: Are you able to understand this?
[00:59:32] vinit shah: Shoes like that.
[00:59:45] vinit shah: Not exactly, but kind of, I think, yes.
[00:59:49] Mithun DJ: Please try to… please try to understand this. In this entire, class that I'm taking.
[00:59:54] Mithun DJ: This is the most important slide, and this is the most important slide. This is called as mod bias variance trade-off. This is also important, because when you build a model, what this diagram is saying, I'm just explaining what the meaning of the diagram is.
[01:00:09] Mithun DJ: There are 4 cases which are possible.
[01:00:12] Mithun DJ: Your model can have low bias, low variance, which is one case.
[01:00:16] Mithun DJ: Which is what it is capturing here. Second case is your model… this is a diagram, wherein you can… your model can have low bias, high variance.
[01:00:24] Mithun DJ: Third case is, this is the case, this is the diagram, wherein your model can have high buyers and low variance, and the fourth case is basically a model which has high bias and high variance, correct? So these are the four cases that we want.
[01:00:38] Mithun DJ: Now, the worst case that you… when you build a model, and when you look at all of these things, and when you see a pattern like this, high buyers and high variance. So this is a big, big problem.
[01:00:49] Mithun DJ: Because, just think of this diagram, and forget the statistics part, just think that you are a shooter or an archer who's aiming, and you're being given a target. This is the bullseye that you're supposed to hit.
[01:01:01] Mithun DJ: Correct? When you basically lift that, lift that bullet, lift that, gun, and when you try to fire, and when you try to aim at this yellow spot, see, barely one of them, is hitting. Many of them are way off the mark.
[01:01:14] Mithun DJ: Right? So this is a problem, correct? Why is this a problem? Neither are you accurate.
[01:01:20] Mithun DJ: you're not at all accurate. Plus, you're scattered everywhere, basically. Therefore, basically, it's a case of high bias, high variance, right? This is the worst scenario.
[01:01:30] Mithun DJ: in, that you can be in. Look at this scenario. This is the best scenario
[01:01:35] Mithun DJ: Because majority of the bullets that you have fired have hit the bull's eye. Almost all of the green dots are in the yellow spot itself, so you are very, very accurate. Low bias means high accuracy. And look at the variance. It is not scattered everywhere. It is still within a narrow space itself.
[01:01:53] Mithun DJ: hitting, therefore it means that there is low variance. So this is an idealistic situation wherein you have low bias, low variance. We want the models to be in this direction.
[01:02:02] Mithun DJ: But typically, our models can be in any of the other three quadrants. It will not be in this. Now, look at this. This is a case wherein you have low bias, but high variance.
[01:02:13] Mithun DJ: Your accuracy is still decent enough, not bad. You can see here, 3, 4, 5, 6, 7. The 7 bullets that you have fired are still hitting the bullside, more or less. There are 3 off the mark. So, bias is low.
[01:02:26] Mithun DJ: But the variance, it is scattered everywhere. These data points are scattered everywhere, therefore we are seeing high variance. Have a look at this third scenario, high bias, because none of them are in the bullseye. You're not able to hit the bullseye at all, it is not…
[01:02:38] Mithun DJ: you're not able to strike the yellow portion at all. So, high bias and low variance. Though you've made a mistake, the mistakes are all basically… it is not scattered, it's all basically here itself. So it's a case of high bias, low variance, but this low variance is also not helping you.
[01:02:57] Mithun DJ: Right? And the last one is the worst scenario. I hope this makes sense.
[01:03:06] Mithun DJ: You said it, yeah?
[01:03:07] Aditya Banda: Yeah, one question. Why is low bias, high variance.
[01:03:12] Mithun DJ: A lot worse.
[01:03:13] Aditya Banda: high variance.
[01:03:15] Mithun DJ: low bias. Accuracy is basically, low bias means accuracy is high, and the high variance means, basically, for different test data set, it keeps its accuracy
[01:03:26] Mithun DJ: keeps varying. On the one test data, it may be 60, next time you may get 50, the third data set, it may be 80. So you do not know for sure whether you should trust the model or not.
[01:03:36] Aditya Banda: Okay, got it. And for the first figure, low bias and low variance, I'm assuming when you say low, it's not absolutely low, right? It's optimally low.
[01:03:44] Mithun DJ: Yes, we don't want a model with 100% accuracy. If it's 100% accurate, then there's something wrong.
[01:03:51] Aditya Banda: So, the first diagram is basically intersection of the graph that you have shown in previous slide.
[01:03:55] Mithun DJ: Yes, yes.
[01:03:56] Aditya Banda: Okay.
[01:04:02] Mithun DJ: Are you able to understand these four scenarios, guys, everybody?
[01:04:05] vinit shah: Yes.
[01:04:20] Mithun DJ: I'm just giving the explanation just in case you forget. You can quickly look at this. These are all standard questions, guys. This is what tells whether a person knows modeling and such things.
[01:04:29] Mithun DJ: not running the code, because I know people are very excited about running code and such things. It's important, very, very important to run the code and give the result, but more than that, do you understand the concept? High bias, low variance means what? A model that has a high bias and a low variance is considered to be underfitting.
[01:04:46] Mithun DJ: This is a case of underfitting problem.
[01:04:49] Mithun DJ: High variance, low bias means what? It is considered as overfitting.
[01:04:53] Mithun DJ: High bias, high variance. A model with high bias and high variance cannot capture the underlying patterns. You're not able to capture the underlying patterns, and is very, very sensitive to training data changes. If you make small changes to the training data, you know, you may get drastically different results.
[01:05:11] Mithun DJ: On average, the model will generate unreliable and inconsistent predictions. You can't rely on it, and inconsistent predictions. Low bias and low variance. A model with low bias and low variance can capture data patterns. One is, it can capture low data patterns.
[01:05:26] Mithun DJ: Therefore, basically, it's highly accurate. Highly accurate models basically have low bias, and it can handle variations in the training data.
[01:05:34] Mithun DJ: We can handle a lot of variations in the training data.
[01:05:37] Mithun DJ: This is the perfect scenario for a machine learning model. So this is an idealistic scenario, or a perfect scenario, where it can generalize well to the unseen data and make consistent, accurate prediction. However, in reality, this is not feasible. So we want to be in the fourth case. Idealistic scenario is never possible in real life, guys, that's what we're saying.
[01:06:02] Mithun DJ: Is this kid,
[01:06:04] vinit shah: Sir, I have one doubt. So, all these models are create… like, in reality, we will have all these kind of models present, right? So, there might be a use case even for these high bias, high variance, and these high bias, high variance as well, right? Not everything could be an optimum model.
[01:06:21] Mithun DJ: So what is optimal?
[01:06:22] Mithun DJ: See, when you're building a model in our fraternity, there's a saying.
[01:06:26] Mithun DJ: In your entire career, just go with this, go with this model. All models that you build in your life are wrong.
[01:06:36] vinit shah: Got it.
[01:06:37] Mithun DJ: Some models are useful.
[01:06:38] vinit shah: cooking.
[01:06:40] Mithun DJ: That's it.
[01:06:42] Mithun DJ: That's a condition.
[01:06:43] vinit shah: So my question is this, right? Are there scenarios where we can, you know, have a model which, say, follows one of these cases, like, say, a high bias and a low variance model? Can there be…
[01:06:56] vinit shah: Some scenarios where such a model is actually, you know, say, useful and more of a quick queen kind of a thing.
[01:07:04] Mithun DJ: So… More, so, 95% of the models that you build will be in the first three scenarios.
[01:07:11] vinit shah: Oh, okay.
[01:07:12] Mithun DJ: In the rarest of the rare circumstance, if you have built a model, basically, very rarely you may get a fourth scenario.
[01:07:19] vinit shah: It's more like a theoretical scenario in the first day, or, like, practical scenarios.
[01:07:22] Mithun DJ: In a real-time scenario, you'll never get a fourth scenario. Something is wrong somewhere. Either it is a sample data, cooked-up data, or you have basically fudged the data, or you're basically artificially inflating the results. Unless you do some JUGARD, you'll not get the fourth result.
[01:07:39] vinit shah: Okay.
[01:07:41] Mithun DJ: Anything which is idealistic is just, you know, it is just, practically not possible.
[01:07:48] vinit shah: Okay, and in these three… in these three, what you're saying, right, what is the one which…
[01:07:52] vinit shah: Commonly, gets created, like, majority of the scenario, which fits of these three?
[01:08:01] Mithun DJ: So, we will have a low bias, which means we'll… typically, with the latest algorithms and such things, you will find one or the other model which is highly accurate.
[01:08:11] Mithun DJ: Same scenario. But he's very inconsistent.
[01:08:15] vinit shah: Oh, God.
[01:08:15] Mithun DJ: most popular scenario.
[01:08:17] vinit shah: Okay.
[01:08:20] vinit shah: Okay.
[01:08:24] Mithun DJ: So I'm just trying to summarize the discussion, guys. A few points I've given.
[01:08:29] Mithun DJ: To summarize, regularization improves the model generalization by reducing overfitting.
[01:08:35] Mithun DJ: Correct? Regularized models learn the underlying patterns, which overfit models
[01:08:42] Mithun DJ: Which overfit models memorize noise in the training data. A overfitting model is actually learning noise. It's memorizing
[01:08:50] Mithun DJ: Nice.
[01:08:51] Mithun DJ: So, typically, you have seen LASSO L1, right? This regularization simplifies the model. So, if you have a very complex model and you want to simplify, somebody may ask you, right, you have 60 variables, how do you simplify? Quickly, you should be able to say it's a L1 regularization case, and it also helps in the interpretability. If you have 60, 70 variables, how will you interpret anything?
[01:09:11] Mithun DJ: Just ask that also, right? Sometimes more, the merrier that follow, that does not follow, right? It improves the interpretability, right, by simplifying the model, by reducing the quotients of less important features to zero.
[01:09:26] Mithun DJ: Got it.
[01:09:27] Mithun DJ: Regularization improves model performance. Yes, this is one advantage, you can improve the model performance. How? By preventing excessive weighing of outliers. So, many times, if you have outliers and systems, automatically, the weight
[01:09:41] Mithun DJ: that will be assigned will be inflated because of the presence of outliers and systems, so you need to correct it, so that's also one more use case for introducing a penalty term and systems.
[01:09:52] Mithun DJ: The fourth point is regularization makes the model stable. We can have stable models. You have seen the… you have seen, basically, when can you have a stable model. The L2 case basically provides a stability, stable model.
[01:10:07] Mithun DJ: It reduces the sensitivity of the model outputs to minor changes in the training dataset.
[01:10:13] Mithun DJ: The fifth point which I'd like to say is regularization prevents models from becoming over-complex. This we have seen. Regularization can help handle high multicolinity. I'm just giving the explanation. High correlation between the features.
[01:10:26] Mithun DJ: Then, basically, regularization introduces hyperparameters, like alpha and lambda. Alpha is the hyperparameter for L1, lambda for, you know, L2. It's under your control. If you give a higher value of alpha, it means more penalty. Lower value of alpha, less penalty.
[01:10:43] Mithun DJ: Correct? And you can control the strength of regularization. It's a highly regularized means, basically, it's a highly disciplined model. You're basically, like, you know, holding a stick and you're beating the child, saying, you better learn this, you better learn this, you better learn this, you're very strict. So.
[01:10:57] Mithun DJ: you have to, you have to just be slightly careful with the value of alpha and lambda that you're giving, correct? If you allow the child to be very flexible, it will not study anything. So, it's a bit of a trade-off.
[01:11:09] Mithun DJ: There is no silver bullet, nobody can say this is the best. You have to experiment, and after 5, 6, or 7 iterations, you'll be able to arrive at the best.
[01:11:18] Mithun DJ: model, right? Regularization promotes consistent model performance. We are trying to encourage the model to be as consistent as possible.
[01:11:27] Mithun DJ: It reduces the risk of dramatic performance changes when encountering new data.
[01:11:35] Mithun DJ: Okay.
[01:12:24] Mithun DJ: Can you see my Google Collab place?
[01:12:29] Mithun DJ: Nobody see this.
[01:12:30] vinit shah: So you can see.
[01:12:31] azad choubey: Yes, please.
[01:13:09] Mithun DJ: Have you worked on the Boston housing data set, guys, everybody?
[01:13:22] Aditya Banda: Not sure.
[01:13:24] Mithun DJ: Okay, I'm just importing some of the libraries here.
[01:13:27] Mithun DJ: Like, brandas, Seabonds, things.
[01:13:52] Mithun DJ: This is the dataset, right? This is how the dataset looks like. I'm just,
[01:13:56] Mithun DJ: working on a very popular data set which is used in the data science community. They've gone to Boston. There are two people, basically, Harrison, I think, Larry and Harrison. It's a pretty old data set. They've gone to a place called as Boston in US.
[01:14:10] Mithun DJ: And they've, tried to find out the aggregate, housing prices and such things, like crime rate, how much is it, is that an industry, whether this is Charles River or not. NOX is nothing but pollution level, 53%, 46%, so on and so forth.
[01:14:24] Mithun DJ: Different localities, okay? R means average number of rooms. Then you have age, distance, RAT, tax, PT ratio, pupil-to-teacher ratio.
[01:14:33] Mithun DJ: B is nothing but blacks, the number of blacks or colored people, and then you have 4.98, which is the percentage of lowest status of population, right? 9% means high.
[01:14:42] Mithun DJ: And MEDV is, I think what? Median Value of Housing Place. So, this is the dependent variable, MedV. So, it is $24,000, $21,600, $34,700.
[01:14:52] Mithun DJ: So the last variable that I have is the independent… sorry, dependent variable, and the rest of the variables are independent. So first, let's build a linear regression model, and then we can see some of the problems. We can see some of the challenges that we can address, okay?
[01:15:07] Mithun DJ: So, as part of the EDA technique, exploratory data analysis, we can do a correlation matrix. You can just see here, this is the correlation matrix, okay? I know that there are almost, like, what, 13 variables, I guess, correct? 13 into 13, 169 values are there. Bit difficult to make sense of this when you just have a correlation matrix.
[01:15:25] Mithun DJ: So, what is the trick that we use? We use, basically, heatmap. We convert this into a heat map.
[01:15:31] Mithun DJ: Heatmap means it'll color code each and, it'll color code the correlation values.
[01:15:35] Mithun DJ: So have a look at this, it'll just color code. With, positive values being blue. Look at this, this is the explanation. Blue means positive, this thing, and then when you have red, it means that it's negative correlation, and if these codes basically are a sort of whitish in color, it means the correlation is weak.
[01:15:53] Mithun DJ: Now, I have given it like this, so if you see, annotation is equal to false, I've given, convert this into true.
[01:16:07] Mithun DJ: Are you able to see this?
[01:16:19] Mithun DJ: Everybody?
[01:16:21] Sonam Manwal: Yes.
[01:16:23] Mithun DJ: So, this is a correlation matrix that we are basically, converting, right? So, what can you understand when you look at this correlation matrix?
[01:16:31] Mithun DJ: So, one of the things that you can basically understand
[01:16:34] Mithun DJ: When you look at these correlation metrics, is… Any ideas?
[01:16:46] Deepak Bobade: Yeah, we can understand whether the variables are positively related or negatively related.
[01:16:53] Mithun DJ: Right, more than that, look at the strength of the correlation. So, this is a…
[01:16:56] Deepak Bobade: Definitely a strength.
[01:16:57] Mithun DJ: Yeah? MedV and all, right? Some of these variables, like 0.7 and 0.74, shows that there is correlation, right? Not bad, there is enough correlation. But this is okay, I don't mind about this, because this is the independent variable being related. But where does the worry come?
[01:17:13] Mithun DJ: So… This is cool warmer. This does not help me. I'll choose, red.
[01:17:20] Mithun DJ: Blue and green, or red.
[01:17:24] Mithun DJ: And alright, and, window.
[01:17:30] Mithun DJ: So that you understand this better.
[01:17:34] Mithun DJ: This is better, because I think green is positive, I guess, right?
[01:17:38] Mithun DJ: Yeah, green is positive, makes sense. And, red ones are basically, no?
[01:17:44] Mithun DJ: Negative correlation. Have a look at the rest of the cells, basically. When you see, suddenly I see 0.91. This is between variables tax and rad.
[01:17:53] Mithun DJ: Right, so… bit of a green flag here, right? 0.73. So there is a bit of a heavy correlation, right? 0.67 strong correlation, 0.72. 0.6 and 0.6 against strong correlation, 0.63. 0.58, 0.76 against strong correlation.
[01:18:09] Mithun DJ: Right? 0.67 and such things when you have… it's an indication that you have strong correlation, right? Strong correlation what? Between the independent variables. How many variables do I have?
[01:18:19] Mithun DJ: So this is a… this is a problem, typical problem of multicol… this is a typical problem of multicollinearity. Look at the number of variables, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13. You have 13 independent variables. 13 independent variables is also… it's still fine.
[01:18:38] Mithun DJ: It is not a very complex model as such. 13 here and there, if you just look at the p-value also, you can eliminate many of these unwanted variables.
[01:18:47] Mithun DJ: So, it may not be a very complex model, but there is some level of complexity. But more than complexity, the bigger headache that we have is basically in terms of multicolinearity, right? So, this is what I can… there is a problem of multicolinity in the dataset that I'm working on.
[01:19:04] Mithun DJ: Let's just go with the flow here. What we are gonna do is, basically, we are going to create our independent and dependent variable.
[01:19:20] Mithun DJ: Right? So, I'm just using the train, test, pretenses things, and here you have a linear regression model.
[01:19:37] Mithun DJ: Model is equal to linear regression, model.fit, and then, basically, we can go ahead and generate some predictions.
[01:19:45] Mithun DJ: We can generate some predictions.
[01:19:50] Mithun DJ: You can just see here.
[01:19:52] Mithun DJ: These are the predictions by underscore red.
[01:19:57] Mithun DJ: So these are the predictions, predicted values.
[01:20:00] Mithun DJ: Okay.
[01:20:02] Mithun DJ: Now, if you want, you can do some evaluation of the model, you can compare the performance of the model.
[01:20:08] Mithun DJ: by looking at the R squared. R-squared is the coefficient of determination. R-squared here is 66%.
[01:20:14] Mithun DJ: That is what I'm getting.
[01:20:15] Mithun DJ: For the tested dataset, okay? It is 66%.
[01:20:19] Mithun DJ: So, just give me one moment, I'll just open up another file.
[01:20:24] Mithun DJ: Wherein you have, linear regression, okay? Let me just open this up.
[01:20:33] Mithun DJ: I've done a detailed analysis there, okay?
[01:20:36] Mithun DJ: I'll just show you some of the results that I got when I did linear regression.
[01:20:46] Mithun DJ: Can I see this case, everybody?
[01:20:52] Mithun DJ: This is another workbook, earlier I had done, almost on the same lines, but I had done extra things. I wanted to show you some things, some graphs, and this is what I've done.
[01:21:00] Mithun DJ: Housing price predictions and rate. Can you see the coefficients here?
[01:21:06] Mithun DJ: Can you see the quotients here?
[01:21:10] Mithun DJ: Are you able to see the questions, guys, for each of the features? Yes or no?
[01:21:14] Mithun DJ: See ya.
[01:21:15] Deepak Bobade: Yes.
[01:21:15] Mithun DJ: See, crime rate and all, it is okay in 0 points. Suddenly, for CHAS, it is getting 2. 2 is still fine. Have a look at pollution level, minus 21 point something.
[01:21:25] Mithun DJ: are able to understand, which means if you're going to use these variables case, what do you think will happen? Your model will be, predicting based on NOX level, that is, the pollution level is given so much of high weightage.
[01:21:36] Mithun DJ: This is a problem. It is like if India scores 300, and Virat Kohli scores 150, and Rohit Sharma scores 130,
[01:21:45] Mithun DJ: What does it mean? In an entire team, only 2… an entire team of 11 players, only 2 people are playing very well. So your team itself is dependent on Virat Kohli and Rohit Sharma, the rest of the people are not playing at all, because the rest of the people, 9 people, their contribution is barely 20 or 30 runs. So that's a problem that you have.
[01:22:02] Mithun DJ: Correct? So you don't want, 21 and such things in relation to others, correct? Here, 20, this is 0.00. There are so many values, you can just have a look at this. Their weightage is 00 and such things.
[01:22:14] Mithun DJ: So, if I were to use this, I'm just getting the feeling that my model predictions will be generated based on the learnings that my model has got based on Charles River and presence of Charles River and NOX. So, this is definitely a candidate for regularization.
[01:22:31] Mithun DJ: I'm just sorting it so that you get a feel for it, right? The top 3 variables, 21, 1, and 1. 1 is still okay, I'm not too worried, but 21 is an eye-catching number, okay? It's a glaring number, we can't… cannot have this. Then you can see chass here, 2 and, average number of rooms.
[01:22:47] Mithun DJ: So, I think the model is just learning based on average number of rooms and your pollution level. Not a good sign, not a healthy sign.
[01:22:55] Mithun DJ: So, you can just look at this. Here, I had built for test data, which is 78%. I was just getting test data, 78%, and then train data, 71%, which is okay, it's almost consistent in that sense.
[01:23:09] Mithun DJ: 20, right? 78, 77, barely the difference is in single digit. That is still okay. Not a… we don't, worry too much if it is in single digit, okay?
[01:23:18] Mithun DJ: Then some predictions and such things are generated. Look at the beauty of this case. Are you able to see this?
[01:23:24] Mithun DJ: There's a package which I've used called as SMS.ols. In one shot, it'll give you the regression output. Have a look at this.
[01:23:33] Mithun DJ: Now, I told you, if you remember, that, you know, if you have borderline cases, like, you know, 12 variables, 13 variables, complexity problem can still be avoided. Why do I say this? This is the model coefficients table. You have, against each variable, you have the coefficients, standard error, T, and p-value.
[01:23:49] Mithun DJ: See, for many of these things, the p-value is bigger than 0.05, which means it is insignificant. For example, indus is insignificant variable.
[01:23:57] Mithun DJ: So, you can kick this variable out. So, your complexity will automatically reduce. 0.98 is bigger than 0.05. So, kick this variable out.
[01:24:05] Mithun DJ: So when you start kicking out these variables, it'll barely come to 9 or 10 variables, right? So this is also one more way in which you can, you know, kick out this variable, right? So if you kick out this variable, many of the variables will automatically get left out, and if
[01:24:21] Mithun DJ: if just in case this age and such variables, age and your index basically are kicked out, your earlier correlation metrics will also look better, because those variables which have a correlation with this will not be so highly related. So these are all some
[01:24:36] Mithun DJ: elementary or rudimentary steps that you and I can do. So I've run this after eliminating. I've rerun the regression model, okay?
[01:24:45] Mithun DJ: So, try to fine-tune the model as much as possible, but…
[01:24:48] Mithun DJ: it still said the conditional number is large, so you can just, you know, this conditional number and such things, you can see 147 into 10 to the power of 4, which means, it's a big, big number. Ideally speaking, the conditional number and such things must be less than 15. It is less than 15.
[01:25:06] Mithun DJ: Which means that your model does not have serious problem of multicollinearity, but if you have such massive numbers.
[01:25:12] Mithun DJ: It means that there's a 1, 4, 7, and then 00. That's a…
[01:25:17] Mithun DJ: big value, so that's why your, your Python prints a message saying the conditional number is large. This might indicate that there are strong multicollinearity or other numerical problems.
[01:25:32] Mithun DJ: So this is the motivation. If you get such messages, basically, then you have to resort to
[01:25:36] Mithun DJ: regularization and justice. Is this clear, guys?
[01:25:43] Mithun DJ: Otherwise, the story ends here.
[01:25:46] Sacheen Adavinavar: Yeah, scale, yeah.
[01:25:59] Mithun DJ: Okay, so this is the same dataset, what do I do?
[01:26:10] Mithun DJ: Let me repeat this procedure, wherein…
[01:26:29] Mithun DJ: Let's give a gap and run this. So, I'm running all the six steps, correct? So, you can see here, this is the LASSO model. Lasso, I've just…
[01:26:37] Mithun DJ: started with a conservative estimate. You can see alpha is equal to 0.1. I do not intend to start with a big number, right? So you can just incrementally go on increasing this, okay?
[01:26:49] Mithun DJ: So, whether you want a big value or a small value, that entirely depends upon you. So, some bit of, you know, trial and error method is there, right? So, note here, if this… if you're not satisfied with this case, you can just experiment with… by giving alpha is equal to 0.2, 0.3, 0.4, and such things.
[01:27:08] Mithun DJ: As you give, a bigger… as you increase, basically, there'll be more shrinkage. More shrinkage means more of the quotients will become zero.
[01:27:16] Mithun DJ: Correct? So, a high value of alpha, basically, it helps you in reducing overfitting. Typically, it is useful when you have too many correlated features, correct? The quotients table, shows which variable lasso would keep and which one it would remove, okay?
[01:27:51] Mithun DJ: to observe here, where is that NOX? NOX value, earlier it was 22. You can see here, it is minus 1.56, which means this model is not actually dependent on only NOX, because there are so many partners of NOX. You can see here, you've, you've almost burnt out the heavy dependence of NOX.
[01:28:12] Mithun DJ: you have burnt out the heavy dependence. Earlier, it was 17 or 22, now it has come down to minus 1 point. So, by and large, if you look at a lot of other variables, it is 1, 2, 3, it is still okay-ish, because I don't think it is learning from one single variable, it is learning from multiple variables.
[01:28:28] Mithun DJ: This is what your laser does. Is this part clear, guys?
[01:28:34] Deepak Bobade: It's… yeah.
[01:28:35] Mithun DJ: Next, basically, I'll just show you one more example wherein, I'll speak about, rich regression, okay.
[01:29:04] Mithun DJ: So, you can see here… see, okay, let's just go back one step. Did you see there is nothing as such which can be forcefully put to zero?
[01:29:12] Mithun DJ: Except this guy. Can you see here? The variable indus is being forced to zero. Can you see here? In the earlier case?
[01:29:19] Mithun DJ: It is not exactly 0, because in the 9th or the 10th decimal, it may be a small value. It is displaying the first 4 or 5 digits here, but it's still pushing it to 0C.
[01:29:28] Mithun DJ: And you notice that this variable Indus was a problematic variable because the correlation and sustains were very high. It was insignificant also. Have a look at age. This is also basically being pushed to zero.
[01:29:41] Mithun DJ: In the case of Lasso.
[01:29:45] Mithun DJ: Okay. Now, in the case of your rigid regression, basically, you're not getting a zero. Even your indus, a variable like indus, it has some 0.25. It has… it is not aggressively grabbing its neck and pushing it to zero.
[01:29:59] Mithun DJ: Correct? So, one of the methods solves the problem of feature selection, the second one solves the problem of,
[01:30:07] Mithun DJ: Multicollinearity, so this is the method that you should ideally speak when you want multicollinearity.
[01:30:13] Mithun DJ: Correct? This is about ridge regression. The third one is, elastic.
[01:30:36] Mithun DJ: Have a look at this.
[01:30:38] Mithun DJ: 00mini 01302 zeros.
[01:30:55] Mithun DJ: Is this clear, guys? It's a combination, which means you can attack both the problems using your elastic net. Let me just show you the code. Code is pretty simple, as I said. What is there in this? Anybody can do this, correct? You can just hear a regularization strength, alpha is equal to 0.1, L1 ratio is 0.5.
[01:31:14] Mithun DJ: If you maintain this as 0, it is rid… Correct?
[01:31:18] Mithun DJ: If you maintain this as 1, it becomes a…
[01:31:21] Mithun DJ: lasso. If you want a mix… since I want a mix, basically, of both L1 and L2, I'm keeping it as 0.5.
[01:31:30] Mithun DJ: Right, so it depends upon what is the strength that you give. By controlling for these parameters, you can just obtain the ideal value.
[01:31:41] Mithun DJ: Is this clear, Liz?
[01:31:51] Mithun DJ: 0.5 is for balance, correct?
[01:32:06] Mithun DJ: Guys, this is what I had, guys?
[01:32:11] Mithun DJ: If you have any questions, you can ask me. If you don't have any questions, we can wind up.
[01:32:15] Mithun DJ: So we have covered, basically, you know, supervised learning, and within that, assumptions of regression, what are the typical challenges and such things that you can find out, regularization, the three types of fittings, the slide is very, very important, that is, your bias variance trade-off.
[01:32:32] Mithun DJ: And then, the four cases, low bias, low variance, high bias, low variance, low bias, high variance, high bias, high variance. The four cases, what is its meaning? That you should be very, very clear with. And we have run four models. One is a typical linear regression model, second one is lasso, then basically you have L1, and then L2 regularization.
[01:33:01] Sushree Dash: I have one question, Mithun. This, L1's, value that's 0.1, when should we change it? I mean, on what basis we should change it?
[01:33:12] Mithun DJ: See, even, let's say you run a L1, regularization, watch out for the quotients.
[01:33:18] Mithun DJ: Correct? If you watch out for the coefficients and you still feel that your NOX value has not changed drastically, increase it furthermore.
[01:33:26] Sushree Dash: We need to increase it, okay.
[01:33:28] Mithun DJ: Yes, yes. So, typically, as a thumb rule, it is not written in any book and such things, but if you want some thumb rules and such things,
[01:33:35] Mithun DJ: You know, thumb rules in a sense, just for your understanding, basically. After you use a L1 regularization, your R-square should be around 0.7 to 0.85, depending upon the level of your,
[01:33:50] Mithun DJ: Depending on the level of your, alpha.
[01:33:55] Mithun DJ: So, anywhere if you get between 0.7 to 85, it's a decent, decent… you can stop there. If it is outside this range, you continue doing this. Similarly, for L2, basically, if your R-square is anywhere between 0.75 and 0.9, depending upon the alpha that you have used, you can go ahead and stop it.
[01:34:13] Sushree Dash: Okay.
[01:34:14] Mithun DJ: And the third case is, basically, I've told you for your elastic net, basically. After doing all of these things, if your value is anywhere between 0.7 to 0.88, depending upon the level of tuning, you can be, you can just, appreciate that you have done a good job.
[01:34:32] Sushree Dash: Thank you.
[01:34:33] Mithun DJ: Thank you, thanks a lot, thank you.
[01:34:52] Mithun DJ: Hey, is there no questions? Shall I take leave, Chris?
[01:34:56] Muni Prakash Ganji: Yeah, this is Muni. I have a small,
[01:35:00] Muni Prakash Ganji: we… we know that L1 and L2, and in… in mix, we have…
[01:35:06] Muni Prakash Ganji: one more thing which we can use it, correct? It's elastic knit. So, it is a best practice, we always can use Elastic Knit in place of L1 and L2, because it is the combination of both, correct?
[01:35:19] Mithun DJ: See, it's like,
[01:35:22] Mithun DJ: See, what medicine that you give depends upon what is a disease that the patient has. So, firstly, you should make up your mind by looking at the correlation metrics and such things, saying that, boss, is my problem of high dimension
[01:35:37] Mithun DJ: Or is my problem of high multicollinearity?
[01:35:41] Mithun DJ: Only if both of these things are a problem, you go and use elastic matter.
[01:35:46] Mithun DJ: If only one of them is a problem, use either L1 or L.
[01:35:50] Mithun DJ: So it's like… If the person has cough.
[01:35:54] Mithun DJ: One medicine. If a person has cold, another medicine. If a person has cough and cold.
[01:35:59] Mithun DJ: right, I would use ElasticNet. So, in case, basically, you do not have… in case, basically, you do not… your dataset does not have that problem at all, then why bother with ElasticNet?
[01:36:15] Mithun DJ: Elastic only for hybrid problems, otherwise no need.
[01:36:21] Muni Prakash Ganji: Okay, by seeing those values, we need to come to conclusion.
[01:36:25] Mithun DJ: Yes, yes, yes.
[01:36:45] Mithun DJ: Okay, guys, thank you so much, have a good day.
[01:36:51] Sacheen Adavinavar: Thank you.
[01:36:52] Mithun DJ: Thanks a lot, bye-bye.
[01:36:53] Dilip S Biradar: Thank you.
[01:36:54] Mithun DJ: Thanks, Spike.
[01:36:56] Neeraj Kumar: Thank you.