# 07 2025-12-20 Linear Regression Theory

course: Module 2 — Machine Learning Algorithms
module: Module-2-Machine-Learning-Algorithms
date: 2025-12-20
type: transcript
video_url: https://personal-learn.armco.dev/files/_Recordings/Module-2-Machine-Learning-Algorithms/07_2025-12-20_Linear_Regression_Theory.mp4

---
[00:25:55] it, or…
[00:25:58] you know, heard about it. So, let us discuss what is linear regression.
[00:26:05] It's just the outline.
[00:26:08] So, first of all, uh,
[00:26:11] What about the word regression? What is, uh…
[00:26:15] Uh, does it hold any meaning, this particular word?
[00:26:18] Or it is just used
[00:26:22] in ML. So, basically, the root of this word, regression,
[00:26:26] I was a word called regressas, which means returning back or going back.
[00:26:32] And how this evolved was the…
[00:26:35] technique called digression.
[00:26:37] Uh, actually, it's, um…
[00:26:40] I'll statistical technique, which is also used in machine learning.
[00:26:44] was evolved by Sir Francis Galton way back in the 19th century.
[00:26:50] And what he observed was that if we talk about tall parents, then they'd have… tend to have
[00:26:56] Tall children, but the children may not usually be as tall as their parents.
[00:27:01] Similarly, short parents also have shorter children.
[00:27:05] But the shorter children may not be as short as the parents.
[00:27:09] Or, in other words,
[00:27:11] Uh, he tried to explain a phenomenon which… which he called…
[00:27:17] Moving towards the mean, or regression towards the mean.
[00:27:22] Which means that the extreme…
[00:27:25] extremities in the data.
[00:27:27] tend to move over as the distribution grows, it tends to move
[00:27:32] More and more closer to the average values.
[00:27:35] And, uh, that's the reason why…
[00:27:37] The descendant of tall parents may not be so tall, and shorter parents may not be short.
[00:27:45] So, very short, and uh…
[00:27:47] The value would always tend to move.
[00:27:50] Or that, uh, it would always be having a regression towards the mean.
[00:27:55] So, this is how this…
[00:27:57] term regression was evolved.
[00:28:00] Um…
[00:28:02] And, uh, in machine learning, linear regression is…
[00:28:06] Uh, very popular technique in which we talk about a dependent variable.
[00:28:12] The dependent variable is nothing but a target variable or the class label.
[00:28:17] And then we have one or more independent variables, which are also called predictors.
[00:28:24] And these are nothing but the features that we use.
[00:28:27] Uh, to, uh, make the prediction.
[00:28:30] So, uh, what is a key idea in regression? So, we try to fit a straight line.
[00:28:37] Or, uh, um, which is a straight line would be in a two-dimensional space, or we talk about a…
[00:28:47] Linear hyperplane in higher dimensions, that
[00:28:50] best describes the relationship?
[00:28:53] Between the independent variables.
[00:28:56] Um, and the dependent variable.
[00:29:00] Now, what is the use of recession?
[00:29:03] A regression is usually, um…
[00:29:06] useful. Uh, when we want to forecast the values.
[00:29:10] Of some particular target variable.
[00:29:13] Or we want to predict our value. Usually, the word forecasting is used
[00:29:18] For example, uh, we want to forecast what would be the temperature for today.
[00:29:23] Or we want to forecast what would be the value of rupee tomorrow.
[00:29:27] or the forecast value of the stock price, and so on and so forth.
[00:29:31] So usually, the word forecasting is used when we talk about a value.
[00:29:37] If we talk about a class label,
[00:29:39] We called it prediction.
[00:29:42] So, regression is often associated with forecasting of a value.
[00:29:47] When we talk about linear regression.
[00:29:51] And it, uh, is also very useful to understand the type of relationship between variables, what kind of positive, negative.
[00:29:59] a relationship it is, and I think we discussed it earlier, which I'll…
[00:30:03] Mention again here.
[00:30:05] Just a second. Just choose my pen.
[00:30:12] So the relationship kind of relationship could be…
[00:30:16] positive. So, if we are talking about…
[00:30:21] Um, two variables, let's say X and Y.
[00:30:27] then there could be a positively near relationship.
[00:30:33] Or…
[00:30:36] there could be a…
[00:30:39] negative linear relationship, or…
[00:30:46] there could be zero correlation, or zero kind of a…
[00:30:51] relationship between the two.
[00:30:52] So, it also helps to understand that kind of a relationship, because we try to…
[00:30:57] find out, uh, the line that fits this kind of given set of data points, and when we do that,
[00:31:04] In that process, we also try to understand.
[00:31:09] Oh, whether the slope… what is the slope of the…
[00:31:12] line, and accordingly, it's going to be positive or negative.
[00:31:16] And so we understand that.
[00:31:19] kind of relationship.
[00:31:20] And, uh, uh…
[00:31:22] And we try to understand how strongly the variables could be related to each other.
[00:31:29] And therefore, how much change in one variable could be associated with the change in the other?
[00:31:34] Uh, which is what is measured by correlation, and also something called R-squared, which we had discussed earlier, which I'll also discuss today.
[00:31:45] So now the assumptions on which linear regression, um,
[00:31:50] is, uh, working, uh, are as follows.
[00:31:54] Uh, so when we are talking about, uh, this form of regression, we are assuming linearity.
[00:32:01] That is, uh, that the input and the outputs
[00:32:05] The inputs and the output are linearly related with each other. So, what does it mean?
[00:32:11] It actually means that…
[00:32:13] If you're talking about a linear relation, there would be a straight line if we are talking about
[00:32:19] Oh, two-dimensional data space.
[00:32:22] If it is a, uh, n-dimensional data space, then it would be a linear hyperplane.
[00:32:29] This is the word that is used, hyperplane.
[00:32:32] Uh, that, uh, separates.
[00:32:35] the… or that represents the relationship between the inputs and the outputs.
[00:32:42] We also assume that all the observations are independent of each other, and that
[00:32:48] There is no dependence amongst them.
[00:32:51] And we also assume…
[00:32:53] that the variance in terms of the error is going to follow the normal distribution.
[00:33:00] What is the normal distribution? This is what is the normal distribution.
[00:33:04] that, uh…
[00:33:06] There will be some values which are very low, some values, I mean, less than
[00:33:12] 10% value is very low, less than 10% value is very high. Usually, all the values will lie in this
[00:33:19] average, um, average range.
[00:33:23] So, we are assuming that the errors are following a normal distribution.
[00:33:31] Now, before we go into what is linear regression, I already explained to you what is correlation earlier, using the Pearson correlation coefficient.
[00:33:40] And I briefly also…
[00:33:42] Uh, just explain again to you.
[00:33:45] Uh, there's something called R-squared that you had also probably seen in the code.
[00:33:51] What it means is that
[00:33:53] Uh… if we try to, you know, R-square tries to measure
[00:33:59] How much of the variability
[00:34:01] uh, is explained by the model.
[00:34:04] Uh, in the dependent variable.
[00:34:07] So, what will be the dependent variable?
[00:34:10] The dependent variable will be
[00:34:12] For example, if we are talking about Y is equal to function of X,
[00:34:17] This is the input variable.
[00:34:21] And this is…
[00:34:23] Or, this is the independent variable.
[00:34:27] And Y will be our dependent variable.
[00:34:31] Because Y is dependent on the
[00:34:34] observations, or…
[00:34:36] Uh, on the values of X.
[00:34:38] So, uh…
[00:34:41] So, R squared helps us to explain how much of the variability or the variance in the
[00:34:48] Dependent variable is explained
[00:34:51] by the model.
[00:34:53] So, a model is good if it captures a huge amount of variability or variance in the dependent variable.
[00:35:00] It is not so very good if it is not able to capture the variability.
[00:35:05] In the dependent variable.
[00:35:08] The values of R squared…
[00:35:10] ranges from 0 to 1.
[00:35:13] But zero means that a model is not able to capture at all
[00:35:18] the variance in the dependent variable, which may be Y.
[00:35:21] And 1 means that the model is fully able to
[00:35:24] capture all the variants that is associated with the
[00:35:29] dependent variable, uh, that could be Y.
[00:35:33] For example, if the value of R squared is equal to 0.85, it means that
[00:35:39] 85% of the variance in Y is being captured.
[00:35:43] by the model. And therefore, uh…
[00:35:46] the, you know, uh…
[00:35:49] Therefore, the… there is a…
[00:35:51] very good or a strong relationship.
[00:35:55] Between the independent variable and the dependent variable.
[00:36:00] The model is good, number one, if it captures 85% of the variance in the dependent variable.
[00:36:08] Secondly, there is a strong
[00:36:11] Uh, relationship.
[00:36:12] between, uh, the predictors or the…
[00:36:16] independent variable and the dependent variable.
[00:36:19] So, R squared is often used to assess
[00:36:23] the models goodness, along with the…
[00:36:26] Uh, other, uh, other metrics.
[00:36:33] So, um, now how… what is the formula to find out?
[00:36:38] The value of R squared. Now, one thing, please note that in the code that you will do,
[00:36:44] Or you would have done…
[00:36:46] R squared is just used as a function.
[00:36:50] And you really don't code.
[00:36:53] it from scratch. However,
[00:36:55] Uh, the formula, uh, is what you should know, which is R squared is equal to…
[00:37:02] Um, 1 minus SSRDS divided by SS total. So, SSRDS is the…
[00:37:08] SS are the, uh, sum of squares residual. This is SSRES, which is residual sum of squares.
[00:37:16] And assess total is the total sum of squares.
[00:37:20] So, what is the residual sum of squares?
[00:37:22] So, if Y hat…
[00:37:25] Y hat…
[00:37:28] represents the predicted
[00:37:30] or the forecasted value.
[00:37:34] Then, and via is the actual given value.
[00:37:38] Because in the training data, we always have the
[00:37:42] a value for the target variable. So, Yi is that value? i is some particular i-th record.
[00:37:50] And Yi hat is the value that has been predicted.
[00:37:54] Then YI minus Yi hat square, that is the difference between the actual value
[00:38:00] And the predicted or the forecasted value, squared.
[00:38:04] And this added for all…
[00:38:06] and samples. If n is the total number of sample,
[00:38:10] In the given data, this is called the residual sum of squares.
[00:38:14] And what is the total sum of squares? SS total is equal to…
[00:38:19] YI minus Y bar. What is Y bar?
[00:38:23] Y-bar is actually equal to?
[00:38:29] Hello?
[00:38:31] I think my, uh, internet is actually giving me some issues, so I will…
[00:38:36] just stop my video.
[00:38:40] I hope you all are able to hear me.
[00:38:45] Okay, okay, okay, thank you.
[00:38:46] Yes, we can hear you.
[00:38:48] Okay, thank you. So, this Y-hat, uh, sorry, Y-bar is nothing but the mean value of Y.
[00:38:57] So, if you are having total n number of samples of Y,
[00:39:01] Then Y-bar would be…
[00:39:03] the summation over all
[00:39:06] the values of Y divided by…
[00:39:11] Total number of values. This will be your Y-bar.
[00:39:13] So, the forecast… this is the actual value, this is the mean.
[00:39:18] The difference between…
[00:39:20] The actual value and the VIN, that is the…
[00:39:23] kind of the variance with respect to mean is what is here, this difference squared added for all N samples.
[00:39:31] So, this is the total sum of squares.
[00:39:34] And this is a residual sum of squares.
[00:39:37] And then we take it 1 minus this.
[00:39:40] And this is what is the R squared value.
[00:39:44] Okay, now let's come back to linear regression.
[00:39:48] So, uh, when, uh, do we want to use linear regression?
[00:39:54] Linear regression is, um, um…
[00:39:57] kind of a very simple solution.
[00:39:59] It is easily understandable.
[00:40:03] And it is very oftenly useful when the data may be noisy,
[00:40:07] Secondly, it may also be useful in filling of missing values.
[00:40:12] So, many a times, regression may be used to fill in missing values.
[00:40:17] Um, and, uh, it is quite computationally less expensive as compared to the other algorithms.
[00:40:25] Uh, it is generally used, of course, for forecasting the value when we know that there is a linear relationship
[00:40:32] between the predictors and the predicted value.
[00:40:37] Or the, uh…
[00:40:39] dependent variable and the independent variables.
[00:40:45] So now, uh, linear regression is of two types, one which is simple linear regression, the other is multiple linear regression.
[00:40:53] So, we'll talk about each of these as we go along.
[00:40:57] So, let's first talk about simple linear regression.
[00:41:00] So, in simple linear regression, we actually talk about
[00:41:05] uh… the…
[00:41:07] dependent variable, and its relationship or its expression.
[00:41:13] In terms of one independent variable.
[00:41:18] So, what is simple linear regression?
[00:41:22] So, in simple linear regression, we have the dependent variable.
[00:41:30] And this dependent variable
[00:41:32] is expressed or found out with the help of
[00:41:35] One single independent variable.
[00:41:41] And that is why it is called simple linear regression.
[00:41:46] So, the expression for simple linear regression is Y hat is equal to beta naught plus beta 1x.
[00:41:55] So, beta naught and beta 1x are the coefficients.
[00:41:58] That we need to find out. Y hat,
[00:42:00] is the predicted variable or the dependent variable.
[00:42:04] And X is the independent variable.
[00:42:08] ERX is just a single variable.
[00:42:11] And beta naught and beta 1 are the coefficients.
[00:42:15] Which are also called the regression coefficients.
[00:42:19] In this case, they are simple linear regression coefficients.
[00:42:22] And beta naught is nothing but the intercept of the line.
[00:42:26] And beta 1 is the slope of the line. So if we have…
[00:42:30] a line like this.
[00:42:32] And then beta naught will be this.
[00:42:35] And beta-1.
[00:42:37] will be the slope of the line. So, this will be the beta 1 value.
[00:42:43] And these are the regression coefficients that we want to determine to be able to find out the expression for
[00:42:49] Y hat, which is the dependent variable.
[00:42:55] So, uh, just some illustration to understand what is linear regression, simple linear regression.
[00:43:03] So, we have this independent variable X, and we have the dependent variable y, and we have
[00:43:09] Uh, some samples, which are given to us, which is what you're seeing in the green color.
[00:43:14] And we want to, uh, in case of a simple linear regression, because there is one
[00:43:20] independent, and one dependent variable. We are talking about simple linear regression.
[00:43:28] We want to find out, um, if you're talking about 2D data, then we want to find out the…
[00:43:33] line that best fits…
[00:43:36] this set of points which are given here.
[00:43:40] So, uh…
[00:43:42] To do that, we try to, uh…
[00:43:47] fit a line in such a way.
[00:43:48] And of course, to fit the line, we have to find out the intercept and the slope of the line.
[00:43:53] And we, uh, try to fit it in such a way.
[00:43:56] that the error
[00:43:59] is minimized. What is the error?
[00:44:02] Suppose, uh, the dash line…
[00:44:06] shows, uh, the regression line.
[00:44:09] That we have found out, the simple linear regression line.
[00:44:13] And the given training points are shown in as the green circles.
[00:44:19] So, uh…
[00:44:21] So, what we want to do is, the error represents the difference between the value that has been taken and the one which is the actual value.
[00:44:29] As per the training sample. So, what we have here…
[00:44:33] This is a value that is given in the training data, but as per this line that has been fit on it,
[00:44:39] If we draw a perpendicular…
[00:44:41] projection, then this is the point on the line that is representing this training data.
[00:44:46] And this training point. So, the difference between these represents the error.
[00:44:52] What is the error? The actual value?
[00:44:55] And this is… this will be the predicted value. So, the difference between the actual value
[00:45:00] And the predicted value is said to be the error. So…
[00:45:04] So, all these projections…
[00:45:06] Uh, can be used to find out the error.
[00:45:10] So, for example, if the training point falls in this…
[00:45:13] Linear line only, then here.
[00:45:16] The error will be 0.
[00:45:18] Because the predicted value and the given actual value coincide, so error is zero.
[00:45:25] So, for example, if we talk about this point,
[00:45:28] We look at the projection of this point. This is the point that we get.
[00:45:32] And we measure the distance, so this is the error.
[00:45:35] So, our simple linear regression will aim to
[00:45:40] identified that, uh, linear equation, or that line,
[00:45:44] That best fits the data so as to…
[00:45:47] Optimally reduce the error.
[00:45:51] So, our aim would be to minimize the error, that is,
[00:45:54] The difference between the predicted and the forecasted values.
[00:45:58] are minimized. This is what is our goal.
[00:46:02] So, uh…
[00:46:04] So, this is the equation.
[00:46:07] Y hat is equal to beta naught plus beta 1x. As I said, beta naught is the intercept.
[00:46:12] Which is this one.
[00:46:14] This is the donor, and the slope is represented by…
[00:46:18] beta 1.
[00:46:21] And these are the parameters.
[00:46:23] that we want to find out.
[00:46:25] Now, this was simple linear regression.
[00:46:28] But we also have multiple linear regression.
[00:46:31] So, multiple linear regression means that…
[00:46:36] We have a dependent variable.
[00:46:38] And this dependent variable actually depends on multiple predictors or multiple independent variables.
[00:46:46] So now, in this case,
[00:46:48] We have multiple features.
[00:46:50] So, our attributes…
[00:46:55] or features?
[00:46:59] are independent variables mean the same thing.
[00:47:06] Maybe, but…
[00:47:10] And the dependent variable.
[00:47:17] The dependent variable.
[00:47:22] is what is found out using these attributes or features or independent variables.
[00:47:29] So, in the case of multiple linear regression,
[00:47:32] odd expression gets modified to this.
[00:47:35] So, in the previous case, R…
[00:47:38] Y hat was nothing but beta naught plus beta 1 times X, because we had only one single.
[00:47:44] independent variable. Now we have multiple independent variables.
[00:47:52] that I use to predict the value of Y hat.
[00:47:54] Which is why, I mean, predict the value of Y, which is called Y hat.
[00:47:59] And it is given as beta naught plus beta 1X1 plus beta2X2 plus beta 3x3tel.
[00:48:05] beta and XM.
[00:48:07] So again, my heart is a dependent variable that we want to predict.
[00:48:11] X1, X2, XT till XN are the independent variables, or…
[00:48:16] That's what I mentioned. These are nothing but also the attributes given to us.
[00:48:21] Or these are the feature values.
[00:48:23] given to us.
[00:48:25] and beta naught, beta 1.
[00:48:27] Beta 2, beta 3 till beta.
[00:48:30] are the regression coefficients.
[00:48:33] Uh, that are…
[00:48:35] Uh, that are to be found out.
[00:48:37] In order to find out this particular equation.
[00:48:44] So now, um…
[00:48:46] Whenever we have to fit a line, whether we are…
[00:48:49] Uh, going to use, uh, multiple linear regression or a simple linear regression.
[00:48:55] Our aim is to minimize the error.
[00:48:58] There are many metrics that I've already covered to measure the error.
[00:49:02] One of them is the mean squared error, or the MSC.
[00:49:06] The mean squared error actually measures the difference
[00:49:09] Between the predicted value and the given actual value in the training data.
[00:49:16] The difference of it, squared.
[00:49:17] And all of them added for all M samples. So, M is the total number of…
[00:49:23] Total number of…
[00:49:26] Training.
[00:49:28] Samples.
[00:49:31] So, we find out, uh…
[00:49:35] the difference between the predicted value and the given value for all M samples,
[00:49:40] square these distant differences.
[00:49:44] add all of them and divide them by M, which is a total number of…
[00:49:48] Uh, values that we are having, uh, in the training data.
[00:49:52] So this cost function is also called a mean squared error, because this is a square of the errors.
[00:49:59] And when we divided by M, we are actually finding out the mean of this error.
[00:50:06] Right? And we want to, uh, optimize the cost function here.
[00:50:10] definitely optimize would mean minimization, right?
[00:50:14] And, uh, that particular, uh…
[00:50:17] line or hyperplane that minimizes the error is the one
[00:50:23] that we use as the best fit line.
[00:50:28] So, let's just look at an example.
[00:50:31] Just to understand, let's say these are the set of values given for X, and these are the set of values.
[00:50:38] given for Y, and we know that Y hat is equal to…
[00:50:42] beta nought plus beta 1x.
[00:50:45] So we have y hat equal to beta not plus beta 1x.
[00:50:49] And in the first case, the value of beta naught is 0.
[00:50:54] And wea 1 is 1.
[00:50:55] So what will we get? We will get Y hat…
[00:50:59] So this will get removed because this is 0, this is 1, so it will be equal to X.
[00:51:03] So, Y hat, or the predicted value of Y, will be equal to X. This is what is given.
[00:51:10] So what does it mean? Y hat is equal to X? It means that…
[00:51:14] If we are talking about…
[00:51:16] Vaivan, hard, then it will be equal to X1.
[00:51:20] And what is X1?
[00:51:22] It is 1. Similarly, Y2 hat will be equal to 2.
[00:51:27] Because this is what then Y3 hat will be equal to 3.
[00:51:31] And Y4 hat will be equal to 4.
[00:51:34] So, these are the predicted, uh, values.
[00:51:37] And the cost function, which is MSC, that is mean of summer squared errors,
[00:51:44] would be 1 minus 1 squared plus 2 minus 2 squares.
[00:51:47] And so and so forth.
[00:51:49] 4 minus 4 squared.
[00:51:51] And then, um…
[00:51:55] the mean of it, which is 1 by 4, and it comes out to be 0.
[00:51:58] So here, the error is 0.
[00:52:01] And the line that is represented will be…
[00:52:05] Uh, Y-hat is equal to X. So, if we were to plot this…
[00:52:09] What will it be? This is X, this is the predicted value, Y hat.
[00:52:13] And we have if…
[00:52:16] Uh, this is one.
[00:52:22] Like this.
[00:52:26] So, we have value 1 and 1, then if X is, uh, 2, then Y is also 2.
[00:52:33] Then in XS3, Y is also 3…
[00:52:36] So, this is the kind of the…
[00:52:38] graph that, uh, or the line that can be found out when.
[00:52:43] We have beta naught equal to 0, and beta 1 equal to 1, or Y hat is equal to X.
[00:52:51] Okay.
[00:52:53] So now, uh, this is the line.
[00:52:56] This one.
[00:52:58] And error here is zero. In the next case, let's assume that beta naught is equal to 0.
[00:53:05] And beta 1 is equal to 0.5. So, equation is y hat is equal to…
[00:53:09] beta naught plus beta 1X, which is equal to…
[00:53:13] Y hat is equal to beta naught is 0, so we'll have 0.
[00:53:17] plus beta 1x is 0.5x.
[00:53:20] Like that. So, 0.5x means, um…
[00:53:24] This will be the line.
[00:53:28] And here in this case, when X is equal to 1, then Y will be 0.5 times of X, which will be 0.5.
[00:53:35] Access to Y will be 1, XS3, Y will be 1.5, XS4, Y will be 2.
[00:53:41] And accordingly, you calculate the error.
[00:53:44] As we calculated here, so it will be 1 by 4.
[00:53:49] Uh, into. And what we want to do, we want to find out the difference between the actual and predicted value.
[00:53:55] So, actual value of Y hat.
[00:53:57] Uh, is equal to, let's say, 1.
[00:54:00] then it will be 1 minus 0.5 square plus…
[00:54:04] 1 minus 1 squared plus 1 minus 1.5 square.
[00:54:11] plus 1 minus 2 squared.
[00:54:13] And whole of it divided by 1x4. And it will come out to be this much there.
[00:54:19] Similarly, in the case 3,
[00:54:21] We have beta-naught and beta 1 both to be equal to 0. So, our equation was…
[00:54:27] beta, uh, sorry, Y hat,
[00:54:30] is equal to beta…
[00:54:32] naught plus beta 1X.
[00:54:34] Now, beta naught is equal to 0, so we are just removing it.
[00:54:38] And beta 1 is also equal to 0, so this also gets removed.
[00:54:43] So our predicted value says Y hat is equal to 0.
[00:54:47] And that's all. And what will be that line?
[00:54:51] So this is going to be that line.
[00:54:53] So, this line represents…
[00:54:57] Uh, the predicted value is always equal to…
[00:55:00] Zero. So, no matter what value of X, Y will always be 0.
[00:55:15] So, any questions, any thoughts so far?
[00:55:25] Um, ma'am, in the formula, 1 by M, what is M?
[00:55:29] Yeah, M is the total number of samples.
[00:55:32] Okay.
[00:55:33] So we have… I'll go back here… yeah.
[00:55:36] Okay.
[00:55:38] So here we have 4 number of samples.
[00:55:41] So, we are dividing by 4.
[00:55:42] Okay.
[00:55:46] Anyone else? Any other questions?
[00:55:51] of course, I'm doing all this manually just for the sake of clarity.
[00:55:55] However, when you write code, you will have functions for linear regression and everything.
[00:56:01] And you'll just be making use of it.
[00:56:03] And, uh, that's it. So, you'll not be required to do all these calculations, but I'm showing because you should know.
[00:56:12] Anyone else?
[00:56:19] So now, once again, we have another example. We have five points, which are what you're seeing here. Again,
[00:56:25] Access the independent variable, and Y is the dependent variable.
[00:56:29] And the same thing, Y hat is equal to beta naught plus beta 1x,
[00:56:34] has been just returned in a parallel form, which is more…
[00:56:38] popular, which is Y is equal to MX plus C.
[00:56:43] Here, beta not and M are same thing. C and beta 1 are same thing.
[00:56:47] Y and Y hat mean the same thing.
[00:56:52] So now, suppose we want to, uh…
[00:56:55] calculate. Let's say that we want to find out this equation.
[00:57:02] So, Y is equal to…
[00:57:05] MX plus C. We want to find out this equation.
[00:57:08] That will be the one that will fit all these points. To find out Y is equal to…
[00:57:13] MX plus C, we have this formula in linear regression where M…
[00:57:18] So I'll write this. I is equal to MX plus C.
[00:57:22] So, this is the formula for M.
[00:57:25] MS summation over X minus X bar. X bar is the mean value.
[00:57:30] So, X-bott is mean X.
[00:57:33] Mean of X.
[00:57:38] And Y bar.
[00:57:40] This is mean of Y.
[00:57:44] So, we have the given value of X minus X bar into Y minus y bar divided by X minus X bar whole square.
[00:57:52] And with this summation over all the values, we find out M.
[00:57:57] It comes out to be 0.4.
[00:57:59] And what are we substituting this to get 0.4?
[00:58:03] We are substituting X bar here.
[00:58:05] What is X bar? That is a mean of all the values.
[00:58:10] So, X bar here is 3.
[00:58:13] You can see here, mean of X is 3.
[00:58:16] And Y bar is equal to 3.6.
[00:58:20] You can see mean of…
[00:58:22] all the Y values is 3.6.
[00:58:28] So, uh, each time, what you have to do is you have to find out
[00:58:32] Uh, for this expression, X minus X bar, Y minus y bar, like that.
[00:58:38] So, for example, if you do X minus X bar here,
[00:58:42] So, X bot is 3.
[00:58:44] And X value is 1. So what… what do you have?
[00:58:48] 1 minus 3, which comes out to be minus 2.
[00:58:51] Then you have year 2, so 2 minus 3, which comes out to be minus 1.
[00:58:57] Then you have 3, so it is 3 minus 3, which is 0.
[00:59:02] Then 3, my, uh… then…
[00:59:04] You have 4 minus 3.
[00:59:06] Which is 1. And then, uh, you have 5 minus 3.
[00:59:11] Which is 2.
[00:59:13] So these are the value of X minus X bar, which we'll substitute here.
[00:59:18] similarly, we have to calculate the value of Y minus y bar.
[00:59:23] So, why are all these values, and Y-bar, I told you, is 3.6?
[00:59:28] So each time, what we have to do?
[00:59:30] Again, do the subtraction.
[00:59:32] So this is 3. So, what do we do when we want to find out Y minus y bar? It is 3 minus 3.6.
[00:59:40] Which comes out to be minus 0.6.
[00:59:43] Then it is 4.
[00:59:45] 4 minus 3.6, which is…
[00:59:48] point forward, then you have…
[00:59:51] 2 minus 3.6, which is minus 1.6.
[00:59:54] Then you have 4 minus 3.6, which is 0.4.
[00:59:57] And then you have 4 minus, uh, sorry, 5.
[01:00:01] 5 minus 3.6.
[01:00:03] Which is 1.4.
[01:00:07] So, this is why minus Y, uh, bar.
[01:00:10] So, in this formula, we also need Y minus y bar. Next, we need X minus X bar whole square.
[01:00:17] So, we calculate the whole square, X minus X bar is minus 2.
[01:00:21] Whole square will be 4, whole square will be 1.
[01:00:25] 0, 1, 4, like that.
[01:00:28] So now we have all these values, X minus X bar, Y minus y bar, and X minus X bar couple square.
[01:00:35] And then we find out also the product of these two.
[01:00:40] That is product of this and this, like, for example, minus 2 and…
[01:00:44] Um, minus 0.06.
[01:00:48] So this is what we do, and like this, we just, you know, find out.
[01:00:54] The final value of M.
[01:00:56] As well as we find out the value of C,
[01:00:59] And then, uh…
[01:01:01] you know, C has been found out to be 2.4.
[01:01:05] Using this… these expression, this expression.
[01:01:09] That we just calculated.
[01:01:12] So, we have, let's assume, based on these points,
[01:01:17] CS2.4, and MS.4.
[01:01:21] So these are the two values, M and C.
[01:01:25] And we substitute them as Y is equal to MX plus C, so M will be 0.4, and C will be 2.4.
[01:01:32] So, we have this Y is equal to MX plus C.
[01:01:36] M is 0.4, this is the value of n.
[01:01:39] And this is the value of C.
[01:01:42] Okay, so we have now found out the linear, uh, equation.
[01:01:47] putting these lines.
[01:01:49] And this is the one with…
[01:01:51] uh, you know, the least error.
[01:01:54] And now, when we found, uh, tried to draw this graph, we have 2.4 as the intercept.
[01:02:01] So this is 2.4.
[01:02:04] And what is the slope? It is 0.4.
[01:02:07] So, uh…
[01:02:10] 0.4 means that if…
[01:02:12] X is equal to 1, then the value of Y would be if x is 1.
[01:02:17] Then Y will be equal to 2.4 plus 0.4.
[01:02:21] Which will be equal to 2.8.
[01:02:24] So, uh, when we have…
[01:02:27] X is equal to 1.
[01:02:31] then Y will be equal to…
[01:02:34] 0.5 times X plus 2.4, which will be 2.8. So you can see here.
[01:02:40] When one, then it is 2.8 is approximately taken as 3. So when Y is 1, this is 3. When this is 2, then according… like this.
[01:02:49] So, uh…
[01:02:51] This is how we fit the line.
[01:02:53] And find out the values of…
[01:02:56] Y and, uh, C.
[01:02:58] Here, in this particular case, because this is a case of a simple linear regression.
[01:03:07] And so, here we have this as the forecasted line. This line…
[01:03:12] Actually, at each and every value of…
[01:03:15] The independent variable.
[01:03:18] Uh, the projection.
[01:03:20] of this grid represents the predicted value. So, we have the predicted value.
[01:03:26] All of these are…
[01:03:28] On the line, simple linear equation that we have fitted.
[01:03:33] And the blue dots show the actual…
[01:03:35] points that are available in the given data.
[01:03:43] And as I already mentioned, the distance between the
[01:03:46] actual value and the predicted value gives us the error, which is what we want to minimize.
[01:03:54] Now, uh, as I already mentioned earlier,
[01:03:59] for, uh, evaluating the performance, we have R-squared.
[01:04:04] And R squared finds out the… helps us to find out the strength of the relationship between the model.
[01:04:10] And the dependent variable.
[01:04:13] For example, R-squared will be equal to…
[01:04:16] Y hat minus y bar.
[01:04:18] Uh, Whole Square.
[01:04:22] And a Y minus Y bar whole square. So, this is the predicted value.
[01:04:26] Y hat. This is the mean value.
[01:04:32] Mean value of Y.
[01:04:34] And this is the given value of Y.
[01:04:40] And this is, again, the…
[01:04:43] This is Vibad, which is the mean of Y.
[01:04:49] So Y bar will be the mean of Y, Y hat will be the predicted value.
[01:04:53] And Y will be the actual given value, substituting these, you can find out R-squared.
[01:05:01] Any questions, anyone, so far?
[01:05:04] Anyone, any questions?
[01:05:11] So I think there is a comment, uh, which says that
[01:05:15] Uh, can we say if R-square is high,
[01:05:18] It guarantees that model is working correctly.
[01:05:22] If R-square is high, it means that the model is able to capture a lot of variance.
[01:05:28] of the dependent variable.
[01:05:32] Which is good for us. So, we can say,
[01:05:35] If r squared is high, we can say that the model is…
[01:05:39] should work well.
[01:05:43] Because it captures a lot of variability of the dependent variable in it.
[01:05:51] Any other comments or questions before we proceed?
[01:05:59] Good news. Can you explain this better a little bit more, how, uh…
[01:06:02] the R-square, so…
[01:06:08] Yes.
[01:06:09] Basically, you said that it'll capture the variance more, right? Uh, but it explained it with an example, if possible.
[01:06:13] Okay, so the example for R-squared will, uh, you will see in the code.
[01:06:18] I'm not sure if you… I think you must have seen already in the code.
[01:06:22] that you did for supervised learning in the last term.
[01:06:27] R squared was calculated.
[01:06:29] Which showed that how good our model is.
[01:06:32] The more… if the value of R squared is high, it means that it is capturing the variance
[01:06:39] In the dependent variable.
[01:06:42] And we want the variance in the dependent variable to be captured by our independent variable. So, our… if our model is such,
[01:06:50] then it is good for us.
[01:06:52] Or do you recollect in the court, having seen R-squared?
[01:06:58] Yeah. Yeah.
[01:07:00] Yeah. So that is a real practical example that you…
[01:07:05] have certain training data,
[01:07:07] And based on the training data, you build a model.
[01:07:10] And then you find out the R-squared,
[01:07:12] For that model, and see how good that model is at capturing the variance.
[01:07:17] In the, uh, dependent variable.
[01:07:22] Is there any other kind of explanation you want, or…?
[01:07:27] If you could declollect that,
[01:07:29] Is it good enough?
[01:07:34] Yeah, should be fine, ma'am, yeah.
[01:07:35] Understood.
[01:07:41] And also, ma'am, is this formula same as the earlier formula that we've seen? 1 minus SS residual by SS total?
[01:07:48] Yeah, it will come out to be the same.
[01:07:51] When you substitute the values, that is 1 minus…
[01:07:55] I assess the residual divided by assessed total. It will come out to be the same thing.
[01:08:00] Okay.
[01:08:02] No.
[01:08:12] Who knows?
[01:08:15] Uh, what we want to see is, uh…
[01:08:20] the year will be the distance of the forecasted and the mean value.
[01:08:28] And this is the…
[01:08:30] distance between the actual value and the mean value.
[01:08:33] And we look at the differences and square it.
[01:08:36] And see the R-squared value.
[01:08:41] So, just manual calculation I wanted to show.
[01:08:45] How we calculate R squared.
[01:08:48] And I think, uh, just now, as it was asked how to…
[01:08:53] calculate. So, we have X, we have Y,
[01:08:55] Then, once again…
[01:08:59] We have these as Y hat as the predicted values. These are the predicted values of i hat.
[01:09:06] So, uh, the difference between, uh,
[01:09:09] the value, and of course, the mean of all values of X
[01:09:13] and Y are shown here, so the…
[01:09:17] Mean of X, uh, would be 1 plus 2 plus 3 plus 4 plus 5, which I have not written here, will come out to be 3.
[01:09:23] And the mean of Y will be 3 plus 4 plus 2 plus 4 plus 5 divided by 5.
[01:09:29] Which comes out to be 6. So this is mean of Y.
[01:09:33] Which comes out to be 3.6.
[01:09:36] Now, we find out the difference between the
[01:09:39] given value of Y and the mean of Y.
[01:09:42] So it is 3 minus…
[01:09:45] 3.6, then it is…
[01:09:47] 4 minus 3.6.
[01:09:50] Then it is 2 minus 3.6.
[01:09:53] then 4 minus 3.6 and 5 minus…
[01:09:57] 3.6. So, like this.
[01:09:59] We calculate Y minus Y bar.
[01:10:02] Then we calculate Y minus y bar square.
[01:10:07] So here we have calculated the square values, and then summed them all up, then we get summation equal to…
[01:10:14] 5.2.
[01:10:17] Then you find out, uh…
[01:10:19] Y-hat, which are the predicted values of Y.
[01:10:24] And then the difference between the predicted and the given value.
[01:10:28] Which is, like, 2.8 minus 0.36, which is minus 0.8, like that. We keep on finding out that
[01:10:36] differences, and then we square them.
[01:10:38] And we find out 1.6 as the 1.
[01:10:42] Now you substitute these values, so Y hat minus y bar whole squared. This is y hat minus y bar whole squared.
[01:10:49] So this is 1.6.
[01:10:52] Divided by Y minus Y bar whole squared. So this was Y hat minus Y bar whole squared, this is Y minus y bar whole square.
[01:11:00] So now you have Y minus Y bar whole square, which is 3 point, um…
[01:11:06] you know, a square of 3 point… sorry.
[01:11:10] Y minus Y bar whole square. Here it is, sorry.
[01:11:14] This is 1.6.
[01:11:16] Uh, this is, sorry, Y hat minus Y bar whole square, and then we have…
[01:11:21] Uh, this one.
[01:11:25] Y minus Y bar whole square.
[01:11:28] So it is coming out to be 5.2, so I've substituted directly, I've not shown
[01:11:33] each and every distance square and added, final value is 5.2.
[01:11:38] So I've just calculated that and substituted, and what I've got is R squared is equal to 0.3.
[01:11:45] Which means that our model is able to capture 30% of the variance
[01:11:50] uh… being represented by the independent variable.
[01:11:56] Which is not a very good thing, just 30% variance is captured only.
[01:12:01] Not much.
[01:12:04] Then we have other performance metrics, which is a root-mean-square error, which is the RMSE.
[01:12:10] The RMSC is, uh, nothing but…
[01:12:14] the actual value, this is the predicted value, why I…
[01:12:19] hat is the predicted value for YI.
[01:12:23] So, the predicted value of YI minus the actual value of YI
[01:12:28] squared, and then this summed over all the samples.
[01:12:35] If n is equal to…
[01:12:36] Total number of samples.
[01:12:42] In the…
[01:12:45] training data.
[01:12:50] So, if… sorry for the…
[01:12:52] bad handwriting. So, N is the total number of samples in the training data, then we divide this whole…
[01:13:00] Y1 by N, which is root mean squared error.
[01:13:04] So this is how we find out the root-mean-square error. This is also used
[01:13:08] As, uh, one particular, uh…
[01:13:12] meetric to find out.
[01:13:14] Uh, the value.
[01:13:16] Any questions, anyone?
[01:13:18] So far.
[01:13:25] Okay.
[01:13:28] Yes?
[01:13:29] Ma'am, previously, we assumed key errors follow a normal distribution.
[01:13:32] Yes.
[01:13:33] Uh, what is the relevance of this concept, or assumption?
[01:13:37] Uh, what are we trying to prove here by assuming that these errors follow normal distribution?
[01:13:45] So, uh, error follows a normal distribution means that
[01:13:49] On an average,
[01:13:51] Uh, you know, all the naturally occurring observations, if we talk about temperature, pressure, they all follow normal distribution.
[01:14:01] Which means that, uh, mostly the value will be around average.
[01:14:06] So we want error also to follow a normal distribution, because
[01:14:11] We want that, uh…
[01:14:13] The error values will not be very low, they won't be very high. There will be some average value all the time.
[01:14:22] Yes. Yes.
[01:14:23] That is why we assume that it is a normal distribution, because
[01:14:26] Uh, let me just share it again.
[01:14:35] So, for example…
[01:14:38] Just a sec.
[01:14:43] this. So, you see that this is very high, this is the highest error.
[01:14:48] But, uh, this will be… this kind of higher run will be very few… in very few samples.
[01:14:54] Which will be less than 10%, according to the normal distribution.
[01:14:59] So this low values will be very low, and this also will be very low. Usually, they will lie in this range.
[01:15:06] And this is what we want, that more or less, the average error will be kind of constant.
[01:15:10] Uh, when we use this kind of, um…
[01:15:14] linear regression.
[01:15:16] Yes, ma'am. Uh, these errors will follow the specific normal distribution when the dataset is large.
[01:15:23] But what if the dataset that we are using is small?
[01:15:27] Then, uh, it is not necessary, they will follow the normal distribution, because
[01:15:31] normal distribution, uh, basically, uh, says given the dataset is huge, or we are taking a large number of data sets.
[01:15:36] Yeah. Yes.
[01:15:39] So, what you're saying is correct, that if we are talking about just 2, 3, 4 examples, like in this case, you can see here,
[01:15:47] Yes, ma'am.
[01:15:48] that probably this is also not falling a normal distribution.
[01:15:52] So, uh, when we are doing regression, we assume that our training data is big enough.
[01:15:57] Yes, sir.
[01:15:58] That it is actually representing the…
[01:16:00] Uh, the data that…
[01:16:03] Uh, we will be using… I mean, that we will be…
[01:16:07] You know,
[01:16:09] That is, for which we'll be making prediction. It will be big enough to represent that kind of use cases.
[01:16:15] So, definitely, the examples here are very small and trivial examples, only to illustrate
[01:16:22] However, our data, when we talk about the training data, should be big enough.
[01:16:27] Yes, ma'am. Understood.
[01:16:28] Yeah, yeah, yeah.
[01:16:31] Yes, uh, I could see a hand raised, I think.
[01:16:36] Yeah, ma'am. Ma'am?
[01:16:38] Yes, Adity.
[01:16:39] Is it… is it safe to say R-square and, uh…
[01:16:43] mean square error slash root mean square error are inversely proportional.
[01:16:50] Um, so if we look at root mean square error…
[01:16:54] This is the formula, see.
[01:16:57] So, in the formula, you have under root here, uh, you have the predicted minus the…
[01:17:04] Actual value, right?
[01:17:05] And then you have the summation out here, and then divide.
[01:17:08] Now, if you compare this…
[01:17:09] Okay.
[01:17:11] with, uh…
[01:17:13] R squared. It is quite different.
[01:17:16] So, in R-squared, you have Y hat minus Y bar.
[01:17:20] And then Y minus Y bar whole square, and then both of them added. There is no square root.
[01:17:27] And all, and there is no mean.
[01:17:29] So, it may not be appropriate to say…
[01:17:33] that there is a direct inverse or a, you know, direct relationship.
[01:17:37] between them.
[01:17:39] We cannot say directly that thing.
[01:17:42] Okay, I was thinking more of, uh, uh, without using formulas, ma'am.
[01:17:47] Mm-hmm.
[01:17:48] in a sense that if I have high R-square,
[01:17:51] Hmm.
[01:17:52] Which means the explainability of variance is more.
[01:17:54] Which means the error is less.
[01:17:57] Therefore, your mean square error is less, and root of mean square error will also be less.
[01:18:02] Yeah, that… that way, that kind of interpretation is fine.
[01:18:07] Okay.
[01:18:08] And that's how we actually use the error and R-squared together.
[01:18:13] And you might be seeing it in code, and you might have seen it, or will see it now.
[01:18:18] Hmm.
[01:18:19] in the code, uh, that we go for models that show higher
[01:18:23] or relatively higher values of R squared.
[01:18:26] Which means that the variance
[01:18:28] Captured is high, obviously, which means that if the variance in the dependent variable captured is high,
[01:18:36] then definitely the error may be low, and whatever is…
[01:18:40] Uh, the measure that we use.
[01:18:42] Whether mean squared error, or RMSC, or MAE, or whatever we use.
[01:18:47] That will give a lower value only.
[01:18:50] Yeah.
[01:18:51] Okay.
[01:18:53] Any other questions, anyone?
[01:18:59] I'm also having the same kind of a thought I'm thinking of is, uh… So, I mean, in general, can we say that as mean square error is how my model.
[01:19:10] How wrong, uh… It is at this moment, and R-square is how much better.
[01:19:16] My model is doing. With the numerical calculations.
[01:19:22] Yeah, definitely. R squared…
[01:19:24] definitely means how good the model is.
[01:19:27] The higher the value of R squared, the better it is.
[01:19:31] And definitely, your RMSC or any other mean square error would be
[01:19:37] Uh, the value, if it is high, means how bad it is.
[01:19:40] So, the two of them, one is good, one is…
[01:19:44] bad. That way, you can think of it. That's okay, yeah.
[01:19:50] Okay, so that's how we can improve SSS to… A cut short MSC as much as and increase the value of R aspire as much as.
[01:20:00] Yes, that's correct. That's correct. What you're saying is correct.
[01:20:01] To get an optimized result.
[01:20:04] that we want to optimally increase our square.
[01:20:10] And we want to decrease
[01:20:13] The error, which can be measured in whatever way, RMSE, MAE, MSC, or whatever.
[01:20:19] That's correct.
[01:20:23] Thank you. And one more, uh, this, as you mentioned, is in the earlier sessions, uh.
[01:20:29] in the coding sessions, uh… Someone has showed us the R-Square and other thing.
[01:20:36] Mm-hmm.
[01:20:41] Hmm.
[01:20:42] I'm not able to recollect. In this session, because it's, uh, the Ashish, which the session he's conducted, he is talking about, more about.
[01:20:46] Yes, yes.
[01:20:47] classification things. Have we not touch borders RSquire and other parameters here in the coding, I guess.
[01:20:52] Okay, no problem, because, uh…
[01:20:55] Uh, during the last session, uh, I was not there.
[01:21:00] So usually,
[01:21:01] Oh, that was a master class session, uh… One of the ways they can.
[01:21:05] Yeah, yeah, no, I'm not talking about the master class. In fact, uh…
[01:21:09] And, uh, I… I think I got to know that in the master class.
[01:21:14] Uh, probably regression was taken up, which I…
[01:21:18] Yes.
[01:21:19] Yeah, so I had… I had not recommended that one. I think probably that got overlooked, because I had not, uh…
[01:21:27] taught regression to all of you.
[01:21:29] So, I wanted, uh, some other topic. However, that's okay.
[01:21:30] Yeah.
[01:21:34] Uh, so we will be taking R-squared, don't worry.
[01:21:38] So, what I was assuming was that when Ashish, in the last turn took
[01:21:43] hands-on on, um, other models in supervised learning.
[01:21:49] So, R squared is not just used for linear regression, it can be used for any…
[01:21:50] Mm-hmm.
[01:21:55] Any supervised learning, that is, any classification model to find out the goodness of the model.
[01:22:01] So, I assume that it was taken. However, even if it is not taken,
[01:22:02] Okay.
[01:22:06] It will be taken up, probably tomorrow.
[01:22:09] Okay, so tomorrow we'll be having the hands-on.
[01:22:12] And in that, we'll include the code on knife, and uh…
[01:22:17] supervised learning.
[01:22:19] Both.
[01:22:21] Uh, have you all, um… I assume, I think you already have done KNRS Neighbor, right?
[01:22:28] Yeah, so, uh, and whenever we are adding on any supervised learning model,
[01:22:35] Then we'll always have, at the end, a kind of a comparative, if that's possible, on that kind of data.
[01:22:42] to see which model is working best on that particular data that we will choose for that date.
[01:22:48] So, uh, we will be using R-squared also in the code, okay? So don't worry.
[01:22:55] Even if it isn't done so far.
[01:23:00] Any other questions, anyone?
[01:23:01] Thank you.
[01:23:03] Uh, how do we figure it out, this, whether, uh…
[01:23:08] linear regression will be appropriate for a…
[01:23:11] Given that, as if…
[01:23:13] Do we have any, uh, criteria or, I mean…
[01:23:18] um… ways to look up the data and choose the model… I mean, choose the regression, not how…
[01:23:26] Yeah, yeah, so that's a good question, Najkumar.
[01:23:30] So, there can be some…
[01:23:32] Some ways, obviously, again, when we talk about machine learning or AI, a lot of things are…
[01:23:38] kind of hit and try, or trial-based iteratively, they are decided.
[01:23:43] So, there can be one or two methods.
[01:23:46] One, you use linear regression and you use other models also.
[01:23:50] And see the performance.
[01:23:53] If the performance of the model
[01:23:57] uh, is coming out, suppose you use linear regression, and the performance metrics, such as the error,
[01:24:03] And the loss and all that is coming out good.
[01:24:08] Then you assume, uh, that if you have used linear regression and the values of the error and all are coming out to be low,
[01:24:15] It means that the data was actually linearly separable.
[01:24:20] Okay, so, um…
[01:24:23] The other ways that you have some, uh, idea about the data.
[01:24:28] If you have some idea about that data, then you…
[01:24:32] probably may know whether it is actually linearly separable or it is non-linearly separable. If you have that kind of domain knowledge,
[01:24:41] Then you can use that to…
[01:24:43] Choose the appropriate.
[01:24:45] classifier, model, or prediction model. That way also can be done.
[01:24:49] But mostly, many a times, because we don't have much idea about that data,
[01:24:55] Then what we do is we try to use different kinds of classifiers, and then choose that classifier.
[01:25:02] That gives us the best, uh, accuracy, I mean, best performance and lowest error.
[01:25:08] And, uh, let's say if that one is the linear classifier, then definitely the data was linearly separable.
[01:25:18] Yep. Um…
[01:25:21] Sometimes we can also visualize the data. If our data is not overly complex, then we can visualize that data.
[01:25:28] We can do a PCA on the data, visualize it, and see.
[01:25:32] How are the classes, whether they are linearly separable or they are non-linearly separable, that also we can do if that is possible.
[01:25:41] We can visualize that data, and then that way also, we can identify.
[01:25:45] Uh, some information about how… what kind of data it is, and…
[01:25:50] Whether it is linearly separable or not, and all that.
[01:25:57] Yeah. Any other questions, anyone?
[01:26:01] Yeah.
[01:26:02] Uh, yes, ma'am. So, you just mentioned PCA, right? So, PCA we do when we want to reduce the dimensions.
[01:26:06] Yes, yes, yes.
[01:26:07] Right? So, if we reduce the dimension, and then we use linear regression, that would give us the same… similar results, right?
[01:26:15] Yes. So, the idea of doing PCA, uh, is…
[01:26:18] Hmm.
[01:26:19] that the overall distribution of the data and the relationships amongst the dependent-independent variable and all,
[01:26:27] will be retained, and that's why we use PC. Otherwise, if…
[01:26:31] PCA would alter
[01:26:33] Uh, the overall nature of the data, then we won't use PC.
[01:26:40] So, which means that should be our go-to strategy, right, ma'am? Like, if… because most of the times, we will have more than 10, 15 attributes.
[01:26:46] Yes, yes, yes.
[01:26:47] In our data. Okay.
[01:26:48] So, uh, if you want to visualize that data, it's a good idea to use PCA.
[01:26:54] Reduce it and visualize it.
[01:26:57] And then see what kind of, uh, you know,
[01:27:00] Now, what is the distribution like?
[01:27:02] And usually, in most of the datasets, EDA is always done as a preliminary
[01:27:07] data analysis, and when you do EDA, then you definitely want to reduce the number of attributes or dimensions.
[01:27:15] And for that, you can use PC.
[01:27:18] Yeah.
[01:27:19] But should I run linear regression on the PCA data, or should I run it on the… entire feature set.
[01:27:27] So, it's always a good idea to do the PCA on the actual data.
[01:27:33] Because, uh, when you use PCA, you have actually reduced the data and made it more compact.
[01:27:40] By reducing the number of dimensions, though the overall
[01:27:44] nature of the data won't change, however…
[01:27:49] Some loss in data would be there, right? Because we are reducing the number of…
[01:27:53] dimensions. So, uh, so…
[01:27:54] Yeah. Mm-hmm.
[01:27:57] you can. I mean, it would be advisable to use
[01:28:02] Regression on the original data, and not on the reduced data.
[01:28:06] However, if you still do it on the reduced data,
[01:28:10] Your results would more or less match.
[01:28:13] Uh, with some tolerance, that is, with some error difference.
[01:28:18] So, definitely, your error will increase.
[01:28:22] If you, uh, do regression using the PCA data,
[01:28:27] And versus that, which you would have obtained, the error that you would have obtained on the original data, that would be definitely lesser.
[01:28:35] So, it's good to use the original data only.
[01:28:38] Uh, for regression.
[01:28:42] Okay. And… do I have to run any kind of a clustering?
[01:28:48] Uh, to understand the data. Before I run any classifier model, or it's not advisable?
[01:28:54] You want to run clustering to understand the data? That's what you're saying?
[01:28:59] Yeah, I mean, there are multiple cases, right, ma'am? Like, if the data is actually resulting in.
[01:29:04] multiple clusters, then I don't think linear regression would be a suitable.
[01:29:08] Answer for that, right? Because then… if there's a line which is splitting the data, and we have more than two clusters, then how it's going to identify, let's say, in this way.
[01:29:20] in this particular graph, there are 4 clusters of data, right?
[01:29:23] One starting just 2 above the line and 2 below the line, but those are, like, in four different quadrants, let's say, in that way.
[01:29:30] Mm-hmm. Okay.
[01:29:32] So, would that… in that case, would that make sense if I run the clustering and then I run linear regression for each of those clusters?
[01:29:39] Uh… okay.
[01:29:40] Individually?
[01:29:42] No, uh, clustering, if you want to run regression for each and every cluster,
[01:29:49] So, that is a specific use case of regression. Now, let me…
[01:29:54] First, take up your first point that, uh…
[01:29:55] Mm-hmm.
[01:29:58] Whether clustering, doing clustering is going to be helpful. So, let's say…
[01:30:03] You have this data.
[01:30:05] This is your…
[01:30:07] I mean, your data space.
[01:30:10] And you have these data clusters.
[01:30:14] probably one cluster here.
[01:30:20] And maybe one cluster here.
[01:30:22] like this. Suppose you're having three clusters.
[01:30:23] Yeah. Yeah.
[01:30:26] And, um, these are the two classes in the data.
[01:30:32] And even when there are clusters, these are linearly separable.
[01:30:36] Therefore, having clusters is… doesn't make it necessary.
[01:30:41] that linear regression is not going to work.
[01:30:44] As long as the data is linearly separable, I mean the classes, it's fine.
[01:30:45] Mm-hmm.
[01:30:49] Even though there are multiple clusters here, you can see that the data is
[01:30:55] linearly separable. That is your first thing.
[01:30:58] Now, the second point about using regression,
[01:31:03] on each of these clusters.
[01:31:05] So, um…
[01:31:06] Yeah.
[01:31:07] That is not directly a use case of simple, I mean, regression.
[01:31:12] But let's say that…
[01:31:16] Within these clusters, these are some points.
[01:31:18] Or you can, you know, just think about…
[01:31:21] Uh, this data space…
[01:31:23] Okay.
[01:31:27] Right, like this.
[01:31:29] And suppose there are some points here, and I'm going to represent them as stars.
[01:31:34] And these points are having some missing values in their attributes.
[01:31:39] And we want to infer those. So, one way could be…
[01:31:43] that since, uh, if we use a simple mean of the…
[01:31:48] attribute that may not be good because these values are very different.
[01:31:53] Right? So what we could do is cluster these points, which I had also covered in the pre-processing part,
[01:31:55] Yeah.
[01:32:00] And then for each cluster, you…
[01:32:01] Yeah.
[01:32:02] do regression, whether it is simple or simple linear regression, or multiple, or whatever it is.
[01:32:10] You do regression and represent
[01:32:12] These set of points with some equation,
[01:32:16] let's say it is Y hat, uh…
[01:32:19] say, A equal to…
[01:32:22] beta naught plus beta 1.
[01:32:25] X. Right.
[01:32:27] And then you represent this one with another equation, say Y had B equal to
[01:32:33] We are not prime plus beta 1 prime X.
[01:32:36] And then you represent this with another one, say…
[01:32:37] Hmm.
[01:32:41] This…
[01:32:43] like this. Now, what you can do is that…
[01:32:45] These missing values, you can infer by…
[01:32:48] You know, these equations.
[01:32:50] So this is, uh…
[01:32:52] a use case for the example that you gave, that you have clusters.
[01:32:57] And, uh, you use regression.
[01:33:01] to fit lines for each of these clusters, and then what are you going to do with it? You can use it to fill in
[01:33:06] missing values in a very good way.
[01:33:09] Okay?
[01:33:10] But if there is no missing value, ma'am, and I'm actually creating a classifier on some, let's say, X.
[01:33:14] Hmm.
[01:33:15] is the value which we are seeing is… can be missing, and I'm doing classifier on Y.
[01:33:19] Mm-hmm, mm-hmm.
[01:33:24] Okay.
[01:33:25] Right? And Z is the function. Okay, so in that case, should I run it individually still? Because I don't have any missing value, but I… want to do the classification on why, which can have, like, any three… which can be any of these.
[01:33:37] three clusters value. So let's say cluster 1 depicts.
[01:33:41] Uh, answer as A1 cluster 2 depicts answer as A2, and cluster 3 depicts answer as A3.
[01:33:46] So, if you're doing a simple
[01:33:49] prediction of values, okay?
[01:33:52] And you don't have any missing values.
[01:33:53] Hmm. Okay.
[01:33:56] Then what you are concerned with is finding out the model for this separation.
[01:34:02] In that case,
[01:34:03] Correct, correct.
[01:34:04] if your data is linearly separable, then you use linear regression to find out
[01:34:10] this equation, and that's all.
[01:34:12] Okay.
[01:34:13] If your model is not linearly separable, then you use some other supervised learning method.
[01:34:18] to represent, uh, you know, this relationship. So, you are trying to find out this function F.
[01:34:24] In this case, this function f is this straight line, right?
[01:34:25] Hmm. Correct.
[01:34:28] But if it is not linearly separable, then you can use some other
[01:34:32] method, for example, if you use a decision tree,
[01:34:37] Then, what you would have is a decision boundary, which will be very non-linear.
[01:34:41] Because what you are having are a set of rules.
[01:34:42] Hmm.
[01:34:45] And these rules are together, when you use them together, they will be…
[01:34:50] Uh, you know, together there'll be…
[01:34:53] representing a very non-linear decision boundary, right?
[01:34:58] Because, let's say you are having some value, let's say X1, X2, X3.
[01:34:59] Mm-hmm.
[01:35:03] Audio 3 attributes.
[01:35:05] And you are doing some kind of a separation.
[01:35:08] on the values of X1, say,
[01:35:11] you know, greater than 5…
[01:35:13] And, uh,
[01:35:15] less than or equal to 5, and here you are having something like…
[01:35:20] may be greater than something…
[01:35:23] say 100, and maybe less than or equal to 100, and here you are doing something like…
[01:35:29] 10 to 100, and here you're doing less than equal to…
[01:35:34] 10 or greater than equal to 100, something like this. Now, this combined…
[01:35:38] If you represent it in a data space, it will come out to be a very non-linear boundary.
[01:35:44] So, if… if your, uh…
[01:35:47] Data is non-linearly separable, then you use some other supervised learning model, but
[01:35:53] Your concern is just to find out this function f, and nothing more than that.
[01:36:00] Yeah.
[01:36:01] Okay. So, I think it's not possible just to… just by looking at the data to assume which algorithm.
[01:36:07] We have to use, we still have to maybe do some kind of an exploratory data analysis on that to understand that first.
[01:36:13] Yeah, so…
[01:36:15] Yes.
[01:36:16] And then only we can make the choice, that whether it's linear regression decision tree or something else.
[01:36:18] Yeah, that's correct. It's always good to actually visualize the data.
[01:36:22] to get some idea on what is the data distribution like.
[01:36:23] Mm-hmm.
[01:36:26] And would there be clusters, or what kind of relationship there may be amongst the points or classes, or whatever?
[01:36:33] And it's usually done, ET is usually done.
[01:36:34] Hmm.
[01:36:36] And, uh, then, while actually finding out or deciding on a model,
[01:36:41] The good idea is… a good idea is to try out some different kinds of models on the data.
[01:36:47] And look at the performance of these models.
[01:36:51] And maybe the R-squared and all other things like that.
[01:36:54] And then identify which particular modeling technique, that is which particular supervised learning method,
[01:37:01] is working best, and then choose that one.
[01:37:07] Yeah, okay. Any other questions, anyone?
[01:37:08] Cool. Thank you.
[01:37:15] Okay.
[01:37:24] So then, um…
[01:37:26] I'll just start…
[01:37:28] something on, uh…
[01:37:30] Logistic regression also.
[01:37:35] I think there is a question
[01:37:38] Yes.
[01:37:39] Yeah, I have no other question, is that, uh…
[01:37:40] Mm-hmm.
[01:37:41] Till now, we have looked for that, um…
[01:37:43] two variables is working, right? Like, X is there, right?
[01:37:48] How it is going to work when we have a multi-dimensional or, like, a multiple variables are there.
[01:37:53] Then how this formula, which we have created, right, how is it going to be created in that?
[01:37:58] Uh, you're talking about regression right now?
[01:38:01] Okay. I'll show you. I had shown you, but I'll show you again.
[01:38:06] Uh, so I'll share the slides again. In the meantime, there was a question that
[01:38:12] How is R2 going to help in the classification?
[01:38:15] So, R2 will help to assess how good our model is at
[01:38:19] Capturing the variants in the dependent variable.
[01:38:22] If the value of R squared is high, it means that it is capturing
[01:38:26] The variability or the variance in the dependent variable to a good extent,
[01:38:32] Which means that our classification model is good.
[01:38:35] If R-square is low, it means that the variance in the dependent variable
[01:38:40] is not being captured properly.
[01:38:43] Uh, or, well, good enough by our classification model, so we'll have to choose some other classification model.
[01:38:50] So, I hope Krishna Kumar, your queries answered.
[01:39:03] Okay, coming to your question, Gunjan.
[01:39:06] I'll just show you.
[01:39:18] I don't…
[01:39:34] So, this is the case.
[01:39:36] So I showed you, first simple, uh…
[01:39:39] Simple linear regression. In simple linear regression,
[01:39:44] We have just one dependent variable and just one independent variable.
[01:39:49] And the predicted value, Y, is represented as…
[01:39:53] We cannot plus beta 1x.
[01:39:54] Yes, yes.
[01:39:56] Correct, yes.
[01:39:58] Where X is just. But if X, uh…
[01:40:01] So this is one attribute, but let's say X is a set of attributes.
[01:40:06] say, X1, X2, X3 till XN.
[01:40:08] Then we use multiple linear regression, which is what you're seeing here. So here, Y hat, which is the predicted value of Y,
[01:40:16] is given as beta naught plus beta 1x1 plus beta 2X2 plus beta 3x3 tilde beta and Xn.
[01:40:22] And this X1 is attribute 1.
[01:40:27] This is attribute to…
[01:40:30] This is attribute 3.
[01:40:32] This is the nth attribute.
[01:40:35] So this is how this gets changed.
[01:40:38] Is that what you wanted to know?
[01:40:39] Uh, yes, but my point is that, ma'am, uh,
[01:40:43] So, example, if I look at this, uh…
[01:40:46] In the real scenarios, when we have to predict something, let's say Y, we have to predict.
[01:40:50] Right? But instead of only a single attribute, we have, let's say, 10 different attributes.
[01:40:57] Hmm.
[01:40:58] Right? And based on the 10 attributes for, let's say, we already… like, it's like a multiple rows and multiple columns are there.
[01:41:06] Sure. Yeah.
[01:41:07] Right? So…
[01:41:08] Then, if you have to train the model, right, to define the QKY relationship we have to define, right? So we have, let's say, dependent variable, we have… independent variable, we have, like, um…
[01:41:17] A, B, C, D, E, F, 6 are the parameter. For example, age,
[01:41:20] your gender, your location, and all. These are there.
[01:41:24] Okay.
[01:41:25] Right, so how it is going to fit in that to predict why?
[01:41:30] Um… I still couldn't understand your question, because…
[01:41:32] Uh-huh. Yeah?
[01:41:35] Uh, let's say you are having some, you know…
[01:41:40] or class label.
[01:41:41] For example, earlier you were given one example long back for the California House
[01:41:45] Mm-hmm.
[01:41:49] Hmm.
[01:41:50] Right? The proximity, so a lot of things are there, right? For the… if you predict the house value,
[01:41:51] Hmm.
[01:41:53] Yeah.
[01:41:54] Right? We have to look for that locality, the proximity, and everything, right? So, multiple parameters are there to identify
[01:41:56] Yeah.
[01:41:59] Yes.
[01:42:00] What is going to be the value? The same, the same case, where we have, let's say, a lot of
[01:42:01] independent variables are there.
[01:42:03] Now, we are identifying something based on that, why should… what is going to be my value of Y?
[01:42:09] Okay.
[01:42:10] So, in that case, how this formula is going to fit?
[01:42:14] Yeah, because, see, whatever those attribute or variables are,
[01:42:19] Those are represented as X1, X2, X3 till XN.
[01:42:22] There can be n number of variables, say…
[01:42:25] Let's say the location of the attribute may be… the location of the property may be one attribute.
[01:42:30] Hmm.
[01:42:31] Maybe the ear…
[01:42:34] Uh, of sale may be one attribute.
[01:42:36] Maybe there might be some other economic factors or something else, maybe the number of bedrooms, maybe another attribute, so whatever those attributes are represented as X1, X2, X3 till XL.
[01:42:48] In this equation. And to use this equation, we'll have to find out the coefficients beta 0, beta 1, beta 2, beta 3,
[01:42:55] Till veto in such a way that these coefficients
[01:43:00] Uh, help to find out, uh…
[01:43:02] regression line.
[01:43:04] Uh, which fits this given data with the least error.
[01:43:08] Yeah.
[01:43:09] Right? So this is how it is going to work, that if you have multiple attributes,
[01:43:14] You don't have a single attribute, then you use multiple linear regression.
[01:43:18] The foundation is still the same, that
[01:43:21] While you're computing these coefficients,
[01:43:24] We try to find out these coefficients, uh, with the aim to
[01:43:29] have the error being minimized by this Y-hat. So, the prediction value
[01:43:35] must be quite close to the actual given value. That is the objective.
[01:43:41] So, everything remains the same, only thing the equation gets changed, and accordingly, the number of coefficients that need to be found out
[01:43:47] varies.
[01:43:48] Okay.
[01:43:50] Yeah. So, the foundation…
[01:43:53] is still the same. So, no matter whether we use regression, which may be simple linear regression,
[01:43:59] or multiple linear regression, or we may be using decision tree, nearest neighbor, or whatever.
[01:44:05] Mm-hmm.
[01:44:06] We use. Our objective is that
[01:44:08] Uh, we want to minimize the loss or the error.
[01:44:13] Right? So…
[01:44:16] If we are… if we know that we want to use multiple linear regression,
[01:44:20] Then we have to find out the coefficients in such a way that the error gets reduced.
[01:44:25] actively reduced.
[01:44:28] Uh, only those coefficients will be chosen.
[01:44:31] Obviously, if you use multiple linear regression in Python, you won't be…
[01:44:35] Uh, hand coding, the finding out, I mean, you won't be actually
[01:44:40] manually computing beta 1, beta 2, and all that, and it will be pre-coded in that function, which it will do.
[01:44:47] And finally, what it will give you will be the optimally best
[01:44:52] modern, uh, multiple linear regression, or simple linear regression.
[01:44:56] But I just wanted you all to have a basic idea of what is happening internally, so I'm showing you these equations.
[01:45:02] Gotcha.
[01:45:03] Uh, yeah. But…
[01:45:05] any model, any supervised learning model,
[01:45:08] Uh, with all the attributes that are given in the data as the input to the model.
[01:45:13] aims to reduce the error.
[01:45:16] on the prediction,
[01:45:18] Uh, to the optimally low… lowest value.
[01:45:21] This is the thing.
[01:45:23] Okay.
[01:45:24] Now, method could be anything. It could be multiple linear, simple linear, it could be decision tree, k-nearest, neighbor, whatever.
[01:45:30] Okay. Noted, man.
[01:45:32] Yeah, yeah.
[01:45:35] Any other questions, anyone?
[01:45:39] Because I thought of starting, uh, logistic regression, though we have it for the next week, but I think
[01:45:46] There isn't much time to start, but I will just introduce you.
[01:45:51] To logistic regression, but before we do that…
[01:45:54] What I wanted to…
[01:45:57] Uh, okay, I think I can see some questions on the chat.
[01:46:02] So, uh…
[01:46:03] I think, uh, I can see a question that how to interpret, uh, between accuracy or any other confusion matrix.
[01:46:11] waste, uh…
[01:46:13] metrics like recall or precision.
[01:46:17] Versus R-square. Okay.
[01:46:19] So, I think, uh, the question is that if we have R-square versus if we have…
[01:46:25] accuracy, precision, recall, and all, then which one to use, or how to use…
[01:46:31] how to use these. So, let me tell you that precision recall F1 score and all,
[01:46:37] are ways to represent the goodness of the model in terms of
[01:46:43] error, or, uh, you know, performance of the model.
[01:46:46] R-square, uh…
[01:46:49] measures the goodness of a model in a different way, because it tries to capture the variance in the
[01:46:56] how good the model is able to capture the variance.
[01:47:00] So, R-square usually complements the usage of
[01:47:03] Uh, recall precision…
[01:47:06] accuracy, F1 score, and all.
[01:47:08] So, accuracy, precision, recall, F1 score,
[01:47:12] False positive rate, false negative rate, and all, they are all…
[01:47:15] You can say one group, and R-square is another group.
[01:47:20] And together, these two…
[01:47:22] can, you know, complement each other to find out the best
[01:47:26] performing model.
[01:47:28] Okay? So this is the difference between them.
[01:47:36] So, that's it.
[01:47:38] Uh, any other questions, anyone?
[01:47:42] Uh, I think I can see another question by Sachin.
[01:47:46] So, when working on linear regression, the question is, when working on linear regression,
[01:47:51] If, uh, we get low MSC, then model predictions
[01:47:55] is closer to actual values.
[01:47:58] Uh, then still, do we have to evaluate root-mean-square-E error also?
[01:48:03] Okay. So, uh, actually…
[01:48:07] I totally while explaining the error metrics, that we have a large number of ways to measure the error, like
[01:48:16] mean square error, root mean square error, mean absolute error, and so on and so forth.
[01:48:21] However, it depends, uh…
[01:48:23] Uh, on the use case, suppose we want to exaggerate the error, we want to measure the error,
[01:48:30] We want to use a metric which exaggerates the error, then we use…
[01:48:35] Uh, mean square error, or root mean square error.
[01:48:38] But if we just simply want to assess the error, then we use mean absolute error.
[01:48:44] So, it's not necessarily…
[01:48:46] It's not necessary that if we have used MSC, and if it is coming out low,
[01:48:51] Then we also have to use RMSC.
[01:48:55] Because if MSC is low, definitely our MSC will also be low.
[01:48:58] Because RMSC is nothing but the root
[01:49:02] of, uh, in a way, root of MSC.
[01:49:05] So, it's not necessary to find
[01:49:07] out RMSE,
[01:49:09] If we have used MSC.
[01:49:11] Okay.
[01:49:17] Okay.
[01:49:18] Yes, ma'am, yeah, thank you.
[01:49:19] Um, any other questions, anyone?
[01:49:25] So before that, I wanted to discuss, uh…
[01:49:28] this thing that I think, uh, in the coming two weeks,
[01:49:33] Uh, I think in the coming week, we are having Christmas, that is midweek.
[01:49:38] And then, uh, the weekend falls on 27th and 28th.
[01:49:43] And then we have, in the next to next week, we have New Year.
[01:49:47] Which, again, falls in the midweek.
[01:49:50] kind of a mid-week Thursday, both of them Thursday only, Christmas and New Year.
[01:49:54] And the Saturday-Sunday are 3rd and 4th.
[01:49:57] So, what I was, uh…
[01:49:59] Uh, thinking was that…
[01:50:05] I could actually, uh, we could have a break on one of these…
[01:50:09] weekends, probably, uh, I like to keep it on the New Year, so you all can celebrate new year.
[01:50:18] kind of extended new year with your friends, family, or whatever you… wherever you want to go.
[01:50:23] So, instead of giving the break on
[01:50:27] 27, 28, let's have…
[01:50:30] the break on the classes on 3rd and 4th.
[01:50:36] I hope that's okay with everyone.
[01:50:39] Um, so that you all can…
[01:50:44] Celebrate New Year in a good way.
[01:50:47] Okay, so I can see.
[01:50:49] I think mostly all of you agree with this, so…
[01:50:52] Uh, the next weekend, we are going to have classes, and next to it, next weekend.
[01:50:58] we'll have no classes, uh, on account of new year.
[01:51:02] And before we wrap up for today, I can see, uh…
[01:51:07] One more comment or question that if…
[01:51:09] I can help to understand multi-collinearity,
[01:51:14] hetero elasticity,
[01:51:16] And all that. So, I think, uh…
[01:51:19] What I can do is that, uh…
[01:51:24] multicollinearity, I will explain. Okay, I'll include, um…
[01:51:28] some concepts to help you explain
[01:51:31] these Krishna Kumar. Right now, I don't have it, but, um…
[01:51:36] I can explain that to you.
[01:51:40] with the help. So, I'll put some material on slides.
[01:51:42] So that it is helpful, too.
[01:51:45] Understand. Okay.
[01:51:47] So, I think, uh…
[01:51:50] Just one second, I can still see…
[01:51:54] Some more comments here.
[01:51:57] Um, okay.
[01:51:59] So, I think that is some question on the assignment. That is one thing.
[01:52:05] And then…
[01:52:10] Um…
[01:52:14] One more comment I can see.
[01:52:17] Uh, that, uh…
[01:52:19] Whether linear regression can handle categorical variables, or it is limited to continuous data.
[01:52:26] So, uh, regression cannot handle categorical data at all.
[01:52:31] Because it relies on values.
[01:52:34] We are going to find out, for example, I showed you…
[01:52:38] We tried to find out the slope, the intercept, and all that.
[01:52:42] So, it is always going to be useful only on numbers.
[01:52:46] And not on categories, okay? So…
[01:52:49] I hope that answers the question of whosoever had asked that question.
[01:52:55] Okay. So, I think so that, uh, with this, uh…
[01:53:00] We'll just come to an end of this session, and we'll continue tomorrow.
[01:53:05] With hands-on on knive and simple linear
[01:53:09] regression, I mean, regression.
[01:53:12] And then the next topic that we'll take in the next week will be logistic regression.
[01:53:18] Okay. So, with that…
[01:53:20] I think we'll break for today.
[01:53:24] Thank you all. Have a great evening. Thank you. Bye-bye.
[01:53:29] Yes?
[01:53:32] Yes, yes.
[01:53:33] I'm… One question. Where do you see that comment or questions? Where do you see that whatever the question being asked, or whatever the comment means, what have you answered?
[01:53:41] actual, uh… In chat, I, uh, I am unable to see that.
[01:53:44] Yeah, we are not able to see all the questions which you are reading right now, recently, no?
[01:53:50] Yes, yes.
[01:53:51] So, I think, am I disabled or will we not visible to us?
[01:53:52] Okay, that's right.
[01:53:53] So, there are two ways. As you all know, that there are some questions that are posted for everyone.
[01:54:00] Got it.
[01:54:01] Right? So, those questions you can read, but some of you write individually also to me.
[01:54:06] No, but, uh, but as per means, uh, as per my view, means whatever the chat is there, whatever the question everybody is asking.
[01:54:14] Mm-hmm.
[01:54:15] Uh, actually, uh, I think we should see that question so that we can.
[01:54:20] Even after seeing that question, we can also imagine that, okay, if this is a question according to that.
[01:54:27] Uh, even something our mind, we can also ask some questions, so.
[01:54:32] Chat should be visible to all of us, ma'am. Uh, this is my request.
[01:54:36] For what, ma'am is telling that, so people… some people are, like, directly, I mean, direct message to the MAM itself.
[01:54:41] Yeah, yeah.
[01:54:44] So, instead of everyone. So, that is what's… yeah.
[01:54:46] Okay, okay.
[01:54:47] So, ev… so, questions return on, um…
[01:54:52] Chat visible to everyone can be seen by everyone, and that's the reason, if you might have noted, I always read the question. I try to read the question.
[01:55:00] Even if it is messaged directly only to me,
[01:55:03] I may or may not say the name of the person, but I read the question and then tell the answer.
[01:55:10] So that you understand what question I'm answering.
[01:55:15] Uh, but it is for all of you, I mean…
[01:55:19] Yeah.
[01:55:20] Okay, yeah, got it, ma'am. Actually. Yeah, in chat, there is option, means it is meeting group, or it is directed to you, okay?
[01:55:27] So, one request to all my colleagues, whoever is asking question.
[01:55:30] Can you please ask question in chat so that we can also understand what question you are going to ask?
[01:55:36] Because, uh, you are asking questions directly, ma'am, so we are, uh, means… Something we are, uh… not knowing what you are going to ask, or sir.
[01:55:47] Okay.
[01:55:48] So… No, not for you, ma'am, I'm asking for our colleague.
[01:55:51] Yes, yes, I… I understood. You are making a request to the other learners.
[01:55:52] Okay. Thank you.
[01:55:56] Even if there are direct messages, what I will always do, which I have been doing now also,
[01:56:03] that at least, even if the question wasn't visible to you in advance, I will read the question.
[01:56:09] Uh, or I'll tell what is the question, and then tell the solution, so that at least you understand what was the question.
[01:56:16] And what answer I'm giving you is for that question.
[01:56:20] So that you can be benefited in a better way.
[01:56:25] a member, uh…
[01:56:26] Yeah, okay, thank you, man.
[01:56:27] Regarding assignment, you missed, I guess. Yes.
[01:56:31] Mm-hmm. Yeah, yeah. So, I think last week I missed a…
[01:56:35] Giving the assignment, uh…
[01:56:38] So, I will formulate an assignment and share this week.
[01:56:41] Uh, so you want an assignment generally on supervised learning, or you want assignment per topic? What way you want?
[01:56:48] It'll be good, ma'am, uh, um, like, the assignments would be, like, we can correlate all those theory and, uh,
[01:56:55] We can…
[01:56:58] Yeah, whatever you have taught, whatever has been taught in class.
[01:57:01] Accordingly, you miss a student, what we've learned that, okay.
[01:57:07] What we have taught, so that even… even if we didn't, uh… Give attention to the class also, but if you give that.
[01:57:16] assignment, so that we can, uh… Go back and we can.
[01:57:20] Uh, see your video, and accordingly, we can, uh.
[01:57:29] Okay, uh, yeah, so I…
[01:57:30] Cover the assignment, or we can complete the assignment, so… I say a student, my intention is that. What is thekr, uh, idea? I don't know.
[01:57:34] No, no, same, same, same only. So, where we can correlate all those, uh, theory, whatever it's taught, and, uh, we can have, uh…
[01:57:43] I mean, it will connect when we will start assignment, right, it will connect theory with our practical, so that is my intentions, yes.
[01:57:52] Oh. So, I will, uh, give you some simple…
[01:57:53] Yeah, yeah, same, sim, since.
[01:57:56] Uh, assignment problem.
[01:57:59] And, uh, I don't think you want it to be… do you want it to be a graded assignment, or you…
[01:58:05] wanted to be on…
[01:58:11] Okay.
[01:58:12] No, ma'am. So, it should be a practice assignment. If that curriculum will have some graded assignment, then that you can plan. But, uh…
[01:58:15] for whatever we…
[01:58:17] have, uh, discussed in the theories that there should be something that we can practice those things in the weekdays, so that it will refresh.
[01:58:26] Okay.
[01:58:27] It will connect, yep.
[01:58:29] Okay.
[01:58:30] Ma'am, you are considering that we are, uh… Already working somewhere and we should be professional, but.
[01:58:40] practicing okay, so I am. Uh, what you can say that I'm considering my student also, so whatever you give the assignment.
[01:58:48] whatever in the 10th class, 12th class, so. I am thinking I want to do that, because I have not worked.
[01:58:56] At anywhere, at any moment on this AI part.
[01:58:59] Mm-hmm.
[01:59:01] Okay.
[01:59:02] So everything is new for me. So, whatever, uh, you are teaching us.
[01:59:05] So, consider as a student and give the assignment so that I can even.
[01:59:10] I don't… Uh, give that tension in class also, but with your assignment, I can go back and see your video, and I can work on that.
[01:59:21] This is my request.
[01:59:22] Okay. Okay, I can give assignment. I don't have any problem, and of course, I assume
[01:59:28] Uh, all of you. In fact, some of you, as you pointed out, Anoop, or…
[01:59:34] Some others may not have worked on ML and AI, and those who would have worked on AIMN may not have actually studied the theory behind it.
[01:59:44] Many a times, it's very important to understand what are we actually doing.
[01:59:48] So, we might be using some Python code or whatever, we might be using some functions without understanding. We understand and use it makes them much
[01:59:57] are used much better and easier.
[01:59:59] So that's the idea why, uh…
[02:00:02] I always start from the basic, because I'm taking regression, I'm going to take all these…
[02:00:07] prediction models and all that. I would…
[02:00:10] Actually, it would be much easier for me to start with NLP directly, but that I don't do.
[02:00:16] Because if I do that, then many of you won't actually… you'll be lost.
[02:00:20] So, uh, that's the whole idea, that everyone is in sync when we talk about
[02:00:27] higher models, so you have this background, and no matter what kind of problem you…
[02:00:31] have at your workplace now or in future, you are able to
[02:00:35] solve it with this background and foundation.
[02:00:39] So, I will give you assignment, um…
[02:00:42] It will be ungraded.
[02:00:45] Uh, and then, um…
[02:00:49] But definitely, I will not be able to discuss it here. That you'll… it will be a self…
[02:00:55] kind of a self… yeah.
[02:00:56] No problem, ma'am. It will be ungraded or something that doesn't matter, okay?
[02:01:00] Yeah.
[02:01:01] We want that as a student, and I should get some assignments so that I can, even I have not given the attention to the class also, so that I can.
[02:01:02] Yes, we are.
[02:01:12] Yeah, so… yeah.
[02:01:13] Work on that. That is the main idea.
[02:01:14] And even… even we can ask some questions around it.
[02:01:19] Yeah, so I will not be able to…
[02:01:20] Yeah, yeah.
[02:01:22] Discuss much on the assignment, because that will be an extra activity, because if I…
[02:01:28] Every week, if I give you an assignment,
[02:01:31] No problem, ma'am. No problem, ma'am. No problem.
[02:01:32] Yeah, I… yeah, why I'm mentioning this, it's good and, you know, safe to mention it.
[02:01:37] Because, uh, because if I discuss hands-on, and I discuss theory,
[02:01:42] And then, additionally, I discussed the hands-on also, then we don't cover what we intend to cover, and it will be your loss only.
[02:01:51] But if you want it as a self, you know, self, uh…
[02:01:54] kind of a self-exercise or something, that way I can give.
[02:01:57] If that is okay with all of you,
[02:02:00] Without the expectation that it will be discussed in class,
[02:02:04] The solution will be discussed, then I can give you. Is it okay with everyone?
[02:02:10] Yeah, ma'am, that's okay. Means, uh, if you're giving 10 questions also, if initially you are giving backup of what you have covered yesterday.
[02:02:19] So, in that, if I, uh… For that question, if I can, uh… Ask for 5 minutes also, so I think it is okay. So, we are okay with that.
[02:02:29] Means, uh, we will not discuss much on what you are giving that assignment.
[02:02:36] Yeah, so that's okay with you.
[02:02:37] That okay at. Please, we want that, okay, what you have taught us, so accordingly, what… yeah, yeah.
[02:02:43] Yeah, but it should be okay with everyone. So that's the thing.
[02:02:46] Uh, tomorrow, there should be no such expectation.
[02:02:49] that if I have given an assignment, that I'm going to discuss,
[02:02:53] The solution of that, that may be really out of scope.
[02:02:57] Because, actually, we are very hard-pressed on time already.
[02:03:01] So, as you know, there's a lot of coverage to be done.
[02:03:05] So, is it okay with everyone? Please let me know over the chat.
[02:03:09] And if it is okay with everyone that you solve the problem as an exercise for yourself,
[02:03:15] Then, I will give.
[02:03:16] Yes, ma'am. Yep.
[02:03:22] Ma'am, as you are, uh, already you are not giving any exercise, okay, so it is already in the same condition with all.
[02:03:27] If you are giving extra exercise, then I think it will be okay with everyone.
[02:03:33] Because I already… you are not giving, so already we are in the same position.
[02:03:38] So, if you're giving something extra, so we will, uh… Try to cover that, and uh… Uh, so it will be beneficial for all, I think.
[02:03:48] Okay.
[02:03:51] Okay, fine. So, uh…
[02:03:53] I will give, uh… I think, uh, I also give quiz, but that
[02:03:58] who is, again, is ungraded one.
[02:04:01] And that is based on the theory, because, uh…
[02:04:04] Many a time, students want to know what kind of…
[02:04:07] Uh, questions.
[02:04:09] Similarlyn, can we have a quick poll? Are you there?
[02:04:18] Uh, who would be the co-host today?
[02:04:23] Hello?
[02:04:26] So, unfortunately, I cannot see who's the host.
[02:04:30] from future ends.
[02:04:33] Uh, the host…
[02:04:35] Yeah, I'm here, ma'am.
[02:04:36] Okay, Severin, could you please post a quick poll?
[02:04:40] On what we were discussing, that if we have an assignment,
[02:04:43] Uh, without discussing
[02:04:45] Uh, the assignment solution here.
[02:04:48] Is it going to be okay or not okay?
[02:04:51] Uh, could you please quickly post it right now, and…
[02:04:55] So that we can have a…
[02:04:58] um, you know, a good idea about what people feel, and then…
[02:05:02] I can give those assignments.
[02:05:08] Uh, we'll just wait for Simran to quickly make the poll.
[02:05:15] So, in the meantime, any other
[02:05:18] comments or queries or question.
[02:05:22] Anyone has, um…
[02:05:24] So, um, in this course, I'll try to cover or maximize the coverage
[02:05:30] And many topics that are not mentioned, I may still cover for you.
[02:05:35] Uh, depending upon the batch, uh…
[02:05:38] understanding and background and all that.
[02:05:42] However, whatever has been promised will definitely be…
[02:05:47] delivered. Something more only you'll get, not less.
[02:05:51] Um, but I request that you all be, uh…
[02:05:55] attentive, and please try to…
[02:05:59] read whatever has been covered, otherwise then you get lost if you don't
[02:06:05] Mm-hmm.
[02:06:06] One question I have not related to the course. Whatever I am seeing, uh… In your background.
[02:06:08] Sure. Mm-hmm.
[02:06:11] There are so many books, how you manage that?
[02:06:15] I mean… Yeah, there are so many books.
[02:06:20] What I am seeing in your background, uh, so I think you have gone somewhere.
[02:06:25] Some… some extent you have gone through that. We are not getting time means how we should get that time so that I can go through.
[02:06:36] So, actually…
[02:06:37] So much books are so much. This is a general question.
[02:06:40] So, actually…
[02:06:43] Whether you go through many books or you don't go through many books,
[02:06:48] But if you are able to grasp whatever you feel you should grasp, that's sufficient.
[02:06:56] Many a times, you actually
[02:06:57] get lost in books.
[02:06:59] So it's not many times very useful to get lost in books.
[02:07:04] I definitely like to read books. I have a lot of books which I read and which I also don't read. I…
[02:07:12] Visually think that I'll read in future.
[02:07:14] And so I keep books with me.
[02:07:18] Uh, okay.
[02:07:19] So, uh…
[02:07:21] It's like that, um, whenever time permits, we read books.
[02:07:28] Whenever time doesn't permit,
[02:07:29] then we don't. But, it's not necessary that to gain knowledge, you have to only read books. There's a lot of literature available.
[02:07:37] And you should be able to grasp whatever is required, and that's…
[02:07:42] Uh, what is important today?
[02:07:45] Right? So, I think the result of the poll is there. I think most of the people agree that if
[02:07:53] There are not solutions or not discussed, that should… that's going to be fine, so I will…
[02:07:58] I will be giving you some assignments, okay?
[02:08:02] So, Ilya, I'm going to end the poll.
[02:08:05] Okay, then, thank you all.
[02:08:06] Thank you, ma'am.
[02:08:08] And have a great evening today. We'll meet tomorrow. Thank you. Bye-bye.
[02:08:09] Uh, yes.
[02:08:15] 2 back.
[02:08:25] Thank you