# 09 2025-12-27 Logistic Regression course: Module 2 — Machine Learning Algorithms module: Module-2-Machine-Learning-Algorithms date: 2025-12-27 type: transcript video_url: https://personal-learn.armco.dev/files/_Recordings/Module-2-Machine-Learning-Algorithms/09_2025-12-27_Logistic_Regression.mp4 --- [00:40:44] shaped curve is maximized. This is called maximum likelihood, and [00:40:50] When we talk about, uh, fitting, [00:40:53] Uh, the data on this kind of a sigmoid function, then we use the concept of maximum likelihood. [00:40:59] And not of residuals, because here it's not applicable. [00:41:05] Uh, because you can see that the points are nonlinearly separable. They are not linearly separable. [00:41:10] So, uh, now the logistic regression, what it will do is that it will use [00:41:18] Uh, this kind of logistic function to predict whether [00:41:21] The class of the mouse is going to be not obese, or it is going to be obese. [00:41:27] But it is using what? A continuous valued function, which is weight. [00:41:32] Weight can take up any value. It could be 0.1, it could be 0.9, it could be 0.001. [00:41:37] So, in a range of, uh, it's not going to be discrete, it's going to be a continuous-valued function, so… [00:41:44] A logistic regression will be applicable on it, and it will use this [00:41:48] To predict whether, uh… [00:41:51] The class level is obese or not obese. [00:41:53] Say, for example, as I already pointed out to you, that [00:41:57] If these are the set of points, [00:41:59] then these mice are the ones that are having [00:42:01] high amount of heat, as you can see here, the value of the weight is higher this side. [00:42:06] So these values are higher, whereas these are the mice which are… [00:42:10] Actually, not obese. [00:42:13] And, uh, what we try to do is that we try to fit a logistic or a sigmoid function on this data. [00:42:21] So, in order to fit what, uh, it means is… [00:42:24] That, uh… [00:42:26] We want to fit this… this kind of a logistic function, or a S-shaped curve, which will help us to predict [00:42:33] The probability that, given a particular weight, [00:42:38] for a mouse, whether it's going to be obese, based on that weight, or it's not going to be obese. So, logistic regression, [00:42:45] Actually, it's going to give probability of the mouse being obese or not being obese. [00:42:55] So, uh, now, uh… [00:42:58] Uh, the thing is that, uh… [00:43:02] Uh, directly, the probability [00:43:04] cannot directly be used. [00:43:06] for classification. [00:43:08] Uh, and uh… [00:43:10] We can just see here, say for example, uh, okay, before we proceed, say, for example, this is the… [00:43:17] mouse having this much weight. [00:43:19] Then, what is happening? We are getting some… [00:43:24] So, we are… we are getting some probability here. [00:43:27] Right, so here, as you can see here, I told you… [00:43:31] that these are probability values, actually. [00:43:34] These are not classes, they are probability values. [00:43:38] So, for example, if we talk about… so, actually, this is a probability zero, this is a probability 1. [00:43:45] And if we talk about a mouse having this much weight, then the probability of it being [00:43:50] obese is going to be zero. [00:43:52] And not always is going to be 1, because this is falling here. [00:43:56] So, similarly, if we talk about a mouse which is having this much weight, [00:44:02] Then the probability of it being obese is going to be 1. [00:44:06] Here, the probability of it being obese is 0. [00:44:10] this is 0, because this probability is zero. [00:44:13] So, anything here will have a probability 0 of being obese. [00:44:17] Anything lying on this line, [00:44:20] is going to have the probability of [00:44:23] Being obese one. So, a mouse having a weight of this much. [00:44:27] is going to have a high probability of being obese. [00:44:31] Whereas a mouse having this much weight? [00:44:34] is going to touch this here, and the probability of it being obese will be 0.5. [00:44:41] So, what are we having? We are actually having probabilities and not classes directly. [00:44:45] However, we use logistic regression for classification. [00:44:49] So then, um… [00:44:51] What we will, uh, do is that we are going to use a range of values, or a threshold, that will help us to decide [00:44:58] So, a threshold on the probability, this is 0, this is 1. [00:45:02] So, we might use this threshold. [00:45:04] say, 0.5. So, anything having a probability greater than or equal to 0.5, [00:45:10] Maybe classified as obese, and anything having a probability less than equal to [00:45:16] uh… 0.5, maybe, classified as not OV, so this is how we are going to use [00:45:23] Uh, probabilities to make [00:45:25] prediction of the class label. [00:45:30] So then, uh, what we actually do is that we have the probabilities, this is 0, this is 1. [00:45:37] And, uh… [00:45:40] So, in order to… so how do we derive this particular sigmoid function? So, in order to derive this, what we do is that [00:45:46] We take the probability and multiply it by… [00:45:50] the weight. [00:45:52] So, when we have probability, when we have a very less amount of weight, which is this much, [00:45:58] And the probability is less, we get a point here. [00:46:01] Then we have slightly higher weight and the same probability we get a curve here. Then we have a… [00:46:06] Slightly higher weight, and slightly more probability we get this point. Then, slightly higher weight, [00:46:12] Slightly higher probability of the mouse being obese. [00:46:15] We get this point. Like this, we keep on… [00:46:19] actually getting all these points. [00:46:21] And as the weight increases, the probability of the [00:46:25] mouths become obese increases, like that. And what we get is this curve. [00:46:31] Right? Now, uh… [00:46:33] So, now what we could do is that, uh… [00:46:38] Uh, this means that we get… [00:46:39] Ma'am, I have one doubt, sorry. Okay, okay. [00:46:41] Yeah, just let me complete, I'll take your down, just… just wait. [00:46:44] Okay, so now, uh, what we are actually doing is we are calculating the likelihood of observing [00:46:52] A non-Obs mouse. So here, if we actually see here, [00:46:56] So, in a way, what we are doing is, when we are trying to fit [00:47:01] this particular line on this data, we are trying to maximize the likelihood [00:47:07] of these points to lie on this particular sigmoid function. [00:47:11] So, we are actually working on likelihood. We want the points [00:47:17] To be most likely lying on this curve, which we have used to fit on this particular [00:47:22] data. So then, uh… [00:47:25] We actually calculate the likelihood of observing the mouse. [00:47:30] two VOBs or not OBs. So if, for example, a mouse which is… [00:47:35] Having such a small amount of weight, let's say this is zero, and this is some unit 1 and 2, and like that, [00:47:42] Maybe in grams. Then… [00:47:46] If a mouse is having very less weight, then… [00:47:49] The likelihood of it being obese is going to be really low. [00:47:53] And the likelihood of it being non-obese is going to be high. [00:47:59] So, here, we are actually just trying to map the likelihoods. [00:48:04] Uh, while we are talking about the sigmoid function. [00:48:06] Yes, Vinit, you had our doubt. [00:48:14] Uh, did someone have a doubt? Yeah. [00:48:15] Yes. Yeah, yeah, Mimi. Sorry, I was trying to go on mute. [00:48:17] Yeah. Mm-hmm. [00:48:19] Uh, so, ma'am, when… so this basically is, you know, based on the function that I'll be calling it, or obese, right? [00:48:24] Yes. [00:48:25] So what happens to these borderline cases, like the one which is not on the line, the… see, in the OBs. [00:48:30] Yeah. [00:48:32] We have the very first point, right, which is nowhere coming in the sigmoid function, yeah. [00:48:34] This one. Mm-hmm, mm. [00:48:36] So these… these will then… not be called out as obese, is it? Or the probability of this coming under obese then goes away, right? Or. [00:48:43] So, yeah, so actually, this is the likelihood curve, right? [00:48:44] How do we read this data? [00:48:49] This is zero probability, this is 1, so if you try to map it here… [00:48:51] Yeah. Uh-huh. [00:48:54] So this will be the probability of it being obese. [00:48:59] Right? Similarly, [00:49:00] Okay, okay, okay, okay, okay, okay. Got it. [00:49:01] Here, you see this, you read it like this. [00:49:03] see this, you read it like this. So, the probability of it being obese is really low. [00:49:09] Okay. Okay. [00:49:10] However, this point is an outlier. So even when the weight is high, it is not obese, which is actually not possible. [00:49:17] Got it. [00:49:18] Any particular mouse which will be. Uh, having high weight will be obese, so this is nothing but an outlier, but the likelihood of the points get. [00:49:27] Coming on this sigmoid is quite high here. Most of the points are coming like this. [00:49:32] Okay. [00:49:33] Okay. Okay. So the outlier, we don't treat them, right? We just kind of ignore or discard them, is it? [00:49:38] No, they would have got treated. This is just an illustration to show. [00:49:42] Okay. Okay. [00:49:43] The likelihood, how we get the max… how we try to maximize the likelihood of the points coming on the sigmoid. [00:49:50] Got it, ma'am. Thank you. [00:49:52] Any other questions anymore? Yeah. [00:49:53] Yeah, one question from my end, uh, ma'am. Is this used only for binary class classification problems? [00:50:01] So the… traditionally, logistic regression is actually meant for binary classification. However, there are variants, and now. [00:50:09] Currently, you can use it in multi-class problems also. [00:50:14] Okay. [00:50:17] Yeah, but when we take it as a… Uh, as a… I mean, as a traditional logistic regression problem, yes, it is. [00:50:25] Primarily meant for binary classification. [00:50:27] Okay, will we see an example of multi-class, uh, classification as well? [00:50:34] Okay. [00:50:35] Yeah, hopefully you should. [00:50:40] Sure, yeah, thanks. [00:50:41] Okay, Adity? [00:50:43] Uh, anyone, I think I saw two hands raised. Anyone else? [00:50:50] I had the same question, actually, like, will it be only applicable to binary classification, or not? [00:50:56] Yeah. Yes, I already told that. [00:50:57] Yep. [00:51:03] So, I already told you that we keep on calculating the likelihood by multiplying the weight with the probability to get this sigmoid function. [00:51:13] And, uh… and we keep on doing it for all the mice, and what we obtain is this. [00:51:20] No, the thing is that… Uh… Once we get all the likelihoods. [00:51:27] Then we multiply them together to get the final likelihood of the data that we. [00:51:33] can see on the line. This is a mathematical way of. [00:51:36] Looking at it. Now… Another way of fitting this is that. [00:51:43] Instead of having the line here, we could actually fit it here also. [00:51:49] Just like if we talk about logistic regression, uh, simple linear regression. [00:51:52] We have these points. We could actually have a line fitting like this. [00:51:57] We could actually have a line fitting like this. [00:52:00] We could have a line fitting like this. Or like this, like, so and so forth, and we try to. [00:52:05] Fit that particular line, which minimizes the error. So here also, we can actually shift this, uh, likelihood curve. [00:52:15] So that we have a curve like this. Or we could actually shift it so that we have another one here. [00:52:22] We have another one here. We have another one here like that. [00:52:26] And so and so forth, many possibilities will be there. [00:52:30] Say, for example. We could actually have something like this as well. [00:52:35] Like that. And finally, we will try to. Use that particular sigmoid curve, that will maximize the likelihood, that is. [00:52:45] What is maximizing the likelihood? It will have the maximum number of points. [00:52:50] That will pass through it. You can think of it like this. [00:52:54] So, even though we can have a curve like this also, however, only one point here is passing through it. [00:53:00] And maybe 3 points here are passing. The rest of them are not passing. [00:53:05] So, we want to maximize the likelihood. And we select that particular curve that maximizes the likelihood, or in other words, we have most number of points. [00:53:13] That will lie on this sigmoid function that we will choose finally. [00:53:20] And that will be the best fit logistic function. [00:53:24] So, uh… So this means that we have a weight, which is a continuous valued variable. [00:53:29] And we are going to… Use it to predict the probability, the probability in turn is going to. [00:53:36] Help us predict the class level based on some threshold that we would have decided for. [00:53:48] Oh. Ma'am, a question on the previous slide. [00:53:49] Example, in this case, that was 0.5. So, uh… Uh, now… [00:53:54] So, this point… [00:53:55] Yes, I request you to please raise your hand, that may be better. Yeah, please go ahead. [00:54:01] Uh, so this 0.5, how would we come up with that number? [00:54:06] No, this is just a range of probability, right? [00:54:07] That… No, I understand that, but, like, you… here we mentioned 0.5, right? That, okay, above 0.5, we are saying. [00:54:09] So, what I'm… hmm… Oh, hmm. [00:54:16] is obese, and below that we are saying is not obese, but how would we say that 0.5 is the correct? [00:54:17] Mm-hmm. Right. So, yeah, yeah, so that is actually a decision boundary we chose… we choose. [00:54:21] number. [00:54:27] So, for example, here I just chose 0.5. It could actually be 0.75 also. [00:54:34] It depends on what is our application. Say, for example, in our experiments. [00:54:40] Let's say that if we choose this. It means the implication is that this much weighted mice will be not obese. [00:54:48] So, for example, in our experiments, we want to treat, let's say, whatever unit this is. [00:54:53] Let's say we want this range. We believe that this much of weight. [00:54:57] makes a mouse to be lean or not obese, then we decide this much. [00:55:03] Okay, so this means that very high weights only will be categorized as obese, whereas. [00:55:10] Uh, you know, relatively less hire will come in, not obese. [00:55:14] So it purely depends on the application, on what we want to do. [00:55:17] So it's… It's written, like, we'll do multiple iterations, then we'll see from where we are getting the best values. [00:55:24] Yeah, you can do that, or it might also be based on some domain knowledge. For example, if there's an experiment. [00:55:30] And it is felt that any particular. Range of, let's say if this range was, say, 0 to. [00:55:36] Psalm 50, uh, grams, let's say. So, if we want that any mice, a mouse having a weight in this particular range. [00:55:39] Mm-hmm. [00:55:45] Should be classified as not obese, then we'll see in this weight what is the probability coming. [00:55:51] So the probability, if it is coming is 0.75, then our decision boundary goes at 0.75. [00:55:56] However, let's say that. Let's say this is some 40 gram. [00:56:01] And we want any particular mouse having weight 0 to 40 to be. [00:56:07] Not obese, then we choose 40 as the decision boundary. [00:56:09] And let's say this is the probability where it is coming. [00:56:13] So, let's say it is 0.3. Then we'll keep the decision boundary 0.3. [00:56:18] So, based on the implication, and probably on the domain. [00:56:23] knowledge, we can decide. Where to put the decision boundary, or where to put the threshold to decide what is the class label. [00:56:32] Usually, domain knowledge is very much useful for deciding that. [00:56:36] Okay? Mm-hmm. [00:56:37] Okay. And another question on this. So, you mentioned that we take the curve where maximum points are covered. [00:56:43] Yes. Mm-hmm. [00:56:44] Right? So… When we are identifying that particular curve, are we just looking at, like, let's say, is obese category, where, like, most of them are. [00:56:53] Actually present, and then we can say that wherever the number is actually below our threshold, we'll consider that it's not obese, or we are considering. [00:57:01] both of the sites. Because we can plot points in both of these things, right? [00:57:09] Uh, I couldn't get you… [00:57:10] I can just calculate is obese, and I can do 1 minus AOB, that should be my is not obese, correct? [00:57:13] Hmm. Mm-hmm. [00:57:15] So, if I'm actually maximizing that whether my SOBs are coming on the curve. [00:57:20] Can I take that particular curve, or I have to maximize both of them? [00:57:24] No, it's maximizing both. So, basically, what are we maximizing? [00:57:28] We are generally maximizing the likelihood of all the points to lie on this sigmoid. [00:57:29] Mm-hmm. [00:57:34] Or the maximum number of points to lie on the sigmoid. [00:57:38] So when I'm saying maximum number of points, it includes those which are here, and it also includes those which are here. [00:57:45] And why can't we actually maximize one of them if the probability of the other event occurring is 1 minus of what we have. [00:57:53] See, 1 minus is the probability. Actually, you can. [00:57:54] On his obese. [00:57:56] Uh, see. What we want to do is, we want to. [00:58:01] Maximacy, for example, more simpler thing is that. Let's say I have these set of points, right? [00:58:03] Mm-hmm. [00:58:08] I want to fit a line. Which best represents these points? [00:58:09] Mm-hmm. [00:58:13] So, I'll try to… Have a line which is nearest to maximum number of points, isn't it? Then only it will best represent. [00:58:22] I cannot have a line like this. [00:58:23] But if… No, ma'am, but if it's not obese also, I convert into is obese, and then I generate a curve, then will it become a straight line? [00:58:29] No, don't. No, no, these are data points. How can you con- these are not probabilities. [00:58:36] These are data points, see, this is a mouse having. [00:58:39] Mouse A having a weight of. Whatever this is, let's say some. [00:58:40] Mm-hmm. Correct. Okay. [00:58:44] Value X1. This is about B having weight of X2. How can you convert them into. [00:58:50] Uh, you know, how can you do a 1 minus on it? This is not probabilities. These are data points. [00:58:53] Okay. Okay, understood. [00:58:56] We want to maximum the, like, we want to maximize the likelihood of the data points to lie. [00:59:04] On this sigmoid, so all data points we will consider, no matter what class it is. [00:59:05] Okay. Understood. And last question is on this shape of the curve, right? You have mentioned that you can make it more vertical, or it can be more horizontally shifted, right? So, in case of linear regression, we were having the slope, and which will make. [00:59:18] Mm-hmm. [00:59:23] The, uh, degree of slope actually more or less, right? So what is the function here which will make it. [00:59:29] Uh, more vertical or more horizontal. Or shifted more horizontally. [00:59:31] See. Hmm. So, in linear, we have a line, which is straight. [00:59:38] Yeah. Right. [00:59:39] This is nonlinear. Right? This is non-linear, and the function I already showed you, it is… This is the function. [00:59:49] Yes. [00:59:50] This is the sigmoid function. So, if you give Z as the input, which will be weight here, what you will get is this function. [00:59:58] If you plot this, this particular equation, what you will get is this function. [01:00:03] So, here we are having a non-linear function. We cannot have. [01:00:04] Right. [01:00:11] So that… Go ahead, Chris, I'm sorry. [01:00:12] Something like a straight line here. And what I'm saying is, hmm… So, what I'm saying is that. [01:00:16] Even a long linear function, say, for example, this sigmoid function. [01:00:22] We can roll like this, we can roll like this, we can draw… we want to. [01:00:27] Oh, sorry, we want to actually… Uh, you know, we can fit the function in any way. For example, here also, in the case of a linear. [01:00:35] Equation, we could have equation here, we could have it here, we could have it here, we could have it here. [01:00:40] However, we don't choose the one that is very far away from maximum number of points. [01:00:46] We choose one which is nearest to maximum number of points. [01:00:50] So, similarly, here also, sigmoid can be drawn in a lot of ways so that it passes from one or more points. [01:00:56] However, we will choose that particular sigmoid function. That maximizes the. [01:01:03] Uh, likelihood of… The points maximum number of points passing through it. [01:01:10] That sigmoid we will choose. [01:01:12] Yes, ma'am, that was my question, that how we are going to draw it. If you see it, the linear regression, if m is equal to 1. [01:01:18] then your Y equal to X, you will get a 45 degree line. [01:01:24] Mm-hmm. [01:01:25] crossing the origin, correct? And then, that if M is going to vary, that line is going to change the shape. [01:01:30] In this case, how are we changing it? I am not able to relate that. [01:01:31] Mm-hmm. [01:01:35] That you gave me one function, and then it is generating one sigmoid. [01:01:37] Mm-hmm. [01:01:40] Now, the shape of that sigmoid is going to change based on certain criteria, right? I'm not able to relate what is that criteria, which is changing the curve, because I plotted that curve once. [01:01:48] And now, you're saying that shape should change. What is the variable which is actually going to change the. [01:01:49] Hmm… [01:01:54] uh, shape. [01:01:55] So, yeah. So basically, if you try to vary here, so we are talking about Z as a… linear function, I mean, it's a function of the values. So if you change those coefficients, it will change. [01:02:04] Mm-hmm. [01:02:08] Z is actually equal to… Beta not plus beta 1X. [01:02:13] Plus beta 2, so. So, yeah, yeah, yeah. [01:02:14] So it's basically log of what we did at linear… linear regression. [01:02:18] So, yeah, so basically, when you change these coefficients, the shape will change. [01:02:24] Okay. Understood. Got it. Thank you. [01:02:25] Okay. Yeah. Yeah. [01:02:30] So now… I think we were here. [01:02:36] So now, uh, the question is that, uh. If we talk about weight and size of the mouse and like this, and. [01:02:44] If we use our equation in the simple form, like Y is equal to mx plus C. [01:02:50] Some question can come to mind that can we not use this equation and then. [01:02:55] just represent the data using such kind of an equation. [01:02:59] So, such kind of equation cannot be used because. [01:03:04] Say, for example, if this is the… You know, this is a function. [01:03:09] Uh, Y is equal to MX plus C. This is your M, uh, sorry. [01:03:15] This is your M, this is your C. And this is your function. [01:03:20] So, in this particular case, what could happen is that. [01:03:23] Okay, if we have a mouse of this much weight. [01:03:26] Uh, then the size could be something. However, uh… If we try to project it in the backward direction, what we could have here is. [01:03:37] Uh, so… The problem with this kind of, uh, you know, linear function fitting a data which is non-linear could be. [01:03:46] That here it is fine that this much weight means this much size. [01:03:50] But as for this… If we try to extend it backward, we could have. [01:03:55] Uh, mouse having negative weight. And also having some size. [01:04:00] Now, it's not possible to have mice which are negatively weighted. [01:04:04] And yet, having sizes. First of all, weight cannot be negative, neither's. [01:04:09] There could be a mouse having a negative weight, and yet having a size. [01:04:13] And also mouse having zero weight, having a size. [01:04:17] So this kind of representation is actually not going to work. [01:04:21] So now, what do we do? How do we actually… So I already told you that this sigmoid. [01:04:25] Definitely representing this data in a better function… better way, but. [01:04:30] No, the value… First of all, the value that the sigmoid gives is probability, and it is constrained in the range of 0 to 1. [01:04:41] That is first thing. And then secondly, we cannot use any, uh, linear operator, we cannot use any linear. [01:04:47] Uh, modeling functions, uh, and all on this kind of… Uh, sigmoid function. So, somehow we want to actually map. [01:04:56] Uh, these, uh… This avoid to a linear form or a linear model. [01:05:04] In such a way that this range of probability 0 to 1, which is confined and constrained to this range. [01:05:11] Also gets converted, and somehow, if we can map it to the range. [01:05:16] Uh, in which any linear function could have. So, any linear function could have values from minus infinity to plus infinity. [01:05:23] And therefore, they are not constrained over the range of this one. [01:05:27] So, if we could somehow map this logistic sigmoid function. [01:05:32] Our logistic function is synonymous. In such a way that instead of 0 to 1 range, it could take up values from minus infinity to plus infinity. [01:05:41] And, uh, it could somehow be, you know, represented linear. It would be really nice for us. [01:05:50] But logistic regression definitely is confined. And it's also not linearly representative table, it's represented by this. [01:05:59] Function, which is a nonlinear function. Which is this one, which is our sigmoid function equation. [01:06:05] So now, what do we do? So, uh… So, what are we going to do is that some, uh… As I said that we want to map. [01:06:15] Uh, it in such a way that it lies from minus infinity to. [01:06:19] Plus infinity, so what are we going to do? We are actually going to. [01:06:23] represent this with the help of. Probability. And, uh… Already, we are having the probability here. [01:06:31] But we are going to use the log odds of. [01:06:34] Probability. Remember, I… Told you log odds. [01:06:39] What is the T equal to probability of the success of an event upon. [01:06:44] The failure of an event. Where success of the event could mean the mouse being obese, then it is going to be the log odds of obesity. [01:06:56] If P means success. And for us, success means mouse being obese. [01:07:05] Then, we are talking about log odds of the mouse being obese. [01:07:10] So if we do that, then probably we might be able to. [01:07:15] map this… Kind of 0 to 1 range to minus infinity to plus infinity if. [01:07:22] We take the log of. Probability of obesity. [01:07:28] If you do that, so how it's going to work? [01:07:32] So, let's now try to transform this, uh… This range, this is our y-axis, which is the probability. [01:07:39] Which is, uh, lying in the range. Sorry, my writing's gone bad here. [01:07:45] is lying in the range 0 to 1. By using, uh, log odds of obesity. How are we going to do it? [01:07:53] So, log odds of obesity, actually. will mean log of P upon 1 minus P. [01:08:01] Right. And, uh, P here means the probability. [01:08:08] Of mouse. Being obese. This is what P means. [01:08:15] And he's going to take up value in the range of. [01:08:19] 0, 2, 1. Right. [01:08:25] Okay. So, let's say that… If the probability of the mouse being obese is 0.5, which I've already told you. [01:08:33] This was some particular. Point here, and the probability of the mouse. [01:08:38] Being obese is 0.5 year. Then, we are going to use the log odd of probability. What is log odd? Log of. [01:08:47] P upon 1 minus P. P here is 0.5, so we have log of 0.5. [01:08:53] Upon 1 minus 0.5. Which is equal to log of… 1.5 of 1.5, which is 1. [01:09:01] Log of 1. Okay, remember, it is natural lock, though I'm writing it as LOG, it should be LN. [01:09:09] Anywhere log of 1. Will be what? It will be 0. [01:09:14] So this means that if we are talking about probability of 0, then we are actually mapping it. [01:09:19] to these set of values, right? So… then… Log of, uh, when probability is 0.5. [01:09:28] Then we have log of 1, which is also… which is going to be 0. [01:09:33] So, now, point 5 is actually going to be centered on 0. [01:09:38] Now, let's try to substitute some other values. Let's say we have probability equal to 0.731. [01:09:45] Which is somewhere, let's say, here. Then, we take log of. [01:09:50] 0.731. So, it is actually log off. [01:09:54] P upon 1 minus P. Which is 1 minus 0.731. [01:10:00] Which is log of 0.73. 1 divided by 0. [01:10:05] If we take it as 0.73, it will be, like, 0.27. [01:10:10] And that 0.27 log of it, it's natural log of it will come out to be 1.99, something like that. [01:10:17] Alright, so this is what we are calculating. And this will actually map to. [01:10:23] Uh, so it is actually Ellen… Uh, so this is what we have, 0.73 upon point. [01:10:33] Oh… 1 minus 0.73, that comes out to be. [01:10:38] This much, which is coming out to be 1. [01:10:42] So, uh, the natural log of 2.717 is 1. [01:10:48] And therefore, we are actually mapping the probability of 0.731. [01:10:53] With these set of values. Which is starting 1. [01:10:58] Similarly, we can go on doing this. If we take a probability of 0.95, again, you have to put in the log odds. [01:11:07] This, so it will be log off. 195 divided by 1 minus 1.95. [01:11:13] Which will be log off. 195 divided by 0.5. [01:11:20] And when we try to do this, then it will come out to be. [01:11:24] Approximately, uh… 2. So, it is actually going to come out to be 1.99 something. [01:11:32] So, then we are going to map this probability to the values 2, like that. [01:11:37] And we can go on doing this. And we can keep on getting these values like that. [01:11:44] Now, suppose we, uh, plug the probability of 1. [01:11:48] And what we have here… Uh, is, uh, log of, again. [01:11:54] P upon 1 minus P. Which will be log off. [01:11:58] One upon 1. Minus 1, so which is log of 1 upon 0. [01:12:05] Now, although, uh, log of. Uh… usually we don't have divide by 0, but in log, we could have it, so log of. [01:12:13] Vara upon zero is actually dog of 1. Minus log of 0. [01:12:20] But log of 1 is 0. And log of 0 is infinity. [01:12:24] So it comes out to be… minus infinity. So, this means that if we are talking about. [01:12:31] Oh, sorry, this comes out to be minus infinity, this comes out to be plus infinity. [01:12:37] So then, if we are talking about the probability 1, we are… then we are trying to map the probability 1. [01:12:44] To the range of plus infinity. So, we can think of it. [01:12:49] That the points which are actually lying on any value. [01:12:53] Somewhere here, would be mapped to… Point set plus infinity. [01:12:58] Similarly, if we… so this means that when we are talking about probability 0.5 to plus 1. [01:13:05] Then, we are actually stretching out this probability of 0 to 1 on the sigmoid to a linear function. [01:13:12] Which is lying from a value of 0 to plus infinity just by using. [01:13:18] Log of. Odds of success. This particular function, this is also called the. [01:13:23] logit function. So, by using this function, we are actually mapping this non-linear. [01:13:29] Uh, curve, which is constrained over 0 to 1 values. [01:13:33] If you talk about 0.5 to plus 1, it is mapped to 0 to plus infinity, and points lying at. [01:13:39] plus 1 probability would… can be thought of lying at plus infinity. [01:13:44] Similarly, if we put the value of probability 0, we'll have log of. [01:13:48] zero upon 1 minus 0. Uh, which is log of… 0 minus log of 1. [01:13:55] And so, it will come out to be minus infinity, yeah. [01:13:58] So the points that are actually lying at probability 0 can be thought of at lying. [01:14:04] At minus infinity. So then, we have actually converted this range 0 to plus 1, where 0.5 to plus 1 gets mapped to 0 to plus infinity. [01:14:14] And 0.520. gets mapped to 0 to minus infinity using the logit function. [01:14:21] So, uh, in logistic regression, we use the logit function. [01:14:25] And we just think of the original samples, which were having the probability 1 of being obese as lying at plus infinity now. [01:14:36] And similarly, those being non-obese. Lying at minus infinity. [01:14:40] So, like this, actually. We have now stressed out and mapped this constraint function to a linear function, as you can see here. [01:14:51] And it's having the range also, which is unconstrained, lying from. [01:14:55] Uh, minus infinity to plus infinity. Okay. So, uh, this I have already explained to you, that if we try to now put in values from 0.5. [01:15:06] In 0, then what we'll get here? Uh, it's like this. So now, in this way, logistic regression uses the logit function. [01:15:16] To, uh, the present. the range, uh, constraint in probability 0 to 1. [01:15:24] To a range which is minus infinity to plus infinity. [01:15:27] And it maps it to a… So, to kind of a linear plane now. [01:15:33] So, uh… So, we are now finally ending up with the log odds of being… of obesity. [01:15:43] And the log of odds of being obesity can be 0, it can be plus infinity, it can be minus infinity, and so on and so forth. [01:15:52] Any questions, anyone? Here. [01:15:57] Up to now. [01:16:07] Okay, there are no further questions. So then, even if we are having this kind of an. [01:16:16] Non-linear squiggly line S-shaped sigmoid function. We can actually, uh, associate it with logistic regression, and then use it to predict. [01:16:28] Uh, values, uh, and, uh. The coefficients actually. [01:16:34] Uh, which are there as, uh… As I already explained to you, they can be actually represented at. [01:16:41] As the log odds graph. As you can see here. So, uh… This is a plot explanation. I tried to, uh… Uh, take up something about the mathematical representation, and also I wanted to explain to you how we map. [01:16:57] Uh, the sigmoid. 2 or linear scale. [01:17:05] Uh, from minus infinity to plus infinity. Okay, Deepak, what question do you have? [01:17:10] Yeah, ma'am, uh, one question, so, like, right now, we understood how to map property 01 to minus entry to infinity. [01:17:16] Mm-hmm. [01:17:17] What is the use case? Like, what we are trying to solve, uh… Uh, is there any… any constraint in terms of. [01:17:23] probability that we can't have it unless we have the mapping, like. [01:17:28] Yes, yes, yes. So, because probability can take up only values between 0 to 1, right? [01:17:36] So, it is constrained over its range of values. However, if we, you know, if we talk about any kind of, uh. [01:17:43] uh… linear… Uh, you know, if we talk about some kind of linear function. [01:17:51] So, we always prefer linear unconstrained functions. Here, the constraint is the value is lying in 0 to 1. [01:18:00] So, uh, we don't want it to be constrained to 0 to 1. [01:18:04] And we somehow want to represent it in a linear fashion in such a way that it is not constrained. [01:18:10] Because probability make the models constraint, and therefore, uh. [01:18:16] Uh, therefore, actually, it… Its usability becomes limited, because we cannot work. [01:18:24] Uh, only on the range of 0 to 1. [01:18:29] So, we wanted to naturally map it to values that are linear. [01:18:33] And why we want to map, because there may be many linear functions. [01:18:37] That may be used. Say, for example, uh… Here, we are only talking about probabilities, but if we use a logit function, then we could map it to. [01:18:50] Uh, such values where, uh, you know, these are more naturally occurring values, right? [01:18:56] We are not having it constrained, and therefore what happens is that there are many linear operators. [01:19:02] That could be used on it. Uh, they cannot be used on probability because its range is 0 to 1. [01:19:09] And the linear functions, they always work on the range of values going from minus infinity to plus infinity. [01:19:16] So we cannot use such kind of operators. We cannot use, uh… Linear models. We cannot do anything linearly on this kind of. [01:19:26] constraint values. Therefore, we need to map it. [01:19:30] To… it's advisable to map it to this range of values. [01:19:35] So, how do we deduce that we should go for this mapping? Are the probability is sort of becoming constrained? If you could probably give an example. [01:19:45] So, for example, uh, see, uh… in future, in deep learning models and all. [01:19:54] Uh, logistic, uh, you know, this logit function is used, and there are. [01:20:00] Basically, what happens is that, suppose. Uh… so, actually, I don't want to go into the math of it. [01:20:08] However, there are many operators. Such as the plus operator, minus operator, and many other things. [01:20:15] That work or, uh, there are, uh, you know, some linear functions that can be used on this kind of range of values. [01:20:23] However, we cannot use them on this range of values. Suppose I want to. [01:20:27] You know, have a function like, uh… Uh, you know, Y is equal to function of X. [01:20:34] Now, a normal function cannot be used on this kind of values. [01:20:37] Why? Because this is just 0 or 1. It is having a range of 0 to 1. Any function, normal function. [01:20:45] We'll… we'll have some set of values that will range now… have a natural range lying from minus infinity to plus infinity. [01:20:53] And we want, uh… You know, this sigmoid function. [01:20:58] To be mapped to a linear function. Because they are most… more simpler to handle. [01:21:05] And they can, uh… they are more intuitive, they are. [01:21:08] Uh, less difficult to understand, but if we want to model it in a linear function. [01:21:14] Then, because linear functions always work on this range. [01:21:19] They cannot range… work on ranges like 0 to 1, they don't work on it. [01:21:23] So then, to map it to a linear function, we have to put it in this. [01:21:28] Range only. And that's why this logic function is used, actually. [01:21:33] In logistic regression. [01:21:35] Just follow-up question. So, we… with minus infinity to infinity, we also have some. [01:21:40] Negative numbers, right? So… How do we understand that we want to, sort of. [01:21:41] Mm-hmm. [01:21:46] Don't want that range, negative range. Right, maybe 0 to… is there any scenario where we only want to keep posted numbers? [01:21:56] And I think you also have given the example about size and weight. [01:22:00] Which doesn't make sense in negatives. So, how do we understand that part? [01:22:01] Mm-hmm. [01:22:06] No, see, those are two different things. So, for the first thing that I tried to tell you about that negative weights was. [01:22:13] That, suppose I have these set of points. Okay, let me go back on that. [01:22:22] Yeah. So, the reason of trying to fit a non-linear curve is that, suppose I had a linear curve like this. [01:22:31] Size is equal to some value of. uh, you know, see and some value of M. [01:22:38] Suppose I represent… This times wait. [01:22:42] size as this kind of a function. So if I have this kind of a function. [01:22:48] Uh, that, uh, represents size as a function of weight. [01:22:53] And this is just a simple linear function, that what would happen is that whatever B that line. [01:22:58] We could actually project it backwards, and we could get negative. [01:23:02] Uh… the weights also. Whereas, it's not possible to have negative weights, if we have this equation or any such similar equation. [01:23:12] So, if a linear equation is used to represent a data which is actually not linearly separable. [01:23:18] Then what would happen is that we could get. [01:23:21] Functions which incorrectly map this data, or. incorrectly represent this data. For example, if this was a. [01:23:30] Uh, equation that we would have obtained where C is .86 and M is 0.7. [01:23:36] Then we could actually have this line, and when we have this line, then… We have weights which are negative. [01:23:42] Yet having size. So, uh, what I wanted to point out here is that. [01:23:47] Sometimes the data is not linearly separable. And we need nonlinear, uh… the presentation, such as the sigmoid here. [01:23:56] Because we cannot have a line like this in this case. [01:24:00] For functions which are really linearly separable, we could have. [01:24:05] A linear function, like… You know, like this. [01:24:09] These set of points, it is fine to have a linear. [01:24:14] Uh, equation representing, like, Y is equal to mx plus C, and it is going to be okay. [01:24:19] But if I use some equation like this to represent these points. [01:24:23] then it won't be correct. Because these points cannot be represented on a line. Most of the points will. [01:24:30] Not fall on that line. That was the thing I was trying to point out. [01:24:35] When I said negative thing. Now, what do you want to ask in this one? Here? [01:24:41] Here, when I'm talking about this. Because here we are trying to. [01:24:49] map these range of values, which is constrained to 0 to 1. [01:24:55] So that we can, you know, use some linear function on it. To use a linear function on it, this range needs to be mapped from 0 to 1. [01:25:03] 2 minus infinity to plus infinity. So here we are using… How do we map it? One way to map it is using the logit function, which is log of p upon 1 minus p. [01:25:14] Which is the basis of logistic regression. So this is what was proposed when logistic regression was proposed. [01:25:22] That the logit function would be used on the probability. [01:25:27] To map it from 0 to 1 range to minus infinity to plus infinity range. [01:25:31] So that some linear. Functions could be mappable on this particular… they could be used in this model. [01:25:41] So here, the negative values are also useful, and the positive values are also useful. Why? [01:25:47] Because, uh, positive values from 0 to plus infinity are actually mapping to 0.5 to 1 probability. [01:25:55] And 0 to minus infinity are mapping to 0.5 to 0. [01:25:59] Because probability will range from 0 to 1, so this… this will map to 0. [01:26:05] This will map to minus infinity. This will map to plus infinity, so this entire range is going to be useful here. [01:26:13] Earlier, what I was saying was that if we incorrectly represent the data using. [01:26:18] A linear equation, then we could. Get something like a mouse having a negative weight and yet having some size. [01:26:28] So, ma'am, uh, the… how do we interpret the data point here? Maybe it could be negative or positive, because I think that's a log of. [01:26:35] The odds that we have. So, how do we read the data points on this graph? [01:26:37] Mm-hmm. This one you are seeing? [01:26:42] Uh, yeah, minus infinity or infinity. [01:26:45] Yeah, so the points that were having a probability 1 of being obese. [01:26:50] Are actually… you can thought of them as being la- as lying on plus infinity. [01:26:56] This axis of plus infinity. Those particular mice that were not obese, or they were having zero probability of being obese. [01:27:06] Can be thought of lying here. At minus infinity. [01:27:10] This is what… this is the way you infer. [01:27:12] Or you map. [01:27:15] Okay, sure, thank you. [01:27:17] Yeah. Okay, any other questions, anyone? [01:27:25] See, uh, you need not necessarily remember the maths, but I always feel that it's good to augment. [01:27:32] Anything that we teach using some math behind it. [01:27:36] So, in logistic regression, basically, we are using… making use of probabilities. [01:27:42] Now, when we are making use of probabilities, then… The value range becomes constrained to 0 to 1, so we are actually mapping it from. [01:27:52] We want to map it to Larsa Infinity 2 minus infinity. Now, how to do that? [01:27:57] To do that, it was thought. Uh, the developers of logistic regression thought to use the logit function, which is nothing but. [01:28:06] Log of 1 up P upon 1 minus P, which is log of odds. [01:28:12] Of success of a particular event. This is, uh… they experimented at find out, found out that if we use probability, we substitute probability in it. [01:28:22] Then one gets mapped to plus infinity minus. 0 gets mapped to minus infinity, so we get this entire range. [01:28:29] And so, that is what was proposed. Okay. [01:28:35] Any other questions, anyone? [01:28:43] So, I've already told you about these. [01:28:53] So, I've already covered. [01:28:58] So this is the final representation, where we have this sigmoid. [01:29:02] And using the log odds. Uh, of being obese, it is getting mapped to this. [01:29:08] Particular, um. Uh, fashion from minus infinitude plus infinity. [01:29:15] So this, I had kept some extra slime. Okay. So, uh, as I already told you. [01:29:21] And you had asked also, but logistic regression. Though it's classically, it's said to be representing binary classes. [01:29:32] However, the… when we talk about. Some popular forms of regression. [01:29:39] Uh, it can be binary logistic regression or multinomial logistic regression. [01:29:43] So, in binary logistic regression, we talk about. Uh, you know, one dependent variable only. [01:29:52] And this dependent variable would take up binary values, like 0, 1. [01:29:57] Spam, non-spam, and all that. However, when we talk about multinomial. [01:30:02] Togistic question, we are actually talking about a variable that can take up multiple, uh, values. [01:30:10] You know, it could be 3 values, 4 values. For example, here. [01:30:13] Uh, the… classes, proficiency level of a speaker. It could be advanced, intermediate, novice, or whatever. [01:30:21] So, multinomial logistic regression can be used for. Examples, which are non-binary in nature, so… Multi-class problems also can be solved using multinomial logistic regression. [01:30:36] So, I already, uh, explained to you binary logistic regression. [01:30:42] And, uh, we get the sigmoid function when we try to substitute these values. [01:30:48] And this is a logistic function of the sigmoid function that we already study. [01:30:53] So, I think, uh… Uh, some of you asked that, can we get multiple, uh, can we use. [01:31:01] Logistic regression to. Uh, classify multi-class, uh, problems. [01:31:08] So, I told you that there are variants of logistic regression that can be used to classify. [01:31:14] Uh, multi-class problem. So here it is a multinomial logistic regression. [01:31:19] And in this case, we have, uh, actually. Uh, this is the function that is used. [01:31:25] So, I don't want to go into too much math of it, but however, we are actually doing a softmax of Z. [01:31:31] And, uh, yeah, and it is represented as E to the power zi. [01:31:37] Divided by E to the bar. ZJ, that j is equal to 1 2K. So, actually. [01:31:44] We are talking about, uh… Values of classes lying from 1 to J, and we are actually taking. [01:31:51] Uh, yeah. If you talk about one particular class, so actually, let's say we are having k number of classes. [01:31:58] Case the total classes. So, if you are talking about, let's say, we are having. [01:32:04] C1, C2, C3 till CK. Then, uh, what we do is we want to find out the softmax. [01:32:12] Of the class, let's say CI. Then, it will be… Sea of. [01:32:20] CI divided by. Summation of all the classes, that is E.T. [01:32:27] ZC1 plus ETPAR ZC2. So and so forth until ZEET. [01:32:33] Sit like this. So, this is the function. Which we use in multinomial logistic regression, and we are able to obtain. [01:32:42] Uh, the… Uh, the probability of. [01:32:47] each of these multiple classes. So this is how logistic regression is also used in a multi-class problem. [01:32:55] Not just in a binary class problem. So, this is a softmax function. I will cover it in more details when we talk about some other applications. [01:33:04] Right now, I've just told you how, uh… How it is used, the softmax, uh, the multinomial classification. [01:33:16] So, uh… I think there's a question that why can't we use any other classification model? [01:33:23] Uh, instead of using multinomial logistic regression. So, uh, see, we are talking about different forms of classification. [01:33:34] So, we talked about decision tree classification, we talked about random forest. [01:33:39] We talked about so many different forms of classification, logistic regression is also one way. [01:33:47] And it is also very popular, so it's important that you understand what is logistic regression. [01:33:52] So, there is no, uh, compulsion, or there's no special thing that why we should use logistic regression. [01:33:58] However, if our data is non-linearly separable, there can be many classifiers. Logistic regression is one of them. [01:34:07] So, we are studying logistic regression as one alternate method. [01:34:12] To classify the given data. Okay. However, classification can be done. [01:34:19] In many ways, this is just one bill. So, these are just some examples of how we do the calculations and all, how we estimate the probability, and then use the log odds of probability. [01:34:32] Like this, uh… So, I'll just skip this, because you'll be doing the code tomorrow. [01:34:39] And in that, uh, we are going to… or use multinomial and binary classification. [01:34:46] binary logistic regression and. Multinomial logistic regression, both. [01:34:52] Any questions, anyone? [01:35:08] So I hope logistic regression is somewhat clear. Uh… see. [01:35:13] The most important thing is you should know where it is used. It is used to. [01:35:18] Solve binary and multiclass problems, both. If you wish to understand the math behind it, this is what. [01:35:26] Uh, is, uh, what I've been explaining. So far. [01:35:33] Okay. [01:35:40] If there are no further questions, then we can break for today. [01:35:44] Um, any questions? Anyone? [01:35:54] I'm just doing a stop share. Okay, there are no further questions, then you can break for today, we can break for today, and we'll meet tomorrow. [01:36:04] for the hands-on session. Okay, thank you all. [01:36:11] Thank you. [01:36:12] Thank you, ma'am. Have a great Saturday. [01:36:13] Thank you. Have a great evening. Bye-bye. Yeah, thanks. And this half an hour is a gift for… from my side to all of you. [01:36:23] Thanks a lot.Thank youThank you