# 05 2025-12-13 Linear and Logistic Regression course: Module 2 — Machine Learning Algorithms module: Module-2-Machine-Learning-Algorithms date: 2025-12-13 type: transcript video_url: https://personal-learn.armco.dev/files/_Recordings/Module-2-Machine-Learning-Algorithms/05_2025-12-13_Linear_and_Logistic_Regression.mp4 --- [00:14:43] Mayukh Ghosh: Yeah, everyone. Am I audible? [00:14:49] Aditya Banda: Yes, you are. [00:14:55] Mayukh Ghosh: Good. Alright, so… I think we have met before, if I'm not wrong. [00:15:03] Deepak Katara: Yes, you know. [00:15:04] Meet Mevada: Yes. [00:15:05] Mayukh Ghosh: Yes, it's my guess, right? I think [00:15:08] Mayukh Ghosh: sometime last month, or maybe a little before that, I don't remember exactly when. [00:15:13] Mayukh Ghosh: But… Yeah, I… some of the names also seem to be familiar. [00:15:23] Mayukh Ghosh: Right, so… So, as you know, this is more of a machine learning [00:15:31] Mayukh Ghosh: So we… I think the last time we met, it was after you did, or were doing the EDA sessions, right? Yes. [00:15:39] Mayukh Ghosh: Now I understand that you have explored machine learning to a… to a considerable extent, obviously not explored completely, that's… it's hard to say that ever [00:15:50] Mayukh Ghosh: So I get that. Now, this is obviously… so we are, again, given what… if you remember last time, so we had… we have two hours, and we want to talk about certain concepts which are important. [00:16:04] Mayukh Ghosh: Like we did last time, and maybe connect a couple of things to business contexts and how [00:16:11] Mayukh Ghosh: things are useful there. Some examples around that, okay? [00:16:15] Mayukh Ghosh: So that's the plan, given the short time we have, okay? [00:16:19] Mayukh Ghosh: Now, I understand that you have gone through regression classification models, clustering, And, [00:16:28] Mayukh Ghosh: Yeah, anything else so far in the ML part? [00:16:39] gunjan bhaiya: A couple of, more concepts was thought about, like, how to find a decision tree, right? The algorithm… [00:16:46] Mayukh Ghosh: including classification, but yeah, I get decision tree, random, forest. [00:16:49] gunjan bhaiya: Yeah, so that was in that factor, yeah, that was getting covered in that video. [00:16:54] Mayukh Ghosh: Considering practical examples. [00:16:56] gunjan bhaiya: One or two. [00:16:57] Mayukh Ghosh: Right, right, okay, alright. Fine. So… [00:17:02] Mayukh Ghosh: I'll try to touch upon a little bit of everything, obviously, yeah, in… so, a little bit of everything in a sense that, obviously, you know the concepts, you have gone through it hands-on, you might have seen a bit of tactical, as you said, right? So, all that is fine, but we'll try to get into some key concepts and how [00:17:17] Mayukh Ghosh: they kind of make or break models, or how they are pretty important, okay? So I'll… [00:17:22] Mayukh Ghosh: Possibly look at it from that angle. [00:17:24] Mayukh Ghosh: And we'll start with some concepts of linear regression, go to logistic, go to other classifications, for that matter. I mean, I think that should be, yeah, enough to… yeah, more than enough to spend the time and more, possibly, okay? [00:17:39] Mayukh Ghosh: So I'll pick up concepts. Again, obviously, the intention is also not to teach you again any of this. I'll pick up certain concepts, and then we'll see where the importance lies and how things go. I'll show you some outputs as well, as we go along. [00:17:52] Mayukh Ghosh: from… on… on Python, okay? And we'll see how certain outputs can be interpreted, or when things go wrong, how to kind of solve them, right? That's very important. [00:18:03] Mayukh Ghosh: So, that's the idea. Now… if I talk of, initially, the regression, [00:18:11] Mayukh Ghosh: Let me… let me open an Excel sheet and possibly jot it down here. I think that's the way we did last time as well. Let me share it. [00:18:26] Mayukh Ghosh: Yeah, I believe you can see. Okay, so there are a lot of sheets here, which some of it I will show you later for some other reason, but right now, we can build a few things here, okay? [00:18:37] Mayukh Ghosh: There's a reason I've opened this sheet, because I'll use it sometime later for some of the other sheets you already see that are there, right? [00:18:45] Mayukh Ghosh: So, one of the main concepts of linear regression, right, if I go back to it and… [00:18:51] Mayukh Ghosh: By the way, any deep learning introduction that has happened, I might… I guess that that might have also happened? Neural networks? [00:18:59] gunjan bhaiya: No. [00:18:59] Mayukh Ghosh: No. [00:19:00] Chandrasekhar Sahu: No, no. [00:19:02] Mayukh Ghosh: So that will come then. Alright. So, see, the… [00:19:06] Mayukh Ghosh: Linear regression, obviously, is the first, [00:19:10] Mayukh Ghosh: model, so to say, or first algorithm, so to say, anyone learns in machine learning, right? And I'm sure it has happened the same way. [00:19:17] Mayukh Ghosh: And… but… but the irony is that in most business problems, right, when you… when you're solving business problems, linear regression is rarely used. Okay, let the truth be said, it's rarely used. [00:19:28] Mayukh Ghosh: very few of business problems will be solved by linear regression. I mean, depends on domain, but if you… unless you go to manufacturing, supply chain. [00:19:36] Mayukh Ghosh: It's even less in other domains, okay? In retail, very less, I would say. [00:19:40] Mayukh Ghosh: Banking, yes, one or two problem statements are there in general, but other than that, yeah, it's pretty much… [00:19:47] Mayukh Ghosh: Dominated… the space of model… modeling algorithms is dominated by [00:19:53] Mayukh Ghosh: classification models, okay? But it's very important to learn or understand linear regression pretty well. Not pretty well, quite well. More than that, in fact. [00:20:02] Mayukh Ghosh: Because everything that follows, which I know, I am sure you know by now, everything that follows starts from there, okay? So you're building a… [00:20:10] Mayukh Ghosh: multi-story department along 10-floor departments, and linear regression is your ground floor, okay? So if the ground floor is weak, or if the ground floor is not built properly, then no matter how many floors you build on that, okay, eventually it will all fall down, right? Or eventually it will be, yeah, weak, and it will be demolished. [00:20:30] Mayukh Ghosh: So that's the, [00:20:34] Mayukh Ghosh: importance it has, or that's the place it has, okay? Now, I'm sure the linear regression you did using the OLS method, right? Am I… [00:20:43] Mayukh Ghosh: My assumption correct? Using the ordinary least square and best fit line, those, those, those sort of areas? [00:20:50] Mayukh Ghosh: Can anyone confirm? [00:20:53] Deepan Kanagaraj: Yes, sir. [00:20:53] Chandrasekhar Sahu: Apparently. [00:20:55] Mayukh Ghosh: Sorry? [00:20:57] Chandrasekhar Sahu: No, I guess, my… [00:20:59] Mayukh Ghosh: So how, how did you do… I mean, so you, you did linear regression, right? [00:21:04] Deepan Kanagaraj: Thanks, Albert. [00:21:05] Mayukh Ghosh: Are you directly moved to logistic? [00:21:10] Deepak Katara: Cover linear regression. [00:21:11] Chandrasekhar Sahu: None of those… [00:21:14] Mayukh Ghosh: So you started with decision tree? What was the first ML model you did? [00:21:18] Ravindra Singh: I think directly we have started Decision Tree and Hiranda Forest. [00:21:22] Mayukh Ghosh: That's it. Directly started Decision Tree, Random Forest, then XGBoost, and all those, right? [00:21:28] Ravindra Singh: That one we haven't. We have only covered two, I think. [00:21:32] Mayukh Ghosh: Okay. [00:21:36] Mayukh Ghosh: So, in machine learning, if I just try to understand the summary a bit, because the information I had is slightly different, but it's important that I hear from you. So, decision tree random forest is… I'm fairly understanding that it's a unanimous thing that all of you did, okay? Any other models you did? Clustering? [00:21:55] Chandrasekhar Sahu: Nope. [00:21:58] Mayukh Ghosh: So, only these two, or anything else you're gonna add on to that? [00:22:03] Chandrasekhar Sahu: My understanding, it is only these two only, my… [00:22:06] Ankit Sood: Only these two. [00:22:07] Mayukh Ghosh: Yeah. [00:22:08] Deepan Kanagaraj: Is that what everyone is saying, or sort of… [00:22:11] Mayukh Ghosh: Yep. [00:22:12] Aditya Banda: Yeah, that's right, only these two. [00:22:14] Deepak Katara: Does that… [00:22:14] Ankit Sood: London Forest. [00:22:16] Aditya Banda: headache. [00:22:17] Mayukh Ghosh: Okay. [00:22:18] Deepan Kanagaraj: No, again, we also covered something like inter… intra-distance and then intracluster distance and all the other things, mostly from theory perspective. [00:22:27] Deepan Kanagaraj: I remember it was still covered. [00:22:30] Ankit Sood: Yeah, Euclidean distance and all the other things. [00:22:33] Mayukh Ghosh: Ukrainian distance me, that would be your clusters and segmentation, right? [00:22:36] Meet Mevada: K-nearest neighbors and stuff like that. [00:22:38] Mayukh Ghosh: beginning and all those. [00:22:42] Ankit Sood: The alien was discussed in the last lecture, but purely from the theoretical perspective. [00:22:46] Mayukh Ghosh: But, okay, but that comes much later. So, linear and logistic were not explicitly done, that's what you were saying, right? [00:22:53] Ankit Sood: Yes. [00:22:54] Mayukh Ghosh: Okay. [00:22:55] Mayukh Ghosh: Okay, mmm… [00:22:58] Mayukh Ghosh: Alright, so I'll not go into that in that detail then, okay, but I understand it's pretty important. So, let me still talk a little bit about it, okay, because I'm sure that that will be useful. I mean, you're gonna connect it back to the decision tree random forest you did, okay? That would be important from that perspective. [00:23:17] Mayukh Ghosh: So you are aware of these terms, right, in decision tree context also. Say, for example, AUC, ROC, Precision Recall F1, these things you know, right? [00:23:29] Nirav Mehta: Yes. [00:23:29] V L Narasimha Rao Kosuri: We do a bus, or… [00:23:33] V L Narasimha Rao Kosuri: I mean, what do they have to teach for? [00:23:36] V L Narasimha Rao Kosuri: You don't have any agenda, just you are checking with us what is covered under… so that you will cover remaining thing. [00:23:43] Mayukh Ghosh: No, not really. I'm not going to cover remaining thing, but that's not possible in a 2-hour thing. [00:23:48] Mayukh Ghosh: I'm checking why… what is… what has been covered to understand the context where I am getting into, and how much you will get of that, okay? Okay. So, yeah, I'm not going to cover anything as a topic, okay? That's not the idea. The idea is to give a wholesome sense of ML in that sense, and if certain things are not covered, I'll obviously, yeah, I will intrude into it. So that's why I was wondering… [00:24:08] Mayukh Ghosh: what has been done, and getting a precise sort of sense out of it, okay? That's the reason of asking. [00:24:15] Mayukh Ghosh: Okay, so… [00:24:17] Mayukh Ghosh: let's talk about those things which you have done, because that is very important in terms of precision recall and the other model yardsticks, for that matter, right? So I'm not going to linear equation, because then that becomes a complicated thing right now, in a short span of time. So, let me go to, then. [00:24:42] Mayukh Ghosh: I'll show you something else, wait, give me a bit of a minute. [00:24:57] Mayukh Ghosh: Okay, here itself I have, I can show you here. [00:25:01] Mayukh Ghosh: So… [00:25:06] Mayukh Ghosh: So these concepts you are aware of, I'm sure? Precision, recall, and then if I go a little bit down the line, the accuracy part, false positives, false negatives, these things you are aware of? [00:25:17] Anup Pankaj: Yes, yes. [00:25:22] Mayukh Ghosh: Right. So. [00:25:25] Anup Pankaj: Hello. [00:25:25] Mayukh Ghosh: This thing, you probably… let's look at. [00:25:28] Anup Pankaj: a business problem. [00:25:28] Mayukh Ghosh: then, okay, given that you know how to understand a classification problem statement, okay? Let's look at a business problem. [00:25:35] Mayukh Ghosh: Say, for example, so this is actually… actually real data. This is done for a manufacturing client. I was doing some consulting for them a year back. So this is the real data on which these… these are real numbers, which went to… so it's not a made-up thing on Kaggle or anything. [00:25:49] Mayukh Ghosh: So here, we try to understand… say, for example, if you look at, say, a decision tree, or a random forest, or any kind of tree models for that matter, right? [00:26:00] Mayukh Ghosh: Now, can anyone quickly tell me? Because I want to make it a more conversation rather than say, I'll just keep on telling a monologue, and then we are nowhere from there, right? [00:26:08] Mayukh Ghosh: Save… [00:26:09] Mayukh Ghosh: how do you understand, or how do you figure out which amount of precision and recall is going to be more important for your model, or which you want to increase more? So what's the understanding behind that? [00:26:25] Mayukh Ghosh: Anyone? [00:26:25] Chandrasekhar Sahu: Please repeat the questions. [00:26:27] Mayukh Ghosh: So I'm saying that precision recall, you know what they are, but I'm saying for a particular business problem, or for a particular problem you're trying to solve. [00:26:34] Mayukh Ghosh: How will you decide which one to emphasize, and why? Because you can't… you can't increase both of them together, right, in general, right? So, how do you take the priority, or how do you make a hierarchy out of them? That, okay, for this model, precision is more important. For that model, recall is more important. So, how do I reach that decision? [00:26:55] Deepan Kanagaraj: And… [00:26:56] Deepan Kanagaraj: Let's say there is a use case where you wanted to identify whether the patient has a cancer or not, right? This is very sensitive to us, right? I don't know which one we will pick it up. So this is one use case. The other use case is either a particular customer will buy that product or not. Probably the [00:27:14] Deepan Kanagaraj: So for us, it's more important or sensitive is to, predict the cancer case rightly than whether the particular customer buy a product or not. So, based on that, probably we'll choose, give weightage to the different one, is what I'm thinking. [00:27:29] Mayukh Ghosh: So you're thinking… so, yeah, your thought process is in the right direction, but you have two problem statements, or you are in a comparative mode, you are assessing them, right? [00:27:39] Mayukh Ghosh: Right, okay, fine. So, anyone else who wants to add to this? Or tell something else? [00:27:44] Chandrasekhar Sahu: If you want to increase the recall, then your projective instance should be more… [00:27:51] Chandrasekhar Sahu: As compared to this false negative. [00:27:54] Chandrasekhar Sahu: So, you need to have more positive incidents, means… [00:27:59] Chandrasekhar Sahu: you can have this from this configure matrix, right? So, the point of this instance would be more, so you can recall it would be more. [00:28:07] Mayukh Ghosh: Okay, okay. [00:28:09] Mayukh Ghosh: Anyone else? [00:28:12] Sonam Manwal: Maybe this recall thing we can consider when… [00:28:17] Sonam Manwal: This false negative is costly affair. [00:28:20] Mayukh Ghosh: Like, hmm. [00:28:21] Sonam Manwal: In that case, we should be considering recall. [00:28:25] Mayukh Ghosh: Yeah, that's very close to what the perfect, or the simplest perfect answer should be, or could be. [00:28:32] Mayukh Ghosh: Yeah, no, the reason I was showing this and I asked this question is also to get a sense of the business connection to precision recall. Anyway, let me go back here, okay? And let's start a little bit from [00:28:45] Mayukh Ghosh: Yes. Okay. Because I would like to build up a little bit, given that you haven't done linear logistics explicitly, I think we have still enough time to make some meaningful stuff there, okay? [00:28:55] Mayukh Ghosh: Because that is important. Then, given that you have done decision tree and random forest, you can easily connect back afterwards, okay? [00:29:04] Mayukh Ghosh: Okay, let me add this quickly. [00:29:06] Mayukh Ghosh: Everyone is aware of this equation. [00:29:10] Mayukh Ghosh: Aware, in a sense, you've seen it in some contexts. [00:29:12] Anurag Krishnam: Yes, yes, yes. [00:29:14] Surya Bobbala: It's… [00:29:14] Mayukh Ghosh: So what does this equation tell me? Anyone, quickly? [00:29:17] Chandrasekhar Sahu: this deadline. [00:29:19] Mayukh Ghosh: Straight line, that's how we sort of identified initially from our school and college and all that, right? [00:29:25] Mayukh Ghosh: This is nothing but a linear regression equation, okay? Even though we learn it from a straight-line perspective when we are young, okay? So, say, for example. [00:29:34] Mayukh Ghosh: Let's say, why is my, [00:29:43] Mayukh Ghosh: monthly spend on food. I don't know whether I discussed this the last day when I took your session. If anyone you can remember, tell me, but I guess it was EDI probably… perhaps hadn't gone into this. [00:29:56] Mayukh Ghosh: And then there's income, okay? [00:29:59] Mayukh Ghosh: Now, what happens is, this E part, usually, initially, we don't have, right? Mathematically speaking, usually we don't have. We have Y equal to X plus C. That's what we know as a straight line equation from our school and college days, right? So what happens here is C is a constant, so C is the value which [00:30:15] Mayukh Ghosh: Y takes when MX is 0, right? Can I say that? Without the E first up? [00:30:20] Mayukh Ghosh: Which means that when I have no income in a particular month, say I have taken a break, not working, or whatever it can be, right? Every month I will not have income in my life, right? So when I don't have an income, does it mean that I'll stop eating? No. So I still have to spend some money on food. [00:30:34] Mayukh Ghosh: And C, your constant, or the intercept, we call it, is a value [00:30:39] Mayukh Ghosh: or the money, minimum subsistence amount, I need to spend on food, irrespective of whether I earn anything or not. I take it from… I borrow it, I… I mean, I take it from my savings, whatever it is, right? [00:30:49] Mayukh Ghosh: Now, that's fine. That is C. I'll keep C separate. It's more about the relation between Y and MX we are interested, right? So, M is the value that connects to X and Y. So, if I say M is 0.2, for example. [00:31:00] Mayukh Ghosh: That means what? 0.2 equal to M, for example. That means, I can say easily, that 20%, roughly, of my… 20% itself, of my, income goes into monthly spend on food? Can I say that? [00:31:17] Meet Mevada: Yep. [00:31:18] Nirav Mehta: Yep. [00:31:19] Mayukh Ghosh: straightforward equation. If I silence the C for a while, right? [00:31:23] Mayukh Ghosh: So… [00:31:24] Mayukh Ghosh: Then say, for example, here I have 100 people, we have, what, how many, 50, 60, 70, 80, whatever. [00:31:30] Mayukh Ghosh: a substantial amount of people in this call right now, right? Say, for example, simplicity, I'm saying that 10 of us, okay, not even 50, 60, 10 of us have the same income. Maybe, what, 50,000 rupees a month, or something, some number, right? [00:31:44] Mayukh Ghosh: By this equation, that means that all of us will be spending $10,000 on food, Right? [00:31:51] Mayukh Ghosh: 20% of that? [00:31:53] Mayukh Ghosh: But do you think that is even possible? Forget 10 people. Do you think even that is possible for 2 people, or 3 people? It's not going to happen, right? That exactly 20% of the income, even if the income is same, will be spent for everyone on food, right? So this straight-line notion that we have in a strictly one-to-one correspondence from a mathematical sense. [00:32:13] Mayukh Ghosh: That doesn't apply in econometrics, okay? This is the economics… regression is the econometric equation, right? It doesn't apply on… in econometrics. And because it doesn't apply in econometrics, we have the E. What you see in line row 16, we have the E part, the error part. And what E does? E balances my MX plus C and Y. [00:32:32] Mayukh Ghosh: Because we know that MX plus C and Y will never be equal, in real time, with real data, okay? For real people, numbers attached to real business, right? So E basically is the error part, which takes care of two things. One is… [00:32:46] Mayukh Ghosh: Missing variables. [00:32:48] Mayukh Ghosh: Okay? [00:32:49] Mayukh Ghosh: I'll come to that. 2 is… psychological factors. [00:32:55] Mayukh Ghosh: Now, what are they? [00:32:56] Mayukh Ghosh: Missing variables here, see, I'm trying to understand my monthly spend on food using income, okay? [00:33:03] Mayukh Ghosh: no… Income is not the only… [00:33:07] Mayukh Ghosh: yardstick or parameter to understand monthly spend on food, right? There can be plenty of other things. How many people I have at home? [00:33:12] Mayukh Ghosh: How many loans I'm paying back, or the credit card bills I'm paying in a month. [00:33:17] Mayukh Ghosh: What are the other expenses I might have? Different cultural and social structure in where I am living, basically. So, plenty of other things, right? So those variables are not there. In this equation, I have only one parameter to judge on, or to assess my monthly spend in food. [00:33:34] Mayukh Ghosh: So, missing variables effectively means that there could have been more Xs, X2, X3, X4, X5, but I don't have them right now. So, that is being, again, kind of captured by the error part, because I don't have information, so it goes into error, or residual, we call it. [00:33:47] Mayukh Ghosh: And psychological factors is not even quantifiable. Even if I have variables here, that's still fine, but psychological factors, I can't qualify. So let's say I want to have biryani every Saturday. Someone wants to have some other food, Chinese every Saturday, for example. Now, they don't cost the same thing. [00:34:03] Mayukh Ghosh: Doesn't necessarily mean that neither of us can't afford the other one, but we have a choice, right, effectively. So that is psychological factors, that is human differences, effectively. [00:34:14] Mayukh Ghosh: So that also is covered in E, and that balances by Y and NX plus C, because these two will never be fully satiated, or fully kind of saturated, ever. [00:34:25] Mayukh Ghosh: Hence, we have the… And this is my linear regression equation. [00:34:29] Mayukh Ghosh: When I have more Xs, X1, X2, X3, X4, what we want to understand is that all the excess [00:34:34] Mayukh Ghosh: all the X… as a group, I… Able to explain [00:34:43] Mayukh Ghosh: Sorry, explain why or not, right? [00:34:47] Mayukh Ghosh: And if they are able to explain, to what extent, effectively. So I want it to be high. [00:34:53] Mayukh Ghosh: If I want my excess to be influencing Y. I want out all the excess, I want to be high. So I say that I have all these variables, but I want to understand, using all of them, that my monthly spend on food is predictable. I understand what it should be, right? [00:35:08] Mayukh Ghosh: So there is a measure here, which, again, I'm sure some of you might be knowing. [00:35:13] Mayukh Ghosh: It's called R-square. [00:35:15] Mayukh Ghosh: Anyone have used it or explored R-Square? [00:35:21] Jagannath Das: Yes. [00:35:22] Mayukh Ghosh: Surely you've heard of it, right? [00:35:24] Jagannath Das: Yes. [00:35:25] Mayukh Ghosh: Yes? [00:35:26] Mayukh Ghosh: Can anyone tell me what is RSquare in this context? [00:35:29] Aditya Banda: Is it the square of errors? [00:35:31] Mayukh Ghosh: No. Just because there is R-square, it doesn't work that way. [00:35:36] Mayukh Ghosh: Yeah. [00:35:37] Jagannath Das: R-square is… R-square is used to… [00:35:40] Jagannath Das: Measure the kind of, measure the how. [00:35:45] Jagannath Das: How good is the linear model is? [00:35:47] Jagannath Das: Hist. [00:35:48] Mayukh Ghosh: It gives you a sense of… [00:35:51] Mayukh Ghosh: The line I wrote, right? That… [00:35:55] Mayukh Ghosh: how to what extent the XS are able to explain why? Or, as you said correctly, right? Yes, yes. [00:36:02] Mayukh Ghosh: So, this value should be as high as possible, right? And R-square will obviously lie in a percentage term between 0 to 100%, or 0 to 1, right, in absolute terms. [00:36:11] Mayukh Ghosh: So the higher it is, or the closer it is to 100%, or 1, [00:36:15] Mayukh Ghosh: the better it is, because then I can say my X's are being able to influence Y to a large extent. [00:36:21] Mayukh Ghosh: And then, for the business, that is very important. I'll tell you a very simple sort of number. Say in 2023 and 2024, I have my variable, which is, say, revenue or sales, I can take, doesn't matter, sales, okay? And maybe 100 million or something for a particular organization, and this is 150 million, okay? From one year, the change has happened. [00:36:39] Mayukh Ghosh: Now… [00:36:40] Mayukh Ghosh: In this case, I have, say, some X variables, or some independent variables, we call them, which you have seen in decision tree also, I'm sure. Say, this is called cost of advertising, maybe something called cost of marketing, usual variables. [00:36:54] Mayukh Ghosh: research and development, or something like that, right? CX1, X2, X3, these three are. So, I might have only these three, in a very simple sense, okay? So, I want to understand that these three, as a group, because these costs have been incurred, the company has already spent money on all this, advertising, marketing, they have spent money, right? [00:37:12] Mayukh Ghosh: In this 2024 cycle and all that. So now they want to understand what is the return on investment, okay? I've spent a lot of money on marketing, a lot of money on R&D, a lot of money to pay people, hire people, all of that, right? Now, at a high-level P&L, I want to understand that what has it caught in terms of [00:37:28] Mayukh Ghosh: return in terms of sales, right? So, say I know that, okay, 50% has increased, which is very good news. I also want to understand that because, given that I have data on these three things in this timeline. [00:37:41] Mayukh Ghosh: As a group, as a whole, how these three are being able to explain the change of 100 to 150. [00:37:48] Mayukh Ghosh: If that… if, say, for example, these three [00:37:53] Mayukh Ghosh: Are able… sorry… are able to… [00:37:57] Mayukh Ghosh: explain it till 1.40. Say… it's saying that, okay, yeah, I get that, the ROI is combined and all of that. [00:38:05] Mayukh Ghosh: I expect, or I get the figure should be 140, but not 150, which is the reality, okay? So prediction is 140, reality or actual is 150. [00:38:14] Mayukh Ghosh: So the gap of 10 is representing my error, or E in that equation. [00:38:21] Mayukh Ghosh: And in this case, my R-square would be… [00:38:24] Mayukh Ghosh: How much? Can anyone tell me? [00:38:30] Aditya Banda: 80%? [00:38:31] Mayukh Ghosh: 80%, 40 out of 50, right? Absolutely. [00:38:35] Mayukh Ghosh: The change is 50, 40, I have been able to explain, and I haven't been, right? So, essentially, in a figurative sense, R-square and error are kind of adding up to 1, can we say that? Right? [00:38:47] Mayukh Ghosh: So, the greater the value of R squared, the lower should be measured, right? [00:38:53] Mayukh Ghosh: And if you go to your tree models, which you have done, so I want you to relate as well. [00:38:58] Mayukh Ghosh: In your decision tree, you have used, I'm sure you have various gining coefficients. You have used an entropy, for example, to build the tree, then prune the tree using hyperparameters, like max depth, this, that, and so many things, right, to understand overfitting, ensure that there is no overfitting, right? [00:39:13] Mayukh Ghosh: So… [00:39:14] Mayukh Ghosh: From that perspective, these… all these models have the same policy, or the same goal, right? And all the models, decision-generated random forest boosting, logistic, linear, KNN, NAVase, SVM, neural network, all of that, right? [00:39:29] Mayukh Ghosh: All of them have only one goal. [00:39:31] Mayukh Ghosh: The commonality is that reduce errors, okay? [00:39:35] Mayukh Ghosh: to reduce errors, and obviously, the process, you will maximize your accuracy, maximize your precision, recall, all of that that you have. So, here, my aim is to reduce this as much as possible. In the process, I will increase this. [00:39:49] Mayukh Ghosh: So one of the main problems when people start off building models, and at an amateur level, we all do that, right? Maybe a decade back, I was in the same sort of boat. [00:39:58] Mayukh Ghosh: Our aim generally, the mindset generally is to increase this as much as possible. [00:40:04] Mayukh Ghosh: Precision, I have to get 90%. Recall, I have to get 100%, whatever, 95%. R-square, 90%. Accuracy, some other number. So, we strive for that. [00:40:15] Mayukh Ghosh: What we should strive for, actually, is to reduce this. [00:40:18] Mayukh Ghosh: And then in the process, whatever I get there, fantastic. But the approach should be opposite. Now, one of the questions that might be coming to your mind is that, okay, fine, what's the problem? I mean, if I do tackle this, this should be low. I mean, whichever way you tackle it, it's absolutely fine, right? [00:40:34] Mayukh Ghosh: But then, It's not exactly the same. [00:40:38] Mayukh Ghosh: So… [00:40:40] Mayukh Ghosh: Why it's not exactly the same is that there can be a scenario where your accuracy, precision, recall, R-square, any of this measure you take, they are all essentially the same, telling you the health or wealth of the model. [00:40:51] Mayukh Ghosh: Any of them you take. [00:40:53] Mayukh Ghosh: they can be increasing, they can go up and up and up, but it doesn't always imply that your errors are down for good, okay? It's something I'll come back to and explain as well, but [00:41:08] Mayukh Ghosh: Anyone resonates with that? [00:41:10] Mayukh Ghosh: Or have you thought of this before? That how the error-to-accuracy relation… our square is something we are doing now, but you have done accuracy in your three models, right? They are very similar things. So, has the resonated ever that how do I tackle it in terms of when I'm building a model? [00:41:24] Mayukh Ghosh: How do I ensure that the model is better, model is improving? Is it about stretching accuracy, or is it about minimizing error, or haven't you thought of that yet, at this point? [00:41:44] Jagannath Das: Minimizing the… [00:41:46] Deepan Kanagaraj: Sony.net. [00:41:48] Mayukh Ghosh: Minimizing the error, okay? [00:41:51] Mayukh Ghosh: Okay, anyone else has an… [00:41:52] Deepan Kanagaraj: Yeah, I magnify. [00:41:54] Jagannath Das: Is it not both? [00:41:57] Mayukh Ghosh: Sorry? [00:41:59] Deepan Kanagaraj: Is it not both? [00:42:01] Mayukh Ghosh: Both, but I'm saying, when you are trying to do it, you have to take one route, right? [00:42:08] Mayukh Ghosh: At a time. [00:42:11] Deepan Kanagaraj: Okay, got it. [00:42:12] Mayukh Ghosh: So then, how do I… because one of the main challenges of building model, right? I mean, I see so many people over the years I've seen that who have done… who have got a degree somewhere, some postgraduate degree, something. [00:42:24] Mayukh Ghosh: And then, yeah, these days, it's all data science, machine learning, not these days, for a long time, in fact. [00:42:30] Mayukh Ghosh: The problem is that when they come for interviews, or when they come for, effectively, say, a project, where they have to really do the work, right? [00:42:39] Mayukh Ghosh: The translation from this knowledge of these things, because these things you know, what is the accuracy position, and you correctly said when you asked in the beginning itself. [00:42:47] Mayukh Ghosh: You know all these things, and I'm not saying that you can't, I'm saying generally, people know this, right? And they understand also how to put them in priority and all that stuff. But when I throw you up really bad data, for example, a really, really bad data, in a sense of massive data, a lot of missing values, outliers. [00:43:03] Mayukh Ghosh: Issues in terms of relationships among variables, and then building multiple models, choosing the best model out of it, put it all in production, so the entire flow of… entire spectrum of things, right? [00:43:13] Mayukh Ghosh: Then, generally what happens, 70-80% just freeze. [00:43:16] Mayukh Ghosh: Okay. [00:43:18] Mayukh Ghosh: Why they just freeze? And I have also been there. It's not that I have been sort of, yeah, blessed with something where I've never frozen, right, in that sense. But we all freeze in the beginning. The reason is that our flow… the flow is not correct in your mind. [00:43:32] Mayukh Ghosh: We learned the concepts, we know the concepts, we connect the concepts as well. [00:43:35] Mayukh Ghosh: When reality kind of hits, and we have to do stuff, right, we don't know what to tackle when and in what order. [00:43:43] Mayukh Ghosh: And the classic example is this one. Should I go for improving accuracy and accordingly try to rectify things? Or should I go and minimize errors? And as Deepan, you mentioned, right, the main thing we all of us think [00:43:55] Mayukh Ghosh: Is that… does it really matter? I mean, I will go ahead and increase our square 80 to 90%, error will go down a little bit, right? [00:44:04] Mayukh Ghosh: But again, at the end of the day, C, it will probably. [00:44:07] Mayukh Ghosh: But now that you know the concepts of overfitting and underfitting, and you know train and test data, right? [00:44:13] Mayukh Ghosh: it's not that simple, right? That I, for a one-shot model, for a single-shot model, I just increase my accuracy and automatically mirror it goes down. [00:44:21] Mayukh Ghosh: And that will be sustaining, or, I mean, that will be sustainable or generalized across my entire data, or across my entire spectrum of work I'm doing. [00:44:29] Mayukh Ghosh: Is there a guarantee of that happening? No, right? [00:44:39] Mayukh Ghosh: Are you getting the point I'm trying to make here? [00:44:43] Deepak Katara: Yeah, just one question I have. So, I think all these parameters or metrics, like decision recall. [00:44:49] Mayukh Ghosh: Anything, yeah. [00:44:50] Deepak Katara: So, if we have that on an optimal level, shouldn't the error come down? Because we are sort of trying to fine-tune model in terms of… [00:45:00] Mayukh Ghosh: Yes, yes, it should, it should, but what is the guarantee that it should be same for the new cycle of data you get, or even the test data, forget new cycle? [00:45:09] Mayukh Ghosh: Does that… is that being ensured when you are stretching your R-square on accuracy? [00:45:16] Deepak Katara: I think in last, practical session, we tried to have a different split, and plot all these, metrics, and see which is the optimal, split. So, would be one way to go ahead with it, I'm not sure. But in real life, there would be a few more, ways as well. [00:45:36] Deepak Katara: But what we have done is, we sort of tried to plot all the metrics, and then in the second approach, we tried to change our split percentage of training and testing data to see how it works. [00:45:51] Mayukh Ghosh: So you tweaked the training test data set in terms of percentage to understand that how the variability is there, right? [00:45:57] Deepak Katara: Yes, yeah, I think that's what we have done. [00:46:01] Mayukh Ghosh: That's one… yeah, yeah, that's fine, but see, again, is there an end to that process? [00:46:08] Deepak Katara: No. [00:46:09] Mayukh Ghosh: You can take any amount, any number of, or any possible number of, combinations of train test, right? [00:46:18] Deepak Katara: Yeah, it doesn't end there. [00:46:19] Mayukh Ghosh: It doesn't end there, right? Say, there are 50 people here, we all take different train tests with different random states, and then at the end of the session, we'll all fight with each other, right? My model is better. [00:46:28] Mayukh Ghosh: Right. [00:46:31] Mayukh Ghosh: So… that doesn't solve things. That's the problem. I mean, it can give you variability. [00:46:37] Mayukh Ghosh: But does it give you… Outcome that is… [00:46:42] Mayukh Ghosh: Acceptable, or outcome that is universal. [00:46:44] Mayukh Ghosh: The answer is no. So, we have to ensure something where… or a process, we have to try where the randomness from a train test plate, for that matter, or any kind of results from… that stems out of that is not there, right? Or is minimized to an extent. [00:47:01] Mayukh Ghosh: So, in any case, that's just a part of what I was talking anyway. So, our idea is to always minimize error. I'm sure some of you have heard of the term gradient descent, right? [00:47:16] Jagannath Das: Yes. [00:47:17] Mayukh Ghosh: Yes. So, gradient… is it… what is gradient? It's the universal… [00:47:22] Mayukh Ghosh: I won't say model, but it's a universal rule or postulation for any kind of MLDL model, right? And gradient reason means that you want to find the… where to find the scenario or the sweet spot where the error's a minimum. [00:47:35] Mayukh Ghosh: Because remember, when you are trying, I'm sure you have tried this in your test data, trend data thing, in your decision trees. [00:47:41] Mayukh Ghosh: You are, at the end of the day, how are you assessing overfitting? [00:47:44] Mayukh Ghosh: By checking the RMSC or MAPE or all these things between your train and test, right? [00:47:50] Jagannath Das: Yes. [00:47:51] Mayukh Ghosh: Right? [00:47:52] Mayukh Ghosh: So… you are, in a way, you are comparing your errors, can I say that? [00:47:58] Mayukh Ghosh: So you are still looking at… looking at it from a perspective where the error should be close from one data to another data? [00:48:06] Deepak Katara: Nope. [00:48:07] Deepak Katara: Yeah, thanks. [00:48:09] Chandrasekhar Sahu: Mayuk, we have not done this kind of, things, So, maybe… [00:48:14] Mayukh Ghosh: You haven't done? Okay. [00:48:15] Chandrasekhar Sahu: It's very, very new to us, but me, it's very new to me, okay, and still we are learning. Maybe some are… [00:48:23] Chandrasekhar Sahu: Into this professional field. [00:48:25] Mayukh Ghosh: Yeah, so they might be a different. [00:48:27] Chandrasekhar Sahu: No, that's good. But whatever you are expecting may not be its, going… it's not. [00:48:34] Mayukh Ghosh: No, no, that's fine, that's fine. No problem. I'll keep it simpler, because I understand, when I heard Decision Tree Random Forest, so I expected that certain things would be known by now, but might be still in process, so then that's okay. [00:48:47] Mayukh Ghosh: I'll come back to this, and more ground level. We got into a different discussion from something else, but we'll come back here, okay? I think that that's better for… that should be the floor level for everyone, which helps, I guess. [00:48:59] Mayukh Ghosh: Alright, so this is R-square. I hope this section, what I just built up here, all these things, up to this, it was clear to everyone? [00:49:08] Chandrasekhar Sahu: Yeah, that's clear. [00:49:10] Mayukh Ghosh: Right. Let's take small steps. I know time is short, but that can't make… I mean, we can't say that I will build steam engine today and aeroplane today, and then also say that I will not fly you because Indigo is on a strike, not in the same class, right? So we'll start with James Watt's steam engine right now. [00:49:29] Mayukh Ghosh: Alright, so, [00:49:33] Mayukh Ghosh: So this is R-squared. Now, in this context of this equation, the linear regression equation we have set up, right, and we are trying to obviously improve my R-square as much as possible. [00:49:42] Mayukh Ghosh: Let me do one thing. [00:49:46] Mayukh Ghosh: Because I didn't expect that we would be getting to this level, so let me show you a couple of slides, okay, quickly. [00:49:52] Mayukh Ghosh: Because that will help, rather than me trying to draw bad diagrams and all of that stuff, okay? [00:50:13] Mayukh Ghosh: Wait. [00:50:31] Mayukh Ghosh: I have to stop the sharing. Otherwise, I'm not even sure where the sharing is going, it's just… [00:50:37] Mayukh Ghosh: Saying you're sharing your screen in the middle of the screen and, you know, creating confusion. [00:50:58] Mayukh Ghosh: Alright, so… Yeah. [00:51:09] Mayukh Ghosh: about that. Right, so… [00:51:12] Mayukh Ghosh: Have you heard of the term? Okay, before I go there, have you heard of the term best fit line? Anyone? [00:51:18] Chandrasekhar Sahu: Yes. [00:51:20] Mayukh Ghosh: Okay, can you say what it is? [00:51:25] Chandrasekhar Sahu: I can say, like, the best fit line is a line where, [00:51:33] Chandrasekhar Sahu: You can see the data point, right? The distance between… the distance from that best fit line is minimum. [00:51:41] Chandrasekhar Sahu: Just know from the theory part, so if you put the best fit line right, the data point, whatever it will be there, around right, the distance would be minimum from that best fit line. [00:51:51] Chandrasekhar Sahu: So… [00:51:53] Mayukh Ghosh: Okay. [00:51:53] Chandrasekhar Sahu: Passily, I… again, oh yeah. [00:51:57] Mayukh Ghosh: No, no, I think, yeah, you're correct. Anyone else wants to add anything to what he's saying? More or less there, I mean, it's not that… [00:52:02] Hemanth: Basically, the line doesn't outfit or doesn't overfit, it goes well with the data. [00:52:07] Mayukh Ghosh: Right, let me show you a good diagram from some website I just quickly opened up. [00:52:16] Jithu Tagore: Like, I thought, like, line passes through most of the data points. [00:52:21] Jithu Tagore: Most of the data points, right? [00:52:22] Mayukh Ghosh: Yeah. [00:52:24] Mayukh Ghosh: See, this is just randomly from some website, okay? This is a diagram. The main idea for us to know here is the diagram, of course. [00:52:32] Mayukh Ghosh: this… this I can say, if it's a scatter plot, and this, all these small circles you can see, you know scatterplot, right? So these small circles are basically my X and Y, height and weight in this case, or whatever it is, right? Now, among… in this scatterplot, if there are innumerable amount, number of data points, whatever they are. [00:52:51] Mayukh Ghosh: I can actually draw infinite amount of lines through these points, right? [00:52:58] Mayukh Ghosh: Any number, right? I don't have a fix that I can only draw one or two lines, right? In a plane, I can draw an infinite number. Now, out of this infinite number, as all of you are saying, say, for example, this line is the best fit line. Why? Because this is closest to most of my points. Can I say that? [00:53:16] Mayukh Ghosh: In an aggregated sense. [00:53:18] Mayukh Ghosh: Or, in other words, the vertical distance, right, between this line and these points, effectively. The vertical distance is the lowest. Can I make this claim, and then say that this is the best fit line for this set of data points? [00:53:33] vinit shah: Yes. [00:53:35] Mayukh Ghosh: Now, vertical distance, some can be positive, some can be negative. [00:53:38] Mayukh Ghosh: So I don't want them to cancel out and get a low number, which is not the true representation. So what I would do is that I would square them. [00:53:46] Mayukh Ghosh: Okay, I can later on root… take a root and get it back to the normal… nominal terms, but right now, I will square them. [00:53:53] Mayukh Ghosh: And I will take… take that the… squared… [00:53:57] Mayukh Ghosh: Distances of all these points to my line, and for whichever line. [00:54:01] Mayukh Ghosh: The sum of squared distances is the lowest. I can easily call that that is the best fit line. With me on this? [00:54:14] Mayukh Ghosh: Right? This is what is a classic… I mean, this is what we call as OLS. [00:54:19] Mayukh Ghosh: Ordinary least square method, which is the way we build linear regression. [00:54:23] Mayukh Ghosh: Okay, the base of machine learning, the absolute ground floor, as I said before, maybe the basement, you can say, not even ground floor. [00:54:30] Mayukh Ghosh: is ordinary least square method for linear regression. [00:54:33] Mayukh Ghosh: Where you are trying to plot two variables. You can do more, but we don't understand diagrams more than 2 or 3D, right? The human brain is not equipped with that. [00:54:43] Mayukh Ghosh: But the machine can, of course, work with them. So we understand the best fit line across variables, where I have plotting… I'm plotting a lot of X variables with respect to 1Y. [00:54:53] Mayukh Ghosh: The same equation I was showing you, Y equal to M1X1, M2X2, M3X3, and so on and so forth, plus C plus E. [00:55:01] Mayukh Ghosh: So that is my linear regression equation. [00:55:04] Mayukh Ghosh: And this is what we… [00:55:07] Mayukh Ghosh: try to sort of solve any continuous data problems with them. Now, what is continuous data problem? Where my target variable is continuous, like sales, profit. [00:55:18] Mayukh Ghosh: something else like, say, revenue, any kind of business problems, whether you know the variable is continuous, it can take any value. It's not a discrete or categorical data. [00:55:27] Mayukh Ghosh: You can then generally try to fit a linear regression equation, which takes this form, and which follows this postulate of best fit line. [00:55:36] Mayukh Ghosh: That's a brief of linear regression equation. [00:55:39] Mayukh Ghosh: Did that make any sense? [00:55:41] Mayukh Ghosh: Obviously, I can't go too deep into it, because it takes… [00:55:46] Mayukh Ghosh: Yeah, it will at least take 4 or 5 hours to do it completely, okay? So I'm just giving you overviews right now, because you have already done decision tree, and also, it's easy to kind of go back and connect if you know the beginning part of it. But whatever little we did in terms of understanding the base of it, did that make sense? [00:56:04] Jagannath Das: Yes. [00:56:05] Deepan Kanagaraj: Yes. [00:56:06] Mayukh Ghosh: Yes? Okay. [00:56:08] Mayukh Ghosh: Now, what I would… what would rather more resonate with your thing is what we call as logistic regression. [00:56:16] Mayukh Ghosh: Anyone here who has had exposure to logistic, or have done, or have… yeah. [00:56:21] Mayukh Ghosh: Yeah, any kind of exposure to logistics from anyone here? [00:56:31] Mayukh Ghosh: Seems not. [00:56:34] Mayukh Ghosh: I don't see a response yet, which I'll take it as a no. [00:56:38] Mayukh Ghosh: That's fine. [00:56:40] Deepan Kanagaraj: Theoretically, we know… [00:56:41] Mayukh Ghosh: Theoretically, no? Okay, okay, tell me what, what you, what you understand theoretically from, for logistic. [00:56:47] Deepan Kanagaraj: Okay, logistic only helps to figure it out only two class problems, and it considers probability to determine the class label, and again, we have something like, logic function to [00:57:02] Deepan Kanagaraj: define that, class variable. Basically, it takes the… [00:57:07] Deepan Kanagaraj: probability, and then it converts into odds, and then, finally, it tried to predict, okay, if anything above 0.5 percentage, then the other class label will get predicted. If it is less than 0.5, then the other class label will predict. So it's basically a binary classification sort of thing. [00:57:24] Mayukh Ghosh: Yes. So, now the question is, to others. [00:57:29] Mayukh Ghosh: Did you understand, or are you aware of what Deepan just said? [00:57:33] Mayukh Ghosh: Or is it not the case with most of you? [00:57:36] Jagannath Das: Yes, yes, we don't. [00:57:38] Mayukh Ghosh: You know? [00:57:39] Jagannath Das: Anyone who doesn't know this base of logistics? [00:57:47] Mayukh Ghosh: Okay. [00:57:48] azad choubey: You can explain, Mayuka a little bit better. Okay. [00:57:52] Mayukh Ghosh: I think what he said is absolutely correct, but yeah, I'll do a repetition or reiteration of that. [00:57:58] Mayukh Ghosh: Let me… I can show you some slides here, which is ready. Linear. [00:58:03] Mayukh Ghosh: couldn't. [00:58:04] Mayukh Ghosh: But that's okay, we are not spending a lot of time on linear. This is more pertinent right now, I guess. [00:58:10] Mayukh Ghosh: So I'll go here. [00:58:16] Mayukh Ghosh: What is the shit? [00:58:25] Mayukh Ghosh: You can… you can see the screen, right? [00:58:30] Mayukh Ghosh: Okay. [00:58:32] Mayukh Ghosh: This screen is a problem. [00:58:36] Mayukh Ghosh: I'll just… I have to… minimize the zoom screen, it's just keep on coming in the middle. Anyway, so… [00:58:46] Mayukh Ghosh: Okay, this is just a supervised unsupervised, which I'm sure you know, so I'm not going to spend any time on this. Now, supervised has regression and classification problem. Regression is what I just… we just spoke about for a while. [00:58:57] Mayukh Ghosh: Where my target variable is continuous, and here, as Deepun mentioned, target variable is categorical, mostly binary classifier. It can be multi-level as well, multinomial. [00:59:06] Mayukh Ghosh: But generally, logistic is binary. We talk about binary models, so to say. All this is fine. I'll go quickly to a back to basics, okay? One quick thing, which I'm sure you guys know, or even if you don't know, it's easy to figure out. [00:59:21] Mayukh Ghosh: Everyone knows what is odds and what is probability. [00:59:27] Mayukh Ghosh: Yeah, even if you don't know, you don't read, you don't understand. [00:59:32] Mayukh Ghosh: Probability of India winning a cricket match and odds of India winning a cricket match is not same, right? [00:59:40] gunjan bhaiya: Yeah, true. [00:59:41] Mayukh Ghosh: Yep. [00:59:42] Mayukh Ghosh: So odds is… the way it is defined, you can see odds can take a value greater than 1. Ideally, it should take a value greater than 1, whereas probability is always between 0 to 1, right? [00:59:54] gunjan bhaiya: So in… can you say that in odds, we are checking the past data, right? Okay, example, if 10 matches happened, example, with the opponent team, right? [01:00:02] Mayukh Ghosh: Right. [01:00:02] gunjan bhaiya: 6, 1, and 4, not also, like, 6 divided by 4, that is so odd. [01:00:06] Mayukh Ghosh: Arts, yes. [01:00:07] gunjan bhaiya: And probability is, like, 50-50. [01:00:09] Mayukh Ghosh: Yeah, probably 6x10, basically, 60%. [01:00:13] gunjan bhaiya: Okay. [01:00:14] Mayukh Ghosh: So, all your order… [01:00:16] gunjan bhaiya: Number and total number of. [01:00:17] Mayukh Ghosh: Total number. So, odds is basically success by failure, probability is success by total. [01:00:25] Mayukh Ghosh: That's the difference. [01:00:28] Mayukh Ghosh: So, now, there is a reason I'm showing this to you, we'll soon come to this. [01:00:32] Mayukh Ghosh: So this is what I was, anyway, talking about. Say, patient having diabetes, not having diabetes, there is a yes, no, this kind of a data you have. [01:00:40] Mayukh Ghosh: 6 by 4, whereas probability will be 6 by 10. And given this probability, the way we're defining total being the denominator, this can't be more than 1. That's why we say probability is always between 0 and 1, right? [01:00:52] Mayukh Ghosh: But if my success… obviously, I don't want people to have diabetes, but if that's what I'm trying to check in this case, say, for example, if my success is more important than failure, or if that is what I'm trying to predict, then I would always want an odds… want the odds value which is more than 1. Can I say that? [01:01:10] Surya Bobbala: Yes. [01:01:11] Mayukh Ghosh: Right? [01:01:14] Mayukh Ghosh: Okay. [01:01:15] Mayukh Ghosh: Now, if I take a log of odds, this is very important, because logistic is built on this slide, in fact, because I think, as Deepan was mentioning, the logic function part and all. The log of odds [01:01:26] Mayukh Ghosh: Here, odds of having diabetes is 6 by 4, odds of not having diabetes is just the opposite, right? 4x6, of course. Then what happens if I take a log? They have the exactly same value of having or not having, the category in one side, in the other way, success and failure. [01:01:41] Mayukh Ghosh: Yes, there is a negative sign. [01:01:43] Mayukh Ghosh: Okay? This is something we will soon use to reduce the logistic equation. Okay. [01:01:49] Mayukh Ghosh: Now… [01:01:51] Mayukh Ghosh: these are some, sort of inputs of this, which is that from a log, how can I get to probability, or from a probability, how can I get to odds? And they are connected, right? Because they have… they share the same numerator. Denominator changes, so we can get to one from the another one, right? [01:02:07] Mayukh Ghosh: All this is fine, what I would rather do… [01:02:12] Mayukh Ghosh: is I would also show you an Excel sheet, which is already done. I don't want to write because we don't have time, to be honest. So I would quickly want to show you something which I did for some other session. [01:02:23] Mayukh Ghosh: in a slightly leadership training for a corporate sometime back, so I can show the same thing to you. [01:02:30] Mayukh Ghosh: Give me a second, I'll share that screen. [01:02:34] Mayukh Ghosh: Yep. Okay. [01:02:36] Mayukh Ghosh: Stop shedding… Okay, something else… Yeah, it's called Class Notes, because… Alright, [01:02:49] Mayukh Ghosh: So, if you can see, this is what we just saw in that Excel sheet as well. And let me take a very easy example, and try to build it from the linear equation we saw. Say, we have data on 1,000 customers from an area, problem statement, okay? [01:03:02] Mayukh Ghosh: We want to know what determines the decision to subscribe to Netflix. [01:03:06] Mayukh Ghosh: Now, obviously, today, in 2025, Netflix is… you don't have a distributor for Netflix, you just pay the money, have a smartphone, smart TV, you watch it there, right? But 15 years back, if you go to US, it was more like our cable TV 25 years back, right? So, you have distributors who are giving it, so they are like the local satellite [01:03:25] Mayukh Ghosh: Channel providers, right? [01:03:27] Mayukh Ghosh: So, in that sort of a scenario, we say, for example, we have a… Data like this. [01:03:33] Mayukh Ghosh: I have only data of age of people, okay? [01:03:36] Mayukh Ghosh: Now, this is basically my Y equal time x plus c, which we just saw in linear regression, right? So, I'm effectively saying that I would try to force-fit this thing on a linear regression equation that we just discussed, okay? [01:03:49] Mayukh Ghosh: The right-hand side, in place of M, we have beta naught, C, we have beta… I mean, M has beta 1 and C is beta naught. That's… that's more terminology in industry sense, but they mean the same thing, okay? [01:04:01] Mayukh Ghosh: And I've taken some random value. Doesn't matter, you can take any value, okay? Well, now, I'll tell you what I'm trying to prove here. That why do I need a separate model for logistic, and why can't I fit it on a linear? One of the more complicated interview questions often asked. [01:04:15] Mayukh Ghosh: is that even on linear regression, y equal mx plus c plus e, then why do you need a separate logic function to build it the moment your target variable, or Y, becomes categorical? [01:04:26] Mayukh Ghosh: Now, it's very intuitively easy to understand. You say, huh, obviously it's categorical, how can I have it in Y this way? It's only a 0-1 kind of thing. Yes, no. [01:04:35] Mayukh Ghosh: But it's very hard to explain. The only way to explain it is to take numbers and do it, okay? So here, my equation is this. I am trying to understand the probability of someone subscribing, which is my success, because subscribing is success and not subscribing is failure for me, right? If I am the person who is giving these people the connection. [01:04:54] Mayukh Ghosh: So, probability of subscription equal to… I can say beta naught plus beta 1 into H plus 0, or MX plus C. Can I say this? Makes sense up to this point? [01:05:05] Ankit Sood: Yes. [01:05:06] Mayukh Ghosh: Yes? [01:05:07] Mayukh Ghosh: I now have randomly taken some values, minus 1.5, 0.00, you can take any value, doesn't matter, okay? Now, we'll try to prove something with this. We'll plot these values here, so the equation becomes this, and I can tell this thing, if the equation is this, line 22, [01:05:22] Mayukh Ghosh: I can also, for the business, right, I can make this conclusion in line 23. [01:05:28] Mayukh Ghosh: Does that make sense? Because business wants to read this line rather than look at this equation, right? [01:05:41] Mayukh Ghosh: Can I say this? [01:05:45] vinit shah: Yes. [01:05:47] Mayukh Ghosh: Yes? Right? So this is how you have to interpret equations, right? Now, let's work around it. Say someone is age 35. Then what happens? I get this equation. Minus 1.5 plus 0.06 into 35. [01:06:00] Mayukh Ghosh: Minus 1.5 plus 1.8, 0.3. Now, if it is 0.3, that means what? That means that a person whose age is 35 in my database has a 30% chance of subscribing to Netflix? Can I say that? [01:06:16] vinit shah: Yes. [01:06:17] Mayukh Ghosh: But the problem is, if that person's age is 25, this becomes 0. [01:06:22] Mayukh Ghosh: Minus 1.5 plus 1.5. So I'm effectively saying. [01:06:26] Mayukh Ghosh: that if you are a young man, who is less than 30 or 25 years old, no matter what you do, no matter how much you pay me, I'm not going to give you subscription. [01:06:36] Mayukh Ghosh: Probability is zero. I'm not giving anyone who is 25 years old any subscription. [01:06:41] Mayukh Ghosh: It's unrealistic, right? [01:06:43] Mayukh Ghosh: And it becomes more unrealistic when you take 45. Probably it becomes 1.2. [01:06:48] Mayukh Ghosh: Is it even possible to have probability of 1.2? [01:06:52] Ankit Sood: No, right? No. [01:06:54] Mayukh Ghosh: It has to be between 0 and 1. [01:06:56] Mayukh Ghosh: If I go less than 25, it becomes negative. If I go more than 40 odd, it becomes more than 1. So, only for a sweet range of age, between, say, 30 to 40 or something like that. [01:07:07] Mayukh Ghosh: This is working. [01:07:09] Mayukh Ghosh: And otherwise, this equation is failing. [01:07:12] Mayukh Ghosh: And this is the reason why a linear, simple linear equation will not work when my target is categorical. Why? Because the right-hand side of this equation has a range, or it can take any value between minus infinity to plus infinity. [01:07:25] Mayukh Ghosh: Whereas, the left-hand side can take only values between 0 and 1. [01:07:29] Mayukh Ghosh: So if I'm trying to equate apples and oranges, I will… I will always have a problem, right? So here, this is my apple, this is my orange. One has a range of anything, one has a range of strictly 0 to 1. [01:07:40] Mayukh Ghosh: And the moment I force-fit an equation sign between these two entities, it will always create chaos like this. [01:07:46] Mayukh Ghosh: Does this make sense? Why it fits? [01:07:52] Mayukh Ghosh: Yes? Alright. Now, what do I do? Okay, failure to ticket. Fine. Now, how to solve this, right? How do I solve this? [01:08:00] Mayukh Ghosh: What do you think I can solve? How do you think I can solve? There are two options, I'll tell you. [01:08:05] Mayukh Ghosh: Given that it's a mismatch, I can only do two things, logically speaking. Either I change or convert something in the right-hand side. [01:08:12] Mayukh Ghosh: to make it comply to left, or I can change anything in the left-hand side to make it comply to right. [01:08:17] Mayukh Ghosh: Now, you think through, and let me know what you think is a better option. Yaditha, you have something to say? [01:08:25] Aditya Banda: Yeah, yeah, but the example that you're given, on the left-hand side. [01:08:29] Aditya Banda: the probability, right? It's not a categorical variable. [01:08:32] Mayukh Ghosh: I have to convert it into probabilities, right? It's, at the end of the day, 0 to 1, basically the range, you know? [01:08:38] Aditya Banda: So, yeah, when you convert… [01:08:39] Mayukh Ghosh: Basically, it's a yes-no, right? [01:08:42] Aditya Banda: Okay. [01:08:45] Mayukh Ghosh: The range is still 0 to 1, right? You take probability because I'm trying to make it continuous in that range, but it can't have any other value, no, outside that. [01:08:54] Aditya Banda: Okay. [01:08:54] Mayukh Ghosh: Yeah. [01:08:57] Mayukh Ghosh: So then what I do? I… [01:09:00] Mayukh Ghosh: Tinker with this part, or I tinker with this part. [01:09:05] Deepak Katara: I think earlier you mentioned about log. [01:09:07] Mayukh Ghosh: So… [01:09:08] Deepak Katara: What do we need to… [01:09:11] Mayukh Ghosh: Yeah, you're right, you have got the… you properly. [01:09:15] Mayukh Ghosh: We will get there, of course. So what we will do is, we will do this. [01:09:18] Mayukh Ghosh: We have to convert prob… we have to change… make sure that my left-hand side is between 0 and 1, or the right-hand side follows. What do I do? [01:09:26] Mayukh Ghosh: I take an exponential of this. [01:09:29] Mayukh Ghosh: exponential of mx plus c, or whatever, beta naught plus beta 1H. [01:09:33] Mayukh Ghosh: Now, the moment I take an exponential, that value will always be positive, right? Exponential can't be negative. [01:09:39] Mayukh Ghosh: So, one criteria is to make it positive. Two things I have to fulfill. Make it positive, and make it less than one. [01:09:48] Mayukh Ghosh: So, first, I take an exponential, make it positive. Now, how do I make it less than 1? [01:09:52] Mayukh Ghosh: I divide that thing by that plus one. Now, plus 1, you can make it plus 2, plus 2, doesn't matter. We have taken a unit value to make it simple. That's the universal way of doing it. So this thing, if I highlight it in the formula bar. [01:10:06] Mayukh Ghosh: This val- this expression will always lie between 0 and 1, can I say that? [01:10:14] vinit shah: Yeah. [01:10:16] Mayukh Ghosh: Yeah? So I'm doing these changes, of course, I'm manipulating, there is no doubt about it. [01:10:21] Mayukh Ghosh: But then, what happens, it's not shown here explicitly, but if you do the right-hand side, left-hand side, up and down, right, and you divide it by 1 minus P, you get this. So basically, this comes here, their unitary method. [01:10:32] Mayukh Ghosh: And you get P by 1 minus p equal to your exponential of mx plus c, okay? You can try it out later on. [01:10:40] Mayukh Ghosh: And then, the moment… to cancel out the exponential, I can do a log to both sides. Exponential goes off. So, again, MX plus C happens, but left-hand side becomes log of P by 1 minus P. And this is my logistic regression equation. [01:10:53] Mayukh Ghosh: Now, what is P by 1 minus P? Just a while back in the slides, we saw, this is nothing but my odds. [01:10:58] vinit shah: mobility, yeah. [01:10:59] Mayukh Ghosh: P by 1 is odds, no? Success by failure? [01:11:02] vinit shah: Yeah, yeah. [01:11:03] Mayukh Ghosh: Right? Yes. So I can say log of odds equal to MX plus C, can I say that? [01:11:10] Ankit Sood: Yes. Yeah. [01:11:11] Mayukh Ghosh: Yes? And that's why I kept the linear here also. In linear, it's a simple Y equal to MX plus C. In logistic, that Y takes the turn of log of odds, because what happens now in this equation. [01:11:23] Mayukh Ghosh: Right-hand side can be any value, minus infinity plus infinity. Given that I have a log outside the probability, this can also take any value. So the initial problem of mismatching the apple to orange comparison is no longer there. [01:11:36] Mayukh Ghosh: So, this is how we approximate logistic regulation equation, and this is super, super important. [01:11:41] Mayukh Ghosh: Anyone who understands logistics, this is the first thing they have to understand. And this is the most important thing, because you need to understand why the equation is what it is. [01:11:50] Mayukh Ghosh: And why… and how we arrive at that equation is very important. [01:11:54] Mayukh Ghosh: And with due respect to decision tree and random forest and boosting and all the other tree models, right? [01:12:01] Mayukh Ghosh: The most rigorous classification model is logistic regulation. There is absolutely no doubt about it. [01:12:06] Mayukh Ghosh: We have moved from it, I understand why it was not explicitly mentioned in the sessions, because recently, or of late, we have moved from it in terms of getting into the GenAI models, or LLM models, for that matter, if you ask me. So that makes sense, I completely get that. [01:12:24] Mayukh Ghosh: But, but, see, the trouble is… [01:12:32] Mayukh Ghosh: But the trouble is… [01:12:33] Mayukh Ghosh: if I… if you want to… see, again, at the end of the day, you have to understand what you want to become, right? If you want to become an ML engineer or an ML expert, for that matter, right? Then this is very important. [01:12:46] Mayukh Ghosh: But if you want to become a data cloud architect, data engineer, data science professional from a GenAI expertise. [01:12:53] Mayukh Ghosh: You might well skip it. It's not required all the time, right? These days, especially. You have high-end transformer models which will do the work, right? So again, depends on where you want to be. But this is critical to understand that this [01:13:07] Mayukh Ghosh: Is the base of logistic, given that you know what is linear. [01:13:10] Mayukh Ghosh: Did this make sense to everyone? [01:13:12] Mayukh Ghosh: How we arrived where we arrived. [01:13:14] Ankit Sood: Yes. [01:13:15] vinit shah: is. [01:13:16] Mayukh Ghosh: Yes? [01:13:16] vinit shah: So, just one question, Mike. So, you mentioned that it's moving away from this logistic, right? So, it's being replaced by what, then? [01:13:23] Mayukh Ghosh: No, nothing in the world can replace logistic, if you ask me. [01:13:26] Mayukh Ghosh: We have just wanted other alternatives, okay? [01:13:29] vinit shah: So what is that alternative? That's what I meant. Maybe. My question was not right. [01:13:33] Mayukh Ghosh: I would say boosting, to an extent. [01:13:35] vinit shah: Okay. [01:13:36] Mayukh Ghosh: I would say ReadyBoost does the job, most of the job what logistic does, Radium Boosting, for that matter, okay? [01:13:42] Mayukh Ghosh: But it is not explainable. [01:13:45] Mayukh Ghosh: logistic is way more explainable, but with a less accuracy. Remember, it's a statistical model, it has so many assumptions, so many rigor, that it often has a… it suffers from low value of accuracy, or precision or recall, okay? [01:13:58] vinit shah: See, I just… [01:13:58] Mayukh Ghosh: Logistic, yes. And that's the reason why people are moving into it, because when you are doing a POC for your stakeholder, right, I'll tell you reality. When you're doing a POC for your stakeholder, trying to sell a project effectively, right? [01:14:11] Mayukh Ghosh: At that point, you can't show a lower accuracy. [01:14:14] Mayukh Ghosh: Because the stakeholder will not understand [01:14:17] Mayukh Ghosh: I mean, most of them will not understand the statistical nuances that led to that accuracy rate. So they will understand that what is the ROI attached with the accuracy of the results you're getting, okay? [01:14:27] Mayukh Ghosh: So, for that matter, sometimes these days, people don't show this, because that doesn't buy them the project. And at the end of the day, you can build a billion project, but if no… there is no buyer, there is no point, right? [01:14:37] vinit shah: Correct. [01:14:38] Mayukh Ghosh: We go to… practicality purpose, we go to boosting these days to show it away, but again, depends. If you're in a bank, for example, I was doing some training for other banks, Access Bank, ICSA Bank, of late. [01:14:49] Mayukh Ghosh: And, I understand that they don't even allow you to use neural networks, okay? RBI has a very clear direction you can't use, anything which you can't explain. [01:14:59] Mayukh Ghosh: Okay, so neural networks, or even boosting for that matter, is not fully explained. There is a bit of ambiguity in the way the work is done, and you can't… you can't see exactly what's happening. [01:15:09] Mayukh Ghosh: So, for them, this becomes very important, because for them, explainable AI is more important than productive AI, okay, if I call it that way. So, it depends on what you were working on, where you were working on, but yeah, people have moved on from this. I have not, because I come from a statistical background, I don't think I'll ever move away from this. [01:15:27] Mayukh Ghosh: But, yeah, plenty of people in the industry have done that, and there are reasons for it. Of course, with the complexity of data we get these days, fitting it to a logistic regression can be hazardous, no doubt. [01:15:41] Mayukh Ghosh: pros and cons, but I would suggest, at some point. [01:15:44] Mayukh Ghosh: If you are really getting into wholesome full-stack data science mode, you must know. [01:15:49] Mayukh Ghosh: That would be my suggestion. [01:15:52] Mayukh Ghosh: Okay. [01:15:55] Mayukh Ghosh: Any questions at this point from anyone? [01:16:03] Mayukh Ghosh: I know you understood what I said, but around that, if anything comes to your mind, maybe we can take a couple of minutes for that, and then I can come back towards the end. [01:16:14] Ankit Sood: So, I have a question, like, you said that its accuracy is not that high, but can we not expect, like, 60-70% as well? [01:16:21] Mayukh Ghosh: That you can. That you can. I'm saying, comparatively, not that high. [01:16:24] Ankit Sood: Okay. [01:16:25] Mayukh Ghosh: Because XDBoost, the way boosting works, you haven't done it yet, but you will see, boosting works when you are building a lot of decision trees, but sequentially. Random forest is something where you are building a lot of decision trees. [01:16:37] Mayukh Ghosh: All simultaneously, right? Can I say that? [01:16:40] Ankit Sood: Meh. [01:16:41] Mayukh Ghosh: But in boosting, it is all sequential. You build one tree, you note down what is the extent of error. [01:16:47] Mayukh Ghosh: And you do something, there are ways to calibrate that, that in the next tree, that error is not repeated. [01:16:53] Mayukh Ghosh: So, in that way, you are actually addressing underfitting and overfitting… [01:16:58] Mayukh Ghosh: all… both together. Random first doesn't do that. Random first is very good to address over 50. [01:17:02] Mayukh Ghosh: But random forest, because it is all simultaneous hundreds of trees, the undefeating part is a bit of a gamble. [01:17:09] Mayukh Ghosh: In boosting, the thing is that it actually systematically reduces both. [01:17:13] Mayukh Ghosh: Okay? And that's why it's very powerful, and we often go to XDBoost and GBA gradient boosting models. [01:17:20] Mayukh Ghosh: get to the high numbers of precision recall, accuracy, F1, whatever you take, right? [01:17:25] Mayukh Ghosh: But then you have to be very smart. So, see, the reason why you should know all of these, or the entire spectrum of things, is because you need to know that it's horses for courses. G2, I'll come to you. It's horses for courses, okay? So, why horses for courses? [01:17:41] Mayukh Ghosh: If you have to understand that your problem statement needs a lot of explainability, and your stakeholder, or the problem you're working on. [01:17:48] Mayukh Ghosh: in a particular scenario, the explainability is of the most importance. You have to always go for a logistic, and then [01:17:56] Mayukh Ghosh: play around it. But if you see that you have to get the project, do a POC, get into the billing board of things, you probably will go for a boosting result, or a random forest result. But [01:18:07] Mayukh Ghosh: that you have to play smartly. Now, the reason why you should be having some idea of all of that, because if you don't know everything, then you don't even have the option to play smartly, because you're then limited with options, right? So, that's the reason we have all these 7, 8 different classification models we generally have. [01:18:24] Mayukh Ghosh: It's not that we are going to use all of them everywhere, but we need to know all of them so that we know that what can be used where. [01:18:31] Mayukh Ghosh: Yeah, you have a question, go on. [01:18:35] Jithu Tagore: Yeah, I'm new to this term, boosting and ST boost. Can you say to learn about that? [01:18:40] Mayukh Ghosh: Yeah, it's good luck. [01:18:41] Jithu Tagore: classification model? [01:18:42] Mayukh Ghosh: It is, it is, it's a bit like random forest only, so it's very similar. You are still building decision trees there. [01:18:49] Mayukh Ghosh: But the only difference is in boosting, as I said, that the trees are built sequentially. You build one tree, okay? See how many… what is the error, extent of error, and you boost them with some weightage we're calling weighing, right? Weight putting weights, so that when you build the next tree, which again has a different set of data, different set of random variables. [01:19:08] Mayukh Ghosh: This particular row for which you got an error, for example, right, will not give an error for the second tree. [01:19:15] Mayukh Ghosh: Something else might be added in the second tree, that will be sorted in the third tree. [01:19:19] Mayukh Ghosh: some error in the third tree will be sorted in the fourth tree. So you build 50, 60, 100 trees, and then, by that time, you have actually sorted plenty of your potential errors. Whereas in a normal decision tree on a random forest, that's very hard to do that systematically, because you're building all the trees together. [01:19:35] Mayukh Ghosh: So whatever is a conservative error at the end of the day, you would have to take it, right? [01:19:39] Mayukh Ghosh: So, boost… that's why the term boosting comes in, that you are boosting. You are boosting one tree to the second tree, boosting more further to the third tree, and you are systematically reducing errors small portion by portion from one tree to second tree, third tree, fourth tree, goes on and on and on, that way. So, that's why the term comes in, boosting. [01:19:57] Mayukh Ghosh: Now, gradient boosting is a direct result of logistic, okay? So, I would suggest, or I would recommend, that before you go to gradient boosting, do a quick revision or a quick study of logistic regression, because they are very well connected, and if your logistic is [01:20:12] Mayukh Ghosh: Quite clear, or pretty clear. [01:20:14] Mayukh Ghosh: Understanding GBM becomes very easy. [01:20:17] Mayukh Ghosh: And most of the unfortunate thing in the industry right now is… I mean, I sometimes get shocks now when I see people in the industry who are working and who are boasting around what they're doing. [01:20:28] Mayukh Ghosh: Is that they're just running codes, getting outputs. [01:20:32] Mayukh Ghosh: Have some idea of how to get to better outputs. [01:20:35] Mayukh Ghosh: And that's it. [01:20:36] Mayukh Ghosh: But the trouble is that if you have to have any longevity for your model. [01:20:42] Mayukh Ghosh: That is not enough. That is remotely not enough. [01:20:46] Mayukh Ghosh: Understanding concepts, understanding connections across models, and what concepts connect each of these models is very important. [01:20:52] Mayukh Ghosh: And logistic regression, like what we just discussed a while back, this thing, what you are seeing in front of your screen, right? The log, the odds, the probability. [01:21:01] Mayukh Ghosh: Boosting is built on odds and probability. So it's very important that you know logistic integration to some extent, so that then when you understand the backstory of boosting, it becomes easier. Because otherwise, these days, people finish boosting in one hour. I see some YouTube program quotes I was seeing where they are doing boosting in 45 minutes. [01:21:20] Mayukh Ghosh: I mean, yeah, so running, from scikit-learn, importing a package, gradient boosting classifier. [01:21:28] Mayukh Ghosh: Putting some hyperparameters, which I'm sure you know, max depth and minimum sample split, and N estimators, this kind of stuff, right, which you did in Adam Forest. [01:21:36] Mayukh Ghosh: You put that up in a bracket in that gradient boosting classifier. [01:21:40] Mayukh Ghosh: The model gives you output, confusion matrix, model gives you classification table, accuracy, precision, recall. [01:21:45] Mayukh Ghosh: You take all of that, you check for overfitting, your model is done. [01:21:49] Mayukh Ghosh: Yes, that indeed takes 45 minutes. [01:21:52] Mayukh Ghosh: But the problem is, all this is fine if it works. [01:21:56] Mayukh Ghosh: If the results are not good enough, what will you do? [01:21:58] Mayukh Ghosh: No one ever tells that in 45 minutes, at least. [01:22:02] Mayukh Ghosh: So you have to prepare yourself in a way where you can ensure that… [01:22:05] Chandrasekhar Sahu: Can we check, like, suppose I build two models. [01:22:11] Mayukh Ghosh: having same accuracy, but I have a different hyperparameter. [01:22:15] Chandrasekhar Sahu: And, I have trained that model with my testing data and training data. [01:22:21] Chandrasekhar Sahu: As we discussed earlier also, we'll be having n number of data sets, it will come, which we have not trained that model. [01:22:28] Mayukh Ghosh: I think. [01:22:29] Chandrasekhar Sahu: So, at the end, who will decide? Like, anyhow, we will going to predict that one, but is there any other methods there, like, that we can predict, like, this model is wrong? [01:22:40] Mayukh Ghosh: Model is wrong, no. That to no, but again, it's not about wrong or right, it's about how to choose the best model, right? [01:22:47] Chandrasekhar Sahu: Correct, in that way. [01:22:49] Mayukh Ghosh: Oh, yeah. [01:22:51] Mayukh Ghosh: So, how to choose the best model is a very complicated exercise. I was doing a session [01:22:57] Mayukh Ghosh: Somewhere else, where it was one hour only on that, how to choose the best model. [01:23:01] Mayukh Ghosh: Now, the first prerequisite for that is you need to know all the models. [01:23:05] Mayukh Ghosh: Okay. [01:23:06] Mayukh Ghosh: Because it's a… if I know only one or two models, and I say that I will choose the best model, it's a bit like that in a class, there are two students, and I am the first boy, right? [01:23:16] Mayukh Ghosh: So, I… that's the first requisite, I need to know at least run 4 or 5 models. If you're doing classification problem statement, you should… minimum is that you run logistic, decision tree, random forest. [01:23:28] Mayukh Ghosh: Boosting. Minimum. Name-based SVM I'm leaving out, because that's optional quite often. [01:23:33] Chandrasekhar Sahu: Huh. [01:23:34] Mayukh Ghosh: But these are minimum, I would say, and then you have to choose the best out of them, right? And that is a complicated exercise. It's not about accuracy, because I'm sure in decision tree context also, you might have encountered this, that when the data is imbalanced, right? Everyone knows what is imbalanced data? [01:23:54] Mayukh Ghosh: Have you heard of the term before? [01:23:56] Surya Bobbala: Nope. [01:23:56] Jagannath Das: Yes. [01:23:58] Mayukh Ghosh: Yes? [01:23:58] Chandrasekhar Sahu: No, no, no. [01:23:59] Mayukh Ghosh: No, no, yes, okay. [01:24:02] Mayukh Ghosh: Alright, now, imbalanced data, you must know. You have to know. Imbalanced data is, in simple terms, is when your target variable, the two, yes and no, basically, the two categories are in imbalance. [01:24:13] Mayukh Ghosh: So, say you go to a fraud data in a bank, okay? The classic example is fraud in bank. [01:24:19] Mayukh Ghosh: So, in fraud in bank, what happens is, you, out of 100 people, who will you fraud? Only one person, maybe 99% of people will not be fraud, right? So, if that is your target variable, detecting frauds, right, in the AML department, the anti-money laundering department. [01:24:33] Mayukh Ghosh: If you're targeting the… trying to predict or understand to figure out the frauds, then out of thousands and thousands of people, only one or two or three are really frauds, right? [01:24:40] Mayukh Ghosh: So in that case, if you have a category, you will… 99.5% maybe is non-fraud. Only 0.5% are fraud, right? And that's a target variable, and it's imbalanced, because one of the categories is underrepresented in a very drastic and gross manner. The other one is absolutely all over the place, right? [01:24:57] Mayukh Ghosh: And in that case, I will show you in the PPT itself, because you brought this up, I would have gone there anyway, that you brought it up. [01:25:05] Mayukh Ghosh: Because I'm… I'm jumping from one PPT to the Excel and all, because given the time, I'm… yeah, I have at least 15 things lined up in my mind which I want to speak about. [01:25:16] Mayukh Ghosh: You can't. [01:25:18] Mayukh Ghosh: I'll just quickly go to… That part, okay? This part. [01:25:27] Mayukh Ghosh: Erito. [01:25:28] Mayukh Ghosh: You're rich, you'll understand. [01:25:30] Mayukh Ghosh: It's not that you have to know logistic for that, it is applicable anywhere, decision tree ran, of course, any classification model. [01:25:44] Mayukh Ghosh: Yes, you can get this PPT, I have… [01:25:48] Mayukh Ghosh: No… yeah, it's not confidential with copyrights, I can give it to you. [01:25:55] vinit shah: I think we were taught something similar to this, [01:25:58] vinit shah: I forgot the topic, you know, something on the decision boundaries, and if you go towards the left. [01:26:03] Mayukh Ghosh: Could be similar, yes, yes, yes. Different terminals? [01:26:06] vinit shah: To go towards the right, yeah. [01:26:08] Mayukh Ghosh: Yes. So this is a real problem. I think… I hope this is understandable to everyone, right? What is the problem here? [01:26:16] vinit shah: Yeah, yeah. [01:26:19] Mayukh Ghosh: Right? So, for this kind of a thing, accuracy is not going to be my… [01:26:24] Mayukh Ghosh: Yardstick for choosing models, right? [01:26:27] vinit shah: Correct. [01:26:28] Mayukh Ghosh: So I can't look into accuracy. Accuracy is actually much less important, and for this precise reason. [01:26:34] Mayukh Ghosh: I have to go to… let me show… you know this, I'll come back to this. [01:26:38] Mayukh Ghosh: You know this? Precision? [01:26:42] azad choubey: Yes, yeah, we are taught it. [01:26:44] Mayukh Ghosh: And recall? [01:26:46] Mayukh Ghosh: Don't look at the first statement, it looks a little bit ambiguous. [01:26:48] azad choubey: Yes, sir. [01:26:49] Mayukh Ghosh: Look at the phone. [01:26:50] vinit shah: Exactly, I was calling out this only, this is what was thought to us, yes, precision and recall. [01:26:54] Mayukh Ghosh: Right? So, we have to understand precision recall. Now, if I go back to the initial example I think one of you were giving about the cancer patient part, right? That's the easy way to understand this. Say, for example… [01:27:07] Mayukh Ghosh: There are two possible errors in a cancer patient scenario, right? The first example is that someone has cancer, but I am saying model is saying they don't have cancer, okay? Second is someone doesn't have cancer, model is saying they have cancer, right? [01:27:20] Mayukh Ghosh: Out of these two, the bigger problem is the first one, where they have cancer, but I'm saying you don't have cancer, right? [01:27:28] Mayukh Ghosh: If I don't give treatment, that person dies, right? [01:27:32] vinit shah: Yes. [01:27:33] Mayukh Ghosh: Whereas the second one also is bad, I mean, you don't want people to just take difficult cancer medicines when they don't have anything, right? But in comparison, the first one is a bigger problem, right? So, here, I need to reduce my false [01:27:48] Mayukh Ghosh: What? Reduce what the more? False positives or false negatives? [01:27:52] Jagannath Das: False negative. [01:27:52] vinit shah: False negatives. [01:27:54] Mayukh Ghosh: False negatives, right? So, I should be increasing recall for this problem statement? The highest… I mean, that should be the… of most importance. [01:28:02] Jagannath Das: Right. [01:28:03] Mayukh Ghosh: This is the first thing. Before you build any model, doesn't matter if it's a logistic or a decision tree or anything else, right? Before opening your Python, Jupyter notebook, or Google Colab, the first thing is to understand the problem statement and understand this thing. [01:28:19] Mayukh Ghosh: Because unless you understand this thing that what is more important to me, recall or precision, every single time, you will go and look at this. I'm sure you have seen this before, right? F1 score, you know? [01:28:33] Mayukh Ghosh: Yes? [01:28:34] azad choubey: Yes, yes. [01:28:35] vinit shah: thought is… [01:28:36] Mayukh Ghosh: So it is basically a harmonic mean of precision and recall, right? Now, it's a good value to get in the middle of it, but if you ask me, in real industry standards, no one cares with your phone score, to be honest, okay? Because [01:28:48] Mayukh Ghosh: people who are relying solely on F1 score, that effectively means that you are unsure whether your recall is more important or precision is more important. And you're trying to go to a middle ground, average ground. [01:28:58] Mayukh Ghosh: And that will not save you in the long run. It will pass your model in the short run, but later on, if that is your chosen yardstick for the model, in 2-3 cycles or sequences, your model will fail. [01:29:08] Mayukh Ghosh: In the production, you've already spent a lot of money on your cloud, or Databricks or Snowflake connection, and all those agreements. You put up your model there, using your F1 score as your, sort of. [01:29:19] Mayukh Ghosh: chosen yardstick. [01:29:20] Mayukh Ghosh: And then the data changes a little bit, your precision recall disbalances a little bit. [01:29:25] Mayukh Ghosh: And… you were done. You spent a few lakh rupees, and you were… [01:29:31] Mayukh Ghosh: You have to answer to the management. [01:29:33] Mayukh Ghosh: This happens very often. [01:29:37] Mayukh Ghosh: So, the bottom line is, make sure that you understand the business problem first. [01:29:42] Mayukh Ghosh: Everything else can wait. [01:29:44] Mayukh Ghosh: And the moment you understand that well, you must understand what is more important, precision or recall. [01:29:50] Mayukh Ghosh: Does that make sense? Why it is so important? [01:29:54] Jagannath Das: Yes. [01:29:55] Mayukh Ghosh: Yes? Okay. [01:29:58] Mayukh Ghosh: This is some numbers, anyway. [01:30:01] Mayukh Ghosh: You know what is ROC? [01:30:04] vinit shah: This was also a thought, yes. [01:30:06] Mayukh Ghosh: ROC curve, right? [01:30:07] vinit shah: support. [01:30:07] Mayukh Ghosh: receiver operating characteristics and all that. It looks a bit like this. [01:30:12] vinit shah: Yes. [01:30:13] azad choubey: Yes, yes. [01:30:15] Mayukh Ghosh: So why… where would I sh- where should I be? [01:30:18] Mayukh Ghosh: In this… in this quadrant, where my mouse is hovering. [01:30:21] Mayukh Ghosh: Where the TPR is highest with respect to FPR, right? [01:30:25] vinit shah: Yeah. [01:30:27] Mayukh Ghosh: So… [01:30:29] Mayukh Ghosh: there is a classic way, I won't be able to go into that. I had it in plan that I would cover that, but I can't, because you haven't done logistic yet, or you haven't done logistic in a full-throttled way. [01:30:40] Mayukh Ghosh: So, there's a connection of ROC with something called log likelihood in the logistic model, okay? So, which you can read up later on, and if you have a doubt or anything, you can come back to me later on, we can discuss, but right now is not the right time for it. [01:30:52] Mayukh Ghosh: But there is a way which… in which we can actually interpret or understand the result of ROC long before we even draw this diagram, okay? Because at the end of the day, machine learning is all about anticipating outputs. [01:31:07] Mayukh Ghosh: It's a very mechanical, boring, and a poor way of doing ML models. If I just run an output and then think, like, oh, Yurek, I have got this, I didn't expect this, right? I mean, that's a very, I mean, naive way of doing things. [01:31:18] Mayukh Ghosh: So, the way for a good modeler, for a good ML engineer, for a good data scientist, is that they run one output, and they can, from that, I mean, run one piece of code or one model. [01:31:28] Mayukh Ghosh: statement, and from that, they can actually anticipate the next 3 or 4 outputs from that output itself. If the concepts are clear, you can do that. So, ROC comes in that context, we'll not get into that. [01:31:39] Mayukh Ghosh: This thing… does this make sense? I'm sure you have talked about AUC as well, right? [01:31:45] vinit shah: They're under a curve, right? [01:31:47] Mayukh Ghosh: The area under the curve, yes, right. Look at these three diagrams, look at them vertically. [01:31:52] Mayukh Ghosh: And tell me whether they make sense. [01:32:07] Aditya Banda: Yeah, I think they do. [01:32:09] Mayukh Ghosh: They do. And given that there are 3 panels, but realistically, you will always be in the middle panel, right? [01:32:16] Aditya Banda: Yeah. [01:32:17] Mayukh Ghosh: the… Right and left are theoretically possible, but that's not going to ever happen in reality, right? [01:32:23] Aditya Banda: There will be some overlap, yes. [01:32:25] Mayukh Ghosh: some overlap. If you look at this small triangle, no? This triangle you can see here, Sorry. [01:32:31] Mayukh Ghosh: This triangle, you can see, you know. [01:32:33] Aditya Banda: Yeah. [01:32:34] Mayukh Ghosh: This triangle is the part of overlap. [01:32:37] Mayukh Ghosh: So, can I say that this triangle represents my FP plus FN? [01:32:45] Aditya Banda: Correct. [01:32:46] Mayukh Ghosh: Right? You have to connect, no, you have to connect concepts, right? So when you look at diagrams, you have to connect with the theoretical concepts. So FP plus FN is that triangle, so the greater that triangle is, my FP plus FN is going up, right? [01:33:00] Jagannath Das: Yes. [01:33:00] Aditya Banda: Yep. [01:33:01] Mayukh Ghosh: And, correspondingly, my AUC is going down. [01:33:06] Jagannath Das: Yes. [01:33:07] Mayukh Ghosh: Right? So, from the confusion matrix, I can actually draw the ROC, or I can connect it back. [01:33:17] vinit shah: Can you comment on that last lane? [01:33:19] Mayukh Ghosh: So I'm saying that this is connecting the Confucian matrix to my… ROC, right? [01:33:24] vinit shah: Correct. [01:33:25] Mayukh Ghosh: So, this way, we have to keep on connecting concepts. I will show you something which is going to be complicated now, I'll not get into that. I'll show you a code file, okay? I'll show you some outputs, and we'll discuss a little bit there, okay? [01:33:39] Mayukh Ghosh: But I hope you are able to make sense of what I'm saying so far, everyone? [01:33:44] vinit shah: Yeah. [01:33:45] Mayukh Ghosh: Yeah. [01:33:47] Mayukh Ghosh: Are you finding it interesting enough? [01:33:49] vinit shah: Just one question, so can the UC become, I don't know, lesser than 0.5? [01:33:54] Mayukh Ghosh: No, no. [01:33:54] vinit shah: The channel was. [01:33:55] Mayukh Ghosh: It won't happen, it won't happen. [01:33:56] vinit shah: 0.5 is the max, least, is it? [01:33:58] Mayukh Ghosh: It's called the AUC for the null model. You know what is a null model? Have you heard of this term before? [01:34:04] vinit shah: No. [01:34:05] Mayukh Ghosh: Null model is a model where it is Y equal to C. The X's are not working. It's basically your independent variables are having no effect on Y. [01:34:14] vinit shah: Okay. [01:34:15] Mayukh Ghosh: Okay? So, effectively, that is what we call as a null model, and this… in that ROC curve, you saw now that red line in the middle that goes, na, 45 degree line? [01:34:25] vinit shah: Got it. [01:34:26] Mayukh Ghosh: That is the line representing the null model. [01:34:29] Mayukh Ghosh: And the blue line is the line representing my model where the X's are working. [01:34:33] Mayukh Ghosh: So that's why we say the farther the blue line is from the red line, the better resume model, the higher resume AUC. [01:34:42] vinit shah: Okay. [01:34:43] Mayukh Ghosh: That's how it all connects back. [01:34:47] Mayukh Ghosh: Okay, I'll show you a quick sort of an output. Obviously, it's a logistic output, but [01:34:51] Mayukh Ghosh: You will get it. It's not that complicated. [01:34:59] Mayukh Ghosh: Okay, so this is an easy data, which I often use, very easy because it's small data, and it's important to understand concept rather than focus on data right now. So it's about admissions data, it's a common data you would find in Kaggle also. It's about people's… [01:35:14] Mayukh Ghosh: people who are trying to get admission for PhD in US or something like that, okay? GRE scores and profile scores and all these things. University ratings, statement of purpose, LOR, CGPA, research, whether they've done research or not. [01:35:27] Mayukh Ghosh: And this is the final chance of admission on the basis of all their performance, okay? Basic stuff. This is my target variable, 0, 1. If they get admission, it's 1. If you don't get admission, it's 0, right? Historical data. [01:35:39] Mayukh Ghosh: import libraries and all that, I'm not going to get into that data, you know how to do EDA? [01:35:45] Mayukh Ghosh: This is fine. [01:35:47] Mayukh Ghosh: encoding… Scaling, trying to split. [01:35:52] Mayukh Ghosh: Logistic regression… [01:35:56] Mayukh Ghosh: you haven't done logistic linear, but have you seen output like this before? Or output table like this in any context of linear logistic? In any context anywhere? [01:36:07] Chandrasekhar Sahu: No. EDS anywhere? [01:36:09] Ankit Sood: Nope. [01:36:10] vinit shah: Got you. [01:36:11] Mayukh Ghosh: Not really, okay. [01:36:13] Mayukh Ghosh: then I'll not, sort of. [01:36:15] Mayukh Ghosh: torture you with this right now. This is… this is… first time it is torture us. It's… it's long. It's long, it's complicated, it uses a lot of statistics, so some idea of stats is very important to understand this table, okay? P-values and all those interval estimates, all those stuff is there, okay? [01:36:32] Mayukh Ghosh: I'm not gonna do that, I'd rather go after this into the main part, which you know, also, is your confusion matrix onwards. Plot the confusion matrix. [01:36:41] Mayukh Ghosh: In this data, I'll give you a bit of a context, okay? So this is the data which has 400 observations, 20% of that is test, 20% of 400 is 80. [01:36:51] Mayukh Ghosh: So I am doing this confusion matrix on my test data. [01:36:55] Mayukh Ghosh: I'm saying predicted is my column, index is my actual, and I'm building them, okay? So this is total 80 observations. [01:37:01] Mayukh Ghosh: You know this confusion matrix, I'm sure you've seen it many times. So, from this confusion matrix, can I tell what is the accuracy? [01:37:13] Aditya Banda: Yeah, I think, accuracy is 33 plus 33 divided by total. [01:37:19] Mayukh Ghosh: Yeah, exactly. 66 by 80, right? [01:37:23] Mayukh Ghosh: Which would be 33 by 40, 33 into 2.5, 82.5, around 82.5, okay? [01:37:30] Mayukh Ghosh: Now, you do this, you can run the classifi… you can run this classification report, which you have done if you are in decision tree also, you have done this kind of thing, and you get this kind of a table, right? You have seen this before? Everyone? [01:37:43] Deepan Kanagaraj: Yes. [01:37:44] Mayukh Ghosh: Yes? Precision recall? It comes in for your ones and zeros and everything? [01:37:49] Mayukh Ghosh: Accuracy F1, that that stuff comes in, right? [01:37:53] Mayukh Ghosh: Now, here you see precision record for this example, is pretty good. [01:37:56] Mayukh Ghosh: Now, before all of this, As someone was saying before, I think it was Deepan, probably. [01:38:03] Mayukh Ghosh: Can you make sense of this line? [01:38:06] Mayukh Ghosh: Everyone. [01:38:08] Mayukh Ghosh: What am I doing here? [01:38:15] Nirav Mehta: Can't do that. [01:38:16] vinit shah: Lying into the tool. [01:38:16] Jagannath Das: hitting the threshold. [01:38:17] vinit shah: probability. [01:38:19] Mayukh Ghosh: Yes, right, so I'll just write it down here. [01:38:22] Mayukh Ghosh: So my actual Y is 0, 1, right? Can I say that? [01:38:27] vinit shah: Yeah. [01:38:28] vinit shah: Yeah. [01:38:29] Mayukh Ghosh: Right? Now, when I try to use this formula I was showing you, right, that equation Y log of P, Y1 minus p equal to mx plus c and all that, right? After that, I can get the P from there, from that equation, right? [01:38:40] Mayukh Ghosh: Yes. Probabilities, which I will call as predicted probabilities, predict drop. [01:38:45] Mayukh Ghosh: So the predicted probabilities will now have to have a cutoff in the middle so that they can be also converted back to 0 and 1, right? [01:38:52] Jagannath Das: Beautiful. [01:38:54] Mayukh Ghosh: The cutoff by default is, as you can see here, is 0.5. [01:38:57] Mayukh Ghosh: So anything… any of the predicted probabilities less than 0.5 becomes 0. Anything greater than 0.5 becomes 1. [01:39:05] Mayukh Ghosh: Make sense, everyone? What we are getting done here, before the conversion matrix, right? [01:39:10] Mayukh Ghosh: You added to your other question? [01:39:11] Aditya Banda: Yeah, so, actual Y, is it between 0 and 1, or it is. [01:39:16] Mayukh Ghosh: It is 0 and 1, that's it. Two categories, yeah. [01:39:19] Aditya Banda: Okay. [01:39:20] Mayukh Ghosh: That's the reason I have to get the probabilities back into the same thing to compare, right? [01:39:25] Aditya Banda: Right. [01:39:26] Mayukh Ghosh: That's why I'm taking the cutoff. [01:39:28] Mayukh Ghosh: To bring it back to this 01 thing. [01:39:31] Aditya Banda: Okay. And the moment I have it as 0, 1 in both actual and both predicted, then only I can draw the confusion matrix, right? [01:39:38] Jagannath Das: Right. [01:39:38] Mayukh Ghosh: Okay, okay. [01:39:40] Mayukh Ghosh: But, remember, for… The 0.5 is the assumptive cutoff I am taking, right? [01:39:48] Jagannath Das: Yes. [01:39:49] Mayukh Ghosh: I'm saying… I'm taking the middle value and doing it. But for every data set, especially for imbalanced data, where one of the categories is highly representing compared to the other. [01:40:01] Mayukh Ghosh: Do you think that 0.5 will always be the best, sort of, to break it into 0 and 1? Possibly no, right? [01:40:08] Jagannath Das: Yes. [01:40:09] Mayukh Ghosh: If your probability distribution, predictive probability distribution is one-sided, 90% of your probabilities are below 0.5, and then if you take 0.5 as your cutoff, then you are not left with enough ones, right? [01:40:21] Jagannath Das: Right. [01:40:23] Mayukh Ghosh: then you cannot tally it back with your actual, and your confusion matrix will be a mess. It will have a lot of false positives, false negatives. Your precision recall goes for a toss, and everything goes for a toss, right? [01:40:33] Mayukh Ghosh: So that means… [01:40:34] Mayukh Ghosh: That this… choosing the right cutoff is very important to get the right precision and recall? Can I say that? Can I make that connection? [01:40:40] vinit shah: Got it. [01:40:41] Mayukh Ghosh: Cheers. [01:40:42] vinit shah: Got it. [01:40:43] Mayukh Ghosh: So, we have to be very, very careful, and we have to be… I mean, this is probably the most important part of the logistic, or any boosting, any classification model, for that matter. [01:40:52] Mayukh Ghosh: Is to get the cutoff absolutely correct. [01:40:57] Mayukh Ghosh: How do we do that, okay? [01:41:00] Mayukh Ghosh: We… now remember, in the ROC, because you guys know ROC, you will be able to understand. In the ROC, or in… for any model, right, the main thing of… the thing of most importance is what? To get a high TP with respect to a low FP, right? [01:41:14] vinit shah: Got it. [01:41:15] Mayukh Ghosh: That is what we're trying to check in the ROC, that TPR and FPR is as far as possible, right? [01:41:22] Jagannath Das: So… [01:41:23] Mayukh Ghosh: Given that, we have a measure, I'm sure you have not heard of this before. It's fairly rare in most cases anyway. [01:41:31] Mayukh Ghosh: I'll get to that quickly. ROC, so you know, and you hear ROC is very good. Anyway, it doesn't matter. [01:41:38] Mayukh Ghosh: We have to… now, the main point here is to identify the best cutoff value. [01:41:43] Mayukh Ghosh: what am I doing? I have run… this is an industry-level code, so I'm running a loop to get it in all and all. I'll also share the code with you, don't worry, you can just explore it with time later on. [01:41:55] Mayukh Ghosh: So this is… I'm getting for a scorecard of logistics, which what I'm doing is that I'm saying that I have only checked the cutoff of 0.5, and I built the confusion matrix, right? [01:42:04] Mayukh Ghosh: I don't want that. I want to look at a cutoff of all possible numbers in… at a gap of 0.1. [01:42:11] Mayukh Ghosh: So I'm saying that first make a cutoff of… cut off at point 1, run everything, then cut off 2, run everything, and after doing all of this, give me for every single cutoff, what are the values I'm getting. [01:42:22] Mayukh Ghosh: It's a complicated… I mean, not complicated code, but it's a very heavy code, takes some time to run. [01:42:28] Mayukh Ghosh: Now, if you see in this case, 0.5 gave me these results. [01:42:33] Mayukh Ghosh: Great. [01:42:34] Mayukh Ghosh: At the moment I go to 0.6, many of these results are actually better compared to 0.5, right? [01:42:40] Jagannath Das: Yes. [01:42:41] vinit shah: Yeah. [01:42:42] Mayukh Ghosh: Which tells me that for this model, 0.5 is not the right cutoff. I have to change my cutoff to somewhere close to 0.6. Exactly where I don't know, I'm getting a sense it's in the zone of 0.6. And then we have a new thing, as I was telling you that you will not know. It's called Uden's index. It's named after someone called Uden. [01:43:01] Mayukh Ghosh: What Utent Syntec does is that it will give you the highest gap between TPR, FPR for every threshold. [01:43:07] Mayukh Ghosh: From a threshold of 3 decimal places from .001, threshold and cutoff are same. 0.001 to 0.099, [01:43:15] Mayukh Ghosh: It will go and check for every single cutoff. [01:43:18] Mayukh Ghosh: Where the TPR minus FPR, what is the num value? [01:43:22] Mayukh Ghosh: And I can go and sort this into descending order. [01:43:25] Mayukh Ghosh: And I'll take the top 5 as head [01:43:29] Mayukh Ghosh: So, this is the kind of value I get. [01:43:31] Mayukh Ghosh: So it tells me that when the cutoff is 0.61, rather than 0.5, the TPR minus FPR is highest, 0.725564. [01:43:42] Mayukh Ghosh: I then go back and convert everything. Again, I change. I increase a 0.5, now I have enough evidence, so I… [01:43:50] Mayukh Ghosh: Take my cutoff as 0.62, and I again build the confusion matrix. Earlier confusion matrix, I had 66 correct, predictions, now I have 68. [01:44:01] Mayukh Ghosh: Does this make sense? [01:44:03] Mayukh Ghosh: The entire flow of things? [01:44:04] Jagannath Das: Yes. [01:44:05] vinit shah: BM. [01:44:06] Mayukh Ghosh: This is what… [01:44:07] vinit shah: Why are we… [01:44:08] Mayukh Ghosh: Yep. [01:44:09] vinit shah: On the cadence index thing, right? Why are we having the… [01:44:13] vinit shah: Top 5, like, is there a… [01:44:15] Mayukh Ghosh: I… at the end of the day, I need only the top, which is only one, the top one. [01:44:20] vinit shah: Yeah. [01:44:21] Mayukh Ghosh: There's a reason for top 5 as well, I'll tell you. It's a little complicated, but I'll tell you. See, look at it, look at it, yeah. We can spend a couple of minutes, this is what I find. [01:44:31] Mayukh Ghosh: The first is 0.62, okay? And difference is 0.72, right? [01:44:37] vinit shah: Got it. [01:44:38] Mayukh Ghosh: Look at the next four. Out of the next 4, the cutoff for 3 of them is around .8586, right? [01:44:45] vinit shah: Correct. [01:44:46] Mayukh Ghosh: And one of them is 0.38, in the opposite direction altogether. [01:44:50] vinit shah: Yeah. But… [01:44:51] Mayukh Ghosh: So here, the threshold is ranging from 0.38 to almost 0.86, which is about 0.5, right? Half of it. Correct. Half of the entire zone, right? [01:45:00] Mayukh Ghosh: But the difference is only changing by 5-6%, right? [01:45:04] vinit shah: Yeah. [01:45:06] Mayukh Ghosh: What can I infer from this? [01:45:08] Mayukh Ghosh: If at all, anything. [01:45:11] vinit shah: Do I have a good class, or good examples, or something? [01:45:15] vinit shah: I don't know, my dataset is better. [01:45:17] Mayukh Ghosh: The… the thing I can probably… [01:45:21] Mayukh Ghosh: interpret is that I don't have a lot of probabilities between 0.38 and 0.86. [01:45:30] Mayukh Ghosh: I don't have a lot of probability between 0.38 and 0.86.86. [01:45:34] Mayukh Ghosh: Because my threshold… difference is not changing a lot. [01:45:37] Mayukh Ghosh: If I had a lot of probability, predicted probability values in this range, then these two numbers has to change drastically. It can't be the same difference every time, right? Similar. [01:45:48] Mayukh Ghosh: So what happens here, it tells me that either I have a lot of probabilities below 0.3, or a lot of them above 0.86. [01:45:57] Mayukh Ghosh: So I don't have a lot in the middle. [01:46:01] Mayukh Ghosh: Which tells me what? Which tells me that either the values are very corner towards 1, or very corner towards 0. [01:46:07] Mayukh Ghosh: The borderline part in the middle is not there a lot. I mean, they are there, but not a lot. [01:46:12] Mayukh Ghosh: So, to get this kind of an understanding of the distribution of probabilities, we look at 5, even though we are going to use the top. [01:46:20] vinit shah: Okay. [01:46:21] Mayukh Ghosh: Okay, so it's just to get a sense, more insights from [01:46:25] Mayukh Ghosh: yeah, the output I can't see, basically. [01:46:29] Mayukh Ghosh: Because remember, here I have only 80. [01:46:31] Mayukh Ghosh: But for a larger data, this 80 might be 8 lakhs, right? [01:46:35] Mayukh Ghosh: So it's not possible for me to go and look at every probability, so I have to assess a lot of things from these kind of outputs, and make sense, and make connections of concepts, right? [01:46:47] Nirav Mehta: Does that mean that if we take 0.38 or 0.85 as cutoff? [01:46:53] Nirav Mehta: One, cool. [01:46:54] Nirav Mehta: TP and, FP would kind of, not change much. [01:46:59] Mayukh Ghosh: Not changed a lot, yes, you're right. [01:47:00] Nirav Mehta: Oh, okay. [01:47:01] Mayukh Ghosh: Yes. [01:47:02] Mayukh Ghosh: So that, that is essentially telling key, anything I take between 0.38 and .86, the change is not a lot, right? [01:47:14] Mayukh Ghosh: Which is where I am… even though I'm not going to do that, I can't go on and take all numbers in between, but that… that tells me that the probability there's a cluster is either this side or that side, so 0.5 definitely is not going to work for this data, right? [01:47:32] Nirav Mehta: Got it. [01:47:33] Mayukh Ghosh: Yeah. [01:47:35] Mayukh Ghosh: Okay, so this is what it is. This improves from point A2 to 0.85, all of that is fine. Now, one important thing here, again, not only logistic, it applies to any model you build. [01:47:45] Mayukh Ghosh: Because confusion matrix comes in every one of them. [01:47:48] Mayukh Ghosh: Is that there are two types of data scientists. [01:47:51] Mayukh Ghosh: who deal with this type of a result, and then take a course of action, okay? The majority of them, 80-90% of them, will be very happy with this result, and with enough good reasons, given that these numbers are very good. [01:48:04] Mayukh Ghosh: No doubt, these numbers are all very good, right? So I'm pretty happy. I go ahead, I build my model, I… [01:48:12] Mayukh Ghosh: pass it and whatever, all that stuff, the next usual stuff, and go for production and all that, right? There's another minority group of data scientists. [01:48:20] Mayukh Ghosh: And I fall in that group, that's why I'm telling you, otherwise I wouldn't have told you. [01:48:24] Mayukh Ghosh: Who will not be very happy with this result, yet? [01:48:28] Mayukh Ghosh: Satisfactory, but not yet final. I'll tell you why. [01:48:34] Mayukh Ghosh: They will… what they will first do, or what I generally first do, is I look at these 12, which are misclassified. No matter how many are classified, I will look at the ones they are misclassified. It can be 10%, 5%, 20%, 40%, I don't care. But who are misclassified? [01:48:48] Mayukh Ghosh: white. [01:48:50] Mayukh Ghosh: look at… you read through this thing, okay? Let me interpret this a little bit for you. Now, out of these 12, there are two types of misclassifications that can happen, okay? The first is what I call, it's a terminology I have given, it's borderline misclassification, the second is [01:49:08] Mayukh Ghosh: Extreme misclassification. [01:49:11] Mayukh Ghosh: What are they? Let's look at this. [01:49:13] Mayukh Ghosh: Out of 12 misproduction, I can see, for example, these are all hypothetical. Say, one of them is where the actual value is 0, but my predicted probability is 0.9, okay? [01:49:24] Mayukh Ghosh: And the second one is actual value is 1, [01:49:27] Mayukh Ghosh: And a predicted probability or predicted value of Y is 0.1. [01:49:32] Mayukh Ghosh: In these two cases, do you think that by changing the cutoff or doing anything, I will be able to make them correct predictions? Is it possible? [01:49:41] Jagannath Das: No. [01:49:42] vinit shah: Nope. No. [01:49:43] vinit shah: Nope. [01:49:44] vinit shah: not… [01:49:44] Mayukh Ghosh: Look at the next two. [01:49:48] Mayukh Ghosh: 55? [01:49:49] Mayukh Ghosh: And 0.4, FQL1, FL0. [01:49:52] Mayukh Ghosh: In the next two, by changing the cutoff a little bit, I can actually make them right classifications, right? [01:50:00] Ankit Sood: Yeah. [01:50:00] Jagannath Das: Yes. [01:50:02] Mayukh Ghosh: Now, what do you think is bigger concern here? The first set of people, or second set? [01:50:09] vinit shah: Concern. Concern as in? [01:50:12] Mayukh Ghosh: Concerned as in… who I should be bothered in terms of result out of this model. [01:50:18] vinit shah: The second part. [01:50:19] Nirav Mehta: Second one. [01:50:21] Mayukh Ghosh: Why do you say second one? [01:50:23] vinit shah: Because it's in your, I don't know, you can control it or something? [01:50:27] Mayukh Ghosh: No, it's actually the opposite. It's the first one I would be very bothered about. I'll tell you why. No, your reason, Vinish, is correct, that I can control it. So that part, I agree, but that will not bother me a lot, I'll come to that. So, see… [01:50:40] Mayukh Ghosh: The first one is where I'm getting massive mistrustifications, right? I cannot rectify it in any way. [01:50:47] Mayukh Ghosh: I didn't… [01:50:48] Mayukh Ghosh: I have to understand, from the EDR, from the previous step itself, that why are these a misclassifications? Let me take an example. [01:50:55] Mayukh Ghosh: Say, I think I took the example before also, maybe in the EDA session, I don't remember, it happened a long time back. [01:51:01] Mayukh Ghosh: say I go to a branch of… branch in a bank, and I'm trying to look at transaction history of a lot of people, okay? [01:51:08] Mayukh Ghosh: And in the branch, for example. [01:51:11] Mayukh Ghosh: I go to, say, BKC in Bombay, okay, and I look at some accounts, account holder names, maybe, Sajidindulkar, I don't know, Shah Rukh Khan, some people like that, and people like you and me. [01:51:24] Mayukh Ghosh: Now, obviously, if I'm trying to build a model, or trying to find the best customers, or who are likely to take a loan, right, I'm trying to build a classification model with all the customers in the bank. [01:51:36] Mayukh Ghosh: Now, in that, you and I are two data points, or all of us are data points, and [01:51:41] Mayukh Ghosh: Mr. Khan and Mr. Tendulgar are also two data points, right? [01:51:45] Mayukh Ghosh: But do you think that their way of… their… the X variables, or the variables pertaining to their features, right? [01:51:52] Mayukh Ghosh: They will be not very similar to our features, right? [01:51:56] vinit shah: Yes. [01:51:57] Mayukh Ghosh: what their daily transaction is, is probably more than our CTC, right? So, given that, yeah, not probably, I'm fairly sure it's more than our CTC, okay? So, given that, I mean, it's a bit of a mismatch, right? [01:52:13] Mayukh Ghosh: So… [01:52:14] Mayukh Ghosh: if I… so this, basically, this predicted probability actual, the first two people are basically my Shahu Khans and Sachin Dendorkas, right? [01:52:22] vinit shah: Correct. [01:52:23] Mayukh Ghosh: So, then the question is, why am I so bothered about this set of people? Here I have 80 data points and 12 errors, which is small, right? But in a real data, I might have 80,000 data points and [01:52:35] Mayukh Ghosh: 12,000 units, 12,000 of these, right? [01:52:38] Mayukh Ghosh: Now, out of this 12,000, I have to understand that how many of them, or what percentage of them, fall in this category. [01:52:45] Mayukh Ghosh: or what I call as extreme mismatches, the tend to look at category, basically. Okay. Why? [01:52:51] Mayukh Ghosh: Because, see, By keeping them in the model and getting them as mismatches. [01:52:56] Mayukh Ghosh: I'm doing a gross disservice to everyone. Why? Because I'm building a model which is not optimal for all of us, and not optimal for these celebrities as well. So I'm trying to find a middle ground [01:53:08] Mayukh Ghosh: Which is approximating neither them nor us. [01:53:13] Mayukh Ghosh: Right? Does it make sense? What I'm saying? [01:53:15] vinit shah: Yeah. [01:53:16] Jagannath Das: Please. [01:53:17] Mayukh Ghosh: Which means that if I have enough number of people here, right. [01:53:24] Mayukh Ghosh: like, maybe, as I said, 12,000, 15,000, enough number of people who are showing this sort of, characteristic. [01:53:32] Mayukh Ghosh: Then can I also come back and say that my data is actually heterogeneous, not really homogeneous, right? [01:53:38] vinit shah: Yeah. [01:53:39] Ankit Sood: Yes. [01:53:39] Mayukh Ghosh: And for a heterogeneous data, I should not be going ahead and building a logistic regression in the first place, I should be doing a clustering in the first place. [01:53:48] Jagannath Das: Beautiful. [01:53:48] vinit shah: Beautiful. [01:53:49] Mayukh Ghosh: Then do a fit separate classification models on the celebrity group and the common people group. [01:53:55] Mayukh Ghosh: So both of them will be serviced properly, and the accuracy will be sustainable. [01:54:00] Mayukh Ghosh: Make sense? [01:54:02] Ankit Sood: True. Yes. [01:54:03] Mayukh Ghosh: And that is something we will overlook the moment we see good numbers here. [01:54:10] Ankit Sood: Hmm, yeah. [01:54:10] Mayukh Ghosh: That's the error 90% of data science people still keep on doing after years and years of exposure. [01:54:20] Mayukh Ghosh: Because we are fascinated by numbers we want to see. We don't want to have a research mindset and delve deep. That's the trouble. [01:54:27] Mayukh Ghosh: And you might ask that why, what happened? Number is good, we can keep on carrying on. No, we can't keep on carrying. What happens is that in the next dataset, the moment the ratio of celebrities increase or decrease, right, then these numbers will not be sustainable, overfitting will happen. [01:54:43] Mayukh Ghosh: But I know that the initial model I built was very nice, fantastic. Then I'll scratch my hair thinking that, yeah. [01:54:49] Mayukh Ghosh: I can't find a reason. [01:54:50] vinit shah: Oh, okay. [01:54:53] Mayukh Ghosh: This is what I see very often, that's why I'm telling you. [01:54:58] Mayukh Ghosh: So, no matter what your accuracy is, no matter what good… how good the numbers are, always analyze your errors. [01:55:04] Mayukh Ghosh: Because that might give you a lot of more insights than you can think of. It can tell a story about a RDA, it can tell a story about the quality of data in the first place. [01:55:14] Mayukh Ghosh: Understandable, everyone? [01:55:16] vinit shah: Yeah. [01:55:18] Mayukh Ghosh: So I guess this is probably a new thought process for all of you? Or have you thought of it, anyone, before? [01:55:23] vinit shah: No, I guess we'll all remember that such an [01:55:26] vinit shah: Garokan are most important than anything else. [01:55:31] Mayukh Ghosh: Yeah, no, I take very catchy names, because then you remember the example, right? [01:55:36] vinit shah: Yes. [01:55:37] Ankit Sood: Definitely. [01:55:37] Mayukh Ghosh: Yeah, if I tell… say, for example, if I tell, [01:55:41] Mayukh Ghosh: if I tell something like, say. [01:55:45] Mayukh Ghosh: I don't know. Raj, Takare and, Ravuram Rajan… [01:55:53] Mayukh Ghosh: you might not be remembering as much as you would remember a Tendulkar or a Shah Rukh Khan, so that's how things goes. But yeah, the main point was that you need to see… you can track back all of these things, you can connect, right? [01:56:05] Mayukh Ghosh: you can connect a confusion matrix to ROC, you can connect that to classification, you can connect that back to your cutoff and your original data, and you can connect this way, from the ERS itself, you can also connect back to see whether your EDA was good enough, whether you had done outlier treatment properly. [01:56:22] Mayukh Ghosh: And if you are done, and you still see problems here, or patterns here, you can possibly go and look at the data itself, which you haven't done or looked properly in the beginning. [01:56:32] Mayukh Ghosh: So at this stage, we can actually spread out our analysis in so many directions that we can have a complete grip [01:56:40] Mayukh Ghosh: On what we have done, and we can actually explain this well to the business in all possible sense. [01:56:47] Mayukh Ghosh: Because otherwise, if you're not confident, right, and you're just showing outputs, then business will eat you, eat you alive. I've seen that happen many times. [01:56:55] Mayukh Ghosh: If it goes to high-level board of directors and all that, which often these models go. [01:57:00] Mayukh Ghosh: They'll call you and say… they'll ask hundreds of questions you don't even know what to answer. [01:57:05] Mayukh Ghosh: And because you're not confident about what you've done in the context of the business, right? You… you are just stammering all over the place. [01:57:11] Mayukh Ghosh: So… [01:57:13] Mayukh Ghosh: To get complete grip, understanding of concepts is absolutely paramount. Running code, anyone can do. You give it to a college guy, two months training of Python, they can run it. [01:57:23] Mayukh Ghosh: Not a big deal. [01:57:24] Mayukh Ghosh: So, that's not the point. The point is, how can I connect concepts? [01:57:28] Mayukh Ghosh: Can I… if I run a model, can I connect… can I… [01:57:31] Mayukh Ghosh: resume what will be my ROC. [01:57:33] Mayukh Ghosh: From that, can I presume what is AUC? From that, can I come back to confusion matrix? From that, can I go back to EDA? And from that, can I understand data better? [01:57:42] Mayukh Ghosh: This chain of sequence, right? If my concepts are clear, I can do it anytime for any dataset. [01:57:48] Mayukh Ghosh: Okay. Make sense? Any questions? [01:57:54] vinit shah: Give us a good, I guess, a good perspective to what… [01:57:58] Mayukh Ghosh: Yeah, see, again, Emil, no. [01:58:00] Mayukh Ghosh: obviously, I mean, 2 hours for a ML masterclass, or whatever you call it, right? It's about discussing or exchanging ideas. I'm like, I cannot go and teach you anything, or that's not the point even, right? Because ML… [01:58:13] Mayukh Ghosh: takes 50, 60 hours to get into any… even scratch the surface well, if you ask me, okay? It takes months to master it, right? [01:58:21] Mayukh Ghosh: It's an ocean, okay? [01:58:23] Mayukh Ghosh: So, the thing we intend to do is discuss ideas and exchange ideas, right? And that's what we're doing. So, effectively, it's to broaden your horizon as much as we can in the short time, because then, if the horizon is broadened, you can yourself figure out a lot of things. You don't need people to sort of tell you every time, right? [01:58:43] Mayukh Ghosh: Okay, one thing, let me… [01:58:50] Deepan Kanagaraj: Mike, one question… [01:58:52] Deepan Kanagaraj: Let's say, somebody from e-commerce gave you a model-building assignment, and then you kind of took 50 features and then built a model for them, and then they kind of productionized it. [01:59:04] Deepan Kanagaraj: And now, the same e-commerce company is saying that, okay, I'm actually adding, 5 or 6 more additional features to the input, what I'm going to feed. Or, let's say, like, they have enhanced their user interface, which is going to capture a lot more additional information now. [01:59:20] Deepan Kanagaraj: So in this situation, as a modeler, will you be going back to the drawing board and then do the same exercise again, or… [01:59:26] Deepan Kanagaraj: You will be just trying to enhance your existing model, what you already built it out. My question is more to do with, are you going to do the same set of exercise one more time, considering the additional six features, or it's an additional rework on the existing ones? [01:59:41] Mayukh Ghosh: If it is 4 or 5 variables, I will not want to rework, to be honest. [01:59:46] Mayukh Ghosh: Okay. If the data complete lot change is invariable as well as the composition of the variables, yes, I have to. I don't have a choice, really. [01:59:56] Deepan Kanagaraj: Okay. It's just adding a few variables, but the existing variables are having similar compositions, I will not want to… I mean, I don't want to build a new model every single time I add two variables, right? [02:00:05] Mayukh Ghosh: So, that is… [02:00:08] Mayukh Ghosh: what I won't do, and that's why there is a trick in the industry, I don't know how many of you know this or have heard of this. [02:00:13] Mayukh Ghosh: There's a thing called creating buffers. Buffer dummies are buffer variables. Have you heard of this? [02:00:21] Mayukh Ghosh: No. So, say, you know dummy variables, right? Encoding, you know, right? Yeah. So, in one-off encoding, I'll give you an example from one-on encoding, you can extrapolate it to other variables also. Say, in a one-node encoding, for example, I'm having a data on some element for the four metros in India, for example, okay? Delhi, Bombay, Al-Qata, and Chennai, okay? [02:00:40] Mayukh Ghosh: Now… [02:00:41] Mayukh Ghosh: I expect that given that some other cities like Bangalore, Hyderabad, Pune, and maybe some other cities were slightly… almost there in terms of the metro area, right? Metro eligibility. So, I expect that in my data. [02:00:56] Mayukh Ghosh: We'll be capturing them in the near future. Some of these may be a Bangalore or a Hyderabad kind of a variable, right? City. [02:01:02] Mayukh Ghosh: So what do I do? Every time I add one more city, I have to go and build everything from one odd encoding and build a model. Does that even make sense for me? The answer is no, it doesn't, right? So I will just keep on building one model all my life, just because I've kept adding cities, right? [02:01:16] Mayukh Ghosh: So, I don't want to do that. So, what do we do? Is that we create something, or we add something called as a buffer dummy. [02:01:24] Mayukh Ghosh: What is buffer telling me? Is that when we are doing the one-off encoding, for example, we are creating variables for each of the cities, right? That's what is one-off encoding we know as, right? [02:01:33] Mayukh Ghosh: So we say that buffer 1, buffer 2 kind of thing, we just put placeholders, effectively, and say that in the pipeline, in the CI-City pipeline, when the data engineering team is kind of putting it back into the sections of the model, right, in the… using the file. [02:01:48] Mayukh Ghosh: At that point, we'll have the instructions in that way, in the automation part, that when this… when you get one or two new categories that comes in, maybe a few times, Bangalore might come two times only, and just new variables starting off, right? Then you just slot it in there. [02:02:03] Mayukh Ghosh: So, we create provisions. [02:02:06] Mayukh Ghosh: the model I'm building now is not affected, but I create provisions so that the entire pipeline works, just inserting them as they go along. So, for a few variables, that is possible. [02:02:19] Mayukh Ghosh: But if, obviously, it's… I cannot create… keep on creating buffer for 20 variables, right? So, 2, 3, 4, 5 is okay. But more than that, yes, I have to bring it back and do it. [02:02:30] Deepan Kanagaraj: Okay, got it. [02:02:32] Mayukh Ghosh: But there is a provision. There is a provision, and which we end up using, because otherwise, as you understand, right, small change here and there every time… yeah, ridiculous, right? [02:02:45] Mayukh Ghosh: Okay. So, yeah, but that's a good question. That's a useful thing to know. I mean, that's more practical work is connected to that, so that way, I think, yeah, it makes sense to know that. [02:02:58] Mayukh Ghosh: All right. So, this was… so I wanted to go into game chart, lift chart, but I'll not go into that, given, we don't have enough time, and that… that's a long discussion, and… [02:03:10] Mayukh Ghosh: Yeah, but I would suggest you guys read up a little bit on logistic regression later on as well, please. And if you have anything, just feel free to come back to me in whatever means it's possible, okay? So I don't have a problem with that. [02:03:24] Mayukh Ghosh: Always happy to help. [02:03:26] Mayukh Ghosh: But the point is, please do go ahead and read up a little bit, is my suggestion. [02:03:31] Mayukh Ghosh: And then gain chart, lift chart, all of this will make more sense. Now that you already know decision tree random forest, which is good, so it will be easier for you to [02:03:38] Mayukh Ghosh: kind of master the rest of the classification models. Starting is always difficult, whatever it is, right? So that I would suggest you do, please do. [02:03:46] Mayukh Ghosh: Because that… that gives you a lot more confidence, even if you're not using it. It's not about using all the time, getting overall confidence is very important, right? [02:03:53] Mayukh Ghosh: And then, after that, gain chart, lift chart, just note down the names. [02:03:58] Mayukh Ghosh: Using precision recall, it happens. [02:04:00] Mayukh Ghosh: So, that part you know, it should not be a problem, and if at all, let me know at that point, we can obviously figure out ways to get that done, okay? So I'm not worried about that, but please get that as a follow-up, is what I would suggest as you kind of embark into this ML world, so to say, right? [02:04:18] Mayukh Ghosh: So, yeah, I would suggest that as an immediate sort of thing for all of you, okay? [02:04:24] Mayukh Ghosh: Of course, boosting, I believe, will happen, neural networks will happen, so definitely do them… I mean, do them well. Those… both are very important, okay? Very, very important topics, because boosting is a bridge of… bridge between classical machine learning and deep learning. [02:04:40] Mayukh Ghosh: Okay? So you understand boosting well, you connect it back with logistic regulation and decision tree, and at the other end, that takes you towards ANN and CNN, RNN kind of models, okay? So it acts as the middle ground, so hence it's very important. [02:04:56] Mayukh Ghosh: So, those things will come up as the next thing and all, and then, just to give you a bit of an idea in terms of [02:05:03] Mayukh Ghosh: the entire flow of things, I would say, okay? [02:05:06] Mayukh Ghosh: I believe you know this, but I'll still go ahead and say, [02:05:14] Mayukh Ghosh: when you go to LLMs and the generative AI models, which is, I guess, the final, sort of, destination of the program, essentially, right? [02:05:22] Mayukh Ghosh: don't jump into it, the program will not allow you to jump into it, I know, but from your own perspective, I'm saying… [02:05:27] Mayukh Ghosh: Don't jump into it. [02:05:29] Mayukh Ghosh: Follow the sequence. [02:05:31] Mayukh Ghosh: Okay. [02:05:32] Mayukh Ghosh: What I mean by sequence is do ML, do DL well. [02:05:36] Mayukh Ghosh: Do NLP well, very, very important, okay? Do NLP well, because NLP and deep learning, the ANN and RNN LSTM, very important. Because if you don't understand RNN LSTM, and your basic embeddings in NLP, I'm sure you have heard of it, embeddings in NLP, these three things. [02:05:57] Mayukh Ghosh: then you won't understand encoder, decoder, and transformers, okay? [02:06:01] Mayukh Ghosh: And if you don't understand encoder, decoder, and transformers, And, [02:06:09] Mayukh Ghosh: you jump into, LLMs and the lanctions and all those, yeah. [02:06:15] Mayukh Ghosh: Models, you will not understand a single thing there. [02:06:19] Mayukh Ghosh: You will know how to run models, but you will not understand, okay? Because understanding transformers and self-attention is most important. I'm sure you've heard of this. There's a very famous paper, you can look it up later on, not now, don't… because you will get confused if you do it now. 2017 paper by Google Brain Team called Attention is All You Need. That's the paper, we changed everything after that. [02:06:38] Mayukh Ghosh: So, that will come, but self-attention is very important, and to understand self-attention, encoder, decoder, RNN, LSTM, and embeddings in NLP, or 2VEC we call it. These 4-5 things are absolutely critical. [02:06:52] Mayukh Ghosh: So make sure when those things happen. [02:06:55] Mayukh Ghosh: you leave no stones unturned. You have to get it completely. Otherwise, yeah, I mean, the program will be, yeah, it will not be a success for you. As simple as that, okay? So make sure that you are very, very focused on those areas when it comes. Before that. [02:07:11] Mayukh Ghosh: Also, spend some time, as I said, on logistic regression and gain chart lift chart stuff, and then connect it back with the tree models. But then later on, those things will also be coming in the same sequence. [02:07:23] Mayukh Ghosh: I will not get into anything, because we are at 8 o'clock anyway, but I can spend a bit more time, we can spend if you are all fine. If you have questions, I think I can take it for a while. [02:07:35] vinit shah: Can you just repeat those 3 or 4 topics that you mentioned, Maya? [02:07:38] Mayukh Ghosh: Right, definitely. So, I would say… [02:07:42] Mayukh Ghosh: RNN, recurrent neural networks, and LSTM is long, short-term memory, the full form. [02:07:48] Mayukh Ghosh: RNN LSTM, obviously, is neural networks, so you need to know basic neural networks for that, but basically, these two in the neural network part. [02:07:56] Mayukh Ghosh: And there is a thing called embeddings in NLP. Embeddings using what to VEC. What to VEC is? VEC stands for vector. What to VEC? [02:08:05] Mayukh Ghosh: So how you are converting word to vectors, basically? Making numbers out of words. So that process is important, because RNN, LSTM, and Word2VEC are the basis for a topic called encoder and decoder. [02:08:19] Mayukh Ghosh: So, you need to know that. Now, encoder-decoder, you need to know, because your transformers are built using encoder and decoder pipeline. [02:08:27] Mayukh Ghosh: Okay? So it's all sequence. [02:08:29] Mayukh Ghosh: And your transformer uses encoder and decoder pipeline in the context of a concept called self-attention. [02:08:37] Mayukh Ghosh: Okay. [02:08:38] Mayukh Ghosh: So, these terms are absolutely critical, because [02:08:41] Mayukh Ghosh: Every single one of them you have to understand well. [02:08:43] Mayukh Ghosh: Then only you will understand how large language models are built. [02:08:53] Mayukh Ghosh: Otherwise… Yep. [02:08:57] Mayukh Ghosh: Like, a lot of the people these days, no? [02:08:59] Mayukh Ghosh: Talking, yeah, we'll have RAG, LangChain, pipeline, this, that, some fancy terms. [02:09:07] Mayukh Ghosh: To get to the back of the fancy terms, these, what I said, is the backbone of that. [02:09:15] Mayukh Ghosh: Okay. Anyone else has any other questions? [02:09:21] Mayukh Ghosh: I hope this was helpful, I know it's very short time, 2 hours, but I hope I could open up a few things for some of you. [02:09:28] vinit shah: Yes. [02:09:29] Deepak Katara: Yep. [02:09:31] Deepak Katara: Just one thing, if you could share some, references. [02:09:36] Deepak Katara: Or study material, right? [02:09:38] Mayukh Ghosh: References for your voice, Deepak, is very low, I don't know why. [02:09:41] Deepak Katara: Okay, so I was asking if you could share some references or study material. [02:09:45] Deepak Katara: The goose… [02:09:46] Mayukh Ghosh: Conferences of which one? [02:09:47] Deepak Katara: On, like, how to go about these topics. [02:09:50] Mayukh Ghosh: Okay. [02:09:51] Deepak Katara: If we go and search on Google, then we would get confused, like, we have to pick and what. [02:09:56] Mayukh Ghosh: Yeah, yeah, that… the trouble with not knowing and doing ChatGPT is very dangerous, right? Yeah, yeah. Yeah, no, no. Fine, I'll keep that in mind, so when I share some of the materials I showed today, I will, along with that, I'll give you a couple of good resources as well, as much as I can remember, yeah. [02:10:13] Deepak Katara: Sure, thanks. [02:10:17] vinit shah: Thank you. [02:10:22] Mayukh Ghosh: Alright, if we have no other questions, I'll stop. Otherwise, if you have questions, I'm happy to answer. [02:10:30] Deepan Kanagaraj: Mike, this is nothing to do with question, but again, what sort of sessions you will be… [02:10:36] Deepan Kanagaraj: joining us, and I know there are multiple [02:10:41] Deepan Kanagaraj: facilitators are taking care of a few topics, right? [02:10:45] Mayukh Ghosh: Yeah. [02:10:46] Deepan Kanagaraj: Is there any allocation that happens with you. [02:10:49] Mayukh Ghosh: Not really, not really right now, to be very honest with you. It depends on a lot of permutations and combinations of availability, time. [02:10:56] Deepan Kanagaraj: Belt Peak. Okay. [02:10:57] Mayukh Ghosh: Okay, nothing I… yeah, nothing fixed at this point, but yeah, like, like, when I came the last time, I didn't know I was coming for this, right? [02:11:05] Deepan Kanagaraj: Oh, okay, okay. [02:11:06] Mayukh Ghosh: Yeah, so it might happen, yes. [02:11:08] Deepan Kanagaraj: Okay. Because sometimes, [02:11:11] Deepan Kanagaraj: missing Durga man session, again, not calling anything, right? So, at times, you can at least go back the material and then look at it. But again, with your session, I think if you are missing anything, probably, it's very hard for us to get that information, because you are [02:11:24] Deepan Kanagaraj: giving a lot of practical flavors to it, right? Which we may not be getting it when we look at the slides. So, that's why, if. [02:11:32] Mayukh Ghosh: Yeah, because you have to recall the conversation rather than the slides, right? [02:11:36] Deepan Kanagaraj: Yeah. [02:11:37] Mayukh Ghosh: That is there always. Otherwise, though, I would have shared slides, and we would have been all doing it without coming here. [02:11:45] Deepan Kanagaraj: Right. [02:11:45] Mayukh Ghosh: That's the difference, yeah. [02:11:47] Deepan Kanagaraj: Okay, so no one knows who is going to present on that particular week. [02:11:51] Mayukh Ghosh: No, right now, I don't know, to be very honest, but yeah, I mean, at some point before, obviously, we'll know. [02:11:57] Mayukh Ghosh: But again, right now, no. I think that's too early to call, yeah. [02:12:01] Deepan Kanagaraj: Okay. [02:12:02] Chandrasekhar Sahu: Tomorrow also, you'll be going to take the session, Mike? [02:12:05] Mayukh Ghosh: No, no, no, not tomorrow. This is the master classes, I think, only for today, for this part, two hours, yeah. [02:12:09] Chandrasekhar Sahu: Yep. [02:12:16] Mayukh Ghosh: Okay, any other questions from not only what I did today, but around ML or anything? Also, please feel free. [02:12:33] Deepak Katara: Premier Swan, if you have any doubts, how do we reach out to you? [02:12:38] Mayukh Ghosh: You can get in touch with Simran, she'll… Okay. [02:12:42] Deepak Katara: Sure. [02:12:44] Chandrasekhar Sahu: Hey, I have a couple of, questions to… Again, I wish to manager. [02:12:51] Mayukh Ghosh: Okay, that you can… I would say that… that park it, that you can do later, but let's first try to finish off if anyone has any other questions in the session-wise, yeah. [02:12:59] Chandrasekhar Sahu: Yeah. [02:13:07] Mayukh Ghosh: Seems like no. [02:13:08] Mayukh Ghosh: Which is okay. [02:13:13] Mayukh Ghosh: Alright, then, in that case, I'll close from my side, okay? Hope this was useful, and hope you sort of enjoyed. That is very important to hear, you know. [02:13:21] Mayukh Ghosh: If you enjoy half the things are anyway done, you will figure it out. Otherwise, things become very difficult going forward. [02:13:26] Mayukh Ghosh: Okay then, alright, thank you. We'll catch up sometime later. [02:13:32] Chandrasekhar Sahu: Thank you, thank you, Maya. [02:13:35] Mayukh Ghosh: I'll have to stop the sharing, and… [02:13:48] Chandrasekhar Sahu: Hello?