# 03 2025-11-15 Data Pre-Processing and Similarity Measures course: Module 1 — Foundations of AI & ML module: Module-1-Foundations-AI-ML date: 2025-11-15 type: transcript video_url: https://personal-learn.armco.dev/files/_Recordings/Module-1-Foundations-AI-ML/03_2025-11-15_Data_Pre-Processing_and_Similarity_Measures.mp4 --- [00:08:19] Hi, Mayuk, we can start. [00:08:28] Okay, yeah, sure, alright. [00:08:39] Excuse me a minute, I'll turn the video on. [00:08:55] Alright. So, yeah, hi everyone, good evening. Um, so my name is Mayuk, as you can already see. [00:09:03] Um, I understand, uh, so you have just begun this program, and. [00:09:10] Uh, it is a progression to generative AI, which is the name of the program, so… Yeah. So what… so I'll give you a brief few seconds intro of myself before sort of getting into the topic. [00:09:21] Okay, so I have been… so I am basically from Calcutta, but I have been in Bangalore. [00:09:28] For over a decade. And, uh… During this entire duration, effectively, I have been doing what I'm doing right now, okay? [00:09:37] So I have been doing… training in different capacities, uh… in a B2C space, in a B2B space, a lot of corporate trainings as well. [00:09:47] And I have headed a program for one of the IAMs before. [00:09:53] Okay? And, uh, yeah, and another program for University of Chicago. [00:09:58] Uh, collaborated program. That was long back, 6, 7 years back. [00:10:01] So, over the last, say, 11 to 12 years, it has been. [00:10:05] pretty much on these lines. Um, so… Uh, given that, I… Also, along the lines, obviously, I have done a bit of consulting also, so… for various places, like, say, the Aditabila Group. [00:10:20] the dining group, many other places, many other… one or two multinationals as well. [00:10:25] Uh, consulting as in training, as well as a lot of project work for them, in terms of… in the space of data science, machine learning especially. [00:10:33] So, that has been more or less the experience. [00:10:38] Uh, so that's a brief of myself, and then, uh, getting into what we are supposed to do today, or we are going to do today as well. [00:10:47] is our understanding of EDA. Now, I know that you have gone through a data preprocessing session before. [00:10:52] Okay? And I have also got a sense of what you have done on that day. Uh, some sense of missing values, maybe outliers. [00:11:00] Um, I understand a bit of transformation scaling as well was part of that. [00:11:04] Uh, anything else? Can any of you confirm, maybe? [00:11:09] If you did any… if you cover any other major topic, because these things, I understand, has been done. [00:11:22] Uh, hello, Mike. Like, uh… [00:11:25] This class will be continued by you or Dr. Dosniwal, like… [00:11:26] Yeah. [00:11:32] No, so I'll be taking today's session, so it's more of a… so, hands-on masterclass kind of a session, okay? [00:11:38] So, I'll give you some ideas. So, if you look at the topic, what we are doing is EDA for ML, okay? [00:11:44] So a bit of idea of how EDA is used for machine learning. [00:11:47] Trying to bridge it a little bit, and then effectively understand the bottlenecks that we often face when you build models. [00:11:54] Okay, so that… that's more or less the idea. [00:11:56] Because it's a 2-hour session, right? So I think. [00:11:58] We'll have to be, uh, discrete as well as brief, and quite crisp in terms of covering these things. [00:12:06] Somoyoke last time Professor has started, uh, undercutting for data preprocessing. [00:12:09] Yeah. Right. [00:12:11] So it's going to continue such sessions, or how is the pattern of this session, or it's… [00:12:15] Yeah, so I will continue, so that's why I'm saying that I have… I have got a download that you… some… these are… these topics were done, right, if I'm not wrong. [00:12:24] Uh, and given that, I will continue along those lines, and I will also possibly introduce a few ideas more from, say, an industry perspective of dealing with EDF or ML. [00:12:34] If that helps you to understand. [00:12:36] So if I understand correctly, you will be conducting practical session, and again, theoretical session will be continued by Professor. [00:12:41] Yes, it's… see, again, it's not really, uh… practical or theoretical hard demarcation, I would say. Yes, we will try to do a bit of practical in the time we have. [00:12:53] But more about the ideas which you already had, okay, in the last session, and you are going to have. [00:12:58] But then again, also bridging certain gaps in terms of. [00:13:02] Yes, you can say practical from a perspective of how they are used in practical sense, or maybe the application, more from that side, I would say. [00:13:11] Okay, so today we will be only focusing on whatever we have learned till date, the practical thing. [00:13:15] And tomorrow… Mm-hmm. [00:13:16] More or less, more or less, yes. I think maybe along the way, one or two things will be introduced. Obviously, that's part of the discussion, right? We generally don't have a strict plan for that, but what you said is more or less a plan, yes. [00:13:22] Hmm. [00:13:26] So tomorrow, again, professor will be continuing last… [00:13:29] Exactly. So, as I think you can see in the chat, right, so it's a masterclass on EDA, effectively, right? From an ML perspective, so… Professor will continue in that pattern what has been done, or what is already going on. So that, that will go uninterrupted. [00:13:48] Okay, yeah. Anyone else has anything before we sort of get started? [00:13:53] Anything you want to ask, you want to add? [00:13:57] Hi, Mike, this is Lokesh here. [00:13:59] Yeah, is the neat tool needed, you know, like Google Collab? [00:14:00] Yeah. Ty Lokesh. [00:14:03] Like, we're going to do EDR right now, right? [00:14:08] Yes, is there any tool? [00:14:09] We are going to discuss it, yes. This is EDA, this is EDA. This is the EDA session, yes. [00:14:10] Okay. Okay, what are the tools required now to do this? [00:14:16] No, you are not going to do a lot. I might show you a bit of it, how it is done, okay, because I know that you have done Python before, right? [00:14:21] So I don't want you to go and do it with myself. That takes time, and your concentration generally is on that, rather than what we are discussing, right? [00:14:29] Uh, so even if I… even if whatever we are doing in terms of the hands-on part. [00:14:34] Then I'll show you some code, some results, and we'll discuss around the outputs, right? So that… that's the main, sort of. [00:14:40] focus, rather than you kind of doing right away here. [00:14:43] Okay, man, thanks, Phil. [00:14:48] Yeah, am I audible? [00:14:49] Yeah. Yeah, G2, you want to say something? Yes, is he what? [00:14:51] Yeah, from my understanding of previous class, uh, I am, like, other than EDA, the class was related to data preprocessing, like, filling missing value and data by… [00:15:04] buying something like that, yeah. [00:15:05] Hmm, hmm. Hmm. Yes. Yes, yes. [00:15:06] These were the… so it's not related to DED, right? EDA is after… [00:15:10] Doing all these things. [00:15:11] Yeah, you are strictly speaking, you are right, but the trouble is the industry these days has clubbed everything into EDA, okay? But you are right. [00:15:19] The preprocessing is subtly different from EDA. One is exploration, one is preparation, right? [00:15:24] Uh, one follows the other, rather than they go hand in hand. So yes, that is right. So you have done preprocessing, I understand missing values, outliers, possibly transformation scaling also you have discussed, right? [00:15:35] So, um, yes, but also EDA is intertwined with those concepts, right? Because you are exploring data to find those patterns, and then doing certain things on that. So in that say. [00:15:45] Since it is interspersed, so we will touch upon those. [00:15:50] And maybe, yeah, with the more focus on the core EDA, if I put it that way. [00:15:56] Okay, okay. Yeah, anyone else, quickly? [00:16:08] Alright, fine. So… What I would rather start off with is… very basic stuff of statistics and how to use it in ADA, because I think. [00:16:17] As you might have already figured out in the previous session and from your previous experience or academic. [00:16:23] background as well, it's fairly conventional to figure this out pretty early. [00:16:28] is that you need a… Fairly decent sort of, uh… Understanding of statistics in terms of, uh. [00:16:39] getting through to EDA, if I put it that way, okay? [00:16:43] So, say, for example, yeah, I will… so… So let's break that two hours we have, okay? So, I mean, again, strictly following it is difficult, but I'll try to. [00:16:53] Follow up on a few concepts, maybe on the first one hour or something. [00:16:56] And then we'll see some outputs, and then possibly look at some of the results in the second half, the later part of the session, okay? So, roughly speaking, that's how we would. [00:17:05] Rather put it. Okay, so, um, yeah, so we have statistics. Now, I'm sure that statistics you had some exposure, you have been… have done, right? [00:17:16] The basic descriptive statistics, like, say, central tendency measures and dispersion, skewness. [00:17:21] correlation, and then possibly, as you know, that once you go to machine learning, the first topic that always comes up is regression, right? Linear regression. [00:17:30] Now, linear regression is a progression of statistics, and linear regression is generally the first model. [00:17:36] Which people learn or build, or get exposed to. [00:17:39] And for that, any model, and being linear, being the first model, the EDA… the understanding of the EDA data preprocessing preparation, feature engineering, all these things club together. [00:17:49] is very important, right? Because 70% of the work typically, as has been stated empirically from the industry standard. [00:17:56] Uh, the time of a project generally goes to EDM, roughly, roughly 60-70%, okay? [00:18:02] And the reason for that is… The data… the more the dimension of the data, right, more statistically rigorous is your exploration of the data? [00:18:12] And hence, more work you have to do in terms of drawing, say, diagrams. [00:18:17] Possibly, uh, finding the missing values individually for each of them, finding the outliers individually for each of them. [00:18:23] And then, also, doing transformations if required. Scaling is a one-shot thing, but these things all. [00:18:29] become multiplicated. When the dimension in terms of columns, or in terms of variables, is more, right? [00:18:37] So that's why the… typically, we end up spending a lot of time in this. [00:18:41] Now, let me do one thing. Let me share the screen. [00:18:44] And I'll… I'll jot down a few things, I'll jot down a few notes here and there, and we'll discuss along those lines a little bit, okay? [00:18:51] Um, and… In that vein, we'll also try to bring out the. [00:18:58] Get into outliers and missing values which you have discussed, and even scaling, which you have discussed, right? We'll go to those things. [00:19:04] via this route. So let me share a blank Excel. [00:19:09] Uh, that should be fine for now. Okay. [00:19:19] Okay. Rather than blank, let me go to this. [00:19:23] Okay, let me quickly spend 5 minutes on this, okay? These are things you know, and you have already done, but this is one way of putting everything together. [00:19:32] So you can see I have put in a couple of columns of data here, okay? Salary and age of people. Very simple. 10, 20, I think, yeah, 20… Or the rows of data. [00:19:41] Saturday in, say, CTC, for example, of. People in Indian currency, you can say, lakhs. [00:19:46] and age of those people. So. Given this. [00:19:51] If you look at these values, you can ignore the bin range for now. [00:19:54] You look at this column A and B, okay, if I highlight them. [00:19:57] Uh, you see, it's going fairly straightforward if I look at it. Obviously, it's 20, so I can look at it manually. [00:20:03] But as I go along, I can see there is one person here, row number 18, who has… who is aged 30. [00:20:09] The salary is 75 plaques, if I put it in lakhs, right? [00:20:15] But there are other people in the group. Age of 35, who has a salary of 20 lakhs. [00:20:19] Uh, maybe someone is 28, is 14, 32 is 19. So there are people around age 30. [00:20:25] But none of the salaries anywhere remotely close to 75, right? [00:20:29] So, this, we understand, is possibly an anomaly, okay? Possibly a difference. [00:20:35] A different data point from the rest of them, right? [00:20:39] Now, the thing is, if I try to fit. [00:20:42] Quickly, these two things, okay? And you know all these things, I'm sure. [00:20:47] Max minimum… mean, median mode. standard deviation, obviously variance is a square of that. [00:20:54] And then quartile is, of course, your third quarter, first quartile, 25… percent range, and percentile is every percent range. Here, I have 99 percentile. If you can see the formula bar. [00:21:05] Which is coming here. So these are the basic things we get when we run a .describe on the Python, right? [00:21:12] We get all these results. Here, I'm getting an Excel using individual formulas, but you do a .describe on the dataset on Python, you get all of these, right? [00:21:19] So, if I see… in this case, from these numbers, for example, okay, if you just concentrate on this column G, for example. [00:21:28] Okay, and knowing fully well that there are 20 data points of. [00:21:31] among these two variables. What a… I mean, as a snapshot, as a first glance, looking at these, and the first sort of set of output you were getting, right? [00:21:40] What all things you can infer from this? about the, say, the distribution of your. [00:21:46] This is salary, right? Distribution of this. data, given that we have salary and age in this manner. [00:21:52] What do you think is, uh… What can I infer from… for the salary variable from this? [00:22:00] What all things, at a glance, without even thinking a lot. [00:22:13] So, uh, one thought is that if we remove the outlier, which is, like, a max. [00:22:18] Right, so that the data is… yeah, the data range will come into a very small range in that case. [00:22:19] the 75. [00:22:24] So that will give the uniformity of data means. [00:22:25] Mm-hmm, uh-huh. Mm-hmm, okay, okay. [00:22:28] So right now, given that we have the outliers included in this, that outlier, maybe, if I say one of them, right? [00:22:35] Now, how do you think that. That particular number, which is, remember, 1 out of 20, 5% of my data, if I am strictly speaking in terms of numbers, right? [00:22:44] 5% being an outlier. How do you think that is influencing all of these, right? From mean to, say. [00:22:51] percentile up to all of these 6, 7… Um, the descriptive statistic yardsticks. [00:23:04] Mm-hmm. [00:23:05] Because, um, because this number is quite high. Right, compared to that average. So, definitely, it is… Uh, influencing that… it is giving some… For a little bit, I would say that. [00:23:13] Different picture. [00:23:14] Hmm. Yeah, in fact, if I add on a little more, you can see the skewness is very high in the positive side, 4.11. [00:23:19] Carlos is okay, you don't have to infer a lot from it, and in the correlation, which I'm sure you know what is correlation, correlation is around 47%. [00:23:20] Mm-hmm. [00:23:26] Okay, which is not good, which is not a great correlation, to be honest. [00:23:29] Sorry, Mike, uh, uh… [00:23:31] I just want to introduce those terms, uh… [00:23:32] So you can see, if I… if I… [00:23:33] not aware of the stumb, literally. [00:23:39] As soon as, uh, Qt is, uh… [00:23:40] Which terms you're not aware of? [00:23:43] Yeah. [00:23:44] Scuerous kurtosis. And correlation you were aware of? [00:23:47] Partially, not that much, yeah. [00:23:49] Partially, okay, okay. Oh, alright, no problem. I'll, I'll, uh, we'll not go into courtesies right now, it's not required right now, but skewness is important, correlation is important, as I said. [00:23:58] So skewness, basically, so you'll see. Uh, a distribution example. For example, this is a distribution, right? If I say I… select all of these 20 rows. [00:24:09] And if I put a line diagram through them, or any diagram for that matter, through them. [00:24:13] And then what I try to understand in that pattern is basically the distribution, I call it in the terminology we use, it's called the distribution of the data. [00:24:22] Right? You agree with me on this? That this is how we… pan it out, or you say that, right? [00:24:27] So it's a distribution. Now, distributions can be of two types. [00:24:28] Yep. [00:24:31] The first… there are two types… broad types of distributions. Everything fits into it. [00:24:36] First is symmetric. The other is asymmetric. Now. [00:24:41] Are you aware of normal distributions? As a… from a statistical perspective, normal distribution. [00:24:48] Continuous normal distributions? You are, right? So, if I go back to, say, normal distribution, everyone is aware I can… I heard only one person say yes. [00:25:00] Nope. [00:25:03] No. Okay. [00:25:05] Sir, you can explain my feel better. [00:25:07] Yeah. Right, fine, because see, again, let me be fairly frank with you, I mean. [00:25:14] Normal distribution, if you're not aware, or if you are, I mean, vague about it, right, or vague about. [00:25:19] things like correlation and skewness, right? Doing easier is hazardous, it's risky, to be honest, okay? Because you will then misinterpret a lot of things, to be fair, okay? [00:25:28] The basic paradigm of statistics is required to do UDA meaningfully. It's not a… see, EDA is not about drawing diagrams or. [00:25:34] just finding some reasons to do. imputations of missing values on outliers, right? In reality, it's much more comprehensive than that. [00:25:42] So we'll get to that. So, yeah, I mean, the more we discuss, we understand that where the starting point has to be, right? [00:25:47] So, normal distribution is a distribution where we say. [00:25:51] in a basic sense, I'm going just very fast and briefly into it, because it's a whole lot of statistics behind it. [00:25:58] Um, it's a symmetric distribution where mean equal to median equal to mode. [00:26:04] Okay, in a normal distribution, approximately at least. Uh, these values have to be equal. [00:26:08] What it means is that it's called a bell-shaped curve. I go here, I'll show you this diagram, okay? [00:26:14] Don't look at this Kurtik and all that, that's not right now relevant for you. [00:26:18] Look at these three… this… especially the red graph, for example, okay? [00:26:23] This red graph looks what? It looks that there is a nice peak in the middle, and it is flowing equally on both sides. [00:26:29] Can I say that from this diagram? And we have a term for that called bell-shaped. I'm sure some of you have heard of it. [00:26:34] Yep. Yeah. [00:26:35] Bell shift curve, right? A normal distribution is a bell-shift curve. [00:26:39] And obviously, all these 3 are bell-shaped in some sense, right? Because their tails are together similar in both sides. [00:26:46] But do they look same? Then no. Because their heights are different, their spread is different, but they're all normal. [00:26:52] Because their mean median mode is generally central, they're here. [00:26:56] Okay. So, typically, we say that when the distribution is normal, the mean, median, mode are going to be approximately same. I mean, theoretically, they say it will be same, but in practice for a long data, I mean, you can't expect them to be exactly the same. [00:27:09] Approximately within 5% of each other, we are pretty happy about it. [00:27:13] So that is a normal distribution. Along with that, for your reading, I'll tell you a couple of things also. Right now, obviously, in one and a half, two hours, I cannot go into stats and go into all of that, to be honest. [00:27:23] But I can tell you, you can just note it down. [00:27:25] Um… you need to also go and revise a little bit of probability theory. [00:27:31] Okay. Basic probability, don't go into axiomatic. Basic probability theory. [00:27:35] And the distributed distributions, that is very important. If you have to understand EDL, you need to know probability distribution. There is no second thoughts about it. [00:27:43] Okay, so priority distributions, there are, again, 20, 15 distributions are there in the stats book. You don't have to go through all of them. [00:27:50] But what you… bare minimum you must know is binomial distribution and normal distribution. [00:27:55] Binomial is a discrete distribution, normal is a continuous distribution. [00:27:59] I'm sure all of you know this difference between discrete and continuous data, right? [00:28:05] Yep. So, right, so binomial is discrete, normal is continuous. Now, why you need to know binomial, why it is so critical. [00:28:13] is that when you go to logistic regression later on, which I'm sure is part of your program, right? [00:28:19] Logistical question? Has to be, right? It's the backbone of data science, right? Huh. [00:28:20] Yes. [00:28:21] Yes, it is. [00:28:22] Yeah, yeah, yeah, yeah. It is, it is. [00:28:24] So, logistic regression actually follows a binomial distribution. So, unless you know binomial, and a term called odds along with that, probability and odds are sort of hand-in-hand, twin brothers, we call them in stats. [00:28:38] So, odds and binomial distribution is critical to understanding logistic. [00:28:43] And so that's why you need to know binomial distribution. Normal distribution, you need to know for anything. You need… even if you need to do EDM meaningfully, you need to know normal distribution. There is. [00:28:51] Um, I mean, yeah, it's very hard, I mean, without knowing this, to be honest. [00:28:56] Okay, so… that is… in brief, that is normal distribution, but you can just… whatever term I said, you can just go up and have a look up. It should not take a lot of time. It's not very complicated, okay? [00:29:07] So, uh, yeah, so this is mean, median motor, same, they are bell-shipped and all, and uh… Did you have any exposure to Hypothesis testing yet? [00:29:18] Not yet, right? It will happen, I am. fairly sure, but yeah. [00:29:22] So, when you learn hypothesis testing also, let me just give you a brief of it, because see, again, at the end of the day, all these things are connected. [00:29:29] Okay. Now… A subject has been built over 200 years with certain connections. That's the 200-year-old subject. [00:29:37] Now, for our sake, we can… we can unwrap them as… we wish, but there are certain connectivities we can't avoid, right? I mean, that's a… if we do that, that's disrespecting the subject in the first place, which is not fair, not right also. [00:29:51] So, when we do hypothesis testing, it should come from… I am sure that other day I was listening that you had. [00:29:57] a brief idea of sampling, right, that was covered in your previous session, right? [00:30:02] how to draw samples. So, probability, then probability distribution. [00:30:06] Then the ways of sampling, and then estimation, then hypothesis testing. [00:30:11] These are basically your five pillars of inferential stats. [00:30:15] Which you must know. If you have to do machine learning properly. [00:30:20] Now, properly is an important word here. There are hundreds and thousands of people doing machine learning in the industry. [00:30:26] Very few do it properly. So, it's important if you learn this way to do it properly, and I would suggest, and because I'm interacting with you here. [00:30:34] That… do it properly. There is… there are no shortcuts in data science, don't take any of them, okay? It doesn't lead you anywhere. [00:30:41] So that's a… More for what you should do next up when you kind of, uh… Do it on your own, a bit of self-study as well. [00:30:51] Okay, let me go back to what we were discussing. [00:30:54] Now, this is skewness is what? So, we are talking about there are two types of data. From there, we get into normal distribution, and we got into all this. [00:30:59] So, skewness is there… there are two types of data. [00:31:02] Symmetric? Asymmetric. What I just showed you in a diagram is a symmetric distribution. A normally symmetric. Why it's called symmetric? [00:31:10] It's symmetry in both sides. It's bell-shaped. Both sides have equally 50-50. In the middle, I'm cutting it at 50%. [00:31:16] And that is where my mean, median mode are. [00:31:19] Left-hand side is 50%, right-hand side is 50%. That's why it's called symmetry. Anything that shifts from symmetry. [00:31:26] is asymmetry. How does it shift from symmetry? Let's go here. [00:31:29] say this… both sides are not equal. Say, the right tail is longer. [00:31:34] Or the left tail is longer, okay? Then it becomes negatively skewed or positively skewed. Now, what does a left-tail longer mean? A left-tail longer means a lot of values are in the lower end. [00:31:44] only a few values at the higher end, okay? And the right tail longer means the opposite, right? [00:31:48] So, they… these are scenarios where you can say skewness. In this case. [00:31:53] My skewness is positive, plus 4. Why it's plus 4? Because everything is around 10, 15, 20. [00:31:59] One value, it's stretching to 75. And that's why it is plus 4. If I make this 75 as minus 75. [00:32:07] Then this becomes negative. It goes to the other side, right? [00:32:10] So that is skewness. Whenever it is not symmetric. [00:32:13] It is cute, then we have to understand whether it's negatively skewed or positively skewed. And how will I get that? [00:32:18] from the distribution. Make sense so far? [00:32:24] Yes. [00:32:25] Yep. [00:32:26] So, one question is that, um… Which means that, uh, just in an ideal case, if the… it is a completely bell curve, right, symmetric. [00:32:32] The scune is near to zero, is that correct, understanding? [00:32:34] Skewness is actually, yeah, theoretically we say, when it is exactly normal, skew units is exactly zero. [00:32:39] So, what we're saying that? Uh, you know, in the better cases. [00:32:44] It should be near to zero, like, um… If something like a plus 5. [00:32:45] Yes, you're right. You're right. The lower the skewness, the lower the skewness is, right, the lower the skewness is, I can say my data is more. [00:32:50] Oh. [00:32:51] Minus 0.5. Because that is better. [00:32:55] The range is slower, and it is well distributed, equidistant from each other, right? [00:33:01] So, as lower skewness means my data is under control, if I say. [00:33:02] Yeah. [00:33:05] But in reality, when you work in the industry. [00:33:08] in any industry you work in, right? Unfortunately, most data are skewed. [00:33:10] Yeah. [00:33:13] You will not face a lot of normal data, because otherwise you have to make it, doctor it, or tailor it, right? That's not going to happen. [00:33:14] Correct. [00:33:20] So most… 80% of the variables you see in real-time variable data, right? [00:33:24] And. So it is par from zero. [00:33:25] They are going to be skewed, okay? So, that is, uh… Ha, skewness will be far from zero, yeah. At least more than plus 1 or minus one, either side. [00:33:30] It is far from zero. [00:33:35] Okay. [00:33:37] So that's what we generally see, generally expect in real-time models or data, because that is real-time field data, right? Say you go for a variable called whatever I have spent in advertising in the last one year. [00:33:47] whatever I've spent in marketing in the last one year. If I try to find that data, and if I put it in an Excel sheet, or take it from a database, right, whatever it is, SQL, Oracle, wherever. [00:33:56] And then I put it in, it will rarely ever be a skewed data… sorry, symmetric data. It will always be skewed. [00:34:03] Either I've done less or more, in some sense, right? [00:34:06] So, generally, we expect it to be skewed in some direction, at least to an extent. [00:34:11] And… and that is where. it becomes a bit challenging for us. [00:34:12] Got it. [00:34:17] Because the more it is towards normality, the less is our work. [00:34:18] Hmm. [00:34:21] Okay, so we, as data scientists, or, I mean, if I put myself in the shoe of a model builder, that's what I want to do here, and that's the experience I want to share, right? [00:34:30] Our mindset is that we want. as many variables as possible to be normal. [00:34:36] To be symmetric. To which the skeweness to be very small, 0.05, 0.1, something like that. [00:34:43] But the unfortunate truth, as I said, is that that doesn't happen. [00:34:46] So then, there's a temptation across data scientists in the community. [00:34:50] to do a lot of transformations, log transformation, that is here, there. [00:34:54] To make it normal. Okay, to make it closer to normal. [00:34:58] But that is something you should be very, very wary of. [00:34:59] Alright. [00:35:02] I'll come to that why. That is actually, you are tampering with data. [00:35:07] There is a very fine… there is a very fine line between manipulating and tampering. [00:35:08] Okay. [00:35:13] Hmm. [00:35:14] And with less… more… I mean, when we're amateur rich, right, when we are starting off, because… 10 years back, I was also starting off, right? I was exposed to these topics. [00:35:22] And I know, because I've traveled the path, that this is very… a very fine line. [00:35:26] And in your initial few projects, you will actually trade over the line. [00:35:31] in order to get the right results, get a high accuracy, because that's what plays in our mind usually, right, for a model. [00:35:36] So, that is not… Huh, yeah. [00:35:37] So, how will we… like, how will we know that, you know, we are either, you know, we are trans… I mean, um… Uh, tampering the data, or we are, you know, manipulating the data and all, because. [00:35:49] in order, you know… yeah, go ahead. [00:35:50] Yes, so that… that… that comes with experience, there is no shortcut to it, but I can tell you some kind of hacks to it, okay, whatever I have seen over. [00:35:58] Yeah, a ticket or so. You are manipulating when you are changing the distribution. [00:36:05] For good, okay? What I mean by for good and not foregoing logic, that is very important. [00:36:11] Now, the example I was giving you a little while back is that you want to make a… skewed data, highly skewed data, like this, skewness is 4, 5, 10, something. [00:36:21] And you want to make it normal, which effectively means you want to bring down the skewness to zero, right? In other words. [00:36:27] So, if I want to do that, then what I do? I can say that, okay, here I have 20 data points, not a big deal, I'll throw this one away. [00:36:33] everything will become skewed, right? But if in place of 20 data point, I have 2 lakh data points, or 20 lakh data points, for example. [00:36:40] Now, it's not possible for me to find these people, right? [00:36:43] in 20 lakh, 30 lakh data points. So… There can be many, hundreds and thousands of people like this who are away from the rest, right? [00:36:50] So if that happens, then I have to understand that what I do is that any operation I do in the EDA, anything. [00:36:56] transformation, any kind of operation, right, imputation. Even your missing value outline imputation as well. [00:37:02] You check the distribution before you do it, you check the distribution after you do it. [00:37:07] You… the two golden rules are, you should not be changing the distribution by much. [00:37:13] And even if you are changing it by whatever fair amount, a little bit, maybe, maybe 20-30% from the original, shifting it, right? [00:37:19] It has to be in the right direction. Now, what is right direction? [00:37:23] towards normality. It should not be in the other direction that you make it more skewed. [00:37:26] Great. So, first thing is that it has to be towards normality, but not by lot. [00:37:32] If it is by lot, it's tampering. If it is by… if it is going the other direction, it is wrong analysis, wrong logic. [00:37:37] So you have to avoid both, and you have to be in the zone to a little bit. [00:37:42] So, and we need to, you know, employ the correct data cleansing techniques and all, right? I mean, all of those, the data cleansing, transformation, and everything, you know, in that, and then the data will be ready for ADA, right? [00:37:46] Yeah, yeah, yeah, 100%. See, EDA, absolutely, absolutely. And that… that is a… Absolutely, because that is the reason why this takes a lot of time. [00:37:57] And this is… we sometimes, these days, I hear that because we have so much innovation in this field. [00:38:02] People say that, ah, EDA will. Who will do, it's easy, and anyone can do. They'll put an intern to do EDI. I went to a company a couple of weeks back. [00:38:10] And it's an analytics company, they do pure play analytics, okay? And there were two interns hired from I Am Bangalore, definitely with pedigree. I understand all of that. [00:38:19] But it's an EDI acre. Now, the trouble is that. [00:38:24] is the trouble, because if, if… Amateurs only are involved in ADA. [00:38:29] Then whoever, if you think that you are reserving the meat of the model building and insights work for the. [00:38:35] senior people, then they will itself make it wrong, because the area is wrong, or idea is dicey. [00:38:40] It's not unsure idea, right? So, this is very, very critical. You need a lot of exposure experience to do it well. [00:38:47] That's why you are called as data scientist. A scientist word. [00:38:48] Understood, yeah. Thank you. [00:38:51] lot of, uh… [00:38:52] Yeah, I don't have enough time today. If I would have got more time, I would have actually proven in this class that it is more art than science, this data science. [00:39:00] Okay, but yeah, at some other time, we'll do that. [00:39:01] Mike, uh, can we… can we… [00:39:03] Yeah. [00:39:04] start the course, because the discussion somewhere, it's not a digestible for me, yeah. [00:39:09] Yeah. Okay. [00:39:10] This is… this is part of the… this is part of the course, Chandrasekhar. This is part of EDA, what we are already doing, actually. [00:39:12] Okie, okay, fine, fine, yeah. [00:39:16] Because this… you can't do EDA without knowing this. Okay, let me be very honest with you. [00:39:17] Okay. [00:39:19] Okay, yeah, yeah, yeah. [00:39:20] Okay, so you can't say that I will do EDF without knowing QNS or without understanding. [00:39:23] No, no, no, not this… not this one. Means, uh, the content of this, yeah. [00:39:28] Yeah, yeah, yeah, yeah. [00:39:29] Ah. I'll tell you, that part you were saying, Ha, that we just, yeah, that's more offshoot of the question he was asking, yeah. [00:39:31] Yeah, yes, yes, yes, yes. [00:39:35] No, no, fine, we'll come back to this, of course, yeah. [00:39:36] So, yeah, yeah. [00:39:38] So, the other part is correlation. As I think someone was saying that your vague idea. [00:39:43] Uh, but anyone here who has… yeah, Sridh, go on. [00:39:46] Um, I had a question, so… [00:39:49] Based on your experience, like, seeing different data, how do you know that, or how do you judge in an early stage that you've just got the wrong data? [00:39:59] Like, the skewness might be, uh, deviating from zero, like, obviously, [00:40:05] No. [00:40:06] Not C, it is… see, wrong data. I don't think you mean wrong data, you mean wrong variable, right? [00:40:08] Yeah, I mean, yes, yes, that… I mean, everything just looks right. Let's say you've got a skewness of 4, it's a normal distribution, it looks… [00:40:14] Hmm… hmm. [00:40:17] But it just doesn't fit the model, or your model is just shit, right? [00:40:22] So, that is happening, but that basically means your data is wrong, right? Or whatever you have collected, you've gone wrong. [00:40:27] Yeah… yeah, see, wrong data. It's a very good question you ask, even if the articulation can be better, but I understand what you say. [00:40:36] Yes, yes, yes. I didn't know how to frame it, yeah. [00:40:40] Yeah. Yeah. [00:40:43] Yes. [00:40:44] But the question is good. No, no, I know, I know where you're coming from, so I understand what you mean, actually, even though you say a little bit different, but. [00:40:45] Uh, see, wrong data? No, but I know what you're saying. [00:40:49] Uh, again, just spending a little bit, because I think this is a very important point you raise. [00:40:55] When you see a data, you need to understand that. [00:41:01] That what problem statement I'm trying to solve, even before getting into ETM, right? [00:41:05] Just because you have a… access to Python or Colab or whatever, right, Jupyter. [00:41:11] And that is free. You don't have to pay… Anaconda, any money to run codes. [00:41:15] That doesn't necessarily mean that you will just open it and writing blah blah blah, anything, okay? [00:41:20] So, that is the last part. You have to devise a plan, you have to understand the data, you have to understand the variables, see whether it is worthwhile even investing time into it. [00:41:30] Then do a POC, then get it approved by the stakeholder, client, be it internal, be it external. [00:41:34] Uh, do a pilot sometimes. Then you go and start your first line of loading data into the Python environment, okay? [00:41:43] And before all of that, it is very important that you know what problem statement you are solving. [00:41:48] What variable is your target variable, main variable you want to predict or figure out or forecast? [00:41:53] And then from that light, once these two things are clear to you, from that light, you look at the other variables in the dataset. [00:42:00] And see, do they make sense? Are there any relation itself? If there is no relation, some random variables are clapped into from the SQL database. [00:42:08] And, yeah, some client or some team in my company, in a product company, for example. [00:42:13] They just say that, okay, do some analysis for this. I mean, that's very vague, right? To do some analysis, I'll draw a bar graph and send it back to you. That is also analysis, right? [00:42:20] So, that's another point. You have to be very clear about it, and you have to see whether it's possible to do something. [00:42:25] And that's where I am sure you are bringing up this wrong data part. [00:42:27] Yes, yes. [00:42:28] It's not about wrong or right data, it's about whether what you are trying to do is possible on this or not. [00:42:32] Got it. Right. [00:42:33] Right. So that is something where, actually, the skewness we have been discussing so far, right, skewness, normality. [00:42:40] These are very important concepts into even assess that. [00:42:42] And even correlation, which is the next thing. So I'll actually take your question to this next concept anyway. [00:42:49] Oh, cool. [00:42:50] And then we'll come back to this, yeah. So, uh, correlation, I'm sure some of you have an idea, as you said, vaguely. [00:42:55] So, correlation basically means. that you have, uh… [00:43:04] Relation, right? I mean, if you ask anyone, even the seasoned professionals, that, okay, define correlation. [00:43:08] They will say relation among two or more variables. [00:43:11] Yes, there is nothing wrong in that, but it is not only that. You have to be a little more specific or precise about what you mean by relation here, right? [00:43:18] It is actually, if you go by… if I remember what I learned in stats. [00:43:22] 15, 20 years back. It is the terminology is the degree of association between two variables. [00:43:29] Okay, the classic example we used to have in our textbooks in stats. [00:43:33] Is, uh, height and weight of people. Okay. There was a different example, which is height and IQ of people, and then short people, especially a lot of girls, protested, and I agree, it was a very bad example. [00:43:45] So we… I mean, even the book in the next edition changed it. I don't know why they kept it in the first place. Uh, so we… It's height and weight of people. Now, if I go to this session now, right now, how many people are there? [00:43:56] 93 people, including me and all, okay? Sir, said around 100 people. [00:44:00] So I… what I do is that I take down the height of 100 people. [00:44:03] I take down the weight of 100 people, okay? I put two columns, like I have put here, hide here, weight here. [00:44:09] Then what I want to understand is that can I find, as a gross level, can I have one single number? [00:44:16] Which will tell me that. a higher height indicates more weight. [00:44:21] or vice versa, a heavier person. Generally, someone who is tall. [00:44:27] Right? That is what I'm trying to find. So that's why it's called degree of association. I am not interested about our height. [00:44:34] Neither I am interested about or wait. I don't care. [00:44:38] Your fat, thin, tall, short, I don't simply care. [00:44:41] You are just one data point, one entry of the hypothesis I'm trying to prove. [00:44:46] Which is, height and weight have a positive relationship. [00:44:49] Or an inverse relationship, or no relationship. There can be 3 things here, effectively, okay? [00:44:55] And this thing we are calling as correlation. Correlation between height and weight. [00:45:00] Now, if the… so the correlation coefficient, there's a way to. [00:45:03] I mean, there is a concept called covariance. I don't want to get there, I mean, then it becomes a stats class, I will not do that. [00:45:09] You can look up later on. There is a concept called covariance, which is used to understand or figure out this correlation formula, okay? [00:45:15] Anyway, however we get it is relation among these two variables. [00:45:19] clubbing them row by row. So, the correlation coefficient of the value lies between minus 1 and plus 1. [00:45:27] The closer it is to plus 1, that means the correlation is strong. [00:45:30] very, very strong, if it is 0.8 and 0.9, so on and so forth. [00:45:34] On the other side, if it is minus 0.7, 0.8 something, then it is very strong inversely. [00:45:40] But if it is hovering around zero. And be it the negative side, be it the positive side, but not too high, 0.123, something like that. [00:45:48] Then there is no correlation. It's a weak correlation, okay? [00:45:52] That's how we define it. Now, we have to add a bit of, uh… we have to… I mean, now sort of put it in the application sense on a dataset, which I'll do. [00:46:02] But first, let me know whether. The basic part of what I just said of correlation is clear to everyone or not. [00:46:13] Yes? My… my pace is fine, I'm not going too fast, too slow, hopefully. [00:46:15] Yes, ma'am. [00:46:16] Yep, yep. [00:46:17] Yes. [00:46:19] All good? Good. Yeah. And I hope you are finding it interesting, enjoying so far? [00:46:20] Should be a good game. [00:46:22] Just… [00:46:26] Yeah, yeah. [00:46:28] Yeah. [00:46:29] That is interesting, that is important, otherwise there is no point, to be honest, if you ask me, yeah. [00:46:30] Yep, yep. [00:46:32] So, uh, so correlation, as I said, negative, positive, let's take an example. Positive, I said, height and weight, we generally expect that, okay, taller guys will have more, sort of. [00:46:42] More weight and more body weight and all that, that, that's positive generally, right? [00:46:46] Uh, negative, say, for example, Deepak, I'll come to you. [00:46:50] Negative, for example, uh… say I take a couple of variables called. [00:46:55] say, uh… Okay, let's take the present day Delhi as an example, okay? [00:47:02] Air quality and pollution. Okay. So, uh, if I have variable A as air quality, and variable B as pollution. [00:47:11] Or whatever, uh, noise, right, in terms of that. [00:47:14] And if I try to find a correlation between these two. [00:47:18] You expect them to be negative? Do you agree with me on that? Air quality, high pollution, low, pollution, high air quality low, right? [00:47:26] Yeah. [00:47:28] Yes. [00:47:29] Yep, so that is where I expect a strong negative correlation. So I have to make sense of this correlation value I get with the variables, how they mean, right? [00:47:35] If I see that the value is something and. [00:47:37] Uh, the variables mean something else. Then, of course, I mean, there is something… yeah, if I go back to the question, I think… Uh, you were asking before, right, Srijat, right? If I got your name correctly, wrong data, right? [00:47:50] Yes, that's… [00:47:51] That is our example of that. that I have a featureless data, I have a… biased data, which is made or built by. [00:48:00] with some agenda, some… something, some foresight, maybe, which is… I'm not even aware of, yeah. [00:48:04] So, it's just, like, it doesn't make sense. Whatever data… [00:48:06] It doesn't make sense, exactly. The variables mean something, the result I'm getting by using them together doesn't make sense, so they are not tallying together. [00:48:14] Got it, got it. [00:48:16] Yeah, meat, I have… I think before that, whoever was raising your hand, I think Deepak, or whoever… forgot… Yeah, dibang on. Me, I'll come to you. [00:48:29] Hmm. [00:48:30] Yeah, just one thing. Yeah, so I just wanted to understand, uh, real-time use case, uh, with coral coefficient, because I think what we have, uh, understood or learned in the previous classes, it comes with the dependent variable we try to figure out… we have something. [00:48:38] Yeah, yeah, yeah, I… yeah. I'll go to that, so I have not gone into regression, that's why I couldn't introduce dependent independent so far, but we'll go there, yeah. [00:48:47] Okay, yeah, sure, we'll take on… [00:48:48] Yeah. Yeah, meat, go on. [00:48:51] Yeah, um, my question might be… it might be meant for further… If you're further along in the class, I guess, but… Basically, let's say if we're building a model with, let's say. [00:48:58] Hmm. [00:49:04] Mm-hmm. [00:49:05] 12 features, and we have a label. Now, uh, what kind of, uh… sort of correlation between these features and. [00:49:10] a label would make sense, like… [00:49:15] Mm-hmm, yeah, exactly, yeah. [00:49:16] Yeah, yeah. Which is pretty similar to what Deepak, I think, was meaning as well. Yeah, it's very similar. We'll go there, we'll come to that, yes. In the next 10 minutes itself, I think we'll be there. [00:49:23] Yeah. Uh, so this is correlation. Now, one concept, which I'm sure you didn't discuss in the previous class, it's rarely discussed these days, and I mean, I'm not sure why, but we. [00:49:33] Generally, for lack of time, we don't go into that these days. [00:49:35] Uh, there's a term called spurious correlation, but I think it is very important you know this. [00:49:41] spurious correlation. If you Google up the term spurious, it means false. [00:49:45] Only the meaning of the term in English dictionary, okay? It's called false. [00:49:49] So it means false correlation. Now, what is false correlation? Okay. [00:49:53] I'll give you an example, okay? A real example. I'm not making it up. You go to New York Times, you look up on Google, you will see a lot of research being done in the US. [00:50:02] Right now, even for the last 5-6 years on this, okay? [00:50:05] And one of the examples is this. There was one person who was building a model, of course, where. [00:50:10] They were trying to understand some predictive features in predictive relationship. [00:50:15] In terms of… What makes ships go down in the Atlantic Ocean? [00:50:21] Okay. So, the target variable, or the variable they were trying to figure out, or understand, or explain. [00:50:27] Is that ships are going down or not? Classic logistic regression problem, obviously. [00:50:32] Yes, no kind of thing. Um, and then… That person has a lot of variables, weather variables, how wind pressure, this, that, some air pressure, all weather things, again, technical things. [00:50:44] To understand that how these things are impacting whether a ship is fine or the ship is in trouble, right? [00:50:50] What this person did is that. He looked up all the data over the last. [00:50:55] what, 50, 40, 50 years, and looked up… this was actually a PhD thesis in Yale University, okay? [00:51:01] 10, 12 years back. And, you know, PHCs have to be presented in front of a panel, right, and where they can be challenged. [00:51:09] So, he was showing this, and he was saying that, okay, I have looked into those days where the ships went down, okay? [00:51:15] And, uh, yes, indeed, the weather was bad. Overall, with all the variables effect club, I can. [00:51:20] conclude that the weather was bad and. Uh, wind and air pressure and all these things, sea pressure was all pretty bad. [00:51:26] So, the ships went down. So I can find a very strong relationship across these weather variables and my ship's. [00:51:32] fit, right? Then one person quickly stood up and said, like, in PhD happens in, if you go to the Ivy League universities, that's what happens, right? [00:51:41] They stood up and said, give me 15 days, I'll… I'll find a contrary result, okay? [00:51:48] And what this person did is that he didn't even look into the days this guy looked into. [00:51:52] He only looked into the days when no ships went down. Every ship sailed fine. [00:51:58] sailed nicely. In 46% of those days, the weather was worse. [00:52:04] in those parameters, the first guy defined, okay? And this is an example of Spurrier's correlation. [00:52:10] Spurious correlation happens when you yourself are subjective. subconsciously biased even before you build a model. [00:52:18] And that is the biggest problem we data scientists suffer from. [00:52:21] Because we are human beings, we are not machines. [00:52:25] So we will all have our incoherent preconceived notions and biases, right? [00:52:29] And that is where, if you see that some correlation is very strong. [00:52:34] But you feel that, like I was telling before to Srida as well. [00:52:37] But I feel that it is… looks unlikely. Double check. [00:52:42] A very good result doesn't necessarily mean a right result, okay? [00:52:45] And in data science, a correct result is more important than a good result. [00:52:49] every single day, in every single facet of data science. [00:52:52] You can get 99% accuracy. I don't care if it is not sustainable. [00:52:56] You give me a 75% accuracy, sustainable over 6 months in production, I'm happy. [00:53:01] You give me 95%, which is sustainable for 7 days, I don't care, I will not take it. [00:53:05] I will not pitch for it, I'll not take it, I'll not make it billable. [00:53:08] Simple, right? That's how it works. So, the spurious correlation effectively is a thing which you need to know, and because. [00:53:15] The other thing it tells us, as a corollary to it. [00:53:18] is that you need to be as unbiased as possible, even if you know that, okay, my sales are driven by advertisement, by marketing. [00:53:25] But you cannot bring that prior knowledge or whatever your ideas are. [00:53:30] When you are building the model, because then unconsciously or subconsciously, you will make the EDA drive towards your. [00:53:37] Beliefs, which is very often, very common, okay? So that is, uh, art, again, as I said before, it… I mean, you have to be conscious, and spurious correlation is a classic example of what we do, or what we get. [00:53:50] When we are not careful with our results, thinking that we already know what we will see. [00:53:51] Yes, yes. [00:53:56] Making sense? [00:53:58] Uh, yeah, but which one was correct? Like, the first one or the second one? [00:54:05] See, it's not about correct. The first one was wrong, right? [00:54:09] Because that guy, he only looked at the days the ships went down. [00:54:13] Right. [00:54:14] And try to conclude that whether bad should go down. [00:54:15] Mm-hmm. [00:54:16] The second guy tried to disprove it by saying that when the ships don't go down, weather sometimes can be very even worse. [00:54:21] So, can I make a conclusion? The answer is no. [00:54:22] Hmm. [00:54:23] No, correct, yeah. [00:54:25] That way. So I'm saying that the correlation was what the first conclusion was not conclusive, if I put it that way, right? [00:54:33] So that is the problem if you take subjective parts of data. If I take a biased sample of a data, this is what happens every single time. [00:54:39] Okay. So, I get a better result than I deserve, or than I should get. [00:54:44] Then I say that I go to my client, I'll tell you reality… in reality what happens. I'll go to my client, my stakeholder, in a meeting. [00:54:51] Where I'll make a 30-slide fantastic PP video on that, okay? [00:54:54] I see this every day, that's what I'm saying. [00:54:57] And then I pitch it up, my results are good. Generally, business people are not statistical, most of them, some of them are, but generally no. [00:55:04] They will say, yes, yes, your numbers are good, yeah, we can have a projected strategy, all these things. They generally are very upbeat about these numbers. [00:55:11] And you put it up in production. Production means you deploy the model, right? You deploy it, your data engineering team comes and deploys them in the CI-CD pipeline. I'm sure you have heard of these terms. [00:55:20] And you put it up on production, and you monitor it. You monitor it using your… Yeah, air cloud systems, airflow, and all this stuff, right? [00:55:29] Now, after… the next set of data comes in, the new set of data comes in for your variables. [00:55:37] you receive a shock. You receive a shock that your model has failed. [00:55:40] Your model has whatever results you gave is not working anymore. It is not true. It is not holding up for new datasets. [00:55:47] Then you go and say. give me some time, I'll have to bring it back to my environment, and I have to check. [00:55:54] Rebuild, your client will say, okay, what happened this, that, blah blah, some meetings up and down will happen. [00:56:01] And then you go back and check it. you are going to send it back to production after 1 month. [00:56:04] That's… remember, you are incurring cost lakhs and lakhs of rupees. [00:56:08] Because you are buying Azure AWS or a GCP, or a Databricks, at least. [00:56:13] So, you do this, you again send it back. [00:56:17] Then again, after 2 runs, it fails. Then your client comes back and says, boss, enough. I have spent 15 lakhs on this, 20 lakhs on this. [00:56:23] Your model is not giving me ROI. We are scrapping the model. [00:56:29] That's the fate of a model. in 3 to 6 months' time. [00:56:33] If you don't do the EDA. Without bias and without thoroughness. [00:56:41] That's the final thing you end up having. So, again, I emphasize. [00:56:49] With this example, again, this is very, very important. [00:56:52] And to do this well, understanding. Basic to intermediate statistics is even more important, otherwise you can't even do ED. [00:57:00] Let me be very honest. Makes sense? Are you able to connect what I'm saying right now? [00:57:08] Yeah, me to go on. Nothing is stupid, go on. [00:57:09] Oh, this might be a very stupid question, but uh… I just… so, as a human… as we all have some biases, right? Is there a way to, sort of. [00:57:18] Mm-hmm. [00:57:20] Quantify that bias in our results. [00:57:22] Yeah, agents haven't been able to do that. Agentic AI has failed there yet. [00:57:27] Yeah, that… That's why I said it might be a stupid question. [00:57:30] No. No, no, it's not a stupid question. See, the question is not stupid. [00:57:35] The thing is that the question doesn't have an answer. That doesn't make the question stupid, right? [00:57:40] So, uh, no, we can't quantify bias, and if… I'm not going to teach you regression, I don't even have the time. I would have loved to, but I don't have the time. [00:57:49] But this question brings me to the mind. The way we generally… I generally at least, start linear regression. [00:57:56] That… Can I quantify. [00:58:01] bias or human preference. If I can, every single day a linear regression equation, I'll fulfill it, completely fit it. [00:58:08] And I'll get 100% accuracy, which is also accurate. [00:58:09] Mm-hmm. [00:58:11] Just because a human mind cannot be biasless or errorless, right? [00:58:16] Then 100% accuracy in a model is not possible, actually, if I think of it from that perspective, what you're saying. [00:58:22] So, yes, no, we can't quantify that, and that is all, whatever we can't quantify, put it as error. We dump it as error. [00:58:23] Mm-hmm. [00:58:28] Basically. And that is the beauty of it, right? If we could have quantified, it'd become very dull, no? [00:58:31] Understood. [00:58:35] are true. [00:58:36] The thrill goes off, yeah. Okay, so this is correlation. Now, here, coming back to this example, right? [00:58:44] It is 46. Now. Notice something carefully. [00:58:47] this 75, I will now change it to 25, okay? I'll change the number. [00:58:54] I will change it to, sorry, 25. See the numbers, how they're changing. [00:59:00] I'll say enter. Correlation has increased from 46% to 83%. [00:59:06] I changed one number, which was… Bad, or… I mean, away from the rest. [00:59:12] Now, the model or the machine is saying. That salary and age are very strongly, positively correlated. Can I say that? [00:59:13] Got it, yeah. [00:59:23] Before, see one number, how it jeopardizes the entire thing. [00:59:27] Is that this strong coordination was reduced by about 40%, right? Half of it almost, right? [00:59:32] So, that is the problem if you have outliers in your data. [00:59:37] Does it make sense now that what potentially keeping outliers might do? [00:59:38] Yes. [00:59:42] Bad outliers, that is. How much it can actually change things? [00:59:46] Yeah, look at the skewness. From 4, it has become 1. [00:59:50] Look at the standard division. From 13, it has become 3. [00:59:55] And, initially, mean and median were 16, 17, and 13. [00:59:59] Now it is 14 and 13.5. What does it mean? It means it is a symmetric distribution, it is normal. [01:00:00] Yes. [01:00:05] They're approximately equal, can I say that? [01:00:10] Right? So I have… and also, evidence is there. Skewness is now very low. [01:00:11] Yeah. [01:00:15] From 4.1, it is 1.1. I've reduced it by 75%. [01:00:19] Right? So, see, I change one number, I change an outlier. [01:00:24] I… overall distribution of my data, I make it healthier. [01:00:27] Everything from central tendency to dispersion. 2 skewness, to correlation. [01:00:33] And because I can do all of this here. [01:00:36] when I do a regression with this, I will have the same effect. [01:00:40] And that is why… EDA, or finding these pre-processing methods properly is very, very important. [01:00:48] Hope that makes sense to everyone. You are getting a… flavor of how bad or how good. [01:00:53] Things can be by just minor proper tweaking. [01:00:58] Yeah. [01:01:00] Yep. So that's why everything is important. Don't look at that, okay, I have 10 lakh data points. [01:01:01] Yes. [01:01:04] Only 20 and 30 are bad. tour that they had no. [01:01:09] It is not about that. You only become a good data scientist if you are a perfectionist. [01:01:14] There are hundreds and thousands of data scientists in the world these days. Everyone is a data scientist, I mean… It's a very fancy, nice term, everyone wants to call themselves that, right? [01:01:22] And there is no harm in that, as long as you are doing justice to the role. [01:01:25] So, the problem is that that doesn't happen. People are rushing. [01:01:29] I learned today, I'll just digress for 20 seconds here, because this is a pet peeve of mine. [01:01:34] I rushed today, I start statistics today, and in 6 months' time, I'm building a… LLM model. It doesn't happen. [01:01:41] You can't. So you have to give it time. It's a serious subject, right? So you have to give it time. And you see one… how minutely and how precisely you have to look at everything that is being happening. [01:01:51] Yeah, go on. I think a couple of questions you have, right? [01:01:55] Yeah. Mm-hmm. [01:01:56] Yeah, so in such cases, when we have, uh, outliers, so what is our strategy around it? Because I think this is a very small data set, we can figure it out. [01:02:03] Yes, yeah, yeah, so I can do manually, right? But in general, I can't, right? Ha, I'll come to that. I'm just showing you a small sample to give you the… show the impact. [01:02:05] Yeah. [01:02:11] Because if I take a large data and show this, you will not see the impact, right? [01:02:14] Right. [01:02:15] I'll do… we'll discuss that, don't worry, yeah. Yep. [01:02:18] But we are playing with the data here, right? Is that right? [01:02:21] Yeah, that's a different question. We are playing with the data indeed. [01:02:30] Yeah. [01:02:31] Now, see, that is a different logic. As I told before, right, you should not be changing it for the sake of changing it, right? To get results good for yourself. [01:02:34] Here, I'm changing it. Why? I can say 25 is the arbitrary change. Obviously, I have to be a little more precise in making it whether to 25 or 30 or 35, right? [01:02:41] But I think there is enough evidence to change it, because 75 is an extreme outlier. [01:02:46] Okay. [01:02:47] If it is a borderline outlier, I would not have bothered. I will actually come to that topic also in the next couple of minutes. [01:02:51] Okay. [01:02:53] But in extreme outlier, yes, it's fine, because that is not making any sense in this case. [01:02:56] 30-year-old guy earning 75 lakhs in this cohort of people. [01:03:00] I don't need that guy. I will throw him away, or I'll keep it, yeah, in a different number. [01:03:01] Okay. You're just saying that, okay, it might be, uh, you know, data error also, that somebody would have [01:03:08] I'll put it by mistake, okay. [01:03:09] Exactly, it's an anomaly, it's anomaly. I am… see, it is… it is very important to understand why am I doing it. [01:03:15] Correct. [01:03:16] I am doing it not because I want the correlation and skewness to be better. [01:03:20] I am getting it as better, great. It helps me. But that is not the reason I did it. The reason I did it is that it's around it. [01:03:28] This has to be very clear in your mind. [01:03:30] You can't manipulate data to… for your own serving of doing nice work, short work. [01:03:38] Got it, okay. [01:03:39] Then it is… you are tampering. But if you have a right logic to do it, in the process, you get a better result, fine. [01:03:45] Yeah. [01:03:46] I think, you know, this is what I asked in the beginning and all. How do we decide that whether, you know, it's wrong data, or it's, you know. [01:03:52] Yeah, I know, you're wrong. Data is a very broad thing, I understand. This is also part of it. [01:03:53] So… Okay. [01:03:57] But yeah, see, wrong data in that sense, but I think also he was also mentioning, right? [01:04:03] It is… it's a complicated thing. I'll take one more example. I know I'm… yeah, we have just one RF just flied. I have only one more hour left. [01:04:12] And we are supposed to do a bit of Python as well, but see, let's see. Uh, I'll get to that. See, the more questions you ask, no, the more examples that come to mind, and which will take more time, unfortunately. [01:04:22] Uh, but okay, we'll talk on that as well. I'll come back to that, so I'll just, uh… Hold on to it, yeah. Um, so this is effectively what we are doing. [01:04:33] Anyone else has any question on this part at this point, precisely on this? Otherwise, I'll just move ahead. [01:04:41] Can you explain a bit, uh, or elaborate a bit on kurtosis? [01:04:46] Kurtosis. Kurtosis is not that… I mean, that important right now, so if you go to this diagram, you can see actually some names are given. [01:04:54] Kurtosis is how picked your diagram is. If it is in the positive courtesy, it means it's leptocardi. It's very sharp. [01:05:00] The peak is very sharp at the top, okay? [01:05:03] But you… you are… your spread is slow. You see, the spread here is low, right? [01:05:07] But in a mesocortic or a normal distribution, it's equal. You are not essentially putting a high peak. [01:05:13] Now, what does a high peak mean here? If I have a high peak here means what? That in the central part of it, I have plenty of data points. [01:05:19] Right? And then it is just… Just putting down a kind of… Falling down both sides very sharply, right? [01:05:27] Same thing in a mesocortic or a normal distribution, it is zero. [01:05:30] There is no kurtosis. The entire distribution, there is no peakedness. It's all equidistant. So, this data point to this data point, the slope. [01:05:38] Whatever the slope is, from this data point to the next data point, the slope is exactly the same. [01:05:43] When that happens, I know I'm getting into linear algebra and all that, coordinate geometry, rather. [01:05:48] But that is how we call, or we understand mesocortic, or normal distributions. [01:05:53] And platicardic is very flat, which looks like a plateau. [01:05:56] And leptococardial positive cortis, that is, is where it is very sharp. So cortosis, you will not have to do a lot of applications of it because this is more from statistical quality control. [01:06:05] So there is a topic in statistics, which is called SQC, or Statistical quality control. Anyone from stats background here will know this. [01:06:12] Uh, in quality control and in design of experiments, we use this a lot to understand distributional patterns in data, okay? [01:06:19] In your machine learning, you will not need to go into that depth, really. Before that, you will start deep learning, so don't worry. [01:06:27] Yeah. So, haha, so this is what, uh, potentially changing this and outlier. [01:06:34] actually can have an impact. Now, one other concept I would want to speak in the outlier sense. [01:06:39] Which I'm sure, again, maybe many of you are not aware. [01:06:42] is a concept, terminology I have given myself, it's my invention, but a lot of people know this, they use different terms. [01:06:48] is a concept, what I want to introduce as bivariate outliers. [01:06:53] Okay. Before that, I have a quick question for some of you. [01:06:58] Um, are you all aware of heat map and scatterplot? [01:06:59] Yes. [01:07:05] Both of it, you know, in Python, you have done, right? So you know heatmap and scatterplot, right? [01:07:10] Now, I have a question for you. If you know heat map and scatterplot, think of it and let me know. [01:07:16] Now, you know, heatmap gives you correlations and all, the combinations and everything, right? So… That is there. Scatterplot also gives you. [01:07:23] two variable relationship and correlation. Now, my question is, I often tend to ask people in interviews as well this question. [01:07:31] Uh, if I draw a heat map. I have already drawn a heat map while I'm doing EDA, okay? [01:07:37] Do you think, after that, I need to draw a scatterplot as well? Or, in other words. [01:07:43] Despite having the heatmap results. Do you think drawing a scatter plot will give me anything extra? [01:08:00] Maybe more readability of the data? [01:08:03] Yes, but do you think from a data handling or data… I mean, taking the data forward, or what is the next step? Do you think we get more insights as such, anything? [01:08:13] I think so, yes. [01:08:14] Or do you think everything is covered? [01:08:16] Anything on the extreme can be seen in this catalog. [01:08:21] Yeah, most of you are correct, actually. You're on the zone, exactly not hitting it, but you're around the zone. [01:08:26] See, that is what I was saying as bivariate outliers. [01:08:29] I go back and I change this guy again to 75. [01:08:35] Take care. And I had a few more data points here, okay? Couple of data points. [01:08:38] This is 50 lakhs, this is someone whose age is 50. [01:08:43] Obviously, can earn 50 lakhs. This is someone who is earning 62 lakhs. [01:08:47] It is… 50. Okay. [01:08:52] Now, if I have this data. I look at salary. Now, you… all of you guys have discussed outliers in the last session, right? [01:08:58] Quickly tell me any one of you, or any two of you, how do you check for outliers? [01:09:04] What are the different processes you… Do, diagrammatic or not, to find outliers. [01:09:09] Usually IQR. IQR. [01:09:10] Boxplot? [01:09:11] Uh, distance from here? [01:09:12] Boxplot, I heard. Anything else? Goes along with it. Anything else? [01:09:17] box product and IQ? Distance… distance from mean, okay. [01:09:18] cleaning method. [01:09:19] Distance from mean. [01:09:23] Anything else? [01:09:24] Yeah. You've done, uh, binning methods. Equip, uh, where they create, uh, frequency zone. [01:09:30] Okay, okay. So these are the different methods, yeah, more or less, that is it, that is what it is. So these are the methods to find outlets, okay? Now, you use all of these on this salary. [01:09:39] Okay. You are using it. Finding all of this is fine. [01:09:42] Do you think right now, given this data, I have changed a little bit, do you think there are any outliers? [01:09:58] Rest only salary column. [01:10:03] Yes, I'm going to pay can be here. Okay. [01:10:07] Okay, if 75 is an outlier, what is wrong with 62? Then that is also very close, no? [01:10:11] Then, we have to consider the 3 outliers, like 50 and above is the outlier, because distance… look at distance from mean is 14, I think 18. [01:10:14] Yeah. [01:10:19] Now it is not 14, now it has gone up a little bit after. [01:10:20] Right, so distance is too high. 18, 18 ATM. It's 18 right now. [01:10:26] Yeah, so you still just, like, a 2X, 3x away, far away. [01:10:27] It is 18.9, yeah, 18. [01:10:32] For me. Yeah. [01:10:33] three outliers, you can say, right? Okay. This is fine. Here, this is how you generally find it, okay? [01:10:38] But one of the things you have to see. [01:10:39] But… but… [01:10:40] is that these two people. So you are saying these two and this one, these three are the outliers, right? [01:10:45] Okay. [01:10:47] Now, if I look at these three as outliers, do you think that these three are behaving similarly? [01:10:53] No. [01:10:54] No, that's… that's worth getting into, is that you have to see the age as well, so… [01:10:56] Yes. See, if I only look at. salary, or the better example is if I only look at age, for example. In this age, do I see any outlets? Forget the last two also. Just look at the first initial data. Do I have any outliers? [01:11:08] Nothing, right? These are all perfectly plausible ages, right? They're all correct. [01:11:13] But, when I… see, this 30 person whose age is 30, he will never be an outlier in the context of age, right? [01:11:23] But, when I add the information that this person is earning 75 lakhs. [01:11:24] Yes. [01:11:27] Then this person becomes an outlier? So this is what we know as bivariate outlets. Finding bivariate outliers actually is more important than finding univariate outliers. [01:11:39] Because if you find your outliers with respect to your target variable. [01:11:43] Say, for example, here, salaries may be my target variable, age is my independent variable, right? [01:11:47] And if I see this guy, 753. And if I keep this person in the data, or, I mean, whatever I do, I mean, individually, if I do only box spots, I will not find this person, right? [01:11:57] So, it is important that you draw scatterplots. Because anything that is in the corner of the scatter, in the top right or bottom left corner of the scatter. [01:12:05] That means that these people are behaving very differently from the rest. [01:12:09] And they, later on, will bring my model's efficiency down. [01:12:15] When we do linear regression. Later on, you will see, we haven't done it yet, so I can't explain from that perspective. There's a concept called heteroscedasticity, which you will learn later on. [01:12:24] Uh, so that is… this is the main source of that. [01:12:27] So we need to be very clear that our bivariate outlets are sorted. [01:12:32] Make sense? Yep. [01:12:33] Sir, one question. So, in skater plot, if X being the higher value and Y also being the higher value. [01:12:40] then I think we should be okay, right? [01:12:42] Then it is okay, it is okay. It's a problem when there is a problem in between them, accordingly. [01:12:47] Okay, got it. [01:12:48] Yeah. [01:12:49] Actually, now you say it, now we say, bye, where yet, since we have only, only two variables. [01:12:56] If there are some multiple variables, then it may depend on multiple accounts. [01:12:57] Yeah, yeah. How do you do it? Right. Very good question. And that, I will answer one… I had a question in mind, but. [01:13:05] your question is actually connected to that. Okay, let me ask you, give you a real-time scenario, okay? What is a real-time scenario? Let me write it down. I've written plenty here, unfortunately. Okay. [01:13:15] This is too much anyway. So this was actually on a similar side of class last time, but I don't think we have that much time today. So, uh… Say, I have 50 variables, okay? [01:13:26] 50 variables in my database. Now, in 50 variables, in a basic regression equation, just to introduce it, I am not doing regression, but you need to know this to get… everyone knows this equation. [01:13:27] Yes. [01:13:40] Yeah. [01:13:41] Yeah, Y equal to MXC, yeah. [01:13:42] Yeah, we learned school, right? State line. So, this is my classic linear regression equation. In that, I just keep adding something called E, or the error, and then it becomes a linear regression equation. [01:13:52] Right now, we'll not get into linear regression, so I will not explain this right now, but to give you a sense of it. [01:13:58] Now, out of these 50 variables in linear regression, one of them is Y, the rest of them are X, X1, X2, so on, so forth, M1, X1, M2, X2, I can extend this equation, right? [01:14:06] So, potentially what happens is that one is my target variable. Target variable is my Y, which I want to predict. [01:14:12] And other 49 are. Independent variables, which are the X's, okay? [01:14:20] And this is where regression comes in. In correlation, we have defined it how? [01:14:22] We have defined, said that degree of association between two variables. [01:14:28] height and weight of people, air pollution and air quality, like that. [01:14:32] In a regression, I am only interested in one variable. [01:14:35] My wife. I don't care about the other 49, to be honest, right now, in this project. [01:14:41] Why I have them? Because I think that they are potentially. [01:14:46] going to explain my target. So I… with the help of this 49, I can understand, explain. [01:14:53] Know how the target is behaving in the past, and how it will behave in the future. [01:14:59] This is, in essence, what we call it as regression. [01:15:03] Okay. Now, before we go to regression, let's come back to EDA, because that's what we are doing. [01:15:07] In the EDA, now I need to know that what. [01:15:10] What ED I'm doing… I mean, I'm doing the EDI for what? That's very important. I cannot go and do random ED, right? I mean. [01:15:16] I don't even know the data, I don't even know the model, and I'm doing EDA means it's a waste of time. It will not lead to anything. [01:15:22] So, before that, as I said, I need to know the problem statement properly. [01:15:24] I also need to know what my models. possibly are, right? [01:15:30] So, given that, I have this set up, right? Now, you tell me. [01:15:34] In the ADA, given that I will going to eventually use them here. [01:15:37] Right? And this is my setup. Listen to the question carefully. [01:15:43] What are the minimum, or what is the minimum number of. [01:15:47] scatterplot, or bivariate plots, you can say. I should be drawing. [01:15:54] In this case. [01:15:58] I think it would depend, so you… maybe you eliminate. [01:16:01] Now, again, it's… what is the minimum number, I'm saying? Maximum can be anything, it depends, of course. But what is the minimum you must check? [01:16:13] Absolutely, 49. Y49 Pallavi? Explain. [01:16:14] 49. [01:16:18] Because one of them is a target variable to, in this case, I just need one. [01:16:23] Yes, you need to understand how the other 49 are connected to the target, right? Individually, one by one. [01:16:24] variable rate, so… [01:16:29] If you don't do it, you don't find bivariate outliers. [01:16:33] Right? Make sense? How you connect them? [01:16:39] How and why this is so important to understand outlets properly and completely? [01:16:45] Yeah. Uh, I had a question. What if… what if the 49 independent variables. [01:16:47] Yep. [01:16:50] So, I guess we're assuming that they are all independent of each other as well. [01:16:54] Yeah, initiative… I mean, that's an assumption. Initially, when I have the data, I will assume that all of them are useful. [01:16:55] Okay, yeah. [01:17:00] But later on, of course, we'll see most of them are not useful, but that's a different thing in the model we'll see. [01:17:02] Okay. So, yeah. That's why I was getting into minimum, so if we eliminate the… features that I didn't have… yeah, yeah. [01:17:07] Oh, I said, no, no, no. Here it is… the question was that 49, I think I'm… I have use of them, yeah. [01:17:15] Yeah, but that's good, that's a good answer, Pallavi. That's absolutely the way… correct way to think. It's not about whether you give the correct answer, it's about how you think through it. [01:17:16] Yes. [01:17:24] Alright, everyone clear with what she said? [01:17:31] Right. So, this… see, this is how you have to frame your EDA. It's EDA, you will learn a lot of things. Sm is in value, SA outlier karo, this, that, blah, blah, so many things. [01:17:41] But when you are really given a data, and you have to do the EDA. [01:17:44] All that is good, but that is not very useful. You have to make your plan of action, right? [01:17:49] And there, the next question that comes in. Is that… is that million dollar question in the industry? [01:17:56] Should you impute outliers first, or missing values first? [01:18:01] Oh, it doesn't matter, really. [01:18:07] Because you have learned the ways to impute before, but now this is the question is, what do you do first? [01:18:15] Like… [01:18:16] Outliers first. Uh, when the… I guess there are multiple ways to, uh, sort of amend the… Uh, missing data, right? So you want to make sure the way you do it, you… the outliers do not affect your data. [01:18:24] Hmm. [01:18:29] Fair point, yes, fair point. Anyone who has a contrary opinion? [01:18:32] Yeah, I would say that we should analyze the data. [01:18:35] Because I think, uh, in this example also, we were checking the deviation and a lot of other things. [01:18:37] Hmm. Yeah. [01:18:41] For sure. So I think that would dictate a lot in terms of what is our approach. [01:18:46] Okay, okay. Anyone else? [01:18:57] See, the two answers I got. I am inclined to agree with both. They are both fine. [01:19:02] There is logic in both, nothing is wrong there. [01:19:05] What Meith was mentioning, I agree. Yes, I would rather do that most of the time. [01:19:11] But before I do that, what you said, I think Deepak's point is important. At the end of the day. [01:19:15] You have to do that first, which is more of a concern for you. [01:19:21] Okay. What is more in terms of percentage, for example? [01:19:26] And how you are imputing, you're thinking of imputing. If you are thinking of imputing or missing values by using mean, median, something like mean, especially mean or something like that. [01:19:34] Then, outlier first, absolutely. Otherwise, you will get a wrong mean, right? [01:19:38] But if you're thinking of binning them, if you're thinking of categorizing them, you're thinking of, uh… Um, interpolation, if you're thinking of other advanced methods, like, say, mice and all, right? [01:19:47] If you do in all those things, then it doesn't matter. [01:19:51] Whatever is graver, the bigger problem, you do that first, because you want to have control of the data as soon as possible. [01:19:58] You don't want to keep the control of the data in. [01:20:01] I mean, keep the control away from you for a longer time by doing other things before that, okay? [01:20:06] Because you need to see real figures, the right figures, as soon as possible. [01:20:10] So that's the reason why the bigger problem we address first. But having said that, if that problem needs that I need to use statistical measures to impute them. [01:20:19] I have to do outlays first. I hope that makes sense. [01:20:29] Yep. Okay, so this is also a very important point to be considered, okay? I'm just… I mean, you know that. [01:20:36] things I am just trying to give you a flavor of. [01:20:37] How you really use them or not use them in real-time work, which is… Eventually, what you want to do, right? That's why you are in the program. [01:20:45] Um… One other thing, I'm sure you discussed in the other day also about scaling, right? [01:20:52] Everyone knows what is scaling a data. Normalization and min-max and all. [01:20:58] Yep. And you also… do you also know about splitting data into train and test samples? [01:20:59] Yep. [01:21:06] I don't think we've gone through that, I'm not sure. [01:21:09] No, no, no, I don't think so. [01:21:10] Yeah, we've gone through that, okay. That is also come… Don't think so, right? Yeah, that's what. So, that comes in feature engineering. I'll just give you a brief of it. It will come when you… before you do the model, it will come. It's actually part of extended EDA in that sense. [01:21:21] data processing, I would say. So, when we are building models, again. [01:21:29] I have a data. Client has given me a data. I have to build a model on that, get some results, insights, share it with them, and yeah, that's how things work, right? [01:21:36] Now, I have this one data, but I cannot go and tell the client, or tell my… business head, for that matter. [01:21:43] That, okay, I have built some model, um… Yeah, seems like it will work next time. [01:21:48] I can't go and tell along those lines. I have to give validation of my work, right? It's very important. [01:21:54] So one of the things we do, the classic thing, it applies to every ML, DL, anything you do, go on. In the next few months of your program, you will keep doing this every time, so I'm just introducing it now. [01:22:04] It is called a splitting of data. Okay, I'm sure some of you have heard of this. [01:22:08] It is called splitting of data, and there is a thing called strain and test split. [01:22:13] So you split the data into train and test samples. [01:22:17] So, say your data has 1,000 rows. Let's take a small example, 1,000 rows. [01:22:21] And you create the first 700 as your train data. [01:22:26] The next 300 is your test data. So what happens? You build your model, you do all of this stuff on your train… I mean, EDI is common, most of it, but when you build the model, you build it on your train data. [01:22:37] You get the results, you test, or you implement those results on the test data. [01:22:43] The other 30% you have kept out separately. And if you see that how well they're fitting, how low is the error there. [01:22:50] Is the error there similar to the error you were getting in the train data? [01:22:54] If it is, then great, you have been able to fit your model on another unseen data. [01:23:00] If it is not, then you have to go back to your trained model of the initial model you have built on the 70% of the data. [01:23:05] And fine-tune it. Redo it. In other words, go and look at the EDA and start from there again. [01:23:12] Model building is a tedious process. You have to do it many times. You do… you go to the end and find out something is not working out, you can start from scratch, sometimes. [01:23:19] Okay. So this train test is a very important concept. I'm sure it will be in the… it will come when you do the models hands-on as well, later on. [01:23:28] Uh, so when you do trend test split later on, but right now, as an overview sense, did you get a sense of what I'm trying to say here? [01:23:29] Yes. [01:23:39] Yes, yes, sir. [01:23:40] Yep. Not complicated, it's very easy, yeah. So we take 70-30 generally because we want to build the model on the bigger data. [01:23:48] But we don't also want to make it so big that the other validation is very small, so we strike a middle ground by taking around 70-30. [01:23:54] Okay. Now… One of the things, one of the big concepts in ETA and feature engineering and data science. [01:24:03] I'll give you a flavor now. We… obviously, we don't have the time, or even you are not even there to sort of discuss this in detail right now. It will come once you do get into ML, at least. [01:24:13] Um, is a con… is a technical term or a concept or a phenomenon, you can say, is called data leakage. [01:24:20] Any of you heard of this term before? [01:24:26] No. [01:24:28] Okay, some yes, some no, I guess. Okay. Data leakage is a very, very, very, very critical thing. [01:24:37] And data leakage comes in… doing certain parts of EDA before splitting. [01:24:42] And doing certain parts of EDA after splitting. Okay, and if you don't follow that, you leak data. Now, what is data leakage? I'll give you a brief. The reason I'm doing is because you have done scaling before. [01:24:53] And the biggest mistake people do around scaling… I mean, see, min-max scaling, normalization, anyone can do. [01:24:59] You put up a standard scalar package, and you run it, right? No big deal. [01:25:03] So that's not where the trouble comes. The trouble comes is when you do the scaling in your process of EDA. [01:25:09] That is important. So, generally. we have to do scaling after we split the data into train and test, not before. [01:25:19] Can anyone think of and tell me why? [01:25:24] Uh, well, so you're normalizing with all of the data, right? [01:25:28] And basically, you're helping the model in that sense. [01:25:29] Right. [01:25:34] So, basically, your… the way you're building your features. [01:25:39] That… that… that process sort of goes into the model building. So now you're… You know, giving some part of information of the testing data to the model there. [01:25:46] Absolutely, that is a fantastic answer. Given that you didn't… you were not aware of this, this is a good answer, knowing it 5 minutes back. [01:25:54] Okay? So, yes, he is absolutely correct. If I do the splitting… sorry, scaling before splitting. [01:26:02] Scaling means what? Let me… let me put it this way. Scaling means what? [01:26:05] That I will… take a standard deviation, take a mean of the distribution across all my variables. [01:26:13] And I will do a X minus mu by sigma. [01:26:15] Everyone on me with me on this? X minus mu by sigma, standard scalar. [01:26:20] Yes. [01:26:21] Right? So we are, in other words. We're using the… Mean and standard deviation to. [01:26:28] scale the data, right? Now, as Meet was mentioning. [01:26:34] If I do it before splitting. I will have only one mean and one standard deviation for the intended. [01:26:38] And I'm trying to do… scale them all together. Then I do the splitting, 70-30 voila. [01:26:44] So what happens is, then whatever is in the 30%. [01:26:48] Those values are already influenced by the mean and standard deviation values which are actually applicable for the entire data, the other 70% part also comes in there. [01:27:00] Right. [01:27:01] Right? Makes sense, right? So that's where we can then say that indirectly at least. [01:27:07] The testing data is not unknown. There is some elements into it which is actually being connected to my initial data. [01:27:14] So if I use that for validation, it is a wrong validation, right? I'll get. [01:27:19] outputs more than I deserve, or better than I deserve, right? [01:27:22] So that is the reason why the scaling has to be after splitting. [01:27:26] I need to scale after splitting. There is another concept here called, I'm sure, again, from the quote perspective, you have seen this before. [01:27:33] That, in scale, how do you scaling? You have done this… you have done Python's coding before, right? [01:27:34] Yes, yes. [01:27:41] Yes. [01:27:42] Yeah. [01:27:43] Yeah. So there is a thing called fit transform, maybe. Maybe you have seen it. I'll also, otherwise I'll show it to you in the next few minutes anyway. You have seen it? [01:27:48] Yeah. Mm-hmm. Yeah. [01:27:51] Yeah, so you know already. So, okay, so nothing new there in that sense then. So, with transform, there is also, you think of it, I'll not tell you the answer now, because that is another longer discussion. [01:28:01] That I will actually, after I split the data, I will fit it on the train. [01:28:06] And with the fit on the train, I will transform both train and test. [01:28:10] Okay? Not fit separately and transform separately. I'll fit only once on train. [01:28:17] With that result, I will transform both train and test. [01:28:19] Okay. I will not give you the answer now. [01:28:22] But this is something which 90% people make a mistake there, okay? [01:28:26] Uh, you think through, once you reach maybe another couple of sessions next week, something. [01:28:31] you will be in a position to assess this, but keep a note of this. [01:28:35] That why we only fit on train and transform both using that. [01:28:38] Train and test, that is something, is a… Thought homework for you, thinking homework, okay? [01:28:45] Uh, yeah. Right. [01:28:46] I'm sorry. [01:28:47] Yeah, Mike, I have one question, sorry to interrupt. So, when you say split the data into 70-30, if you have n number of variables out there, right? [01:28:55] So, what are the criteria, so how do we split the data? [01:28:58] No, it is row-wide split. It has nothing to do with variables, row-wise. [01:29:01] Okay, okay. So we randomly split based on 70-30. [01:29:02] Always. Yes, yes, random. It's a random script, right? [01:29:07] Because if I make a buyer split on my own split, then I will end up with spurious correlation, right? [01:29:08] Okay. [01:29:12] Alright, got it. Thank you. [01:29:15] Yeah. [01:29:16] But, yeah. So, it is random, yes. It's a machine will do it using Python's random packages. There is a package for everything. You don't have to do anything. [01:29:22] These days, with ChatGPT and with the package information we have in Python especially. [01:29:29] Coding has become school kids. job, to be honest, if you ask me, okay? [01:29:33] No one bothers to be very frank with you. [01:29:37] I mean, that too, you can… anyone can do, but the main… Uh, make or break is not there. Main break or break is somewhere else, which is actually where we are discussing now. [01:29:48] Okay, so this is more one thing I wanted to talk about scaling, because you already know scaling, okay? [01:29:52] Uh… okay. What? [01:29:54] I'll suggest one question. I missed that the last topic, so if you could just… The fit training… [01:30:00] Which one? Which one? Ah, so that's a question. I haven't given the answer, so I thought… I told that you go back and think on this a little later, okay? So the question is. [01:30:10] That when I am scaling it after splitting, as we discussed, right, the scaling should be after splitting, right? [01:30:15] So when I'm doing that, generally we use this fit and transform to scale it, right, in the code. I think many of you said that you know this part, right? [01:30:25] Yeah, I'm… I'm aware, I'm not sure if anyone has this. [01:30:26] Or you don't know? Okay, okay, okay. [01:30:28] No… [01:30:30] I'll touch base, don't worry, I'll touch base. I'll show it to you, don't worry. [01:30:31] It's good if you can touch base, Mark, so that… [01:30:34] So, uh… okay, I'll show it to you, and then come back there. Maybe, maybe in the last 15-20 minutes, we'll do a bit of those things. [01:30:40] Huh, so… okay, we'll get there. So that is part of the scaling concept I wanted to get into. [01:30:46] One thing I wanted to add there. Um, one example I want to take to give you a sense of how to approach an outlier scenario, okay? [01:30:55] This is a variable called age, okay? for a banking transaction. [01:30:59] My minimum is, uh, 18, max is 105, okay? [01:31:04] Mean, uh, mean can be 50, median maybe. Doesn't matter, I'm just putting random numbers. Mean median is not very important here. [01:31:14] What is more important is 99th percentile. Which I'll put it as 85. Everyone knows percentile, right? What it means? [01:31:21] Uh… And equal to 100. [01:31:27] I have this data. 100 people, small data, 100 rows. [01:31:33] Banking sector, where I'm looking at age of people who are dealing with the bank, transactions, transaction data, customer data. [01:31:41] Um, this is the age, one of the variables is age. [01:31:44] Uh, one of the independent variables. I'm trying to make sense of it, trying to find outliers there, right? [01:31:49] And this is the information I have. Okay. [01:31:54] you got this after running describe function, you got this in Python. [01:31:58] Now, think of it logically, think of it step by step. [01:32:01] What's the first thing you will do once you see this? [01:32:13] Right, so if there are, you know, missing things, or, you know, missing data in that. [01:32:17] Missing note, missing lot, missing is different. I'm talking about only from outlier perspective. [01:32:23] IQR means you will draw a box plot, right? [01:32:24] Uh, IQR, just for my currency. Yeah. Yeah, yeah. [01:32:27] Now, one question that comes to my mind, I wanted to get to that question also via this example. [01:32:32] In Boxtroad, as you know, right, we have this example, we have this formula, right, that IQR 1.5Q1 minus 1.5 IQR and all that, right? [01:32:41] And Q3 plus 1.5 AQR. Now, one question let me ask you. [01:32:42] Mm-hmm. Yeah. [01:32:45] Why it is 1.5 and not, say, 2.5 or 0.5? [01:32:51] Have you thought of this? Have you asked this? [01:33:02] I'm gonna repeat again. [01:33:03] We have to think more. Now, see, I'm… I'm… Sorry? [01:33:04] Okay, repeat again what is… what's the question? [01:33:09] The question is, in IQR formula, we have no Q1 minus 1.5 IQR for box plots, no? [01:33:18] Yeah, yeah, oh, okay. [01:33:19] And Q3 plus 1.5 IQR. So my question is why it is 1.5, not something else. [01:33:20] the distribution, uh, expecting… [01:33:25] the data for 1.5… [01:33:27] like, distribution is like that, I've seen it now. [01:33:33] Yes, you are right. So what he says, G2, is actually correct in a sense, because in normal distribution. [01:33:38] We have this thing called mean plus minus 3 sigma. You know this? That. [01:33:43] There's a thing called mean plus minus 3 sigma, where 99.73% of. [01:33:47] observation should lie. Everyone aware of this? Theorem? No? Okay. [01:33:54] Hmm, okay. That is intrinsic to outlier treatment, this theorem, actually. Okay, I'll tell you why. [01:33:55] I think yes. [01:34:02] What's the name of the theorem, Mayung? Sarek? [01:34:03] Okay, um… It's called Chebyshev's Inequality, but we don't use that term, it's more pure statistics term. [01:34:10] What we say is, and let me write it down here for your reference. [01:34:14] Um, mean… Plus, minus… 3 into SD. [01:34:20] Okay. So what does it mean is that in the left-hand side of the normal distribution, I go to this curve. [01:34:26] So this red line, as we discussed before, is the normal distribution, right? [01:34:29] So, the left-hand side is my… Negative side, the right-hand side is a positive side. Here, everything is zero, the center part, mean, median, mode, all are same, right? [01:34:38] In this left-hand side and the right-hand side, because they are equal, the tails are equal in both sides, right? [01:34:43] What I'm saying is that the mean being. a small number, or a zero, or a central val, whatever, it can't be zero also all the time. Whatever the central value is. [01:34:53] From that, if I take the standard deviation into both sides, plus 3 standard deviation in the right-hand side. [01:34:57] And minus 3 standard division in the left-hand side. [01:35:00] Okay. I will reach a number. Say, for example, in a distribution, mean is 50. [01:35:06] So, is, uh, SD is… 10, okay? Then my mean plus 3 standard deviation is 80. [01:35:15] And the mean minus this energy of 20. So I can say there is an 80-20 range. [01:35:20] From this formula, can I say that, both sides? [01:35:21] No, no. [01:35:24] Right? So I am saying, effectively, Chabishev says, or this inequality says, that if my distribution is normal, obviously this holds when it is normal, otherwise symmetric skewness distribution matrix will not hold. [01:35:35] Okay? If my distribution is normal. Anything which is outside this range. [01:35:42] is a potential outlier. Because empirically proven, I'll write it down here, this is absolute golden basics of stats, you must know. [01:35:50] 99.73% of observations. will lie between. [01:35:57] this… or in this range, in this range of mean plus minus 3 standard deviation. [01:36:03] Only 0.27% empirically lies outside. And anything that is outside, which is the 0.27%, we say that they are potential outliers. [01:36:13] And that is why I was earlier referring to this thing, that we always have a… Tendency of trying to convert a skewed distribution to normal distribution. [01:36:23] Why? Because if I can do that, I can easily find out, I don't have to draw box plot this, that, blah blah, every… all… all those… drilling kind of things. I'll just do this calculation. I'll use this formula, I'm done. [01:36:36] A lot of less work? Are you getting what I'm saying here? [01:36:37] Yes. [01:36:39] Yes, yes. It's actually a shortcut. [01:36:41] Yes? So that's why the temptation. Ah, that's why the temptation to get everything to normal. [01:36:47] Because then I can use all of these. But that is not a good temptation to have, unfortunately. [01:36:54] It's a temptation of having biryani every day, not good for your health, right? [01:36:57] So this temptation is not good for your model. [01:37:00] If it is there already, great, fantastic. But if it is not, don't try to do that. Don't try to over-engineer a data for your own sort of, yeah, easy work. [01:37:09] Okay, that's not good. Then the production thing I told before, no? [01:37:14] Your budget will be overrun, your project will be scraped in 3 months, that will happen, eventually. [01:37:18] Short-term happiness, long-term problem, okay? So, this is what is the… Jabyshev's inequality for normality. So, if I can prove it's a normal distribution, outlier treatment is very easy. [01:37:30] But, as I said before, most of the distributions are not normal. [01:37:34] they will be symmetric… asymmetric. Some skewness up, down, somewhere, right? [01:37:39] I'll… to make it a little more easy here, I'll make the mean as… To make it very clear, make it as 45, no problem. [01:37:46] Now, if I come to this. I know that this is not a normal distribution, right? Can I say that? [01:37:47] Yeah. [01:37:56] I don't even have to draw a plot. I look at these numbers, I already know. [01:38:00] Great, yeah. [01:38:01] Right, you don't have to draw a plot for everything. See, the smart data scientist doesn't write 100 lines of code. [01:38:05] They would like 20 lines of code, but they can infer 5 inferences from one output. [01:38:10] And the other guy will infer every… output 1 inference. That's not going to help, right? You have to understand that one output, I can actually draw 10 inferences from this. [01:38:18] Which saves me writing 10 lines of code, right, at the end of the day. [01:38:23] So I know it's not normal, it's symmetric. Sorry, it's skewed, right? [01:38:26] The first thing I have to see here, or think of here. [01:38:30] Is that who are these people? Forget numbers, forget max mean 9900, anything. [01:38:37] Who are these people? Why it is important to know who are these people? [01:38:40] Because my context of data is very important. If I think that these people are Indians. [01:38:45] or South Asians, India, Pakistan, Bangladesh. These kind of countries where we leave, the geography where we live, for example. [01:38:53] Or you go to certain other countries, South America, for example, Mexico. [01:38:57] certain North African countries. Southern European countries, for example. There. [01:39:03] Maximum of 105, do you think that an Indian bank will entertain someone who enters and says, I'm 105 years old? [01:39:11] No. [01:39:13] No. Even if he's saying the truth, they will say, what's Abhjakasoja, right? Simple. [01:39:18] So, that is something we are not used to. [01:39:21] We don't see people roaming around at 100 years old, right? [01:39:25] But in certain other countries, you go to Japan, you go to New Zealand, you go to Scandinavia, for example. [01:39:30] Their average life experience is 94. 92. This is only 10 more than that. [01:39:36] Ever life is explicitly, India is 70. It is 1.5 times that. [01:39:41] So I need to know who is this person. Is this person from. [01:39:45] Tokyo? Is this person from Copenhagen? Or is this person from Mumbai? [01:39:51] That is critical, because I know, I mean, in reality, I know a person in UK, he passed away a couple of years back. [01:39:57] He was 101 when he was doing his own banking and post office work. [01:40:00] And then the very next year, he went and enrolled himself into an old age home. [01:40:05] So, things are different in different parts of the world. [01:40:07] So I need to know how improbable this value is. [01:40:12] If it is in our part of the world, it's absurd. [01:40:15] If it is in their… some part of the world, it is not that absurd, it's possible, okay? [01:40:21] That's one thing. I have to first understand that, is the data wrong or not? Or is the data. [01:40:24] Improbable or not, right? So for that, I need to know the demography. [01:40:29] Done. Whatever it is, I find that, okay? Now, the next thing, let's try to understand this. [01:40:33] Now, you have, in outlier treatment, you must have heard of a term called capping off or winderization. [01:40:41] No, no, Sandia. [01:40:42] Are you aware of this term? [01:40:43] So, what outlier treatments have you learned so far? [01:40:47] How to treat outlays. I'm not saying finding, treat. [01:40:54] None, I think. [01:40:56] Okay. So, see, the point is, might be… might be… might be covered later on. [01:41:01] Finding is the first part. Then you have to do something to it, right? [01:41:06] I will find, and then I'll be happy with that, right? [01:41:11] So, what are the treatment? The main treatment is called windsurization or the primitive term when we were learning 10, 15 years back, it was called capping off. What is capping off? [01:41:21] I say that my 105 is improbable as a name hot. [01:41:23] So what I'll do is that I, rather than imputing it with a mean or a median. [01:41:28] I will impute it with the nearest number. A percentile where it makes sense, maybe the 99th percentile, which is the 85. [01:41:35] So I'll replace 105 by 85, rather than by 45. [01:41:39] Because imputing outlier by a central value doesn't make sense. [01:41:43] Central Value somewhere in Delhi, outlier is in London, right? [01:41:47] So, that imputation gives a lot of wide-ranging changes. Distribution is more tampered than manipulated, right? [01:41:54] So what I do is that we do capping off. [01:41:56] So here, if I use that method, the industry uses capping off. [01:42:00] Widely, okay, for outliers. Say, for example, I have 100 observations. [01:42:06] Let me make it 500. 100 will not make sense. [01:42:09] 500 observations, 500 people, okay? So the last 1 percentile will have 5%, can I say that? [01:42:15] 5 people in the last percentile? [01:42:20] Yes. Yes, yes. [01:42:21] Makes sense so far? Able to follow? Yeah. So, say this 5… I'll put the age of these 5 people. Say 87. [01:42:28] Um, maybe 90? 92? 95. And our maximum, 105, okay? [01:42:33] Now, I want to cap it off at 85. [01:42:36] I have to do something to outline, I have to capital. [01:42:39] Let's do the deviations. 87 minus 85 is 2. [01:42:42] 90 minus 85 is 5. 92 is 7, 95 is 10. [01:42:48] This is 20. So this is the deviation, the net deviation. If I change all of this to 85? [01:42:55] Right? So I add them up, 14, 24. [01:43:02] 44. This is… adds up to 44. 44 is the total deviation across 500 observations. If I do outlier treatment. [01:43:12] I do 44 by 500. Okay. 44 by 500, if I put an equal to sign here. [01:43:20] It is 0.08, not even 0.1. this much, I am deviating by doing a groundbreaking outlier treatment. [01:43:31] Do you think that this variable, this much deviation. [01:43:34] will have any impact in this variable becoming an important variable in my analysis later on? [01:43:40] Or a non-important one. Do you think it will have any impact or any change? [01:43:41] Yes. [01:43:45] No, it won't happen. [01:43:46] It won't. It won't. It will not impact anything, but here. [01:43:52] This 75 will have an impact. As we saw. [01:43:58] So you have to choose your outliers. It's a bit like you have, like we say, you know, you have to choose your battles. You don't go and want to… fight with everyone in the world, right? You just choose your battles, conserve your energy. [01:44:08] Same applies for data scientists. Don't. fight every battle. [01:44:12] minutely, just because you found outliers, there is no brownie point or Nobel Prize for doing anything to them, okay? [01:44:18] If you find outliers and then see that they are borderline outliers. [01:44:21] Don't do anything to them. You're just wasting your time. It will not lead to any tangible result later on. [01:44:27] But if you see their extreme outliers, they are going to have a. [01:44:30] Uh, detrimental effect on the rest of the thing, like we saw here. [01:44:35] Then, yes, you go on Winsorize, change cap off, categorize dummy variable. [01:44:39] hundreds of things you can do. But don't do it without logic. [01:44:44] Just because someone said that outlier, if you find, you treat. [01:44:47] Nothing like that. There's a step in between. Where you think that… is it worth doing or not? [01:44:52] So that is very, very critical. Don't be a process-driven data scientist. That is not good. [01:44:59] Be a thinking data scientist, okay? I hope I'm making sense, are you able to understand? [01:45:07] Yes, yeah. [01:45:08] Nope, taking the same example anyway. Yep. So this is… This is another thing in terms of outlet I wanted to mention. [01:45:16] This is very, very critical. We don't want to over… I mean, see, do smart work. At the end of the day, that matters, right? [01:45:24] And… this is smart work. You don't want to… just because you have, as I said, right, you have a. [01:45:29] notebook, and yeah, writing quotes don't cost you money. [01:45:34] It doesn't mean that you will just keep on doing that, right? I mean, it doesn't make it randomly mechanical. [01:45:40] All right, um… Okay, so what we can do, I… can think of 5 other things they want to discuss, but. [01:45:49] We don't have time for that, so I'll… I'll not get into those things, and I am sure that those things will be covered at some point anyway. [01:45:56] Um… Any questions anyone have so far? [01:45:57] Yes. [01:46:04] based on the session that you gave today, right, and based on whatever basic understanding we have, are there any topics that you would suggest [01:46:10] Maybe from a basic perspective that can help us [01:46:13] in the further classes, because… [01:46:15] You knowing this might… [01:46:17] give us a better insight. [01:46:18] Yeah, so I will say there are 3 topics which you should really concentrate on. [01:46:24] Okay. One is statistics. Two statistics, 3 statistics. [01:46:30] If that helps you to understand what I mean. [01:46:33] Statistic is big. [01:46:34] Okay, so what in statistics, yeah, like… [01:46:37] I guess we have spent our school, college in statistics, and we don't forget anything, we don't remember today. [01:46:40] No, no, no, no, no, no, no. You cannot now… now the… now you are into a different zone, right? [01:46:48] Yeah. [01:46:49] You cannot now study statistics like you studied in college. [01:46:51] Exactly. [01:46:53] Got it. [01:46:54] Okay, because that will not help. Okay, so you… I'll spend a couple of minutes, because I think this also is a good question you ask. [01:46:58] See, when you are learning statistics, for example, or studying again, revising, whatever it is, right? [01:47:10] Yeah. Yeah, yeah. [01:47:11] It is at this level, it is not about knowing the definitions of the formula. No one cares. No one cares, to be honest, right? Everyone knows or Googles up, right? [01:47:14] synthetic. [01:47:15] So, what is that what you can bring in as a human being, which Google or AI or Copilot or ChatGPT can't? [01:47:22] Right? That you have to think of, right? So, from that perspective, when you are learning statistics, or going back to revised statistics. [01:47:29] Think of it, every concept you learn or revise, right? Every particular concept, there are two approaches you should do. [01:47:35] First is that try to see where all it fits. [01:47:38] in your machine learning journey. Once you get into ML, you will get that sense, right? [01:47:42] And second, how are these concepts. As they go along in the sequence, right? [01:47:48] How are they connected with each other? And more importantly, how are they disconnected with each other? [01:47:55] That is the approach when you're learning, everything has to be there. [01:47:58] Definition formula kit in hoga, that is college, or done. [01:48:02] I hope… I don't know whether it made sense, you have to tell me. [01:48:04] No, no, that's true, because today's date, if somebody tells me to buy out something, I may not even put the efforts to do it. [01:48:08] Yeah, yeah, yeah. [01:48:09] So, keeping that in mind, I don't know. Is there a subtopic that you can suggest, or we just pick it up and try to understand? [01:48:16] I think if you say, if I say… Concentrate on certain things, right? If you pinpoint, right. [01:48:21] I would say that it's very important that you. [01:48:24] are very clear with your concepts in. hypothesis testing. [01:48:29] Okay. [01:48:30] Absolutely critical, because if you don't understand hypothesis testing well, right. [01:48:34] You will not understand. The most important parts of regulation, or time series. [01:48:41] Because it is all built on hypothesis testing. So, I see these days there are people who are in the industry, I've heard a lot of people saying that, I think, Abi, we have transformers, we have generative AI, LLM model, this, that, Langchin, llama, blah blah, whatever. [01:48:55] A stats for NHA. If anyone says that, or if any of you hear that, use both your ears. [01:49:02] Enter from here, exit from here. That's absolute rubbish. [01:49:03] Okay. Got it. [01:49:07] If you don't know stats, you don't know data science. [01:49:10] Cool stuff. So, that is the first thing. You have to make sure that the hypothesis test, probability distributions, estimation theory. [01:49:19] These things are clear, to some extent. You don't have to be a statistician, but somewhat clear. [01:49:23] Then only you can go into ML, otherwise ML itself will be vague for you. You will now only learn how to run a code and get an output, but that is not data science or machine learning. [01:49:31] So, to understand concepts, and when a model fails, right, what we see every day. [01:49:36] If the model fails. When a model fails, how will you solve it? If you don't know the concept of how it is running in the backend, right? [01:49:42] If it is running smoothly, you know the code, you run it, you get results, fantastic, you give a go for a party. [01:49:47] But if it fails, then you have to bug… you have to solve it, right? [01:49:50] And to solve it, you need to know where it failed, and to understand that, your conceptual clarity and sequencing of the model framework is very important. [01:49:58] Otherwise, so, okay, you fail, then I don't know. Someone come and help me. [01:50:03] You can't become a property… as I say, again, quote unquote, proper data scientist. [01:50:09] You can go and tell anyone. Who cares? Yeah, Suda, do you have a question? I think some hands I can see his wrist. [01:50:16] Yeah. [01:50:17] I thought he, uh, sorry to interrupt, I heard you tell three things. One was hypothesis testing, one was something with probability, and one more thing. [01:50:20] Probability distribution, which is, in fact what we touched upon, normal and binomial. [01:50:24] Okay, and the last one? [01:50:28] Okay. [01:50:29] estimation, theory of estimation. There are a couple of concepts there. Anyway, you can note it down. [01:50:32] One is point estimate… estimation, one is interval estimation. [01:50:36] point and interval. [01:50:37] And the third interval, the third thing is maximum likelihood estimation. [01:50:42] Because your logistic regression uses the algorithm which is maximum likelihood estimation. So, if you don't understand that, you don't understand logistics also. [01:50:50] Okay. [01:50:53] Sorry, go ahead, uh-huh. [01:50:54] So, linear logistic… yeah, in a linear logistic is all statistical models. No, if you don't understand statistics, then how will you know statistical model, right? The name itself defies it, right? [01:51:01] Got it. [01:51:04] And even time series forecasting. It's the most complicated thing. [01:51:07] And for that, your stats and linear regression both have to be top-notch clear. Otherwise, that. [01:51:12] that doesn't go in. It takes years for people to master it. [01:51:16] Okay. [01:51:18] Thank you. [01:51:19] But that's a lot later, don't worry. Yeah, okay. Um… Fine, we have, what, I don't know, 15, 20 minutes. [01:51:26] 15 minutes left. Okay, so should I then go and show you some codes, or you want to discuss a few things? You let me know what helps you, I don't know. We have… the constraint is time, you tell me. [01:51:37] I just have one very quick question, um… [01:51:38] Yeah. [01:51:40] So, I've seen you've touched this point a lot in this class, right? That saying that we have to differentiate ourselves from the AI model. [01:51:49] So, I just want to understand the current state of it. Let's say if I, uh… [01:51:50] Hmm. [01:51:54] Hmm… [01:51:55] server to my data, and just tell a reasoning model like Opus or Jackson. [01:52:00] If I, you know, do something with my data. [01:52:01] Hmm. Hmm. [01:52:03] So, like, what's the current status? So, where is the AI model lagging, and, like, what's the differentiating factor, uh, precisely? [01:52:10] I will… I will… so, it's more about the evolution and where we are and what to do, right? From that perspective, you're asking, right? [01:52:17] Correct, correct. [01:52:18] I'll get to that at the end, okay? Towards the end, we'll spend a couple of minutes on that, but one thing to now, just one thing I call… I mean, I caught from what you said is that you upload your data on ChatGPT and all. Please don't. [01:52:29] No, no, I won't, I'm just saying, if I do give a connection, like, a local model or something like that, I know they're gonna steal it. [01:52:32] Please don't. [01:52:36] If it doesn't start with… [01:52:38] No, it is… so it's not about stealing. See, you… I don't know how many of you know this. [01:52:39] Yeah. [01:52:41] There is an entire case that is running against HSBC. [01:52:45] Because one of their employees did this. And they're a bank. [01:52:50] And the fine in European region is 20 million euros. [01:52:57] If you do this. don't do. Don't even think. [01:52:58] Oh my god. [01:53:02] If you get caught, right. I will not tell you the consequences, you can understand, possibly. [01:53:11] Alright. I meant more of, like, if you have, like, a local… [01:53:18] Yeah. [01:53:19] No, no, I know you told it lightly. I know, I know, you told it lightly, but I mean, yeah, I still… that kind of stuck in my ear, because every time I hear that, I thought that. [01:53:21] Got it, got it. [01:53:23] I mean, this is very critical to be sure that we are not getting into that. [01:53:27] Got it, got it. [01:53:28] Yeah, because there is an entire topic of… called responsible and AI ethics, which is very important these days, right? Because people. [01:53:35] Thousands of people in the world are using ChatGPT, right? Because that's free. [01:53:38] But how many know how to use it? No one. [01:53:41] 1% of them, and that's the entire problem. Anyway, um, okay, I'll stop the sharing. Let me do one thing. [01:53:53] Let me show you a quick, sort of, because we are supposed to do a bit of coding, so I thought of doing, but again, I know the timing is a… I mean, the time is very strict, couple of hours only. [01:54:03] Um, so… Yeah, I still have a few things in mind, but I think maybe we can see a bit of code if that works. What do you guys suggest? [01:54:04] Statistic is big. [01:54:14] Good, good. Let's see some examples. [01:54:17] Yeah. I'll just show you a little bit anyway, we don't have much time, so we'll have to… we'll only have to be satisfied with a little bit. [01:54:27] Um, let me see which… what I want to do, because I have 4 or 5 different notebooks on EDA. I want to show you at least one. [01:54:35] some part of it, at least. Oh. [01:54:40] Would you be able to share your notebooks later? [01:54:42] Yeah, yeah, yeah, that I can, no problem. Yeah, see, again, that's a good point, because notebooks, so if I share, you will yourself understand what is happening. [01:54:51] I mean, you know, but still, I would like to show you once, just to go by the decorum of the class or the session, uh. [01:55:00] content as such. Uh… yeah, I would rather… because I was talking about scaling, let me show you that piece of the code a little bit, okay? [01:55:09] Because that is something I told you, but I didn't show you exactly that fit transform well apart. [01:55:16] And stop the sharing. Give me a minute, I'll go there. [01:55:22] I actually also prepared a short sort of PPT for you guys. [01:55:27] Which was more in the, uh… [01:55:33] along, say… so the PPT is about how you do missing value and outlier. [01:55:37] in different industries, what are the main best practices across all the major 4 or 5 industries? [01:55:43] But I don't think I will even have any time to show it to you as well. [01:55:50] that you could share, right? [01:55:51] Uh… I will later on, maybe I'll give it to you, yeah, not a problem. [01:55:55] Because that is just plain English, even if I don't read through, you will understand most of it. [01:56:06] Yeah, yeah, they'll share it with you, don't worry, yeah. [01:56:07] You can share it to, uh, insurance people. [01:56:08] Okay, so this is… this is, uh… okay, don't look at this, this is NLP. I'm not going to show you NLP right now. [01:56:15] I'm going to go to IDA. [01:56:21] So, this is what I was talking about, scaling, right? Let's look at the scaling piece, at least, in whatever time we have. [01:56:27] I'm not going to run anything. We don't have time for running and all those… Uh, uh, thing. [01:56:32] So this is the data, don't worry about the data, it's a very basic sort of data. So, see, this is what I was mentioning. When you… so, you know this thing, standard scalar package, you have heard of this before? [01:56:43] This… no, alright. So, see, let me go to the top then, not a problem. [01:56:47] So, see, what happens? Okay, in fact, I am… I am importing there itself. [01:56:51] So, you have learned the concept of scaling, right? So, this is the data, which is, uh, sales data. You can see some variables here, which are. [01:56:59] Put the screen a little away, yeah. So these are… you can see the variable names are very easy. It's a… it's a retail data. [01:57:05] price item type, some stationary data, right? Fat content, weight. [01:57:10] Um, visibility of item in a store. Outlet number ke hai, and all that stuff. [01:57:15] some retail columns are there. outlet establishment here when it was built, size of outlet in medium, small, and all. [01:57:23] How much sales they are doing, that's basically my target and all that, right? [01:57:27] So this is a very straightforward retail kind of data. [01:57:30] Uh, I don't think so, I will bypass this part, but I don't think you have done dummy variables or one-hot encoding yet, right? [01:57:39] Yeah, that is feature engineering, don't worry, we'll skip it for now. [01:57:42] So that is more about how to convert your. [01:57:46] string variables or categorical data into number data, right? That is what is one-order encoding or dummy variables. [01:57:53] That's an entirely separate discussion, uh, which is… which takes a bit of time. [01:57:58] Uh, okay. So this part I'll skip. Um… A lot of encoding methods I've used here. [01:58:07] Yeah. Okay, fine. So we go to scale the data. Now, in scaling, the two scaling methods you have learned so far, I am sure you have learned standard scalar and min-max scaler. [01:58:16] Right? Yes? So what is the difference in standard scalar, I do X minus mu by sigma? [01:58:20] Yes. [01:58:24] And in min-max killer, they say X max minus X, mean by X max kind of thing, right? [01:58:30] Everyone knows that? Now, tell me one thing quickly. [01:58:35] I hope it is fine if we go overtime by 10 minutes. [01:58:36] Yeah, done none. [01:58:37] Nope. [01:58:39] Because I think we will, then 10-15 minutes, I guess, yeah. [01:58:40] Yes. [01:58:44] So, uh, because some things I… I mean, I will feel irritated if I don't touch them, so I will have to. [01:58:49] Fortunately. So, uh… min-max, so the range… After you do a min-max scaler. [01:58:56] the range of the values should be between minus 1 and plus 1, right? [01:59:05] Okay, now I have a question for you. Now that you know. [01:59:09] That, in a normal distribution, what is an outlier? We have discussed it, right? That 99.73% and. [01:59:15] all that stuff, right? That, uh, anything mean plus minus 3 sigma and all that, right? [01:59:19] Now, given that you have that knowledge. Now, think through, connect the dots, and tell me. [01:59:25] That if in a data, or in a variable, okay, listen carefully, if in a variable, which is normal. [01:59:30] First condition, variable is normal. Second thing, I have already done outright treatment. [01:59:37] These two things have happened. Normal variable, outlier treatment done. [01:59:40] Now, I want to scale the data. scale that variable. Then, for that variable, I mean, everything, but that variable result I'm seeing after scaling. [01:59:48] What do you think should be the value. Within which my scaled numbers or scaled results should like. [01:59:57] post-scaling numbers. [01:59:58] Peace. [02:00:04] I mean… [02:00:05] Sorry? No, it's not about mean. I'm looking at a number. I'm looking at which number to number, the range of a number in which it should like. [02:00:11] after I do scaling to this, when it's normal and outlier has been taken off. [02:00:15] Minus 1 and plus… [02:00:16] Yes. [02:00:18] So, [02:00:19] 0 to 1… [02:00:20] No. No, you can't just randomly tell why you are saying all this, you have to give a reason, no? [02:00:26] How you're getting there? [02:00:30] Depending on the columns we are selecting, uh, like, [02:00:36] If we have 3 columns, and once we have scaling one column, we will check [02:00:41] Other two constraints, I will be selecting that range for the straight CRV. [02:00:44] Okay. Okay. [02:00:49] Okay. Let's say, okay, fine, I'll come back from there. Let's say I have this date, okay, I have this column. Say, for example, I have this wait column, item weight, okay? [02:00:57] Or let's take sales, easy to take sales, right? We have sales, yeah. Outlet sales. [02:01:02] Now, if I want to scale this data, I have other columns. I'm scaling all of them together. [02:01:06] But if I want to scale the values of this column. [02:01:09] Now, how do you think, if I use a standard scaler, for example, right, how do you think. [02:01:14] that these values all become small, right? After scaling, we see they become very small. [02:01:17] 0, 1, 2, 3, some very small numbers, right? [02:01:20] How do you think we reach that number? What is the formula for that? [02:01:21] Okay, thank you. [02:01:25] We might usually take a log of it or something. [02:01:26] For, uh… [02:01:29] No, no, there is nothing to do with log here. Log is transformation, remember. [02:01:33] Ray. [02:01:37] Oh, okay. [02:01:38] Scaling and transformation are different things. Don't get confused. Very common, people get confused. Transformation is something where we do log reciprocal, boxcogs. [02:01:41] square root. In transformation, what happens? I change distribution after I transform. That's why it's called transformation. [02:01:49] Okay. [02:01:50] In scaling, I scale it down, but distribution doesn't change. [02:01:52] Okay, cut it. [02:01:53] Um, so for… for standard scaling, I think plus 3 minus 3, maybe? [02:01:59] Yeah. Oh, yeah, yeah, I was thinking if we go with the mu plus 3 sigma minus 3 sigma. [02:02:00] Absolutely. I… I'm not sure why you said maybe, but… see why it is minus 3 plus 3. Exactly, exactly. You have the right reason. See, let me add a column here, okay? [02:02:11] I am saying, mean… Plus, minus 3 into ST. [02:02:17] Sorry. Sd. Now, what, again, the only problem, why are you not getting it? Also, I should have realized it, is that you don't know standard normal distribution. You know standard normal distribution? No, right? [02:02:30] Because we were not aware of normal distributions, I guess you don't know standard normal answer. [02:02:33] So there is a distribution called standard normal, and please, please read through these things. I said probative distribution, binomial, normal, and standard normal as well. It will come with it. [02:02:41] So standard normal distribution says this thing, what you are doing in your. [02:02:45] X minus mean by SD, no? This is… we call a Z. [02:02:50] You have seen this before, Z equal to this, which is my standard scalar, right? [02:02:51] Mm-hmm. [02:02:54] This follows something called standard normal. Why it is called standard normal? [02:02:55] Yeah. [02:02:58] Here, my mean is 0. And ST is 1. [02:03:03] in the distribution sense. So when I plug in these values. [02:03:08] in this plus minus 3 sigma. What happened? 0 plus 3 and 0 minus 3, right? [02:03:18] That is why, if I write it. For a variable. [02:03:24] Which is normally distributed, okay? Screen is coming. Yeah, for a variable which is not… okay, it is taken as a for loop for anyway. [02:03:33] So I'll put a comment. stuff here. For a variable which is normally distributed. [02:03:41] And outliers are taken care of. Post-scaling values. [02:03:47] If it is standard scaler, of course. Should always… sorry. [02:03:53] Always lie between minus 3 and plus 3. If you see any values going outside minus 3 and plus 3. [02:04:02] For a normally distributed variable, approximately at least, and when outlets are taken care of. [02:04:07] That tells you your outlier treatment was wrong. Not taken care of properly. [02:04:12] This is how we connect concepts. I have done outright treatment. Now, who will tell me whether it's the right treatment I've done for this variable or not? [02:04:19] I come here, I… via scaling, actually, I get a sense. [02:04:23] That whether my method was the best possible method or not. [02:04:29] Yes. [02:04:31] Can you repeat that, sorry? [02:04:32] Making sense? Yes? You have to… Yeah, so I'm saying that if I don't see this, that means what? [02:04:36] Distribution is normal given to me. I can't do anything. [02:04:39] The only thing I have done myself is out. [02:04:40] Hello. [02:04:45] So, the only thing I am doing otherwise here myself is outlier treatment, right? [02:04:50] So, if I have done a blunder, or if I have done some mistake, I have done here. [02:04:54] That's the only place where I could have done a mistake. [02:04:56] If I see these values outside the range, that means the outlier treatment I have used for this variable is not the best possible treatment. I have a better treatment to do. [02:05:05] criteria. [02:05:06] Yep. So this is how you have to connect, because at the end of the day, see, your model, your work, you have to find in the EDS step itself. [02:05:13] Some other method to reconfirm that the process you have done before is the correct one, or the best one or not. [02:05:19] This is how we do it. We can actually use scaling. [02:05:23] For at least… if it's normal, it's better. If it's skewed, there are different ways to do it, which we'll not get into now. [02:05:29] If it is normal at least, we at least know that my outlet treatment has been done properly, or I missed it, or I overlook it, or whatever, right? [02:05:36] It can be many things, but… I have to… actually, I can use scaling to do. [02:05:41] Avoid data leakage, which we discussed before, as point number one. [02:05:44] And second is, actually, in certain cases, it can actually give a sense of how my outliers or missing values have been done. [02:05:51] So, in that way, scaling is not a process which is limited in itself. [02:05:55] It can actually lead to… or help you in assessing other things in the EDA log. [02:06:01] Which will make your EDA more robust, more sort of confident about it. [02:06:09] Yeah. [02:06:10] Makes sense, everyone? Yep. This is how you have to do it, because this is how it becomes interesting, otherwise just covering topics one by one, learning the code by. [02:06:18] If I am subjected to that, I'll sleep. I'll be very honest with you. [02:06:21] So, unless I'm connecting concepts and going back and forth, there is no thrill of doing it. And if there is no interesting thing, I'm like, yeah, why would I… Spend my weekends and learn some, yeah, run-of-the-mill stuff, right? [02:06:37] This is why, and this is also helpful in your work. [02:06:39] Your EDA is then much, much more robust. So, encoding may not go, and even missing value outlier may… the coding stuff, I'll not go. I'll share these codes with you, don't worry. [02:06:49] Later on, you can have a look into it once you are a little more into all of these things. [02:06:54] Um, I want to show you a couple of… one more thing. [02:06:57] Which is, I think, very important, even before you get to EDA. [02:07:00] And apart from that, I have one question someone asked about the AI journey stuff kind of thing. I'll address that. [02:07:05] So these two things will do. Uh, I will… I will open a different… data set. I will open a real sample data, and I'll… I'll ask a question from that to you. [02:07:22] Okay. Excuse me a second… [02:07:33] Okay, uh, stop the sharing. Yeah, I'll just share the different screen. [02:07:41] And I'll tell you what I want from that. [02:07:46] Okay, some… actually, this is feedback that has come up, alright. [02:07:49] Okay, if you want to, you can fill it up, then we'll come back to this. We'll need another 10 minutes, I hope that's fine. [02:07:59] You first fill it up quickly, as soon as you possible, then I'll come to this question I have. [02:08:21] Done, no? Take it. Okay. So, see, this is the data. [02:08:25] If you look into it, it has… I'll just zoom it out a bit as well. [02:08:30] So there are some variables. Row number one basically is a variable column, okay? You can see age, gender, own home, merit, so on, so forth. [02:08:37] Now, these variables… Uh, this is a data of an online manufacturer, online, uh, retailer, a businessman who deals on… in online business, okay? [02:08:47] Now, what happens is that these are the customer… this is his customer data, okay? Now, what is his customer data? [02:08:52] Age of customer, old, middle, young, gender of customer. [02:08:57] Whether they have their own home, for that matter, okay? [02:08:58] And you are also looking at. salary of the customer, location, I'll come back to. [02:09:03] We're looking at salary of the customer, number of children they have in the household. [02:09:06] Uh, history of their buying in frequency in the last one year. [02:09:10] Number of catalogs you have sent to them? Okay. And… Also, uh… little bit in terms of… Also, rather the important part is amount spent. So, how much amount. [02:09:25] They have spent on your product so far. Should be in dollars, it's too small to be in rupees, of course. [02:09:32] should be in dollars, okay? So, this is what it is. [02:09:34] About the history… sorry, the one thing I left out is location. Location is… A competitor's business who has a physical store. I am doing online. [02:09:44] This competitor is actually having a physical store. What it means by a physical store is that, say, for example, I'm selling furniture. [02:09:49] I… online, I'm selling furniture on Amazon, for example. [02:09:54] Now, one of my customers, who is staying in a certain place. [02:09:56] He or she has a physical store of furniture just below his or her apartment. [02:10:03] So, they will rather go there and have a look rather than. [02:10:06] Buying online without having a look, right? On the other hand, if that… another person has a furniture store 3-4 kilometers away from his place. [02:10:14] He will have to take a… he will have to drive, or take a car, or auto, or a bus to go there. [02:10:20] Spend some time, come back. He will say, I'll do online. [02:10:22] Okay, so that's why this location far or close is very important for my business, okay? [02:10:28] Now, this is obviously data for regression models and all, which we are not going to do right now, but I have a question for you. [02:10:33] Clients will generally give you this kind of datasets. [02:10:37] And they'll say, give me analysis, measure of predictive modeling karke do. [02:10:40] They will not tell you what is target variable, what you do, and all. You have to frame it. Problem statement, you have to frame. [02:10:44] What I want you to do is to frame a problem statement, okay? Now, looking at these 9-10 variables, customer ID co-op chur sake, it is not very useful, although IDA. [02:10:54] But the rest of it. you think of it and tell me. [02:10:57] Out of these 10, which you think should be the why? [02:11:01] Or the target variable. It can be only one way, right? The rest of it will all be X or independent. [02:11:06] Which you think is logically… should be my target variable here, which I want to understand or explain. [02:11:15] amount spent, amount spent. [02:11:16] from thinking from the businessman's perspective. [02:11:17] Yup. [02:11:19] Okay, amount spent. Everyone thinks amount spent? [02:11:22] Yeah, like, from a revenue perspective, amount spent is important. [02:11:23] Okay, thank you. [02:11:26] Location can also. [02:11:28] salary. [02:11:29] Hmm… Okay, now we are getting into everything one by once. [02:11:34] Okay. Let's try to… see, again, I know that you are more in… I mean, you're thinking in a different way, different direction, and you are giving the answers, which is all fine. [02:11:42] Let's try to break it down in logic. Let's look at the first 7 columns. 1, 2, 3, 4, 5, 6, 7, up to children, okay? [02:11:51] Let me select it. Think that you are the businessman. You own this business, okay? [02:11:58] Do you think that you are the businessman, do you think that predicting any of this will be of any help to you? [02:12:04] Yeah. [02:12:05] Yes, maybe the salary… higher the salary, higher the spent part. [02:12:09] No, no, you are predicting it, you are not saying that I will… If you… even if you predict it, can you give this guy a more salary? [02:12:12] No, no. [02:12:14] No. [02:12:16] Say, can you say that? Old kumay yang banadunga. [02:12:19] Nay. [02:12:20] No. [02:12:22] Nate. [02:12:25] So, I can't do any of that, right? So these are none of my concerns. [02:12:32] Got it. [02:12:33] I am going to use them to explain that kon kitna melebe karjkar raha. [02:12:36] Okay. [02:12:37] Simple. How they're spending on my product. Right? Catalogs is something I sent. [02:12:43] I send my stock list to them on emails or whatever, right? [02:12:45] Uh-huh. [02:12:46] Now, I sent 6 mei ju, baraju atra, beheju. [02:12:50] Got it. [02:12:51] It's eye control 100%, right? Whereas history and amount spent. [02:12:54] Or something, I don't control 100%, either of them, but I want to control more of it, right? [02:13:00] Yeah. [02:13:01] So, in our data. When you have to form a problem statement, your target variable is never which you don't control at all. [02:13:12] Okay. [02:13:13] Or also never which you control fully. It's always that variable which you have some control, but want to improve your control. [02:13:14] is, yeah, yeah. [02:13:18] That is always going to be your target. Because business wants to know result about only those things. [02:13:29] Yes. [02:13:30] Yeah. [02:13:31] The rest don't care. Make sense? How you try to find it from, say, 100 variables, how to do? [02:13:33] That's the logical way of doing it. Because client will say, I don't know, I will put their hand up and say, I don't know, I've given you data, you do. You are a data scientist, you do. [02:13:41] That's… that's what we hear every day. So, that is where you have to logically eliminate and figure out what to do. [02:13:51] Was it understandable? That's actually a very critical part of the entire beginning itself. [02:13:55] is. Is. [02:13:57] Yes? Good. That was one thing I wanted to show you at the end of it, not EDA, really, but even before EDA, and probably more important. [02:14:03] Okay. Um, alright, I think I will stop here. [02:14:07] No, so, the answer is amount, is it? [02:14:09] Yeah. I'm serious amount, you are correct, yes, yes. [02:14:10] So, when you say that I have some control and not some control, too, how does a sum control come up in the amount spent? [02:14:16] some control to have, because if I say… some control as in the fact that I know that these people have been buying from me, so they'll probably continue buying from me. [02:14:23] Okay, okay. [02:14:24] Great. But I want to increase their buying amount. [02:14:27] Okay. [02:14:28] Tomorrow I want them to buy for 200, right? [02:14:31] Uh, okay. [02:14:33] So there you want to increase control in that sense. [02:14:34] Okay, okay, got it. [02:14:38] Okay. Now, one thing I would say, in the meantime, if you have any questions, just think of it quickly. One minute I'll spend on the initial question I have on the AI journey and stuff, right? [02:14:46] Um… see, again, obviously, I am sure you guys know this thing, and you are in this program, which is a Gen AI program, so you will… you will reach the LLM at some point. [02:14:56] Now, how you reach, how many times you stumble in the middle. [02:15:00] How many things you overlook, put under the carpet. [02:15:03] It's a different question, but it will reach, right? [02:15:07] So, I mean, that all of us do, there's no point, really, I mean, there's no point cribbing about it, that's always there. [02:15:14] Um… But… see, the journey is that in the last 7 years, I'll come back to you, Sriat and Navishek, don't worry, I'll come back to you. Keep your hands raised, I will remember, okay? [02:15:24] So, uh… In the last 7-8 years. [02:15:27] this field, especially the last 3-4 years, has changed drastically. [02:15:32] This program of Gen AI would not have been there if it was before COVID, right? This concept would. [02:15:36] It was not even existing, to be honest. So all of this came, if you go back 120 years, we had only t-test and z-test in statistics. [02:15:45] There was a guy called Dalton, there was a guy called Irving Fisher, who was doing. [02:15:50] practices of two-sample t-tests. Taking two datasets and trying to find out who is better, who is comparable to each other, and all that stuff, right? [02:15:58] Then in the early 1910s, they needed more variables. They could store more data, they could have more variables in that sense. [02:16:04] They went into something called ANOVA, another statistical concept, analysis of variance, okay? I'm sure some of you know of it. [02:16:10] That was for 3 variables, multiple and over. Then that was not enough. First World War came in. [02:16:14] They wanted to optimize results. In Europe, there was no food, there was nothing, no rationing. [02:16:19] So they said, okay, I have to optimize food for every family, how much meat they will get, how much rice they will get. [02:16:25] Yeah, if you have, you can drop off. We are done with the session, I'm just adding to the AI part. [02:16:29] So, they have to optimize, so they brought in regression. [02:16:33] Regulation is optimizing, finding Y with respect to X, okay? [02:16:38] That was there. Then in the… it was going on like this in the Second World War as well, when I'm talking about practical world impact of data science, or these ML models, okay? [02:16:46] Then in the 1970s, 73 to be precise, we had the oil shock. [02:16:50] Where the Middle East said to America, I will not sell you oil. [02:16:53] No matter what you do, I don't care, I'll not sell you oil. Simple, okay? But America said that, I mean, they knew that they needed oil. At that point, they were in a sort of cold war with Russia as well. [02:17:03] So they were in a tricky position, Ronald Reagan. [02:17:06] So he employed his Federal Reserve's chief economist, Milton Friedman, and brought up monetary policy. [02:17:12] opened up an institute called SAS, which is the precursor of Python. [02:17:16] Statistical analytical software, I'm sure some of you know. [02:17:20] It's 50 years next year. And the couple of interns who worked there in the beginning were named Bill Gates and Steve Jobs. [02:17:27] Okay, they were there in SaaS for a couple of years. [02:17:30] So, uh, anyway, so this boom came up, and in the 90s, when data storage became very cheap. [02:17:36] We got into proper machine learning, tree models, boosting, backing, decision trees. [02:17:41] clustering, uh, forecasting. And even your deep neural networks and NLPs and all come a little later, 2000s and all. [02:17:50] Why? The only reason this data became… storage became cheap. [02:17:53] Our phones store 2GB data. In 1990, storing 1GB was costing half the price of a Beetle car in New York. [02:17:59] Okay. So, that's the reason it all started. Then we have social media, Facebook, Twitter, WhatsApp, a lot of data coming in. We can scrape… web scraping is possible. [02:18:08] Then NVIDIA came up with GPUs in 2011, right? [02:18:12] And then the entire thing. sprang out, essentially. The multiple complicated deep learning models, non-numeric data, text, image, video. [02:18:21] Speech, all of this came up. And in 2017, we have this paper, seminal paper from Google Brain Team, called Attention is All You Need. [02:18:29] I'm sure you have heard of it. You can Google up, you will get the paper also. [02:18:33] It's public domain paper. And that brought up a concept called transformers, which I'm sure you will have towards the end of this program. [02:18:40] When it is there, maybe after a few months, of course. [02:18:43] Please make sure you understand transformers well. Because if you don't understand transformers well, you will not understand LLMs anything. [02:18:50] Because that is the gateway to large language text models. [02:18:55] And now we are here. We are into high-end LLMs. We are not building LLMs, no one can really, to be honest, except unless you are Google, Microsoft, Amazon, Apple. [02:19:04] Because it costs 40, 50 crores to build a LLM, okay? And not many companies have that power to build one model. [02:19:08] So, we are doing pre-trained LLMs. Google has built, Microsoft has built. [02:19:12] This model is already trained. We are using it as a kind of a cast on our data, and getting results out of it. That's other companies are doing, right? [02:19:21] And that is where we are. As I speak in 2025, and then agents came up in the last year. Agent, I will not talk a lot about it, because it's still experimental. [02:19:28] Agents have been failing most of the times, then passing, to be honest, okay? So it is still a thing developing. It's still not there. It started, right? [02:19:37] So, given that, this is where we are. So you need to be very good with your machine learning, you need to be very good with your deep learning and your NLP. [02:19:44] Natural language processing. If you have to survive and thrive and. [02:19:48] be successful in what we call as data science and AI today. [02:19:53] These 3 are the blocks, especially ML is your ground floor, DL is your first floor. [02:19:59] Along with the NLP. If your ground floor first float are weak, no matter how long a building you build, it will fall off. [02:20:04] I hope that gives you a sense of idea. Whoever asked that question, I forgot. [02:20:12] Okay. Cool. All right. Anyone has anything else? [02:20:13] a long sprint. [02:20:17] I think some of you raised your hands, sorry, yeah, go on. [02:20:20] Yeah, maybe I had to doubt, uh… [02:20:22] Actually, I would just need some clarification. So you showed a dataset. [02:20:23] Yep. [02:20:26] And we did learn, like, quite different concepts till now. [02:20:30] And I'm pretty sure, like, later we'll try to… we'll have a way to glue it together, but right now, can we just take us through the journey of how you would… [02:20:38] go about, like, from the raw data, uh, and then… [02:20:42] like, all the steps to come up with the model, finally. [02:20:43] Gotcha. Got it. Till the model you were saying, right? [02:20:46] Yeah, there'll be a lot of back and forth, but still, like, if you can give a broad… [02:20:47] Hmm. Yeah, yeah, yeah. So, so, see, first you get the data. [02:20:52] Then the second thing is what I just showed you. You find… you form the problem statement on the target variable, right? [02:20:58] Once you are done with that, obviously, in discussion with your stakeholders, you cannot do it alone. [02:21:04] Once you are done with that. You try to start the data. The first step, generally, is to run a describe function. [02:21:11] If I talk in Python terms. run a describe function. [02:21:15] you get these results I was showing you on Excel, right? All these statistical results. [02:21:19] Get a sense of how the data is looking. [02:21:22] What the distributions look like. If I had a little more time, I could have shown you some Python output on the describe function. [02:21:27] But, uh, right now, we don't have… maybe we'll catch up with some other time and we'll do that. [02:21:33] Look at distributions, try to understand the distributions, how good they are, how bad they are from those results, okay? [02:21:38] That will give you a lot of output. And remember, when you're doing ED8, take a pen and paper. [02:21:42] Beside you. With you. Don't… every single step you run, and all the insights you get from there. [02:21:49] 2, 1, 2, 3, 4, write it down. If you don't write it down later on, it all will be all over the place. You will not remember what you did where, and what connects to whom, what doesn't connect to whom. [02:21:58] It's all a fish market, right? So, right, keep writing down. Then, once you do that, you've. [02:22:04] Go to the univariate analysis. raw histograms, draw box plots if required. [02:22:10] Do the unified analysis, cross steps if required for categorical data, all of this stuff, okay? [02:22:16] Then do bivariate scatterplots, heat maps, and multivariate pair plots if required. [02:22:19] to the plotting part, from univariate to multivariate, okay? [02:22:23] get more insights, correlate them with your. Describe function insights, see where you are matching, fantastic. If you're not matching, re-look into it. [02:22:31] This part is done. Then you go into outliers and missing values. [02:22:36] Treat them, figure them out. If required, treat them. If not required, keep them as it is, doesn't matter, right? [02:22:41] Deal with that part. Once you are done with that part. [02:22:44] You go and try to see if any. Anumalese transformation, any changes are required in the future engineering sense, right? It's not mandatory. Every time you will not require that. [02:22:54] But you see, if required, you have to do it. [02:22:56] And then you come to the splitting of data. [02:22:59] Which I was telling you earlier, split the data into train and test. [02:23:03] Before that, there's a thing called dummy variable creation, which you haven't done yet, which is converting string or categorical, or object variable to numeric data, which you will… I'm sure you will be learning soon. [02:23:13] Maybe tomorrow itself. Um… Which is part of feature engineering, which is very important, because unless you put it into numbers, you can't use it, right? So, obviously, there is no point keeping it as objects or string. [02:23:22] Once you're done with that. split data into train and test. [02:23:26] Then do scaling on that, scaling on the numeric variables. [02:23:30] For the categorical data, you already have dummy variables. [02:23:33] scale the numeric, coming variables on categorical, concatenate them. [02:23:37] On the train data, and the test data, separately. [02:23:41] Now, this trend data you have after concatenation of. [02:23:44] categorical dummies, and your numerical scaled. This is the data which is ready for your first iteration of model. [02:23:52] This is the entire step you have to go through. [02:23:57] Yeah, actually, that is a good one. [02:23:58] Does it make sense? [02:24:01] Yeah, and then after the model, there's another… another war. This is the first war, First World War, Second World War will start after that. [02:24:04] Hmm. [02:24:08] Okay, uh, then we run the test, right? And then we get the, uh, fetration, or… [02:24:13] No, no, no, no, no, you first run the train, okay? You first run the full model on train. [02:24:21] Oh… [02:24:22] And then you have to do feature selection, there are plenty of things. Any model will have a lot of yardsticks to check, this criteria, that here, there. [02:24:25] You have to do a lot of up and down in the model as well. [02:24:27] And if it doesn't work, you have to go back to EDA, come back to model, up, down, up, down, many times. Once you are finalized that this is your train model, which is good enough, fixed. [02:24:36] Then you compare with test. Not initially. [02:24:42] Also, we do run the prediction, right, to understand whether… [02:24:47] Okay. [02:24:48] Later on, later on, later on, later on. At least make sure that your variables are fixed for your test, for your train. [02:24:50] Okay. Okay, okay. [02:24:53] Yeah. Yeah, G2, go on. [02:24:59] Yep. [02:25:00] Hello? Yeah, this is indeed related to the topic you thought, but I have faced an issue. When the data was around millions, [02:25:05] I wasn't able to load the data. [02:25:06] Hmm. [02:25:09] Yeah, before I told you we're up. [02:25:10] Yeah. How many GB? [02:25:14] In Colab, in collab, eh? [02:25:17] Attending a competition? [02:25:18] No, then you do one thing, then you take a licensed version of collab. It costs 1,000 rupees a month. [02:25:23] Okay. In the license version of Collab, you will have a… they will give you a GPU, 4 GPUs, parallel four parallel computing. You can use that. [02:25:33] uh, like, uh, told before that, uh, we… if we can channel it and go, like, 10,000 or 50,000, do we need to go with that approach? [02:25:42] Or, uh, be… [02:25:43] You can, you… You can, but see, the only thing is that you just… please make sure that your 50,000 you are taking, the sample you are taking. [02:25:50] is exactly representing your population. From a distribution perspective of across all the variables. [02:25:57] That is very important. Otherwise, there is no point. [02:26:03] So, uh, so even if we chunk the data, you know, come to the conclusion for aDA, right? [02:26:07] You can. You can do it on a smaller data on ED and everything on the model also, but that will be a pilot. You cannot say that I'm done for the entire thing. [02:26:16] For the entire thing, you have to use either a GPU, or you have to probably use a cloud support. [02:26:24] So, so the industry… [02:26:28] you know. [02:26:29] Your voice is breaking a little bit, I guess. I don't think if it's for others, but I think it's breaking a bit for me. [02:26:33] Yeah, I will raise the question in the… [02:26:38] So, [02:26:41] I have some internet issue now. I will raise the question. [02:26:45] Okay, okay, sure, yeah, I think that's why it's breaking a bit, yeah. [02:26:52] My con… can't we use data reduction techniques in this, I mean, instead of, you know, having a whole big set. [02:26:59] Maybe, you know, we can have a small side there. [02:27:00] Yeah, so… Yes, Sarat, yes, yes, of course, yes. But that also has its own plus and minus. [02:27:08] Which is a discussion for another 30 minutes. [02:27:14] I can't go into it right now. I can't go into it right now. [02:27:18] But, yes, but yes, the answer, short answer is yes, but with a lot of caveats. [02:27:24] It's not a resounding universal yes, no. [02:27:31] Okay. Anyone else has anything as a short question rather than a discussion point, because we really don't have time for that. [02:27:39] Yeah, hi Mike. So, question is related to… so, when do you make sure that your model is accurate? [02:27:40] Yeah. [02:27:45] So it depends on some set of iterations we do, right? [02:27:46] Hmm… Yeah, not only iterations, it's about some yardsticks you have to maintain. Say, for example, on certain level of accuracy, you have to. [02:27:49] So… [02:27:54] Attain, right? A certain level of error, you have to. [02:27:55] Yeah. Mm-hmm. [02:27:57] bring down it too, right? In those yardsticks, you have to comply, basically. [02:28:02] But again, these yardsticks are not fixed for every model, right? [02:28:03] Okay. Yeah, yeah. [02:28:06] So they will be changing from model to model. [02:28:10] So that is not… so, see, that's the beauty of data science, right? [02:28:14] So, what one model you do today, you will never do ever in your life again. [02:28:15] Hmm… Yeah. [02:28:18] Everything is unique. So every yardstick is also unique. [02:28:28] Okay, um, I hope you… knew a few things, you tried to… what I tried is that you… for you to connect a few things. I hope that has happened to some extent, at least, if not entirely. [02:28:41] Uh, and more than that, what I hope is that you sort of. [02:28:44] Enjoyed the part. That is most important for me, okay? You enjoyed the introduction, and. [02:28:49] sort of take some leaves out of it, okay? [02:28:53] I hope that has happened. [02:28:57] Yes, thank you. [02:28:58] Thanks. [02:28:59] Yeah. It is, thanks. [02:29:01] Thank you. [02:29:02] All right. Ah, okay, all right, good, good to know that. Thank you. Then, all right, we'll log off for now, and yeah, thanks, and all the best. [02:29:07] Thank you. [02:29:10] Thank you. [02:29:14] Thank youThank you