# 04 2025-11-16 Hands On Data Pre Processing

course: Module 1 — Foundations of AI & ML
module: Module-1-Foundations-AI-ML
date: 2025-11-16
type: transcript
video_url: https://personal-learn.armco.dev/files/_Recordings/Module-1-Foundations-AI-ML/04_2025-11-16_Hands_On_Data_Pre_Processing.mp4

---
[00:24:41] here you can see, uh, like, for imputer also, they have given a…
[00:24:46] specific, uh, like a pack, specific function is given for, uh, like, a computer, then if you want to do PCA, then a specific function is given. If you want to do, like, if you want to, like, normalize the dataset to have, like, we want to… each of our columns to have,
[00:25:05] Zero mean and unit variance. So, uh, so otherwise, we would have to, like,
[00:25:11] Write code for each of these functions separately for doing PCA, for doing preprocessing, for imputation also.
[00:25:19] So, all of these things are, like, they have implemented it, and here we are just shown 3 of them, but even then, all the machine learning models are also, like,
[00:25:30] pre… like, inside this skill and package. So, it's a very handy package if you want, uh, like, it's a set of standard if you want to do machine learning.
[00:25:41] Then for plotting the graphs, there is a famous package matplot, and uh…
[00:25:47] And to, like, make the… these graphs aesthetically pleasing and all, to make them look better, there is a library Seabone.
[00:25:57] So, these are the libraries used.
[00:26:01] Uh, library packet is used.
[00:26:05] Uh, if people are… if you have low, uh, loaded the files, we can then proceed with the code. Have you people loaded the file?
[00:26:15] Yeah, thank you for the confirmation. Swagat.
[00:26:16] Yeah, it got loaded.
[00:26:19] Okay, so…
[00:26:20] Uh, just a second, uh, any questions, anything, any queries so far?
[00:26:26] Because…
[00:26:27] Uh, ma'am, when I upload the CSG file, uh, it's actually loading the California housing data set.
[00:26:36] Yeah, I uploaded this, uh, CSV file. But when I opened the sample data folder.
[00:26:41] It's showing the California housing data.
[00:26:42] No, no, no, no, you don't have to open the sample data folder. When you upload it, just run this first block, loading the dataset. It will, uh, here, it is like a… Sam, you are opening a sample data folder, this is another file, these are from, by default, from Collab.
[00:26:58] Okay, okay. Thanks. Thanks for explaining.
[00:27:01] Any other query?
[00:27:03] Uh, can you please share the Excel again? I think I'm not able to download that.
[00:27:21] I think it is available in LMS.
[00:27:28] So the link is there, um…
[00:27:30] on the LMS, it's also there on the chat.
[00:27:33] For those who would not see it on the LMS.
[00:27:36] On the link of all the files.
[00:27:42] Yeah, it's LMS, actually. I think KMS is being returned, uh, it's LMS.
[00:27:48] Yep.
[00:27:50] Also, in LMS, I think the file name has a, uh… Parent, this is one maybe we might have to rename that, so, because the name given in the.
[00:27:58] Okay.
[00:27:59] Python file is… A little bit different, yeah.
[00:28:01] Okay, that might be, uh, that might be due to…
[00:28:04] Probably due to, uh, download and then…
[00:28:09] Maybe multiple, so you can please rename it so that you don't have any.
[00:28:12] problem.
[00:28:13] Yes, after renaming it works.
[00:28:18] Anyone else? Any other query or issues?
[00:28:20] Uh, do we… So we need to upload Python notebook as well, or we would be creating the one right now.
[00:28:30] So, you can just, uh, upload and open this Python code.
[00:28:36] Uh, so you won't be writing the code, uh, line by line here?
[00:28:37] Okay.
[00:28:41] Because the code file is very large,
[00:28:42] Mm-hmm.
[00:28:44] And, uh, we can definitely have you write statement by statement, but then…
[00:28:49] Many of the concepts will not be covered.
[00:28:52] So, what we do is that we pre-prepare the code for you,
[00:28:56] And we walk you through the code.
[00:28:59] And, uh, any queries you have, so that helps us to do better coverage.
[00:29:06] Yeah.
[00:29:09] In the chat, also, I've shared the CSV file.
[00:29:12] If anybody wants to directly download.
[00:29:18] Shall we begin now?
[00:29:19] Yes, ma'am.
[00:29:22] Sorry to interrupt, uh, it is just showing California housing and the other.
[00:29:23] Sorry, sorry to interrupt.
[00:29:24] Yes, ma'am.
[00:29:27] File, it is not showing.
[00:29:28] So if… so if you have a… if you go over, like, if you click over this folder and do an upload over from a local system, wherever the file is,
[00:29:37] So, it will finally show a file like this, below sample data.
[00:29:44] If, uh, similar to mine.
[00:29:48] And when you click over it,
[00:29:49] Okay.
[00:29:52] Okay, okay. So, issue of, uh, name is there.
[00:29:53] When I renamed it worked.
[00:29:55] And when you, like, open it, that it will also open in the…
[00:29:59] collab like this.
[00:30:04] Uh, yeah, kindly rename the file, that might be due to…
[00:30:08] download and upload on LMS that…
[00:30:10] Maybe twice it was downloaded. So please rename it to the…
[00:30:14] Name as in the code.
[00:30:17] Otherwise, it's not going to work.
[00:30:22] Okay, shall we start now? The coding?
[00:30:26] Sorry, Dilip here, I just need help. I think I messed up with the folder. I am not able to go back.
[00:30:32] So, can I share, and…
[00:30:34] Uh, so the link to the lab materials, it's there on the chat, Dile.
[00:30:40] You can… no problem, you can go over to the link again.
[00:30:44] And re-download the folder.
[00:30:46] Is it okay?
[00:30:47] I'm sorry, I wasn't talking about the link or the folder. When I upload the sample data.
[00:30:52] on the… on the folder. Just a second. That's it.
[00:31:05] Uh, so Dilip, we are already sharing something.
[00:31:10] Uh… I think it may disrupt the flow of the class, if you'll now…
[00:31:16] Prior to present.
[00:31:27] Uh, I request Simran to please, uh…
[00:31:30] Uh, resolve the issue with the lip.
[00:31:36] Sure, ma'am.
[00:31:52] Well, uh, like, let me show you once again, like, how we will… so this is, like, a normal… I've, like, deleted my file, so how do I load it, uh, whenever I'm in collab notebook?
[00:32:04] I will go to this, uh, files here.
[00:32:07] Uh, Ashish, I request you to just close this, and you can just open it right from scratch for the benefit of people who haven't…
[00:32:16] Uh, open. You can just…
[00:32:18] do it afresh.
[00:32:21] Okay.
[00:32:42] So, uh, I hope my screen is, uh…
[00:32:44] tab is visible.
[00:32:48] Yes.
[00:32:49] Yeah. So, uh, this is, like, a fresh notebook, it is not, uh, so, it is not connected, otherwise, uh, it would have shown here connected. It is, like, connect. So, what I do is, when I'm on this folder here,
[00:33:01] When I moved to my first package, if I will directly run this command,
[00:33:06] It will show me an error, because file is not loaded.
[00:33:11] Or it is this path file does not exist in path. Because right now, the file is not loaded. So what I will do is, there are certain options here on the… there are, like, 5 options.
[00:33:21] Below the key, there is this option, Files.
[00:33:25] I open it. Right now, there is a folder sample data. I have not… we will not go inside sample data. This is by default in Colab.
[00:33:34] So, we will not touch it.
[00:33:36] We'll go this upward here, this first icon here, which is telling to upload to session storage.
[00:33:46] So, when I click over it, uh, it shows me my local files.
[00:33:51] And from there, I, like, upload this auto-MPG.
[00:33:55] Once it is uploaded, below the sample data folder, you will see AutoMPG file.
[00:34:02] And because it's like a CSV file, it can be directly… if you click over it, it can be directly loaded in the Coolab, and it will look like this.
[00:34:13] And now, when you, like…
[00:34:15] run this block. So, it runs successfully.
[00:34:24] So, uh, is everyone, uh, ready?
[00:34:27] Okay, Rajkumar, what is the issue?
[00:34:31] I'm seeing the file got uploaded at the bottom, but it's not showing under.
[00:34:36] sample, below the sample data.
[00:34:38] But are you… when you run this block, are you able to get this output?
[00:34:42] No, no. It's showing below, but not the top at the top.
[00:34:44] Yeah, maybe…
[00:34:48] So, if your file is there, like, you can copy-pa, you can do… can do a copy path.
[00:34:49] When I'm…
[00:34:55] No, it's not shown over here. shown at the bottom.
[00:34:59] at the bottom, like…
[00:35:01] And the files itself, the disk is there, right?
[00:35:04] Where you can see at the top of the disk, I have… Below, below.
[00:35:08] So your file is inside this, na? It's inside sample.
[00:35:15] And you'll put…
[00:35:16] No, no, no, 69.5 is complete. It's very, I'm able to see.
[00:35:23] Uh, look, this is a…
[00:35:24] It's not in the top, actually.
[00:35:25] And… good.
[00:35:26] So, um, I request Rajkuman, we also have the pre-run, uh, output file. Uh, this is a PDF file.
[00:35:35] Just for reference and for sake of all the other learners,
[00:35:41] Uh, if I request you to, uh, probably you can go through. Sims and I request you to…
[00:35:46] Uh, please, uh, coordinate with Raj Kumar. I think he's facing some issue in loading up the file.
[00:35:55] So that we can start… is it okay, Rajkumar, so that, uh…
[00:35:58] Yes. Fine, fine.
[00:35:59] You can… yeah, you can open the PDF which is containing the output, uh, whatever, and all the code, everything is there.
[00:36:07] So, you can go through, and offline, you can…
[00:36:11] Uh, trial run and see.
[00:36:12] Okay.
[00:36:13] Okay.
[00:36:15] Yeah.
[00:36:16] Okay. I love it.
[00:36:17] Others are… are there anyone, uh, are somebody able to load it successfully, running for others?
[00:36:25] So, uh… so…
[00:36:26] Yep.
[00:36:27] Yeah, it's working fine. I think we can start.
[00:36:28] Okay, cool.
[00:36:29] You can…
[00:36:30] Yeah, it's just… this weekend run. Yes, we're gonna start, actually.
[00:36:32] Okay, okay.
[00:36:33] Running from… it's running for me.
[00:36:34] Yeah. So, for, uh, those, uh, some, uh, people,
[00:36:36] If it is, uh, not loading, or you are facing an issue, I request you to please open the PDF file.
[00:36:44] That is a pre-run output. It contains code and output.
[00:36:47] It will help you to understand what it is, so that we can start the session.
[00:36:51] And offline, you can fix your issue.
[00:36:54] Okay, Ashisha request you to start, please.
[00:36:57] Okay. So, this is the first code block, which is actually running, and these all are, like,
[00:37:05] This is all, like, text, this is not, like, Python block. This is the first block, which is, like, running.
[00:37:11] So, uh, so basically, I… in order to load this file, CSV file,
[00:37:17] I'm first importing Panda's package, and import Panda's as PD, and then…
[00:37:24] NumPy is not just required here, but yeah, it has been written here, because in future, it might be needed. And, uh, yeah, and sometimes warnings appear in…
[00:37:33] in the code, like, due to, like, a obsolete version of package we are using and everything, so not to get confused and everything, so, uh, we, we have suppressed the warning using import warnings.
[00:37:47] And so, in the entire notebook, uh, warnings don't appear anywhere. So, this is this thing. Now, the actual loading of the file. So, this is the function readcsv.
[00:37:57] You just give it the path of your file. So here, it is in the, like, the… since the file is in the current working directory, so just I give the name of the file.
[00:38:08] So, otherwise you will have… otherwise, the full path, wherever the file is, we give it here in the string.
[00:38:14] And then, uh, it is a variable, it, uh, here I have named this data frame, but it can be, like, named anything.
[00:38:21] So, basically, this variable will now contain my data frame when I do pd.readCSV, and when I run over it,
[00:38:30] I get, uh… so, it is showing me first 5 values, and it is showing me last 5 values. First 5 head values and last 5 tail values.
[00:38:40] So, just by running it, we can see that there are
[00:38:44] 398 rows and 9 columns.
[00:38:46] And because there are only 9 columns, so it is, like, totally visible in this notebook, but in some dataset, there can be, like, 19, 20 columns, so sometime you will have to, like, scroll over it to get… see all the columns and everything. And uh…
[00:39:01] So, yeah. So…
[00:39:03] Yeah, so if, uh, this is 3989 column, this information
[00:39:09] can also be, like, printed if I do a… on this variable, if I do a df.shape,
[00:39:15] Again, it's the same information, so it's a two-dimensional data.
[00:39:19] where, like, I have 398 samples, and for each sample, there are, like, 9 columns.
[00:39:26] as such. And, uh…
[00:39:28] Then, like, the first thing, whatever we do is, when we have a tabular data, normally,
[00:39:35] We… in Pandas, we do a df.info. It… the df.info
[00:39:41] It gives me, uh, like, a count of non-null values and the data type of each of the
[00:39:48] columns which are there in my, uh, like, data frame.
[00:39:52] So, you see, uh, here it is telling me that there are 398 entries. Again, the same thing.
[00:39:56] And there are 9 columns, and for each of the columns,
[00:40:01] It is telling me, uh, like, there are non-null entries. So, basically, if there… if there would be… have been some missing values,
[00:40:09] Uh, like, some missing values would be there. So, in some entries, uh, the non-null values would be… have been less.
[00:40:18] But right now, in the dataset, there are no null values, so it is showing, uh, that, uh, for all columns.
[00:40:25] There is, like, all values are non-null, so that's why, for all the value, the entry is same.
[00:40:32] Right now. Then, it is…
[00:40:35] showing me the data type for each of these, uh, feature… of the columns.
[00:40:42] So, uh, the first observation is, like, that there are no missing values in the dataset, but
[00:40:49] The other, uh, information is that the other thing to point out here is,
[00:40:54] That sometimes, this can be like a… like, it can be a misnomer, uh, if we just go by what it is saying. Because, you see, uh, here,
[00:41:05] The car… it is saying that car name is an object data type. Object is basically, like, it is saying it's a categorical data type, because car name is in string.
[00:41:14] But for other data types, like origin, uh, like I have shown in the data card that
[00:41:20] Origin had only 3 values, 1, 2, and 3, so it is basically categorical.
[00:41:25] But because, since it is stored in numbers,
[00:41:28] Pandas is, uh, like, treating it as an integer, integer type.
[00:41:34] So, uh, that means we can't directly rely on, like, Pandas to, like, be totally sure of that, uh,
[00:41:43] whatever it is, like, written as integer here, it may be a categorical data type also. So, the first thing to note here is this, because
[00:41:54] As such, this, uh,
[00:41:56] this… even this MPG high, there are only two values for MPG high, 1 and 0. But because these are, uh, these values are stored in numbers, 1 and 0, so it is treating it as an integer.
[00:42:07] So, in order to, like, to be pretty sure that, uh, like, these values which, uh, like, which are, which of them are categorical and which of them are not, what I've done is,
[00:42:19] So there are 9 call… so there are 9, uh, columns.
[00:42:23] And for each column, I have printed the unique values associated with each column.
[00:42:31] So, you see, cylinders has 5 unique values only. Displacement has 82 horsepower is 94.
[00:42:36] Uh, weight is in 351.
[00:42:39] So, uh, among them, uh, there are, uh, like,
[00:42:43] We can say, like, how the dataset usually treats is that, uh, displacement,
[00:42:49] Horsepower.
[00:42:51] acceleration, and uh… like, one more values, uh, there. So, yeah, there are… I will, like, show in the, yeah. So, displacement, horsepower, weight, and acceleration, these are the…
[00:43:07] numeric columns, because there are so many unique values for that, we are considering.
[00:43:13] Other than that, all of them have, like, a limited number of unique values, uh…
[00:43:16] So, these are, uh, basically categorical features.
[00:43:21] So, one thing, uh, so the first thing we, like, encountered is that this can be deceptive.
[00:43:27] the data type shown here, because it is only treating, uh, like, a categorical variable and object data type to only the variable which is in string. So, you see the car name here.
[00:43:39] carnivision string, that's why it is giving its object.
[00:43:44] Other than that, for every data type, because… for every column, it is stored in numbers, so it is saying either integer or float.
[00:43:51] So, float is obviously, if something is float, we will say it's a continuous numeric feature.
[00:43:56] But for integers, it can be, like, there are only some few unique values.
[00:44:02] So, uh, so basically, uh, what we will go with, that displacement, horsepower,
[00:44:10] weight and acceleration. So, displacement is 82 unique values, horsepower is 94, weight is having 351, and acceleration 95.
[00:44:19] Hmm. So these are, like, four numeric columns for further analysis, and rest of all, are categorical.
[00:44:26] And MPG high is categorical, but it's a target variable, it is not an input feature that we consider. It's a, like, that's something we want to predict.
[00:44:33] So, this… this separates it out from other…
[00:44:38] columns.
[00:44:40] So, is this thing clear to everyone? Like, then we can proceed to the missing value analysis.
[00:44:48] Okay, so…
[00:44:51] Yeah, so, uh, the first thing is missing, uh…
[00:44:53] value analysis, and in order to do it, because there are, as such, no missing values in the dataset,
[00:45:00] But we want to do a missing value analysis, so, uh, for this, uh, just for the sake of learning here,
[00:45:07] We are randomly inserting missing values here to show how it works.
[00:45:13] But otherwise, in a dataset,
[00:45:15] like, if in the starting we'll do a DF.info, or we see that the missing values, then we proceed with these strategies.
[00:45:24] Um, in this dataset, there are no such missing values, so we manually
[00:45:27] insert them.
[00:45:29] So, these are my four numeric columns. The studies which I are showing here, they work only for numeric column, for inserting missing values, so that's why I've chosen only these four numeric columns. Others are categorical.
[00:45:46] So, I just created a copy of that original dataset here.
[00:45:49] And, uh, like, this is np.random.seat written here, so it's for fixing randomness, so that, uh…
[00:45:55] Because we are randomly inserting missing values in each column, so at every run,
[00:46:02] some other entry can, like, be missing, can be shown as missing, but we want it for every run.
[00:46:09] that the randomness to be fixed, that if column… if row 1 is showing… so basically, but by fixing the random C to be mean is,
[00:46:19] Uh, that if, uh, for this sample 0 here, displacement is showing NAN,
[00:46:25] Now, for multiple runs of this, whenever I run it again and again, uh, this particular entry should show NN only.
[00:46:33] and others should not show. So, that way we can fix that, uh, the randomness, and it will be easier for us. Otherwise, if it would have been not been there,
[00:46:43] Uh, but right now, since it is written, so it… yeah, so you can see that, uh,
[00:46:49] We are not able to control the randomness now, that NAN can occur at any of the… any of the… so NAN is right now here.
[00:46:57] And I… when I run it again…
[00:47:00] And then we'll be at some other position. Right now, in this first 20 column, NAN is not even here.
[00:47:05] So, yeah, so in order to control the randomness,
[00:47:09] We are doing, uh… we usually fix a seed. We usually give our integer number here. It's np.random.seed, and you can give any number. It's not just, like, 42, but if you give 42,
[00:47:21] You will see the same
[00:47:24] any n values, like, uh, in the notebook shown here.
[00:47:27] Which I'm showing.
[00:47:29] So yeah. So, this is fixing of random seed. And then, uh, what I'm doing is…
[00:47:34] Assist, uh, sorry, uh, oh.
[00:47:36] how this 42 will, uh, interpret in…
[00:47:40] in the code.
[00:47:42] Like…
[00:47:44] Like, it's just an integer number, I can… could use 2000 also, something else would have occurred. I'm saying that by using 42,
[00:47:53] If you will also use 42. So, basically, the results will be same in your notebook also, that this first, uh, this first sample will have, in the displacement, will have NN here.
[00:48:03] Mm-hmm.
[00:48:05] Uh, let me add here…
[00:48:07] So, what happens is that…
[00:48:08] Even when we have the same code and the same dataset,
[00:48:12] If you run it multiple times, what can happen is that randomly, uh…
[00:48:17] Based on different, uh, running or implementation things, which we'll discuss in more details as we go along in the lectures.
[00:48:26] Uh, the output could be a little different.
[00:48:31] Not substantially different, but it could be different. For example,
[00:48:35] Uh, you calculate, uh, the percentage or performance, it can vary.
[00:48:39] Slightly. So, what can happen is that sometimes if you want to present these results,
[00:48:45] Then every time, you'll have to run and see it again and again.
[00:48:49] So if you want the results to be consistent, though substantial difference may not be there.
[00:48:55] But if you want the results to be consistent,
[00:48:58] Uh, across all your runs, across multiple people,
[00:49:02] Then we just put a certain seed. It's called a random seed. It could have any value.
[00:49:09] And once you put a fixed seed, then…
[00:49:12] The data that is, uh…
[00:49:14] being chosen to do the run.
[00:49:16] is going to be chosen in a certain way always, and therefore, the results that you are going to have
[00:49:22] Uh, will be, uh…
[00:49:24] will be consistent.
[00:49:27] No matter how many times you do the run.
[00:49:29] So that's the reason of, uh…
[00:49:31] putting that seat. If you don't put the seed also, it will run.
[00:49:37] And, uh, the results will be…
[00:49:38] Uh, correct only, only there'll be some minor differences.
[00:49:42] That's the reason. Is it okay?
[00:49:44] Yeah, man. Yeah.
[00:49:54] Nothing.
[00:49:55] No arbitrates are arbitrary.
[00:49:56] Oh, Mom, when we are choosing this, uh, value of the seat, 42, like, are we considering anything while choosing the value, or… Okay.
[00:49:57] You can take it 100, zero, anything, but whatever you choose, that will be fixed.
[00:50:00] Okay. Okay.
[00:50:03] Yeah. Okay, Ashish, over to you.
[00:50:06] So, uh, so basically, in this code, what I will do is, for each of the four columns,
[00:50:12] I will loop through them, and for each of the columns, randomly 10% of the
[00:50:17] their values will be, uh, will be, like, we will make them NAN. Like, we will call… we will, like, make them missing value.
[00:50:28] So, this is the actual code where this happens. So, I'm looping through my numeric calls, which is, like, the name of these columns. There's a list.
[00:50:37] And then, uh, whatever is the length of my data set, so if you see…
[00:50:44] Uh, dear, I do a DF10, it is the same thing, 398.
[00:50:50] So, among 398, I'm selecting a size of, like, 10% size of it. That's why this is 0.1.
[00:50:57] And then I'm randomly, from this, uh, like, uh, point, 10% of them.
[00:51:03] So, I'm… for this random… from this, uh, like, 10% of the size, I'm randomly selecting indices,
[00:51:10] For, uh, for which… which will be, like, treated as missing. So, once I have those indices, I will… for those values, I will insert NAN. So, this is the code for that, and that's why, after, like, running this code,
[00:51:24] Uh, you see that some values are NAN.
[00:51:28] So, because the choice of what will be the, uh, what will be the 10% values that have, uh, that will be chosen, that can differ, that's why, uh, like, if you use some other seed value, it will show you some
[00:51:42] As a result, uh, here. So, yeah, so this has changed that, uh, there are 42, this was an right now, NAN for displacement is on 16th, so yeah. So, 10% of the value will still be treated as NAN, but…
[00:51:55] Uh, their locations can differ. So, in order to, like, be consistent with everyone, when we present results, we choose the same value, so even when they run our code, they see the same
[00:52:07] output as V, as we do.
[00:52:09] So yeah, so this is the first thing that we are creating, uh, like, for each of the numeric column, we are creating, manually inserting 10% entries in each of the column as missing values.
[00:52:23] So, now, if you will do again, and, like,
[00:52:27] df.df underscore 10.info.
[00:52:33] So, now you see that earlier, there were only 395 non-linear entries, but for the 4 columns, now, because 10% have been missing, there is only 359 non-null, it is saying.
[00:52:47] For horsepower, it is also saying 315 and non-nell.
[00:52:49] So now, Pandas is able to detect that there are missing values. Uh, earlier there were none.
[00:52:55] you know, some doubt was there.
[00:52:58] Yeah, uh, so there were… there were no missing values earlier, so where it is putting this NEN.
[00:53:05] So, this… so consider this, this is a column.
[00:53:09] It is taking all of these, like, 398 samples are there of this column.
[00:53:15] So, randomly selecting some indices from the C398, 10% of indices, and over that, it is like putting this, uh, overwriting this NAN, whatever its true value, it is over that, it is overwriting NAL.
[00:53:28] And this is doing for all the four columns.
[00:53:31] So these are not the missing values, but it's just overwriting the numbers.
[00:53:35] It's overriding.
[00:53:36] Yeah, yeah, overwriting over the actual values. We have manually created them to, like, teach the imputation strategies. There were no missing values in this dataset originally.
[00:53:48] So,
[00:53:49] Uh, let me add here, Ashi, sorry to interrupt.
[00:53:50] So, what is the purpose of that?
[00:53:52] So this whole exercise is on data pre-processing, specifically
[00:53:58] Uh, filling missing values.
[00:54:00] So, uh, the dataset that we chose, uh, which was ideal for this, is this particular data, as Ashish has discussed.
[00:54:09] But, uh, this data doesn't contain any missing values, then how do we illustrate to you that, uh, what code to write to…
[00:54:19] Uh, you know, do missing value imputation. So, for that, we are artificially creating some missing values.
[00:54:26] In the given data, and then we will…
[00:54:29] See, our exercise requires different percentages of missing values, like maybe 10%, 20%, 30%, like that.
[00:54:38] So, according to what we want to illustrate, what we will do is we'll create artificially those many missing values.
[00:54:46] And then, uh, illustrate to you, with the help of code, whatever you learned in theory,
[00:54:52] that how missing value imputation will be done,
[00:54:56] Uh, using the different strategies that you have studied.
[00:55:00] And, uh, when you vary, uh, this with, uh, different percentages of missing value,
[00:55:07] How these different techniques get impacted and all that.
[00:55:10] So, to do that, we'll actually need different amounts of missing values.
[00:55:15] Uh, and so on and so forth, so that's why we are artificially creating them.
[00:55:21] Actually, in a real-world data,
[00:55:24] Most of the times, you do have missing values,
[00:55:28] And, therefore, uh, you'll need to treat them. So, this is actually a way to…
[00:55:34] show you how to treat missing values by creating those many missing values and then treating them.
[00:55:40] Is it okay?
[00:55:44] Yeah.
[00:55:45] Okay, okay, got it. Thank you.
[00:55:46] Uh, month out, I see. Uh, we are inputting… missing values, uh, for the part of code, which is below the random seed line, right?
[00:55:55] Yeah.
[00:55:57] Note we are dealing with the, uh, seed which we are putting in.
[00:56:03] Yeah, that's separate.
[00:56:04] That's separate, right? Okay.
[00:56:07] The actual part is this only, uh, this is just for fixing the seat, yeah. Only this part is for it.
[00:56:12] Mm-hmm, okay.
[00:56:16] Okay, so… yeah, so we have now DF underscore 10, and…
[00:56:21] which is, uh, like, having 10% of the values have been…
[00:56:26] made missing in each of the four columns of this numeric.
[00:56:30] And now, since we… the data said there are missing values.
[00:56:33] We can discuss the, uh…
[00:56:35] It'll extend the impression strategies which are applied when they are missing values.
[00:56:40] So, like, there are four strategies.
[00:56:43] to how to replace these missing values. So, the first is that you replace the missing values by taking mean of each of the columns.
[00:56:53] So, these studies are applied column-wise, so over the entire column, you take the…
[00:56:58] whatever the mean of the remaining values, and…
[00:57:00] For each of the NAN values, you replace it with the mean.
[00:57:04] Similarly, you can replace the NAN values, the missing values, with the median for each of the numeric column.
[00:57:10] And, uh, you can also replace it with its mode.
[00:57:14] the most frequent value. And this mode value can also work for if there are categorical features where there is missing. It can also work for that thing.
[00:57:23] But mean and median, uh, they usually work with.
[00:57:26] numerical… they are… they will work with numeric columns only.
[00:57:29] And the last strategy is, like, class-wise mode.
[00:57:33] So, the thing is, so basically, classifies mean that, uh, because there are two classes here, so we will take a subset of the class, uh, either 1 or 0, and within each of the subset,
[00:57:45] Uh, we will find the class… we will find the mode, and for those particular, uh, like, the subset where, uh,
[00:57:51] we will, like, fill with the mode for that particular… with whatever the mode for that particular class.
[00:57:59] So this is classifies mode imputation, and these are, like, standard.
[00:58:02] So, yeah, so…
[00:58:05] So, we begin with the code here for how to do them. So, uh, uh…
[00:58:12] So, for, like, standard strategies like mean, median, and mode, so SKlearn already has a simple imputor, like a…
[00:58:21] Gladys function defined here.
[00:58:24] So, I'm calling it, uh, that only. So, I begin here, uh, I have imported simple imputor, and I'm telling it that the strategy is mean.
[00:58:34] And this is just because there are so many, like, we are doing mean, median mode, everything, so I'm creating a copy of it.
[00:58:41] every time. So, DF underscore 10 was the… where we have introduced 10% values as missing. So, this is, uh, now for, uh, because mean will be here, so I'm creating a copy of it, DF underscore mean underscore 10.
[00:58:56] So yeah, so imputermine is the…
[00:58:58] class object here.
[00:59:02] And on it, I am calling a function fit underscore transform.
[00:59:06] So, fit underscore transform is like, it will doing… it will be doing two things. It will be fitting, and it will be transforming. So, basically, by fitting means…
[00:59:15] in SKLN, fitting has various meaning, but here, in imputation, it means that
[00:59:21] You first, uh, you first go through the column,
[00:59:25] you first will go through the column and find for each column, the mean.
[00:59:30] So, this is what we mean by fitting. And then, when you know the mean, then when you will do the transformation, that for all the missing values, you will
[00:59:39] introduce the mean value, that is called transformation. So, both of these operations are performed in a single function, that's why it's called fit underscore transform.
[00:59:49] And you are providing it for all the columns, all the columns and all their entries you are providing, so that's why I've written error.
[00:59:58] This is the data frame name, and because I want to do the, like, imputation only on the numeric column, the four columns on which I have introduced, uh, the missing values. So, these numeric columns,
[01:00:10] Uh, so these are the numeric columns.
[01:00:15] I have, like, used here, uh, so, for, uh, so basically, DF…
[01:00:23] underscore 10 has, like, all the 9 columns, but we want to fit only on these 4 columns.
[01:00:29] So, that's why we have written here that, uh, only do for numerical… so it will basically do only for the numeric columns.
[01:00:37] So, similarly, this was for mean.
[01:00:41] Just, you have to specify by a string that now you want to do a median imputation.
[01:00:47] So, it will do a median imputation.
[01:00:50] And, uh, yeah, and it will, like, store it in this, uh…
[01:00:54] data frame area, uh, just a minute.
[01:00:57] Yeah. So, okay, so sim… and then, again, with mode, uh…
[01:01:02] Uh, the name is not Mode Air.
[01:01:06] It's the name most frequent.
[01:01:08] And it will… and once again, for each of them, for mean, median mode,
[01:01:13] I'm creating separate data frame, because I will visualize them later.
[01:01:18] So that's why I separated data frames are created for each of them when we are imputing them. And the last is mode within each class.
[01:01:24] So, with the mode within each class, like, there is no standard function, so that's why I've written, uh, some code.
[01:01:31] So, small section, uh, like, logic is written here, so basically, uh, we are looping
[01:01:37] two unique values of the target variable, because we are doing a mode within each class. So the class variable is mpg high, the target variable.
[01:01:45] And if you do, uh…
[01:01:48] And if you do, like,
[01:01:50] Uh, if you, uh, like we have already shown also that there are only two unique values.
[01:01:54] Okay, so it is because its function has not been introduced.
[01:02:00] Okay, so, uh, there are only two valves. So, basically, this line, it only returns a list with two values, because there are only two unique values in the target variable, which is mpg class variable. So, it will, uh, loop through these two values.
[01:02:14] And it will create a… it will create a mask for, basically, when the value is 0, it will create a mass that the true values are those values where the mask is… where the class label is 0.
[01:02:26] And then it will, for only that particular subset of dataset, it will find the mode.
[01:02:30] And, uh, replace them.
[01:02:33] replace the missing values with the mode for each column.
[01:02:37] So, and then it will loop through one and do it similarly.
[01:02:40] So, these are the four strategies, and it will give me four, like, for imputing the missing values here.
[01:02:48] So, uh…
[01:02:50] Yeah.
[01:02:51] As is, could you please repeat one more time, uh, the fourth point?
[01:02:54] Yeah, so the fourth point is, uh, like, I have, again, I am on, again, imputing by mode, but in order to be, like, more smart about it, what I want to do is that I want to have more class-wise.
[01:03:08] So, basically, like, as I'm saying, that there are two class… there are two class labels in the dataset.
[01:03:16] only two class levels, 0 and 1. So you take a subset of the full dataset, where the class level is 0,
[01:03:24] And for that subset, for each of the columns, you find the mode.
[01:03:29] So, and then, for that mode, you replace… in that subset, you replace the value with that mode. And similarly, the other subset will be when the class value is 1.
[01:03:40] And that subset, for each of the columns, you find the mode, and whatever the values are in that subset, you replace them with the mode. So, this is being done here.
[01:03:55] So… so we proceed, uh…
[01:03:57] Next, can I proceed?
[01:04:03] Okay, okay.
[01:04:04] Yeah, please carry on. I think people will ask if required.
[01:04:06] Okay, so, uh…
[01:04:08] So, this has, like, this is, like, we have created so many data frames, uh, like, for mean, we have created this, DF underscore mean, uh, 10, DF underscore median, 10.
[01:04:18] and everything. So, uh, like, but how do we know that, like, even the impression strategies, like, uh, like, working well or deviating and everything?
[01:04:28] So, these are just, like, you have created it and, uh, like, these are tabular data, how to even see that something, like…
[01:04:36] Uh, in order to proceed further, we would like to have some visualization to, like, be sure about which of the strategies, like, not deviating further. So, what we do is… so basically, as ma'am told, PC, she will cover later, but right now, we, uh, I have used PC only as a visualization tool, because I have to… it can do… because if, for, if there are, like, more than…
[01:05:00] If there are multidimensional dataset, we have 9 features here, and if we consider the 4 numeric features also, so…
[01:05:07] Because I'm doing only 2D visualization, so PCA can help in that, for doing 2D visualization of the data. So that's why I have used, uh, considered it only as a visualization tune right now. So, uh…
[01:05:21] it, uh, so basically, uh, so yeah, so we have… so what I will do is, I will do a PCA visualization for the original data set,
[01:05:30] And for each of these imputed datasets, uh, like, which are by this variable name here.
[01:05:37] So, uh, uh, this code is mainly this… all of this code is for plotting purpose only.
[01:05:45] Most of this code is, like…
[01:05:46] Dealing with a plotting purpose. So, basically, I have the data frame, whatever, because there are now five, uh, data frames, the original data frame, and all of these data frames by this name, like this, uh, mode, uh, mode by class, so there are so many data frames.
[01:06:02] So, uh, that's why I've written a generate function, which will take the data frame.
[01:06:09] It will take the name is just a string value, nothing about it, and numeric columns on which I want to do PCA.
[01:06:16] Uh, and, uh, yeah, so this will be a…
[01:06:21] my numeric columns. Y is my…
[01:06:26] Last label.
[01:06:28] And then, yeah. So, this, all this code is related to
[01:06:31] like, obtaining the PCA.
[01:06:35] PCA thing. Yeah, so, uh, the first, I will show you what, by visualization, what we mean is,
[01:06:41] So, if you see here, this is the origin… these are, like, 2D visualizations. So,
[01:06:47] So, the two actions of the PCA which are shown here is PC1, PC2. These are the, like, the axes, the coordinate axes, where the maximum variance of the data is there.
[01:06:57] So, and this is for the original dataset. So, what we will do is, for each of our imputed data frames also,
[01:07:04] We will plot this PCA, kind of, and see how much we are differing from the original one.
[01:07:10] If we are differing, uh, very much, that means the amputation strategy is not that good. Otherwise, our strategy is, like, similar to it, we can say. So, basically, it's only for visualization. As such, we don't have to know about PCA, and other thing,
[01:07:26] We can say is that, uh, since there are four numeric columns, so there are four PC components, and the first component is having 80% variance of the data, the second is having 16… so…
[01:07:36] These are usually by descending order. And the last one are the, uh, like, having very, very close variance. So, we have…
[01:07:43] Uh, just I wanted to add, sorry to interrupt, Ashish.
[01:07:46] Uh, so here, uh, what, uh, why, uh, this explained variance,
[01:07:53] Uh, we can understand that the first principal component is majorly
[01:07:59] capturing, uh, the variants within the data as high as 80.07%.
[01:08:06] So it's a very important component.
[01:08:09] Those components that are not capturing a lot of variance in the data, like PC3 and PC4, are not
[01:08:16] Significantly doing that,
[01:08:19] If we want to ignore it, we can also do that. That, uh…
[01:08:25] is one thing that is being shown here.
[01:08:27] In addition to that, the loading matrix, which is underneath the explained variance,
[01:08:33] also shows that if we are talking about the different attributes,
[01:08:38] Then, uh, you know, uh…
[01:08:40] How much of makeup we have for this…
[01:08:43] So that is also an important thing.
[01:08:46] Uh, to note…
[01:08:48] And we want to actually do missing value imputation, or we want to fill in missing value,
[01:08:56] In such a way, uh, that the overall data distribution
[01:09:01] Or the meaning of the data doesn't change. That is the whole objective.
[01:09:05] Otherwise, it's very, uh, easy, right? We can just fill in some value.
[01:09:10] Uh, and just do something, but we don't want to do that.
[01:09:14] We… what we want is that the overall data distribution, if there are class labels or class distribution,
[01:09:22] Uh, then, uh…
[01:09:23] That should not change.
[01:09:26] Additionally, as I mentioned, that loading matrix, uh, actually shows, uh, the amount of correlation
[01:09:32] Between the original variables that were there,
[01:09:36] And the principal components, or the PC1, PC2,
[01:09:41] And all that that have been created. For example, if you see…
[01:09:45] Uh, in the loadings matrix, displacement.
[01:09:49] Then, it has 0.3…
[01:09:52] Or, uh, 0.53%, uh…
[01:09:54] For positive correlation,
[01:09:56] with, uh, displacement.
[01:09:59] So, first principle component is positively correlated.
[01:10:03] with correlation factor of 0.53.
[01:10:06] With the attribute displacements. Similarly, the second component has positive 0.25.
[01:10:12] Whereas the third component is negatively correlated with displacement, which is minus 0.45.
[01:10:19] And then fourth one is PC4 and minus 0.66, so like that.
[01:10:23] So this is showing the correlation between the principal components that we have found out.
[01:10:28] And the attributes that have been considered.
[01:10:31] However, what is correlation? How we calculate it?
[01:10:35] What is principal component analysis? How we…
[01:10:37] Uh, you know, find it out. All that I will cover in theory.
[01:10:41] Why this is being shown here? Because it's primarily for visualization purpose only.
[01:10:47] that we are using here.
[01:10:49] So, that's why Ashish is showing it. So, right now, you can just assume.
[01:10:55] That principal component helps us to…
[01:10:59] Uh, visualize the data in a 2D or 3D graph.
[01:11:02] How that does, I will show… tell you, and similarly…
[01:11:08] Uh, we have these, uh, loadings matrix that can help us to find out the correlation. Similarly, how much variance is being captured, that also can be found out.
[01:11:17] So, any questions, anyone, on this part before we start, uh…
[01:11:18] Yes.
[01:11:21] The rest of it, okay, Lokesh, what do you want to know?
[01:11:24] Yeah, ma'am, like, um, basically…
[01:11:27] What we are trying to find in visualization, like, uh, the displacement and horsepower are more influencing than
[01:11:32] The weight and acceleration, right?
[01:11:36] Uh, not necessarily. We can…
[01:11:37] For the target variable.
[01:11:40] Yeah, so we may not necessarily say that.
[01:11:43] Because if you look at the different principal components, which is what we are using to do the visualization,
[01:11:50] PC1, PC2, PC3, and PC4.
[01:11:53] Okay.
[01:11:54] So, displacement shows a positive quite a high positive correlation, which is 0.53.
[01:12:00] Whereas if you, uh, look at, uh, the acceleration,
[01:12:05] Uh, that will also have… so all of them have some positive… so if you look at PC1,
[01:12:09] The acceleration is, like, minus 0.39.
[01:12:13] So, that's quite, uh, some value of negative correlation. It cannot be ignored.
[01:12:19] So, you cannot say that if we talk about PC1, then…
[01:12:23] Acceleration is not important. No, we cannot say that. Only thing is…
[01:12:27] that the, uh, displacement horsepower and weight are positive…
[01:12:32] Blue correlated, whereas acceleration is negatively correlated, okay?
[01:12:36] If you look at PC2,
[01:12:38] Uh, definitely horsepower is showing, uh, a negative correlation of minus .02.
[01:12:45] Now, since the number is very small,
[01:12:48] Therefore, we can say that, yes, this is a little…
[01:12:51] insignificant, this much of, uh, negative correlation.
[01:12:56] Other than that, mostly either the positive or negative correlation is there, and…
[01:13:01] Uh, some number, significant number is there.
[01:13:04] Is it okay?
[01:13:05] Yeah, thanks for that one.
[01:13:07] Oh, man, what is PC1, PC2, and…
[01:13:12] So forth.
[01:13:13] So, uh, PC1, PC2, PC3, and PC4 are the principal components. They are…
[01:13:19] that we obtain after doing PCA.
[01:13:24] Okay.
[01:13:25] Okay, so I'll be telling you what is PCA, however…
[01:13:30] You can assume that we are trying to
[01:13:31] Map the multiple features or attributes to
[01:13:36] Lesser number of attributes using PCA, so we can visualize.
[01:13:40] Okay?
[01:13:43] Okay, Abhishek Dhaban, Devan?
[01:13:45] Uh, what is your question?
[01:13:46] Um, the SPCA related to imputation strategy also, like, uh…
[01:13:50] we did four, uh, invitation strategies, right?
[01:13:54] So, is there any relation between that, or this is, like, completely independent?
[01:13:59] No, uh…
[01:14:00] Is it done to understand, or…?
[01:14:02] Yeah, so right now, we are just taking them as independent totally.
[01:14:06] Why? Because, uh, in addition, see, sometimes visualization of the data also helps to fill missing values.
[01:14:14] For example, when we…
[01:14:16] Fill missing values, and we visualize the original data
[01:14:20] And the data, after doing the…
[01:14:23] Uh… the imputation, then the overall distribution should
[01:14:29] More or less remained the same.
[01:14:31] So, we wanted to show that, that's why we are doing PCA.
[01:14:35] So, it's not related to filling of missing values, it's more related to visualization only right now.
[01:14:44] Okay, uh, Deepak, do you have a question?
[01:14:48] That's one thing, uh, probably we are going to go through it later on, but at this point, how do we categorizing it so that it's easy for us to go through this exercise?
[01:14:56] Uh, no, I didn't get you. Can you please come back?
[01:14:57] So how… Yeah, how do we categorize, uh, these PC1, PC2, PC3 flow? So, what does it, like, how do we interpret it at this point?
[01:15:06] So, uh, see, interpretation, uh, I will explain to you broadly.
[01:15:14] Uh, in the theory session. However, principal components try to capture
[01:15:19] The variance in the data.
[01:15:22] So, the major, uh, variance, uh, directions it tries to capture.
[01:15:27] Say, for example, uh…
[01:15:31] So, Ashish, I wanted to show the… can you please go up?
[01:15:35] Uh, yeah, right here.
[01:15:37] So, you can see, uh, the amount of variance that, uh, is there.
[01:15:49] Oh, sorry. Uh, so you can see the amount of variance in the data.
[01:15:54] For example, this first principal component is capturing 80% of the variance in the data, right?
[01:16:00] So, uh, so, uh, the interpretation is that we want to
[01:16:05] See, we want to actually… we are having large number of attributes.
[01:16:09] Now, if you want to visualize the data,
[01:16:12] Then, in that particular case, uh, so many attributes, how we will represent. It's not possible.
[01:16:19] We can visualize the data if it is two-dimensional or three-dimensional.
[01:16:23] Maximally three-dimensional, right? Humans cannot…
[01:16:26] have more visualization.
[01:16:28] And then 3D, they cannot visualize.
[01:16:31] So, we need to map them.
[01:16:34] too, uh, fewer number of…
[01:16:37] features. So, for that, PCA is used.
[01:16:41] And, uh, in PCA, what happens is, mathematically, we try to
[01:16:46] capture the variants across different directions in the data.
[01:16:50] So, that's an alternate way, complete alternate way, different way of representing the same data.
[01:16:57] So, the first principal component represented as PC1
[01:17:02] got just 80% of the variance. Second one's 16, like that.
[01:17:06] Okay?
[01:17:08] So, we are trying to, uh…
[01:17:12] map one, uh…
[01:17:15] One way of representing the data, that is the original way with values and attributes,
[01:17:20] into a complete new way, where we have just principal component that capture the variance in the data.
[01:17:26] across different orthonormal directions.
[01:17:29] This is what we are doing.
[01:17:33] Is it okay, Deepa?
[01:17:34] Thank you.
[01:17:36] Also, for the PC counts, can be, like, more than or less than 4 also, right? Like, PC, like, PC 1, 2, 3 only?
[01:17:43] But in addition to that, you can change the percentage of that.
[01:17:46] I see, right?
[01:17:47] So… so I think we, uh, chose four attributes, is that correct, Ashish?
[01:17:51] Yeah, yeah, for attribute, that's why 4PC.
[01:17:54] So… yeah, because we have chosen here four attributes, that's why we are having four, um…
[01:17:59] principal components out here.
[01:18:03] Okay?
[01:18:05] Okay, alright, uh, the PC counts depends on the number of, like, activity you are going to choose for this analysis, right?
[01:18:10] Correct. Correct, correct.
[01:18:15] Okay, so I think we can proceed now.
[01:18:24] Okay, so, uh…
[01:18:25] So, we'll start, uh, then, that, uh, we just use PCA for visualization, and uh…
[01:18:30] And we are, like, you can see PC1, PC2 here, 80%, 16% is the variance they have captured.
[01:18:37] And each of these PCO and PCO2, these are like vectors. You see, it's a four-dimension vector, consisting of all these input features.
[01:18:46] So, assume that these are two vectors on which we will map our original data along these axes.
[01:18:54] This is basically 1x's, this coordinate X PC1.
[01:19:00] Which is a linear combination of all the input features, and it is capturing the maximum variance. That's why we have considered that. And the second best variance is in the PC2 vector. That's why I've considered that. So basically…
[01:19:12] Apart from its mathematical understanding, what is there, it is majorly used also to visualize, because usually the first two axes are the one where maximum variance is captured, and if you want to do, we can include a third component also, and do a 3D visualization.
[01:19:29] So, yeah. So, that's why this is the original data, and I have, like, we have class-separated them. So, basically, these are data points,
[01:19:38] 0 is for class, uh, like, blue is for class 0, and orange is for class 1.
[01:19:43] along this new coordinate axis, PC1 and PC2.
[01:19:46] So, this is the visualization of the original data, how it looks. So, here you can see that, uh,
[01:19:52] Both the classes are even separated. There is some mixture in between, but the class zero is… the majority of the Class 0 component, they are, like,
[01:20:03] On the right side, on, uh, of it, and some of them are mixing with Class 1 in the middle.
[01:20:11] So, this is, like, how the original dataset visualization looks.
[01:20:17] So, what we will do is, uh, so this, uh, is the function, uh,
[01:20:22] where I told that these are numeric columns, this is the original data frame, and this is just a string to… and this will output something like this. So what we'll do is,
[01:20:31] For all the four imputations we did, we will get this visualization and compare with the original visualization, where there was no imputation. So, that's what the code is about now. So, uh, so first, the mean imputation, so the data frame was this.
[01:20:48] And, uh, so this is the immune imputation.
[01:20:52] This is like an extra parameter, because I will do a pair, uh,
[01:20:57] like a pair of visualizations with the original data frame, so that's why this is supplied here.
[01:21:01] So, at the end, you can see, like, uh, yeah, so first you see that the variance. So, obviously, since
[01:21:09] We have, like, imputed some values, so the variance, uh, like, in original dataset, the first component was having 80% variance.
[01:21:17] Now, variance has reduced, it is now 73%, and PC2 is also, like, reduced. If you've seen the original,
[01:21:27] Oh, just a minute, uh…
[01:21:29] Where it is.
[01:21:34] Yeah, in the original, it was 80% and 16%.
[01:21:37] Here, uh, the…
[01:21:40] Uh, first one has reduced, and it has, like, uh, like, accumulated in the other one. So, it has risen from 16% to 17%, so yeah.
[01:21:48] So, first component is having lesser of variance now, because imputation has…
[01:21:53] slightly affected it. And when we see the plot, this is the original, and this is the mean imputation.
[01:21:59] So, uh, uh…
[01:22:01] In plot as such, like, if we look very closely, then only some, uh, like, you can see some differences. Otherwise,
[01:22:11] like, pinplot the distribution more or less looks similar to it. So, this is, uh, and more so, it's because we have only done imputation of only 10% for each of the four columns.
[01:22:23] That's also the reason for this.
[01:22:26] And then, uh, similarly, I'm doing a, uh, like, PC visualization for the median imputed.
[01:22:33] And here, a slight reduction, like, for mean, it was 73%, here it is also, like, approximately 73%.
[01:22:40] 72.98, it is sync.
[01:22:43] So yeah, so if you see the visualization again, so… yeah, so, uh, this was the… on the left is the original one, and on the right is the median imputed.
[01:22:53] And, uh, yeah, so, uh, you see, it is also, like, more or less, uh, looks, uh, like, similar, but, but, uh,
[01:23:03] You will see some more values are, like, merged inside the…
[01:23:07] like the orange one. So, in the left, the blues are, uh, like, not that much merging, but in the right, the median, some more values are, like, have been merged inside the one. But yeah, but the… it… it's not that much observable right now, because it's, like, only 10% of the values are inserted, so most of the distribution
[01:23:30] Looks similar. And even the variance is, like, originally it was 80% variance captured, here also, it's, uh, 73% is still captured in the first component.
[01:23:39] So, yeah. So, then for mode imputation,
[01:23:44] Where we do our mode imputation.
[01:23:46] So here, uh, there's a slightly more reduction, only 70% has been captured.
[01:23:53] Uh, compared to the 80% original.
[01:23:56] And here is its visualization. Uh, and uh…
[01:24:01] Yeah, so this is… it's visualization. It's slightly, uh… slightly more deteriorating than original one, but…
[01:24:08] We'll have to look very closely here, because, uh, right, as I again said, only 10% of them have been imputated, so…
[01:24:16] So, it is very hard, as such… yeah, but here, it's very visible, so if you see…
[01:24:25] The PC… the original one, the…
[01:24:27] The oranges are, like, uh…
[01:24:29] The blue ones, the class 0, has been extended very much inside the orange one, in the mode imputation. So, some differences visible in the mode imputation.
[01:24:40] Because variance has reduced further, it's only 70% variance now, 10% difference from the original.
[01:24:46] So, that's why it is able… it is kept… along this PC1 direction,
[01:24:51] The variance captured is less, that's why it is more congested in this PC1 direction now.
[01:24:57] And next, again, by mode by class. So, you see, the mode by class is able to retain, it is showing 76% variance. It's the highest.
[01:25:06] The median was saying around 30… the median and mean was saying around 73.
[01:25:10] And it is able to, like, capture 76%. So, it's the best for our case.
[01:25:16] And, yeah, so this is the original one on the left, which is captured along the x-axis, the PC1 direction, it is captured 80% of the variance.
[01:25:25] And here, it has captured 76% of the variance along PC1 direction.
[01:25:30] So, uh, here also…
[01:25:31] Just a small deduction in variance, and that's why their model looks similar.
[01:25:36] And we can say that among all of these fours,
[01:25:38] for… the maximum is variance based off by mode by class.
[01:25:44] And it looks… this imputation looks…
[01:25:47] more similar, and we will use… so we can conclude that we will, in our case, in our dataset, mode by class is working well, as compared to the other imputation strategies.
[01:25:59] Yeah, so, uh, any questions?
[01:26:03] to this point?
[01:26:10] Yeah.
[01:26:11] Well, yeah, hi, is this, uh… So, like, uh, to calculate this, like, mean, median, and mode, like, we have to write the different logics for all these things, right?
[01:26:15] We have to, uh, come again, please, Roman?
[01:26:16] Look, all these things. Yeah, uh… For the main, like, medium, and more, like, okay, so… all these 3-4 types, like, we have it here, okay? Like, for those 4, like, we have to write the separate logic, the separate code will be there, right?
[01:26:31] No, no, so, so, yeah, so let… so, yeah. So, so this was the… okay, so this is the code, this is a generic code which will work for everyone.
[01:26:35] I can source with the code, yeah.
[01:26:40] So, basically, you give the data frame, you… it's just a string, the numeric columns on which you will do the PCA.
[01:26:49] Mm-hmm.
[01:26:50] So, if we start here, so you'd, uh, take the subset of the numeric columns on which PC8 is X. Y is my, uh, the target column, the MPG high column.
[01:27:00] then this is a standard scaler for which I will do, uh, like, it is introduced to do, uh, people… it will, for each of the columns, it will do, it will, uh, transform them to have a zero mean and unit variance.
[01:27:13] for each column separately. So, when we do PCA, we first standardize the data, so that's why it's, like, used here. So, what it is doing is, it is calling an object of standard scalar, and again, the same thing, fit underscore transform.
[01:27:29] So, it will find for each column the mean and variance, and then it will transform each column.
[01:27:34] So, it will subtract from each value its mean and divide by variance. So, this is how, uh, so basically, it is something like that. Suppose A is…
[01:27:43] values A in one column, and if I find my mean is mu for that column, overall column,
[01:27:51] And, uh, suppose my variance is where… my variance is this. So yeah, so, uh…
[01:27:56] this kind of operation is happening for each of the entry, like, this… I'm showing for one value. It will happen for the entire column entries, and that… then I get a scaled version. I will call it a scaled version of my…
[01:28:10] like, data, and this will be given to the PCA…
[01:28:15] function which is indiscular. So, PCA will take the scale data, and it will do the, like, the PCA computation mathematically over the whole thing which are there.
[01:28:25] And from that, I have, like, used that… from that, I have, like, used that PCI component. It can then, if you will call this variable here, uh, explain variance ratio, it will tell you the variance also. It's like a list which, uh,
[01:28:40] for… and it will tell you variants for each of the components, which I have printed here.
[01:28:43] That's why it was showing here the, like this, this first thing, the first print statement.
[01:28:51] And then, uh, each of these components, as I told, these are also vectors.
[01:28:56] So, that's why the loading meta- uh, from after that, each of these components, we have shown as a matrix form.
[01:29:03] And, uh, shown their relation with the original features, how they are a linear combination of the original features. So, basically, this all part is, like, uh…
[01:29:12] theoretical, depending on, uh, like, how PCI works, but yeah. So, what we'll do is, eventually,
[01:29:19] We will, uh, like, use, uh…
[01:29:24] Uh, here, it is like, maybe it's here.
[01:29:27] loading speed… yeah, it's not here. So, basically, after that, this graph, the code is for that graph, this graph, this bar graph, which is showing for PC1 only, the country. So basically, this bar graph is nothing but…
[01:29:40] This PC1 column, all the values which are written here, 0.5, 3.5, 4.39, they are shown in that bar graph.
[01:29:46] In a form of paragraph.
[01:29:49] So, this code is here. Uh, this part is here.
[01:29:55] Just load my… yeah, so this part is here, for showing that bar graph. Now, after that, since I have my, uh, both of the… both of… I have my PCA components,
[01:30:07] I can, uh, I can call my, like, I can call the… I have called… I've used another function to, uh, like…
[01:30:14] plot the PC1 and… to plot the data and transform them in the directions of PC1 and PC2. So, it is happening here.
[01:30:22] So, basically, there is a function called byplot, which is, uh, inside this, another function.
[01:30:28] And, uh, so basically, like I've said, that, uh,
[01:30:32] If I'm doing a mean median mode of imputation, so I'm giving the original data also, original reference will not be done here, because I have to plot two subplots.
[01:30:42] to show the… on the left side, I am showing the original plot also, on the right side, I am showing the, uh, plot for the imputation one.
[01:30:50] So, that's why, uh, there is an if-and-else here, that if original reference is none,
[01:30:55] Then do this. Otherwise, uh, do this. So, if original reference is none,
[01:30:59] You just directly call this byplot function, which is plotting that, uh…
[01:31:04] plotting that graph for PCA.
[01:31:07] And you give it your, uh…
[01:31:09] like, piece, uh, like, that transform data, the PC components, and it will
[01:31:13] plot… it will transform data along that, and it will plot.
[01:31:16] And if there are, uh, like, if the original reference is not none, that means I am comparing with the both of them, like the original data… original unimputed data and the imputed data also. So then, at that time, this part of code is run.
[01:31:32] But yeah, so this is all we made in a generic function, so you can… we can call it for
[01:31:36] all the data, all the 5 cases, the original one, and the importation ones.
[01:31:37] Okay, Ashish, I think you can take over from here.
[01:31:47] Okay, so, yeah, so…
[01:31:48] Okay, we got it.
[01:31:49] So, much of the code here is, like, for plotting only, and that's why not initially, I did not discuss it in detail. Yeah, so, but there is one single function in which you just, uh, like, change the parameters, and you will get plot for, uh, different of the cases.
[01:32:07] So, this was, uh, for, uh, this was how we do and did this.
[01:32:13] Yeah, so, uh, in the second part of it, what we did was, initially,
[01:32:18] What we did was we only imputed 10% of the values.
[01:32:22] And we see the variance is not that much affected, so…
[01:32:26] In the second case, the strategies are same, the visualization is same, but the only difference we are doing is
[01:32:32] We are creating copies where each of the… we are imputing 30% of the values in each of the columns.
[01:32:38] So, right now, if I run it again…
[01:32:43] many ways, because now 30% of the values are imputed, so you will see, like, many values are imputed. Otherwise, 10% very less values are imputed. Now, when we will, like…
[01:32:53] do this same impression strategies. The strategies are actually same.
[01:32:58] that you do by mean. Everything, like, this is basically repeated of this code. The only difference here is…
[01:33:04] Uh, that I have used now. Earlier, it was 0.1, I have now used 0.3. Otherwise, the whole code is
[01:33:10] just the same, and the imputation strategies and the code is also the same. Obviously, the name is different, because now I need to compare. So, this was done, now you will… now I will, when I will show you the visualization, you can clearly see that if you are having…
[01:33:25] Larger values is missing.
[01:33:28] Uh, how these imputation strategies, uh, like, fail or, uh, does not match up to the original distribution.
[01:33:36] So, uh, is this the second particular, then I will show you the, uh, like, the visualization for this.
[01:33:43] So, in the second part, I've done nothing, uh, nothing has changed, just that we have imputed 30% of the values up instead of 10%.
[01:33:51] Is this okay?
[01:33:52] Uh, sorry to interrupt, just to add here, we want to see the impact of increasing, let's say, in a data set, if there was
[01:34:01] as high as 30% of missing values, then what will be the impact of…
[01:34:05] Using mean, median, mode, or class-wide, uh, median,
[01:34:10] for filling of missing values, and so on and so forth, which I had also covered in theory.
[01:34:15] And I'd shown you different graphs in theory as well, so you will see the same impact in code also.
[01:34:21] Uh, what happens when different amount of missing values are there, and that's why we are doing all this. Over to you, Ashish.
[01:34:28] Okay, so let's see the visualization again. So, the function is same, so that's why I've not written it again, so…
[01:34:36] The same analyzed single PC function is used. And, uh…
[01:34:40] Yeah, so this was the original one, 80%, 16%, just for recap. This was the original variance in the unimported data, and…
[01:34:48] This was its biplot. And then we are seeing the… for the mean embedded data, the first.
[01:34:54] So, you see, uh, this time,
[01:34:57] It's only 62% variance in the first axis, PC1, and 19% in the second axis.
[01:35:03] And now, uh, you can see the… there is a much difference in there, uh, like, PC, uh, like, visualization of their PC components.
[01:35:12] So now the difference has emerged, because, like, as high as 30% values are missed in each of the columns, so therefore, uh, it's, like, the variance has shrunk in PCA. If you see the left graph, the XX is the PC1Xs, there is, like, it is… it is varying enough to a larger width.
[01:35:32] in the Texas, but here it has shrunk, because now it is only 62%.
[01:35:37] Uh, variance is being captured.
[01:35:39] And similarly, for the, uh, if we do for median one, it is also 30, 30, uh…
[01:35:44] So basically, in the dataset, what we observed is either you do for 10%, either you do for 30%. The median and mean, the percentage of variance being captured in the first component is almost same. It is also, like, 62%, just a mile. Same pattern was observed in the, like, for a 10% also.
[01:36:02] So, just a minor reduction from INT.
[01:36:06] But yeah, but originally it was 80% of the variance, now it is only 62%, so that's why the difference is stark here, and uh…
[01:36:15] And the distribution has changed. And similarly, for mode… so mode is performing the worst. It is, like, only 54% of a variance is captured.
[01:36:26] And that's why, like, it has shrunk more right now in the PC1 direction, the very less variance is captured. Uh, and, uh, uh, yeah, so, uh,
[01:36:40] you can… by shrunking the, like, uh, yes, yeah, so this is here, and last is the mode… class-wise mode data.
[01:36:50] So here, there is, uh, you will see that there is, again, like, similarly in mode imputation, there is improvement. So, mean and median were 62%.
[01:36:58] mode was only 54%, but classifies mode, even for if I'm doing 32% data, 30% data is, like, randomly given, assigned missing values.
[01:37:08] It is still able to capture 71% variance. And remember, the original variance was 80% captured in the original data. So, we can say that class-wise mode is, like, working well in our case.
[01:37:20] This is, like, a conclusion from…
[01:37:23] this visualization for both 10%, 30%.
[01:37:27] So, hm.
[01:37:29] We can conclude that in our dataset,
[01:37:31] There is, uh, that, there is a class separation in this data, which is shown from this dataset, but…
[01:37:39] If we were to choose a, like, a strategy for missing value, like,
[01:37:45] Mode by class is working.
[01:37:47] The best in our case. So, this finishes the part one of our, like, uh…
[01:37:52] this notebook.
[01:37:54] So, can we consider, like, more by class always will be good, I mean?
[01:38:02] No, no, it, uh, like, depends, because here, the data is categorical, like…
[01:38:06] Because some cars are having high efficiency, some are less.
[01:38:11] So, maybe if there is no categorical variable, the target variable is categorical here, so that's why it's working, like, good here, and it will, like, depend on the dataset used.
[01:38:21] Uh, let me add here, Raveen, generally speaking, if we are, um, having class, uh,
[01:38:28] Separated data, then this will be better. This will work better.
[01:38:32] Okay, ma'am.
[01:38:33] Generally speaking, obviously,
[01:38:36] Uh, definitely based on the specificness of the data and all things also may vary, but generally, yes.
[01:38:41] Okay, thank you.
[01:38:44] Shikar, do you have a question?
[01:38:46] What about the probabilistic method?
[01:38:48] That we talked about. Will it help more?
[01:38:51] Yeah, so… yes, probabilistic method also can be used, but we haven't illustrated that one here.
[01:38:59] That's not very commonly used, uh…
[01:39:03] So, you estimate the probability of the class,
[01:39:07] Uh, the attribute value.
[01:39:11] Uh, for a… based on the class, uh, distribution, yeah.
[01:39:13] That is not yet illustrated here.
[01:39:16] So these are the most common methods that are used.
[01:39:20] Uh, and as you can see here, if we use a class-based, uh,
[01:39:25] Uh…
[01:39:27] feeling of missing value.
[01:39:29] Uh, that works the best.
[01:39:31] holding probabilistic never have given better results.
[01:39:36] Yeah, probabilistic method could have been better. However, we have not illustrated it here.
[01:39:42] Because, uh, there are categorical attributes, there may be categorical attributes also.
[01:39:48] So, probability will work primarily on numbers only.
[01:39:52] So, it is also restricted in some way.
[01:39:55] So, that is also there.
[01:39:57] Uh, we'll see if, uh, that can be added.
[01:40:02] later in some other session.
[01:40:05] So, right now, we just tried to concentrate on…
[01:40:09] The most common methods.
[01:40:13] Dwarakish, you have some, uh, question?
[01:40:16] Um, so, uh, here we have considered only mode by class, right? Like, should we also consider, like, mean by class, or will that have any impact?
[01:40:25] If we can consider definitely having a mean versus mean by class,
[01:40:31] Uh, maybe… so mean by class may be better.
[01:40:34] But still, it will not be as good as mode, which may not be as good as medium by class.
[01:40:40] Okay, you can do it as a self-exercise also, just try to…
[01:40:45] change the code slightly and try to do that.
[01:40:48] Uh, it will be better than simple meal.
[01:40:55] Anyone else?
[01:40:56] Okay, thank you.
[01:41:01] Okay, so I think, uh, Ashish, you can take the second part, which is on… primarily only on EDA.
[01:41:08] Okay, so now, uh, like, uh…
[01:41:12] Now, what we'll be using is the original data frame,
[01:41:16] DF, because it had no missing values, and we are doing now ED analysis.
[01:41:21] And this is, like, we can… we can treat it as separate, because we had created multiple copies of this dataset to show missing values, but since originally there were no missing values, and now we are showing the ED analysis, so that's why we are using… we'll work on the original data frame DF only.
[01:41:37] For this second part.
[01:41:39] So, uh… so, yeah, so the first type of EDA analysis,
[01:41:44] Which we are showing here, it's basically count plot, and count plot is usually for…
[01:41:51] the columns which are categorical, because it shows, like, a bar graph for various categories. So, what I've done is, like,
[01:42:00] Uh, so basically, uh, what this code… this function is, uh, there, it is, uh, what it's doing is, it is taking the data frame, and you
[01:42:10] provided the column name for which you want the, uh…
[01:42:13] like, the count plot. So, uh, all of these, most of these things are annotation details here.
[01:42:19] And there is one, uh…
[01:42:22] like, there is one separate, uh, for if the column name is origin, there is one separate if condition, because the thing was that the original values were 0, 1, and 2, but I wanted to, like, to present for presentation 1, 2, and 3, so that's why I…
[01:42:38] this, uh, is separate.
[01:42:39] Otherwise, the main line which is working here is this, sns.conplot for SNS is for Seabourn.
[01:42:46] import C1 as SNS?
[01:42:49] You give it your data, uh, you tell it which column you want the column plot, and it will do this for you.
[01:42:55] So, yeah, so if first I've given the cylinders. So, it is showing me, like, uh, like, it can be seen that the number of standard 4 is the highest, and…
[01:43:06] 3 and 5 are the lowest. So, basically, it's just the same thing, the unique values, but it is shown in a graphical format right now.
[01:43:14] The number of unique values in the categorical column.
[01:43:18] Similarly, for model year,
[01:43:20] So, model years are from 1970 to 1982.
[01:43:26] And we are seeing the count of the cars for each of the model year.
[01:43:29] And, uh, 73 is having the highest of the cars in our dataset.
[01:43:35] Similarly, the data frame and the column is origin, the origin country, so we can see that maximum cars are from USA.
[01:43:44] And, uh, followed by Japan and Europe, they are, like,
[01:43:50] more or less similar values. So, yeah, so this is a…
[01:43:55] So, basically, for categorical feature, the first EDA, what we can do is, like, we can do a COMP plot. Then we
[01:44:03] proceed for the numeric features. For numeric features, what we'll be doing is, it's called pair plot, for… we will take pair for each of the numeric features. We'll consider them in pairs, and see, uh, how, uh, they vary with respect to each other.
[01:44:20] So, here again, I'm taking the data frame, and uh… Seabon already has this function called pairplot, so you take the data frame, you tell it the numeric… you tell it the columns for which you want the pair plot, so it… numeric columns is a list of all the numeric columns, the four columns.
[01:44:37] And, uh, because it will create a 14… 4x4 matrix for 4 columns.
[01:44:43] So, on the diagnosis,
[01:44:46] On a diagnosis will be the… the columns will be same, 1, 2, 3, so it is, uh, so that's why it's, like, given that for the diagonal, consider the knowledge density estimation, like,
[01:44:59] Consider the same variable, but it's, like, continuous estimation of only that same variable. So, it will be clear by, like, the
[01:45:09] graph. So, uh, so this is like the, uh, uh, like the… just the plot of the displacement, because it's a diagonal value, just the plot of the displacement showing in continuous form.
[01:45:19] And similarly, the plot of horsepower. And other than that, for each of these cells here,
[01:45:28] Just…
[01:45:29] Uh, yeah, so for each of these cells here, it's a pair plot. So, displacement is on the y-axis, and
[01:45:32] the horsepower is on the x-axis for the first one, and similarly.
[01:45:37] Just, I wanted to add here, uh, Ashish, sorry to interrupt. So…
[01:45:42] You can see here that if we look at the individual plots, like displacement, it is, uh…
[01:45:48] Uh, lab skewed plot, if you go to horsepower, this is also similar, so it has a right tail, both of them, if you could please scroll down, Ashish.
[01:45:59] the plot for weight also shows, uh, skewedness on the right one, but it's not very much tailed.
[01:46:08] And, uh, surprisingly, acceleration is a good bell-shaped curve, which shows a very good distribution, normal distribution.
[01:46:17] Now, if you go up, uh, let's look at the relationship. So, we can see here that displacement versus…
[01:46:25] Uh, the second one is, uh, horsepower.
[01:46:26] Horsepower, horsepower.
[01:46:28] So, this shows kind of a linear, positive linear relationship, where when displacement is increasing,
[01:46:35] So, when horsepower is increasing, displacement is also increasing.
[01:46:39] And also, I think it's obvious that when the horsepower of a…
[01:46:43] engine increases, displacement will be more.
[01:46:48] Similarly, the second variable, uh, you can point them, Ashish. The second one also.
[01:46:55] You can see, second one is, uh…
[01:46:56] Displacement and weight.
[01:46:57] Okay.
[01:47:00] Uh, second one is weight, you're saying.
[01:47:04] Uh, so when, uh…
[01:47:06] Uh, here, what we can see here, that, uh…
[01:47:10] Uh, the…
[01:47:12] As the weight increases, uh…
[01:47:15] Uh, as per this graph, the displacement is also increasing.
[01:47:19] Which is, again, a linear relationship between the two.
[01:47:22] However, the important thing to notice with the fourth one, which is the acceleration.
[01:47:27] Versus the displacement, or…?
[01:47:30] So, uh, the relationship, uh, is not very clear.
[01:47:34] Similarly, you can note here the one important kind of relationship that I'd like to show is between the
[01:47:42] Acceleration and the horsepower, which is the…
[01:47:45] Uh, in the second row, the extreme, uh,
[01:47:48] Right one, yeah. So…
[01:47:51] the rightmost one in the second row for the horsepower.
[01:47:55] Yeah, so you can see here that, uh, when the acceleration is increasing,
[01:48:02] Uh, the horsepower is, uh, decreasing.
[01:48:05] So, this is one negative, uh, relationship that we can see here. So, like, this pair plots are actually…
[01:48:12] drawn to understand the relationship between each pair of attributes that is a significance, and it's important.
[01:48:20] component of EDA, too.
[01:48:24] illustrate this kind of relationship.
[01:48:25] Oh, okay, I think we can continue. Over to you, Ashish.
[01:48:29] Okay, so…
[01:48:31] Yeah, so, observation, ma'am, has, like, clearly explained, uh, like, sufficiently explained, uh, what is the observation.
[01:48:38] So, I just moved to the second pair plot. So, this pair plot is also similar, but uh…
[01:48:45] the same pair plot of numeric column, but I have done a class-based coloring in order to, like, illustrate the difference better, so…
[01:48:52] It is the same pipload, but now, since there are two classes, so we can see the data in separate into classes by color.
[01:49:02] So, it is, uh, apart from that, it is, like, just the same periplot. And it can show us more, uh, insights, like, yes, as we see that displacement and
[01:49:13] horsepower, they are correlated, positively correlated, when one increase, the other increase. But we can see that, uh…
[01:49:19] for the class, uh, for the class level…
[01:49:24] For red means one. So, uh, like, for high efficiency, it's like, the upper values only, and for, uh, for low efficiency, it's like the lower values of displacement, 10 horsepower.
[01:49:37] So, here, it is also, like, separated by class. Their correlation is also, like,
[01:49:43] like, different… there is a difference in class of them.
[01:49:47] And apart from that, if we see that if we just see the displacement, just the…
[01:49:55] Unimodal, the distribution of the displacement, if there was no class-based color separation,
[01:50:03] It appeared that there was one peak here, and some peaks were here. Uh, one separate… so it was appearing as a multimodal, uh, multimodal distribution.
[01:50:11] And because when I do a class-based distribution, it has now separated that this, for class 0, there is one peak, and the other peak is for the values in the class 1.
[01:50:24] So, this is kind of, like, more, uh, like, analysis you can get by doing a class label-based separation of these pair plots.
[01:50:34] So, other than that, the normal thing that one, uh…
[01:50:38] the correlation between displacement and horsepower, the correlation between displacement and weight, that things are same. So, yeah, so this is, uh, this, and uh…
[01:50:50] Then, uh, like, because we have noted that there is correlation and novel. So, why not, uh, to see the actual values
[01:50:57] The heat map of all these numeric features, that what is the, like, the… in numbers, how much is the correlation?
[01:51:05] which exists.
[01:51:07] So, for this,
[01:51:10] What we do is, we take our, uh, we take all our numeric columns, and we also, uh…
[01:51:16] take the target variable for that also, we will show the
[01:51:21] correlation. So, uh…
[01:51:23] Yeah, so what we'll do is, we will, uh, first…
[01:51:28] In the Pandas data frame, so this is numeric columns on which I want to show correlation.
[01:51:33] In Panda's data frame, there is already a function by the name .corrr.
[01:51:39] If you give it the data frame on which we want to do correlation, obviously, uh, all this should be numeric values only, on which it will do correlation.
[01:51:48] And, uh, yeah, so these all are numeric values. MPG high is categorical, but since it is, like, considered
[01:51:54] like, numeric-only 1 and 0.
[01:51:57] It becomes a special case of, like the Pearson correlation, and it is, uh,
[01:52:02] Handaz is able to handle that. Otherwise, if it would have been, like, more than two values and everything, it would have thrown an error.
[01:52:08] So, right now, we are able to do this.
[01:52:11] And, uh, yeah, so this is the correlation matrix, uh, and, uh, again, Seabond has a…
[01:52:18] like, again, a visualization function called heat map, in which, if you give it a correlation matrix,
[01:52:25] It will plot this kind of nice heat map, uh, where we can
[01:52:28] see the values and, uh, like, where color grading is there, how much is the value. So, uh, just as we have seen in the pair plot, in displacement is having very high correlation, 0.89 with horsepower.
[01:52:43] And 0.93 with the weight.
[01:52:46] But a negative correlation with acceleration, and even with a target value, there is a negative correlation.
[01:52:53] So, basically, as the engine, we can say that the engine of the…
[01:52:57] of the car, the displacement measures that is increased, the efficiency of the fuel is, like, decreasing.
[01:53:04] So, this tells us that. And similarly, like, we can see the correlation values for every other thing here.
[01:53:12] Yeah, so we can see for the target variable, uh,
[01:53:16] Uh, that displacement, horsepower, and weight, all of them are negatively correlated.
[01:53:21] As all this increase, the efficiency, fuel efficiency decrease.
[01:53:26] And with acceleration, it is having a, uh, like a mildly positive correlation. So, basically, it's only 0.3… high positive is considered, uh, above value of 0.5. So, yeah, but a positive correlation with acceleration.
[01:53:40] Here. So, this was for the, uh…
[01:53:43] Uh, like the correlation heatmap for the continuous features.
[01:53:47] So, uh, well, shall we begin with the next part?
[01:53:51] Ma'am? Okay.
[01:53:52] Yeah, I think… yes, yes, you can proceed.
[01:53:54] Okay, so then this was like the, uh, we are doing, uh, two, we are taking two features at a time, or…
[01:54:01] And taking numeric features. So, now we will do unit analysis, uh, analysis on each feature separately, not their pairwise analysis.
[01:54:09] Again, this analysis on numeric columns.
[01:54:13] And, uh, what we are showing is, for each column, we are showing the histogram.
[01:54:19] Its knowledge density plot. So, a KD plot was also visible there in the pair plot, in the diagonal values. It's just the same thing.
[01:54:26] And on the third is the box plot. So, for each of the values, uh, we can see the histogram and the KDA plot is normalized value of this histogram.
[01:54:37] And it's box spot for each of the four, uh, numeric features.
[01:54:41] Here. And, uh, like, uh…
[01:54:44] Uh, these, uh, we can see from this plot that, uh, like, this is also, uh,
[01:54:50] I observed earlier and noted earlier, that there are peaks here.
[01:54:54] So, it is called… displacement is a multimodal distribution, and we have shown classwise also.
[01:54:59] Horsepower is also multimodal, and left skewed, as I'm told, and acceleration?
[01:55:05] is having our, like, normal… kind of normal distribution.
[01:55:09] And the box plot, it can show us some more thing about data. So, basically, we see that
[01:55:16] And in the box plot of displacement,
[01:55:18] Uh, the median value's there, and apart from the whiskers, the end values of this box part, there are no such values.
[01:55:26] But in, uh, for horsepower, there are some, uh, values which extend these, uh,
[01:55:30] whiskers. So, basically, these are very ext… so, in horsepower, there are some extreme values also, which are not part in the quantile range.
[01:55:38] So, which are outside the, uh, quantile range. So, this is for horsepower.
[01:55:43] And similarly, you see in weight, there are no such, like, extreme values, weak, uh, or outlet type of values, but in acceleration, again, there are some…
[01:55:53] outer values.
[01:55:57] So, apart from that, all of them are, like, apart from them, again, what we do is…
[01:56:03] We do this univariate analysis
[01:56:05] Uh, but by color-based coloring, because we can do this because we have only, like, uh, we have class separated data.
[01:56:12] And there are only two classes here, so visualization also does not look that much confusing.
[01:56:18] So, the code is almost the same as above, just like you will see.
[01:56:24] separate plot for each of… by the classes.
[01:56:28] So that's why we are seeing two box plots here, because now there are… for task 0 and class 1.
[01:56:34] Uh, on the attribute displacement and everything.
[01:56:38] So, this, uh, like, can tell us further that, uh,
[01:56:43] Uh, uh, that in displacement, uh…
[01:56:48] If you were seeing only the normal one without separating class, there were no outliers. But for class one, uh, within its quantile range, if you consider them separate by class,
[01:57:01] There are some values which are in Class 1 also that are, like, outside of the range.
[01:57:07] And, uh,
[01:57:08] Uh, just to add here, the, uh, this box plot helps us to find out the minimum and maximum values. You can point there, Ajish.
[01:57:17] So, this is the maximum value, and then…
[01:57:20] The viscurs show the minimum and the maximum value, so the lower and the higher one will show that.
[01:57:26] Then we have, uh, the first quartile and the third quartile, uh, being shown and being separated by this median value.
[01:57:34] So this is a median value, then you have lower one, uh, will be your…
[01:57:39] of Q1, or first quartile, and the second one will be your third quartile.
[01:57:44] Then you have the min-max values.
[01:57:46] And, uh, the height of the box will show the amount of variance in the data. If the variance is…
[01:57:53] Higher the height will be high, whereas you can see, for example, for class 0, the height is very less, means the variance within the data is very less.
[01:58:02] You can see for Class 1.
[01:58:04] And then the circles are points that show anomalies. So, this is how you explain a box plot.
[01:58:11] It's quite a useful plot to…
[01:58:14] illustrate the…
[01:58:16] different, uh, things or different, uh, features about, uh,
[01:58:22] I mean, the different properties of any particular attribute.
[01:58:28] So, uh, after that, uh, like, what I'd like…
[01:58:33] So, written here some observation and summary of all the, like, the observation ma'am told. The same thing is written here.
[01:58:40] And finally, there is another… the last univariate plot, it is called, like, violent plot. So, violent plot, what is, uh, does is, it, uh,
[01:58:52] It has inside it the box plot also, and it also has that univariate, like, the distribution for each variable in it. So,
[01:59:01] Both, uh, both two things are combined here in the violin plot. That's why you are seeing the… this… this is the, uh, like, the, uh…
[01:59:09] plot of that single variable, and this is its box plot for displacement. So, if all the things are… it combines both of these things together in one plot.
[01:59:19] So, here, also, like, so basically, a violin plot, box plot, everything's, like, in Seabourn, uh, they have given exact functions for it, so as such, no code is required. There are direct functions in which you give the data frame, and you tell the column for which you want the violin pet or box plot.
[01:59:36] Similar, uh, and you get the plot.
[01:59:38] And this thing is, like, in loop, because I want violent plot for each of these four of them.
[01:59:45] So, these are the viral plots. So, basically, the observation will more or less will be same, because we have, like, individually also plotted these distributions, the box plot.
[01:59:55] So, observations remain same.
[01:59:57] And the last one is that, uh, these violin plots, again, we have plotted violin plot, but because we are insisting also on class-based separation, so again, the vinyl node, but class-based coloring.
[02:00:09] So, you see that now,
[02:00:10] Uh, for each of the class, we are getting separate violin plots for each of the features, the four numeric features.
[02:00:17] So yeah, so, uh, like, with this, we, like, finish with the, uh, like, the, uh, this notebook today.
[02:00:18] Yeah, over to you, Ashish. You can continue, please.
[02:00:25] Any questions, anyone?
[02:00:33] So, I'll request all of you to please keep pace, uh…
[02:00:37] Uh, with whatever is being taught in theory and hands-on, everything will be correlated. Additionally, also, please go and go back and have a recap
[02:00:48] On the hands-on also, because it's help you… it's going to help you code the future.
[02:00:55] Uh, thanks. So, because we are sequentially building up on the concepts.
[02:01:00] Any questions, anyone? Any queries?
[02:01:07] Okay, Abhishek, uh…
[02:01:08] Oh. Thanks.
[02:01:09] Followed by G2.
[02:01:11] So, ma'am, uh, here we are doing bivariate trotting, basically to understand the correlation, and then based on that, we'll drop or keep the attributes, right?
[02:01:20] And univariate is to identify outliers, and based on that, we'll remove the outliers, is that correct?
[02:01:27] So, univariate means considering one variable. So, if we want to look at the intricacies,
[02:01:33] of how any particular variable is behaving,
[02:01:37] We take univariate analysis.
[02:01:39] The purpose can be anything, looking at outliers, looking at, let's say you have a box plot, then you look at the lowest, highest values, how they are varying.
[02:01:49] Let's say I've got 5 different attributes,
[02:01:52] And I want to see how their lower and higher values are changing.
[02:01:56] So, I can do those 5 plots all on the same graph.
[02:02:00] And then compare the lowest and highest value in a very quick way. I can also have a look at the anomalies, I can look at the mean
[02:02:08] I'm a median values, and so on and so forth.
[02:02:10] So, if you are concentrating only on one attribute, use univariate.
[02:02:15] If you want to look at a combined distribution of…
[02:02:19] a pair of attributes, then you choose a bivariate.
[02:02:23] analysis. The purpose could be anything, as I told you.
[02:02:27] Okay?
[02:02:32] Is it okay now?
[02:02:36] Yeah, thanks. Uh, Jeetu, uh, you have a question?
[02:02:39] comparable. Uh…
[02:02:43] Mm-hmm.
[02:02:44] I would like to, uh… can you explain why this violin plot for, uh…
[02:02:47] see anything like mean, medium, or mold, or…
[02:02:52] A lot of the distribution bank alone.
[02:02:55] Why this viral platform?
[02:02:57] Can you brief it for me?
[02:03:01] Uh, Jito, your voice is breaking, however…
[02:03:05] Uh, I believe you want to understand the violin plot, is it correct?
[02:03:08] Yeah, yeah.
[02:03:12] Uh, okay, uh, Ashish, would you want to add on this?
[02:03:16] So, violent plot, like, it's just a combination of the box plot you're seeing, and the histogram, the knowledge density estimation we did. So, it's just combining both of them in one plot, and, like,
[02:03:29] It's nothing separate from that, so that's why it's, like, added in the end, and that's why, also not explain the observations on it, because these can be, like, manually seen there, but because it's, like, used in, like, analysis, that's why it was, like, added here.
[02:03:45] But uh… but these things, like the…
[02:03:48] data distribution and everything, these are, like, it captures both of things, like the box plot analysis also, and the distribution analysis also.
[02:03:57] I can add here that, uh, what is happening is, along the central line,
[02:04:03] Uh, we are trying to kind of mirror the density of, uh, so the density plot is getting mirrored, as you can see here, both the sides of this…
[02:04:13] central axis, uh, shows the distribution or the…
[02:04:16] Uh, density of the points.
[02:04:19] And then the frequency of the data points, uh, is shown at different values.
[02:04:25] Right, so if we are having wider area…
[02:04:28] then it means that more number of points are there. If it is narrower, it means lesser number of points are there.
[02:04:34] for that particular value, because these are nothing but the distribution of points, which are mirrored along the central line.
[02:04:41] Okay, uh, so this is what you primarily see.
[02:04:47] So, this is the overall shape of the violin, what it shows, the distribution.
[02:04:53] Then, uh, what you are seeing in the center would be the dots would show the median points.
[02:05:01] And, uh, then, uh, you also see the interquartile range here.
[02:05:07] With the help of the… these grid lines.
[02:05:12] Uh, that is what you are seeing here.
[02:05:14] And, uh, I think pretty much, uh, this is it. Then you are also having…
[02:05:22] I think, uh, outliers and all that stuff, so this is how, uh…
[02:05:26] More or less, this is what the violent plot…
[02:05:29] is representing.
[02:05:31] Okay.
[02:05:33] Anything else?
[02:05:36] Is it okay, Jeetu?
[02:05:40] Okay. Yeah, yeah.
[02:05:44] Anyone else? Yes, Dwarkish?
[02:05:48] Um, how do we visualize anomalies in this? Like, for box plot, uh, we were able to see the anomalies, right? So, uh, how can I interpret anomaly from this one?
[02:05:58] Uh, so… if you look at the lines that are extending,
[02:06:03] From the rest of the data distribution,
[02:06:07] So, you can see these pointy sides.
[02:06:10] These are potential outliers here. You see the data distribution?
[02:06:14] Oh…
[02:06:15] Yeah, yes. These are the potential outliers, because the data distribution is lying more on the… this, uh…
[02:06:22] this fatter part, you can see the…
[02:06:23] You can call it kind of the flesh of this…
[02:06:26] Yeah, uh, you can see the cards are there, you can see the…
[02:06:30] broader part, this is the major bulk of data distribution.
[02:06:34] But, uh, the whiskers here are the pointy parts.
[02:06:39] And these are what the anomalies would be.
[02:06:41] Because very few points are there. As I already told you, this shows the count or the distribution of the points. I mean,
[02:06:47] frequency of the points.
[02:06:55] Any other questions, anyone?
[02:07:04] Okay, so if we don't have any questions, then shall we wrap up today?
[02:07:12] Okay, then, uh, we'll wrap up here. Thank you so much. Thank you all. Have a great day. Thanks, Ashish.
[02:07:14] Sure, yeah.
[02:07:20] Thank you, bye-bye.
[02:07:21] Yeah, we are able to.
[02:07:22] Thank you.
[02:07:23] And go to him.
[02:07:24] Yeah. Thanks.
[02:07:25] Thank you.
[02:07:26] everyone.