# 13 2026-01-17 Hands On Unsupervised Learning

course: Module 2 — Machine Learning Algorithms
module: Module-2-Machine-Learning-Algorithms
date: 2026-01-17
type: transcript
video_url: https://personal-learn.armco.dev/files/_Recordings/Module-2-Machine-Learning-Algorithms/13_2026-01-17_Hands_On_Unsupervised_Learning.mp4

---
[00:41:01] Okay.
[00:41:02] Yes, ma'am. Uh, so to, uh, like… Um, reduce the computational overhead, so that's why we will be dropping out the duplicate rows from our dataset.
[00:41:15] So, uh, drop.duplicates. automatically help us to, uh, you know, remove the… it will identify all the duplicate rows, and it will remove.
[00:41:25] the duplicate rows and keep only the original or one row for the multiple, uh, same entries.
[00:41:33] So, and then, uh, as we can see, um, after dropping out the duplicate.
[00:41:39] Rose, we have, uh, 13,543, uh, we are left with 13,543 rows only, or recalls.
[00:41:48] So, um… Here, we are checking what are… The unique, uh, values.
[00:41:54] In our class column. So, uh, what are the labels?
[00:41:58] That are assigned to the, um… dry bean. So, we have, uh, 7, uh, unique.
[00:42:08] Values in the class column. And, uh… Value counts, uh, as we have used it before. So here you can see the clear user value comes. What it does is.
[00:42:17] It, uh, takes the unique value from the… from that column, and.
[00:42:23] count how many rows has. Uh, that unique value.
[00:42:28] So, uh, this, uh, Dermicin class. is, uh, there in 3,546 rows. And similarly, for the other class… for the other class labels, we have the count.
[00:42:40] Like, how many rows contain that class label. So, uh, this is the use of value counts.
[00:42:46] Uh, in our data frame. So, uh, now we'll plot, we'll, uh, see the histogram.
[00:42:56] plots for. R16, uh, different, uh, columns.
[00:43:03] which we have in our dataset. So what we are doing is, we are extracting the unique classes.
[00:43:09] from using the unique function. Um, from the class column, and we are, uh, extracting those, and we are, uh, like, saving it in a unique class.
[00:43:18] Data, uh, sorry, uh, variable. So, um, this unique class will contain a list.
[00:43:25] Of the unique values of the. Uh, class column.
[00:43:32] Then we are, um… using SNS. So, uh, as initially, I told you, CBON is used to visualize the data. So, SNS, we have given the, uh, acronym.
[00:43:44] to that, uh, seaborne, uh, as SNS. We imported it using SNS acronym.
[00:43:53] So we are using SNS, and from that, we are using the color palette, um… method. Uh, so what this, uh, does is it… So, uh, like, how we will differentiate between different classes, it will just do that. So, what will the… what will be the palette of our, uh.
[00:44:13] these plots. So we are defining it here. So, number of colors, like, end colors will define how many colors we are using, so it will be the length of the unique classes, which is 7.
[00:44:28] Uh, then we are, like, uh… creating a dictionary from it, and we are passing it to class colors, because.
[00:44:35] Uh, why we are doing it, because later we will see here we are using class colors here, so the palette only takes the dictionary value, so that's why we are converting our palette into dictionary.
[00:44:49] Then we will extract the numerical columns from the dataset.
[00:44:54] So, we will, uh… So, this is the command that we are using to do that. So what we are doing is we are dropping.
[00:45:02] Out the class column, because it was categorical in nature, and after dropping it, we are just extracting all the columns, because rest of the columns were numerical.
[00:45:12] And, uh, that's how we'll create a list of numerical columns.
[00:45:20] Now we are initializing the subplots, so we will create a 4 cross 4 grid of the plots.
[00:45:26] Uh, because we have 16 columns in the dataset, so 4 across 4 will do our job.
[00:45:33] So, uh, how we are doing it, uh, we are, like, from Matplotlib, we are calling the function subplots.
[00:45:40] And here we are telling how many… Uh, like, for that grid.
[00:45:46] Uh, how many, uh, columns will be there, and how many rows will be there. So, 4 rows and 4 columns we need.
[00:45:51] Or four cross four grade we need. And what is the figure size of each subplot?
[00:45:59] And, uh, this will return the overall figure. And the, uh… So, here we are, uh, like.
[00:46:07] AX variable will store the subplot location, so AX will have… Like, subplot 1, subplot… or we can say subplot 0, subplot 2, subplot 3 subplot.
[00:46:20] Plot, uh, 4. So, like this, it will have the values, and the figure will have the overall figure.
[00:46:25] a location. So now, what we are doing is, we are calling it on the… like, we are iterating it over the variable, and variable will store the.
[00:46:36] numerical column, and the subplot will store the. location of the, uh… for each, uh… numerical column. So, uh, here, what we are doing is, from numerical columns.
[00:46:49] we are iterating over the numerical columns, and uh… At each iteration, we are, uh, like, assigning it to the.
[00:46:57] variable, variable. And, uh, we are also, like, we have, uh… ex.flaten. So, why we had flattened it? Because, uh… let me show you.
[00:47:51] display what we have in AX. So, as you can see.
[00:47:55] AX is a, uh, to, uh, like, a 2D array.
[00:48:01] Right? Because we have an array, and then again, we have… this. So, uh, this is how AX stores the location of each subplot. So.
[00:48:10] Uh, dot fla… like, this function, what it will do is, uh, it will make it easy.
[00:48:18] So we have to iterate over every location. So we cannot, uh, like, it will be very difficult to iterate over this 2D array. So, uh, we'll convert it into a 1D array, and then we'll iterate over it.
[00:48:28] So, after using this. my AI, like, after using flatten.
[00:48:37] Uh, method on… My AX will… be converted into a 1D array. And, uh, it will be easy for us to iterate over the 1D area, so that's why we are using flatten function here.
[00:48:53] So, in this manner, we will iterate over each column and each location for that column.
[00:48:58] And then we are calling the hist plot function from the CBON, um… library, and we are giving it the, uh, DF beans.
[00:49:11] data, and we are telling that X. So, uh, X will tell, like, what should go on the x-axis, so variable, like, the column name.
[00:49:21] And, uh, what location? So, AX is equal to… subplot is telling that area should be at location 0.
[00:49:28] And, uh, this, uh, hue is used to tell, like, uh, what, uh, like, how we want to differentiate every, uh.
[00:49:37] the subplot in the, like, the histograms in the… in one subplot.
[00:49:44] So, we want… we want to differentiate using it… using the class func… uh, like, the class labels.
[00:49:51] And, uh, palette we have already described using the class colors, and then, uh, alpha.
[00:49:56] Okay? And… Uh, okay, so we are also, um… Like, you can see this, uh… dashed line over the plot. So, this tells us the mean of the area.
[00:50:11] Uh, column, which we, uh, calculated before also using the describe function.
[00:50:18] So, I want to, uh… Like, know where my main.
[00:50:21] is for the overall data. So that's why I am also, like, over that… over my plot, I am drawing a line.
[00:50:29] And to draw that line, this code is being used. So what I'm doing is… Um, for my… for my that, uh, like, for my area column.
[00:50:39] So, variable contains my first column right now in my first iteration.
[00:50:43] So, I am calculating the mean. So, mean function is used to calculate the mean of the overall data.
[00:50:49] of that column. So, on x-axis, I want to, like, plot that line.
[00:50:55] On the x-axis where my mean of the overall data is.
[00:51:00] Color is M, so this is used to define the color of my line, which is this magenta color.
[00:51:06] Label should be mean, line style is dash line style, so that's why it is, uh, dashed here.
[00:51:12] And the line width is 2. So this is the meaning of this one-line code.
[00:51:18] And, uh, these are… like, to plot my overall figure, I'm using plt.tyte layout and plt.showto.
[00:51:26] or display my whole figure. So, in this manner, the histograms will be plotted for each column in the dataset.
[00:51:35] And as you can see, uh, every subplot. has, like, it is showing histograms for each class also. So here, this dark brown color.
[00:51:46] As you can see, it is here. Uh, it is overlapping, so this is for the saker class. Bombay class has the orange one.
[00:51:54] So, Bombay glass is lying over here. Uh, and the mean is here for the overall data. So, in this manner, the classes are being distributed in this.
[00:52:04] subplot.
[00:52:08] Uh, so just wanted to add here, if you can just.
[00:52:12] Store out slightly. For example, if you look at the perimeter, you can see here that there are different varieties of beans have different kinds of perimeters.
[00:52:23] For example, you can see that the dermas are.
[00:52:26] Uh, the seeker has a very low perimeter, so these must be very small.
[00:52:32] Beans, uh, sizes. This variety must be. Beams of very small sizes, whereas if you look at.
[00:52:39] Uh, Dharma Sun. They have very large parameters, comparatively.
[00:52:44] Whereas if you look at the total number of beans, that is the.
[00:52:50] Number of samples belonging to these kinds, such as Seeker, they are quite high.
[00:52:54] As compared to the samples of Dermason, which is quite low.
[00:52:59] Uh, you can also see parameters which are in between, like that of Kali.
[00:53:03] Uh, similarly, if you look at the… Solidity, I think you can go to that, but you can focus on that.
[00:53:11] Yeah, so if you look at this particular graph.
[00:53:12] Yes.
[00:53:14] Then you can see here that, again, there is a peak.
[00:53:18] Which is probably around, uh… Oh, so if you look at the color, probably Dharma sun.
[00:53:24] and Seeker. They are quite, uh… They are very compact and solid, whereas the others.
[00:53:31] Specifically, the one that is shown in the blue color, I think the horos.
[00:53:35] is quite low. So, similarly, you can see the different kinds of, uh.
[00:53:40] attributes being shown here for the different varieties, like what makes up these varieties, how these are different.
[00:53:48] Or how similar these are. Uh, you can look at the extent, you can look at the.
[00:53:53] Uh, other features as well. So, this helps you understand the different varieties.
[00:53:58] Of the beans, or could be any data type that you are considering.
[00:54:02] So I thought I'd just add that here. Sorry to interrupt, Asmita, you can please carry on.
[00:54:08] Thank you, ma'am, thank you. Uh, so, uh… Now, we are looking… so, here we looked at the histograms of, uh, various columns.
[00:54:19] But now we will be looking at, like, um.
[00:54:22] how, like, the area is distributed over these, uh, different classes. So, uh, to plot that, what we are doing is, we are again taking a 4 cross 4 grid, so that our 16 columns will fit in.
[00:54:39] So, uh, similarly for each variable and for each location, we are, uh.
[00:54:45] plotting a strip plot. So, why it is called a strip plot? Because here we can see the range.
[00:54:50] Of that column for that particular. class. So that's why… so it's like a strip, like, it is distributed as a strip.
[00:54:58] hair. So that's why it's called a strip plot. So we are, uh, again, um, passing our, uh.
[00:55:04] Uh, data set. There, then, um… Right, on y-axis, we want our… this time, we want classes on the y-axis, the class names on the y-axis, that's why we have.
[00:55:17] Uh, you know, past that for the y-axis variable.
[00:55:20] Um, again, the location of the subplot, the palette, and alpha. And we are printing it. So, uh, what this tells us here, we can see that.
[00:55:29] Uh, area. So, for our column area. And for class Bombay, we can see that the dry beans have higher values.
[00:55:40] For area, for the Bombay class. So maybe they are, like, bigger beans. So, Bombay class have bigger beans.
[00:55:47] And similarly, the perimeter is also, uh, have the high range. So it, like, uh, from 1200 to 2000, it is… the values are ranging between this. For seeker, the perimeter ranges between, like, 600.
[00:56:01] And it is under 1,000. So, uh, this… these plots, uh.
[00:56:06] help you visualize what is the range for each class for that particular.
[00:56:11] column. So, for area, as we can see, for horos, uh, it is less than… Like, uh… this 1 lakh, I guess.
[00:56:21] If I'm… yeah. So in this manner, these plots will help us identify the range for that class and for that particular variable. What is the range of values that class.
[00:56:35] holds. So you can visualize it easily from these strip plots.
[00:56:44] Then we are displaying the correlation matrix. So, as I… told you here, uh, area and perimeter. So, of, uh, like, obviously.
[00:56:56] And, uh, like, it's… a common, or… it is implicitly… we can implicitly derive that if a class has a higher range.
[00:57:06] For area column, it will have a higher range for parameter column, so they are directly proportional. Area and perimeter is directly proportional. So, correlation also tells us that.
[00:57:18] Uh, so, to display that, we will be using the heat map from the Seaborn.
[00:57:22] library, heat map function. And we are passing the, uh, all the numerical columns, like the data frame for all the numerical columns, and we are calling in the function coral.
[00:57:34] correlation, uh, and. Uh, then, uh, like, these are basically that we want the, uh, white tick labels, like, we want these labels here.
[00:57:47] And we want the annotations, so this is… naught is… annotations is true, that we want, uh, on our graph, we want all the values of the correlation that we have calculated, and CMAP is the color map.
[00:58:00] we are using CoolBomb color map, and we are plotting it.
[00:58:05] plt.show, we applaud… we are displaying our plot. So here you can see that area and perimeter has a correlation of 0.97. So the diagonal values will always be 1, because correlation with itself.
[00:58:18] a column, uh, it will calculate correlation of a column with itself, it will… it is always 1, so the diagonal values are of no use for us.
[00:58:26] But the other non-diagonal values, we have to look… we have to, like, have, uh, have a look at them.
[00:58:34] So, area is correlated. like, uh, it has a correlation value of 1 with, uh… convex area column.
[00:58:41] So, uh, they are highly correlated. As you can see from the value also. Also, parameter and area is correlated, and major axis length.
[00:58:52] Uh, and, uh, areas correlate… highly correlated with minor axis length.
[00:58:57] So, and also with the… equidimeter equivalent diameter.
[00:59:05] column. So, it has a value of 0.98. And similarly.
[00:59:10] Uh, all the other… the correlation values of all the other columns to the, uh… Different to… is shown here in the correlation heat map.
[00:59:33] So, um, now we are using another type of visualization, which are… which is known as pair plots.
[00:59:40] So, uh, pay plots actually, uh, what it does is… Again, it, uh, calculates.
[00:59:48] The, uh, it, uh, plots the, uh… Every plot against.
[00:59:56] itself and the other columns in the dataset. And if… Uh, we have a strong curved.
[01:00:04] Uh, plot or a curve plot, or a linear plot, which means that those columns are correlated, or they have.
[01:00:13] some kind of relationship between them. So here we can see in the pair plot that the perimeter and area.
[01:00:22] Uh, have this, the, uh, like… the dataset, uh, if we plot perimeter and area, like, area on the.
[01:00:30] x-axis and perimeter on the y-axis, so we can see that the, uh… this scatter plot has a.
[01:00:37] has a curve here. So, which means that they have a relationship between them, a positive relationship or a negative relationship, that we don't know, but there is some relationship between the.
[01:00:48] dataset. And in a similar manner. aspect ratio and eccentricity is also.
[01:00:55] uh… relate, correlated. So, pair plots, um… I have… we have, like, uh… So, to visually spot correlations, so, uh, as we, uh, plotted the heat map.
[01:01:09] And, uh, so to visually, uh, or, like, to see it as in the, uh, scatterplot form, we use the pair plots.
[01:01:18] So, tighter linear curve trends. If the data set, uh, represent or depicts this kind of, uh… behavior, which… that means that the two columns are correlated.
[01:01:34] So that's why we have plotted this spare plots. So, the, uh… The columns which are not correlated.
[01:01:42] For them, the dataset is, like. Uh, here you can see aspect ratio and area. So, they are not correlated columns, so that's why they're, uh.
[01:01:51] The data points are scattered across, or they are clustered between different classes, because we have used hue as a, like, for hue, we have used class as a label, so that's why, uh, here.
[01:02:04] The data points have formed clusters. So these are the pay plots for all the columns.
[01:02:17] Now we are, like, uh, we have seen. In a text form, like, how our classes were distributed through the dataset, like, using the value count function.
[01:02:27] To visualize it, uh, in a bar plot form.
[01:02:32] So, that's why we are using a countplot here.
[01:02:35] To visualize the distribution of each class throughout our dataset.
[01:02:39] And, uh… This is the distribution of each class in the dataset. So this is the count plot.
[01:02:57] So now we'll prepare our data, uh, for k-means clustering, or to use the algorithm.
[01:03:03] So, by preparing the data, we mean that. Um, it will, if I scroll up.
[01:03:09] Uh, where we have displayed the data.
[01:03:16] Yes. So here you can see the area, uh… has values ranging in thousands, like 28,395.
[01:03:27] But the perimeter, uh, have values. in the parameter column, we have values which are in hundreds.
[01:03:35] So, uh, and similarly in major axis lens. So, we have.
[01:03:40] the range of the values in each column is very different, and it, like.
[01:03:45] changes drastically, like, some, uh, columns have values in 1000, and some are in one digit also. So, um, to… Like, so, it will not… prepare our dataset accordingly, or if we'll not scale our dataset, what will happen, the algorithm will get influenced by the values which are.
[01:04:05] which have a higher range. So, uh, basically what I'm trying to say is, like, my k-means will be influenced by the column.
[01:04:12] like, area, because it has higher range in that, because it will not be able to differentiate between the ranges, like, the columns has different ranges.
[01:04:21] So, to, like, uh… To overcome this, to overcome, uh… Uh, this mistake that, uh, my algorithm doesn't.
[01:04:30] make a mistake by the range of the values. What we will do, we'll prepare our data set, so what we are using.
[01:04:38] a standard scaler to scale the values between the range 0 to 1.
[01:04:44] Uh, for each column of the data. So, firstly, what I'm doing is I'm mapping.
[01:04:52] My class variable to numbers. So, my class has categorical.
[01:04:57] values, uh… Uh, in the column, but I want to, like, uh… turn it into a number, like, from 0 to 6. So I want to map.
[01:05:07] that. So for that, I'm using here, uh… So this code is mapping my unique classes to numbers. So what I am doing is I'm, uh, iterating over the.
[01:05:19] class, and I'm extracting the index and the class name.
[01:05:24] Uh, from this… uh, data frame, and then I'm, uh, like… extracting it in index variable and the class variable, and then I'm creating.
[01:05:38] a dictionary out of it. So, in this one line.
[01:05:42] It is a little bit complicated, but in this one line, what I am doing is, I am extracting my class and the index and storing it in index variable and CLS variable, and then I am creating a dictionary out of it.
[01:05:54] So, uh, and this dictionary is being stored in the variable class map.
[01:05:59] And this class map is then… I'm… what I'm doing is I'm creating a new column in my data frame.
[01:06:05] With the name class number, and then I'm, uh, passing this dictionary.
[01:06:11] the dictionary that I have just created. Uh, I am pass… I'm, like, this class number will have the values of this dictionary.
[01:06:20] So, um, now we will prepare our features and target variable, feature variable and target variable, so X will store all the features, which are the numerical columns of the dataset, and Y will store the.
[01:06:33] target variable, which is the class. So, uh, for… so how I'm doing that? I'm dropping class and class number, because now I have class also, and class number also, which I have just created here.
[01:06:46] So I'm dropping that, and uh… I'm only extracting the numerical variables in X, and I'm storing the numerical variables in X.
[01:06:56] And for, uh… so Y2 contains the true labels, uh, of the dataset, and which is the class number that I just created here.
[01:07:07] this column. So Y will have… Y true will, uh, store.
[01:07:12] The true labels of the dataset. So, using standard scalar, um… function, I will be scaling my, uh, full data set, and, uh.
[01:07:24] It will be stored in Xscale. So, standard scalar is Z-score transformation.
[01:07:28] Um, we have a min-max scalar also, which uses the minimum and maximum value to transform our data, and a standard scaler uses the standard deviation and mean values of the data to transform the.
[01:07:42] data into, uh, in the range of 0 to 1.
[01:07:48] So XScale will contain the scaled values, or the normalized features of the data frame.
[01:08:01] So now, we will be determining, um… our optimal number of, uh, like, optimal number of clusters using the ELBO method. So, I'll be using.
[01:08:13] Um, I'll be importing, uh, k-means library from the, uh, sklearn.cluster library.
[01:08:21] And so I will be checking, uh, the range of K from 2 to 35.
[01:08:28] So this is the range that I'm defining for which I'll be checking the within-cluster.
[01:08:34] Uh, errors.
[01:08:35] Uh, just sorry to interrupt, Rasmita, I think there's some, uh… Question on the chart that in the heat map.
[01:08:43] Uh, if the area goes off the convex area also almost goes up, is the.
[01:08:48] Is this suggesting, uh, redundant data analysis? Uh, no, it is not suggesting redundant data analysis.
[01:08:57] Whatever, uh, uh… features are positively correlated.
[01:09:03] Uh, they will, uh, behave similarly. In this case, it may be area and convex area, which may be actually meaning two different things.
[01:09:12] A convex area might probably, uh… Be mathematically different from the simple area.
[01:09:19] And that's why the two have been considered separately here in the dataset.
[01:09:24] And, uh, when one goes up, the other also goes up, which means that there is a positive correlation.
[01:09:31] So this is what is important to… In far from the heat map.
[01:09:36] Uh, that there can… The attributes that behave similarly. Now, this information can be used later.
[01:09:45] Uh, in some ways, for example, if we have to reduce the number of.
[01:09:49] attributes than those which are behaving very similarly. Can be reduced. So, let's say that if, um, in our analysis for us.
[01:09:59] Uh, area and convex area. Even though they are, they may be calculated mathematically differently, but if they are meaning.
[01:10:08] Some similar, ah. notion, then what we can do is, based on the similarity.
[01:10:16] Uh, in their behavior, we can just ignore one of these. So, in a way, what you are saying is.
[01:10:21] Maybe correct in some cases that one may be treated as redundant if that.
[01:10:27] If two attributes are behaving similarly, then we can reduce the number of attributes.
[01:10:31] On that basis. So, for this purpose, also, correlation helps.
[01:10:37] So, uh, if we have, uh… Attributes which are very similarly behaving, we can.
[01:10:43] Uh, consider one or few of these and ignore the others.
[01:10:47] Okay. So, this is what I wanted to add. Maybe, Asmita, you can continue from here.
[01:10:53] Thank you, thank you so much, ma'am.
[01:10:55] Okay, ma'am. So, uh, I'll be checking for the, uh, for the values 2 to 34.
[01:11:03] Uh, because this is, like, uh… Open bracket situation here. So my values will be checked for 2 to 34.
[01:11:13] For… to… for the K range of 2 to 34.
[01:11:16] Then, you know, I'm initializing a list, inertia, which will store, uh, for each, uh, uh, K, uh.
[01:11:24] it will store the within, uh, uh… clusters sum of square errors.
[01:11:30] So, uh, I'll be iterating over the range. And then, uh, k-means is called, like, this library, which we have.
[01:11:39] imported. So we are calling this function k-means, and we are telling it how many clusters we want. So, uh.
[01:11:48] hair will be… if for the first iteration, it will be 2, for the second, it will be 3.
[01:11:54] So for each value, it will be calculated… k-means will be calculated.
[01:11:59] And, uh, also, uh, so that the k-means will.
[01:12:03] run properly, we are running k-meins 10 times, and then we are, like, uh, calcul… like, uh.
[01:12:10] calculating the, uh, within sum of square errors. And random state is used, uh, to, like, rep… for reproducibility.
[01:12:19] So, if you'll change this, maybe, uh, because, uh.
[01:12:23] G-means is a… it uses some values, uh, differently, so, uh, for every… if we are not using random state, for every iteration, it will calculate a different value.
[01:12:33] for the within cluster sum of square errors. So 42 is being used for reproducibility.
[01:12:39] And then we are fitting it on the scale data that we had of the numerical values, and we are appending.
[01:12:46] Um… so append is used to, uh, append to the, uh.
[01:12:52] list, and we are appending the, uh, error. to the inertias, um.
[01:12:59] variable that we have initialized here. So, in this manner, we will calculate all the errors for each K.
[01:13:06] And then we will plot. it using this figure, we'll be calling the figure function, then we'll plot.
[01:13:13] Uh, on the x-axis, this is a very simple plot. We'll plot the K range on the x-axis.
[01:13:20] the errors on the y-axis, and we are, like, this, uh, dotted.
[01:13:24] line we are plotting, so that's why marker of O.
[01:13:28] Then the… so this is basically how we are labeling our.
[01:13:34] x-axis, so for x-axis, I have number of clusters, for y-axis, I'm labeling it using inertia within cluster sum of square errors.
[01:13:41] And the title of the whole plot is Elbow Method.
[01:13:44] Uh, we have X sticks, so… Each stick, what it is, so it is… the range of the, um… Okay, and we are, uh… so this… this grid is being shown, so it… why it is…
[01:14:00] There, because we have set it true, and then we are displaying our.
[01:14:04] plot. So in this manner, this plot has been created.
[01:14:09] So, from this, as we can see, after 7.
[01:14:14] my values are gradually decreasing, but we cannot say, like, uh, still we… we are not sure, because it is not a clear elbow.
[01:14:22] So what we will be doing is, we'll be, um… Uh, calling our k-means using 6, 7, and 8 clusters.
[01:14:31] And then we will see the validity using the intrinsic and extrinsic.
[01:14:38] Uh, validity, uh… metrics.
[01:14:43] So, um, here… Sorry.
[01:15:38] Okay, so if we'll, uh, see… For every K, we can see we have, like, displayed here the, uh.
[01:15:46] Errors for every K also, and after 7, we can see that.
[01:15:50] The error is, like, decreasing. But not as abruptly as it was decreasing before 7.
[01:15:59] So we will take 6, 7, and 8 as k values, and we will evaluate on that.
[01:16:05] So, here, from the metrics library, we are calling out all the.
[01:16:10] mattresses that we want to… that we want, uh, or we want to calculate or display.
[01:16:19] For, uh, each K. So we have silhouette score Kalinsky, Herba score, Davis Bolden score, random score, adjusted random score and adjusted mutual Information Score, homogeneity Score, and Completeness Score.
[01:16:33] So we have all these, and we will, uh, like, compare.
[01:16:37] Uh, for every K, we'll compare these values, and we'll see which one is better.
[01:16:42] So, we are initializing our number of clusters as 7, and in the same manner as we have done it.
[01:16:49] For the elbow, uh, graph, we are calling it.
[01:16:53] Like, we are calling k-means on the n number of clusters, and we are now.
[01:16:57] Um, like, predicting the labels also. Like, the clusters also. So we are storing it in the ViPredict.
[01:17:06] label. So, how many labels we will have here? Uh, it will depend on the number of clusters that we have defined here. So we'll have 7.
[01:17:13] predicted, uh, uh, labels in the YPredict. So, YPredict will have 7 unique values.
[01:17:22] in it. And then, uh, using these, which we have imported.
[01:17:26] these functions, we'll, like, calculate the interest… intrinsic measures.
[01:17:32] So, intrinsic measures don't use the actual labels of the data, so it only, uh.
[01:17:38] It only tells us about the completeness of the cluster, or the compactness of the cluster.
[01:17:45] But extrinsic measures uses the… they compare the. both the actual labels and the predicted labels that we have predicted using the k-means. So that's why here Y2 values have been used, but in the intrinsic measures, we are not using Y true values, the true labels of the dataset.
[01:18:04] And we are printing them. So here you can see the… Different values that.
[01:18:13] we have calculated. That, uh, like, the function has calculated, and it has been displayed here.
[01:18:23] So, here is a little, like, ma'am, as ma'am, uh, told you, she will, like, uh, thoroughly, uh.
[01:18:32] give you a brief about what are these matrices, but here also you can find out what.
[01:18:38] each matrix is used for.
[01:18:46] So now we, uh, we'll be visualizing, uh, the clusters that we have just predicted, or we have just, uh… identified using the k-means clustering.
[01:18:55] Uh, just, uh, sorry to interrupt, Asmita. Any questions, anyone, so far?
[01:19:02] Yeah, yeah, I think I've shared in the chat as well. So, we have done EDA, and we figure out.
[01:19:03] Any questions? Yes, please.
[01:19:08] Like, we have done some Instagram plots and everything.
[01:19:12] But how does it correlate how, like, how, uh, why it is important for clustering? I… That's what I wanted to understand, because we have done EDA, we understood how it is, uh… the class is plotting with respect to different feature, but what did we…
[01:19:28] uh, extracted out of that step. That is something I wanted to understand.
[01:19:33] So, let me add here, Deepa, that the EDA that we showed.
[01:19:37] Was just kind of a preliminary analysis. It doesn't add… it may or may not add any direct value to the clustering process itself.
[01:19:48] It is just giving a general idea about the data and the different features.
[01:19:52] Uh, in the data. So, it may not directly be relevant to clustering.
[01:19:57] What we just added that to show you. Uh, that how EDA can be used.
[01:20:04] Uh, to understand what are the different attributes, like, what are the range of values.
[01:20:10] And, uh, what may be the… Correlations between them and all that.
[01:20:16] Uh, so, it is just like a data pre-processing kind of a thing.
[01:20:20] But it's not directly related to the clustering itself.
[01:20:24] Uh, what is directly related to clustering itself are the plots that you are going to see here.
[01:20:31] In which PCA has been used to reduce the number of attributes.
[01:20:34] Uh, to, uh, uh, to a 2D scenario. And then the plots have been shown.
[01:20:42] To see what the clusters look like for different values of K.
[01:20:47] Okay.
[01:20:52] Any other questions, anyone?
[01:20:58] Mm-hmm.
[01:20:59] Ah, yes, ma'am. Uh, just wanted to know, uh… In clustering, it's unsupervised, right? So, where, uh… Why are we using true labels here, as in?
[01:21:03] Yeah.
[01:21:08] What color are you doing for True Labels here?
[01:21:09] Sorry, so… Yeah, yeah. So, actually.
[01:21:13] Since we want to understand and learn clustering. So, good way to do is, just for the learning purpose.
[01:21:20] In the real life, you may not have any labels.
[01:21:23] Uh, given to you, but right now, we have used a dataset in which the class labels are given.
[01:21:29] So that when we do the clustering. We can actually use the extrinsic.
[01:21:35] Uh, clustering, um… Uh, measures… To understand how good the clustering is.
[01:21:42] So, actually, the class levels may not be… Uh, even available in the real world data where you actually don't know much about the data, it may be a new data.
[01:21:53] But we chose a data which has the class labels just to.
[01:21:57] Uh, understand how to assess. The clusters that have been formed, because the clusters that have been formed.
[01:22:03] Can be compared with these class labels, because these are nothing but the groups or the.
[01:22:09] clusters that are available in the data and that information is.
[01:22:12] Explicitly conveyed. By these class levels.
[01:22:17] So that is a use case why we have used this kind of data.
[01:22:21] But you can use a data that doesn't have any class label.
[01:22:25] We purposely used it. Just to show how good our clusters will be, because.
[01:22:30] We'll have the class label so we can compare our clusters with those.
[01:22:35] Uh, class labels that are given. That's the reserve.
[01:22:41] Yeah, okay, got it.
[01:22:45] Any other questions, anyone? Yes, Shivans?
[01:22:47] Oh, hey ma'am. One quick check, ma'am, uh, to understand better the inertia.
[01:22:54] Uh, with respect to the intrinsic measures and extrinsic measures.
[01:22:55] Mm-hmm. Mm-hmm.
[01:22:58] So, when we have jotted down these scores, is this directly represents initia, or… some other correlation.
[01:23:06] No, in… yeah, so actually… There are different extrinsic and intrinsic measures.
[01:23:13] I haven't yet taught you, and I'll be teaching them to you.
[01:23:17] So, they are used to measure the goodness of clustering. Inertia is one extrinsic.
[01:23:24] Uh, cluster validity index. That is used to compare.
[01:23:28] The given class labels in the data. With the clusters that you have formed.
[01:23:35] So, if the clusters match with the given class labels.
[01:23:39] Uh, then, um, the inertia, accordingly will be assessed. So, it is nothing but a.
[01:23:46] extrinsic cluster validatory index, okay? But there are other measures, tester validity indexes, like, for example, the.
[01:23:55] Uh, intrinsic ones are more popular, like, as I said, the done index, Davis Golden.
[01:24:01] cell heart and all. I will be telling you what these mean in the my theory class.
[01:24:02] Yeah.
[01:24:07] As of now, you just assume that they convey the goodness of the cluster.
[01:24:12] Inertia is an extrinsic measure.
[01:24:16] Okay. Got it.
[01:24:21] Anyone else? Any other questions?
[01:24:26] Uh, okay, Smitha, you can carry on from here.
[01:24:30] Okay, ma'am. So we'll visualize now the clusters that we have formed using k-means, um, through, uh, by doing the.
[01:24:40] PCA, uh, analysis, or the PCA decomposition. Uh, so because we have, uh… 16 columns in our dataset, so we cannot plot all the 16 columns in one plot, and.
[01:24:53] Uh, like, see how our clusters have been formed. So we are using PCA for two components, so we are reducing our dimensionality to 2, so that we can.
[01:25:05] plot it on a 2D graph. Um, so what we are doing is, we are, like, extracting the class labels.
[01:25:11] So, uh, we'll compare both the. the actual labels, and we'll see how.
[01:25:18] Our data set is divided on, or grouped on the actual labels, and.
[01:25:23] We will also see side-by-side, like, how our data.
[01:25:28] after using the k-meen. So, how k-means has divided the dataset?
[01:25:34] So we'll compare both. The two labels also, and the labels that are being formed using k-means.
[01:25:40] So that's why we are extracting the class labels, then we are calling PCA.
[01:25:45] Uh, and we are defining the number of components that you want to use.
[01:25:49] And, uh… this is used, like, on the numerical columns.
[01:25:55] On the scale numerical columns, we are calling it the fit transform.
[01:25:59] And we are, uh, like, uh, extracting it in XPCA variable.
[01:26:05] Uh, so… Uh, so again, we are extracting the unique class variables, the class to color, so this… I have already, uh… Like, described by how we are doing it, we are taking the index, the class label.
[01:26:21] And we are forming a dictionary out of it.
[01:26:22] And then, uh, we are using this as Vitru named function. We are extracting it in Y2Named function.
[01:26:31] So, uh, we are using subplots. So here, as we have formed… initially, we have formed a 4 cross 4 grid, now we are forming a 1 cross-2 grid, because only two graphs that we have… we have to display.
[01:26:44] So the first plot… will be off the data set.
[01:26:50] grouped by the true bean labels, and the second, uh, graph will be of the dataset, grouped by k-means cluster labels.
[01:26:59] So what we are doing is, um… we are calling axis 0, which means on the zeroth location, I want to scatter… I want to plot a scatter plot.
[01:27:07] Uh, on the XPCA value. Uh, so, uh, as we have only two, uh, we have taken only two components, so we have a zeroth component and the first component.
[01:27:19] So, on the x-axis, we have our zeroth component, and on the y-axis, we have our.
[01:27:23] first com… the first component, the second component. And C is used, uh, CMAP is used, it is telling the color map of the… of our plot, and then, um, what we are doing is we are giving the title, the X label and the Y label to the plot.
[01:27:45] And, uh, similarly, uh, uh… Here, we are telling, like, the color.
[01:27:52] how we want to, uh, like, how we want to cluster… show the clusters.
[01:27:57] So, what we want to do is, we want to, uh, label the clusters using the class labels.
[01:28:04] Like, each cluster should have a different. color, and it should represent the class of that cluster. So this handles will do that.
[01:28:14] Then we are plotting our legend also. So, Legend will be plotted using this handles, um, variable.
[01:28:22] The title will be class, the location will tell where I want to place my legend.
[01:28:29] Uh, this will tell that, uh, I want to place my legend outside of the plot.
[01:28:34] The font size of the legend. The title font size, and frame.
[01:28:39] If I want to frame my legend or not. So, in a similar manner, I have plotted the, uh… plot for the… for the data, which is… which has been grouped using k-means.
[01:28:50] And then, we are displaying it. So, here you can see the legend is outside of the plot.
[01:28:56] And it has a frame on, and every, uh… every data point which has been clustered into one cluster has.
[01:29:02] a different color for each class. So, this, uh, right… the left side of… the left side plot.
[01:29:10] is the plot where we have the actual labels.
[01:29:14] Of the data, and the right side is the plot of the clusters that are formed using the k-means.
[01:29:20] So, as you can see here, Bombay and Kali. So, they are.
[01:29:23] They belong to different… these data points belong to different cluster.
[01:29:27] But using k equals to 7. Our k-means is not, like, differentiating these two classes.
[01:29:35] It is not being able to differentiate between these two classes. So here, it has formed one cluster.
[01:29:41] Uh, which is cluster 6. Out of these data points.
[01:29:46] So it was not able to differentiate. well between the Bombay and the Cali class.
[01:29:51] So, in this manner, you can visualize. the plots.
[01:30:01] So here, we have also displayed our variance, so… A principal component 1 has captured 55% of the total dataset variance.
[01:30:11] And, uh, PCA component 2 has captured 26%, and total, like, um… combinedly, they have captured 81, or approximately 82%.
[01:30:22] Uh, of the dataset variance.
[01:30:31] So now, we will, uh, evaluate. RK means using k cluster. The code is similar.
[01:30:38] Uh, we are calling. RK means on, uh, sick… we are… we have defined 6 here for the number of clusters, and we are calling k-means.
[01:30:47] And then we are calculating our… both the measures.
[01:30:51] And then we are displaying it. And in a similar manner, like I describe… I explained above.
[01:30:58] We are applying PCA. To, you know, visualize the… dataset, the white… this is the group data using the actual variables, or, sorry, actual labels, and this is the data when K is 6.
[01:31:13] So here, as. Uh… By Bunya class and the Kali class.
[01:31:22] is not being, uh, the k-means is not differentiating it as two different classes, and also.
[01:31:28] For, uh, this Dermecin. and SEER class. They are also grouped together.
[01:31:34] So, k-means is also not being able to differentiate between.
[01:31:38] Um… Dermussen and Siraclas.
[01:31:42] So, with cluster… when we are defining our number of clusters as 6, so 4, like.
[01:31:49] two, uh, four classes are not being, uh… identified as different classes.
[01:31:58] So, in a similar manner, we will evaluate the k-means on, uh, when we'll define k is equals to 8.
[01:32:06] And…
[01:32:07] And Desmita, I have one question here on the, uh, when the… Game instills, I mean, it is 6.
[01:32:15] Yes.
[01:32:16] So it is, uh, uh, creating one more new cluster there, if you can see the green color one.
[01:32:21] green color, yes.
[01:32:22] So, which, um, which is not there in the, uh, the previous one, I mean, the… other, uh, how we have to interpret that one.
[01:32:30] The true… in the true. Yes.
[01:32:33] Uh, through one. How we have to interpret that.
[01:32:37] So, uh, what my, uh, so what K-Means is doing when we have taken K is equals to 6, so these are the outliers in the class.
[01:32:47] Hmm.
[01:32:48] like, uh, as they are very far away from this, like, in the true, uh, when we have grouped the data set using the true labels, you can see. So, uh, here also, the.
[01:32:57] The data points which are dense, they are very close to the centroid of the.
[01:33:04] of the overall data set, which is there in this cluster. But the data points which are far away, we can.
[01:33:10] we can say that they are showing a little.
[01:33:11] Mm-hmm.
[01:33:13] Like, if we want to, like, put it in this class, we can also put… because maybe, uh, the distance of this data point would be much closer to the.
[01:33:23] centroid of this cluster. So that's why what it is doing is… so these, uh… data points which are.
[01:33:31] acting as an outlier to the pink class, which is.
[01:33:35] Horos class. It has created a cluster for that.
[01:33:36] Huh?
[01:33:41] For those outliers. The k-means has done that.
[01:33:44] Okay. And also, it has reduced the number of clusters compared to the.
[01:33:47] Okay? Number of clusters, because we have defined 6… Yes.
[01:33:52] Yeah. Um, hmm.
[01:34:04] So when we, uh, when we define our number of clusters to be 8.
[01:34:09] Um… we can see… The K-Means has done a good job.
[01:34:15] In, like, differentiating the classes, as here, the… Barbuana and Kali class has been, like.
[01:34:27] is, uh, K-Means has, like, made, or has identified these as two different classes.
[01:34:34] And also, similarly, for Dermacin and Sira. K-Means has identified.
[01:34:39] This has, uh, two different classes, but all… but as we can see, the gray, uh… The cray color here, which is cluster 5, which should be actually being the, uh, this, uh.
[01:34:53] Horo's class, but they are being, like. Uh… it has made a new cluster out of it, which is cluster 5.
[01:35:03] Uh, just to add here, we want our cluster, uh, the algorithm.
[01:35:07] To pick out clusters which are more like the zoo clusters that are available.
[01:35:13] We don't want to unnecessarily artificially create clusters or even merge clusters.
[01:35:20] So I think that is the point that Arthmisha is trying to make.
[01:35:23] That there might be clusters getting lost, or there might be clusters getting merged, or unnecessary getting splitted.
[01:35:30] If the value of K is either too small.
[01:35:32] Then they get mushed, just to accommodate the number of clusters.
[01:35:37] As whatever number is decided. And if K is too large, then it will split across the clusters.
[01:35:44] The two clusters and unnecessary create more number of clusters, so…
[01:35:55] Yes, so, uh, with cluster value, uh, number of cluster value equal to 8.
[01:36:00] Uh, the k-means is doing a better job in differentiating two different classes, but also it has created.
[01:36:08] Like, if, uh, we don't… we didn't have… Um, the true labels, uh, it has created a.
[01:36:15] new class out of these, uh, points. Which is cluster 5.
[01:36:22] Yeah. So now, we will be comparing all the measures that we discussed before for all the values of K, 6, 7, and 8, and we will see.
[01:36:33] Like, how, uh, they have… the cluster… clustering has been performed.
[01:36:38] For these values of K. So, as we can see, the silhouette value should be high.
[01:36:44] So, a higher silhouette value is, uh, indicates the goodness of the cluster, of the clustering algorithm.
[01:36:51] So here, the silhouette value for 6 is high.
[01:36:54] So, uh, we can say that, uh… Within the cluster, or the compactness, or the completeness of the cluster, uh, has been, uh… Like, uh, is good when we chose the value.
[01:37:09] of K as 6. But comparatively, when we chose K value as 8.
[01:37:16] it decreases, like, 7, also, it degraded, and for 8 also, the value is very small compared to 6.
[01:37:27] So, here we can see, like, for Davies and Balden also.
[01:37:29] Um, the higher value… the lower value, uh, is, uh, good.
[01:37:35] So, for 6, the value is lower, so we can say when our k value is 6, the goodness of the clusters.
[01:37:42] is good, but compared to when my k value was 8.
[01:37:49] Um, so it has not performed well, the clustering has… is not as good as when it was 6.
[01:37:55] And similarly for Kelsinki and Harbis, uh, measure also.
[01:38:00] But, uh, as we move ahead, like, if we see the homogeneity.
[01:38:05] of the clusters, so homogeneity tells us the purity of the cluster. If, uh… the clustering that has been performed, and each cluster contains only one class. So, uh… Homogeneity tells us that.
[01:38:19] So for, uh… so, uh, like, the homogeneity, for homogeneity measure.
[01:38:24] My, uh… when my value was 8 for the number of clusters, the clustering is better.
[01:38:31] So, um… The goodness is better for value K.
[01:38:35] But, uh, it is not good for. Uh, the value 6.
[01:38:40] And, uh, similarly, the adjusted mutual info for the adjusted mutual info measure.
[01:38:47] when my value was 8, my clustering is good.
[01:38:52] But, uh, as you can see, when the value is 6.
[01:38:56] It is not as good. So in this manner, we can compare the.
[01:39:02] measures, or the matrices.
[01:39:14] Any questions?
[01:39:31] Um, ma'am, this is the… Last.
[01:39:32] I think you can carry on. People are…
[01:39:34] Yeah. So, I think, uh… We have in this way, actually.
[01:39:39] Uh, taken different values of K, and then compared.
[01:39:44] The goodness of the cluster with different values of K.
[01:39:48] And, uh, then, uh, finally decided the correct value.
[01:39:52] And, uh, with that, we can form the clusters. So, this is what we want to illustrate.
[01:40:00] Uh, in this, uh, particular hard zone. So, please feel free to ask any questions that you have on the code that are.
[01:40:09] has been illustrated, or any other thing.
[01:40:16] I'm, uh, wondered out, uh, you haven't covered this, uh, like, yeah, the different metric that we measure, but…
[01:40:22] In this case, we already had the data, uh, starting, uh, the class already defined, but let's say we are doing it
[01:40:30] completely unsupervised, then what will be, like, one single thing? Will it be homogeneity always, or…?
[01:40:36] Do we always have to look at everything together and then come to a conclusion on which…
[01:40:41] It's always, uh… It's always… okay, it's always better to use multiple.
[01:40:47] indices, and if the class labels are not given.
[01:40:52] Then, in that particular case, it is a good idea.
[01:40:55] To compare the intrinsic… using intrinsic cluster validity indices.
[01:41:01] Such as Silhardt and other indices. Okay, so extrinsic ones are the ones that you use.
[01:41:09] When the class labels are given to you. But in case they are not given to you.
[01:41:13] Then use the intrinsic ones. So, it may be true, as you are saying, that in many real-world data sets.
[01:41:20] Uh, you may not be given any information about the class labels or groups that are existing in the data in all such cases.
[01:41:28] You may actually prefer to use the intrinsic ones that actually look at the.
[01:41:33] Uh, how the data points are distributed in… within the clusters and across the clusters.
[01:41:40] So, such as the Silhard score and the. Uh, Davis folder index and other.
[01:41:47] characteristics, right, and extrinsic ones are, like. entropy, or… Other measures that inertia and all these things.
[01:41:58] So if you don't have any information, use the intrinsic.
[01:42:02] Once. Okay.
[01:42:12] Ma'am, uh, is there any…
[01:42:13] Okay.
[01:42:16] Ma'am, is there any criteria limit for this?
[01:42:18] scores, ma'am, for each one, like,
[01:42:23] Like, we are seeing, like, k equal to 678.
[01:42:26] Like, these three points, like, um…
[01:42:28] So, is that, uh, for example, if it has been below 1…
[01:42:33] So, it has… we can surely, we can tell that it has been perfectly, uh, the cluster has been perfectly separated like that, ma'am.
[01:42:45] No, there is no standard, uh… thumb rule or no standard, uh, you know, convention that can be utilized.
[01:42:54] To find out, uh, you know, to infer very clearly. So, as I told right from the beginning.
[01:43:02] That this whole process of clustering is comparatively. Uh, you know, quite, uh… arbitrary and all that, so… It's very difficult to say that, okay, this value comes out to be correct or something. You will actually have to iterate, see.
[01:43:19] Uh, if you have some external, um, constraints, that also you may need to consider and all that thing.
[01:43:25] So, other than that, there are no thumb rules that you can use that, okay.
[01:43:28] This we can use and find out which is a good value for K or… Something like that. Unfortunately, we don't have.
[01:43:35] Okay, ma'am, thanks for that.
[01:43:42] Yeah, I can see some other hands raised, maybe you can ask.
[01:43:58] So, I mean, similar to the question asked, Loki, which is, uh, the Silute score is about… is the cluster goodness.
[01:44:09] Uh, and as you… Explained… Asmita explained this to… value near to the 1, is that good cluster and near to the 0, it's a cluster.
[01:44:21] So the value which we have derived now is 0.3 is still… Low to, I mean… If we see the plus one range and zero range.
[01:44:32] So is this a dataset, uh… issue, or we can do something… More to improvise the selfie score here.
[01:44:42] Yeah, so usually, uh, though, you know, ideally the value of plus one can be.
[01:44:47] Considered a very high and nice value. However, uh, in real-world data sets, you may not expect it as high as plus 1 or, you know, 0.8 or 0.9.
[01:44:59] Because the data might have certain. Uh, constraints, or my biases, or something like that, so therefore.
[01:45:04] Yeah.
[01:45:07] Uh, relatively high value, so… You just have to look at the relativeness of the values.
[01:45:14] To conclude what is high. So, what you are saying is correct, that 0.3 is actually not very high.
[01:45:21] But, uh, looking at the dataset. We can, like, we are not having higher values in that.
[01:45:29] So, we can expect that to be… Uh, taken as a high value.
[01:45:34] So what you're saying is a valid point, but unfortunately.
[01:45:39] All these, uh, machine learning algorithms, they don't have a.
[01:45:42] You know, very clear-cut and non-fuzzy kind of. Or a crisper solution, 2 things.
[01:45:48] They're kind of fuzzy and like that, because the data is a real-world data.
[01:45:49] Okay.
[01:45:54] Okay.
[01:45:55] No. Nope.
[01:45:58] I'm going to also explain, uh, like, we started by…
[01:45:59] undersura.
[01:46:03] clustering the different kind of coffee beans here.
[01:46:06] But, uh, like a real-world application of this, or what would have been the need to go into this clustering, like, some application…
[01:46:15] example, if you have test for this.
[01:46:19] Uh, you mean to say some applications of clustering?
[01:46:21] No, not just clustering, but let's talk about just this example, uh, where we took the dataset of segregating the coffee bean, right?
[01:46:26] Okay. Mm-hmm.
[01:46:28] So, what could have been the need of…
[01:46:32] Doing this in this case, just to understand the business use case.
[01:46:35] Okay. Okay. Say, for example, you know, you have, uh, let's say that you're having a field where you're going coffee.
[01:46:44] And you want to identify what is the… Uh, variety of coffee beans, uh, that you can obtain.
[01:46:52] Then what you could do is take a sample from the field.
[01:46:55] And take note the attributes that are given there.
[01:47:00] And then do the clustering and see. Uh, in what ratio the different, or, you know.
[01:47:05] Uh, what kind, uh, different varieties are growing and in what kind of ratio they are growing there. So that could be one application.
[01:47:12] There could be a number of applications. Let's say that.
[01:47:17] Different farmers, they are getting their coffee beans. Uh, and giving it to some factory that is processing the beans.
[01:47:24] No, the factory cannot… You know, identify manually, or it's difficult to manually look at so many beans.
[01:47:31] So there could be a machine learning or an AI system that could actually.
[01:47:35] Uh, look at the attributes, measure the different attributes of the different beans.
[01:47:41] Automatically, and then cluster them to. you know, segregate the different kinds of beans that would result in different varieties of coffee.
[01:47:48] Okay, yeah.
[01:47:54] Any other questions, anyone?
[01:47:56] I'm searching here. Uh, so, uh, in this example, right, so we used the elbow method for finding the optimal K value, right?
[01:48:09] So, are there any other… I mean, other way of finding it, or this is the only one?
[01:48:16] So, the other method is that, uh. If you have some background information.
[01:48:21] About what may be the approximate number of, uh.
[01:48:25] Uh, clusters, then you can try that, uh, you know, you can try around that value.
[01:48:30] Approximate value. But other than that, uh… This is the most popular one, and uh… I think this is the one which is popularly used.
[01:48:37] Hmm.
[01:48:40] Uh, if you have domain expertise or domain knowledge, then you can use that value, or.
[01:48:47] You can iterate around that value to find out the most optimal k value.
[01:48:52] There aren't many other methods that are scientifically available.
[01:48:53] Okay.
[01:48:56] Other than this, which is the most common one?
[01:49:04] Any other questions, anyone?
[01:49:17] Okay, if there are no further questions on this, then shall we wrap up for today?
[01:49:32] Okay, in that case, uh… Yes, any question?
[01:49:35] No, I'm just saying, yeah, we can wrap it up then.
[01:49:36] Anyone? Oh, yeah, yeah.
[01:49:39] Okay, Asmita, thank you very much, uh, for, uh… Taking us through this, uh… Uh, hands-on. And thank you all. We can now break. Thank you.
[01:49:51] Have a great evening, thank you, bye-bye.
[01:49:53] Thank you, ma'am. Bye.
[01:49:54] Thank you, ma'am. Bye-bye.