# 11 2026-01-11 Session course: Module 2 — Machine Learning Algorithms module: Module-2-Machine-Learning-Algorithms date: 2026-01-11 type: transcript video_url: https://personal-learn.armco.dev/files/_Recordings/Module-2-Machine-Learning-Algorithms/11_2026-01-11_Session.mp4 --- [00:51:38] And if the ordering of the points changes, the clusters must not change. They should remain the same. [00:51:45] Then the clustering method should be scalable in the sense that. [00:51:47] As the number of points increase. It should not happen that the clustering method fails and it is not able to cluster. It should be able to cater to huge quantities of data. [00:51:59] Then, uh… The clusters that are formed. [00:52:03] Could be of arbitrary shapes also, so our method must be such. [00:52:08] That it is able to. Identify the true clusters without. [00:52:13] Uh, you know, um… Uh, having some bias towards spherical or. [00:52:19] oval or circular clusters only. And so and so forth. There are so many properties or requirements. [00:52:28] That a good clustering method should satisfy. For example, I already told you clusters. [00:52:34] If there are arbitrary shaped. I'll show you shortly what that means. [00:52:40] Then we should be able to identify that. Then our cluster should be able to deal with noise and outliers. [00:52:46] It should be able to deal with dynamically. Uh, coming in data points. [00:52:53] And then it should be able to handle different types of attributes, and it should. [00:52:59] Not happen that only one specific attribute type is only possible to be clustered. [00:53:04] then it should be insensitive to the order in which the points are presented, and so and so forth. The clusters form should be scalable, interpretable. [00:53:13] Uh, because now, currently, we are… we also focus on interpretability in AI. [00:53:20] So, there are a lot of requirements. This is just some important ones that. [00:53:26] Any clustering method should satisfy. So that it is, uh, it becomes part of a good clustering method. [00:53:33] Okay, so I will, uh… Explain to you what each of these points mean as we go along. [00:53:41] Now, whenever we talk about clustering, there are two ways in which the data is actually. [00:53:47] organized. One is called a data matrix. Another is called the dissimilarity matrix. [00:53:53] What is data matrix? Let's say we have n data items, and each has P attributes. [00:54:00] So, we have got 1, 2, 3 till N. [00:54:04] Uh, rows in our given data. And each row has… Uh… 1, 2, 3, till P, number of attributes or features in it. [00:54:16] This is what is meant by N cross P. [00:54:19] And suppose we organize it. In a matrix which is N cross P in nature. [00:54:24] That is, we are organizing the data in such a way that we are having. [00:54:29] Uh, table or a matrix in which there are n number of rows where n is the total number of data points. [00:54:35] And P number of columns. Where P is the number of attributes or features that we are having. [00:54:43] Then we call this kind of organization of the data or the data structure as data maxes. [00:54:49] And it is object by variable, object by attribute structure. [00:54:53] The other is the dissimilarity matrix. What is this similarity matrix? We organize the data elements in a N cross N fashion. [00:55:02] What is N cross N? We have data element 1, 2, 3. [00:55:07] Till N. These form the rows of the matrix. [00:55:11] And then, the columns in the matrix. Are also the data element 1, 2, 3. [00:55:18] Tell N. So we organize them in an N plus N matrix, and what is the entry inside the cells? [00:55:25] This is the distance between the first point and the first point, second point and the second point. [00:55:30] So, the diagonal elements will be zero, because it measures the distance of 1 with respect to 1. [00:55:38] Whereas the second, uh… cell that you can see here refers to the distance of 2 with respect to 1. [00:55:45] And the one here is one with respect to 2. [00:55:49] Like that. So, the rest of the elements are the distances between. [00:55:54] Uh, the various, uh… Uh, points, data points. For example, this one would be 2 with respect to 1, this is 1 with respect to 2, this is. [00:56:04] 3 with respect to 1, this is. Uh… to… 2, 3 like that. [00:56:11] And because the distance of. 2 with respect to 1 will be equal to distance of 1 with respect to 2 by the property of symmetry. [00:56:19] That I've discussed earlier, therefore it will be sufficient to populate. [00:56:23] Anyone half of this matrix, either the lower side or the upper side. We don't need to have both. [00:56:29] Because they are going to contain the same values. [00:56:34] Okay, Deepak, what question do you have? [00:56:36] Oh, ma'am, it was around the slide which you had for requirement for clustering. So, for clustering, I think, uh, we are going to cover, like, what is the right ways to. [00:56:41] Mm-hmm. [00:56:46] Hmm. [00:56:47] But what is the data pipeline looks like? Uh, if I want to, sort of, reach to the point where. [00:56:52] Suppose I have to take the decision whether the clusters are good or not. [00:56:56] So, from the scratch. how should I approach? From the point I have the data. [00:56:57] Mm-hmm. [00:57:02] Yeah, so I'll show you methods. To form the clusters and then to assess the goodness of the cluster, so you'll have to wait for some time. [00:57:12] Because this is what you're going to study here. [00:57:14] So, once you have the data, you're going to form the clusters, and then you're going to assess the clusters. [00:57:20] With the help of some, uh, indices or measures. [00:57:25] That will help to identify whether the clusters formed are good enough or not. [00:57:30] So we will cover that, okay? [00:57:31] Right. So, I think what we are going to cover is the mathematical, uh, ways to figure out whether the clusters are good or not. [00:57:40] But suppose I need to do clustering, so suppose, for an example. [00:57:45] I need to figure out what features or, you know, what columns do I need to consider for the clustering. [00:57:52] Mm-hmm. [00:57:53] Like, likewise, what are the things I need to do before I reach to the point where I'm clustering, and then. [00:57:58] Uh, doing this mathematical calculation. To see the quality of the cluster. [00:58:00] Okay. So, what you need to do… yeah, so what you need to do before starting with the clustering. [00:58:05] preparations. [00:58:11] You need to apply data preprocessing techniques. Which you already have studied. For example, if the data is containing missing values, you need to resolve it. [00:58:19] Suppose the data is having noise. Then you need to, uh, do something to handle the noise. [00:58:26] And so and so forth. So these already you have. [00:58:29] Studied. Let's say we don't want to include outliers for the process of clustering. [00:58:36] Then you need to denoise the data, remove the outliers. [00:58:40] Like that. So, whenever you do any kind of machine learning on the data. [00:58:45] The data must be good quality data. A good quality data is one that doesn't have missing values, that is not very noisy. [00:58:53] And, uh, it is consistent and all that. If required, it needs to be normalized to bring the different. [00:59:02] Uh, attributes to the same range, and so and so forth. So, all that you need to do. [00:59:07] Once you are ready with all that, that is data preprocessing part is done. [00:59:12] Then only you'll need to do clustering. Also, you might need to. [00:59:17] Derive feature vectors out of the given data. And that will be dependent on the data type. So, you will be studying different methods as we go along. [00:59:26] Say, for example, if you have simple numeric data, you don't need to do anything. [00:59:30] But let's say I've got text data. Then we'll need to re-present the text in some way. [00:59:36] So that it is, uh, understandable by the computer. [00:59:39] And once we present the text in that way, then we can do the clustering. [00:59:44] So, all this will be required to be done on the data. [00:59:47] Then you will cluster the data, and after forming the cluster, you will assess the goodness using certain. [00:59:53] matrix and measures, which you will be studying. Okay. [00:59:57] Asha, ma'am. Thank you. [01:00:01] So, as already, uh, discussed earlier, we have classification and clustering. [01:00:08] So, the meaning of the English word classify means to grow. [01:00:12] And cluster also means to grow. But as I already discussed, classification is a supervised learning process. [01:00:19] Because here, we are given the class labels. So, we already know the group information, or we already know what are the. [01:00:26] Uh, groups that are available in the given data. [01:00:29] Whereas, if we are doing clustering, then we have no idea about the group information. [01:00:35] And we arrive at it at the end of the process, so this is the difference between. [01:00:39] Classification and clustering, though. Or grammatically, classification and clustering means the same thing, which is grouping. [01:00:50] No, there are many algorithms or many approaches. To perform clustering. Say, for example, partitioning, hierarchical approach, density-based approach, and so on and so forth. [01:01:02] And each of them will, uh, have a large number of algorithms in them. [01:01:07] So, we'll be, uh, discussing some of these approaches as we go along. [01:01:12] Say partition. So, uh, in partitioning, what we do is that we. [01:01:17] Uh, partition, or, uh… Uh, you know, construct different subsets out of the data. [01:01:25] Or we form partitions out of the data. And we evaluate these partitions. [01:01:32] As the groups in the data in such a way that. [01:01:35] Uh, you know, we want to reduce the inter-cluster or the intra-cluster distances, which are also called the errors. [01:01:43] Uh, within the data. So, partitioning means. [01:01:48] Dividing the given set of points into partitions or subsets. [01:01:52] In such a way that the points within each subset or the. [01:01:56] Uh, partition. are very similar, and there are a lot of methods, many algorithms that fall in this category. Say, for example, k-means is one very, very popular. [01:02:08] partitioning, uh, based algorithm. So, in case of partitioning, let's say we have some n number of data points given to us. [01:02:20] And the whole idea is to group the points. [01:02:24] And the number of groups, let's say we form are, let's say, K. [01:02:28] Well, the number of groups formed will be much, much less than. [01:02:32] The total number of points given to us, or K will be much, much less than. [01:02:36] And in the worst case, the number of clusters will be equal to the number of data points. [01:02:42] That is, there will be one data point per cluster that will be formed. [01:02:46] So here, K refers to the total number of clusters. N is the total number of. [01:02:51] Data points that are given to us. [01:02:56] case the number of clusters that we form. And the special feature of partitioning is that when we form the cluster. [01:03:04] Then each cluster. Must have at least 1 data element in it. [01:03:11] We cannot have clusters that are empty. So, each cluster must have at least one data element in it. [01:03:19] And each data element, or each data point. will belong exactly to one cluster. [01:03:26] Not less, not more. Now, what do you mean by not less, not more? [01:03:32] I said in partitioning. The clusters are formed in such a way. [01:03:36] That each data point belongs exactly to one cluster. [01:03:42] Not less, not more. Not less means… That there cannot be points that can be left unaccommodated, or in other words, whatever the points are. [01:03:52] That we are going to use for our clustering. [01:03:57] Must be part of one cluster. So they cannot be just left, uh, without getting. [01:04:04] Uh, without becoming part of any clustering. So, in partitioning. [01:04:09] We have to make sure that all the points that are given to us form part of some. [01:04:14] Cluster. But, at the same time, when each point must be part of one cluster. [01:04:21] They cannot be part of multiple clusters, so each point. [01:04:24] Must belong to only 1 cluster. And that's why I said. [01:04:29] That partitioning aims to identify clusters. In such a way that each point. [01:04:38] belongs to exactly one cluster, not less, not more. [01:04:40] Not less, because we cannot leave any point without putting it in a cluster. [01:04:46] Each point must be accommodated in a cluster. However… Each point cannot be accommodated in multiple clusters, it can be part of only one cluster. [01:04:59] And, uh, so, alternatively, we keep on. Uh… putting the points, reallocating them. [01:05:06] Until our clusters are stabilized, and this is how we form the clusters. [01:05:13] And then we… I'll be discussing the algorithms in partitioning as we go along. Right now, I've just told you what do we mean by. [01:05:21] Uh, clustering. Now, the thing is that we need to form. [01:05:25] K number of clusters, as I already discussed. So, how do we decide what value to… for K to use? [01:05:33] So let's say we are given these set of points, which are shown in the blue and. [01:05:36] circles and the red squares. And we need to decide, because, uh… We may not have some prior information about the data, so we need to decide how many clusters to form. [01:05:49] Let's say if the total number of clusters is 1. [01:05:52] Then all the points would be placed in one particular cluster, as you can see in this figure. [01:05:58] And all of them are shown in the same color, which is green. [01:06:01] So, irrespective of what should have been the cluster. [01:06:05] uh… type, all the points are accommodated in only one cluster. [01:06:10] And when we are doing a cluster, we consider one objective function, which we try to either maximize or minimize. [01:06:17] Usually, we try to minimize. And it is the error function that we try to minimize. In this case, when k is equal to 1, the error is 873. [01:06:27] Now, let's say we took K is equal to 2. In this case, the objective function value comes out to be 173.1. See, earlier it was. [01:06:36] 873, so it was almost 8 times. So, the error was almost 8 times if the objective function is measuring the error. [01:06:45] And now it has come down to. almost 1 year, and there are 2 clusters that you can see here. [01:06:52] And obviously, these two clusters actually resemble the. Uh, the actual natural clusters that were given to us. [01:07:03] Suppose if K is taken as equal to 3. [01:07:06] Uh, then we can see here. That the value of the objective function has reduced still. [01:07:12] And, uh, in this particular case. You're the cluster… one cluster is correct, whereas the other bigger cluster has been actually broken down into two clusters. [01:07:22] And so and so forth. So what we try to do is… That we try to identify the correct value of the number of clusters to be formed. [01:07:32] By drawing a graph, which is a value of the objective function versus the value of K. [01:07:38] And what we do here is that we keep on plotting it, and then we try to identify. [01:07:43] The knee of the curve, so wherever the knee comes. [01:07:48] That is a point where the value of the objective function is optimally low. [01:07:52] It could be low further also, as you can see here, as we keep on increasing the value of K. [01:07:57] The values reducing may be up to 6, after which it is completely plateauing. [01:08:03] However, the impact of increasing the number of clusters. [01:08:08] Versus the gain in the objective function is not that very much. It is just this much very. [01:08:12] Very less gain is there, however. The number of clusters is increasing. [01:08:18] Uh, to a very high value. Therefore, we only choose the optimally high value, uh, sorry, low value. [01:08:24] Which is probably 2 in this particular case. So this is how we decide the correct value of K. [01:08:31] Uh, when we form the clusters. So, uh, is it okay? Any questions so far? Anyone? [01:08:41] Mam, I just have one question. [01:08:43] So, all the data points should be considered, or in… that also we can kind of take a subset of the data points? [01:08:44] Mm-hmm. [01:08:53] Mm-hmm. [01:08:57] All of them. [01:08:59] Okay. [01:09:00] Uh, when we are doing partitioning. Whatever points we have, we need to consider them. Until, unless you have a specific area. So, if… If your clustering is focused only on a subset. [01:09:04] Then you start with that subset only, but whatever the points in that subset, you will have to accommodate all of them in customs. [01:09:12] Okay, fine, thank you. [01:09:17] So, for example, here is, uh… Uh, an example where we are having. [01:09:24] To, uh, groups. So, let's say we chose K equal to 2. [01:09:28] And we are going to use k-means, which I'm going to discuss shortly. [01:09:32] What does K means? So, when we are going to form two clusters, so initially what we have done is we have taken the set of points. [01:09:41] And randomly assign them to clusters. So, let's say there are two clusters, C1 and C2. [01:09:48] Some of these points, which are shown in the red color, are assigned to cluster C1 randomly. [01:09:53] Then to C2, then to C2, then to C1, then to C2, then to C1, then to like that. [01:10:00] Right? So, randomly. Uh, some points are assigned to C1, the others are assigned to C2 without any. [01:10:07] Got any specific consideration. Once the points have been assigned, what we do is. [01:10:13] That we try to find out the mean of all the points that are assigned to cluster C1. [01:10:18] So, to find out the mean, what you'll have to… for cluster C1, we'll have to consider all the points which are there in cluster C1. [01:10:25] So these are the points. Plus, these are the points, plus this is a point. All of them we use to find out the mean. [01:10:33] And use the meal as, uh, the cluster meal or the cluster centroid, that is a. [01:10:38] Center, uh, point in the cluster, so the mean has come out to be these values, as you can see here. [01:10:45] And then, um… Uh, similarly, you find out the mean of the points. [01:10:53] In the cluster C2, what are those points? These are the points, plus these are the points, plus these, plus these. [01:11:00] If you find out the mean, what you will find out is this. [01:11:02] So, uh… In the initial round, we randomly assigned points to cluster C1 and C2. [01:11:09] And found out the mean of the points that have been assigned. [01:11:12] Two clusters, C1 and C2, even though randomly, and use the mean as the cluster center or the centroid. [01:11:19] Which is what it is called. Now, this is the plot. [01:11:22] Uh, of these points, as you can see here. [01:11:25] Cluster C1, and then C2. Now, so this is C1, and this is C2. [01:11:31] Now, what are we going to do in the second iteration? [01:11:34] Uh, so, for example, this is another set of points. [01:11:39] Uh, what we will do is, in the first iteration, we will. [01:11:42] Try to assign the points. Uh, let's say in this case. [01:11:47] We have chosen K equal to 3. So, we'll initially randomly assign the points to 3 clusters. [01:11:55] And you can see that. Uh, once these points have been assigned to these clusters. [01:12:01] The cluster centers are K1, K2, and K3, respectively. [01:12:05] Okay. Which are the mean of the points? Now, what are we going to do? [01:12:10] You're actually going to now find out the distance between. [01:12:14] Each and every point with respect to each and every cluster center. Let's say if. [01:12:19] This is a point. Once these cluster centers are identified. [01:12:23] By the initial round, in the second iteration, we are going to find out the distance of this point with respect to all these 3 cluster centers. [01:12:31] And see whether this point will still belong to cluster K2. [01:12:35] Or it will move to K1 or K3 based on the… on its distance. [01:12:39] With respect to that point. Just a sec. [01:13:38] Yeah, sorry for that. So, like this, we try to find out the distance of each and every point with respect to. [01:13:45] Uh, each and every cluster center, or the centroid. [01:13:50] So, similarly, we find out the distance of this point. [01:13:55] This point with respect to… All cluster center, this one, like that, we repeat this process for all the points. [01:14:04] So, let's say we consider this point, then we find out the distance of this point with respect to all cluster centers. [01:14:08] So it's becoming a cluttered up, so I'm not showing. [01:14:12] However, we do this exercise for all the points. [01:14:14] And it is possible that in the first round, if the point was associated, let's say, to cluster K1. [01:14:20] It might move to K2 or to K3. So, you can see here, for example. [01:14:27] Uh… When we are assigning the points, the assignment is shown in the three colors, one in the red color, the light green, and the dark green color. [01:14:39] Now, uh, what we'll do is, once again, iterate and find out the distance of each and every point with respect to each and every cluster center and see. [01:14:48] Whether it will continue to belong to these clusters or not. [01:14:52] So, now what you can see here… That these were the old cluster centers which are written as K1, K2, and K3 old. [01:15:00] And once the points have been assigned, you can see that. [01:15:04] These points are assigned to. This particular cluster, say K2, like this. [01:15:09] Then, uh, K3? And these points are assigned to K1, like this. [01:15:16] So, what we'll do, we'll once again, find out the mean of all these points. [01:15:20] Uh, which are… which are assigned to cluster K2, and update K2. [01:15:26] So the value of the cluster center, or the mean might change. [01:15:29] Similarly, here. Again, we'll find out the mean of all the points that are assigned to K1 and K3 also. [01:15:35] And based on that, we will find out. new cluster centers. So these are… these are your initial cluster centers, given old, K2Old, and K3 old. [01:15:44] After the first iteration, where the points. Uh, the distances have been… Calculated again in the second iteration. [01:15:52] And based on the distances, the points have been reallocated to. [01:15:57] Uh, uh, the… Uh, to different cluster centers. [01:16:01] Then what is happening? The new cluster centers are shown as K1. [01:16:07] K2, and K3. So you can see here now that the cluster centers have changed, and accordingly, the assignment of points might also change. [01:16:17] Okay, so I can see some hands raised. Uh, what question do you have? You can ask me in the order in which you raise your hand. [01:16:26] Um, uh, like, could you just, uh, explain, like, how did we come up with K1L, K2 whole, and K3O? [01:16:33] Okay. So, for example. [01:16:37] I'll come to this figure when we found out K1, K2, and K3. [01:16:41] So what we do is that… Let's say the points that you are seeing in the red, dark green, and light green are actually assigned to these 3 clusters. [01:16:50] So this is point P1, let's say P2. P3, like that. So what we do is we find out the average. [01:16:58] Of the attributes, let's say attribute A1. For all the points. [01:17:03] Like this. If there are, say, n number of points assigned to this cluster, then. [01:17:11] These are the, uh… This will be? [01:17:16] The, uh, mean attribute value for Ath attribute. Then we'll have. [01:17:21] Let's say, be a… another attribute called B. So, we'll find out the average value of all that. [01:17:28] So, whatever be the number of attributes, looking at all these points, by considering all these points, we'll find out the mean value. [01:17:36] Of all these points, and that mean value. Will come out to be 0.k1. This is how we find out the cluster centers. [01:17:47] And this is based on the initial random assignment. [01:17:48] K1, K2, and K3. Is it okay? [01:17:50] Yes. Now what we do is, once we find out the. [01:17:51] Okay. [01:17:52] cluster centers, then we try to again find out the distance of each and every point with respect to each of these cluster centers. [01:17:59] And see whether some point moves across some clusters. [01:18:02] cluster. And if a point now becomes closer. To some other cluster center, then it is moved from one cluster center to the other one. [01:18:12] Like that, we go on iterating, and each iteration, we. [01:18:14] Update the cluster center by finding out the new mean. [01:18:19] Is it okay? Okay. Uh… [01:18:26] No, ma'am, I just had one question. Is the K1 represented as gray, just a mismatch in color, or does that have a different significance? Because… [01:18:33] K2 is in red, K3 is in this dark green, but 0.1 are in light green, but this is in gray. [01:18:36] Ah, yeah. No, it should be in green only. [01:18:42] Yeah. Yeah. So, actually, why this is shown in gray is because now I see… shown the new point. [01:18:43] Okay, fine. It's just a misrepresentation, right? Okay, thank you. [01:18:54] In the actual color. So, actually, the old points I'm showing in dull color. [01:19:00] Dull red. [01:19:01] No, no, ma'am, uh, what I mean to say is K1, if you see, right, it is, like, dark gray, but the points of K1 are very light green. [01:19:08] Yeah, so there is a mismatch between these. Yeah. Yeah, yeah. [01:19:10] Ah, yeah. I just wanted to check that, or did that significant… signify something else? I just wanted to clarify. [01:19:15] No, no, no. Yeah, it should have been light green. [01:19:18] Yeah. Okay, any other questions, anyone? [01:19:23] Yeah, one question, uh… So basically, what we are going to do is, we are going to have K for different values, and for each value, we are going to do this assignment, uh. [01:19:38] No, I couldn't understand what different values means what? [01:19:39] Taking the K means, is that the case? [01:19:40] So, uh, suppose I take K is equal to 2. [01:19:44] And then, uh, I do random assignment, do this exercise, and then probably for one point, I'll stop. [01:19:45] Mm-hmm. [01:19:53] I don't, like, what is the criteria to stop, uh… You know, these reassignment, maybe the distance, uh, or something else. [01:20:01] Ah, okay. Yeah, go ahead. [01:20:02] Yeah. And then I take K is equal to 3, then do the assignment, and, uh, you know, go through this process again. So, is that the… A process? [01:20:14] Yeah, so, uh, let me tell you. There are two things that you do. [01:20:18] One, there was this graph. Uh… which shows the value of the objective function. [01:20:26] Versus the value of K. Right? So here you keep on varying the value of K and find out the value of this objective function. [01:20:34] And when you vary the… vary the value of K. [01:20:37] For each value, you do this process that I'm showing here. [01:20:41] Right. Now, what is this process? Let's say if I choose the value of k equal to 3. [01:20:42] Yeah. [01:20:48] Then, if I choose a value K equal to 3, then what I'll have to do, I'll keep on have to. [01:20:53] Uh, reassigning the points. Uh, like this, as you can see here, that… In the initial round, we'll randomly assign the points to some. [01:21:04] 3 random cluster centers. Then, after the random assignment, we'll find out the mean of the points and update the cluster center. [01:21:12] Then again, find out the distance of the point with respect to these. [01:21:15] 3 cluster centers, and see if one point… if some points move across clusters. [01:21:21] And if they do, we move them. Then again, we find out the mean. [01:21:25] And then, after finding out the mean, we once again find out the distance. [01:21:29] And we go on iterating like this till substantially there's no movement of points across the clusters. [01:21:35] When there is no movement, substantial movement of points across the clusters, we. [01:21:40] conclude that more or less, our clusters are stabilized. [01:21:44] So, this whole thing we do. For each value of K. [01:21:49] Okay. [01:21:50] Sure. And in previous slide, uh, we have seen some error. So, what that error is, is that the same that we have read before, like SSE or MSN? [01:22:01] And it's like… [01:22:02] So, usually, this error will be what? See, suppose… Let's say I'll show it to you with a… Neurodiagram… Okay, let's say these are the clusters, okay? [01:22:16] These 3 boundaries are showing the 3 clusters, and these are the cluster centers. [01:22:22] So, as per this diagram, I'm talking at this point. [01:22:26] What it means, the center of the cluster is this one, of this cluster. [01:22:31] Of this cluster. So, if we are having the center of the cluster at some point, then what we want is… As per the process of clustering, we want the points to be. [01:22:44] Compactly placed around the cluster center. So, this is the distance. [01:22:49] Of the points that belong to this cluster versus the cluster center. [01:22:54] Right? These are the distances. Of the points which are there in the same cluster versus the cluster center for that cluster only. [01:23:02] So, if the cluster is compact, then what do we expect? [01:23:06] We expect that the points, the cluster center. If these are the points, then the cluster center should be somewhere in between this only, right? If the. [01:23:18] Cluster is compact. Or, in other words, the distance between the cluster center and each of the points in the cluster should be. [01:23:25] Less. And this is what is the error. The error shows the distance or the difference between the points versus their cluster center corresponding cluster center. [01:23:36] Because if the cluster is compact, then the points will be very near to the cluster center. They shouldn't be very far away. [01:23:45] Okay, so this is the error. The distance between the cluster center. [01:23:50] For the respective cluster and the distance of the points that belong to. [01:23:54] Uh, that particular cluster. So, we do this for each of the clusters, and then sum all of these errors. [01:24:02] To find out the final total error. And we try to minimize the objective function means we try to minimize the cluster. [01:24:09] Uh, the error in each of the clusters. So that the sum total error. [01:24:14] Gets minimized. Or in other words, all our clusters are compactly packed. [01:24:20] Is it okay? [01:24:21] Yeah, so the… we look for objective function number, as well as the error. [01:24:26] Both, right? Okay. I see. [01:24:28] So, objective function is the error only. See, this is the error, let's say error even. [01:24:34] Is the error for this particular cluster? Which measures the distance of each and every point with respect to the cluster center. [01:24:41] E2 is the error which is there for this cluster, which is nothing but the. [01:24:46] Some total of the difference of the cluster center with respect to all the points in that cluster. [01:24:53] Then error E3 is the error for this cluster, which is measuring the distance of each and every point, which is assigned to this cluster. [01:25:01] From the cluster center for this corresponding cluster. Then the objective function will be the sum of E1 plus e2 plus E3. [01:25:09] So, we are trying to find out the… some total of all the errors across all the… Uh, clusters. This is the objective function. [01:25:18] Yeah, I think in previous lecture also, we read something about some of SQUAD errors. [01:25:23] Right? Is that the same? Okay. [01:25:24] Yeah, it is some of square errors only. Here, it… that could be anything. It could be sum of square error, it could be mean absolute error, whatever you are choosing. [01:25:34] But some of squared error is one of the most popular objective functions we use. [01:25:39] Where we, uh, find out the difference, square it, and add it, and do this. [01:25:44] For all the clusters. Okay, so that could be one objective function, very popular objective function. [01:25:52] Okay. [01:25:55] So now, after this second alteration, we have found out the new cluster centers, like this. [01:26:03] This and this. Now, once again, we are going to find out the sine and mind of points with respect to these clusters. [01:26:09] You can see specifically the boundary points, see? This… this is a boundary point to this cluster. [01:26:16] This is a boundary point. These are boundary points to their respective clusters. [01:26:21] Now, after finding out the new cluster center. We will see whether the boundary points are reassigned or not. The rest of the points which are within. [01:26:30] Uh, you know, within… which are closer to the cluster center may not get reassigned. [01:26:36] So, this is the movement of the clusters that I've shown by the arrow. [01:26:41] So, this old K3 gets moved to new K3? [01:26:45] Old K1 gets moved to new K1, so this is new K1, this is new K2, this is new K2. [01:26:53] Now we'll once again look at the assignment of points. You can notice. [01:26:56] That, uh, these points that I've shown in the red color are the boundary points. [01:27:02] Uh, this one is a… was earlier assigned to this cluster, K1. [01:27:07] This point was belonging to K2. And this point is also belonging to K1. [01:27:13] Now, after doing this new calculation of the new cluster centers. [01:27:17] We will see. See, this point was actually belonged earlier to cluster K1. [01:27:24] Now, it is reassigned to cluster K2, because the distance between. [01:27:29] K2, with respect to the distance between K1 is less. [01:27:33] This is more. So, this point is no longer assigned to this. [01:27:38] It is now assigned to this cluster. Similarly, if you look at. [01:27:43] This point and this point. So these… both these points are assigned to… Cluster K3, because the distance of these points are now closer to K3 rather than to K2. [01:27:57] So, like this, the points might keep on getting assigned. [01:27:59] You can see here, these points have… which are shown in. [01:28:03] The bigger, bolder shapes are reassigned. Right, so these 3 points are reassigned to new cluster centers. Accordingly, now new means are found out. [01:28:13] And again, K1, K2, and K3 will move to new locations. [01:28:18] And so these are the new locations for K1. [01:28:21] K2 and K3. Once again, we'll then find out the distances of each and every point assigned to this cluster with respect to. [01:28:29] The new cluster centers and see if points move across cluster, and we go on. [01:28:33] Doing this till there is substantially no change. When there is no change, we. [01:28:38] Take these clusters as final. So this is a kind of a summary of what we have seen so far. [01:28:45] So, in k-means, we decide to have k number of clusters, and we use the means, so the name. [01:28:51] Signifies two things. Case the number of clusters means is the average of all the points belonging to that cluster, which is used as the cluster center. [01:29:01] Right. And what do we do? We have these points. [01:29:06] Let's say these are the points that you are seeing in this color. [01:29:09] The dark and the light shades here. And this rectangle shows the data space that you are having. [01:29:17] And these red points are too random. Uh, clusters that are cluster centers which are assigned. [01:29:24] Initially, and case taken as 2. Now, once the assignment of the cluster centers is done, we try to find out the distance of each and every point. [01:29:33] With respect to each and every cluster center. And assign the points to its nearest cluster, and these are the clusters that are formed. [01:29:41] That you are seeing here. With the cluster centers somewhere here. [01:29:45] And here. Now… Once this assignment is done. [01:29:50] The actual means is now calculated. Earlier, it was just random allocation of cluster centers. [01:29:56] So in the next iteration, we are going to find out the mean of all the points that are assigned to the cluster. [01:30:02] And based on the mean, the cluster center might vary. So, you can see here, this cluster center has moved to this point. [01:30:08] And this cluster center is slightly moved. And we have the new Fluster center. Once again. [01:30:14] We'll find out the distance of each and every point with respect to each and every cluster center. In this case, there are two clusters. [01:30:21] And then reassign points. Which are closer to the other cluster centers, so now the reassignment of points has been done. You can see here. [01:30:29] That this cluster, this point has moved from this cluster. [01:30:32] It's no longer assigned to this, it is moved to this, as you can see here. [01:30:36] Then, once again, we find out the cluster center. [01:30:38] And go on doing this till. No further movement of points, substantial movement of points happens. [01:30:45] And now we say that the clusters are converged. [01:30:48] In real world data, where we have millions of points. [01:30:52] Sometimes it is possible. Uh, that there might be… thousands or lakhs of iterations of this kind. [01:30:59] So, in order to curtail the number of computations, we might put a cap on the total number of iterations. [01:31:05] That we want. So this is a very popular method to stop the process of. [01:31:10] Clustering. K means clustering, that will put a limit. [01:31:14] On the maximum number of alterations that we want to go ahead with. [01:31:17] Say, for example, if the number of iterations is 1000. [01:31:21] Then these points will be iterated 1,000 times, and then it will stop, no matter what is the objective function value. [01:31:28] And this is, uh, this may be the case when there are so many points that it becomes really difficult. [01:31:34] to converge, uh… through the… to take this convergence through the entire process. [01:31:40] So we put a cap on the total number of alterations. So this is how. [01:31:48] One quick question. [01:31:49] Okay, means progresses. Uh… Mm-hmm. [01:31:50] So when we cap, uh, at, let's say, 1000, [01:31:53] Mm-hmm. [01:31:54] And, uh, if there are still the points of a lot of distance in the clustering, I mean, uh, further from cluster, [01:32:02] So, does it impact the result? How does it impact the result? [01:32:03] Mm-hmm. [01:32:07] Yeah, so actually, the idea of putting a cap on the number of iterations is. [01:32:12] That you want to stop at that point, you want to stop the clustering process. [01:32:16] From becoming a kind of an infinite… Kind of infinite process. [01:32:22] Obviously, when you're putting a cap on the number of. [01:32:26] iterations. Definitely, you are not going with the best. [01:32:30] You're going with some optimally best answer. So, to save the compute resources, you have put a cap on the number of iterations. Definitely. [01:32:40] The clusters that are formed up to that point may not be the best clusters. [01:32:45] But whatever those are, you go ahead with that only. [01:32:48] So, that could mean that some of the points are not. [01:32:52] Very close to the cluster in which they are placed, but that's the way it is, that if you want to save on the number of computation. [01:32:58] Then you just stop the process with some number of alterations in mind. [01:33:03] And obviously, when you are putting a cap, it is assumed that you have some idea that. [01:33:08] You know, after some number of alterations, your answer will be. [01:33:11] Uh, within some threshold. Of goodness. [01:33:15] And then you can also, if you want to have a. [01:33:19] Ideal, uh, number of iterations. Uh, then what you could do is that have some different numbers of, uh, of the cap limit. [01:33:30] And see how good the clusters formed are, and then decide on the final cap. [01:33:37] That's how it goes, okay? Okay, any other questions, anyone? [01:33:47] So now, um… Uh, so when we are talking about k-means. [01:33:53] Uh, as I already told you, the initial cluster center, and they are called centroids also, are often chosen randomly. [01:34:01] And, uh, therefore, the clusters that are. Uh, produced in the next iteration could vary substantially. [01:34:07] So, clusters generally could vary from one run to the other. [01:34:12] And the centroid that is used to assess the goodness of the. [01:34:15] points, uh, is nothing but the mean of the points within the cluster. [01:34:20] And the closeness of the points. Which helps us to find out the objective function or assess the error. [01:34:28] Uh, is the distance of the points with respect to the plaster center? No, usually if we are talking about numeric data, Euclidean distance is often a very good measure. [01:34:37] To find out that particular distance, or the closeness. [01:34:41] Now, uh, having said that. Uh, this is the sum of squared error, as already discussed. This is one of the most popular objective functions. [01:34:50] So, what do we do? We find… try to, uh… When we talk about the use of sum of squared errors to find out the goodness of the cluster. [01:34:59] We find out the distance of each and every point. [01:35:02] With respect to the, uh, cluster center. So, assume that the cluster center for the IAT cluster is given by MI. [01:35:11] Okay, so MI is the representative or the centroid. [01:35:15] For the IX cluster and access any other point belonging to the IH cluster. [01:35:22] So, X is any point belonging to cluster CI, so we tried to find out the distance between. [01:35:26] The centroid, that is, uh… Cluster center MI with respect to all points X within the cluster. [01:35:33] And we submit some such… all the distances for all X for the cluster CI. [01:35:39] And then do it for all K clusters. So this is what is the… sum of squared error that we find out the error within the cluster, and then. [01:35:47] Find out the sum of errors across all clusters. [01:35:50] And this gives us the total sum of squared error, or the objective function value. [01:35:55] Which I already explained earlier, for example, error 1 plus error 2. [01:36:00] Plus error 3 will be the total error if we are having 3 clusters, which is. [01:36:05] i is equal to 1 2K, where k is equal to 3. [01:36:08] This is what I'd explained earlier. So, these are just the steps, I'll skip over them, I've already explained to you. [01:36:16] Uh, for those who are interested to know the complexity, uh, computational complexity. [01:36:21] Of k-means, um… So, it is order of, uh, V goes stands for order of. [01:36:27] NKID, so N is the total number of points. [01:36:31] Which we want to cluster, case the number of. [01:36:33] Cluster centers, i is the number of iterations, and D is the number of. [01:36:37] Attributes or dimensions in the data. So, uh, the number of clusters is fixed. [01:36:43] i is the number of alterations, we might fix that also, and the number of attributes are also. [01:36:49] fixed given a particular set of points. I mean, data. [01:36:54] So the complexity of this particular clustering method is. [01:36:58] Actually, proportional, or is of the order of N where the N is the total number of points that we want to cluster. [01:37:06] So, this is the complexity. Now, having discussed K-means. [01:37:11] Uh, do you foresee any issues or any problems? [01:37:15] In cavings. [01:37:21] Anyone, any idea? [01:37:23] Too many attrition. [01:37:27] The number of iterations. [01:37:31] Okay. [01:37:32] As a part of that, bringing the points closer, closer to the K, right? Maybe that might end up, uh, I don't know. [01:37:40] Having wrong set of data clustered. [01:37:44] Wrong set of data clustered, I can't understand what that would mean. [01:37:51] Yeah, wrong clusters. [01:37:52] Uh, you… you want to say wrong clusters? [01:37:53] not scalable. [01:37:54] Okay. Anything else? [01:37:58] not scalable. [01:38:06] Let's come to this, you know, very intense, it's, uh… [01:38:07] Oh… So, uh, I could see some… [01:38:13] Computationally very intensive, and I could see someone, um, right, overfitting could be there. [01:38:19] Okay, overfitting could be an issue. But overfitting could be an issue in. [01:38:24] Overfitting or whatever, uh, term you want to use. [01:38:29] Could be an issue with any, uh, any clustering method. [01:38:32] Uh, I'm talking about the specific issues. That we can foresee. [01:38:40] Uh, with k-means clustering. So, I could see here money, say, a never-ending loo. [01:38:46] Or if the initial centroid goes wrong, then the whole clustering may go wrong. That's what Sushri says. [01:38:53] Okay. Okay. [01:39:00] Okay, let's discuss these answers one by one. So, first of all. [01:39:05] Talking about k-means. So, you see that in K means… Okay, it's nothing but an input that is given to the process. [01:39:16] We need to know the correct number of clusters in advance. [01:39:21] In some datasets, we may not know what is the correct value of K. [01:39:25] And in such cases, we like to actually have to do what? [01:39:29] Do the clustering. This is the objective function. [01:39:36] We have to have to do clustering for different values of K. [01:39:40] Right. We'll have to do this, and then we'll have to find out the knee. [01:39:44] And then user me as the final value, or the final value of K. [01:39:48] And if you randomly give K as the input, it might go wrong. [01:39:52] So, the number of clusters must be known in advance. [01:39:56] Which may not be known, given that the data is new data. [01:39:59] Okay, so this is one particular issue. That is a very important issue. [01:40:05] That could be there. Uh… If you are using chemist, I could see, uh, Lokesh, uh, saying that. [01:40:13] There could be an issue with the scaling. So, wrong clusters would be formed, is that data is not normalized. [01:40:21] So, that is correct, that clusters may go wrong. [01:40:24] If the data is not normalized, so we need to normalize the data. [01:40:29] But again, that may be true for any clustering method, right? We need to. [01:40:32] Uh, normalize the data, or we need to scale. [01:40:36] The data… Uh, if you have to form the clusters correctly. [01:40:42] Yes, that is an issue, but it does hold for many others. [01:40:46] clustering methods also. So, but for this particular clustering method. [01:40:52] Deciding the correct value of K is one particular thing. [01:40:56] Uh, that, uh, needs to be taken care of. [01:41:01] Then, I think Deepak has the comment that if data preprocessing is not done properly, then we could have noisy data or outliers. [01:41:08] Okay, so handling of outliers. Uh, that we need to discuss, and then secondly, Sachin wants to say it may be sensitive to initial centroid values. [01:41:18] So, uh, definitely it is correct that it may be sensitive. [01:41:24] 2 initial values. Because even though, uh, we may iterate, however. [01:41:32] Uh, the differences, uh, between the initial centroid and the final centroid may not be that much. [01:41:39] So, alterations were not substantially change the centroid. It will move the centroid to some extent. [01:41:47] Okay, I think there was some… Point about outliers. [01:41:54] As I said there, as I said right in the beginning. [01:41:58] That, uh, in the process of partitioning and k-means is a partitioning-based method. [01:42:05] We need to accommodate all the points in the cluster. [01:42:08] Now, what would that mean? In, uh, in terms of our lives. Any comments? [01:42:14] Anyone? We have to accommodate all the points. [01:42:21] In the clustering process, and let's say we are having outliers. [01:42:25] Then what happens to famines? [01:42:33] So, obviously, let's say if, uh… And just show it with an example. Let's say these are the data points that we are having. [01:42:46] And, um… Um… [01:42:57] And then… [01:43:05] And then this may be another cluster that we are having. [01:43:15] Let's say this is the data distribution. You are actually having 2 clusters, one shown in the circles. [01:43:20] And one shown in the process. Okay, and there are some noise points, as you can see here, this is one noise point, this is 1 noise point, this is one noise point. [01:43:31] The noise has not been treated somehow. And it is there. Now, what would happen is that because k-means relies on the idea of means, what is means? [01:43:39] It is the average of the values of. All the points. [01:43:44] As a result, if there is noise. Let's say these two points are noisy, or I will say, let's say. [01:43:51] This is not a point, maybe there are two points here. [01:43:54] Then what would happen? Uh, because these two points are very far away, and. [01:43:59] As for any partitioning method, each and every point needs to be accommodated in a cluster. [01:44:05] Therefore, what would happen? The mean might change substantially. Let's say if this is the mean. [01:44:10] It should have been here, but… But it might actually go here. [01:44:17] This might go here, because some of the points are very far away, so the mean might change. [01:44:22] Similarly, if we have. You know, some points here, and this is the main cluster. [01:44:28] Then the mean of this cluster might… it should have been here. [01:44:31] But it might go here. It might get swayed because of these noise points. [01:44:37] Which are foreign. So… The K-means process is not robust to noise, it is not able to handle noise. [01:44:44] And it should be… noise should be handled as part of preprocessing if we are expecting our data to be noisy. [01:44:52] Alright, so this is… Uh, this is what will happen. [01:44:58] And the mean value will definitely get influenced by the outlier. [01:45:03] And it would change. It is not going to be where it should be, it is going to drift away. [01:45:09] So, k-means is actually. sensitive to noise. So, this is one issue. [01:45:14] One issue was that, uh, we need to know the correct value of K, then it is sensitive to noise. [01:45:21] Then definitely, uh, it may be sensitive to the order of. [01:45:24] Order in which the points are shown. There are many other issues also. Let's say. [01:45:31] If we talk about the initial cluster sentence. So, the initial choice of the cluster centers might actually impact the process of clustering. [01:45:40] Uh, substantially. Let's say these are the set of points, and the natural clusters are shown in red, blue, and green colors. [01:45:50] And when we do k-means clustering. Uh, then it is possible. [01:45:56] That's some of the points. Uh, actually, are incorrectly assigned. For example, these points. [01:46:02] Should I be in this cluster? But they are actually assigned to this cluster. [01:46:07] So, uh… and this is the case when the number of clusters is correct. If k is equal to 3. [01:46:13] Let's say if k is equal to 3, it is still possible. [01:46:17] That if we had the cluster center, so initially, probably we had the cluster center somewhere here, here, and here. [01:46:23] Let's say if the cluster centers are not assigned correctly right from the beginning, there's one cluster center here. [01:46:29] Let's say one is here, and one is here. [01:46:32] Then the clustering that would be formed will be like this. [01:46:37] Now, what is wrong with this clustering? The clustering is wrong because it is unnecess. [01:46:41] So really, merging two clusters, and it is splitting one large cluster. That should not have happened. [01:46:47] Actually, the clustering processor should be able to identify the correct cluster, which is what… which is this, this, and this. [01:46:54] It should not merge these, it should not split the larger cluster. [01:46:58] So, it is possible that if the initial. Plus the centers are grossly incorrect, then the whole process will definitely go incorrect. [01:47:08] So then, um… So, to find out the… Uh, correct cluster, global clusters will have to restart with some different random. [01:47:18] Uh, seats for the cluster center, and then we'll have to repeat the process. [01:47:24] Uh… So, any questions, anyone, so far? [01:47:32] Remember to mention earlier that, uh, to identify K, [01:47:35] Right? There's an objective methodology there, right? We have to lower that. [01:47:36] Mm-hmm. [01:47:39] So, if you go with that approach, I identify two or three clusters, then it is not going to give the optimal solution. [01:47:44] Okay, we have to go with the 2 or 3 cluster. [01:47:49] So, what I told you was, this is the objective function that we use, let's say, SSE. [01:47:53] Hmm. [01:47:55] And this is the value of K. Then we try to find out the knee of the curve, let's say this is K is equal to 3. [01:48:01] Okay? [01:48:02] So now, with this, we have arrived at the correct number of clusters, let's see. [01:48:07] Hmm. [01:48:08] But I'm talking about the cluster centers. If the initial choice of the cluster center. [01:48:12] was grossly incorrect. Then, it will never… it is… there is a possibility that it will never converge to the correct cluster centers, even at the end. [01:48:22] Say, for example, here you can see here that if the initial choice of cluster centers was this. [01:48:27] Hmm. [01:48:29] Then the final clusters would also look like this, more or less like this. [01:48:32] So here, what is happening? Two clusters have got merged, and one bigger one has got split. [01:48:39] So this is incorrect, just because the… Initiate cluster center choice even when k is equal to 3. [01:48:44] Was very incorrect. It should have been where? It should have been this. [01:48:48] This and this, somewhat nearby this. But what has happened? The choice has. [01:48:55] It's such that there are two cluster centers randomly chosen here. [01:48:59] And 1 was chosen here. Since they are so incorrect that even after it trading. [01:49:04] There is a possibility it will never conulge. This is what I'm saying. [01:49:08] Okay, got it. [01:49:11] Yeah, and this is also when K is correct. [01:49:14] If K is wrong, then this will further. uh, you know… increase the chances of the whole process going wrong. [01:49:22] So, there's a second step. After identifying K, there's a second step to identify the right centroid. [01:49:23] So, uh… [01:49:31] Actually, there is no standard way to identify the right value of. [01:49:38] You just have to see, unfortunately or fortunately, you doing AI generative AI and agentic AI is more or less iterative. [01:49:47] You do something, you see the result, then change it, then do something like that. [01:49:51] And there is no scientific way to. Find out that what should be the correct random, I mean, seed value that you take initial. [01:49:58] Okay. [01:50:00] So, it can always happen. That with whatever random seed value you start with. [01:50:06] Your final objective function will not be minimized. To the value, it should have been minimized. And therefore, your. [01:50:14] Clusters may not go correctly. However, there are different ways. For example. [01:50:17] Uh, visualization is one simple way, which can help you to. [01:50:22] You know, decide what initial value. Beautiful, awful cluster center to use, right? If you visualize, then you can probably see what the clusters look like, and that can help you to decide. [01:50:34] Somewhat the value, correct value of food. Like that, we try to do. [01:50:39] But even if we select that, uh, it is still prone to error if some new data comes in, right? [01:50:46] Okay, so now we are at this… till this point of time, we are assuming a static data set, which is this. [01:50:53] If you're talking about a dynamic dataset where points keep coming in, then this whole process will not work. [01:50:59] Why it will not work? Because it is relying on a static set of data. [01:51:04] Right, if there are new. points, then the whole dynamics of the clusters might change. [01:51:10] And that is one challenge that I had mentioned earlier. [01:51:13] That the traditional key means is not applicable to dynamic data, because if. [01:51:17] new points keep coming in, then the clusters will never get formed, because these iterations will keep on. [01:51:23] Happening, right? The cluster centers will keep on changing. [01:51:27] Again and again, based on the new ones. So, it's not applicable to a dynamic data set, it's only applicable to static data set. [01:51:36] Okay, so is there any other, uh, clustering method which is, uh, there if the data is dynamic? [01:51:44] So, unfortunately, even though. You know, the, you know, uh… To us as humans, it looks very… Uh, non-trivial, it looks very simple. [01:51:56] And easy to understand, interpret, and… You know, thing. But unfortunately, there is no foolproof algorithm to such a simple process as string. [01:52:06] Uh, and the properties of a good clustering that I mentioned. [01:52:11] Actually, all of these properties are not available in any single algorithm. [01:52:16] So, therefore, if we talk about dynamic data. In their basic forms, no clustering algorithm is able to cater to dynamic data. [01:52:23] However, we can do some things. To make it applicable to some extent. [01:52:29] For example, let's say my data is dynamic. Let me see if I have a… Okay. Let's say my data is dynamic, and I'm measuring it at time t2, T3. [01:52:41] T4, T5, T6. [01:52:46] Like that, it is coming in at different times, okay. [01:52:50] Now, I know that the traditional. Uh, k-means is not catering to dynamic data, so what do I do? I bucket the data. [01:52:58] Let's say I… Consider, let's say. [01:53:04] Poor samples are considered. As one data, you know, as one single static data, or a, uh, we call it snapshot of the data. [01:53:13] Okay, so we consider it as one single snapshot. [01:53:16] And apply ka-mines on this. So, when we apply k-means, all the points, or all the data points. [01:53:23] That are coming in these 4 time samples. Are taken together, and… Let's say some… Let's say if K was equal to 3. [01:53:32] And we got some values, K1 is equal to… Whatever be the dimensionality. [01:53:38] Some, uh, point P1. And K2 was equal to P2. [01:53:43] And K3 was equal to P3. Now, what we do is that. [01:53:48] After, uh, since new data points keep coming in, then what we can do is. [01:53:53] There are so many variants, then what we do, we consider the next snapshot. [01:53:58] And look at all these points. And then see whether P1, P2, P3. [01:54:03] uh, you know, changes, or they remain the same. [01:54:06] So, it might drift to P1-. E2-P3 dash. [01:54:11] And then, again, consider the next bucket. So, we bucket some time stamps together. [01:54:18] And then apply key means to it. Again, bucketing can be done in multiple ways. [01:54:21] Here are the two buckets, or the multiple buckets are disjoint. [01:54:25] We could also have a bucket like this. Then we could have a sliding window, so next bucket is like this, next is like this, like that. [01:54:33] So we have some common points, and we have some new points. [01:54:36] So it totally depends how we want to do, but there are different variants that are available or that can be possible. [01:54:42] Okay? [01:54:44] Yeah, but by this method, we might lose, uh, like, some insights which were there in the previous timelines, right? [01:54:53] So, we use the sliding window, and we are, uh, from… we are in between 6 to 10, right? [01:54:58] So we might not have data of 1 to 5, right? [01:54:59] Mm-hmm. [01:55:02] And, uh, it might lead to some issues, right? Some inconsistencies, probably. [01:55:10] Yeah, so the thing is that… When you are considering the centroids, then you are actually considering the summary of the points, right? These centroids are nothing but the mean of all the points. [01:55:21] That you clustered in the previous timestamp or snapshot. [01:55:25] So, in a way, these clusters… Centers that you are having from the previous snapshot. [01:55:31] Are we presenting the points that you had in that. [01:55:35] previous snapshot. However, you are right that you cannot have the detailed. [01:55:40] Points every time. You cannot have that, and obviously. [01:55:44] To have something, you have to lose something else. So, in this case, you want to bring in dynamism in the data. [01:55:50] Then, obviously, certain data points will become obsolete over point… over… Some timestamps. And the newer timestamps will have. [01:55:59] more, uh, you know, emphasis on this process, but that's the way it goes. [01:56:04] And you will see, eventually, in generative AI and everything, this is how it goes. [01:56:08] That some points go out of focus, and some points become in focus because you are having dynamic data. [01:56:15] Had you not got dynamic data, then all are equally important. [01:56:22] Obviously, let's say… let's say if we consider the weather of a place. [01:56:25] When we look at the last 5 years, we don't look at 50 years. [01:56:30] Right? We don't look at 50 years data, we look at 5 years data to infer what is going to be the weather for today, tomorrow. [01:56:37] Or in other words, the very older historical data is. [01:56:44] Just treat it as a little bit obsolete, but if you are interested to look at long-term patterns. [01:56:48] Then you look at 50 hundred years data, right? [01:56:49] Perfect. [01:56:52] Okay. [01:56:53] Okay, so it depends on how much, uh, in future we want to predict, right? It's up to that thing. [01:57:01] So, there are two things. One is prediction. If we do prediction. [01:57:04] Then we want to consider the recency of the data. The example that I gave, say, 5 years of, uh. [01:57:10] Weather data we need to consider, maybe to predict what is going to be the weather today or tomorrow, like that. [01:57:16] But if I'm interested to know the long-term trends, let's say I want to know. [01:57:21] Let's say these are ears. And over the 100 years. [01:57:25] Probably, if I'm talking about temperature as one parameter. [01:57:29] And I want to see the trend. In the temperature at a particular location over last maybe 100 or 50 years. [01:57:36] Then, I'm actually looking at trending of the data. So, long-term patterns would be data trends. [01:57:42] If you are interested to find, then we need to consider all this data, 100-year data. [01:57:47] But if I have to predict the weather for today or the temperature for today, it's no good looking at 100-year data, right? The dynamics of the place would have changed. [01:57:56] So, we look at last 5 years or 3 years of the data, typically 5 years too. [01:58:01] predict the value. Okay, so this is a forecast, or a long-term. [01:58:07] And one thing is the predict. Prediction of a current value. Two are different. [01:58:16] Okay, any other questions, anyone? [01:58:19] I have one small question around. So, the k-me depend on the initial centered that we take. [01:58:21] Mm-hmm. Yes. [01:58:25] And I think we just learned that it came might fail if the initial centroid is wrong, because it might lead to wrong calculation or. [01:58:32] Yeah, if it is grossly incorrect, yes. If it is slightly incorrect, it's okay, it will converge, but if it is grossly incorrect, it will never converge, that's correct, what we are saying. [01:58:42] Yeah, but centered is basically, mathematically, main of the values of. [01:58:49] Yes. [01:58:50] point that we have. So, the only way it can go wrong is where we have outlier or noisy data, right? Is there any other scenarios, where it might… [01:58:56] Yeah, yeah, definitely. As I told you. In the first iteration, then, when we start the first iteration. [01:59:04] Okay, when we start the first iteration, I just have points, right, like this. [01:59:10] And let's say we decided that we are going to go with, say. [01:59:14] You know, 5 number of clusters, or maybe 2 number of clusters we want to go with. [01:59:21] Fine. Now, we have these points, and we know two cluster centers are. [01:59:24] Required, if I have no other background information about the data. [01:59:29] Actually, I could start with… so, initially, 2 random points in this data space, if k is equal to 2 are assigned as the cluster centers. [01:59:39] So, I might have a cluster center here and another one here. [01:59:43] And this might be very… incorrect, because it is possible that if the data distribution is like this. [01:59:50] Then it is possible that one cluster center should have been here and one another here, so that the points get clustered like this. So my. [01:59:59] allocation of the initial random selection of the. Cluster center was grossly incorrect. [02:00:05] And so it will never go correct if it… these are very far from the real one. [02:00:09] Actual ones. That is what. [02:00:10] It's very tightly coupled with the value of K as well, so… [02:00:14] Yeah, so first of all, if K goes wrong, everything goes wrong, no matter how good you start. [02:00:20] Now, the next thing, if you start with the correct value of K. [02:00:24] Then also, and the cluster centers are. Very much incorrect. [02:00:28] Then also, your process will go wrong, even if case correct, and the initial random. [02:00:33] seeds that you start with is incorrect, then also it goes wrong. [02:00:40] Yeah. Yeah, and there, yeah. [02:00:41] Okay. [02:00:43] Let's say after we identify… so… [02:00:46] In the previous slide, you shows that, okay, clusters are wrongly placed. [02:00:49] based on the visualization, like, uh… [02:00:51] Okay, the suboptimal cluster situation. [02:00:54] But how to look for the mathematically? [02:00:56] Mm-hmm. [02:00:57] So it's like, okay, we have to look for all these three cluster centroid, and we have to identify distance. If distance is coming down, [02:01:02] Which means that we are… [02:01:04] not optimally corrected the cluster, is it correct? [02:01:10] See, that's what I showed to you. So, the objective function is the sum of squared errors, right? [02:01:16] Let's say if you are having 3 cluster error over cluster 1, 2, and 3. [02:01:21] Okay. Now… Even if I chose the suboptimal cluster center, this error might still go down. [02:01:29] Isn't it? And therefore, even after I would have conver- uh, done. [02:01:34] The total number of alterations, let's say I had 1,000 alterations, I did all these thousand. [02:01:39] And the error might still be going down, but still the clusters form will… can go incorrect only, because I started grossly incorrect. [02:01:47] It should have been where? It should have been here. [02:01:50] Here and here. Because the placement of my initial random cluster was so incorrect. [02:01:55] That after, um… Uh, you know, going through 1,000 iterations also, it cannot drift. [02:02:02] So very far. It… I mean, these points cannot drift to 2 points here and 1 point here. That will not happen. [02:02:11] And therefore, even if the error is reducing. Or the objective function is becoming optimally low still. [02:02:18] The final clusters may still be incorrect only. Because that value is going down within this cluster, isn't it? [02:02:27] So, what would that mean? Some of the boundary points might move from this cluster to this cluster, or to this cluster. That is all it will mean. [02:02:36] Kidding, yeah. [02:02:38] And so, our final answer may still be incorrect only, and it sounds really very absurd. [02:02:46] But it is true. That's how it proceeds. And it's not just true for k-means. [02:02:50] It is true for most of the clustering methods. [02:02:52] So you have to be very careful. If you are doing a clustering, you have to make sure. [02:02:57] That whatever you're choosing, your choice and everything goes right. [02:03:01] Otherwise, you will never be having the correct clusters, and you might even never know that these are incorrect. [02:03:08] So, maybe one question. So, do we need to have visualization of the data to sort of eliminate. [02:03:14] Uh, the chances of going wrong. [02:03:18] Yeah, it's always very good to visualize the data. It's always a good thing. [02:03:23] And… [02:03:24] Right, because mathematically, it's very decisive, right? Even though the errors are coming down, you know, we might have around. [02:03:29] Hmm. Yeah, that's true. What you're saying is correct, but what to do? This is how it happens. [02:03:35] And then, uh, in the next slide, which I'll discuss on the next turn, we'll still have other scenarios. [02:03:41] Where, even if you start with the correct cluster centers. [02:03:44] And the error is going down, then also it can be incorrect. I'll tell you those scenarios also. [02:03:50] So… so K-Means has a lot of… You know, flaws. Not just k-means, most of the clustering methods. [02:03:58] Uh, we are discussing k-means, that's obviously one thing. [02:03:59] Got it, yeah. [02:04:04] Ma'am, yeah, yeah, uh, one on the iteration part, [02:04:05] Any other questions? Yeah? [02:04:10] Mm-hmm. [02:04:11] So, uh, you told us that, uh, the, uh, by finding the knee of the curve, right, we could actually understand what is a good value for the K. [02:04:16] Mm-hmm. Mm-hmm. [02:04:19] But in case of iterations, how would we understand, like, this is the number of iterations that should go? [02:04:30] Okay, so let's say you… As I said, there is no foolproof solution to anything. [02:04:34] Many a times you just decide on a number. [02:04:37] But if you really need to decide on some good value. [02:04:41] Of the cap on the number of alterations. So, what you do, first you do this. [02:04:46] With some number of alterations, so you decide the correct value of K. [02:04:51] After deciding the number of correct value of K for that value of K, again, you have. [02:04:57] The objective function, right? And the number of alterations. [02:05:01] The cap on the number of alterations, and then look at the objective function, and then choose. [02:05:06] Where it goes optimally minimal. It might be a knee or it might be something, whatever it is. [02:05:11] So, that's how you do it. But there is no standard way to decide that. [02:05:16] Most of the times, how data scientists do it, they just. [02:05:20] Decide on some number, and that's all. And that number might actually depend on the amount of compute they have, they feel they are available. [02:05:21] And that should be the value, right? [02:05:30] whatever compute is available with them, based on that. [02:05:34] They might decide, okay, they have this much of runtime available, accordingly, so many iterations. [02:05:40] Okay, okay, so it basically depends on our computation power. [02:05:41] And that's it. That's how it was. [02:05:45] Okay. So… [02:05:46] Yeah, yeah, because you're putting the cap. Because of the compute only. [02:05:53] Mm-hmm. Mm-hmm. [02:05:54] So, one last question. So, this is a kind of funny question. So, there is a possibility of going everything wrong, then why are we [02:06:02] Then why are we doing what? [02:06:03] Why are we, like, taking this key means into consideration? Like, I mean, the whole result could go wrong, right? [02:06:12] Yes. randomly available. [02:06:13] And yet, we are still, you know, working or using this key main. So, what are the, you know, strong things which are, you know, [02:06:20] Uh, making us, uh, use this K-Mains clustering. [02:06:26] Yeah, so K means clustering is one of the most popular clustering methods. [02:06:30] Uh, surprisingly, even after having so many flaws. And the most important advantage of this method. [02:06:38] I would say… is this. I will just take you through it. [02:06:44] Just a second… [02:06:53] The complexity. [02:06:54] this. Order of N. It has the lowest complexity, even though. [02:06:57] It looks very complex that every time you have to find out the distance of one point with respect to other centroids, and then. [02:07:04] Hydrate and do go on. But it is order of N, which is the number of points, finally. [02:07:10] And, uh, air. And though it looks surprising, but it is still much, much less compute-intensive as compared to other methods, which I'll discuss. [02:07:20] As we go on. So this is one advantage, and then the. [02:07:25] Uh, other important point is very, very intuitive. Like, you know, it's very easy to think that, okay, we have so many clusters, and we have centroids, and we try to find out. [02:07:35] The cluster centroid and look at the distances. So, it's very, uh, interpretable. [02:07:40] this method. So that is also one very big advantage of kings. [02:07:44] And as you go along, in the next turn, I will discuss… it has so many other flaws also. [02:07:50] And one by T1, start thinking that what's the use using it, just like you said. [02:07:56] But still, it is useful, and believe me. There is no foolproof clustering method that is available. [02:08:02] None of them has all the properties of a good clustering method that I discussed. [02:08:06] Not all, even 2 or 3, having 2 or 3 is also difficult. [02:08:12] So then we have to go with what we have at hand, and that's all. [02:08:16] That's the thing. Better to have something rather than to have nothing. [02:08:22] I would say this is what is. The answer to your question. [02:08:27] Yeah. Any other questions, anyone? Uh, I think there was a comment by someone that if, uh. [02:08:36] If there are noise that the centroid might get drifted, so that is correct that I already discussed. [02:08:45] I think, uh, Krishna Kumar has a question that K-means is. [02:08:50] Uh, means more for recommendations rather than decision. Uh, I would, uh… So, what exactly do you mean by recommendation rather than decision? [02:09:02] Uh, if you want to say recommendation on the clusters? [02:09:03] Um, is it, uh… Yes, ma'am, uh, is it, uh, the practical use case? Is it more related to segmentations? [02:09:11] Recommendations which we are applying. Other than, uh, a decision coming out of it. That was a question. [02:09:13] Mm-hmm. No, no, nothing like that. So, the clusters that will… Um, my gaming's also are used to make decisions. [02:09:25] It's nothing that it is only used for recommendation. It's only that we are discussing k-means, and you are looking at the flaws, then you are feeling that it is not so good. [02:09:33] But, as you go along, you will see that there are other methods, they are having their own limitations. [02:09:40] And because of those limitations, we still may choose k-means, and we might end up thinking that, okay, out of the lot, this is better. [02:09:50] Okay. [02:09:51] Okay. [02:09:56] Okay, any other quick questions before we wrap up for today? [02:10:05] Okay, then, if there are no further questions, then we can wrap up for today, and we'll continue in the next one. That will be the next weekend. [02:10:13] I think today you'll have two classes, one in the morning, one in the evening. That's a constraint only for this particular weekend. [02:10:21] And, uh, the other weekends will still go the way we do. [02:10:22] computationally, uh, sensitive. [02:10:25] Okay, thank you all, have a great day, thank you, bye-bye. [02:10:27] Thanks, Matt. [02:10:28] Thank you. Thank you, ma'am. [02:10:33] Thank you.