# 06 2025-11-23 Similarity Measures course: Module 1 — Foundations of AI & ML module: Module-1-Foundations-AI-ML date: 2025-11-23 type: transcript video_url: https://personal-learn.armco.dev/files/_Recordings/Module-1-Foundations-AI-ML/06_2025-11-23_Similarity_Measures.mp4 --- [00:05:48] aditya shrivastava: Good morning, all. [00:06:03] Neeraj Kumar: Hi, good morning, all. [00:06:07] Gautam Sharma: Hey, good morning, everyone. [00:10:29] Durga Toshniwal: A very good morning to all of you, and welcome to today's session. [00:10:33] Durga Toshniwal: So today we'll be continuing [00:10:36] Durga Toshniwal: with what we had started yesterday, so I'll just take a minute to share my slides, and we'll go on from there. [00:11:03] Durga Toshniwal: Just a second. [00:11:32] Durga Toshniwal: So, yesterday we had started off, [00:11:39] Durga Toshniwal: With similarity measures. [00:11:42] Durga Toshniwal: So, we discussed, what do we mean by… Data. [00:11:49] Durga Toshniwal: So, data is something that has, properties like attributes. [00:11:54] Durga Toshniwal: And it represents… Some notion from the real world. [00:12:00] Durga Toshniwal: Just a second, I'm actually facing some issues. [00:12:06] Durga Toshniwal: So, one thing I want to point out here is that, data, so attributes, in a data set [00:12:14] Durga Toshniwal: Maybe of two kinds. [00:12:18] Durga Toshniwal: They may be numeric or categorical. So, numeric attributes are those which contain numbers as the values in them. [00:12:26] Durga Toshniwal: Say, for example, I say, [00:12:29] Durga Toshniwal: Some temperature. So, temperature is a number. [00:12:33] Durga Toshniwal: And it behaves like a number. [00:12:35] Durga Toshniwal: Whereas certain attributes are categorical in nature, categorical means that they represent some categories. [00:12:43] Durga Toshniwal: So, they can be numbers also, but their meaning is still a category. For example, if we talk about binary numbers 0 and 1, so 0 and 1 are numbers, definitely, but [00:12:54] Durga Toshniwal: Truly, they are not numbers. They represent two categories, something like binary value 0, something like binary value 1. For example, I gave you yesterday that if, let's say, the voltage is less than or equal to 5 volts. [00:13:10] Durga Toshniwal: then the Boolean value representing it could be 0, and if the voltage value is greater than 5 and less than 10, then the values can be 1. So, 0 and 1 are numbers, but they are actually representing a range of values. [00:13:27] Durga Toshniwal: So, such kind of data is categorical in nature. [00:13:30] Durga Toshniwal: Again, other examples could be male, female, so that could be 0 or 1. [00:13:36] Durga Toshniwal: Or, it could be zip codes, or PIN codes. So, PIN code is a number, but it's truly not a number, it's actually representing a set of values. [00:13:47] Durga Toshniwal: In this case, a set of locations. [00:13:51] Durga Toshniwal: So, this is about numeric and categorical data. [00:13:55] Durga Toshniwal: Then, [00:13:57] Durga Toshniwal: The other way we could categorize attributes is discrete and continuous. Discrete, again, are those that have a finite set of values. [00:14:06] Durga Toshniwal: Like, for example, anything which is having a finite set, let's say the value could be 1, 2, 3, or 0, 1, 2, 3, till 9, something like that. [00:14:17] Durga Toshniwal: Whereas, as discussed yesterday, continuous valued attributes are those that have an infinite value set in a given range. Say, for example, weight, height, and all that. [00:14:29] Durga Toshniwal: So, all these can take up any value. Even between 0 and 1, there could be infinite values, like 0.99999, then 0.9111, like that, so many could be there. [00:14:43] Durga Toshniwal: Then we discussed about similarity and dissimilarity. [00:14:47] Durga Toshniwal: Yesterday, and similarities, also said to be the proximity or likeness between different data elements. [00:14:56] Durga Toshniwal: So, yesterday I had not discussed, the properties, [00:15:00] Durga Toshniwal: So, whenever we measure similarity, we need some measure or metric. [00:15:07] Durga Toshniwal: to quantify. [00:15:09] Durga Toshniwal: similarity. [00:15:11] Durga Toshniwal: So, what are the properties that any particular measure should satisfy [00:15:18] Durga Toshniwal: For it to be usable to assess the similarity are as follows. [00:15:22] Durga Toshniwal: Let's say we talk about D as the distance, and P and QR as the two data elements. [00:15:28] Durga Toshniwal: Then, the measure should be such that the distance between P and Q should be equal to the distance between Q and P. And this is called the property of symmetry. [00:15:39] Durga Toshniwal: That is, [00:15:41] Durga Toshniwal: We can discuss it like this, that if we are talking about the images of Alex and Bob, and if Alex looks like Bob. [00:15:51] Durga Toshniwal: Then Bob should also look like Alex. It should not happen that Alex looks like Bob, but Bob has nothing similar to Alex. This is the property of symmetry. [00:16:02] Durga Toshniwal: Then we have the property of a constancy of self-similarity, that is, if we are talking about a distance of A with respect to A, then it will be 0. [00:16:13] Durga Toshniwal: Which means that, if you are talking about Alex and Bob example again, then, Alex should look more like himself, rather than what he looks as compared to Bob. It should not happen that Alex looks more similar to Bob than to himself. [00:16:31] Durga Toshniwal: So, this is the property of, self-similarity, constancy of self-similarity. [00:16:38] Durga Toshniwal: Then we have the property of positivity. [00:16:41] Durga Toshniwal: That is, the distance between any two points can be greater than or equal to zero, but it can never be negative. [00:16:48] Durga Toshniwal: The distance between two points P and Q will be 0, which means that P and Q are exactly similar to each other. [00:16:56] Durga Toshniwal: This is called the property of positivity. So, distance will always be 0 or positive. [00:17:02] Durga Toshniwal: And the distance will be 0 when P is exactly equal to Q. [00:17:07] Durga Toshniwal: So… Which means that there are two data elements which are exactly similar to each other. [00:17:15] Durga Toshniwal: Then the third, then the, next property, that is the fourth property, is said to be triangular inequality. I think we all have read it, maybe in class 8th, 9th, like that. So if we want to assess the similarity or the distance between B and C, which are two points. [00:17:35] Durga Toshniwal: then the distance between B and C will be less than or equal to the distance between B and A [00:17:41] Durga Toshniwal: and CNA. [00:17:43] Durga Toshniwal: So, just imagine that we have 3 points, ABC, [00:17:48] Durga Toshniwal: you can follow my cursor, and they are kind of vertices of a triangle, then I think all of you already know that if we have, any particular side, then the length of any particular side will be less than or equal to the length of the other two sides. [00:18:06] Durga Toshniwal: addition of the other two sides. That is, distance between A plus B and distance of A and C. So, that will always be greater than or equal to distance between B and C. [00:18:21] Durga Toshniwal: So, generally speaking, what it means is that, let's say that if we have Alex and Bob, and Carl. So, let's say that if, Alex looks like Bob. [00:18:32] Durga Toshniwal: and Alex also looks like Carl, then it should not happen that Bob and Carl don't match. If Alex looks like Bob, and Bob looks like Carl. [00:18:44] Durga Toshniwal: or Alex looks like Bob and Alex also looks like Carl, then Bob and Carl should also look like each other. It should not happen that Alex looks like Bob and Alex looks like Carl, but Bob and Carl don't, match each other. [00:18:59] Durga Toshniwal: So, these are the properties, property of symmetry, property of self-similarity, property of positivity, and positi… property of triangular inequality. [00:19:11] Durga Toshniwal: And if a measure satisfies all these properties, then it can be used to assess the similarity. [00:19:19] Durga Toshniwal: Between, data elements P and Q, and it's also said to be a metric. [00:19:26] Durga Toshniwal: So now there are a lot of, [00:19:29] Durga Toshniwal: measures or metrics that are popularly used to assess the similarity between data elements. Suppose we have numeric data. [00:19:39] Durga Toshniwal: And, the distance is to be measured between query and candidate, Q and C. [00:19:46] Durga Toshniwal: And the query and candidate are represented by these two curves out here, one red, one black. [00:19:52] Durga Toshniwal: So, we can use Euclidean distance to find out the similarity between the two. The similarity will be measured by looking at the perpendicular distance, or the distances at each and every sample. [00:20:11] Durga Toshniwal: And how is it represented? It is represented as, under root of… [00:20:15] Durga Toshniwal: The difference between the kth value of data point P and the kth value of data point Q, [00:20:23] Durga Toshniwal: The difference squared, and added, and then rooted. [00:20:27] Durga Toshniwal: So, if there are n number of attributes, then we do this for all N attributes. So, k will take the value from 1 to N, [00:20:36] Durga Toshniwal: And for each attribute, we'll find out the difference in the values of P and Q. [00:20:41] Durga Toshniwal: Square it, and add it together, and then under-root it. [00:20:45] Durga Toshniwal: This is how Euclidean distance is measured. It was proposed by scientist-mathematician Euclid, Many, many years back. [00:20:54] Durga Toshniwal: Say, for example, I'm sure all of you are very familiar with this formula. [00:20:59] Durga Toshniwal: if K is equal to 122, that is, if n is equal to 2, [00:21:06] Durga Toshniwal: Then we have just two attributes, X1 and, X and Y. So the two points are, like, X1, Y1, X2Y2, and the distance between them would be X2 minus X1 squared plus Y2 minus Y1 squared rooted. [00:21:22] Durga Toshniwal: So, these are two points, and this is the way. If there were more number of attributes, the difference between all of them would be squared and added, and then the whole thing would be [00:21:34] Durga Toshniwal: rooted. This is the Euclidean distance. [00:21:40] Durga Toshniwal: So, this is just an example. If we have 4 points, P1, P2, P3, and P4. [00:21:46] Durga Toshniwal: And they are given here in this matrix, like P1 is 0, 2, P2 is 2, 0, like that. So we are talking about points in the two-dimensional plane, and they are represented as X comma Y. [00:21:58] Durga Toshniwal: and we want to find out our distance, then we can do that with the help of the N cross n matrix. What is N? N is the number of points. So N cross N means we have 4 points, represented as 4 rows, and the 4 points also being represented as 4 columns. [00:22:17] Durga Toshniwal: And then we find out the pairwise distance between the points. So the distance of P1 with respect to P1 will be 0. [00:22:24] Durga Toshniwal: By the property of self-similatory. [00:22:27] Durga Toshniwal: So, all the diagonal elements will be 0, because they represent the distance of P2 with respect to P2, P3 with respect to P3, and P4 with respect to P4. [00:22:37] Durga Toshniwal: Also, the values, [00:22:41] Durga Toshniwal: either above the diagonal or under the diagonal would be symmetric, and hence only one part needs to be populated. So, for example, if we look at the distance of P2 with respect to P1, then it will be same as P1 with respect to P2, which is 2.8 [00:22:59] Durga Toshniwal: To it. [00:23:00] Durga Toshniwal: Like that. So we keep on calculating and populating one side of the diagonal matrix, and then use it. [00:23:07] Durga Toshniwal: So, from this set of values, we can conclude that P3, the distance between the pair P3 and P2, is 1.414, and this is the closest pair. [00:23:19] Durga Toshniwal: Now, if we look at this figure, we can conclude also that P2 and P3 are the closest pair of points, and the rest of the points are further away from P3 and P2. [00:23:30] Durga Toshniwal: Accordingly, the distance is also coming out to be lesser. [00:23:34] Durga Toshniwal: So, this kind of a matrix, which is a n cross n matrix of data elements, is said to be a distance matrix. [00:23:42] Durga Toshniwal: Okay, so generally, we have something called the Minskow-Whisky distance. The Minskow-Whisky distance is a generalization of Euclidean distance. So, in Euclidean distance, we had the difference squared and then rooted. [00:23:57] Durga Toshniwal: Now, if we have… if we take the distance between PK and QK, and raise it to the… [00:24:04] Durga Toshniwal: Raise it to the power R. [00:24:07] Durga Toshniwal: And then do this for all the attributes from k is equal to 1 to N, sum this, and then take the rth root, that is, raise it to the power 1 by R. [00:24:18] Durga Toshniwal: Then, the distance is set to be the Minsko Whiskey distance. [00:24:23] Durga Toshniwal: If the value of R is equal to 2, then what would we have here? PK minus QK squared? [00:24:30] Durga Toshniwal: And then summation of K is equal to 1 to N, and all of this summation would be rooted. So if R is equal to 2, then this Minskow-Wisky distance is set to be Euclidean distance. So R could take any value, like R is equal to 1, R is equal to 2, R is equal to 3, and so on and so forth. [00:24:51] Durga Toshniwal: So, accordingly, when R is equal to 1, what we have is said to be the L1 norm. [00:24:58] Durga Toshniwal: When R is equal to 2, then what we have is L2 norm, and it is same as Euclidean distance. [00:25:06] Durga Toshniwal: Similarly, we have L3 norm, Linfinity norm, and all that. [00:25:10] Durga Toshniwal: The L1 norm is also called the Manhattan distance or the Hamming distance. [00:25:16] Durga Toshniwal: And it just measures the… it takes the mod of the difference between P and Q. So if we are talking about n number of attributes, then we take the mod of the difference between k attribute of P and Kth attribute of Q. [00:25:31] Durga Toshniwal: We take the difference and sum it all up. [00:25:34] Durga Toshniwal: And, raise it to the power one by one, where r is equal to 1. So, it is… it is same as the summation of the mod of differences. [00:25:44] Durga Toshniwal: So, this is the L1 norm. [00:25:46] Durga Toshniwal: If we talk about L2 now, then R is equal to 2, so we take the difference, square it, add it for all the attributes, and then root it. [00:25:55] Durga Toshniwal: So we do a square root. So, this is L2 now. [00:25:59] Durga Toshniwal: L1, L2 norms, they are very popularly used to find out the similarity between points. [00:26:06] Durga Toshniwal: So now, once again, we have the same set of points, which is P1, P2, P3, and P4, as you're seeing here. [00:26:13] Durga Toshniwal: And, we again have the distance matrix, that is a 4 cross 4 matrix, where the rows are all 1, 2, 3, 4, P1, P2, P3, P4, columns are also P1, P2, P3, P4. [00:26:25] Durga Toshniwal: The diagonal, as I already discussed, will be all zeros, because it's the distance of P1 with respect to P1, P2 with… like that, with respect to P2, P3 with respect to P3, and P4 with respect to P4. [00:26:39] Durga Toshniwal: And then we find out the difference between P1 and P2 using the L1 norm. So if you look at P1 and P2, so P1 is 0 and 2, and P2 is 2 and 0. So we take the difference and take the mod of it, so it will be, like, 0 minus 2 mod plus 2 minus 0 mod, so it will come out to be 4. [00:27:01] Durga Toshniwal: Then, similarly, we find out the other distances. Suppose we find out the distance between P1 and P3, it will be 0 minus 3 mod plus 2 minus 1 mod, so it will be 3 plus 1. [00:27:13] Durga Toshniwal: And that comes out to be 4. [00:27:15] Durga Toshniwal: So you can see here, P1 and P3. [00:27:18] Durga Toshniwal: have a distance of 4. So, like that, it is L1 norm that is getting calculated. It's also called the Hamming distance, or the Manhattan distance. If we talk about L2, which I had shown you earlier also, again, the diagonal matrix will be 0, and we'll look at the difference between P1 and P2. [00:27:38] Durga Toshniwal: And then, square it. [00:27:41] Durga Toshniwal: for all the attributes, add it, and then do the under root. So here, we'll take, for example, P1 and P2, so it will be 0 minus 2 whole square, plus 2 minus 0 whole square, and then root. So that is 4 plus 4, that is 800 root. So that is 2 under root 2, that comes out to be 2.82 root. [00:28:01] Durga Toshniwal: Like that, we calculate. [00:28:03] Durga Toshniwal: Okay, I'll just stop here. Any questions, anyone, so far on what we have discussed? [00:28:10] laxmi sahu: Ma'am, when we are calculating the distance between P1, P2, P3, and P4, what is the purpose? Like, can we say that the way we are doing yesterday, calculating the checkout coefficient, and also check the similarity and dissimilarity of two documents? [00:28:30] laxmi sahu: Are we considering these points, P1 to P4, as some source of information, and then comparing them? [00:28:38] Durga Toshniwal: Yeah, exactly. Let's say P1, P2, P3, P4 represent the coordinates, let's say you have an object, let's say a square object. [00:28:48] Durga Toshniwal: So, its four vertices are represented, let's say, by P1, P2, P3, and P4, and we want to, [00:28:56] Durga Toshniwal: No, I'll not give this example, sorry. Let's say that, the location of a person is represented by latitude-longitude. [00:29:06] Durga Toshniwal: Which are nothing but X and Y. So, we have 4 people, P1, P2, P3, and P4, 4 persons, and we want to find out which 2 people are standing very close to each other. [00:29:17] Durga Toshniwal: Then we'll find out, to do that, we'll have to find out the distance between each pair of persons. So that will be, like, distance of P1 with respect to P2, with respect to P3 and P4, then P2, P3, P2, P4, and P3, P4. [00:29:32] Durga Toshniwal: Right? And then find out which is the least distance. For example, I did earlier. [00:29:38] Durga Toshniwal: So, in this way, we can find out the distance between [00:29:41] Durga Toshniwal: Persons. FP1, P2, P3, and P4 are persons. [00:29:47] laxmi sahu: Okay, and any industry application of this? [00:29:50] Durga Toshniwal: Yes, yes. There are a lot of applications. Suppose I'm given a set of points, I have a data set, and I want. [00:29:58] laxmi sahu: want to fall. [00:29:58] Durga Toshniwal: Form groups out of it. [00:30:00] Durga Toshniwal: So, I want to group the data. So, how will I group? We group the data based on the similarity, because groups represent similar data elements. So, in that [00:30:11] Durga Toshniwal: We will need to find out the similarity between the data elements, so that similar points can be grouped together. [00:30:18] Durga Toshniwal: This is useful in a lot of ways. For example, suppose I've got some 100 documents. [00:30:24] Durga Toshniwal: And I… it's very difficult every time to look at each and every document. Instead, what I could do is organize them into a set of similar documents. So, for that, I'll need to find out the similarity between all the documents, and then put all the similar ones in one group. [00:30:43] Durga Toshniwal: Now, instead of looking at each and every document, we can just look at the group. [00:30:48] Durga Toshniwal: Right? [00:30:49] laxmi sahu: Jack, yeah. [00:30:51] gunjan bhaiya: Ma'am, when you say, grouping, right, so there are multiple points out there, like, example, distance between multiple points can vary, right? So, example, 2.82 with… between P1 and P2, similar P4 and P3 can also be there. So are we saying different, based on the distance, we can club all these pointers together? [00:31:11] Durga Toshniwal: Yes, so based on the distance, we can find out which points can be grouped together. In this case, there are just 4 points, because I was already illustrating the calculation. [00:31:21] Durga Toshniwal: However, think of a dataset having thousands of points. Now I want to find out which points are similar so that I can group them. Then I'll have to calculate the similarity between the points, and based on those values, we can group them. [00:31:35] gunjan bhaiya: But we also have to look at one dimension of against which point we are trying to do similarity, right? [00:31:41] gunjan bhaiya: just by looking number, like, example 2.82, we cannot say, okay, all 2.82 will come together, because it is from… from P1, what is the. [00:31:50] Durga Toshniwal: Yeah. [00:31:50] gunjan bhaiya: 2.8. [00:31:51] Durga Toshniwal: Yeah, so, yeah, yeah, you are right. So, suppose we do a pairwise distance, then the distance between P1 and P2 is something, let's say, 2.828. Now, if the distance between P1 and P3 is also close to 2.828, then probably P1, P2, and P3 are close together, so that we have to do. [00:32:11] gunjan bhaiya: Okay, so we have to consider from the point 1, and we have to build, like, put the base as a P1, right? Then we have to identify nearby that, correct? [00:32:19] Durga Toshniwal: Yes, that's good. Got it. Yeah. [00:32:23] Durga Toshniwal: Anyone else? I saw a hand raised. Any other question? Anyone? [00:32:31] Rhichik Saha: No, Mom, actually, you answered it just now, so it's fine. [00:32:34] Durga Toshniwal: Oh. [00:32:37] Durga Toshniwal: Okay, so then we talked about similarity between binary vectors yesterday only, so I'm just going to take a recap. [00:32:46] Durga Toshniwal: Offered. So we can represent a certain type of data in form of binary vectors. For example, P and Q are shown here. And then we measure the similarity. [00:32:57] Durga Toshniwal: With the help of simple matching coefficient, or the Jacquard coefficient. [00:33:02] Durga Toshniwal: So, the simple matching coefficient, is generally used whenever the data is, symmetric in nature. That is the occurrence of 0, 0 and 11. [00:33:14] Durga Toshniwal: As a match, both are equally important. [00:33:17] Durga Toshniwal: So in simple matching coefficient, we take, matches of 00 plus matches of 1 1, divided by some total of all kinds of matches, like M01, M10, M00, plus M11. So that is simple matching coefficient. [00:33:32] Durga Toshniwal: And, then we have, and this is an example illustrating simple matching coefficient. [00:33:39] Durga Toshniwal: Then we also have a variant of simple matching coefficient, which is called the Jacquard coefficient, which only considers the matches of 1-1. [00:33:47] Durga Toshniwal: And, in this case, it is M11 divided by M0110 plus 11, and we don't consider the 00 case here. [00:33:56] Durga Toshniwal: Because in certain cases, the absence of values, 0 is to zero. [00:34:01] Durga Toshniwal: Can actually, deceptively measure, can deceptively lead to, [00:34:08] Durga Toshniwal: you know, vectors being considered as similar, and I gave you an example. For example, if we have a shop in which there are some hundred items, and if there are two persons buying some set of items. [00:34:23] Durga Toshniwal: Then, definitely, any normal person may buy some 5 items or 10 items out of 100 items. So, if we have a vector in which, the row is the item… row represents the person's purchase. [00:34:38] Durga Toshniwal: And columns are nothing but all the items in the shop. [00:34:42] Durga Toshniwal: The one that is bought is given a value 1, and that that is not purchased by the person is given a value 0. [00:34:49] Durga Toshniwal: Then, in this binary vector, because there might be 100 columns representing the 100 items that are there on the store, only 5 or 10 may be bought. So, out of those 100 values, just 5 or maybe… [00:35:04] Durga Toshniwal: 5 to 10 of them may have a value 1 standing for the purchase, and all would be 0. Test of them would be 0. So, if we have two such binary vectors, then [00:35:16] Durga Toshniwal: The 0-0 match would be, like, 90 out of 100. [00:35:20] Durga Toshniwal: If 10 are the maximum items that are purchased by 2 individuals, then, in the binary vector representing the sale of the items, out of those 100 elements, 90 would be 0 for both of these people. [00:35:36] Durga Toshniwal: And since there are so many 0 to 0 values, therefore, the vectors would come out to be similar. [00:35:46] Durga Toshniwal: In spite of the fact that the items purchased may be very different. So, even if the items purchased are very different. [00:35:54] Durga Toshniwal: For these two people, but because 90% of the items were not purchased, and that stands for a zero in the first person binary vector, and the second person's binary vector representing their [00:36:08] Durga Toshniwal: And the items that they purchased. So those M00 matches could result into a high amount of similarity between these two people's purchase profile. But actually, if we think about it, we are interested to look at what they purchased. [00:36:25] Durga Toshniwal: And whether whatever was purchased by person P1 and P2 matches. [00:36:30] Durga Toshniwal: But what have we got here? We have got the similarity between P1 and P2's items that were not purchased, because we are… the matches between 0s to zero, that is, not purchased and not purchased for both of these people, are quite high. [00:36:47] Durga Toshniwal: As a result, the similarity between their sale vectors come out to be very high. [00:36:52] Durga Toshniwal: But this similarity does not represent the similarity in the items they bought. Rather, it represents the similarity in the items that they did not buy. [00:37:02] Durga Toshniwal: And therefore, this value is just a deceptive value. [00:37:07] Durga Toshniwal: Oh. [00:37:15] Durga Toshniwal: Sorry, I got muted. [00:37:17] Durga Toshniwal: So, so therefore, we have something called a Jacquard coefficient. [00:37:24] Durga Toshniwal: In the Jacquard coefficient, [00:37:31] Jithu Tagore: What mirror plugin. [00:37:34] Durga Toshniwal: Yeah, sorry. So, then we have something called the Jacquard coefficient. The Jacquard coefficient actually only considers the 1 is to 1 matches, and does not consider 0 to 0 matches at all. [00:37:57] laxmi sahu: Mom, you're on mute. [00:37:59] Durga Toshniwal: Yeah, just a second. [00:38:10] Durga Toshniwal: Yeah, sorry for that. [00:38:12] Durga Toshniwal: So then, so the Jacquard coefficient only looks at the 1 is to 1 matches, and it does not consider 0 is to 0 matches for all such cases where we don't want the 0 to 0 match to sway the similarity. [00:38:29] Durga Toshniwal: Here, in many cases, we are interested to look at only one-to-one match, rather than considering 0-0 match. [00:38:37] Durga Toshniwal: And and the very common example is that of the purchase of items, where we are interested to only look at what was purchased by A and B, or person P1 and P2, and how much similar that is. [00:38:51] Durga Toshniwal: So, in such case, it should be the Jacquard coefficient that should be used. Otherwise, the not purchased versus not purchased will come out to be very high, and that is not, what we wanted to measure. [00:39:05] Durga Toshniwal: So, definitely, based on what we want to measure, we can choose simple matching coefficient versus the Jacquard coefficient. [00:39:16] Durga Toshniwal: So here's an example, of simple matching coefficient versus the Jacquard coefficient, which I had [00:39:23] Durga Toshniwal: Probably shown you yesterday also. [00:39:27] Durga Toshniwal: But anyhow, I'll, just go… we'll go through it once again. So, we are having two binary vectors, P and Q, [00:39:35] Durga Toshniwal: And then, so the first case is that of 0 to 1 matches, which are shown in the red color here. So there are two instances where 0 and 1 occur in P and Q. [00:39:47] Durga Toshniwal: Then, if we look at 1 and 0, then again they are shown in green color, which are 2. Then M0 to 0 matches are 7. 1… [00:39:56] Durga Toshniwal: 1, 2, 3, 4, 5. [00:40:02] Durga Toshniwal: So it should be 5, actually, and not 7 here. And then M11 matches our 1, [00:40:09] Durga Toshniwal: So, it is only shown as 1. So, when we calculate the simple matching coefficient, we substitute the value. So, it will be 1 plus 5 divided by 2 plus 2 plus 1 plus 5. [00:40:20] Durga Toshniwal: So, it comes out to be in the denominator 10, and in the numerator, it will come out to be 6. So, it will be 6 by 10, or 60% will be the simple matching coefficient value. So, you can see that, using the simple matching coefficient, it shows a high amount of match between P and Q. [00:40:38] Durga Toshniwal: If we use a Jacquard coefficient, then we are looking at M11 matches only, and in the denominator also, we are not considering M00 values. So then what we'll have, M11 is just 1, [00:40:50] Durga Toshniwal: then divided by M01 plus 1… M10 plus M11. So, 01 will be 2, plus M10 will be 2, and M11 will be 1. So, it will be 1 by 5, and the match is just 20%. So, if we use simple matching coefficient. [00:41:07] Durga Toshniwal: then the match comes out to be 60%, and if we use jackard, then it comes out to be 20%. Now, let's have a look at these two vectors once again. We look at P and Q, and we visually see how much similar are they. If, we talk about 60% match. [00:41:25] Durga Toshniwal: Then 60% match means that they should be matching to a great extent. [00:41:30] Durga Toshniwal: But what can we see here? We can see here that actually the match is not that very high, because PNQ bought, they have purchased if these represent purchase values. Then, the third item was purchased by PNQ, [00:41:47] Durga Toshniwal: Both together. [00:41:48] Durga Toshniwal: And other than that, there were some items purchased by P, not purchased by Q, or purchased, not purchased by Q, and purchased by P, and rest of them are all not purchased by both. [00:42:00] Durga Toshniwal: So, simple matching coefficient, because it considers 0, 0 to… I mean, 0 to 0 occurrences also, so the similarity is just reflecting the 0 to 0 matches, which are, like, 60%. [00:42:13] Durga Toshniwal: Almost 60%. Whereas, the similarity should not have been 60%, it should have been much lower. This we can find out by visually inspecting PNQ. [00:42:25] Durga Toshniwal: And… [00:42:27] Durga Toshniwal: So, Jacquard coefficient is a better representation if you are talking about buying of items by person P and Q. However, in some other cases, simple matching coefficient may be a better representative. For example. [00:42:41] Durga Toshniwal: If we talk about state, there's a state called P and a state Q, and these are the different districts inside P, and these are the different districts in the queue. [00:42:52] Durga Toshniwal: Let's say we want to find out, and 1 represents the majority of females, or… and 0 represents the majority of males in the different districts. Then, in that particular case. [00:43:05] Durga Toshniwal: The consideration of 0-0 and 1 is to 1 match, both are equally important, because we are interested to find out which are the districts where there are male-to-male or female-to-female, majority. So, in such cases, both the occurrences of 0-0 matches and 1-1 matches are equally important. [00:43:24] Durga Toshniwal: And in that case, simple matching coefficients should be used. [00:43:28] Durga Toshniwal: So now, having talked about simple matching coefficient, the next measure is cosine similarity. I already discussed with you yesterday. [00:43:37] Durga Toshniwal: that we fire. Let's say we are representing two, documents with the help of, binary vectors. [00:43:45] Durga Toshniwal: And then we try to find out the cosine of the angle between these two, which we can find out by doing the dot product between D1 and D2, divided by mod of D1 and mod of… I mean, product of mod of D1 and D2. [00:44:01] Durga Toshniwal: And let's say these are the two documents, D1 and D2. [00:44:05] Durga Toshniwal: Where the columns represent the different words that are occurring inside the document D1 and inside the document D2. If the word occurs 3 times, then the value is 3. If it occurs 2 times, it is a value is 2. If the word 1 occurs 1 time, then in D2, the value is 1, like that. [00:44:24] Durga Toshniwal: And then we can find out the dot product, as you are seeing here, this calculation. [00:44:30] Durga Toshniwal: So what we'll get is the cosine of the angle between D1 and D2, which is coming out to be 0.33. [00:44:37] Durga Toshniwal: And then we can find out the, angle by doing inverse, or we can just keep this value as such also, cosine [00:44:46] Durga Toshniwal: Between D1 lead 2 is 0.33, which means that, the similarity in the documents is 0.33, so it's not very low and it's not very high. [00:44:58] Durga Toshniwal: So this is how we actually find out the similarity. So, cosine similarity is a very popular metric that is used when we talk about documents. So it's very, very popularly used in documents. [00:45:13] Durga Toshniwal: For document similarity. [00:45:16] Durga Toshniwal: Then yesterday, we started out with the confusion matrix. [00:45:20] Durga Toshniwal: And the confusion matrix actually shows, [00:45:24] Durga Toshniwal: Provides us with a mechanism to measure the similarity. [00:45:28] Durga Toshniwal: So, let us look at the confusion matrix. So, in the confusion matrix, the rows represent the actual values, and the columns represent the predicted values. However, in other… in some other literature, you might see these reversed. That is predicted on the [00:45:46] Durga Toshniwal: Row side, and the actual on the column side. [00:45:50] Durga Toshniwal: So, you can do that, you can write it that way also. Accordingly, these values will change. [00:45:57] Durga Toshniwal: the values inside the confusion matrix will change. So, let's assume we are having the actual given values as rows, and the predicted values as the [00:46:06] Durga Toshniwal: Columns. And then we are having two classes, positive and negative. [00:46:11] Durga Toshniwal: Then, if the actual class is positive, and the predicted class is positive, then what we have is true positive. [00:46:18] Durga Toshniwal: If the actual class is negative, and the predicted class is also negative, we call it true negative. [00:46:24] Durga Toshniwal: Then, if the actual class is positive, is negative, and it's predicted as positive, then we call it false positive. [00:46:33] Durga Toshniwal: If the actual class is positive, and it is predicted as negative, then it is a false negative. Why false negative? Because it's not negative, but it is predicted as a negative. False positive means it is predicted as a positive, but it's actually a negative. [00:46:50] Durga Toshniwal: So now, let's take the different, [00:46:53] Durga Toshniwal: performance metrics that can be calculated using the confusion matrix. So, for most of the classification problems, we'll have to find out the confusion matrix, and then use it to look at the performance of the given model. [00:47:09] Durga Toshniwal: So, for example, the first thing that we are studying is precision. [00:47:20] Durga Toshniwal: So, what are the total number of positive predictions? They are represented by this column, right? I hope you can see my cursor. So, these are the positive predictions. Now, positive predictions are made up of the values that are positive and predicted as positive. [00:47:36] Durga Toshniwal: So, those are true positive. [00:47:39] Durga Toshniwal: But there are some values that are actually negative, but have been predicted as positive. So, they are called false positives. So, the total prediction consists of two positive values and the false positives, which are actually negative, but predicted as positive. [00:47:55] Durga Toshniwal: So, the denominator, that is, the total predicted positive, will be true positive plus false positive, and the numerator is nothing but the true positive value. So, accordingly, precision will be TP divided by TP plus FP. [00:48:10] Durga Toshniwal: So, precision talks about how many, true positives are there out of the total positive predictions. [00:48:18] Durga Toshniwal: So, so we can say that with how much precision, is our model working? So, when we say that means that actually predicted positives are how many out of the total positive predictions. [00:48:35] Durga Toshniwal: So, [00:48:38] Durga Toshniwal: So now, when we talk about precision, precision will be high when the denominator and the numerator are equal. When they are equal, then the value will come out to be approximately 1. [00:48:52] Durga Toshniwal: And when can the denominator and numerator be equal? They can be equal if the false positives are approximately zero. [00:49:00] Durga Toshniwal: So, false positives mean, the posit… the values that are negatives, but have been predicted as positive. So, if the negatives that are predicted as positives are very low, then we can have precision to be very high. [00:49:15] Durga Toshniwal: So that the denominator and numerator become approximately equal to 0. [00:49:21] Durga Toshniwal: And, and, so false positives will be, in that case, very, low. [00:49:29] Durga Toshniwal: So, for example, if we talk about, you know, spam detection, this is the example that we were discussing yesterday, that we have email spam detection, and this, actually, the email spam, tries to identify, an email to be a spam or a non-spam. [00:49:48] Durga Toshniwal: So, if we are talking about a false positive. [00:49:53] Durga Toshniwal: Then a false positive is one, which is actually not a spam. It is non-spam, but has been classified as a spam. [00:50:02] Durga Toshniwal: And, definitely, [00:50:06] Durga Toshniwal: So, if we are talking about false positives in this case, then we are talking about non-spam emails that are classified as spam. [00:50:16] Durga Toshniwal: And, if the false positive rate is very high, that is, a lot of non-spams are classified as spam emails, then all such emails will be blocked, and they will be filtered out. [00:50:29] Durga Toshniwal: And as a result, the user who's using this email server may actually not receive many important emails that were legitimate emails, but have been classified as false positives. [00:50:44] Durga Toshniwal: That is, they have been classified as, as spam emails, whereas they were not spam, they were legitimate emails. [00:50:53] Durga Toshniwal: So, in such cases, the false positive rate, if it is high, then it will very severely impact the email server's performance, because a lot of legitimate emails will be filtered out. [00:51:07] Durga Toshniwal: And as a result, the user will get, will not be able to receive a lot of emails which are legitimate, and they may be very critical. [00:51:18] Durga Toshniwal: Right, so, this is how precision works. [00:51:23] Durga Toshniwal: Oh… So, I think I could, see some question from, [00:51:30] Durga Toshniwal: Lokesh, that, how to reduce the false positives to make precision better. [00:51:36] Durga Toshniwal: So, how to reduce a false positive? That will depend on how you are actually designing your model. So, this you will study when you… we'll discuss classification. [00:51:49] Durga Toshniwal: That how, in a classifier, how it is tuned so that the false positives become lesser, and the precision becomes higher. [00:51:58] Durga Toshniwal: So, by adjusting the… [00:52:01] Durga Toshniwal: hyperparameters in the model, you could, fine-tune precision so that it becomes higher, okay? So there are a number of parameters that are used in designing a classification model. [00:52:15] Durga Toshniwal: And all these, metrics, performance metrics, are usually used in classifiers. So when we'll design a classifier model, we will study [00:52:24] Durga Toshniwal: How to tune the model so that the precision is high. When it is high, means false positives are low. [00:52:32] Durga Toshniwal: Okay. [00:52:34] Lokesh R: Yes, Claude. [00:52:35] Durga Toshniwal: Yeah. [00:52:37] Durga Toshniwal: So, next thing is recall. So, initially, in precision, we talked about the predicted positives. [00:52:44] Durga Toshniwal: Now, the next measure talks about the actual positives. So, this is the first row out here. This is the actual positives. So, actual positives means true positives, plus positives that are declared as negative. So, they are what? False negatives. [00:53:00] Durga Toshniwal: So, the actual positives are made up of true positives plus false negatives, because false negatives are actually positives that have been declared as negative. So, the… accordingly, the measure is TP divided by TP plus FN. [00:53:15] Durga Toshniwal: Because, we are talking about two positives divided by the actual positives. [00:53:20] Durga Toshniwal: Now, if we want the recall to be high, then this means that the denominator should be approximately equal to the numerator. When the denominator is approximately equal to the numerator, then in that case, false negatives should be approximately equal to zero. So, the false negatives can be approximately equal to zero. [00:53:41] Durga Toshniwal: When they are approximately equal to zero, then this means that there are very few positives that are declared as negative. [00:53:52] Durga Toshniwal: So then, so, now, again, recall can be very important in certain applications. [00:54:01] Durga Toshniwal: Because, for example, if we are talking about a diseased person. [00:54:08] Durga Toshniwal: So, if we are trying to… we are having a test which finds out whether the person is actually, just a second, excuse me, I'm… [00:54:48] Durga Toshniwal: Yeah, sorry for that. [00:54:50] Durga Toshniwal: So now, suppose we are looking at the possibility of a person having a disease or not having a disease. In such cases, the false negative rate should be very low. [00:55:04] Durga Toshniwal: Why? Because if a person is, positive, and this positive is actually declared as a negative, for example, if we talk about corona, then a person having corona [00:55:16] Durga Toshniwal: If he's declared a negative, then he would not be quarantined, and he would mix with other people, and this will lead to spread of corona in a further fashion. So, if in such cases, where we are trying to [00:55:32] Durga Toshniwal: find out whether the person is having a disease or not having a disease. Recall is very important, because we want [00:55:39] Durga Toshniwal: The false negative rate to be very low, because positive should come out to be positive only. [00:55:50] Durga Toshniwal: So then, we have precision and recall. [00:55:53] Durga Toshniwal: Precision is measuring, the predicted… the total positive divided by predicted pos… predicted positive, and recall measures the total true positive divided by the actual, positives. [00:56:08] Durga Toshniwal: And both are important in different ways, as we explained. [00:56:12] Durga Toshniwal: So, if we are talking even in case of a diseased person, say, having corona. [00:56:19] Durga Toshniwal: And if there is a high false positive rate. [00:56:22] Durga Toshniwal: So, if there's a high false positive rate, means the person is not having corona, and is actually declared as a false positive, means he's not having corona, but is dictated to have corona. Then what would be the consequences? The consequence would be that the person would be quarantined for 14 days. [00:56:42] Durga Toshniwal: In spite of not having corona. So, the after-effect of having false positive rate to be high, when we are talking about a diseased person, or a non-diseased person, in that case, if a healthy person is declared diseased, it could have some, critical, impact on that person. [00:57:02] Durga Toshniwal: If we are talking about certain types of diseases. [00:57:06] Durga Toshniwal: Contrary to this, if the false negative rate is very high, then what would happen in the previous case, a healthy person was declared having corona, and therefore had to go through quarantine unnecessarily. [00:57:22] Durga Toshniwal: Whereas, if we are talking about false negative rate to be very high, then it means that a person who's actually positive is declared as a negative. Or in other words, if a person is having a corona, and he's declared not to have it, then what would happen? [00:57:39] Durga Toshniwal: That person will be allowed to go to stay in his home, and he would mix with his family and other friends and other circle, and a lot of people will get corona. So, in this case, precision, having high precision is important, because we don't want a healthy person to suffer. [00:57:56] Durga Toshniwal: Similarly, having a high recall is also very important because we don't want a diseased person, to be free and affecting other healthy persons, right? So this means that both precision and recall both are important. [00:58:15] Durga Toshniwal: But which one do you think is more important, whether it is precision or is it recall? Which one should be treated as more important, if we talk about some disease? [00:58:28] Durga Toshniwal: Anyone, any answer? [00:58:30] Durga Toshniwal: That whether it's precision that's more important, or recall. [00:58:35] Durga Toshniwal: So, Deepak says false negatives. [00:58:37] azad choubey: Yeah, recall… [00:58:40] Deepak Bobade: In case of Corona. [00:58:42] azad choubey: Yeah, I think it's different. [00:58:43] Deepak Bobade: Don't record yet. [00:58:44] Durga Toshniwal: Recall is more important. [00:58:48] Durga Toshniwal: Okay. [00:58:49] Durga Toshniwal: So this means that… [00:58:51] Anup Pankaj: But ma'am, without precision, how we can recall? First, we have to find out that who is [00:58:56] Anup Pankaj: disease or not, so the… I think the important… first important is precision. [00:59:03] Durga Toshniwal: Okay. [00:59:03] aditya shrivastava: It's all important. [00:59:05] Anup Pankaj: Both are important, but… [00:59:07] Durga Toshniwal: Both are important, recall is important, precision is important, case by case. [00:59:13] Durga Toshniwal: So, we have got 4 answers now, both Precision, recall, Case like this. [00:59:21] Dwarakesh T P: So it could also be something that is dependent on whether it is a contagious disease or not. [00:59:27] Durga Toshniwal: Yeah, let's… we are talking about Corona right now, so you can take a call. [00:59:31] Durga Toshniwal: Which one do you want to see? [00:59:32] Dwarakesh T P: it could be recalled. [00:59:34] Durga Toshniwal: It could be recall. [00:59:36] Dwarakesh T P: Which means that… [00:59:37] Durga Toshniwal: False negative rate should be low, which means that a deceased person should not be declared [00:59:44] Durga Toshniwal: Healthy, but false positives could still be there. [00:59:49] Dwarakesh T P: Yeah, but, like, if it is, in the first case, it is affecting only one person, but if it is the… in the other case, it is affecting a bigger set of people. [01:00:00] Durga Toshniwal: Yeah, but a person who's not having corona, let's say is confined to a quarantine center, which was being done in earlier… in the earlier waves, then that would be very bad, and he would have… he would get definitely corona. [01:00:19] Durga Toshniwal: When he stays with corona-impacted people. [01:00:23] Durga Toshniwal: He's likely to have. [01:00:25] Durga Toshniwal: Well, so we, had all kinds of answers that recall is important, precision is important, both are important, and case by case. So the answer to this question is definitely. [01:00:38] Durga Toshniwal: We want recall to be high, because we don't want false negatives, because we don't want people having corona [01:00:46] Durga Toshniwal: To, you know, to be declared healthy and mixed with others. But at the same time, we want precision also to be high, because we don't want false positives, or in other words, we don't want healthy people to be declared positive and be quarantined unnecessarily. [01:01:04] Durga Toshniwal: Or in other words, we want recall to be high and precision to be high, or we want both to be high. [01:01:11] Durga Toshniwal: Right. Now, the question is whether precision and recall, can both of them be high at the same time? [01:01:21] Durga Toshniwal: Can they be hired the same time? Any thoughts on this? [01:01:27] Sonam Manwal: Nope. [01:01:29] Durga Toshniwal: They cannot be. [01:01:31] vinit shah: It can be if the number of… maybe we have done some multiple tests, and we have come up to where we now find out our model is much better, or something of that sort. [01:01:41] vinit shah: Over a period of time. [01:01:43] Sonam Manwal: No, it… [01:01:44] Durga Toshniwal: decision. [01:01:44] Sonam Manwal: If we will make it a little more lenient, then false positive will increase. If we will make it in strict, then false negative might increase, no? [01:01:54] Sonam Manwal: So both can't be stayed together. [01:02:02] Durga Toshniwal: Okay. [01:02:03] Durga Toshniwal: So I think the… Oh… the… answered. [01:02:11] Durga Toshniwal: Oh. [01:02:13] Durga Toshniwal: So, the answer of the question is. [01:02:21] Durga Toshniwal: So, I think I could see, something on the comment can be depending on training data and the model. No, it doesn't depend on training data, and I mean, it does, but any wave precision recall given a model and a training data, which one would be high or low, and how it would impact? [01:02:38] Durga Toshniwal: So the answer to this question is, ideally, both precision and recoil are important, and both must be high, but [01:02:46] Durga Toshniwal: But the truth is that when one is high, the other will be low. [01:02:50] Durga Toshniwal: Why, they can't be high at the same time? Both cannot be high at the same time, because they have a kind of an inverse relationship. So, for example, if we, go for a high precision. [01:03:04] Durga Toshniwal: High precision means false positives will be very low. [01:03:08] Durga Toshniwal: When false positives, when the model aims at false positives to be very low. [01:03:13] Durga Toshniwal: It means that, the model is very conservative, it doesn't want to announce positives as positives, so that the false positive rate may go very low, right? So the model is very, very careful to declare positive as positive, so that, in an attempt to reduce false positives. [01:03:32] Durga Toshniwal: So the false positive rate, for it to be low. [01:03:36] Durga Toshniwal: The model is very, very conservative on declaring positives as positive. As a result, fall positive is low, and the precision is high. [01:03:47] Durga Toshniwal: Now, if you look at recall, if you want a recall to be high, then the false negatives should be low. If the false negatives are low, means… [01:03:56] Durga Toshniwal: that the positive rate should be high. This means that if we want a recall to be high, then the model is overly lenient. What it does, it keeps on declaring everything to be positive, or mostly everything to be positive. As a result, what happens is that [01:04:19] Durga Toshniwal: As a result, what will happen? The false negatives rate will go low, because, mostly everything will be positive only. So when the fa… [01:04:32] Durga Toshniwal: So, when the false positive rate will be low. [01:04:39] Durga Toshniwal: So, sorry, I got muted. So, when the recall will be high, then what would happen? It will be the false negative rate will be… which will go low. [01:04:51] Durga Toshniwal: And the model will be very lenient, it will keep on declaring everything positive. When precision is high, a false positive rate will be low, and the model is very conservative in declaring positives as positive. [01:05:04] Durga Toshniwal: So, so this means that the model is very conservative and strict in one case, which is precision, and very lenient in the other case, which is recall. [01:05:15] Durga Toshniwal: So, therefore, precision recall cannot be high at the same time, because in one case, it requires the model to be very strict. In the other case, it wants the model to be linear. [01:05:24] Durga Toshniwal: So, the model cannot be strict and lenient at the same time, and therefore, precision and recall cannot be high at the same time, and there needs to be a trade-off. [01:05:34] Durga Toshniwal: Right. So, again, this is just a similar discussion that I already said, that when we have high precision and low recall, one… this is one case, the other is high recall and low precision. [01:05:48] Durga Toshniwal: So, when we have high precision, and we have low recall, then definitely, as I said, that model is very strict. [01:05:55] Durga Toshniwal: And, very few false positives are there. [01:05:59] Durga Toshniwal: That is, those positives, which are, you know, which are actually negative, and recall is low. [01:06:09] Durga Toshniwal: And similarly, if we have high recall and low precision, then in that particular case. [01:06:14] Durga Toshniwal: what would happen is that when we have high recall, then it means that the false negatives may be low. So, this is just what I discussed. And so, there needs to be a balance, or a trade-off. [01:06:30] Durga Toshniwal: Between precision and recall. [01:06:34] Durga Toshniwal: And, so how do we actually obtain this, balance? So, this can be done, with the help of, something called the harmonic mean between true positive and false positives. [01:06:50] Durga Toshniwal: Which is nothing but the F1 score. [01:06:54] Durga Toshniwal: So, what is the F1 score? F1 score is actually the harmonic mean between precision and recall, so it is actually 2 divided by 1 upon precision plus 1 upon recall. [01:07:07] Durga Toshniwal: And when we open the process, it becomes 2 times precision deco divided by precision plus deco. So this is the expression we get, which is nothing but the harmonic mean of PNR. [01:07:18] Durga Toshniwal: So, F1 score is 2 times precision into recall, divided by precision plus recoil. [01:07:25] Durga Toshniwal: Okay, so this is what is F1 score. So F1 score tries to, maintain a balance between precision and default, which is what we want. [01:07:39] Durga Toshniwal: So, now, we talked about, [01:07:43] Durga Toshniwal: Precision and recall in a lot of details, and we also discussed F1's code. [01:07:48] Durga Toshniwal: The next thing of importance is accuracy. [01:07:52] Durga Toshniwal: Okay, Jeetu, what question do you have? [01:07:57] Jithu Tagore: Like, discussing about the model, precision and recall, but in the formula, I can see only the, the parameter that is differentiating the precision and recall is. [01:08:12] Jithu Tagore: false positive and false negative. So, it's not dependent on any other parameters, but the model, predicting false positive and false negative can be, high if the model is very good, right? [01:08:29] Jithu Tagore: So, it can have high values also, right? [01:08:32] Bhaskara Veera Kumar Rayavarapu: Yes. [01:08:34] Durga Toshniwal: You wanted to see? [01:08:38] Jithu Tagore: Am I clear? [01:08:41] Jithu Tagore: Sorry, I guess… [01:08:41] Durga Toshniwal: There's some noise, actually. [01:08:45] Jithu Tagore: Okay, if the model is working fine, then the false positive value will be low. [01:08:52] Bhaskara Veera Kumar Rayavarapu: longer. [01:08:52] Jithu Tagore: This negative value would also be low, right? [01:08:56] Durga Toshniwal: So, just a second, G2, I cannot hear because there's a lot of noise in the background. [01:09:02] Durga Toshniwal: I think it's Fahaskar K. [01:09:04] Anup Pankaj: Oscar, can you mute, please? [01:09:09] Durga Toshniwal: Yeah. [01:09:10] Durga Toshniwal: So, what I could understand, G2, is that if a model is good, what you're saying is that false positive and false negative both can be low at the same time. Is that what you're saying? [01:09:22] Jithu Tagore: Yeah, correct, correct. [01:09:24] Durga Toshniwal: So I explained… [01:09:25] Jithu Tagore: Sri Lanka. [01:09:27] Durga Toshniwal: Sunny. [01:09:28] Durga Toshniwal: Yep. [01:09:29] Durga Toshniwal: So, I explained to you the situation, how if we have false positive, high, I mean, precision high or recall high, then in such cases, what will be the impact? If we want precision high, then we want false positive to be low. [01:09:48] Durga Toshniwal: When we want false positives to be low, then what happens? The model is very, very cautious to declare positives as positive, right? It might declare them as negative, right? But it will not declare them as positive, because it wants the false positive rate to be very low. [01:10:06] Durga Toshniwal: Approximately zero. So, if it wants the false positive rate to be approximately zero, it means that the, [01:10:15] Durga Toshniwal: what could happen is that the positives are not being declared as positives very easily, but negatives can still be there. So, when false positive rate is low, then false negative rate will be there, because the model can declare everything as negative, right? [01:10:32] Durga Toshniwal: Why it will declare as negative? Because it wants to reduce the false positive rate. [01:10:37] Durga Toshniwal: And so, when false positive rate is low, false negative rate will be high. When this is high, then recall will go down. [01:10:45] Durga Toshniwal: So, when position is high, recall is going down. Now, let's talk about the case of recall. When recall is high, it means that the false negatives are, is approximately zero. [01:10:56] Durga Toshniwal: When false negative is zero, what is happening is the model doesn't want to declare, you know, it doesn't want to declare the… what it does is that it doesn't want… sorry, it doesn't want false negatives. [01:11:12] Durga Toshniwal: So, it's overly lenient. It will declare everything to be positive only. [01:11:18] Durga Toshniwal: So, if it declares everything to be positive, then what will happen? False positive rate will go high, because there will be many negatives which are also declared as positive only. So, false negative rate will go down, but false positive rate will go up. [01:11:35] Durga Toshniwal: So, if the false positive rate goes up, then precision will go down. So, when recall is low, precision is going to be high, and when precision is high, recall is going to be low. Is it okay, G2? [01:11:49] Jithu Tagore: Yeah, yeah, I caught it now. [01:11:52] Durga Toshniwal: Yep. [01:11:55] Durga Toshniwal: Now, let's talk about accuracy. [01:11:58] Durga Toshniwal: So, in case of accuracy, what we are having is, again, we look at the… Oh… [01:12:05] Durga Toshniwal: We look at, this model, confusion matrix. [01:12:10] Durga Toshniwal: Accuracy means, the correctness of the model. So, correctness means two positives. [01:12:16] Durga Toshniwal: Plus two negatives. [01:12:17] Durga Toshniwal: So, these are the correctly announced positives and negatives. So, we have this in the numerator, true positive plus true negative. [01:12:25] Durga Toshniwal: divided by some total of all declarations, which is 2 positive plus false negative plus false positive plus 2 negative. So, these are the total number of predictions that we are having in the denominator. [01:12:38] Durga Toshniwal: So, then the accuracy is 2 positive plus 2 negative, divided by some total of all the predictions. [01:12:46] Durga Toshniwal: So now, if we are given a model, we talked about, you know, we talked about precision and recall having kind of inverse relationship. [01:12:56] Durga Toshniwal: Then what can we say about accuracy? [01:13:01] Durga Toshniwal: So… Oh, how do you think accuracy… [01:13:04] Durga Toshniwal: Will be balanced, or, in what cases, accuracy can be high. [01:13:14] Durga Toshniwal: So… [01:13:16] vinit shah: When the positives and negatives that are being found out are more than false positives and false negatives. [01:13:23] Durga Toshniwal: Yeah, so accuracy will be high when the false positive rate and the false negative rate are both approximately zero. [01:13:32] Durga Toshniwal: That is true, all the positives announced are two positives only, and all the negatives announced are two negatives only. In that case, the false positive rate will be approximately 0, and false negative rate will be 0, and accuracy will be 1. [01:13:48] Durga Toshniwal: So, when the correct predictions is approximately equal to total number of predictions, then accuracy is going to be high. [01:13:55] Durga Toshniwal: Obviously, this is a very ideal case, because in any real-world application, the false positive rate and the false negative rate may not be zero, that two at the same time. [01:14:06] Durga Toshniwal: So, achieving an accuracy of 100% or equal to 1 is actually… Oh… very, [01:14:14] Durga Toshniwal: It's more of a theoretical thing rather than a practical one. [01:14:20] Durga Toshniwal: Now, the next thing is that we are having F1 score, and we are having, accuracy. Then, which one to use? Whether we should use F1 score, or whether we should accuracy, and how the two are related. So, the answer to this is that if we talk about accuracy. [01:14:40] Durga Toshniwal: Then the accuracy actually, tries to measure [01:14:49] Durga Toshniwal: Sorry. So, when we talk about accuracy, actually, we are talking about a balanced dataset. So, in a balanced dataset, the class distribution [01:15:00] Durga Toshniwal: Say the positive and the negative class distribution is more balanced. That is, there might be equal number of samples having, class 0 and those having class 1 may be equal. [01:15:12] Durga Toshniwal: So, in such cases where both the classes are equally important, and they occur equally, almost equally frequently, then what we consider is accuracy to be of more importance. [01:15:27] Durga Toshniwal: Then the next thing is, when do we use F1 score? So, we use F1 score, when, we are concerned with false negatives and false positives. [01:15:39] Durga Toshniwal: And, we want to find out a kind of a balance between, or a trade-off between precision and recall, because we won't want both of them to be high, but they cannot be high at the same time. So, in such cases, where [01:15:57] Durga Toshniwal: The false negatives and false positives, both are costly to have. [01:16:02] Durga Toshniwal: then we want to have F1 score to be optimally high. In that case, accuracy may not be that very important. It may be F1 score that may [01:16:11] Durga Toshniwal: more important. So, F1 score may be more important when we are having a skewed class distribution, or when we want to strike a balance between precision and recon, and accuracy may be more important when we are having a uniform class distribution. And [01:16:28] Durga Toshniwal: The false positives and false negatives, [01:16:32] Durga Toshniwal: Both are equally important, and both may be approximately zero. [01:16:38] Durga Toshniwal: So then… Then we have something called sensitivity. What is sensitivity? [01:16:46] Durga Toshniwal: Sensitivity, actually, is similar to recall only. [01:16:50] Durga Toshniwal: And what was recall? Recall means… [01:16:53] Durga Toshniwal: The total positives divided by the actual positives. [01:16:57] Durga Toshniwal: That is TP divided by TP plus FN. [01:17:01] Durga Toshniwal: That is recall. [01:17:03] Durga Toshniwal: Then we have something called sensitivity. The sensitivity also means [01:17:08] Durga Toshniwal: The true positives divided by true positive plus false negative. [01:17:12] Durga Toshniwal: And, so recall and sensitivity are actually having the same definition. [01:17:18] Durga Toshniwal: And, sensitivity also measures the proportion of the actual or the true positives. [01:17:24] Durga Toshniwal: That are correctly identified out of the actual positives that are given. [01:17:31] Durga Toshniwal: Right, so, so, recall and sensitivity are the same things, only names are different. [01:17:39] Durga Toshniwal: In certain applications, we call sensitivity as to be of more importance, so there we do not talk about recall and precision, we talk about sensitivity. [01:17:49] Durga Toshniwal: Which is meaning the same thing as recall. [01:17:53] Durga Toshniwal: Then we have another thing called, [01:17:56] Durga Toshniwal: Specis- specificity. So, specificity, actually. [01:18:02] Durga Toshniwal: measure. So, this is something different. So, what is different about specificity? [01:18:08] Durga Toshniwal: So, specificity actually measures the rate of true negative divided by total negative values. So, true negative divided by true negative plus false positive. [01:18:22] Durga Toshniwal: So, these are the actual negatives which are there. 2 negative plus true negative plus false positive. So, false positive means something which is declared positive, what is actually negative. [01:18:34] Durga Toshniwal: So then we have, specificity that talks about true negative divided by true negative plus false positive. And we have precision, which is talking about true positive divided by true positive plus false positives. [01:18:49] Durga Toshniwal: So, so the difference between position and specificity is in the denominator, where we are talking about true positive in one case and true negative in the other. [01:19:02] Durga Toshniwal: Okay. [01:19:03] Durga Toshniwal: So then, many a times. [01:19:07] Durga Toshniwal: It is the sensitivity and the specificity that are of more importance. [01:19:13] Durga Toshniwal: Or in other words, it's the recall. [01:19:15] Durga Toshniwal: And the specificity that are given more importance. [01:19:19] Durga Toshniwal: So, recall talks about the, positives, the actual positives divided by the… [01:19:27] Durga Toshniwal: the declared positive, divided by the actual positives, whereas specificity talks about the two negatives divided by the actual negatives that are available in the… that are there in the model. So the focus is on the lower row. [01:19:44] Durga Toshniwal: in the confusion matrix, which is false, which is two negatives divided by false positives plus two negatives. That is, two negatives are TN, and negatives which are declared as positives are the FP. So we have TN divided by TN plus FP, which is specificity. [01:20:05] Durga Toshniwal: So, generally speaking, specificity means the two negatives, which are actually identified as negatives only. [01:20:13] Durga Toshniwal: For example, if we are talking about healthy versus deceased person, then the percentage of healthy people who are identified to be healthy. [01:20:22] Durga Toshniwal: And not having illness would be specificity. [01:20:28] Durga Toshniwal: So now, we talked about precision recall, then accuracy, then sensitivity and specificity. [01:20:36] Durga Toshniwal: The sensitivity is also said to be [01:20:39] Durga Toshniwal: The true positive rate. So what is true positive rate? [01:20:43] Durga Toshniwal: True positive rate means true positive. [01:20:47] Durga Toshniwal: divided by true positive plus false negatives. This is the true positive rate. [01:20:52] Durga Toshniwal: And then we have the true negative rate. True negative rate means true negatives divided by false positive plus true negative. [01:21:00] Durga Toshniwal: So, the sensitivity and specificity represents the 2 positive rate and the 2 negative rate. [01:21:09] Durga Toshniwal: And we are… many times, we are interested in both of these, the true positive rate and the true negative rate for a given model. [01:21:16] Durga Toshniwal: We want the positives to be declared as positives, and the negatives to be declared as negatives. Accordingly, our TPR and TNR need to be high. [01:21:29] Durga Toshniwal: So, this is the same thing, just written another time. So, we have false positive rate, which is the false positives divided by [01:21:36] Durga Toshniwal: are two negatives plus false positives. So, what are two negatives? Two negatives are those which are negatives and, you know, identified as negative, and false positives are those which are actually, negative, but defined as positive. [01:21:54] Durga Toshniwal: So it is true negative plus false positives, and true negative plus false positives, these two things actually make up of actual negatives. [01:22:04] Durga Toshniwal: He's too. [01:22:05] Durga Toshniwal: So this is what we have, false positive divided by actual negative. [01:22:09] Durga Toshniwal: Then we have false negative rate, where we have the total false negative values divided by true positive plus false negatives. So, we have false negatives plus [01:22:20] Durga Toshniwal: Those, negative values, so we have, false negatives, that is, positive values that are announced as negatives. [01:22:31] Durga Toshniwal: So, we have, in the denominator, false negatives divided by actual positives. Why? Because we are having true positives, and we are having false negatives, so we are actually looking at actual positive, rho. [01:22:45] Durga Toshniwal: So, this is what we are having. So, FPR is given as false positives divided by the total number of negatives, and FNR is given by false negatives divided by the actual number of positives that are there. [01:23:00] Durga Toshniwal: Is it okay, everyone? Any questions? Anyone on this? FPR, FNR, or anything else? [01:23:09] Nirav Mehta: I'm not here. [01:23:09] Deepak Katara: Yes. [01:23:11] Deepak Katara: On F1 score, we want it to be balanced, high or low, what… sorry, I didn't miss that. [01:23:19] Durga Toshniwal: What, you're saying? F1 score? [01:23:23] Deepak Katara: Should be, like. [01:23:23] Durga Toshniwal: Hi, VIN. [01:23:26] Deepak Katara: Yeah. So… [01:23:26] Durga Toshniwal: Yeah, so if you're talking about unbalanced datasets. [01:23:31] Durga Toshniwal: Where, where one class may be a majority class and the other may be [01:23:38] Durga Toshniwal: minority class. In such, classes. [01:23:43] Durga Toshniwal: it will be the F1 score that will be more important. [01:23:48] Durga Toshniwal: So, if you're having skewed class distribution, it's better to look at F1 score, because the false positive, false negative rates may be very different depending on the majority class and the minority class. [01:24:02] Durga Toshniwal: But if our class distribution is balanced, that is, the total number of samples belonging to the positive and the negative class are almost equal, then we should go for accuracy, because it talks about true positive plus true negative divided by some total of all. So it considers both of these things. [01:24:22] Deepak Katara: Yes. [01:24:22] Durga Toshniwal: So… Yeah, so balanced data sets, you go for accuracy. [01:24:27] Deepak Katara: Asking, how do we interpret the value of F1 score? [01:24:31] Durga Toshniwal: How do you interpret? So, I told you that it is a harmonic mean between precision and deco. So, it is 2 times precision plus… so it's… [01:24:40] Deepak Katara: I think it's here. [01:24:44] Deepak Katara: Is she doing? [01:24:50] Durga Toshniwal: Yeah, here it is. So it's just a harmonic mean. 2 upon 1 by P plus 1 by R, which becomes 2PR divided by P plus R. [01:25:00] Durga Toshniwal: So, it's just a harmonic mean between P and R. [01:25:03] Durga Toshniwal: That's the way we think of F1 schools. [01:25:08] Deepak Bobade: So, ma'am, the overall goal should be to keep the F1 score higher, is that correct? [01:25:14] Durga Toshniwal: Yeah, F1 score can be high. [01:25:17] Durga Toshniwal: The goal can be to keep F1 score high. [01:25:20] Durga Toshniwal: Specifically if you are having an unbalanced dataset. In a balanced dataset, accuracy also can be high. [01:25:28] Deepak Bobade: Okay, that means, F1 score and accuracy both can't be high. [01:25:32] Durga Toshniwal: Thank you. [01:25:32] Deepak Bobade: day. Okay. [01:25:34] Durga Toshniwal: They can be optimally high at the same time, but they may not be exactly high at the same time. [01:25:40] Deepak Bobade: Okay, so inverse relationship, basically. [01:25:43] Durga Toshniwal: Yeah. [01:25:44] Deepak Bobade: Okay. [01:25:46] Nirav Mehta: So, ma'am, I had a similar question. On F1 score. [01:25:51] Nirav Mehta: If it's… let's say if it is high, that means the data is queued. So, what is the next step? Do we do anything, on the data, or we kind of discard that one score and go back to the accuracy score, as you mentioned? [01:26:07] Durga Toshniwal: So, again, it depends on how you want to treat the data. [01:26:11] Durga Toshniwal: In many cases, if you have a skewed class distribution, then you can treat the data so that the skewness is removed. So there are many algorithms where you do upscaling or downscaling. [01:26:24] Durga Toshniwal: So, when we… or you can call it oversampling and undersampling. So, if we do oversampling, then what happens is that the minority class, more number of samples can be added. [01:26:36] Durga Toshniwal: So that both the classes, become equally, you know, equally prominent. Then there is also algorithms to do undersampling, where the majority class is undersampled, so that the samples are both reduced. [01:26:52] Durga Toshniwal: Okay, so… [01:26:54] Nirav Mehta: Thanks, man. [01:26:55] Durga Toshniwal: Yeah. [01:26:57] Durga Toshniwal: anyone else? [01:27:05] Durga Toshniwal: Any other questions, anyone? [01:27:24] Durga Toshniwal: So, we talked about, sensitivity and specificity. [01:27:29] Durga Toshniwal: And about FPR and FNR. [01:27:36] Durga Toshniwal: So, let's talk about the relationship between sensitivity and specificity. [01:27:42] Durga Toshniwal: Which is very important in, many cases. [01:27:46] Durga Toshniwal: So, let's assume that we are having a test, a laboratory test. [01:27:51] Durga Toshniwal: And, and the decision on the test would be positive if the value is 50 or higher. [01:27:59] Durga Toshniwal: For the parameter which is being measured in the test, and if it is lower 250, then the sample is qualified as a negative sample. [01:28:08] Durga Toshniwal: So anything, 50 or higher would be, positive, or too positive. Anything less than 50 would be 2 negative. And accordingly, we have, [01:28:20] Durga Toshniwal: The number of points and the test value, which is… we are seeing, let's say, the range of values could be 0 to 100. [01:28:27] Durga Toshniwal: And accordingly, we see the distribution of two positive and 2 negative. [01:28:32] Durga Toshniwal: Around the decision boundary, which is 50, in this case. [01:28:36] Durga Toshniwal: No. [01:28:37] Durga Toshniwal: Let's say that we move the decision boundary towards left or towards right. So, two positives are shown in this green shaded area, and two negatives are shown in the [01:28:49] Durga Toshniwal: blue shaded area. So, when we move the decision boundary towards right. [01:28:55] Durga Toshniwal: Then what would happen? Here, if you are moving the decision boundary towards right, then, then there would be intersection. [01:29:03] Durga Toshniwal: because the decision boundary is moving towards the right, then what would happen? Some of the two negatives would be qualified as false positives, and some true positives will become false negatives. So, an intersection comes into picture. When we are moving the [01:29:20] Durga Toshniwal: decision boundary, either left or right, and therefore, we have these intersection sections. If we move it towards the right, then what we have are false positives, otherwise we have false negatives. [01:29:33] Durga Toshniwal: So now, let's assume that, we have moved the decision boundary to the value which is less than, the minimum value, in this particular plot. So it's, let's say the minimum value is 20, then we have moved the decision boundary to a point which is less than 20. [01:29:53] Durga Toshniwal: When we did that, now anything, as I said, Oh. [01:29:58] Durga Toshniwal: Here. Anything to the right of the decision boundary is true positive. Anything to the left of the decision boundary is true negative. [01:30:06] Durga Toshniwal: Now, if we look at this particular example, we have moved the decision boundary, up to a value of around 20, and as per the definition, anything to the right of the decision boundary falls in the true positive. [01:30:21] Durga Toshniwal: class or category. So, since the decision boundary is on the extreme left here, so anything to the right of it becomes true positive. In this case, all the values [01:30:34] Durga Toshniwal: R towards the right side, and therefore all values, whether they are true negative, two positive, and all, they're all declared as true positives only. [01:30:43] Durga Toshniwal: Or, in other words, if we talk about this test, which is finding out, the… [01:30:50] Durga Toshniwal: the positively affected people versus healthy people, then true positive means every… everybody who undergoes this test will be announced to be diseased, because he's announced to be positive. [01:31:03] Durga Toshniwal: Because everything is going to be true positive if the decision boundaries moved way up to the left side. So when everybody is declared as being deceased, then the test which is conducted has no meaning. It is actually useless, because no matter whether we [01:31:22] Durga Toshniwal: Oh, sorry. [01:31:25] Durga Toshniwal: So, no matter whether we conduct the test or not, everybody's going to be come out, to be coming out as positive only. So, we don't actually even need to conduct the test, we can simply say whosoever is getting tested will be positive for sure, and so the test is rendered useless. [01:31:43] Durga Toshniwal: Now, let's talk about the other case, where we'll talk about the significance of specificity. [01:31:50] Durga Toshniwal: So now, in this particular case, you are going to move the decision boundary towards the right cent. [01:31:56] Durga Toshniwal: So, as we move it, the value of the true positive will decrease and true negative will increase, which I've already explained to you. [01:32:04] Durga Toshniwal: So, [01:32:06] Durga Toshniwal: So, now, if the decision boundaries move to a point which is 100 or higher value, then anything to its left now becomes true positive. [01:32:15] Durga Toshniwal: So, any, anything to the left of the decision boundary will be true positive only. [01:32:22] Durga Toshniwal: And in such cases, anything that undercores the test [01:32:27] Durga Toshniwal: which is shown by this particular region, will be actually, declared true negative only. Why? Because anything to the right is true positive, and anything towards the left is going to be true negative only. [01:32:42] Durga Toshniwal: So, since all the values are coming towards the left, so anything that undergoes such a test is going to be negative only. Everything is negative. [01:32:54] Durga Toshniwal: So, so in, other words, again, such kind of a test where the decision boundary is so high that everything becomes negative is also going to be useless only. [01:33:06] Durga Toshniwal: So, we don't want a case where we have, all positives or all negatives. We don't want. We want the decision boundary. [01:33:15] Durga Toshniwal: To be correctly designed so that there are some, [01:33:19] Durga Toshniwal: Two positives, some, two negatives and all. [01:33:27] Durga Toshniwal: Oppo. [01:33:29] Durga Toshniwal: So, here's the definition I've just put it for sake of, completeness. So, true positive will be true positive divided by, the true positive plus false negatives, and then specificity will be 2 negative divided by true negative plus false positive. [01:33:45] Durga Toshniwal: Right, so this means that if you look at this particular portion, we have true positive divided by a true positive plus false negative. So these are the two positive values. [01:33:57] Durga Toshniwal: Then, [01:33:59] Durga Toshniwal: And accordingly, we have the two negative values, and so on and so forth, right? So, it is very easy to note that when sensitivity goes high, so if this will go high, then in that case, false negatives will go low. When false negatives go low. [01:34:18] Durga Toshniwal: Then, in that particular case, what will happen? If false negatives go low, means everything is becoming positive. [01:34:25] Durga Toshniwal: If specificity is high, then false positives are low, means everything is becoming negative. So, when specificity is high, sensitivity will be low. When sensitivity is high, then specificity will be low. [01:34:40] Durga Toshniwal: And we don't want both of them to be low, because we want both of them to be optimally high at the same time. [01:34:47] Durga Toshniwal: Whereas, there may be situations, okay, so anyhow, so depending on, what kind of values we want, we may set our threshold, based on our requirements, whether we want the sensitivity to be high, or specificity to be high, and so on and so forth. [01:35:08] Durga Toshniwal: The next thing that we are having is ROC curve. What is ROC curve? [01:35:13] Durga Toshniwal: ROC means Receiver Operating Characteristic. [01:35:17] Durga Toshniwal: And, this is a very strange full form. However, [01:35:22] Durga Toshniwal: It, derives its origin in 1941. [01:35:27] Durga Toshniwal: When it was… this term was coined. [01:35:30] Durga Toshniwal: And originally, it was designed by the operators of the military, radar receivers. [01:35:37] Durga Toshniwal: And, therefore, because they were receiving the characteristic value, or the value that is coming, so they called it receiver operating characteristic. [01:35:48] Durga Toshniwal: Right, so now, again, when we talk about ROC, [01:35:55] Durga Toshniwal: We are talking about plotting the value of, true positive rate. [01:35:59] Durga Toshniwal: As a function of the false positive result, right? [01:36:03] Durga Toshniwal: So, I'll show you the graph in a few seconds from now. [01:36:08] Durga Toshniwal: So here we have false positive rate, which is false positive divided by 2 negative plus false positive. We have sensitivity and recall, defined as, [01:36:19] Durga Toshniwal: True positive divided by true positive plus false negatives. [01:36:25] Durga Toshniwal: So then we have this ROC curve. [01:36:28] Durga Toshniwal: And the ROC curve is said to be… sorry. [01:36:33] Durga Toshniwal: So, if you're talking about a random classifier that doesn't, that doesn't care about positive and negative classes. [01:36:41] Durga Toshniwal: then it's going to have a decision boundary, which is similar to this dashed central line that you are seeing here. Anything that is positive, or anything that is negative, and is classified as negative, it doesn't bother to do much. [01:36:54] Durga Toshniwal: Such a kind of classifier is a random classifier, and [01:36:59] Durga Toshniwal: And it's not a very good classifier, because the true positive rate and two negative rate are almost equal. We don't want this. [01:37:08] Durga Toshniwal: Right? So we want a situation where, true positive rate should be high. [01:37:15] Durga Toshniwal: But this is not sufficient because the true negative, the… sorry, the true positive rate is high, but the false positive rate will be low. So this graph shows the true positive rate versus the false positive rate. [01:37:32] Durga Toshniwal: So, when the false positive rate is zero, then the true positive rate will be very high. [01:37:40] Durga Toshniwal: So this will be the point, if you can follow my cursor, that is when false positive rate is 0, then 2 positive rate will be 1. [01:37:47] Durga Toshniwal: Or it will be very high. [01:37:49] Durga Toshniwal: And, and like that. So, the ideal graph would be a right-angled triangle, or a rectangular, kind of a thing. [01:38:00] Durga Toshniwal: So that will be the ideal graph when we draw the ROC curve, which shows true positive rate versus true… sorry, true positive rate versus the false positive rate. [01:38:13] Durga Toshniwal: Anything that is non-rectangular as this dashed line, which is, actually lower to the central dashed line, is very bad. We want… we don't want a classifier which is, which is even worse than a random classifier. [01:38:28] Durga Toshniwal: We want a classifier that is better than a random classifier, and accordingly, you see these green, orange, and blue curves. [01:38:36] Durga Toshniwal: So the green curve shows a classifier that's slightly better than a random classifier, then the orange one shows still better, and the blue one is the best out of all these. Okay, so we want a classifier, that, is having true positive rate. [01:38:54] Durga Toshniwal: And to false positive rate to be optimally high. [01:39:00] Durga Toshniwal: Then, if we talk about the, area under the curve. So, here we talked about only the curve itself, we didn't talk about the area under the curve. [01:39:11] Durga Toshniwal: So, the area under the curve, I've shown here with the help of, some, so this graph. So, you see this is too negative, this is too positive, both are disjoint, so… and the decision boundary is in between. [01:39:24] Durga Toshniwal: When the decision boundary is in between, then, what happens? [01:39:29] Durga Toshniwal: Two negatives are qualified as two negatives, and two positives are qualified as two positives, and what you'll have is this kind of a ROC curve. [01:39:39] Durga Toshniwal: Which, okay. [01:39:41] Durga Toshniwal: And in this ROC curve, you can see that, the value is approximately 1. [01:39:48] Durga Toshniwal: If we talk about the decision foundry, which is slightly towards the right side, then you can see the intersection, and accordingly, you can calculate the area under the curve. [01:39:59] Durga Toshniwal: And this area under the curve will be lesser than the area under the curve in the first case, and it is, say, maybe around 0.8 to 0.9%. [01:40:09] Durga Toshniwal: Then, if we have moved the decision boundary further towards the right, we see a situation where there is an overlap. [01:40:17] Durga Toshniwal: And when there's an overlap, then what would happen? The, [01:40:23] Durga Toshniwal: So, when there's an overlap, then the area under the curve will reduce. [01:40:28] Durga Toshniwal: So, area under the curve will maximize when the true negative and true positive are separate. [01:40:33] Durga Toshniwal: As they start overlapping, as the overlap increases, the area under the curve decreases, and we are going… we are going towards a worse situation. [01:40:43] Durga Toshniwal: Then, the final situation is where the two positives and two negatives coincide with each other. [01:40:49] Durga Toshniwal: When they coincide with each other, then we are having, the worst combination, and we actually represented by this kind of a linear graph. [01:40:59] Durga Toshniwal: Okay, so these are the cases that are important. [01:41:03] Durga Toshniwal: Usually, a very good value would be 0.921, or 0.8 to 0.9. A fairly good value would be 0.7 to 0.8. [01:41:11] Durga Toshniwal: And then some poor value will be, like, 0.6 to 0.7, that is, we are talking about area under curve, and of course, if it is lesser than that, then we are not considering it at all. [01:41:23] Durga Toshniwal: So then, this is how we consider, true positive, rate and false positive rate, and AUC ROC curve. [01:41:32] Durga Toshniwal: Any questions, anyone? [01:41:37] Durga Toshniwal: Any questions, anyone, on this? [01:41:43] vinit shah: Ma'am, AUC is calculated in a certain way, like, sorry if there was a slide I… [01:41:48] Durga Toshniwal: cure. [01:41:49] Durga Toshniwal: I shouldn't. [01:41:58] Durga Toshniwal: Yeah, so AUC is the area under the ROC curve. [01:42:03] Durga Toshniwal: AUC's AD undercurve. [01:42:05] Durga Toshniwal: And it is the area that you can see shaded, which is the ROC curve, the receiver operating characteristic, and whatever area is covered by that will be the AUC. [01:42:17] vinit shah: Okay, and the inference that you have created over here, right, when the TN and the TPs are two separate classes and not overlapping with each other, that's when we get this AUC. So, keeping this TPR and FPR as a function. [01:42:29] vinit shah: We need to refer to the previous slide, is it? Like… The graph over there. [01:42:34] Durga Toshniwal: Yeah, so actually, what this means is that [01:42:37] Durga Toshniwal: When the decision boundary is central, and two positives are two positives only, and there's no overlap between the two positive and true negative, then the area under the curve will be maximized, right? There will be hump, or a distribution of the negatives, and there will be a separate distribution of the positives, and there's no overlap. [01:42:57] Durga Toshniwal: So, in this particular case, when we take the area under the curve, it will be sum total of the area under the true negative curve and true positive curve. [01:43:05] Durga Toshniwal: And so, it will be the maximum one, which is what you're seeing here. It will go up to 0, 1. [01:43:12] vinit shah: And I understood that part. [01:43:13] Durga Toshniwal: Huh. [01:43:13] vinit shah: What I mean to say is, the previous slide was talking about TPR as a function of FPR, right? And these graphs is talking about TNNTP. So, I'm just trying to see, does the TNNTP related to the FPR in any way, or what? [01:43:28] Durga Toshniwal: See, what I showed you here was that, this graph Yeah, so this is, so… [01:43:38] Durga Toshniwal: So, this graph talks about the true positive rate, which is TP related to TP, and this is the false positive rate, right? Now, the two are related in the sense that if we were talking about this kind of a situation. [01:43:52] Durga Toshniwal: So, ideal would be these two get separated out. [01:43:56] vinit shah: Correct. [01:43:57] Durga Toshniwal: Okay? [01:43:57] Durga Toshniwal: And, when they get separated out, and if we calculate the area, that will be maximum. [01:44:03] Durga Toshniwal: And then here, what we are talking about, in ROC is actually, we are talking about, the false positive rate, and false positive rate is 1 minus the specificity only. [01:44:17] Durga Toshniwal: Okay, and the 2-positive rate is also the sensitivity, or the recall. So, in a way, we are actually looking at, this TP, TR, and all that only. So, we are looking at 2 positive and true negative. Both are of importance. [01:44:32] vinit shah: So 1 minus TNR. Specific T would be 1 minus TNR for me, right? True negative ratio, okay. [01:44:39] Durga Toshniwal: False positive weight is 1 minus specificity. [01:44:42] vinit shah: Yeah, and specificity is TNR, right? Or… [01:44:45] Durga Toshniwal: Yeah, specificity is TNR. [01:44:47] vinit shah: Okay. Got it, got it, man. [01:44:50] Durga Toshniwal: And therefore, all this is of importance, right? [01:44:53] vinit shah: And what's the full form of AUC? Sorry if it was… [01:44:56] Durga Toshniwal: Area under curve. Area under curve. So, if these are two different curves, you draw the area under… you calculate the area under these two full two humps, then it will be maximum. [01:45:07] Durga Toshniwal: As the overlap increases, the area under the two curves will decrease because of the increasing amount of overlap. If the overlap increases, that is… what it means is that the fuzziness between the two negative and two positive is increasing. [01:45:23] Durga Toshniwal: Or, in other words, some true negatives are going into false positives, and some true positives are going into false negatives. [01:45:31] Durga Toshniwal: So, the overlap is increasing, and when the overlap is increasing, the area under the curve is decreasing because of this overlap here. And then the final worst situation is when two negative and two positive overlap. [01:45:45] Durga Toshniwal: When they overlap, then what happens? The area is the least, because we just have one hump. So if you calculate the area under this curve, it will be just this much. If we have an intersecting two humps. [01:45:58] Durga Toshniwal: then the area will be larger than the lowest case, which will be slightly more. So this area keeps on increasing as the overlap decreases. And what does the overlap mean? It means that some true positives are announced as [01:46:12] Durga Toshniwal: False, as false negatives, and true negatives are defined as false positives. That is the overlap. [01:46:22] vinit shah: Got it. [01:46:23] Durga Toshniwal: Yes. [01:46:24] Deepak Bobade: Oh, fault. [01:46:25] Durga Toshniwal: Yeah. [01:46:25] Deepak Bobade: Yeah, the first one is kind of ideal case, right? It won't happen in real life. [01:46:31] Durga Toshniwal: Exactly, exactly. So, this is the ideal case where all two positives are announced two positives only, and all two negatives are announced to negatives. It wouldn't happen. Practically, it wouldn't happen. [01:46:43] Deepak Bobade: Take a look around there. Yes. [01:46:46] Durga Toshniwal: And that's why achieving this, area under curve as 1 will be very difficult. [01:46:55] Durga Toshniwal: So your ROC curve would be, you know, it wouldn't be a rectangle or something like that. It will be having some smooth slope or something like that. [01:47:07] Durga Toshniwal: Because the value of 1 will not be reached. That is, area under curve will not be 1. So accordingly, this will be a smooth curve, not a kind of a step function. So this, if it has a value 1, then it will be like a step only. [01:47:26] Deepak Bobade: Okay. [01:47:27] Deepak Bobade: Okay, but there would be trade-off, right? [01:47:30] Deepak Bobade: We'll have to either, go for little increase in sensitivity, or a little increase in… Specificity. Specificity, yeah. [01:47:39] Durga Toshniwal: Yes, yes, that's correct. And accordingly, these regions would overlap based on what we are trying to attend. [01:47:46] Deepak Bobade: Yes, they will, yeah. [01:47:48] Durga Toshniwal: Yeah, and definitely we want a reduced overlap only. We don't want to… [01:47:53] Durga Toshniwal: a higher overlap. And if the overlap is very high, then we achieve this case, which is a random classifier, where the probability of classifying positive or negative is 0.5. [01:48:05] Durga Toshniwal: So if you are designing a classifier which is having a probability 0.5 of any class, positive or negative, is as good as not designing anything, because without designing also, you can say that the probability of occurrence of a class will be 0.5 only. [01:48:20] Durga Toshniwal: Right? I can… we can generally say, if you are having two classes, 0 and 1, then 0.5 will be the probability of having one class, and 0.5 will be the probability of another class. So this is the worst case, because classifier doesn't do anything. [01:48:35] Durga Toshniwal: We didn't require a classifier for this kind of a situation. [01:48:39] Deepak Bobade: Yeah, it just picks random, right? [01:48:41] Durga Toshniwal: Yeah, it's just 0.5. So, 0.5, we don't need a classifier to say 0.5. We can do it ourself also. [01:48:49] Deepak Bobade: ourselves also, yeah. We can… we can pick a finger, and then, yeah, we can do that. [01:48:54] Deepak Bobade: Yeah. [01:48:56] Durga Toshniwal: Any other questions, anyone? [01:49:03] Durga Toshniwal: So, so it's a little bit confusing if you think about it, so… [01:49:08] Deepak Bobade: Ma'am. [01:49:09] Durga Toshniwal: Sensitivity, yes. [01:49:11] Deepak Bobade: Yeah, over here, in the previous slide. [01:49:15] Deepak Bobade: So, you said the false positive rate and two positive rate should be optimally high, right? But in this case, if we look at the square, the right angle, right, above, 0.8, [01:49:31] Deepak Bobade: 0% to positive rate. [01:49:34] Durga Toshniwal: Yes. [01:49:36] Deepak Bobade: Okay. [01:49:50] Durga Toshniwal: Yeah, sorry for that. [01:49:52] Durga Toshniwal: Yeah, no, I said that true positive rate should be high, not false positive rate to be high. We want true positive rate to be high. We don't want false positive rate to be high. [01:50:03] Deepak Bobade: Gotcha, gotcha. [01:50:04] Durga Toshniwal: We want true positive rate to be high, we want true negative rate to be high. But we don't want false negative rate to be high, we don't want false positive rate to be high. [01:50:15] Deepak Bobade: Okay, so TPR and TNR, should be optimally high. [01:50:18] Durga Toshniwal: Yes. [01:50:19] Deepak Bobade: FPR, and then, what is this? FNR? [01:50:22] Durga Toshniwal: FNR. FNR. [01:50:23] Deepak Bobade: the, the. [01:50:25] Durga Toshniwal: This should be low. [01:50:26] Durga Toshniwal: So you just think about it, that, you know, if there's an alarm that is weeping, let's say that there's a faulty situation. [01:50:34] Durga Toshniwal: And in the faulty situation, there is an alarm that is set to blow, and let's say it's a manufacturing unit, it stops. So if the false positive rate is very high, every other time, the alarm will keep blowing, and the manufacturing will just stop. [01:50:53] Durga Toshniwal: So, this means false positive rate should not be high, because unnecessarily some corrective action is getting, done. [01:51:01] Durga Toshniwal: False negative rate also should not be high, because sometimes it is a false negative that might raise some kind of a trigger or alarm. [01:51:09] Durga Toshniwal: And and in that particular case, again, that might impact certain set of activities. So we don't want false positive rate to be high, we don't want false negative rate to be high. We want true positive rate to be high, or true negative rate to be high. [01:51:24] Deepak Bobade: Got it. [01:51:25] Durga Toshniwal: Right. [01:51:28] Durga Toshniwal: Okay, any other questions, anyone, before we wrap up for today? [01:51:38] vinit shah: Ma'am, just a request, if it's possible, right, for all these, sensitivity, recoil, precision, accuracy, and all. If you can have some examples, right, it could give a bit more context to the mind, like, if you can just have a table or something as a 2-minute catch-up in the next class, if that's possible. [01:51:53] Durga Toshniwal: Oh, I'll include what I did include on each slide, if you would have had a look. Let me just do a reshare. [01:52:03] Durga Toshniwal: So, actually, in… in each of these, I had included, if you would have read… [01:52:11] Durga Toshniwal: For example, we start out with this. [01:52:15] Durga Toshniwal: See? [01:52:16] Durga Toshniwal: Here, for each of these, I've given examples. [01:52:21] Durga Toshniwal: So… [01:52:22] Durga Toshniwal: They are not all together, but for each, there is an example, whatever I was discussing, like sensitivity, then specificity. [01:52:31] Durga Toshniwal: Then, precision, recall, Everything was there, actually, in the respective slides. [01:52:40] vinit shah: Yes, ma'am. [01:52:41] Durga Toshniwal: Yeah, it was there, actually, but you want it tabularized or something? [01:52:44] vinit shah: Yeah, just one tabular single slide, right? So that if we just look at it, we can, kind of just refresh it, looking at the example. I guess that's easier to remember than these, [01:52:54] vinit shah: Formulas, or whatever these are called. [01:52:57] Durga Toshniwal: No, these formulas are very important. You need to understand, and also, you may not memorize, but you need to have a good idea about them. [01:53:07] vinit shah: Okay, well. [01:53:08] Durga Toshniwal: Yeah, all through you will be using precision decal, sensitivity, FPR, FNR, all that stuff you'll be using. [01:53:15] vinit shah: Yes, sir. [01:53:16] Durga Toshniwal: So, these are important matrix. Okay, I'll try to do it next time, okay? [01:53:20] vinit shah: Thank you. [01:53:22] Durga Toshniwal: Thanks. [01:53:23] Durga Toshniwal: Anything else, anyone? [01:53:27] Durga Toshniwal: Okay, then I. [01:53:29] laxmi sahu: One question of, when we are expecting assignment and that grouping and all sorts of things, do we, like, group if [01:53:37] laxmi sahu: I can get any timeline idea. [01:53:40] Durga Toshniwal: You're talking about the formation of the groups? [01:53:43] laxmi sahu: Yeah, or maybe the assignment, and like, do we have any… [01:53:49] laxmi sahu: Or timetable, sort of, that in this week, we are covering this part, and this month… [01:53:54] Durga Toshniwal: Curriculum, yeah, curriculum. Yeah, yeah, that's ready, I'll share it with you. [01:53:59] Durga Toshniwal: So, it is… it's going to be, I mean, approximately correct in the sense that sometimes we move weeks ahead, or go back, depending on the pace and all that, but a rough curriculum is ready. I will share it in this week. Sorry for that, I wanted to share last week also. [01:54:15] Durga Toshniwal: But there were some things I wanted to add in that, that's the reason. I'll do that. [01:54:20] laxmi sahu: Oh, no problem. Thank you. [01:54:24] Durga Toshniwal: Okay then, thank you, everyone. Thank you, have a great Sunday. Thank you, bye-bye. [01:54:29] Pawan Misra: Thank you, man. [01:54:31] Krishnakumar MS: Give a… [01:54:32] Durga Toshniwal: Thank you.