# 06 2025-12-14 Supervised Learning Contd

course: Module 2 — Machine Learning Algorithms
module: Module-2-Machine-Learning-Algorithms
date: 2025-12-14
type: transcript
video_url: https://personal-learn.armco.dev/files/_Recordings/Module-2-Machine-Learning-Algorithms/06_2025-12-14_Supervised_Learning_Contd.mp4

---
[00:54:06] will be the total number of samples having
[00:54:09] That particular value divided by those having that particular class labeled.
[00:54:15] So, AI case, the number of instances having
[00:54:18] Attribute AI, belonging to class CK.
[00:54:22] Now, for example, we have to find out
[00:54:25] The probability of status equal to married.
[00:54:29] This is the status equal to married.
[00:54:31] Given class label, no.
[00:54:34] So, given class label merit, uh, given stat, uh, class label, no.
[00:54:40] How many are having status equal to married?
[00:54:43] So, out of class label, know how many are married.
[00:54:46] So this is one.
[00:54:48] Here it is married.
[00:54:51] then this is another one, class label is no, given status, uh…
[00:54:56] As statuses mattered.
[00:54:57] Then we have another one.
[00:54:59] Class label is no.
[00:55:01] Uh, and the status is married.
[00:55:06] And then we have no.
[00:55:07] class label, uh, is married.
[00:55:09] So, how many of them are having no?
[00:55:13] Uh, when, um…
[00:55:17] How many out of those having no are having status equal to merit?
[00:55:21] So, those having no are how many? 1, 2, 3, 4, 5, 6, 7.
[00:55:27] So, total 7 samples are heavy.
[00:55:30] Class label, no. Out of those having no, how many are having status equal to merit?
[00:55:36] We counted them 4, 1…
[00:55:38] 2, then 3…
[00:55:41] And then 4. So, this probability is 4 upon 7.
[00:55:47] Similarly, the probability of having
[00:55:50] A refund equal to yes, given, class label is yes.
[00:55:54] That is, how many samples are having cheating equal to yes?
[00:55:58] There are 3 samples having cheating equal to yes. 1.
[00:56:02] 2, 3. Out of these,
[00:56:05] How many are having, uh, refund equal to yes?
[00:56:10] So, refund is, yes, how many out of these three?
[00:56:15] Refund is no here, refund is no here, refund is no here.
[00:56:17] So, it's going to be 0 on 3, which is 0.
[00:56:22] So, like this, we actually calculate these values.
[00:56:26] And then use them to infer the class label.
[00:56:30] So, let's say now, in our question,
[00:56:33] Uh, in this problem,
[00:56:36] We are given an unseen record with refund equal to no.
[00:56:40] Marital status equal to single.
[00:56:43] An income equal to 70, then what should be the class label?
[00:56:47] So, to calculate this, we will use the…
[00:56:51] Bayes theorem, and what does it say? Suppose probability of C given…
[00:56:56] A1, A2, A3 till AN.
[00:56:59] is going to be probability of…
[00:57:03] A1, Q1C…
[00:57:05] into probability of A2 given C, like that.
[00:57:10] And so on and so forth, into probability of C.
[00:57:13] Do I do it by probability of A?
[00:57:17] Like that. So, let's say… and in this case, we are having two class labels, yes and no.
[00:57:24] So, here in this case, we have to find out the class label.
[00:57:29] that the unseen record,
[00:57:30] will be having, will we know?
[00:57:33] Given the values A1A2, A3, whatever. So, what is it?
[00:57:37] refund.
[00:57:39] Equal to no.
[00:57:41] Marital status, I'm writing in brief, is single.
[00:57:46] And… income.
[00:57:49] Equal to 70. Okay.
[00:57:52] will be equal to…
[00:57:55] Probability of having…
[00:57:58] So, we are going to now subtit this pi.
[00:58:02] refund.
[00:58:05] equal to no given.
[00:58:07] class label is no.
[00:58:10] into…
[00:58:12] probability of marital status equal to single.
[00:58:17] given class level is no.
[00:58:20] into probability of…
[00:58:23] Uh, income.
[00:58:25] Equal to 70, given class label no.
[00:58:29] divided by, uh, yeah.
[00:58:31] And then into probability of class C.
[00:58:34] into probability of class no.
[00:58:38] divided by probability of these value pair combinations. What is it? Refund?
[00:58:44] Equal to no.
[00:58:46] Marital status equal to single.
[00:58:48] And income equal to 70K.
[00:58:51] Is it okay, everyone? Any questions?
[00:58:54] Anyone on the calculation?
[00:58:57] Then I'll proceed to solve it.
[00:59:02] So now, the probability…
[00:59:05] That refund is no given class label is no.
[00:59:09] So, the total number of records having class labeled no R, as I already calculated, 1, 2,
[00:59:15] 3, 4, 5, 6, 7.
[00:59:20] So, probability…
[00:59:22] off, so I'll write on the next slide.
[00:59:27] Probability of defund.
[00:59:31] equal to no given class label no.
[00:59:34] will be equal to…
[00:59:36] Total are 7, right?
[00:59:39] 2, 3, 4, 5, 6, 7.
[00:59:42] And out of these 7, how many are having refund equal to no?
[00:59:47] So, out of this…
[00:59:49] This is having refund, yes, so we cancel this.
[00:59:52] Then, out of this, this one is having a no, we count this.
[00:59:56] Out of this, again, defund equal to no, one more.
[01:00:00] then… this is having a defund, yes, so we don't count this.
[01:00:05] then refund equal to no, another one.
[01:00:08] Currently, this is yes, so we don't count it.
[01:00:12] Then refund equal to no another one. So, how many? 1, 2…
[01:00:15] 3 and 4. So, probabilities 4 into 7.
[01:00:20] sorry, 4 divided by 7.
[01:00:22] Then the probability of
[01:00:25] Marital status equal to single.
[01:00:27] given class level, no.
[01:00:29] How much is that? Total are 7 already?
[01:00:33] Out of these 7, how many have class labeled single?
[01:00:36] 1…
[01:00:38] then… 2?
[01:00:42] Then…
[01:00:45] 3, because divorced and single are in the same category, right? So we count it.
[01:00:50] So this is one. This is two.
[01:00:54] Then… this has 3.
[01:00:58] And that's all.
[01:00:59] First one is… first one is yes, ma'am.
[01:01:02] Sorry.
[01:01:03] Referred, yes. First one is refund, yes, right? We are looking for refund also, right?
[01:01:09] That is true, but now we are only considering… remember, we have considered
[01:01:14] The possibility or the probability of refund independent of marital status, independent of taxable income.
[01:01:21] So, we will look at their values independently only.
[01:01:24] We are not going to relate refund equal to yes, so we will…
[01:01:28] Uh, you know, we will not consider.
[01:01:32] Here, we are considering, as I said, in Naivee's…
[01:01:35] We consider the possibility of having attributes independent of each other.
[01:01:39] So, first of all, we consider defund equal to no.
[01:01:45] So…
[01:01:46] Okay? Then, now, we forget about refund and taxable income, and only consider marital status.
[01:01:50] So, out of those having…
[01:01:52] Uh, class label, no. We are having how many? One.
[01:01:55] 2… and 3…
[01:02:00] Okay, okay.
[01:02:03] Yeah, yeah.
[01:02:04] That's it. So it's possibility… okay? So, this is how the independence works.
[01:02:08] And then…
[01:02:09] Okay, actually, I was confused with the no. I thought that no was refund, but it was cheating, actually.
[01:02:13] Given no ass cheating, right?
[01:02:14] Yeah, yeah, yeah. This is class level. I'm right with you, so…
[01:02:20] And then we have income…
[01:02:22] Uh, equal to 70K.
[01:02:25] given class level, no.
[01:02:27] So, out of these, how many are having 70K? Remember, we are discretized.
[01:02:32] So, if we consider this…
[01:02:34] Uh, this one is having 125, so it will not come in this category.
[01:02:40] Then this one will… I think it was greater than 80K.
[01:02:43] And, uh, less than 80K was another category.
[01:02:47] So then, out of this, this will also not come.
[01:02:52] Then, here we have…
[01:02:54] 70K, so this will be 1.
[01:02:58] then this is also higher than 80.
[01:03:00] Then, another one is 60, so this mice will also come in this category, same category as 70K.
[01:03:07] 75K will also come in the same category. So again, we are having three samples out of a total of 7.
[01:03:16] Okay.
[01:03:18] So, these are the three. Then, probability of having class level no is how much we already calculated. It is 7 on 10.
[01:03:27] And the probability of having this value-pair combination.
[01:03:30] refund, uh, equal to no.
[01:03:35] For my marital status equal to single.
[01:03:37] And, uh, income.
[01:03:40] equal to 70K. What is the probability of this? Let's calculate that.
[01:03:45] refund equal to no marital status, single.
[01:03:48] And income 70.
[01:03:50] The refund is no marital status, single.
[01:03:54] But income is not…
[01:03:57] 70K, so we don't consider this… we don't consider this.
[01:04:02] Uh, here we consider this refund is equal to no.
[01:04:05] Madrid felt a status singer, and…
[01:04:09] Income is $70K, so this is one.
[01:04:12] Uh, one, uh, value pair combination which matches.
[01:04:17] then this doesn't match.
[01:04:19] This doesn't match? Okay, this one we considered.
[01:04:22] Refund is no. Marital status is marriage, so we don't consider this. Again, this is gone.
[01:04:29] refund, uh… this is refund equal to yes, so we don't consider this.
[01:04:33] Then we have refund equal to no.
[01:04:36] And, uh, this is marital status is marriage, so we don't consider this.
[01:04:40] And that's it. So, the probability of having this value per combination
[01:04:46] is 1 on 10.
[01:04:50] However, usually, uh, and the expression was what? This was our expression.
[01:04:55] Right? The expression was…
[01:04:58] We put all these probabilities in the numerator, whatever we got.
[01:05:03] into probability of no.
[01:05:05] So, actually, many a times, we just ignore the calculation of the denominator, because
[01:05:10] Whatever be the class label, this is going to be fixed.
[01:05:15] So, dividing it by the same factor for all the class labels will not matter much.
[01:05:20] So, we can ignore this also. So, I'm going to ignore this denominator.
[01:05:24] From the calculation, and only going to consider the numerator.
[01:05:28] Because even if this class label was, yes, this will still be the same.
[01:05:34] So, I'll just multiply all these
[01:05:36] These four probabilities together, which comes out to be…
[01:05:40] Uh, if I talk about the probability of having the class labeled,
[01:05:44] No, given these value pair combination. Refund equal to no.
[01:05:49] Marital status equal to single.
[01:05:52] income is equal to 70K.
[01:05:54] So, and then this becomes equal to…
[01:05:57] Probability of…
[01:05:59] refund, uh…
[01:06:02] So, let me just write this.
[01:06:04] 4x7.
[01:06:06] into 3x7.
[01:06:09] into 3x7.
[01:06:12] into 7x10.
[01:06:14] And this gets answered.
[01:06:16] So what we are going to have here is…
[01:06:19] How much? 94s are 36.
[01:06:24] So, we are having 9436, and…
[01:06:27] You have 77 ja 49 into 10.
[01:06:31] So, this is the probability.
[01:06:33] of having class labeled,
[01:06:35] No. This is for the class label known.
[01:06:38] Is it okay, everyone?
[01:06:40] Any questions? Anyone?
[01:06:47] Ma'am, one question, uh, here it is given that income is equal to 70,000. Then why we are considering below 70 also, 70, 60, 75 also.
[01:06:55] Yeah. Hmm.
[01:06:59] Because if you recollect, when I was doing the decision tree,
[01:07:02] We had discretized this taxable income.
[01:07:05] Right? It is a continuous valued…
[01:07:07] a variable. It can take any value in our range.
[01:07:11] So, we need to discretize it so that
[01:07:13] Uh, it is having certain, uh, discrete set of values, which we can count.
[01:07:18] Otherwise, if we talk about this range, let's say from 60 to 125, the values can be infinite, right?
[01:07:25] It could be 70.1, 70.01, 70.001, like that.
[01:07:30] So, to make it a finite set,
[01:07:32] We use the decision boundary.
[01:07:35] of greater than, uh, equal to 80K.
[01:07:41] Uh, as one category.
[01:07:42] And rest of it as another category.
[01:07:46] Okay, okay.
[01:07:47] So, yeah, so 70, 60, 75 all come in this rest category.
[01:07:54] Yeah, I just used the same example.
[01:07:55] Okay, it was already categorization.
[01:07:57] So that, uh…
[01:07:59] Uh, you know, we could follow it in a similar way, so…
[01:08:03] Here, the discretization is following this rule.
[01:08:07] greater than or 80K will be in one category, all the rest, less than…
[01:08:12] 80 will be in the other category. Based on that only, I've done this count.
[01:08:17] So, whenever you are using knife-based, uh, classifier, you'll have to first discretize
[01:08:23] The, uh, attributes, if they are continuous valued.
[01:08:26] The rest of them are all discretized, because these are having two values, these are having…
[01:08:30] certain set of values, right, which are finite and can be counted.
[01:08:35] Otherwise, we cannot find out the probability.
[01:08:40] Anyone else? Any other question before I proceed?
[01:08:46] Okay, now that we have calculated the probability for no,
[01:08:51] We can go ahead and calculate the probability for ES also in a similar way.
[01:08:56] So then, now, suppose the class label is yes.
[01:09:00] Then, what are we going to calculate the probability of refund?
[01:09:04] equal to no Q1 class label yes.
[01:09:09] Probability of marital status equal to single.
[01:09:13] Last label equal to yes.
[01:09:16] Probably of income.
[01:09:19] Equal to 70K, given.
[01:09:21] Class label, yes.
[01:09:24] And then probability.
[01:09:27] of class level is. These are the things. I'm ignoring the calculation of denominator.
[01:09:32] Because it's going to have this combination once again, so it's calm.
[01:09:36] So now, let's look at the yes cases. This is one yes, this is one yes, this is one yes.
[01:09:41] Probability of refund equal to no.
[01:09:44] When filing of income taxes equal to yes.
[01:09:47] Is there, uh…
[01:09:49] Such cases. So, when…
[01:09:51] Uh, the cheating is yes.
[01:09:56] Filing of income tax… uh, sorry, refund is equal to no.
[01:10:00] Refund, no.
[01:10:02] Reef or no.
[01:10:03] And then, here also refund.
[01:10:05] So, the probability of having refund equal to no
[01:10:09] Given the class label, cheating on filing of income taxes, yes, is 3 on 3, which is equal to 1.
[01:10:18] Now, let's calculate a marital status single. So…
[01:10:22] Because single and divorced came in the same category, so I'm going to count this as 1.
[01:10:28] Then this has 2, this has 3.
[01:10:31] So, it is again 3 on 3, which is 1.
[01:10:36] And then income equal to 70K given.
[01:10:39] Uh, class label, yes. So, income here is…
[01:10:42] This was the discretization.
[01:10:46] So, if the income is greater than or 80K, it comes in one category, otherwise it comes in the 70K category. Here, income is…
[01:10:53] greater than, uh, or equal to…
[01:10:56] 80, it is 95, so…
[01:10:59] We don't count this. Here, also, it is greater than 80. Here, also, it is greater than 80.
[01:11:04] So this means the probability of income equal to 70 given
[01:11:08] cheating on filing of income taxes, yes.
[01:11:11] is 0 on 3, which is equal to 0.
[01:11:15] And probability of having GS is 3.
[01:11:18] About 10. Now, if you multiply all this, because of 1, 0,
[01:11:22] This whole expression of probability of having
[01:11:25] Class label, yes.
[01:11:27] That is cheating on income tax, filing of income taxes, yes, given.
[01:11:32] refund equal to no.
[01:11:35] Marital status equal to single.
[01:11:38] Earned income is equal to…
[01:11:41] 70K will be equal to the product of all of these, so which is 1 into 1.
[01:11:48] Into 0, into 0.
[01:11:50] And I'm going to ignore the denominator calculation, so this I'm not going to do.
[01:11:55] So it's coming out to be zero.
[01:11:57] So this means that for this particular unseen record,
[01:12:01] the probability of having a class labeled yes is 0.
[01:12:05] And in the previous case, the probability was 36 on 490, so definitely,
[01:12:10] The probability of cheating on filing of income tax
[01:12:14] of no is higher than that of yes, because this is 0.
[01:12:18] So, final, class label.
[01:12:21] Uh, so the class label that we will infer
[01:12:26] will be equal to no, because the probability of no
[01:12:30] Probably of having no.
[01:12:32] is greater than the probability of having yes.
[01:12:36] This is how KniveBase calculates the probability.
[01:12:39] Is it okay, everyone? Any questions?
[01:12:40] Okay, yeah.
[01:12:43] Yeah, sure.
[01:12:44] Uh, yes, yes. So, uh, when will this probability come into picture?
[01:12:48] Uh, so we are already predicting on the basis of the available features, right?
[01:12:54] Mm-hmm.
[01:12:55] And then we are calculating this probability. So, what is the use of this probability, the inferring, uh, like what we inferred, right? The no class.
[01:13:04] Yeah, so this is another way of inferring the class label.
[01:13:08] For example, you use the decision tree.
[01:13:11] Where you used, uh, you know, value pair combinations, then the path on the tree was traversed.
[01:13:17] And based on that, the class label was inferred.
[01:13:20] Then another wave was k nearest neighbor, where you took the K number of nearest neighbors to look at the class labels of those.
[01:13:27] And based on the majority vote,
[01:13:30] Uh… so you used to infer the class name.
[01:13:35] Now, this is a third way to infer the class label, where you're using probabilities
[01:13:42] to infer the class label. However,
[01:13:44] Uh, the only thing is that if we are given a data, then how do we decide?
[01:13:49] Which particular class… classifier to use.
[01:13:52] So, usually, when we have non-linearity in the data, as I already discussed,
[01:13:58] For example, I told you that if we are having a
[01:14:03] person having the, you know, some medical disease.
[01:14:06] Then, uh, there can be many external factors, and therefore, uh,
[01:14:10] uh, classifier, which is purely based on the attribute-value combinations.
[01:14:16] may not be able to infer.
[01:14:17] So, in all such cases,
[01:14:18] Okay.
[01:14:19] Probabilistic, uh, classifiers are more useful.
[01:14:23] So, whenever you are given a data,
[01:14:25] either you should have some background information about the data,
[01:14:30] Otherwise, what you can do is try to use…
[01:14:32] Different classifiers on the data,
[01:14:34] And then see which one is working best.
[01:14:37] Okay, like that, you can… so this is another way to infer the class label.
[01:14:42] So, okay.
[01:14:43] So, it… yeah, go ahead.
[01:14:45] Infer, as in we are predicting, right?
[01:14:47] Yes, correct. We are predicting the class label.
[01:14:50] Based on the probabilities, because the probability of having the class labeled
[01:14:55] No was 36 on 490.
[01:14:58] Which is greater than the class level probability for class label yes, which is 0.
[01:15:03] Therefore, we are predicting the class label as no based on the probability.
[01:15:13] Yes.
[01:15:14] No, it would be considered the whole dataset over here, right? So, let's say if we…
[01:15:15] put in one row, right?
[01:15:17] to predict what, uh, what we think, uh, class, right? Whether he will cheat or not, right?
[01:15:24] Yeah.
[01:15:25] So, at that moment, how will we, like, implement this thing?
[01:15:28] to predict.
[01:15:29] Yeah, yeah. So remember I told you there are two kinds of classifiers. One, eager learners.
[01:15:36] Another, lazy learners.
[01:15:38] Eager learners are those that build a classifier model and keep it ready.
[01:15:42] to make a prediction, whenever unseen record is presented.
[01:15:46] Lazy learners, they don't do anything. They just sit with the training data.
[01:15:51] And wait, uh, for the, you know, unseen record to be presented.
[01:15:56] So, this falls in this category of lazy learners only. I'm telling you, lazy learners right now.
[01:16:01] Okay. Okay, okay, okay, okay, right.
[01:16:03] So there is no model building.
[01:16:06] Whenever there's an unseen record, so this record was presented,
[01:16:10] Now, what… what should it do?
[01:16:11] So, what will it do? It will try to calculate the probability of occurrence of all the classes possible. Here, there are two classes, no and yes.
[01:16:20] And then it found out that the probability for class no is higher, so it…
[01:16:23] predicted the class label as no, that's all. But there is no border building here.
[01:16:28] Okay, got it. It's lazy, yeah, that's what I forget.
[01:16:32] Yeah. Yeah.
[01:16:34] Any other questions, anyone?
[01:16:44] So, uh, now…
[01:16:46] Uh, just a few points about the knife-based classifier.
[01:16:50] So, if you recollect, in the KnRS neighbor, which was another lazy
[01:16:55] Lana-based.
[01:16:58] classifier. We had, uh…
[01:17:00] So this K and N was actually not robust to noise.
[01:17:03] If there was noise, it would impact
[01:17:06] the class label. But, if we are talking about knife,
[01:17:10] a base classifier, because if there are noise points, then the probability
[01:17:15] of the noise points,
[01:17:17] Uh, you know, uh…
[01:17:19] having an impact on inferring the class label, or predicting the class label,
[01:17:24] will be very low. Therefore,
[01:17:27] Knife-based classifier is actually very much robust to noise.
[01:17:31] So if we have noisy data, we can think of using knife-based classifier.
[01:17:37] Uh, then, um…
[01:17:40] It can also be used, the knife-based classifier,
[01:17:43] For, uh, filling in missing values. Suppose we had missing values in the data,
[01:17:49] Then we can use this to predict the value or the class label that should be there.
[01:17:54] And we can use it.
[01:17:56] Then, uh, knife base is also very robust to irrelevant attributes. Certain classifiers, they don't work very well.
[01:18:03] When there is redundancy, or there are attributes that are actually not relevant to our
[01:18:08] Uh, prediction. In this case,
[01:18:11] Even if there are irrelevant attributes, they will, uh…
[01:18:15] not play any major role.
[01:18:18] In the prediction, because the probability of the irrelevant attitude will be almost equal,
[01:18:23] Given any class label.
[01:18:25] Therefore, only those which are important to the classification will play a role, and they will help to predict the class level. So, therefore,
[01:18:33] Knife base is also robust, too.
[01:18:36] Not just robust to noise, but also robust to relevant attributes.
[01:18:41] Uh, and then one important thing is that knife is, uh…
[01:18:46] assumes independence.
[01:18:49] However, the attributes may not be independent of each other all the time.
[01:18:54] If the attributes are not independent,
[01:18:57] Then we should not use knife.
[01:19:00] Okay. So, uh…
[01:19:02] I'll give you a typical use case of NAI, for example, I want to categorize the word Amazon.
[01:19:11] So, uh…
[01:19:14] So, what could be the categories in which Amazon could fall?
[01:19:18] Suppose there's a word, Amazon, that is being encountered.
[01:19:22] What could be a class label for this word? Anyone?
[01:19:26] I'm sure all of you will quickly say something.
[01:19:35] Got it. So, the first option that would come to mind is e-commerce.
[01:19:41] website, or e-retailer, or whatever.
[01:19:44] But some of you might think of something else also, and what can be that?
[01:19:45] Eco.
[01:19:50] Yes, yes, so there's a river, and there's a rainforest area.
[01:19:51] River. River.
[01:19:55] Which is called the Amazon rainforests and all that.
[01:19:58] So, uh, the probability… so…
[01:20:02] Let's say we are using Amazon in a sentence.
[01:20:05] Let's say I say I purchased, um…
[01:20:09] Lot of good things on Amazon.
[01:20:10] Then, based on the context, the probability of Amazon belonging to e-commerce will be higher,
[01:20:16] As compared to it being a river.
[01:20:18] However, in some other sentence, if I say,
[01:20:21] The Amazon River is very long.
[01:20:25] Then, in that particular case, based on the context, the probability of it being
[01:20:29] e-commerce website will be very low, and that, it being river or a rainforest,
[01:20:34] Rainforest area will be.
[01:20:37] Uh, rainforests will be high.
[01:20:38] So, like this, we can use probabilistic
[01:20:41] Uh, based, uh, probability-based prediction.
[01:20:45] To find out the class level.
[01:20:48] However, there can be certain cases where…
[01:20:51] The assumption of independence may not hold. For example, here,
[01:20:55] The probability of Amazon being in e-commerce or being a rainforest, they are all independent.
[01:21:02] Uh, but let's say if we are having a time series data,
[01:21:06] Uh, where we are having, let's say, some value.
[01:21:10] Which is B, uh, sorry, here we are having tying.
[01:21:14] And we are having some value that is being measured.
[01:21:17] And, uh, let's say this is time 0, time T1, time T2.
[01:21:24] Like that till time t, and here, that…
[01:21:27] Second value actually may depend on the first value. First may depend on the previous, like that.
[01:21:31] If we are having this kind of dependence, then we will not use NiveView.
[01:21:35] Because Naive assumes independence in the attributes.
[01:21:39] Where in this particular case, the values at time T1, T2, T3, tilt t, and…
[01:21:44] Maybe they're different, uh…
[01:21:46] Values that may… that may be considered.
[01:21:49] Just a second.
[01:21:52] Sorry for that.
[01:21:58] Yeah.
[01:22:10] Okay, so, uh, accordingly, based on the use case, we need to decide whether to use NI or not to use NAI.
[01:22:17] So, uh…
[01:22:21] I've discussed what is Knive, how we use…
[01:22:24] Now, uh, another thing that I would like to discuss is K-Fold cross-validation, which is not directly related to any category.
[01:22:32] But it's another… it's a way of performing
[01:22:35] Uh, validation of a classification model.
[01:22:38] So I think in your hands-on, you have already discussed, you have already been…
[01:22:43] taken through the usage of how to decide K.
[01:22:49] Where you look at some performance metric, like,
[01:22:53] accuracy, precision, recall, right?
[01:22:56] anything of that kind.
[01:22:58] Versus the value of K,
[01:23:00] And then, probably, you look at the knee of the curve.
[01:23:05] Say, here it is, the knee, and you use that value of K,
[01:23:09] to decide.
[01:23:11] Uh, you know, how to build a classifier, or…
[01:23:15] how to break the classifier into training data and test data, how to…
[01:23:21] utilize that. Let's say the training data, uh…
[01:23:25] here was 0, then it was, like,
[01:23:29] 80, uh, sorry, 20, 30, 40, 60…
[01:23:33] 80, 100. So, wherever.
[01:23:36] The performance was best, optimally best.
[01:23:39] You took that value to…
[01:23:41] Uh, I saw, uh, training data, amount of training data that you are going to choose.
[01:23:46] to build a model, and rest of it would be your…
[01:23:49] test data. So that was one way to identify that value based on the knee of the curve. Another way is to use K-Fold cross-validation.
[01:23:58] So, what is careful cross-validation?
[01:24:00] In careful cross-validation, what you do is that you divide, like, given training data into folds.
[01:24:06] Polls are all equal in size. You can think of…
[01:24:09] Your training data, like this rectangular sheet of paper, okay?
[01:24:14] like this, and you actually fold it into equal…
[01:24:19] sized phone. And then…
[01:24:21] Uh, you have different iterations.
[01:24:24] In which you choose one fold at a time.
[01:24:26] For example, in first iteration,
[01:24:29] You have this… this is your paper.
[01:24:32] You folded it into 5 parts.
[01:24:34] One fold, you choose as the test data, rest of the four you choose as the training data.
[01:24:40] Now, what can happen is, in this particular case,
[01:24:43] Now, if we use a certain…
[01:24:45] subset of the training data to build them, uh, to use as the test model.
[01:24:50] It is possible, and the rest of it to train the model.
[01:24:54] Then, it can happen that…
[01:24:56] certain samples might be lost during training.
[01:25:00] And when being used as tests, they can affect the test performance. So,
[01:25:06] To have a fair distribution of the training versus the test samples,
[01:25:10] In each iteration, one of the folds will be used. So, in the second iteration,
[01:25:14] Some particular fold is used, which is different from the other one, and rest of them will be used for train.
[01:25:20] Then here, one fold will be used. Then here, one fold will be used.
[01:25:24] Then here, one fold will be like that.
[01:25:27] So, every time, different folds will be used for
[01:25:30] Training versus test, and then the average of all these.
[01:25:33] Uh, the average performance of all these will be used.
[01:25:38] As the performance of the, uh…
[01:25:40] final model that is built.
[01:25:44] So, this is what is K4 cross-validation.
[01:25:46] That you divide the given training data into K number of folds,
[01:25:50] And then you use these folds.
[01:25:52] Uh, to iterate.
[01:25:55] on a model training, as well as prediction.
[01:25:59] And then, uh, you give the final performance. So, this is how you use…
[01:26:05] K-fold cross-validation. This is another way
[01:26:08] to choose, uh…
[01:26:11] how… which data to use for training, and which one to use for test.
[01:26:15] So, here again, this is the same one.
[01:26:18] So, for each of these folds, whenever it is used in one first iteration, some performance is there, then
[01:26:25] Like, for each iteration,
[01:26:28] Uh, there will be different performances that will be obtained. This performance could be accuracy,
[01:26:35] And or it could be…
[01:26:37] precision…
[01:26:40] It could be recall…
[01:26:43] It could be F1 score, or whatever we are interested in, or all of that, and then use that.
[01:26:50] average of all these to infer the performance.
[01:26:53] of the final mod.
[01:26:55] So this is how K-Fold cross-validation is used.
[01:26:59] So, is it okay, everyone? Any questions so far on K4 cross-validation, or on NIVE?
[01:27:06] Before I start with the…
[01:27:08] a new portion.
[01:27:11] Any questions, anyone?
[01:27:12] Okay, thank you, thank you.
[01:27:16] Yeah, if we could, uh…
[01:27:20] Sorry?
[01:27:23] Uh, one question, why do we take the average link?
[01:27:25] We take the average because in this, we are using K number of folds. Here, k is equal to 5.
[01:27:35] to give a…
[01:27:36] fair idea about what the performance is.
[01:27:39] Suppose we infer the performance of the model we are building,
[01:27:43] Using this as the training data, and then this as the test data.
[01:27:47] then it is possible that this particular fold may be having certain class as the majority class.
[01:27:54] And so on and so forth. So this value of performance…
[01:27:58] may not be the correct, uh…
[01:28:00] Reflection of the overall performance of the model.
[01:28:04] similarly, if one particular fold, this is used as the test data, and the rest of it as used
[01:28:09] is used as a training data.
[01:28:11] And then the performance may be biased either on these samples or on this particular
[01:28:17] test sample set. So, therefore, to make it
[01:28:21] unbiased, no matter what portion of the…
[01:28:24] training data and the test data we are choosing to find out the performance.
[01:28:29] whatever be the number of folds, the overall performance will be the average of
[01:28:35] All the independent performances, so that if there are any data biases or any other such issues,
[01:28:41] They get removed by, uh…
[01:28:44] doing this, uh, averaging.
[01:28:48] Is it okay?
[01:28:51] Anyone else? Any other question?
[01:28:52] Yeah, sure, thank you.
[01:28:56] Yeah, uh, like…
[01:28:57] How was it different, too?
[01:28:59] the earlier, uh…
[01:29:01] kind of method wherein we use to kind of, let's say,
[01:29:04] train interest.
[01:29:06] VINs. Yeah, make bins, right?
[01:29:10] Uh, how is this different, you are saying?
[01:29:16] So, in the earlier method, what you do was, let's say that…
[01:29:21] You… suppose we are talking about training data, so…
[01:29:24] You use 10% of it, 20, 30.
[01:29:29] 40 up to, maybe.
[01:29:30] 80, 90, 100%, whatever.
[01:29:32] And this is your performance.
[01:29:35] And, uh, initially, the performance is going to be very low.
[01:29:40] And as you go, it is going to be higher and higher.
[01:29:43] So, you kind of choose the…
[01:29:46] knee of the curve to decide how much of it will go in the test set.
[01:29:47] Yes.
[01:29:51] And how much of it will go in the training set.
[01:29:55] This was the method that was discussed earlier. This is another way of doing the…
[01:30:01] partitioning into training and testing. That's it.
[01:30:05] Both are equally popular.
[01:30:11] Any other questions? Anyone?
[01:30:12] Yeah, yeah, the J cross-fold validation is to, uh, find out the hyperparameter, right?
[01:30:21] So here, K4 cross-validation actually decides
[01:30:26] How much of, uh…
[01:30:28] Uh, so, uh, here we actually don't decide
[01:30:33] exactly, like, 40, 60, 80-20, like that, we don't decide how much to put in training and test.
[01:30:40] It automatically, based on the value of K, it automatically makes those many folds.
[01:30:47] And then iterates on all the folds to do the training and test.
[01:30:51] And then gives the performance. Only thing that we need to decide is the correct value of the number of folds.
[01:30:57] So, in this case, what are we going to do?
[01:30:59] If these are the number of folds.
[01:31:02] Then, let's say we have k equal to 0, 1, 2, 3, like that.
[01:31:07] till, let's say, some number of folds we decide.
[01:31:10] And this is the performance.
[01:31:12] So, to decide the correct number of folds, initially,
[01:31:16] With K number of folds, the performance will be something.
[01:31:19] Then, with the one number of folds, it will be something. Initially, it will be very high.
[01:31:25] Because all the data is being considered. Slowly, it will decrease, and it will become stable.
[01:31:31] So, we don't also need unnecessarily higher number of folds, because they are not adding any value.
[01:31:37] to our training and test purpose. So, we will go for those many folds.
[01:31:41] That gives an optimally high amount of performance.
[01:31:44] Increasing number of folds beyond this limit,
[01:31:47] This value of K is actually not going to help us
[01:31:51] improve the performance. It only adds to more and more complexity in terms of the number of computations.
[01:31:58] So, we don't increase beyond this value.
[01:32:01] So, this will be the number of folds that we will choose.
[01:32:04] to, uh, you know, to have for our model-building purpose.
[01:32:12] So then this is also a hidden trial that will start from some value, and then we'll go until this curve becomes flat.
[01:32:18] So, fortunately or unfortunately,
[01:32:22] If we talk about ML, AI, DL,
[01:32:26] Generative AI, or anything for that matter.
[01:32:29] This hidden trial will be there all through.
[01:32:33] Okay, so the… what you're saying is correct, that we actually are not…
[01:32:34] Okay.
[01:32:37] Doing anything other than iterating or doing kind of a…
[01:32:41] not exactly hit and trial, we are having a scientific way to decide
[01:32:45] What value of K2 chose?
[01:32:48] And to do that, fortunately or unfortunately, or coincidentally,
[01:32:53] It's an irony that we, you know, we want to choose the correct value of K to build the model.
[01:32:58] But even before choosing the correct value of K, we have to iterate on all possible or…
[01:33:04] Certain set of possible values of k.
[01:33:06] And use all of these to build a model.
[01:33:09] And then find out the correct or optimally correct value of K, and then use it to build up.
[01:33:14] final model. So, before even building the final model,
[01:33:18] We have to build all possible models using different values of k. So, we have to do so much.
[01:33:24] To use correct value 1 value of k, we have to actually…
[01:33:28] Iterate on all possibilities.
[01:33:30] or set of possibilities, and then choose the correct value of K, and then use that to, again, build the final model.
[01:33:37] So, what you are saying is correct, that we have to do…
[01:33:41] much more to choose one
[01:33:43] correct value. That is the thing. Until and unless we know that correct value, which obviously…
[01:33:48] We may not know.
[01:33:50] So then, to find out the correct value, we have to use all the values, then see what is giving the best performance, and then use that as the
[01:33:59] value for… of our choice.
[01:34:07] Any other questions?
[01:34:09] Yeah, yeah.
[01:34:10] I have a question here. So, uh, is it good to assume that all of these iterations are.
[01:34:16] creating different versions of the model. Right? Uh, basically, Performance 1, Performance 2, all of them are.
[01:34:27] Yeah.
[01:34:28] More performance of different models. So, at the end of the training of, let's say, the five-folds, which model do we use for.
[01:34:34] let's say productionizing the solution.
[01:34:37] So, here in this particular case, it is the value of k which is important.
[01:34:44] Not any one particular one out of this.
[01:34:45] Hmm.
[01:34:47] We want to choose those many folds.
[01:34:51] So, what we have to decide here, the question that we hand is which K to use.
[01:34:56] So, if I… if I have k equal to 5,
[01:34:59] Then I'll make all these iterations and see the average performance.
[01:35:04] Which will be for K equal to, if this was 5, then whatever the value of performance is,
[01:35:10] This is what I've got. Then I repeat this process with k equal to 4, with k equal to 6, with k equal to 8, with k equal to 10, like that.
[01:35:19] And then choose that value of K,
[01:35:22] Uh, which gives the best overall this average performance.
[01:35:26] So, we are not choosing one model out of these models. We are actually
[01:35:31] Deciding on the optimally best value of K.
[01:35:35] Because what is this K doing?
[01:35:37] Here, if K was 5…
[01:35:39] Then what are we doing here? We are actually splitting.
[01:35:42] Uh, the training test into 20s to 80, you can see this, right?
[01:35:46] If K was something else, accordingly, this split ratio will…
[01:35:51] differ. So, in a way, by doing all these folds, we are actually doing this.
[01:35:53] Mm-hmm.
[01:35:56] Also, my question is a bit different. So, let's say we finalize that k is equal to 5 is giving us the best cumulative performance.
[01:36:05] Okay.
[01:36:06] Um, but while doing the prediction, let's say, on.
[01:36:11] On the production environment. We have to use one of.
[01:36:14] these five models, right, where it has been trained on a particular.
[01:36:20] Uh… test set and tested, uh, training set and tested on the table.
[01:36:26] So, okay, there are two, three things in here. One, number one,
[01:36:31] We are given a single dataset.
[01:36:35] Usually, which is labeled.
[01:36:36] And we have to decide how much of it
[01:36:40] to keep for training.
[01:36:43] And how much of it to keep for test?
[01:36:46] Okay, that is the first question that we have to…
[01:36:48] Find out. Now,
[01:36:49] Hmm.
[01:36:50] Uh, it can happen that if we just randomly decide how much to put in training and test,
[01:36:56] Then the model that we build may not be optimally the best model.
[01:37:00] So, then we have to scientifically decide how to split the given single dataset into training and test set.
[01:37:06] So, this careful cross-validation, or the knee-based
[01:37:11] method helps us to decide how much of this data, this single data, to keep in training, how much to keep in test.
[01:37:19] Let's say we are using K-Fold cross-validation.
[01:37:22] The second question is that our… after deciding
[01:37:26] the… after, you know, finalizing value of K.
[01:37:30] Let's say we finalize k equal to 5.
[01:37:33] Then, what will be the performance?
[01:37:35] Right? So what will be the performance? The performance will be this.
[01:37:40] Which is the average performance over all the folds?
[01:37:41] Mm-hmm.
[01:37:44] Because we cannot consider the performance
[01:37:47] over any particular fold as the final performance, because
[01:37:51] It is specific to a fold. It is not generic.
[01:37:55] So, we give this as this one.
[01:37:58] The average performance, because…
[01:38:00] It is possible that a certain part of… subset of the data, when used as the test data, might give some very high performance, because this
[01:38:10] test set will be very similar to the rest of the training set.
[01:38:14] As a result, the performance, maybe, say, 95%.
[01:38:17] Whereas a certain section may contain certain samples which are not there.
[01:38:22] Largely not there. In this case,
[01:38:24] This performance may go to 80%.
[01:38:27] Here, it might be a mix, so it might go to, let's say, 85.
[01:38:31] And here, it may go to 90.
[01:38:33] And here, maybe it is again 90.
[01:38:34] Okay.
[01:38:35] Then, which one to use as the model performance? We can't use 95, but we can't use 80 also.
[01:38:41] So, we take the average one.
[01:38:45] Is it okay?
[01:38:49] Okay.
[01:38:50] Okay. Yeah, alright.
[01:38:52] Mm-hmm.
[01:38:53] Uh, ma'am, one question. So, suppose we are on the, uh, step where we need to figure out training and testing, so one way is, uh.
[01:39:00] Yes.
[01:39:01] K-cross board validation. What are the different ways that we can figure it out? Because I'm sure there would be few others as well.
[01:39:06] And what are the parameters on which we take the decision?
[01:39:10] No? Huh.
[01:39:14] Hmm.
[01:39:19] Yeah.
[01:39:20] Because for an example, I'll just add, uh, so it seems that we are sort of creating multiple folds, so I believe it also comes with the cost, right? So, suppose… To avoid the cost, then probably we might need to look for something else. So, yeah, on that aspect.
[01:39:25] Yeah, so now the thing is that, broadly, there are these two methods which I already showed you.
[01:39:30] One, that you… you decide on some split ratio.
[01:39:35] Training test split ratio, and for each of these ratios,
[01:39:39] You find out the performance.
[01:39:42] And then choose Start Training Test Ratio.
[01:39:45] Which gives the optimally best performance.
[01:39:48] Because, uh, or you…
[01:39:52] Actually, these two are finally mapping to the same thing. When you decide the number of folds,
[01:39:56] Then, also, you are deciding the train-test ratio only.
[01:40:00] Only thing is that how many folds…
[01:40:02] We can go on… let's say if there are 100 samples in the training data, K could be 100.
[01:40:08] Right? Which means, worst case is, there's one point in one fold.
[01:40:13] But we don't want that, because we don't want to find out the building of a model for 1-1 point each, right? We cannot have 100 models.
[01:40:22] So, we have to use optimally low value,
[01:40:25] It should be not very low, not very high. So, highest value, if there are 100…
[01:40:30] Total training samples.
[01:40:33] Then K could be 100, and the lowest is…
[01:40:36] So, we are having 100 folds.
[01:40:39] Each containing 1-1 point.
[01:40:41] But we don't want that, right? We don't want our model to be built on one point.
[01:40:45] then it will be overfitting. Or we have…
[01:40:48] K equal to 0, where there are no folds, all of the data is going into training.
[01:40:54] And all of the data is also going into test.
[01:40:58] Then also, it is, uh, you know,
[01:41:00] overly generic, because…
[01:41:02] We don't want that. So, we want a value of k.
[01:41:06] Which is high to some extent. If it is overly high, too many computations will be required.
[01:41:11] Say, if size of training data is 100, and we are keeping value of Ks also equal to 100,
[01:41:17] One model will be built.
[01:41:19] per data point.
[01:41:21] so many models will be built.
[01:41:23] You don't want that, and this is a very small example.
[01:41:26] usually training samples will be in thousands or lakhs, like that.
[01:41:31] So those many models we cannot build that will be very compute-intensive.
[01:41:35] Similarly, we cannot have just one model or two models also.
[01:41:39] So, to save the number of computations, yet also to retain the performance,
[01:41:44] Within a certain threshold that is satisfactory for us,
[01:41:48] We choose careful cross-validation as one alternate.
[01:41:52] Where you decide where you look at the performance versus the number of folds, and…
[01:41:58] Choose the best value of foods.
[01:42:00] But these are the two most popular methods. I'm…
[01:42:03] at least as far as I think.
[01:42:05] I'm not aware of any other method.
[01:42:08] Which is used for dividing the given data into training and test data.
[01:42:12] However, for other hyperparameters, there are other methods also.
[01:42:17] I'm only right now concentrating on dividing the data into training and test set. That's all.
[01:42:22] But there can be many other hyperparameters that we will do as we go along.
[01:42:27] And for them, other methods we'll use.
[01:42:30] It's okay.
[01:42:33] Yeah, on splitting the data into training and tests, the percentage, it can.
[01:42:37] Go to n number of split rate, 40, 60, it could be 30, 70.
[01:42:42] Yeah. Yeah.
[01:42:47] Yes.
[01:42:48] You can… so there could be multiple, uh, scenarios. So, how do we constrain ourselves?
[01:42:49] To, suppose, 4 or 5, so that… There's less competition.
[01:42:54] Yeah, so then, again, for everything, there's a cost and there's some optimality.
[01:42:59] You have to strike a balance. Let's say you have a lot of compute resources available with you.
[01:43:05] Then you go over the whole range,
[01:43:07] And see which gives you the best performance, most optimal performance.
[01:43:12] The other possibility is you are hard on compute.
[01:43:16] Then you have to restrict yourself to a set of probable set of values, and then see out of that which one is best, and choose that only, whether it is
[01:43:24] Final West or not.
[01:43:26] This is how you do.
[01:43:31] Okay, Pallavi, you raised a hand, do you have a question?
[01:43:37] Uh, yes. Uh, so this is not regarding this validation, but… A bit before, I've missed out also.
[01:43:45] Um, can you, like, in what situation may base classifier be better than precision tree or random?
[01:43:55] Yeah, yeah. Sorry.
[01:43:56] Forest. If you have, yeah.
[01:43:57] Yeah, so I actually told you that situation.
[01:44:05] So, for example,
[01:44:06] Uh, we have a set of training points, right? So…
[01:44:12] I gave you while doing the decision tree, let's say you have two samples.
[01:44:18] One is having a refund, no single $70K.
[01:44:22] And let's say there's another sample I'm adding here, which is no single.
[01:44:28] And so, 70K, and here's a yes.
[01:44:31] So, in this case, what is happening? The attribute values are the same, right?
[01:44:37] For record this, and the one that I've added. They are the same.
[01:44:40] But one is having a class labeled no, one is having a yes.
[01:44:44] Now, if you were to draw a decision tree, as we discussed during the decision tree portion,
[01:44:49] It's not possible to draw a decision tree.
[01:44:52] For such a data, where the values of the attributes are the same, but the class labels are different.
[01:44:58] We can't really draw a decision tree.
[01:45:01] But then, how do we actually predict the class label?
[01:45:05] So, to predict the class table, here we use probabilities.
[01:45:09] And see what is the probability for a no versus what is the probability for a yes.
[01:45:15] And accordingly, predict the class label based on whichever probability is higher.
[01:45:19] So this is a very, um, typical use case.
[01:45:23] Where attribute value pair combinations are not unique.
[01:45:28] Right? Not unique means?
[01:45:30] It's not a unique means that for a certain value per combination like this,
[01:45:36] There will be only one particular class label associated with it.
[01:45:40] But in this case, for the same value pair combination, both the class labels are associated, so value pair combinations of this kind
[01:45:48] are not unique to… in association with a certain class label.
[01:45:53] Right? Because multiple class labels
[01:45:56] are being associated with these set… same set of values.
[01:46:00] then how do we infer whether to…
[01:46:02] The class label is no or yes, because the values of the attribute are the same.
[01:46:07] So, in that case, we look at the probability of occurrence of the class labels given these values.
[01:46:12] And then use that probability…
[01:46:14] Which is calculated over the entire training data.
[01:46:18] To find out the class link.
[01:46:20] So these are… this is the use case.
[01:46:22] Is it okay, Palla?
[01:46:26] Um, okay, yeah.
[01:46:28] So, value pair combinations, if they are all unique.
[01:46:32] Then you don't necessarily used to… may use prob- uh, probabilities. You can use decision tree random forest, and no.
[01:46:40] But, uh, in many cases, the values of the attributes are not unique.
[01:46:47] Two, in association with the class label for the same values,
[01:46:51] There can be class labels C1 also in some training data… in some training sample.
[01:46:56] And certain other training sample for the same attribute value, it could be class label C2 also.
[01:47:02] So, uh, in such cases, it's better to use a knife.
[01:47:10] Okay, any other questions, anyone?
[01:47:19] Okay, so the next topic that I'll be, uh, beginning will be, uh, regression, and I want to begin it on a…
[01:47:27] Uh, I'll do a stop share here. I want to begin it afresh on a fresh turn.
[01:47:32] But before we break for today, uh…
[01:47:36] Uh, what I want to say, uh, indicate to you was that in the hands-on session,
[01:47:42] The way that we follow is that we pre-prepare a…
[01:47:46] quote for you, and in the hands-on,
[01:47:50] There could be actually two ways.
[01:47:52] One way is that you write the code,
[01:47:56] Uh, at the time of the session, then you write it
[01:48:00] Line by line, and then you complete it, and then you compile it, run it, and see.
[01:48:06] That can be one way, but given the complexity of this course, and we have a lot of concepts to cover,
[01:48:13] So, and this is not a programming course, we…
[01:48:16] we assume that programming is known to you.
[01:48:18] So, therefore, if we do it that way,
[01:48:21] Then, even for a single topic, which are, you know, the ML topics that we started out with, it might take multiple turns to do even one concept.
[01:48:29] So, therefore, we do all the exercise, and…
[01:48:32] background homework of preparing the code.
[01:48:38] And providing the code,
[01:48:40] the help file to the code of what we are doing.
[01:48:43] And the pre-run copy of it, along with a PDF containing the…
[01:48:48] output.
[01:48:50] for you to refer to everything ready for you, so that in the…
[01:48:55] live session, we can maximize…
[01:48:59] Uh, the transfer of knowledge, and you can…
[01:49:03] go through the code, and then ask…
[01:49:05] what, uh, you know, what is not understandable, which particular line of code, or what instruction, or what function.
[01:49:14] or library, or whatever you want to know.
[01:49:16] That you can actually ask, instead of…
[01:49:19] Zeroing down on writing one line of code at a time, and then wasting time.
[01:49:25] Uh, because most of you already know Python, or you would have learned it as a bridge course before starting this course.
[01:49:32] So we don't want to waste your time, we want to… it to be most fruitful.
[01:49:37] So I hope it is okay with everyone.
[01:49:39] Uh, because otherwise, we just spend time writing the code here and never finishing it up, even for…
[01:49:46] one classifier or one concept.
[01:49:49] Is it okay, everyone? Any…
[01:49:51] Any inputs or any thoughts or anything else that you want to say?
[01:49:55] On it.
[01:49:58] Yeah, that's fine, ma'am.
[01:49:59] Yeah, ma'am, I have an input, actually. So if that file that is going to be shared during the session will be shared in advance, then we can actually go through that code in advance, right? And then we can discuss upon it.
[01:50:10] During the session, we are actually going through line… code line by line, right? So… If we are knowing it in advance, then we can actually ask doubt during the session as well.
[01:50:20] Okay, so, uh…
[01:50:23] It's provided, uh, I think, a day in advance or so, but we can provide it more in… more in advance.
[01:50:30] For those who are interested to go through, you can go through. We can do that, no issues on that.
[01:50:38] So, any other inputs or anything?
[01:50:42] Uh, because as we go along, the code will become more and more complex.
[01:50:46] And, uh, it will, uh, the… and the length of the code will also increase a lot. So, we cannot really…
[01:50:53] afford to waste our time writing
[01:50:55] code in real time here.
[01:50:58] Uh, and for the more complex deep learning models, actually, you'll have to go to PyTorch also.
[01:51:04] And I hope all of you are familiar with PyTorch, and if you are not, please go ahead and…
[01:51:09] Um, just to…
[01:51:11] You don't need to know…
[01:51:13] like, any specific complex thing?
[01:51:16] But you need to know how to use Python and PyTorch.
[01:51:19] So, we'll be doing all the DL models in PyTorch as they become more and more complex, so…
[01:51:24] That also you should be familiar with.
[01:51:27] Uh, and, uh…
[01:51:30] There are sessions on… in the previous patches for those, uh…
[01:51:34] Learners who are not familiar, they were arranged also, some queue sessions.
[01:51:39] So, that also can be done if some of you would want. You can let us know.
[01:51:47] We'll be doing DL in this month, or it will be done in next month?
[01:51:51] No, I think the curriculum is shared with you, so this month you won't be doing, because actually in DL, everything will be
[01:51:58] Uh, based on supervised learning only.
[01:51:59] Okay.
[01:52:01] So, you should know what is supervised learning.
[01:52:02] Okay.
[01:52:03] So, first, I'll be taking you through the concepts of supervised learning.
[01:52:07] Supervised learning. Then, I'll be taking you through…
[01:52:12] some basic concepts of NLP.
[01:52:17] deep learning, and then NLP. So we'll progress like this, that we are having the basic ML-supervised learning primarily.
[01:52:24] Right now. After that, you will be going to the deep learning models.
[01:52:29] Then, after the deep learning models, I'll take you through NLP concepts, the basic concepts, and then we'll start with the advanced
[01:52:36] concepts. And more or less, the outline is the one that is given in the curriculum.
[01:52:42] The dates may vary slightly based on some national holidays,
[01:52:47] or let's say that it is possible that in certain portion,
[01:52:51] uh, certain batch takes more time.
[01:52:54] or less time. So, depending on the pace, also, it varies a little bit, plus or minus.
[01:52:58] And all that. And there might be some doubt resolution sessions also, which I usually keep based on the request of the learners that
[01:53:06] Okay, after some n number of sessions,
[01:53:08] You all want to have some doubt resolution session.
[01:53:11] So, in such cases also, I keep doubt resolution sessions.
[01:53:15] So, in between, if you have any queries,
[01:53:18] Uh, because sometimes you may not go through the, uh, the portion.
[01:53:22] In that week, you might take it next week or so. So…
[01:53:25] to resolve this, we also have this. So, it varies a little, plus-minus.
[01:53:30] But broadly the same flow.
[01:53:32] And the timeline will be followed.
[01:53:38] Any other questions, anyone?
[01:53:40] Before we break water down.
[01:53:41] Will there be any, uh, assignments or something given, like, because, like, it's like, we are just learning theory and, uh…
[01:53:47] Uh, it might be that if you can give some… some examples that we can try out ourselves.
[01:53:52] And, uh, come up with some doubts, or whatever we face.
[01:53:58] So, what we have, uh, done for the previous batches that usually…
[01:54:05] based on the theory,
[01:54:07] Some questions we give.
[01:54:09] that is based on that theory that they have learned.
[01:54:12] And the hands-on covers the…
[01:54:15] kind of thing that you should know in terms of coding.
[01:54:18] However, if you wish to have coding assignments,
[01:54:21] I can give you, but it will be hard to discuss the coding assignment here in class.
[01:54:27] That can be a pure self-study kind of a thing.
[01:54:30] Why? Because if I discussed that, then the portions that we have to cover in theory and hands-on.
[01:54:36] As such, the course has a vast amount of
[01:54:39] coverage that is required, right? It's huge. The curriculum is very…
[01:54:45] very large. So, if we spend time over discussing of…
[01:54:51] projects or assignments, quoting assignments that I give you.
[01:54:54] That will make it a little difficult. If you wish, I can give you, but that have to be purely self-study in the sense that
[01:55:02] You can go through and try to…
[01:55:04] solve it, and then maybe you can discuss yourselves or with some other friends and see.
[01:55:08] the solution, something like that.
[01:55:10] If that works for you.
[01:55:11] Yeah, that can be also helpful, yeah.
[01:55:12] Yeah, and then maybe in the doubt clearing class, we can… if we have any specific doubts, we can…
[01:55:17] Uh, ask.
[01:55:20] Okay, so that can be done, that some toy data along with some…
[01:55:24] Yeah, something.
[01:55:25] problem, assignment, you're coding example can be given that you can do on your own.
[01:55:33] to just have an assessment of what you're learning, or something like that.
[01:55:37] Okay, we can do that.
[01:55:41] Yeah, we're on the same topic. Uh, I think, uh, I also have one suggestion.
[01:55:46] So, whatever we discuss in the class, we sort of… Uh, understand it, but can't go into nitty-gritty details of it. So, after the session, if we could sort of share.
[01:55:57] Uh, what do we need to read after it, and what are the topics? Probably give some sort of assignment that.
[01:56:04] Probably help us visualizing, like, what we have learned. If there are any gaps, we can ask as a doubt, or we can sort of.
[01:56:11] Because we understand the topic, but, uh, to make a.
[01:56:16] Uh, you know, to make a connection, what we have learned so far, so that's really hard.
[01:56:21] Supposing our real life, if we sort of encounter with some use case.
[01:56:24] How do we deal with it? So, that clarity is still, I believe.
[01:56:30] I'm not able to, uh, sort of learn that so far.
[01:56:33] So that would help. If we have some sort of assignment, study material, some sort of case study that you could provide, that we can go by ourself.
[01:56:41] Uh, so I think, uh…
[01:56:43] Uh, there are a lot of material is also shared by…
[01:56:48] Simran. Simran, I hope you are here. Can you please add onto it that contains…
[01:56:54] theory and some, uh, use cases, or typical use cases of the theory that you do.
[01:57:01] Uh, I think Simran will be able to comment on that. I hope you all are getting, because she…
[01:57:07] I review it and…
[01:57:10] We share, that is one thing.
[01:57:12] Secondly, uh, related to, uh, mapping with some real-world situation.
[01:57:18] So, what we can do is, uh, though, because you are early on in this course,
[01:57:23] I know in the other batches, people start…
[01:57:26] mapping it to what they are encountering at their workplace.
[01:57:30] Uh, but what we can do is that we can, uh…
[01:57:34] You know, during the doubt resolution session, or otherwise.
[01:57:38] Uh, you can think of, uh…
[01:57:40] your, uh, you know, at your workplace, what kind of problems you are having.
[01:57:45] In terms of, uh, of this particular course.
[01:57:49] And such kind of examples can be discussed. However,
[01:57:53] What you need to read extra is purely dependent on you.
[01:57:57] And let me caution you, because if you go on online, or if I give you some… I already gave you a list of books,
[01:58:04] And otherwise, also, you can find a lot of material and a lot of use cases.
[01:58:09] of whatever each topic, whatever I teach to you.
[01:58:13] And if we, uh, if you, you know,
[01:58:16] Uh, week by week, if you keep on delving,
[01:58:20] Too much into that, I think you'll get more confused and lost. This is based on my experience with…
[01:58:25] Over-teaching so many batches, and also teaching at IIT.
[01:58:29] that too much material also makes you feel confused and lost, because then you don't know
[01:58:35] You are actually just thinking what to do and what to… which dot to connect to which.
[01:58:42] So, better is that… and, uh, you know, there are so many things you practically get
[01:58:47] you know, you think of reading something you don't understand, you read something related to it, then something, something.
[01:58:53] And the main focus area is gone.
[01:58:56] So, I would not recommend you
[01:58:58] to, uh, read too much.
[01:59:00] Until unless, you know, you know that you have that much of time, and…
[01:59:06] All those things. Given the fact that you all are working…
[01:59:09] And time constraints are always there.
[01:59:12] Therefore, I will like you to focus on what I'm covering, because it's going to be useful in our future
[01:59:18] Uh, everything is actually sequentially related, so it will be related. If you know this concept, it will be related to…
[01:59:25] The next… next set of things in deep learning, in…
[01:59:30] Generative AI and all that.
[01:59:31] But self-study on…
[01:59:34] learning more topics actually has no…
[01:59:36] limit to it.
[01:59:39] And it has a high probability of you getting lost, so…
[01:59:43] Uh, I think I'd like to refrain from that. I already gave you the list of books.
[01:59:49] If you wish, lot of material is there.
[01:59:52] Based on your tying,
[01:59:54] And your interest, you can read.
[01:59:58] So, therefore, Deepak, it will be difficult for me to tell you you need this more and that more.
[02:00:04] Okay?
[02:00:07] Yeah. However, as I think somebody else also pointed out, that…
[02:00:11] Some of you want to exercise some coding based on the concept done.
[02:00:17] that kind of questions or assignments that we can do. They'll be ungraded.
[02:00:22] Not everyone who was… whosoever wants to do it can do it.
[02:00:27] And it is only for self-study and self-exercise, and that's all.
[02:00:31] And there's no pressure, uh, that, okay, if I don't do, then I'm not going to get that much of marks and all that will not be there, given that all of you are already
[02:00:41] pressured by being in…
[02:00:44] a job at doing this course, which is in itself quite heavy.
[02:00:51] Anyone else, any other comments, anything? Any suggestions?
[02:01:02] Okay, if there are no further questions or comments, then we can wrap up for today.
[02:01:03] Yep. Yes.
[02:01:09] Thank you all, have a great day. Thank you, bye-bye.
[02:01:16] Thank you.
[02:01:22] Thank you