# 01 2026-01-24 Introduction to Deep Learning course: Module 3 — Deep Learning & NLP module: Module-3-Deep-Learning-NLP date: 2026-01-24 type: transcript video_url: https://personal-learn.armco.dev/files/_Recordings/Module-3-Deep-Learning-NLP/01_2026-01-24_Introduction_to_Deep_Learning.mp4 --- [00:53:45] that the predicted value is very much close to the actual expected value. [00:53:50] Okay. And here, uh… [00:53:54] The input is flowing in the forward direction. [00:53:57] So, we have the input, then… [00:53:59] This goes in the forward direction. In the forward direction, like this. [00:54:03] So, the signal is flowing from the left side to the right side in the neural network. [00:54:10] And, uh, then the final outcome can be calculated. [00:54:15] So now, how do we get, um… [00:54:17] the outcome. So, first of all, uh, what we should, uh, understand is that we have something called the activation function, which is represented as the F here. [00:54:28] And, uh, if we just took the input, X1, X2, XC till XN, and did a… [00:54:35] product of these with the weights, and added them all together. [00:54:39] And then calculated the output. [00:54:43] Then what we would get is just a linear summation. [00:54:45] Uh, of, uh, the inputs. [00:54:49] However, what we want is, uh, not just a linear combination, because… [00:54:54] A linear combination renders the output [00:54:58] To be quite, uh, uh, you know, static kind of a thing. [00:55:03] We want to understand the nature of our data could be non-linear. [00:55:08] And therefore, what we want to have is a nonlinear representation. [00:55:12] And to have that kind of nonlinear representation, what we have is called the activation function. [00:55:20] Which is represented as the F here. [00:55:23] Okay, so we have the activation function. [00:55:27] The, uh… the inputs are multiplied by the weights, and the summation of these is fed into the activation function. [00:55:34] And this activation function is what? [00:55:37] Uh, adds the non-linearity. [00:55:39] to the, uh, given, uh, input data. Otherwise, uh, without this, it would have been just a linear combination. [00:55:47] And then, uh, this, uh, after having… [00:55:51] Uh, being fed into the activation function, what we get is the output, which is a non-linear. [00:55:56] combination of the weighted sum of the inputs. [00:56:00] Okay, I can see some hands raised. What questions do you have? [00:56:06] You can ask me in the order in which you raised the hand. [00:56:08] Uh, I'd like to go first, ma'am, Vinit here. [00:56:11] Mm-hmm. [00:56:12] So, what is the importance of this W1, W2? I think you used the word weighted or something, if I heard correctly. [00:56:18] Yes. [00:56:21] Yes. [00:56:22] Uh, just keeping the other previous lessons in mind, right, where we again had the same layers where we used to give inputs, and it used to read, and it used to come up, right? Dusted learnings, and… [00:56:33] those things. So, what does the weighted part add value over here, or… I don't know, is this just a… [00:56:40] A constant that is used, and we might be talking about it later. [00:56:44] Yeah, so the magic of deep learning… [00:56:48] is dependent on weights. [00:56:49] So, what we want here is that… [00:56:52] What we want is that we have a combination of inputs, [00:56:56] And given these inputs, we want the outcome [00:56:59] To be as close as to the expected outcome. For example, if we have a set of images, [00:57:05] And they are representing some, uh, safe routes. [00:57:08] say, apples, oranges, and bananas. [00:57:11] So, what we want is that… [00:57:13] Whatever set of images we are giving, the prediction made on them is… [00:57:17] Correct. It should not happen that apples are predicted as bananas, and bananas are predicted as oranges, something like that, right? [00:57:26] Uh, now, uh, to make that prediction in a proper way. [00:57:30] Uh, what we… so, what we do is that we multiply the different images [00:57:36] I mean, the different inputs, [00:57:38] With weights. Okay. [00:57:39] Okay. [00:57:40] And then we try to combine these inputs. [00:57:44] Uh, which are multiplied with weights. [00:57:47] And then, I'll just explain to you, uh, how it works. [00:57:51] And, uh, we want this combination to be given as an input to this deep learning model in such a way [00:58:00] that the output is as close as to the input. [00:58:03] Or the error is minimized. This is what we want. [00:58:06] So, this is why we… and… [00:58:10] Actually, if we just give the inputs, we don't have anything to worry, right? [00:58:13] Let's say I just take the input 61X2, XC till XR, and give them… [00:58:17] to the model. [00:58:19] And then the model is going to just predict [00:58:22] something based on the input received, but then what can I vary to reduce the error? [00:58:28] We don't have much to vary. [00:58:30] We want something with which we can play so that [00:58:34] We can reduce the error. So, we have induced these weights, [00:58:38] We are going to vary these weights in such a way. [00:58:40] that the output is much close to the… [00:58:44] expected output, and the error is reduced. [00:58:48] So, this is a very broad overview of what weights… why weights are useful. [00:58:53] However, I'll take in a lot of details how these weights are chosen and what role do they play. [00:58:59] stoking? [00:59:00] Okay, and just one last question around that, ma'am. So, when I get the final output, uh, does it… [00:59:07] say that, uh, this particular variable is, uh, this was the impact of it based on this weight, or… [00:59:13] Uh, like, in my final output, will the WN come out in any way, or…? [00:59:18] That'll only be when we start changing the values to see the differences. [00:59:23] Yeah, so actually, when we… [00:59:25] So, the outcome is just the outcome. Let's say I want to predict [00:59:29] From a set of images, what are apples and what are oranges? [00:59:34] Correct. [00:59:35] Initially, my images were 150, 50 were apples and oranges. [00:59:39] In the outcome, I got maybe a mix of 30-70. [00:59:43] Okay. [00:59:44] So, what I get is just the outcome. [00:59:45] And so, I will not get any weight or anything. I'll get the outcome. [00:59:50] Okay. [00:59:51] Now, the thing is that the error on the outcome is high, because it was supposed to be 50-50, now it is 30-70. [00:59:57] Okay. [00:59:58] So, we want to reduce this error, so what we will do is, we will… [01:00:01] I'll explain to you, uh, we'll back-propagate the error. [01:00:05] And then revise the weights. [01:00:08] Okay. [01:00:09] And then, after revising the weights, do this, uh, process once again, combining in weighted sum of inputs, taking it through the activation, [01:00:18] And then, once again, do the prediction. [01:00:20] Now, let's say from 3070, it has moved to 4060. [01:00:23] So our result has improved. [01:00:26] So, again, what we will do is back-propagate the error once again, change the widths, and then again, see what is happening. [01:00:33] It is… there is a possibility that by doing all this process, again, iteratively, [01:00:37] Again and again, we'll get… [01:00:39] quite close to the expected output, say 50-50 was the input. [01:00:43] Maybe we get 45, 55. So our error is much reduced now. [01:00:47] So, this is the whole motive that we keep on defining the weights. [01:00:51] So that we can actually do… we have a control over the inputs, right? If we… I just give the input, I don't have any control. [01:00:59] What I'll get is a static output, and I cannot refine it. [01:01:03] But suppose I want to refine it, [01:01:06] Then, in that particular case, [01:01:07] I should have something to play around with. [01:01:10] In a way, weights are doing this, but it's not just weights, it's also activation function. [01:01:16] There are hyperparameters, all that I will, uh, explain to you as we go along. [01:01:22] Right now, I'm just explaining to you some basics. [01:01:24] Okay? [01:01:25] Okay, got it. Thank you. [01:01:26] Yeah. Who was the second one to raise the hand, followed by Deepak? [01:01:36] Okay. Okay, Deepak, what question do you have? [01:01:37] Oh, ma'am, my question was answered, yeah. [01:01:38] Yeah, just one quick question. So, we earlier discussed about the MH, and you have given some example. [01:01:44] Mm-hmm. [01:01:45] So, what would X1, X2, X3 would be if we take the example of that image? [01:01:51] Yeah. [01:01:53] So, X1, X2, X3, again, uh… [01:01:54] Uh, that we have discussed earlier. [01:01:55] I'll just explain to you in the next few slides, it could mean a couple of things. [01:02:00] It could either mean attributes, or it could mean images themselves. [01:02:04] So, I'll explain both of these, uh, what, uh, both of these mean. [01:02:05] Sure. [01:02:10] I'll just explain to you, just allow me a few minutes, and it will be more clear to you. [01:02:16] And just one more thing, is the activation function also called transformation function, something like that? [01:02:21] So, the activation function helps to transform the input. [01:02:26] in, uh, you know, it helps to map it to a… [01:02:29] Uh, to a, you know, different dimension set. [01:02:32] Or, uh, to a different way in which it is perceived. But it is not directly called transformation function. [01:02:38] It is just activation function, which is the standard terminology for it. [01:02:45] Okay. So now, let's see what is happening here. We have these inputs generally represented as X. [01:02:53] And then we have the weights, which are represented as W. [01:02:56] So, we have X… [01:02:58] And then we have these weights, W, we are multiplying these weights and applying a function over it. [01:03:03] And this is the output that we are obtaining. [01:03:06] Now, uh, how… [01:03:08] How, uh, more can we explain this? So, I'll just explain to you. [01:03:15] So then, uh, since we have… [01:03:18] these inputs as X1, X2… [01:03:22] X3… so on and so forth, till XN. [01:03:26] So, we can generally write any input as Xi. [01:03:29] So, we have an input XI. [01:03:32] And then we have these weights, W1, W2, W3 till WN. [01:03:36] So, we can generally write… [01:03:40] VATSWR. So, XI is multiplied by WI. [01:03:43] And then, all these are summed together. [01:03:47] So, we have the summation, and how many total we have, and in number of inputs, so I is equal to 1 to n. [01:03:53] So, this is the summation. [01:03:55] And all these are fed to… [01:03:57] processing unit, which does nothing but uses this function over them. So, we have this activation function. [01:04:04] applying over them. So this is what is happening, and what is… we are getting is the Z. [01:04:09] Which helps us to make the prediction of Y. [01:04:14] Now, let's try to understand the dimensionality of each of these. [01:04:19] So, we have, uh… [01:04:21] Here, the input which is in the form 1 cross n, [01:04:25] Okay, so this is 1 cross in. [01:04:28] Then we have the weights. [01:04:30] Which are N cross 1. [01:04:32] And then we have the output, which is… [01:04:35] Uh, one cross one. So, this is actually… [01:04:38] Uh, if you look at it, this input is of the form 1 cross n. [01:04:43] Then weights are of the form N cross 1. [01:04:46] And when we multiply these two, [01:04:50] What we will get is 1 cross 1, obviously. [01:04:54] Now, what implications does this have? I'll explain it to you. [01:04:58] Uh… [01:05:01] So, uh, actually, if we look at each of these inputs, [01:05:06] As for this, the input takes up the dimension D is the dimension, [01:05:11] So, input, if we talk about the input 1, input 1, [01:05:14] As the dimension 1 cross in? [01:05:19] Just a second, my display is actually giving… [01:05:22] digital pad is giving me some problems here. [01:05:25] So… so we have the input. [01:05:29] And this input is actually belonging to… [01:05:32] The dimension 1 cross in. [01:05:35] So, what could it mean, actually? [01:05:39] Uh, so the input can be thought of as, uh… [01:05:43] Here, uh, one refers to rho, and n refers to the column. [01:05:47] So, our input 1, actually, [01:05:50] could be organized in form of… [01:05:52] Uh, vectors having… [01:05:54] And number of attributes. [01:05:56] So, each input… [01:05:58] Or your HX1. [01:06:00] If X1 is the input, [01:06:02] could be represented by this combination of A1, A2, A3 till AN. [01:06:09] So these are, uh, the attributes and number of attributes, or generally speaking, XI. [01:06:16] who take up the firm. [01:06:18] A1, A2, A3 till A. [01:06:21] So, this will be having values like… [01:06:24] AI1, AI2… [01:06:26] And AIN, like that. [01:06:29] Okay? And then, each XI would be multiplied by a weight and cross 1. So, what is this having? [01:06:35] This is having a dimensionality of 1 cross n. [01:06:40] And then each of the weights would be actually… [01:06:43] and cross 1. So if we have weight I, [01:06:45] Then, it is a N cross 1. [01:06:48] So, what is it having? N number of rows, and one column. [01:06:53] So, we can have this, like, [01:06:56] This is your row 1, row 2, row 3. [01:07:00] Like that, till row 1. If this is… [01:07:03] Uh, having a dimensionality and cross one. [01:07:08] And, uh, then, uh… [01:07:10] So, this is another thing. Now, what we… [01:07:13] want to do is that we want to… [01:07:16] Multiply these two, so when we multiply Xi and… [01:07:20] So this is the dot product. [01:07:22] So, what will it represent? It will represent a dimensionality of… [01:07:28] Uh, one cross in… [01:07:30] times dimensionality of N cross 1. [01:07:33] Which is the dimensionality of 1 croissant, which is nothing but your output. [01:07:37] And this output will be these fades. [01:07:40] Alright, so it will be a product of, uh, let's say this is containing R1, R2, R3 till RN. [01:07:47] Then, what will be the product of these two? [01:07:51] I'm just representing this product here. [01:07:53] So, it will give AI1… [01:07:57] dot R1 plus… [01:07:59] AI1? [01:08:01] dot R2. [01:08:03] Thus, I3.r3, and so on and so forth, till AIN. [01:08:09] dot RN. So, this will be the output, which will be fed. [01:08:14] into the, uh… [01:08:16] activation function. Now, this is a visualization, or this is a representation, where I'm saying… [01:08:24] that each input [01:08:26] is being represented by a… [01:08:29] Uh, 1 cross n vector. [01:08:31] Our real-world, uh, inputs are usually represented… may be usually represented, or… [01:08:38] Complex data may be represented by vectors, which may be n-dimensional in nature. [01:08:44] However, you can also think of this… [01:08:47] As a single numbers. [01:08:50] So, both could be valid. [01:08:53] Lear to my input, each input is just a single number. [01:08:56] Right? I can think of it as… [01:08:59] let's say different attributes. [01:09:01] Say, you are talking about house… [01:09:04] house prices, so maybe age of the house, maybe one attribute? [01:09:08] Then maybe location? [01:09:11] Maybe number of, uh, bedrooms? [01:09:16] Uh, then, uh… [01:09:18] number of, maybe, uh… [01:09:20] the… maybe the amenities that are available. [01:09:26] And so on and so forth. So each of these is a number, say age may be something, say… [01:09:31] 10 years. Location, maybe some particular area. [01:09:35] And so on and so forth. So here, [01:09:37] The inputs are not comprised of vectors, they are just single numbers. [01:09:43] Right, so you can think of it in any of these ways. [01:09:47] So, you can think of it as single numbers also. [01:09:50] In which case, you are just having… [01:09:52] Here, it's a 1 cross n dimensionality, so X is having a dimensionality of 1 cross n. [01:09:58] So, our inputs are what? [01:10:01] Uh, they are like X1, X2, X3 till XN. [01:10:06] Single numbers, then you are having weights. [01:10:10] that have a dimensionality of N cross 1. [01:10:13] So you have W1, W2, W3 till WN. [01:10:17] And then, this gives us the output. [01:10:19] Which is Z, which is a 1 cross N. 1. [01:10:23] So, both are, uh, just… [01:10:27] representations. [01:10:29] of a similar concept. [01:10:31] Okay, I can see your hand raised, uh, who is it? Can you just ask? [01:10:36] The question? [01:10:40] Is there any question, anyone? [01:10:43] Neeraj, do you have a question? [01:10:44] Uh, yes ma'am. Sorry to disturb you, ma'am. [01:10:48] Yeah, no problem. [01:10:49] But, you know, actually, uh, whatever the data goes in the… as an input, it will… it will be in the form of metrics, or… [01:10:58] Yeah, usually we think of it as a metrics only. [01:11:02] Uh, if… if you're talking about single numbers, like X1, X2, X3, are single numbers, say… [01:11:08] Age is one attribute, location is one attribute, so actually, they are not metrics. [01:11:13] Uh, but together, they can make up one matrix, right? So this together, all the inputs together are a matrix. [01:11:21] But in the previous case, where I said that [01:11:24] Each input in itself is comprised of so many features. [01:11:28] Then in that case, each input is a vector, and all of them together make up a matrix. [01:11:33] Is it okay? [01:11:35] Uh, ma'am, but what will be in the form of images? Suppose that, uh, I have images just like a banana, I want to compare with, uh, just like that, uh… [01:11:44] It is, uh… [01:11:48] What is? [01:11:49] It is, uh, actually, it is a banana, or it is… [01:11:50] Orange. [01:11:53] But what would be the form of images, man? Uh, what would be the… [01:11:58] Suppose that if, uh, input these images… [01:12:02] Mm-hmm. [01:12:03] I didn't want to… [01:12:04] So, what will be the input like? That's what is your question, right? [01:12:06] Yes. Yes, ma'am. [01:12:09] Okay, let me see if I have space… [01:12:15] Okay, sure… [01:12:19] So, actually, images are represented in form of numbers. [01:12:25] So this image, let's say this is an image of some object. [01:12:30] Okay, we are having, say, an apple, or maybe we are having an orange, or something we are having. [01:12:34] Okay? So, you already may be knowing that, actually, image is actually represented like a grid. [01:12:42] And these… these grid is having a lot of cells. [01:12:47] And based on the resolution, [01:12:50] The number of cells increase. [01:12:51] This is called the resolution of an image, and that's why when we buy a cell phone, we look at [01:12:57] High resolution camera. [01:13:00] or the number of pixels that are represented. [01:13:02] Okay. Now, these… [01:13:04] Each of these is a pixel. [01:13:06] Which is a small area on the image. [01:13:09] And these pixel, when we are actually representing the image, then the color, let's say it's a grayscale image, let's say. [01:13:17] Then, each… [01:13:19] Pixel will be represented by a number. [01:13:23] Say number 51. [01:13:25] So, this number is actually indicating the amount of grayness. Then the next pixel will have another number. [01:13:31] The third one will have another number, like that. [01:13:35] And each of these numbers [01:13:36] We present the amount of blackness covered in the… contained in that particular pixel. [01:13:42] Okay, so probably the orange will have… [01:13:45] a boundary, and these… this boundary will be represented by dark pixels like this. [01:13:53] Okay, these may be dark black colored. The boundary of the… [01:13:57] image. Whereas all other pixels may be lighter colored. [01:14:01] And maybe the interior may be contained by a lighter shade of gray. [01:14:06] So, it will have some other pixel values, and these together will make up [01:14:12] Uh, the object, which is the orange. [01:14:14] So, for the purpose of our algorithms, any algorithm, [01:14:19] This will be represented. [01:14:22] In form of a vector. [01:14:24] Pixel 1, pixel 2, pixel 3. [01:14:26] Then, let's say if there are so many pixels, then… [01:14:30] Pixel 256, and each pixel will be having some value, say… [01:14:35] 21, 51, so on and so forth, 9, 91, whatever is that value of the grayness. [01:14:43] So, this is how the, uh, image is represented, because [01:14:46] Any computer system, it cannot understand anything other than numbers. [01:14:51] So, we need to represent it in numbers. Now… [01:14:55] This is our input. So, this may be our input 1. [01:14:58] Another image will be… [01:15:01] So… [01:15:04] An image will be represented similarly as input 2, and so on and so forth. [01:15:08] Okay, so these are… these are nothing but vectors. [01:15:11] And this is… these are the vectors that I talked about. So this… [01:15:15] One image will be one vector, another image will be another vector, third image may be third… [01:15:21] vector, and all of these are inputted here. So, this is one image. [01:15:24] So this is one vector, this is another image, another vector, like that. Is it okay? [01:15:31] Uh, yes ma'am. Thank you. [01:15:40] Okay, Aditya, what question do you have? [01:15:42] But ma'am, um, in this, one more thing occurs, like, uh, if there is a wrong data, or, like… [01:15:50] Uh, you explained it, uh, that it will convert it in the vectors, right? [01:15:56] Rather numbers. So, uh… [01:16:00] Like, there are multiple images of banana. [01:16:03] Mm-hmm. [01:16:04] Right? In that, uh, um… [01:16:06] Mary? Yeah, yeah, please go on. [01:16:08] Yeah, yeah, it's okay. Uh, so… [01:16:13] So what I'm asking is, uh, that… [01:16:16] Sometimes, uh, it may read in a wrong way also, or we can… [01:16:22] give a wrong input, uh, wrong learning to the AI, or deep learning to… [01:16:28] or a model, uh, whichever we want them to learn. [01:16:34] from our data, right? So, if they learn it in a wrong way, then the output will be also wrong. Like, if I… [01:16:35] Sorry. [01:16:42] Give them the vectors like banana, uh, the vector to an orange. [01:16:47] then it will be, uh, the output will be wrong, right? [01:16:52] They'll… by their learning? [01:16:53] Yeah. [01:16:55] Because it… yeah. [01:16:56] Yeah, yeah, so what… haha, go on. [01:16:58] Because they always, uh, learn from our, uh, inputs. [01:17:02] They always, uh… [01:17:04] Even the EIA, or right now, whatever AI are there, they always learn from our, uh, images, or whatever inputs, and then… [01:17:13] They give us the output. [01:17:18] So, what you're saying is exactly correct, and this is what is the example I gave earlier also. [01:17:25] That, let's say that there is an animal-like cat. [01:17:27] Okay. And the algorithm has no other mechanism but to learn by what it gets. [01:17:34] So, let's say there are images of cats. [01:17:35] Now, uh, what happens is that, let's say there are images of cats and dogs. [01:17:41] Now, it can always happen that because of some fuzziness in some cat image, [01:17:46] Uh, the algorithm will, uh, uh, you know, mistake it for a dog. [01:17:51] Because, let's say the… [01:17:53] Uh, you know, the image was in such a fashion. [01:17:56] Uh, that, uh, certain features were not getting captured correctly. [01:18:02] And so what you are saying is that whatever learning, if it is, let's say, image data, [01:18:07] And if there are certain images which are fuzzy, and which are not clearly identifiable, [01:18:13] Then the learning done also can be wrong. [01:18:16] And that's correct, actually, because… [01:18:19] Uh, only humans can be very precise and… [01:18:22] good at, uh, prediction or learning and all, but algorithms actually have to be… [01:18:27] Feeding on what they get. If what they feed on is [01:18:31] non, uh, is fuzzy. [01:18:32] Am I right? [01:18:34] Then the answer can be wrong. [01:18:37] So then, how we can improve it? Uh, means… [01:18:40] What is the way to improve, or what is the correct way? [01:18:44] So the, uh… so the thing is that that's what the deep learning is all about. You will be… [01:18:50] Learning more and more complex models which help to [01:18:54] make the learning more and more better. [01:18:58] See, if we think about it, so far, we discussed all the simple machine learning models which are based on shallow learning. [01:19:05] So, uh, what I'm trying to build for all of you are a stack of algorithms where you go more and more complex to improve the learning only. [01:19:14] So, what can you do, like, for example? [01:19:16] If you were to use, uh, you know, uh, simple neural network, [01:19:21] Uh, then you will be able to… [01:19:25] opt-in, let's say it is having some deep layers in it, you'll be able to obtain better [01:19:30] Um, results, better prediction on the image. [01:19:33] Then, if you want it still better, then you have a more complex model. [01:19:37] Say you have, uh, let's say that… [01:19:40] You are using CNN model. Then in CNN, there are so many architectures, [01:19:45] And then you learn some specific use to make it better, use some specific architecture, like that. [01:19:52] So you're making the models more and more complex. [01:19:55] So that, uh, the learning that they do are much more and more better. [01:19:56] Okay. [01:20:00] So, this is how we can improve. However, obviously, [01:20:03] There's still a lot of scope, and there are no foolproof [01:20:06] algorithms. However, the state-of-the-art generative AI models [01:20:12] They have a very good, um… [01:20:14] Uh, you know, the very good prediction, and… [01:20:17] Uh, very good, uh, uh, representation of the knowledge. [01:20:22] So, this is how we are going from machine learning to AI. [01:20:26] Uh, to deep learning, to generative AI, [01:20:29] And then making use of those generative AI in agentic AI. So, this is how it is improving. [01:20:35] Okay. Thank you, man. [01:20:37] Okay? [01:20:38] Yeah. Okay, Ankit, and then Abhishek? [01:20:44] Ma'am, my question is related to the X1, X2, X3, right? So, you have given two examples. [01:20:50] In one example, we have… we are using X1X2 as an input. [01:20:53] Mm-hmm. [01:20:54] And in another example, we are using X1X2 as attributes of a certain input, so what is it exactly? I'm getting a little confused there. [01:21:02] Okay, so what I wanted to show here is the intricacy, or the… [01:21:07] complexity, you could go from. [01:21:10] So, these X1, X2, X3 could be single numbers. [01:21:14] Like, for example, this age, location, number of bedrooms, or… [01:21:18] amenities or whatever, they could be just one number each, they are not vectors. [01:21:24] Right? So they are just one numbers. [01:21:27] But we could also have inputs like the image. Image could be represented in form of a [01:21:33] vector, which are nothing but the number of pixels. [01:21:37] that we are having. So, each input actually could be a vector. [01:21:41] But both are possible. [01:21:43] Each input could be a simple number. [01:21:46] Or each input could be a vector. [01:21:48] This vector is representing an image. [01:21:51] And both are correct. This is what I wanted to say. [01:21:55] So, your data could be simple numbers going into it, [01:21:58] One number, just simple one number, there is no vector or matrix, so to say. [01:22:03] Here, each of these are… [01:22:07] Vectors in themselves, the inputs, there could be one image going here, right? [01:22:12] And this image is represented by a… [01:22:14] biovector, then there's another image as an input, and this could be our vector, like that. [01:22:19] So this is what I wanted to see. [01:22:21] Is it okay? [01:22:23] So, to summarize, in this example. Uh, for the housing index, I think you talked about. [01:22:30] Ideally, we are actually fetching one entire row as an X1. [01:22:36] Correct. Now, that row can even have… can have only one column, or it can have more columns. If it has more columns, it will become a vector. [01:22:42] And if it has a single column, then it's a scalar, basically, because it has only one attribute. [01:22:46] Yes, yes. Right. [01:22:51] Okay. Abhishek, do you have a question? [01:22:55] I mean, uh, right now you gave an example for image and, like, how it's input as a vector, but earlier you were also mentioning that the later layers will… [01:23:03] Uh, like, try to learn shape when I just write. So, can you relate the two, uh, like, how exactly does that happen? Is it… does it happen because of the activation function, or…? [01:23:13] What happens afterwards? I'll… [01:23:14] I will explain to you, yeah. [01:23:15] Okay. [01:23:16] Yeah, I'll explain to you. Just, uh… [01:23:20] Um, as we go along, I'll explain to you how the changes happen, and also tell you what our activation functions, different types of them, everything. [01:23:27] We'll cover. Okay. [01:23:31] So then, uh, what I told you so far in the previous was the activation function. [01:23:36] Working on the weighted, uh, summation of the inputs. [01:23:40] However, we also have something called the bias. [01:23:43] Okay? So, this bias, actually… [01:23:48] Uh, so when we have the weighted sum going here… [01:23:50] Then the bias is also inputted here. [01:23:53] So we have the bias working here. [01:23:55] So, uh, the activation function, actually, [01:23:59] Uh, is working on bias plus the weighted sum. [01:24:05] And then, uh, the bias is also because this expression [01:24:09] Uh, comes out to be, like, uh… [01:24:12] Uh, 1 cross n… [01:24:14] And cross… [01:24:17] Sorry. So these are… [01:24:21] the vectors are 1 cross n. [01:24:25] multiplied by a vector of N cross 1. [01:24:28] So we get a 1 cross 1 here. [01:24:32] So the bias is also 1 cross 1. [01:24:36] So, uh, bias is useful. [01:24:38] In many ways, sometimes we want to provide some offset. [01:24:42] Uh, to the summation, then we can add a bias. [01:24:46] Sometime the product of the input and the weights might come out to be zero, in which case the [01:24:51] When the input is 0, the output will always be 0 only. [01:24:56] can be 0, so to avoid that also, bias is added. So, bias is added for different things. [01:25:00] We'll study in more detail as we go along. [01:25:04] So now, having talked about this, what we have are inputs. Inputs propagate in this fashion. [01:25:11] Right? In the forward direction. [01:25:14] They are getting multiplied by weights, then they… [01:25:17] Then we have the bias and the summation. [01:25:20] activation function like this. [01:25:23] Then we have the output. [01:25:24] coming out here. What we're talking about is a form. [01:25:28] In a forward, we have… [01:25:31] What direction in which data and the information? [01:25:34] And then… [01:25:36] The end of the forward, we have the prediction. [01:25:39] Which is the value… [01:25:42] But, uh… [01:25:45] To break the friction, we'll have… [01:25:49] actually learn the… [01:25:50] So, we have the value of N… [01:25:52] suites, right? We have to learn these values. [01:25:55] It follows the following. [01:25:58] that initially, when we don't have… [01:26:01] Here, we just run… initialize these widgets? [01:26:06] So, initially, the weights are randomly… [01:26:09] And the product of… [01:26:12] turn the… [01:26:13] I am… [01:26:14] Ma'am, your voice is breaking. Is it only happening for me? [01:26:15] Yeah, your voice is breaking, ma'am. [01:26:19] Okay, sorry. Let me just, uh… [01:26:22] switch off my… [01:26:24] And probably you can just… [01:26:28] Is it better now? [01:26:31] Hello? [01:26:35] So… [01:26:36] Yeah, speak something, and then only we can… [01:26:37] Yeah, yeah. Oh, okay, I'll proceed. [01:26:40] Please let me know if it's breaking. So what I'm… [01:26:41] No, ma'am, it's still, still happening. [01:26:45] Okay, let me just… [01:26:46] Yes, ma'am still breaking me. [01:26:47] I'm just… [01:26:57] Just a minute. [01:27:55] So, I just switched off… [01:27:57] You can just let… [01:27:58] Still the same, still the same. [01:28:00] If you are able to properly. [01:28:03] No, still the same. [01:28:05] Yeah, this is breaking. [01:28:06] Oh, thank you. [01:28:07] celebrating. [01:28:12] Okay, just once again… [01:28:55] Okay, once again, Anne. [01:28:56] I'm starting. Please let me know. [01:29:02] So, what we are doing in the… [01:29:05] total for… I switched off my video. [01:29:08] That's a great problem. [01:29:12] No, I think you could try rejoining, uh, it is still not working. [01:29:19] So, I'll do that. [01:30:15] tried to change my memory, and please let me know. [01:30:18] has improved. [01:30:19] Okay. [01:30:23] Um, still same for me, I think… [01:30:26] Okay. [01:30:27] Ma'am, is there any, uh, loose connections, or any wire you are using? Like, [01:30:34] Maybe. [01:30:36] Yeah, so… it's certain… [01:30:41] As soon as you say now, uh, after each word, right, it stops over there. [01:30:49] Uh, okay. [01:30:53] Uh, I'll, uh, is it better now? I've tried to reconnect to… [01:30:58] Yeah, it is good, yeah. [01:30:59] Yes, ma'am. [01:31:00] Okay. Okay, okay, thank you. [01:31:01] It's better now. [01:31:04] So I think it was not the network. [01:31:08] You're… this… the last part was good. I tried to just reconnect. [01:31:16] Uh, so the training process would involve the following. First of all, since we don't have an idea about the widths, [01:31:24] So, what we will require is to randomly initialize the widths. [01:31:33] Okay? And then, uh, for each forward pass, what we would do is that, uh, the… [01:31:40] Weights would be multiplied by the input. [01:31:44] And, uh, this, uh, weighted sum would be passed on to the activation function, of course, with the bias. [01:31:51] And the output would be obtained. [01:31:54] Then the comparison between the actual target [01:31:58] a value that is actual expected value. [01:32:00] And the predicted value would be done, and the error would be calculated. [01:32:05] Uh, using some kind of loss function. [01:32:08] Say, for example, sum of squared error, or whatever, mean absolute error. We have studied so many loss functions. [01:32:15] Now, this error would actually… [01:32:18] Uh, no, be back-propagated. [01:32:20] Or it would be passed backwards. [01:32:22] So, in the forward… so, initially, we did a forward pass. [01:32:26] Here, we found out what is the error between, uh… [01:32:30] The predicted value and the actual value. [01:32:32] Expected value, and the error is actually back-propagated. [01:32:36] And once it is back-propagated, then what would happen is… [01:32:40] that the contribution of these weights to the error would be identified. [01:32:45] And once this contribution would be identified, the weights would be [01:32:49] changed. And then, once again, after changing these weights in the next iteration, the weighted, some new weighted sum will be [01:32:57] Um, passed on to the activation function. [01:33:01] And, uh, the new output could be calculated, and then, again, the difference between the expected output and the [01:33:09] New output is calculated. [01:33:11] And whatever the error that comes in is back-propagated, and this is done iteratively again and again. [01:33:18] Till, uh, the expected outcome and the actual outcome, the outcome of the algorithm, [01:33:24] to quite some extent. [01:33:26] or the loss is minimized. [01:33:28] What is loss? Loss is the sum of the error that is obtained on each and every [01:33:33] record, or each and every, uh, uh, training data that is passed to the algorithm. [01:33:40] So, this is how, uh… [01:33:42] the forward pass and the backward pass. [01:33:45] is done, in which the errors are back, propagated, [01:33:49] And the back propagation of the errors are used to refine the model and give better results. [01:33:55] And, uh, the rates are iterated. [01:33:59] So that the loss gets reduced, and uh… there's something called gradient descent. [01:34:05] That helps to adjust the weight. [01:34:07] So, uh, gradient descent is an algorithm that helps to… [01:34:12] Um, uh, to, uh, adjust the weights in such a way that, uh, [01:34:17] The error, uh, you know, it… [01:34:22] keeps on decreasing, or it gets minimized. [01:34:26] And uh… if it gets mini-wise, then the gradient descent will work further to reduce it more. [01:34:32] Like that. So it, uh, iteratively keeps on… [01:34:35] Reducing the error, we'll discuss in a lot of detail how gradient descent does it. [01:34:41] As of now, I want you all to know that gradient descent actually [01:34:45] tries to, uh, adjust the weights in such a way that the error [01:34:49] is reduced, okay? [01:34:51] And then, uh, the new weights are used in the forward pass. Next forward pass. [01:34:57] The error is found out it is back-propagated, and so on and so forth. [01:35:02] So now, uh… [01:35:05] Now, we have, uh, many architectures. [01:35:08] In, uh, which are made up of neurons. [01:35:11] The simplest one is a multi-layer perceptron, or the MLP, in which there is a single hidden layer. [01:35:16] We have the input layer, output layer, and a single hidden layer. [01:35:20] Okay? And, um, again, the number of neurons that must be there in the input layer, output layer, hidden layer, each hidden layer, [01:35:29] Again, can be, uh, is variable. We need to actually adjust [01:35:34] A lot of parameters, like the number of hidden neurons in the hidden layer, [01:35:39] In the input layer, in the output layer, and there are so many other parameters. [01:35:43] that we need to optimize, and uh… [01:35:46] We will be discussing all these as we go along. [01:35:49] So now, uh… [01:35:52] What we discussed is the… [01:35:54] the forward pass. [01:35:56] And uh… in the forward pass, what is happening is that each layer [01:36:02] is actually going to do a non-linear transformation. [01:36:06] On the input data, [01:36:08] Uh, in such a way that it is mapped from one input vector space to another [01:36:13] input vector space. [01:36:15] Right? So, it is going from one space to the other. I'm already… I have already explained to you. [01:36:21] For example, initially, [01:36:24] The data was, say, 1 cross n. [01:36:28] Here, our initial input was… [01:36:31] Uh, of the form, one cross in. [01:36:35] Then, after applying the weights of N cross 1, what we obtained is an output of 1 cross n. [01:36:40] So, we have mapped the data out, transformed it from a 1-cross N to a 1 cross… [01:36:45] One dimensionals. [01:36:49] So, in this process, [01:36:50] We are actually transforming the data from one particular dimensionality to another, [01:36:55] Uh, vector space or dimensionality space. Additionally, [01:37:00] We are also doing a non-linear transformation on the data with the help of the activation function. [01:37:05] Because we just don't want… [01:37:07] are in, uh, output to be a linear combination of the input, and hence we use the… [01:37:12] Uh, activation function. [01:37:14] Okay. So now, uh, let's see how all this works. [01:37:19] So, let's say that we are having, uh… [01:37:24] We are having a given… [01:37:26] artificial neural network, and uh… [01:37:32] Uh, it, uh, actually is having some multiple layers in it. Say, for example, there are two hidden layers. [01:37:38] One output layer and one input layer. [01:37:41] And, uh, our input is nothing but a 3D input, because it is of that dimension. [01:37:47] 1 cross 3. Then, we are actually… [01:37:50] Uh, now, uh, as I already explained to you, we are having a fully connected neural network. [01:37:55] So, the input from each… each input is actually going to each and every neuron in the next layer. [01:38:02] So, each input is connected to all other neurons in the next, or nodes in the next layer. [01:38:08] So, what we will have here is a… [01:38:10] 3 cross 4. Transformation, then we have a 4 cross 4. [01:38:15] And then we have a 4 cross 1. [01:38:18] So, this is how… what we'll obtain is the final output. [01:38:22] So, I'll explain to you in more details. [01:38:25] So, our, uh… [01:38:27] Uh, this A and N is actually… is nothing but it's… [01:38:31] It's a kind of a chain of metrics multiplications. [01:38:34] Along with non-linearities applied to it with the help of activation function. [01:38:39] So, let me see how these… let me explain to you how these matrix multiplications work. [01:38:45] So, how this matrix multiplication works is… [01:38:48] Let's say our initial data is, like, [01:38:51] represented in one cross, uh, three form. [01:38:54] So, this is our input. [01:38:57] So, uh, we are having 3 inputs, uh, X1, X2, and X3. [01:39:03] X2 and X3 here. [01:39:06] And, uh, we are actually, uh… [01:39:09] representing these. [01:39:11] Okay, um… [01:39:15] So, I can see that there is a question. [01:39:18] That says that regarding the ANN model, is there a recommended upper limit on the number of input attributes? [01:39:26] Or does it depend on the dataset size? [01:39:29] So there is no limit on the number of attributes. [01:39:32] And the size of the data, it just depends on our data, and that's all. [01:39:39] Okay. [01:39:43] So, coming back to this, we have the input, which is in the form of a 1 cross 3. [01:39:48] metrics, so we are having, uh, uh… [01:39:51] 1 cross 3, then we are putting a transformation with the help of a weight matrix, which is a 3 cross 4 in nature. [01:39:58] Because we are having 3 inputs. [01:40:02] And we are having, uh… [01:40:03] four neurons in the next layer, so we are mapping the three inputs to a… [01:40:09] Uh, 3 cross 4 form. [01:40:10] Then in the next layer, [01:40:12] We are having, again, 4 neurons, so we are having a 4 cross 4. [01:40:17] And then, uh, the output weights are 4 cross 1. [01:40:20] And finally, what we are having is a 1 cross 1. So, we can think of it… [01:40:25] Uh, as a chain of matrix multiplications, [01:40:29] Uh, here, and uh… [01:40:31] What we are obtaining is from 1 cross 3, [01:40:36] We are getting transformation. [01:40:37] to 1 cross-one format directly with the help of these mid-matrices, like W1, W2, W3, and all. [01:40:44] So, I'll explain this in more details as we go along. [01:40:48] So here, we are actually… [01:40:50] transforming, uh, first, uh, [01:40:53] Uh, the… [01:40:55] three-dimensional input to a 4D space. [01:40:58] Then, uh, and to do that, we are using a matrix of the form 3 cross 4. [01:41:04] Then, we are again taking this 4D space to another 4D space with the help of another weight matrix, which is 4 cross 4. [01:41:12] The initial one was 3 cross 4. [01:41:14] And from there, we are finally wanting to reduce it to a 1D space, so we, again, [01:41:19] transform it with the help of another weight matrix, which is a 4 cross 1 form. [01:41:27] And, uh, so here, the dimensions of the matrix, [01:41:32] represent the input. [01:41:34] And then the different dimensions, uh… [01:41:37] represent the, uh, of the weights, represent the transformations that we are doing it. [01:41:43] Uh, with the help of different weight mattresses, okay? [01:41:47] So this is how, uh, our, uh… [01:41:50] these transformations on mappings work. [01:41:53] So I think I saw handwrist, uh… [01:41:57] Is there a question? [01:41:59] Mm-hmm. [01:42:00] Yes, ma'am. So, uh, for the weight, or, like, in terms of the definition, what I was considering is, for each attribute, we would have some sort of weightage or significance. So, I was associating this weight to that attribute. [01:42:12] But in here, I think audio explained what I understood is. [01:42:17] So it depends on the next layer, like, number of nodes that we have in the next layer. [01:42:22] So that is something I'm not able to understand. [01:42:25] Where does the weight actually we need to associate? Is it the next layer of. [01:42:34] Actually, uh… [01:42:35] Notes, or that we have as input. [01:42:37] So, it's the same thing that we discussed. So, let's say we are having these inputs here. [01:42:43] Okay? [01:42:45] Because each input is going to contribute to each of the next level neurons. So, X1 is… [01:42:53] Actually, going to contribute to this node. [01:42:56] This node, this node, and this node. [01:42:59] So, it will be multiplied by different weights. [01:43:03] Okay, four different weights. [01:43:05] Uh, so that the contribution of this [01:43:09] goes to each and every node. So, input 1 or X1 will contribute here. [01:43:15] to each of these, then… [01:43:18] X2 will contribute also to each of these. [01:43:21] And X3 will contribute also to each of these. [01:43:24] Accordingly, what you're seeing here… [01:43:27] This represents the contribution [01:43:29] Let's say if this matrix is, uh, W11, W12, [01:43:34] W13, W14. [01:43:37] So, this is the contribution of X1. [01:43:41] Okay, to node 1, Node 2, Node 3, Node 4. [01:43:45] Then, this is X2. So, this is a… [01:43:49] Contribution to X22 node 1, node 2, Node 3, Node 4. [01:43:53] Then node 1, this is the contribution of… [01:43:56] The third input to all four nodes. [01:43:59] And then these get added, the… [01:44:01] let… this is our node 1. [01:44:03] This is our node 2. [01:44:05] This is our node 3, and this is our node 4. [01:44:10] N1, N2… [01:44:14] N3, and N4. [01:44:15] And these are added, so what you get is, at N1, contribution, [01:44:19] Weighted contribution of X1, X2, and X3. All these 3 are combined together. At N1 again, [01:44:26] At entry again, and these are different widths that are multiplied and added together to get [01:44:32] A combination of X1, X2, and X3 are 101, at N2, at N3, and N4. [01:44:39] Okay? Like that it goes. [01:44:44] I will… yeah, I'll explain to you with a numeric example. I think that will be more clear. [01:44:45] Yeah. And… What happens on the next layer, which is… [01:44:49] Yeah. [01:44:51] So, uh… [01:44:52] Thank you. [01:44:53] So I'll just, uh, take a… [01:44:55] Example, in which, uh… [01:45:00] I'm going to consider… [01:45:03] some values for… [01:45:05] all of you. [01:45:07] So, I'll take a slightly simpler… [01:45:10] neural network. [01:45:13] So, this is the network that I'm going to use. So, it's going to have 3 inputs, say X1… [01:45:19] X2, and X3. [01:45:23] Okay, and then there is one hidden layer only, just for the sake of simplicity. [01:45:28] I'm making it as one hidden layer and one output layer. [01:45:40] So it's becoming a little cluttered up, so I'm not making all the connections. [01:45:44] Or maybe I can make… [01:45:46] And then all of these will go here. [01:45:49] Like this. Okay. [01:45:52] So then, this is our input, X. [01:45:56] And let's say I'm taking a very simple example. [01:46:01] So I'm taking very simple numbers and a small… [01:46:05] metrics only, just for the sake of clarity. This is your… [01:46:08] input matrix X. [01:46:11] And then you have our weight matrix, which is our 3 across 4 metrics. This is a 1 cross 3 matrix. [01:46:17] So, the weights may be like this. [01:46:22] Just some simple numbers I'm taking. [01:46:26] Okay? And then we have the next set of weights, which are going on the output side here. [01:46:32] So we have 1, 2, 2, 1. [01:46:35] Just another rate matrix. [01:46:38] And then what we'll have is the output. [01:46:40] So, what it will look like… [01:46:42] So, first of all, it will be… [01:46:45] Going like this… [01:46:47] simple matrix calculation. So, we are multiplying it like this. [01:46:52] The usual way, plus 2 into 2 plus 3 into 2. [01:46:57] Then the next entry would be 1 into 1. [01:47:00] Plus, um… [01:47:03] 1 into 1 plus 2 into 1… [01:47:06] plus 3 into 1. [01:47:09] Then the next entry would be 1 into 2. [01:47:12] plus 2 into 3… [01:47:15] plus 3 into 1. [01:47:18] And then 1 into 1, plus… [01:47:20] 2 into 1 plus 3 into 3. This is the standard way. [01:47:26] And then this followed by 1, 2, 2, 1. [01:47:30] So what we will have here is… [01:47:33] Uh, 1 plus 4 plus 6, so this is 11. [01:47:39] And then here we are having… [01:47:41] 3 plus 2 plus 1, so it is 6. [01:47:45] And then we'll have 3 plus 6 plus 2. [01:47:48] So it will be 11, and then this is 12. [01:47:52] And then it is 1, 2… [01:47:54] 2 and 1. [01:47:56] And then final outcome will be what? [01:47:59] 11 into 1 plus 6 into 2. [01:48:03] plus 11 into 2. [01:48:05] does 12 into 1? [01:48:07] And that will make it 11 plus 12 plus 22 plus 12. [01:48:12] So that will be 57. So this will be your output. [01:48:16] This is how it is working. [01:48:18] So, if we talked about this example that I showed to you earlier, [01:48:23] So now… [01:48:26] Here, if we think about it as the notes, this is… [01:48:29] These are your N1, N2, N3, and 4. [01:48:32] And then… [01:48:34] Here is your, let's say, N5, N6, N7. [01:48:39] and N8. Then what is happening? [01:48:43] The contribution of the weighted sum of what was coming out. [01:48:47] Here. So now, this is your… [01:48:51] And 5… this is N6. [01:48:53] This is N7. [01:48:55] And this is N8. So, what is happening here? The weighted… [01:49:00] Some of all these three were going as input to X1. [01:49:04] Now, what will happen? Again… [01:49:06] Whatever is coming out of here, which is… [01:49:10] Uh, which was what? [01:49:12] This was… let's say if the weights were, uh, W11… [01:49:17] times X1 plus… [01:49:20] Uh, this was W21 times X2. [01:49:24] plus W3 1… [01:49:27] times X3. This was what was going here. [01:49:30] Now, this is going to go here. [01:49:32] Then this is going to go here, this is going to go here, this is going to go here. [01:49:37] So the… all these components… [01:49:40] Which are similar weighted… [01:49:42] sums will go here, and make up N5, N6, N7, N8. [01:49:47] So, this is a contribution of N1, this is the contribution of N2. [01:49:52] This is a contribution of N3, this is the contribution of N4. [01:49:57] And it's a little complicated to write, but I can write and show you also. [01:50:02] Just a second, I think there was a question along… [01:50:08] Just a second, yeah. [01:50:10] So, what will be, uh… [01:50:13] N5 having the contribution of N1, 1. [01:50:18] Uh, would be, like, for example, X, uh… [01:50:23] X1 times W11. [01:50:28] Plus… [01:50:34] Thanksgiving is a little bit of a problem. [01:50:37] Then we have X2 times W21. [01:50:43] Thus, X3 times… [01:50:47] W31. Let's say… [01:50:49] This is what you obtained at N1. [01:50:53] Okay? Now, this will be multiplied by the further weights. [01:50:58] whatever weights you are having, this is the input going here. Now, [01:51:02] Let's say at the next stage, the weight was W21. [01:51:06] So, this will go here. [01:51:08] This part will be what I'm writing here. [01:51:11] The contribution of N1. [01:51:13] So, I'm talking about what is going at N5. [01:51:17] Okay? This will be the contribution of N1. [01:51:23] Then talking about contribution of N2 will be… [01:51:26] X1 times W21. [01:51:28] plus X2 times W. [01:51:31] Uh, 2-2 plus X3 times… [01:51:35] W32… [01:51:37] times… [01:51:39] So, this is, let's say, the second layer, first one. [01:51:42] Then it is a second layer. [01:51:44] Second one. [01:51:46] This was the contribution of N2 to N5. [01:51:50] Then we have the contribution of N3. [01:51:53] 2N5? [01:51:55] And this will look like… [01:51:59] X1… [01:52:01] Okay, my… [01:52:07] Because you've done writing pad is not… [01:52:08] working properly. That's why I'm having the most problems, yeah. [01:52:12] So then we have X1 times W31. [01:52:16] plus X2 times… [01:52:19] W32 plus X3 times W33 [01:52:24] times… [01:52:26] W23. [01:52:29] Then, similarly, we have the contribution of N4. [01:52:32] So, all these together are the contributions. This is the contribution of N1. [01:52:40] Then this one is this one. [01:52:41] Then, the third one… [01:52:50] It's going again and again. It's giving me a lot of problems. [01:52:54] So this is your N2. [01:52:56] This is your N2, this is your N3 contribution. [01:53:00] And this is your N4 contribution. [01:53:03] So, I think there was a question on what will be the next looking like. This is what it will look like. And you can see, [01:53:09] That in such a simple ANN with just two layers, 3 input and one output, [01:53:14] It is becoming so very complicated, and this is what… [01:53:18] what you're getting at N5 only. [01:53:23] Then, similarly, we can do it at N6 contributions. This will represent N6. [01:53:27] N7N8. Is it okay? [01:53:29] Any questions on this? [01:53:33] And once all of this calculation is done, then we apply activate function, right? [01:53:39] Yeah, so the activation goes… [01:53:41] At this point. [01:53:43] Add this one. All this summation is done, and on the summation, [01:53:49] Activation is applied. [01:53:51] Okay. So, on each layer, it will be done. [01:53:54] Yes, actually… [01:53:58] So, you know, start equation. [01:53:59] So you can see here… [01:54:03] This is your activation. [01:54:05] And it is on the submission of all of these inputs. [01:54:09] Okay? [01:54:11] Okay, thank you. [01:54:15] So, in such a simple… [01:54:18] in, and with such few inputs, [01:54:21] Such few hidden layers and per hidden layer, such few neurons. It looks so complicated. [01:54:28] And what we are trying… and then this is the forward pass, and then the error will be back propagated. [01:54:34] And these weights will be changed. [01:54:36] And you know that these weights are actually contributing [01:54:40] Uh, and making… I mean, are helping, uh, the two… [01:54:46] to, you know, map the contribution of each of the inputs to each of the hidden [01:54:53] Um, uh, node or neuron in the hidden layer. [01:54:56] And therefore, the role of weights is very important, and we do different iterations. [01:55:02] To keep on, um, changing these weights such that [01:55:05] The final outcome, which will be the activation done on… [01:55:10] This together, then this together, [01:55:13] will be very much similar to the expected output. So, this is what we are doing. [01:55:19] In here. [01:55:21] So, this is just, uh… [01:55:23] What I wrote, I will… for… [01:55:26] more clarity, I'm just going to write in the form of just notation. [01:55:31] whatever I wrote for those… [01:55:33] who wish to understand it. [01:55:37] In this way, let's say these are the weights. [01:55:40] So I was actually… [01:55:42] writing about these only. [01:55:51] Then we have W31, W32. [01:55:54] W33W34. [01:55:57] So these are our inputs, X1, X2, X3, 1 cross 3. [01:56:01] Then, this is 3 cross 4. [01:56:04] And this was the… [01:56:06] input. [01:56:08] X1, X2, X3. [01:56:14] And then this was our next layer. [01:56:19] each of these are connected like this. [01:56:21] In a fully connected layer. [01:56:24] Like this. [01:56:27] And then, this is a, uh… [01:56:30] 1 cross 3, then we are going at 3 cross 4. [01:56:34] And let's say, uh… [01:56:36] We are having, uh, contributions of each of the input going to each of the neurons, N1, N2. [01:56:43] N3 and N4. [01:56:45] So, this is what is at N1. [01:56:48] I mean… [01:56:50] I'll not say it like this. [01:56:53] And let's say the next matrix will be [01:56:56] of this fall? [01:56:58] If this is our output. [01:57:00] And these all are going into the output, like this. [01:57:04] So it's becoming a little cluttered up. [01:57:08] But anyhow… [01:57:10] And then what we'll have is output. [01:57:12] So then, the products that will be [01:57:15] will go like this. [01:57:17] plus X2 times… X2 times W21. [01:57:24] Thus, X3 times W31. [01:57:28] Then we have X1 times… [01:57:31] W12? [01:57:34] Thus, X2 times W22? [01:57:37] plus X3 times W32. [01:57:41] Then we have X1 times… [01:57:45] W13, this is this and this. [01:57:49] plus X2 times W23. [01:57:52] plus X3 times… [01:57:55] W33. [01:57:57] So, I did this with this row. [01:57:59] And then finally, we have X1 times W1 fourth. [01:58:04] plus X2 times W. [01:58:08] 24 plus X3 times W34. [01:58:15] It's just open. Right. [01:58:18] Now, this is what you're getting at node… [01:58:23] N1. [01:58:24] Which is weighted contributions of X1, X2, and X3. This is what you're getting at N2. [01:58:31] This is what you're getting at N3. [01:58:33] And this is what you're getting at N4. [01:58:37] And then suppose we are having the next layer. [01:58:42] Like this. [01:58:45] Then, finally, what we will get is… [01:58:49] This multiplied by this. [01:58:52] So, it will be… [01:58:54] X1W11? [01:58:56] plus X2W21. [01:58:59] plus X3… [01:59:02] W31. [01:59:06] So it's becoming very laggy, that's the reason I'm not able to get it. [01:59:10] Right, okay. X11? [01:59:15] So, plus X2 times W21. [01:59:20] I don't know what's the issue with the pen? [01:59:24] Yeah. Plus… [01:59:26] X3 times W3 1. [01:59:29] This whole thing, multiplied by T11. [01:59:33] Plus, whatever contributions are coming from [01:59:37] Uh, this next one. [01:59:40] I'm just writing it as N2. It's not N2, just… just for sake of clarity. [01:59:45] Like this, we get… and finally, what we get is a… [01:59:49] 1 cross 1. [01:59:51] matrix, which is our output. [01:59:54] Okay, so this is how it works. [01:59:56] I haven't yet shown you the activation functions, but I'm just showing you the simple [02:00:01] metrics transformations that are happening. [02:00:04] Is it okay? Any questions, anyone? [02:00:08] Yes, ma'am. [02:00:09] Yeah, go ahead. [02:00:12] Okay, uh, correct me if I'm wrong, okay? [02:00:15] So, uh… so on the very right, right, that would be our, uh, class which we would be predicting, right? [02:00:22] Yes. [02:00:23] And then, uh, 1 cross 3 is our input, correct? [02:00:26] Yes. [02:00:27] So, that means if we talk about the training part, right, we would be giving features to the left, right, and we would be expecting certain, uh, value. For example, on 1 to 10, maybe probably 4. [02:00:40] We might have assigned to a certain thing, right? [02:00:43] Mm-hmm. [02:00:44] Uh, depending on the features. And then, now, what we are trying to do is, we are trying to, uh, change the weights, right? [02:00:50] Mm-hmm. [02:00:51] Uh, so that the output comes as 4, because we chose, right, for a given feature, that would be 4, right? [02:00:56] Hmm. [02:00:58] Correct, correct. [02:00:59] That's how it's… so we are changing the weights, correct? [02:01:01] Mm-hmm. Yes, yes. [02:01:10] Yes. [02:01:11] And then once we reach to a point where expected value, right, and the given value is same, uh, we stop it, right? Then that is the model we want to have, right? And then… [02:01:14] Uh, we would, uh, give some more, uh, unknown data, probably, right? And we would try to predict what it might give. [02:01:21] Yes. Exactly. [02:01:22] Is that the… is that correct? Okay. [02:01:24] Got it, got it. [02:01:25] Okay, so, uh, also here, 1 cross 3, right? [02:01:30] Uh, when we say 1 cross 3, uh, I am, uh, supposing that whatever arrows we are connecting over there would have a weight on it, right? [02:01:38] Which arrow? These arrows. [02:01:41] Yes. [02:01:42] These arrows, right? So, this would be the weight. So, initially, there would be some random weights, right? [02:01:46] Yes, yes, that's correct. [02:01:47] And then, similarly, when we connect N1 to the end node, right? [02:01:53] Yes. [02:01:54] There would also be some, uh, random, uh, awaits, right? [02:01:59] Yeah. [02:02:00] These are the weights which will be initially randomly assigned. [02:02:03] Yes. [02:02:04] randomly available, right? So, it could be anything, and then we will just update those weights, uh, until we get the expected result we are expecting, right? [02:02:11] Yes, exactly, exactly. [02:02:12] Okay, okay. [02:02:14] Okay. [02:02:15] Okay, and then what about the bias part? I have seen in the previous slide. [02:02:19] Yeah, so bias is actually, uh, added again at this level. [02:02:24] So, when you are having the summation, [02:02:26] Along with the summation, you'll have a bias. [02:02:30] As you can see here… [02:02:34] So, what role does it play, actually? [02:02:36] Yeah, so see, the bias is, as the name suggests, it's going to… [02:02:41] Uh, add some offset [02:02:45] Okay, okay. [02:02:46] to this. Suppose you want… let's say your expected output is still not coming, [02:02:51] In spite of iterating on these weights, [02:02:54] It's not, uh, still coming out. [02:02:57] your outcome or your output is still not coming, uh, similar to your expected outcome. [02:03:04] Mm-hmm. [02:03:05] Then you have a bias to play around with. [02:03:06] So that this input to the activation function may be improved so that your [02:03:12] Actual output and expected output match. [02:03:16] So, it is something like seeding or something, right? We are just trying to… [02:03:21] Okay? [02:03:22] Yeah, exactly, yeah. You are trying to offset this weighted sum with something else. [02:03:26] So that, with a desire to get a better output. [02:03:30] That's the whole thing. [02:03:31] Okay, so weights were not enough, so we added one more extra parameter as we… [02:03:35] Okay. [02:03:36] Yeah. Yeah, yeah, and as you will see here, these are also not enough. Additionally, [02:03:40] You are going to have different types of activation functions, you're going to have different kinds of optimizers, [02:03:46] Your, uh… so you can see the amount of uncertainty in this whole process. [02:03:51] Yeah. [02:03:52] The number of neurons that you can [02:03:54] Keep in the hidden layers. Itself is variable. [02:03:57] You can vary that, that is a hyperparameter. [02:04:00] Yes. [02:04:01] Then the number of layers itself is also a hyperparameter. [02:04:04] Yes. [02:04:05] The bias is a hyperparameter, the activation function is a hyperparameter. [02:04:09] Then the number of iterations you use to iterate on the weights [02:04:13] is a hyperparameter. [02:04:15] Then you have other things like epochs and so many, so everything is variable. [02:04:20] Just to keep on trying to… [02:04:23] nudge the model to move towards the output. [02:04:28] And why we are doing all this is because our data is very complex. [02:04:31] And, uh, the algorithm… [02:04:34] You cannot just, you know, [02:04:36] With fewer number of parameters or hyperparameter, it may not be able to give the desired result. [02:04:43] Okay, so that's why we want more and more flexibility, and to create more flexibility, we are actually… [02:04:49] Putting in more and moreables. [02:04:52] or parameters in the model which we can play around with. [02:04:55] That's the whole idea. [02:04:57] Okay, so I'm assuming most of the parameters are just limitation to our computation power, right? [02:05:05] Yeah, so definitely, deep learning models have a huge number of hyperparameters. [02:05:12] compute-intensive, yeah. [02:05:13] Which makes them very, uh, you know, yeah, compute intensive. [02:05:15] Yeah, I feel that, like, the 3 mattresses over here is more than, like, it's taking a lot of, uh, power. [02:05:24] Yeah. [02:05:25] And then you're saying, uh, we could, uh, there is no limit, so we could go to, I mean, n number of, uh, uh… [02:05:31] neurons and a number of, uh, layers at that same time, right? [02:05:35] Yes, yes, that's correct. [02:05:37] Yeah. [02:05:38] Okay. Okay, so… okay. So, I'm… I think the combination part of where we combine, uh, multiple, uh, [02:05:46] Hyperparameters? Yeah. [02:05:49] activation function. [02:05:50] What was that again? Uh, non-linearity thing. No, no, the non-linearity, we added activation functions, yeah, so we might be combining them together, right? You said… [02:05:55] The whole layer was for that particular activation function, right? [02:06:00] Yes, yes. And at different layers, we could have different activation functions, also. [02:06:05] So, the combination, again, is a hyperparameter. [02:06:08] Okay. [02:06:09] That, uh, different layers could have different… [02:06:12] activation functions, and the combination could actually, uh, [02:06:17] Together result in a better output. [02:06:20] And, uh, the thing is that whatever we are seeing right now looks very… [02:06:26] tedious, and uh, you know, it has so many… [02:06:29] Hyperparameters, and so many things. [02:06:32] But this is just the most simple form of it. [02:06:35] Okay. [02:06:36] That is the irony, that what we are thinking today is quite complicated. [02:06:40] that at each node, there are contributions of each input, then these are again getting weighted. [02:06:45] Going forward, this is blah blah, then we are having… [02:06:48] Activation functions, [02:06:50] We are having, uh, you know, optimizers. You're having so many things. [02:06:57] And this looks complicated, but this is more simple. [02:06:59] And when you learn other models, more and more complex deep learning models, [02:07:04] This only is the basic thing. [02:07:07] Okay? So… [02:07:08] Okay, got it, yeah. Thank you. [02:07:10] Yeah, I could see another hand raised, uh… [02:07:13] Is that a question? [02:07:15] Yes, ma'am. So, uh… [02:07:17] Just for the clarification, if there are, let's say, uh… [02:07:23] Mm-hmm. [02:07:24] Uh-oh, n number of layers. So, does the, uh… [02:07:25] Each layer has a different activation function, or it will be the same across those layers. [02:07:31] And, uh, that was one question. [02:07:34] The second one was, uh, if we do a backpropagation, [02:07:38] Uh, do we change the bias as well? Uh, uh, for the back propagation? [02:07:43] Or it stays same and we only adjust the weights, uh, as per the, uh, desired outcome. [02:07:51] Okay, so coming to your first question, the activation is on a per-layer basis. [02:07:57] So, for different layers, you could have similar or different activation functions based on your [02:08:03] requirement, or whatever suits you, okay? So there is no necessity [02:08:07] that all the layers would be having the same activation function, not… [02:08:11] Not required. It could be same, it could be different. [02:08:15] Okay. [02:08:16] So, the activation function is parlayer. [02:08:18] Fine. That is your first part. Second part is about the bias. [02:08:22] Yes, the biases also are adjusted in backpropagation. [02:08:26] So, nothing is fixed here. [02:08:28] Okay. [02:08:29] And that is the… [02:08:30] Uh, you know, that is the… [02:08:32] uh… what should I say? That is a… [02:08:35] most important part. [02:08:37] of the, uh, deep neural… I mean, deep neural networks, or… [02:08:42] Generally speaking about the deep learning models, everything is variable only. There's nothing fixed. [02:08:48] So, uh, so yes, bias also can be iterated in backpropagation. [02:08:54] Uh, and it can be varied, too. [02:08:57] bring about better output. That's the whole idea. [02:09:02] Okay? Yeah. [02:09:04] Any other questions, anyone? [02:09:09] So it's quite interesting if you look at it. It's, you know, a lot of maths going into it. [02:09:16] And uh… just keep on thinking, just if you visualize that there are, let's say, two hidden layers. [02:09:23] Then you talked about the contribution of the inputs going to each node in the first layer. [02:09:29] Now, these will get weighted [02:09:33] By the second layer, say if it was second hidden layer, [02:09:37] Uh, and the weights are, like, W21, I should say, capital W21, or capital 111. [02:09:43] These will further get weighted. [02:09:46] And then their components would get added. [02:09:49] And they would go to the next layer. [02:09:51] And if there are more layers, then that would go on and on. [02:09:55] So, the initial inputs are a very small portion of it, which will be weighted, will go here. [02:10:01] at N1, N2, N3, and N4. [02:10:03] And that small portion will again get weighted. [02:10:06] And, you know, those will propagate further, more and more, more and more like that. [02:10:11] And therefore, when we finally arrive at the final outcome, if there are… we are having multiple layers, [02:10:18] So, it will be a… a lot of different combinations of X1, X2, X3, which will get added together. [02:10:27] Uh, to get the final outcome, which will be your Y. [02:10:31] So, uh, in this example, we just have single layer, so you can look at [02:10:35] this combination. It will be untickable. [02:10:38] If we talk about, you know, another layer, [02:10:41] Having N5, N6, N7, N8, [02:10:43] Then N1N2, N3, and N4. [02:10:46] Uh, if these are the ones that, you know, we just… I'm simply saying this [02:10:52] Uh, they will get, uh, weighted and added together, all of these. Say, W11 multiplied by this, added with W… [02:10:59] 1 to then multiply it by W13, multiplied by W14, all of these added. So you'll get a very fancy full [02:11:08] expression. And then you'll… [02:11:09] have similar products for all four. [02:11:12] neurons, if there was another layer. [02:11:16] And so on and so forth, like that. [02:11:19] So, any last few questions before we wrap up for today? [02:11:28] Sorry, probably a dumb question. How is the… [02:11:31] layers, hidden layers, uh, decided as to, let's say, in this example, 1 is 2, 3, where the kind of input layers. [02:11:39] How was, uh, let's say, if it were a real kind of example, and so forth. [02:11:44] So, how it was decided that it would have 4 nodes in the… [02:11:49] kind of, uh, second layer. [02:11:52] Yeah, so the best answer I can give you is… [02:11:56] That in… if you talk about deep learning models, there is no straightaway answer that, okay, 4 or 3 or 5 or 10 will be good. [02:12:04] All you have to do is just iterate. [02:12:07] So, let's say if I have to decide, first of all, the number of layers I want to choose, [02:12:13] Then I'll have to iterate on the number of layers and see that is a hyperparameter in itself. [02:12:17] Then, in each layer, [02:12:20] How many neurons or how many nodes that we want is also a hyperparameter. There's no fixed solution. [02:12:27] So, we cannot say that we can go with 5, 10, 15, 20, or something like that. [02:12:32] We'll have to iterate and see how many per layer are giving the best result. [02:12:36] And that's why, because there are so many hyperparameters, you can think of a plane. [02:12:41] Or you can think of, uh… [02:12:43] A set of variables, which is very high. [02:12:46] So, the number of neurons per layer is a variable. [02:12:49] Number of layers itself is a variable. [02:12:52] And all these you iterate, and that's why… [02:12:55] Your deep learning models are very, very compute-intensive, because you have to iterate all these. [02:13:00] to arrive at the best answer, that is, what will be the best number of players, what will be the best number of [02:13:05] Neurons perler, and all that, but there's no fixed answer. [02:13:09] You just have to iterate and see. [02:13:15] Okay, any… anything else? [02:13:22] Okay, if there are no further questions, then we can stop here and continue on this tomorrow. [02:13:23] Thanks. [02:13:28] Okay, well, thank you all, have a great evening. Let's meet tomorrow, then. [02:13:31] Thank you, bye-bye. [02:13:32] Thank you, bye-bye. [02:13:35] Thank you.