# 02 2026-01-25 Deep Learning Concepts Intro

course: Module 3 — Deep Learning & NLP
module: Module-3-Deep-Learning-NLP
date: 2026-01-25
type: transcript
video_url: https://personal-learn.armco.dev/files/_Recordings/Module-3-Deep-Learning-NLP/02_2026-01-25_Deep_Learning_Concepts_Intro.mp4

---
[00:06:59] Durga Toshniwal: Good morning, everyone. We'll just wait for 2-3 minutes, so… More people will join.
[00:09:10] Durga Toshniwal: A very good morning to all of you, and welcome to today's session.
[00:09:14] Durga Toshniwal: So, today, we'll be continuing with the discussion on deep learning, so let me just share my slides, and we'll go on from there.
[00:09:35] Durga Toshniwal: Just a second.
[00:09:52] Durga Toshniwal: Okay.
[00:09:54] Durga Toshniwal: So, yesterday we started out with the introduction to deep learning and, the relevance of the term deep in deep learning, which refers to a large number of layers in the neural network.
[00:10:06] Durga Toshniwal: And that's why the term deep is coined. Additionally, the word also draws upon the concepts from the real world, that is, human brain, which comprises of neurons. The neurons in the brain are connected.
[00:10:22] Durga Toshniwal: To each other, all neurons similarly here.
[00:10:26] Durga Toshniwal: The… Neurons or the nodes are organized in layers.
[00:10:32] Durga Toshniwal: We have the input layer, the output layer, and in between them, hidden layers, which can be one or more hidden layers. And the, the outside world can interface with the neural network.
[00:10:45] Durga Toshniwal: With the help of the input or the output. So, the input is accepted, add the input layer.
[00:10:52] Durga Toshniwal: The hidden layers are not directly accessible, but they interact with the input layer.
[00:10:57] Durga Toshniwal: And then, the data is passed to them, it propagates through them.
[00:11:03] Durga Toshniwal: After undergoing some transforms, and finally the output is generated, and it's available at the output layer.
[00:11:11] Durga Toshniwal: We also talked about the differences between machine learning and deep learning, though of course, deep learning is a specific subset of machine learning.
[00:11:19] Durga Toshniwal: So, when we talk about, handcrafted features, we are actually talking about, machine…
[00:11:31] Durga Toshniwal: And, however, when we talk about deep learning models, the model actually learns the feature on its own, and then utilizes them to make the prediction.
[00:11:41] Durga Toshniwal: And, deep learning is very much applicable on complex data, such as audio, video, and all that.
[00:11:48] Durga Toshniwal: And the deep learning algorithms learn from huge amounts of data via the experience gained from the data, rather than a lot of manual tuning, which is done only in the case of machine learning.
[00:12:03] Durga Toshniwal: So, and the more the number of layers, or the deeper the network is.
[00:12:08] Durga Toshniwal: The more the understanding of the abstract
[00:12:14] Durga Toshniwal: Abstractness of the object that we want to recognize is, or whatever we want to predict.
[00:12:27] Durga Toshniwal: And we talked about some analogies in case of deep learning. Say, for example, if we talk about images, then the initial layers would actually identify the image boundaries or edges.
[00:12:40] Durga Toshniwal: the intermediate layers would actually identify or connect the edges into some parts, and then the rest of the layers, which are more deep, would actually identify and connect these parts into objects like that. So, basically, deep learning involves step-by-step
[00:12:59] Durga Toshniwal: Learning or understanding of the data.
[00:13:02] Durga Toshniwal: And, it's able to learn the complex pattern.
[00:13:06] Durga Toshniwal: And represent this complex pattern with the help of large number of layers, or with the depth.
[00:13:13] Durga Toshniwal: Using hierarchical representation of the concepts, which I'll just now discuss. For example, in case of objects, or maybe features on the face.
[00:13:23] Durga Toshniwal: We represent them in, in case, like, in a fashion of a hierarchy.
[00:13:29] Durga Toshniwal: Such that we have the boundaries, then the facial,
[00:13:34] Durga Toshniwal: The facial parts, like eyes, nose, ear, and then all of them connected to make up the face.
[00:13:41] Durga Toshniwal: We also talked about other problems that got associated with
[00:13:47] Durga Toshniwal: Shallow learning versus deep learning, so where we discussed that shallow learning relies on rules.
[00:13:53] Durga Toshniwal: So, something like rules, and such rules are static in nature.
[00:13:58] Durga Toshniwal: But, many times, there are a lot of dynamic conditions which are very different from the expected,
[00:14:07] Durga Toshniwal: Expected, results.
[00:14:09] Durga Toshniwal: And we talked about the image of a cat, where the cat, looked a little, more fluffier because of a large number of
[00:14:19] Durga Toshniwal: a very long fur on it, or the image was blurry, or the image was not a regular image of a cat standing, it would be a cat lying down, and so on and so forth. In all such cases, the standard dual sets fail.
[00:14:33] Durga Toshniwal: And, the machine learning, or the shallow algorithm is not able to.
[00:14:38] Durga Toshniwal: identify the organism in the image. And in such cases, it is the deep learning algorithms that come into picture, because they, obtain an understanding of the concept
[00:14:53] Durga Toshniwal: That is being represented by the image if we are talking about image data.
[00:14:59] Durga Toshniwal: So we… I gave you several examples, including the one that you're seeing on the slide.
[00:15:06] Durga Toshniwal: And,
[00:15:12] Durga Toshniwal: And then, we also talked about how, for example, how images are perceived by shallow learning algorithms versus deep learning algorithms. So the shallow learning algorithms would actually, in massage.
[00:15:28] Durga Toshniwal: an image in form of a lot of pixels, where the value of the pixel is just a number. So, it would be a large number of numbers that could be organized in a vector form.
[00:15:40] Durga Toshniwal: But there is no context or meaning associated with these numbers, which are nothing but representing the grayness in the image if the image is grayscale, or it would be RGB intensity values if it is in RGB. That is red, green, and blue colors.
[00:15:58] Durga Toshniwal: Whereas the deep neural networks or the deep networks would actually try to understand the concept represented in the image by looking at the edges or the boundaries, then combining them into shapes, then adding the shapes together and making up the problem.
[00:16:17] Durga Toshniwal: Then we talked about the different applications, domains where, ANNs are useful.
[00:16:23] Durga Toshniwal: And, some of the important ones included bioinformatics, image processing, pattern recognition, signal processing, prediction, and so on and so forth. There are lots of them.
[00:16:35] Durga Toshniwal: Then we talked about what is the basic, makeup of ANN and how they resemble the human brain.
[00:16:42] Durga Toshniwal: In which we discussed that the human brain is actually made up of neurons, which comprise of nucleus, of course, and
[00:16:51] Durga Toshniwal: So, neurons are connected to each other with the help of exon, and the connections between two, between the, two neurons is represented by synapse.
[00:17:04] Durga Toshniwal: So, similarly, in a ANN, artificial neural network, we have got neurons, which are also called nodes, and nodes are connected to each other with the help of edges, which are similar to the synapse.
[00:17:18] Durga Toshniwal: And just like in the brain, all neurons are connected to all others. In this particular case, in case of ANN, the neurons of one particular layer are connected
[00:17:30] Durga Toshniwal: To, the next layer.
[00:17:33] Durga Toshniwal: They're not connected one to all, but they are connected from one layer to the other. And if all the neurons in a preceding layer are connected to all other neurons in the successive layer, then it is called a fully connected layer.
[00:17:51] Durga Toshniwal: Then we discussed how, inputs, when they are given, propagate through the ANNs, or neural networks.
[00:17:58] Durga Toshniwal: And we discussed that each of the edges from the input layer to the hidden layers are actually containing widths, and the inputs are actually combined with the widths, and then these are transferred to the hidden layer nodes.
[00:18:17] Durga Toshniwal: So, for example, here, what you can see here, that there is an input layer, and the input layer is comprising of n number of inputs, like X1, X2, X3 till XN, and they are getting multiplied by the weights W1, W3 till WN.
[00:18:32] Durga Toshniwal: And then, from there.
[00:18:35] Durga Toshniwal: They are going to the processing unit, which contains the activation function after getting summation.
[00:18:42] Durga Toshniwal: And we obtain the output, and the output can be, the prediction on the class label Y.
[00:18:50] Durga Toshniwal: Just a second.
[00:18:55] Durga Toshniwal: So then, as we discussed yesterday, these X1, X2, X3 till XN could be single numbers.
[00:19:02] Durga Toshniwal: or they could actually be associated with vectors, n-dimensional vectors. So, for example, if we are having the input being represented in form of a n-dimensional vector.
[00:19:17] Durga Toshniwal: Then in that particular case, each of these X1, X2, X3 would be n-dimensional vectors.
[00:19:23] Durga Toshniwal: So, we are having X generally in the form of 1 cross N, and W is in the form of N cross 1, and when we multiply the two, what we get is a 1 cross 1.
[00:19:34] Durga Toshniwal: And this is what is the output that we get, and this is a single number or a prediction.
[00:19:42] Durga Toshniwal: So 1 cross n, means that we are having, a row Say, excise a row.
[00:19:50] Durga Toshniwal: Or a vector having a single row.
[00:19:53] Durga Toshniwal: And it is having how many? N number of columns. So, the number of columns are C1, C2, C3 till CN.
[00:20:00] Durga Toshniwal: And these are the different attributes.
[00:20:03] Durga Toshniwal: Say, for example, a simple example could be, let's say we are measuring the temperature.
[00:20:11] Durga Toshniwal: Oh.
[00:20:14] Durga Toshniwal: I'll put it on. So, let's say we are measuring the temperature of a place. So, temperature at a place, P1.
[00:20:23] Durga Toshniwal: Could actually be represented by the temperature
[00:20:27] Durga Toshniwal: Taken over Day 1, Day 2, Day 3 till?
[00:20:31] Durga Toshniwal: Let's say it was 1 month, so, it was day 30.
[00:20:36] Durga Toshniwal: So that could be your first input, which is your X1. Then, the temperature for some other place is also measured over this month.
[00:20:44] Durga Toshniwal: D1, D2, D3 till D, 30, and it will have values like some temperature 20, 21, something, something 19. This will also have some values, like that. So this could be your X1, this could be your X2, like that. We have got some n number of…
[00:21:02] Durga Toshniwal: locations or places over which the temperature has been measured for a month. This is just an example. Now, this is given as an input, and then, the weight, weights are considered in the N cross 1 fashion.
[00:21:16] Durga Toshniwal: And the two are, multiplied.
[00:21:20] Durga Toshniwal: And then the… whatever we obtained is summed together with the help of the summation, and then fed into the activation function. So we have X multiplied by V8,
[00:21:32] Durga Toshniwal: And all these Xi's, WIs are multiplied and added from 1 to N, and this input is given to the activation function.
[00:21:42] Durga Toshniwal: And this is a one-cross one. So, in a way, we are transforming the given data from one dimensional plane to some other dimensional plane. Initially, it was a dimensional plane, which is one cross end, and what we obtain is one cross one.
[00:21:59] Durga Toshniwal: So, we did some explanations in this regard, for example, representing an image and then using it for the calculation.
[00:22:06] Durga Toshniwal: So, so far, we had discussed about only the product of the input and the weights and the activation function, but there's also something called bias, so that we can have a better control over what is the output of the activation function.
[00:22:22] Durga Toshniwal: So the bias term is, used. The bias, actually, you can think of it as, that it gets applied to each and every node.
[00:22:31] Durga Toshniwal: Or, in general, it gets applied to the entire layer.
[00:22:36] Durga Toshniwal: And then, since we talked about the information getting flowing from left to right, right, the inputs getting multiplied, the weights going to the summation, then going to the processing unit containing the activation function, then going towards the output, so all this is called the forward pass.
[00:22:56] Durga Toshniwal: However, initially the rates are assigned randomly, and we want to assign them in a better fashion so that the output is similar to the actual output.
[00:23:08] Durga Toshniwal: So what we have is, for example, we randomly initialize the weights, and then do the dot product between the input and the weight matrix.
[00:23:18] Durga Toshniwal: And then we pass it to the activation function, and then the output is generated. And the generated output is compared with the actual output which is available with us in the training data, and see how much of a difference is there.
[00:23:32] Durga Toshniwal: So, the difference is nothing but the error in the prediction, and this error is back-propagated.
[00:23:38] Durga Toshniwal: To, towards the layer which is the closest to the…
[00:23:43] Durga Toshniwal: output layer. This layer would propagate it further downwards, then further downwards, like that, till it reaches the input, layer.
[00:23:53] Durga Toshniwal: That is the first hidden layer. Or, in other words, we actually back-propagate. Backpropagate means, whatever is propagating is flowing backwards from the right to the left.
[00:24:06] Durga Toshniwal: In this case, it is the error, and the summation of errors over all the records in the training data is in totality called the loss. So there are a large number of loss functions, and we want to minimize the loss, which is nothing but minimization of the errors, and for that, we will perform the backward pass.
[00:24:26] Durga Toshniwal: In which the error will be back-propagated, and then the weights will be altered, and then, once again, forward pass will happen.
[00:24:33] Durga Toshniwal: and the prediction is done again, then the prediction is compared with the actual value, and then we see whether it is correct or not. If it is correct, fine. If it isn't, then again back-tropagate the new error, and so on and so forth.
[00:24:47] Durga Toshniwal: So, this is what it discussed. Then we talked about the MLP, which is the multi-layer perceptron.
[00:24:53] Durga Toshniwal: So, the multi-layer perceptron comprises of one hidden layer, an input layer, output layer, and it's,
[00:25:00] Durga Toshniwal: It's one of the simpler forms.
[00:25:04] Durga Toshniwal: Then, So the artificial neural network, or the ANN, actually does a nonlinear transformation of the input
[00:25:15] Durga Toshniwal: And additionally, it also changes the dimensionality of the input from one vector space to the other, as we already mentioned. Initially, it was 1 cross n, then it got multiplied by N cross 1, and finally, it got mapped to a 1 cross 1 space. So, from a 1 cross N, this is the input.
[00:25:35] Durga Toshniwal: This is the weights.
[00:25:38] Durga Toshniwal: And this is the output.
[00:25:39] Durga Toshniwal: So, from here, it is getting transformed to a new dimensional plane, which is 1 cross 1. So, this is one thing. Secondly, the…
[00:25:49] Durga Toshniwal: Data is also nonlinearly transformed, which I will discuss with you shortly, how it becomes nonlinear. As of now, you can remember that the activation function is nonlinear, and so when the data passes through it, then it also becomes nonlinear.
[00:26:07] Durga Toshniwal: Then we talked about several examples of how, the forward propagation or the forward pass happens.
[00:26:16] Durga Toshniwal: And, I gave you examples in which the… so basically, when we try to flow… pass this information from the input layer to the output layer via the hidden layers, it is leading to a chain of matrix multiplications, as we discussed yesterday.
[00:26:33] Durga Toshniwal: Where we have the input, which is a 1 cross N, then we have the hidden layer weight matrix, which is 3 cross 4, or N.
[00:26:40] Durga Toshniwal: Then we have the next one, which is 4 cross 4, and then 4 plus 1, to bring it back to the original dimensionality. And finally, the answer is 1 by 1.
[00:26:51] Durga Toshniwal: So, I showed to you yesterday how inputs play a role. As you can see here, if X1, X2, X3, these are the inputs to this process. Then they get multiplied.
[00:27:01] Durga Toshniwal: And then add it together. For example, if we talk about
[00:27:05] Durga Toshniwal: node N1, N2, N3, N4, then N1 would contain, part of X1, it would also contain part of X2, and then part of X3.
[00:27:16] Durga Toshniwal: So, we are having these parts here, so all together, this makes up the first one, so…
[00:27:23] Durga Toshniwal: We go like this.
[00:27:25] Durga Toshniwal: Then, similarly, we can, find out for the other nodes also. So this is what I've written here also, X1W11 plus X2W21 plus W3… X3W…
[00:27:38] Durga Toshniwal: 31 and W21, and so on and so forth. So, at each node, there will be a weighted
[00:27:44] Durga Toshniwal: Component of the initial inputs, and all these may be added together, and then something, something can be done.
[00:27:54] Durga Toshniwal: This is a very simple example that I had given you, so I request you all, please, not to draw on the screen.
[00:28:01] Durga Toshniwal: I can see that somebody's drawing.
[00:28:08] Durga Toshniwal: Now, the issue is that when I do a stop share, then I'll have to… this will stick on my slides, and it will come on each and every slide.
[00:28:31] Durga Toshniwal: Now, it's going to come on each and every slide, till I disconnect, and then I have to rub it with an eraser.
[00:28:38] Durga Toshniwal: So then I think some of you had a question that, okay, at N1, N2, N3, and N4, these are the kind of calculations that happen to the, with respect to X1, X2, X3.
[00:28:51] Durga Toshniwal: So, a portion of X1, that is X1.w11, reaches node N1. Similarly, X2W21 reaches N1, and X3W31 reaches N1, and the sum total is what we obtain at W1.
[00:29:05] Durga Toshniwal: at N1. Similarly, we do it for N2, N3, and N4. And once we obtain that, and we have another hidden… suppose we have another hidden layer, then we can take the corresponding weight matrix, and then multiply
[00:29:18] Durga Toshniwal: the components of N1, N2, N3, and N4 with T1, and what we'll finally have is what is going to be obtained at the next node.
[00:29:29] Durga Toshniwal: Whatever it is.
[00:29:32] Durga Toshniwal: So this is how the calculations are done.
[00:29:34] Durga Toshniwal: And now we'll proceed with the further things.
[00:29:38] Durga Toshniwal: Okay, before we start today's portion, are there any questions?
[00:29:44] Durga Toshniwal: Anyone, any questions up to now?
[00:29:57] Durga Toshniwal: So I'm just switching off my video so that we don't face any problems like we did yesterday.
[00:30:03] Durga Toshniwal: Okay, if there are no questions, I assume that it is quite clear to all of you, and we'll proceed.
[00:30:08] Ravindra Singh: If you don't mind, ma'am, can you put a little bit on propagate things? Like, how we will give the input again?
[00:30:17] Durga Toshniwal: Poe?
[00:30:18] Ravindra Singh: appointment.
[00:30:18] Durga Toshniwal: error.
[00:30:19] Ravindra Singh: Yeah.
[00:30:20] Durga Toshniwal: Error will backpropagate, you're saying?
[00:30:22] Ravindra Singh: Yes, yes.
[00:30:24] Durga Toshniwal: Yeah, so actually, the back propagation of error is a long derivation.
[00:30:28] Durga Toshniwal: Which I will cover in the subsequent lectures, because the back propagation of error, that calculation and the updation of the weights on that basis is a very, it's quite time-consuming and long, so it will take at least one session.
[00:30:43] Durga Toshniwal: So today, yeah, so today I'm not covering, as of now, you just assume that the error will be back propagated, and then this will be used to recalculate the weights.
[00:30:54] Durga Toshniwal: And then, once again, the data will be, moved in the forward direction, in the forward pass, and then the output will be generated, comparison between the actual output and the expected output will happen.
[00:31:08] Durga Toshniwal: And then the loss is, or the error is again back-propagated, so this will go on and on. However, exactly how it happens, I will explain with the math in some of the upcoming lectures.
[00:31:20] Ravindra Singh: Okay.
[00:31:21] Durga Toshniwal: Yep.
[00:31:22] Durga Toshniwal: Any other questions, anyone?
[00:31:28] Durga Toshniwal: Okay, if there are no further questions, we'll just proceed.
[00:31:40] Durga Toshniwal: So then,
[00:31:43] Durga Toshniwal: So, whenever we want to make a prediction using an ANN, what we require is the input.
[00:31:49] Durga Toshniwal: Then we also need to have an idea about the weights and the activation function.
[00:31:54] Durga Toshniwal: That we are going to use. Additionally, we should also use, know what would be the value of the bias.
[00:32:00] Durga Toshniwal: And, then, what happens is, once we, initialize the weights randomly, Then we use it.
[00:32:08] Durga Toshniwal: To make a prediction using the activation function and all these things that has been decided. And if the predicted value is not very close to the initial value.
[00:32:19] Durga Toshniwal: Then we just, backpropagate and, try to iterate on different things, like, the weights, the initial weights, what they can be, and so on and so forth.
[00:32:32] Durga Toshniwal: Okay.
[00:32:33] Durga Toshniwal: So, we have, something like this, that we have the input layer, then we have weights, then we have the hidden layer, another hidden layer, and then we have the output.
[00:32:51] Durga Toshniwal: I think I'll do a stop share and then share again, because there's this green mark coming all over.
[00:32:57] Durga Toshniwal: It's creating a… Nuisance on the slides.
[00:33:01] Durga Toshniwal: I'll just remove it, and then we can discuss it back again.
[00:33:08] Durga Toshniwal: Give me a minute, please.
[00:33:16] Durga Toshniwal: he's alone.
[00:33:30] Durga Toshniwal: Okay, I'm just sharing it back again.
[00:33:50] Durga Toshniwal: So then, if you talk about a fully connected layer.
[00:33:56] Durga Toshniwal: And you have one hidden layer, Another hidden layer.
[00:34:01] Durga Toshniwal: And then you have the input layer.
[00:34:04] Durga Toshniwal: And the output layer.
[00:34:07] Durga Toshniwal: This is what you are having.
[00:34:09] Durga Toshniwal: And the input is having 3 notes.
[00:34:13] Durga Toshniwal: And this is having 4 nodes.
[00:34:16] Durga Toshniwal: And,
[00:34:17] Durga Toshniwal: So, what we'll have to do is we'll have to choose this based on the 3 nodes we are having, a 1 plus 3 matrix. Then, to allow them to get multiplied, we'll have to represent it in a 3 plus 4 fashion.
[00:34:30] Durga Toshniwal: And after that, we will have, 4 cross 1.
[00:34:36] Durga Toshniwal: And then this will be 1 cross 1. This is how we will do the dot product.
[00:34:42] Durga Toshniwal: Okay.
[00:34:43] Durga Toshniwal: Now that we have, talked about how the input propagates, and then we also discussed that
[00:34:51] Durga Toshniwal: The nonlinearity in this whole process is brought about by…
[00:34:56] Durga Toshniwal: Activation functions. And we discussed yesterday that if there are multiple layers, then we could actually have multiple activation functions, which may be same or different.
[00:35:05] Durga Toshniwal: So, let's assume we are having activation, function alpha 1, okay?
[00:35:12] Durga Toshniwal: So, so what we have is, the first activation function. This first activation function,
[00:35:21] Durga Toshniwal: It's going to, utilize the data that it is receiving.
[00:35:26] Durga Toshniwal: Now, what is the data that is getting received?
[00:35:29] Durga Toshniwal: So it goes like this.
[00:35:33] Durga Toshniwal: So, at the first hidden layer, this is the first hidden layer, we have for, the…
[00:35:45] Durga Toshniwal: So we have Y, which is the expected outcome, as a function of
[00:35:52] Durga Toshniwal: So, if we talk about this layer, let's say,
[00:35:56] Durga Toshniwal: Just a second. So, we are talking about, let's say, the last one.
[00:36:01] Durga Toshniwal: The last layer.
[00:36:03] Durga Toshniwal: So, here, we are talking about the output here. So, we are having Y, and Y is a function of what? It is a function of…
[00:36:11] Durga Toshniwal: Whatever is coming here, along with the weight 3.
[00:36:16] Durga Toshniwal: So, what is coming here? So, here…
[00:36:19] Durga Toshniwal: The input is coming, and the activation function will be used on it, which is your alpha 2.
[00:36:27] Durga Toshniwal: So, we'll have function, function of, X times, sorry, alpha 2 times weight 3.
[00:36:36] Durga Toshniwal: So, what we are having here, we are having the activation, Function, alpha 2.
[00:36:43] Durga Toshniwal: sorry, we are having the product as alpha 2, and then we are multiplying it by the weight matrix, which is W3.
[00:36:51] Durga Toshniwal: Now, if we look at alpha 2, actually, So, what is Alpha 2?
[00:36:57] Durga Toshniwal: Alpha 2 is nothing but…
[00:36:59] Durga Toshniwal: So, alpha 2 is obtained here. So, what is it, actually? Function of… what we are getting here is alpha 1 times W2.
[00:37:09] Durga Toshniwal: So, this is our W2.
[00:37:11] Durga Toshniwal: And we are having inputs as alpha 1, so it will be alpha 1 times W2.
[00:37:17] Durga Toshniwal: So now, if we, look at, alpha 1 itself, then what is Alpha 1?
[00:37:26] Durga Toshniwal: Alpha 1 is… so, we are having
[00:37:29] Durga Toshniwal: This is alpha 2, this will be alpha 1.
[00:37:32] Durga Toshniwal: So, alpha 1 is nothing but a function of what it receives, dot with whatever weight matrix is. Weight matrix is W1 here, and what it is receiving is the input X here.
[00:37:45] Durga Toshniwal: It is receiving the inputs X only.
[00:37:48] Durga Toshniwal: So, it will be X.W1.
[00:37:53] Durga Toshniwal: So then, we have, already, as I told you, we have Y will be a function of…
[00:37:59] Durga Toshniwal: alpha 2 times W3.
[00:38:02] Durga Toshniwal: But what is Alpha 2?
[00:38:04] Durga Toshniwal: Alpha 2 is function of alpha 1 times W2, so this becomes function of…
[00:38:11] Durga Toshniwal: Function of… what is alpha 2? It is function of alpha 1 times W2.
[00:38:17] Durga Toshniwal: times W3. This is what it is, and this is your alpha 2.
[00:38:22] Durga Toshniwal: But Alpha 1 again is coming here, so what is Alpha 1?
[00:38:26] Durga Toshniwal: Alpha 1 is what?
[00:38:28] Durga Toshniwal: Alpha 1 is function of X times W1. X is the initial input, so it will be
[00:38:36] Durga Toshniwal: F off?
[00:38:38] Durga Toshniwal: X times W1.
[00:38:41] Durga Toshniwal: Times.
[00:38:42] Durga Toshniwal: W2?
[00:38:46] Durga Toshniwal: times W3.
[00:38:50] Durga Toshniwal: And then, if we try to open this.
[00:38:53] Durga Toshniwal: Or we can just think of it like this. So, the expression that we obtained is this one.
[00:39:01] Durga Toshniwal: We have, activation function being applied at the dot product of X and W1, then the product of X dot W1 and W2,
[00:39:12] Durga Toshniwal: is used to feed into the next activation function, which is again F. And the third one will draw upon what we have arrived at.
[00:39:22] Durga Toshniwal: with the… with the activation function earlier, along with weight. So, this is how we obtain this expression. This is the expression that you see, okay?
[00:39:34] Durga Toshniwal: And, this is how the, nonlinearity brought in about the… brought in into the neural network by activation works.
[00:39:46] Durga Toshniwal: Okay.
[00:39:48] Durga Toshniwal: So, we, saw that, The ANN,
[00:39:56] Durga Toshniwal: Perform some nonlinear transformations on the input.
[00:40:00] Durga Toshniwal: And,
[00:40:02] Durga Toshniwal: So, it does two things. One, it maps or transforms the data from one dimensional space to some other. As I mentioned, 1 cross n is actually mapped to 1 cross 1.
[00:40:13] Durga Toshniwal: So, here we are doing a transformation on the data, and this transformation is taking the data from one particular vector space to the other. That is one thing that
[00:40:23] Durga Toshniwal: The second thing is that, suppose we are doing a classification problem, and we want to differentiate between the different classes with the help of a decision boundary.
[00:40:34] Durga Toshniwal: So what happens is that,
[00:40:38] Durga Toshniwal: This data, usually, which is used in ANNs, may be nonlinear in nature.
[00:40:45] Durga Toshniwal: So, if it is nonlinear in nature, then the… Input data.
[00:40:51] Durga Toshniwal: directly is not, separable in its given form, because it is nonlinear, and it is not, so it is not linear, it is nonlinear, so we cannot use it directly and separate it out.
[00:41:03] Durga Toshniwal: So, what we could do as a solution is that we could perform some nonlinear transformations on the data in such a way that the input data is projected from the original space to a new vector space.
[00:41:20] Durga Toshniwal: And in this new vector space, then we can use the decision boundary to separate you know, the…
[00:41:28] Durga Toshniwal: So, to separate the different classes that are available in the data. So, it will be more clear when I explain to you diagrammatically right now.
[00:41:39] Durga Toshniwal: So, what we are having is a visualization of the data. Let's say our data is having two classes, say, C1 and C2.
[00:41:47] Durga Toshniwal: One is represented in the blue color, another is represented in the red color.
[00:41:52] Durga Toshniwal: And now we are given this data, and this data, as you can see, is not linearly separable. It is nonlinear in nature.
[00:41:59] Durga Toshniwal: And, So, we want to actually… so, how should we solve the problem?
[00:42:07] Durga Toshniwal: Of this kind of data, where the decision boundary is not a linear one, it is a nonlinear one.
[00:42:14] Durga Toshniwal: then what should we do? So, the solution, too, is that
[00:42:19] Durga Toshniwal: One solution to it is that what we do is we project this data into a higher dimensional space.
[00:42:26] Durga Toshniwal: And how do we do it? We do it with the help of nonlinear transformations. And then, when we project it to a high-dimensional plane.
[00:42:35] Durga Toshniwal: Then, the… The data gets separated in such a way that it becomes linear, linearly separable.
[00:42:43] Durga Toshniwal: And, you can see here that, the green boundary
[00:42:47] Durga Toshniwal: Shows that the data has been, separated linearly with the help of this particular kind of boundary.
[00:42:56] Durga Toshniwal: Okay?
[00:42:58] Durga Toshniwal: So, now what happens is that, initially, the data was of a lower dimension.
[00:43:03] Durga Toshniwal: And it was not linearly separable. Then we actually fed it as an input the activation function.
[00:43:10] Durga Toshniwal: And, then what we tried to do is map it to a very high-dimensional space.
[00:43:16] Durga Toshniwal: So, what would happen? The points would get, sparsified, they would move away.
[00:43:23] Durga Toshniwal: from what their original locations were. So when you are mapping it to a higher dimensional, space, then the points would actually move, like what you are seeing here.
[00:43:34] Durga Toshniwal: So, you will have…
[00:43:36] Durga Toshniwal: you'll perform a nonlinear transformation on the given data. The points will become more sparse and linearly separable with the help of this kind of decision boundary. And now when we… and now the points are separable.
[00:43:50] Durga Toshniwal: And if we try to project these two, we'll be able to, you know, project them and identify this kind of an image.
[00:43:58] Durga Toshniwal: separation, or a boundary. So, here we had a nonlinear boundary.
[00:44:04] Durga Toshniwal: Then we wanted to make it linearly separable. To make it linearly separable, we mapped the data into a much higher dimensional space.
[00:44:13] Durga Toshniwal: Such that the… the data now becomes linearly separable with respect to each other, that is the different classes. And then now, this one is given as input.
[00:44:24] Durga Toshniwal: 2R… Algorithm?
[00:44:27] Durga Toshniwal: And when this is given as input, then the algorithm is actually able to identify the nonlinear boundary with the help of the understanding it has gained with the
[00:44:39] Durga Toshniwal: Deep learning model, okay?
[00:44:41] Durga Toshniwal: So this is how, the…
[00:44:46] Durga Toshniwal: The decision boundaries, which are complex, can help to reproduce the complex data, or to work on the complex data in a better fashion.
[00:44:58] Durga Toshniwal: So, to summarize, artificial neural networks are very flexible, and
[00:45:05] Durga Toshniwal: They can model very complex functions, and therefore, they're also called universal function approximators, because they can easily model some of the simple and complex ones, both.
[00:45:20] Durga Toshniwal: And, due to the increase in the computational power recently, because of GPUs being available and other things, so, ANNs are becoming very, very, very popular.
[00:45:35] Durga Toshniwal: Okay.
[00:45:37] Durga Toshniwal: So, I already told you, we have an input layer, let's say this is the input layer, then we have the hidden layers.
[00:45:47] Durga Toshniwal: Like this, let's say.
[00:45:49] Durga Toshniwal: These are the hidden layers, and this is the output.
[00:45:53] Durga Toshniwal: Now, the thing is that, if we have a set of input, and we want to create an output.
[00:45:59] Durga Toshniwal: then what can we do? We can initialize the weight and all, and after initializing the weight and, having a fully connected layer.
[00:46:08] Durga Toshniwal: and I'm not showing all the connections, then what are we doing? We're trying to actually propagate the weighted
[00:46:16] Durga Toshniwal: Some of the products of the input and the weights, we take this, and then we propagate it further, and so on and so forth, till the output matches with the actual output. Okay.
[00:46:27] Durga Toshniwal: So now, suppose there is error in the output, then do we have something over which we can exercise control and change it?
[00:46:36] Durga Toshniwal: So that we can obtain a better output, right? We want to, when we have, actually, propagated data into the neural network, we get an output.
[00:46:48] Durga Toshniwal: Which is the prediction that has been done using the neural network. But whatever the value is,
[00:46:55] Durga Toshniwal: So, what we want is that, it will be giving some output, but what we want is that if we can control this output in another… in some other way, in addition to what we have studied. So,
[00:47:08] Durga Toshniwal: So if we could have, more control, then it will be nice, so that we can actually try to…
[00:47:14] Durga Toshniwal: change the output so that it becomes better, okay? Say, for example, in this figure, you see a box, it has an input, and it has two outputs.
[00:47:27] Durga Toshniwal: And it has a couple of knobs over it. Here, 5 knobs are shown.
[00:47:31] Durga Toshniwal: So, if we want the output to come out properly.
[00:47:35] Durga Toshniwal: Then we… what we can do is adjust these knobs.
[00:47:39] Durga Toshniwal: In such a way that the output is very much similar to the input that has been given to the
[00:47:46] Durga Toshniwal: to the algorithm. Or in other words, we want to exercise some control in such a way that we can vary the inputs slightly so that the output we get is much better.
[00:48:00] Durga Toshniwal: So then, here you can see these are knobs, but actually, we don't have knobs, right, in our ANN. What we have are parameters, okay? So these are also called hyperparameters. We want to
[00:48:14] Durga Toshniwal: You know, have a control on the hyperparameters so that we can tune them to obtain good output, okay?
[00:48:22] Durga Toshniwal: So then, for example, if I gave you an example yesterday where we had images of apples and oranges, say 50-50 images, each apples, oranges, respectively, and we give it as an input to this process.
[00:48:36] Durga Toshniwal: And then see what has been predicted. Probably, if, let's say the distribution of apples versus bananas was 50-50, in the output, the ratio is coming to 30-70 apples to bananas. So this is incorrect. Then how do we solve this problem? Of course, we can vary the weights.
[00:48:55] Durga Toshniwal: And the bias, but this is only one control. But we want other controls also, because in the real world, the neural networks can be very deep. Then, how do we exercise this control if we are limited to just one or two hidden layers?
[00:49:10] Durga Toshniwal: So, if we have a free hand on how we can iterate to obtain a good output, then what we can do is that we can, you know, exercise control with the help of these knob-like structures.
[00:49:26] Durga Toshniwal: And these actually refer to the different parameters, or the hyperparameters that we can vary in order to obtain a good outcome, right?
[00:49:36] Durga Toshniwal: so then… Just a second.
[00:49:49] Durga Toshniwal: Just a second.
[00:49:53] Durga Toshniwal: Oh.
[00:49:57] Durga Toshniwal: So we… we… what we will do in the forward pass is that we'll feed the neural network with some image containing an object, say the object is an apple. So, we feed it to the
[00:50:09] Durga Toshniwal: ANN.
[00:50:11] Durga Toshniwal: And,
[00:50:12] Durga Toshniwal: When the output comes out, probably, let's say the ANN inferred data orange image has been… image of an orange has been given as input to it.
[00:50:24] Durga Toshniwal: Right? Or in other words, it gives an error. So, similarly, there could be error on other predictions being made. We sum them all together, and what we have is called the
[00:50:34] Durga Toshniwal: Some total of error. And the… we have something called cost, where cost refers to, when we talk about a neural network.
[00:50:45] Durga Toshniwal: Then the, the amount of times incorrect answers came
[00:50:51] Durga Toshniwal: from the neural network, we call it cost of the… that particular, function or method. So, if there is high error, then the sum total of errors coming to each and every,
[00:51:03] Durga Toshniwal: Node will be high.
[00:51:05] Durga Toshniwal: So, or in other words, there'll be high amount of, error. And if, you know, the errors are lower, then,
[00:51:16] Durga Toshniwal: what we have is a low overall loss, low loss, or low error, okay? So we have a function called… so now, how do we actually, use the cost function or the loss?
[00:51:30] Durga Toshniwal: To iterate upon it in such a way that the actual values and the predicted values are quite close together.
[00:51:39] Durga Toshniwal: So how do we do that? How do we adjust the values and, you know, go further? The answer is that we can easily use a cost function, that the cost function may not be very simple and generic over multiple data sets or runs. So, what do we do?
[00:51:57] Durga Toshniwal: So, what we do is that, you know, we have,
[00:52:03] Durga Toshniwal: In order to reduce the error, we have something called gradient descent. So, gradient descent is a, is a method that tries to minimize the cost involved.
[00:52:14] Durga Toshniwal: by adjusting the different hyperparameters, okay? So, we just performed the gradient descent, and the gradient descent, if it leads to reduction in the loss, which is the overall error, then we say that gradient descent is going in the direction of reducing the errors, and then.
[00:52:41] Sacheen Adavinavar: Mami went on…
[00:52:44] Durga Toshniwal: Yeah, sorry for that.
[00:52:46] Durga Toshniwal: So then, the gradient descent actually helps to minimize the function.
[00:52:51] Durga Toshniwal: And it actually does so. What it does is that when the errors are back-propagated, the gradient descent will help us to minimize the loss.
[00:53:00] Durga Toshniwal: And if it goes in the direction of minimization, it will go on and proceed to minimize it further, and then…
[00:53:12] Durga Toshniwal: And then, again, it, you know, minimized it further and further.
[00:53:18] Durga Toshniwal: We, what we obtain is an optimally low value of the, cost.
[00:53:23] Durga Toshniwal: Or the loss.
[00:53:24] Durga Toshniwal: So, gradient descent is a method, is a mathematical method to find out the, to optimize the loss to a minimal value. So, it is an optimally minimal value. It is not the final, or it is not the least value.
[00:53:41] Durga Toshniwal: So, gradient descent helps us to do that.
[00:53:44] Durga Toshniwal: And, so… In a neural network.
[00:53:48] Durga Toshniwal: So, how gradient descent works, I will explain to you separately in another lecture, because it is… it involves a lot of things. As of now, you can assume that we want to minimize the loss
[00:54:00] Durga Toshniwal: That is the error. So what is error? Error is the difference between the predicted, value of the target variable versus the actual given value.
[00:54:11] Durga Toshniwal: And when we sum total the errors for all the records in the training data, what we have is the loss. And we want to minimize the overall loss, which is the sum total of the error on that particular model.
[00:54:24] Durga Toshniwal: So, to minimize the loss, we should have some method, and the method that we use is gradient descent. So, gradient descent helps us to minimize the loss.
[00:54:33] Durga Toshniwal: And, yeah, so gradient descent could actually lead to minimization or maximization, any one thing.
[00:54:41] Durga Toshniwal: So, we have to tune it in such a way that gradient descent is leading to the minimization of the loss, not to the maximization of the loss.
[00:54:49] Durga Toshniwal: Right? And then, and what it does is gradient descent, that it's going to adjust the parameters that are available, in the neural network.
[00:55:01] Durga Toshniwal: And by adjusting the parameters, we are expecting that the gradient descent is going to move downhill. When it… we say move downhill, it means that it is going to move in a direction where the loss is decreasing.
[00:55:17] Durga Toshniwal: And so on and so forth, and then finally, the final optimally minimal loss is obtained with the help of gradient descent, and then we use that particular model, and the values of the hyperparameters at that point in time are the final hyperparameter values.
[00:55:38] Durga Toshniwal: So then we have, just a second…
[00:55:55] Durga Toshniwal: Okay, so then what we are doing here… so, we are having, in the forward pass, we are having these inputs, for example, X1,
[00:56:05] Durga Toshniwal: X2, and X3. These are the inputs.
[00:56:08] Durga Toshniwal: Then, these inputs are weighted with the weight matrix containing weights W1, W2, W3, 10W, and
[00:56:17] Durga Toshniwal: And it's a fully connected layer, so all the inputs are going to
[00:56:22] Durga Toshniwal: All other notes, like this, like this, like this, like this, like this. And these receive weighted, inputs, and then sum them up.
[00:56:30] Durga Toshniwal: And then,
[00:56:32] Durga Toshniwal: This for… then the activation function works here, maybe alpha 1. Then again, the weights are exercised.
[00:56:41] Durga Toshniwal: And then, finally, we come to this point where, again, another activation function might work on it, and then finally, we have the output.
[00:56:50] Durga Toshniwal: And so on and so forth. So, the output that we are obtaining is, let's say, Y hat.
[00:56:58] Durga Toshniwal: So, there will… there might usually be a difference between the actual value
[00:57:04] Durga Toshniwal: And a predicted value. So why is the actual value?
[00:57:08] Durga Toshniwal: Of the, for the prediction, and Y hat is the predicted one.
[00:57:15] Durga Toshniwal: And the difference between Y and Y hat is said to be the error.
[00:57:19] Durga Toshniwal: Right? And the sum total of errors that we obtained in this case on X1, X2, and X3, because there are only 3 inputs right now. So, the sum total of all these is called the loss.
[00:57:32] Durga Toshniwal: And, what we try to do is that, we try to… so this was a forward pass, everything was flowing from this direction to this, to this, like this. Now, what we'll do, we'll look at the difference between the actual value and the predicted value.
[00:57:48] Durga Toshniwal: Which is called the loss, and we'll propagate the loss in the backward direction.
[00:57:53] Durga Toshniwal: And what we'll try to do is we'll try to study
[00:57:56] Durga Toshniwal: What led to this loss, or this error? So, how do we do it?
[00:58:03] Durga Toshniwal: What we tried to do is that,
[00:58:06] Durga Toshniwal: So, let's say… I'll once again write these. This is X1, this is X2, this is X3.
[00:58:14] Durga Toshniwal: These are some weights, W1, W2, W3, like this.
[00:58:19] Durga Toshniwal: And then we get the weighted sum here, and then we have the activation function working on them.
[00:58:26] Durga Toshniwal: And then the outputs with some other weights are there.
[00:58:30] Durga Toshniwal: Like this. And then we have another activation function, and then some other weights, like this, and finally, what we have is the output.
[00:58:38] Durga Toshniwal: And then the output is Y-hat, and we compare it with the given output.
[00:58:44] Durga Toshniwal: And C, what is the difference? And that difference is the error that we are going to back-propagate.
[00:58:51] Durga Toshniwal: Okay.
[00:58:53] Durga Toshniwal: So then, here, to backtropagate, what we do is that we calculate
[00:59:00] Durga Toshniwal: The slope of the last layer, that is the…
[00:59:04] Durga Toshniwal: So what we do is we try to find out the error.
[00:59:09] Durga Toshniwal: Right? And this error is back-tropagated, so we cannot find out what is happening here right now, because we are still at the output layer. So, to see what is causing the error, we first back-tropagate it to whatever is nearest to it.
[00:59:25] Durga Toshniwal: So, this layer is nearest, so the error is back-propagated to this particular layer.
[00:59:30] Durga Toshniwal: Once it gets back propagated, then what, what happens is that, this layer again tries to compare, let's say the output at this point is Y hat prying, it tries to compute the difference
[00:59:46] Durga Toshniwal: between the actual given value and the predicted value. And if there's a error, then what it will do, it will say that I have not created this error. This error has been created because of the inputs
[01:00:00] Durga Toshniwal: That I've received from the previous layer. So, what it will do, take this difference and back-propagate it to the next previous layer.
[01:00:08] Durga Toshniwal: This previous layer, what it will do is that it will look again at its own prediction, let's say this, and the actual value, and see what is the difference, which is the error here. And then what it will do,
[01:00:21] Durga Toshniwal: It will use this error, and then once again propagate it, because it says that what we have is a result of what is getting inputted here.
[01:00:30] Durga Toshniwal: Then, what will happen? The error has, back-tropicated to the last layer. Now, again, the weights will be readjusted.
[01:00:38] Durga Toshniwal: Okay, and after readjusting the weights, again, the forward pass will happen. That is, the weighted inputs, new weights are there, so new weighted inputs are inputted, then activation is used, then more weights are used, like that.
[01:00:54] Durga Toshniwal: So this is the forward pass, once again. Then again, the errors will be found out, and these errors will be back-tropagated from the output layer up to the initial input layer.
[01:01:05] Durga Toshniwal: And this goes on and on till the expected output and the actual output match, okay?
[01:01:11] Durga Toshniwal: And this is done with the help of, back propagation, working on, along with the gradient descent, okay?
[01:01:19] Durga Toshniwal: So, any questions so far?
[01:01:22] Durga Toshniwal: on this. Anyone, any questions?
[01:01:24] Deepak Katara: So, one is, we are adjusting weight, weights on… on basis of weight propagation, but how do we initialize weights?
[01:01:35] Deepak Katara: Like, what are the criteria?
[01:01:37] Durga Toshniwal: Yeah, so actually, we just… the basic algorithm, or the algorithm in its basic form just says that we randomly, initialize the bits.
[01:01:50] Durga Toshniwal: So, we just randomly initialize the widths. However, there are optimizers that are available to optimize the widths, but the general convention is we just randomly assign them.
[01:02:03] Durga Toshniwal: That is how it happens.
[01:02:05] Deepak Katara: But that might increase cost to a very extent if we don't initialize it right, because there would be, you know, multiple cycles where we will have delta of outcome that we should expect versus what we are getting, right?
[01:02:20] Durga Toshniwal: So what you're saying is correct, that if the weights are randomly initialized, then they may be very different from… and therefore, the output that we obtain may be very different from the expected or the actual given output.
[01:02:35] Durga Toshniwal: And to make this difference smaller between the given output, that is actual output, and the predicted output, we have to adjust the weights. And if the weights that has been assigned randomly are very different, then we'll need a lot more iterations.
[01:02:54] Durga Toshniwal: to converge to weights that are good enough for us. So that is true. That's why it is compute-intensive, actually.
[01:03:01] Durga Toshniwal: But in the basic form, there are no special algorithms to decide what the weights are.
[01:03:09] Deepak Katara: Yes, and one more thing, it's just not weights, which is variable. We have bias, we have this activation function. So these two things are also variable. So how do we know that we need to fine-tune only weight, not bias or activation function?
[01:03:25] Durga Toshniwal: So, actually, when we are doing the forward pass.
[01:03:28] Deepak Katara: So, we have the bias, we have the weights.
[01:03:32] Durga Toshniwal: We have the activation functions. There are many other hyperparameters also, like epochs, iterations, and so on and so forth.
[01:03:40] Durga Toshniwal: So, during the forward pass, what you will do is you'll try to iterate upon these.
[01:03:46] Durga Toshniwal: iterate upon the hyperparameters, like the number of neurons in the layer, number of layers, number of iterations, and so on and so forth. And then, when the output is generated, you compare it with the given output and see the difference, and then you'll back-propagate the
[01:04:03] Durga Toshniwal: these. And then, accordingly, when you will backpropagate, the backpropagation will primarily involve, setting the weights and the bias to a correct value.
[01:04:17] Durga Toshniwal: So, backdropagation is usually only to correct the
[01:04:21] Durga Toshniwal: errors so that they become optimally minimal. And usually, we, in the back propagation, we look at weights and biases. But in the forward pass, you usually look at bias weights, you look at, so in the forward direction, you take some weights and use them.
[01:04:39] Durga Toshniwal: Okay, you take some weights and you use the bias, to propagate
[01:04:45] Durga Toshniwal: the information further. What you iterate upon here is other hyperparameters, like number of layers, number of neurons per layer, number of epochs, number of iterations, optimizers, and so on and so forth.
[01:05:00] Durga Toshniwal: And in the backward pass, or the back propagation of errors, you look at weights and biases. But there are a lot of hyperparameters, and we need to actually,
[01:05:11] Durga Toshniwal: Use all of them in the optimally best fashion.
[01:05:15] Deepak Katara: One more thing, so in one of the lectures on transform learning, right, what we have understood, that not all the layers are exposed, there are multiple layers, you know, which are hidden, and we just fine-tune last couple of layers.
[01:05:33] Deepak Katara: So, if that's the case, right? So…
[01:05:37] Deepak Katara: there might be a possibility we end up, you know, getting wrong model, because we are not getting access to a lot of layers, it's just few layers which, are exposed. Like, I think we got an example of VCC, Inception, all these.
[01:05:55] Deepak Katara: Right. So in that scenario, like, how do we, sort of, get that right.
[01:06:02] Durga Toshniwal: So, actually, transfer learning is a completely different concept.
[01:06:07] Durga Toshniwal: So, in this, what happens is, as you might have studied, that you actually train a model.
[01:06:12] Durga Toshniwal: Based on some particular, data in some particular domain, okay?
[01:06:18] Durga Toshniwal: And, once the model is obtained.
[01:06:20] Durga Toshniwal: You will generalize it in such a fashion that it is also useful on some other data of some other domain.
[01:06:28] Durga Toshniwal: So, we build a model there.
[01:06:31] Durga Toshniwal: And then, after building the model, we actually generalize it. So usually, we don't generalize the model so much, but here we generalize it to the extent
[01:06:40] Durga Toshniwal: That, it is also helping in prediction of some other domain, related data.
[01:06:47] Durga Toshniwal: And when we are doing a generalization, then we actually don't change the entire model, we just change the last few layers.
[01:06:56] Durga Toshniwal: in such a way that it is able to do a prediction on that particular domain data that we want to do the transfer learning for. Say we are having, let's say, a model which is applicable on images that has been taken up from the
[01:07:15] Durga Toshniwal: which are the images of some buildings and cars, so images on the road and the surroundings. Let's say we want to do a prediction on that. The prediction could be prediction on the…
[01:07:27] Durga Toshniwal: objects that are there in the image. The objects could be car, it could be building, and so on and so forth.
[01:07:34] Durga Toshniwal: So, this is one particular domain. Now, once the model is built.
[01:07:39] Durga Toshniwal: What I want to do is that I want to do… use transfer learning in such a way that the model is also applicable on some other domain images. Say we are taking images not from the road, rather we are taking images, let's say, from
[01:07:55] Durga Toshniwal: some other area. Let's say we are having some, you know, we are having, instead of the road and the surrounding buildings, we are having a, let's say some industrial area.
[01:08:10] Durga Toshniwal: So, the kind of objects that you will find in the industrial area may be very different from the normal residential area. So, the learning that has been done
[01:08:19] Durga Toshniwal: should be useful on the other domain data also. So to… for it to be useful, what we try to do is we just try to manipulate the last few layers
[01:08:31] Durga Toshniwal: instead of redoing the entire thing, if we redo the entire model, then there's no transfer learning happening. So, we use the partial model that has been made, and on the rest of it, we try to, you know, work on the last few layers.
[01:08:49] Durga Toshniwal: So that the error minimization happens only there, and it can be used for transfer learning. However, a point to note is that when we use transfer learning, then, in that particular case.
[01:09:02] Durga Toshniwal: The accuracy and the other matrix may not be very high, because after all, what we want to use is the advantage of transfer learning, but we are not, you know, looking at precision recall accuracy per se only.
[01:09:23] Deepak Katara: Sure. Last one question, probably it can be answered later on, but I think one of a member in our team also asked the same question. So, how did we evolve from linear regression to nonlinear regression to deep learning? Because these have some common grounds in these three, so…
[01:09:43] Deepak Katara: If you could answer, like, what is a correlation? How do we sort of evolve, and in which problem we go with regression, and which problem we go with deep learnings?
[01:09:53] Durga Toshniwal: open.
[01:09:54] Durga Toshniwal: So, regression is a, you should say that regression is a very less sophisticated method.
[01:10:03] Durga Toshniwal: Again, regression could be linear or nonlinear, both. It could be univariate, multivariate also.
[01:10:10] Durga Toshniwal: So, regression is used in problems that may not be very complex. For example, if you are talking about some language models, or NLP kind of data, or we are talking about images and all that, in all such cases, we are
[01:10:29] Durga Toshniwal: In all such cases, the… the complexity in the data is very huge.
[01:10:37] Durga Toshniwal: Say, for example, if you talk about natural language, that means that the language spoken by humans. Obviously, humans have… are very intelligent, and what they perceive.
[01:10:49] Durga Toshniwal: that is very difficult to obtain using any simple model. So, if you're having relatively less complex data, say numeric data.
[01:10:59] Durga Toshniwal: Maybe less complex, then we can use the simpler models, like linear or nonlinear forms of regression.
[01:11:08] Durga Toshniwal: Or logistic regression for classification and all that. However, if we are generally looking at complex data, it is good to use deep learning models.
[01:11:18] Durga Toshniwal: So this is a very,
[01:11:20] Durga Toshniwal: It's a generic way of looking at it.
[01:11:24] Durga Toshniwal: Again, if you look at deep learning.
[01:11:28] Durga Toshniwal: So deep learning also involves some linear combinations. So in the linear model, whatever you… I mean, in the regression model, if it is linear, you're doing completely linear way. Non-linear, you are using some other things.
[01:11:43] Durga Toshniwal: So, in the, deep learning model also, if you look at, this part, the summation of…
[01:11:52] Durga Toshniwal: The, the summation of the products.
[01:11:56] Durga Toshniwal: Of the weights and the inputs.
[01:12:00] Durga Toshniwal: This part, this much.
[01:12:02] Durga Toshniwal: This is not… nothing but a linear combination of the weighted inputs. So, up to this point, it is only simple… kind of simple linear combination only.
[01:12:13] Durga Toshniwal: And the nonlinearity comes with this function f only.
[01:12:18] Durga Toshniwal: Which is not there in case of regression.
[01:12:22] Durga Toshniwal: So, this is how you use it. If you have complex data, try to use deep learning, it will give you better results. If your data is relatively simpler, use simpler models like regression.
[01:12:35] Durga Toshniwal: Okay.
[01:12:36] Deepak Katara: And can we call regression as a subset of deep learning, or separate things?
[01:12:43] Durga Toshniwal: Actually, regression, is primarily a statistical method. It's also used in machine learning, AI, and all.
[01:12:53] Durga Toshniwal: So, basically, we cannot call it as a subset of deep learning. Regression may not be called a subset of deep learning, but it is a part of ML, generally.
[01:13:05] Durga Toshniwal: AI and ML, it is a part of AI and ML.
[01:13:08] Durga Toshniwal: But not of deep learning, because in deep learning, we have nonlinearity getting incorporated in the model with the help of activation functions.
[01:13:19] Durga Toshniwal: So, the two are not directly comparable that way.
[01:13:24] Deepak Katara: Chill, I think…
[01:13:27] Durga Toshniwal: Any other questions, anyone, before we proceed?
[01:13:30] Jithu Tagore: Well, am I audible?
[01:13:33] Durga Toshniwal: Yes, yes, you are.
[01:13:34] Jithu Tagore: Ma'am, you have said two terms. One was transfer learning, and another one was fine-tuning. Like, what is the difference between both? Like, are all the weights, being updated during transfer learning, or
[01:13:49] Jithu Tagore: And fine-tuning, or in some layers of phase to…
[01:13:53] Jithu Tagore: For, and, some layers of a standard, after, balance layers are being updated, like, like, that, is that the difference? I don't know.
[01:14:06] Durga Toshniwal: Okay.
[01:14:07] Durga Toshniwal: So the thing is that if we, change, all the set of weights.
[01:14:13] Durga Toshniwal: then you can think of… then what are you doing? You're building a full new model, isn't it? If I change this… these set of weights, then I… I like W1, W2, W3, then I change W4, W5, W6. If I change all of these, then I'm doing the entire model back again.
[01:14:30] Durga Toshniwal: If I'm doing the entire model back again, then there's no transfer learning that is happening.
[01:14:35] Durga Toshniwal: So, in case of transfer learning, a partially, partially, built model is utilized.
[01:14:43] Durga Toshniwal: And a part of it is modified, which can involve changing of weights and biases in some layers.
[01:14:52] Durga Toshniwal: In such a way that the output, becomes closer to the expected output.
[01:14:59] Durga Toshniwal: Which may be from some other domains.
[01:15:02] Durga Toshniwal: So we don't do, rework on all the widths. We just work on some layers and their widths.
[01:15:10] Durga Toshniwal: Okay? So, if you are not, if you're not talking about transfer learning.
[01:15:15] Durga Toshniwal: Then we are talking about fine-tuning, that is, let's say we are having just, I'll go to this example.
[01:15:31] Durga Toshniwal: Just one second.
[01:15:36] Durga Toshniwal: Yep.
[01:15:38] Durga Toshniwal: So, what is happening here? There are multiple layers, right? Input layer, output layer, and maybe 3 hidden layers. This is hidden layer 1, this is hidden layer 2, and this is hidden layer 3.
[01:15:50] Durga Toshniwal: So, what is happening here? The initial, hidden layer is working on some edges and boundaries. The next one combines some of these to obtain some shapes. The third one does a further combination on these.
[01:16:04] Durga Toshniwal: So, this is kind of fine-tuning on what is coming from the previous hidden layer.
[01:16:10] Durga Toshniwal: Okay, so this is an example of fine-tuning. However, in transfer learning, Probably.
[01:16:18] Durga Toshniwal: The weights and biases and rest of the stuff.
[01:16:21] Durga Toshniwal: Of the last few layers have changed.
[01:16:24] Durga Toshniwal: With the objective of obtaining a good output, or a, you know, the output which is very similar to the
[01:16:32] Durga Toshniwal: Expected output.
[01:16:34] Durga Toshniwal: For the new domain, on which we are going to use this partially learned model. So, this is the difference.
[01:16:43] Durga Toshniwal: Stokely?
[01:16:46] Jithu Tagore: Ma'am, I have similar, like, I have done an application called… used a model called Yolo V5. So, 4, 5, 10 networks.
[01:16:56] Jithu Tagore: So, for improvement, I have phased 10 layers, 10 above, then continued the training for 5 more approx. Like, that, that second stage is called, transfer learning, right?
[01:17:11] Durga Toshniwal: I couldn't get you. You have used a 10-layer, in a nut.
[01:17:15] Jithu Tagore: Architecture. All of Wi-Fi model, yeah.
[01:17:18] Durga Toshniwal: Okay.
[01:17:19] Jithu Tagore: Like, for image training, it's an image classification model. I have trained for 10 epoch first.
[01:17:26] Jithu Tagore: Then, after getting the model, I will freeze the 10 epochs, then I will, train for, 5 more epochs. So, I'm, the second time I'm doing, like, transfer link, right?
[01:17:42] Jithu Tagore: Is that okay?
[01:17:43] Durga Toshniwal: So, in the second, in the second set, you are going to use some new data.
[01:17:50] Durga Toshniwal: Right? Some other domain data.
[01:17:53] Jithu Tagore: Is that correct?
[01:17:55] Jithu Tagore: Yes, ma'am.
[01:17:56] Durga Toshniwal: Yeah, so that's okay. That is a form of transfer learning.
[01:18:00] Durga Toshniwal: Because I haven't yet taught EPOC, so I have not mentioned it, but yes, you can do it like that. That you're partially… I mean, you have shown the initial data on it sometimes, and then after that, you show the other data.
[01:18:15] Durga Toshniwal: And then you train it for the other data, right?
[01:18:19] Jithu Tagore: So…
[01:18:20] Durga Toshniwal: That's fine. That's one way to explain transfer learning model.
[01:18:25] Durga Toshniwal: So, actually, there are many methods of doing transfer learning. One, you can do fine-tuning, which may be on some layers, or you can use pre-trained model, and then you can, you know, run it for some epochs on the initial data, and then run it for some other number of epochs on other data.
[01:18:45] Durga Toshniwal: Or you could actually, you know, you could have some feature extraction method, and then you use some other kind of data. So there are a lot of ways to achieve transfer learning.
[01:18:58] Durga Toshniwal: And this is one of them, whatever you said.
[01:19:02] Jithu Tagore: Okay, ma'am. Now it's clear.
[01:19:06] Durga Toshniwal: Any other questions, anyone?
[01:19:13] Durga Toshniwal: So these are very important concepts. I request all of you to pay attention, and if you have any doubts, please do ask, because the whole of the deep learning will be based on, you know, fully connected neural networks.
[01:19:27] Durga Toshniwal: And we are going to, in the latter part of it, we are not going to discuss how this works. We'll assume that this is working.
[01:19:36] Durga Toshniwal: Okay, so, now, what I'll do is that,
[01:19:41] Durga Toshniwal: I'll just do a stop share, and I'll share another set of slides. Allow me a minute.
[01:19:50] Durga Toshniwal: Just a second.
[01:20:02] Durga Toshniwal: Just a…
[01:20:38] Durga Toshniwal: Okay, so I hope… My screen is visible now.
[01:20:43] Durga Toshniwal: So, having talked about, the neural networks.
[01:20:49] Durga Toshniwal: And having talked about weight, spices, and activation functions, the first thing that you should know, is what is activation function. And I told you that
[01:21:00] Durga Toshniwal: If we just multiply the weights and the biases, what we are having is just a linear combination of the weighted inputs, and there's no non-linearity that is happening in it.
[01:21:10] Durga Toshniwal: And this non-linearity is actually brought about with the help of activation functions. So there are a variety of activation functions, and we can use one or more of them in our
[01:21:20] Durga Toshniwal: Neural network.
[01:21:22] Durga Toshniwal: So, the first one that we are going to discuss is, the stepwise function, okay?
[01:21:29] Durga Toshniwal: So, first, Discuss the step function.
[01:21:32] Durga Toshniwal: It's also called the threshold function or the step function. More generally, it's called the step function.
[01:21:37] Durga Toshniwal: So, why is it called the step function? I will just explain to you.
[01:21:42] Durga Toshniwal: It is called a step function because, see here, it takes the shape of a step.
[01:21:48] Durga Toshniwal: Just like this.
[01:21:50] Durga Toshniwal: So it's just looking like a step.
[01:21:52] Durga Toshniwal: And what does it mean?
[01:21:54] Durga Toshniwal: It means that, if FX is the output, Access the input.
[01:22:00] Durga Toshniwal: And FX is the output, so this whole output.
[01:22:05] Durga Toshniwal: dysfunction.
[01:22:06] Durga Toshniwal: It's a function, which may be a function.
[01:22:09] Durga Toshniwal: the input.
[01:22:11] Durga Toshniwal: So if we have a value, 0… When X is 0.
[01:22:17] Sacheen Adavinavar: Ma'am, your voice is better.
[01:22:21] Durga Toshniwal: Voice… thing you're saying?
[01:22:22] Sacheen Adavinavar: Like, like yesterday, yeah.
[01:22:25] Durga Toshniwal: Just, I don't…
[01:22:43] Durga Toshniwal: Okay, just let me know if it is breaking further.
[01:22:48] Sacheen Adavinavar: No, ma'am, that's fine.
[01:22:50] Durga Toshniwal: Okay, thanks.
[01:22:51] Durga Toshniwal: So, the value of the function f will be 0. If X is less than 0, and it will be 1,
[01:23:00] Durga Toshniwal: If X is taking a value greater than or equal to 0.
[01:23:04] Durga Toshniwal: So this is what is the step function. So, you can see here, this is X,
[01:23:09] Durga Toshniwal: And this is FX.
[01:23:12] Durga Toshniwal: This is 0, this is 1.
[01:23:15] Durga Toshniwal: If X is taking up a value which is less than 0, so these are the values which are less than 0, say this is minus 1, minus 2, minus 3, and so on and so forth, minus 9, like that.
[01:23:27] Durga Toshniwal: So, if X is having a value which is negative, or less than 0, then the step function will give an output of 0.
[01:23:36] Durga Toshniwal: And for any value which is equal to 0 or greater than 0, say 1, 2, 3, till 9, 10, like that, then the value of the step function is going to become 1.
[01:23:51] Durga Toshniwal: So, this is how this function is defined.
[01:23:55] Durga Toshniwal: And, the threshold in this case is 0, because this is the point where the step comes in.
[01:24:02] Durga Toshniwal: Right, so this function is also called the threshold activation function, or generally the step function, and it is quite a simple activation function.
[01:24:13] Durga Toshniwal: And the behavior is, taking in the values, X, as the input.
[01:24:21] Durga Toshniwal: Which is continuous valued input, and generate a binary out of it based on the threshold.
[01:24:27] Durga Toshniwal: In this case, our threshold is… this is the threshold, and this is equal to zero.
[01:24:33] Durga Toshniwal: If it was anything else, then the step would be at some other threshold. If the threshold was… if the step would be like this, say I'm calling it F prime X,
[01:24:46] Durga Toshniwal: 0 or 1. If X is, let's say.
[01:24:49] Durga Toshniwal: less than 2, or X is greater than or equal to 2.
[01:24:53] Durga Toshniwal: So then, the step in that case will be here.
[01:24:56] Durga Toshniwal: like this.
[01:24:58] Durga Toshniwal: And the rest of it will be…
[01:25:00] Durga Toshniwal: 0 only. So, we could have any threshold. The most common one is 0 only, as the threshold. And you can see that this function is nonlinear in nature, because what is it doing? Up to this point, it is transforming all the numbers, continuous valued numbers.
[01:25:18] Durga Toshniwal: To 0, and after that, to 1. So this is how this activation function is incorporating the linearity and transforming the given values
[01:25:27] Durga Toshniwal: From minus infinity to plus infinity to 1 or 0, depending on the threshold.
[01:25:34] Durga Toshniwal: So then,
[01:25:35] Durga Toshniwal: Generally speaking, if we talk about some threshold theta, then FX will be equal to 1 if x is greater than or equal to theta, and it will be 0 if x is less than theta.
[01:25:47] Durga Toshniwal: Right. So, and, X is usually the input to the activation function.
[01:25:56] Durga Toshniwal: And FX is the output.
[01:25:59] Durga Toshniwal: of the activation function, and theta here is the threshold that we are going to use. In the previous case, theta was equal to 0.
[01:26:10] Durga Toshniwal: So, the, usually the input to the activation function would be what? The input to the activation function would be summation of
[01:26:20] Durga Toshniwal: Xi?
[01:26:23] Durga Toshniwal: Producted with WY.
[01:26:26] Durga Toshniwal: Sorry, I equal to 1 2M.
[01:26:29] Durga Toshniwal: So this is the input, which is the weighted, sum of the inputs that is going to the activation function, okay?
[01:26:37] Durga Toshniwal: So…
[01:26:38] Durga Toshniwal: This is the input to the activation function, and the threshold could be 0 or anything else, and the step would work.
[01:26:45] Durga Toshniwal: Like this. And this is called a unit step because the value is 1.
[01:26:52] Durga Toshniwal: So now, so, what happens is that, in this particular case.
[01:27:01] Durga Toshniwal: Any, so any neuron here.
[01:27:05] Durga Toshniwal: In our neural network, that is actually getting a value 0 is set to be activated.
[01:27:11] Durga Toshniwal: Whereas those which are getting a, output, I mean, which is having a value which is less than zero, are said to be deactivated, right?
[01:27:21] Durga Toshniwal: Or in other words, the nodes or the neurons can get activated or deactivated based on this particular step function that is used, or the activation function that is used.
[01:27:33] Durga Toshniwal: You can think of it like this. You can just correlate it with the switch, an electrical switch. Let's say we are having an electrical switch.
[01:27:43] Durga Toshniwal: Like this.
[01:27:45] Durga Toshniwal: If the value of the current is…
[01:27:49] Durga Toshniwal: Greater than or equal to zero.
[01:27:52] Durga Toshniwal: Then the switch would actually transform in a…
[01:27:55] Durga Toshniwal: Closed circuit form. And the current would flow through it.
[01:27:59] Durga Toshniwal: if… The switch is receiving a value of the current.
[01:28:05] Durga Toshniwal: Which is less than zero.
[01:28:07] Durga Toshniwal: Then it will continue to remain in the open firm.
[01:28:13] Durga Toshniwal: This is what it means.
[01:28:15] Durga Toshniwal: Or, in other words, let's say…
[01:28:21] Durga Toshniwal: Let's say if we are having multiple, you know, notes like this.
[01:28:28] Durga Toshniwal: These are receiving some inputs.
[01:28:30] Durga Toshniwal: And, the value of the input, this is X1, this is, say, X2, this is X3, and say this value is minus 1, say this is 0 and this is 1, then what will happen? When this propagates further.
[01:28:47] Durga Toshniwal: because of the activation function here, being applied here, then, what would happen? Anything going out of here.
[01:28:56] Durga Toshniwal: with this, value of 1, this will be minus 1 times weight W1, like that. So this part, actually, this value, because of the activation function, will be rendered to 0.
[01:29:08] Durga Toshniwal: Or, in other words, we can think of it as a soft button, which is turned off.
[01:29:15] Durga Toshniwal: Right? Because of this input. So whenever there's a negative input, the output will be zero. Or in other words, the contribution of this particular input is not going to be there on any of the
[01:29:27] Durga Toshniwal: Notes which are further there.
[01:29:30] Durga Toshniwal: So we can think of it as, you know, deactivation of this node, or a soft removal of this node.
[01:29:37] Durga Toshniwal: And the other nodes that have a value of current greater than or equal to zero will act like a closed switch, and the current will flow. I'm just giving an analogy, and they will remain activated.
[01:29:51] Durga Toshniwal: In this particular, neural network.
[01:29:56] Durga Toshniwal: I could see some hands raised. Are there any doubts or any questions?
[01:30:01] Deepak Bobade: Yeah, yeah, ma'am, so in the context of, deactivated, right? So that you're saying minus 1 into W1, that value won't be there, right, when, it goes to the final output?
[01:30:14] Durga Toshniwal: Yeah.
[01:30:15] Deepak Bobade: So, why it's not, like, instead of, you know, minus 100W, and why it's not, like, 0.1?
[01:30:22] Deepak Bobade: Would that be a wrong representation?
[01:30:27] Durga Toshniwal: Which one? 0.w1.
[01:30:29] Deepak Bobade: Yeah, on the, layer, right? On the second layer.
[01:30:34] Durga Toshniwal: Okay, let me… let me see if I have space.
[01:30:38] Deepak Bobade: Just stop.
[01:30:40] Deepak Bobade: It's on the top right, minus 1.w.
[01:30:42] Durga Toshniwal: Yeah, yeah, I know that. I… I'm looking for some more space to…
[01:30:47] Durga Toshniwal: take this up fully. Anyhow, I'll do it here.
[01:30:50] Durga Toshniwal: So, let's say, you have this… F, working on… XRWI, right?
[01:30:57] Deepak Bobade: Yeah.
[01:30:58] Durga Toshniwal: So, at any particular node.
[01:31:02] Durga Toshniwal: If we are having,
[01:31:06] Durga Toshniwal: Let's say we are having this X as the input.
[01:31:09] Durga Toshniwal: And one is minus 1, one is, say, 2, one is 1, something, something. It could be 0 as well.
[01:31:16] Deepak Bobade: Okay. And then you are having some weight metrics.
[01:31:19] Deepak Bobade: C'est…
[01:31:21] Durga Toshniwal: as I took earlier, a similar example, if I take something something, Okay.
[01:31:27] Durga Toshniwal: Now, when we do this product.
[01:31:30] Durga Toshniwal: Let's say we talk about only one particular XIWI at only one particular node, so what it obtain will be minus one?
[01:31:40] Durga Toshniwal: dot 2 plus… 2.2… Plus 2.2, right?
[01:31:47] Durga Toshniwal: And on this summation, the activation.
[01:31:50] Deepak Bobade: 1.2, ma'am. Last one was 1.2.
[01:31:53] Durga Toshniwal: Sorry. Okay, that's fine.
[01:31:55] Deepak Bobade: Yeah.
[01:31:56] Durga Toshniwal: Okay?
[01:31:56] Deepak Bobade: Now the activation function is going to work on it.
[01:32:00] Durga Toshniwal: So when it works on it, it's going to actually work on minus 2 plus 4 plus 3.
[01:32:07] Deepak Bobade: Oh, okay, got it.
[01:32:09] Durga Toshniwal: And now this activation function is a step function, this becomes 0, this becomes 1, this becomes 1.
[01:32:16] Durga Toshniwal: Okay, and final output will be 1 only, because it's a step.
[01:32:20] Durga Toshniwal: So, whatever it's receiving from this particular node, which had a minus 1,
[01:32:27] Durga Toshniwal: is rendered useless. Or in other words, you can think of that node, if this was the connection, sorry, I have this here. You can think of this node as getting disconnected.
[01:32:37] Durga Toshniwal: From this one also, this one also, this one also, this one also, because it's… input is not getting… going, to go anywhere, it's not going to get contributed.
[01:32:47] Deepak Bobade: So you can think of it as an open switch.
[01:32:51] Durga Toshniwal: Right? And this value will not flow, across to the succeeding layer.
[01:32:59] Deepak Bobade: Okay, so overall thing is, the function is acting once we are getting that, multiplicated matrix, right? I thought, it's happening in between.
[01:33:09] Deepak Bobade: So once we get the, multiplicated metric, we could run over that F of function… F function, right? Depending on…
[01:33:18] Durga Toshniwal: Remember, remember, what we are talking about is an activation function. Where is activation function working? It is working
[01:33:27] Durga Toshniwal: At this point.
[01:33:30] Durga Toshniwal: Zinter.
[01:33:31] Deepak Bobade: Hmm. This is F.
[01:33:33] Durga Toshniwal: Okay, outside of some of these…
[01:33:35] Deepak Bobade: Huh?
[01:33:36] Durga Toshniwal: So, this F is this step function.
[01:33:40] Deepak Bobade: Hmm, hmm.
[01:33:41] Durga Toshniwal: Got it, got it, yeah.
[01:33:44] Deepak Bobade: Got it.
[01:33:45] Durga Toshniwal: Yeah. Any other questions, anyone?
[01:33:48] Durga Toshniwal: I think I could see one question on the chat which says that, how do we know which weight causes more loss?
[01:33:55] Durga Toshniwal: So, actually, I will, show you the math of back propagation.
[01:34:00] Durga Toshniwal: Which will take some time, maybe one session time, which I'm not shown right now, by the help of which you can decide which weight is causing more loss. As of now, it will be difficult for me to explain, so I have not yet covered it. I will cover.
[01:34:16] Durga Toshniwal: Okay, Lokesh.
[01:34:17] Lokesh R: Okay, ma'am, thank you.
[01:34:19] Durga Toshniwal: Any other questions on this step function?
[01:34:25] Durga Toshniwal: Okay, so this was the simplest activation function.
[01:34:29] Durga Toshniwal: No,
[01:34:32] Durga Toshniwal: What happens is that, in this particular case, one or more neurons may get activated. When we say activated, it means that they will output a value of 1, right?
[01:34:44] Durga Toshniwal: And, there is no such restriction that only one neuron can be active. Any number of neurons can be… can remain active if they trigger to a value of 1.
[01:34:55] Durga Toshniwal: And each neuron, the value is calculated independently.
[01:35:00] Durga Toshniwal: So,
[01:35:02] Durga Toshniwal: let's say if we are having 3 neurons, A, B, and C, and each of them have an input which is greater than theta, where theta is the threshold, then all of them will output 1.
[01:35:13] Durga Toshniwal: Okay.
[01:35:14] Durga Toshniwal: So,
[01:35:16] Durga Toshniwal: So this is how this is going to work, and these neurons will remain activated if they are outputting a 1, and they will be deactivated or considered disconnected if they are outputting a value 0.
[01:35:39] Durga Toshniwal: Sorry, I got muted.
[01:35:44] Durga Toshniwal: Okay.
[01:35:46] Durga Toshniwal: No.
[01:35:48] Durga Toshniwal: Let's say we talk about a multi-class problem.
[01:35:53] Durga Toshniwal: So, if we are talking about a multi-class problem, then we have multiple classes that need to be predicted. For example, we are talking about a handwritten image identif- and digit identification.
[01:36:06] Durga Toshniwal: When we are talking about digit identification, then we are talking about the digit, with the class label 0, then 1, 2, 3, 4, 5, 6, 7, 8, 9. So it's no longer binary in nature.
[01:36:21] Durga Toshniwal: And, what we are expecting are, the output of 10 different classes, right?
[01:36:29] Durga Toshniwal: Now, if you are using a step function, what would happen if the, if the input is 0?
[01:36:37] Durga Toshniwal: The output is going to be 1. If the input… so let me just draw that curve again.
[01:36:44] Durga Toshniwal: So this is our step function.
[01:36:47] Durga Toshniwal: This is 0, this is 1, 2, 3…
[01:36:56] Durga Toshniwal: So, no matter what, Comes out, what is given as input, the output will be 1 only.
[01:37:04] Durga Toshniwal: And if it is negative, only then it comes out to be 0. So for any input, if the digit is a 1, handwritten 1, output will be 1. If it is a 10, if it is a 9, sorry, if X is equal to 9, then also the output of 9 will be 1.
[01:37:22] Durga Toshniwal: Or in other words,
[01:37:24] Durga Toshniwal: So, as we know that using a step function, any number of neurons could be triggered or become activated.
[01:37:33] Durga Toshniwal: So that is fine, as long as it is okay with the problem that we are going to solve. But in some cases where we are talking about multi-class problem, we don't want all the neurons to be activated
[01:37:46] Durga Toshniwal: At the same time. We only want one output to be generated.
[01:37:53] Durga Toshniwal: In this case.
[01:37:54] Durga Toshniwal: all the outputs, if these are the 10 inputs going to the process, all of them would generate one only as the output. So, all are triggered, or all are activated, and therefore, we cannot identify the classes with respect to each other.
[01:38:09] Durga Toshniwal: Because all neurons are active. But we only want one neuron to be active at a time, so that in the multi-class problem, we can identify 1 from 2, or maybe 8 from 9, and so on and so forth.
[01:38:23] Durga Toshniwal: So then…
[01:38:25] Durga Toshniwal: So this means that, in many cases, we don't want the activation, to work on, on the concept of binarization. And we would want, the activation to be based on some other concept.
[01:38:42] Durga Toshniwal: Right? And, we just don't rely on, binary kind of, activation of, or triggering of the…
[01:38:52] Durga Toshniwal: activation function. So then, we have some other, activation functions.
[01:38:58] Durga Toshniwal: That use some other form of,
[01:39:02] Durga Toshniwal: Some other form of, you know, functioning.
[01:39:05] Durga Toshniwal: So… Let's talk about the next.
[01:39:08] Durga Toshniwal: function, which is a PLT, or piecewise linear threshold function. So, earlier was a step function, now we are having a PLT function, which is the piecewise linear threshold function. So, what is a piecewise linear threshold function?
[01:39:24] Durga Toshniwal: It takes the shape of what you are seeing here.
[01:39:27] Durga Toshniwal: This is the shape.
[01:39:29] Durga Toshniwal: Okay, this is X.
[01:39:31] Durga Toshniwal: This is FX.
[01:39:35] Durga Toshniwal: And… Okay, so now, in this particular case.
[01:39:41] Durga Toshniwal: The value of X, FX will be 0 if X is less than or equal to alpha. So this is alpha.
[01:39:51] Durga Toshniwal: And the value… if the value of X is less than alpha, this
[01:39:57] Durga Toshniwal: Then the output is going to be 0.
[01:40:00] Durga Toshniwal: So, FX is 0.
[01:40:02] Durga Toshniwal: VIN?
[01:40:03] Durga Toshniwal: X's… Less than or equal to alpha.
[01:40:07] Durga Toshniwal: Now, the next condition is FX is equal to 1, this is 1.
[01:40:13] Durga Toshniwal: If X is greater than or equal to beta.
[01:40:17] Durga Toshniwal: Or B, whatever you want to call it. Let's call it B.
[01:40:20] Durga Toshniwal: So, FX is going to be 1.
[01:40:23] Durga Toshniwal: if X is greater than or equal to B.
[01:40:28] Durga Toshniwal: And in between these two, FX is equal to… MX plus B.
[01:40:37] Durga Toshniwal: Okay?
[01:40:38] Durga Toshniwal: So, where it could be C or B, whatever it is. So, it is, in this case, MX plus B.
[01:40:46] Durga Toshniwal: And in this case, X will be, less than B, and it will be greater than alpha.
[01:40:53] Durga Toshniwal: So here, X is greater than alpha, and it is less than B.
[01:40:58] Durga Toshniwal: So, this is the case.
[01:41:00] Durga Toshniwal: These are the values of X, and for that value of X, FX will be MX plus P, this equation.
[01:41:07] Durga Toshniwal: So, what you are seeing here… It, these three pieces…
[01:41:12] Durga Toshniwal: FX is equal to 0 when x is greater than, less than or equal to alpha.
[01:41:17] Durga Toshniwal: FX is equal to 1 when X is greater than or equal to B.
[01:41:22] Durga Toshniwal: And FX is equal to MX plus B when x is greater than alpha and less than B.
[01:41:29] Durga Toshniwal: So, these are linear in nature. This much is linear, this much is linear, this much is lean.
[01:41:35] Durga Toshniwal: But… If we put them all together, the whole becomes a nonlinear function.
[01:41:42] Durga Toshniwal: Again, and brings about the non-linearity in our model. Okay. So, it… so this model is called piecewise linear, because there are three pieces, 1, 2, and 3.
[01:41:54] Durga Toshniwal: And in themselves, they are linear. This much is linear, this much is linear, this much is linear. But when they are combined together, then they become a combined nonlinear function. So within that interval, it is linear. Otherwise, in totality, it is nonlinear. This is called a PLT, or the piecewise linear function.
[01:42:13] Durga Toshniwal: Okay.
[01:42:17] Durga Toshniwal: So, now, actually, if we had the step function, it would be like this.
[01:42:24] Durga Toshniwal: Here, if you are having threshold 0, and this is a value 1, this is X,
[01:42:30] Durga Toshniwal: And this is FX.
[01:42:32] Durga Toshniwal: So, in a step function, there would be a sudden jump on the threshold to a value 1.
[01:42:39] Durga Toshniwal: Whereas the… the PLT is kind of a soft version of this, so instead of having an abrupt step.
[01:42:47] Durga Toshniwal: Suddenly coming in, it's having a shape like this.
[01:42:51] Durga Toshniwal: So, the increase in the value from 0 to 1
[01:42:56] Durga Toshniwal: actually is not having an abrupt… there's no abrupt change. It is happening between alpha and B values, so it is a gradual change.
[01:43:06] Durga Toshniwal: Right, so it is gradually changing from 0 to 1, rather than a sudden change from 0 to 1.
[01:43:13] Durga Toshniwal: So, now the question is that how does a sudden change hamper the processing, whereas how would piecewise linear threshold be better? It will be better, because if we try to differentiate here, the slope will come out to… it is not differentiable.
[01:43:30] Durga Toshniwal: Alright, but if we look at this, then it is differentiable, because there's a finite slope out here.
[01:43:37] Durga Toshniwal: Right, so, so therefore, and we always want differentiable functions are always preferred.
[01:43:45] Durga Toshniwal: Because when we try to use gradient descent.
[01:43:48] Durga Toshniwal: On the functions, then we need to have them differentiable.
[01:43:54] Durga Toshniwal: And, the stepwise function is not differentiable.
[01:43:58] Durga Toshniwal: Right? It also allows the neurons to, respond in a gradual fashion, rather than in a trigger-non-trigger mode.
[01:44:08] Durga Toshniwal: And these are more interpretable also.
[01:44:10] Durga Toshniwal: Therefore, PLT, which is a, kind of smoother or a soft form of the initial one.
[01:44:18] Durga Toshniwal: That is the step function.
[01:44:20] Durga Toshniwal: Is better than the step function itself.
[01:44:23] Durga Toshniwal: And it is also non-linear.
[01:44:30] Durga Toshniwal: Now, the piecewise linear threshold, actually, as I already mentioned, it looks something like this.
[01:44:37] Durga Toshniwal: This is zero.
[01:44:39] Durga Toshniwal: And then, this is alpha.
[01:44:41] Durga Toshniwal: So, alpha, and this is B, value B. So, it goes… Oh.
[01:44:49] Durga Toshniwal: like this.
[01:44:52] Durga Toshniwal: And becomes like this, right?
[01:44:55] Durga Toshniwal: So, again, of course, it is a software version of this function, which is our step function.
[01:45:01] Durga Toshniwal: This is our step function.
[01:45:03] Durga Toshniwal: This is 0, this is 1.
[01:45:05] Durga Toshniwal: And it takes up… the values…
[01:45:10] Durga Toshniwal: of 0 for X less than 0 equal to 0, or 1?
[01:45:14] Durga Toshniwal: So, this is definitely a software version, that is true.
[01:45:19] Durga Toshniwal: But… For values which are less than alpha, it still gives zero result.
[01:45:24] Durga Toshniwal: So, all the neurons that would be activated by an input which is negative will actually be dropped out, or they would be deactivated, and their input will not have any weightage in the further calculation, because anything less than alpha
[01:45:42] Durga Toshniwal: Is going to give an output of 0.
[01:45:45] Durga Toshniwal: So, although it does away with the problem of this gradual shift from 0 to 1 by replacing it with this kind of a gradual slope, which is, like, alpha X plus P,
[01:45:58] Durga Toshniwal: Oh… But the problem that it still faces is this part.
[01:46:03] Durga Toshniwal: Where?
[01:46:04] Durga Toshniwal: the neurons that receive this kind of input are deactivated from further processing. So.
[01:46:12] Durga Toshniwal: We have variants of PLT, like ReLU, that will be used.
[01:46:17] Durga Toshniwal: And it's a special case of PLT.
[01:46:20] Durga Toshniwal: Then we also have leaky relu, hard sigmoid, hard tanage, and so on and so forth, which are also other, variants of
[01:46:29] Durga Toshniwal: ReLU, of, PLT. So, we are going to look at relu now.
[01:46:35] Durga Toshniwal: So, what is value?
[01:46:38] Durga Toshniwal: So, in ReLU, what we have is, ReLU is, nothing but rectified linear activation.
[01:46:48] Durga Toshniwal: So, this is the R.
[01:46:50] Durga Toshniwal: And this is the E, and this is L, and this is unit. So, it is rectified linear activation function of a unit.
[01:47:00] Durga Toshniwal: Okay, as it is called, value.
[01:47:02] Durga Toshniwal: And what it is going to do.
[01:47:05] Durga Toshniwal: It's also a form of piecewise linear function only, but it will input, it will take an input, and it will output a value which is actually going to be equal to itself.
[01:47:19] Durga Toshniwal: So the function of a value is fx is equal to max of 0 or X.
[01:47:28] Durga Toshniwal: What does it mean?
[01:47:29] Durga Toshniwal: It means that if X is equal to minus 1,
[01:47:33] Durga Toshniwal: then FX will… of minus 1, Will be a max of…
[01:47:39] Durga Toshniwal: 0 or minus 1, which will be 0.
[01:47:42] Durga Toshniwal: If FX is equal to minus 2,
[01:47:45] Durga Toshniwal: F of minus 2 will be max of…
[01:47:48] Durga Toshniwal: Minus 2 and 0, it will be 0.
[01:47:51] Durga Toshniwal: Like that.
[01:47:53] Durga Toshniwal: So, X will keep on outputting 0 as long as… sorry, FX will keep on outputting 0 as long as X is negative. Now, let's look at, 0.
[01:48:05] Durga Toshniwal: Then, F of 0 will be max of 0, which is 0.
[01:48:12] Durga Toshniwal: And then if X is equal to plus 1, then F of plus 1.
[01:48:16] Durga Toshniwal: Will be equal to max of…
[01:48:19] Durga Toshniwal: 0 and 1, which will be 1.
[01:48:21] Durga Toshniwal: If X is equal to 10, then F of 10 will be equal to…
[01:48:26] Durga Toshniwal: Max of 0 and 10, which is 10.
[01:48:29] Durga Toshniwal: So here, the output is equal to input provided FX is greater than or equal to
[01:48:35] Durga Toshniwal: It is greater than zero.
[01:48:38] Durga Toshniwal: So, so the graph will be like this.
[01:48:41] Durga Toshniwal: Here we have 0, here we have minus 1, minus 2, like that. Here we have plus 1, plus 2, plus 3, like that. So, as long as values are negative, what we have will be 0 here.
[01:48:54] Durga Toshniwal: At 0 also, it will be 0, but for positive values, the output will be equal to the input. So, for a value of 1, this will be 1. For a value of 2, it will be 2. For a value of 3, it will be 3. For a value of 10, it will be 10, like that. So, this is the function.
[01:49:12] Durga Toshniwal: Okay.
[01:49:14] Durga Toshniwal: So now, this is also a modification of stepwise and then PLT only. In PLT, we had an additional thing, which was like this.
[01:49:22] Durga Toshniwal: So we just removed this.
[01:49:24] Durga Toshniwal: And, anything any positive output is equal to itself.
[01:49:31] Durga Toshniwal: This is what is relevant.
[01:49:33] Durga Toshniwal: So, in this case, it will output a zero only when there's a negative input. Other than that, for all inputs which are greater than zero, the same value as the input will be outputted.
[01:49:46] Durga Toshniwal: And this is also, again, piecewise linear. This is linear, this is linear. But together, this whole thing is non-linear.
[01:49:55] Durga Toshniwal: Okay. Are there any questions so far?
[01:49:59] Durga Toshniwal: Anyone, any questions?
[01:50:10] Durga Toshniwal: Okay.
[01:50:13] Durga Toshniwal: So now, why would a ReLU be more useful? This is what the ReLU looks like.
[01:50:18] Durga Toshniwal: Here it is X, and here it is FX.
[01:50:23] Durga Toshniwal: So, relu, is nonlinear in nature, definitely.
[01:50:27] Durga Toshniwal: Then,
[01:50:29] Durga Toshniwal: So, what happens is, in this particular case, see, the problem of, large activations, which was there because of positive inputs in case of step function, as well as partially in case of PLT function, would be solved by relieve.
[01:50:45] Durga Toshniwal: Because ReLU gives our output, you know, of 0 when its value are either negative or zero.
[01:50:53] Durga Toshniwal: So this means that if 50% of the inputs are zero or negative.
[01:50:58] Durga Toshniwal: Then, around 50% of the nodes will get deactivated.
[01:51:03] Durga Toshniwal: And therefore, the activation is sparser than what it would be in case of the stepwise function, because in that case, a large number of activated nodes would be there, and then it will be difficult to differentiate the different classes in a multi-class problem.
[01:51:23] Durga Toshniwal: And then it is comparatively less computationally intensive.
[01:51:29] Durga Toshniwal: Now, these were the advantages. What can be the disadvantages with ReLU? One is the dying ReLU problem. What is a dying ReLU problem?
[01:51:39] Durga Toshniwal: If there are many neurons that are receiving negative inputs that are lying… that is lying here.
[01:51:45] Durga Toshniwal: Then what would happen is all these neurons will actually be kind of deactivated or switched off.
[01:51:52] Durga Toshniwal: Or you can think of it, like this, that if, you know, if… let me draw it here.
[01:52:00] Durga Toshniwal: So, let's say… These are inputs coming to… These four neurons, X1, X2,
[01:52:07] Durga Toshniwal: X3 and X4. This is plus 1, this is, let's say, minus 1, this is minus 2, and this is minus 3.
[01:52:15] Durga Toshniwal: These are the inputs, and then we talk about the next layer.
[01:52:20] Durga Toshniwal: Here, then each of these inputs are going here, like this, right?
[01:52:26] Durga Toshniwal: So then,
[01:52:28] Durga Toshniwal: And then each of these inputs are going here, like this, these are going here, then these are going here. I'm not drawing the entire thing. So, anything that is, going from… from this node.
[01:52:41] Durga Toshniwal: Because the input to this is minus 1, so it will be minus 1 times whatever the weight is, right? So therefore, the quantity will come out to be negative only. This negative quantity, wherever it goes.
[01:52:55] Durga Toshniwal: So, what will happen? Anything contributed by this node, by this node and this node. Practically.
[01:53:01] Durga Toshniwal: all these nodes are rendered useless, so there can be too many neurons, that are kind of switched off in a soft fashion because of the output being zero.
[01:53:13] Durga Toshniwal: When, the input is, less than zero, right? So here, what could happen? This could stop completely.
[01:53:23] Durga Toshniwal: And therefore, it could, create a problem, which is called the dying value problem, okay?
[01:53:29] Durga Toshniwal: So…
[01:53:31] Durga Toshniwal: It can also happen when the weights get updated, then the neurons might get negative input, and they get, kind of switched off.
[01:53:43] Durga Toshniwal: Also, another problem with this is that it is not zero-centered.
[01:53:47] Durga Toshniwal: this… This particular part is not zero sensor center, rather it is eccentric with the zero, right?
[01:53:55] Durga Toshniwal: So, this also can, affect the optimization process.
[01:54:02] Durga Toshniwal: So, any questions so far, anyone?
[01:54:11] Durga Toshniwal: So I'm going to do a stop share.
[01:54:14] Durga Toshniwal: And, I'd like to know,
[01:54:18] Durga Toshniwal: How many of you have started forming the groups?
[01:54:22] Durga Toshniwal: Are there many of you who have started forming the groups? If you have any questions on whatever I taught you today, you can ask, then we'll talk about the groups.
[01:54:37] Durga Toshniwal: Anyone who has formed the groups.
[01:54:42] Durga Toshniwal: So, if it is becoming difficult to form the groups, then I will request Simran, I will tell you the typical size to cut the list into,
[01:54:52] Durga Toshniwal: bins of equal size, that will be… the size will be the number of
[01:54:56] Durga Toshniwal: Members per team, it should be.
[01:55:01] Durga Toshniwal: So that I am going to do. If any of you have formed the groups, you can intimate right now on the chat.
[01:55:08] abhilash daniel: Okay, this is Abhidash.
[01:55:11] Durga Toshniwal: I have made a group of 3 folks.
[01:55:14] Durga Toshniwal: Okay, that's good that you made a group of 3 folks. All, if the size finally comes out to be 5, then you'll have to accommodate 2 more.
[01:55:22] abhilash daniel: Yeah, so that should be good.
[01:55:23] Durga Toshniwal: That's it, yeah.
[01:55:25] Durga Toshniwal: So, what I'll, do is I'll request Simran to float a Google Sheet.
[01:55:30] Durga Toshniwal: In which you can mention the names of the people who have got together to form a group. It could be whatever the number is.
[01:55:38] Durga Toshniwal: say, 5 or less than 5. I'll tell the exact number, how many will be there in a team, and then accordingly, you can either add more members or reduce the members, whatever.
[01:55:50] Durga Toshniwal: And default is that if we are not able to find a team or make a team, then we'll just cut out the list into equal number of members per team, and we'll divide it into teams like that.
[01:56:04] Durga Toshniwal: Okay.
[01:56:05] Durga Toshniwal: So this is what we are going to do.
[01:56:09] Durga Toshniwal: So, Simran, please request, the Google Sheet with the, with the input for the
[01:56:17] Durga Toshniwal: There's the team member names and all that.
[01:56:21] Durga Toshniwal: And then we'll see how many teams have got formed. And accordingly, in the next turn, we can finalize the teams, okay?
[01:56:29] GenAI Batch-2 Manager: I'll email it to them by today, and put it on the LMS as well.
[01:56:32] Durga Toshniwal: Yeah, yeah, thanks for that.
[01:56:34] Durga Toshniwal: Okay, so I think we can finish forming the groups now. Any questions, anyone?
[01:56:41] Durga Toshniwal: Any doubts, any queries on whatever we discussed or otherwise?
[01:56:51] Abhishek Acharya: No, ma'am, audio.
[01:56:53] Chandrasekhar Sahu: One question, like,
[01:56:56] Chandrasekhar Sahu: So, we are, we are having, like, a lot of activation functions. So, how this theorem will going to know, like, that which activation functions
[01:57:05] Chandrasekhar Sahu: Will be good for…
[01:57:08] Chandrasekhar Sahu: getting the right output. Like, we are having random input, and also random, activistance layer, so it will be a lot of,
[01:57:15] Chandrasekhar Sahu: Calculations it will do for each and every layer, sir.
[01:57:19] Durga Toshniwal: What you're saying is exactly correct.
[01:57:22] Durga Toshniwal: Because we actually don't know what will be the inputs, what are the activation functions, what are the weights, and everything, and everything needs to be iterated and adjusted.
[01:57:33] Durga Toshniwal: And so, also, activation function. So, usually, we need to iterate on the different activation functions until and unless we know the specific use case. For example, we know that we are having a multi-class problem, then we should not use the step function.
[01:57:49] Durga Toshniwal: or the PLT function, then we can use the other functions, like that you can do. But otherwise, there's no standard way to decide which function to use. You have to use it and see how the performance is coming. If it is coming good, that's fine. If it is not coming, reject it and use something else.
[01:58:06] Durga Toshniwal: But that's the one leaving, yeah.
[01:58:10] Chandrasekhar Sahu: Okay, ma'am. Thank you.
[01:58:11] Durga Toshniwal: Yeah, and that's why deep learning is very compute-intensive.
[01:58:15] Durga Toshniwal: Very, very confusing tense.
[01:58:19] Durga Toshniwal: Any other questions? Anyone?
[01:58:27] Durga Toshniwal: Okay, if there are no further questions, then we can break for today.
[01:58:31] Durga Toshniwal: And we'll continue in the next turn.
[01:58:35] Durga Toshniwal: Okay, we'll continue with what we are doing. Okay, thank you all. Have a great day, thank you, bye-bye.
[01:58:42] Abhishek Acharya: Thank you, have a nice day, bye.
[01:58:44] Durga Toshniwal: Thank you.
[01:58:44] Akshat Pandya: Thank you.
[01:58:46] Neeraj Kumar: Yeah, my bad.
[01:58:47] Chandrasekhar Sahu: Bye.