# 03 2026-01-31 Deep Learning Concepts

course: Module 3 — Deep Learning & NLP
module: Module-3-Deep-Learning-NLP
date: 2026-01-31
type: transcript
video_url: https://personal-learn.armco.dev/files/_Recordings/Module-3-Deep-Learning-NLP/03_2026-01-31_Deep_Learning_Concepts.mp4

---
[00:13:03] Durga Toshniwal: A very good evening to all of you, and welcome to today's session.
[00:13:08] Durga Toshniwal: So today, we'll be discussing, deep learning.
[00:13:13] Durga Toshniwal: Which we had started in the last time.
[00:13:16] Durga Toshniwal: We'll be containing with the concepts, so…
[00:13:19] Durga Toshniwal: Let me just share my slides.
[00:13:22] Durga Toshniwal: And we'll, start…
[00:13:25] Durga Toshniwal: From where we had left. Also, the lab materials, the link for the lab materials, shows the materials that will be used tomorrow.
[00:13:34] Durga Toshniwal: So tomorrow, we'll be having hands-on on the deep learning that we'll be discussing today.
[00:13:40] Durga Toshniwal: Okay, so let me just share my slides, and we'll go on from there.
[00:14:04] Durga Toshniwal: So, if you all recollect, in the last turn, we had been discussing the various concepts, and we had stopped at the point where we discussed what is back propagation, and I told you that whenever we try to calculate the output.
[00:14:21] Durga Toshniwal: So, the output is, can be represented as Y hat, as you are seeing here. And then, the comparison between the actual given output versus the, predicted output is done.
[00:14:34] Durga Toshniwal: And, the difference is said to be error.
[00:14:37] Durga Toshniwal: When the error is totaled over all the samples in the training data, it is called loss.
[00:14:43] Durga Toshniwal: So, this error is actually back-propagated.
[00:14:47] Durga Toshniwal: To whatever is nearest to the output. So, to the output, the first hidden layer is the nearest, so it is back-propagated to the first hidden layer. The first hidden layer, in turn, draws upon its input from the previous hidden layer. So, the
[00:15:03] Durga Toshniwal: So the, errors are then back-propagated to the next hidden layer in the, backward direction.
[00:15:12] Durga Toshniwal: And then this will again back-replicate it to the, up to the point it reaches the input.
[00:15:19] Durga Toshniwal: So this is how the error will be back propagated, and the purpose of backpropagation is to find or update the weights in such a way
[00:15:29] Durga Toshniwal: That, when we, again, have another forward pass, then the expected, value and the predicted value match, in a better way. And the accuracy and the performance of the classification or prediction, improves.
[00:15:48] Durga Toshniwal: So, this was the whole idea of backpropagation.
[00:15:52] Durga Toshniwal: So, we'll be now studying further things. Just a second.
[00:16:02] Durga Toshniwal: Okay.
[00:16:03] Durga Toshniwal: So now, when we talk about back propagation, as I told you, that, in backpropagation, we are going to
[00:16:10] Durga Toshniwal: Back-propagate the errors, and this will involve finding out the derivative.
[00:16:16] Durga Toshniwal: of the cost function or the loss. So, J, the cost function here is represented as J, it is also called loss.
[00:16:24] Durga Toshniwal: It's the same thing.
[00:16:26] Durga Toshniwal: And, the derivative of the loss with respect to… or the rate of change of loss with respect to the weight and the biases is calculated when we backpropagate the loss, or the error. So, here, what we are doing is that we… actually, when we back-propagate, see, what is happening is that
[00:16:45] Durga Toshniwal: As we discussed at length in the last turn, that the inputs, are, multiplied by weights, and the weighted,
[00:16:55] Durga Toshniwal: some of the different fractions of the input will go to the different neurons in the hidden layer. Then again, they will be multiplied, those what is received at a neuron. Again, it will be weighted
[00:17:09] Durga Toshniwal: And like that, it goes on and on. So, finally, the output that we obtain here
[00:17:16] Durga Toshniwal: Actually, depends on, on these weights, whatever these weight mattresses are.
[00:17:22] Durga Toshniwal: It depends on these weight mattresses, and also on the bias.
[00:17:27] Durga Toshniwal: So, so, this means that if we want to improve this process, we'll have to actually iterate and find out the best possible weights so that
[00:17:37] Durga Toshniwal: the predicted output and the expected output match, and the loss is minimized. So, for that purpose, we find out the rate of change of loss with respect to weight, and with respect to bias.
[00:17:49] Durga Toshniwal: So, actually, what we do is that we actually find out the loss, or the cost derivative with respect to the weight matrix. Similarly, this with respect to the bias matrix. However, when we do that, actually, the derivative is done element twice. As a result, what we have here is
[00:18:08] Durga Toshniwal: Derivative of J with respect to derivative of W, JKKL. So, what is WJKL? It means that we are considering the, weight at the Kth neuron. So, K here, if we talk about WJKL,
[00:18:26] Durga Toshniwal: K refers to the KF neuron.
[00:18:30] Durga Toshniwal: This is the weight from the KF neuron.
[00:18:33] Durga Toshniwal: And this is there in L-1th layer.
[00:18:38] Durga Toshniwal: And the weight on… at the Kth neuron in the 1 minus… L-1th layer is actually going to…
[00:18:47] Durga Toshniwal: the JF neuron.
[00:18:49] Durga Toshniwal: Okay, so it's going to the JF neuron.
[00:18:53] Durga Toshniwal: In Ella Tlare.
[00:18:57] Durga Toshniwal: So, there are two hidden layers. This is layer, let's say, this is one particular hidden layer like this, another hidden layer like this. This is L-1th layer, this is leth layer.
[00:19:09] Durga Toshniwal: And, in this, we have, let's say this is the Kth neuron, and here, this is the, jth neuron.
[00:19:18] Durga Toshniwal: So, what it means is that the weight that is going… that is in between these
[00:19:27] Durga Toshniwal: From the Kth neuron up to the jth neuron.
[00:19:30] Durga Toshniwal: From L-1th layer to the leth layer, this is what is represented by WJKL. Actually, it will be a full matrix, but it is shown element-wise here.
[00:19:40] Durga Toshniwal: Similarly, BJL is the bias at the jth neuron in the 11th layer.
[00:19:46] Durga Toshniwal: this layer.
[00:19:47] Durga Toshniwal: And J, of course, is the cost function, as we already discussed.
[00:19:52] Durga Toshniwal: And, as I already mentioned, we are interested in finding out
[00:19:56] Durga Toshniwal: The derivative, which is done at the matrix level. However, when we do that, it is actually calculated element-wise.
[00:20:06] Durga Toshniwal: So then, back propagation will allow us to find out
[00:20:11] Durga Toshniwal: How, the… how each weight in the neural network contributes to the loss or to the overall error, because we are back propagating the error and finding out the rate of change of error with respect to the weights.
[00:20:27] Durga Toshniwal: So, therefore, batch propagation will help us to assess this. So, this means that we will be able to find out
[00:20:35] Durga Toshniwal: That… oh.
[00:20:40] Durga Toshniwal: So, we'll be able to find out the,
[00:20:44] Durga Toshniwal: how much each and every weight is contributing to the overall error. So, we will know which weight is contributing, how much, and all that, but
[00:20:54] Durga Toshniwal: Back propagation will help us to estimate this… this thing, that the weight is going to contribute to the overall error, how much it will. But it doesn't show any path to optimize the weight.
[00:21:08] Durga Toshniwal: So, the weight optimization cannot be done by backpropagation. Backpropagation is only a way to find out or to assess the contribution of each and every weight towards the
[00:21:21] Durga Toshniwal: overall loss or error or cost, they are, synonymous. But, because the error is because of the
[00:21:31] Durga Toshniwal: widths, which are not the optimal widths, but that optimization is not provided by backpropagation, and we have to do something extra.
[00:21:40] Durga Toshniwal: So, what is that extra?
[00:21:43] Durga Toshniwal: So, this means that we will need some extra or some other algorithm that will help us to optimize the weights
[00:21:53] Durga Toshniwal: In such a way that the predicted output and the given output match to the best possible fashion. And there are a large number of optimization methods.
[00:22:04] Durga Toshniwal: That will be used along with backpropagation, so that when the backpropagation
[00:22:10] Durga Toshniwal: finds out the rate of change of error with respect to each and every weight. The optimization method is going to optimize each and every weight so that this error
[00:22:23] Durga Toshniwal: gets reduced.
[00:22:25] Durga Toshniwal: Okay, I could see a hand raised. Who had a question?
[00:22:28] Aditya Banda: Yes, ma'am. So, when you back-propagate, do you predict at every layer, and then calculate the loss or cost at every layer?
[00:22:37] Durga Toshniwal: Yes.
[00:22:37] Durga Toshniwal: So I showed you, actually, here.
[00:22:41] Durga Toshniwal: So the back propagation is done layer to layer, like this.
[00:22:46] Durga Toshniwal: So, the output is actually first backpropagated to whatever, it's nearest to. So, nearest… if this is hidden layer, this is hidden layer 1, this is 2, and this is the input layer, and this is the output layer.
[00:23:01] Durga Toshniwal: So, the output layer is actually nearest to hidden layer 2, so it will backpropagate the error to hidden layer 2.
[00:23:07] Durga Toshniwal: When it… yeah, and then… but Hidden Layer 2 actually knows that it derives its input from Hidden Layer 1.
[00:23:15] Durga Toshniwal: Because it's not directly getting any input. It is only feeding upon what it is getting from Eden Layer 1. So, it again, in turn, back-propagates this
[00:23:26] Durga Toshniwal: in the backward direction to hidden layer 1. And this goes on and on till the input layer is reached, because finally, it is the weighted inputs that go to the hidden layer 1.
[00:23:37] Durga Toshniwal: So this is how it is done.
[00:23:42] Durga Toshniwal: Any other questions, anyone?
[00:23:48] Durga Toshniwal: Awesome.
[00:23:49] Durga Toshniwal: So then, as I already mentioned, that back propagation identifies
[00:23:55] Durga Toshniwal: the contribution of the weights towards the loss, overall error, or the cost. However, backpropagation doesn't help us to optimize the weights, and we'll need to use some other algorithms.
[00:24:09] Durga Toshniwal: And, for, and, so there are functions, that help us to… optimization functions that help to evaluate, what the set of weights, how they are going to impact
[00:24:26] Durga Toshniwal: you know, the loss function. The loss function is also called the objective function. The objective function generally means any function that we want to maximize or minimize.
[00:24:37] Durga Toshniwal: In our case, we want… the objective function is usually the loss or the cost function.
[00:24:43] Durga Toshniwal: And because it is the loss, we will like to minimize it.
[00:24:46] Durga Toshniwal: So, this particular function will actually, look at or try to assess any candidate solution and give a score, how useful it is, right? So the loss function will always give loss in form of a single number, which will be a scalar value.
[00:25:06] Durga Toshniwal: And, this value can be used to assess how good or bad a candidate weight matrix is. So, the possible weights could be ranked, and then compared.
[00:25:19] Durga Toshniwal: So, here we have the loss.
[00:25:21] Durga Toshniwal: J, with, which is a function of weights and bias.
[00:25:25] Durga Toshniwal: And, what is this? This is the lamb… this is lambda or loss.
[00:25:30] Durga Toshniwal: That is the difference between the predicted value, this is the predicted value.
[00:25:36] Durga Toshniwal: of Y?
[00:25:39] Durga Toshniwal: And this is the given value, or the actual value.
[00:25:45] Durga Toshniwal: of Y, and the difference between these two is the loss, and this difference is found out for all M samples. These are the number of samples in the training data.
[00:25:56] Durga Toshniwal: And then the average loss is found out by doing the summation and divide… dividing it by M.
[00:26:04] Durga Toshniwal: So now, we look at the different optimization techniques that are available and commonly used. There are a large number of algorithms that are possible, however, I'm going to talk about some of the important ones here. So the first one, is random search.
[00:26:23] Durga Toshniwal: So, random search, is a very simple, kind of, solution.
[00:26:31] Durga Toshniwal: And what it does is that, it just randomly looks at, some possible weights.
[00:26:40] Durga Toshniwal: And then, what it, does is that it tries to optimize those, and then.
[00:26:57] Sacheen Adavinavar: You got muted, ma'am.
[00:27:07] Durga Toshniwal: Oh, sorry for that.
[00:27:12] Durga Toshniwal: So, in random search, the whole idea is that
[00:27:16] Durga Toshniwal: What the optimization algorithm will do, it will just try out some random set of bits.
[00:27:24] Durga Toshniwal: And what it will do is that out of these random set of widths, it is going to keep a track
[00:27:30] Durga Toshniwal: Of which set of weights is working best, and it's going to choose that.
[00:27:35] Durga Toshniwal: And therefore, What it will do is that, once it chooses some particular set.
[00:27:43] Durga Toshniwal: It will go on defining that, and it will assume that this is the best solution. So, in random search, a random set of weights are used, they are just iterated upon.
[00:27:53] Durga Toshniwal: And the solution is taken as the final solution, without any consideration of any other kind of
[00:28:01] Durga Toshniwal: Betterment or anything.
[00:28:04] Durga Toshniwal: So, if we try to draw an analogy, we can think, let's say, if we think about a person who's actually walking on a hilly terrain with a blindfold on, and this person is supposed to climb down the hilly terrain.
[00:28:20] Durga Toshniwal: Then, the person in… if we… if he or she follows random search, then what the person will do? He… he will just try to randomly…
[00:28:34] Durga Toshniwal: Walk in one particular direction with the hope to reach the bottom.
[00:28:39] Durga Toshniwal: And the random direction that is chosen is just followed as the best possible solution. So this strategy definitely is not a very good strategy, because the solution that it leads to may not be the best solution at all.
[00:28:55] Durga Toshniwal: And so… We'll look at other strategies.
[00:29:05] Durga Toshniwal: Then the next strategy is a random local search. So what does a random local search, do?
[00:29:11] Durga Toshniwal: The random local search just looks at
[00:29:15] Durga Toshniwal: Instead of making a purely random choice, it is just going to look at a local, random local solution.
[00:29:22] Durga Toshniwal: So, the earlier… this one, the difference between random search and random local searches, that in random search, it looks at a random, solution on a global, scenario.
[00:29:37] Durga Toshniwal: Here, what it will do, it will select a random, it will do a random search locally, and then just think that this is the best possible solution and choose it.
[00:29:49] Durga Toshniwal: So, it's going to be computationally intensive, because it's going to search in the local neighborhood.
[00:29:56] Durga Toshniwal: And, and going to choose the solution as the best one, and still, this could never be the best one.
[00:30:03] Durga Toshniwal: For example.
[00:30:06] Durga Toshniwal: If we again take the same analogy, that we have a blindfolded person who's trying to do, climb down a hill.
[00:30:14] Durga Toshniwal: And the person just extends, so what is a random local search? So, let's say the person is standing here.
[00:30:23] Durga Toshniwal: So, the person extend, extends his foot randomly in a particular direction.
[00:30:30] Durga Toshniwal: locally. So, this is his vicinity. He just locally extends his foot.
[00:30:35] Durga Toshniwal: If he feels that the foot is going in the downward direction, he proceeds to… he proceeds in that direction. So, actually, the direction should have been this, but locally, the person saw that this is also going downward, so he chose to go here.
[00:30:50] Durga Toshniwal: And what will he reach? He'll reach a local minimum, but not the global minimum.
[00:30:56] Durga Toshniwal: The global minimum would be this.
[00:30:59] Durga Toshniwal: which he's supposed to do, that is climb down the hill. So, for some time, he will go in the downward direction, and then will actually go upward rather than going downward. So, random local search is also not a very good
[00:31:12] Durga Toshniwal: Solution, because it's just randomly looking at the vicinity and selecting the, solution, some particular solution, in the hope that it is a optimally best solution.
[00:31:26] Durga Toshniwal: Then, the next strategy is following the gradient.
[00:31:30] Durga Toshniwal: So, following the gradient means that…
[00:31:34] Durga Toshniwal: So the… in this particular case, the best direction in which the algorithm will proceed will be the one in which the rate of change of the weight vector will be maximum, or the steepest.
[00:31:49] Durga Toshniwal: So, the gradient, or the rate of change of the loss will be the highest, or in other words, this is going to be the steepest descent, and this will be chosen as the direction to be followed.
[00:32:05] Durga Toshniwal: So, if we return back to this solution, back again… Then…
[00:32:11] Durga Toshniwal: Let's say, again, the person is doing a hill climbing.
[00:32:15] Durga Toshniwal: Sorry, is, climbing down the hill, and the person is standing here.
[00:32:23] Durga Toshniwal: And then, or let's assume, yeah, so that the person is standing here, And, so…
[00:32:33] Durga Toshniwal: What, the blindfolded person will try to identify which particular… he tries to extend his foot and tries to feel
[00:32:43] Durga Toshniwal: That the gradient in which direction is the steepest.
[00:32:48] Durga Toshniwal: And then he will choose that particular direction to go ahead, with the steepest descent.
[00:32:56] Durga Toshniwal: So, in this particular case, when the foot is extended here, this is the… this is a very steep gradient as compared to this one. So, the person will start walking in this direction, and follow the steepest descent, and will reach very quickly to the, optimal solution.
[00:33:16] Durga Toshniwal: So, this is following the gradient.
[00:33:20] Durga Toshniwal: So, this was about some optimization techniques. There are names of these algorithms, which I'll go… which I'll be discussing shortly. These are the strategies that are followed very commonly.
[00:33:34] Durga Toshniwal: Soono?
[00:33:36] Durga Toshniwal: The gradient, actually.
[00:33:38] Durga Toshniwal: is nothing but the slope, and it tells us the direction in which our function is going to show the steepest rate of change, right? But…
[00:33:51] Durga Toshniwal: So, there are a few things that we discussed. The first thing was batch propagation. Batch propagation helps us to identify whether the weights that are being used are
[00:34:03] Durga Toshniwal: Are contributing… how much they are contributing to the error or the loss.
[00:34:08] Durga Toshniwal: That is number one. But they provide no…
[00:34:12] Durga Toshniwal: mechanism to improve or optimize those widths. So that is done with the help of the optimization functions.
[00:34:19] Durga Toshniwal: The optimization functions help us to choose
[00:34:23] Durga Toshniwal: The, direction to proceed where the rate of change of the slope or the descent is the steepest.
[00:34:33] Durga Toshniwal: Right, so this is what the gradient descent is going to… algorithm is going to help us do. So, it will help us to proceed in that direction. That will take us fastest.
[00:34:44] Durga Toshniwal: To our optimal solution.
[00:34:48] Durga Toshniwal: However.
[00:34:49] Durga Toshniwal: How much to go at a time, or what is the step size that must be chosen? That it does not help us with.
[00:34:57] Durga Toshniwal: So, so, this means that if we are, if somebody is walking.
[00:35:04] Durga Toshniwal: And that person is actually blindfolded
[00:35:07] Durga Toshniwal: then how much, in what direction he or she has to proceed that the optimization algorithm will tell, with the help of… and the gradient descent will show. But what is the…
[00:35:20] Durga Toshniwal: how big a step that person should take at a time, that will be decided by something called the learning rate. So, the learning rate helps us
[00:35:32] Durga Toshniwal: to, figure out how slow or fast we will move towards the direction of the optimal weights. So, this is the learning rate which is used in conjunction with gradient descent.
[00:35:47] Durga Toshniwal: To, to see how fast we can obtain our best possible solution.
[00:35:53] Durga Toshniwal: So this means that, let's say that if we are actually talking about
[00:35:59] Durga Toshniwal: the analogy that we had. Now, in this case, let's assume that we have got a…
[00:36:08] Durga Toshniwal: Just a second. So we have got a…
[00:36:11] Durga Toshniwal: Hill that has been actually organized into some steps.
[00:36:16] Durga Toshniwal: And let's assume that somebody has to climb up the hill.
[00:36:20] Durga Toshniwal: So, the person who has to climb up the hill.
[00:36:23] Durga Toshniwal: Has to decide how much she has to climb at a time.
[00:36:28] Durga Toshniwal: So, definitely, the steeper the hill is going to be.
[00:36:33] Durga Toshniwal: The more the gradient is going to be.
[00:36:36] Durga Toshniwal: And the more the gradient is going to be, the higher the step size will be. So, if the person decides to proceed like this, then the step size will be this much. He has to come up to this point to climb up this step.
[00:36:50] Durga Toshniwal: Then the gradient has reduced slightly, and as the gradient has reduced slightly, the step size has also reduced. So, as the gradient goes on decreasing, the step size also goes on decreasing.
[00:37:03] Durga Toshniwal: You can think of it like this.
[00:37:06] Durga Toshniwal: Let's say that we are having an object that is moving with an acceleration.
[00:37:14] Durga Toshniwal: off… Let's say 100 meters per second.
[00:37:20] Durga Toshniwal: And, the wheels are rotating. Let's say the object has… it's a car, and it's having a rotation.
[00:37:29] Durga Toshniwal: Which is happening, one rotation per second.
[00:37:34] Durga Toshniwal: Right? So, one rotation is happening in one second.
[00:37:38] Durga Toshniwal: So, the time to rotate.
[00:37:41] Durga Toshniwal: Here's one second.
[00:37:43] Durga Toshniwal: So this means that the acceleration is 100 meters per second, and the time to make one rotation is 1 second, let's say.
[00:37:53] Durga Toshniwal: Then the distance that will be covered by the wheel.
[00:37:58] Durga Toshniwal: Will be equal to 100 into 1.
[00:38:01] Durga Toshniwal: That will be 100 meters in one rotation.
[00:38:06] Durga Toshniwal: So, this will be the distance, or this will be the minimum distance that will be covered by one rotation of the wheel. Similarly, here in gradient descent, when the person is doing a hill climbing.
[00:38:19] Durga Toshniwal: Then?
[00:38:20] Durga Toshniwal: Based, given a gradient, there will be some minimum step size that has to be chosen, and that will be decided by the learning rate.
[00:38:29] Durga Toshniwal: So let's now, to make things a little more interesting, let's look at these troughs, or these parabolic
[00:38:38] Durga Toshniwal: containers.
[00:38:40] Durga Toshniwal: And there are… there is a ball inside. The final objective is that the ball should actually move
[00:38:47] Durga Toshniwal: To the lowest possible, point in the trough, or the container, which is this, in each of these.
[00:38:56] Durga Toshniwal: And the ball, the ball could actually move towards the left or the right.
[00:39:02] Durga Toshniwal: Such that, the objective is satisfied, that is, it reaches the lowest spot.
[00:39:07] Durga Toshniwal: So then, how would gradient descent, help the ball to go or reach the lowest step?
[00:39:15] Durga Toshniwal: First of all, the slope is calculated at any particular position of the ball. Let's say this is the current position, so what will be the gradient? That is… the gradient is nothing but the tangent, so this is the gradient.
[00:39:30] Durga Toshniwal: of the ball at the current position. Let's say number one position is the current position.
[00:39:36] Durga Toshniwal: And after the gradient is calculated, if the gradient is negative, then the ball will move towards the right direction. However, if the gradient is positive, it will move towards the left position. So in this case, the gradient you can see is a negative gradient.
[00:39:53] Durga Toshniwal: Because, It's showing a line moving downward.
[00:39:58] Durga Toshniwal: So this is a negative gradient. Therefore, the wall will be actually moving towards the, it will be moved towards the right direction. So now the ball will move towards the right direction, and it will reach here, such that this is the second step.
[00:40:15] Durga Toshniwal: Now, again, the gradient will be calculated here.
[00:40:18] Durga Toshniwal: When the gradient is calculated here, it comes out to be positive. So, whenever the gradient is positive, then the wall is made… is thrown in the left direction. So, this is thrown in the left direction, so that it reaches here, and once again, the gradient is calculated, it comes out to be negative.
[00:40:36] Durga Toshniwal: And then it is thrown.
[00:40:38] Durga Toshniwal: Towards the right side, in such a way that finally it comes to the lowest position. So, this is how it will proceed.
[00:40:46] Durga Toshniwal: So, first, what it will do is, it will come towards the… it will be thrown… because the gradient is negative, it will be thrown to the right side, or I can show it here. Initially, the… let's say the ball was here, the gradient was calculated, it came out to be negative.
[00:41:01] Durga Toshniwal: like this.
[00:41:03] Durga Toshniwal: So it was thrown to the… Right, direction?
[00:41:07] Durga Toshniwal: At the right direction, we found out the gradient, it is positive, so again, it was shown to the left side. The left side, the gradient is negative, so it was shown to the… thrown through the right side.
[00:41:18] Durga Toshniwal: And this continued till it finally reached the…
[00:41:24] Durga Toshniwal: lowest position. So this is how the gradient, descent, helps.
[00:41:29] Durga Toshniwal: To, you know, help us to reduce the weights, or optimize the weights.
[00:41:36] Durga Toshniwal: So that the loss reduces. In this case, we are just, you know, to gamify a bit, we are talking about a ball which is moving in a container in the downward direction.
[00:41:50] Durga Toshniwal: Right. So then, the question is that, of course, the ball, as we said, as per the algorithm, it is going on, moving from left to right, right to left, like that, till it reaches the lowest position of the
[00:42:04] Durga Toshniwal: Of this container. But how much will the mall move at a time? How will that be decided?
[00:42:11] Durga Toshniwal: So, to decide that, what we have is called… something called the learning rate. So, learning rate is a very important parameter.
[00:42:19] Durga Toshniwal: So, if the learning rate is very, less.
[00:42:22] Durga Toshniwal: then what will happen is, so this learning rate is something that I told you earlier.
[00:42:28] Durga Toshniwal: That is the amount, or the step size. So, in this case, given a gradient.
[00:42:36] Durga Toshniwal: The learning rate will be, deciding how much the step size will be. Because, given a gradient, the gradient into the learning rate will actually give us the step size.
[00:42:51] Durga Toshniwal: So, if the gradient is given already, then the learning rate is the one that will decide how much is the step size. So, let's assume that the gradient of this container is already given, the ball is starting from this position.
[00:43:05] Durga Toshniwal: And if the learning rate is very less, then what will happen? It will proceed with a very small step size, and
[00:43:13] Durga Toshniwal: That step size is so small that the movement of the ball is going to be parallel to the
[00:43:19] Durga Toshniwal: trough for the container itself. So, let's see if the ball started here, then it will move a very small step, go like this, move another very small step like this, and it will go on moving parallelly to the… in the direction of the
[00:43:35] Durga Toshniwal: In the… parallel to the, to the wall of the container.
[00:43:41] Durga Toshniwal: And this kind of a descent is going to be very, very slow and time-consuming.
[00:43:46] Durga Toshniwal: And, it will take a huge amount of time for it to reach the lowest position. So this was the case when learning rate was very, very less.
[00:43:57] Durga Toshniwal: Accordingly, step size was very less. The other, solution, to this problem is that
[00:44:05] Durga Toshniwal: Let's say the learning rate was very less, resulting into a very small step size, so let's now choose a learning rate which is quite high.
[00:44:14] Durga Toshniwal: Accordingly, the step size will also be high, because given a gradient, then we multiply it with the learning rate to get the step size. Now, assume we are having a very high learning rate, then what happens is, the step size becomes so large that it cannot be accommodated this side. So, the ball will start moving.
[00:44:33] Durga Toshniwal: in the opposite direction. Now, in this opposite direction, what will happen, the gradient will be calculated, and if it comes out to be positive, then it will be moved towards the
[00:44:45] Durga Toshniwal: opposite direction, so it will be moved like this. From here, we'll once again find out the gradient, it comes out to be negative, it will be pushed to the opposite direction, like that. So, what will happen? The wall started out here. Because of the large learning rate leading to a very large step size, it proceeded
[00:45:04] Durga Toshniwal: In such a way that it started moving away from the lowest position. The objective was to reach this
[00:45:10] Durga Toshniwal: Then it moved here, then it moved here, then it moved like this, and it will finally go out.
[00:45:17] Durga Toshniwal: So, our learning rate should not be very high.
[00:45:20] Durga Toshniwal: Because it will lead to… The ball moving in the opposite direction.
[00:45:26] Durga Toshniwal: And the steps, becoming too large.
[00:45:30] Durga Toshniwal: If the learning rate is too slow, then the gradient descent is going to be very slow, and it will be very time-consuming to reach this lowest position.
[00:45:40] Durga Toshniwal: So, we want the learning rate to be optimally high, not very high, but not very low. So, it should be low, but not very low. It should be high, but not very high.
[00:45:50] Durga Toshniwal: So, this is what the learning rate should be.
[00:45:53] Durga Toshniwal: And, because the learning rate is very, very critical, because it decides the step size. And, this step size is going to decide how much the… at a time, the gradient descent is going to, proceed.
[00:46:08] Durga Toshniwal: So, how much of the weights will be adjusted, or how much updating will happen? So that is what is decided by this, so it's very critical.
[00:46:19] Durga Toshniwal: Okay, Satin, I can see you have a question.
[00:46:23] Sacheen Adavinavar: Yeah, ma'am. So, for example, if the input is a CAT image, and then.
[00:46:30] Sacheen Adavinavar: it as a dog. So then, like, while the BAP propagation, right? So, to find…
[00:46:36] Sacheen Adavinavar: So will it look, for the… in the layers, like, suppose it has created a multiple layer, right?
[00:46:44] Durga Toshniwal: for example.
[00:46:45] Sacheen Adavinavar: base part layer, or, shapes of the, the email, I mean, the animal, or…
[00:46:53] Sacheen Adavinavar: So, those parts, it will look for the, the error, and then, makes the changes there, I mean, corrects it, or how is it going to happen?
[00:47:03] Durga Toshniwal: Yeah, so let's say if it is an image, and it's an image of a face or an animal, which is to be recognized, a dog or a cat, and there are some hidden layers. So what will happen? Let's say the image is that of a dog.
[00:47:17] Durga Toshniwal: And after the neural network proceeds in the forward direction, the prediction is not that of a dog, it is of some other animal. This means that the predicted value and the
[00:47:31] Durga Toshniwal: actual given value do not match. So, like this, all the training samples are, you know, used in the forward pass, and the sum total of the error is back propagated.
[00:47:43] Durga Toshniwal: When it is back propagated, it will be back propagated layer by layer until the input is reached. So, in this process, you remember what was happening in the forward pass?
[00:47:54] Durga Toshniwal: from the input, there was a weighted, you know, each of the neurons was receiving a weighted sum of all the inputs, right? And then, like that, it was happening for all the neurons, so…
[00:48:07] Durga Toshniwal: If it was neuron 1, 2, 3, and 1 hidden layer, then one would be receiving the weighted input of all the… I mean, weighted sum of all the inputs, like that.
[00:48:18] Durga Toshniwal: Now, because these weights are the ones which are deciding how much of input 1 is going to neuron 1, how much of input 2 is going to neuron 2, how much of input 3 is going to the respective neuron, and so on and so forth.
[00:48:33] Durga Toshniwal: And then, again, whatever is received is neuron 1 in hidden layer, 1 will be propagated a fraction of it by the weights in the next layer to the next neuron in the next hidden layer.
[00:48:46] Durga Toshniwal: So, these weights are the ones that are finally making up what we are getting at the output.
[00:48:52] Durga Toshniwal: So, when the error is back-propagated, these weights are adjusted.
[00:48:58] Durga Toshniwal: Right? And this is what we are discussing, that how the optimal weights are decided with the help of the optimization algorithm. And of course, in this process, when the weights get adjusted, and once, in one iteration, the weights get adjusted and the forward pass happens, then what will happen?
[00:49:19] Durga Toshniwal: Based on the new widths, suppose there are multiple layers, the first layer is picking up boundaries of the face. The second layer is connecting the boundaries into face features, the third one is connecting these face features into the face itself.
[00:49:34] Durga Toshniwal: If that is happening in the different layers, then when the weights will change, this will also change, and supposedly, it is supposed to improve.
[00:49:45] Durga Toshniwal: Okay.
[00:49:46] Durga Toshniwal: And again, then, once the forward pass happens, then the predicted value and the actual given value is compared.
[00:49:53] Durga Toshniwal: And if it doesn't match, and there's still a lot of error, then again it will be back propagated, weights will be changed, then again, forward pass will be happening, and this whole process of identifying boundaries, then connecting them to get features, then connecting features to get the face will be repeated.
[00:50:11] Durga Toshniwal: Like that. It keeps on happening, till…
[00:50:14] Durga Toshniwal: Our neural network is able to capture
[00:50:17] Durga Toshniwal: the correct organism, right? So this is how it proceeds.
[00:50:26] Durga Toshniwal: Any other questions, anyone?
[00:50:31] Durga Toshniwal: So, I can see some comments, so…
[00:50:36] Durga Toshniwal: I think Aditya has a question that how is the difference between the loss function of mean square error, RMSC, and what we see formula we see here? I could not understand your question, Aditya.
[00:50:51] Durga Toshniwal: What?
[00:50:52] Aditya Banda: There was a slide where you showed the formula for loss function.
[00:50:57] Aditya Banda: And generally, in classification models, we say RMSC or Mean Square Error as loss function, right? I'm wondering if both are same.
[00:51:07] Durga Toshniwal: Yeah, they are the same. Let me take you back to that slide.
[00:51:13] Aditya Banda: I think it's the second or third slide. Yeah, this one.
[00:51:16] Durga Toshniwal: This one, yeah.
[00:51:19] Durga Toshniwal: Yeah, so this is actually just a generic formula, which says the difference between the… this is Y hat is the.
[00:51:27] Aditya Banda: Predicted value, this is the given value.
[00:51:29] Durga Toshniwal: So, the difference between these two, and this difference summed for all samples, and the average taken. That's it.
[00:51:36] Durga Toshniwal: So, if you look at, let's say, sum of squared errors, what do you do? You calculate, let's say you are having Y hat as a predicted value, Y is the given value. Take the difference, do a square, then add it for all the attributes, say.
[00:51:53] Durga Toshniwal: k equal to 1 to, let's say, A are the total number of attributes, then do a square root of it.
[00:52:00] Aditya Banda: That's what you do.
[00:52:01] Durga Toshniwal: So, it's the same thing written here, but this is a generic formula.
[00:52:05] Durga Toshniwal: And the specific case could be SSE, MAE, whatever you want to use.
[00:52:12] Durga Toshniwal: Okay? And, this is written here. The loss or the error is written as a function of weight and bias.
[00:52:21] Durga Toshniwal: Why? Because… These, the… the…
[00:52:25] Durga Toshniwal: The output actually depends on weight and bias.
[00:52:30] Durga Toshniwal: So here, directly, I've written the output. I haven't written output as a function of weight and bias. However, the output is always a function of weight and bias only, isn't it? Because
[00:52:40] Durga Toshniwal: The inputs are getting multiplied by the weights, and getting added with the bias, and this gets propagated.
[00:52:48] Durga Toshniwal: In the forward pass to give the output.
[00:52:51] Aditya Banda: bet.
[00:52:52] Durga Toshniwal: I will show the derivation also, if it is of interest to all of you, how Y hat is actually a function of W and B, but this is a generic expression, but it is the same thing that you have studied, okay?
[00:53:09] Durga Toshniwal: Then I think there's one more question.
[00:53:15] Durga Toshniwal: Oh, okay, it's a comment or a question. I think, Sachin, you want to say that if input… I think we already discussed that, so I'm not going to go over it every day.
[00:53:26] Durga Toshniwal: So, let's now proceed.
[00:53:33] Durga Toshniwal: Oh.
[00:53:34] Durga Toshniwal: Now we'll talk about the different, variants of radiant descent that we just discussed.
[00:53:41] Durga Toshniwal: So, first is the mini-batch gradient descent, which is a very popular one, which is very commonly used.
[00:53:47] Durga Toshniwal: So, mini-batch gradient descent is, used. So, there can be cases where the training data is containing a large number of samples, say, millions of them.
[00:53:58] Durga Toshniwal: So, in all such cases,
[00:54:04] Durga Toshniwal: It may actually not be very useful to compute the loss over the entire set of the training data.
[00:54:12] Durga Toshniwal: in one go, as a single parameter update. So what we try to do is that we divide the given training data into batches.
[00:54:21] Durga Toshniwal: And then, we try to calculate the gradient descent per batch.
[00:54:27] Durga Toshniwal: And then take, you know, use that, instead of using it over one full entire training set.
[00:54:34] Durga Toshniwal: So, this is called mini batch gradient descent. So, we proceed by using batches
[00:54:41] Durga Toshniwal: on the training data. That's why it is called mini-batch.
[00:54:45] Durga Toshniwal: The other is the SGD, or the stochastic Radiant Descent.
[00:54:49] Durga Toshniwal: So, it's actually a special case of minibatch.
[00:54:54] Durga Toshniwal: If the mini batch has only one single sample in it.
[00:54:58] Durga Toshniwal: then what we have is the stochastic gradient descent. So I told you that the training data is divided into batches, and when the batch size is not very small, I mean, it is some finite, but it is not equal to 1, then what we have is a mini batch.
[00:55:15] Durga Toshniwal: But, if the size of the…
[00:55:19] Durga Toshniwal: If the size of the batch is equal to one sample, then this form of gradient descent is called stochastic gradient descent. It is very compute-intensive, but sometimes what happens is that the entire training data may not be available with us
[00:55:35] Durga Toshniwal: At any point in time.
[00:55:38] Durga Toshniwal: Now, the question is, how can it be that retaining data may not be available with us? Can there be such a situation? Yes, there can be. For example, if the data is getting generated continuously, let's say there's a sensor which is continuously measuring something.
[00:55:53] Durga Toshniwal: So, if a sensor is continuously measuring some value, then the training data will not
[00:55:59] Durga Toshniwal: you know, cannot get, cannot be available at a time. It is getting… it's just coming in continuously.
[00:56:08] Durga Toshniwal: So, in the… in such cases.
[00:56:10] Durga Toshniwal: We cannot use the simple gradient descent, neither can we use batch gradient descent. It would be more applicable to use the stochastic gradient descent, because
[00:56:20] Durga Toshniwal: The stochastician descent considers one training sample at a time. So, as the data keeps coming in.
[00:56:27] Durga Toshniwal: The samples are used one… one by one, like that.
[00:56:31] Durga Toshniwal: So…
[00:56:33] Durga Toshniwal: Definitely, because the samples are used one by one, stochastic gradient descent, or SGD, is going to be quite…
[00:56:41] Durga Toshniwal: compute intensive. Sometimes it's also called the online gradient descent.
[00:56:45] Durga Toshniwal: But, it may not be very useful always, but if it is essential, we do use it.
[00:56:54] Durga Toshniwal: Okay, so that was about the gradient descent variants, like minibatch, stochastic gradient descent, and all. So I told you the optimization approaches, now I'm just showing you the names of the optimization algorithms.
[00:57:08] Durga Toshniwal: So, we have common ones like Adam Optimizer, very popular. Then we have Ada Grad, then we have others like Adamax, RMS Prop, and so on and so forth. I will tell you
[00:57:21] Durga Toshniwal: about… there are so many more. So, I will just cover one or two here, which are very common. So, first is the ADAM optimizer, which is a very, very popular optimizer.
[00:57:31] Durga Toshniwal: So, Adam Optimizer is an optimization algorithm.
[00:57:35] Durga Toshniwal: And it is used in place of the stochastic gradient descent, which is a classical way to keep updating the widths, iteratively by taking one sample at a time.
[00:57:47] Durga Toshniwal: So, what it does, it considers some… two different extensions of the stochastic gradient descent. So, what it does, one, it uses the concept of adaptive gradient descent, or ADAGRAT. ADA grad is also one gradient descent algorithm. So, what is Adaptive Gradient Descent algorithm?
[00:58:08] Durga Toshniwal: What it does is that, the ADA grad is going to maintain a, you know, learning rate per parameter basis. It's not going to have a generic learning rate.
[00:58:21] Durga Toshniwal: So, it will adaptively look for a learning rate.
[00:58:25] Durga Toshniwal: optimal learning rate per parameter, so that the performance of the algorithm improves.
[00:58:34] Durga Toshniwal: Specifically, we are talking about some complex problem, like language… natural language prob… NLP problems, or computer vision problems on such cases.
[00:58:46] Durga Toshniwal: Because the neural networks may be very deep, and the gradients might become very less or sparse. In such cases, instead of generally taking a learning rate, per parameter, learning rate is taken, and this algorithm is called the ADA grid.
[00:59:03] Durga Toshniwal: Or the adaptive gradient algorithm.
[00:59:06] Durga Toshniwal: The other algorithm is the root-mean-square propagation, or the RMS prop. The RMS prop, again, also maintains learning rate on a per-parameter basis. However, these get adapted
[00:59:20] Durga Toshniwal: on the basis of an average value. And what is this average value? This average value is the average of the
[00:59:30] Durga Toshniwal: Magnitude of the recent gradients.
[00:59:35] Durga Toshniwal: So, as the weights are changing, so once the weights change, the gradients, you know, change. So, whenever gradients are changing, some recent, gradient magnitudes are kept, and the average of this is used.
[00:59:53] Durga Toshniwal: To calculate the learning rate. So…
[00:59:57] Durga Toshniwal: So since these… because the average value keeps on changing, so, RMS prop would adaptively keep on changing the learning rate, for each and every parameter. So, Adam Optimizer actually combines these, two.
[01:00:14] Durga Toshniwal: Ada Grad and RMS prop, that it does maintain a per-parameter learning rate.
[01:00:20] Durga Toshniwal: And, it also, uses the adaptiveness in the RMS prop.
[01:00:28] Durga Toshniwal: And this is how it combines these two.
[01:00:32] Durga Toshniwal: I think we already discussed about the activation functions, important ones, like the step function.
[01:00:39] Durga Toshniwal: Which is also called the threshold function. Then we talked about piecewise threshold function.
[01:00:46] Durga Toshniwal: Or the PLT function.
[01:00:49] Durga Toshniwal: I think we did not discuss the sigvoid function, so I will discuss sigvoid. It's a very, very important,
[01:00:56] Durga Toshniwal: Activation function.
[01:01:01] Durga Toshniwal: So, earlier we talked about, functions that had, very, you know…
[01:01:12] Durga Toshniwal: A very steep change, or an abrupt change.
[01:01:17] Durga Toshniwal: in them, like the piecewise, like the step function. The piecewise, step function was a little gradual.
[01:01:25] Durga Toshniwal: However, they were not smooth at all. They had sharp, pointy, edges.
[01:01:32] Durga Toshniwal: Where, the steps were changing.
[01:01:35] Durga Toshniwal: So, the sigmoid function actually aims to have a smooth transition.
[01:01:41] Durga Toshniwal: wherever the steps are changing. So, I'll just show you what it means. Before that, I can see
[01:01:47] Durga Toshniwal: There's a comment or a question by Shikhar that should we be decreasing the step size every time we move from the center?
[01:01:55] Durga Toshniwal: Or past the zero scope, so that we are not in the loop of coming back to where we start.
[01:02:02] Durga Toshniwal: Okay. So, the question is, I'll,
[01:02:08] Durga Toshniwal: So, we keep on decreasing the step size, actually, I think the question is… Closer to this… 1.
[01:02:24] Durga Toshniwal: Just a second.
[01:02:29] Durga Toshniwal: So, here you can see in the figure that what is happening is that
[01:02:34] Durga Toshniwal: When the ball is moving from the upper direction to the lower one, Then… Like, here to here.
[01:02:45] Durga Toshniwal: So, the step is something. This is the size of the step is S1, this is S2, and so on and so forth. It's automatically decreasing. Actually.
[01:02:55] Durga Toshniwal: This step size is decreasing because what is step size? Step size is equal to gradient times learning rate.
[01:03:04] Durga Toshniwal: If gradient is fixed, then the learning rate changes. Let's assume the learning rate is same, okay? Once it is decided, it's going to be something. But you can see here that the gradient is changing, and it is actually decreasing.
[01:03:19] Durga Toshniwal: At this point, the gradient is really very high, you can see it is very steep. So, as the ball is descending, the gradient is reducing, as you can see here.
[01:03:29] Durga Toshniwal: So, the learning rate is fixed, the gradient is reducing, so the step size is going to be smaller and smaller, so that the ball comes to the zero position.
[01:03:40] Durga Toshniwal: I hope it's okay, it's clear.
[01:03:48] Durga Toshniwal: Okay, so I… would not… get any response? Anyhow, I have actually addressed the question.
[01:03:55] Durga Toshniwal: So, let's talk about the sigmoid function. So, the sigmoid function is actually having a smooth step.
[01:04:01] Durga Toshniwal: It's not having an abrupt step like the PLT or the step function, which have a sudden, change.
[01:04:10] Durga Toshniwal: Okay, now, this sigmoid function is nonlinear, and therefore, when, combinations of this are made, they are going to be nonlinear also. And sigmoid function is, very good for classification purpose.
[01:04:26] Durga Toshniwal: And the range of the values that will be outputted from the sigmoid would be 0 or 1. So, as you can see here, it is a smooth function, it doesn't have an abrupt step, first of all, and the output will be maturing to either 1 or 0. So, the value will be 0,
[01:04:45] Durga Toshniwal: for, negative, values of X, and 1 for positive values of X.
[01:04:54] Durga Toshniwal: So, so now, one important thing to note is that, this is the Y value, and this is the X value.
[01:05:05] Durga Toshniwal: So, as the value of X increases.
[01:05:08] Durga Toshniwal: Then the sigmoid is not going to respond very nicely to the changes in X, because
[01:05:15] Durga Toshniwal: what is going to happen is, here, the change is very, very gradual. So, as the value of X increases.
[01:05:23] Durga Toshniwal: the… the response on the sigmoid becomes very slow. It will change very, very… in a very small fashion.
[01:05:31] Durga Toshniwal: Similarly, if the value of X
[01:05:37] Durga Toshniwal: So, this was when it was increasing. When the value of X decreases, again, the change is going to be very slow, because slowly it will approach 0.
[01:05:47] Durga Toshniwal: That is the reason.
[01:05:49] Durga Toshniwal: the…
[01:05:50] Durga Toshniwal: The sigmoid function, sigmoid X, is given as 1 upon 1 plus e to the power minus X. So I'm just going to show you.
[01:05:59] Durga Toshniwal: This… let's say we are having FX.
[01:06:02] Durga Toshniwal: Which is a sigmoid,
[01:06:08] Durga Toshniwal: function of X, which is 1 upon 1 plus e to the power minus X.
[01:06:14] Durga Toshniwal: So now, How the sigmoid always has a value minus 1?
[01:06:20] Durga Toshniwal: or 0, or sorry, 0 or plus 1. The range of values will be either 0 or plus 1, and that's all.
[01:06:29] Durga Toshniwal: So how that is happening, I'll just show you. Let's say when X is tending to infinity, that is, it is a positively large value.
[01:06:39] Durga Toshniwal: then what will happen? The sigmoid of X is going to be
[01:06:44] Durga Toshniwal: 1 upon 1 plus e to the power minus infinity.
[01:06:50] Durga Toshniwal: Sorry, what has happened.
[01:06:51] Durga Toshniwal: Yes.
[01:06:54] Durga Toshniwal: No, but, if we talk about…
[01:06:59] Durga Toshniwal: the power minus infinity, how much would that be? That is going to be 0. That is a standard thing. Sorry, my digiting writing pad is giving me a little bit of a problem.
[01:07:09] Durga Toshniwal: So, e to the power minus infinity actually is 0 or 10s to 0.
[01:07:14] Durga Toshniwal: So then, in this case, when
[01:07:17] Durga Toshniwal: Extends to infinity, then sigmoid function.
[01:07:22] Durga Toshniwal: of such a value of X will be 1 upon 1 plus e to the power minus infinity, which will be 1 upon 1 plus 0, which is going to be 1.
[01:07:34] Durga Toshniwal: So, which means that…
[01:07:35] Durga Toshniwal: When the value of X tends to infinity, the value of the sigmoid X tends to 1.
[01:07:42] Durga Toshniwal: Before that, it is very, very slowly approaching 1.
[01:07:46] Durga Toshniwal: Similarly, when X is tending to minus infinity, Then,
[01:07:53] Durga Toshniwal: Then what will happen? The sigmoid function
[01:07:59] Durga Toshniwal: on x will be equal to 1 upon 1 plus e to the power minus… minus infinity, right? So this is going to be 1 plus 1
[01:08:09] Durga Toshniwal: To the bar.
[01:08:12] Durga Toshniwal: positive infinity. Now, what is e to the power infinity tending to?
[01:08:19] Durga Toshniwal: So… E to the power infinity, Actually tends to what?
[01:08:28] Durga Toshniwal: E to the power infinity, actually.
[01:08:30] Durga Toshniwal: Oh.
[01:08:32] Durga Toshniwal: Sorry?
[01:08:34] Aditya Banda: e to power… e to the power of infinity, I think it tends to infinity.
[01:08:39] Durga Toshniwal: it tends to infinity, you are right. So e to the power infinity will tend to infinity only, and what we'll have here for the value of
[01:08:49] Durga Toshniwal: Sigmoid on X.
[01:08:51] Durga Toshniwal: Will be won upon.
[01:08:53] Durga Toshniwal: 1 plus infinity.
[01:08:56] Durga Toshniwal: And this will be 1 upon infinity, which will be 0.
[01:09:00] Durga Toshniwal: So, when… X tends to minus infinity, actually. The sigmoid on X will tend to 0.
[01:09:08] Durga Toshniwal: which is what you are seeing here, it is just tending to 0. Actually, it will never touch 0, it will touch 0 at minus infinity. Here, it will not touch 1, it will tend to 1.
[01:09:19] Durga Toshniwal: at plus infinity. So this is how this sigmoid function box.
[01:09:27] Durga Toshniwal: The next is the tanage function, which is actually a variant of, you can say, sigmoid function.
[01:09:33] Durga Toshniwal: So, what is a tanh function? It is 2 upon 1 plus e to the power minus 2x minus 1. This is the standard formula, mathematical formula. The shape is very similar to sigmoid. It's a smooth
[01:09:45] Durga Toshniwal: step function. However, It is centered on origin.
[01:09:51] Durga Toshniwal: So if you look at the previous case, the sigmoid is not centered on origin.
[01:09:57] Durga Toshniwal: It is just lying like this. However, tanh is centered on origin, and sometimes it is useful because a zero-centeredness
[01:10:06] Durga Toshniwal: can help, to, you know, make the model, inputs that may be the negative, positive, neutral, like that. So, that zero-centeredness can be useful sometimes.
[01:10:21] Durga Toshniwal: And if we try to look at the relationship between tan H and sigmoid, then tan h is equal to 2 times sigmoid of 2H minus 1.
[01:10:29] Durga Toshniwal: And the range of tanh is minus 1 to plus 1.
[01:10:33] Durga Toshniwal: So, let's see how this happens.
[01:10:37] Durga Toshniwal: So, tan H as, is actually equal to…
[01:10:40] Durga Toshniwal: As mentioned here, it is 2 times sigmoid of 2x minus 1.
[01:10:45] Durga Toshniwal: So, sigmoid of 2x means 1 plus e to the power minus 2x, because here, minus 2x is there, minus 1.
[01:10:56] Durga Toshniwal: This is a formula, so this is substituted here. So if we try to look at this.
[01:11:02] Durga Toshniwal: Let's see what happens when x tends to infinity.
[01:11:05] Durga Toshniwal: So when x tends to infinity, tan h
[01:11:09] Durga Toshniwal: X will be equal to 2 times 1 plus e to the power minus infinity.
[01:11:16] Durga Toshniwal: We already know a minus 1.
[01:11:19] Durga Toshniwal: And so this is going to be 2 times 1 upon 1 plus 0, because I already told you, e to the power minus infinity tends to 0.
[01:11:29] Durga Toshniwal: And this minus 1, so it will be 2 minus 1, and so it is 1.
[01:11:34] Durga Toshniwal: So, when X tends to infinity, 10HX will tend to 1, which is what you are seeing here as per the graph.
[01:11:41] Durga Toshniwal: Now, let's look at what happens when x tends to minus infinity. When x tends to minus infinity, then 10H of X will be equal to 2 times 1 plus e to the power
[01:11:55] Durga Toshniwal: Minus, minus infinity.
[01:11:58] Durga Toshniwal: Minus one.
[01:12:03] Durga Toshniwal: Sorry, there will be a 2 here, okay, yeah, that's fine.
[01:12:10] Durga Toshniwal: So then,
[01:12:14] Durga Toshniwal: E to the power minus minus infinity is 2 times 1 upon 1 plus e to the power infinity.
[01:12:21] Durga Toshniwal: Minus 1.
[01:12:24] Durga Toshniwal: And e to the power… sorry, this is infinity… e to the power infinity always tends to infinity, so it will be 2 times 1 upon 1 plus infinity.
[01:12:35] Durga Toshniwal: Minus 1, but 1 plus infinity is also infinity.
[01:12:39] Durga Toshniwal: Minus 1, so this will be 0 minus 1, so it will be minus 1.
[01:12:45] Durga Toshniwal: So you can see here, as x tends to minus infinity, the value will tend to minus 1. And accordingly, the range of values will be minus 1 to plus 1 when x tends to minus infinity.
[01:12:57] Durga Toshniwal: 2, placentil.
[01:12:59] Durga Toshniwal: This is how Tanage works.
[01:13:02] Durga Toshniwal: Any questions, anyone?
[01:13:09] Durga Toshniwal: It's just very simple match, nothing much in it.
[01:13:15] Durga Toshniwal: Okay.
[01:13:17] Durga Toshniwal: Then we have the relue function, which is also… it's a very, very popular function, and what does relue say?
[01:13:24] Durga Toshniwal: That it says that the value is max of
[01:13:28] Durga Toshniwal: 0 and Z. If Z is the input, or if X is the input, then it will be X here.
[01:13:36] Durga Toshniwal: So, whatever be the max value, the output will be that only.
[01:13:40] Durga Toshniwal: So if, let's say, we are having some negative value, if X is equal to, say, minus 1, or minus 2,
[01:13:49] Durga Toshniwal: Then… R of X will be equal to max of…
[01:13:55] Durga Toshniwal: 0, minus 2, which will be 0.
[01:13:58] Durga Toshniwal: If X is equal to minus 1,
[01:14:01] Durga Toshniwal: then the value of relu will be, again, max of 0 and minus 1, which will be 0.
[01:14:08] Durga Toshniwal: If X is equal to 0, again.
[01:14:11] Durga Toshniwal: R of X, which is a relu value, will be max of 0, which will be 0.
[01:14:17] Durga Toshniwal: And if X is equal to 1, then R of X will be equal to max of 0, 1, which will be 1.
[01:14:26] Durga Toshniwal: And if X is equal to 2,
[01:14:28] Durga Toshniwal: Then, relu function will be max of…
[01:14:32] Durga Toshniwal: 0, 2, which will be 2, like that. So what can we see? For negative values, it is 0, and for positive values, it's equal to the value itself.
[01:14:42] Durga Toshniwal: This is what the relu is, okay? It's… relu full form is rectified linear activation.
[01:14:49] Durga Toshniwal: And, it is also a piecewise, kind of piecewise linear function.
[01:14:55] Durga Toshniwal: Piecewise linear, but overall it is nonlinear. The two pieces get added together to make a nonlinear, function, so these add to the nonlinearity.
[01:15:08] Durga Toshniwal: So, if the input is positive, the output will be the same as the input. If it is negative, the output will be 0.
[01:15:16] Durga Toshniwal: And very often, ReLU is used as a default activation function.
[01:15:21] Durga Toshniwal: And once the model is… gets improved, then some other functions are tried.
[01:15:27] Durga Toshniwal: Because it outputs, it actually acts like a switch, the relu function.
[01:15:34] Durga Toshniwal: If the input is negative or zero, it's going to open, and no signal is going to flow through it. But if the value is positive, then whatever is the input will be the same as the output. So this is what the value does.
[01:15:49] Durga Toshniwal: then the variant of relu is leaky relu. So, in the relu, we had just zero values for negative…
[01:15:56] Durga Toshniwal: Values of X. In leaky value.
[01:15:59] Durga Toshniwal: there will be some percentage of X, for example, 0.01 of X, that will be outputted for negative value. So, the curve will be something like this or this, and this, of course, is the same.
[01:16:13] Durga Toshniwal: So, leaky value will, allow non-zero values as an output for negative values of X.
[01:16:21] Durga Toshniwal: And, so, this will, you know, prevent
[01:16:27] Durga Toshniwal: The function from behaving like a switch.
[01:16:31] Durga Toshniwal: So, as I told you, in the case of a simple reliew, we have kind of a switch kind of a thing. When value is equal to zero or negative, this switch will open. It won't allow any signal to flow through it. When it is positive, then it will
[01:16:47] Durga Toshniwal: Behave like a closed circuit.
[01:16:50] Durga Toshniwal: But a leaky value?
[01:16:52] Durga Toshniwal: Leaky ReLU, the name is derived that some small amount of current, just like if we have a leaky switch, even if it is off, some small current will… might flow through it. Similarly, we have leaky relu.
[01:17:06] Durga Toshniwal: So, this leaky value is not going to ever allow an open switch. Some small amount of current will always keep flowing, and so…
[01:17:14] Durga Toshniwal: It will always be taken as a closed switch, and it will never…
[01:17:19] Durga Toshniwal: be something like this, so it will never have a…
[01:17:24] Durga Toshniwal: So it will never render the neuron, you know, inactive. The neuron will never be deactivated, it will always remain in the active state because the input is going to flow through it, whether it is small or it is large, whatever it is.
[01:17:46] Durga Toshniwal: Then comes the softmax function. The softmax function, again, is a very, very important form.
[01:17:52] Durga Toshniwal: a function, what it does, it actually…
[01:17:58] Durga Toshniwal: Converts a vector of numbers
[01:18:02] Durga Toshniwal: into a vector of probabilities. And what are these probabilities? Probability value is proportional to the value that is there in the vector.
[01:18:12] Durga Toshniwal: This is what it does. It is designated as Sigma.
[01:18:15] Durga Toshniwal: Z, so this is the vector, so it is designated as Z arrow on it, ZI, so the I-th value is actually given as E to the power ZI divided by the summation of
[01:18:30] Durga Toshniwal: all the values of J1 to K. That is, if there are K number of values in the vector.
[01:18:37] Durga Toshniwal: Then E to the power Z are all added together, which becomes the denominator.
[01:18:42] Durga Toshniwal: Okay, I could see a hand raised, who has a question?
[01:18:49] Pallavi Chakravarty: Hello, ma'am.
[01:18:51] Durga Toshniwal: Yes.
[01:18:51] Pallavi Chakravarty: Yeah, so, I understand that for… I will go, like, to go back to that, tanh function and sigmoid function. So, both of them have, like, the upper and lower limit closer to, like, for tanh, it's minus 1 to 1, and for…
[01:19:09] Pallavi Chakravarty: Sigma, it's 0 to 1.
[01:19:11] Durga Toshniwal: Yep. Oh.
[01:19:13] Pallavi Chakravarty: So, why, like, what's the difference between the two? Like, how should we decide which activation function
[01:19:21] Pallavi Chakravarty: will be better for both of them. Like, for both of them, the gradient is kind of similar, right?
[01:19:26] Pallavi Chakravarty: So…
[01:19:28] Durga Toshniwal: Actually, they are not similar in the sense that if you look at this particular function.
[01:19:35] Durga Toshniwal: It is thought zero center, it is taking up a shape like this, and most importantly, the output range is 0 to 1. So, suppose we want the output to always be positive, then we must choose sigmoid.
[01:19:49] Durga Toshniwal: But, if we are wanting the output to also take up negative values, then we must use 10H. We should not use sigvoid, because that will always give us 0.
[01:19:59] Durga Toshniwal: To one value only. It will never output a 1. So it depends on what we want the output to be like.
[01:20:05] Durga Toshniwal: Accordingly, we use the… The best function.
[01:20:13] Durga Toshniwal: Okay? And this… this function is actually, as you can see, symmetric about 0.
[01:20:20] Durga Toshniwal: So the… the nature of this function is different, even though the shape of the sigmoid and tanet is similar, but the function in itself is different.
[01:20:32] Durga Toshniwal: Okay?
[01:20:35] Pallavi Chakravarty: Okay.
[01:20:38] Durga Toshniwal: Yeah, so I could see… There is some question.
[01:20:44] Durga Toshniwal: Shikar wants to know, I think that's done already.
[01:20:48] Durga Toshniwal: And who else? Do we have any other questions?
[01:20:52] Durga Toshniwal: So Aditya wants to know whether we want to…
[01:20:57] Durga Toshniwal: use, whether we use the sigmoid for a… For a multi-class problem, so, the sigmoid actually…
[01:21:11] Durga Toshniwal: What is happening?
[01:21:13] Durga Toshniwal: I just… Lost, yeah. So the sigmoid actually outputs a 0 or 1. It can be,
[01:21:21] Durga Toshniwal: It is primarily used for a binary class problem.
[01:21:25] Durga Toshniwal: But it can be used for multi-class problems also, it can be used… you will see this tomorrow. You have to just wait till tomorrow to see.
[01:21:36] Durga Toshniwal: How to choose activation functions, and how it all works. So, you'll actually have to wait for a little more time.
[01:21:45] Aditya Banda: Okay.
[01:21:46] Durga Toshniwal: Yeah. So, the most typical one that is used in a multi-class problem is the softmax.
[01:21:53] Durga Toshniwal: That's the most popular one, if you are talking about a multi-class problem, because what it does is that it converts a vector of numbers into a vector of probability, where the probability can be the probability for that class.
[01:22:06] Durga Toshniwal: So, let me give you some examples.
[01:22:12] Durga Toshniwal: Sorry, just a second.
[01:22:15] Durga Toshniwal: So, this is the formula. So, let's say that I'm having a vector like this.
[01:22:21] Durga Toshniwal: So let me say that my Z actually is a vector, containing these 3 values. Now.
[01:22:29] Durga Toshniwal: I want to use this vector.
[01:22:32] Durga Toshniwal: To find out the probability that each of these values I mean, probabilities,
[01:22:40] Durga Toshniwal: I mean, I want to obtain a vector containing probabilities. Obviously, these probabilities can be used
[01:22:46] Durga Toshniwal: as the probability for a class. And the value will be proportional to the value of the vector, of the element of the vector itself. So how it will work?
[01:22:56] Durga Toshniwal: So, let's say I want to talk about the first element.
[01:22:59] Durga Toshniwal: So, the value of the first element will be equal to So, we have here…
[01:23:09] Durga Toshniwal: If we talk about the first element, we will raise it to the power 1.
[01:23:14] Durga Toshniwal: divided by all the elements. So, how many elements are there? One.
[01:23:20] Durga Toshniwal: 2, and so this is 3.
[01:23:23] Durga Toshniwal: Now, if we calculate these, then what we'll have here is e to the bar 1, obviously.
[01:23:30] Durga Toshniwal: We all know it is 2.718.
[01:23:33] Durga Toshniwal: If we raise it to the part 2, it comes out to be 7.389.
[01:23:39] Durga Toshniwal: And if we cube it, it comes out to be 20.
[01:23:42] Durga Toshniwal: 0.086… So, the value of…
[01:23:50] Durga Toshniwal: this… Will be 2.718 divided by summation of all these.
[01:23:58] Durga Toshniwal: It is the power one, I'm not writing the individual values, you can just sum them up.
[01:24:04] Durga Toshniwal: And if we calculate, then what we will find out is it comes out, the probability comes out to be…
[01:24:12] Durga Toshniwal: 0.33.
[01:24:17] Durga Toshniwal: Then, if we do this for the second one.
[01:24:21] Durga Toshniwal: Then it will come out to be 0.09.
[01:24:25] Durga Toshniwal: And… For the third one, it's coming out to be… 0.245.
[01:24:32] Durga Toshniwal: I've just substituted the values as per this formula. For the second one, it will be E to the power 2 divided by E, the same denominator, third one EQ.
[01:24:43] Durga Toshniwal: So, when we do that, what we are getting, we are getting a vector like… 0.033, 0.09.245.
[01:24:53] Durga Toshniwal: So these are the probabilities that we are getting, and this could be related to the probability of a particular class. And in this problem, there are three classes.
[01:25:03] Durga Toshniwal: And accordingly, we get the probability of occurrence of these three classes.
[01:25:08] Durga Toshniwal: So this is how the softmax function is typically used for a multi-class problem.
[01:25:16] Durga Toshniwal: So…
[01:25:18] Durga Toshniwal: I've already shown the calculations, so this is a question how activation function must be chosen. First of all, let me tell you very clearly, there's no standard
[01:25:28] Durga Toshniwal: Way to decide which
[01:25:30] Durga Toshniwal: activation function is to be used, but if we want a compute very less compute-intensive, we can use reliew, because it acts like a open switch when the values are negative, otherwise it's like a closed switch with output equal to input, like that.
[01:25:50] Durga Toshniwal: If we want a classification problem, then sigmoid will be used, it could be a binary classification problem, or a multi-class problem, we can use the other functions, the softmax function, and so on and so forth.
[01:26:05] Durga Toshniwal: ReLU is very commonly used, though.
[01:26:08] Durga Toshniwal: Okay.
[01:26:10] Durga Toshniwal: So, this was about the activation function and other stuff. Now, the most important, one of the most next important thing is hyperparameter tuning.
[01:26:19] Durga Toshniwal: So, hyperparameters are the different parameters that can be used to
[01:26:24] Durga Toshniwal: Control the learning that we… that is involved in a neural network.
[01:26:29] Durga Toshniwal: And there are a large number of hyperparameters, which together control this process of learning.
[01:26:35] Durga Toshniwal: So, we are going to talk about these hyperparameters, and
[01:26:41] Durga Toshniwal: These are tuned in such a way that the loss function.
[01:26:45] Durga Toshniwal: gets, optimally reduced or minimized. So let us see what are these hyperparameters.
[01:26:53] Durga Toshniwal: First of all is Epoch.
[01:26:55] Durga Toshniwal: So, what do we mean by epoch?
[01:26:58] Durga Toshniwal: So, epoch means, first of all.
[01:27:03] Durga Toshniwal: So, epoch means that we are talking about the number of times
[01:27:09] Durga Toshniwal: The learning algorithm, or our classification algorithm, is going to see
[01:27:17] Durga Toshniwal: The entire training data, or the entire data is going… training data is going to be presented how many times is what is epoch.
[01:27:26] Durga Toshniwal: And we generally say that epoch comprises of one complete pass.
[01:27:32] Durga Toshniwal: Over the training data.
[01:27:34] Durga Toshniwal: And usually, to train a deep neural network, we would require multiple epochs, okay? So when we show the full training data to our neural network, what we have is one epoch. We'll typically require multiple epochs.
[01:27:51] Durga Toshniwal: So this is… an e-book.
[01:27:54] Durga Toshniwal: Then we have something called the batch size. Now, what is batch size? So, these things are often confused by many people, so it's important to know what is all… what are all these hyperparameters.
[01:28:06] Durga Toshniwal: So, let's say we have a training data.
[01:28:10] Durga Toshniwal: Then, that training data is actually split into groups called batches.
[01:28:15] Durga Toshniwal: And each batch is actually presented in such a way that it undergoes one iteration.
[01:28:22] Durga Toshniwal: And then, multiple iterations make up one epoch. So, this is the relationship between batch iteration and epoch.
[01:28:30] Durga Toshniwal: So, a data set is decided… is divided into batches. Let's say we are having some thousand samples in our training data, and the batch size is 100, then the total number of batches that we are going to have is 1000 by
[01:28:45] Durga Toshniwal: Number of batches will be?
[01:28:49] Durga Toshniwal: Thousand by?
[01:28:52] Durga Toshniwal: 100, so there'll be 10 batches.
[01:28:55] Durga Toshniwal: So, each batch is going to be presented so that there'll be one iteration per batch, so…
[01:29:03] Durga Toshniwal: There'll be 1 iteration per batch, so if there are 10 batches, then the number of iterations will be 10.
[01:29:09] Durga Toshniwal: per epoch. So, in one epoch, there'll be 10 batches, and accordingly, there'll be 10 iterations. If they have 5 epochs, then the total number of iterations we'll have is 5 into 10 equal to 50. So, this is how batch size works.
[01:29:27] Durga Toshniwal: Now, there are different types of batch sizes, depending upon what gradient descent algorithm we choose. For example, batch gradient descent.
[01:29:36] Durga Toshniwal: So, in this case, the batch gradient descent
[01:29:41] Durga Toshniwal: involves, the entire set as one batch. So, all training samples are presented at one go, and there'll be only one update per epoch, because there aren't multiple iterations, there's going to be only one iteration, and one batch for the entire training set.
[01:29:59] Durga Toshniwal: Then we have mini batch gradient descent. The mini batch gradient descent will have, some small, some batches made out of the given training data. The size would be typically from 16 to 512, and this is the one which is most widely used, mini batch gradient descent.
[01:30:17] Durga Toshniwal: And based on the batch size, the number of iterations will be happening.
[01:30:22] Durga Toshniwal: Buddy Epoch.
[01:30:24] Durga Toshniwal: Then we have the stochastic gradient descent. As explained, it will have the batch size of 1. That is, one training sample will be, resulting into one batch. And so, per sample, there'll be one update in the weights and all those things.
[01:30:42] Durga Toshniwal: So now, how does batch size, you know, determine, what is getting impacted?
[01:30:52] Durga Toshniwal: So, if we have a very small batch size, then some, you know, there'll be some implications of having a small batch size one. The batch can easily fit in the memory.
[01:31:04] Durga Toshniwal: And the update… Will be faster, because the batch size is going to be small.
[01:31:10] Durga Toshniwal: And, since the batch size is small, therefore.
[01:31:14] Durga Toshniwal: It can help in generalization, because the entire data will not be treated at one go.
[01:31:22] Durga Toshniwal: Now, conversely, The large batch size means
[01:31:28] Durga Toshniwal: That the memory requirements… requirements will be very high, because the entire training data has to be placed in memory and taken as one batch.
[01:31:40] Durga Toshniwal: And then, in this case, the update will definitely be more stable, because it's working on the entire data. But, the generalization may not be that very good, because we are just doing one iteration.
[01:31:55] Durga Toshniwal: So… Now, the typical batch sizes, commonly used are 32, 64, and 128.
[01:32:03] Durga Toshniwal: Usually, the batch size will depend on the… what is the training data size, what is the available memory, so that the batch can fit into it. If we have a GPU, then the batch size could be large. If you're a simple CPU, then it needs… it has to be small.
[01:32:21] Durga Toshniwal: And accordingly, the model complexity would vary.
[01:32:24] Durga Toshniwal: So, generally speaking, we could just have a simple intuition, where we could explain what batch size means. Batch size means how much of data the model is going to look at before, you know, look at and before the learning is complete.
[01:32:43] Durga Toshniwal: So, batch… batch size is going to control the amount of data the model is going to look at, and
[01:32:52] Durga Toshniwal: And use it to make updates before the entire learning is over. So you will just shortly understand what is the meaning. I'll show you.
[01:33:03] Durga Toshniwal: First, you need to know what is iteration.
[01:33:06] Durga Toshniwal: So, often, epoch, iteration, batch size, all these batches, they're all confused. Epoch means, when the entire training data
[01:33:18] Durga Toshniwal: Is, shown.
[01:33:21] Durga Toshniwal: to the, classifier. Then what we have is one epoch.
[01:33:26] Durga Toshniwal: Then batch is the size, of the training samples that are taken at one go.
[01:33:33] Durga Toshniwal: What is iteration, then? So, iteration means…
[01:33:37] Durga Toshniwal: It is one complete update of the model parameters, primarily weights and biases.
[01:33:45] Durga Toshniwal: Using one single batch. So, as I told you, the training data is divided into batches.
[01:33:51] Durga Toshniwal: One batch is presented to the model in one iteration.
[01:33:56] Durga Toshniwal: And what does an iteration involve? The iteration actually involves a forward pass.
[01:34:02] Durga Toshniwal: Right, so the input data, which is organized in form of a patch.
[01:34:07] Durga Toshniwal: passes through the entire neural network and makes, helps to make predictions. These predictions are then used to calculate the loss, that is the comparison between the given value and the predicted values done.
[01:34:23] Durga Toshniwal: And then, this loss is back propagated, so backward pass also happens.
[01:34:28] Durga Toshniwal: And when the backward pass happens, then the gradient is calculated with the help of
[01:34:35] Durga Toshniwal: The methods that we already discussed. When the gradients are calculated, then these
[01:34:42] Durga Toshniwal: gradients help to update the weights, right? So that, the… when the weights get adjusted, then the output gets adjusted so that it becomes, nearer to the actual given output.
[01:35:00] Durga Toshniwal: And the weights are adjusted using whatever optimizer has been chosen. Say, for example, a simple gradient descent that we talked about, or maybe ADAM, or whatever.
[01:35:13] Durga Toshniwal: So, this entire cycle is said to be one iteration. So, one iteration involves forward pass.
[01:35:20] Durga Toshniwal: Then it involves the calculation of the loss. Then it involves the back propagation of the loss.
[01:35:27] Durga Toshniwal: to, find out the gradients. Then the gradients are used.
[01:35:32] Durga Toshniwal: To update the weights.
[01:35:35] Durga Toshniwal: So, up to the point where the updation of weight happens.
[01:35:39] Durga Toshniwal: This whole cycle is said to be one iteration.
[01:35:42] Durga Toshniwal: And one iteration is done over one batch, right?
[01:35:47] Durga Toshniwal: So, in an epoch, there might be multiple batches.
[01:35:51] Durga Toshniwal: For using one batch, one iteration is done.
[01:35:55] Durga Toshniwal: So, as many batches there are in one epoch, so many iterations will be there.
[01:36:01] Durga Toshniwal: And then, if there are multiple epochs, then we need to calculate the total number of iterations, that is the number of iterations per epoch into the number of epochs. That will be the total number of iterations.
[01:36:15] Durga Toshniwal: Okay?
[01:36:16] Durga Toshniwal: So, here is another example, iteration versus epoch, which is very much confused.
[01:36:23] Durga Toshniwal: By many. So, epoch is one full pass over the entire training data, whereas iteration means
[01:36:29] Durga Toshniwal: It is one batch that we are looking at, not the entire data. And it also involves the whole process of forward, backward, I mean, forward pass, loss calculation, back propagation, and then update. So this whole thing is iteration.
[01:36:48] Durga Toshniwal: So, for example, if we have a dataset, training dataset of 1,000 samples, and the batch size is 100,
[01:36:55] Durga Toshniwal: Then, the number of iterations.
[01:36:58] Durga Toshniwal: Per epoch will be equal to the number of batches, which is 1000 by 100, which is equal to 10.
[01:37:05] Durga Toshniwal: So, this means that, Paripo, there'll be 10 iterations. Okay.
[01:37:13] Durga Toshniwal: So, iterations are important because they help, they, they cover, the updation of weights.
[01:37:23] Durga Toshniwal: So, the more number of iterations, the more number of times the weight will be updated, and the more compute-intensive our process will be.
[01:37:32] Durga Toshniwal: Obviously, if there are too many iterations, the weight will be updated again and again in such a way that the predicted output is very close to the given output for the training data. Or in other words, our classifier will get overfitted.
[01:37:51] Durga Toshniwal: On the training data, because the weights are getting, customized too much on the basis of the training data.
[01:38:00] Durga Toshniwal: If there are very less iterations, then what we'll have, the loss will be high, and the model will be underfitted.
[01:38:07] Durga Toshniwal: So this is the outcome of…
[01:38:09] Durga Toshniwal: you know, wearing the weights. I could see a hand raised. Anyone is having a question?
[01:38:19] Durga Toshniwal: Is there a question anymore?
[01:38:24] Durga Toshniwal: Okay, I'll proceed.
[01:38:26] Sacheen Adavinavar: And, how to choose the, batch size, ma'am? So you explained that,
[01:38:31] Sacheen Adavinavar: 16 and 32, right? So, they're, like, if you are choosing, more in the single epoch, so that is going to be, like, faster, and then,
[01:38:45] Sacheen Adavinavar: This one, this slide.
[01:38:48] Sacheen Adavinavar: So, first of all.
[01:38:50] Durga Toshniwal: Epoch is one full pass.
[01:38:53] Durga Toshniwal: on the… Entire data. Okay.
[01:38:58] Durga Toshniwal: So, this means the full data.
[01:39:02] Durga Toshniwal: Now, EFOC is made up of… is broken down into patches.
[01:39:07] Durga Toshniwal: Batch size could be something, say 32 samples, 64 samples, 128 samples, like that.
[01:39:15] Durga Toshniwal: So now the… your question is, how do we decide which batch size to choose?
[01:39:19] Sacheen Adavinavar: Right.
[01:39:20] Sacheen Adavinavar: Yeah, yes.
[01:39:21] Durga Toshniwal: So there is no standard solution to decide what is a batch size. What you will have to do is you'll have to iterate with different batch sizes.
[01:39:30] Durga Toshniwal: And, this is what is the process of hyperparameter tuning, and then take up that value that gives the minimal loss.
[01:39:40] Durga Toshniwal: Yes.
[01:39:41] Durga Toshniwal: These are the common batch sizes that are taken.
[01:39:46] Aditya Banda: Can you go back to that slide where you wrote epoch and batches? Yeah. When you say one epoch, epoch is equal to one full pass, do you mean full forward and backward pass, and updating the weights and biases?
[01:40:03] Durga Toshniwal: So, when I say one full pass, what it means is that
[01:40:08] Durga Toshniwal: The entire training data is shown
[01:40:12] Durga Toshniwal: to the, classifier, or the model. This is what generally is one epoch, okay? So here, we are not talking about…
[01:40:22] Durga Toshniwal: making batches, or we are not talking about iterations and all, but we are actually… yeah, so iterations are a part of epoch only, so one full pass will involve everything, but the data is taken in one go. You can assume that the full data has been seen by the
[01:40:41] Durga Toshniwal: Model. This is what is one epoch. So, so you see that, when the full data has been training data, it's actually training data.
[01:40:54] Durga Toshniwal: So, when the full training data has been shown to the model, it means that it has learned
[01:41:01] Durga Toshniwal: What was available in the training data?
[01:41:04] Durga Toshniwal: Now, the thing is that it is assumed, though of course showing the entire training data.
[01:41:09] Durga Toshniwal: Means… should have meant that the model would have learned it entirely, but the thing is that
[01:41:16] Durga Toshniwal: It doesn't learn it entirely, even though it sees it.
[01:41:21] Durga Toshniwal: So, therefore, usually, there are multiple epochs.
[01:41:25] Durga Toshniwal: So, we keep on showing the entire, training data to the model so that, with an aim to improve its performance. So, as it is shown again and again.
[01:41:38] Durga Toshniwal: its performance increases. It's just like, you know, a human. You take a book and try to memorize some concepts, even though you would have memorized
[01:41:49] Durga Toshniwal: the book once.
[01:41:51] Durga Toshniwal: your, you know, replication will be there, but it may not be so very good. Now, once you see the book again, it may improve. Then again, it may improve further. But if you go on doing it again and again.
[01:42:05] Durga Toshniwal: then the generalization will be lost, so any generic answer cannot be answered. I mean, you won't be able to answer.
[01:42:12] Aditya Banda: Similarly here, if there are too many epochs.
[01:42:15] Durga Toshniwal: then the model becomes overfitted, because then the output, the predicted output, we tend to have it exactly same as the human output. And if there are lesser number of epochs, then it is underfitted. So this is the whole thing, okay?
[01:42:31] Aditya Banda: Sure.
[01:42:32] Durga Toshniwal: Yeah, so here…
[01:42:34] Durga Toshniwal: So then, we can just think of an analogy, where one iteration would actually mean, looking at or practicing one set of questions. One epoch means the entire question bank.
[01:42:49] Durga Toshniwal: And training means, we keep on solving the full question bank till we would have learned whatever we wanted to learn in a
[01:42:59] Durga Toshniwal: good way.
[01:43:02] Durga Toshniwal: So now,
[01:43:04] Durga Toshniwal: since we talked about epoch, batch size, and iteration, so I also thought we'll mention training. So, training is the…
[01:43:12] Durga Toshniwal: process of making the neural network learn the, you know, the entire data. And this not just involves the forward pass, it involves weight updations and everything that goes into the training.
[01:43:26] Durga Toshniwal: So the training will definitely involve multiple epochs, because, as I explained to you, as the model sees the training data more and more, it improves, its performance improves.
[01:43:40] Durga Toshniwal: So, for example, if we talk about the dataset and the batch, then the training dataset, if it has a size N, then it comprises of all the training samples. If we talk about batch B,
[01:43:52] Durga Toshniwal: then this batch B will be a portion of this N, which is a small chunk of the training data, and the batch size B will be the number of samples in one batch.
[01:44:01] Durga Toshniwal: So, here's an example, which I already, I think we have covered this earlier. Data set size 1,000, batch size 100. So, number of batches will be 1000 on 100, which will be 10.
[01:44:16] Durga Toshniwal: So, iteration is just like, it's similar, it's also called a step, and it will involve one forward pass, one backward pass, loss calculation, weight updation, using
[01:44:28] Durga Toshniwal: One batch of data.
[01:44:31] Durga Toshniwal: Okay, so often iterations are also said to be steps in many libraries, so you can, you know, take it like that.
[01:44:40] Durga Toshniwal: So this is what is iteration.
[01:44:42] Durga Toshniwal: And what is epoch? Epoch is full pass over the entire training data, and it could involve multiple iterations. So if you talk about the number of iterations per epoch.
[01:44:53] Durga Toshniwal: That will be the data size divided by batch size, and then the total number of iterations, that will be number of epochs into iterations per epoch.
[01:45:04] Durga Toshniwal: So, for example, if the number of epochs is 5,
[01:45:08] Durga Toshniwal: Then, the total number of iterations… so iterations per epoch will be how many? They will be thousand upon hundred, because the batch size is 100.
[01:45:18] Durga Toshniwal: So, number of batches will be 10, so the number of iterations will be 10.
[01:45:23] Durga Toshniwal: And there are 5 epochs, so it will be 5 into 10 total 50 iterations.
[01:45:28] Durga Toshniwal: In, this 5e box.
[01:45:33] Durga Toshniwal: So this is just the relationship, made a little more simple. We have the entire training process, and the training could involve multiple epochs.
[01:45:43] Durga Toshniwal: In each epoch, we have multiple batches. Say, if training data has samples equal to 1,000, as we mentioned.
[01:45:52] Durga Toshniwal: And the batch's batch size is, say, 100, Then, total number of batches…
[01:46:03] Durga Toshniwal: will be equal to 1,100, which is equal to 10. So, there'll be 10 batches.
[01:46:09] Durga Toshniwal: per epoch.
[01:46:14] Durga Toshniwal: So, in one epoch, there are 10 batches. Second epoch, 10 batches, third epoch, like that.
[01:46:19] Durga Toshniwal: If the total number of epochs are taken to be equal to 5, then there'll be 1010 batches.
[01:46:27] Durga Toshniwal: per epoch.
[01:46:28] Durga Toshniwal: Now, when we have one batch, this batch is presented to
[01:46:33] Durga Toshniwal: what we call one iteration. So, per batch, there'll be one iteration. So, for 10 batches, 10 iterations. For next 10, there'll be next 10. So, the total number of iterations will be the number of, number of batches into number of epochs. So, here.
[01:46:51] Durga Toshniwal: number of… Iterations, as per the formula also that was given, will be.
[01:46:58] Durga Toshniwal: Number of batches.
[01:47:02] Durga Toshniwal: Body book?
[01:47:05] Durga Toshniwal: into number of… Epochs.
[01:47:10] Durga Toshniwal: Which is equal to, in this case, 10 batches per epoch into 5 epochs, so it is 50.
[01:47:17] Durga Toshniwal: So this is the relationship between training data, epoch, batch, iteration, And,
[01:47:25] Durga Toshniwal: This is how they work.
[01:47:27] Durga Toshniwal: Any questions on any of these? Epoch, iterations, badge, and all that?
[01:47:33] Durga Toshniwal: Anyone?
[01:47:44] Durga Toshniwal: Then, there are some other hyperparameters, like the number of layers. So, this is also a hyperparameter.
[01:47:51] Durga Toshniwal: So, the number of input and output layer are one each, so that is fixed. But what is, varying is the number of hidden layers, which is the number of layers.
[01:48:04] Durga Toshniwal: So, number of hidden layers are the number of layers that we are talking about, and this can vary.
[01:48:09] Durga Toshniwal: Then the next hyperparameter is the size of the layer, so the number of neurons that are contained in a layer is the size of the layer.
[01:48:17] Durga Toshniwal: And again, this can be varied so that the loss gets reduced.
[01:48:21] Durga Toshniwal: So you can see here, these are the hidden layers, hidden layer 1.
[01:48:25] Durga Toshniwal: Hidden led to… 8 under 3, so this is 8 under 1, this is 8 under 2.
[01:48:32] Durga Toshniwal: 3, you don't have 4, hidden there, 5, like that.
[01:48:36] Durga Toshniwal: So these are all 5 hidden labs.
[01:48:39] Durga Toshniwal: And they are having some number of neurons, which is the size of the neck.
[01:48:43] Durga Toshniwal: Say, 4 in this case, 5 years, 6 years, like that.
[01:48:51] Durga Toshniwal: The next thing is loss functions. We have already discussed the loss functions, which are the mean square error, mean absolute error.
[01:48:59] Durga Toshniwal: cross-entropy loss, I will show you what it means.
[01:49:03] Durga Toshniwal: So, these are the formula. You already know what is mean square error, or the…
[01:49:07] Durga Toshniwal: L2 loss, which is the difference between the given value and the predicted value squared.
[01:49:15] Durga Toshniwal: and then summed over on all n samples, and then averaged out. This is the mean squared error. So, error is squared, and then the mean of it is found out by dividing it by N.
[01:49:28] Durga Toshniwal: The next thing is the mean absolute error, which is also called the L1 loss. So, MA is the absolute difference between the
[01:49:35] Durga Toshniwal: a given output and the predicted output, and this done over all samples divided by the total number of samples. So, this is the average error. So, this is the mean average error.
[01:49:46] Durga Toshniwal: Then we have the mean bias error, so here, the difference between the…
[01:49:50] Durga Toshniwal: the biases for, the, the difference between the, sorry, the predicted… the given value and the predicted value statement. It is not the absolute difference, it is just a simple difference.
[01:50:06] Durga Toshniwal: All of this is added, and then the mean is found out.
[01:50:10] Durga Toshniwal: Then we have the cross entropy loss, or the negative log likelihood, and the formula is given here. So, if Yi hat is the predicted value, and Yi is the given value, so it is the log of Y hi hat times Yi, plus log of 1 minus Yi hat.
[01:50:29] Durga Toshniwal: into 1 minus YN. So this is the cross entropy loss, it is a very popular loss.
[01:50:35] Durga Toshniwal: formula that is used.
[01:50:40] Durga Toshniwal: So I think with this, I will wrap up here. Any questions on whatever I've done? Because we'll do back propagation in the next time in more details.
[01:50:54] Durga Toshniwal: Any questions, anyone?
[01:50:59] Durga Toshniwal: Is it becoming clear? Are there any… Problems…
[01:51:04] Durga Toshniwal: Anything that you want to say or comment?
[01:51:08] Durga Toshniwal: I hope.
[01:51:09] Muni Prakash Ganji: If you do some calculations with this, some max, it will be helpful to us doing understanding.
[01:51:15] Muni Prakash Ganji: Because for me, the eco's…
[01:51:17] Muni Prakash Ganji: You were saying the echo equal to number of echoes. I'm getting some doubt in there.
[01:51:23] Muni Prakash Ganji: The calculation of the epoch?
[01:51:26] Durga Toshniwal: Sorry.
[01:51:26] Muni Prakash Ganji: Yeah, that means, suppose, If we do some iterations until the complete dataset, we're calling as one echo, correct?
[01:51:34] Muni Prakash Ganji: And in that, depend upon the batches, again, we are dividing into the
[01:51:40] Muni Prakash Ganji: Some echoes. Like, one echo equal to number of echoes you were saying. Like, like, fire.
[01:51:45] Muni Prakash Ganji: it goes into 50 iterations equal to 50, we are saying correct. If you do some maths, it will be helpful to us in understanding.
[01:51:53] Durga Toshniwal: Yeah, so I'll just show you, at least for this, I already showed you an example. See, EPOC is the full training data, but full training data may not be…
[01:52:03] Durga Toshniwal: shown to the model at one time. Usually, it is divided into groups called patches.
[01:52:11] Durga Toshniwal: And one batch is passed to one iteration. Okay, so this means that per batch, there will be one iteration.
[01:52:19] Durga Toshniwal: Bye.
[01:52:20] Durga Toshniwal: So, if we are having in one epoch, we are having 10 batches, then in one epoch, we'll be having 10 iterations.
[01:52:27] Durga Toshniwal: But let's say we are having, in total, 5 epochs.
[01:52:31] Durga Toshniwal: And, so the total number of iterations, how many there will be, there'll be 10-10 iterations per epoch, so the sum total of all the iterations will be 10 into 5, which is 50.
[01:52:44] Durga Toshniwal: Okay, so the batches will not be 50, the batches will be 10 only, because we are dividing the given data into groups. The what is the given data? The given data is 1 epoch, and one epoch is divided, let's say, into 10 groups, then there'll be 10 batches only. They won't be 50.
[01:53:04] Durga Toshniwal: But the iterations will be 50, because there are 5 epochs, and per epoch, there'll be
[01:53:11] Durga Toshniwal: 10 iterations, or 10 batches, like that. So I had shown you that calculation.
[01:53:16] Muni Prakash Ganji: Indonesia.
[01:53:16] Durga Toshniwal: relationship also.
[01:53:18] Durga Toshniwal: Yeah.
[01:53:22] Durga Toshniwal: So, what I will do is, we will also do a running example.
[01:53:28] Durga Toshniwal: Okay? In which I will show you, using some activation function, how we actually, find out the output, and then we do a back propagation of the error, and all that I will show you. But to show that, I will need to
[01:53:45] Durga Toshniwal: Show you the derivation of backpropagation.
[01:53:50] Durga Toshniwal: So I will do that in one of the forthcoming tons. I will show you how back propagation
[01:53:55] Durga Toshniwal: helps to update the weights, how they are updated using backpropagation. It's a little complex, I mean, it's not complex, but it's a derivation after all, so it will take some time.
[01:54:06] Durga Toshniwal: And you can understand, and then this will be used throughout. It will be used in deep learning, it will be used in NLP, it will be used in all the models.
[01:54:18] Durga Toshniwal: the back propagation of errors and the weight updates. That will be used everywhere. So I will explain to you once, and you can understand it, and then
[01:54:28] Durga Toshniwal: You can then assume that you're understood, and it's going to be used like that.
[01:54:33] Durga Toshniwal: In the further models.
[01:54:35] Durga Toshniwal: Okay.
[01:54:36] Durga Toshniwal: I think I saw a question… on something.
[01:54:45] Durga Toshniwal: So I think there was a question on some…
[01:54:48] Durga Toshniwal: Some detailed notes or a reference book.
[01:54:51] Durga Toshniwal: So, there are many standard books on deep learning. I can, actually, in the start of the…
[01:54:57] Durga Toshniwal: course. I had given you a list of books, including books on deep learning also. You can consult
[01:55:05] Durga Toshniwal: Books on machine learning, deep learning, all are listed there.
[01:55:09] Durga Toshniwal: In the, I think, first or second lecture I had given you. So you can go back and please have a look.
[01:55:15] Durga Toshniwal: already I gave you. I think Adity has written about some particular book.
[01:55:21] Durga Toshniwal: For deep learning and hyperparameter and all.
[01:55:24] Durga Toshniwal: You can also use this.
[01:55:29] Chandrasekhar Sahu: Ma'am, like, we'll be going to cover, TensorFlow also,
[01:55:35] Durga Toshniwal: So, you'll be using Tensor starting tomorrow.
[01:55:39] Chandrasekhar Sahu: Okay.
[01:55:40] Durga Toshniwal: Yeah.
[01:55:41] Durga Toshniwal: As I told you right in the beginning, that you'll need to know tensors also.
[01:55:46] Durga Toshniwal: Because deep learning you do with Tensor, that's much easier. Tomorrow, you'll be using it, and now onwards, you will be using it for all the hands-on.
[01:55:57] Chandrasekhar Sahu: Okay, ma'am, thank you.
[01:56:00] Durga Toshniwal: So, I think there's a question, is which one is better, LU or Sigmoid? Actually, it depends. It depends on what we want to do such in. So, if,
[01:56:09] Durga Toshniwal: You know, we want the output to be used in a classifier, then we can use sigmoid, because it gives a 0 or 1, kind of a binary classifier.
[01:56:20] Durga Toshniwal: Value is, more useful, it's more simpler.
[01:56:25] Durga Toshniwal: And, less compute-intensive. So, it depends what you want to do with the output. So ReLU is more useful when we just simply want to see what is the predicted output, not necessarily for a classification problem.
[01:56:41] Durga Toshniwal: But it is compute, less compute intensive, it is much simpler than Sigmoid.
[01:56:46] Sacheen Adavinavar: Okay, it is basically based on the use case.
[01:56:49] Durga Toshniwal: Exactly, exactly.
[01:56:52] Durga Toshniwal: In fact, needless to say that all the deep learning and the generative AI models
[01:56:59] Durga Toshniwal: There is nothing standard, no… no written rules, nothing like that. These are all thumb rules. You need to decide which model to use when, how many layers, what is the number of neurons, but everything is just kind of tentative only.
[01:57:15] Durga Toshniwal: There's nothing fixed solution. You look at the use case, and accordingly you decide what to do.
[01:57:23] Durga Toshniwal: Any other questions, anyone?
[01:57:36] Durga Toshniwal: Okay, so if there are no further questions, we'll wrap up for today.
[01:57:41] Durga Toshniwal: You can still, you know.
[01:57:44] Durga Toshniwal: think about whatever we have learned and all that, and you can ask in the subsequent lectures also, because this really very important part. It will be used throughout the rest of this course, so you need to understand it very, very nicely.
[01:58:00] Durga Toshniwal: Okay then, with this, we'll wrap up if there are no further questions.
[01:58:07] Sacheen Adavinavar: Okay, thank you.
[01:58:09] Durga Toshniwal: Thank you, thank you all, have a great evening. Bye-bye. We'll be tomorrow.
[01:58:14] Durga Toshniwal: Thank you. Bye-bye.
[01:58:16] Krishnakumar MS: Thank you.