# 09 2026-02-22 Basics of NLP course: Module 3 — Deep Learning & NLP module: Module-3-Deep-Learning-NLP date: 2026-02-22 type: transcript video_url: https://personal-learn.armco.dev/files/_Recordings/Module-3-Deep-Learning-NLP/09_2026-02-22_Basics_of_NLP.mp4 --- [00:00:12] gautam Jha: San Pieces. [00:00:36] gautam Jha: Thank you, sir. [00:00:39] gautam Jha: Chairman, sir. [00:00:45] gautam Jha: This reform express derail nahiwi. [00:00:54] gautam Jha: Kashi. Kashi hogi. [00:01:02] gautam Jha: Problems tahun gi. [00:01:23] gautam Jha: It's Ramadan? [00:01:28] gautam Jha: Crooked. [00:01:29] gautam Jha: website, milieu. [00:01:36] gautam Jha: Pointed your way. [00:01:40] gautam Jha: That is my name. [00:01:54] gautam Jha: Austin Javier. [00:02:06] gautam Jha: Sorry, bidla. [00:02:10] gautam Jha: piece of respect, bro. [00:02:27] gautam Jha: It's gonna take [00:02:39] gautam Jha: Do you want to die? [00:11:34] Durga Toshniwal: A very good morning to all of you, and welcome to today's session. [00:11:38] Durga Toshniwal: So, I'll be just starting in one or two minutes, so… [00:11:42] Durga Toshniwal: Let's see a few more people join. [00:12:48] Durga Toshniwal: So, welcome, and I'll just share my slides, and we'll go on from there. [00:12:55] Durga Toshniwal: Allowing you a minute. I'll just… Shift. [00:13:17] Durga Toshniwal: So, I hope my screen is visible to everyone. [00:13:24] Durga Toshniwal: So, let's start with the agenda for today. We are actually going to study recurrent neural networks, or briefly called RNNs. [00:13:33] Durga Toshniwal: They are very important, components, in deep learning. [00:13:39] Durga Toshniwal: So the agenda would be to… Discuss sequence models, discuss about [00:13:45] Durga Toshniwal: applications of RNN, the theory of RNN, and… [00:13:50] Durga Toshniwal: gradient descent as it is applicable to RNNs. [00:13:54] Durga Toshniwal: So, first of all, what are sequence models? So, sequence models in machine learning are those, [00:14:02] Durga Toshniwal: Or those models, which accept sequence as inputs, and sequence as outputs. [00:14:08] Durga Toshniwal: So, what can be a sequence? [00:14:12] Durga Toshniwal: Sequence could be a sequence of numbers. It could be a sequence of text. [00:14:16] Durga Toshniwal: For example, if we are speaking something, then whatever we are speaking, that text is a sequence, and that sequence would be fed into a sequence model. It could be an audio clip where somebody is speaking. [00:14:30] Durga Toshniwal: So, so, one, you know, the voice… [00:14:35] Durga Toshniwal: Or the audio will always have a sequence of whatever is spoken. [00:14:41] Durga Toshniwal: Similarly. [00:14:42] Durga Toshniwal: Video is a sequence of frames. Then we have, as I said, sequence of numbers, which could be time series data. For example, it could be stock data, or some other weather data, like that. So all these are examples of sequences. [00:14:57] Durga Toshniwal: And recurrent neural network, or RNNs, are very important and popular algorithms that are used to handle sequences. [00:15:08] Durga Toshniwal: So now, if we talk about applications of sequence models, so, if we talk about, speech recognition, so the input to this kind of a data is speech. Just a second. [00:15:25] Durga Toshniwal: So, if we talk about the input, the input here is a speech. [00:15:32] Durga Toshniwal: And what we get output is, suppose we want, because we want to, see what is, you know, the text that is being spoken about in the audio. [00:15:45] Durga Toshniwal: Which is the recognition of what is being spoken. Say, for example, there's a speech signal, which is an audio, which is going inside, and what you are getting is the… [00:15:54] Durga Toshniwal: text conversion of it. So, we get a text transcript, which says how cold is it outside, which is what was spoken, like that. So, in this case, both the input as well as the output are sequences. The input is a sequence of what has been spoken, so it's an audio clip, and the output is the text. [00:16:13] Durga Toshniwal: The, conversion of the, the output is the text conversion of that audio. [00:16:19] Durga Toshniwal: So, again, that is a sequence. [00:16:24] Durga Toshniwal: Then, another example of sequence model is in sentiment classification. So what we give here is an input of words. [00:16:32] Durga Toshniwal: Say, for example, this movie is fantastic, I really like it because it's so good. So, here, what we have is an input text, and based on that, the output here is going to be the sentiment. [00:16:46] Durga Toshniwal: Which is represented in form of a 5-star rating. Say, for example, if it is very good, then there will be a rating which will be high, say 4-star. [00:16:57] Durga Toshniwal: So, like that. [00:17:03] Durga Toshniwal: Just a second. [00:17:07] Durga Toshniwal: Then what we have… I think I lost my pointer. [00:17:13] Durga Toshniwal: Then, another form of application of a sequence model is, in, video. [00:17:22] Durga Toshniwal: So, we have a video, which a video is nothing but a set of frames. So, you can see that there's someone, which is shown running in the video, and the activity of running is represented by these [00:17:35] Durga Toshniwal: By these frames. [00:17:38] Durga Toshniwal: And this, sequence of video frames serves as the input to the sequence model, which could be an RNN. And the output is, activity recognition for that particular video. For example, the activity can be running here. [00:17:56] Durga Toshniwal: Another example of a sequence model is an image captioning. So here, what we have is an image, which is an input. [00:18:03] Durga Toshniwal: So, what we give is a single input here, which is an input, and what we get, as output is a sequence of words. [00:18:12] Durga Toshniwal: Right, so here, what we are having is, an output, which is in the form of a sequence of words, which is the description. [00:18:20] Durga Toshniwal: For that particular, [00:18:23] Durga Toshniwal: image. For example, if we talk about the Indian flag, then the sequences, the output is representing the description, which says this is the Indian flag. [00:18:36] Durga Toshniwal: Then we have language translation as another application, where the input is a sequence, which is in form of some language, and the output is also a sequence, and it is in the form of another language. [00:18:48] Durga Toshniwal: So, for example, you can see that the input is a sequence in form of Hindi, and the output is another sequence which is in form of English. [00:18:59] Durga Toshniwal: Now, let's talk about RNNs. [00:19:02] Durga Toshniwal: So, RNNs are very powerful, and they belong to the category of algorithms that have some internal memory. [00:19:12] Durga Toshniwal: So far, we talked about simple artificial neural networks, and such simple neural networks do not have any memory, and we discussed them also. All they have is an input, and they have hidden layer, which consists of neurons, and then we have output. [00:19:28] Durga Toshniwal: So, there is no internal memory. However, there are many deep learning applications that require the usage of some form of internal memory, which the recurrent neural network is having. So, RNNs were initially proposed in 1980, [00:19:44] Durga Toshniwal: However, they were not put to much use until, recently, around a decade back or so. [00:19:52] Durga Toshniwal: And, there, the use of RNNs, has become now, of late, quite popular, and of course, it requires computational power. [00:20:02] Durga Toshniwal: So, RNNs are very popular because of their internal memory. RNNs can remember the input data to some extent. [00:20:12] Durga Toshniwal: And that's why RNNs are very popular and very much used [00:20:17] Durga Toshniwal: for making prediction for sequences. We discussed sequence models. RNN is a sequence model because it takes as input a sequence. It can take as input a sequence. It can also take as input a single, [00:20:32] Durga Toshniwal: number, or a single data also, but its, maximum potential is in prediction that involve inputs in form of sequences. [00:20:46] Durga Toshniwal: And since they have memory, so they can remember the sequence, and therefore, their prediction comes in very nicely when we are having these sequences, which cannot be remembered by other forms of [00:21:00] Durga Toshniwal: artificial neural networks. [00:21:04] Durga Toshniwal: Okay, just a second, I'm having some issue with my display. [00:21:09] Durga Toshniwal: Just allow me a moment. [00:21:20] Durga Toshniwal: Just a second. [00:21:28] Durga Toshniwal: I'm having some issue, and… My cursor is also gone. [00:21:34] Durga Toshniwal: What's going wrong? [00:21:50] Durga Toshniwal: So I… I think my system is Rosen. [00:21:56] Durga Toshniwal: Oh. [00:21:59] Durga Toshniwal: I'm not able to do anything. [00:22:07] vinit shah: Maybe you can restart it, ma'am? [00:22:10] Durga Toshniwal: Yeah, I wanted to do a stop share, but… [00:22:14] Durga Toshniwal: So I can't do anything because I can't find my cursor is… Not visible to me. [00:22:21] vinit shah: And just force restart on. [00:22:23] vinit shah: Forced power shutter. [00:22:25] Durga Toshniwal: Yeah, so I actually don't want to do a shutdown of this laptop. [00:25:00] GenAI Batch-2 Manager: Hi, everyone. So ma'am is facing some issues with her laptop, as she mentioned. So just give her a couple of minutes. It may take some time, she's trying to sort it out. She'll join as soon as it's… [00:28:44] Durga Toshniwal: Yeah, sorry for that. Actually, my laptop just was hanging. So, we'll continue from where we left. Sorry for that. Just allow me to share my slides. [00:28:54] Durga Toshniwal: Just a second, please. [00:29:06] Durga Toshniwal: Yes, I'm doing that. [00:29:44] Durga Toshniwal: So, we had already covered part of it, so I'm just going to skip through what we covered. [00:29:55] Durga Toshniwal: Okay. [00:29:56] Durga Toshniwal: So now let's talk about how RNNs work. So I'm going to… [00:30:01] Durga Toshniwal: Explain all that. Okay, sorry, just a second, this is not what I wanted to share. [00:30:07] Durga Toshniwal: Sorry for that. [00:30:16] Durga Toshniwal: Just a second. [00:30:28] Durga Toshniwal: I'm just reassured. [00:31:05] Durga Toshniwal: Okay. [00:31:06] Durga Toshniwal: So, now let's see how recurrent neural networks actually work. So, for that, let's consider, the stock prices. [00:31:15] Durga Toshniwal: And we want to see how the stock prices change over time. So, for here, we have a price of a stock, as you can see here. [00:31:26] Durga Toshniwal: So, here is the price on the… [00:31:29] Durga Toshniwal: Y-axis, and this is the days on the x-axis. And we want to use [00:31:36] Durga Toshniwal: This time series data. It is called time series because we are measuring the price over every day, day 1, day 2, day 3, so it is measured with respect to time. [00:31:47] Durga Toshniwal: Here, the price of the stock is measured with respect to time, and hence it is called a time series data. [00:31:53] Durga Toshniwal: And we want to use this data for making a prediction. [00:31:57] Durga Toshniwal: So we can see here that for the first four days, that is the first day, second day, third day, fourth day, we can see that the price was going up. [00:32:06] Durga Toshniwal: And, And then we want to predict what will happen in future. [00:32:11] Durga Toshniwal: And also, we might also have some other stock price, for another company. [00:32:17] Durga Toshniwal: In which case, we can see, as we can see here, there's much more data. So here, we just have data for [00:32:24] Durga Toshniwal: 1 day, 2-day, 3-day, 4 day. These are… this is the data. Whereas for other, stock price, which is shown in the blue color, we have data over 1, 2, 3, 4, 5, 6, 7, 8 days, so it's almost double the data that we have. [00:32:42] Durga Toshniwal: And, yet we want to use this longer prediction, longer time series to make a prediction. [00:32:49] Durga Toshniwal: So, we have now got two stock, data based on two companies, one shown in the blue color, which is a longer one, and one shown in the red color, which is a shorter one. And we want to use our neural network to predict the stock prices for each of these blue, as well as the red. [00:33:09] Durga Toshniwal: Stock data for two different companies. [00:33:14] Durga Toshniwal: And, therefore, the problem at hand that we have is that the neural network must be able to handle different amounts of sequential data. So this is sequential data, you can see that it is coming. [00:33:27] Durga Toshniwal: Inform us of sequence, this is the blue, stock price, I mean, data represented by the blue graph, data represented by the red graph, again, is a sequence, and both of these sequences are of different lengths. [00:33:43] Durga Toshniwal: So, a normal neural network may not be able to cater, first of all, to sequential data, because it doesn't consider the relationships in the data, in the input data. [00:33:55] Durga Toshniwal: It just considers data as input data. Moreover, the data may be in form of sequence that may have different lengths, so it is… the normal neural network also cannot cater to this kind of [00:34:07] Durga Toshniwal: Variable input sizes. [00:34:12] Durga Toshniwal: So, now we want to use this to make a prediction on the next day, which may be the 10th day, or the 9th day, or whatever. [00:34:20] Durga Toshniwal: So, [00:34:23] Durga Toshniwal: In one case, we are having some 5 preceding days. In the other case, we may are having some 9 preceding days, and we want to predict the data for the 10th day, or for the [00:34:36] Durga Toshniwal: Sixth day, like that. [00:34:39] Durga Toshniwal: So, our neural network must be flexible in terms of the amount of sequential data that we want to use to make it… to make the prediction. It could be 9, 10 days, it could be 6 days, it could be 11 days, and so on and so forth. [00:34:53] Durga Toshniwal: So now, the form of neural network that is able to handle sequences of variable lengths are recurrent neural networks, or RNNs. [00:35:04] Durga Toshniwal: So we are going to now use, RNNs to solve this problem. [00:35:10] Durga Toshniwal: And, just like any other neural network, RNNs also will be having weights and biases. So, this is just an example of a recurrent neural network that you're seeing on the slide. It will have weights like W1, B1, biases, weights like W1, bias like B1. [00:35:28] Durga Toshniwal: Another weight, W3 virus. [00:35:32] Durga Toshniwal: B2, and so on and so forth. There'll be an input. [00:35:36] Durga Toshniwal: And there'll be an output. [00:35:39] Durga Toshniwal: So, all these will be there. [00:35:42] Durga Toshniwal: Additionally, there might be something else that I'll discuss. [00:35:49] Durga Toshniwal: So now, as we just discussed that… [00:35:55] Durga Toshniwal: There are weights, there are biases. [00:35:59] Durga Toshniwal: And… [00:36:00] Durga Toshniwal: In addition to the input and output node, we have these, you know, nodes, which are the nodes inside the hidden layer, and they are also having the activation function. [00:36:13] Durga Toshniwal: So you can see this is a value activation function. [00:36:17] Durga Toshniwal: And activation functions are also required, just like any other neural network in RNNs also. [00:36:25] Durga Toshniwal: Additionally, there's something called a feedback loop. So, this feedback loop actually did not exist in the [00:36:33] Durga Toshniwal: previously studied artificial neural networks. So now, what we have is an input, what we have is an output, what we have are weights, biases, and all. Additionally, we have an input that is going from the output of the activation function [00:36:51] Durga Toshniwal: 2… The input side, And this is preceding the addition of the bias, so… [00:37:00] Durga Toshniwal: Here we have the input. The input is proceeding like this. We have the feedback from the activation, which is getting multiplied by weight W2, then it's getting added to the input, which has got multiplied by weight 1, and the two get added together, and then this passes to the, bias. [00:37:20] Durga Toshniwal: And then the summation goes to the… [00:37:22] Durga Toshniwal: activation function. So this is the big difference which is there in recurrent neural networks, and how this will help us, we will just, see. [00:37:34] Durga Toshniwal: So, for now, [00:37:44] Durga Toshniwal: Sorry, I got muted. [00:37:49] Durga Toshniwal: So, so although this neural network is actually, looking as though it's only taking a single input, which is this, however, we'll shortly see [00:38:02] Durga Toshniwal: that the neural network is capable… this neural network, which is the recurrent neural network, is actually capable of taking as input a sequence, which I'll just show you shortly. [00:38:13] Durga Toshniwal: So, [00:38:15] Durga Toshniwal: The feedback loop is the one that is actually, though it is looking as though there's only one input, however, the feedback loop is going to help us [00:38:25] Durga Toshniwal: To use this kind of a neural network on sequential data. [00:38:30] Durga Toshniwal: And, make predictions using the sequential data. [00:38:35] Durga Toshniwal: So, let's see how this is. Let's say if our input is 0, [00:38:41] Durga Toshniwal: And there is a weight that is getting multiplied, so this output will be 0, it will be summed with 0, and this is the value for a zero, the output [00:38:51] Durga Toshniwal: Will be a 0 multiplied by weight, and so this output will be 0. [00:38:59] Durga Toshniwal: So now, we can actually now run, [00:39:04] Durga Toshniwal: the other part of the stock data, so maybe we took 0 as the input. Maybe now we can take some other value, say 0.75, [00:39:13] Durga Toshniwal: And feed it as the input, like that, and keep on doing it. [00:39:18] Durga Toshniwal: So… Let's now consider specific cases and run it on the recurrent neural network and see. [00:39:25] Durga Toshniwal: Let's assume that any such data, say this one, is actually… can be represented in form of segments. [00:39:33] Durga Toshniwal: Right? And each of these segments could be characterized by one of these [00:39:38] Durga Toshniwal: Say, the value for, yesterday, so there are 3 values in it, yesterday, today, tomorrow. We have to predict tomorrow based on yesterday and today. [00:39:48] Durga Toshniwal: So that, value for yesterday could be low, and today also could be low, and based on that, tomorrow could also be low. So this we'll be predicting. [00:39:58] Durga Toshniwal: Similarly, yesterday could be low, today is a little higher, then tomorrow could be still higher. [00:40:04] Durga Toshniwal: Then yesterday could be very high, today could be slightly low, and then tomorrow could be… [00:40:10] Durga Toshniwal: Very low or zero? [00:40:12] Durga Toshniwal: Or yesterday, today, both could be high, and then finally, today also could be high. So these are the different trends that could be there in the stock data, and we are going to discuss [00:40:26] Durga Toshniwal: and study using these in the recurrent neural network. Because any stock price data, say, for example, the one that we had drawn, something like this, will actually be made up of trends like this only. Say, for example, this part will match this. [00:40:44] Durga Toshniwal: And then this part, Well, this segment will match this. [00:40:49] Durga Toshniwal: then we might have this again matching this. Then we might have this, which may be matching this, like this. So, any data could actually be broken up into such kind of segments, so we'll study how these segments could be used as an input to the RNN. [00:41:09] Durga Toshniwal: And, we'll run the RNN, so that it feeds upon yesterday's and today's data, and it will predict the tomorrow's price. [00:41:22] Durga Toshniwal: So, when we are calling these, prices as low, medium, and high, we'll have to substitute some values to it. So, we are putting low as 0 and high as 1. [00:41:34] Durga Toshniwal: Like that. [00:41:38] Durga Toshniwal: And we are going to now make these values run through the neural network, which you are seeing here. Once again, assume that the neural network has already been trained, and the weights and biases are finalized. As we discussed yesterday also. [00:41:54] Durga Toshniwal: When I was talking about activation, that, here, of course, we… when the… [00:42:02] Durga Toshniwal: the RNA is trained over the input data, the weights and biases are not finalized, but once the training is complete, then weights and biases are finalized. Let's assume [00:42:14] Durga Toshniwal: These weights, like W1, B1, W3, B2, and W2, are all finalized, and the model has been trained on the input data, which is a sequence. And now it is ready to make a prediction, and we are going to use this [00:42:28] Durga Toshniwal: Kind of data to make the prediction, where the input is for yesterday 0, today is zero, and what we want to predict is the value for tomorrow. [00:42:41] Durga Toshniwal: So now, and of course, there's a feedback loop that will also be utilized in this process. [00:42:49] Durga Toshniwal: So, initially, we'll enter the yesterday's and today's values. [00:42:56] Durga Toshniwal: So let's first input this yesterday's value here. [00:43:00] Durga Toshniwal: So, this will be serving as the input. [00:43:04] Durga Toshniwal: So, yesterday's value is what? It is zero. So, yesterday's value is input, sudden. [00:43:12] Durga Toshniwal: So, yesterday's value is 0, and this will be inputted. [00:43:16] Durga Toshniwal: To the neural network. [00:43:18] Durga Toshniwal: Alright, so we'll have a zero here. [00:43:23] Durga Toshniwal: No. [00:43:24] Durga Toshniwal: We are going to propagate this 0 through the neural network. So now, the zero will go here. [00:43:30] Durga Toshniwal: So, we have that equation as we had yesterday, so X could be so here, [00:43:42] Durga Toshniwal: This M will be our weight, and C will be our bias. So, we'll substitute the value so we can find out the X value for the activation function. [00:43:53] Durga Toshniwal: So, X will be equal to… The yesterday's value? [00:43:58] Durga Toshniwal: I'm just writing like this, times W1 plus bias. [00:44:03] Durga Toshniwal: Yesterday's value, as you see here, is 0, so it's 0 times weight 1 is 1.8. [00:44:09] Durga Toshniwal: Less bias is 0. [00:44:12] Durga Toshniwal: So, that comes out to be 0. So, X is equal to input, and that is 0. So, X is the input to the activation function. [00:44:21] Durga Toshniwal: So the input to the activation function is 0, which is here. [00:44:26] Durga Toshniwal: Now, what is the activation function? The activation function here is a value. [00:44:32] Durga Toshniwal: So, this function is a relu. [00:44:34] Durga Toshniwal: And the function of relu is… FX is equal to… 0, X. [00:44:43] Durga Toshniwal: Now, what is our X? X is equal to 0, so F of 0 will be equal to max of 0, which will be 0. [00:44:52] Durga Toshniwal: So the value of F0 will be 0, which is Y. So our output is going to be Y. So this is a point, input is 0, X is 0, Y is also 0. [00:45:03] Durga Toshniwal: Okay, so now this, this is the output. [00:45:07] Durga Toshniwal: That we are going to get. [00:45:10] Durga Toshniwal: And, [00:45:12] Durga Toshniwal: So this is a value of Y1, right? So, what we have is, X is equal to 0, and F of X is equal to max of 0, which is 0, and this is a value of Y1. [00:45:29] Durga Toshniwal: We are saying Y1 because it is a sequence, and we'll be having other outputs also, so we'll call this one Y1. [00:45:36] Durga Toshniwal: And others we might call Y2, and so on and so forth. So now we are having X is equal to 0, and Y1 is equal to 0, and this will be propagated again through these widths. So now what will we have? We'll have W3 times Y1, [00:45:54] Durga Toshniwal: Because Y1 is the one that is coming out of the activation function. [00:45:59] Durga Toshniwal: And, then we'll add the bias B3. [00:46:03] Durga Toshniwal: So, W3 is 1.1, so we are having 0 times 1.1. [00:46:09] Durga Toshniwal: Plus, bias is 0. [00:46:12] Durga Toshniwal: So, what we are having, again, is 0. [00:46:16] Durga Toshniwal: So, our output is going to be 0 here only. [00:46:20] Durga Toshniwal: Okay. [00:46:20] Durga Toshniwal: So, I've already showed you these calculations, it is 0. [00:46:25] Durga Toshniwal: plus bias, which is going to be 0. So, the predicted value [00:46:30] Durga Toshniwal: When we take as input, this is the value we took as input, this is the yesterday's value that went inside, and when we took that, we predicted today's value, which came out to be 0, right? This is equal to 0. [00:46:43] Durga Toshniwal: So today's value has come out to be 0. [00:46:46] Durga Toshniwal: So, so what we are having here is what we have predicted, which is today's value, which is zero, and the actual value is also 0. So our prediction for today is correct. But remember, we are not actually interested in predicting the value for today. We are interested in… [00:47:02] Durga Toshniwal: Using yesterday and today, and making a prediction for tomorrow. [00:47:11] Durga Toshniwal: So then, So, we are going to use both of these values. [00:47:18] Durga Toshniwal: Yesterday and today. [00:47:20] Durga Toshniwal: to make the prediction for tomorrow, and so what we'll go… what we will be doing is, whatever we predicted for today, right now, we'll just ignore this, okay? This prediction will happen. [00:47:31] Durga Toshniwal: But we'll just ignore, because we are only interested in this value. We are not interested in yesterday and today. [00:47:38] Durga Toshniwal: And we are fed, right now, only yesterday's value. [00:47:42] Durga Toshniwal: Now, what are we going to do? [00:47:44] Durga Toshniwal: So, we are now going to utilize this feedback that we hadn't used so far. [00:47:49] Durga Toshniwal: So, this feedback will take the output, whatever the value of output is. Remember, X was the input, which is equal to, 0. [00:48:01] Durga Toshniwal: and Y1… Was the output of the value, which was equal to 0. [00:48:06] Durga Toshniwal: So this Y1 value, which is here, for X equal to 0, it is equal to 0. [00:48:12] Durga Toshniwal: So… [00:48:14] Durga Toshniwal: So, now we are going to use this feedback loop, and this value will actually be multiplied by minus 0.5. So, Y1 times [00:48:23] Durga Toshniwal: Weight 2. [00:48:25] Durga Toshniwal: Will be fed, and it will be summed with this input. [00:48:29] Durga Toshniwal: Which will be our X value. [00:48:31] Durga Toshniwal: Which is yesterday's value, so I'm putting a… X yesterday into 1.8. [00:48:40] Durga Toshniwal: So… This is what is coming from here. [00:48:43] Durga Toshniwal: Yesterday value? [00:48:45] Durga Toshniwal: times 1.8, and this is the output Y1. [00:48:50] Durga Toshniwal: So, this is Y1 times. [00:48:53] Durga Toshniwal: minus 0.5. [00:48:55] Durga Toshniwal: Okay, so this is what we are having. [00:48:58] Durga Toshniwal: And Y1 is 0. [00:49:00] Durga Toshniwal: And this is multiplied by a weight of minus 0.5. [00:49:04] Durga Toshniwal: So, it will come out to be 0, and yesterday's value is also 0, and it's multiplied by 1.8. So, what we'll have here? [00:49:13] Durga Toshniwal: Finally. [00:49:16] Durga Toshniwal: So, these are just the steps that I already showed to you. We'll have I1W2, which we just calculated. [00:49:23] Durga Toshniwal: Right. And, this is, this is the input. [00:49:29] Durga Toshniwal: And it is, going to be multiplied by W1, and then… [00:49:35] Durga Toshniwal: This Y1W2 will be added along with input times B1. [00:49:40] Durga Toshniwal: So now, [00:49:42] Durga Toshniwal: So far, what we had was only yesterday's value, and with the help of yesterday's value, what we had got is this Y1W2. [00:49:50] Durga Toshniwal: Not… now, what we are going to do is… [00:49:53] Durga Toshniwal: we are actually going to input today's value here. So today's value is 0, we have again inputted it here. So now we'll have today's value multiplied by the weight W1, [00:50:05] Durga Toshniwal: Then, what we are having from here, as we are concentrating on the feedback part, we had Y1 times W2. Y1 was getting generated because of yesterday's value. [00:50:19] Durga Toshniwal: Right? We did it already. [00:50:22] Durga Toshniwal: And we multiply it by the weight W2. So now, what will we have? We have today's value, which is 0. [00:50:29] Durga Toshniwal: W1, which is 1.8. So, this is what we are obtaining here. This is coming here, this part. [00:50:36] Durga Toshniwal: And then Y1 times W2. Y1 is equal to 0. Weight W2 is 0.5 minus 0.5. [00:50:44] Durga Toshniwal: So this part is the one that is coming from this feedback loop and going here, into the summation. So this is going into the summation, and the two are getting added. [00:50:55] Durga Toshniwal: So, when they are getting added, what we are getting here, finally. [00:51:00] Durga Toshniwal: is 0 only, because we have 0 times 1.8, plus 0 times minus 0.5, so it will come out to be 0. [00:51:09] Durga Toshniwal: So, this is just the same thing shown here, so I'll just skip it. [00:51:14] Durga Toshniwal: So, we are going to have a zero here, which will be inputted. So, this will be, zero. This part will be zero, as you can see here. [00:51:23] Durga Toshniwal: This is coming out to be 0, plus this is coming out to be 0. So the input to this bias is going to be 0. [00:51:31] Durga Toshniwal: And when we add the bias to it, then also it will come out to be… so if I add the bias to it, then this bias is zero, so… [00:51:40] Durga Toshniwal: Overall, the input to this activation will be zero only, okay? When we have inputted. [00:51:46] Durga Toshniwal: Yesterday's value, and today's value. [00:51:50] Durga Toshniwal: Okay. Now, the problem with, this looks okay, but the only issue is that [00:51:56] Durga Toshniwal: Here we have to remember that what we are getting as an output to the activation function was a component [00:52:04] Durga Toshniwal: Of, what was fed, from the yesterday's value, right? [00:52:11] Durga Toshniwal: And what we are having right now. [00:52:14] Durga Toshniwal: It's something that we are getting from the, today's value. [00:52:18] Durga Toshniwal: So, this is today's value. [00:52:21] Durga Toshniwal: So, we'll have to actually remember all the time where is the today's value, where is the yesterday's value. Now, if we have a sequence of 10 points, then we have to remember where is yesterday, where is today, and all that. [00:52:34] Durga Toshniwal: And therefore, it becomes quite clumsy to remember which particular value is at the input at any particular point in time. [00:52:44] Durga Toshniwal: So, what we'll do, instead of doing all this, we are going to unroll, the feedback loop. [00:52:51] Durga Toshniwal: Right? And when we unroll the feedback loop, then what it involves is, it involves actually making a copy. [00:53:03] Durga Toshniwal: Oh. [00:53:06] Durga Toshniwal: So, when we are unrolling, we are actually making a copy. [00:53:10] Durga Toshniwal: Of, the neural network. [00:53:13] Durga Toshniwal: for each input value, right? So we are actually, in a way, when we are talking about unrolling, we are kind of unfolding. [00:53:23] Durga Toshniwal: And [00:53:26] Durga Toshniwal: So, what are we going to do? For each and every feedback, we are going to make a copy of this entire neural network, one for one input. [00:53:35] Durga Toshniwal: Okay, so now, what it will look like, because we have two inputs, one is yesterday's, one is today's, so, what we have done, we have taken, this input, we have taken this recurrent neural networks part. [00:53:52] Durga Toshniwal: And whatever we have. [00:53:55] Durga Toshniwal: What we do is that we just replicate it, because we have two inputs, so we are just replicating it as two [00:54:02] Durga Toshniwal: Two components, right? [00:54:05] Durga Toshniwal: And here, the way it spices and all, just we are replicating as such. So, we are just making a copy. [00:54:12] Durga Toshniwal: But, the difference is that [00:54:15] Durga Toshniwal: Previously, we were having a feedback that was feeding from the output of the activation function to the input of this bias. [00:54:23] Durga Toshniwal: So, if we are making a copy, we don't really have to do this any longer. [00:54:27] Durga Toshniwal: And what we will do is, because we have made a copy of this, we'll feed the output from here as an input to this. [00:54:35] Durga Toshniwal: So earlier, what was happening, you can see here, this feedback was getting fed into the summation, which is going as an input to the bias. So, what have we done? We have just connected… disconnected from here, and connected it here. [00:54:52] Durga Toshniwal: So now, it's going as an input to the summation here. And now we have got this, recurrent neural network unfolded or unrolled. [00:55:01] Durga Toshniwal: One for each input. So this was for the first input, which is the yesterday's input, and then this is for the today's input. And together, we'll be predicting, using this whole structure, we'll be using tomorrow's, the prediction for tomorrow's value. [00:55:17] Durga Toshniwal: And whatever will be predicted here. [00:55:21] Durga Toshniwal: And this… Is actually today, so when we took yesterday's input, what we predict is today's value. [00:55:28] Durga Toshniwal: We actually are not interested in today's value, so we are just ignoring this. [00:55:33] Durga Toshniwal: And then this is fed as the input, and this is to this activation function. And here is the today's input along… so the impact of yesterday's input is coming here. This is today's input, and what we get is the output, which is the, tomorrow's output. [00:55:54] Durga Toshniwal: So, this is how, the yesterday's value and today's value both are making a… [00:56:02] Durga Toshniwal: an impact. And what we are having is a new network that is actually having two inputs, instead of one with a feedback. Now, there are two inputs with one connected to the other. [00:56:16] Durga Toshniwal: That is the output of the activation function connected from the first unit to the second bar. [00:56:24] Durga Toshniwal: So now, what we are having are two outputs, as I already explained, one for today and one for tomorrow. [00:56:30] Durga Toshniwal: We are not interested in today's input, I mean, today's value as an output, so we'll ignore that, and we are only interested in tomorrow's value as the output. [00:56:41] Durga Toshniwal: Then there'll be two inputs, one for the yesterday. [00:56:45] Durga Toshniwal: And one for the today, that will be fed. [00:56:48] Durga Toshniwal: And, [00:56:50] Durga Toshniwal: Simply, when we are having this yesterday's value inputted, all the maths will be done, and it will be propagated to the output, as well as [00:56:59] Durga Toshniwal: So this is getting propagated to the output, and what we are getting is the prediction for today's value. We'll just ignore that value. [00:57:08] Durga Toshniwal: And, we'll not look at it, because we already have that value, right? We are having it here. And what we really want to predict is the tomorrow's value, so we'll consider tomorrow's [00:57:18] Durga Toshniwal: We'll just consider, Whatever we are getting as an output of the value. [00:57:24] Durga Toshniwal: We'll feed it, and then we'll get the summation. [00:57:28] Durga Toshniwal: And then we'll get the final value. [00:57:31] Durga Toshniwal: So, the impact of yesterday's value is actually coming here. [00:57:36] Durga Toshniwal: Isn't it? Because this was the input. [00:57:40] Durga Toshniwal: from yesterday. This is getting multiplied by weights, all that, getting fed to the value, then the value's output is dependent on the yesterday's input. So, this yesterday's input [00:57:52] Durga Toshniwal: It's actually coming in this form, and becoming a part [00:57:56] Durga Toshniwal: of the input which is going along with today's input. So… [00:58:00] Durga Toshniwal: Here, the sequence plays a role, because we are considering yesterday's value also, along with today's value, to compute the final output. [00:58:10] Durga Toshniwal: Both feet. [00:58:12] Durga Toshniwal: So now we'll just go on inputting the values and see. So, as calculated earlier, yesterday's value inputted here. [00:58:20] Durga Toshniwal: And then, multiplied by the weight. [00:58:23] Durga Toshniwal: The bias is added, and what we get is a zero. [00:58:28] Durga Toshniwal: Here. [00:58:32] Durga Toshniwal: Just a second. [00:58:37] Durga Toshniwal: So, we already had a zero here. [00:58:40] Durga Toshniwal: Oh. [00:58:41] Durga Toshniwal: My display's actually giving me a lot of talk to. [00:58:48] Durga Toshniwal: And so, this is… Yeah, so this is coming in. [00:58:55] Durga Toshniwal: As an input here. [00:59:00] Durga Toshniwal: Once again, [00:59:01] Durga Toshniwal: I hope I'm not facing the same issue. [00:59:07] Durga Toshniwal: Yeah, here is my cursor. [00:59:12] Durga Toshniwal: Yeah, here it is. [00:59:14] Durga Toshniwal: So then, we can see this, it's just repetitive. [00:59:19] Durga Toshniwal: Then we feed it as the input. This input goes to the input of the summation. Actually, my… Yup. [00:59:27] Durga Toshniwal: So this is going here as the summation. [00:59:31] Durga Toshniwal: And finally, the output is cutting. [00:59:34] Durga Toshniwal: Predicted. [00:59:36] Durga Toshniwal: All these steps I've already showed you. [00:59:39] Durga Toshniwal: So, the summation comes out to be 0 from here. It is 0 times 0.5, so this is 0. This is 0 times… this is today's input, it is 0 only. 0 times 1.8, this is 0. Summed with 0, so input will be 0. And what will be the output? Output is… [00:59:57] Durga Toshniwal: FX is equal to max of… [00:59:59] Durga Toshniwal: 0 comma X. So, it is, in this case, max of 0, which is 0. So, output is again 0. So, what we have predicted, finally, is the value. [01:00:13] Durga Toshniwal: So this is… this was Y1, which was predicted on yesterday's value, and then based on today's prediction, we get Y2. [01:00:20] Durga Toshniwal: And then, this is multiplied by weight and bias. [01:00:24] Durga Toshniwal: So, it… so Y2 is here, multiplied by W3, added 2 bias. So, this was 0. [01:00:31] Durga Toshniwal: Predicted value 0. [01:00:33] Durga Toshniwal: Weight was 1.1, bias was this, so what we are getting 0. So this is the prediction for tomorrow. And you can see here, this predicted value for tomorrow is actually consistent with the original observation for tomorrow, that was 0. So we have predicted the value for zero, which is tomorrow, and it is correct. [01:00:53] Durga Toshniwal: So, the recurrent neural network, with the help of this feedback loop, has considered the input for yesterday's value, the input for today's value, and it has predicted the value for tomorrow, and it has predicted it correctly. [01:01:11] Durga Toshniwal: Now, similarly, we could actually run, the other segments also on this neural network. [01:01:19] Durga Toshniwal: These are the different scenarios that we… should consider. [01:01:24] Durga Toshniwal: And when we run it, what we will get are the correct values predicted for tomorrow, right? [01:01:30] Durga Toshniwal: So, for example, if we feed instead of this, which has 000, if we feed this as 111. So, if we feed yesterday's value as 1, then what will we get? This is 1 times 1.8, so it will be 1.8. [01:01:45] Durga Toshniwal: Plus bias is 0. [01:01:47] Durga Toshniwal: then the value of 1.8 will be fed here, and this is the value. So… -Oh. [01:01:55] Durga Toshniwal: FX will be equal to max of… 0, comma, X. [01:02:01] Durga Toshniwal: And this will be… [01:02:02] Durga Toshniwal: 1.8 plus 0, so it will be 1.8. So the output, which is Y, is 1.8. [01:02:09] Durga Toshniwal: This 1.8 will be fed into, so we'll just ignore this. [01:02:15] Durga Toshniwal: So, we'll go with this, so 1.8. [01:02:19] Durga Toshniwal: Into minus 0.5. [01:02:21] Durga Toshniwal: Let me see if I have space here, so I don't have spacing. [01:02:25] Durga Toshniwal: So then, it will be how much? It will be? 0.9, minus 0.9. So, this value will be minus 0.9. [01:02:35] Durga Toshniwal: Just a second. [01:02:40] Durga Toshniwal: So… [01:02:46] Durga Toshniwal: So this value is minus 0.9 here. [01:03:01] Durga Toshniwal: So, this value is minus 0.9. [01:03:04] Durga Toshniwal: And 1 was the input here, and it is multiplied with the weight, 1.8. [01:03:10] Durga Toshniwal: So this becomes 0.9. [01:03:12] Durga Toshniwal: Sorry, 1.8 here, so… This is 1.8. [01:03:18] Durga Toshniwal: 1 and 2? [01:03:19] Durga Toshniwal: 1.8. [01:03:22] Durga Toshniwal: So this becomes 1.8, then 1.8 along with minus 0.9, and a zero bias means the input will be 1.8. [01:03:32] Durga Toshniwal: Minus 0.9. [01:03:35] Durga Toshniwal: So I'm facing a lot of issue with my… Writing pad. [01:03:40] Durga Toshniwal: It's giving me problem every second. [01:03:56] Durga Toshniwal: So, we have as input, 1.8 here. [01:04:01] Durga Toshniwal: Then we have minus 0.29. [01:04:04] Durga Toshniwal: Then, this is the bias, which is 0. [01:04:12] Durga Toshniwal: And what we are getting, the input here will be 0.9. So this… what input is going here is 0.9. And what will be the output Y2? [01:04:23] Durga Toshniwal: So Y2 will be equal to function of X, which is 0.9. [01:04:28] Durga Toshniwal: And because of value, what will it be? Max of… [01:04:35] Durga Toshniwal: I'm just… Facing a lot of issues, sorry for that. I don't know what's going on. [01:04:52] Durga Toshniwal: I'm just sharing again. [01:05:13] Durga Toshniwal: So, this value, when we talk about relu, then Y2 will be equal to function of 0.9, which is max of [01:05:24] Durga Toshniwal: 0, 0.9, so this will be 0.9. [01:05:29] Durga Toshniwal: So now, this 0.9 will be the output, Y2, [01:05:33] Durga Toshniwal: So this is going to be 0.9. And as for this equation now, 0.9 Will be multiplied by 1.1. [01:05:43] Durga Toshniwal: So, W3 times Y2. [01:05:47] Durga Toshniwal: This is W3 times Y2. [01:05:50] Durga Toshniwal: plus V2. [01:05:52] Durga Toshniwal: V2 is 0. [01:05:54] Durga Toshniwal: W3 is how much? 1.1? [01:05:57] Durga Toshniwal: So, I'm just writing 1.1. Y2 is 0.9. [01:06:02] Durga Toshniwal: So, what we'll get here? [01:06:04] Durga Toshniwal: Will be approximately 1. [01:06:10] Durga Toshniwal: So, the value that we'll get here is 1, which will be the predicted value once again. [01:06:16] Durga Toshniwal: Let me go. [01:06:18] Durga Toshniwal: Really facing a lot of issue. [01:06:20] Durga Toshniwal: So, actually, the predicted value will be 1. It's not going to be 0, as you are seeing here. [01:06:26] Durga Toshniwal: So this value will be 1. [01:06:30] Durga Toshniwal: Just a second. [01:06:36] Durga Toshniwal: Actually, what is happening is my display, my write… digital writing pad is getting discurrent. So… [01:06:45] Durga Toshniwal: So, this will be 1. [01:06:48] Durga Toshniwal: The predicted value will be 1.1 times 0.9 plus 0, so this will be approximately 1. [01:06:55] Durga Toshniwal: And the output will be 1. And this predicted value for output should be 1, which is the correct value. [01:07:04] Durga Toshniwal: Okay, any questions so far? [01:07:06] Durga Toshniwal: on this… Anyone, any questions on what we discussed so far on RNNs? [01:07:17] vinit shah: So I just wanted to understand, keeping the time series data, so this just keeps looping, is it? Like, if I want to look at the data for the last one year, so is it, like, iterative that last one year's data [01:07:28] vinit shah: one year, then the next day data keeps adding, adding, adding for me to get tomorrow's data. [01:07:35] vinit shah: Like, here we are just looking at two inputs right here. [01:07:37] Durga Toshniwal: Yes, yes. [01:07:38] vinit shah: What if I want to do the same thing for one year? [01:07:41] vinit shah: So, will the last one-year data then become input to the one-year minus 364 days, and then that to the next one, and that to the next one? [01:07:48] Durga Toshniwal: Exactly. So, let's say that if you have one year of data, so, and let's say you want to use 364 values to make a prediction for the 365th one. If that much length you are having, then you'll need to have those many number of units. [01:08:05] Durga Toshniwal: unrolled. [01:08:06] Durga Toshniwal: So, you'll have 364 units to predict the 365th value. So, each of these will keep on going as the input to each unrolled unit. So, for example, here, there are two preceding days, so we are using two days to predict the third one. [01:08:22] Durga Toshniwal: But if we are having 364 days to predict the 365th one, we'll need to use so many units. [01:08:31] Durga Toshniwal: So, what you're saying is correct. [01:08:34] vinit shah: Okay. [01:08:36] Aditya Banda: Continuing on that, example, ma'am. [01:08:39] Durga Toshniwal: If we have 364 units. [01:08:43] Aditya Banda: after I get an output from value function, why am I propagating it further? I… that's all I need, right? I need Y1, Y2, and Y3 to continuously feed into the next unit. [01:08:54] Durga Toshniwal: No, I didn't get your question. Can you come back again, sorry? [01:08:57] Aditya Banda: So, in this example that you're showing on your slide. [01:09:00] Durga Toshniwal: After I get Y1. [01:09:02] Aditya Banda: There's no point in me propagating further down the network. So I can use that Y1 in the next loop, get Y2, use Y2 in the next loop, and so on. [01:09:16] Durga Toshniwal: So you mean to say that if I have Y1 here, then I feed it here? [01:09:20] Aditya Banda: Right, right. [01:09:22] Durga Toshniwal: So I know today's value. [01:09:24] Aditya Banda: So, why am I propagating it fully, the… entirely through the neural network in the first loop? [01:09:33] Durga Toshniwal: The second part of the first loop doesn't make sense, right? After I get Y1. [01:09:37] Durga Toshniwal: Yeah. [01:09:38] Durga Toshniwal: So, the thing is that… Here, when you are actually having Y1, [01:09:44] Durga Toshniwal: So, what are you doing when you are, let's say that when you are having two values. [01:09:50] Durga Toshniwal: I mean, when you are having two such units, then you are having Y1 here, okay? And then you are inputting this Y1 here. [01:09:59] Durga Toshniwal: Right? So this is, you know, today's value that you have predicted. [01:10:05] Durga Toshniwal: Right? [01:10:06] Durga Toshniwal: But actually, we are not interested in the predicted value. We are interested in having the actual value being fed into it. [01:10:16] Durga Toshniwal: Yeah. Because the prediction might have some error also. In this case, it is not having. [01:10:22] Durga Toshniwal: But assume that there is some error, so whatever is predicted, if that is fed into it, then what will happen? That prediction's error will propagate. So we don't want the predicted value to be fed, we want the fresh value to be fed. [01:10:36] Durga Toshniwal: Then how will we do it? [01:10:38] Durga Toshniwal: When we are feeding the fresh value inside here, so you can think of a sequence. Let's say I've got a sequence of 3 numbers. This is a sequence of 2 numbers, this one and this one. So I'm feeding the sequence like this. This is the input one. [01:10:52] Durga Toshniwal: This is the input 2. If I had another input, input 3, like that. [01:10:57] Durga Toshniwal: Right? So I'm independently feeding the inputs. If I'm independently feeding the inputs, then how will I get the impact of the sequence? [01:11:06] Durga Toshniwal: That is the preceding value in the sequence, how will I get an impact of it on the current value? To get the impact, what I'll have to do is, I'll have to consider a portion of it [01:11:19] Durga Toshniwal: Right? So a portion of this input will actually be propagated here. [01:11:25] Durga Toshniwal: Along with the inputs, the fresh inputs impact. [01:11:29] Durga Toshniwal: And both will get added together as an input to the activation function for the second unit. Now, the output will actually depend not just on the current input, it will also depend on the previous input. So, this will be, yeah. Now, I have another input here. [01:11:46] Durga Toshniwal: So, I want this input to also consider the impact of the previous two inputs. [01:11:53] Durga Toshniwal: So, the previous two inputs, impact is considered jointly in this output, so which I'll feed it here. So, like this, what I'll have is the last [01:12:03] Durga Toshniwal: You know, whatever the number of last few inputs are, that is coming as a single number and getting fed into this unit, current unit. [01:12:13] Durga Toshniwal: This is how it is making an impact, is it okay? [01:12:17] Aditya Banda: Yeah. [01:12:18] Durga Toshniwal: Okay. [01:12:19] Durga Toshniwal: Okay, Deepak, do you have a question? [01:12:22] Deepak Katara: Yambo. [01:12:23] Deepak Katara: So, one thing, so do we go through all the nodes, and suppose if I take a node of yesterday, get the output Y1 or Y1-? [01:12:34] Deepak Katara: So, do I need to make an adjustment in terms of weight, or bias, or activation function, so that I can feed accurate value to today, or basically in the feedback loop? [01:12:47] Deepak Katara: Because I think you have also mentioned, if Y1 is, has more delta compared to the actual value, then it means we are feeding the wrong values to the next node. [01:12:58] Deepak Katara: Right? So, how does the cycle go? Do we make the correction, and reduce the error, and then feed into the next loop, or… [01:13:08] Deepak Katara: We just pass it, pass it on, whatever we get. [01:13:11] Durga Toshniwal: So, remember that we discussed about backpropagation, right? So, that backpropagation is applicable to any neural network. [01:13:21] Durga Toshniwal: Here, at the start, I told you. [01:13:24] Durga Toshniwal: Just like any other neural network, a recurrent neural network will also start with random weights only. [01:13:31] Durga Toshniwal: And these random weights will get… keep on getting updated with every iteration, till the loss is within the tolerance we want it to be. [01:13:41] Durga Toshniwal: Once that happens, the weights and the biases will be finalized, and what you're seeing here are the finalized weights and biases. So, I'm not showing the backtropagation here. [01:13:53] Deepak Katara: Okay, but it has to be done. [01:13:55] Durga Toshniwal: before we start using the recurrent neural network or any other neural network. So that is the first step. That is learning the model, right? So learning the model means learning the parameters, which are the weights, biases, and other things. [01:14:10] Durga Toshniwal: the activation functions, and all those stuff that we already talked about. So, after the learning is complete, what you're seeing here are the final bits. [01:14:20] Durga Toshniwal: Okay. [01:14:21] Deepak Katara: And I think one of the questions that has been asked, that if we have more, nodes or past headal layer to be feed into, so, like, what is the threshold? Because in that case, it will become more and more complex, right? If we have more, hidden layer in the past, and that we need to feed into further, layers. [01:14:41] Deepak Katara: Is there any threshold to it? Like… [01:14:43] Durga Toshniwal: No. There is no threshold into how many recurrent neural network units could be unfolded. [01:14:50] Durga Toshniwal: No, there is no… none. But there are after-effects of it, which I'll discuss shortly, okay? So there is an impact of increasing the number of units, and I'll discuss it. [01:15:02] Durga Toshniwal: As we go along. So what, whatever you all are saying is actually correct, that as we have a long sequence, more and more unfolding will be there, and more the number of unfolding, the complex this recurrent neural network will become. [01:15:18] Durga Toshniwal: And and that will have some impact. And what that impact will be, I will be talking about shortly, but one simple impact that you all can think about it is in this example. Let's say I have a sequence of 364 days. [01:15:33] Durga Toshniwal: And I have to predict 365th day, right? So I start out, I don't have… okay, let me consider this space. Let's say I have, you know, I'm just talking about a recurrent neural network as a box, and I'm just writing it here. [01:15:49] Durga Toshniwal: As usual, this pen is creating me problem. [01:15:54] Durga Toshniwal: Just a second, it's just coming up. [01:16:00] Durga Toshniwal: You know, what's gone wrong with it today. [01:16:10] Durga Toshniwal: Okay. [01:16:11] Durga Toshniwal: So, we have one recurrent neural network unit here. [01:16:16] Durga Toshniwal: And… [01:16:21] Durga Toshniwal: Okay, so this is our input one. [01:16:25] Durga Toshniwal: Then we have another… I'm just showing these as boxes, I'm not showing the intricacies. Then I've got input 2, input 3, like that. I've got 365 inputs, and there's a feedback. So I'm just simply showing it like this, obviously. [01:16:42] Durga Toshniwal: It has all the details in it, and this is a final input, and this is the first input, and this is the second, like that. So, you can see here that if we have, [01:16:53] Durga Toshniwal: 364 days, and we are predicting the 365th value. [01:16:58] Durga Toshniwal: Then, the impact of the… [01:17:01] Durga Toshniwal: very initial values has got diluted to quite some extent because it's getting propagated over 364 units. [01:17:10] Durga Toshniwal: So, this is one after effect. [01:17:13] Durga Toshniwal: off. [01:17:14] Durga Toshniwal: Of using RNN. However, I will explain it more in details as we go along, okay? [01:17:21] Durga Toshniwal: So now let's have another, data where we are having three preceding numbers, and we can use it to make tomorrow's prediction, which is the fourth value out here. [01:17:33] Durga Toshniwal: Then, in such a case, we'll keep on unrolling the recurrent neural network more and more number of times, because previously we had two inputs, so we were using, you know, yesterday's input here, this… this one. [01:17:48] Durga Toshniwal: So, if we have a day before… day before yesterday also, then we'll need another unit. So now what we are having, this is the day before unit value, this is, then yesterday's value, this is today's value. [01:18:03] Durga Toshniwal: And we are going to calculate whatever we are having. Say, for example, this day before yesterday is having a value 1. [01:18:11] Durga Toshniwal: So it will be 1 into 1.8 plus 0. So what will propagate inside 1.8? [01:18:18] Durga Toshniwal: And the output will be a max of… [01:18:21] Durga Toshniwal: 0, 1.8, which will be 1.8. [01:18:25] Durga Toshniwal: Right. [01:18:26] Durga Toshniwal: And then this will be multiplied by 1.1, so we are not going to consider this part, we are going to consider the feedback part, which is multiplying 1.8 with minus 0.5. [01:18:42] Durga Toshniwal: So my… again, my… This thing is gone. [01:18:55] Durga Toshniwal: So now it is 1.8 again. [01:18:58] Durga Toshniwal: Times. [01:19:01] Durga Toshniwal: minus 0.5. [01:19:03] Durga Toshniwal: So it becomes minus 0.9. [01:19:06] Durga Toshniwal: And then yesterday's value is 0.5, so this is 0.5. [01:19:13] Durga Toshniwal: And then 0.5 of 1.8 is 0.9. [01:19:17] Durga Toshniwal: So this is 0.9 and minus 0.9, so it becomes 0 added with 0, so 0 is the input, and max of 0, so 0 will be the output. [01:19:27] Durga Toshniwal: Then 0 will be multiplied by minus 0.5, so it becomes 0 out here. [01:19:32] Durga Toshniwal: And then today's value is 0.5, so this is 0.5. [01:19:36] Durga Toshniwal: And this multiplied, By 1.8 becomes 0.9. [01:19:42] Durga Toshniwal: And these three summed together becomes 0.9, and the output will be max of 0, 0.9. [01:19:52] Durga Toshniwal: Which will be 0.9. [01:19:54] Durga Toshniwal: And then this 0.9 will be multiplied by 1.1, so it will approximately become 1. This will be added, and this becomes 1. So our final output is 1, and the expected output is also 1. [01:20:06] Durga Toshniwal: So, now, because we had 3 previous values, so we needed to unroll it 3 times for day before yesterday, yesterday, and today. [01:20:14] Durga Toshniwal: So this is how we keep on plugging more and more values, and the previous values make an impact by propagating like this, as you can see here. [01:20:24] Durga Toshniwal: Till the final value is reached. [01:20:27] Durga Toshniwal: So, these are just the calculations I just already showed to you. [01:20:31] Durga Toshniwal: And finally, the predicted value comes out to be 1, and this is the same as the expected value, which is also 1. [01:20:40] Durga Toshniwal: So, this is how recurrent neural works. Now, one thing important to notice, because we are unrolling the same unit, so the weights, biases, like W1, B1, W2, W3, B3, or B2 will all be shared. [01:20:56] Durga Toshniwal: in all the enrolled units, so this is going to be the same. They are not going to change. So, we are not, having… we are not increasing the number of parameters here, because more the number of weights and biases, more the number of parameters in the, neural network. Here, the parameters are shared, and so [01:21:15] Durga Toshniwal: Parameters are not increasing that way. [01:21:19] Durga Toshniwal: So, if we are rolling it 3 times, all of these will be used 3 times. [01:21:24] Durga Toshniwal: So, same weights will be used thrice, same biases will be used thrice, and so on and so forth. [01:21:30] Durga Toshniwal: So, all weights and all biases will be used as many times the unit is unrolled. And so, we are not increasing the number of parameters by the number of times we are unrolling the unit. They still remain the same. [01:21:44] Durga Toshniwal: Okay. [01:21:46] vinit shah: One question, then. [01:21:49] vinit shah: See, the advantage that I see with this kind of time series data is that I have a way to know what are all the historical data, right? [01:21:58] vinit shah: Now, shouldn't I then be able to manipulate the biases and weights so that I get the, I don't know, the closest weight or bias that kind of leads to all this data? [01:22:10] vinit shah: like, I don't know if I'm putting it in the right way, but say, for example, my weight 1 is my first input, right? Now, the second input, I already know what was the data. But if my weighted bias did not give me closer to this data, shouldn't I then change this to such a way that my next data is [01:22:28] vinit shah: kind of closer to this Y output, or whatever it is, and then Do it so that [01:22:34] vinit shah: after my 364 data that is already available to me, I get something that is [01:22:41] vinit shah: Closer to the values that are already available, and there is no changes or high diversion on those values? [01:22:50] Durga Toshniwal: So, what you're saying is correct, but remember that I told you right in the beginning, let's say I have a 365 unit, you know. [01:22:59] Durga Toshniwal: I have a 364 unit. [01:23:02] Durga Toshniwal: RNN out here, okay? There are 364 units in it. [01:23:07] Durga Toshniwal: Okay, so I don't train it unit, I mean. [01:23:12] Durga Toshniwal: I don't train it one by one. [01:23:15] Durga Toshniwal: whatever is the output, right, these outputs are coming out parallelly, and the final one is also coming, and there are feedbacks going like this. So this whole is trained together. [01:23:27] Durga Toshniwal: Okay, you don't train it that you get today's, whatever, you know, day before's, or maybe the first value, so you don't train this unit separately to get a correct output here, then you train it, because they don't work separately. [01:23:43] Durga Toshniwal: I'm just showing you unrolling unit by unit for the sake of clarity. However, this is done as a single thing. [01:23:51] Durga Toshniwal: I already know, yeah, I already know what sequence length I have. So those many units will be unrolled, and as I told you, for any neural network, the weights and biases are initially randomly fed. [01:24:04] Durga Toshniwal: then these random weights are actually stabilized, and they are iterated upon by the process of backpropagation. The weights are changed till the loss [01:24:15] Durga Toshniwal: What is a loss? The difference between the predicted value and the actual value, so the loss at the output is optimally minimized. Till that happens, we keep on changing and updating the weights and bias. [01:24:28] Durga Toshniwal: And when the loss is within the threshold, or within our tolerance, we stop, right? So, all this, whatever you are saying, is done before making use of this particular RNN unit. [01:24:42] Durga Toshniwal: So, we don't do that. We take a single unit, then find out the weights and bias, then take the… unroll it, take another, you know, part, then do it. We don't do it like that. All this is done together, it is trained together as a single. [01:24:57] Durga Toshniwal: recurrent neural network made up of so many units, then finally weights and biases and all are finalized, then it is used for prediction. Right now, I'm showing you it's used for prediction. [01:25:08] Durga Toshniwal: And we assume that the weight W1, B1, W3, B2, and W2 are all already learned. [01:25:18] Durga Toshniwal: Is it okay? Does that clarify? [01:25:21] Durga Toshniwal: Winit? [01:25:22] vinit shah: Maybe I'm still not able to wrap my head around it, so what you… what you told, I understand it. What you said is, initially, I've already, done the backcourt propagation, and I've finalized the weights and the biases. [01:25:31] Durga Toshniwal: And only after that, I'm introducing all this thing into the neural network so that, you know, it then starts giving the. [01:25:37] vinit shah: like output into the subsequent input, and we get the output. That part, I understand. So the only thing that I'm kind of trying to understand is. [01:25:45] vinit shah: how are we leveraging the information which I already know, right? Like, what I'm trying to say is, if I'm already creating a system with the weights and bias, based on my tolerance and my weight propagation and all that. [01:26:02] vinit shah: how is this then adding value? So, is it your same… is it, like. [01:26:08] vinit shah: it's only these addition to the outputs that are going to give me a better predicted value tomorrow, because if I've already back-propagated, found out the ideal weight and ideal bias, then why do I have to get into the… [01:26:23] Durga Toshniwal: Okay. [01:26:25] Durga Toshniwal: So, remember, what we have is the training data. [01:26:29] vinit shah: Then we have the test data. [01:26:31] Durga Toshniwal: And when we have the unseen data. [01:26:34] Durga Toshniwal: The weights and biases are found out using the training data, right? [01:26:40] Durga Toshniwal: And once they are found out, then the test data is used to see how the performance of the model is. [01:26:47] Durga Toshniwal: Once these two have been used, then I'm going… then we assume that the model is stabilized, all the weights, vices, everything that we had, found out by learning this training data, so… [01:27:00] Durga Toshniwal: See, in any machine learning, AI, deep learning model, learning has to be done. Learning is always done on the training data. [01:27:07] Durga Toshniwal: Once the training data is utilized to do the learning, these weights, biases, and all are found out. [01:27:14] Durga Toshniwal: Once they have been found out, the model is ready to make prediction. And then the unseen data is fed into it, and the prediction is done. What you are seeing is this part. [01:27:25] Durga Toshniwal: That you are feeding to it a sequence. [01:27:27] Durga Toshniwal: For which you want to predict the fourth value. [01:27:32] Durga Toshniwal: So… [01:27:33] Durga Toshniwal: I think what you were thinking is that probably you're missing, that these weights, biases, and all are found out using the training data. [01:27:42] Durga Toshniwal: Once the training data has been used, all these are found, stabilized, and once this is done, then the model is used for making a prediction of the unseen data. And this is what we are doing right now. [01:27:55] Durga Toshniwal: Say, for example, if you remember any machine learning model, and we did a lot of hands-on, say, for example, if we have a KNN model. [01:28:04] Durga Toshniwal: Then we first find out the value of K. This all we do with the help of the training data. Then, once we finalize this value of K, then we'll build the final model using this K, and then we'll input to it some new data and use it for making prediction. [01:28:20] Durga Toshniwal: Here also, we are going to use the recurrent neural network, and to use it, we'll have to find out the weights, biases, and all. We'll find it out using the training data. Once we have found it out and finalized the weights, bias, and all in such a way that the training loss is optimally low. [01:28:37] Durga Toshniwal: And so, the model is ready, then we use it to make a prediction. We are making predictions here. We are not training the model here. [01:28:47] Durga Toshniwal: Is it okay, Vinet? [01:28:48] vinit shah: Yes, yes, yes, it did clarify. Thanks, ma'am. I had not died down. [01:28:51] Durga Toshniwal: Okay. Any other questions, anyone? I think, Lokesh, you wanted to say that you had a doubt that if we are having zero weight spices, and all, then how are we doing the learning? [01:29:02] Durga Toshniwal: So, what I want to say here is, remember, the weights, biases, and all have been found out earlier during the training phase. So, the weights and all are learned based on the training data. They could be zero, they could be anything. In this example, some of the weights and biases are zero. [01:29:18] Durga Toshniwal: Okay? Is it okay, Lokesh? [01:29:21] Lokesh R: Yeah, man, I'm getting it. [01:29:23] Durga Toshniwal: Yeah. Oh, now let's proceed. [01:29:30] Durga Toshniwal: Oh, boy. [01:29:31] Durga Toshniwal: So now, as we have been mentioning, and I think some of you have pointed out, that if our sequence is pretty long, then so many unrolled, you know, units will be there, and the recurrent neural network will become very complex, and it will be difficult to train. So this problem… [01:29:49] Durga Toshniwal: Is there with recurrent neural networks, that the longer the sequence, the more complex they become. [01:29:56] Durga Toshniwal: And the problem that arises is that of a vanishing gradient or exploding gradient. So we'll… I'll just describe these problems, how they come in, when sequences are long, how do vanishing or exploding gradient come into picture. [01:30:11] Durga Toshniwal: Okay, so to make things more clear, first of all, we'll try to understand, how the exploding gradient or the vanishing gradient problems come. [01:30:23] Durga Toshniwal: And, for that, we are going to, you know, primarily focus on weight 2. So, there was weight 1. [01:30:30] Durga Toshniwal: And there was a weight 3 here. There was a bias 1, there was a bias 2, and there was a weight 2. Because weight 2 is involved in this feedback, I'm only going to consider weight 2 out here, right now, for explaining the exploding and vanishing gradient problem, okay? [01:30:47] Durga Toshniwal: Because this is the one that is contributing to propagation of the values in the future recurrent… I mean, the further succeeding, units of the recurrent neural network. So, now let's… [01:31:01] Durga Toshniwal: Talk about, the role of weight, too. [01:31:04] Durga Toshniwal: Now, I will be deriving how back propagation works, maybe in one next turn or so. However. [01:31:11] Durga Toshniwal: When we are doing backpropagation, I told you, what we try to do is we backpropagate the error. [01:31:18] Durga Toshniwal: Or the loss. What is the loss? The loss could be measured in any of the error functions, say, for sum of squared errors. So, you already know what is sum of squared error. [01:31:30] Durga Toshniwal: I have already discussed with you, so loss is getting measured with the help of this function, which is a sum of squared error. [01:31:37] Durga Toshniwal: And, when we try to propagate, so what we try to do is, first of all, we try to find out the impact of [01:31:46] Durga Toshniwal: the different weights and biases on the loss. So, there will be, what will be the impact of weight 1 on the loss. So, that will be the derivative of the loss with respect to weight 1. [01:32:01] Durga Toshniwal: But wait 1 is where? It's not here right now, right? So, weight 1 is, nearer to the input. So to find that out, we could actually work it out like this. [01:32:13] Durga Toshniwal: that… [01:32:14] Durga Toshniwal: This is a loss, so if you have this, any recurrent neural network, say, for example, this was the one we had input. [01:32:25] Durga Toshniwal: So again, it's becoming very laggy, so I'm not able to show you. So this is our input. [01:32:31] Durga Toshniwal: And then we have… [01:32:40] Durga Toshniwal: Sorry for that, actually. Today, I'm facing a lot of issue. [01:32:48] Durga Toshniwal: So this is the input, then this is the weight one, then… [01:32:54] Durga Toshniwal: We have our bias 1, this is the summation of bias 1. Then this is our activation function, it could be value or whatever it is. And here we have… [01:33:05] Durga Toshniwal: we have a weight 2, sorry, weight 3 and bias 2, and this feedback is going… [01:33:13] Durga Toshniwal: to the next unit via weight 2. So, actually, we are trying to find out the impact of weight 1. [01:33:21] Durga Toshniwal: On the loss. And to do this, we'll actually need to look at… so here will be our output. [01:33:28] Durga Toshniwal: So, since the output is actually closest to the predicted value, so we'll find out the gradient, or differentiate the loss with respect to the predicted value. But the predicted value actually depends on Y2 [01:33:41] Durga Toshniwal: So, we are differentiating the predicted value with respect to Y2, but Y2 is dependent on X2, right? If there are two units here, right? [01:33:50] Durga Toshniwal: I don't have enough space, but we had Y1 here, and we had Y2 here. [01:33:56] Durga Toshniwal: So, this Y2 is actually dependent on this input, which was X2. [01:34:01] Durga Toshniwal: And then this X2 will depend on weight 1, [01:34:05] Durga Toshniwal: Here is weight 1, like this. So, and we'll keep on going further down, so we'll keep on adding. So, what is happening here is that, [01:34:16] Durga Toshniwal: So, these are the gradients that we'll keep on finding out to back-propagate the error. This is called back-propagation of the error, because the error is impacted by the predicted output. This predicted output is impacted by the value of Y2, [01:34:32] Durga Toshniwal: Y2 is actually found out because of Y1, so we'll find out the gradient of Y2 with respect to X2. X2 is impacted by weight 1, so we'll do our differential of it with respect to weight 1. [01:34:44] Durga Toshniwal: And so and so forth. [01:34:46] Durga Toshniwal: So now, what will happen is that… and you remember that we had, I had already showed you that step size [01:34:55] Durga Toshniwal: In the gradient descent function is equal to the gradient times the learning rate. [01:35:03] Durga Toshniwal: This we already did. Now, this is the gradient that we are talking about. [01:35:08] Durga Toshniwal: And we are going to, you know, use these gradients in the gradient descent function, which will help us to optimize the widths, okay, and minimize the loss function. The loss function could be sum of squared error, or it could be anything. [01:35:25] Durga Toshniwal: So, like this, what is happening is that, when we are trying to minimize our, [01:35:32] Durga Toshniwal: what will happen? The step size, the amount of updation we'll do on the weights will be dependent on the step size, right? So if the step size is small, the small amount will be updated, small amount will be updated like this. Accordingly, the gradient will keep on changing. [01:35:50] Durga Toshniwal: Right. Now, let's talk about a vanishing or the exploding radiant problem. First, we'll talk about the exploding radiant problem with respect to weight W2. [01:36:01] Durga Toshniwal: So, since this weight, W2, is impacting the input that is going to the next, unrolled unit, so if, let's say, this weight is greater than 1, let's say it's equal to 2, then whatever the input here gets multiplied by 2, right? [01:36:16] Durga Toshniwal: So, if you are having an input 1, and this weight will actually have this input, so input to the activation is input 1. It will get multiplied by 2. [01:36:28] Durga Toshniwal: And this is the way to, and then it will be fed here. [01:36:32] Durga Toshniwal: Then what will happen is that [01:36:35] Durga Toshniwal: For the next unit, this will be again multiplied, this input will be again multiplied, it will be propagated further by multiplying by 2, so another 2 will come into picture, and so on and so forth. More and more number of 2s will keep on adding. [01:36:50] Durga Toshniwal: Because of these weights, W2, Getting chained in this whole cascaded [01:36:56] Durga Toshniwal: recurrent unfolded units. And because there are 1, 2, 3, 4 units, so it got multiplied 4 times. So if I had n number of units, it would be 2 to the power n, because weight is 2. So the input is actually multiplied by 2 to the power [01:37:12] Durga Toshniwal: 4, where 4 is the number of recurrent neural network units. [01:37:18] Durga Toshniwal: And 2 is the weight W2. [01:37:20] Durga Toshniwal: So, whatever the weight W2, that will be actually raised to the power number of times unrolling is done on the recurrent neural network unit. [01:37:29] Durga Toshniwal: And so, if, let's say, whatever is the input, it will get, in this case, because number of units are 4, so it will get multiplied by 16. [01:37:39] Durga Toshniwal: Now, if it is 364, then it will be 2 to the power 364, by which it will get multiplied. So, this input 1 is getting magnified more and more. [01:37:50] Durga Toshniwal: And as a result, the gradient of it, so this is the input, if there are 50 units, then it will be, like, 2 to the power 50, you can see here. And so the impact of this weight 2 [01:38:02] Durga Toshniwal: is becoming more and more. So… [01:38:05] Durga Toshniwal: So, this means that the input is getting multiplied by a very big number, depending on the number of times the unit is unrolled, right? And therefore, this input is becoming very large, exponentially large. And this large number is actually going to, because this [01:38:23] Durga Toshniwal: input is getting multiplied by such a big number, so it's actually going to impact the gradient also, which I showed to you earlier, in one way or the other. As a result, the gradient is also going to become very large. You can see here. [01:38:38] Durga Toshniwal: So, what is happening? Our input is getting multiplied by, let's say, 2 to the bar 50, which is a very big number. [01:38:44] Durga Toshniwal: And this input, or whatever, W2 to the power, whatever number of times this unit is, it is unrolled. [01:38:52] Durga Toshniwal: And this is going to impact what the output, and this output will be impacted by X1. [01:38:59] Durga Toshniwal: So, and then this X1 is actually integrated with respect to weight 1, and so this X1 is now no longer X1, it has become 52 to the power 50 times of X1. Such a big number. As a result, this whole gradient [01:39:15] Durga Toshniwal: has become very, very large. It has become… [01:39:18] Durga Toshniwal: How many times? It has become 2 to the power 50 times of what it should have been. [01:39:23] Durga Toshniwal: Right? It's become huge. So the gradient has become very huge, and this is called the exploding gradient problem. [01:39:30] Durga Toshniwal: Right, so, so therefore, now what will happen is that when the gradient will become very large. [01:39:37] Durga Toshniwal: Then what will happen? We can no longer take small steps to adjust the weights and biases. So, if we talk about this figure, what we want to do is we want to take small, small steps [01:39:52] Durga Toshniwal: right, you see the gradient, so that we can actually achieve this. Now what is happening? We may be at this point, and now the gradient might change in such a way that it becomes so large. [01:40:05] Durga Toshniwal: that we are actually going somewhere here. So, we are actually going away [01:40:10] Durga Toshniwal: We are actually going like this. [01:40:14] Durga Toshniwal: So, I already explained to you gradient descent earlier. We want to proceed like this so that we can approach this point. We are actually going more and more away from this point, because the gradient is becoming larger and larger. This is called the exploding gradient problem, and in such case. [01:40:31] Durga Toshniwal: the loss will never be minimized. We cannot minimize it, because the step size is becoming very large. What is the step size? [01:40:40] Durga Toshniwal: Step size is equal to gradient times. [01:40:43] Durga Toshniwal: learning rate. So, gradient is becoming very large, and so this step size, what you can see here is becoming very large. So, this was a step size [01:40:52] Durga Toshniwal: then… so this is the step size, so it will go on increasing, then this will be the step size, and we'll actually just going far and far. It will go on going far like this, because steps will be more and more. [01:41:04] Durga Toshniwal: But what should have been step size, so that we can reach this point? The step size should have been slightly small, so that… and small, small gradients would make us [01:41:15] Durga Toshniwal: Lead to this optimally low value. [01:41:18] Durga Toshniwal: So this is the problem. When we are having, you know, a large number of units in the recurrent neural network, and a weight 2 is large, then a huge number will [01:41:30] Durga Toshniwal: come into picture, and this will result in us making very large steps. You can see here, they are so large, we'll keep on, and then we'll overshoot the minimum without preaching it, and it will go on moving away and away. [01:41:44] Durga Toshniwal: And the function will never converge, and the loss will never be minimized. This is called the exploding gradient problem. This comes into picture when the weight is greater than 1. When weight is greater than 1. Now, let's say that we want to limit the weight to a value less than 1. [01:42:04] Durga Toshniwal: If we want to limit it to a value less than 1, then what we'll be having is a vanishing gradient problem. Earlier, we were having exploding gradient problem. [01:42:12] Durga Toshniwal: Now, what is the vanishing gradient problem? [01:42:16] Durga Toshniwal: Once again, we'll look at this setup, where we are having a large number of [01:42:20] Durga Toshniwal: unrolled recurrent unit, neural network units, RNA units, and now weight is less than 1, so it is 0.5. [01:42:28] Durga Toshniwal: Now, what will happen? Just like earlier, we were multiplying input into [01:42:34] Durga Toshniwal: The weight to raise to the power number of times it is unrolled. [01:42:38] Durga Toshniwal: Right? So here, whatever is the input, that will get multiplied by 0.5 times the number of times the unit is unrolled, here it is 5 times, so it will be 0.5 raised to the power 0.5, right? [01:42:51] Durga Toshniwal: So now, just imagine if we have a large number of units, say we had 50 units once again, so it will be 0.5 raised to the power 50. So 0.5 raised to the power 50 will be very, very, very small number. It will be so small that this will be approximately 0. [01:43:08] Durga Toshniwal: When this is approximately 0, then the input is multiplied by 0.5 to the power 50, or approximately 0. So the impact of this input 1 becomes 0. [01:43:21] Durga Toshniwal: And this is called a vanishing gradient problem. In this case, what will happen is, again, the impact of weight 2 will come on the gradient. [01:43:29] Durga Toshniwal: Right? This was the gradient here, the equation was written out here. [01:43:34] Durga Toshniwal: So now, when we consider this. [01:43:37] Durga Toshniwal: Earlier, this was getting multiplied by a huge number. Now, this weight one will get multiplied by approximately 0, and this gradient will become 0. [01:43:47] Durga Toshniwal: When this gradient becomes zero, then what will happen? It will take infinite number of steps to actually reach… go from here up to this point, right? Because the gradient is zero, and what is the step size? Step size is equal to gradient times [01:44:04] Durga Toshniwal: learning rate. So this gradient is 0, so step size is 0. [01:44:08] Durga Toshniwal: Then, how will we go proceed to move to this point, which is the global minimum of this function? So, this is the vanishing gradient problem. Both exploding gradient problem and the vanishing gradient problem [01:44:23] Durga Toshniwal: are important problems in recurrent neural networks because of the work they have… way they have been designed. And as the number of… as the length of the sequence increases, the number of units connected [01:44:35] Durga Toshniwal: or unrolled, I should not say connected. The number of times the recurrent neural network unit is unrolled becomes very large accordingly. [01:44:45] Durga Toshniwal: Based on the value of the weight 2, either we'll end up having exploding gradient problem, and if it is very small, less than 1, then we'll have the vanishing gradient problem. [01:44:56] Durga Toshniwal: Alright, so this is all about recurrent neural network. Any questions? [01:45:08] Durga Toshniwal: Any questions, anyone? [01:45:21] Durga Toshniwal: So, recurrent neural networks are definitely quite popular, but they do suffer from these limitations. One, that the input keeps fading out. [01:45:31] Durga Toshniwal: it will either fade out, or it will become too large. Any… the impact of any particular input will either fade out if the weight 2 is very small, or it will get amplified hugely if it is too large, and it will also result in exploding and vanishing gradient problem. [01:45:48] Durga Toshniwal: Respectively, for very large or very small ribs. [01:45:51] Durga Toshniwal: Still, in spite of all those facts, recurrent neural networks still continue to be used, and then there are variants of it. [01:45:59] Durga Toshniwal: That improve upon this. So, I'll discuss in future terms. I'll also discuss the derivation of backpropagation, how backpropagation actually updates the weights. [01:46:10] Durga Toshniwal: That also I'll show you using a neural network in the… one of the next terms. [01:46:22] Durga Toshniwal: Any questions, anyone? [01:46:30] Durga Toshniwal: So, if there are no further questions, then we can break for today. [01:46:35] Durga Toshniwal: Thank you all, have a great Sunday. [01:46:37] Durga Toshniwal: Thank you. [01:46:38] vinit shah: Bye-bye. [01:46:39] gunjan bhaiya: coupon. [01:46:40] Nirav Mehta: Enduro. [01:46:42] Durga Toshniwal: Thank you.