# 08 2026-02-21 Deep Learning and NLP Contd

course: Module 3 — Deep Learning & NLP
module: Module-3-Deep-Learning-NLP
date: 2026-02-21
type: transcript
video_url: https://personal-learn.armco.dev/files/_Recordings/Module-3-Deep-Learning-NLP/08_2026-02-21_Deep_Learning_and_NLP_Contd.mp4

---
[00:07:39] Shikhar Gupta: Hello.
[00:07:42] Shikhar Gupta: Looks similar.
[00:09:15] Durga Toshniwal: A very good evening to all of you, and welcome to today's session.
[00:09:20] Durga Toshniwal: So, I'll just be starting in another 1-2 minutes, just in case few more people join.
[00:10:36] Durga Toshniwal: Let me just share my slides.
[00:10:39] Durga Toshniwal: And, we'll go on from there.
[00:10:44] Durga Toshniwal: Just a second, I'm just sharing.
[00:11:00] Durga Toshniwal: I hope my slides are visible to all of you.
[00:11:09] Shivansh Sharma: This one.
[00:11:10] Durga Toshniwal: Yeah, thanks so much.
[00:11:12] Durga Toshniwal: Okay, so we had discussed neural networks, in some of the previous turns.
[00:11:18] Durga Toshniwal: And I'm going to discuss in more detail the role of activation function. So, we all discussed
[00:11:24] Durga Toshniwal: That neural network, in their basic form,
[00:11:29] Durga Toshniwal: If we talk about weights getting multiplied by inputs, and so just let me choose my… But…
[00:11:42] Durga Toshniwal: So here, let's say we are having some inputs like X1 and X2.
[00:11:47] Durga Toshniwal: And we are having some weights, like W1, W2… W3, like that.
[00:11:55] Durga Toshniwal: Then what are we going to have?
[00:11:57] Durga Toshniwal: At this point, we'll have… X1?
[00:12:03] Durga Toshniwal: What's happening with this pen? I don't know. X1W1.
[00:12:08] Durga Toshniwal: And then, we'll have… X2W2.
[00:12:13] Durga Toshniwal: And something like that.
[00:12:15] Durga Toshniwal: So, this is nothing but a linear sum.
[00:12:19] Durga Toshniwal: Of weighted inputs that are being received at the input nodes of this neural network.
[00:12:26] Durga Toshniwal: This… these are the input notes, this is the output… these are the output nodes, and these are the hidden layers.
[00:12:34] Durga Toshniwal: So, what we are having just is a simple weighted sum of the inputs, so this is just a linear sum only.
[00:12:41] Durga Toshniwal: Then, how is it that nonlinearity is coming in a picture in a neural network?
[00:12:51] Durga Toshniwal: So, what happens is that we have got activation functions that work at each of these hidden layers.
[00:12:59] Durga Toshniwal: And when such activation functions work.
[00:13:04] Durga Toshniwal: These are the ones that bring in nonlinearity.
[00:13:07] Durga Toshniwal: to our weighted inputs or inputs. So, we'll just illustrate how neural networks are actually going to play a role.
[00:13:17] Durga Toshniwal: in bringing about nonlinearity. So, let's assume if we have this kind of an input.
[00:13:25] Durga Toshniwal: And, just a second…
[00:13:29] Durga Toshniwal: So we have this kind of an input, as you are seeing here, and
[00:13:34] Durga Toshniwal: We want, the inputs to pass through the neural network, and the expected output has to be something like this, as you can… as you are seeing here.
[00:13:43] Durga Toshniwal: So, in order to do that, in order to achieve this, neural networks will be useful, and other
[00:13:50] Durga Toshniwal: Other forms of classification may not be able to do classification and prediction. So, let's take an example.
[00:13:59] Durga Toshniwal: To see how it works. Let's assume that, we are going to, have a drug being tested.
[00:14:08] Durga Toshniwal: And this, drug is being administered
[00:14:12] Durga Toshniwal: As a test case, two, three sets of people.
[00:14:16] Durga Toshniwal: And these 3 sets of people are being given 3 different dosages. So, this set is being given a low dosage.
[00:14:27] Durga Toshniwal: Then, another set is there, which is being given, medium dosage.
[00:14:32] Durga Toshniwal: And there's a third set which is being given high dosage.
[00:14:36] Durga Toshniwal: And we want to see, which of the dosage, whether the low dosage, or the medium dosage, or the high dosage.
[00:14:43] Durga Toshniwal: Is effective in treating a particular ailment.
[00:14:48] Durga Toshniwal: in the population.
[00:14:51] Durga Toshniwal: Right, so now, and then we want to use this data to make a prediction model, which will be able to predict
[00:15:00] Durga Toshniwal: that depending on a new kind of dosage, which is different from this low, medium, and high, if there's a new kind of dosage, then what is going to be its effectiveness in future? So this is the predictive model that we want to build using this particular data.
[00:15:18] Durga Toshniwal: So now let's assume that while, the test is being done.
[00:15:23] Durga Toshniwal: What we saw is that when a low dosage is actually administered to the set of people, then the efficacy of the drug is very low.
[00:15:37] Durga Toshniwal: Or it is not effective at all. Hence, if we were to draw a graph between the dosage, this is the dosage, and this is the efficacy of the
[00:15:47] Durga Toshniwal: Drug. Then, with low dosage, so this relates to low dosage.
[00:15:53] Durga Toshniwal: So, with low dosage, the efficacy is zero. So, this is what we are showing.
[00:15:59] Durga Toshniwal: With this graph.
[00:16:01] Durga Toshniwal: Now, if we talk about medium dosage, medium level dosage, which is, like, if this is 0, then it comes to 0.5.
[00:16:10] Durga Toshniwal: If it is medium dosage, then it was found out that the drug, becomes very effective at… with medium dosage. So, therefore, the efficacy of the drug is high, which is approximately 1. So here, efficacy is 1, and the dosage is medium.
[00:16:30] Durga Toshniwal: This is what is indicated in the graph. And for the third set of people who were administered with high dosage, these set of people were administered with high dosage.
[00:16:40] Durga Toshniwal: And then it was found out that once again, at high dosage, the drug was not effective. And therefore, if this was 0, and this was 0.5,
[00:16:51] Durga Toshniwal: So, this is one.
[00:16:53] Durga Toshniwal: So, at a dosage which is high, which is 1. So, the range of the dosage is from 0 to 1.
[00:17:00] Durga Toshniwal: Then, if it is 1, then again the efficacy of the drug becomes 0. So, it was 1 at 0.5, it is… it was 0 at low and high dosages. So, this is the data now we are having.
[00:17:13] Durga Toshniwal: And we want to model this data.
[00:17:17] Durga Toshniwal: To make use of it in future, right?
[00:17:20] Durga Toshniwal: So now, once we have this kind of data, you can see here, this is the dosage.
[00:17:26] Durga Toshniwal: This is the efficacy. This is low dosage. This is medium dosage.
[00:17:32] Durga Toshniwal: And this is high dosage.
[00:17:34] Durga Toshniwal: And we want to see whether, if we are having, let's say, some new dosage, In future.
[00:17:46] Durga Toshniwal: Then… Father, that… will be… Effective or not.
[00:17:59] Durga Toshniwal: This is what we want to predict.
[00:18:04] Durga Toshniwal: And for that, we have this data that is available with us, and we want to build a prediction model using this data so that we can test in future if we have a new dosage which is not low, high or medium, then what will be the impact on the efficacy of the drug?
[00:18:22] Durga Toshniwal: So now, what we'll try to do is, we'll try to fit some curve on these given points.
[00:18:28] Durga Toshniwal: So, if we try to fit the curve on these points, then what will happen is that no matter how many times we try to fit it, if we try to fit a straight line, then it will pass through two of the points, but not through the third point. For example, if we try to fit a line like this.
[00:18:49] Durga Toshniwal: which, let's say this is, this is one set of points, this is another set of points, and this is another set of points, then it is passing through these two set of points, but the third one is left out, so it's not passing through the third one. Now, if we try to draw a straight line, which passes through these
[00:19:08] Durga Toshniwal: Then, these tools… This is covered, but this is left out.
[00:19:14] Durga Toshniwal: So, what we are having is, no matter what we do,again, if we try to draw a line that passes through these two.
[00:19:21] Durga Toshniwal: Then this is left out.
[00:19:23] Durga Toshniwal: Which means that if we try to fit a straight line to these, set of points, then always two set of points are covered, and the third one will be left out, so we will not be able to
[00:19:36] Durga Toshniwal: fit a straight line onto these three sets of points. So, we want to use something, we should use something different.
[00:19:43] Durga Toshniwal: To use something different, what we will suggest is the use of neural network. So, neural networks are very good at fitting nonlinear curves on the given data.
[00:19:57] Durga Toshniwal: So…
[00:19:58] Durga Toshniwal: What the neural network could do is, it could fit something like this. It's a nonlinear decision boundary.
[00:20:05] Durga Toshniwal: And it is fitting the given data nicely. So this kind of swiggle, or a curved line, is one.
[00:20:12] Durga Toshniwal: That can be fitted using a neural network, but not by anything else.
[00:20:19] Durga Toshniwal: So now, as per this squiggle, we can see that
[00:20:23] Durga Toshniwal: Any dosage which is very low.
[00:20:26] Durga Toshniwal: is going to have a low efficacy. Any dosage which is high is also going to have a low efficacy. Anything which is near to a medium
[00:20:35] Durga Toshniwal: medium dosage is going to have a high efficacy. So, all the data that we have is fitting quite nicely by this squiggle that has been
[00:20:48] Durga Toshniwal: Drawn using a neural network.
[00:20:52] Durga Toshniwal: So, as we can see here, for low dosages.
[00:20:56] Durga Toshniwal: the value is zero, as it is supposed to be, because when the drug was administered to a set of people, and low dosage was given, then the efficacy was zero. Similarly.
[00:21:09] Durga Toshniwal: If the medium dosage was administered to another set of people, then the efficacy was quite high, which is approximately 1.
[00:21:20] Durga Toshniwal: And when the dosage was very high.
[00:21:24] Durga Toshniwal: And it was administered to another set of people, the efficacy was again low. So what we are getting is a squiggle like this, and this we are going to get with the help of a neural network.
[00:21:36] Durga Toshniwal: Now, if we have more complex data.
[00:21:39] Durga Toshniwal: Then this kind of more complex data, as you can see here, that we… again, there is some other drug.
[00:21:45] Durga Toshniwal: In which the data looks to be more complex, where for a medium, you know, for some value of dosage, the efficacy is high.
[00:21:57] Durga Toshniwal: For some others, it is slow like that.
[00:21:59] Durga Toshniwal: So, again, in this case also, a very complex, a more complicated.
[00:22:04] Durga Toshniwal: Data set is there, and hence, a more complicated decision boundary will be required, and it will be something like this kind of a double-harmed squiggle, or this kind of a curve that will be fitted, and that we can do again with the help of a neural network, and this is not possible with
[00:22:23] Durga Toshniwal: Other simple, linear or nonlinear classifiers.
[00:22:28] Durga Toshniwal: So now let's see how this happens, because as I already told you.
[00:22:33] Durga Toshniwal: That if you are having a neural network like this, as I showed to you earlier.
[00:22:39] Durga Toshniwal: And then, you know, these nodes are receiving a weighted sum.
[00:22:46] Durga Toshniwal: So, a weighted sum of something, like, for example, if we have two inputs, X1 times weight 1 plus X2 times weight 2, let's say this is going to some hidden node, then this is just a linear sum, there's no nonlinearity into it. Then, how are we getting this squiggle, which is actually a nonlinear curve?
[00:23:07] Durga Toshniwal: So now, to… see how we get it. We are going to use this particular sample data.
[00:23:13] Durga Toshniwal: Of drug dosage administration and the efficacy versus the dosage that we already have derived. And we are going to use this.
[00:23:21] Durga Toshniwal: Now, let's assume, we are having this kind of a structure with us. This is nothing but a neural network.
[00:23:28] Durga Toshniwal: So, what are its constructs? So, what we are having here, in this neural network, there are different components.
[00:23:36] Durga Toshniwal: So, first of all, this is the input node.
[00:23:41] Durga Toshniwal: This is the output node.
[00:23:43] Durga Toshniwal: And this is the hidden layer. So these are the nodes in the hidden layer.
[00:23:48] Durga Toshniwal: So, you all already know what are these.
[00:23:51] Durga Toshniwal: So this whole thing, whole construct is a neural network. So, it will have…
[00:23:56] Durga Toshniwal: Input neuron, which takes as input the dosage value.
[00:24:01] Durga Toshniwal: Then, it has an output neuron, And this neuron, It's actually going to…
[00:24:09] Durga Toshniwal: Give the output, which is the efficacy.
[00:24:12] Durga Toshniwal: And then… What we are going to get, finally.
[00:24:17] Durga Toshniwal: Is this kind of a squiggle, which is used to make prediction.
[00:24:22] Durga Toshniwal: So that if we have a dosage which is not this, not this, or not this, which were used in the experiment, let's say if we have a dosage somewhere here.
[00:24:32] Durga Toshniwal: of this value, then we can find out what will be the efficacy. So, based on this curve, we can find out that the efficacy will be something like a 0.8 or 0.9, something like that.
[00:24:42] Durga Toshniwal: So this, this kind of a green squiggle can be created with the help of a neural network, kind of neural network that you're seeing here. So then, as I mentioned already, so these are the nodes
[00:24:54] Durga Toshniwal: They are represented as these squares. I already told you this is the input node, and this is the output node. This takes dosage as input, this generates efficacy as the output.
[00:25:08] Durga Toshniwal: the output would be, you know, something of this kind. This is what we are expecting it to be.
[00:25:13] Durga Toshniwal: Now, these are what? These are the nodes in the hidden layer. So, these are the nodes in the hidden layer. I'm just writing it as a check.
[00:25:21] Durga Toshniwal: And I'll just talk about what is there inside it.
[00:25:25] Durga Toshniwal: And then, in between the nodes, as you all already know, there are weights and biases. So, we have these edges, these are the edges that are connecting these nodes, like this, like this, like this, right?
[00:25:40] Durga Toshniwal: So, these edges actually are containing certain numbers on them.
[00:25:45] Durga Toshniwal: Right, so these are the numbers, or the values, that stand for the weights.
[00:25:50] Durga Toshniwal: These are the weights, maybe this is weight 2. Then this is some bias, this is another bias, and these are further weights. So what we are having are numbers. These numbers represent the different parameter values.
[00:26:04] Durga Toshniwal: And these values are estimated when the neural network is actually finalized, so when we try to fit the data.
[00:26:13] Durga Toshniwal: On the neural network.
[00:26:15] Durga Toshniwal: Then, what we get are the final weights. Remember, we had discussed earlier that whenever,
[00:26:23] Durga Toshniwal: Neural network is used, and a data is fit onto it, initially, random weights are assigned to it.
[00:26:30] Durga Toshniwal: Random weights are assigned to the different edges in the neural network, and then what happens is that, using these random weights, output is generated, and the difference between the actual output and the predicted output is seen.
[00:26:48] Durga Toshniwal: And this is said to be the error, and this is calculated with the help of the loss. And then that error is actually… so if this was a neural network.
[00:26:58] Durga Toshniwal: Let's say this is a very simple neural network.
[00:27:01] Durga Toshniwal: Like this, it is a fully connected neural network.
[00:27:05] Durga Toshniwal: So then, if it is predicting some value Y,
[00:27:08] Durga Toshniwal: Y hat, and this was the actual value Y. These two are compared, and with the help of the loss function, the difference between them is C. This difference is actually back-propagated.
[00:27:20] Durga Toshniwal: And when it is back propagated, it goes right up to the input layer, and then what happens is that these weights… some weights were there, right? So these weights are actually adjusted in such a way that once, now, the new weighted input
[00:27:37] Durga Toshniwal: passes through the neural network, then the value of the Y hat is more closer to the actual value of Y. That is, the prediction becomes more closer to the value of I, and this loss actually decreases. And we go on doing this till finally the weights are stabilized, and the loss is optimally reduced.
[00:27:57] Durga Toshniwal: So, this is what is done in back propagation. So, here, what we are, seeing are these widths, which are the parameters.
[00:28:06] Durga Toshniwal: Weights and biases on the network, neural network.
[00:28:11] Durga Toshniwal: And, these, values, as I said, these parameters, that is, the weights and biases, are estimated when the neural network is fitted onto the data.
[00:28:23] Durga Toshniwal: So, as I said, that initially some random weights and all are assigned, then with the help of back propagation of errors, what we do is that we try to reduce the loss, and we go on iterating till the loss is optimally reduced. At this point in time.
[00:28:39] Durga Toshniwal: Whatever weights and vices we are seeing.
[00:28:42] Durga Toshniwal: all this process is already complete. That is, the neural network has been fitted to the data, and the weights and the biases are already finalized, okay? So you assume this so that we have our neural network ready with us.
[00:28:58] Durga Toshniwal: So, I'll just stop at this point. Any questions, anyone?
[00:29:04] Durga Toshniwal: Regarding whatever we have covered so far.
[00:29:14] Durga Toshniwal: Is it okay? No questions?
[00:29:20] Muni Prakash Ganji: Yeah, as what we learned means, it's definitely in our neural networks, the system will calculate the weights as per the best possible
[00:29:29] Muni Prakash Ganji: Waiting, so that we'll get a optimal solution.
[00:29:33] Muni Prakash Ganji: Like, medium, example, which are already short.
[00:29:37] Muni Prakash Ganji: Yeah. Beautiful.
[00:29:38] Durga Toshniwal: Beautiful.
[00:29:39] Durga Toshniwal: That's correct, and this is done with the help of the back propagation of the loss.
[00:29:44] Durga Toshniwal: So, that, derivation, I… also, I will show you.
[00:29:48] Durga Toshniwal: But, as I already told you, that whenever we are trying to stabilize or fit a neural network onto our data, then these weights are, you know, there are several iterations in which the loss is back-propagated, and then the weights are adjusted, then again, forward pass happens.
[00:30:06] Durga Toshniwal: Then again, loss is calculated, then again, it is back-propagated, and this goes back and forth till there is an optimally low value.
[00:30:14] Durga Toshniwal: For the, loss. So those weights and vices are taken as the final weights and vices. So what you are saying is correct. In this case, it will be, like, the output will be 0.5. That will be the best output here.
[00:30:28] Durga Toshniwal: Anything closest to it will be the best.
[00:30:34] Durga Toshniwal: Any other questions or comments, anyone, before we proceed further?
[00:30:41] Durga Toshniwal: Okay, so let me, proceed further from here.
[00:30:45] Durga Toshniwal: Now, we already know that if we are to fit any straight line, then we have an equation like Y is equal to
[00:30:51] Durga Toshniwal: Mx plus C. Like, this is the standard form, Y is equal to MX plus C. Here, I've just written it in the form X equal to MZ plus C. And, we have to estimate the value of M and C like this, right? So this is said to be the slope, this is the slope.
[00:31:10] Durga Toshniwal: And this is the intercept.
[00:31:13] Durga Toshniwal: So, I'll try to explain how neural network is working and how activation functions are.
[00:31:20] Durga Toshniwal: Playing a role in obtaining a nonlinear output with the help of, this kind of a equation.
[00:31:29] Durga Toshniwal: So now, just a second.
[00:31:41] Durga Toshniwal: So, here also, as I already explained, actually, these parameters, which are the weights, And the bias values.
[00:31:52] Durga Toshniwal: So, these weights and biases are actually unknown.
[00:31:55] Durga Toshniwal: when, initially, when we try to fit the model onto the data, and that's why they are mentioned as question marks, right? And then, when we actually fit the data, then we try to, as I mentioned, that we try to iterate upon the
[00:32:11] Durga Toshniwal: We try to see, back-propagate the loss that is obtained, that is a difference between the actual output
[00:32:19] Durga Toshniwal: which is Y, and the generated output, the difference between them is the loss.
[00:32:27] Durga Toshniwal: And this loss is actually back-tropagated on the neural network.
[00:32:32] Durga Toshniwal: And then, like this, it goes on. So, these weights and biases would be calculated in this process, and whatever is final is actually retained.
[00:32:43] Durga Toshniwal: So now, after doing all this, let's assume that our weights are already finalized.
[00:32:50] Durga Toshniwal: Okay, and we have used back propagation for this. So,
[00:32:59] Durga Toshniwal: So now that we have fit the neural network onto this particular specific data, we have found out the weights and biases. How we do it is by using backpropagation, which I'll explain to you in detail and derive and show you.
[00:33:13] Durga Toshniwal: In one of the upcoming terms. For now, you just assume that these are the final weights and biases that we have obtained.
[00:33:21] Durga Toshniwal: Now, our still… our question is, this one, that, we still have to estimate M and C, so… and find out then the value of X, and so on and so forth. We are going to do it.
[00:33:36] Durga Toshniwal: But before that, I would like to actually show to you that, I had explained to you that this is the input node.
[00:33:45] Durga Toshniwal: And this is the output node.
[00:33:48] Durga Toshniwal: And these are what nodes in the hidden layer.
[00:33:52] Durga Toshniwal: This one?
[00:33:54] Durga Toshniwal: And this one. These two are the nodes in the hidden layer.
[00:33:57] Durga Toshniwal: And what you can see inside them is this kind of a…
[00:34:01] Durga Toshniwal: curve inside them, right? So these are identical curves, however,
[00:34:06] Durga Toshniwal: Because these two nodes are belonging to the same layer, so the same activation function is actually working on them. So this kind of a curve is nothing but the activation function.
[00:34:20] Durga Toshniwal: And this is a primary building block in a neural network, because, as I already told you, that if I have some input.
[00:34:29] Durga Toshniwal: and I just multiply it by some weight, and then I have another input, X2, and multiply it by another weight. This is just a linear weighted sum. There's no non-linearity in it. So how is it coming? It is coming with the help of these activation functions, and I'll shortly show you.
[00:34:45] Durga Toshniwal: How they work.
[00:34:47] Durga Toshniwal: So, as of now, what I want you to know is that using this kind of a curse.
[00:34:54] Durga Toshniwal: Oh… Which are the… primary building blocks for activation functions. What we will, what we can obtain is
[00:35:04] Durga Toshniwal: That we can actually fit this data.
[00:35:08] Durga Toshniwal: You know, fit the squiggle-like of a curve.
[00:35:13] Durga Toshniwal: We can actually fit a squiggle, like a curved line onto this data.
[00:35:19] Durga Toshniwal: With the help of what we call is… This kind of a curve.
[00:35:25] Durga Toshniwal: Okay,
[00:35:43] Durga Toshniwal: Sorry that I got muted. So, what is happening is that…
[00:35:51] Durga Toshniwal: So, what is happening is that what we are obtaining is, you know, this kind of a…
[00:35:57] Durga Toshniwal: These kind of curved lines, which are the activation functions. And these are the primary building blocks so that we are able to fit this kind of a curved line
[00:36:08] Durga Toshniwal: Which is a squiggle line onto this data, right?
[00:36:12] Durga Toshniwal: How we will do?
[00:36:14] Durga Toshniwal: So, these activation functions, actually, you can see here, are reshaped.
[00:36:23] Durga Toshniwal: So, and I'll show you how they are reshaped, and then twisted, and all that is happening. And finally, the squiggle is getting obtained. You will eventually see that this portion of the activation function is actually mapping to this.
[00:36:38] Durga Toshniwal: And this portion is actually mapping to this. And then these two added together, actually. So when this will be added to this, what you will get is something like this, which is the squiggle that we want.
[00:36:54] Durga Toshniwal: Gotcha. And, as you can see here, this this went superimposed
[00:37:02] Durga Toshniwal: Like this, this is what we obtain.
[00:37:06] Durga Toshniwal: Right, so then the two are summed together. What we are getting is this green squiggle that is fitting the data.
[00:37:13] Durga Toshniwal: Okay, so,
[00:37:16] Durga Toshniwal: As you can see here, if I try to actually… so you might be wondering how we are getting this, actually. When we try to add this superimpose this and this, then what I'll get is something like this, a line which is higher up, right? I'm adding something.
[00:37:33] Durga Toshniwal: Then, if I add this and this, I'll get something here.
[00:37:38] Durga Toshniwal: Then, if I…
[00:37:40] Durga Toshniwal: add, let's say I come to this point and this point, so this much of it is negative. So, what will happen? The squiggle will start going down, so like this.
[00:37:50] Durga Toshniwal: Similarly, if I try to add this much, and this is some negative portion, so it will be going down here. So, like this, so we'll actually get something like this.
[00:38:05] Durga Toshniwal: Similarly, this has a value this much, and this is something negative, so that much, when I subtract from it, I get this point.
[00:38:13] Durga Toshniwal: Then there's a value this much.
[00:38:16] Durga Toshniwal: This much, and we subtract this.
[00:38:19] Durga Toshniwal: then what we get is this curve. Like this, we go on getting this, and we superimpose it. Now, the question is, how do we actually, from these activation functions, how are we getting these, which we superimpose, to obtain the nonlinear swigger, is what I'm going to explain to you.
[00:38:36] Durga Toshniwal: First of all, there are different kinds of activation functions
[00:38:41] Durga Toshniwal: That we discussed, right? Say, for example, I discussed with you Oh.
[00:38:51] Durga Toshniwal: So… This is Relyu.
[00:38:55] Durga Toshniwal: You might be knowing, I already discussed. Anything negative is 0. Anything positive is equal to the value itself. Then we have tan h.
[00:39:05] Durga Toshniwal: Right, so that is something like this. And what we are having here, this is a soft plus function.
[00:39:15] Durga Toshniwal: This is soft plus. It is called a soft plus activation function.
[00:39:26] Durga Toshniwal: So, this is also quite a popular, activation function, and we are going to make use of this, soft plus activation function.
[00:39:41] Durga Toshniwal: So…
[00:39:43] Durga Toshniwal: Whenever, okay, one thing I want to point out is that whenever we finalize upon a neural network, we have different hyperparameters.
[00:39:57] Durga Toshniwal: Which I already have covered in the previous turn.
[00:40:00] Durga Toshniwal: And the choice of the activation function is also a hyperparameter, so we can choose ReLU, we can choose softmats, we can choose SoftPlus, we can choose any particular function.
[00:40:12] Durga Toshniwal: That is useful for us.
[00:40:15] Durga Toshniwal: Here, we are using the soft plus function, okay? Let's assume we chose to use the soft plus function.
[00:40:22] Durga Toshniwal: And then, we'll proceed to work on this.
[00:40:31] Durga Toshniwal: So,
[00:40:33] Durga Toshniwal: So, in this simple problem, where we have 3 kinds of dosages, low, medium, and high, we are trying to solve this problem to obtain a squiggle-like output.
[00:40:44] Durga Toshniwal: Using a very simple, neural network, which has just one input neuron, One output neuron here.
[00:40:53] Durga Toshniwal: And it has just one hidden layer.
[00:40:57] Durga Toshniwal: And in the hidden layer, there are just two neurons here, so this is one neuron.
[00:41:02] Durga Toshniwal: In the hidden layer, this is one neuron, and this is another… this is another neuron. So these are the two neurons.
[00:41:10] Durga Toshniwal: In the hidden layer.
[00:41:12] Durga Toshniwal: And the activation function is a softmax function.
[00:41:15] Durga Toshniwal: And other than that, we have weights and biases that have been adjusted. So, these are… these are some weights, this is a bias, this is a weight bias. Weight, weight like that. This is a bias. All these have been found out.
[00:41:29] Durga Toshniwal: All these are parameters that have been already tuned with the help of backpropagation, as I already discussed.
[00:41:37] Durga Toshniwal: Now, what we are going to do is, we are going to try
[00:41:41] Durga Toshniwal: to put the different dosage values and see what kind of output gets generated. So, first of all, we'll start with a low dosage, right? And what is the low dosage going to be? The lowest value of the dosage will be, like, the input dosage
[00:42:02] Durga Toshniwal: We'll take… As 0, and we'll put it into the… neural network.
[00:42:10] Durga Toshniwal: Right, so we'll put the dosage of 0 in the neural network and see how it works.
[00:42:17] Durga Toshniwal: Okay.
[00:42:19] Durga Toshniwal: So, I think there's a question. Lokesh wants to know that how does the neural network decide the initial values for weights and biases before training begins?
[00:42:30] Durga Toshniwal: So, so, actually, this initial values are decided with the help of different kind of optimizers that
[00:42:41] Durga Toshniwal: That suggests certain values that can be started with.
[00:42:45] Durga Toshniwal: So, these are random values that, every particular algorithm
[00:42:51] Durga Toshniwal: optimizer algorithm will suggest, and then later these will be iterated upon, and then they'll be changed. So this is how we actually start with these values.
[00:43:06] Durga Toshniwal: So now, what we are inputting…
[00:43:09] Durga Toshniwal: to the neural network is a dosage 0, which is here. So this is our soil.
[00:43:15] Durga Toshniwal: So this is the dosage 0. We put it as an input here.
[00:43:21] Durga Toshniwal: Now, we are going to see how this is going to work.
[00:43:25] Durga Toshniwal: So, let's see, when we put a zero here, then, this input actually
[00:43:33] Durga Toshniwal: this input is going to go, towards the hidden layer, right? And now we are going to calculate
[00:43:41] Durga Toshniwal: How, this dosage is going to impact
[00:43:45] Durga Toshniwal: So, as we already know, we had a function like x is equal to… MZ plus C.
[00:43:52] Durga Toshniwal: So, using this equation, And we'll put Z as the dosage here.
[00:43:58] Durga Toshniwal: And then we have the weight here.
[00:44:01] Durga Toshniwal: And we have the bias here that we'll substitute here to calculate the value of FX.
[00:44:07] Durga Toshniwal: So, if X is equal to, sorry, dosage is equal to 0,
[00:44:14] Durga Toshniwal: then X value would be equal to… Dosage.
[00:44:19] Durga Toshniwal: Times this weight, Which is minus 34.4.
[00:44:25] Durga Toshniwal: plus… This C, which is the bias here, so it is 2.14.
[00:44:31] Durga Toshniwal: So, dosage is 0, so it's 0 times minus 34.4 plus 2.14, and this is equal to 2.14. So, we'll have the value of X equal to 2.14. When dosage is equal to 0, X is equal to 2.14.
[00:44:51] Durga Toshniwal: Now, this, soft plus function is like this.
[00:44:57] Durga Toshniwal: The soft plus function follows this equation. FX is equal to log of 1 plus e to the bar x.
[00:45:04] Durga Toshniwal: So, what we'll do, we'll…
[00:45:06] Durga Toshniwal: substitute the value of X, so what will… so it's actually natural log, it is ln, so I'll write it ln.
[00:45:14] Durga Toshniwal: 1 plus e to the power 2.14, like this.
[00:45:19] Durga Toshniwal: So, when we put this value.
[00:45:23] Durga Toshniwal: We are going to, get, if you try to calculate the natural log.
[00:45:30] Durga Toshniwal: What we'll get is, 2.25.
[00:45:34] Durga Toshniwal: So if you substitute this value out here, what you'll get is…
[00:45:39] Durga Toshniwal: 2.25. Okay, that is E to the power 2.14.
[00:45:44] Durga Toshniwal: If you'll recollect, this function is, like, if you recollect, we used to have functions, for example. I'll show you here, we had function for sigmoid.
[00:45:55] Durga Toshniwal: If you recollect, which was fx equal to e to the power X upon e to the power X plus 1. And for relu, the function was fx is equal to max of
[00:46:06] Durga Toshniwal: 0 comma X, like that, there were different functions. So, here, the function that we are having for soft plus is FX is equal to natural log of 1 plus e to the power X.
[00:46:19] Durga Toshniwal: This is the function that we are using. And when we put a dosage of 0,
[00:46:26] Durga Toshniwal: Then, the value, so when dosage was 0,
[00:46:32] Durga Toshniwal: X came out to be 2.14, I just calculated it.
[00:46:36] Durga Toshniwal: There's 0 times 30… minus 34.4 plus 2.14, and FX, which is equal to natural log of 1 plus e to the power 2.14, will come out to be, if we calculate it.
[00:46:50] Durga Toshniwal: It will come out to be 2.25.
[00:46:54] Durga Toshniwal: Okay.
[00:46:56] Durga Toshniwal: So then, we have this value. Now…
[00:46:59] Durga Toshniwal: What are we going to do? We are actually going to,
[00:47:03] Durga Toshniwal: you know, draw this out here. So, we have the dosage here.
[00:47:09] Durga Toshniwal: Let's see.
[00:47:11] Durga Toshniwal: So this is just the same calculation. So we have X equal to 2.4.
[00:47:16] Durga Toshniwal: Which is, given to the soft plus function, which is ln plus,
[00:47:22] Durga Toshniwal: into ln of 1 plus e to the power X.
[00:47:26] Durga Toshniwal: Which is this, and which, so this was X, this is FX.
[00:47:33] Durga Toshniwal: So this is FX, which we calculated, and we found out to be.
[00:47:37] Durga Toshniwal: I told you it was coming out to be how much?
[00:47:41] Durga Toshniwal: 2.25.
[00:47:43] Durga Toshniwal: So now… When X is equal to, so when dosage is 0,
[00:47:51] Durga Toshniwal: X is equal to 2.14, and FX is equal to 2.25. So, let's see on this activation how it comes out to be. When dosage is 0, X is 2.14.
[00:48:04] Durga Toshniwal: And FX is 2.25.
[00:48:07] Durga Toshniwal: So, if this is 2.14 somewhere, Then, FX was 2.25, so… If this is 2,
[00:48:18] Durga Toshniwal: then 2.2… Sorry, 2.25, sorry, 2.14. So this value may be slightly more than 2, and if this is slightly more than 2, then FX will be 2.25, so it will be somewhere here, like this.
[00:48:33] Durga Toshniwal: So this will be the value that we'll be having.
[00:48:38] Durga Toshniwal: Right? So it's just generally shown, like this.
[00:48:43] Durga Toshniwal: So, assuming this is, these values…
[00:48:49] Durga Toshniwal: 2.14, and assuming this value is 2.25. Okay, so what we are having is this point.
[00:48:57] Durga Toshniwal: Now, let's calculate further.
[00:49:01] Durga Toshniwal: Okay, another thing is, that, as I said, when dosage is zero.
[00:49:08] Durga Toshniwal: X is 2.14, and Y is 2.25. So here is the dosage. This dosage is 0, so this is the point here.
[00:49:17] Durga Toshniwal: And, Y, which is the efficacy, is how much? It is 2.25. So, if this is 1, this is 2.25, then this value will be here for 0. So this is the point that we just derived, and it corresponds to this point here.
[00:49:34] Durga Toshniwal: So, this is dosage 0, and this is the value of the efficacy.
[00:49:40] Durga Toshniwal: So now, let's derive for some other values also.
[00:49:43] Durga Toshniwal: So I'll just show you.
[00:49:46] Durga Toshniwal: Let's say now we are having… Dosage.
[00:49:52] Durga Toshniwal: Which is equal to 0.1.
[00:49:55] Durga Toshniwal: So, if you're having a dosage, then x-axis… For activation function, Will be how much?
[00:50:08] Durga Toshniwal: It will be… X will be equal to… dosage.
[00:50:14] Durga Toshniwal: times minus 34.4, which was the weight, plus the bias was 2.14C.
[00:50:21] Durga Toshniwal: This is 34.4, this is 2.14. I'm just substituting these values, and in this case, dosage is 0.1.
[00:50:31] Durga Toshniwal: into… Minus 34.4.
[00:50:35] Durga Toshniwal: plus 2.14. So, what are we having here? It is minus 3.44 plus 2.14.
[00:50:43] Durga Toshniwal: which comes out to be minus 1.3. This is the value of X.
[00:50:48] Durga Toshniwal: And then FX is going to be F of…
[00:50:52] Durga Toshniwal: Minus 1.3, which will be log
[00:50:56] Durga Toshniwal: 1 plus e to the power minus, e to the power X, which is minus 1.
[00:51:02] Durga Toshniwal: And if we calculate this, this will come out to be 0.24.
[00:51:06] Durga Toshniwal: So, so, what is happening here?
[00:51:11] Durga Toshniwal: In the previous case, when dosage was zero.
[00:51:17] Durga Toshniwal: X was coming out to be 2.14, and Y was coming out to be… 2.25, something like that.
[00:51:24] Durga Toshniwal: And now, when dosage is 0.1,
[00:51:26] Durga Toshniwal: Then X is coming out to be minus 0.3, Y is coming out to be 0.24. So, what does it mean? So, if we think about this activation function like this.
[00:51:36] Durga Toshniwal: This was our activation function, right?
[00:51:41] Durga Toshniwal: And we had, for X2.14, if this is 1, this is 2, this is 2.14, we had a value of 2.25, somewhere like this.
[00:51:52] Durga Toshniwal: This was the value.
[00:51:53] Durga Toshniwal: When dosage was zero.
[00:51:56] Durga Toshniwal: Now, when dosage is 0.1,
[00:51:59] Durga Toshniwal: So, 0.1 means, it will be somewhere, 0.1 means X is equal to how much? Minus 1.3.
[00:52:06] Durga Toshniwal: So this is minus 1, this is 1, this is 2.
[00:52:09] Durga Toshniwal: And this is zero, of course.
[00:52:12] Durga Toshniwal: And this is minus 2.
[00:52:14] Durga Toshniwal: So it is minus 1.3, means somewhere here. When X is this much, then Y is how much? It is 0.24.
[00:52:22] Durga Toshniwal: 0.24 will be, like… Somewhere here.
[00:52:26] Durga Toshniwal: So, this will be the value.
[00:52:28] Durga Toshniwal: So, on this curve, of the soft plus function.
[00:52:34] Durga Toshniwal: When we are having dosage 0, we are having it here. When we are having dosage equal to 0.1, we are going to this point.
[00:52:45] Durga Toshniwal: And… Like that, if we have a dosage of 1, then also we can calculate, let's see…
[00:52:59] Durga Toshniwal: So… Let me take the next one.
[00:53:04] Durga Toshniwal: Let's say now, if dosage… Is equal to 1.
[00:53:11] Durga Toshniwal: Then, x-axis… for activation.
[00:53:17] Durga Toshniwal: function.
[00:53:19] Durga Toshniwal: is going to be how much? X will be equal to?
[00:53:23] Durga Toshniwal: 1 times minus 34.4.
[00:53:27] Durga Toshniwal: plus 2.14.
[00:53:30] Durga Toshniwal: Which will be minus 34.4 plus 2.14.
[00:53:35] Durga Toshniwal: Which is equal to minus 32.26. This is your X. This is when dosage was 1.
[00:53:42] Durga Toshniwal: And now, FX will be equal to function of minus 32.26.
[00:53:48] Durga Toshniwal: which is equal to natural log of 1 plus e to the power minus 32.26. And if we actually substitute this value, it will come out to be LN1 plus 9.765. So I just pre-calculated, so I'm just writing, but you can calculate and see also if you wish.
[00:54:08] Durga Toshniwal: And this value is so small that it is approximately 0. So it will… you can say it is, 1, which is, again, approximately 0.
[00:54:18] Durga Toshniwal: So then… VIN?
[00:54:21] Durga Toshniwal: dosage… Is equal to 1.
[00:54:25] Durga Toshniwal: X will be equal to how much?
[00:54:28] Durga Toshniwal: X is minus… 32.26.
[00:54:33] Durga Toshniwal: And Y is equal to 0.
[00:54:36] Durga Toshniwal: Now, if we again try to draw this…
[00:54:39] Durga Toshniwal: Like we were trying to do previously.
[00:54:43] Durga Toshniwal: Then, what were the values?
[00:54:45] Durga Toshniwal: When dosage… was 0. X was… 2.14.
[00:54:55] Durga Toshniwal: And Y was 2.25, which is what I had drawn here.
[00:55:00] Durga Toshniwal: This is around 2.14, this is around… 2 point.
[00:55:04] Durga Toshniwal: 25.
[00:55:06] Durga Toshniwal: So this was the curve when dosage was 0.
[00:55:09] Durga Toshniwal: Then, with dosage equal to… 0.1.
[00:55:14] Durga Toshniwal: X was equal to how much,
[00:55:18] Durga Toshniwal: In that case, X, when dosage was a 0.1, was minus 1.3. I just calculated, you can check your notes, and Y became equal to 0.24. So, when X was minus 1.3,
[00:55:35] Durga Toshniwal: this much, minus 1, this is minus 1, this is minus 2, this is plus 1, this is plus 2. Then Y was, like, this much, very less.
[00:55:45] Durga Toshniwal: And now, when dosage…
[00:55:57] Durga Toshniwal: Just a second, I think I just did something.
[00:56:06] Durga Toshniwal: Sorry for that, I'll just maximize it.
[00:56:10] Durga Toshniwal: And when dosage is equal to 1,
[00:56:14] Durga Toshniwal: Then the value of X is how much?
[00:56:17] Durga Toshniwal: We just calculated it is as less as minus 32.26, and Y is 0, which means that X is, like, somewhere very far away. I'm… I cannot draw because this is minus 2, so minus 32 will be very far, assuming this is minus 32.
[00:56:36] Durga Toshniwal: then Y will be 0 somewhere here. So…
[00:56:41] Durga Toshniwal: So, what we are having, we are having a function that is slowly increasing Through this point, like this.
[00:56:49] Durga Toshniwal: And this is the region of this function that actually We are using.
[00:56:56] Durga Toshniwal: This is when dosage was zero.
[00:56:59] Durga Toshniwal: And this was when dosage is equal to 1. This was when dosage was equal to 0.1. So this is the region of the activation function that will be utilized
[00:57:10] Durga Toshniwal: In our calculations, in our further calculations. Okay.
[00:57:17] Durga Toshniwal: So now, as I already mentioned to you, this is the region that, you know, we are having here. And now, if we try to plot these values that I just showed you.
[00:57:29] Durga Toshniwal: These are the values. Y is equal to how much? Y is initially with dosage equal to 0, we are having it as 2.24, then 0.24, then 0.
[00:57:40] Durga Toshniwal: So if we try to plot these values out here.
[00:57:44] Durga Toshniwal: on this. So, when we are having our dosage as low as zero.
[00:57:49] Durga Toshniwal: then the value of Y was, how much? 2.25, something like that, right? So with 0, Y was having a value 2.25. With 0.1, Y was having a value of how much? See here?
[00:58:04] Durga Toshniwal: With 0.1 dosage, Y was having a value of 0.24.
[00:58:10] Durga Toshniwal: So, with 0.1, it was having a value of 0.24, somewhere like this. Then, as we go on decreasing the value, and add the value of dosage equal to 1,
[00:58:21] Durga Toshniwal: The efficacy was zero, that is, output was zero. So this is the kind of curve that we are getting.
[00:58:27] Durga Toshniwal: Right?
[00:58:28] Durga Toshniwal: when we are actually, utilizing. And what is this portion, actually? This corresponds to this region.
[00:58:36] Durga Toshniwal: That I just showed you, right? This is the value, this is the value of X, and this is the value of FX. So when dosage was 0, X was equal to
[00:58:48] Durga Toshniwal: 2.14.
[00:58:50] Durga Toshniwal: 2.14, and this was 2.25. So we start out with this region, and we go up to the point where Y becomes equal to 0, and X becomes minus 32.26. So this is the region of the activation function that is picked up.
[00:59:07] Durga Toshniwal: And… It's going to be?
[00:59:10] Durga Toshniwal: put here. I've just tried to plot the Y values. These are the Y values, FX values, okay?
[00:59:17] Durga Toshniwal: Now, let's see how this will be utilized. We'll see in further sections.
[00:59:23] Durga Toshniwal: So, this is the range.
[00:59:25] Durga Toshniwal: On a relatively narrow range.
[00:59:28] Durga Toshniwal: Of the activation function that has actually been sliced out.
[00:59:33] Durga Toshniwal: When we change the dosage, From 0 up to 1.
[00:59:38] Durga Toshniwal: And accordingly, X varies from, 2.14 to…
[00:59:43] Durga Toshniwal: minus 32.26, and Y varies from 2.25.
[00:59:50] Durga Toshniwal: to zero. These are the things that are shown here.
[00:59:54] Durga Toshniwal: Okay, I can see a hand raised.
[00:59:58] Durga Toshniwal: Who are the questions?
[01:00:01] Aditya Banda: Yeah, yes, ma'am, just one question. When you train the model, that graph is plotted for all the dosage values in the training data. Is that… is that correct?
[01:00:13] Durga Toshniwal: So, in this case, because, what we did was that a physical test was administered on 3 sets of people.
[01:00:22] Durga Toshniwal: So, the only data that we initially had was this one. However, now we are trying to just see
[01:00:30] Durga Toshniwal: And we want, using these three values, we want to fit the model in such a way that we can get the prediction.
[01:00:37] Durga Toshniwal: Right.
[01:00:38] Durga Toshniwal: So, and to do the prediction, we are going to use the activation function, and I'm going to show how the activation function helps to derive this kind of a squiggle. However, what you're saying is that if we have some training data, then all of it will be used, right? This is what…
[01:00:57] Durga Toshniwal: Yeah, yes, yes. In this case, our training data is just these set of points, because this is all we have.
[01:01:05] Aditya Banda: Okay.
[01:01:07] Durga Toshniwal: Any other questions anymore?
[01:01:11] gunjan bhaiya: Ma'am, couple of questions.
[01:01:13] Durga Toshniwal: So this, we have done.
[01:01:15] gunjan bhaiya: for the one node of the one layer, right? Similarly, it will be calculated for each and every node based on the weightage and
[01:01:22] gunjan bhaiya: Bias, right?
[01:01:23] Durga Toshniwal: Yes, exactly. And for the…
[01:01:26] gunjan bhaiya: Each and every value's inputs, which is going to be there, like 0.11 and everything, so on, right?
[01:01:30] Durga Toshniwal: Yes, yes. So, for example, this was very simple. Suppose I have two inputs, X1 and X2.
[01:01:37] Durga Toshniwal: And there are 3 neurons, and it's fully connected. There'll be 3 sets of weights, then there might be another layer, like this, fully connected.
[01:01:48] Durga Toshniwal: Like this, right?
[01:01:50] Durga Toshniwal: And then there might be one output.
[01:01:53] Durga Toshniwal: Yeah. So, here there'll be one activation function, here there'll be one activation function, then there'll be weights here and vices, weights here and vices, weights and vices, and then this will be the output.
[01:02:04] Durga Toshniwal: So…
[01:02:06] Durga Toshniwal: All of this will be done at each and every node. I just, for the sake of illustration, I've taken a very simple example to make things more easy to understand.
[01:02:16] Durga Toshniwal: Yep.
[01:02:17] gunjan bhaiya: Got it. And other question that, when you put a formula, formula is Y equal to MX plus C, right?
[01:02:23] Durga Toshniwal: Yeah, I'mtrying to illustrate with that.
[01:02:25] gunjan bhaiya: Yeah, but when you put the dose as an input, right.
[01:02:29] gunjan bhaiya: First point, you're deriving X, but at least if the formula, Y equal to MX plus C, which means that it should be Y, right? But you're saying it is a deriving as an X, and your activation function coming as a Y.
[01:02:41] Durga Toshniwal: Yeah. Actually, activation function is taking the value of X, but I'm not having any X here, what I'm having is just the dosage, so I want to map the dosage to the input X.
[01:02:53] Durga Toshniwal: To do that, I'm using this, because in this, we are having these parameters. So, actually, the idea here is that whatever is the input that is going… so, what we want to do is just, let's say there's an activation function here.
[01:03:09] Durga Toshniwal: So, this activation function doesn't work directly on this input, right, which is the dosage.
[01:03:15] gunjan bhaiya: Correct. It's going to work on X1W1 plus X2W2.
[01:03:20] Durga Toshniwal: And bias, if any.
[01:03:23] Durga Toshniwal: So…
[01:03:24] Durga Toshniwal: So the input to the activation function will be what, if I talk about input X, which is the input to the activation function, it will actually be this.
[01:03:35] Durga Toshniwal: If these are the weights and the inputs, right?
[01:03:38] gunjan bhaiya: Yeah.
[01:03:38] Durga Toshniwal: So, in this case, we have only one input, which is the dosage.
[01:03:43] Durga Toshniwal: But this dosage is not the input to the activation, because the activation function receives a weighted sum of the inputs. So this is the input.
[01:03:53] Durga Toshniwal: Plus, if there's a bias, then that will also be added. So, this is the input to the activation function. So, I want the input to the activation function, right? Then only I can calculate the output of that activation function. Because, any output
[01:04:08] Durga Toshniwal: of the activation function. In this case, it is…
[01:04:11] Durga Toshniwal: this much. Where X is the input. So this is the input I'm calculating.
[01:04:18] gunjan bhaiya: Got it.
[01:04:19] Durga Toshniwal: Yeah.
[01:04:21] Durga Toshniwal: And, this is a simple example where you have just one weight and bias. Actually, it will be a matrix and all that, as I showed you earlier.
[01:04:28] gunjan bhaiya: Correct.
[01:04:29] Durga Toshniwal: Yeah, but for the sake of understanding, it will be easier to take a simple example and then to see.
[01:04:38] Durga Toshniwal: So now, let's talk about what we were discussing. So…
[01:04:43] Durga Toshniwal: We already discussed the case of, you know, the different dosages.
[01:04:51] Durga Toshniwal: Just a second.
[01:04:58] Durga Toshniwal: Oh, please.
[01:05:00] Durga Toshniwal: So I was just looking at what I covered from my notes, sorry for that.
[01:05:06] Durga Toshniwal: Yeah, I'm already covered this.
[01:05:10] Durga Toshniwal: Right.
[01:05:12] Durga Toshniwal: So then we are actually utilizing this slice
[01:05:17] Durga Toshniwal: Out of, the activation function, which is a soft plus function.
[01:05:24] Durga Toshniwal: And, we are trying to map it to some Y values, which are these values, already, as I showed you.
[01:05:35] Durga Toshniwal: And then,
[01:05:37] Durga Toshniwal: Here, the X coordinate is obtained using this, this, and the dosage. And then, accordingly, we get the Y values, and then this is the slice that we are utilizing. And then, what we are getting is
[01:05:53] Durga Toshniwal: this. Now, what are we going to do?
[01:05:55] Durga Toshniwal: Once we obtain the Y values, which I already showed you, see here.
[01:06:01] Durga Toshniwal: Paul, for example, I told you that if dosage is zero.
[01:06:09] Durga Toshniwal: then X is going to be 2.14.
[01:06:12] Durga Toshniwal: And Y is going to be 2.25, right? Similarly, the dosage is 0.1.
[01:06:19] Durga Toshniwal: then X is going to be,
[01:06:22] Durga Toshniwal: Something, I don't remember what it was, but anyway… Oh… Deposit…
[01:06:29] Durga Toshniwal: Let me just check. If it is, this, then the value would be, like, minus 32. That was with dosage equal to…
[01:06:40] Durga Toshniwal: 0. X was, like, sorry, dosage was 1, and X was minus 32.26,
[01:06:51] Durga Toshniwal: and Y was equal to 0. Here, Y was equal to how much? I just showed you, you can check it also.
[01:06:59] Durga Toshniwal: Yeah, it was minus 1.3, I recollect now. And Y was equal to 0.24. These were the values. With this, we had drawn this graph, right?
[01:07:09] Durga Toshniwal: Now, what we will do is that… We have, weight here.
[01:07:14] Durga Toshniwal: So we are now going to multiply all the Y values, which are B's values, With minus 1.3, okay?
[01:07:24] Durga Toshniwal: So now, when we multiply these values by minus 1.3, then what will happen?
[01:07:30] Durga Toshniwal: So, at,
[01:07:35] Durga Toshniwal: Okay, so… So, let's say that I have a dosage 0.
[01:07:42] Durga Toshniwal: When I had a dosage 0, I had X equal to 2.14.
[01:07:46] Durga Toshniwal: I just wrote it in the previous slide, and FX, which is equal to Y is equal to 2.25.
[01:07:53] Durga Toshniwal: No.
[01:07:55] Durga Toshniwal: we have to multiply by the next weight, which is minus 1.3. So I'll… let's say I'm calling it Y prime, which will be Y into
[01:08:07] Durga Toshniwal: Minus 1.3. So, it will be 2.25 into minus 1.3.
[01:08:14] Durga Toshniwal: And that comes out to be…
[01:08:16] Durga Toshniwal: How much it comes out to be?
[01:08:19] Durga Toshniwal: minus 2.93.
[01:08:21] Durga Toshniwal: So actually, this point, after getting multiplied by minus 1.3, is getting mapped somewhere here. If this is 1, this is 2, this is 3.
[01:08:31] Durga Toshniwal: then it's coming to somewhere here now. This is the new point.
[01:08:35] Durga Toshniwal: From here, it is getting converted to this, and actually, the length of this is less, this is 2 and this is 3, I don't have enough distance, so…
[01:08:45] Durga Toshniwal: I'm just, drawing it here.
[01:08:48] Durga Toshniwal: Okay, so then… So, we are multiplying 2.25 with minus 1.3, getting minus 2.
[01:08:57] Durga Toshniwal: 9-3, which we are going to plot. Okay.
[01:09:01] Durga Toshniwal: No.
[01:09:03] Durga Toshniwal: The next, what we can consider is dosage. So, with dosage 0, we had,
[01:09:10] Durga Toshniwal: X equal to 2.14. I'm writing so that you don't miss out.
[01:09:17] Durga Toshniwal: And it doesn't create any confusion. And Y hat, I already… Y', I already calculated as minus 2.93.
[01:09:26] Durga Toshniwal: And then, if we have dosage equal to 0.1, then X was minus 1.3,
[01:09:33] Durga Toshniwal: And FX, which is equal to Y, which I calculated, is 0.24.
[01:09:38] Durga Toshniwal: And now, what will be Y prime? It will be… 0.24 times minus 1.3.
[01:09:45] Durga Toshniwal: And this will be minus 0.312.
[01:09:48] Durga Toshniwal: So, with dosage equal to minus 0.1,
[01:09:52] Durga Toshniwal: your Y prime is equal to minus 0.312.
[01:09:57] Durga Toshniwal: Okay, this is the case.
[01:09:59] Durga Toshniwal: Then, if we take dosage equal to 1,
[01:10:02] Durga Toshniwal: then X is equal to, minus 32.26, something like that. And FX, which is equal to Y, is equal to 0.
[01:10:12] Durga Toshniwal: And so, Y prime will be 0 and 2.
[01:10:16] Durga Toshniwal: Minus 1.3, which will be 0. Now, let's see what happens, what is the impact?
[01:10:21] Durga Toshniwal: Earlier, This was, the dosage.
[01:10:27] Durga Toshniwal: This is the efficacy.
[01:10:30] Durga Toshniwal: So, when dosage is zero.
[01:10:35] Durga Toshniwal: So this is 1, this is 0, this is 0.5, something like that.
[01:10:40] Durga Toshniwal: And this is 1, this is 2. So when dosage is… was 0 earlier, what we were having Y is 2.25. So this is 2. Let's say this is 2.25.
[01:10:51] Durga Toshniwal: And then when dosage was 0.5, then Y was 0.24.
[01:10:55] Durga Toshniwal: So, it was somewhere here.
[01:10:57] Durga Toshniwal: When dosage was 0.1.
[01:11:02] Durga Toshniwal: Sorry, point 1 will be here.
[01:11:04] Durga Toshniwal: So, 0.1, then Y would be 0.2 for something like this here.
[01:11:09] Durga Toshniwal: And then, when, dosage was 1,
[01:11:13] Durga Toshniwal: then the value would be Y. So, actually, if you try to draw it, it would come out something like this.
[01:11:19] Durga Toshniwal: Now, when we are multiplying it by the next weight, which is minus 1.3,
[01:11:26] Durga Toshniwal: This point goes to minus 2 point.
[01:11:29] Durga Toshniwal: 93. So if this is 1, this is 2, this is 3, it comes out to be here. So this one gets mapped to this point.
[01:11:37] Durga Toshniwal: And then, with a dosage of 0.1, which is probably here, it goes to minus 0.312.
[01:11:46] Durga Toshniwal: So, minus 0.3 will be here somewhere.
[01:11:50] Durga Toshniwal: So, this point… so this point gets… becomes this one here.
[01:11:55] Durga Toshniwal: And then, with the dosage 1, it is still 0, so now this becomes something like this. So, you can see, remember.
[01:12:06] Durga Toshniwal: That the activation function that we were using Was something like this.
[01:12:13] Durga Toshniwal: We were using a part of it, as I had shown in this red box, see here?
[01:12:19] Durga Toshniwal: This is the portion.
[01:12:21] Durga Toshniwal: Some portion out of it was sliced, and this was mapped to this graph, this part of it.
[01:12:28] Durga Toshniwal: And now, after multiplying this part with minus 1.3, what we are getting is
[01:12:36] Durga Toshniwal: So, this gets flipped. It just gets flipped, and how does it become like this? So, it's like, you know.
[01:12:43] Durga Toshniwal: Taking it, and then just putting it like this.
[01:12:48] Durga Toshniwal: Just the other room.
[01:12:50] Durga Toshniwal: Like this. So, we are actually…
[01:12:53] Durga Toshniwal: Just taking it and inversing it and putting it… so it is getting sliced, then it is getting mapped here, then… so this slice, which is relatively small, is getting stretched and put here, and then it is getting flipped and put it… getting put on the other side.
[01:13:11] Durga Toshniwal: This is what is happening when we talk about the first activation function. So I could see a hand raised. What is the question? I think Deepak, or who had it?
[01:13:21] Deepak Katara: Yes, ma'am. I was trying to understand how did we derive minus 1.3, because including minus, the only sole purpose that I could figure out is to flip the graph, or to create the admittedling visa.
[01:13:34] Durga Toshniwal: Or do we have?
[01:13:36] Deepak Katara: So, like, how do we conclude?
[01:13:39] Deepak Katara: Number and, the sign.
[01:13:41] Deepak Katara: Which we need to not take notice.
[01:13:42] Durga Toshniwal: which number you're talking about? These numbers?
[01:13:45] Deepak Katara: Minus 1.3, before…
[01:13:48] Durga Toshniwal: These are all the parameters only. These are the weights.
[01:13:52] Durga Toshniwal: These are the weights, the next weight, so if you have, you know, this kind of a network, then you'll be having weights here.
[01:14:01] Deepak Katara: And then you'll be having waits here?
[01:14:05] Durga Toshniwal: So, this is weight 3.
[01:14:08] Durga Toshniwal: Right? And I told you that weights and biases are actually
[01:14:12] Durga Toshniwal: finalized, these are parameters that are finalized after backpropagation happens. So whenever, let's say there's an input
[01:14:23] Durga Toshniwal: So, I'll draw it again. So, this is, let's say, an input.
[01:14:28] Deepak Katara: So it's not on the forward pass, it's after the backpropagation and on the correction.
[01:14:33] Durga Toshniwal: So… so what happens, actually? I'm not drawing a fully connected one, you can just assume, okay? Okay, I'll draw it.
[01:14:41] Durga Toshniwal: Something like this, okay?
[01:14:44] Durga Toshniwal: So now, when input goes in the forward pass, it will propagate this layer, this layer, and then an output will be generated. It's Y-hat.
[01:14:52] Durga Toshniwal: But what is the actual output? Why? So, we calculate the difference between these two, call it loss. Whatever the formula is, we use it and calculate loss. Now, this loss is actually back-propagated like this.
[01:15:07] Durga Toshniwal: So it goes to the layer it is nearest to, so this layer. But this layer receives its input from this layer, so it gets propagated back.
[01:15:17] Durga Toshniwal: Up to the input. Then what will happen? These weights, which are here on these edges, they will get adjusted.
[01:15:24] Durga Toshniwal: then these weights will get adjusted, these will get adjusted in such a way, with an attempt to reduce the loss. So now what will happen? After the first iteration is over, in the second iteration, new value will be found out of the output.
[01:15:40] Durga Toshniwal: Again, the difference between this new value and the actual value will be calculated. New loss will be calculated. If this loss is lesser.
[01:15:50] Durga Toshniwal: then it is better, but still it will be back-propagated, and like this will go on doing till loss gets optimally reduced. And in this process, the weights also get… weights and bias get adjusted. So this is the way how weights are finalized when a neural network is built.
[01:16:09] Deepak Katara: I actually assumed that it's just a forward pass, but I think it's after The correction of the width.
[01:16:14] Deepak Katara: Yeah. Okay. Just last one question. So…
[01:16:18] Deepak Katara: Is our expectation that the graph that we have of efficacy and OJ should be in sync with whatever the output that we are getting? Because right now, it is just the mirror image of
[01:16:32] Deepak Katara: graph of… versus FDC.
[01:16:34] Durga Toshniwal: So, we haven't yet derived the entire squiggle. It's only part of it.
[01:16:40] Durga Toshniwal: See, I've only covered one part of…
[01:16:44] Durga Toshniwal: If you go back to this neural network.
[01:16:51] Deepak Katara: Okay, so we'll calculate other nodes and.
[01:16:53] Durga Toshniwal: Yes, so this is the entire neural network, right?
[01:16:56] Deepak Katara: I just covered this much.
[01:17:00] Deepak Katara: Okay, got it.
[01:17:01] Durga Toshniwal: So now we'll be talking about the rest of it.
[01:17:07] Durga Toshniwal: Yeah.
[01:17:09] Durga Toshniwal: Any other questions, anyone?
[01:17:11] gunjan bhaiya: Yavin.
[01:17:11] Durga Toshniwal: You.
[01:17:12] gunjan bhaiya: So, in the second… after the first layer, we have given the weight 1.3.
[01:17:19] gunjan bhaiya: I assume that bias is considered zero, right?
[01:17:22] Durga Toshniwal: Yeah, so here, actually, yeah. So, bias comes in whenever we are having a neuron.
[01:17:28] Durga Toshniwal: Oh…
[01:17:29] gunjan bhaiya: So we only have one layer, so that's why it's a 0. Only simply multiplication.
[01:17:34] Durga Toshniwal: Yes, yes, that's correct.
[01:17:35] gunjan bhaiya: Okay.
[01:17:37] Durga Toshniwal: Any other things?
[01:17:38] Shikhar Gupta: How do we decide the initial weights?
[01:17:42] Durga Toshniwal: Yeah, so as I said, that initially, the weights are just randomly assigned. Then there are optimizers that will optimize, provide some optimal optimizations over the weights.
[01:17:54] Durga Toshniwal: Based on the, back-propagated loss.
[01:17:58] Shikhar Gupta: The first… the initial weights are just random.
[01:18:02] Durga Toshniwal: Yes, they are just randomly assigned.
[01:18:05] Durga Toshniwal: Okay. Yeah.
[01:18:07] Durga Toshniwal: Okay, now that we are done with the upper part.
[01:18:10] Durga Toshniwal: We will actually start with the lower half, and with this, we have obtained Oh… this.
[01:18:20] Durga Toshniwal: This is what we have obtained.
[01:18:22] Durga Toshniwal: Which is nothing but a stretched, inverted.
[01:18:26] Durga Toshniwal: this part of this slice of the activation function. Now, let's go to the
[01:18:33] Durga Toshniwal: So this is all… this is what we obtained, right? This… this is a stretched one.
[01:18:39] Durga Toshniwal: And it is a flip part of this slice.
[01:18:44] Durga Toshniwal: Now we'll go, and we'll calculate…
[01:18:47] Durga Toshniwal: the rest of it, right? So, we are also having another,
[01:18:53] Durga Toshniwal: node, which is a lower node. We are going to use this. Again, we have that function, which is,
[01:19:00] Durga Toshniwal: X is equal to… Mz plus C.
[01:19:05] Durga Toshniwal: So, we are going to use this.
[01:19:07] Durga Toshniwal: And in this particular case, yeah, our value is minus 0.52.
[01:19:13] Durga Toshniwal: So our, value of X would be, like, Dosage times… Minus 2.52.
[01:19:24] Durga Toshniwal: plus 1.29.
[01:19:26] Durga Toshniwal: This is the… And this will be what? This… this X is our, coordinate of the activation function.
[01:19:36] Durga Toshniwal: So, now… Let's say if we are having dosage, Equal to zero.
[01:19:48] Durga Toshniwal: Then, the activation function value.
[01:19:57] Durga Toshniwal: Is going to be how much?
[01:19:59] Durga Toshniwal: It will be X equal to dosage times minus 2.52.
[01:20:05] Durga Toshniwal: plus 1.29.
[01:20:07] Durga Toshniwal: And that will be how much? That will be 1.29.
[01:20:13] Durga Toshniwal: And then, we have to calculate FX.
[01:20:16] Durga Toshniwal: Which is function of 1.29.
[01:20:19] Durga Toshniwal: Which is equal to 1 upon… Sorry, which is actually… Oh, it is…
[01:20:27] Durga Toshniwal: 1 plus e to the power 1.29.
[01:20:31] Durga Toshniwal: And this becomes equal to how much? 1.53.
[01:20:35] Durga Toshniwal: So, this is 1.53 means.
[01:20:38] Durga Toshniwal: that if we talk about this activation function, the value of X was 1.29. So if this is 1, then it will be somewhere here, 1.29, and the value of FX is 1.53, so if this is 1,
[01:20:52] Durga Toshniwal: So it will be somewhere here, like this.
[01:20:55] Durga Toshniwal: So this is the point that we have chosen.
[01:20:58] Durga Toshniwal: And accordingly, for a dosage 0, the value of
[01:21:02] Durga Toshniwal: X is 1.29, and FX is 1.53.
[01:21:08] Durga Toshniwal: Okay So, it's going to be 1.53.
[01:21:12] Durga Toshniwal: Okay.
[01:21:13] Durga Toshniwal: No.
[01:21:17] Durga Toshniwal: -Oh.
[01:21:22] Durga Toshniwal: And then, oh…
[01:21:29] Durga Toshniwal: So, we are going to, similarly calculate The rest of the values.
[01:21:35] Durga Toshniwal: Okay, so this is just the illustration of what I already calculated.
[01:21:40] Durga Toshniwal: Right, so what we are getting is, this value, out here.
[01:21:45] Durga Toshniwal: Which is, with dosage 0. We are getting…
[01:21:50] Durga Toshniwal: X equal to 1.29, and then, with this, we are getting a value of 1.53.
[01:22:03] Durga Toshniwal: So…
[01:22:04] Durga Toshniwal: If dosage is 0, then since the output is 1.53, so we are mapping it, let's say if 2 was somewhere here, then this is 1.53. Okay.
[01:22:16] Durga Toshniwal: Now, let's calculate for other dosages. So, if we put dosage equal to… Point over.
[01:22:23] Durga Toshniwal: Then, what we'll have… do I have a… yeah, I'll put it here.
[01:22:27] Durga Toshniwal: So… So, earlier, When dosage was zero.
[01:22:33] Durga Toshniwal: Then, our, X was calculated to be 1.29.
[01:22:39] Durga Toshniwal: And Y was calculated as 1.53.
[01:22:42] Durga Toshniwal: And if I talk about, the… Activation function.
[01:22:51] Durga Toshniwal: So, if this is the activation function that we are using.
[01:22:56] Durga Toshniwal: Right? So, we have the func- and this is 1, this is 2.
[01:23:01] Durga Toshniwal: This is minus 1, this is minus 2.
[01:23:04] Durga Toshniwal: Then when dose was 0, X was equal to 1.2, and Y was… and this is 1, this is 2.
[01:23:11] Durga Toshniwal: Like that, so it is, like, this much.
[01:23:15] Durga Toshniwal: We are at this point.
[01:23:17] Durga Toshniwal: Now, when dose is… dosage is equal to… 4 into one.
[01:23:23] Durga Toshniwal: So, we'll do these calculations again. So, X will be equal to… dosage…
[01:23:30] Durga Toshniwal: Times minus 2.52, which was the weight.
[01:23:34] Durga Toshniwal: Plus 1.29.
[01:23:37] Durga Toshniwal: So that will be, like, 0.1 into minus 2.52.
[01:23:42] Durga Toshniwal: 2952.
[01:23:45] Durga Toshniwal: That's 1.29.
[01:23:48] Durga Toshniwal: And that comes out to be minus 0.252.
[01:23:52] Durga Toshniwal: Plus 1.29, which is 1.038.
[01:23:57] Durga Toshniwal: So, this is the value of X.
[01:24:00] Durga Toshniwal: Now we'll substitute this value and get FX.
[01:24:05] Durga Toshniwal: Which is why, actually.
[01:24:07] Durga Toshniwal: So, this will be… log 1 plus e to the power 1.038.
[01:24:13] Durga Toshniwal: And this will come out to be…
[01:24:19] Durga Toshniwal: 1.2… Oh.
[01:24:23] Durga Toshniwal: 8… 2.
[01:24:25] Durga Toshniwal: Let's see?
[01:24:26] Durga Toshniwal: Which is 1.341. So, this is why.
[01:24:31] Durga Toshniwal: So now, when X is equal to 1.038, so…
[01:24:36] Durga Toshniwal: When X is equal to… earlier X was how much? 1.29. Now it is 1.03, so it's approximately 1 only. Then value is… earlier it was 1.5, now it is 1.3.
[01:24:48] Durga Toshniwal: So it will be somewhere here.
[01:24:51] Durga Toshniwal: here, somewhere.
[01:24:52] Durga Toshniwal: This will be the part.
[01:24:54] Durga Toshniwal: Okay.
[01:24:56] Durga Toshniwal: And, when… dosage.
[01:25:00] Durga Toshniwal: Will be equal to 1.
[01:25:02] Durga Toshniwal: Then X will be equal to dosage… Times 2.52.
[01:25:11] Durga Toshniwal: plus 1.29.
[01:25:13] Durga Toshniwal: So, this will be equal to… 1 into minus 2.52.
[01:25:21] Durga Toshniwal: Plus 1.29, which is equal to…
[01:25:25] Durga Toshniwal: minus 1.23. So this is a value of X.
[01:25:30] Durga Toshniwal: And then Y, which is equal to FX, will be equal to…
[01:25:36] Durga Toshniwal: 1 plus e to the pawn minus this.
[01:25:39] Durga Toshniwal: Which comes out to be… 1.292.
[01:25:45] Durga Toshniwal: Which comes out to be 0.257.
[01:25:48] Durga Toshniwal: So, if we talk about dosage of 1, then X is equal to how much? Minus 1.2.
[01:25:54] Durga Toshniwal: So, access this much around.
[01:25:56] Durga Toshniwal: And FX is 0.2.
[01:25:59] Durga Toshniwal: So it must be somewhere here.
[01:26:01] Durga Toshniwal: So, if this was the soft plus function, then out of this, this is even a narrower slice that we are having this time.
[01:26:10] Durga Toshniwal: Earlier, the slice was, like, 2 point something up to minus 32.26. Now the slice is going from how much? X is equal to…
[01:26:22] Durga Toshniwal: X is equal to 1.29, so this is… 1.29.
[01:26:28] Durga Toshniwal: And… Sorry.
[01:26:31] Durga Toshniwal: And here, X is equal to minus…703
[01:26:36] Durga Toshniwal: 1.23, so this is the only distance.
[01:26:39] Durga Toshniwal: And in, if we talk in terms of Y, it is 1.53.
[01:26:46] Durga Toshniwal: 2… 0.257. So, such a narrow slice we have.
[01:26:52] Durga Toshniwal: Now, if we actually map these values on the dosage plot, what are we going to get?
[01:26:59] Durga Toshniwal: This was our dosage plot.
[01:27:01] Durga Toshniwal: Right?
[01:27:02] Durga Toshniwal: And this is the zero dosage, this is the 0.5 dosage, this is 1 dosage, and this is probably 1, 2, 3 here, 1, 2, 3, something like that. With a zero dosage, we were getting what? We are getting a value of 1.53.
[01:27:20] Durga Toshniwal: So this is 1.5 somewhere.
[01:27:23] Durga Toshniwal: This was the line.
[01:27:24] Durga Toshniwal: Then, with 0.1, we were getting,
[01:27:29] Durga Toshniwal: Why we just calculated, and it came out to be 1.34.
[01:27:34] Durga Toshniwal: So, with… 0.1, we were getting 1.34.
[01:27:40] Durga Toshniwal: So it was somewhere here.
[01:27:44] Durga Toshniwal: And then, with dosage equal to 1, which is here.
[01:27:48] Durga Toshniwal: We were getting something like 0.257.
[01:27:52] Durga Toshniwal: So, it must be somewhere here.
[01:27:54] Durga Toshniwal: So what we are getting is a line like this, right?
[01:27:58] Durga Toshniwal: And this is… If this was our activation function.
[01:28:04] Durga Toshniwal: Then a very narrow slice of it was actually getting utilized.
[01:28:09] Durga Toshniwal: Like this, and it was getting… What?
[01:28:12] Durga Toshniwal: on this. So, this is the activation function. It is getting stretched, and it is being put here, and this is what we are having.
[01:28:21] Durga Toshniwal: And we already have, another…
[01:28:24] Durga Toshniwal: Activation, I mean, because of the previous node, we already have a graph like this, which is there.
[01:28:31] Durga Toshniwal: It's actually approaching 0.
[01:28:33] Durga Toshniwal: And now we have got this one, this line now.
[01:28:38] Durga Toshniwal: So… Just a second.
[01:28:45] Durga Toshniwal: So you can see here that when we are trying to plug the different values of the dosages, we are getting all these points.
[01:28:52] Durga Toshniwal: Based on these calculations which I just showed to you.
[01:28:56] Durga Toshniwal: And what we finally get is this orange curve. No.
[01:29:00] Durga Toshniwal: We have got this orange curve, and this orange curve, actually, this one.
[01:29:05] Durga Toshniwal: It corresponds to this small little slice of the activation function. It has got so much stretched out, And…
[01:29:13] Durga Toshniwal: Here, you can see it was like this.
[01:29:16] Durga Toshniwal: it has actually, this has got flipped, and it has become the other way around. So it has got stretched, and it has got flipped.
[01:29:25] Durga Toshniwal: these two things have happened to it. Now, the next thing that we are having is…
[01:29:30] Durga Toshniwal: Again, the multiplication of it with…
[01:29:33] Durga Toshniwal: 2.28, right? So, now we are, going to do this to obtain The new set of values.
[01:29:43] Durga Toshniwal: Okay.
[01:29:47] Durga Toshniwal: So, I'll show you the calculations.
[01:29:51] Durga Toshniwal: So… So here, what are we having? 2.28.
[01:29:57] Durga Toshniwal: So, earlier, we already had, when dosage was zero.
[01:30:02] Durga Toshniwal: Then our X was 1.29, and Y was 1.53.
[01:30:07] Durga Toshniwal: So then, Y prime will be 1.53 times 2.28, which is the weight, which is 3.48 at 4.
[01:30:16] Durga Toshniwal: For dosage equal to 0.21, We had X calculated X as 0.
[01:30:23] Durga Toshniwal: 1.03, Y is 1.341.
[01:30:28] Durga Toshniwal: And… we are scaling, Y to Y prime, which is 1.
[01:30:34] Durga Toshniwal: 341 times 2.28, which becomes 3.05748.
[01:30:41] Durga Toshniwal: And then, when dosage is actually 1, we calculated excess minus 1.23, And, Y as…
[01:30:53] Durga Toshniwal: 0.257. And now we are scaling Y.
[01:30:58] Durga Toshniwal: So, this is our Y. We are scaling it with 2.28.
[01:31:02] Durga Toshniwal: And what we are getting is 0.586. Okay.
[01:31:06] Durga Toshniwal: So then, if we try to draw this graph, earlier.
[01:31:11] Durga Toshniwal: When dosage was 0, this is 0.5, this is 1, and this is 1, 2, like that.
[01:31:20] Durga Toshniwal: So, the previous values of Y was… with dosage 0, it was 1.5. With dosage 0.1, it was 1.3.
[01:31:32] Durga Toshniwal: Somewhere here.
[01:31:34] Durga Toshniwal: And with dosage 1, It was, like, .257.
[01:31:42] Durga Toshniwal: So, with dosage 1, it was 1.256, somewhere here. So, this was a line that you had previously.
[01:31:51] Durga Toshniwal: That was derived from the slice.
[01:31:54] Durga Toshniwal: by flipping and stretching of that activation function. Now, we are further…
[01:32:01] Durga Toshniwal: putting it to this set of new values. So the new Y prime becomes 3.4884.
[01:32:08] Durga Toshniwal: So now it is going to…
[01:32:11] Durga Toshniwal: 3.4 rated 4. This is 3, this is 4, so it's somewhere here.
[01:32:18] Durga Toshniwal: And then the next value with 0.21 is 3.05.
[01:32:24] Durga Toshniwal: So, with 0.01, it is somewhere here.
[01:32:27] Durga Toshniwal: And with 1, it is 0.586.
[01:32:30] Durga Toshniwal: So, it's going to be somewhere here.
[01:32:33] Durga Toshniwal: So, this is the new line that we get.
[01:32:37] Durga Toshniwal: After we are doing a scaling with a factor of 2.28, okay? So now, and we already have a previously obtained
[01:32:47] Durga Toshniwal: One like this.
[01:32:48] Durga Toshniwal: And now we have this and this that we have obtained, okay?
[01:32:55] Durga Toshniwal: Then, what are we going to do?
[01:32:59] Durga Toshniwal: So, this is the old one, and this is the new one that we have opted. This is after this part has got over.
[01:33:09] Durga Toshniwal: Now, what we have is, we have… after doing all this, we have to sum these two.
[01:33:15] Durga Toshniwal: Together.
[01:33:17] Durga Toshniwal: So, if we sum these values, So… I have these values here.
[01:33:23] Durga Toshniwal: If I sum them up, let's say at dosage equal to, I'll show you these 3 points at least.
[01:33:29] Durga Toshniwal: With dosage equal to zero.
[01:33:32] Durga Toshniwal: Y was having a value, finally.
[01:33:35] Durga Toshniwal: Minus 2.93, after all these scalings and all. And, this new YVOS 3.
[01:33:45] Durga Toshniwal: 4884. And finally, it came up to 0.558.
[01:33:52] Durga Toshniwal: So now, with dosage equal to zero, what are we getting? 0.558. So it must be somewhere here. This is the most recent one that we are getting after adding this value and this value.
[01:34:06] Durga Toshniwal: With dosage equal to 0.1, Y will be equal to how much?
[01:34:12] Durga Toshniwal: You can just… Add these two values.
[01:34:16] Durga Toshniwal: So, what are we having? A dosage equal to 0.1.
[01:34:23] Durga Toshniwal: The new value is… 3.05748.
[01:34:29] Durga Toshniwal: And the previous one was how much?
[01:34:32] Durga Toshniwal: Oh… Let me just think… What was the previous one?
[01:34:41] Durga Toshniwal: It was, I think.24, something like that.
[01:34:45] Durga Toshniwal: Just let me check it.
[01:34:47] Durga Toshniwal: I think we have blur somewhere.
[01:34:55] Durga Toshniwal: So…
[01:35:06] Durga Toshniwal: Yeah, with dosage equal to 0.1, this is minus 1.3.
[01:35:12] Durga Toshniwal: And Y is 0.24.
[01:35:22] Durga Toshniwal: minus 1.3.
[01:35:26] Durga Toshniwal: And it was 0.24.
[01:35:29] Durga Toshniwal: And this we are going to add with a new value, so we'll have something like what?
[01:35:34] Durga Toshniwal: It will be, like, 3 point… 2, 9 something.
[01:35:40] Durga Toshniwal: So then… At 0.1, we'll have 3 point something.
[01:35:46] Durga Toshniwal: So, where is this 3? 3 is here, 3.2.
[01:35:49] Durga Toshniwal: So…
[01:35:52] Durga Toshniwal: this will be the value, so this curve will come out like this. And if dosage is equal to 1, then what would we have as output? Y was 0, and in this one, we have 0.586. So, the output is equal to…
[01:36:10] Durga Toshniwal: 0.586, which comes out to be here.
[01:36:15] Durga Toshniwal: So, the squiggle that we obtain is something like this.
[01:36:20] Durga Toshniwal: Right? This is what we have obtained.
[01:36:24] Durga Toshniwal: Now, so this squiggle, is something like what we should have obtained.
[01:36:30] Durga Toshniwal: However, you can see that it is not ending at 1 year.
[01:36:36] Durga Toshniwal: For a dosage of 1, the value was 0.
[01:36:40] Durga Toshniwal: And for a dosage 0, the value was 0. However, what we are having is something higher up.
[01:36:47] Durga Toshniwal: So that is actually not correct.
[01:36:50] Durga Toshniwal: Now, if you recollect, what do we have now?
[01:36:57] Durga Toshniwal: Oh.
[01:37:01] Durga Toshniwal: So, we also have something like minus 0.58. So, there is a final summation with this, so this will also be added, and then what we will get is the final result. So, in this calculation, which I was showing you here.
[01:37:21] Durga Toshniwal: So, this value was coming out C minus, plus .586. So, when, this value is finally, that, final output.
[01:37:34] Durga Toshniwal: which I'll call this, will be 0.586 minus .586, right?
[01:37:44] Durga Toshniwal: And so, it will be 0 here.
[01:37:46] Durga Toshniwal: Like this, actually, this, graph will have minus .586 offset added to it, and it will move from here to here.
[01:37:58] Durga Toshniwal: do here.
[01:38:00] Durga Toshniwal: So, this is how.
[01:38:02] Durga Toshniwal: we'll get it. So, these are the steps. This was one of the…
[01:38:07] Durga Toshniwal: graphs, this was another one. When we added, I showed you… here, I showed you the addition of all these values, and by adding this and this, what we obtain
[01:38:21] Durga Toshniwal: Is this… But, actually, what we should have obtained should have come here, and started out from here.
[01:38:29] Durga Toshniwal: So, we'll finally obtain the… Final graph with this… getting added.
[01:38:37] Durga Toshniwal: And then this squiggle is actually moving to the position that it should have moved.
[01:38:42] Durga Toshniwal: So… Here's what is the magic of the activation function. We just had simple soft plus two functions.
[01:38:51] Durga Toshniwal: By using certain weights and biases, we selected a slice out of it, This slice, actually, was,
[01:39:03] Durga Toshniwal: Was actually something like this.
[01:39:06] Durga Toshniwal: It was something like this. And then this slice was actually, stretched
[01:39:14] Durga Toshniwal: Form of this one. And then it was flipped like this.
[01:39:18] Durga Toshniwal: By the multiplication with minus 1.3.
[01:39:22] Durga Toshniwal: It was slightly stretched more, because it is not minus 1, it is 1.3.
[01:39:27] Durga Toshniwal: So it was stretched more, and it was flipped in a opposite direction.
[01:39:32] Durga Toshniwal: Whereas this part of the input actually derives its nonlinearity from such a small slice of the activation function. It's only this much. If you cut it out, what you'll obtain is only this much. This actually was stretched to something like this.
[01:39:50] Durga Toshniwal: And then after, Getting this, it was, multiplied by 2.28, so it was magnified something like this.
[01:40:00] Durga Toshniwal: And then…
[01:40:01] Durga Toshniwal: this we had obtained, and this one we had obtained, we added it together, so we were getting a squiggle which was slightly higher, then this was offsetted by minus something, and we got the final squiggle. So this is a example where we can see how these activation functions play a very important role
[01:40:21] Durga Toshniwal: In fact, these are the basic blocks for this nonlinear kind of a output, otherwise we wouldn't have obtained, because what we are having here is just a linear sum of the inputs only. But these activation functions make the final result nonlinear.
[01:40:40] Durga Toshniwal: So, and this is a very small example.
[01:40:44] Durga Toshniwal: Now you can actually imagine.
[01:40:47] Durga Toshniwal: That you have a huge, neural network, let's say there are some number of inputs, there are lots of hidden layers with different numbers of neurons, like this.
[01:40:59] Durga Toshniwal: So, you can actually imagine how many…
[01:41:02] Durga Toshniwal: You know, such, what a role these activation functions
[01:41:07] Durga Toshniwal: will be playing, so I'm not showing the whole thing because of the clutter, but you can think about it. There'll be activation function at each of these. There'll be weights and biases every time.
[01:41:19] Durga Toshniwal: So every time some portion of the activation function will be sliced, it will be stressed, or it will be shortened, then it will be flipped, and all those flips and all will get added, and what you will obtain is the final function.
[01:41:35] Durga Toshniwal: So this is how activation functions work, and they add nonlinearity, and this is what is the beauty of the neural network, which can be only obtained using activation functions.
[01:41:49] Durga Toshniwal: So, this is what I wanted to cover today. Any questions?
[01:41:56] Nirav Mehta: Yeah, when you say, the activation function is large, can you kind of elaborate a bit more on that?
[01:42:05] Durga Toshniwal: Sorry, what did you say? When the activation?
[01:42:07] Nirav Mehta: Function is sliced, when we say the activation function is sliced, I did not get completely. What do we do when we slice it, and why do we slice it?
[01:42:18] Durga Toshniwal: Okay, I'll just show you.
[01:42:23] Durga Toshniwal: Just a second.
[01:42:37] Durga Toshniwal: So… When dosage was zero.
[01:42:41] Durga Toshniwal: X was taking… getting a value of 1.29, right? We calculated. This was the lower part I'm showing you.
[01:42:49] Durga Toshniwal: Similarly, upper one also can be shown. Then Y was coming out to be 1.53, I had shown this calculation here.
[01:43:00] Durga Toshniwal: Right?
[01:43:01] Durga Toshniwal: This was the calculation. When dosage is zero.
[01:43:04] Durga Toshniwal: Then, X is 1.29, and Y is 1.53. This is what I've written here.
[01:43:12] Durga Toshniwal: Then.
[01:43:13] Durga Toshniwal: So, when I try to take these values, these are the X and Ys for what? The activation function.
[01:43:19] Durga Toshniwal: So, when X is equal to 1.29, somewhere here, then Y is 1.53 this, okay, this value.
[01:43:30] Durga Toshniwal: And what is this? This corresponds to the lowest dosage. So this point corresponds to what we are having for the lowest dosage.
[01:43:39] Durga Toshniwal: For the highest dosage, which is 1, so there's nothing higher possible beyond 1 for the dosage, value of X is going to be minus 1.23, which is this.
[01:43:50] Durga Toshniwal: And the corresponding value of Y is 0.257.
[01:43:54] Durga Toshniwal: It just is… Now, this soft plus function is like this. It's… it's going on like this.
[01:44:02] Durga Toshniwal: But for this particular neural network, out of this whole
[01:44:08] Durga Toshniwal: So you can see this is the soft plus function, but we are not using the entire function. We are taking a slice of it, which ranges from
[01:44:18] Durga Toshniwal: What part? I'll write it here. The coordinates which we are choosing is X is equal to 1.038.
[01:44:27] Durga Toshniwal: And Y is equal to 1.341.
[01:44:31] Durga Toshniwal: This is corresponding to dosage 0.
[01:44:36] Durga Toshniwal: So, this is the minimal point.
[01:44:38] Durga Toshniwal: I mean, this corresponds to dosage 0. For dosage 1, it is X is equal to minus 1.23.
[01:44:46] Durga Toshniwal: Sorry. And Y is equal to 0.257.
[01:44:52] Durga Toshniwal: Just a second.
[01:45:03] Durga Toshniwal: Sorry, just a second, please.
[01:45:05] Durga Toshniwal: Yeah. So, so these… these are the points that we are choosing. So, isn't this… this part a slice?
[01:45:14] Durga Toshniwal: This one, this is what we are using, this slice.
[01:45:17] Durga Toshniwal: And the rest of it, we are not using.
[01:45:20] Durga Toshniwal: This is what the slice is that I'm talking about. This slice is actually then
[01:45:26] Durga Toshniwal: Stretched and put like this on the dosage curve. This was the dosage.
[01:45:33] Durga Toshniwal: And this is what is the efficacy.
[01:45:36] Durga Toshniwal: And this is stretched further, and it is put like this.
[01:45:41] Durga Toshniwal: It is also flipped, and then further it is multiplied by a weight of 2.23 and stretched even more.
[01:45:48] Durga Toshniwal: Is it okay?
[01:45:50] Nirav Mehta: Yes, ma'am. Thank you.
[01:45:53] Durga Toshniwal: Yep.
[01:45:55] Durga Toshniwal: Okay, I could see that, Lokesh, you have a question, that in this example, we have two neurons creating the curves. How do we know which neuron, like, we are plotting two neurons?
[01:46:07] Durga Toshniwal: And which one is actually contributing more to the final decision boundary?
[01:46:13] Durga Toshniwal: So… Yeah, so Lokesh, when we are having a full neural network.
[01:46:19] Durga Toshniwal: We are only concerned with the final output.
[01:46:22] Durga Toshniwal: which neuron is contributing, how much, and what is it actually contributing? That may be difficult to infer, like, like the inference I'm showing, because I took a very simple neural network just to illustrate
[01:46:38] Durga Toshniwal: How?
[01:46:39] Durga Toshniwal: how these work, right? But if there were multiple neurons, then there would be so many, slices which might be stretched, or which might be, you know, flipped, or something, something would be happening. So it will be difficult, you really cannot know.
[01:46:57] Durga Toshniwal: And, therefore, it will be difficult to find out which is contributing how much, because if you also recollect, suppose you have a neural network like this.
[01:47:08] Durga Toshniwal: Right?
[01:47:09] Durga Toshniwal: So, this one is getting the inputs after weighted something, sum, and this one will be getting
[01:47:17] Durga Toshniwal: this weighted inputs weight, weighted form, right? So, then final contribution of what are the inputs and what are the weights and all, it's difficult to infer.
[01:47:29] Durga Toshniwal: So, that may not be possible.
[01:47:33] Lokesh R: Okay, ma'am, I'm getting it now, yeah.
[01:47:37] Durga Toshniwal: Okay. Any other questions, anyone?
[01:47:50] Durga Toshniwal: I think,
[01:47:56] Durga Toshniwal: If these are the only questions, then we can break for today.
[01:48:04] Durga Toshniwal: So, I think there's a question by Pavan. What is the soft plus activation function, and how is it different from ReLU or TANH? So, I actually showed it right in the beginning, Pavan.
[01:48:16] Durga Toshniwal: But I'll show it again.
[01:48:18] Durga Toshniwal: For you.
[01:48:21] Durga Toshniwal: So, let me go back on the slides just a second.
[01:48:30] Durga Toshniwal: So, right in the beginning, I had actually shown to you what I…
[01:48:35] Durga Toshniwal: So I have those written already here.
[01:48:38] Durga Toshniwal: Just a sec.
[01:48:58] Durga Toshniwal: Somewhere I wrote it right in the beginning. That's what I'm trying to search.
[01:49:06] Durga Toshniwal: Where has it gone?
[01:49:15] Durga Toshniwal: Otherwise, I'll rewrite if I don't find it.
[01:49:26] Durga Toshniwal: I remember having written it somewhere. Yeah.
[01:49:31] Durga Toshniwal: So, this is the soft plus function, right? I had also written the former name.
[01:49:38] Durga Toshniwal: I think part of my writing is actually gone.
[01:49:42] Durga Toshniwal: So, soft plus, the function is FX is equal to natural log of 1 plus 2E to the power X.
[01:49:51] Durga Toshniwal: Okay, if we talk about RELU, Output is max of…
[01:49:57] Durga Toshniwal: 0 and X. I wrote that all, I don't know why… where it has gone. Similarly, you have also got tan edge, which I had written
[01:50:07] Durga Toshniwal: In the beginning, somewhere.
[01:50:10] Durga Toshniwal: So… And, softmax, and there are so many other functions. The formula I'd given you earlier also.
[01:50:20] Durga Toshniwal: Right, for example, if we talk about sigmoid, Then sigmoid…
[01:50:27] Durga Toshniwal: I had told you, we'll have FX…
[01:50:30] Durga Toshniwal: equal to e to the power X upon 1 plus e to the power X.
[01:50:36] Durga Toshniwal: Relyu, I already told you.
[01:50:38] Durga Toshniwal: So, these are the different activation functions, right?
[01:50:42] Durga Toshniwal: So, I… and this is Softmax. This… sorry, soft plus.
[01:50:49] Durga Toshniwal: Is it okay?
[01:50:57] Durga Toshniwal: Okay, are there any other questions, anyone?
[01:51:04] Lokesh R: Ma'am?
[01:51:05] Durga Toshniwal: Yes?
[01:51:07] Lokesh R: Ma'am, why we are using only the sort plus activation function, ma'am, instead of relo and then?
[01:51:13] Durga Toshniwal: You can use any. I just wanted to illustrate.
[01:51:16] Durga Toshniwal: I just chose SoftPlus for it. You can do the same set of exercise using ReLU, leaky relu, or any other function.
[01:51:24] Durga Toshniwal: Okay, this is just an example.
[01:51:28] Durga Toshniwal: Just to make, yeah, just to make,
[01:51:32] Durga Toshniwal: Some, you know, example which is understandable.
[01:51:37] Durga Toshniwal: That's why I just chose this, but there is no other reason.
[01:51:44] Durga Toshniwal: There's no particular reason for that, okay? You can use anything.
[01:51:49] Lokesh R: Okay, ma'am, thank you.
[01:51:51] Durga Toshniwal: Yeah.
[01:51:57] Durga Toshniwal: So, any other question?
[01:52:00] Durga Toshniwal: Anyone?
[01:52:02] Durga Toshniwal: So you can just think of… just, you can try to visualize in your mind that you have a complex neural network.
[01:52:10] Durga Toshniwal: In which you are having different kinds of activation functions, or same or different, For different layers.
[01:52:18] Durga Toshniwal: And each of these might be sliced.
[01:52:21] Durga Toshniwal: Small slices, big slices, then these slices might be stretched.
[01:52:25] Durga Toshniwal: And then these, might then be flipped, or they might be written like that. And then all those flipped and stretched or shrinked versions of the slices might be superimposed.
[01:52:39] Durga Toshniwal: And what you might get is, you know, the final, the final output.
[01:52:46] Durga Toshniwal: say, that's how you actually get it.
[01:52:51] Durga Toshniwal: So, okay.
[01:52:55] Durga Toshniwal: So that's it, and I can see, comments by Sivanch.
[01:53:00] Durga Toshniwal: -Oh.
[01:53:01] Durga Toshniwal: Okay, then?
[01:53:03] Durga Toshniwal: If there are no further questions, then we can stop here.
[01:53:09] Shivansh Sharma: Okay.
[01:53:10] Durga Toshniwal: Thank you, thank you all, have a great night.