# 07 2026-05-31 T5 Theory Contd course: Module 4 — Generative AI & LLMs module: Module-4-Generative-AI-LLMs date: 2026-05-31 type: transcript video_url: https://personal-learn.armco.dev/files/_Recordings/Module-4-Generative-AI-LLMs/07_2026-05-31_T5_Theory_Contd.mp4 --- [00:13:49] So, uh, good morning. Is, uh, Ms. Simran here? [00:14:00] Hello, hi, this is Akanksha [00:14:01] Okay, Mr. Kansha, are you, uh… are you the batch manager? [00:14:07] Yes. [00:14:08] Actually, I need some contact details for future staff. Could you please share the [00:14:13] email ID through which [00:14:15] Like, regular and, uh… [00:14:18] regular communication can be received. [00:14:21] Yeah, sure. [00:14:41] Do you have any specific queries [00:14:44] Rachel? [00:14:45] Yeah, regarding attendance and projects, uh, actually, I was… [00:14:48] Not in a situation to attend it to some, uh… [00:14:53] I'm very sorry for that [00:14:54] medical situation, so that is why I wanted to clear… [00:14:57] Like, provide some sort of context behind my attendance. [00:15:03] Okay. [00:15:04] And… [00:15:05] So, have you… have you been marked incorrectly in any of the classes? [00:15:08] No, I have not been marked incorrectly, but I need to provide some justification for the [00:15:14] absentees, absent dates. [00:15:15] Okay. [00:15:17] Sure, sure, sure. [00:15:18] You can raise a ticket also, if you raise the ticket, it will, directly come to the taxi, and we will solve… we can solve it. Have you raised any tickets? [00:15:28] No, I have not tried the ticketing system yet. [00:15:32] Yeah, to… please, I request you to raise a ticket regarding that [00:15:36] On the new, uh, LMS, right? [00:15:42] Yes, on taxi. [00:15:45] Yes. [00:15:46] on Tixie. Uh, I'm not sure if I'm, uh, if I have access to that portal. [00:15:52] Okay. [00:15:58] Yeah, hello, uh… [00:16:00] Good morning, Mom. [00:16:01] Good morning, everyone. [00:16:04] Uh, so I think there was some concern regarding, uh, learner, uh, about some of… somebody about change in timing, so if you all recollect right from the… [00:16:14] First turn, we had pointed out that there could be some changes in the timings. These are temporary only. [00:16:21] Uh, because of the availability of, uh, the expert or of myself. [00:16:26] Uh, then there might be changes. However, once your, uh… [00:16:30] Uh, this encoder-decoder model is done, that is, the T5 is done. [00:16:35] That is, T5 is done, uh, in the… in this particular code. [00:16:41] Uh, that is T5 theory, and T5, uh… [00:16:46] So, we are, uh, talking about encoder-only, decoder-only, and encoder-decoder models, and their hands-on will be there. [00:16:53] So, currently, you're finished, uh, the hands-on for the first one, which is the encoder-only. [00:16:58] That is part. Uh, after you would have finished the 34… [00:17:02] GPT, which is a decoder only, and encoder-decoder, which is T5 model. [00:17:09] You would probably finish it this weekend, or, I mean, today, or maybe the next turn, Max. [00:17:15] Then you'll have the hands-on for these two. [00:17:18] And after that, uh, the next of the portion, your timing will get, uh, uh, again, [00:17:23] Uh, probably reverted to the same, or whatever best suitable, okay? So, this is just a temporary… [00:17:32] Shift in timings due to availability purposes. [00:17:35] Okay, so this is not a permanent change, it is only for some time. [00:17:40] Uh, uh, so… [00:17:44] I think, uh, I'll just request, uh, Mr. Rohit, if he's there. I think he must have joined already. [00:17:51] Let me just check it. [00:17:54] So, Mr. Somi Sovani, are you there? [00:18:02] I think you must be joining. [00:18:05] Shortly. [00:18:07] Uh, he hasn't joined yet. [00:18:09] So once he joins, [00:18:11] Uh, we can, uh, start the session. [00:18:14] The slides for your last turn have been posted on LMS. [00:18:22] For those who want to refer to it, okay? [00:18:27] So I'm just, uh, placing a call to the expert to join. [00:18:32] In the meantime, um… [00:18:35] After you would have done your theory on the last model that is encoder-decoder model, [00:18:43] Then, I will take you through the project, uh, capstone project details. [00:18:50] what all you are supposed to submit, how you are supposed to submit, so we'll have a session on that after… [00:18:55] Uh, this all is done. [00:18:57] Uh, so… [00:19:01] Mr. Somia, are you there already now? [00:19:06] Any other questions, anyone, in the meantime? On anything? [00:19:12] Uh, good morning, ma'am. Actually, uh… [00:19:15] Good morning. [00:19:16] So, there have been some emails from Fletcher's team about some career assistance. [00:19:21] Mm-hmm. [00:19:22] And I was, uh, wondering, like, for whom is it applicable? Is it applicable for… [00:19:26] Only the people who have appeared for the diagnostic test, or is it applicable for every student who is enrolled in this program? [00:19:33] Or is it… okay. [00:19:34] Uh… yeah, carry on. [00:19:36] Yeah, that's my main query, how, like, how does it work out? [00:19:41] This placement assistance. [00:19:42] So, okay, so, uh, I think, uh, the Futurance team will be able to address this better way. [00:19:49] Future and stream, can you please come in? [00:19:54] Yes, ma'am, I'm here. [00:19:55] Yeah, can you please address the query? The… [00:19:58] Career assistance, uh, is it open to everyone, or…? [00:20:02] Yes. [00:20:03] Yes, ma'am, it is open to everyone [00:20:05] Okay. [00:20:07] Okay, fine. So, I think that answers, yeah. [00:20:14] Ma'am, what will be the, time given for the project submission for the capstone project [00:20:21] Yeah, yeah, so that will be around, uh… [00:20:24] More than a month, you will get, but actually, your projects and all will be finalized, uh… [00:20:30] Halfway through, so… [00:20:32] Uh, actually, the whole team gets around… [00:20:37] 4 months to do the project. [00:20:38] But after you complete the theory, [00:20:41] A full one month, you get… [00:20:44] to only work on the project. Other than that, once your team is finalized and your project topic is finalized, then you get [00:20:52] time anyway, and complete free time is 1 month, slightly more than a month that you will get. [00:20:59] Okay. [00:21:04] Mm-hmm. [00:21:05] And the projector… sorry, ma'am, to come between, uh, like, the projects, uh, are decided by, like, us, or, uh, like, [00:21:10] From your side, the suggestions will be given, and… [00:21:14] Like, and when it will be, uh, when we have to start it. [00:21:18] So, as I said that once your generative AI, I mean, text LLM portion gets over, [00:21:26] The theory and hands-on. [00:21:28] Then I will be introducing to, uh, the project, about the project to you. [00:21:34] And the project topics, I can give you some… [00:21:37] Um, suggestive list of topics. [00:21:41] And you can choose your own topic also, or you can choose any other topic for that matter, okay? [00:21:48] But I… I will take a separate session for all this, so presently, you need not worry. [00:21:54] Okay. [00:21:55] I will address all your queries related to project in that separate session. So, let us start with the portion for today. I'd like to invite Mr. Soum. [00:22:03] Mr. Samuel, I'll request you to enable your video. [00:22:09] Yeah. And to take on from your welcome, Mr. Sameer, please, um… [00:22:14] take over from here to cover the… [00:22:17] list of the stuff. Thank you. [00:22:20] Thank you, ma'am. [00:22:23] Let's get started. Just give mea sec. [00:22:24] Uh, Mr. Song, your, uh, I think voice is coming a little low. Can you please increase the volume? [00:22:31] Just a second, ma'am. [00:22:35] I think, I am having some issues. Just, just a second, sorry [00:22:38] Yeah, but… yeah, I think with this, it may be better. [00:22:59] Is it any better now? Or is it still not [00:23:06] Okay. [00:23:10] So, let me get started. [00:23:21] So let… let me get started with today's lecture. I'll set it up. I actually ran into some setup issues today [00:23:28] So, just give me a second. [00:23:40] So [00:23:41] Yeah, that was why [00:23:54] So, if I'm asking [00:24:03] So, I hope my screen is visible to everybody, and I am audible now [00:24:12] Am I? [00:24:13] Yes, yes. [00:24:15] Okay. [00:24:16] Sure. [00:24:18] Let me just start with this one. [00:24:22] So if I… So yesterday, we had discussed about [00:24:29] the T5 model, and we had taken… let's take a quick overview over it. So, this is the architecture of the BERT and the GPT models that we have seen, and this is putting it together. So, this is a encoder-only architecture. This is a decoder-only architecture, where [00:24:45] It uses autoregressive generation to generate one token after another [00:24:50] Here we have a BERT model, where it takes all the tokens and looks at both past and future context to generate a contextualized open embedding. [00:24:59] T5 uses both an encoder and a decoder, and so the encoder model here is the bottom part of this, the encoder stack. The encoder stack takes a input, it looks, it creates a contextualized embedding for the entire input [00:25:14] Basically, it looks at both the past and the future context, which is bi-directional attention, as we have already looked at. So it takes the bidirectional attention and [00:25:24] It tries to produce a set of in-context or contextualized embeddings for the input. And once that is done, this T1, T2 to T, and this one is the token embeddings, or the contextualized token embeddings for all the input tokens given to the T5 models encoder stack [00:25:39] The… these contextualized open embeddings are then used by the decoder. So the decoder takes an input of the decoder BOS or beginning of sequence token [00:25:51] Along with the encoder stacks outputs. So these inject the context about what the decoder needs to generate into the decoder. So this this mechanism here is cross attention. And as we already looked at cross attention needs to be bidirectional [00:26:06] Because [00:26:07] That is… the decoder needs to know what is the context, or in what context it needs to generate the future tokens [00:26:15] So it is very important for the cross attention to be bidirectional, because if it is masked, then the decoder will not have the entire context. [00:26:24] However, the decoder's own self-attention needs to be [00:26:28] But needs to be masked, because the decoder cannot look at its own future like while decoding tokens in an autoregressive manner, the decoder cannot look at the future tokens, and hence it needs to be masked here [00:26:40] So this is an overview of the T5 architecture. And as we had looked at the training task that T5 uses is span corruption, where 15% of the tokens of every input sequence are dropped. [00:26:52] And all of the consecutive drop tokens are replaced by a single sentinel token. So, for example, here, thank you for inviting me to your party for inviting is dropped and replaced with a sentinel token X. [00:27:03] Last is replaced by the Sentinel to open wide, and this is the input to the model. The output that is expected from the T5 model for the span corruption task is [00:27:12] The Sentinel token, followed by what the model thinks should be replacing the Sentinel token. [00:27:18] And at the end, another central token. So for this x for inviting Y last another sentinel token Z, which signals that all the sentinel tokens have been decoded [00:27:31] So, this is how the decoder is going to be trained. So, for example, the input sequence thank you x, meet your party y weak is tokenized and these tokens are then sent into the encoder. The encoder produces the encodings for the contextual encodings for all of these tokens. So one embedding per token [00:27:51] These set of embeddings are fed into the decoder, along with the decoder start token. The output that is expected from the decoder is X for the decoder start token, and then we append X to the decoder start token. So we give the second sequence as decoder start token in X, and the expected output is 4 [00:28:07] And then we append 4, and the expected output for the entire sequence now is inviting, and so on and so forth. So the loss is the loss function that is used is cross entropy loss, where cross-entropy loss basically is the probability of the ih token being the next token as predicted by the model, and as [00:28:26] It is, actually. So P of I is the actual probability, which should be 1 for the expected token and 0 for all other tokens, and Q of i are the model outputs where Q is the probability of the ih token being the next token across the across all the tokens in the model's vocabulary [00:28:42] So this is the prostropic loss. And this loss is calculated for each of the outputs from the model and back propagated through the entire decoder and encoder stack [00:28:53] And this is how the unsupervised training for DT5 model is done. So as we had already looked at, T5 model and all other models we have seen so far use transfer learning. So first, there is an unsupervised pre-training, and then there is a discriminative fine-tuning. So this fine-tuning is always going to be supervised. We have a labeled data set [00:29:11] So for example, we could have a text classification or text summarization or question answering or NMD data set where we have a set of inputs and the expected output or the label for that particular [00:29:23] input. So, in text classification, we would want to classify, like, for example, it might be a sentiment analysis problem, where we want to classify the text as either positive sentiment, negative sentiment, or neutral sentiment, or in-text summarization. We could give an entire Wikipedia article, or [00:29:39] A particular news article and ask the model to generate a small abstract for it. [00:29:45] But then there is also the question answering task and NMD. So the T5 model was fine-tuned over all of these tasks. The general language understanding benchmarks, test classification, the CNN and Daily Mail data sets, text summarization and Stanford question answering data set and the WMT English to French, English to German, and English to Romanian data sets [00:30:03] We had already looked at how the training was done and how the fine-tuning was done. The training over the entire C4 data set, which was the data set that was used for the unsupervised pre-training, which was about 750 GB in size [00:30:19] was done over 2.19 steps, and it processed a total of about 34 billion tokens. After this pre-training was done, then the fine-tuning step used the, the base model that was trained here, and it fine-tuned them over the training datasets of each of these benchmarks [00:30:35] And then it was validated over the validation datasets of each of the benchmarks. So, for example, if we take the Stanford question answering data set, the model was trained for about 2 par 18 steps or fine tuned for about 2 par 18 steps, processing about 17 billion tokens across the Stanford question answering data sets [00:30:52] Training dataset. And after that, it was the model that was being trained. As it was being trained was stored every 5,000 steps. So for these 2 par 18 steps, or about 2 lakh 40,000 steps, a copy of the model was stored every 5,000 steps [00:31:07] And when we get to these, like, after we have saved the model for all of the 5,000 steps, we run the validation dataset over all of the saved models, and whichever model gives us the best possible output, that is what is chosen to be the published model for that particular task. [00:31:24] And this is how T5 was trained. So looking at the architecture for T5, we can see that it has a BERT-like encoder, which is an auto encoding encoder where we give it a set of inputs and there is an input embedding, and we give it [00:31:39] We don't give it any positional encodings here, and we had discussed that this is because it uses relative image, where the relative positions I attended to as a part of the multi-head attention itself, and we do not need to inject any particular trained [00:31:54] positional encodings. Bird here would have injected sinusoidal position encodings while Gpt 2 trains a set of positional encodings on itsown while the model is being trained, while for the T 5 we train a set of values that are used as used to attend to relative positions during the [00:32:13] During the models attention computation. And it is the same for the decoder as well. The key component here being the cross attention, where the encoders and output is fed into the decoder and [00:32:25] The decoder then looks at the encoder's output, and the decoder beginning sequence tokens, whichever sequence we have fed into it, and it computes whatever is going to be next in this particular sequence that whatever is the output from the decoder is then appended to the input, and it runs. So for any given input, the encoder runs once [00:32:44] And produces the, the per token encodings which contain the context. [00:32:51] of the input. The decoder looks at the context of the input and runs across attention and injects that context into the current decoding sequence. It also has a self-attention by which means it understands [00:33:04] its own decoded sequence. It understands what it should be decoding next, based on the cross-attention, and it also understands what it has already decoded, or what sequence it is already processing using the self-attention [00:33:16] The self-attention is masked, because the decoder cannot look into its own future tokens. It needs to have context only from the past decoded sequence, not from any of the future tokens that might appear. So we need to mask this self-attention. However, the cross-attention need not be masked [00:33:31] We need the entire context in order to determine what to decode next. [00:33:35] And this softmax with linear layer here, this is the language modeling head, which is very similar to what we had looked at during GPT-2. The GPT-2 model also has its own language modeling head [00:33:45] Over here, and this language modeling head is then used and [00:33:49] This is the this language modeling head then gives us the set of logits across all of the vocabulary of the probabilities of every token being the next token in that sequence. [00:34:01] So, the decoder's LM head works in a very similar way. It takes the decoder output and it predicts across all of the vocabulary tokens which token is going to be next. The probability of each token being next. And we sample from that and get the next token [00:34:16] We append that to our current decoder input, and we run the decoder again. [00:34:21] So [00:34:22] We had also looked at the input tokenization and looked at the unigram algorithm with sentence piece [00:34:29] As the tokenization algorithm [00:34:31] For the T5 tokenizer. Sentence fees is very different from, the tokenizer that we have looked so far, like, looked at so far, in that BPE and WPM use a very small vocabulary. They just start with the initial alphabets, and they [00:34:46] learn a set of merges that is they at every step during the tokenizers training, they learn how they can merge these smaller tokens or the constituent alphabet into larger tokens [00:34:58] So that they are able to express all of the words in the training corpus or training data set. [00:35:05] However, for sentence piece, it is the other way around. We start with a huge initial vocabulary, and then we try to reduce the vocabulary down into the size that is required. Also, we had looked at that we had looked at BPE and WPM, and they always start by taking the words [00:35:22] One by one. BP and WPM both assume that the words are going to be separated by spaces. Sentence fees [00:35:30] takes the entire input as a raw set of stream. It does not need to do a pre-tokenization in that. It does not need to separate the words out, and then tokenize each of the words individually. [00:35:40] What is… what it does is it takes the entire stream of input. So sentence piece is very important in that it addresses the fact that not all tokens or all languages are going to be using spaces to separate out words. [00:35:54] languages like Mandarin, Chinese, and Japanese don't really use any of these spaces to separate words. So sentence piece address is a very, very commonly phased problem. [00:36:05] And also, the other part is that it starts with a huge initial vocabulary that we had looked at during this example. So, for example, for BP or WPM, we would only have the constituent F of alphabets as the initial vocabulary, so it would consist of, like, the letters H, U, G, P [00:36:24] And B and S as is initial vocabulary. However, for the sentence piece model using the unigram tokenization algorithm, we start with all possible tokens. Like, for example, HU could be a possible token because HU occurs here. And similarly HUG occurs here, so it is also part of the initial vocabulary [00:36:44] This huge vocabulary is then shrunk down to whatever the vocabulary size that is required is. So for example, if we take that the vocabulary size here is 10, then it starts out with more than 10 tokens and it keeps dropping tokens [00:36:58] At every stage to shrink the vocabulary down to the expected vocabulary size. And how does it drop tokens? Well, it calculates a loss over every possible tokenization that is for all of the words in the corpus. [00:37:13] And this cumulative loss is then calculated after dropping the token as well. Once the losses are calculated across the candidate tokens to be dropped, we can then accumulate we can then look at the delta or the change in loss when each token is dropped. [00:37:30] And based on that, we decide which token to drop. For example, in this in this particular example [00:37:35] We calculate the loss over the tokenization of the basic vocabulary hug, pug, pun, bun, and hugs using the initial vocabulary that we had derived over here as the probability of the word times the negative log of the probability of the tokenization [00:37:52] The probability of the tokenization is simply calculated as the probability of each of the tokens occurring [00:37:57] in the in the data set [00:37:59] So once we calculate the loss over the basic vocabulary, we then look at candidate tokens that could be robbed. For example, if we wanted to drop PU from the vocabulary, the loss doesn't change, so the delta is 0. However, if we wanted to drop the token HUG from the vocabulary, then the loss changes [00:38:16] And it changes by 23.5. And similarly, we look at all of the candidate tokens that we want to drop, and we see that none of the tokens, none of the other tokens give us a delta that is better than 0. So no change in the loss. So that is why we drop the token [00:38:33] Sorry, we dropped it open PU, and we do not drop the token HUG, because that dropping the HUG token would increase the loss by 23.5. So then we recompute the frequencies after dropping the tokens, and then we recompute the tokenizations, the scores for the tokenizations and the loss [00:38:52] And this process starts over again. And this is how SPM shrinks its vocabulary down into the expected vocabulary size. [00:39:00] We had also looked at how the tokenization algorithm is then used [00:39:06] Using the B2B algorithm for tokenizing any input sequence. So we had looked at this GIF [00:39:13] Here, and we had seen that, it calculates the loss at every given, at every given token. So it starts with [00:39:21] For example, here [00:39:23] It, starts with the character underscore, and then it looks ahead and sees all possible tokens that are part of the vocabulary and calculates the score. The initial scores are all infinity, so whenever it finds a better score, it updates these [00:39:36] So, for example, it [00:39:38] Starting from underscore, we tokenized till TELL, which is a part of the vocabulary. But [00:39:45] If we go any further, the tokens, there are no further tokens in the vocabulary. For example, underscore T-E-L-L underscore is not a part of the vocabulary. Similarly, underscore T-E-L-L underscore M is not a part of the vocabulary, and so on and so forth [00:39:57] So we can only update the scores till 10, starting from underscore. Then we start from T, and look at every possible [00:40:04] open, from T, and if there are any tokens, we update the scores for that. And similarly, we continue so on and so forth. We look at… basically, we are looking at all of the possible tokenizations of this stream. Tell me about Qualcomm. [00:40:17] And then we are looking at the best scores that are possible. So, once we have found that the best scores are minus 11.6, starting at underscore, and starting from this character underscore, it is 19.1, starting from this, it is 26.2, and starting from this, it is 39.6. [00:40:34] There are no better possible scores across all of these rows. And we also record what is the best possible score that we can achieve from the current token. So we record this minus 39.6, we record where is the beginning of this token that allows us to achieve this loss. In our case. [00:40:51] The beginning of this token is underscore, so this underscore right here, right? So anything… so we remember that from underscore to M is, like, the best score here is one token, and then we backtrack our way similarly. So once we get here [00:41:07] Then we see, okay, what was the best possible score before this? So for T, the best possible score is minus 26.2, and we remember where did we get this from? So, looking at this, we can see, okay, the minus 29.2 is from this underscore right here. [00:41:19] So underscore about is another token and so on and so forth. We keep backtracking, and that's how we get the tokenization underscore tell as one token underscore me as one token underscore about as the third token and underscore Qualcomm as the fourth token [00:41:34] Here, these underscore basic, underscore characters basically represent spaces [00:41:38] So [00:41:41] This was the example that we looked at yesterday, and how the tokenization would be done for a particular input. Once the tokenization is done, we then derive the input representation, which in our case just consists of the token embeddings. We look at the the the embeddings for each of the tokens in the vocabulary and use that as a lookup table to [00:41:58] get the particular embeddings for the tokens that we have just tokenized. And that gives us the input representation for the T5 model. The T5 does not add any positional encoding to the token embeddings. Instead, it uses a relative position embedding. So [00:42:13] We now take a look at what is relative position and how it is used. So, before we do that, is there any doubt, or anybody has any questions regarding what we had looked at so far? [00:42:35] Okay, so if there are no questions, we can continue. Just a second. I just need to [00:42:58] I see that somebody has raised a hand. Is that Neeraj, do you have a question? [00:43:03] Nothing. Uh, yes, yes, I have one query. [00:43:06] Uh, as you say that it is 15% is dropped, uh, tokens, and it is, uh, done by randomly, or it is some… [00:43:15] No? [00:43:16] Specific rules are there. [00:43:17] So, it is not done randomly, as we looked at, we calculate the loss over the tokenization for the given corpus. So, for example, here we have this corpus as the corpus that we want to learn a tokenizer for. And we start with the initial vocabulary, which is all possible sequence of tokens that could be there, or every substring that is there here in this particular vocabulary [00:43:39] We then calculate the, the tokenization scores for each of the possible tokenizations. For example, if we take the word per [00:43:47] And this vocabulary, we have 3 possible ways of tokenizing it. We can tokenize it as singular characters, P, U, and G. [00:43:53] Or [00:43:54] We could tokenize it as characters P, U, and G. So this is one token, this is another one, or consequently, we could also have it as P and UG, because both of those tokens are a part of our initial vocabulary. [00:44:07] So P and Eugen [00:44:08] We calculated a score for all possible tokenizations. So, for example, the score for, tokenization PUG is probability of the token P times probability of the token U times probability of the token G [00:44:19] So, what is the probability? Well, it's basically the number of times P occurs divided by the sum of probabilities across… the sum of the number of times all these tokens appear. So the sum of the number of times all of the tokens appear is 210. So that's constant for all of these, and the probability of P occurring is 17, because P occurs 5 times here [00:44:38] B occurs 12 times here, so total 17, and so we get the score here. And for PUG, we, we take the score as 0.007, and again, we compute the score for P and UG as 0.007. And similarly, we do it… so, right now, there are three possible tokenizations. We take the one that gives us the best score [00:44:55] When two of them give the best score, we take the one that occurs first. So since PUG occurred first, we take that one as the tokenization. So this is the tokenization, along with all of the scores. Now, we want to drop tokens. Now, the base premise is that we want to drop those tokens that do not affect the loss over this entire corpus [00:45:12] The loss is simply the probability of this word. So the word HUG, hug, occurs 10 times. So the probability of that word in the corpus times the negative log of these four of that tokenization, which is here, 0.0714. And we do that for all of these tokens, so we get a loss of 168.9 [00:45:30] 169.8, sorry. So this is the base loss. We do not want this loss to increase, right? So what we want to do is we want to drop the tokens in a way such that this loss is like we get to our target vocabulary and at the same time this loss does not become very high [00:45:46] So, for example, if we take the word, the token PU, we could tokenize it as P and UG as well. And the score for that tokenization was exactly the same, 0.00777 that we have here. [00:46:00] So since we can drop the token PU and it would not lead to any increase in loss, the delta is zero, we dropped. But why do we not drop HUG, for example, because if we were to drop HUG, the loss increases to 193.3. We don't want the loss to increase. We do not want to lose representability. The loss basically the loss of the tokenization [00:46:19] gives us the representability of that organization, or using the particular vocab that we are currently using, how well can I actually represent my input corpus? So that is why we need to ensure that the loss does not increase much further, and at the same time, we want to get down to our target vocabulary, because we start from a very huge initial vocabulary [00:46:41] And so this is how the tokens are dropped. It's not dropped randomly [00:46:46] Yeah, got it, thank you. [00:46:48] Does… yeah [00:46:53] Sorry, Neeraj, you were saying something then [00:46:56] Ah, yes, uh, got it. [00:46:59] Yeah. [00:47:00] Thank you. Thank you. [00:47:02] Yeah, so basically, this is how we want to get down to our target vocab 3. So in all previous examples we have taken, let's say, our target vocabulary to be 10. We would right now drop TU and then calculate the loss over this tokenizations for the words hug, bug, pun, bun, and hugs again [00:47:18] We would calculate the scores and the losses, and then we would look at other candidates that we can drop without increasing the score, or increasing the score as minimally as possible. [00:47:27] So [00:47:29] This was all about the input tokenization. And now we get to relative position self attention. What does what does relative position mean? Well, so far we have looked at positional encodings as absolute position. So, for example, in BERT or Gpt 2, we have a position encoding for every absolute position of the context that the model understands [00:47:48] So, for example, if the context was 1024, we would have 1024 to 1 embeddings indexed from 0 to 1023, and each of these numbers represents a position. So 0 is the first position, one is the next position, and all of these are absolute positions. [00:48:02] So [00:48:05] The model is basically learning absolute positions as an embedding or representing each of these positions in a vectorized format. [00:48:15] Where, T5 learns relative positions instead. So what is a relative position? Okay, given the query, for example, if I take the query token to be the sixth token or given a query as six [00:48:26] I look at any token preceding it as the relative distance, a negative relative distance from it, and any token after it as a positive relative distance from it. So if I were having a query token 6, and I were looking at 9 as the key token, then what is the relative distance between these 2? Okay, the distance is 3. So 9 is 3 tokens away from 6 [00:48:48] Or one is 5 tokens behind six. So that is how the relative position works. Relative positions look at how far each of the tokens are from each other while computing the Q times K transpose, basically the QK dot products [00:49:05] of all of the query and key vectors. So if I take the query vector for the six token, and if I ever to have a relative attention of it with respect to the key vector of the first token, so if I were using the first token as my key and sixth token as my query. [00:49:21] I would say, okay, I need to inject a relative position of 5 tokens behind. Or if I were using the key for the ninth token, I would say, okay, I have to inject the relative position of about three tokens starting from the current coin ahead of it [00:49:36] So this is, in a sense, what the model is trying to achieve. It is trying to look at all of the queries relative to all of the keys, and how far away they lie. Instead of learning absolute positions, and then looking at, okay, so if… if this were to be done in BER [00:49:51] The query vector for the six token would have the positional information about the sixth token, and then the key vector for the first token would have the positional information regarding the first token. So when we compute the QK dot product for these two vectors. [00:50:08] They would inherently have the relative distance as 5. They would… it would basically mean that the model understands, okay, this is a key for the first token, and this is a key for the sixth token, so they are about 5 tokens apart, and that is what the attention would then try to pay, like, that is what the attention computation would then try to consume or understand [00:50:31] And, I see there's a query in the chat. Is it for the encoder side of T5? No, it is for both. It is for both the encoder and the decoder. So if we look at the architecture here [00:50:42] We can see there is no positional encoding for encoder or for decode like BERT uses a positional encoding for the encoder model here, and GP2 also has a positional encoding that is injected into the decoder only architecture here. We see that over here we don't have any positional encodings that are injected to the input tokens of either the encoder or the decoder. All of it is handled during the self-attention of the encoders and the decoders [00:51:08] So coming back to relative attention [00:51:13] So, this is the relative positions for [00:51:22] So, in one sentence, the token might be at second position, in another, it might be at fourth position. Sorry, I'm… what was the query? Sushri? [00:51:33] So, in, in, absolute position, we have the exact position where we are, which we are keeping. But in relative position, it's the difference between that position and, the query or the key positions [00:51:49] So, if, suppose, in one sentence, the token is coming at the second position, and in another sentence is coming at 4th or fifth position. So how, in that case, the relative positioning will work? Because it's like [00:52:03] So, so when the… so when these sentences are going to be consumed one at a time, or for example, these sequences of tokens are going to be consumed one at a time [00:52:13] So when we are consuming the sequence of tokens which has the token occurring at position 2. [00:52:19] the relative attention, so the relative attention first of the relative attention for every token with respect to another token is going to be different. So, for example, the token here is occurring at the second position, but the relative distance of this token from the third token is just [00:52:34] But it is 5 with respect to the 6 token, and behind, or similarly, ahead and behind. So, the distance of this query, this key from, this query, for example, is one ahead. That is the relative distance here. So, in essence, let me take this [00:52:50] This next one here. So this is how the Q times k transpose relative attentions would look like. So these are the query positions for all the tokens, and these are the key positions. So when you are computing the query and key [00:53:05] similarities, or Q times k dot product for each of the query and key vectors, that is the time that you inject the position, the relative position into them. So, the relative position of the 0th query, like, with respect to these keys, would be 0, 1, 2, and so on and so forth [00:53:20] And the relative position of this query with respect to the 0th key, because the query is second, so it's basically the redivision of this query with respect to the previous token, which is going to be minus 1 here. This would be 0, and so on and so forth. So when you're processing these two sequences, this entire computation would be done again [00:53:37] And so the relative positions for each of the tokens would change again, and that is how you would get the positional information for all of those tokens relative to each other for every particular computation of the attention [00:53:52] Or for every token that you're computing the attention for. [00:53:55] Does that answer your question? [00:53:58] Yeah, but in… while doing positional encoding, we add that to the vector. Here, we are not adding anything to the vector. So, how the model will know which position it was [00:54:10] Yeah, so coming to that, we will be adding a bias, so to speak. So, we will be adding the position, the positional, the relative positional encodings to this entire q times k transpose matrix during the computation of the attention. We'll get to that [00:54:26] So right now, we were just looking at what relative positions are and how we can think of them compared to when we are looking at a query and a key matrix. So when we are computing the attention, we take the dot product of each query with every key. And in a sense, what the absolute positional encodings are doing [00:54:44] is very similar to relative position, right? Because when you are computing the queries and the keys, the Q times K dot products, you already have the positions injected right? But if you don't already have the positions injected, you can simply just add another relative position encoding to this particular [00:55:05] Value. This predictor, the dot product value, the output of that dot product. You could just add another value to it, and that injects the positional information. How that is injected we'll be looking at, but right now this is like how the value that is to be injected is calculated [00:55:21] is this. So [00:55:23] Okay. Okay. [00:55:25] Right? [00:55:27] So here we have the relative position that we are calculating of every query with respect to every key, because we need to know, like, when… what we are paying attention to and what position what we are paying attention to is in with respect to us. So if the query here at the sixth token [00:55:45] If this sixth token is paying some attention to, let's say, the first token, it needs to know, okay, the first token is five tokens away from me [00:55:51] So it needs that that information is going to be injected via like added into the dot product as a bias. Over here with absolute positions, that would already be computed because the q and k vectors would already have their respective positional encodings added to them, right? They would already know about their positions. So when we compute the dot product [00:56:09] That positional information would already be a part of that dot product. Here, the positional encoding is injected into that dot product after the dot product is computed. [00:56:18] So, the relative positions for every query with respect to all the other keys in the sequence [00:56:23] would look something like this. So for the 0th token, the query is corresponding to the 0th token with respect to all of the other key positions, like, for example, let's say there are a sequence of 10 tokens, then the relative distance of the 0th query token would be 0 to 9 [00:56:38] The relative distance of the first query token would be minus one to eight, and so on and so forth. For the last token. [00:56:44] The later positions would be minus 9, minus 8 to 0, because it is the query, the 10th tokens query is at a relative position of 0 from the 10th token's key because it's the same token. [00:56:56] So, on the diagonal, we always have relative distance to be 0. And anything to the future [00:57:02] of our current token is going to be a positive distance or positive relative distance away from it, and anything to the past of our token is going to be a negative relative distance away from it, right? Because we are looking at the past. [00:57:13] So how does relative position self-attention then inject this positional inject a value corresponding to these positions into it? Well, the model learns of a particular set of embeddings for relative positions [00:57:30] The relative position is defined as key minus query position. That is the distance between the key and query tokens. So the relative distances are split into two categories, close and long distances. So the model basically tries because if we were to learn a single embedding for all of these positions, we would need to learn a lot of embeddings. Let's say there were 10 [00:57:47] 24 relative positions, because the context length of the model is 1024. We would have to learn 1024 separate encodings for past and for the future. So that would basically mean we have to learn 2048 values in order to be able to differentiate between the past and the future tokens [00:58:03] In the encoder, for example. So instead of this, the model uses something called bucketing. It defines a set number of buckets and then tries to bucket these relative distances into the set number of buckets. There could be different number of relative distances falling into a particular bucket [00:58:22] But there is… so [00:58:24] for… in order to bucket it, it… it defines close and long distances. For closed distances, where the model says, okay, if the tokens arevery close together, then they need their own positional encodings, because for every single closed distance, let's say the model defines 5 as its close distance [00:58:40] So it would say any token falling 5 behind or 5 ahead should get its own positional encoding, or it should get a particular value for it. But this is because it is more important to pay attention to tokens that are in close proximity. So the context of any that any token gets from any other token [00:59:00] would be the highest when it is in very close proximity with that other token. So that is how the English language or any other language works, right? So when you're speaking in a sentence, the contextual meaning of any word in that sentence is generally defined by a few words ahead or a few words behind [00:59:15] So that context keeps shifting across the sentence. So anything that is in close proximity should be getting really good attention, because [00:59:24] It needs to be able to differentiate between each and every position in close proximity, whereas the farther out we get, we do not care how far ahead it is, or how far ahead, like, it might be, let's say, 100 words ahead, or 110 words ahead. We might say, okay, this entire sequence, 100 words or 110 words ahead [00:59:41] can receive equal amounts of attention, because it's just far away. So that is the intuition behind [00:59:47] defining close and long distances. So the model defines close and long distances for every particular closed distance. We get half the buckets for closed distances and half the buckets for long distances. So, for example, in this model, we have 32 buckets. [01:00:01] So for the encoder model, these 32 buckets will be split into 16 buckets for the past and 16 buckets for the future, and within these 16 buckets for the past, we would say any distance of 8 up to 8, because 16 buckets, right? So by 16 by 2. [01:00:16] So the model defines any distance, any relative distance up to 8 in the past as a closed distance, and any relative distance more than 8 in the past as a long distance. Similarly, in the future as well. It says any token appearing up to 8 tokens away from the current token has its own closed distances [01:00:32] And any token appearing more than 8 tokens away from the current token is a long distance. So it learns specific positional encodings for each of these buckets. So in essence, what it is doing is for all close distances, 1, 2, 3, 2, 8 [01:00:49] It learns a specific value for those close distances. And as we get farther out, we simply start [01:00:56] we simply start bucketing these together. So initially, let's say the model buckets 8, 9, 10, and 11 into a single bucket, and learns one positional encoding for this bucket, and then it takes 12, 13, 14, 15, 16 into another bucket, and so on and so forth, till the max distance. The max distance that is defined by this model is 128 tokens [01:01:16] So how do we bucket it? It uses this logarithmic formula that is defined here and buckets by 2 times 1 plus log of relative position, the current relative position keep t minus query position into two by n buckets [01:01:28] Divided by log of max distance into 2 by n buckets [01:01:32] And this is the formula that it uses to define how many tokens are going to be bucketed together. [01:01:38] So given any relative position, what is the bucket that this falls in is calculated by this formula for any particular relative position that is less than n buckets by 2. It gets its own [01:01:51] bucket. So, each token from 1 to 8 gets its own bucket. So, half the buckets have exactly one token, and the remaining half have more and more tokens falling into them as they, you know, as the relative distance increases. So, graphically, we can look at [01:02:07] As follows [01:02:08] For the T5 encoder model, we take the 32 buckets and split into 16 buckets for the future tokens and 16 buckets for the past tokens. Within these 16 buckets, we say eight is the eight distances are reserved for closed distances or eight buckets are going to be reserved for closed distances. So from 0 to 8 [01:02:27] Each bucket has a singular distance, and it has its own unique positional encoding that is learned right [01:02:34] And for every bucket after this right from let's say the ninth bucket, it takes the token token or the position 8 9 10 and 11, the 4 tokens fall into the same bucket. Then this takes 12, 13, 14, 15, and 16 5 tokens fall into this bucket. About 8 tokens fall into this bucket about, I think, 12 or 13 tokens fall into this bucket, and so on more and more tokens or more and more relative distances [01:03:00] fall into consequent buckets. So the 8 buckets, 1, 2, 3, 4, 5, 6, 7, and 8 buckets reserved for long distances have more and more tokens falling into them as the number of distances increase, or as the relative distances increase [01:03:15] And for any bucket that is in the closed distance, it has exactly one token or one distance in it. So bucket zero attends to position the relative position 0. Bucket one attends to relative position one. Bucket two attends to relative position two. Bucket 3 to [01:03:31] Native position 3, and so on till bucket 7. So, 8 buckets attend to relative positions 0 to 7, and the remaining 8 attend to relative distances 8 and above [01:03:41] So initially, a few number of tokens or fewer number of tokens fall into a given bucket. And as we get farther out from our current token, or as we get as we get away from our current token, more and more tokens or more and more relative distances fall into the same bucket [01:03:58] So this is how the encodings are learned. And then each of these bucket has its own representation of its distance. So each bucket learns a value that it injects into the Q times k dot product to signify how far away [01:04:14] This q times k this query in key vector are going to be. So the farther away they are, the more of these query and key vectors get the same distance information. The closer they are, the more the model ensures that they get they are able to differentiate between these [01:04:31] So the model can differentiate between the tokens 0 to 8, which is very close to each other. So any token that is 8 into the future or 8 into the past is going to have its own unique attention for the attention. And [01:04:47] Anything apart from that, like, once we get more than 8 tokens into the future, then they are bucketed together, and more and more tokens fall into the same bucket as we get farther and farther out. So this is defined by a logarithmic function that we have just taken a look at. And similarly, for the decoder, for the decoder, since we only need to look at the past [01:05:03] All 32 buckets can be reserved for fast. Anything in the future does not need to be paid attention to, so we know [01:05:10] We need not learn any max distances for future tokens, because they are anyway going to be masked out. [01:05:16] The only relative distance need to be, like, the only distance that need to be learned are going to be the past relative distances. So all the 32 buckets are going to be reserved [01:05:25] for the past tokens. And so we get 16 as our close distance. So from token 0 to 15, each of them has its own relative position embedding [01:05:37] And for tokens 16 and above, they are again bucketed using the same format that we looked at [01:05:42] where just the end buckets increase to 32 instead of 16, right? [01:05:46] So, that is how the encoder and the decoder bucket various relative distances together. So putting it all together, the attention computation looks as follows. We take the query key and value vectors that are computed, right? And we add [01:06:03] So the input the input token embedding is multiplied by the query Wq. Or the query weights. The same token embedding is multiplied by Wk, or the key weights same token embedding is multiplied by the [01:06:17] WV or the value weights to get the Q, K, and V vectors for each of the tokens. The Q, K, and V vectors, when we are multiplying and taking the dot product, or the maximal for these Q times K transpose, we have this alpha. This is going to be the relative distance that we just compute [01:06:33] So how do we compute the relative distance? We take the Q times K transpose, and then we bucket these tokens, the Q and K tokens together by first computing how far away they are. So let's say we have this 0, right? And [01:06:46] Anything to the past is negative, minus one, minus two, minus three, minus four, and so on and so forth. Anything to the future is positive, 1, 2, 3, 4, so on. Next, we take each of these relative distances, and we apply our bucketing formula to it, and determine which bucket does this fall in [01:07:02] So anything in the future is going to be taking the upper half of the bucket. So let's say we have 32 buckets, and bucket 16 onwards is future, and 0 to 15 is past. So the relative distance 1 here will fall into the 16th bucket, because it's a closed distance in the future. Any… anything up to 8 is a closed distance in the future [01:07:19] So, bucket 16, 17, 18, 19, 20, 21, 22, and 23 is reserved for it. And anything in the past, so bucket 1, 2, 3, 4, 5, 6, 7, and 8 is going to be reserved for close distances. [01:07:31] And then as we get, so the ninth token is going to get into the ninth bucket, the 10th token, if there were one here, would get into the same ninth bucket. The 11th token would also get into the ninth bucket and the 12th token would also get into the ninth bucket and so on and so forth. So for each of these relative [01:07:47] positions that we calculate based on the query and key vectors that we have [01:07:51] We first apply the bucketing formula, and we calculate which bucket that falls in. That can look something like [01:07:59] For example, that could look something like this right here, right? When this is for the encoder. So the future tokens are not going to be masked out. So the bucket 0 to 15 is reserved here and bucket 16 onwards is going to be reserved for the future tokens. So once you calculate the buckets for each of these tokens, the bucket might look something like this [01:08:19] Then we take, for example, when we are computing this Q times k transpose matrix to this value, we add the alpha value corresponding to the 17th bucket, or the token embedding, or sorry, the position embedding for the 17th bucket is going to be added here. The position embedding for the first bucket is going to be added here [01:08:37] The position embedding for the second bucket is going to be added here, and so on and so forth. [01:08:42] So that is basically going to that is basically how relative positions are attended to while computing the attention itself. It is simply added as a bias. This the value to be added is calculated by calculating the buckets based on the relative distances. So the first [01:08:57] The first [01:08:59] Step is to calculate the relative position of each of the queries with respect to each of the keys or the query positions can be referred to as context and K as memory positions, because if we are using KV caching, then K and V are going to be in memory. So these are memory positions, these are context positions [01:09:15] And the relative pollution is basically going to be the k position minus the query position or context position. And using this relative position matrix that we have just calculated, we apply our logarithmic formula and get the buckets that we [01:09:29] That these tokens must fall in, these relative tokens must fall in. And then, based on the bucket that they fall in, we calculate which value amongst the 32 values for each of the buckets that we have, which value must be added to this particular relative distance. So this alpha is calculated, like, basically taken as a lookup [01:09:47] From the bucketing matrix that we've just calculated and we add that to the Q times K dot product. So Q times k transpose plus alpha. This is the new factor here during relative position self attention. And so we update the bias, and then we add it [01:10:02] While we do the Q times k transpose Matmel, we scale it optionally, and then we add the relative attention that we have here, the relative attention bias, right here. We mask it if it's a decoder, if it's the encoder, we don't mask it. We then calculate the softmax function, and we take [01:10:19] the multiplication of this output, the attention scores after this, with the values to get the output for that particular head. This is done in every single head. Every single head learns its own representation for all of the 32 relative position buckets. [01:10:35] So each head has 32 values corresponding to its own positional encoding for all of the 32 buckles. [01:10:44] So [01:10:45] Any doubts or any questions regarding this relative position or self-attention [01:11:08] Yeah? [01:11:09] I have one doubt, Somir. So, 16th bucket is my key bucket, right? Where my key token is present, right? [01:11:15] Here in this example [01:11:16] It's… it's not a key, so anything in the future is going to be… so, because we have 32 bucket [01:11:23] Sorry, query bucket, not key bucket. [01:11:26] So, the, the buckets don't correspond to queries or keys. The buckets correspond to relative positions of the query and the key. [01:11:34] So the first thing we do is we calculate the relative position of the query with respect to the key. So this is basically the key position minus the current query's position. [01:11:43] Right? So if the key is to the future, if the key token is to the future of it, then the key, the relative position is positive. If the key token is to the past of the query token, then the relative position is negative. [01:11:54] Okay. [01:11:56] the… these buckets don't correspond to query or key tokens. They correspond to the relative positions of the query and key tokens. [01:12:10] Yeah. [01:12:11] So, if… suppose there are multiple tokens which are present in one bucket, will there… the attention head for that, that those all, tokens would be similar or same, something… I mean [01:12:18] The attention score will not be the same, because the attention score is going to be affected by the dot product of those queries and keys, right? So the key tokens for these tokens, let's say token positions 9, 10, 11, and 12, the token positions that fall closer together, right? So the all of these are going to be part of the same bucket. Let's say the ninth bucket in our case [01:12:37] So, the position encoding for that ninth bucket is going to be the same. The model is essentially saying, I don't want to differentiate between the positions of these four tokens because they are close together, and so they are relatively close to, like, they're close together and [01:12:53] They are relatively far away from my current query as well, so I need not pay that much attention to how far away, exactly how far away they are from me. I just know they are at a distance of, let's say, 10 meters, give or take away from me. Like, if we would think of it in human terms, we would think. [01:13:08] This for anything that is very close to us, we know the exact distance, how far away it is going to be. So let's say from my home, I know, okay, the nearest bus station is, like, 5 kilometers away, and I know all of these bus stations that fall within these 5 kilometers. But anything farther out, I might know the general area that they are in, but I might not know their exact position [01:13:29] The model here is also doing a similar thing. It is saying for anything that is closer to me, I will pay attention to its exact position, how far away it is from my current query. But if it is far enough away, I just need to know the general area or the general distance that it is away from me, not the exact distance that it is away from me. [01:13:48] Okay. [01:13:49] That allows it to reduce the number of the number of embeddings that it must learn for positions. So if all of the positions were paid attention to exactly [01:14:00] Then the model would have to learn 2048 values for this relative positions to work for a context length of 1024. Right? But right now it just learns 32 values for every head. [01:14:13] Okay, got it. Thank you. [01:14:14] Yeah. [01:14:15] So let us take a look here [01:14:19] on how these are calculated right now [01:14:22] So… [01:14:24] Should have, yeah. [01:14:28] So, what we essentially can do is pre-compute these values. These values right here, right? These can be pre-computed. This update bias method is done generally offline [01:14:39] pre-computed, and we have the bias matrix ready for us for every single head. So all we then need to do is just add that bias matrix while computing the the attention, so that the inference is faster rather than having to do all of this because for any given model, let's say if the context of 51 [01:14:57] I can just calculate this entire matrix for all of the 512 tokens, and based on how, based on my current query and key length, I can just take a small part of it. So let us take a look at that, right? So for this model, the T5 model, the maximum context length is 5 to inflict. So I just take that sequence length here [01:15:15] And I calculate the [01:15:19] The relative positions of all of the key positions with respect to the query positions. So let me just put a big point here, and [01:15:20] I think it's better. It's much better. Please contact me. [01:15:30] Uh, Mr. Songi, if you could, uh, please enlarge slightly for better viewability. [01:15:35] Yeah. [01:15:37] Yes, yes, ma'am [01:15:41] Is it better now? [01:15:46] Yeah, yeah, it's better, it's better. [01:15:53] So I'll put a break point here [01:16:08] Oh, sorry [01:16:28] So, if I were to look at, for example, the query positions and the key positions, I would have positions from 0 to 511, because the sequence length for this model or the max context length for this model is 512. So when I calculate the relative positions, what I have is the relative positions of all the keys with respect to all of the queries [01:16:46] Right? So, let us take, for example. [01:16:55] So, the diagonal is zero, and for all of the, all of the positions to the past of me, I have a negative distance, and to all of the positions to the future of my current query, I have a positive distance. So for the 0th query, the 0th k has a distance 0. [01:17:11] The next key one is the first key is one distance of one to the future of the current query. So that's why 1, 2, and so on and so forth. [01:17:20] For the last query, the query corresponding to the 512th token, I have a relative distance of 0 for the current key, like the 512 key. But for any key before that. So for the 51 key, I have a relative distance of minus one relative distance of minus 2, and so on and so forth [01:17:37] So if I, for example, if my query were now only 10 tokens long, then all I would need to do is [01:17:43] Give me the [01:17:46] relative queries for 10 tokens, right? I just take, I just slice it by 10 tokens, and I get my relative distances out of it. So all I now need to do is calculate the buckets for the entire query here, and then replace those buckets with the corresponding values for each of the heads. So I have a per head matrix [01:18:02] For all of the 512 positions, and I can simply just slice it when required by the amount of tokens that I currently have, and add that into my computation to get the relative distance positions as a part of the attention computation. So now let us take a look at how the bucketing is going to be done. [01:18:20] So if bidirectional attention is enabled, then I only have half the buckets available in each distance right? So, for example, if it is the encoder model, then the attention is bidirectional, and then the number of buckets I have available reduces by half, because I only have 16 buckets available to me for both 16 for the past and 16 for the future. [01:18:37] If it is not bidirectional, if it is masked, I have all the buckets [01:18:42] just reserved for the past opens, right? So then relative positions become just the current relative positions, and I don't need to update the number of buckets currently right? [01:18:52] And then, for all of the closed distances, I have an exact representation. So for any distance that is [01:18:59] close to me, which is number of buckets by 2. So, let's say it is an encoder model, and I have number of buckets by 2. [01:19:06] So, my current number kit is 16. Then, anything that is a distance of 8, less than a distance of 8 away, is a small or close distance away from me. And anything that is greater than 8, or distance, or relative distance of 8 away, is going to be far away from me. So, if it is a large distance, or a far distance away from me. [01:19:25] Only then I need to bucket it, because anything that is closer to me, I don't bucket. I have… [01:19:30] A single token, a single relative distance in that particular bucket for any closed distances, right? So that is how I do the logarithmic bucketing only for larger distances right here. And this is the formula that we have just had a look at. So max exact plus log of relative distance divided by the max exact [01:19:49] divided by the log of the max distance by max exact times the number of buckets minus the current, like, or the max exact distance that I have. So, max exact distance is always half of the current number of buckets that I have. So [01:20:02] Once I do the logarithmic bucketing available for me, I can just see what the buckets are, right? And I… these buckets is basically then used to calculate my exact positional embedding. So then let us take a look at what these buckets might look like [01:20:26] Right? [01:20:36] So now, if I take [01:20:39] latest buckets [01:20:41] You can see [01:20:43] the relative [01:20:45] Polition has changed into relative buckets. I have computed the buckets for each of the relative positions that I just calculated, so I can still take relative positions, for example. [01:21:00] In the position… [01:21:04] related position. Yeah. [01:21:08] So, the relative positions zero [01:21:11] fall into bucket zero. The relative position 1, falls into bucket 1. The relative position 2 falls into bucket 2. [01:21:18] And so on and so forth. But the relative distance 500 falls into the last bucket [01:21:23] Bucket 15, right? Because I can only represent exact distances till 8. So let's take, for example, that a relative position is till [01:21:35] 9. And let us take relative buckets till 9. [01:21:42] So you can see, the relative pollution till 8 have their own buckets, right? 0 to 7 has its own bucket. And for the future, I take bucket 17, 18, 19, 20, 21, 22, 23. [01:21:55] All of these have its own bucket. Anything that is close to me, 8 tokens close to me, up to seven from the current distance, if it is less than [01:22:04] 8, or these 8 tokens have their own buckets, bucket 0, bucket 1, bucket 2, bucket 3, and so forth, right? [01:22:11] But for anything that is farther away from me, so let's say if I were to take a relative distance of 15 [01:22:24] You can see that [01:22:25] Sorry, not related. Okay, it's relative distance [01:22:33] Right? So, for relative positions after 8, right, till here, I get [01:22:41] 4 of these relative distances falling into the same bucket, 9, 10, 11, and 12 fall into the 8th bucket. [01:22:45] 1314 fall into the 9th bucket. And similarly, over here, you can see that this 24, these tokens here, 9, 10, 11, 12 here, as well fall into the same bucket. So since I have buckets 16 and above reserved for the future, that is why the bucket here is 24, and in the past, we have the bucket 8 [01:23:05] So so 8 plus 16 is 24 because 16th bucket into the future is going to be reserved for the future tokens, right? So that's 9, 10, 11, 12, these 4 tokens all fall into the same bucket right here, right? [01:23:19] And any further tokens fall into the 25th bucket and so on and so forth. [01:23:25] So, as we get farther away from our current query, we get similar amounts of relative position attention being paid to it, right? So all of these four relative positions receive the same positional encoding from D model [01:23:41] The model trains is the same value to for every single head in the competition, it trains a singular value. So every head gets its own unique value for the relative position over here. [01:23:55] And for all of the closed distances, it gets an exact relative position. So anything that is up to a distance of 8 away from the current token, or the current query, it gets a relative position that is unique to it. And any farther they are bucketed together [01:24:09] Using logarithmic formula. So the farther away we get, the more tokens fall into the same bucket till [01:24:15] the context length is achieved [01:24:19] So that is how we get this matrix right here. [01:24:24] The relative positions here [01:24:27] So, related to positions 59, 510, 511, all fall into the 15, right? [01:24:33] And similarly, over here into the future, they all fall into the 31st. [01:24:39] So once I have these buckets pre-computed, all I need to do is look up the appropriate bias that I need to add and just add that to the model for every single thing. Every head has 32 positional values or positional encodings that it adds based on the relative buckets [01:24:55] And the tokens fall into the relative distances that the tokens fall into. So [01:25:03] Any doubts? Any questions regarding how this is computed? [01:25:21] So, basically, once we have our buckets, or the relative buckets that each of these distances fall into [01:25:29] We can simply calculate the bias or the positional encoding that we need to add by using the weights that are already present in the model, right? Every bucket, the encoder relative position bias basically has, is present in the model, and it… every single head [01:25:44] Has its own representation for that particular relative distance. So every head has 32 positional encodings. These positional encodings are bucketed. So once we calculate the bucket, we simply calculate the value that must be added for that head. So essentially, here, the encoder buckets are sequence length across sequence length that we have just seen [01:26:01] These are 512 for us, 512. Well, sequencing is 512 we look up the the particular value for all of the heads, right? So, for example, for bucket 0, if there are 12 heads, there would be 12 encoders, right? [01:26:17] So we just look up the encodings for the for the bucket 0, and we add them here. We just write them here, and then we add the we we take up the [01:26:28] the encodings for bucket one, we take up the encodings for bucket 2, bucket 3, bucket 4, and so on for all of these buckets. So, in essence, this sequence per sequence then becomes sequence per sequence head, because every val… every single value here is replaced by the encodings for all of the heads, right? [01:26:45] And all we then need to do is transfer this out, so that we have N head cross n sequence cross n sequence, where every head has its own sequence length, cross sequence length values that it must add. And all we need to do is slice this by the current sequence length and add it to the head [01:27:02] So, that is how the pre-computation of this update by a step is done. This update bias is basically the function that we just looked at. Compute the relative position buckets, look up the per-head encodings for each of those buckets, and transpose it and keep it ready to be added [01:27:18] just by using a slice operation. Slice it according to the number of query positions that we have. Let's say it's an 8 across 8 matrix or a 15 plus 15 matrix. We just look up the relative biases for the first 15 values for every single head, and when that head is being computed, we just add the sequence length sequence length matrix into the Q timesk transpose matrix that we would have calculated from this Macmull [01:27:42] During the computation of that attention fields. [01:27:44] So we, the scaling is also optional. The T5 model does not apply any scale. So this root D division is not done for the T 5 model. You just do a softmax of Q times K transpose plus alpha, and the masking is optional. The encoder does not mask the decoder uses the same lower triangular mask that we have looked at for the GP2 model [01:28:03] Because it is causal, right? The self-attention for the decoder is causal. The cross-attention is not [01:28:09] So, the encoder self-attention. Again, let's take the same example. The animal didn't cross the street because it was too tired. So [01:28:17] for every single query. The context of that query can be determined by the tokens to the future and to the past. [01:28:26] It can receive attention from a token that was in the past, right? Animal can pay attention to it, even though it is the future of the animal, right? This is because it needs to build a comprehensive, cohesive understanding of what the… what the context [01:28:41] for the word is, right? What does it refer to? It refers to animal [01:28:46] And similarly, for multiple heads, it learns the subject-verb relationships between different parts of the sentence. So it refers to the animal, it didn't cross the street, so this particular head here is paying attention to didn't [01:29:00] This head here is paying attention to it, was too tired, and so on and so forth. So, different heads learned different parts of the sentence, pay attention to different, relationships between the words, like, subjects, the verbs, the adjectives, and all the components of the sentence. Different, different heads pay different attention, different parts, [01:29:19] of that sentence, we see different kinds of meaning based on the attention that it receives from the other parts of that sentence. So each head is basically learning a different relationship within that sentence. So that is the reason we have multiple heads [01:29:35] And that is why the encoder must be bidirectional, because any particular word could have attention from the past and from the future. Like, it can receive attention from too tired, because it was too tired, or it could also… what does it refer to? It refers to the animal [01:29:51] And that is why we need a bi-directional self-attention for the encoder, because we need a contextualized token embedding for every single input token. This context is then fed into the decoder [01:30:02] in order for the decoder to be able to understand what it must produce next. So what must be the output that the recorder produces? It is a completely different domain that it might be decoding into, for example, for neural machine translation, the encoder can encode English, and the decoder may decode into German. So [01:30:18] The English context is fed in a vectorized format into the decoder. [01:30:23] So, the first part of the decoder is the self-attention. The decoder must also understand what it has already said, right? So for example, in the English to German translation using the T5 model, the animal didn't cross the street because it was too tired. The German translation is thus tear uberk daser nicht well es zu muder war [01:30:42] So, when it is, let's say, decoding S, it, like, so when it has decoded till here, it must understand what it already has decoded previously, right? So, for example, when the beginning of sequence token is given [01:30:55] The self-attention can only pay attention to beards, right? And then, as… as the decoder autoregressive will decodes one token after another, only the tokens that have already appeared in the past can be paid attention to by the current token, right? [01:31:11] So, for example, when we are taking S, for example, S can only receive attention from Zoom, Mode, and more. It could not [01:31:20] receive any attention from TIA, which is, like, animal here pays attention to it. The corresponding equivalence here would be animal, this tier here, paying attention to S. But it doesn't, because when tier was recorded, S is to the future of that. The decoder cannot pay any attention to any future tokens [01:31:36] It can only pay attention to a token that is to the past of the current token. So when it is decoding war, it must now decode end of sequence EOS. So it has the entire context, it knows what it already has decoded, or what it already has out in the past. And based on that, it can then deport whether the next token is going to be end of sequence or some other token, or whichever else is going to be [01:31:59] So the first token here is always beginning of sequence. [01:32:03] And the encoders [01:32:05] The encoders, encodings, or the contextual embeddings are fed into the decoder to be able to… for the decoder to be able to understand what it must then produce, right? [01:32:16] So, the self-attention is masked because the decoder cannot pay attention to anything that is to the future of its current decoding. It can only pay attention to the currently decoded sequence. [01:32:27] Right [01:32:29] So here, again, the multiple heads are doing the same thing. Each head is paying attention to a slightly different part of the same sentence, right? [01:32:37] So for example, here [01:32:40] As what are being decoded and S is 8, zoo is I think mode refers to too tired and zoo was or something like [01:32:53] That is what it means in German. So here it can… different heads are paying attention to too tired or so this green head here is paying attention to too tired, or this is paying attention to walls, the pink head here, and so on and so forth. Different heads are paying attention to different parts of the sequence, and the decoder is causal [01:33:10] It can only pay attention to parts of the sentence that already exist. So in the past of the current sequence. So, S cannot receive attention from anything that occurs before it. It can only receive attention from anything that occurs after [01:33:25] So [01:33:26] The cross-attention [01:33:28] like, in a way, feeding the decoder with the context for the next token, right here, right? So, for example, this is another example of French to English. Which is, I am a student in English. So, the encoder encodes this French into a sequence of [01:33:45] Tokens. So every token has its own contextual embedding. These contextual embeddings are used to produce encoders like these encoder, these contextual token embeddings produced by the encoder are used to produce the decoders K and V for the cross attention [01:34:04] So how does this happen? How does this work? Well, let's take a look [01:34:07] So, right here, if we look at the T5 model, like, the cross-retention here [01:34:14] The [01:34:17] Yeah. [01:34:18] So first, we run the encoder [01:34:20] Let me just remove this quick point here. [01:34:25] Thank you. [01:34:28] Right here. [01:34:30] We run the encoder, and we get a set of encoder hidden states, or the output from the encoder is basically referred to as the encoder hidden states. These are the per-token contextual embeddings or the vectorized context of the input sequence [01:34:43] These are fed into the decoder so that along with the decoder start token ID, so that it can decode. [01:34:50] the next sequence [01:34:54] So, this T5 decoder here, it is supplied with the decoder input IDs, which is the decoder start token ID and the encoder hidden states, or the encoder output [01:35:03] So how does it then use this to compute the output right here? So let's say this is the decoder block. [01:35:11] The decoder block first runs a causal self-attention, which is masked, and then it runs a cross-attention [01:35:19] So how are the encoder params or the encoder hidden states used to compute this cross attention? The encoder hidden states are passed into the cross attention matrix [01:35:25] And they are used to calculate DKV for that particular cross-attention, right here. So these encoder outputs, right, for example, if you take [01:35:37] The key input is the encoder hidden states, and the decoder block then [01:35:43] compute this right [01:35:49] carry input times KW and KV input times VW, right here [01:35:54] Okay? [01:35:56] It provided its cross-attention. So, all you're doing is you're taking the encoder's hidden states, and you're using that to compute the keys and the values for the current decoding sequence. The query is calculated from the decoder's [01:36:11] own context, right? So the query refers to what the decoder knows about the sequence and the key and the values are the context or what the decoder receives as context from the encoder, right? [01:36:25] So these per-token embeddings are used to calculate the keys and values. So let's say the 3 tokens were fired into the encoder. The encoder produces output of 3 cross 768, where 768 is the embedding dimension right [01:36:39] These are then passed into the decoder for every layer in the decoder during the cross-attention computation, the K and V are is produced using this encoder output as the encoder outputs Machmann with the cross attentions [01:36:55] key and value weights. And this key and value is the way that the decoder understands what it must decode next, because the query is calculated from the decoder's current output sequence. So the decoder knows, okay, I have done this, or I have just [01:37:10] decoded this sequence. And what I must decode next is then fed into by using the key and value as a lookup, or context lookup, so to speak. So [01:37:21] The decoder understands what it must decode next based on the context that is injected into it from the encoder. [01:37:29] the keys and values of this computation is the way that the encoders context is injected into every layer of that decoder during the cross-retention step. The decoder then continues decoding token by token, and at every stage, it understands what must be next based on its own context during the self-attention [01:37:51] And the injected context during the cross-retention. [01:37:55] So, any questions regarding this? [01:38:09] Okay. [01:38:10] So, if there are no further questions, we can take a comprehensive look at this, how the cross-attention is computed. So, for the input embeddings matrix X, we create the matrix query key and value as Q of X times Wq, y times wk, and y times WV [01:38:28] Here Y is the encoder's output or the encoder hidden states. And Wk and WV are the keys and value weights for each of the cross-attention layers for each of the decoder layers. So every layer has its [01:38:41] weights, key weights and value weights. The query matrix, the query vectors are computed per token based on the decoder's current input and the key and value matrices are computed from the encoder's contextual output [01:38:58] So the decoder's own input is used and the encoder's contextual output is used to inject more context into the decoder's current input or its current understanding of the state that it is in. [01:39:12] So, after these query key and value weights are multiplied with the decoder inputs and the encoder outputs, we get the query vectors, the key vectors, and the value vectors. These query key and value vectors are then used to compute the cross-attention [01:39:27] Similar to how self-attention is done, which is basically soft marks a few times k transpose by root D times V. Again. [01:39:33] The T5 model does not scale, does not use this root D factor to scale. It only uses the the relative position embedding, and that relative position embedding is already injected as a part of the decoder's self-attention. So it does not need to inject any more position embedding into the cross attention. The decoder already the the decoder's queries [01:39:54] Already now have the positional encodings embedded into it, and the encoders values already know about their relative positions, because it was computed as a part of the encoder's self-attention as well. So the cross attention block does not need to use any relative positional embeddings. They are only used during the self-attention of the encoder and of the decoder as well [01:40:14] So, by the time we get to cross attention, the encoder output, the contextual encoder output and the decoder's own current input already has the position information baked into it as a part of the self-attention computation of both of these units. So during the cross attention, we just need to inject the context [01:40:30] We already know about the positions in both the states. So we just need to inject the context about the decoder's current state and what it needs to do next based on the encoder's output context. [01:40:42] Right? So, putting it all together, cross-attention might look something like this. The animal didn't cross the street because it was too tired, was the input. And das Strasser nicht, well, it's zoom war is the output that it decodes into [01:40:57] So here, as you can see, the cross-attention during cross attention animal, the encoder output context for animal is used to inject context into the decoder's sequence as [01:41:10] last year, so the highest attention is being the highest gross attention is being paid by animal to Dastier, which is the animal tier in German means animal. [01:41:22] So this is how cross-attention is used to inject the encoder's context into the decoder, so that the decoder knows what to do next, or how to decode any further. So as you can see here, what did the animal not do? The animal didn't cross the street, so some amount of attention, cross attention, is paid by the animal to Strasser as well [01:41:41] which which is street in German [01:41:43] And it also pays some amount of attention to mire, which means too tired [01:41:49] And so on and so forth. So every single cross-attention head pays attention to different parts of the sentence, and similar… in a similar manner to how the self-attention works [01:41:59] And in addition to that, it is using the encoder's output context to inject some context or some understanding of the of the output state into the decoder [01:42:12] Any questions, regarding this cross-attention computation and how it is achieving the context carryover from the encoder to the decoder [01:42:35] Okay. [01:42:36] So if there are no further questions, yeah, so we can. So this is how the encoder decoder attention cross attention works. [01:42:44] So, first, the encoder runs on the input sequence. It generates a per-token contextual embedding. These per-token contextual embeddings, along with the decoder's beginning of sequence token, is passed as the input to the decoder. The decoder runs in an autoregressive manner. Every single output is appended to the previous decoder input [01:43:01] And the decoder is run again. This is, this process is continued until an end of sequence token is generated by the decoder. So during the decoder's computation. [01:43:12] I think we have a question [01:43:14] Okay, so during the decoder's computation [01:43:18] The decoder first computes a self-attention so that it knows what it has already output, what, so what is its current state, and then it runs a cross-attention. During the cross-attention, the encoder outputs inject some amount of context into the decoder, so that the decoder knows what to do next. [01:43:37] So for example, during cross-intention, the encoder's output corresponding to the token animal is paying some attention to the German equivalent, which is tier [01:43:47] It is also paying some attention to Strasser es Zumudevo as well. So [01:43:53] The animal, the animal was [01:43:56] tired [01:43:58] And it didn't cross the street. [01:44:01] So the same parts of the sentence being paid attention to in English during the encoded self-attention are being paid attention toby the cross-attention as well. It is injecting the English context into the German language so that the decoder can decode [01:44:18] into German what it must what the contextual meaning of the English sentence was. And that is how this neural machine translation example is working using the T5's encoder-decoder cross attention [01:44:31] Right. [01:44:33] So coming to the benchmarks, the T5 model performance was benchmarked against BERT base across the following blue benchmark tasks, which was multi-genre natural language inference, the semantic textual similarity benchmark, the Microsoft Research Paraphrase Corp [01:44:49] And the textual entailment corpus of the GLUE benchmark, which is the language understanding. It was also run on the Stanford sentiment analysis tree bank [01:45:00] And the particular spores that the T5 base model was able to generate, was the new state of art. It beat the bird base in all of the benchmarks and in the average, GLUE score as well [01:45:12] The T5 model also performs better than board base on the Stanford question answering dataset, outscoring bird-based by quite a significant margin in F1 score. [01:45:22] So, we can have a look at how this is running and we can put it all together as [01:45:28] I'll close the mark [01:45:40] As you can see, first we tokenize the inputs. These token embeddings, so these tokens are used to create per token token embeddings for the context of the entire sentence. This is generated by the encoder. This, along with the decoder beginning of sequence token, is passed to the decoder and the decoder generates in autoregressive minor. It generates token by token starting with decoder BOS and [01:46:04] as it generates tokens, those are appended to that particular decoder's input, and it keeps on running till it generates an end of sequence token. Once the end of sequence is generated, we stop the generating. So I have set a maximum of 40 tokens to be generated, but just 6 were required, because the 7 token was in EOS, so it exists [01:46:20] from that [01:46:21] So, the output from the model can be looked at here. These are the tokens that we get as an output from the model, and these tokens are de-tokenized to get us this sentence, which is the German equivalent of [01:46:34] Tell me about Qualcomm. I asked the model to translate English to German and tell me about Qualcomm, and I got the output. Which is, tell me about Qualcomm in German. [01:46:44] So putting it all together, this is how DT5 model works. [01:46:49] Yeah. [01:46:51] We can take a few questions [01:46:54] If there are any [01:46:55] Thanks. [01:46:56] Any? [01:46:57] Yeah, or if I need to explain… sorry, sorry, ma'am. [01:46:58] Yeah, I [01:46:59] Yeah, any questions, anyone? [01:47:00] Yeah, I have one for the future rent, so I don't see the module force recording, on IITR, dashboard [01:47:09] So, can I know when they will be updated? [01:47:14] Your Future and co-host, can you please come in? [01:47:18] Yes, ma'am, I'm here [01:47:19] Yeah. [01:47:20] Okay, so you are asking for Module 4 [01:47:23] Yes [01:47:24] Okay. [01:47:28] Yes, I'm checking. [01:47:29] Basically, yesterday is today's recording, and then… and then other module for recordings. [01:47:34] Okay, it's not [01:47:35] Uh, since when, uh, Deepak, can you please tell since when, um, the recordings you feel are not there? [01:47:39] I only see Module 3's recording, so, the date for the Module 3's end recording, let me tell you right [01:48:07] I see 14 to 2026. That's it. [01:48:10] Sorry, what do… what… what did you do say? [01:48:14] I see 14 to 2026, the whole tab for Module 4 is missing. [01:48:20] I see. [01:48:21] Alright. Alright, alright. [01:48:24] Okay. [01:48:26] I will get it uploaded by tomorrow, Yodi [01:48:29] Will that be fine? [01:48:30] Okay. Thank you, yeah, yeah, that will work, yeah. [01:48:32] Okay, thank you, thank you. [01:48:33] Yeah, please, uh, get it done, uh, ASAP. [01:48:37] Yes, ma'am, yes ma'am. [01:48:38] Uh, uh, Pallavi, do you have a question? [01:48:44] Yes, so this is just, like, to clarify on T5, so when we generate the input tokens, input token embeddings, the same token is passed to the decoder, so… so then the decoder uses the self-attention mask to [01:49:00] calculate, like, its own tokens, and also uses the cross-attention. Additionally, where we do not mask the future tokens, and together they… after that, we generate the final token from decoder. [01:49:15] Yeah. [01:49:16] Is that right? [01:49:17] Yeah. [01:49:19] Okay. [01:49:20] So we can have a look at the architecture here, once more [01:49:24] The encoder runs a single time. It generates a set of context embeddings for every single token that was input to the encoder. These are supplied to as inputs to the decoder along with the decoder's current sequence. Initially, it is beginning of sequence and as and when it generates tokens, those are added to it [01:49:43] Keeps on running till end of sequence token is encountered. First, it runs the self-attention, which is causal or masked, and then it runs the cross attention. This cross-attention here, the query is generated using the decoder's current input, right? So the output from the cross attention self-attention here is used to generate the [01:50:00] query for the cross-attention, and the key and value for this cross-attention is generated using the encoder's output states. So this is the connection where we get the context, or the encoder's contextual understanding of the input into the decoder's sequence [01:50:15] This is where it is mixed. And this is the key component that allows the decoder to understand the domain output from the encoder, or the context from the encoder, and it then produces a set of a vector [01:50:31] That is its understanding of the of the current sequence, I think, or the next value of the current sequence, which is then used by this classifier here to output probabilities or logits from which we can sample the next token [01:50:48] That is the sequence in which it works. [01:50:51] Okay, thank you. [01:50:56] So, does that answer your query? [01:51:01] Yes, yeah, thank you. [01:51:02] Okay, any other questions, anyone? [01:51:10] Okay, there are no further questions, then we can wrap up for today. I think the portion that was intended was [01:51:18] T5 and, uh, we… that has been covered by Mr. Sawyer. [01:51:22] Uh, okay, thank you all. [01:51:26] Thank you, Mr. Somani, for taking out the time. [01:51:28] Thank you, thank you all, have a great day, thank you, bye-bye. [01:51:33] Thank you. Thank you, everyone.