# 06 2026-05-30 GPT Theory Contd and T5 course: Module 4 — Generative AI & LLMs module: Module-4-Generative-AI-LLMs date: 2026-05-30 type: transcript video_url: https://personal-learn.armco.dev/files/_Recordings/Module-4-Generative-AI-LLMs/06_2026-05-30_GPT_Theory_Contd_and_T5.mp4 --- [00:31:50] So, this is how it works. This, uh… [00:31:52] For example, without giving… this is an example with takes. Without giving hashing and with key caching. [00:31:56] So, for without kicking, we, at every step, we must compute the entire thing. We must compute the first query, the second query, the first key, the second key, and then compute the Q timesca passwords, right? [00:32:05] And same for values as well. But with key caching, we need not it. Once we have computed the first many, [00:32:11] We can cache the K value. [00:32:13] Then the next query, we only compute this query and cache this. And for the next query, we only compute this query and this key, and we catch this. [00:32:19] And so on and so forth. So instead of computing all of this and all of this at every step, [00:32:24] We only compute one query vector and one key vector at a mix. That reduces the computation by quite a huge factor. [00:32:30] Uh, but of course, at the price of some women. [00:32:33] So, uh, this was the overview for KVC action that we discussed last week. [00:32:38] Any questions regarding this? [00:32:54] Yeah. Um… [00:32:56] So, once this is done, once the gaming caching is done, this allows us to speed up the model sentence by a huge factor. [00:33:05] This saving of… or committing of the key values to a cache, uh, of course, needs some memory and has some memory overhead, but… [00:33:13] At the same time, it allows us to save on a lot of computations, which, uh, send these models are anyway, uh, like, [00:33:19] they add memory bounds, saving on computation does help quite a bit. [00:33:23] So, I'll… yeah. [00:33:27] So, let us take a look at the code that we had taken last week as well. [00:33:30] So, this… this is how it changes between Kcash. [00:33:35] Let me just pick this one, right? [00:33:38] So, I'll… this is how the… [00:33:41] the computation for multi-retention might change in the case of KVCH. So, when we have no KV cache here, [00:33:47] We need the entire sequence. [00:33:49] At every step, right? But if we do have a Kiwi cache, [00:33:53] We need not compute the entire n-sequence rate. N-sequence will always be one after the first query, or after the first sequence is when the prompt processing is done, and end sequence will always be 1. [00:34:03] So, we will compute new query, new key, and new value. [00:34:06] And we just append that. [00:34:09] to the current cake, right? [00:34:11] So for the first time, when we are running it, [00:34:13] we will have to create a KV cache. And for the next time, we just keep stacking it. So, end sequence, and weverse end sequence plus 1 is going to appear. [00:34:22] And in the first run, we create the KVCAR shape, right? [00:34:25] So, right, this is the first sequence, because then KV cache will be none for the first [00:34:31] Uh, because we don't have a KB cache for the first query. [00:34:33] Over here, at every single inference, we have to supply the entire sequence. N sequence is going to be, uh, initially, let's say, it's 5 tokens, then it will be 6 tokens, then 7 tokens, then 8 tokens, then 9 tokens. [00:34:45] Whereas in this case, end sequence for the first query will be 5 tokens. Then for the second, uh, second… [00:34:50] pretty. It'll just be one token, the new token that was generated. Then for the third query, it will just be one new token that was generated, because the five… [00:34:59] tokens initially versus are committed to KVCatch. Then the first token after that is also committed to KVCANS. So KVCash keeps increasing at every single… [00:35:06] iteration. So, for example, if I have framed [00:35:09] Let's see… cheap. [00:36:11] Oh, sorry, yeah, old cake, it should be. [00:36:35] You can see, for the first one, the prep, [00:36:37] is 5. [00:36:39] And this is 6. The next one, the prev, is… [00:36:43] Again, 5, 6, 5. And it continues increasing, right? [00:36:47] 6, 7. So these are for all the layers. For each layer, you have [00:36:52] KD cache for each layer, right? [00:36:54] So for each layer of the decoder, you have a key catch. This is for the first layer, this is for the second layer, the third layer, and so on and so forth. [00:36:59] Then, for the next token, the previous is 6, and we append 1. Then 7, then 8. [00:37:05] And so on and so forth. So at every single step, the KV cash increases by a single token, which is the number of tokens, like, the most recent produced token. [00:37:15] So, that allows you to… [00:37:24] That allows you to, um… [00:37:26] Basically, save on these computations. So, as you can see, the cache is constantly increasing. Here we have 44 tokens at the end, right? [00:37:34] So these 44 tokens, we would ideally, uh, without KV catch, what would have happened is the first query would have been 5 tokens. [00:37:41] And then, during the next query, instead of computing just the sixth open, we would be computing the entire [00:37:48] sequence, which is the previous 5 plus the current one. So instead of just computing 6, uh, like, just computing 1 here, and appending it to get 6, we would have to compute the entire 6, 6 plus 7, 68 here, and 6 plus 768 here. And similarly, for the subsequent query, instead of just computing 1, [00:38:04] And appending that to 6, you get the 7. We would have to compute the entire 7 plus 768 once more. And similarly, for the next sequence. [00:38:12] When we come down here, uh, as the… as the number of tokens predicted increases, like, for example, in this… in this example, we have 44 token speed predicted. So, as these number of tokens keep increasing, [00:38:22] Uh, during this run of the decoder, we just computed the [00:38:27] last tokens, key and value, and committed that to the previous key and value to get from 43 to 44. [00:38:33] But, uh, without KVCAS, we would have to recompute this 44 plus 768, uh, into two. [00:38:38] vectors, because one for KV and one for key, and one for random. [00:38:42] So, that is a huge overview. [00:38:44] And this allows us to save [00:38:46] a lot on the… [00:38:50] Oh, on the computation. So, this is very important. [00:39:01] So… [00:39:05] I hope this is better. [00:39:13] So, yeah, um… [00:39:16] Using this KV cache, uh, we are able to save on a… [00:39:19] lot of computation that allows us to run the model much faster, uh, without Kiwi caching. With KV cache, we get about, uh, 2 or 3x speedup in terms of number of tokens predicted per second. [00:39:30] when we use KB Cache versus when we don't use KVCH. Oh, and this factor just increases. [00:39:36] Um, as the context length of the model increases, as more and more tokens keep getting generated. So, we need to [00:39:41] Uh, we only need to compute one token at every step here, whereas here, we need to compute all of it, right? [00:39:47] So, this n sequence plus n-embed n-sequence continues to increase. Here, n-sequence gets pegged at 1 after the first query. [00:39:55] When we need to compute the KV for the entire sequence, uh, in the first one, right? [00:40:03] So, this is… [00:40:09] Yeah. So this was about KV caching, um… [00:40:12] This is one of the most important, uh… [00:40:15] or the foundational concepts of inference on any of these decoder-only models, where we need to increase the number of tokens generated or token throughput. [00:40:25] And then coming to fine-tuning. So, apart from, uh, just text generation, GP2 model can be fine-tuned for many specific down-speed tasks. One of them, for example, is summarization. Another one could be text entailment. Text entailment basically means that, uh, given a hypothesis and, uh, explanation for it, we need to classify as whether the explanation explains the hypothesis. [00:40:47] Or, like, sorry, when the hypothesis explains the, uh, the… [00:40:50] the passage, or the hypothesis does not explain the passage, or whether the hypothesis is not related to the passage given as well. So that is to explain ailment. [00:40:57] Um, and one… one other task could be, for example, that I have taken here, which is summarization. [00:41:03] So, let us take a look at how we would use GPT-2 for, uh, like, fine-tuning, uh, for the summarizations. [00:41:10] Uh, the output from the GPU model, uh, is… [00:41:14] basically, if it went to a language modeling, right? This language modeling head is the classifier that we had spoken about, right? But for the next sentence prediction, or next word prediction, all we do is, uh, [00:41:23] take the last, uh, the logics of the last token that was input to the model, and ask the classifier to take the next open, based on these logs, right? So… [00:41:32] Oh, this is what, uh… [00:41:35] This is the part of the model that we are fine-tuning or changing. When we are fine-tuning it, uh, we train a different modeling, or a different classifier, which is a different feed for our network and softmax combination. So, let's take these summarization example. For example, if I were to give this article to GPT-2, I would want this summary out of it. [00:41:52] So this can be my training aspect. I take all the Wikipedia articles on the website, I give it the article, and the summary is the abstract of [00:42:02] of that specific article. [00:42:04] So, how do I prepare the model? So, the training dataset is basically the article, [00:42:08] then a special token called Summarize, and… [00:42:11] This is… this is a new token that I would add to the vocabulary, which is basically one of the added tokens in the vocabulary. Uh, GP2 has a couple of added tokens that are not being used, so I could basically train one of those to use the… to, uh… [00:42:23] substitute as the summarized token, and then I would give the article summary. And similarly, article 2… summarize article 2 summary, and pad it to the maximum length of whichever, um, [00:42:32] like, this would be full of padding tools. So, whichever has the longest, whichever article is the longest, uh, [00:42:38] that would have just the article and the summary tokens. The rest of them would have some padding appended to it. [00:42:43] And at every given point in time, I just, uh, I just get this, uh… [00:42:48] Article 1 tokens, and summarize as the input to the model. And I get output [00:42:53] Uh, as… and then, uh, based on this, I get an output. So I… [00:42:58] predict… like, I feed that output to the classifier, and I predict the next token. And I keep doing that till, uh, it gives me an end of sequence, which… at which point I should basically have the article summary here. [00:43:08] Right? So, that is… that is how I would fine-tune the GP2 model. [00:43:13] Uh, for this particular task of [00:43:17] Uh, summarizing the article. Here, [00:43:19] This summarization is just done by training a different final modeling head, or a language modeling head. [00:43:24] Right? We do not need to retrain the entire GPT-2 architecture, transformer decoder architecture. I don't want to fine-tune or propagate all of the loss. In fact, as we had seen in, uh, and discussed in the previous lecture, that if we were to propagate this loss throughout the model, it would basically lead to the model forgetting or unlearning a lot of the basic [00:43:44] Um, a lot of the basic knowledge that it had learned during the unsupervised, uh, [00:43:49] pre-training. Instead of… and that is the reason that, uh, this is called a discriminative fine-tuning. Discriminative is basically how much information we want to backpropocate into the, uh, [00:44:01] into the modules layers. Every layer learns at a different learning rate in the discriminative fine-tuning. Here, I would only want to train the language modeling head for this task. [00:44:10] So, I propagate none of the errors back into the model. So, as we had seen the loss function, [00:44:15] Last time, right? Uh, let us take a recap, a quick recap of the loss function. [00:44:29] Yeah. [00:44:31] So, uh, this was the discriminatory fine-tuning, uh, right? Where LDAS was the loss on the fine-tuning dataset, and Lambda would control how much of this loss. [00:44:42] is back-propagated into the model. [00:44:44] Great. If we were to, like, this was the loss function for the unsupervised pre-training, and using that, we can calculate LLC. [00:44:51] on this particular corpus, and if I were to propagate a lot of this loss, if lambda were to be too large, then the model would basically unlearn a lot of the things that learned during this, uh, unsupervised pre-training. So, we need to be very careful about this primitive fine-tuning, and what amount of loss we want to propagate back into the transformer itself, rather than just the final feed-forward layer, which is [00:45:10] this HLN and softmax combination, which is HLN is going to be the, um, like, sorry, HLN times W, which… where W is what we want to train. [00:45:19] Uh, these are the weights that you want to train for a feed-forward layer that would then use the softmax activation to give us probabilities of, uh, for the next output from this. [00:45:29] So, we need to be very careful about, uh, the, uh, the discriminative parameter, or lambda, as we… [00:45:35] see. Otherwise, uh, we can… [00:45:38] unlearn quite a lot of the information that the model has learned during its unsupervised training. [00:45:43] And we don't want that to happen. So, that's the reason we only want to train the language modeling head, uh, the LM head. We switched that out for a new head, and we get a summarization. [00:45:52] And similarly for all of the other examples that we had looked at, text entailment, or centers classification, or any of the other things. So the Jupyter model can be fine-tuned, uh, with different language modeling heads for different tasks. [00:46:02] Uh, and the application domains for these files could be completely different. [00:46:08] So… [00:46:14] coming to the next slide, um… [00:46:16] GP2 also demonstrates some amount of in-context learning. Well, what is in-context learning? So far, we have seen that models learn via training, like, explicit training steps. [00:46:26] Where we first have the unsupervised pre-training, and then we have the task-specific fine-tuning. Uh, this fine-tuning helps the model express some task-specific domain, like, for example, summarization, or, um, [00:46:38] entailment, or sentence classification, or so on and so forth. But [00:46:44] We also have seen that GPT or other large language models as well, uh, can perform some amount of learning without any stochastic gradient descent. Like, we don't need to backpropagate any of the losses. [00:46:56] Just by giving them some examples, they're able to give meaningful outputs. For example, I could take GPT-2 and give it the input, thanks, [00:47:04] to mercy. Hello, to Bonjour. Mint, to mend. [00:47:07] order to question mark, and ask it to what question mark is, and it would give the output looter. [00:47:13] Which means that it understood that, okay, uh, thanks, who's mapping to Mercy, which is the French word for thanks. Hello, was mapping to Bonjour, which is the French word for hello. [00:47:20] Mint is not implemented here, which is a French word for mint. So, otter. I should give the French word for otter, which is looterite. So, this… [00:47:28] is called in-context learning. So, there are a couple of types of in-context learning. [00:47:33] zero shot, few shot, and one-shot learning. [00:47:35] So what is zero-shot? Well, Zero Shot is where I don't give it any of the examples. In this, I have given it a few examples, right? So this is a few short learning. [00:47:43] But if I wanted to, for example, do zero-shot learning from the model, what I would do is I would just say translate English to French, order to question mark. [00:47:52] or, for example, translate English to French, cheese to what is the output? Question mark. This would be my prompt. And the model would then predict the output. [00:48:00] So, this is the hero short learning, where the model is given no [00:48:03] context, or no example of what it needs to do. It just is given a task, and like a task description, and then the task gets [00:48:11] So, this is zero-shot. In one shot, I give it exactly a single example. So I would say translate English to French, [00:48:19] Uh, sea otter Leutoderma, what is cheese? Cheese? For example, and ask it to put a question mark. [00:48:24] So, in addition to task description, in this example, we are giving one example. [00:48:28] And that is one-shot learning. If I were to give it a few examples, like I have given it in either this one, [00:48:34] or here where I gettranslating issues assigned C, or to little peppermint to mention O-V, plush hereafter, giraffe plush. [00:48:41] So, what is the output for cheese? This is, uh… [00:48:45] an example of few-shot, where I asked the model to take, like, basically learn a mapping, or do some sort of reasoning, where, okay, what is the pattern, or what is the output going to be? [00:48:55] from the given pattern, or the given examples. So this is for few-shot learning. Uh, GPT demonstrates all of these. It can do zero-shot for very simple tasks, it can do one-shot for some slightly more complicated tasks, and it can also do a few shots for [00:49:08] really complex tasks. So, that allows… GPT has demonstrated a lot of in-context learning capacity, and the size of, of course, the larger the model, the more the in-context learning that it can perform, based on the large amount of general knowledge that it has learned during the unsupervised pre-training in various languages. [00:49:24] So, this is quite, uh… this just highlights the power that we have with respect to in-context learning, or any of the other search mechanisms, where you can just give the model a few examples of what is the task that must perform, and then performs that task without [00:49:40] having to do any sort of fine-tuning, or without having to do any sort of gradient descent. [00:49:45] Black propagation of errors, or any of the, uh, the usual learning mechanisms that we have for ML models. [00:49:52] So, this was about in-context learning. Uh, let us take a look at the benchmark. So, how do we say that the output is meaningful? Well, one of the most basic benchmarks for, uh, [00:50:01] the causal language modeling is… or the next token next word production is perplexity. It is a very widely used metric for evaluating the model performance of large language. [00:50:13] It measures the uncertainty of the model split. So, [00:50:15] an ideal model output should be 1. [00:50:18] or probability of 1 for whatever the next word is, and 0 for everything else, right? [00:50:23] That is the ideal output. [00:50:25] But it does not always happen. [00:50:28] we get a very high probability for the next word. [00:50:31] at very low probability for all of the other ones. So, in that sense, the model highlights a confidence that it has, or how much… with how much confidence is it predicting the next word. [00:50:41] So, this is the uncertainty. So to speak, of the entropy of the model's predictions. So, [00:50:46] Perplexity basically measures how uncertain or how inconfident the model is. If the model is very confident, then the publicity is lower. [00:50:53] A low publicity, conversely, highlights a higher confidence and better performance on the tasks. [00:50:58] While a higher perplexity demonstrates that the model is very, very much uncertain and is struggling a lot more with the input. [00:51:05] So, how do we define this entropy or perplexity mathematically? Well, this is the mathematical equation for publicity, which is… [00:51:12] e to the power of minus 1 by n, where n is the total number of words in the model vocabulary. So 50,000257 for GPT-2, for example, but N is basically the total number of words that the model knows. [00:51:22] Because, uh, the model can only predict the probability of the words or the tokens that it knows. So, [00:51:27] when we want, uh, to wait across all of, uh, to… [00:51:32] all of the probabilities for every token, we want a probability, and we want to weight the model's confidence of that probability. We must also divide by the number of tokens, right? So, this is the N, the model vocabulary. And then, uh, we take the summation of the logarithm of [00:51:47] Each token, given the context, right? So, the P of Xi given X1 to Xi minus 1, [00:51:54] for all I belongs to the model vocabulary, uh, gives us the probability that how confident the model is that the next token is XI, [00:52:01] based on the current sequence. [00:52:03] So, this is also… this can also be defined as the conditional probability of the IATS token being the XIH token being the next token, given all of the X1 to XM-1 context. [00:52:13] So, this is how we define it, and this formula can equivalently [00:52:18] again, equivalently be written as the below equation. [00:52:22] Great. Which is the nth root of the product of 1 by P of Xi for all I belongs to the model vocabulary. [00:52:32] So, um, any doubts about this? [00:52:46] Okay, um… [00:52:48] So this is the public safety measure. It basically tells us how confident the model is in its output. [00:52:53] If the perplexity for, uh, like, if, for example, the probability was 1 for a certain [00:53:00] output, and 0 for everyone, for every other output, then the model would be completely certain. [00:53:06] And the publicity would be the lowest. [00:53:08] And conversely, if the model is very uncertain, and it's 1 for everything, then publicity would be a very high number. [00:53:14] So, the low publicity indicates a higher confidence and better performance of the model. If the publicity is very high, then that means that the model is struggling with the input. It cannot represent… it cannot predict with any sort of confidence what the next token [00:53:26] is supposed to be. So, how do we evaluate this, given a dataset? Well, let's have a look. [00:53:32] Uh, if the model were not limited by context phase, we would just keep evaluating the publicity by autoregressively factoring in the sequence. For example, on this sequence, we would first [00:53:41] compute the probability of, uh… [00:53:44] face, given, okay, then we would compute the probability of is given phase, or given is, startup given, uh, er, and so on and so forth. So, we would keep autoregressively increasing the context, right? [00:53:56] So, for every token, we take all the preceding tokens and compute the probability, uh, that the model believes this [00:54:02] token, or this word, is the next word based on all the previous context. [00:54:06] But, in real life. I mean, practical applications, we have seen that model has a max context length. It can only represent so many tokens, because it does not, uh, it cannot differentiate between any positions after its context length. So if the context length of, for example, GPT-2 is 1024, it only knows positions 1, 2, and 1024. It cannot [00:54:24] understand opposition 1025, because it has no position embedding for that particular position. [00:54:30] So we have a constraint on the maximum number of tokens that the model can process, or the maximum amount of tokens that it can understand. For this, [00:54:37] the GPT-2 model, uh, has a 1024 maximum context length, and we must break down the sequence into smaller chunks. [00:54:43] Well, there is another problem here. [00:54:45] So, for example, if I… in this example, if I were to break down the sequence at based, [00:54:51] Where the model can only take up to base as its context. [00:54:55] Well, then what happens? We get the perplexity for hugging, for face, for is, a word based [00:55:01] till base, we have a context. For base, we have the context till everything before base, right? Hugging Face is a startup, is the context for base. [00:55:09] But the moment we get to in, [00:55:11] the entire context is… [00:55:13] dropped, right? If I split here, [00:55:16] the entire context was dropped. [00:55:18] So, essentially, the perplexity of the model for IN would be very high, because it does not know [00:55:24] anything that comes preceding it. And this puts skew our output perplexity by [00:55:30] a very large margin, and it would be of no use to us. [00:55:33] So then, how do we overcome this? Well, we instead use a sliding window strategy. [00:55:38] So, once we get to based, we drop hugging. [00:55:41] And take in into a context. So, in still has most of the context. Instead of having Hugging Face as a starter base, in has the context face is a startup based. [00:55:51] So, it still has some of the context, and it can still… the model can still predict in with some sort of confidence. [00:55:57] Right? So, this sliding window allows us [00:56:00] to continue… continuously keep sliding across the… across the input sequence of tokens that we want to compute the publicity over. [00:56:07] or the dataset that you want to public… compute the publicity over, and [00:56:11] still keeps some context about what is going on. So, it can still predict in with us [00:56:17] a good amount of confidence, and that can still give us a workable perplexity fit. [00:56:23] So this is about publicity. [00:56:25] That would be it for GPT-2. [00:56:29] So, any doubts regarding GPT-2? [00:56:33] Any questions? And we'll take a short break, uh, 5 minutes, just if anybody has any questions. [00:56:39] So, please feed me across. [00:56:43] So, one point is that, whenever the concept of publicity, right? The… [00:56:51] Yeah. [00:56:52] If the value is low, it is high confident, right? The value is predictable, correct? [00:56:56] Yeah. [00:56:57] So, how that is different than temperature, right, which is setting it, right? Because temperature also to define the randomness. [00:57:02] Right? The content. So how both are correlated with each other's? [00:57:09] So, temperature is completely different. Temperature is a sampling parameter that is used when we want to get the next token. [00:57:12] Correct? [00:57:15] Uh, in a given sequence. When we have a set of logits, we apply a temperature to it. A lower temperature allows the model to be more deterministic, a higher temperature allows the model to be more imaginative. [00:57:24] Correct. [00:57:26] So what happens is, uh, when we are taking the, uh, the softmax of the logits to give the probabilities, right? So we have a set of logits which the model outputs. The logits are basically [00:57:36] So, let me come back to this slide, right? [00:57:39] Mm-hmm. [00:57:40] Right? So, these are the logics, right? These are the softmax of the logits gives us these probabilities, right? This classifier gives us logic, which is the LM head, this gives us the logits, and then we apply the softbox. [00:57:50] To give us the probability. Uh, so before applying the softmax, we apply the temperature. [00:57:53] Okay. [00:57:55] That allows us to control how much of a distribution is going to happen in the model. So, if we apply a higher temperature, [00:58:02] what happens is that we bring a lot of the values closer together. So if you have ever looked at the softmax or the [00:58:09] Mm-hmm. [00:58:12] like, the softmax or the soft plus function, so to speak. [00:58:16] Uh, just let me bring up a pen. So, the graph for that looks something like this, right? [00:58:21] between 0 and 1. This is where it is, right? [00:58:22] Yeah. Yeah. [00:58:25] So, for any values that are very high, [00:58:28] or very low, right? [00:58:29] Mm-hmm. [00:58:30] the output, the probabilities are very close together. [00:58:35] Right? [00:58:36] Okay. [00:58:37] So… so if my input, uh, the input to the softmax function, [00:58:41] Where the output from the classifier, which is basically the input to the softmax, that gives me this, right? Softmax. [00:58:47] Here, if my output from the classifier was a very high number, [00:58:51] for most of it. Then, [00:58:53] The… the output… [00:58:55] probabilities would be very closely grouped together. [00:58:59] But, if I were to divide these values, the input here, if I were to divide these x by a large number, let's say 10, [00:59:06] for temperature, I would get these values here. [00:59:09] Right? Somewhere here. [00:59:11] So the difference between the [00:59:14] probabilities would be quite high, and it would be easier for me to sample [00:59:18] And understand whichever, uh, output the model wants a higher probability for. [00:59:24] Okay. [00:59:25] So, this is where temperature comes in. This value that we divide by in order to get these, shift these to here. [00:59:28] Mm-hmm. [00:59:30] This is called temperature. So if temperature were to be very low, [00:59:35] like, for example, if temperature is 0.1, and my output is here, [00:59:38] then it becomes 10 times more, and it gets you. [00:59:41] So, if temperature is low, the model is more deterministic. We get a lot of probabilities close together, and uh… [00:59:47] the model behaves in a more deterministic way, or a more greedy way, so to speak. [00:59:48] Okay. [00:59:50] But if the temperature is very high, we get a large amount of candidate tokens. [00:59:53] Okay. [00:59:55] Right? So, these large amount of candidate tokens that we can sample amongst. So, we can then pick out [01:00:00] some token with a higher confidence. And that gives the, uh, the model the ability to get [01:00:06] different outputs, or, like, a vast variety of more imaginative outcomes. [01:00:10] So this is how temperature is applied. [01:00:11] Got it. Got it. [01:00:12] Yeah. [01:00:13] Right? Uh, when we are, yeah, when we are doing perplexity, uh, we just take the output from the classifier. We do not apply any temperature. We just take the softmax of whatever the classifier outputs. [01:00:21] Mm-hmm. [01:00:23] We don't apply any temperature whatsoever. So, temperature is 1.0 in that case. [01:00:28] Because dividing by 1 gives us the same output. [01:00:29] Got it. Yeah, yeah. [01:00:31] Right? Yeah. [01:00:32] Thanks. [01:00:51] So, any other doubts? [01:01:02] Yeah. [01:01:03] Uh, hello. Uh, for Fabric City, uh, what is the range of values? [01:01:08] So, the perplexity can be really high as well. It could be infinity also, if the model is, like, see, you can see the probabilities here. The only constraint is that the output from the model, the sum of all the probabilities can be at max 1. [01:01:21] Because you can't have a sum of probabilities higher than 1. That's the only constraint. So, if the model were very uncertain, [01:01:26] then, uh, you would get a very large amount of very small probabilities, taking logarithmic outputs of [01:01:32] small numbers would give you quite high values. [01:01:36] in the negative, since it's anywhere minus, so you get… [01:01:39] quite large values, and publicity would be very… it could explode. [01:01:42] So, uh, value increasing means the model is performing… [01:01:47] poorly. Yeah. [01:01:50] Yeah. [01:03:05] Um, so if there are no further doubts, we can move on to the T5, or text-to-text transfer transformer. [01:03:11] Uh, this is an encoder decoder code. [01:03:14] So, we should be taking this up next. If there are any doubts regarding GPT-2, [01:03:17] Please be told us. [01:03:26] Okay. So, let's move on. Uh, let's get to Transformer, uh, encoder-decoder transformer models, where, uh, we have [01:03:34] T5 as the example for this kind of model. T5 stands for text-to-text Transfer Transformer, so… [01:03:40] It's, uh, that's, that's how this was named. [01:03:46] Uh, what does, uh… [01:03:48] transfer learning mean. So, we did take a look at this as well. So, during transfer learning, we have a pre-training, uh, we do pre-training for a model, and uh… [01:03:59] This pre-training is done on an unlimited dataset. [01:04:02] And this knowledge is then [01:04:04] basically self-supervising, because the unlabeled dataset is, in itself, a self-supervised task. The next token is going to be the [01:04:12] label for the current, and so on and so forth. [01:04:15] So, uh, we then fine-tuned the model on a smaller label dataset for any specific downstream tasks. [01:04:22] So, we have looked at GPT, Bert, Albert, Proberta. All of you… all of which use the same transfer learning methodology. [01:04:27] Password learning is basically learning across a large domain, or unlabeled set of data, and then transferring that information to a specific supervised task. [01:04:36] on a label dataset for some specific downstream task that we want to achieve might be sentence classification, might be text entailment, might be anything. [01:04:44] Right? That we could want to use. So… [01:04:48] We have looked at a lot of examples of transfer learning. [01:04:51] But we have a variety of models. [01:04:53] Albert was a light version of Bert Robertaly optimized BERT. BERT was, uh, [01:05:00] bidirectional encoding representations from transformers, and GPT was the decoder-only model that we just had a look at. [01:05:07] So, all of this have the same training, same training methodology of pre-training, followed by specific fine-tuning. [01:05:12] But D5 is a text protects encoder, decoder, transformer model. It stands for, uh, like, text-to-text basically means that it takes text input and produces text output. [01:05:22] Transfer is the methodology that's used for training, transfer learning, hence the word transfer. [01:05:27] And transformer is the architecture. It's an encoder-decoder-transformer architecture. Hence, T5. [01:05:33] So, it was first proposed in this paper, Exploring the Limits of Transfer Learning with Unified Textwriting transformer in RFL et al. in 2019. [01:05:42] It has, uh, in its base format, it has 220 million parameters. [01:05:46] $110 million for the encoder, and about $110 million for the recorder. [01:05:52] So, what is text-to-text? [01:05:54] What does text-to-text mean? Well, it tries to pose every NLP task as a text input, text output task. [01:06:00] So, for example, neural machine translation can be, uh, [01:06:03] convert it into a text-to-text task as… [01:06:06] translate English to German. English text to German text. [01:06:09] Let us take, for example, this vector. [01:06:15] If I were to give it the input, translate English to German, tell me about Qualcomm. [01:06:19] This is a text-to-text input, so this is text input, and I get output as [01:06:25] Zalin see me Uber Qualcomm, which is telling me about Qualcomm in German. [01:06:28] So, this is what is meant by text-to-text. I give a text input, I get a text out. [01:06:34] Even any sort of, uh, [01:06:36] integral task is expressed as a text-to-text task, here for T5 in the model. [01:06:41] The other one is linguistic acceptability. [01:06:44] where we have a sentence, and if it is acceptable or it's not acceptable, right? [01:06:49] Um, or similarity, so to speak, of, like, linguistic acceptability is whether the, uh, [01:06:54] text is linguistically correct, it's grammatically correct or not. [01:06:58] So, I can just give COLA sentence, which is, uh, caller linguistic acceptability sentence. I give a text, and I ask it to give whether it's acceptable or it's not acceptable as the output. So, I give [01:07:10] color sentence text, and I get acceptable as the output or not acceptable as the output, that this is grammatically correct, it's linguistically acceptable for the language that I want to. [01:07:18] I want the model to predict, or… [01:07:20] It's not acceptable for the language that I want to call it. Semantic textual similarity is basically whether two sentences are similar to each other or not. So, I can do STS sentence A sentence B, [01:07:31] And I get output as, uh… [01:07:34] I explored why, which is an output, which is how similarities are. This is basically, uh… [01:07:39] a number, X is a number, and Y is a number, so I can get something like 2.3, 4.5. [01:07:44] Something like that. That gives me the, uh, similarity between sentence A and sentence B. [01:07:49] So, ST sentence A, sentence B. This is another text-to-text task. I gave it a text input, and it gives me a text output, like, [01:07:56] the number here to the text output. It's given as a string. So, this is semantic textual similarity. Text summarization, where I give summarize text, and I get the summary as the output. [01:08:05] So this is… [01:08:08] This is the T5's methodology, the training methodology. Express every NLP task that it wants to train for as per text-to-text task. [01:08:14] And then we do transfer learning. Pre-train on a large unlabeled dataset, and fine-tune on a specific downstream task. [01:08:23] So, this is what T5 is all about, translate English to German, that is good, that thus is good. [01:08:29] call us into the course is jumping well, not acceptable. [01:08:33] The course is jumping well, well, it doesn't make any sense, because the post cannot jump. [01:08:37] force is not a person. [01:08:38] So, that's why it's not acceptable, it's grammatically incorrect. [01:08:42] Um, the next one. [01:08:45] STL sentence 1, the rhino graze on the grass. Sentence 2, a rhino is grazing in a field. These are quite similar, so we get 3-point data. [01:08:51] Uh, summarize state authorities. [01:08:55] Right, correct? [01:09:01] Summarized creatures, dispatch emergency crews Tuesday to survey the damage after living within Mississippi. Well, 6 people were hospitalized after a storm hit the area. [01:09:09] That is the summary of that input text. So… [01:09:13] defies rise to its… tries, uh, to… [01:09:16] express every NLP task, whether it be semantic textual similarity, whether it be linguistic acceptability, whether it be translation, neuronic translation, or any of the other tasks that it trains for, it tries to express them as a text-to-text task. [01:09:29] take a text input and convert it into a text output. [01:09:33] Uh, well, how's it different from GPT and what? [01:09:35] So BERT is an auto-coding model, it is an encoder-only architecture. GPT is an autoregressive model, it is a decoder-only architecture. There's no encoder in GPT. [01:09:43] And the inputs are directly decoded. [01:09:46] Uh, T5 is an encoder-decoder architecture. It is both auto-encoding and autoregressive. The encoder for T5 is similar to BERT, which is autoencoding, and this is, uh, this produces the context for the input. This context is then given to the decoder, which is autoregressive. That is, it produced one token at a time. Initially, it's given the output from the encoder and a beginning of sentence token. [01:10:06] And the decoder then runs autoregressively. [01:10:09] generating one token at a time, till it, uh, encounters an end of sequence. [01:10:14] So, that is the difference between GPT and BUT on, uh, and T5. [01:10:19] on the highest level, that the architecture for DeFi is an encoder-decoder architecture, which consists of an autoencoding encoder and autoregressive decoder. [01:10:27] So, the decoder is… can be thought of as similar to GPT, but there is a key difference that we'll get to shortly. [01:10:33] And BERT is similar to the encoder part of T5V. [01:10:38] So, what is the training objective for T5? Well, [01:10:41] the T5, uh, authors made of the paper, they came up with a new training objective. Bird trained for mass language modeling, GPT trained for causal language modeling. T5 trains for span corruption. [01:10:51] we will have a look at what's called spiral corruption is, and how it is achieved shortly. [01:10:55] So, recapping the architecture, we saw that BERT [01:10:59] For every given token, it can take… [01:11:01] context from the past and the future, because [01:11:04] For any given word, when we want to create a context, or the meaning of that word depends on the context from both the past and from the future in a given sentence. [01:11:12] The meaning of a previous word might change based on something that happens later in the sentence. [01:11:17] So, that is… that is the reason why BERT needs to take a look at both the past and the future. [01:11:22] to give a contextual token embedding for each token. [01:11:26] Whereas for OpenAI's GPT, or any other, uh, [01:11:29] autoregressive decoder-only model. It cannot have a look at the future, because it is an autoregressive model. At the current time step, it does not know what the future is. [01:11:37] So, it cannot take any hint from future context. It does not know what the future is going to look like. [01:11:43] So given… even it can only have the context about even and anything before it. [01:11:47] For E2, it can take the context E1 and E2. [01:11:51] from… for EN, it can take the context E1 to EN to predict EN plus 1. [01:11:56] And so on and so forth. So… [01:12:00] Now, coming to T5's architecture. Well, it consists of an encoder stack, and then a decoder stack on top of it. So, as we can see, the encoder stack is similar to BERT. It generates a contextual token embedding. [01:12:09] For every token input, it generates a token, uh, embedding output. So, for even, we get D1, for E2, we get T2, for EN, we get TN. And this token embedding is contextual, because it can take context from both the past and from the future of the entire input, because it needs to [01:12:23] generate a context-based [01:12:25] Text embedding, uh, vectorized text embedding for every input text. [01:12:29] Once this text embedding is then passed to the decoder stack, it injects its context that it knows about the input, [01:12:37] into the decoder. This is also done [01:12:40] from the, uh, from past and from the future. [01:12:42] But once the decoder is running, the decoder can only inject its own [01:12:48] context in an autoregressive manner. That is, it can only inject from [01:12:51] past to the future. It cannot know anything about the future, because when the decoder stack is running, it generates one token at a time. [01:12:58] So for every given token, it can only have the context of itself and anything to do the past of it, not anything in the future. [01:13:05] this part, where the encoder stack is [01:13:08] Injecting its context into the decoder stack, this is called cross-attention. This is the mechanism of the encoder-decoder architecture where, uh, whereby the encoder gives its context to the decoder. So, for example, given an English sentence, [01:13:22] The encoder can encode the meaning, the contextual meaning that it understands of this English sentence in the tokens T12D and in the output vectors T1 to T, and corresponding to tokens E1 to EN. [01:13:33] Once this is done, [01:13:35] this entire set of embeddings, uh, for every token, [01:13:38] And, uh, one-to-end is supplied to the decoder stack. [01:13:42] And a beginning of sequence token is due. [01:13:44] is supplied as well. The decoder understands, okay, I need to begin [01:13:47] my current sequence, and the context for the output is present in T1 to TN. [01:13:53] So then, it produces an output of a D1. And then at the second autoregressive decoding step, we get VOS and T1 as the decoder input, along with the entire encoder output. [01:14:05] So, again, it understands its current context in [01:14:08] in only the past to future things. So, you can only get the current context from anything that is past. [01:14:14] So, it gets its current context from all the tuples it has generated till this step. [01:14:18] And then, also from the encoder's outputs themselves. So, this decoder, for example, could be running in general… in German language. This could be trained to understand or generate the context for German language. [01:14:31] But it takes an input of the contextual embedding of English language. So the encoder needs to [01:14:35] come to a common, representable token embedding, whereby any of the input languages that this model has been trained for, the encoder [01:14:44] and generate a common token embedding. [01:14:47] that contains the context of all of those languages that might be present in the input. This context, then, [01:14:54] can then be supplied to the decoder, which [01:14:56] By which the decoder can then specialize in a particular target language, or maybe a set of target languages, and generate based on whatever the input context was. [01:15:05] So this is a very high-level overview of how T5s, or encoder-decoder-transformer architectures in general, achieves [01:15:13] something like neural machine translation, or any of the other tasks that they might be trained for, like polar or linguistic susceptibility, or semantic textual similarity, and so on and so forth. [01:15:23] The key component here being the cross-attention mechanism, whereby the encoder is allowed to supply its context to the decoder stack. [01:15:31] So, let's take a look at the T5 training task, which is spank corruption. What does span corruption mean? Okay, we take the input, and we randomly sample and drop out 15% of the tokens in a sequence. [01:15:41] And these… these tokens that got dropped out are consigators. All consecutively, like, not all 15% consecutive, but a set of them are consecutive. [01:15:48] And all consecutively dropped off tokens are replaced by a single token representing the span. [01:15:53] So this is the span corruption. We give an input paragraph, we sample randomly 15% of the tokens, [01:16:00] and drop them out. And for all of the sequence of tokens that are consecutive, that are dropped out, we replace it with a single token. [01:16:08] called a special Sentinel token. [01:16:09] Right? We have a set of Sentinel tokens, like T5, I think, has about 99 Sentinel tokens, so within a span, we can corrupt, uh, like, within a given, uh, paragraph, we can corrupt about 99 spans. [01:16:18] These internal tokens are special tokens, and they are added to the model vocabulary so that the model can understand what these are. [01:16:24] I'm given this input, the target sequence is all the dropped-out spans. So, let's take a look here. [01:16:30] for example, thank you for inviting me to the party last week. This is the input. We randomly sampled 15% of the tokens here for inviting and large, for example. [01:16:38] And we drop them up. So, for inviting is replaced by Sentinel token X. [01:16:42] Last is a place, since it's only a single token, this is also considered a spam. If we were also, uh, for example, dropped out, then last week would [01:16:49] be represented by buying, but your only last has dropped out, so we represent that by some point of NY. This is the model input. [01:16:56] Given this model input, the training task is to predict [01:17:00] this sequence of output from the decoder. [01:17:03] X for inviting, Y last, and a special sentinel token, red. [01:17:08] Great. So, the target sequences, all the drop-down spans of tokens, delimited by the corresponding sentinel tokens in the input sequence. [01:17:16] And an extra essential token at the end of the target sequence. [01:17:19] The model is trained to predict the target sequence, given the span-corrupted input sequence, and this is done by using cross-entropy loss. [01:17:26] What is prostate property loss? Well, given an observation PI, and this is the true observation, and QI, which is the predicted observation from the model, uh, cross-entropy loss is defined as, uh, the summation across all of the target [01:17:39] tokens, I, PI logQi. [01:17:42] the summation of minus PI log Qi. [01:17:45] So, this is the, uh, like, how likely, uh, the output is. So, [01:17:50] For example, here, uh, for the first output, given this input, for first output, [01:17:56] P of I for X is going to be 1, and P of I for [01:18:00] every other token in the vocabulary is going to be zero. [01:18:03] Because the true observation should contain X here, right? So the probability of X occurring here is 1, [01:18:09] And the probability of X occurring anywhere, like, for any other token occurring here is all zero. [01:18:14] That is the true. [01:18:17] localities, like, the true probability. [01:18:19] And QI is going to be the logits from the decoder. So, [01:18:23] we take the IX logic corresponding to X, [01:18:26] as the QI observation here, and we take, uh, and then we compute minus sigma [01:18:32] PI log Qi. Well, since PI is 0 for all the tokens, all of them are ignored, and it's 1 here, so it becomes minus sigma log Qi. [01:18:39] For token X, or… [01:18:41] So, minus sigma log QX, for example, for open X. [01:18:45] And that would be the loss that is computed for the output protection of token X. So let us take a look at the rating. [01:18:51] So, let's take… this was the input sequence. [01:18:53] Thank you, X, for, like, me to your party, 5 weeks. And the target output should be X, given BOS output should be X. [01:19:02] for inviting. Why? Last. [01:19:04] So, given the decoderStack token, BOS, [01:19:08] but should be X, given the start to open X, output should be 4. [01:19:12] given all of this, ARPU should be inviting, given all of this, outputshould be wire, given all of this, output should be last. [01:19:18] given all of this, output should be Z, and so on and so forth. And at the end, we get an end-of-sequence token. [01:19:24] So, given Z, output should be end of C. So these are the target tokens that we want, and these are the decoder input tokens. [01:19:30] And based on this, we compute, based on these target tokens, we compute the cross anchoropulus. [01:19:35] Uh, when we provide the decodest factor token, the logics, ideally, should contain 1 for X and 0 for everything else. It doesn [01:19:43] real-life situation. So, uh, it would be some very high value for X, and some very low values for all of the tokens in an ideal output file training. If it is not, we [01:19:52] learn the, uh, we calculate the loss using the cross entropy formula, and then we backpropagate this loss throughout the layers. [01:20:00] So, the decoder output states are passed through the language modeling head, which is basically the classifier, which gives us the logits. [01:20:06] And these con… these hidden states are connected to logits, and then these logits are then, uh, [01:20:13] calculated the loss on, uh, based on the, uh, [01:20:16] cross-entropy formula, and the loss that is calculated is then back-propagated throughout the net. [01:20:22] So, how does it work? We take this, we take the target sentence, [01:20:27] we run it through the encoder. The encoder generates a set of encoder outputs. These encoder outputs at the decoder start to open, is provided to the decoder. It, let's say, generates X, or some other token. [01:20:37] We calculated the loss on this, and then backpropulated. [01:20:40] Then, given decoders not open in X, and the encoder's outputs, the decoder should ideally predict 4. [01:20:48] then given the decoder… the encoders' outputs and decoder start to open, X for the decoder should predict inviting, and so on and so forth. So… [01:20:55] But at every given stage, the decoder's inputs are the encoder's outputs, and all of the previous tokens. [01:21:02] And… [01:21:04] technical distracted. So… [01:21:08] Any questions regarding the training task and how the T5 model is being trained? [01:21:13] our cross-entropy loss, or any other architecture questions. [01:21:31] Okay. [01:21:32] Just give me a second. [01:22:01] Yeah. So then… [01:22:03] This was the first step of transfer learning. [01:22:06] Where we do the unsupervised [01:22:08] Uh, pre-training, and then coming to the supervised fine-tuning. [01:22:13] So, the T5 model was fine-tuned for, uh, various down-screen tasks. These downstream tasks were tasks such as text classification on the general language understanding dataset, blue, [01:22:23] or take summarization on the CNN or Daily Mail datasets, where, given a large amount of articles, it had to generate a summary or the abstract for each of those articles. [01:22:33] And… [01:22:36] Uh, question answering on the Stanford question answering dataset, and neural machine translation from, uh, [01:22:41] the WMT English to French and German to Romanian, uh, like, English to French, English to German, and English to Romanian data science. [01:22:48] So, this, uh, the fine-tuning loss was also calculated using the same cross-entropy function that, uh, was used during the unsupervised pre-training. [01:22:57] Uh, and for each of the fine-tuning, uh, like, for each of the fine-tuning datasets, uh, the… [01:23:04] the fine-tuning was done over 2PAR18 epochs, or steps, and a checkpoint was stored every 5,000 steps. [01:23:09] Then, over the validation dataset for each of these, uh, training tasks, [01:23:13] the checkpoint amongst these 5,000, uh, like, 5… amongst these, uh, checkpoints stored at 5,000 square, the checkpoint whichever gave the highest validation data set results, were used to publish the [01:23:23] results in the paper. [01:23:25] So, uh, for example, um, take… [01:23:29] taking, uh, the glue of text classification task, for example. We would take the blue training set and the clue, uh, and the glue, uh, validation set. [01:23:38] On the glue training set, we would train the model over 2,62,144, 2 power 18 steps, and then the checkpoint would be stored at, uh, every 5000 ships. [01:23:48] Then, we would run the validation datasets over all of these checkpoints that were stored. [01:23:52] Uh, and whichever checkpoint gave the highest, uh, validation [01:23:56] Oh, it's good. FN Sport, which was… whichever was highest, that checkpoint would be published, and that would be the output of, uh, the T5 model on that particular, uh, [01:24:06] that same task. And such as, uh… and as such, all of these training steps, or fine-tuning steps, were repeated for each of the trainee decides. [01:24:15] on each of these tasks. So, desk summation, training, data sets, neural machine transmission training datasets, and question-answer training datasets. [01:24:21] different, like, the model was fine-tuned for each of these tasks, and a lot of checkpoints were stored, and from each of these checkpoints, whichever one gave the best result on the validation dataset, [01:24:30] on the corresponding validation dataset is the one that was used to package the results. [01:24:35] Um, well, what was the dataset used for, uh, training, like, the unsupervised pre-training? [01:24:41] It was the C4, Colossal Clean Prod Corpus. It consists of a common crawl web extracted text dataset, but it has been cleaned. [01:24:48] Uh, the cleaning process basically involves, uh, deduplication, which is removing all the duplicate content, discarding incomplete sentences, and removing any noisy content or offensive content from the, uh, [01:24:59] common crawl web extracted, uh, text dataset. [01:25:02] So, after the cleaning, this dataset, the C4 colossal green product called this dataset, uh, is about 750GB, or 745 GB in size. [01:25:13] Uh, the T5-based encoder and decoder details are as follows. Each of these base encoder and decoder architectures have 12 layers, and have… and use 768 as their embedding dimensions, and each of them have 110 million parameters, for a total of, uh, [01:25:27] 220 million parameters for the T5 base model. [01:25:31] So, this is, uh, this is the overview of the training steps here. [01:25:37] We pre-train the model. We have a bird-based sized encoder and a bird-based sized recoder transform, under 10 million parameters each. [01:25:45] Uh, we train this on the C4 dataset, which is the colossal Crawl Clean dataset. [01:25:50] Um, so these… these pre-training, uh, was done over 2 part 19 steps, and it processed approximately 34 billion tokens. [01:25:57] Uh, and then it was fine-tuned over, uh, the glue or the CNNDM, or the SPOD, or the Super Glue, or the WMD14, ENDE, which is the English to German dataset, or the English to French dataset, or the English to Romanian dataset. Each of these, uh, tasks were, uh, [01:26:14] they basically trained over 2 power 18 steps, and every 5,000 steps, we save a checkpoint. [01:26:21] of the model, like, a snapshot of the model. [01:26:23] And evaluate all of these checkpoints on the validation dataset for this particular, uh, [01:26:29] training dataset, fine-tuning dataset. [01:26:32] And whichever one gives the best performance on that particular fight during dataset, uh, fight during validation dataset, is then used to publish the results for. [01:26:43] So, during the fine-tuning, about 17 billion tokens. During each of the fine-tuning, [01:26:48] So for glue, we process about 17 billion tokens. For CNN, we process about 17 billion tokens, and… [01:26:53] So, for each of the fine-tuning tasks, for each of the domain-specific tasks that we are fine-tuning for, we process about 17 billion tokens, and the pre-training itself is done over 34 million. [01:27:03] And the output is checkpointed and showed every 5,000 steps during the fine-tuning. [01:27:08] And that model is then validated to get the best possible results across all of the validation datasets. [01:27:14] And those are the models that are published. [01:27:16] So, this is about, uh, this is the overview of the transfer learning, uh, methodology that is used by the T-Size, uh, authors to train the T5 model with text-to-text transfer transformer. [01:27:29] So, any doubts so far? [01:27:39] Okay, um, so let us take a quick overview of the previous architectures, BERT and GPT. So, as we had seen, for BERT, we have the inputs. We do the input embedding, we add the positional encodings, and then we perform the bidirectional edits. [01:27:53] followed by a fleet forward. All of this is done by a number of times, like, the number of layers, which is, let's say, 12 for bird-based. So, there are 12 such layers, and the output is then the contextual token embeddings. [01:28:04] For GPT, we take the input, uh, we add the… we embed the inputs using the token embedding matrix, and as a lookup table, and then we add the positional encodings, pass this to multi-attention. Here, it is causal, or masked multi-attention. [01:28:18] And then the feedforward, and this is done for the number of layers, uh, as well, that are present in the GP2 model, which varies by small, medium, large, and so on and so forth. [01:28:27] And then, at the end, we have a language modeling rate. This language modeling head gives us logits, from which we can sample the next token for the given input. [01:28:34] So, this is the encoder-only architecture, this is the decoder-only architecture. [01:28:40] Uh, so how does the encoder decode architecture work? Well… [01:28:42] This is how it works. So… [01:28:44] this encoder [01:28:46] this architecture is similar to BERT. You can see that there is no difference. So, BERT is… [01:28:51] the bird encoder is basically the same as [01:28:54] this encoder here, right? The T5 uses a very similar encoder architecture. There is a difference. [01:29:00] Well, you can see that there are no positional embeddings that are added to the input here. Bird uses sinusoid position embeddings, whereas T5 uses a special kind of personal embeddings called Trulative position embeddings. [01:29:11] These are added during the attention computation itself. We shall have a detailed look at how this relative attention, uh, embeddings are calculated and how they are added, and what they are trying to achieve. [01:29:21] So, bird-based had a personal encoding, which was basically a scientific position embedding that was calculated by interpolating the dimensions two at a time, and applying the sine and cos theta formula to each of those, uh, basically a rotation. [01:29:35] for every single, uh, position. [01:29:39] Well, and GP2 used pre-trained position encodings, which were trained during the model train. [01:29:44] Over here, neither the T5 encoder, neither the T5 decoder use the, uh… [01:29:49] the pre-trained or a fixed set of sinusoidal positional embeddings. [01:29:54] Instead, they rely on preventative attention, or relative distance attention, so they calculate the relative distance between the attention itself and inject the positional information based on the relative distances during the attention competition. [01:30:06] So, uh, we take the inputs, which is, uh, for example, translate English to German, tell me about Polycom, would be the set of input tokens. [01:30:14] Uh, we calculate the, uh, input embeddings. [01:30:17] So, we take the input embeddings here, and this is, again, similar to BERT and GPT-2, which is basically, we take the token embedding matrix as… and look up the tokens that we have as inputs, uh, and we get the input embeddings. [01:30:29] This input embeddings is passed into the encoder layer, or the set of encoder layers. [01:30:35] Uh, at each layer, we do the bidirectional multi-attention. Again, this is bidirectional because it needs all of the contexts for the input. [01:30:41] Like, it needs to understand what is the sentence trying to say. Uh, so, it needs to understand the context of each of the tokens and generate a contextual token embedding for [01:30:50] Each and every token that is given as an input. [01:30:53] So, that is the reason that it uses [01:30:55] bidirectional engine, similar to both. [01:30:58] Then it is passed on to the feedforward, and this is done for the number of encoder layers. So, it's 12 for the T5 base model, so there are 12 encoder layers, so this is done 12 times, and then the output that we get at the end of those 12 layers [01:31:11] is going to be a set of token embedders, uh, for each of the input tokens. So, if you had any input tokens, we have end token embeddings here. [01:31:19] Then, the decoderly comes. We pass these encoder outputs as inputs to the decoder. [01:31:26] Along with the decoder start token. [01:31:28] So, why this shift at right? Well, shifted right because if you have a look here, [01:31:32] Essentially, this output is shifted right as an input, right? Decoder start token shifts all of the output. [01:31:39] Right, right. One step right. [01:31:41] Hence why we call it input-output shift right, but it's basically, we supply the, uh, decoder [01:31:48] start open here, along with the encoder's outputs in here. [01:31:52] And for the decoderStar to open, uh, and all of the subsequent tokens, like, it predicts a particular token based on the decoder start token and the context that is injected. [01:32:01] from the encoder, using the encoder's outputs, which is the per-token contextual token embedding. [01:32:06] So, we inject context from the encoder into the decoder. [01:32:10] Yep. I'm the, uh, supplier decoder decoder start token, or BOS. [01:32:18] This BOS open is then used to predict the next token, which is an appended, and then the subsequent is BOS plus the token predicted, and again, [01:32:24] This is constant. The encoders outputs remain constant. The encoder runs once, and the decoder runs autoregressively multiple times till it encounters EOS. [01:32:33] Right? And, well, encoder uses two types of pitch. So, there is a key difference between this and the GPT-2 encode. The GPT-2 encoder used mass multi-hand attention. [01:32:43] But here, we have something called cross-attention as well. [01:32:46] Whereby we eject the encoders' context into the decoder, along with its own context. [01:32:52] So, for example, if it were a translation task, and the encoder has encoded English sentence, it gives the encodings, or the context for the English sentence, [01:33:01] as a context to the decoder. And then the decoder is recording in German, for example. So, [01:33:07] The decoder has context for German. [01:33:10] because of this self-attention. [01:33:12] And it also gets the context from English because of this cross-attention. [01:33:16] So, that is how it is able to [01:33:18] get a detailed [01:33:20] view of what it should generate next, based on the… whatever the encoder has understood about the English language, [01:33:27] generates the context. This context is injected into the decoder. Decoder understands the German language, and also decodes, uh, one by one. [01:33:34] open by to open, and teach understanding the current German context and the total English context, and sees what it should generate next, so that it can translate the given English sentence or the English context into German. [01:33:47] So, here, this is… [01:33:49] This is the key component, which is multi-head attention, cross-attention. [01:33:53] And this is similar to GPT-2's mass multi-gate attention, there is no difference here. We just have a causal mask, because this should only be left to right. [01:34:02] Because as the decoder is decoding tokens one by one, it can only look at the tokens that it has previously decoded. It cannot look at the future tokens. [01:34:09] So, which is why this is vast. [01:34:11] And, uh, then the encoder supplies its, uh, context using the multi-head cross-attention. This is not masked, because [01:34:19] the complete context must be looked at by the decoder in order to generate the next decoder, so… [01:34:25] This multi-head attention, or cross-attention here, is not matched. [01:34:28] The only attention in the encoder and decoder that is masked is the decoder's self-attention. [01:34:33] So, that's why this mask is optional. The mask is applied on this, [01:34:38] But not on this. And then we have the feed-forward disability to this, and it goes on. And at the end of all of the decoder layers, which are 12 decoder layers that we have, [01:34:47] Uh, we supply the outputs from the decoder. [01:34:50] into a classifier, or the language modeling head, and we calculate a set of logits. These logits are sampled, and we get the next token from them. [01:34:57] And it continues till we get… till we encounter a… [01:35:02] end-of-sequence EOS, and at that point, the decoder starts. [01:35:05] So… [01:35:07] Uh, any questions regarding how this encoder-decoder architecture is working? [01:35:12] And what each of these components are trying to achieve? Any questions? [01:35:25] I have one question. [01:35:26] Yep. [01:35:28] So, in the decoder, for its own self-attention, we are masking the forward tokens, right? We are only looking at the tokens that we have already used, like [01:35:38] previous tokens. But for this cross-attention, why are we not doing the scene? [01:35:46] Okay, um, so, for example, let's take I wanted to translate, tell me about Qualcomm into German, right? [01:35:51] So, I give the decoder BOS, and a beginning of sequence, and I gave it all of the English context. [01:35:58] If I were to use a causal mask at this point, [01:36:01] I would mask all of the English context out, right? Because I only have one token that… that is beginning of sequence, right? [01:36:07] So, the decoder has to mask out everything. [01:36:09] Which is, at this point, everything in English is to the future. [01:36:13] And the other thing is the number of tokens that a decoder predicts would also be different from the number of tokens that the encoder has processed, because the same sentence could be much longer in English versus German. [01:36:24] So, uh, that is the reason why, like, you can see, you can think of it this way, like, we need to understand what the English sentence means in order to convert it into [01:36:34] the German variant, right? It's not always a one-to-one mapping. It's not like, tell me about Qualcomm has… tell has a singular word in German, me has a singular word in German, about, has a singular word in German, Wolkov has a singular word in German. [01:36:46] Great. Uh, there could be… [01:36:48] a set of words that could be a single word in the target language, in the neural machine translation task, for example. [01:36:54] So, it needs to understand the context of what is being said in the sentence, in the input sentence. [01:37:00] Uh, and it needs the entirety of that context to generate the next group. It is a context-aware generation. [01:37:06] So, that is the reason why this cross-attention is not much, because if you were to mask this, then we would essentially be reducing or not… [01:37:14] essentially not providing any context to the decoder itself from the encoder. [01:37:19] The decoder must use its current, uh, like, within its current context, it must know [01:37:24] that, okay, I have predicted this so far, so I have, like, let's say, tell me about. So, Urzalan V. [01:37:30] I think, is the equivalent for that. So, in German. [01:37:35] So, the decoder must look at German as, okay, I have said this so far, what do I say next? [01:37:40] What do I say next? Comes from the English context. Okay, I need to say about Qualcomm. So, it must know what happens in the future in English in order to generate what [01:37:49] happens next in German. [01:37:51] So, uh, I hope that clears it up a bit. [01:37:55] Yeah, thanks. [01:37:59] So, yeah, that is the reason why cross-attention is never masked, and it basically is injecting all of the context that has, uh, that the encoder has produced. [01:38:08] into the decoder, so that the decoder can give a more context of edge generation of the next book. [01:38:17] So, let's look at the input tokenization. [01:38:20] Uh, T5 uses a Unigram algorithm, or called sentence piece, uh, so… [01:38:25] Setence fees is a library that implements a Unidram tokenization algorithm, uh, as its tokenizer. [01:38:31] So, say, uh… [01:38:33] we basically call this as a sentence-based tokenizer, so… [01:38:36] T5 uses unigrams, intense piece tokenizer as its tokenizer… algorithm, tokenization algorithm. It addresses the fact that not all languages use space-specific words. Since T5 is a textbook X transformer, and it supports a variety of [01:38:48] languages as its input. We cannot always rely on spaces to separate words. For example, so far, the tokenizers that we have looked at, BP and WPM, [01:38:57] always first go word by word. So words are defined by spaces, but that's not true for all languages. [01:39:04] like, for Chinese and Japanese, we do not have any spaces subjected words. So, [01:39:08] We cannot rely on spaces to separate words in order to tokenize words and sentences. [01:39:14] So, sentence piece addresses this, and it treats the input as a raw stream of bytes. [01:39:19] Which is basically a raw stream of characters that it includes all the spaces, uh, and everything. [01:39:24] Everything that is input to the model is used as a set of characters that it needs to use to tokenize. [01:39:30] And it then uses a unigram algorithm to construct the appropriate vocabulary. [01:39:34] This is done by computing a set of unique words, which is the initial vocabulary. Uh, so all the symbols that could be used to write the words. [01:39:40] is the set of initial vocabulary, which is much larger than the BP. In BP, we only use the unique symbols that could be used to write the [01:39:47] or entirety of the input columns. [01:39:50] Here, we use all of the symbols, all combinations, all possible symbols that could be used to write words. [01:39:55] Uh, without any, uh, without any conspirate on them being unique. [01:39:59] So to speak. Like, for example, NP and WordPiece, we only use the letters, like, hugs. [01:40:05] or H-U-G-S, and we only use singular letters. Here, we could use multiple letters as well. So, HU could be, uh, initial, uh, part of the initial vocabulary. [01:40:14] instead of just HUG and S for the word hugs. [01:40:17] So, compared to BP and WordPiece, Unigram works in the opposite direction. BP and WordPress start with a very small initial vocabulary, or minimal vocabulary, and try to merge and learn merges [01:40:27] Uh, so that they can expand their vocabulary and tokenize the sequence better. [01:40:32] The Unigram algorithm used by sentence piece works opposite. What it does is, it takes [01:40:37] a very huge additional vocabulary, which consists of a set of all symbols that could be used to write all of the words. [01:40:44] And then it… it… [01:40:46] starts dropping tokens from this vocabulary till it reaches the decide vocabulary site. So it starts from very large vocabulary and goes down to a smaller vocabulary. BPE and WordP start from very small vocabulary, and learn merges [01:40:58] And the frequency, like, these modules are scored by… differently, uh, in different elements. WordPiece uses, uh, [01:41:04] normalized frequency, where it uses the frequency of the pair divided by the frequency of each of the parts of that pair. And BPE, for example, uses the frequency of the pair [01:41:12] it's a… there's no normalization there. [01:41:14] So, Unigram calculates a loss instead, so it drops tokens, and which token is going to be dropped is decided by whichever one causes the minimal loss over the course. So, it's basically looking for explainability. [01:41:25] What Unigram algorithm tries to do is, okay, what is the minimum number of tokens? I can explain this entire corpus with. [01:41:32] And in that process, it starts with a very huge, uh, set of vocabulary, and then it computes a loss. [01:41:38] The lowest loss means that dropping that token has the lowest impact, or the lowest impact at the explainability of that corpus based on these, uh, input tokens. [01:41:46] So, for each symbol in the vocabulary, the algorithm computes how much the loss would be. [01:41:50] If that token were to be dropped, and then tokenized… and then tokenization were to be done across the entire purpose. [01:41:56] So, if the loss is the… like, whichever token gives the lowest loss, if the loss is higher, that means that we use a lot of expressibility by dropping that token, so we should not drop it. If the loss is very low, it means, okay, we can [01:42:07] you know, drop this token, and still not be that [01:42:11] badly, uh, hit in terms of explaining the purpose. [01:42:14] So, the symbols that have a lower effect in the overall loss are the best candidates, and the symbols that havethe largest loss [01:42:21] We don't do that. So, a subset of these symbols, controlled by a hyperparameter, uh, P, which is the, like, there is a threshold below which we cannot [01:42:31] below which we should not drop, so that is controlled by a parameter to the organizational algorithm. And once that is in, we can drop the symbols, as in when we want, till we get to the vocabulary size. [01:42:43] So, what is the loss function itself? The loss function is, uh, so this is the loss function here. [01:42:49] So, given a word W, or consisting of T1 to TN, [01:42:53] Uh, which is the tokenization of this word, at NT1 and TN are all tokens in the vocabulary. [01:42:57] and sees the corpus of all the words in the given document. [01:43:00] So, the loss is computed over all the words, or all the tokenizations of that particular purpose of words. [01:43:07] W belongs to C. [01:43:09] Frequency of W times [01:43:11] a negative log probability of them. What is the probability of double? Probability of W is probability of T1 to TN, which is basically the multiplication of the probability of each of these tokens. So P of T1 times P of T2 times P of T3 times P of TN. [01:43:25] And how… how is the probability of each of these tokens? [01:43:28] Well, it's basically the frequency of the token divided by the total amount of, uh, frequency. So, given T in the vocab, in the current vocab, the frequency of… the sum of frequency of all the tokens [01:43:40] is used as a normalization factor for the current frequency. [01:43:43] frequency of the current token divided by the sum of frequency of all the tokens in the, uh, particular, uh, vocabulary. [01:43:49] is the probability of that token. [01:43:52] And for any given, uh, word W in a purpose of, uh, in a corpus of word C, we do a loss function as sigma W belongs to C, frequency of W, which is the [01:44:02] frequency of that set of characters occurring in that entire, um… [01:44:06] times the negative log of the probability. [01:44:13] So, let's take an example and look at how this loss function is used to drop tokens. [01:44:18] So, let's take the same example that we have taken so far. HUGPUGPUN, B-U-N, and H-U-G, all occurring, like, H-E-G occurs 10 times, PUG occurs 5 times, PU occurs 12 times, PU occurs 4 times, and HUGS occurs 5 times. [01:44:31] The splits in the vocabulary, uh, initially will be, for example, uh, if I were to take the BP algorithm, I would just use the unique [01:44:38] alphabets here. So H, U, G, P, [01:44:41] N, V, [01:44:44] and S. That would be my initial vocab, right? [01:44:47] But here, that's not the case. I take all possible, uh, sets, or all possible sequences of, uh, letters that could [01:44:54] be possible in this… [01:44:56] compass. So, of course, we have the singular letters, H-U-G, but then we can also have HU and UG as our initial vocabulary. [01:45:04] And then we have PUG, and PUG already is present, so we just take PU in this case, and then PUN. [01:45:10] PU is already present, so we just take UN in this case, and then we have B and UNN are already present. [01:45:16] So we just take BU, and UN is there, so we don't take that. And similarly for hugs, we have HUG as the new sequence. [01:45:23] Like, it's U… and HUG was already there here, but HUG, the three-letter word, is a new sequence that we learned from this. [01:45:31] And it could be GS, and also, uh, UGS could be another, uh, input, like, initial vocabulary part. [01:45:39] So, we calculate the frequencies for each of these slopes. So the frequency for H, for example, is 15. Frequency for U is 36. Frequency for G is 20, and so on. So we calculate the frequencies in the initial frequencies of the initial vocablet. [01:45:52] And then we consider, for example, we take the word P, UG in the corpus that we have. [01:45:58] And we consider possible tokenizations. Well, we could, given our initial vocabulary, we could tokenize it as P-U-G. [01:46:05] Or they could tokenize that PU and G, or we could tokenize it as P and UG. [01:46:10] But then we have to calculate a score for each of this. The score is basically a frequency of W times this, right? So, uh… [01:46:17] the score is, like, the probability of each of these tokens. [01:46:21] occurring individually. So P times, uh, like, the probability of P times probability of U times probability of G, which is [01:46:28] this way, right? The probability of that tokenization. [01:46:31] is 0.003, which is the sport. And the probability of this tokenization is 0.007, and this one is also 0.007. [01:46:39] Um, so at this point, we can take either one of these tokenizations to be the best, because this has the highest score. [01:46:47] But since they have the same score, we can take the first one, or PUG, as the tokenization. [01:46:51] And then tokenizing all of these other words, and calculating the tokenizations with the highest scores for each of them. [01:46:57] they get this as the tokenization for the corpus. So, hug is tokenized as HUG. [01:47:02] Which is the signal token that we have, called is tokenizes PUNG, partners tokenizes PU and N, B, partners tokenizes B-U-N-N, and HUGS is tokenized as H-U-G and S. [01:47:13] So, now we need to start dropping tokens, right, from our vocabulary in order to get down to 10 tokens. How do we draw? Which token do we drop first? [01:47:22] Well, we compute the loss of dropping the token. [01:47:26] For example, if we were to drop the token, um… [01:47:29] like, the initial loss over this purpose is… [01:47:32] going to be here. And then, if you were to drop the token, let's say PUG, uh, UPNUG, so… [01:47:38] we can drop PU, right? Because we have another tokenization with the same score for the word PUG, so we can drop PU. [01:47:45] So, when we drop TU, the only tokenization that is going to be affected is going to be PEUG and PEUN. [01:47:53] And both of these, uh, the tokenization score remains the same, and third is removing the PU token will give the exact same loss that we have currently. So the delta loss is zero. [01:48:01] However, if we consider removing HUG from the vocabulary, uh, then the tokenization for two [01:48:08] of the words changes. So, for hug, it becomes H-U-N-G, and for hugs, it becomes H-U-N-GS. [01:48:15] And the scores for HUG would be, like, [01:48:18] the score that you have created. For example, here, we should be probability of HU times the probability of G, and probability of HU times the probability of GS would be so on and so forth. [01:48:26] So, now, when we compute the new laws, the loss is 183.3, so that's a positive increase in the loss. The delta is 23.5. [01:48:32] So, we cannot drop HUG, but we can drop PQ at this stage, because it gives the lowest [01:48:38] loss. Like, it… the loss remains exactly the same. There's no delta. There's no increase or decrease in the loss. [01:48:44] So, at this time, we drop PU. [01:48:47] I'm the new vocabulary becomes this, and then the new frequencies become this. And we… [01:48:52] continue calculating the losses, and uh… dropping tokens. [01:48:57] I still really vocabulary is exactly what, let's say, 10 in this example, so we'll keep talking. [01:49:03] Any doubts regarding this? [01:49:15] So now, let us take a look at how the tokenization actually works. This was about how the tokenizer was trained. [01:49:21] Or how do we get the vocabulary and the scores for each of that? So, if you look at the T5 tokenizer here, [01:49:34] So if you take the look at the tokenizer for this model, the T5 model, you can see that every single token has a score attached with it. [01:49:42] So, this token has a score of this, this token has a score of this, this token has a score of this. [01:49:49] So this is basically the, uh… [01:49:51] The negative log probability of each of these tokens, right? So, for P, like, in order to calculate this, [01:49:57] We can take the negative block probabilities and add all of them in order to get the negative log probability of the tokenization. So, this is what we want to maximize, right? When we are tokenizing. So, instead of just giving the probability P , they give the negative log probability [01:50:10] So that is what we have here. [01:50:13] Which is all given here. Thank you. [01:50:18] So, we have a score for each of the tokens. [01:50:20] And when we want to tokenize, how do we do it? Well, we use the Bitter B algorithm, which is a dynamic programming. [01:50:26] Basically, we'll take a detailed look at it. [01:50:30] So, let's say I were to tokenize, tell me about Qualcomm, right? [01:50:34] How does that work? Well, um… [01:50:36] First, we replace all the spaces with the special character, uh… [01:50:41] Which is basically an underscore that we want to do. [01:50:43] So, we replaced this with a special character, a special space that, uh, the tokenization algorithm actually understands, right? For example, this character is the character that we replace all the spaces with. [01:50:54] And then we initialize the best course that we can have, and then we compute the best course based on all of the possible tokenization. So, given any input sequence, we compute all the possible tokenizations that could be [01:51:07] Uh, done for that input sequence. [01:51:09] And we compute whichever one of those tokenizations gives the best score. [01:51:14] The one that gets the best score, the part for that is stored. Path of I stores the length of the final token based on the best path to text item. [01:51:22] So, it stores the best parts. [01:51:25] And how is… what are the best scores? Well, whichever one has the best [01:51:28] So, if the spore is greater than the current base score, we update. [01:51:31] Whenever we find a new tokenization, that gives us a better score. [01:51:34] The scores are initially initialized to minus infinity, so initially they all get a higher score. [01:51:40] But when we are exploring more and more tokenizations, all possible tokenizations of an input text, [01:51:46] Uh, as and when we find a better tokenization for the text, we update the score. If the score is higher than the current score, we update that score, and we update the best path to the current one. [01:51:55] Once we have all the best scores and best paths, [01:51:58] With us, we then proceed forward. [01:52:01] So, let us take a look at this, right? Let's take the input, tell me about Qualcomm. [01:52:07] So, how does this work? So, we start at the index underscore. So, this is the start index and end index of any given tokenization. So, given [01:52:14] If we take only… if we start tokenizing here, and we only tokenize underscore as a single token, [01:52:21] the score from minus infinity becomes minus 2. [01:52:23] And then, if we take underscore T, [01:52:26] as, uh, for example, if you take underscore NT as the token, then the score becomes minus 8.4, and TELL, so we continue doing it. So, we get here, right? [01:52:36] And how long do we do it? We do till whatever is the maximum token length that we have when we have the… when we have trained the tokenizer. So, let's say after TELL, [01:52:46] Um, there is no better [01:52:48] Um, like, the maximum length is over. We cannot go any further, because there is no token that is longer than [01:52:55] underscore T-E-L-L underscore something in this entire corpus. [01:52:59] So, then we go, uh… [01:53:02] So then, we take the next token, right? The next character, and we start tokenization there. [01:53:08] So, from underscore, we tokenized as much as possible. [01:53:13] So, underscore… underscore T, underscore E, underscore T-E-L-L, T-E-L-L, right? So, this was the maximum that we could go, no further tokens were found in [01:53:19] put any of this. Underscore was not in the purpose. Underscore TELLM was not in the purpose, underscore T-E-L-L-M-E was not in the purpose, and so on and so forth for everything. So those cores for all of these still remain minus input. [01:53:32] Then we start that tokenization at the next character. [01:53:35] If, at any given point here, we find that starting a tokenization here gives us a better score than what was here, we would update that. [01:53:42] So, starting the tokenization at T, we could only update the score for one of these, right? We could not go for any of, like, we did not update here anyway, because the best personal minus 8.4, but in this row, [01:53:52] We could update only this one. There is no other tokenization that is possible. [01:53:56] So then, starting at E, we do… we explored all possible tokenizations, and we update the scores. Same with L. [01:54:04] L, and then starting with the next underscore, we see, okay, underscore, [01:54:09] underscore is a token, and the score is here. Underscore M is not a token, but underscore ME is a token, and the score can be updated, because that was minus infinity, so we can update this. [01:54:19] And so on and so forth. So we keep doing this. So, this way, we keep populating the entire table. [01:54:25] for all possible token admissions. At this point, we have the best course. So, if I were to take this token, the best score is this. [01:54:33] Best quote is this, best quote is this, and best quote is this. [01:54:36] So, now I can backtrack. I know that starting [01:54:39] At this point, my best score [01:54:42] engage with this one, right? [01:54:44] So, underscore Qualcomm becomes one tool. [01:54:48] Then, starting at this point, I know that getting to this point gives me the next best point. [01:54:53] So, underscore about. [01:54:55] So, that is the next booklet. And then, going back, underscore me is the third token, and then from here to the end, which is underscore tell, is the last one. [01:55:04] So this is how we backtrack and retrieve the token ID. [01:55:07] And since we have gathered them from end to start, we need to reverse the entire set. [01:55:13] So, since we, at this point, have tokens underscore Qualcomm, underscore about, underscore me, underscore tell. [01:55:18] So we just reversed them to get underscore tell, underscore me, underscore about, underscore Qualcomm as the input tokenization. [01:55:24] for tell me about Qualcomm. And this is how the, uh… [01:55:27] sentence-based organization using the Unigramm algorithm is done. [01:55:31] at any given point, we want to minimize these scores. So, we want to ensure that, uh, sorry, maximize the scores. That is, we want to ensure that for all poss… amongst all possible, uh, tokenizations, [01:55:41] of tell me about Qualcomm. The one that we use has to have the best possible score. [01:55:50] Any doubts regarding this? [01:56:02] So, in the code, all we are saving is this, uh, this last best course array. [01:56:08] And, of course, the length of the token, so that we know where to backtrack to, because given this index, I need to store this index, okay? The best tokenization score [01:56:15] Starting here, I get from… by starting from here. So, that's why I store the best, um… [01:56:21] the best paths. This… these red arrows that we see here, [01:56:25] stored in the best parts array, and the best sources stored here. So, starting here, I know that I get the best score [01:56:31] of minus 39.6. If I start 9 tokens behind, which is minus 36.2. [01:56:36] if I… then starting here, I get the best score of minus 26.2. If I start about 5, uh, 6… [01:56:42] steps behind, which is at underscore about. [01:56:45] And then I know that if I start, uh, like, if I start post steps behind, I get the best score of [01:56:50] on this one, minus 19.1 for the 2. Underscore ME. [01:56:55] And, uh, from here, I just want to start, because I know, okay, there's no further thing, I just need to go back. So, here, the best. [01:57:00] So, in the best parts array, the value corresponding minus 39.6 would be 9, the value corresponding to 86.2 would be 6. [01:57:09] the value corresponding 19.1 would be 3, and the value corresponding this would be 4. [01:57:14] So that I can backtrack it. [01:57:17] So, this is all about the, uh, input tokenizations. [01:57:22] on, uh, how the tokenization is done for any given input text using the VW algorithm for PGNEGRAM sentence piece. [01:57:28] tokenization. This is a very different way of tokenization compared to BP and Gold Peace. [01:57:34] Uh, so WordPiece and VPE both start with the smallest possible purpose, and try to get to the largest possible purpose. [01:57:41] Uh, whereas here, we start with the largest possible purpose, and then [01:57:46] try to keep dropping tokens with minimal loss of explainability over the entire purpose till we get to our desired vocabularies. [01:57:54] And all of that is done based on this loss margin. [01:57:57] So… [01:58:00] Are there any doubts, um, regarding this tokenization algorithm, and how it is achieved? [01:58:14] Okay. [01:58:23] This… [01:58:52] So, coming to the input representations for DT5 tokenizer. [01:58:56] Uh, the input to T5 is tokenized using, uh, [01:59:00] the sentence-based model and the B2B algorithm to give us a list of tokens. These tokens are then converted to token emittings. This is similar to how it is done for BERT as well. [01:59:09] Fortify the vocabulary size is 32,128 tokens. Uh, the token embeddings are stored as a lookup table, and they are indexed by the token ID. This is similar to BERT and GPT, so given a sequence of tokens, let's say 100, 200, 200, 400, we would take [01:59:22] the, uh, the open embedding methods, and use it as a lookup table, uh, and we would get the index 100, 200, 300, and 400. Uh, so these vectors, the vectors corresponding to these indices on that matrix, [01:59:34] And give us that open embedding. For the… [01:59:38] for the input to the T5 model. [01:59:40] So, once we have these input embeddings with us, um, we do not add any positional embeddings. This is a major difference compared to BERT. [01:59:48] Which adds BERT and GPT, which add a fixed position embeddings, BERT uses sinusoid position embeddings, while GP2 uses [01:59:56] Um, trade position embeddings. [01:59:59] T5 does not. T5 uses relative positional emittings during the self-attention computation itself. It basically bases the position on the relative position of any two tokens while it's computing the Q times k transpose matrix. [02:00:14] So, this is a very major difference between T5 and BERT and GPT models. So, it uses relative [02:00:20] positional embeddings, uh, during attention to inject the positional knowledge about all of the tokens. [02:00:26] Whereas BERT and GPT and other models use [02:00:30] specific position embeddings for every single position of the entire… [02:00:35] So, here the input representation basically just consists of the token emittings. Unlike BERT and GPT, which has token embeddings, and the added position embeddings as well. [02:00:45] So, uh, what is relative position? Well, relative position is how far two tokens are from each other. So, absolute position, like everybody else, use [02:00:54] GPT-2 at BERT, so for position 0, you have a position vector. For position 1, you have another one, and so on, until all the L-1 positions that are present in the, uh, context width of L. So, there is a position vector for each of these positions in, uh, [02:01:08] in Burton GPT-2. Bird is a sinusoidal approachability to calculate it, to be able to trains these vectors. [02:01:14] Here, uh, we just use the relative position. So, for example, for a given token i equals 6 for a query corresponding to the token i equals 6, [02:01:22] Uh, when we are calculating the key. [02:01:24] Uh, the Q times K transpose dot product, we would say that, okay, the key, uh, corresponding to token 1 is at a relative distance of minus 5, and the key corresponding to token 9 is a relative distance of 3 away from the [02:01:38] current position, or the current queries position. [02:01:40] So this is the relative position. This is the concept that is used by the T5 tokenizer, uh, sorry, T5 [02:01:47] encoder and decoder models to inject position, uh, [02:01:52] positional knowledge about the tokens. Instead of using absolute positions, they use relative positions. [02:01:57] Um, so how this works, we will take a look at, uh, in detail, but it is… it would require, uh, the next session, because, uh, if I are starting in a way, we just have about 10 minutes or so remaining. [02:02:09] So, that, uh, would basically spill over to the next section, uh, next session, so we would… [02:02:15] that's… it would be better to continue. So, to have better continuity throughout. [02:02:19] So, if there are any questions, um, anything, uh, that [02:02:23] Remember the last time you can have a question and answer, or, like, a QA session for 10 minutes or so. [02:02:46] Now, if there are any questions regarding the tokenization algorithm, the, uh, the sentence-based algorithm that we just had to look at, and how the tokenization is being done, [02:02:55] Um, the code workflow that we looked at, so… [02:03:02] Any questions at all regarding that? [02:03:16] I think we need some time to understand [02:03:22] Uh… yeah. [02:03:23] this a little bit. That's the… [02:03:25] Sure, sure, sure. Um, if there are no questions, then we can end the session here. Uh, we should… [02:03:32] Yeah. [02:03:33] Yeah, one thing on the same note, so if you could update the slides, then it would really help [02:03:40] Yeah, sure, sure. I'll share the slides, you know, what we have covered. I'll share the slides. [02:03:44] Yeah, I can find it for the last session as well. [02:03:47] Uh, sure, sure. I will update this one. [02:03:50] And, uh, Jenny, badge manager. [02:03:52] So, yeah [02:03:54] I raised the one request, I think it's almost one month back. [02:03:59] Uh, we didn't update that the material in there. [02:04:03] Right? Because I'm not able to access the material in a week-wise manner. [02:04:07] As we, in an old, uh, site, we are able to… [02:04:11] And a new one, I think, it's very tedious right now to follow up. [02:04:16] what is ongoing. [02:04:19] So, can you please work on that? [02:04:25] generate a batch manager. [02:04:29] Uh, yes, yes, sure. [02:04:31] You are saying, yes, yes, but nothing has happened. [02:04:35] Actually, I have escalated this issue to the tech team already. [02:04:38] Ma'am, but I think, uh, [02:04:41] current scenario, how much time it will take. [02:04:45] It takes only if someone fixed that, it takes only 1 or 2 days. [02:04:50] And last one month, I'm continually saying that. [02:04:52] Yes, I understand. [02:04:53] That is not the right way. You are using a root key IIT name. [02:04:56] Please make sure what you are doing. [02:04:59] Yes, sure. Just give me some time. I have already connected to the tech team. [02:05:04] I don't know why the delay is happening. [02:05:05] Thank you. [02:05:07] You are able to access the recordings by, uh, via the… [02:05:10] Recordings, recordings I am able to access. [02:05:14] Okay. [02:05:15] But, ma'am, last year… [02:05:16] lesson I missed due to some urgency at my… [02:05:18] Uh-huh. Uh huh. [02:05:19] But right now, I need to only need to… [02:05:24] Is this…? [02:05:25] Um, in a messy state right now. [02:05:26] Okay, okay, okay. [02:05:27] So I need to, first of all, go to the… all the recordings. [02:05:30] Mm-hmm. [02:05:31] And then I'll search the document, so… [02:05:33] Trying to understand and makeup. [02:05:34] Okay, so you only need the documents. [02:05:36] Yes, I only needed documents, material. [02:05:39] If you don't, uh, able to… [02:05:41] do that now, so provide me the weekly wise deep learning. [02:05:44] whatever we covered right now. Four glass, one. [02:05:48] Okay, okay. [02:05:51] Yeah, sure. I will be connecting you after the class. [02:05:54] Yeah, hello ma'am, I have one query regarding also that. [02:05:58] Yes, please. [02:05:59] Yeah, frequently, your changes in time of the Sunday classes, just like earlier it was a 10.30, and now it is 11.30, because it is very difficult to manage these times. [02:06:09] Uh, actually, the timing is confirmed by the professor only. [02:06:14] So… [02:06:15] Yes, ma'am, uh, but you need to understand about our Asuna. [02:06:17] I understand. Okay, I will discuss it with the professor once. What was before your… [02:06:24] What was the before timing before this? [02:06:25] Yeah, it was, uh, yeah, it was a 10-10, uh… [02:06:30] Um, it was dead. [02:06:32] Okay, so you all want… [02:06:33] And before that, it was also a 9.30, it was just, uh, gradually increasing the time, and now it is 11.30. [02:06:42] Okay, okay, I understand. Okay, we… I have to discuss that with the professor only. [02:06:45] Yeah, sure, yeah. Thank you. [02:06:49] Oh, thank you. Yes, yes. [02:06:50] Hello? Ma'am, uh, for me, for me, my recordings are not visible, like, uh, for the past one month, the previous class recordings, I'm not unable to see. [02:06:59] Uh, it's the link, or anything? [02:07:00] You're not able to see the recordings. [02:07:02] Yeah, I still, uh, URL, like, IATR.swichens.com? [02:07:09] Haha, yes. [02:07:10] Uh, yeah, but I am unable to see the previous class recordings, like, for the past one month recording, I'm unable [02:07:16] turn this in the front of them. [02:07:18] Okay, okay, okay, okay, wait. [02:07:21] Okay, I'll also connecting with you after the class. [02:07:24] Okay, thank you. [02:07:26] Okay, thank you. [02:07:28] Okay, is the class over? [02:07:32] Uh, yeah, uh, we, we, uh… [02:07:36] That was just 10 minutes remaining, so, uh, we have to take up a more complex topic. [02:07:40] That's why I think it's ready to spill over to Tomorrow's lecture. [02:07:44] Okay, sir. [02:07:47] So, thank you, thank you, Patrick. [02:07:53] No, no. [02:07:55] Thank you.