# 05 2026-05-24 GPT Theory course: Module 4 — Generative AI & LLMs module: Module-4-Generative-AI-LLMs date: 2026-05-24 type: transcript video_url: https://personal-learn.armco.dev/files/_Recordings/Module-4-Generative-AI-LLMs/05_2026-05-24_GPT_Theory.mp4 --- [00:13:20] It doesn [00:13:21] And produces a per-token text embedding for each of the tokens. So given a set of tokens, a sequence of tokens, even E2 to En [00:13:29] But, looks at the context for each of the tokens and understands what those tokens or words mean and then generates a contextual textual embedding. So T1 is the contextual embedding for the event, the input E1 for input E2, it is T2 and for EN it is tn [00:13:46] So, the input to BERT is a sequence of tokens. This is obtained by the tokenization of the input text. We had taken a look at the tokenization last week and how the workpiece model tokenizer works [00:13:58] And [00:13:59] Sorry, yeah. And, the output, as we had taken a look last week, was a set of token embeddings. Each token gets its own token embedding. So given a sequence of n tokens, the output from BERT consists of n token embeddings [00:14:13] And each of these representations are conditioned on both the past and the future context simultaneously, because as we had looked at last week, the meaning of each word [00:14:23] changes based on what we already know and what we might know in the future of it. Like, yeah, for example, shoot could have different meanings in different sentences, and some of those contexts might be derived later on in the sentence as well. So that is the reason that it needs to take a look at both the past and the future [00:14:42] condition the output on both the past and the future. This is called bi-directional attention. We had taken a look at this as well, and how bidirectional self-attention works. So this was a quick recap of BERT. Now, we should look at a new task today, which is causal language modeling. [00:14:58] What… what does causal language modeling mean? Well, in a sense, it just means predict the next word based on a given set of inputs. So given a set of inputs, for example, thou shall, the output would be the next sequence that the next word that might occur in this sequence. For example, thou shalt not [00:15:16] can be a completion. So this is causal language modeling. Why causal? Because it is cause and effect. [00:15:24] As we can see, why is it cause and effect? Well, because the output the output gets calls from the inputs. So the inputs affect the output, so to speak. That's why it's causal. [00:15:38] This is the task that GPT-2 aims to perform. [00:15:43] We can think of it as this very simple black box. We give it a set of inputs, thou shall, and these are input tokens, for example, or input words, as we can think of them. Thou shall, this is a trained language model, and its task is only to predict the next word, so thou shall not, for example. [00:16:00] Hey, the outputs here, all of these outputs are basically all the words that the model knows, and the probabilities are each of the probability of each of these words that the model knows being the next word in this current sequence, or the next token in this current sequence [00:16:15] And this is the task that is causal language modeling. So in this case, a well-trained model might give a weightage of, let's say, 40% to not. And this is the highest amongst all of the predictions. So we say, okay, thou shall not is the next word in the sequence. And then [00:16:35] In order to proceed any further, we must then append not to this sentence and give thou shall not as the input to get the next word. So this is causal language modeling. [00:16:46] Well, what is GPT? As we have already looked at, GPT stands for Generative Pre-Trained Transformer. Generative means that the model intends to generate the next word in a sequence. That is why the word generative and pre-trained. Well, it's already trained on the language modeling task [00:17:01] Right, causal language modeling task that we just looked at. And transformer, it uses the transform architecture, but it uses the transformer decoder architecture. This is the decoder only model. It does not have any encoder like BERT is an encoder-only model, it does not have a decoder. This model is a decoder-only model. It does not have any encoder [00:17:18] And it uses the transformer decoder architecture, and that's why it's a generative pre-trained transformer. [00:17:23] So the the GPT-2 model has in the in its largest form has about 1.5 billion parameters. The model preceding it, Gpt. One only had about 117 million parameters. There are various versions of Jp2 available, small, medium, large, and extra large. The Excel variant is 1.5 billion parameters [00:17:46] In size [00:17:47] And… so, what is GPT? Well, GPT, as we already discussed, uses transformer decoder blocks. It is a decoder only architecture, and it is an autoregressive model. What does autoregressive mean? As we just looked, for example, given the sequence thou shell, we predict the next word not. Now, how do we get any further [00:18:05] Well, we append not to the inputs, and then we run it entirely again, and that's thou shall not is the next input, and then it gives something. It gives some other word as the next word in this sequence, and that word is then appended, and it goes on this way. So you can think of this as a time series [00:18:21] Right? So auto regressive models and time series is a type of statistical model that we use for analyzing the time series data and predicting what the next data might be. It is based on the idea that at any given point in a time series, the current value can be explained by all of its previous values. [00:18:40] That is autoregression. So, auto regression basically means that given all previous values, the current value is explained by everything that occurs in sequence before it. and that is what we are doing not then generating a next word and then appending that word and then generating a new word [00:18:56] keep appending N plus 1 words. So, that's why auto regresses. It outputs one token at a time, and each token is then appended to the input sequence, and this combined sequence becomes the next input at the next model iteration, right? So this is [00:19:12] high level overview of what Gpk does, and to see it in action we can see, for example, if we give the model the Gpt 2 model an input of recite the first law of robotics. Let's say it might give us an output that says a robot [00:19:27] may not disobey its master tax. Like, it will continue doing that. But yeah, so this is how GPT works in a black box, this is the black box. This is the input, and we get a single token at every iteration. We take this token that is out, that the model outputs [00:19:42] And then we append it to the input [00:19:45] Right here. And then we generate another token, and we append it, and we generate another token, and we append that, and it goes on this way. [00:19:53] So, this is GPT-2. The differences between GPT-2 and BERT on, so GPT-2 is a decoder-only model. BERT is an encoder-only model [00:20:01] GPT-2 uses autoregressive decoding, but is an auto-encoding model. It generates a contextual text embedding for each token that is supplied to it as an input. What is the training objective? Well, the training objective for GPT-2 is causal language modeling. We shall take a look at what causal language modeling is in, like. [00:20:20] In detail of what causal language modeling is, on a very high level, we have just seen what causal language modeling is. It's just predicting next word in a sequence. And BERT uses masked language modeling, right? The first training task for BERT was masked language modeling, where we were masking a few of the input tokens, and then we were asking BERT [00:20:38] to generate an encoding for all of the tokens, including the master token, and predicting what that mask token might be based on… based on the output embedded, right? So, and then evaluation tasks, we have Glue and perplexity. Glue is the general language understanding, evaluation [00:20:56] And Bert had various other ones like Squad and named entity recognition, and so squad is a Stanford question answering data set. We had taken a look at all of these in detail last week. [00:21:08] So these are on a very high level, these are the major differences between GPT-2 and BERT. [00:21:14] So taking a look at the architecture itself, BERT is a transformer encoder. It… it produces contextual embeddings based on both the past and the future. So if we look at token E1, you can see it gets all of its context from everything that occurs after it, E2 to EN [00:21:29] For E2, it gets its context from both the past, that is even, and the future, that is E3 to E, and EN gets all of its context from everything that occurred in the past [00:21:39] And it goes on for each of these transformer layers, and we get a so TRM transform layers. This is one transformer layer, this is another transformer layer, and so on [00:21:47] like, we had taken a look last week, birth has 12 layers. [00:21:51] And after all of these are run, we get a contextual token embedding T. One, T. 2 Tn, where this embedding is basically the vectorrepresentation of this input word by like that is also containing its context right [00:22:06] In the sequence, we inject context from both the past and the future. In GPT-2, we only take any token only gets its context or its meaning from any token that occurs previously, because it is an autoregressive model, right? Like, any current value can only be explained by its past values, even in, even as we are speaking [00:22:26] We can only construct or compose sentences by appending tokens based on what is already present, like what we are already speaking. Once we start speaking a sentence, we keep appending new words based on the context that we receive [00:22:43] We cannot predict what we might see in the future and then inject context from the future as well. So similarly for GPT-2 as well, we do not take any context from any of the future tokens. So even as just itself, E2 has even an E2 because the sequence still E2 is known at the time step of E2. [00:23:00] E3 will have E1, E2 and E3, and EN has all the past tokens for its context. So each token only gets its context from anything that occurs to the past [00:23:11] This is the major difference between the masked language modeling and the causal language modeling. So we can only take context from the past [00:23:22] So, let's come to the training task. The first objective, which is causal language modeling. This is unsupervised. We just give the model a sequence of unlabeled tokens even to U1 to UN. And the objective is to maximize the likelihood that for every given token [00:23:39] Based on the preceding tokens, the output of the model is that given token. So, for example, for every token UI, we give it UI minus k to UI minus one, where K is the model's context. So as we've taken a look last week, the every model has a limited context. The context length for [00:23:57] Birds, for example, those fighters. The context lane for GPT-2 is 1024. It cannot take an input of more than 1024 tokens. So [00:24:06] Within the context. Let's assume for this to currently be 0. Like, we start from the beginning of the context to the current token, and everything fits within the context. So from U0 to UI-1, we supply the input, and the output from the model is… should be UI, right? [00:24:23] So we use the log likelihood function to maximize this likelihood, right? The L of U for any given corpus of U from U1 to UN is that the submission for every summation across all of these tokens [00:24:38] Given all preceding tokens in the sequence, the probability for the current token should be the highest. This is the causal language modeling task. And this is the training task, basically. So for UN, we only give u1. There is no preceding context. For U2, we are the model needs to be trained in a manner such that the probability of U [00:24:55] Based on U1 should be maximal, and then for U3, we give u3, the probability for U3 based on U1 and U2 should be maximum for UN. Assuming U1 to UN fit in the context [00:25:07] the probability of UN should be maximal based on the inputs u1 to Un minus one. This theta represents the model's parameters that we are training right the network parameters, so to speak. [00:25:20] So, this is the loss function for the causal language modeling task. We need to maximize this likelihood for every single token in a given corpus. This is unsupervised because there are no labels, there are no human labels. The model trains itself on predicting the next token in a sequence. [00:25:37] Also, given every preceding token, it should be able to maximize the probability for the current token. [00:25:43] So [00:25:44] This is achieved by the transformer decoder with n layers as follows. Initially, we give so this gather is basically the the table lookup that we had seen last night by WE is the token embedding, and Wp is the positional embedding. So WE is a size vocabulary size, which is the maximum number of tokens that this model can recognize. So WE for [00:26:06] GP2 has 50,257 embeddings. So, because the vocab size for GPT-2 is 50,257 tokens, that is, it can recognize 50,257 unique tokens [00:26:18] For each of these tokens, there is an embedding [00:26:20] And given corpus U [00:26:23] of input tokens we gather the embeddings, gather the embeddings basically for, let's say, this token, this corpus consists of token 100, 101, 200, 201, and 301. So, we take the 101st, the 100 and 101st [00:26:40] or token embeddings from this table WE. We use this as a lookup table. So we simply index it by 100, and whatever embedding that is, we take that, then 101, we take that 200 [00:26:50] 201, and let's say 301. So, that is the token… the input token embedding. This gadder layer is doing just that, getting the input token embeddings from the model weights and for the given input tokens. And then we add the positional tokens. The positional embeddings. The positional embeddings help a model recognize what position each token occurs in, because [00:27:11] The meaning of a word can also change based on where it occurs in a given sequence, and the model needs to distinguish between each different unique position. So the position 0, 1, 2, and so on and so forth. So for every position in the context [00:27:27] In the supported context, let's say 1024 is this model's context lens. So we have 1024 unique position embeddings. 0, 1, 2, 3 until 1023. So we add, let's say current in our current example, we only had five input tokens, right? So we had 100, 101, 200, 201, and 301. So that's 5 input tokens [00:27:46] So for each of these input tokens, we give the pollution embeddings for the 0 width index as for the first token as the zeroth index and the second token as WP of 1 and 2 and 3 and 4. So 0, 1, 2, 3, 4, the first five [00:28:01] positional embeddings from this are taken and added to the gathered token embeddings [00:28:08] And this forms the input to the transformer [00:28:11] And then, once we supply this input to the transformer, we run all of the transformer layers for model. The the number of layers depend based on the size of the model like Gpt 2 small has a lower number of layers compared to extra large, which has a much larger [00:28:27] stack depth for the transformers. So let's say we have n layers, then we run the transformer DIH transformer for I belongs to 1, 2 N. [00:28:36] And the output from this is then [00:28:38] multiplied by the… basically, we have a linear layer, which is also known as the LM head layer. We will take a detailed look at what this layer is, what it does, and this gives us an output of 50,257 tokens [00:28:53] Which give us the probability for each of those tokens being the next token in this current sequence. So this is how we achieve the causal language modeling. Are there any doubts or any questions so far? [00:29:09] Is there any model which uses both the encoder and decoder of the transformer? [00:29:14] Yeah, there are a few models. For example, the MT5, the T5, and the Baht models both use encoder and decoder both [00:29:24] Can you please explain further function? [00:29:29] Yeah, so this is basically taking the model weights which contain the token embeddings for each of the unique tokens. So as I said, let's say the vocabulary for this model is 50,257 tokens. So this is a table [00:29:45] Of 50,257 encodings, right? And each of them are indexed by the token index. So, the first token is token 0. So, the first token embedding in this matrix, so to speak, is that [00:29:58] And then the first opened. The second token is basically… the second token in the vocabulary is at the second index. The third token is at the third index, and so on and so forth. So let's say if I had 100, 101, and 200 as my token IDs, as my input, basically, once I tokenize the input, I got the token IDs as 100, 101, and 200 [00:30:16] So how do I get the token embeddings for each of this? I just treat this as a lookup table right? And I take the index 100, the vector at the index 100, and the vector at index 101, and the vector at index 200 based on my input token IDs. That is what this gather function. It is [00:30:32] gathering the token embeddings from a lookup table. [00:30:37] I just wanted to say that on the question about the other… [00:30:41] Uh, transformer-based models, which are, uh, encoder-decoder type. [00:30:47] So, we will be covering T5 in the next weekend. [00:30:51] Okay, so we will be… as we have been, uh, covering encoder only, now decoder only. [00:30:58] And then we'll be talking about encoder-decoder, which is T5 in the next turn, okay? [00:31:05] Okay, ma'am, thank you. [00:31:06] So I just wanted to add here, yeah, please carry on. [00:31:09] Yeah, so [00:31:12] Does that [00:31:13] Or gadolia [00:31:21] Okay. Any other questions? [00:31:29] Okay, so we can proceed then. [00:31:34] Then, coming to the second training task, which is once the model has been pre-trained, we can fine-tune it. This is called discriminative fine-tuning. It is discriminative because each layer in this transformer now, like over here, we are using the same learning rates for each of the layers. We wanted the entire model to learn equally. [00:31:50] In fine-tuning, we do not want the model to learn like we do not want the transformer layers itself to learn or change their weights by a lot, because that risks us forgetting the information that we had learned during the unsupervised pre-training stage. So we use a different parameter, a lambda parameter, to control how much information [00:32:08] Is actually backpropagated back into the, the transformer models layers themselves, and how much is used to train the last layer, which is basically this the modeling head that we had spoken of, which is a feed-forward network consists… it's basically a classifier network [00:32:24] With softmax as its activation function [00:32:27] So [00:32:29] For any labeled downstream tasks. We want to maximize the log probability for each of the pair of inputs, X comma y. Now, what is the labeled downstream task? Well, there can be various different tasks. For example, there could be sentence classification, there could be text entailment [00:32:44] Entailment is where we are given two sentences, sentence A and sentence B, or a premise and a hypothesis, and we want to predict or basically tell whether this premise is explained by this hypothesis, or it is not explained by this hypothesis, or the premise and hypothesis are not related to each other. [00:33:01] So these are the three output classes that we might have, like hypothesis explains premise. Hypothesis does not explain premise and premise and hypothesis are not related. [00:33:10] So, for… so that is one of the examples of a downstream task, right? And then there could be textual similarity, where we have text one, and so how between text one and text 2, which one is similar to each other? So we must run it on text1, text 2, and text 2, text one, because either way, it can be… as we already know, this GPT-2 model cannot take any, context from the [00:33:29] So, for text to, let's say if text 2 was the one that was actually like providing context for text one, then this would not suffice. So we need to provide text 2 first as well. So this is one example of textual similarity. And then multiple choice answers, where for the context, which is the question, or the paragraph that we have in the question [00:33:50] And the question itself, and then each of the multiple choice answer. Answer one, answer two, answer n. We want to predict whether this answer is the correct answer for the given context, like the paragraph and the question. [00:34:01] So, these are… and there are various other downstream tasks as well, but these could be a few of the downstream tasks that the GPT-2 model can be used for, apart from, of course, next token protection, which is basically generation task. So [00:34:17] For each of these tasks, we need to fine-tune the model. And how do we fine-tune it? We give it a labeled set of inputs where X comma y, where X is the input and Y is actually the class. So X is the sequence of tokens x1 to XM, and y is the class. So, for example, let's take the entailment [00:34:34] So X1 to XM contains the hypothesis and the premise. And Y is the class whether the hypothesis explains this premise does not, or they are not related. So we have a corpus of a sequence of such x or y pairs or a dataset C [00:34:50] And then each of these sequence of tokens in C is passed through the transformer model, and we get the transformer outputs, right? HL where the outputs from the nth layer of the transformer. We pass this into the feedforward layer here, this feedforward WTE, we only train this [00:35:06] So [00:35:07] We train a set of HW to predict Y, based on X1 to XM, or the output of X1 to XM from the transformer right? So we take a… and we do a softmax, because it's classification, so we need a probability of each of the classes. So we do a softmax activation [00:35:23] And we use the negative log likelihood function again. So that's log of we need to maximize the probability that the wired class is the answer for X one to right? So X comma y belongs for every pair of [00:35:37] supervised labels, XMA belongs to C. We run the transformer on X1 to XM. We get the outputs, we run the classifier layer, which is the speed forward network, and a softmax activation function, and we get the probability of Y, and we need to ensure that this [00:35:54] Loss is maximized. Now, how do we back propagate this back into the model? Apart from this loss, we also take the basic loss, the model's loss as well, which was this [00:36:04] equation here and we multiply it by a parameter lambda. This lambda parameter controls how much information is back propagated back into the model. So this is the final loss function for the discriminative fine tuning. If this lambda is very large, it might so happen that during fine tuning, the model might forget [00:36:21] the, a lot of general knowledge that it might have learned during the unsupervised free training. So this lambda must be controlled in order to ensure that we only train the modeling head or this last layer rather than the entire transformer itself [00:36:37] for fine-tuning, because the fine-tuning specifically means that the model has all the base knowledge and some extra knowledge that it requires to perform a given downstream task. [00:36:46] So [00:36:47] Yeah, any questions about this? [00:36:54] Okay. [00:36:55] So, yeah, so this is the cumulative loss function, LS based on the discriminative fine-tuning and the unsupervised pre-training task loss together, right? [00:37:06] So, these are the input structures for various fine-tuning tasks, as we saw, text classification, text entailment, text similarity, multiple choice question answering, and so on and so forth. So GPT can be applied to a vast variety of domains, apart from just next word prediction [00:37:21] And it can be fine-tuned for each of these tasks using this discriminative fine-tuning text or training task. [00:37:30] So, what was the dataset, and what were the training details that were used for each of these for each of these, right? Let's talk about the unsupervised pre-training first. So the unsupervised pre-training for GPT-2 model was done on the web text data set which contains about 40 GB of text, which was corrected from over 8 million web pages [00:37:53] These webpages itself were selected as the highest rated links from heavily voted or highly upvoted Reddit posts. The criteria for the Reddit post were that they would have received at least three karma points before December 2017, which is when this dataset was created. [00:38:12] So the premise being that humans who are posting and people finding that helpful and giving them karma points or upwarding their answers would basically mean that the outbound links are good, right? So, these outbound links [00:38:27] the text from each of these links was collected and combined together into a web text data set which can which is about 40 Gb of text, and the total number of web pages such selected were 8 million, and all the duplicate links and all of the Wikipedia articles were removed from this dataset as well, since this would lead to overfitting, it was found that a lot of models were overfitting [00:38:48] If, all of these links were kept as a part of the dataset. [00:38:53] Now, GP2 models were available in 4 sizes, small, medium, large, and extra large, as we had [00:39:01] already discussed. So GPT-2 small contains about 120 million, 170 million parameters. [00:39:07] The media model contains about 345 million parameters and double the decoder layers of the small models, 24 decoder layers, and the embedding dimension changes as well. This embedding dimension is basically the dimension that the token embeddings have, right? So each of the tokens, as we had seen this layer WE [00:39:24] This layer here, which has token embeddings for each of the tokens. Well, what is the number of dimensions of the embeddings for each of those tokens? So for the small model, it is 768. For the medium model, it is 1024. For the large model, it's 1280, and for the extra large, it's 1600. So more embedding dimensions allow it to represent much more context or much more meaning in those token embeddings [00:39:48] So, GPT-2 small has about 117 million parameters, Medium has about 345, large is about 762, and extra large has about 1.5 billion parameters. And the number of decoder layers are 12, 24, 36, and 48, with extra large being the deepest [00:40:03] In terms of the decoder layers compared to small, which only has 12 decoder layers [00:40:08] So, let's look at the GPT and BERT architecture. Well, this is the encoder only architecture. We provide inputs, andthese inputs are tokenized, and we get a sequence of input tokens. These input tokens are then converted into input embeddings using the WE weights as a lookup table [00:40:27] These input embeddings are then further provided some context with their position context, right? So the position encodings are added to these input embeddings. And this final input representation is then sent into the transformer encoder layers. So each of these encoder layers run [00:40:46] bi-directional self-attention here, and a feedforward, and it goes on for n times, where n is the number of layers, the bird-based model has 12 layers, so this goes on for 12 times, and we get a set of contextual token embeddings as output [00:40:58] with one token embedding representing each of the inputs. And the GPT 2 is different from this architecture. It is a transformer decoder. We use masked multi-head attention instead of bidirectional multi-head self-attention, which was used in BERT [00:41:13] This masking is allowing us to control the flow of information. So, as we had looked at, BERT can take context from both past and from the future. GPT-2 should only take past context for the current token. It cannot take any future context for the current token, because [00:41:30] That would mean that we, as we are speaking, we also know what is going to happen in the future, which is quite wrong, and does not happen. So that's why this masking allows us to control the flow of information, or to ensure that it only flows from left to right and not the other way around [00:41:47] So this is the major difference between GPT-2 and BERT. And of course, the input embeddings and positional embeddings are created using the input tokens. This input token is this input representation is then sent into the transformer decoder layer [00:42:02] And there are n such layers, where n for the smallest model, it's 12 layers, and for the largest, it's about [00:42:07] 48 layers, and the outputs from this GPT-2 model are then supplied into a linear layer, which is the layer that we had looked at here, this WE, which is a classifier network that predicts the token probabilities which so the output from this is basically 50,257 probabilities after the softmax [00:42:28] And we then sample or basically select a token from those probabilities. That selected token becomes the next token or the output text, and this process repeats again because it's autoregressive. So one token at a time. [00:42:42] That is our GPT-2 works, and the major difference between the architecture being the bidirectional versus the masked multi-head self-attention. This is the major difference. [00:42:54] Let's look at the input tokenization. Bert used a wordpiece model. Gpt uses a byte pair encoding model as its tokenizer. BPE starts with a small vocabulary of tokens which includes basically [00:43:06] the initial alphabet or all of the unique letters in each of the words in a given document. [00:43:13] and the special tokens, such as VOS, beginning of sequence, end of sequence, mask, pad. These 4 5 tokens are special tokens, and we have the initializer bit. And then it starts combining these initial alphabets in using merges, and it learns how to build the unique words in the input vocabulary [00:43:32] So given a large set of [00:43:34] or large corpus of data. We train the tokenizer to build all of the words of that corpus of data using this initial alphabet. And this is done by computing the set of unique words, and then building the initial vocabulary, and then merging until the desired vocabulary size is reached [00:43:50] The merging rules merge any 2 existing tokens into a new token, and the how how are these mergers selected? Because at any given point there could be a large number of possible merges. We simply take the most frequent current merge [00:44:05] This is a very significant difference from word peace model where it not only took the most frequent current pair, but it weighted this by the frequency of each of the pairs. So, as we had seen last week, the score, the merging score in WPM model was based on frequency of current pair divided by the frequency of [00:44:22] A, and frequency of B for a pair AB, right? So it was not simply the most frequent current pair. It was the most frequent current pair where each of the input each of the parts of that pair were occurring less frequently. BPE has no such constraint. It simply takes the most frequent pair at every given stage of merging [00:44:41] So let's take an example and understand how BPE works. So we'll take the same input corpus [00:44:47] And that we had looked at last week, which is hug occurring 10 times, fog occurring 5 times, fun occurring 12 times, BUN occurring 4 times, and hugs occurring 5 times. This is exactly the same as last week's example. We shall also keep the, the target vocabulary size the same. Let's say we want to merge until only, like, 10, in the vocabulary size [00:45:08] So how do we start? [00:45:09] Well, we consider the initial alphabet as each of the unique letters in these words. So the unique letters are H, B, H, G, N, P, Q, and s, right? So that is our initial vocabulary and our splits, or so to speak, each of these words are converted into their letters H U G P U G [00:45:30] P-U-N, B-U-N, and QGS [00:45:33] These are the initial splits [00:45:36] Yeah. So now let's assume vocabulary size 10. So we need to train or we need to learn merges because the current vocabulary size is only 7. We need to learn three merges in order to achieve a vocabulary size of 10. So how do we learn the 1st merge? We take each of the candidate merges [00:45:54] And see, what is the frequency of their occurrence. So, HU, HG, BG, BH, BN, PS, PU, UG, and so on and so forth, all the possible merges. We compute what is the frequency of occurrence of each of these merges, and this [00:46:10] dataset in this corpus. So let's consider UNG. So UNG occurs, let's say, 10 times here, 5 times here [00:46:18] and 5 more times here. So it occurs 20 times in this vocabulary, and this is the most frequent currently. So we merge it, and we learn the merge rule as u, G, to UG, and our vocabulary becomes B, G, H, N, P, S, U, and this new token, which is UG [00:46:36] And we also merge the tokens, usually in each of these plates here. So HUG becomes HUG and PUG where UNG are merged, and again HUG again, the UG is merged. So our current corpus becomes this [00:46:53] And now, the next most frequent… now we again see what is the next most frequent pair. Now we find that H and UG is the next most frequent pair, which occurs a total of 15 times, 10 and 5, 15 times. And there is no other pair which occurs as frequently, so we merge HUG into HUG, and we update the corpus and the vocabulary. So now we have [00:47:13] UG, UN, and HUG as well, right? So [00:47:17] Now we can see that the merge rules are UG to UG, UN to UN, HUG to HUG, and the vocabulary becomes BGHNPSU UG, UN, and HUG, right? [00:47:29] So this is our tokenizer [00:47:32] And how do we tokenize from any given word? So, we first take, for example, let's take the word BUG. We split it into B, U, and G, and then we combine UG, because that is our first merge rule, UG is UG. [00:47:47] And we combine you. And, for example, mug, we split into MUNG. Now, M is not in the vocabulary, right? So this becomes an unknown token. This is another point of difference from the word piece model where the entire token MUG would be classified or tokenized as an unknown token [00:48:07] BPE does not do that. It tokenizes as much as it can, and whatever it cannot, it classifies as an unknown token. So only M is unknown, and U and G are further merged using the rule UG. [00:48:17] Now, let's say we had to, tokenize the word thug, T-H-U-G. So, we would first split it into letters, T-H-U-G [00:48:26] And then we would see, okay, the most frequent merge that we have learned is UG. So we would 1st merge UG [00:48:34] And we get Ug here. And then we see that, okay, now can we merge any further? Well, Hug, we have a rule for that, and there is no other more frequent rule that can be applied here. So we take Hug, and we apply that, and we get the tokenization of tag as unknown token and HUG [00:48:51] So this is the [00:48:55] Byte pair encoding, VTE model that [00:48:57] GPT-2 users are the tokenizer [00:49:00] Any questions regarding the tokenizer or the input tokenization for GPT-2? [00:49:15] Okay, so if there are no questions, we can proceed. [00:49:20] The input representation, which is basically taking the token embeddings and the position embeddings, adding them together to pro to create an input representation [00:49:30] The input to GPT is tokenized during the BPE tokenizer to give us a list of tokens. Then this list of tokens are converted into embeddings using the gathered layer that we had seen. So for every token in the BPE model's vocabulary GPT has a unique token embedding, which is that token's representation in the model's vector space. So for GPT, the vocabulary size is 50,000 to 57 tokens, and each of these tokens has its own token embeddings. These are stored as a lookup table and they are indexed via the token ID. [00:49:57] So once we have the tokenizer, which tokenizers are given input into a list of input tokens, we can then gather these [00:50:06] token embeddings for each of these input tokens, and then we take the position encodings. The position encodings help GPT express the sequence of these words or the position at which a given token or a given word appears in a sentence [00:50:21] And this allows the model to represent sequence information, right? So that is also a very important part of input representation, because the same token occurring at two different points in a sentence must be discriminated against. Both of them cannot be the same. So that's that's how [00:50:38] this position embeddings are very important. And there is another point of difference between BERT and GPT. BERT use sinusoidal position embeddings, whereas Gpt uses trained position embeddings. These are learned during the model training. These are not calculated using the [00:50:54] the [00:50:55] this formula that we had seen last year, last week. So the scientific position embeddings for any given set of dimensions D can simply be calculated as this formula, right? For Gpt 2, it does not hold. We cannot calculate the position embeddings. These position embeddings are learned as a part of the model's pre-training itself. [00:51:16] They are not calculated and injected into the model separately. [00:51:21] So [00:51:22] The overview of GPT's input representation can be seen as, let's say we have a position embedding matrix here, and we have the input token IDs here, and there is a vector lookup table for each of these token IDs. So token ID 38, we simply go to the vector lookup table, we take the index 38, and we say, okay, this is the [00:51:40] Vector for to open one, which is ID 38 11571 is this next token, so we take the vector for that token, and similarly, so on and so forth, till the vector of the eighth token. And then we take the position encoding. Now, as I said, the context length for this model is 1024 [00:51:55] And since, let us consider the GPT-2 small model, then the encoding dimensions would be 768. So the position embedding matrix would be of size 10 to 4 rows plus 768 dimensions or columns. We simply take the first 8 [00:52:10] Because there are… we have 8 input tokens. We take the first 8 position ids, and we add them to their corresponding vector token IDs, right? And we… after this addition, we get the embedded tokens, or the embed tokens with the position embedding [00:52:25] So this is the input to the model, which is of 8 tokens for a 768 [00:52:31] So, this is the input representation [00:52:34] I hope it's clear, if there are any questions, please you can ask them about the input representation to the GPT-2 model. It's quite similar to bird's input representation. There is the only difference is there's no there's no sequence information, right? Bert had two kinds of sentences, sequence A and sequence B, so there was a token type ID [00:52:53] GP2 has no to open type IDs, it only has position encodings and token embeddings. So you add them together to get the input representation for this model. [00:53:11] We could also take a, [00:53:14] food, and see how this works. So [00:53:19] Let's see… we have to share that. [00:53:21] So this is a token encoding. This is the position encoding. So WTE plus WPE given a set of input tokens, we simply calculate this. [00:53:32] So [00:53:35] Telets. [00:53:37] Okay, actually, let me stop this here. [00:53:43] Let me put a breakpoint first [00:53:51] Yeah? [00:53:52] Yeah, hello? Yeah, before we are proceeding, any questions, any doubts? [00:53:56] I think, uh, the code example is quite interesting, but before that, any doubts on the theory part or on the code? [00:54:04] whatever is being done, any questions, any doubt, [00:54:07] Anyone? [00:54:19] Yeah, thank you. [00:54:22] So, we can also look at the entire sequence first. So we get this, this is the input sequence. Tell me about Qualcomm, for example. We give this to the [00:54:34] Tokenizer here, the encoder. This uses a BPE tokenizer, and it produces a set of input tokens. The input ID is here, and we shall look at the input IDs, what these are once we get the inputs anyway passed in here. So, we have the input IDs [00:54:49] We ensure that these input IDs and the number of tokens that we want to generate. So for example, for this example, I'm limited to 40 here. So this should be less than the context length of our model. If it is greater than that, then we can't generate these many tokens, so we need to, you know [00:55:04] assert out, or basically error out. And the output IDs, or the output token we get from the generate function, right? So the generate function here, it runs in a loop for n tokens to generate, and it calls the GP2 model for each of, the inputs. The first input is the input IDs, the next input, well, we append the next ID, right? We keep appending the generated token [00:55:27] So once we get the logits from GPT-2, we do a simple greedy sampling here. We just take the maximum one, and we do it. Just a second. [00:55:35] Give me a second, I'm getting a call, I'm sorry. [00:56:09] Yeah, so sorry about that. [00:56:17] Yeah. [00:56:22] If we see the inputs here, this was the set of input tokens for [00:56:28] the sentence, which was, tell me about Qualcomm, right? So, tell me about Qualcomm, we sent it into the, input tokenizer, and then we got these set of input tokens. This input tokens were sent into the model here, GPT two [00:56:44] And this GP2 model here, before we run it, we need to convert these input tokens into token embeddings before we run the transformer block. So the model has a set of weights, WTE and WPE, right? So [00:57:01] Yeah, so as you can see, the… the shape of WTE is 50,000 to 57, which is the number of tokens that this model can [00:57:12] Oh, understand, or the vocabulary size for this model, which is 50,257 tokens. Each of these tokens has [00:57:19] our vector of embedding size 768. So that's why 50,257 plus 768 is the shape of this. [00:57:26] And if we do the same for the position encodings in C, that we can only represent 1024 tokens within our context [00:57:35] Any more than this, we don't have opposition embedding to discriminate between them [00:57:41] So [00:57:42] What would be, for example, the input representation for this token? Well, let's calculate it. Let's take WTE off [00:57:49] 24446, right? [00:57:54] This… this is the vector representation. This consists of 768 values. This is the vector representation for the token ID 24446. And since this is the 1st position, which is the zeroth position, we can say that we need WPE [00:58:11] all zero, right [00:58:14] So this now, if you were to add both of these together [00:58:23] So this would be the input representation for the very first token, right? For the very first token, this [00:58:30] set of 768 values here, right? These set of 768 values. This is what the model knows. Okay, this is. This means [00:58:38] tell, for example, right? So, this is the input to this model here [00:58:45] And let us now just do next, right? [00:58:48] And let me print X, for example. X right so you can see that the last value here will basically match this value. We just calculated it [00:58:58] But that's all and the first value here, or the first few values here, you can see would match these first few values right here. [00:59:05] Great. [00:59:07] So, this is what token embeddings and position embeddings are used for. This is the input to the model. So the first token, the second token, the third token, and so on. So let's say if I do [00:59:24] You can see, we have 4 representations, right, for each one, for each of the tokens [00:59:29] Right here. This would be sent into the transformer block [00:59:34] And what that does will soon… we'll take a look at in the code and the lecture as well, and how this runs. [00:59:41] So we can also take a look at the tokenizer itself, and how that works. [00:59:52] So this is the tokenization part here. Let me come here [00:59:58] This VPE function is going to be called by [01:00:01] Yeah, this is the encode function that we have called right here. [01:00:05] Right? The encode prompt where we supply tell me about Qualcomm as a part of this. This encode function first basically let me put a breakpoint here as well. [01:00:17] And let me put a point right here as well. [01:00:30] So, this is done. Let's come here [01:00:47] So, let us come here [01:00:49] the second to that. In the meantime, are there any doubts regarding the input representation and, how that is done? [01:01:08] Okay, so this is the second [01:01:12] So this is what we have taken a look at just the input token IDs, the vector for each of these tokens, the positions, and adding them, we get the embedded tokens. [01:01:22] So [01:01:25] It's just loading all of the way, it takes a bit of time for that because it's a bigger model. [01:01:32] Yeah, there we go. [01:01:34] Yes, right? So, if you see, the text is, tell me about Qualcomm and [01:01:39] We first convert this text, so we just need to get the, all of the unique textual tokens, like, the self.pattern is similar to what we had seen last week, where, for the BPE model, we were ignoring or converting a few of those [01:01:53] characters like the E with the tilde on top. We just wanted the base character. This does the same thing here, and let's take the first token. So, within… for token in each of these, so let's take [01:02:05] So the current token is only 10, right? We consider each of these words one by one, so we take 10 first [01:02:11] So this token is then supplied into the token.encode into UTS, and we get a UTS-8 encoding for this. [01:02:22] Okay. Right. [01:02:24] Thank you [01:02:25] And now, we would go into this BPE token exchange, right? [01:02:32] Okay, right here. [01:02:34] So, we get to encoder.py 62, which is this name [01:02:39] Now, this, we can ignore this part because currently the cash would be empty. This is just to speed up the tokenization itself, but we can ignore this. And let us take let us see what this get passed does, right? And how we tokenize first [01:02:53] Sorry. [01:02:55] So [01:02:56] From the current token [01:02:59] Yeah. [01:03:01] From the current token, we convert this into this. T-E-L-L, which was the first step of tokenization. If you remember, right here [01:03:11] Yeah, the first step of tokenization, we were doing bug to BUG, and now we want to merge, right? So similarly T-E-L-L, right? This is the first step. [01:03:22] Then we come in here, and we continue merging till we cannot merge anymore, right? So let's see how that works, right? [01:03:29] While true [01:03:31] 5 grams of this [01:03:36] Right. [01:03:37] So, what is the first diagram? We take… so here, right here, what was this? We were given a set of merges, the pairs, so let me just print the pairs, for example [01:03:47] all possible pairs. T-E-L-L. We get T-E-E-L, and LL, right? So these are now these have been sorted based on the occurrence of these merges. So we need to merge LL first and then TE and then EL if we were to merge it as such. So let us take LL in the first [01:04:08] example, right? So the first comma second of this bigram would be L and L, because it's Ln. So let's say this we first and the second right so, as we can see here [01:04:20] Now [01:04:21] We continue merging, like, we need to merge LL in the first word, right? So this is… this is the code for merging it. And once we are done with this, we would get a new word, which would be LM. So let us see how that happens. [01:04:37] Right? [01:04:40] So… [01:04:42] As you can see. The new word would be LL Merced. [01:04:46] Which is what we are doing here, right? [01:04:49] BUG, we merge this, and we get the new word to be BNUG. We continue merging so on till we cannot merge anymore. [01:04:56] So, that is how we would continue merging this. Still, we cannot merge anymore, right? So [01:05:02] That would be done right here. While true. So once this is done, we would return this word. So let's continue. [01:05:20] See, what's the new word now? [01:05:24] So, see, we merged E as well, because we had a merged rule for E and LL in our merges. So if you see here, you'll have a rule for LL [01:05:36] And we have a rule for E and LL. [01:05:38] Right here. So, we have merged, and we continue merging again [01:06:00] So now we have finished our mochi. We could also merge T and ELL [01:06:08] Yeah. [01:06:10] We can see here that we have merged this one as well, and TELL is our token. [01:06:15] So TELL is going to be the first token. This word could be merged subsequently so that the entire word is a singular token because we have a token called tell in our vocabulary, right? [01:06:26] So, this is how the tokenizer works. We continue merging till we cannot merge [01:06:34] So similarly, like, for example, this we were merging U and G, then we merged HUG, and now we cannot merge anymore because it's unknown and HUG. So we stop here. So this is what we did. We merged ELL, we merged LL, then we merged E and LL, and then we subsequently merged T and EL [01:06:51] And that's how we got the first input token to be 10. And this continued for the subsequent words as well. [01:06:58] tell me about and Qualcomm. Each of these words were tokenized separately, and we got the sequence of input tokens that we just had a look at [01:07:07] So [01:07:08] Any doubts? [01:07:14] Okay. [01:07:16] So, coming to self-attention, once we have fed the input into the GPT decoder, it understands the context of each word using the causal or causal multi-headed self-attention. Now this masked or causal multi-head self-attention is different from bidirectional. This is the key component that allows us to ensure [01:07:34] That, the model only takes context from the left and not from both the left and from the right, which is from the past and from the future. We only need it from the past. So, this masked multi-head self attention is the major difference between an encoder and a decoder model, and this allows us to relate each word in a sentence to all the word that preceded it [01:07:53] Strictly preceded, not succeeded. BERT allows all the words preceding and succeeding to provide context for any given word. GPT ensures that each word in the sentence only gets context from all the other words preceding it in a sentence, and learns the causal relationships and meanings of these words. [01:08:10] So, it runs the mass multi-exelf attention. This is very similar to what we had seen last week for BERT, where we have the input X. These are the input embeddings. The input embeddings that we just prepared, we just had a look at. So, for example, 4 cross 768 is the size of X [01:08:28] Then we have the model weight, the query weights, the key weights, and the value weights. We multiply this to get the query matrix, the key matrix, and the value matrix Q, K, and V. And then we compute the masked multi-head self-attention. So how do we do that? First, we take the dot products of [01:08:43] Each of these Q and K vectors as Q times k transpose. So this Macmal basically is taking the dot product of each of the queries with each of the keys [01:08:53] So this is implemented for efficiency as a math month, Q times k transpose, and then we do the scaling by root D, which is the D is the number of head dimensions. So 768 by 12 in our case, and under root of 768 by 12 is the scaling [01:09:09] Then we add the mask. This mask is the is why we call it causal or mask multi itself attention. How this works, we shall take a look at shortly. So this masking allows the model to mask all the tokens that appear to the future of any given token [01:09:25] So, if we had a sequence, let's say, tell me about Qualcomm. When we are processing the word tell, we can only take context from tell. We cannot take anything from me about and Qualcomm. We cannot allow me about our Qualcomm to pay any attention to tell. So this is what this mask is doing [01:09:41] For Tell, it is masking me about in Qualcomm. Then, when we are processing the second word. So, now we have tell me as a sequence, right? So tell me [01:09:50] We can take… me can get attention from, like, me can pay attention to tell, and tell can pay attention to itself, and tell can pay attention to me, but it… none of the attention for any of these tokens can be received from any… or derived from anything in the future, which is About and Qualcomm. That is not allowed [01:10:05] Right? So [01:10:07] That… that is another… so the masking is also done there, and this masking allows us to mask out about in Qualcomm, and when we are processing, tell me about, we need to mask out Qualcomm. And finally, when we process tell me about Qualcomm, well, at this point, we have the entire context. So, every token can pay attention to every other token. [01:10:25] So this masking allows us to ensure that no word gets any future context. It can only take context from its past [01:10:34] The softmax function is applied to this. The summation of the Q times k transpose by root D plus mask, and this of max then zeros out all of the values that were like that had been masked out. [01:10:46] And then we do the, we multiply by the values, or scale by the values to get the attention scores for or basically the output from the attention for each of the heads. And these are then concatenated, and then we have an output projection just like both. The major difference between BERT and GPT being this mask. [01:11:02] This mask converts the attention from bidirectional to causal. [01:11:07] So let's take a look at bidirectional self-attention that we had seen last week. So, when we were processing the token 8, it could derive attention from anything that was to the past and to the future. So, the animal didn't cross the street because it [01:11:23] Now, when animal is being processed, it is in the future, but animal can still pay attention to it. Whereas, if we look at causal self-attention, animal cannot pay any attention to it. It can only derive attention from anything that appears to the future, which was it was too tired [01:11:39] So it can derive attention from walls, it can derive attention from too and tired. So [01:11:45] The context of it, or the attention that it gets, can only be from anything that was like that appears to the future of it, right? When animal is being processed, it does not know that there is it in the future, so it cannot pay any attention in the future, right? So [01:12:00] This is so in a very visual sense, this is how we can look at bidirectional self-attention and mass self-attention, right? Nothing in the future can pay attention to a current token, right? [01:12:14] Was it was too tired. It pays attention to was too tired. It pays attention to too. Because when tired is in processed, was being processed, two is being processed, it has already occurred in sequence. When animal is being processed, it has not. So, no such connection from any [01:12:28] Future token to any past token, right? None of these can occur. Whereas for bidirectional it can [01:12:36] So, is, any questions about this, visually? We shall take a look at what the mask is and how it is created now. So any questions regarding this? [01:12:53] Okay. [01:12:55] So, this masking step is the core component which basically converts this bi-directional self-attention into masked multi-head self-attention. The masking is very important to allow context to to control the flow of context from the future [01:13:10] We cannot do that. So [01:13:13] So, let us take a look at how master attention works. So consider an input of four tokens. Robot must obey orders and we assume that we are the second step where we are processing must. So robot and must have occurred in sequence. Obey orders are to the future. [01:13:30] We… we need to mask the last two tokens out, the obeying orders need to be masked out to prevent them from affecting the attention source since they appear in the future of the current decoding step, we cannot allow them away and orders to pay any attention to robot must, right? [01:13:46] Essentially, all the feature tokens must be zeroed out in the attention matrix. The Q times k transpose matrix is the attention matrix, and we need to zero out all future tokens for every single row of that matrix. So, for the second row, when we are processing must, we can, for example. [01:14:01] Must can pay attention to about 20% to robot and 80% to itself, but it cannot pay any attention to obey and orders. There can be no flow of attention from must to obey and orders. So we mask both of them out [01:14:15] 0 and 0. [01:14:16] So let's take, for example. [01:14:19] Let's take this example and look at it in detail. So now we have this, we have the queries robot must obey in orders and we have the keys. Robot must obey orders, robot must obey orders, for example. So Q times k transpose, right? So before softmax, these are the scores that we have. Queries times keys, we get these scores [01:14:35] So here, each row is robot must obey orders, robot must obey orders, where these are the queries, and these are the keys. So when we are processing robot [01:14:45] Right? We do not know, must obey orders. So, there can be no [01:14:49] no flow of attention from, like, robot to must obey our orders at this step. Right? So that's why we mask all of this out. [01:14:59] Then, when we come to the second decoding step, we have robot and must. So must can pay attention to robot, must can pay attention to itself. But must cannot pay any attention to anything that occurs [01:15:10] in the future, which is obey an ordinance. So, these two must be masked out in this Q times k transpose matrix. So, minus infinity, minus infinity, right? At the third step, robot must obey. Again, we need to mask this out, because orders has not yet occurred in sequence order [01:15:27] Right? So we cannot allow any flow of attention from robot must obey to any of like these. Right? So that's why we [01:15:36] Zero that out. And then, coming to the final step where we have all the tokens. Robot must obey orders, robot must obey orders. So now everything is known. So we need not zero out any part of this, the last row because everything now has occurred and anything can pay like any token can pay attention to any token. So [01:15:55] robot must obey orders. Orders can pay attention to robot, it can pay attention to must, it can pay attention to obey, it can pay attention to itself, right? [01:16:05] So, that's why there is no mask here. So once we do the softmax, softmax of minus infinity is 0. So there is no attention being paid from robot to mass, robot to obey and robot to orders, 0, 0. And this upper triangular [01:16:19] matrix is basically zeroed out [01:16:21] And this is the muscle, right? We must mask anything that occurs in the future, which is the upper triangular matrix of Q times k transpose. Q times k transpose is always a square matrix because Q, if there are n queries and n keys, then there are n keys, because both of them are calculated from the same input x right? So if this is N cross [01:16:43] embedding dimension, and this is also n cross embedding dimension, and this is also N cross embedding dimension. So the output for each of these query key and value matrices is always N cross a dimension, like the attention dimension, n cross attention dimension, n cross attention dimension, and n cross attention dimension [01:16:57] So if you do a MacML of Q times k transpose, we always get an n plus n square matrix. So the upper diagonal of that, the upper triangle of that matrix across the major diagonal must be zeroed out [01:17:10] To allow only the past to pay attention, to be paid attention to. We cannot allow flow of attention from the past to the future. We can only allow it [01:17:21] In a single direction, left to right, no left or right to left, right? [01:17:25] So [01:17:28] Any doubts regarding this? [01:17:38] Okay. [01:17:40] Yeah. [01:17:41] Uh, no, I was thinking something, but I will ask later on. [01:17:44] Yeah. [01:17:47] Okay. [01:17:49] Sure. So, this is how the attention masking works. Once this is done, so once we multiply this by the value matrix, right? So we know that this like the value matrix also has rows for each of the tokens. The value matrix has robot must obey, and orders [01:18:06] So when we multiply this at the first step of decoding, it can… a robot can only pay attention to itself. [01:18:12] So the first value can only have the corresponding [01:18:16] Context for robot only, whereas the last row robot must obey orders [01:18:23] How does this paying attention to robot must and obeying. So, the value, the value row for orders will have the context from the entire sentence. And how much it will have will be controlled by these attention scores, right? So this attention masking allows us to ensure that there is no future context. If we had not done this, then we would have [01:18:41] some attention being paid by robot to like must obey and orders which all appear in the future. Robot does not know about any of those tokens, so it cannot pay any attention to any of those future tokens. [01:18:54] So this is very important, and this is the core idea that allows causal language modeling and token generation to proceed by a transformer decoder architecture. [01:19:06] So, once this masking is done, we take a softmax and we multiply by the value matrix, which is again the same. So we say this is the input x we multiply it with the queries, keys and value weights to get the queries, keys and value matrices, and we do the q times k transpose by root D, mask it, take the softmax to get the attention scores, as we have just seen [01:19:25] The all of the future context is removed from the attention scores. [01:19:30] And then we multiply that by the value matrix. This is the in-context, or the encoded dimensions of that contain the meaning for each of those tokens. So, for the value representing robot, it only contains robot, whereas for the value representing orders, it contains context from robot must obey orders. So it has a context [01:19:52] a complete context. And that complete context allows it to gain a deep understanding based on the attention scores of what is being done, or what the model is trying to trying to represent [01:20:03] Right? So once this is done, so this entire equation is done per head. Each head runs the Q times k transpose by root D, masking, softmax, and the value multiplication. Once we have… so, this is the output from each of the heads. Let's say there are 8 heads, so Z0 would be the output of the… for the first head [01:20:21] Of this softmax times V, and Z1 would be the output from the first head, and Z7 would be output from the eighth head. So Z0 to Z7 are all the outputs from all the heads. These are concatenated together, and then these are projected back into the model dimension, right? Because this is all a linear projection. [01:20:41] The input model dimensions are projected into the head dimensions using the QK and V matrices, right? So these now need to be projected back into the model's dimension, so that the model can further process these and understand what was trying to be represented by the attention itself [01:20:55] So for that, we have the output rejection similar to how BERT does it, which was WO. And once we do that, then we get the output from the attention layer. [01:21:06] This output [01:21:08] The output here. This output is after the add-in norm is done after the output projection. So anything so Q, K transpose by root d plus mask softmax times V is done per head concatenate across all the heads project back into the model's dimensions [01:21:25] And into the model's vector space so that the model can, again, understand what was being represented by the attention operation itself. And then we do the addition and normalization, and then the feedforward [01:21:35] And then it proceeds so on and so forth. For each of the transformer decoder layers. [01:21:41] Right. So [01:21:44] This is all about past self-attention. Putting all of it together across multiple heads, we can see, let's again take it [01:21:54] So anything that occurs to the future of it, it can only pay attention to was too entire because anything that where it is in the future cannot like animal cannot pay attention to it, didn't pay attention to it, was can't pay attention to it because it is to the left [01:22:10] of walls. So, different heads are again learning different parts or paying attention to different parts of the sentence, right? So it receives attention from walls, it receives attention from tired in a different head, like this this head, this blue head, is paying attention to tired [01:22:26] A few heads are paying attention to what was a few heads are paying attention to two. So this allows the model to represent what it was, okay? It was too tired, so [01:22:38] Accumulating across all of the heads allows us to get the complete context. [01:22:42] If you just consider a single head [01:22:45] Well, we don't get much context about, okay, it was, but we don't get much context, right? Was refers to it, but what, what, what was the state that was being referred to that can be done by accumulating across all of the heads, like [01:23:02] So the concatenation and the output projection allows the model to get context from the cumulative context from all of the heads together. This blue head allows it to get context. Okay, it was tired. And how tired? It was too tired, right? So [01:23:15] This blue and green head together allows that context to be injected, right? And again, it is unidirectional. Anything that it [01:23:27] occurred when it was being persist, only that token can be paying any attention to it. So animal didn't cross cannot be paying any attention to it, because when animal was being processed, it had not yet occurred. Was too and tired, can pay attention to it, because when these were being processed, it had already occurred [01:23:42] Right? So again, the the lower triangular rule that we had. So the upper triangular is all masked out for the Q times k transverse matrix. [01:23:52] So. [01:23:53] Now, coming to the model output itself, the output from the final decoder layer consists of a set of vectors. Each input token has a corresponding output vectors. This is similar to what if we give N input tokens, we get N output vectors. Now, from the n output vectors, each output vector has a dimension of 768 for the GPT-2 small [01:24:11] Model [01:24:12] Right? So, we take the vectors across each of these. Now, let's say our input was beginning of sequence help Prince Mayuko. Well, we send that as an input. We get and we get a corresponding [01:24:24] output for each of these years. Now, we only run the classifier on the last token because we need to predict the next token, right? So only for the output vector or the output embedding, or not embedding, but the output vector for this Mayuko is [01:24:39] sent to the class file and so on to the to the sampler, which then allows us to get, okay, help Prince Mayuko, the output is Keith. So the next the next token in the sequence is help Prince Mayuko keep piece would be the next, for example [01:24:54] And so on. [01:24:56] So this is then fed into a language modeling head. The language modeling head, as we had seen, was this language modeling head was this part, the softmax and linear layer. The linear layer 1st takes the input or the output from the decoder transformer decod [01:25:13] And it projects it into the vocabulary dimension, right? The output from the tokenizer or from the sorry, the output from the token. Sorry, I'm sorry the output from the decoder itself is N cross 768, right? Because 768 is the embedding dimension of the [01:25:31] The decoder. So, n tokens as input 768 output. 768 vector dimension for each of the n tokens, say N cross 768, this is fed into the LM head, which projects it into the vocabulary. 50,257, because if you want to select the next token, we need to know, okay, what is the probability for each of these tokens right [01:25:52] So for token zero, for token one, for token two, and so on. So this n cross 768 is fed into this, we get an N cross 50,257, where each of these rows, each of these N rows, contains the probabilities of the tokens for the next tokens for each of them. Now, since we already know the token still till end [01:26:14] We only know the n plus 1 h token. We take the last the 1 cross 50,257 probabilities. We sample across these, or we take the maximum one in this 50,257 probabilities as our next token, right? [01:26:31] So, that is what is being achieved here [01:26:33] The output consists of a vector of dimension 768 for each of the input tokens from the transformer decoder, and this is then fed into a language modeling head, which provides us with 50,257 probabilities for each of the tokens in the model's vocabulary [01:26:50] We select a singular token and we append that to our input. And this new input is then fed into the transformer decoder. And it continues till we encounter an end of sequence, or we run out of context [01:27:03] But mostly we encountered an end of sequence and we stop there, right? So for example, help Prince Mayuko keep the append keep into the sequence here. We run that, help Prince Mayuko keep, we get, let's say peace, we append that to the input sequence, we run that and we get help Prince Mayuko [01:27:21] peace in the kingdom. And after this, we get an end of sequence, so we stop generation, right? [01:27:27] So [01:27:30] Yeah, this is how the model output is processed. Any doubts regarding this? [01:27:43] So, since the model is auto-regressive at every step, we append an input token. The output token to the inputs, existing inputs, and we create a new set of inputs. These new set of inputs are passed into the model to create a new output token, and it continues till an end of sequence is sampled from the classifier that we have [01:28:01] Okay [01:28:02] So now, we come to another very important concept, which is called KV caching. [01:28:07] So this is based on an observation. The observation is that given any sequence of input tokens X1 to Xn, the self-attention block computes a 3 sets of vectors. The query vectors, the key vectors, and the value vectors in order to predict whatever the next token, xn plus 1 is that we had just seen, the query vectors are multiplied with the key vectors as a dot product [01:28:29] Which is basically Q times k transpose matrix. This Q times k transpose attention matrix is then multiplied by the values V1 to Vn, the value vectors. So this allows us to get an in-context representation of each of the tokens [01:28:45] Or, so to speak, an attention-weighted representation of each of the tokens, which is then used by the subsequent layers of the decoder, and this goes on till all the decoder layers process all of the tokens. And after that we have a language modeling head [01:29:01] This is a classifier that classifies or gives us probabilities for each of the [01:29:07] of each of the tokens being the next token, and [01:29:11] that token is the XN plus 1 token after sampling, right? But in order to do this, we are… the core concept here is that we are computing blocks of vectors, Q, K, and V, for each of the input tokens X1 to Xn. Now, when we process the next input, which [01:29:28] We append Xn plus 1 to the sequence and we process this entire input as to get the next open Xn plus 2. What we then do is we compute a set of input vectors, or we set we compute a set of query key and value vectors for each of these input vectors, x1 to Xn, once again [01:29:44] And we have Q1, QN plus 1, and then we have K1 to Kn plus 1, and then we have V1 to VN plus 1 in order to predict Xn plus 2. Now, what we see here is that these queries, keys, and values, they are being recomputed, like Q1 to Qn is being recomputed for [01:30:00] The only new vector here is QN plus 1, and then k1 to Kn is being recomputed [01:30:05] The only new vector is kn plus 1, and v1 to VN is being computed, and the only new vector that we are computing here is VN plus one [01:30:14] So we are doing a lot of replicate computation to predict [01:30:19] Each token, right? And for XN plus 2, we would only need to get a new vector QN plus 2, Kn plus 2, and VN plus 2. [01:30:27] we recompute the entirety of all of these vectors, right? So if we take a look at this, all the previous QK and V vectors are computed each time a new token is taken into consideration. [01:30:39] The key observation here is that when we are actually computing the next token, we only need the last row of the attention matrix, right? If you look at the attention matrix here [01:30:50] We only need this in order, because the output corresponding to each of these rows, right, the, this is corresponding to robot, this is corresponding to must, this is corresponding to obey, and this is corresponding to orders in this example, right? Now, in order to take the next token, we only need the attention from orders. We don't need robot must obey. The output is, of course, n cross 768, where n is robot [01:31:13] N0 is a robot, N1 is must, N2 is obey, and N3 is orders. [01:31:20] But when we are actually running the classifier, we are anyway taking just the last layer, the last row which corresponds to orders, as we had discussed here. So [01:31:30] We are running the classifier only on the last output of the decoder. So, given a next sentence prediction task, we actually only need the queries. But in order to get the context, or get the context from all of the previous vectors here. [01:31:46] We need the keys, and we need the values, because this score is going to be multiplied by the values next, right? We need the keys to compute the scores itself, because orders is being computed by taking the dot product of the orders vector with the robot vector [01:32:01] Which is this one, and then the dot footer of the order vector with the must vector to get this score, and then the drop rate of the orders vector with the obey vector to get this code, and the orders… the dot product of the orders vector with itself to get this code. [01:32:17] So we need the key vectors [01:32:20] And we need the value vectors, because this is going to be multiplied by the value vectors corresponding to each of the past tokens. So as long as we have the current query, and all of the previous fees and values, we can always predict the next token, right? So we need not recompute [01:32:35] All of these queries, all of these keys, all of these values at each step. [01:32:40] This is the core idea behind KV caching. So visually, if we see… so this is a random sentence, and the sky is blue. So if we want to predict the next token, right? So [01:32:50] We need all of the keys and all of the values [01:32:54] To get context from the past. So what is going to be next after blue? Well, it could be full stop, it could be comma, it could be anything. But what is it? Well, in order to predict that, we need all of the past keys and all of the past values, but we only need the current query, which is blue. [01:33:10] Right [01:33:12] So [01:33:14] Is… are there any questions about this? Because this is quite it's tricky, but we need to understand this in order to understand how to speed up the inference and what KV caching is. [01:33:25] Any questions regarding this? [01:33:37] Okay, so basically KLE caching is based on these two observations that we only need the current query and we need all of the previous keys and all of the previous values. Taking all of the previous keys, all of the previous values K1 to Kn, v1 to VN [01:33:53] And then we compute for Xn plus 1, we compute the key Kn plus 1, the query qn plus 1, and the value VN plus 1. So given we already have K1 to Kn and V1 to Vn, we already have all the key vectors now, K1 to K plus 1 after we compute k n plus 1, and v one to v n plus 1, after we compute V n plus 1. [01:34:13] And we only need the query q n plus 1 anyway, because [01:34:16] All of the past tokens we are not passing through declassify. The outputs from the decoder correspond to each of the past like robot must obey or Bos key like help prints myuko only we only need the output from Mayuko. We do not do not need the output from each of these [01:34:33] Because we are anyway not going to use that. We only need the last sequence. [01:34:38] So [01:34:40] This is the core idea behind KL caching. So this allows us to speed it up by a lot because [01:34:46] If we look at this here, if we are preventing the end computations for each of the queries, keys, and values, and this… because each of these computations are quite heavy MACMLs, and they require a lot of memory bandwidth, these are quite memory-bound, and [01:35:02] They make inference quite slow. We need not… and they also waste a lot of compute resources, which would [01:35:07] explored the, the cost per token. [01:35:11] So, KD caching allows us to cache or save K1 to KN, V1 to VN during the input processing stage. So when we are processing help, Mayuko or robot must obey orders, for example, we save the keys and values for each of those tokens. [01:35:26] Into memory, or we just save them somewhere. And when we are then processing the token, the output tokens, one by one, we just compute the current query key and value. We append the current key and value to all the previous keys and values, and then we take the current query, and we compute the Q times k transpose by root D plus mask [01:35:44] Softmax times V, and we get the in-context [01:35:49] And like all of the current context based on this. We get all of the current context based on just the current query and all the previous keys, and all the previous values. And this in-context representation can then be used by the model to predict whatever is going to happen next. [01:36:06] We are not losing any information, the output remains exactly the same [01:36:10] The model allows much faster inference in order for this to proceed, right? So [01:36:19] If you can think of it, for example, when we are processing this, when we are given the sentence, this is a random sentence and the sky is blue. When we are processing this, we compute all of the keys, all of the values, all of the queries, we do the complete Q times k transpose and everything. We only need the last row of that in order to [01:36:37] The blue is always required. The query corresponding to blue is all is required, along with all of the keys and all of the values. [01:36:43] Once we have computed these KV embeddings, for example here during processing, we have computed the K and V embeddings for all of this. So once we have the next token here, right, we can just append the queries or the keys and the values for this next token and the current query and we can the current query is alsothe query corresponding to this token and we continue [01:37:03] Right? So, essentially, we cache the KNV vectors for while we are doing the input processing for generating future tokens. And as we keep generating tokens, we keep appending these the Kn plus 1 and v n plus 1 vectors to the existing cache. So we always have [01:37:20] all of the tokens available to us [01:37:24] So [01:37:25] So, yeah, I have a query now. [01:37:26] Yeah [01:37:27] Yeah. [01:37:28] So when you mention that, uh, because of this, right, it will consume a lot of memories, because it's always going to consume a lot of details, right? [01:37:32] process again and again, everything sentence, right? [01:37:38] Yeah. [01:37:39] Then… [01:37:40] what other different strategies [01:37:41] are there, which can… to optimize this memory efficiency? Is there any history to that? Because if it's, like, a long context, Windows are there, right? [01:37:49] Yeah. [01:37:50] then it is going to take a good amount of memory to consume before the caching, right? So what are the different strategies out there which will optimize the memory? [01:37:56] Um, without impacting the performance of [01:37:59] Uh, caching. Is there any strategies that are there, or is it going to be, like, that only because we have to use this, it will follow the same [01:38:05] amount of memory. [01:38:16] Yep. [01:38:17] Yeah, so KD caching itself does consume a lot of memory because we need to save all the K and V values. So as the context window gets longer and longer, the amount of memory required to save all of these KV vectors increases, as you said, right? So for a 1024 context, we need to save a maximum of 1024 kV vectors [01:38:27] Yeah. [01:38:36] Mm-hmm. [01:38:37] But if the context is 1 lakh, for example, or current models have context in like, I think greater than even a few hundred thousand or sometimes even a million. So we need a lot of memory to save the K and V vectors. But if we wanted to save on this memory [01:38:45] While not [01:38:47] of giving away much inaccuracy. Well, we always have… it's always a trade-off. We always need to give, like, trade-off between memory and accuracy. So, there is, there are multiple techniques that allow us to reduce the amount of memory that the KV itself takes. For example, there are quantization strategies you can [01:39:05] Quantize the KV cache before committing to memory [01:39:09] And then you can load it back, decentize it, and use it. [01:39:15] You can always quantize CDCash, but that… there is always a minor quantization loss, no matter how complex the quantization strategy, you could use something as simple as round to nearest, or you could use something as complicated as, I guess, the latest turbo quant paper, for example, from Google. [01:39:29] So all of those have impacts. The impact might vary based on the quantization strategy being used and how aggressively we are quantizing it, but there is always going to be some impact on accuracy. [01:39:38] Okay, so… so, if it is a case, right, [01:39:42] You are saying that, but I understood that? There's a no way to reduce, or only about, like, a couple of, like, quantized or dequantize, but that has required a lot of computation power, right, to do the quantize and quantize, right? So it's, again, like, um… [01:39:53] To save the memory, we have to go increase the computation [01:39:57] logic as well, right? So, are there different ways also, apart from the KV caching, which is more efficient than KV caching, or this is the only GPT caching is the only way, right, as of now? [01:40:13] Yeah. [01:40:14] We always need the entire context. So, for predicting any future token, we always need entire current context and all the past context. So the current context is in the center value token and the past context are in the K1 to Kn and V1 to VN vectors [01:40:28] So we cannot drop a few of these in order to we cannot drop some of these values because we will be dropping some context. So even in a sliding window type mechanism, or there are more advanced varieties, for example, there is a key def algorithm, and there are all of those [01:40:44] Where you can selectively drop a few of the key and values in order to save on memory. If you are running out of it. But there is always an impact on accuracy. If you drop context, there is going to be an impact. It won't be exactly the same as processing all of them together, because remember, the transformer is [01:41:03] Yeah? [01:41:04] To have all the queries, all the keys, and all the values, right? That is how the training is done. Its training is done in batch processing. So it has Q times k transpose Q, K, and V are always Q1 to Q and K1 to KN, and V1 to VN. So if we drop some context out of it [01:41:17] We always have some impact on the output, right? But with KV caching, if we do not drop any of the keys, any of the values, and we only use the current neth query vector, QN plus 1 at here [01:41:30] Yep. [01:41:31] We get exactly the same output. I shall demonstrate it as well. Shortly, we get exactly the same output using KV caching as we get without using KV caching, and recomputing all of these values, right? But if we were to save… but it does take a lot of memory, and if we wanted to save on memory, we could use quantization and other techniques [01:41:46] But and various techniques of quantization that are quite complex and quite simple techniques, but all of them have some impact. We don't always get exactly the same output. [01:41:53] Got it. Yeah, thanks. [01:42:01] Yeah. [01:42:03] So [01:42:04] Let us see, right? [01:42:07] Yeah, so this is about KD caching. And KL caching allows us to, of course increase the speed of the model, the inference speed of the model, while of course it does take a bit of memory penalty because K and V need to be saved. But there is always a significant speed up in the generation of a new token [01:42:24] When we use KV caching, as we do without KV caching, right? So this is the core concept [01:42:31] in the prompt processing, or when we give the input [01:42:35] the input processing stage. The KNVs are calculated, and these are cached in memory. The queries are also calculated, and the Q timesk transfer attention. All of it is done. We get the N cross vocab size result. We take the last value out of this, like, from the model output, we take the last row, we sample the next token [01:42:53] And when we are processing that next token, we compute the query only for that token, along with the key and value for that token. The key and value is appended to these key and value matrices here [01:43:04] And we get the… also, the… the pink ones are already cached, the blue ones are currently calculated [01:43:10] So after we append it, we then compute the attention scores [01:43:14] At this point, we only get the attention for the current token, right? Like, the attention scores, few times k transpose is going to be 1 cross n now, not N cross N. So, and we need not mask anything, because for the current token, everything preceding it is in context, right? We need not mask anything, because when we're processing the n plus 100 token, there is no future context [01:43:32] So, we need not drop any context. We need all of the context. Everything is preceding the current token, right? [01:43:39] So we compute that. We get a 1D attention score now, because it's just one token, one query. So we have one plus the head dimension here, and then we compute [01:43:51] The transformer output, the transformer output also is now 1 plus 768, and then we do the classifier, or the LM head [01:44:14] key query, singular key and singular value vector is calculated. The keys and values are appended to the existing KV cache, and the process repeats itself. So, yeah, this is what I have demonstrated here by the means of these quick videos. So the video here is without KV cach [01:44:30] And the video here is with KV caching. You can see that the differences is at every stage [01:44:36] We are computing queries, we are computing keys, we are computing the entire attention matrix here, right? So we compute this query, we compute this query, and we compute everything, right? [01:44:46] But with KV caching [01:44:48] We compute a query, and we cache it. We cache the keys. We compute the next query, right? So query 1, key 1 is cached. The blue means it's committed to cache. So, the second one [01:45:00] This is cached. It's committed to cache. So when you're compressing the third token, we, you know keep incrementally adding rows into the key caches and the value caches. So [01:45:11] I hope, like, this is clear. If there are any doubts, please do ask. [01:45:22] Okay, so now, let us take a look at how it is done in the code it says [01:45:29] So [01:45:31] This is a model. This does not use any KD caching here, right? [01:45:32] Okay, there are no doubts, then you can proceed, Mr. Solani. [01:45:38] can see, when we process inputs, right, we give all the inputs to the current model, WT of inputs range line of inputs. We append the token that the model has predicted. We append this to the [01:45:50] Uh, just a second, uh, Mr. Samani. [01:45:54] Yeah? [01:45:55] Can you please enlarge, uh, the… [01:45:56] Uh, the screen slightly so that… [01:45:59] the code is visible. [01:46:02] Yeah, is [01:46:06] Yeah. [01:46:07] It's coming out a little small, yeah. You can slightly enlarge it more. [01:46:08] For better clarity. Yeah, that's better. [01:46:10] Yeah, please, continue. Thanks. [01:46:16] I hope it's now it's visible now and it's clear [01:46:23] So, yeah, we take the outputs, we take the logics of the outputs from the model. And now we have sampled. Right now we're using greedy sampling, which it's very simple. We just take the largest probability at each step. It's just the largest argument argmax [01:46:40] From all of these, from the last one, see again this gives an output of N cross N vocabulary, N sequence cross n vocab. This is the output from the GPA2 function, right? We take the… we do the input representation, we run the transformer block, the transformer block itself gives us an input of N sequence embed, which is n tokens cross 768 and [01:47:01] In our case, it's GPT-2 small. We then run the LM head layer, and we get this, right? In sequence where n vocab is 3,257. We only take the last [01:47:13] the last row of this, because we only need the… each row gives us the probabilities for the next token. But in a given sequence, like, tell me about Qualcomm. I already know what occurs after tell, I know what occurs after me, I know what occurs after about, and I only need to know what occurs after Qualcomm. So I only take [01:47:31] The row corresponding to Qualcomm, which is the last one, which is logits of minus one gives me the last row, and I take the maximum argument which gives… which tells me what occurs after Qualcomm. What occurs after Qualcomm, whatever that is, let's say Qualcomm's, so apostrophe S, I append that to the current input [01:47:47] Okay, see, Interst next ID, append it to the current input, and I process all of this again, right? So there is… and if you take a look at the transformer block itself, right here, right? [01:47:58] This is the multi-head attention, and then this is the feedforward. So if I take a look at the multi-head attention here, you can see I am computing the entire thing all over again, right? At any at every stage, I have N sequence plus n embed [01:48:13] N sequence plus enabled like which is 3 is this Q, K, and V in in this particular model what we get is the the query keys and value weights are composed as a singular weight matrix. So that is why we need to do some reshape operations and some split operations [01:48:30] But, in context, I have query, key, and value. [01:48:36] vectors. [01:48:38] Pardon me. So I've query key and value vectors for each of the tokens. N sequence is the number of tokens. So for each of the tokens, I have query key and value vectors, and each vector has an embedding dimensions, right? [01:48:49] So right here. You can see that I'm recomputing at every stage. So when I do tell me about Qualcomm, and I then get the next token, and I supply that in, I get tell me about Qualcomm's I reprocessed tell me about Qualcomm here, right? So how does KV caching help? [01:49:03] Alright, let's first… okay, let's… sorry, I need… let me just remove all of the breakpoints here. [01:49:30] Yeah, so how does scaling action happen? [01:49:33] Well, let us take with KV caching. I give the same input, tell me about Qualcomm. I do the same pre-processing, and I go all the way. I just have an extra KV cache here. Initially, there is no KV cache, right? Because [01:49:49] there is nothing that has been processed at this stage. [01:49:53] There is no KV cache currently [01:49:58] Then, I process GPT-2, right? I give all of the inputs and I get an output KV because when I'm processing the inputs, I also create the initial KV cache [01:50:08] How is that created? Well, let us take a look. I get the inputs here, and initially, KV cache is none, because I have passed none here. So, initially, KV cache is none [01:50:18] There is no KV cash. When this is done, it knows, okay, I need to create the KV cache, right? If there is no KV cache, I need to create a KV cache. [01:50:28] So [01:50:31] I do the input processing, and then I pass, I pass a KV cache block in here, right? Transformer block. [01:50:40] This KV cache block is going to be populated by this transform block. So let's take a look at this. So again, I'm passing the KV cache block into the multi-head attention as well. And the multi-head attention here [01:50:53] So initially, KD cache is going to be null. So all of this is so it's going to be none. This is not going to be done. Initially, the current cache is going to be null as well. So I need to create kv cache here [01:51:07] Right, so, right here. KDash is initially none [01:51:13] And, I have passed in… where is this? Yeah, let me go in here. Yep. So GPT2. The KV cashier, I have passed in an empty block. This is to be populated. So I go here, I go into this, and I populate new query, new key, and no new value [01:51:30] Using the QKV that have just calculated, right? This is, again, end sequence plus NM-bed. [01:51:35] Right, and I add that into new query, new key, and new value. I add that, and I take the current cache, which is QKV of 1 and QKV of 2. [01:51:45] Because there are 3 of these, right? The queries and keys. So I take the keys and values. 0 is the query, 1 is the key, and 2 is the value. I take this, I append this to the current cache, and I return it. [01:51:57] Initially, this is none, so this is not done. I get the current cache, and this is going to be returned out, right? [01:52:03] And that is how I populate the KV cache. [01:52:07] If you look at this here, you can see this is without KiwiCache. And if I look at this [01:52:15] Okay, Rikash. [01:52:17] So, it's 1.12 iterations per second, right? Which is 1.12 tokens per second, so to speak, of I generated 40 tokens and the output is Qualcomm's new 835 SoC. So tell me about Qualcomm's new 835 SoC. Qualcomm's SoC is [01:52:35] That is designed to be a better processor than the current this if it is designed. This is the output that I get from the model for 48 tokens. It continues on. I just stopped it at 40. So if you take it with KVCache as well, so we can have a look here. [01:52:47] We get the same input. [01:52:50] We run for the same number of questions, 40, it's the same model. And [01:52:56] All I'm doing different is that at every given stage of the attention of the multi-head attention, I'm saving the cache values as the current cache. If I already have some KV cache, I take that and I add the old cache and the new cache together. [01:53:11] Like the new cache that I have just created the end sequence here is going to be one in subsequent iterations here, because if you see here [01:53:19] the inputs. I append inputs only with the next, right? [01:53:24] So [01:53:27] Right here. [01:53:30] And, yeah. So I split it into QKV, I append my current cache with the old cache at every given stage. So this appending is basically this, blue and pink lights. The pink lines are already computed [01:53:44] Which is oral cache, which I already have. In the first stage, since I don't have any cash, I create all of the new cache, but when I do have old cash, I just append the new cache to the old cache. At every stage, I append this blue, which is the new cache, to the old, which is [01:54:01] The purple purple cache [01:54:03] If you look here, it has sped up by quite a bit from 1.12 to about 2 iterations, 1.6 iterations. [01:54:11] And this is also… currently, I'm running on a device that's quite constrained to memory. This just has 16 GB, but if I run it on a device with, let's say, 32 GB memory, the same code, I get about 2 to 3x speed up. So instead of one, two iterations per second, I get something like [01:54:27] 3, and here I get something like 9 or 10 iterations per table. So, but as we can see, the KV cache has sped up my [01:54:43] The more memory available there's going to be less swapping, and that allows you to run much faster. But even on this quite constrained device [01:54:52] Because [01:54:53] quite a bit faster than just running without cache, okay? So [01:54:58] Yeah, this is about KD caching [01:55:02] If you have any doubts [01:55:05] Regarding the code, or regarding how KV caching is working, please do let me know. We can… we can go into a bit more detail about this. It's quite important, because this allows the inference to run much faster, and this allows a lot more methods of [01:55:24] Reducing memory or speeding up inference by KV cache management techniques. [01:55:41] Okay. So, if there are no doubts, we can proceed. Regarding the fine-tuning. So, for example, we have just taken the unsupervised training as the first part [01:55:53] Uh, just to… sorry to interrupt, uh, Mr. Savani. [01:55:56] So we have around, uh… [01:55:58] Uh, probably… [01:56:00] a few minutes left, so… [01:56:03] Uh, if this is the last topic, then it is fine. Otherwise, probably… [01:56:08] uh… with some questions, we can conclude and continue next turn. It is up to you. [01:56:14] To take the call, uh, if you would like to continue and finish, or… [01:56:19] We'd like to defer it to the next team. [01:56:22] Sure, I think we can just refer it because we have the perplexity and the fine-tuning part remaining. We can defer it to the next lecture and take some questions or doubts that have cropped up during this lecture. [01:56:36] Yeah, so I'll request all of you to… [01:56:40] ask any questions, any doubts, or anything. [01:56:44] Because these are very important topics that are being covered. [01:56:48] And by an expert. [01:56:51] who has a very… [01:56:53] very, very strong background. [01:56:56] In… in this storm room. [01:57:00] So, uh, I think, uh… [01:57:01] If you have any questions, any doubts? [01:57:04] You can ask. Otherwise, we can break for today. You can go through… [01:57:08] the contents I will be sharing with you. [01:57:11] And on the next turn, uh, next theory turn… [01:57:17] Next to return, uh… [01:57:19] The remaining portion we will cover. [01:57:21] And, uh, then, uh, we'll, uh, I mean, we'll have a hands-on, and we'll… in the next story, [01:57:28] Turn will cover this remaining portion and start with the T5 model. [01:57:32] So, shall we wrap it here, everyone? [01:57:37] If no further questions are there. [01:57:44] Okay, so I think we'll just wrap it up here for today, then. [01:57:49] Uh, and uh, I think we all must thank Mr. Somani for taking out the time and taking us through in so much detail. [01:57:57] Uh, for this particular topic. [01:58:00] Uh, thanks, thank you, Mr. Sumani, and we'll request you to continue with the remaining part in T5 in the next 30 sessions, probably next Sunday. [01:58:09] Okay, thank you all. Thank you, everyone. Bye-bye. [01:58:13] Thank you, ma'am. Thank you, everyone.