# 03 2026-05-17 LLM BERT course: Module 4 — Generative AI & LLMs module: Module-4-Generative-AI-LLMs date: 2026-05-17 type: transcript video_url: https://personal-learn.armco.dev/files/_Recordings/Module-4-Generative-AI-LLMs/03_2026-05-17_LLM_BERT.mp4 --- [00:19:51] A very good morning to all of you. [00:19:54] And welcome to today's session. [00:20:00] Uh, so today, we'll be starting with a very important and interesting portion. [00:20:05] So, uh, after having done all kinds of, uh… [00:20:09] You know, things including LLMs, transformers, so sorry, transformers. [00:20:14] So, we'll be now using the transformers to, uh, in LLMs. [00:20:19] So… so we'll start with Bert. [00:20:21] Uh, and for, uh, going ahead, I'd like to invite, uh, our expert for today. [00:20:27] Um, Mr. Sami Somani. Mr. Sami, I request you to please enable your video. [00:20:37] Yeah, uh, very good morning to you, Mr. Sormier, and welcome to today's session. [00:20:42] Uh, so let me introduce Mr. Soumya quickly. [00:20:46] So, Mr. Savani is a machine learning engineer at Qualcomm. [00:20:50] Specialization in efficient on-device, [00:20:54] inference for LLMs. [00:20:56] And as part of Qualcomm, AI Engineer Direct SDK, that is Q&N, and Generative Inference Extension, [00:21:04] That is GenAI CPU, uh, backend team. [00:21:07] Uh, he primarily focuses on optimizing model performance. [00:21:12] for enabling cutting-edge AI capabilities on edge devices. [00:21:18] And he brings, along with him a strong research and engineering expertise, which is… [00:21:22] Uh, shaped by his master's degree in computer science from Witzpilani. [00:21:27] Hyderabad campus, and a research-focused internship from IIT Routke. [00:21:33] Uh, he's worked on a multitude of projects, uh, and he's doing the cutting-edge, state-of-the-art work at Qualcomm. [00:21:42] So, with that brief, uh… [00:21:44] introduction, I'll try… I'd like to invite Mr. Sameer to take on from here. [00:21:51] Over to you. [00:21:55] Good morning, ma'am. Thank you very much. [00:21:56] Good morning. [00:21:57] took me just to share my screen. [00:22:00] Just a second. [00:22:09] Um, do let me know if my screen is visible. [00:22:15] Yeah, it is visible. We can see it. [00:22:20] Oh, sure. [00:22:22] Thank you. So, today we'll be starting with BERT. So, this is, uh, this is part of the series for Transform Models for NLP, and we'll be looking at, uh, the encoder-only model, which is BERT today. Um, so BERT stands for Bidirectional Encoder Representations for Transformers. [00:22:40] what… we will have a detailed look at the predecessor to BERT, the immediate predecessor to BERT, and to BERT itself, and a few of the BERT derivatives. [00:22:50] Um, so let us look at the quick outline. [00:22:54] Um, first, you'll be going through the background. [00:22:57] like, the predecessor model to work, and then we'll give a brief introduction to BERT. Uh, with that, we'll be exploring, uh, what BERT is. [00:23:06] Which is basically an encoder-only model in Transformers, and how BERT is used for text processing. [00:23:11] Um, what self-attention means, and how it is bi-directional for this model, and we'll look at a few other BERT derivatives, like Roberta, Albert, and what problems or what, um, [00:23:24] like, what was the motivation behind developing these models? And then we'll be looking at the benchmarks that the BERT models are tested on. [00:23:33] These benchmarks are standard benchmarks that are used for all NLP and natural language understanding, uh, [00:23:40] models. They were used prior to BERT as well, but BERT was, uh, given the newest state-of-the-art performance on these benchmarks. [00:23:47] So, let's look at a quick background. The immediate predecessor to Bird [00:23:53] So, can, like, the state-of-the-art model before BERT was embeddings from language models. [00:23:57] Um, what… it is also known as ELMO. What ELMO does is it, uh, produces contextualized word embeddings from text. [00:24:04] So, instead of using a fixed embedding for each word, uh, like word to egg or glove uses, it looks at the entire sentence, and it understands the context of a given word, and then it assigns an embedded vector to that word. [00:24:15] So, in most languages, it's quite difficult to understand the meaning of a word, because this meaning, or vectorized embedding, can change based on [00:24:25] where the word occurs in a sentence, and what precedes it, like, what the context of that word was. For example, you could shoot somebody with a camera, [00:24:32] Which is clicking a picture, or you could shoot somebody with a gun. [00:24:35] It's only clear by looking at the entire sentence, or the entire phrase that you understand what this should [00:24:42] the first two. So, of course, the embedding for shoot should be different, uh, in… [00:24:48] you know, in sentences where it is being referred to for clicking a picture, and shooting somebody [00:24:53] is entirely different, uh, meaning, so it should have an entirely different embedding vector. So, [00:24:59] El more embeddings, uh, like, what Enmo does is it takes a sentence, it, uh, tokenizes the sentence, and looks at the tokens, and then assigns tokens, basically words, and then it assigns a embedding vector to each of these words. [00:25:12] So, let's has an embedding vector, stick has an embedding vector, 2 has an embedding vector, and so on and so forth. So, this… [00:25:19] This is called, uh, text encoding. [00:25:23] So, how does EMO work? Well, Elmo consists of L layers of language models, which are… L layers of these… the number of layers refers based on the, uh, the number of parameters in the model. [00:25:34] Um, the base Elmo model had 93 million parameters. It consists of bile elements. Biolums are bi-directional language models. These input tokens are processed by, uh, [00:25:45] character levels here in first, before the input is then, like, these input embeddings that are produced by the CNN is then fed into the Elmo itself. [00:25:53] And, uh, these… the LMO model takes these token embeddings, understands the context of the sentence, and assigns a output text embedding. [00:26:01] for each of the, uh, for each of the tokens that are given, uh, as an input to it. Every layer of the Elmo model captures different information from the source segments. [00:26:10] And the final embeddings are computed as a weighted sum across the entire layer. [00:26:14] So, what this means is that, uh, different layers of Elmo can capture different parts of a sentence and different, uh, [00:26:22] meetings that might be assigned to different words from different parts of the sentence. Like, sentence, verb, uh, sentence has verbs, actors, and actions, so… [00:26:30] Different layers can understand these complex relationships between the actors, the verbs, and the actions. And, uh, the final embeddings are computed, like, the final embeddings for the sentence are computed as a weighted sum across the layers for every token. [00:26:43] In a given sentence. Now, [00:26:45] Um, what… what does the bidirectional, uh, nature of this bi-LM? What does bi-LM mean? Well, BILM, the bidirectional nature means that, uh, for any given token in a sequence of tokens T1 to TN, or for any given word in a sentence of, uh, in a sentence of n words, [00:27:03] Um, the forward element looks at every word or every token preceding itself. [00:27:08] And understands what this word might mean based on everything before it. So, it tries to maximize the probability of TK appearing given T1 to TK minus 1. So, looks at everything in the past, [00:27:18] and understands what this word means. [00:27:20] And similarly, uh, the backward LM looks at the future. It takes the current token and everything that appears after it till the end of the sentence and understands what this current token means. So this is the bidirectional nature of the element. There's a forward pass, [00:27:33] Which they explain… which reads left to right, and the backward pass which reads right to left. And understands [00:27:40] of the meaning of every topic in a sentence, both from the past and from the future context. [00:27:48] So, input tokens are produced by a character-level CNN, as we discussed, so we will not go much into detail on this, because this… [00:27:53] This is a lecture predominantly focused on bird, but on a very high-level, what this does, it takes every character and assigns, uh, like, every character in a token, and assigns a token embedding based on that, based on those characters. These are to capture the engram dependencies. [00:28:09] Uh, because the CNNs, the filter width might vary. [00:28:12] Um, it, like, there are multiple CNNs. Each CNN can have a token window of, let's say, 2, 4, 7, and so on and so forth. So… [00:28:20] like, the N in the engram varies for every CNN, and that allows it to, uh, give a proper, uh, token embedding. This token embedding is then fed into the, um, [00:28:30] Elmo model, and it outputs, uh… [00:28:33] output contextualized token embedding. [00:28:35] Um, every layer of the Elmo model captures different information, as we saw, and the… [00:28:40] created some that we refer to. So, for example, we take this. There are, uh, two layers, let's say, for this, uh, [00:28:46] this demo ELMA model. The forward and both of these layers has a forward pass and a backward pass. [00:28:53] So, our forward one looks left to right, and the backward one looks right to left. So, given any input token TK, [00:28:58] And its input token embedding XK from the CAT level CNN. [00:29:03] The XK is tied to the layers of the Emerald model. Um, every token… [00:29:08] goes through all of the layers, and they get an output embedding. The hidden state from… so these, uh, LMs are nothing but, uh, LSPMs, um, and the hidden state of these LSPMs are taken as the output for each of the layers. [00:29:19] And similarly, we take the backward LM, and we take the hidden state from that. [00:29:23] then these hidden states are concatenated across the layers. So this one is concatenated with, um… this is concatenated with this, this is concatenated with this, this is concatenated with this. [00:29:32] And we get, uh, these outputs. And then we take a wicked sum across all of these weights. These weights are task-dependent, these are trained, [00:29:41] Uh, based on what is the downstream task that this Elmo embedding is being used to achieve. [00:29:45] So, for this token, PK, we get these outputs. These 6 outputs are then weighted by the, uh, [00:29:52] by the task weights, so… [00:29:55] For the 08 layer, for the first layer, for the second layer, we have, uh, [00:29:58] waits for a given task. We take the weighted sum, [00:30:01] Which is, uh, multiply this, uh, weight by this entire vector, this weight by this vector, and this weight by this vector, and we take an average across all of this, the weighted sum. [00:30:10] Um, and then… [00:30:12] We do, uh, scaling. The scaling factor is also trained, and that output is basically the output of the token. So that is the contextualized text embedding for the input token, TK. [00:30:23] And this is how EDMO works. So Elmo is a task-specific representation, and a downstream task needs to learn how to take these output token embeddings and infer what they actually mean. [00:30:35] So, uh, now, this was the, uh, background to Bert, and then coming to BERT itself, [00:30:41] BERT is bidirectional encoder Representations from Transformers. Like Elmo, it is also bidirectional, that is, it allows any representation to have context from past and from the future of the occurrence of that given token. [00:30:55] It is an encoder-only transform model, hence it's called encoder. [00:30:59] And representation means, uh, like, it gives a bidirectional… it uses a bidirectional self-attention to represent the context of every token. So that is, uh, that is why it's called bidirectional encoder representation using transformers, from transform. [00:31:13] Um, you know, yeah, Gunjin? [00:31:18] Yeah, so, um, uh, if you go to the previous slide, just before just doing that… [00:31:22] Yeah. [00:31:23] The Elmo model, we're saying that, like, this one, like, backward and forward workings together. What it means that, um, [00:31:28] How do you explain earlier that, um… [00:31:30] In the forward pass, it look for the future [00:31:34] content which is coming up. And a power buyer is looking for the what is already there in the backward, right? [00:31:38] Yeah. [00:31:39] And in the backward, LM, it is checking what is coming future, right? [00:31:42] Yeah. [00:31:43] So, in this ALMO model, both are working together. [00:31:46] Yeah. [00:31:48] So, uh, okay, so at the same time, when the inputs were given, [00:31:52] it subsequently, or as I parallelie, is checking forward and back-end [00:31:56] backward alum both together, is that correct? [00:31:59] Yeah, you could think of it as a forward pass, where you take the input sequence, you give it T1, T2, T3 as the input to the forward element. [00:32:05] Yeah? [00:32:06] And you get an output from that, and you concatenate across that. And then you take the backward element, for example, that is another LSTM, like, it's basically a bidirectional LSTM model. [00:32:16] Uh, it can be done, like, given an input sentence, it can be run pilotally, or you could think of it as running left to right and right to left, um… [00:32:24] sequentially. Either way, the hidden states at each [00:32:29] output are concatenated after the entire sequence has been processed by both the forward and the backward end. [00:32:36] Okay, so it's not necessarily it's going to be parallel, it's going to be sequential also, right? [00:32:40] So let's say, once you complete the forward pass, then start the backward pass. Is that completely correct? [00:32:44] Yeah, you can think of it that way as well. It's basically a bi-directional LSPN. It's… or a bi-directional RN. [00:32:50] Okay? Because as the arrow shows, like, it is something… [00:32:55] let's say one step is getting, like, hidden layer is passed, it is going to the backward, then from backward again, it's coming to, like, forward, like, this is, like, zigzag, it's coming up, but it's, like, complete forward, and then… [00:33:04] Maybe that's what? [00:33:05] No, there is no exact year. Like, given TK, it's looking… [00:33:08] behind it. K-1. [00:33:09] Yeah? [00:33:10] So, this is basically the output from the current one is fed to the next one, right? Because it's a recurrent neural network, uh, if you're familiar with RNNs, you know that the hidden state is passed across that. [00:33:20] So, it's called a backpropagation through time. So, given T1, it produces the output hidden state, um, and then the hidden state is passed over to K plus 1, TK plus 2, so on and so forth. [00:33:31] For the backward LM, you start the other way around. So you give TN, [00:33:35] Mm-hmm. [00:33:36] And, uh, then you get GIFTN minus 1, TN minus 2, and so on till TK plus 1, and then you give TK. [00:33:41] So, that's the only difference there. [00:33:44] Okay. And you mentioned that in the Alma model, there's 93 million [00:33:48] layers are there, is that correct? Or tokens are there? [00:33:51] 93 million parameters, the number of parameters, yeah. [00:33:52] Pat on the good. [00:33:54] Okay, got it. Okay. [00:33:57] Uh, yes, Rachel? [00:34:00] Yeah, can you go back to the previous slide? So, I have one question, regarding the [00:34:02] Yeah? [00:34:07] Yeah. [00:34:08] The input tokens are processed by a character-level CNN, right? So, since this is a vision-based model, right? So, I mean, how exactly it works? [00:34:13] it's… it's not a vision-based model. It's a input text to output text embedding model. [00:34:18] It takes text as input, like, as we saw here, it takes text as inputs and produce embedding vectors. A contextualized embedding vectors. [00:34:25] for each of the input tokens. [00:34:29] So… [00:34:30] It is a, different than what we have [00:34:34] the Cnn model for the vision. [00:34:37] It uses convolutional layers, uh, convolutional layers are frequently used in, uh, [00:34:43] in vision models, as you say. But convolutional layers are also used in NLP tasks. [00:34:48] Okay. [00:34:54] Uh, yeah, so continuing to this, um… [00:34:57] The input to BERT is a sequence of tokens. Uh, these tokens are obtained by the tokenization of a given input text. So, given a sentence, the tokenizer, uh, for BERT model runs on itand produces a sequence of tokens. These sequence of tokens are fired as an input to the BERT model. [00:35:12] And the output that we get from the BERT model is a set of contextualized token embeddings for each of the tokens. [00:35:19] So, given a sequence of input tokens, BERT also produces a set of token embeddings, one corresponding to each of the input tokens. So if you have n input tokens, you get n output token embeddings. [00:35:29] Each of these token embeddings have the context, or the contextual meaning of that given token in the entire sequence. [00:35:36] So this is also a text encoder model, um, and the maximum sequence that BERT [00:35:41] Like, the bird-based model can process at a given time is a sequence of PHY12 tokens. So, if you have a text that's larger than 512, we have to split it down into chunks of 512 tokens. [00:35:51] And if it is smaller than Phyto, then you can tag it to FITOL tokens. So, have a fixed size input to work. [00:35:57] So, uh, but the maximum number of tokens that BERT can process in a singular pass is 512. [00:36:04] Now, let's look at, uh, for example, Elmo model and BERT model side to side. So, given an embedding, or, like, given an input token even, it takes and produces an output embedding, uh, T1. [00:36:15] for that given token. [00:36:17] So, so on and so forth. So, for every single, uh, like, at every… so these are the transformer layers, there can be N transformer layers in BERT. So, for every layer, you can see that it takes [00:36:28] the entire sequence into consideration, including the past and the future. For example, even you can see everything appearing after even is taken into consideration to pro… to give [00:36:38] and embedding for that token. While anything, like, for E2, anything from the past and anything from the future is considered. So it is bidirectional, similar to how Elmo works, where you have a forward LM and you have a backward LM. Um, so as you can see here, the backward LM starts with EN, [00:36:53] and goes backwards till P1. And the forward element starts with E1 and goes forwards. [00:36:58] Till Ian. And the output from both of these are concatenated and then weighted across the entire number of layers of ELMA to get the output, uh, [00:37:07] token embedding for each input token. [00:37:10] Now, Elmo uses a concatenation of forward and backward elements, as we just saw, and BERT are jointly conditioned on both past and future contexts simultaneously. [00:37:18] As we saw here. So, this is where self-attention diverges from [00:37:23] audience. Uh, but can process the entire sequence simultaneously. You give E1 to EN together, [00:37:29] the total transform layers, and… [00:37:32] This self-attention mechanism takes care of allowing you to get, uh, context from both past and the future. [00:37:39] While Elmo, for end to have to run it, [00:37:41] at n steps, for example. So, for each step, you process a singular token. [00:37:46] And the outputs across those are then taken to be the token embeddings. [00:37:53] So, what was Bert trained on? What is the training task? BERT basically has two basic training tasks. [00:37:57] One is mass language modeling, the next is neck sentence protection. [00:38:01] So, what is the masked language modeling task? So, given an input sequence, [00:38:06] 15% of those are input tokens in the input sequence are masked. [00:38:11] Like, approximately. So, basically, what it does, it takes 15% of those input tokens, [00:38:16] 80% of these, 15% tokens are replaced with a masked token here, and 10% of them are replaced with a random token, so out-of-place token. [00:38:25] And the remaining 10% of these 15% are kept unchanged. And this is then fed as an input to the word model. So, for example, if we take this as an input sequence, and let's say improvisation was selected as one of the tokens to be masked. [00:38:39] Randomly. So, we mask it, and then we feed this as in the masked sequence as an input to BERT. [00:38:45] We get the output, we get a token embedding for every single token. And then, [00:38:50] we train a classifier, basically, to tell us what this… what the output embedding to this mask [00:38:58] token embedding might actually signify. And the task itself is that [00:39:05] These mass opens output embedding, after going through the classifier, should predict that improvisation [00:39:09] was the token that this must open represents. So, this is the mass language modeling task. It is asking Bert to predict what a given token was. [00:39:17] Given a sequence, it's asking Bert to learn the entirety of the sequence, and see, based on the context from the past and from the future, [00:39:25] what should be the token that appears here? And this is one of the training tasks of Word. [00:39:31] Use the output of the masked words position to predict what the masked word was. [00:39:34] And this is mass language modeling. And the next one is… next sentence prediction. [00:39:38] So, what does the next sentence prediction? Well, [00:39:41] given, uh, given a sequence of sentences, [00:39:44] For example, the man, uh, master Penguin, are flightless birds. Sentence A and sentence B are two of these sentences. We must predict [00:39:52] Whether sentence B occurs after sentence A or not, like, whether it makes sense that sentence B is a continuation to sentence A. [00:39:59] So the input is the two sentences are separated by a special separator token, set here. [00:40:04] So we… and the starting button was always CLS to Bird. [00:40:07] Then, uh, given any input sequence, we randomly select a split, [00:40:12] For 50% of these splits, we replaced the split B with some other sentence split B, so that, uh, we have a negative case. And for the remaining 50%, we use the original split. So that's the positive case. [00:40:23] Uh, the output from the bird, uh, the CLS token here, [00:40:27] we'll learn about this in more detail later. This yellow list open here gives the sentence embeddings. [00:40:32] Okay, so the… [00:40:35] classifier is then trained on the sentence embedding to see whether the sentence embedding makes sense or not. [00:40:41] So, if it does, then the classifier takes this and predicts whether the sentence V was next in the sequence or not. So, whether sentence B was [00:40:49] Uh, whether it makes sense for sentence B to occur after sentence A. So, is this a continuation sentence or not? So, that is next prediction. [00:40:57] These are the two tasks that Bert was trained on. [00:41:00] And the loss function for BERT was basically the sum of the masked language modeling likelihood and the next sentence prediction likelihood. [00:41:08] So that was the training loss function that Word was used. It needed to maximize the, uh, the likelihood for mass length, the negative block likelihood for mass language modeling, and for next sentence prediction. [00:41:19] So, the sum of these two losses was used at this training loss function for BERT. [00:41:24] Uh, the dataset that Bert was trained on was the English Wikipedia dataset, which consisted of 2,500 million words. [00:41:30] And, um, the books Corpus dataset, which consisted of 800 million words from about 7,000 unique published books, [00:41:37] for various genres. [00:41:40] The word-based model had 6 transformer blocks, the embedding dimension was 768, the maximum sequence length was 512, and the total number of parameters was 110 million. [00:41:51] For the smallest word-based model. [00:41:53] So, at a very high level, this is how the encoder architecture in BERT looks like. [00:41:59] you are given, uh, like, you are given a… [00:42:03] given any input sequence, you're given a sequence of tokens by the tokenizer. CLS, I am fine, except, for example, might be a sentence, and this would be the output from the tokenizer for that sentence. [00:42:12] Um, then each of these tokens are, uh, [00:42:16] convert it into input token embeddings. [00:42:18] using the token embedding, um, [00:42:21] clear of the BERT model. [00:42:23] And then you, uh, you proceed to do the, um, [00:42:27] multi-head self-attention here. [00:42:30] And the outputs of each of those heads are concatenated, and then you project it outwards. You do, uh, addition and normalization. [00:42:36] Andthen you proceed to the feed-forward layer, and then you do another addition normalization, and you proceed to the next layers. Um, so… [00:42:44] You can, like, for 6 layers, you do this entire process, uh, from self-attention to normalization to feed-forward to normalization six times. [00:42:51] And the output is the contextualized text embeddings for each of those tokens. [00:42:56] So, this is a very high-level visualization. We shall now look at, uh, [00:43:00] at each of these steps in detail. The tokenization step, the embedding step, [00:43:05] the, uh, attention, the feed-forward, and then the output. [00:43:09] So, how does the input tokenization work? What is the tokenizer that BERT uses? BERT uses a word-based tokenizer. [00:43:17] as the WordPeast organizer starts from a small vocabulary, uh, including the special tokens like separator, CLS, mask, and so on, that are used by the model, and the initial alphabet. The initial alphabet is basically, uh, all of the unique letters occurring in the entire, uh, [00:43:33] training vocabulary. So, it then identifies subwords by adding hash-hash prefix to each of these, uh, each of the letters that occur within a word. So, for example, given a word, word, [00:43:45] It would be split as W, Hashash 4, Ashash R, and Hashash D. The hashhash signifies that this is not the beginning of a word, but [00:43:53] an intermediate word, like… [00:43:56] All others, um, somewhere inside a word and not at the beginning of a word in this instance. [00:44:03] So, what we stoked as the first trains by taking these special tokens and the initial alphabets, and then identifying the subwords present in the word. [00:44:13] And next, what it does is, now, given these tokens, it needs to learn, how do I merge these tokens into the word word. [00:44:20] Right, so if word were given, how do I tokenize this word into these tokens? So, for this, we need to run the merge rules. [00:44:26] The merge rules are learned based on a score, which is computed as frequency of the pair divided by frequency of first times frequency of the second token. [00:44:33] appearing in the entire purpose. [00:44:36] So, how does this all look when put together? Let's take an example and look at it. [00:44:41] So, for example, if I were given the [00:44:43] help us hog occurring 10 times. So, I am given a document which has hug occurring 10 times, hug occurring 5 times, pun occurring 12 times, [00:44:52] One occurring four times, and hugs occurring five times. [00:44:56] So, this is my entire corpus. I first… [00:44:59] begin by splitting the words into their, uh, initial alphabet. So, for example, hug is split as H, hash as you, hash as G. Similarly, pug, we get P, U, N, G, and pun, P, U, and N. And pun, B, U, and N, and hugs, H, U, G, S. [00:45:12] So, now, from this, my initial vocabulary would be all of the unique tokens amongst all of this. [00:45:18] So, the unique tokens amongst all of this are BHPGNS and U, where G, N, S, and U occur only in the middle of any of the words. They are not beginning, uh, letters for any of the words. As we can see here, there is no word in our vocabulary that begins with G, N, S, or U. [00:45:35] they only begin with H, P, or B. [00:45:39] And that's why BHP are the only letters in Anish vocabi that do not have the hashtags. [00:45:45] So, uh, the initial vocabulary can consequently be, uh, derived from this corpus after the splitting as this. [00:45:51] Next, we now need to learn the merger rules. [00:45:56] that how do we merge this initial vocabulary into the token part? [00:45:59] fun, fun hugs, and so on. So how do we merge these tokens to get the words that we desire in our vocabulary? [00:46:04] So, this is done by considering all of the combinations of, uh, [00:46:11] well, of tokens that can be merged. For example, B can be merged with G, B can be merged with N, B can be merged with S, and B can be merged with U. Similarly, HD, HN, HS, and HU. PG, PN, PS, and PU, and so on and so forth. [00:46:23] We cannot merge B and H, because both of them are beginning letter words, because B and H simultaneously cannot be beginning of any particular order. So we only take hashh. [00:46:34] As are, uh, this thing. [00:46:36] But we can always merge G and N, because GN or UN, for example, might be occurring somewhere. [00:46:43] Uh, in the same… in the middle of a world, so… [00:46:45] After taking all of these candidate merges, let's, for example, consider the candidate merge U and G. [00:46:51] But UNG together is president, uh, is present 20 times in our vocabulary. So, UG occurs 10 times in HUG. [00:46:57] usually occurs 5 times in perk, and usually occurs 5 times in pucks. So, for a total of 20 times in our vocabulary. However, when we consider the individual frequency of U and G in our vocabulary, we get that U occurs [00:47:10] about 10, 15, [00:47:12] 27, 31, 36 times. [00:47:16] And G occurs 10, 15, [00:47:20] And… 20 times. So, when we use our metrics, which is frequency of pair, which is… [00:47:27] 20 divided by frequency of first, 36, times frequency of second, 20. [00:47:31] We get the score for this candidate merge, UNG, as 1 by 30 cents. [00:47:35] And, uh, similarly, we consider various other pairs, but the one that gives us the best score currently is G and S. When we calculate the score for GS, it turns out to be 1 by 20, which is the maximum amongst all of the candidates at this stage. So, we merge [00:47:49] GNS into GS. [00:47:52] And so now, our vocabulary also includes a token. [00:47:55] GS. So, yeah. [00:47:58] It also includes a token.js here. [00:48:00] And the new splits, after merging G and S. [00:48:04] where, for example, we had GS together, we merged this into GS now. So the new splits become this. [00:48:11] And this way, we continue learning merges till we reach our desired vocabulary sites. Let's say our desired vocabulary size was 10 for this current example. The final vocabulary after all of these merges are learned would be BHP, G, N, S, U, G, S, HU, and HUG. [00:48:27] So… [00:48:29] Now, with this vocabulary, we can tokenize, uh, like, all of these words. For example, let's take the word hugs and tokenize it. [00:48:36] Um, we try to maximize the length of the token. [00:48:40] So, starting with Edge, Edge is a token. Well, HU… [00:48:43] HU is also token. HUG, HUG is also token. So we take HUG as the longest subsequence, as our first token. [00:48:51] And then we have S, hashh S. So this is our token here. [00:48:54] And so bugs, for example, we have B, but then there is no BU. [00:48:58] in our vocabulary. So, we split Act B, and then we get [00:49:02] UGS. Now, UGS is not in our vocabulary, so we need to further split it into U and GS. [00:49:09] So, we get BUGS as… [00:49:12] the output tokens for the word bugs. [00:49:14] at any step during the tokenization of a word, if we cannot find this [00:49:18] like, if they cannot find any split that satisfies this. [00:49:23] then the entire world is tokenized using the special token unknown. [00:49:28] UNK. [00:49:30] So, this is how the birds input tokenization works. [00:49:34] Any doubts? [00:49:43] So, after the input tokenization, we then go to the input representation. [00:49:49] So, we fill it the input sequence to… [00:49:52] the WordPase Organizer. Word-based organizer is already pre-trained, it just does this tokenization step. [00:49:57] Uh, for both. Which gives us a sequence of tokens. [00:50:02] These sequence of tokens are converted into token embeddings. The token embeddings is basically a set of, uh… so for every [00:50:07] token in Words vocabulary. [00:50:10] BERT has a vocabulary of size 30,522 tokens. That is the, uh, that is the model's vocabulary, or the tokens that it understands, the unique tokens that it understands. Any, uh, sequence can be represented as a combination of this set of unique tokens. [00:50:25] So, [00:50:27] Every one of these tokens has a unique token embedding, uh, for the model. [00:50:32] So, given our list of tokens, [00:50:35] We use the, uh, this pre-trained token embeddings that are present in the model suites to get the token representations [00:50:43] token embedding representations, initial token embedding representations. These are non-contextualized for [00:50:50] the input to the BERT model. So for every token that the word piece models vocabulary, uh, outputs, [00:50:55] Uh, in our input sequence. We take, uh, we use this embedding, uh, [00:51:00] matrix that we have in, um… [00:51:03] in the model, and we look it up as a lookup table via the tokenID. So, for example, token 0 will… [00:51:10] will have the, uh, token representing token 0 as its, like, token embedding representing token 0 as it's embedding. [00:51:17] So, the next part is segment embeddings. So, since BERT is also trained on the next sentence prediction, uh, [00:51:23] task. It needs a way to distinguish between sentence A and sentence B. So, this is called the segment embeddings. So which segment is this? Is the segment A, or is the segment B? [00:51:32] So these embeddings are present [00:51:35] how to distinguish between the sentence pairs for certain tasks, such as nest in this prediction. [00:51:40] If we are only supplying a single sentence to bird as input, we only use the segment A as B segment embeddings. [00:51:47] And then, third thing that BERT requires is a way to distinguish between the positions of each of the tokens. So, the same token may appear multiple times in a sentence. The same word may appear multiple times in a sentence. [00:51:58] But Bert needs to know that this [00:52:00] Each of these occurrences is different from the previous occurrence. So for that, we need a way to express the position of word in a sentence. [00:52:08] And for that, Bert has position embeddings. [00:52:11] Bird uses sinusoidal position embeddings as its position embedding. [00:52:15] So, [00:52:16] So, just a second, sorry to interrupt. [00:52:18] Any questions so far, uh, because it's a very important concept. [00:52:29] Yeah, ma'am. In the previous slide, [00:52:31] When you say the vocabulary size is 30,52, right? [00:52:35] Tokens only. So… [00:52:36] Yeah, yeah. [00:52:39] In general, we have seen that in the… [00:52:41] The latest models, in billions, millions tokens are there, right? So… [00:52:45] If something is coming up in between, then how this bird will take care of other words, the vocabularies. [00:52:51] So, uh, I think for the millions that you, uh, refer to for newer models, that's the context size. That, for BERT, is 512. [00:52:58] The vocabulary size is… so, basically, when you train a model, you are given [00:53:05] training data. You also train the tokenizer on that training data. So, you need a way to express, uh, like, for us, we are the alphabet, right? We can express any word as a set of alphabets, or a set of letters from the alphabet. [00:53:16] So, similarly for the model, this 30,522 tokens are all… [00:53:21] It requires to express any complex sequence. So, given any input, it should be able to use its vocabulary. It knows, it understands these letters, these formation of letters. [00:53:31] And it can then express this input sequence as a combination of those. For example, [00:53:36] like, once we have learned this as our vocabulary, the vocabulary sizes in this example is 10. [00:53:40] But once we have learned this, uh, as our vocabulary, we can then represent any of the letters in our corpus, and hopefully any of the letters that we encounter in our validation dataset as well, using this vocabulary. So, for example, once is tokenized as B, U, G, and S. [00:53:55] Together. So this is the input organization for this word, bugs, which is not present in the original vocabulary, for example. [00:54:02] Right. So… [00:54:03] Okay. [00:54:04] The vocabulary size here refers to the number of [00:54:08] Um, letters, or the building blocks that the model understands to represent complex input sequences. [00:54:14] Okay. [00:54:17] It's separate from the context size, which is the maximum number of tokens that any model can process. [00:54:23] So, for BERT, it can only process 512 tokens at any given point. So, if you have a sequence that's 1024, you take 512, and then you need to shift it [00:54:31] For the next 5 months. There are various ways of how you shift, because, uh, you need to pass embeddings from the previous one to the next one as well. [00:54:37] Yeah. [00:54:38] things can change. So, but yeah, that is the drawback of the original BERT encoder model. You only have sideways. [00:54:44] For your context length. [00:54:45] Okay, Mark. [00:54:48] Um, I'm good. [00:54:51] Sure, can you go back to the previous slide? [00:54:53] Yeah. [00:54:55] So, here we… you mentioned in the step, right, that we will be reducing the size. So, in the fourth step, we were creating the splits, and then we reduced the size, right? And we introduced HU, HUG, and GS. And the size we made [00:55:10] You mentioned desired vocabulary size. [00:55:11] Yeah, desired vocabulary size is an upper element, that's it. That is basically set using, uh, the maximum number of points. So, [00:55:19] After a given point, uh, your tokens can keep expanding, but they do not consecutively add your Expressionism to that vocabulary. So, like, for example, with Bohab was found that 30,522 tokens are enough, uh, to represent [00:55:32] input sequences, or diverse set of input sequences. [00:55:35] And, uh, we can, of course, keep merging [00:55:38] post that as well, but we do not get enough representability out of it, and we also need to train a model to understand those many tokens, which adds to the cost. [00:55:46] So, that's the reason that the vocabulary size is set, uh, to, uh, let's… for this example, we set it to 10. So we only… [00:55:54] learn merges till the vocabulary size reaches 10. So, since we started with an initial vocabulary of 8, we only learned 2 merges, HU and GS. [00:56:03] Yeah. [00:56:04] HUG, so for example, like that. So, this particular one… oh, sorry, we started with 7C, we only learned 3. [00:56:08] G-S-H-U-N-H-U-G, as, uh… [00:56:11] Uh, noches. [00:56:17] Yeah. [00:56:18] Yeah, so for BERT, when they would have trained it, so 30,522 was the number they came up with [00:56:21] Yeah, for the… yeah. [00:56:29] So… [00:56:30] Ultimately. And it was never expanded after that, or it was just in the first iteration, it was 30,522, and now it might be something else. [00:56:31] So, BERT model currently also uses our tokenizer with 30,522. Subsequent models have used different tokenization strategies, like, [00:56:39] But what pieces use here. There is also tick tokens, sentence fees, byteware encoding, and other tokenization algorithms. [00:56:48] Okay. [00:56:49] So, subsequent models used different tokenization algorithms, so it's a different vocabulary size and different merge size as well. [00:56:56] Okay, got it. [00:57:01] So, yeah. [00:57:04] Now, uh, so, coming back to input representation… [00:57:08] given any input sequence, we, uh, tokenize it, and if we have two sequences, like sequence A and sequence B, we… [00:57:14] tokenize it into, uh, we tokenize it and separate it with a separate token. [00:57:19] If they have a single one, it ends here. So this is the tokenizer, uh, tokenizer output. [00:57:25] After that, we then, uh, convert this to token embeddings, uh, using the input token representations for each of the tokens in the model's vocabulary. So we have, like, the embedding for CLS, embedding for my, and dog, and so on and so forth. Then we take the sequence embeddings. [00:57:40] And we add that to, uh, the… [00:57:43] token embedding. So, for sequence 8, like, the first sentence is sequence A, so we are embedding A, [00:57:48] from the segment templates. And the second sentence is sequence B. So we add D-segment embeddings, embedding B, corresponding to embedding B into the second set of tokens. [00:57:56] And then, finally, Bert needs to represent the, uh, [00:58:00] the place, or the, uh, the position at which each of these tokens are appearing. So, it adds position embedding. So, every position [00:58:08] from 0 to 511, which is 5 world positions that both can process at any given time. [00:58:13] So, it takes, uh, the position embedding corresponding to position 0, adds that to the first position. [00:58:19] then E1 to the second position, E2, and so on till B times for the 11 positions. [00:58:23] And this gives, uh, input, uh, this gives us the input representations that we then feed in. [00:58:29] to the, uh, the decode, like, the encoder, uh, layers of the model. [00:58:34] Right? So this, this part here. [00:58:38] We take this input sequence and convert it into an input representation, or input embedding. So this [00:58:43] is the input embedding that we then feed into the, uh, the decoder layers, right? [00:58:49] So, this is the input representation. [00:58:52] Um, now, what is this position embedding? Uh, like, as we saw, uh, it was a sinusoidal position putting. How is it calculated? So, given the dimensionality of the model, where DD is 768 for birth, [00:59:03] Um, we… Science World position embeddings are obtained as, uh, [00:59:07] using this formula. So, like, for the I-th position, where I can be 0, 1, 2, 3, so on till 5.11, because it can take 5 quills, input, uh, representations. [00:59:17] So, each of the, uh, IAT position, the position embedding is calculated using sine of [00:59:23] I buy 10,000 to the power of 2 by D. [00:59:26] two into one byte. So for the… for the first set, it's… [00:59:29] 2 into 1 by D, and for the last little bit D by 2, giving us a total of D, uh, dimensional. [00:59:34] position, uh, representation for this model. [00:59:38] Now, why do we use sinusoidal precision embeddings? Well, for any fixed offset from a given position. So, given the i-th position and the I plus k-th position, where k is of the fixed opposite, we can easily represent the pKth position, i plus k position, [00:59:52] as a linear transformation, or a rotation of the initial position. [00:59:57] PI. So it… it helps the model to understand [01:00:01] the deviation between the positions as simply a rotation between, uh, [01:00:05] the key positions that [01:00:07] are present between PI and PI, basically. [01:00:10] So, this allows the model to attain to relative positions quite simply. [01:00:15] Because PI and PI plus K only differ by the rotation of K. [01:00:20] Or the linear transformation matrix that can be, uh, [01:00:24] And that can be used to rotate this IH position by k times to get the PI plus K. That's the reason to use sinusoidal positioning units. [01:00:32] So… [01:00:34] we can have a quick look at the code and how that works. So, for example, given this [01:00:40] our sequence tell me about Qualcomm. [01:00:43] Um, let me just shorten this to Qualcomm. [01:00:46] Okay, and let me go to… [01:00:49] encoder freeway. [01:00:52] As we were just seeing that BERT encoder uses this. So, the first step would be to clean the input sequence, because it needs a representation that is static, like, [01:01:02] queue can be capitalized, Q can be small, so it converts everything to lowercase. [01:01:06] It normalizes any special characters that might be present. [01:01:10] So, this character, this Unicode character, is split, uh, using this normalization, NFT normalization export. It's called normalization from canonical decomposition. [01:01:19] It is split into the base character, and then the accent to that base character. And subsequently, in the next step, this accent is stripped. [01:01:25] So, this is cleaned. [01:01:28] to only give us the base character. So this allows us to do our input cleaning and use our, uh, use the vocabulary that we know. So, for example, if I were to give Nino as the input, [01:01:37] Um, the tokenizer would internally first make it Nino, and then make it smaller case, Nino. [01:01:42] Similarly, for Manchester Munich in German, for example. [01:01:48] Uh, then it splits it into a set of, uh, set of words, and then each word is, uh, then split using the word piece tokenize. [01:01:56] So, let me just go here. [01:02:19] Give it a second. [01:02:20] And, uh, all of you should be able to understand it very nicely. Any questions, anyone? [01:02:25] Uh, maybe you can enlarge it slightly, uh… [01:02:30] Okay. Um… [01:02:31] Yeah, for better visibility. [01:02:32] Let me just open it in our… [01:02:39] Is this better? [01:02:43] Yeah, yeah, yeah, sure. That's better. [01:03:15] If I were to go here, and… [01:03:18] bring the basic tokens. [01:03:21] tell me about Polycom, which was my input sequence. [01:03:24] was first converted into lowercase and cleaned, and then each of these words was split. [01:03:30] by the basic tokenizer. We have tell me, about, and call form. And then we come to WordPeach tokens. [01:03:36] So, looking at the vocabulary for BERT, for example, [01:03:39] tell, we can see, would be present, um… [01:03:42] As a word. [01:03:44] me would be present as a word again. [01:03:47] And about it also good. [01:03:49] So, each of these can be singular tokens, but there is no wall form token. [01:03:54] So this needs to be split. Let's look at how that split happens. [01:03:58] Next. [01:04:03] let's say I can… [01:04:16] And is it open? [01:04:18] I'm doing… [01:04:20] I mean, is it open? [01:04:23] about is it open? [01:04:27] coming to Qualcomm, right? [01:04:30] So, this needs to be tokenized. [01:04:32] So, let's see how that tokenization happens. [01:04:37] Okay. Next. [01:04:40] when it started as a name, joincast. So let's… [01:04:43] Consider the first substring, greedy by first rate, sorry. [01:04:49] The first substring that we see is the entire substring. We consider this, is this a part of the vocabulary or not? We consider, and we find whether this is a part of vocabulary. If it is a part of the vocabulary, then it will be tokenized as is. [01:05:00] But since this is not, it has to be split up. [01:05:06] Great. So, see, again, we have to split it. [01:05:09] So, we consider the next one. [01:05:11] is Qualcomm the singular M? We remove the M. Is this part of the vocabulary? [01:05:17] We continue doing this. [01:05:24] we get to QUA. [01:05:28] And, like, we have kept looking Q-U-A-L-C, is it a part or not, and so on and so forth, we keep reducing it. [01:05:40] So, if you look at this… [01:05:43] We have finally found QU as a part of the vocabulary. [01:05:47] So we use that. So, if you, for example, look here… [01:05:50] You will find that QU is a part of the vocabulary. [01:05:53] So, it has tokenize the first Qualcomm. It has tokenized the first part of that word as QU. [01:05:59] And next, if we keep continuing it. [01:06:04] If you keep continuing this, uh, we would get to… [01:06:08] the input tokens. [01:06:12] So, for example, this would be the comparator. [01:06:15] Just be good. [01:06:19] Yeah. So, this would be the set of tokens that you would have gotten for Palfun. Tell me about, and this would be a set of purpose for both. [01:06:27] This is how it works. QU, and next you would find me. [01:06:30] A, as the next open. [01:06:32] So the EL is the next open, and then you would find, uh, COM as the next open, and then a singular M as the last open. [01:06:39] So that would be this set of tokens. [01:06:44] And this is how the tokenization works. [01:06:48] Now, let me remove this, and let's see how the input representation has done, right? [01:06:54] So, it may then go… so this is the input representations. We need to combine the word emittings, the position embeddings, and the token type embeddings. [01:07:02] OpenType is basically the segment embeddings. [01:07:05] So, this is how that is done. [01:07:08] Um, if I were to put a breakpoint here, for example. [01:07:15] Let's start again. [01:07:18] missed. Now, you can see, [01:07:21] my input IDs would be the… [01:07:23] the sentences… [01:07:26] sequence that I have. So this is the CLS open, this is the separated open. This is the tell me about and [01:07:31] Qualcomm talks. [01:07:33] So for each of these tokens, it'll look up in this params.embittings. This is the… [01:07:40] So this matrix is the embedding matrix. If you look at the shape of this, [01:07:51] Excellent. [01:07:52] she explored, yeah, everything so forth. [01:07:55] Yeah. [01:08:00] It was… yeah. [01:08:08] Yeah. This is 30,522 and 768. [01:08:12] So, $30,522 is the number of input tokens, maximum number of input tokens, and 768 is the dimensionality for both, right? [01:08:20] So, uh, let's say… [01:08:24] Right? So if I were to convert this into a token embedding, all I would do is… [01:08:33] And this would be the embedding that I have. [01:08:35] as an input to the model for that, right? [01:08:38] And similarly, we have the position emittings. For position IDs, position IDs are simply 0 to n, which is where N is the sequence length, and similarly, we have token type items. [01:08:47] So, we can have a look at token type IDs as well. It will be of size 2. [01:08:57] Sentence A and sentence B, each having dimensionality of 7680. [01:09:08] So… [01:09:11] Is this part clear? [01:09:14] Um, any questions? [01:09:26] So, this was about input representation. Now, let us come to self-attention. So, the next, uh, building component of the BERT model [01:09:34] is the cell potential. Once we have fed the input representation into the BERT encoder, it understands the context of each word using the multi-hit self-attention mechanism. [01:09:43] This self-attention mechanism is where it relates each word in sentence to all the other words in the sentence, and learns the relationships and contextual meaning of each of these words. [01:09:53] And it is because of this self-attention mechanism that it is able to build a contextual representation of each of the input tokens in a given sentence. [01:10:01] So, given an input embeddings matrix X, [01:10:05] We create three methods. Like, this is the input, uh, this is the input representations. X is the input representations that we just saw. [01:10:11] So we, uh, we create the query, key, and value metasys by multiplying the input matrix. [01:10:18] With the, uh, the weights. The query weights, the key weights, and the value weights, right? WQ, WK, and WV. [01:10:25] So, this is how, uh, the matrices are, like, the QK and V matrices are computed. [01:10:31] And then we proceed to the self-attention computation. [01:10:34] So, what the self-attention does is it first computes the dot products of Q and K. [01:10:40] And then these dot products are scaled by the dimension, uh, the underroot of the dimension. So, D768. So, for example, that could be under root 768, or the head dimension 768x12. [01:10:51] And then the softmax function is applied, and then we obtain the attention scores by multiplying it with the value matrix. [01:10:57] So, what… what is this query? What is this key, and what is the value? So, given any input sequence, the query is the token itself. [01:11:04] The key is the comparison amongst all the other tokens. [01:11:08] And the value is the re-weighting of a given token based on the context that it got from [01:11:15] all of the other tokens. [01:11:17] So, for example, let's take the sequence, tell the animal didn't cross the street, because it was too tired. [01:11:24] So, here, we can see that it [01:11:26] the context for it comes from animal. What was it? The animal? [01:11:32] The basic question that, uh, we need to ask in order to understand this [01:11:35] sentence is, what does it refer to? See, it was too tired, but there must be some attention that is paid to it by too tired, because what was it is too tired. [01:11:44] And what does it refer to? Well, it refers to the animal. [01:11:48] So this representation is one of the heads of the BERT model. [01:11:52] when we supply this… the animal didn't cross the street because it was too tired as an input. This is… this is… this is how, uh, attention might look like. [01:11:59] for a given head. So, what does it do? [01:12:03] we first convert each of these tokens into an input representation. The input representation is then multiplied with the query key and value weights in order to get the query key and value matrices. [01:12:13] the query is the token itself, this token. It is the query. [01:12:18] The keys are all of the other tokens that it is comparing against. It is bi-directional, because as we can see, anything that appears to the future of it can also pay attention to it, and anything that appears to the past can also pay attention to it, right? [01:12:30] it can receive attention. Like, when animal occurs, animal doesn't know what is in the future of it, but animal still pays attention to it. [01:12:37] Great. So, it's bi-directional. [01:12:42] Anything from the future, and anything from the past is paying attention to it. So, the context for it is [01:12:47] achieved by asking two questions. [01:12:50] what was it, and what does it refer to in this case? And this is how the self-attention mechanism [01:12:56] allows for, um, allows for a contextualized representation, or contextualized token, uh, embedding representation. [01:13:04] By understanding what [01:13:07] a given token means in a particular context. [01:13:10] So then, the next question is, why multiple heads? [01:13:13] Why could we not use a single head? Why do we need to do this per head? [01:13:17] Great. So, what is the concept of heads? Well, each head basically does this Q times K transpose where it would be softmax times V as its computation, but each head is actually understanding different parts of the sentence. [01:13:30] Great. [01:13:32] Every sentence is composed of a subject, a verb, and an actor, and various other parts. So we need to learn the different relationships between different, um, [01:13:42] like, actors and different verbs. So, one headquart, for example, be learning subject-verb relationships, while another head could be focusing on the active-work relationships. [01:13:52] So, this is the core idea behind using multiple heads in parallel instead of a single head. Each head can attend to different contextual meanings, different parts of a sentence, and understand various different contextual meanings. [01:14:02] Uh, of the relationships, of the words in the sentence. [01:14:06] So, this attention of QKV, as I said, is… this is computed per head. [01:14:11] So, this is, like, a graphical way or an illustration to help visualize how it works. This is the input matrix, the input, uh, token embeddings. [01:14:20] This is then fed into the, uh, first decoder block, for example. So, these are the key and value matrices for each of the heads. So, this is… [01:14:28] like, every head has its own set of WQ, WK, and WV weights, for example. [01:14:34] There are several… let's say there were 8 hits, so we have 8 query key-value matrices. We multiply X, and we get QKV, QKV, QKV as, uh, the query key and value representations for the [01:14:45] input for every single head. [01:14:49] And then we do this attention computation using these Q, K, and V as our input, and we concatenate across the outputs. Like, output of attention is concaten across all of these heads. [01:14:58] So Z0 to Z7 are concatenated on this… the first axis. [01:15:02] And that is the output of the attention spouse. [01:15:06] These attention scores are… so you can even think of heads as subject vector spaces. [01:15:12] And these query key and value matrices, as predictions into those vector spaces. [01:15:16] The input is projected into each head's vetted space. [01:15:19] Each head understands a part of the sentence, or a different relationship between different parts of the sentences. [01:15:25] And then all of these relationships are concatenated, and then projected back into the model square space, where the model can then understand what, uh, the, like, [01:15:34] how these different parts come together to make the sentence. That is all done by the value mix, right? Like, once we multiply by the value, [01:15:42] And we concatenate, uh, we then… we, like, [01:15:46] Even after we have multiplied with the value, we are still in the heads vector space here, right? Because, uh, Z0 is only for head 0. Z1 is for head 1, and Z7 is for head 8, so… [01:15:55] We concatenate across this, and we project back into the model-select space before we can proceed further. [01:16:00] So, that is how the multi-head self-attention works. That is, uh… [01:16:05] that is the core idea behind it. Like, each head can attend to different parts of the settings. Now, let's visualize how different types might be getting different parts of the sequence. [01:16:14] So, let's take the same example. The animal didn't cross the street because it was too tired. And let's take, um… [01:16:19] the parts of the sentence that different heads are paying attention to. So, for example, if we look at this head, [01:16:25] And the token 8, it receives attention from the animal, and [01:16:30] That's it. So, this head is mostly on the subject-verb relationship. So, like, what is the subject? What is it? It was the animal. [01:16:38] or different head, for example, learns what was the action that was being done. So, for example, did it. [01:16:44] So this… this is… this head understands that it did it. It didn't do something. [01:16:49] So, another head might understand that what did it not do? Well, it did not cross the road, as we see here. [01:16:54] cross… it receives the most attention from Cross, for example. And why did it not cross the road? Well, because it was too tired. [01:17:02] was too tired. Here, this green head. Uh… [01:17:06] understands what it was. It was too tired. [01:17:09] So, this is how the various heads all come together to provide context for each of the tokens in a given sync. [01:17:17] Any questions? [01:17:24] Okay. So, as we saw, uh, the QKNV, uh, matrices are, uh, like, the QK and V vectors for each of the tokens are formed by, uh, a linear projection into the heads vector space. [01:17:37] Then we go into the scale.product potential, which is basically this formula. [01:17:43] Um, the scaling here is under root D, that's why it's scaled. [01:17:46] So that's the scale.predict Attention, then concat across all of these heads. [01:17:50] And we get the linear, uh, like, this is the output, which is back into the model vector space. [01:17:56] Like, after we have concluded all of this, these are still separate vector spaces. We must then come back into the models vector space that the model… the rest of the model can then understand. [01:18:05] So, which is this input vector space. [01:18:08] And this is the prediction back into that, right? [01:18:11] So, this is the part of the model, like, these are the multi-attentions. The linear predictions… [01:18:16] to queries, keys, and values. [01:18:18] And then the attention amongst all of these 12 heads that are present in the base BERT model. [01:18:24] Then we concatenate across this, and we project back into this input embedding vectors. [01:18:28] And we get the output as this. This output is then normalized, and the residuals are added. [01:18:34] And after that, we go to the feed-forward layer. [01:18:38] Which is the basic, uh… [01:18:40] transformer architecture. The forward layer then does an objection, followed by an activation, then the down projection, and then, uh, it is again normalized, and the residuals are added. [01:18:50] And then we go to the next, uh, decoder transform… no, sorry, encoder transform layer. [01:18:54] Here, and we should do this for 12 players. [01:18:57] So, the core idea, the core, uh, contextual [01:19:02] imposition is coming from the multi-head self-attention. The multi-head allows you to [01:19:08] weight different parts of the sentence, and allow you to understand the relationship between the various [01:19:13] parts, like, the subjects, the verbs, and various other complex relationships that might exist in any given sentence or document. [01:19:22] So… sweet. [01:19:24] And why do we need self-attention? Why not recurrent neural networks? Well, because self-attention has a lower [01:19:31] time complexity than, uh, recurrent neural networks. Recurrent neural networks are sequential. As we saw in the ELMO model, to process N, input tokens, we need n steps. At each step, you can only process a singular [01:19:41] open. And each of these tokens are processed in O of ND squared. This is the time complexity of a recurrent neural network. [01:19:49] But for self-attention, uh, we saw that we only need a singular operation. We feed the entire sequence into self-attention operator. [01:19:57] And it then does the, uh, it then computes the attention spores for each of these, uh, tokens. [01:20:03] Corresponding to each of the other tokens. Like, we have just visualized for it, but such attention scores also exist for each of the other tokens, right? Because… [01:20:11] can receive, uh, all attention from all of it, right? All of the tokens before and all of the tokens after it as well. [01:20:20] So, that is why there is no sequential dependency in self-attention, and the time complicity for every self-attention layer is O of N squared d. [01:20:29] But N is the number of tokens that we are supplying as an input, and D is the dimensionality of the module itself. Y n squared? Because, uh, we have… this Q times K transpose is O of N squared, right? This is n plus n equals. [01:20:42] And, uh, yeah, that is the idea behind, uh, the core idea behind self-attention. [01:20:49] And then coming to the model output itself, the output from BERT is a set of embedding vectors. Each input token [01:20:56] gets its corresponding embedding vector. [01:20:58] as the output, besides the CLS token is a special token during training, uh, it is trained… the model is trained in such a way so as to give a sentence embedding as a part of the CLS token vector output, right? [01:21:11] So, the entire sentence's meaning or context is held within this one token. [01:21:17] And every token itself gets its own, uh… [01:21:19] open a presentation or a contextualized token embedding as the output for… [01:21:24] from the BERT model. The, uh, the hidden size, or the output size, is basically 768. [01:21:30] So each of these vectors are 768 dimensions. So given N input tokens, you have an output of n cross 768 from both. [01:21:40] So, as we were just saying, the sale list open is a special token where, uh, it provides the context for the entire sentence. [01:21:49] Uh, the output emitting vectors, uh, give the contextualized token embeddings for each of the input tokens. [01:21:56] And that allows us to train a variety of models to perform various different downstream tasks, as we will now take a look at. [01:22:03] The embedding vector corresponding to CLS token, for example, can be used as an input to a classifier. The classifier can perform various tasks, such as sentiment analysis, where it classifies the sentence as positive, negative, or neutral. It can do spam detection, where it classifies spam or not spam, and various other downstream tasks. [01:22:21] And the individual token's output, uh, can then be used for various tasks, such as named entity recognition, based on the output, uh, embedding for each of the tokens. We can train a model to tell us whether or not [01:22:32] They're trying to classify to tell us whether, uh, this is a named entity or not. [01:22:36] So, various other such downstream tasks can be performed by using the per-open embedding vectors, or the sentence embedding vector. [01:22:46] So, let us look at the downstream tasks, like the architecture for a downstream task, for example. [01:22:52] Um, so if I were to give a sentence as an input and receive, uh, [01:22:56] the outputs, the open embedding outputs, and I wanted to do something like sentiment analysis or spam detection, I would only use the sentence embedding, which is the embedding corresponding to the CLS token. [01:23:05] I take this embedding, and… [01:23:18] You could also use it for, um… [01:23:21] defining a similarity matrix between two sentences, whether two sentences belong to the same topic or not, or you could see how close to sentences are together using cosine similarity scores and various other such [01:23:34] tasks can be performed by using this sentence embedding output from Word. [01:23:39] And, uh, for example, if I wanted to use the PERT token embeddings for [01:23:44] named and named entity recognition. What I would do is train a classifier, [01:23:49] to tell me, uh, whether this input embedding [01:23:51] is, uh, named entity or not. Named entity could be a person, could be an organization, could be a geographical location, could be [01:23:59] anyone else, or, like, other, for example. [01:24:03] So, this is an example of a named entity recognition based on the word-based model. [01:24:07] given an input Mr. Trump's tweets began just moments after a Fox News report by Mark Tobin, a former reporter for the network, about Protestant Minnesota and elsewhere. What are the named entities in this sentence? So, I take the sentence, I provide this as an input to BERT, [01:24:22] And, uh, use the per-token outputs from BERT, and train a classifier on it to tell me whether parts of this sentence are named recordings or not. [01:24:31] So, uh, the output from that classifier gives us the probability that an input, uh, embedding was a named entity, or what type of named entity it was. For example, Mr. [01:24:42] from these 3 tokens are a person. [01:24:46] Right? Um, a Fox News, this is an organization. [01:24:49] Mike Tobin is a person. Minnesota is a geographical location, and no other token refers to a name like me. So, it's all classified as other, other, other. [01:24:59] So, this is an example of, um… [01:25:02] name directory definition. Another downstream task could be, for example, question answering. So, given a reference text and a question, [01:25:09] Bird must point out which part of this text, um, is the answer to this question. So, for example, if I give a text, [01:25:17] Bert Large is really big. It has 24 layers, and an embedding size of 1024 for a total of 3140 million parameters. [01:25:22] Altogether, it is 1.34 GB, so I expect it to take a couple of minutes to download your Parab instance. So, how many parameters does BertLage have? [01:25:30] Is my question. So, how do I provide this as an input? I use the segment embeddings of BERT, for example. [01:25:36] So, I use CLS, and then I put my question separator, then I put my reference next. [01:25:41] Segment A embeddings are added to the question, and segment B's embeddings are added to the reference text. [01:25:48] And then, how do we, uh, how does Bert answer it? Well… [01:25:52] deeper token output emittings are, uh… [01:25:54] fed into two classifiers. The first classifier is the start span classifier, where it classifies whether a particular token [01:26:03] is the starting of an answer or not. So Word needs to highlight a spine of text containing the answer. This is represented as a predic… as a prediction of which token marks the start of the answer. [01:26:12] And for every token in the text, we pass its per-token embedding into the start [01:26:16] in the start answer sequence classifier, and whichever token has the highest probability of being started, the answer, we take that as the start of the answer, or start of the span of text containing the answer. [01:26:27] And the second classifier is the end token classifier, or end of text classifier. [01:26:32] So, what is the end of this particular answer? So, whichever token has the highest probability of being the end of a spine of text, we take that as the end of the text. So, in our question, this bird model highlights 340M. [01:26:46] Uh, between spans 27, uh, 27 to 13 as its answer, 3140 million. [01:26:52] So, this is how our question answering task can be achieved by using Word. [01:26:57] And various other downstream tasks can be done, but these are a few examples. [01:27:03] Any questions? [01:27:13] Um, so we can quickly take a look at a BERT model output and how it works. [01:27:19] For example, so here, I have given, tell me about Qualcomm as the input. We have already seen the input tokenization and how that works. [01:27:27] And then we have seen how the inputs are processed and combined with the word emittings, the position emittings, and the segment embeddings to give us input embeddings. These input embeddings are fed into the transformer block, so we have 12 blocks here. [01:27:41] or 12 layers in the BERT model that I'm using. So… [01:27:45] each of these blocks run in or not. The output from each of the transformer blocks is fed as input to the next transformer block. [01:27:52] Uh, that is what this is achieving. Each transformer block has a multi-head tension followed by residue validation and layer normalization operations, and then it has a feed-forward network. The feed-forward network, as I said, is… it's simply an upper projection with an activation and a down projection. [01:28:08] And, uh, then there's another normalization. [01:28:10] And this is done for the number of transform blocks we have, which is 12. And then this X, the output that we get, [01:28:16] That is the contextualized token embrace. [01:28:18] So this is the sequence output. And then we have the special CLS token state. [01:28:23] Uh, which is then pooled via a special pooler layer that is also part of the base model. [01:28:29] Uh, for our CLS, or our… [01:28:30] sentence embedding output, so pooled output and sequence output. [01:28:34] Sequence output is the point token embedding, pooled output is the embedding per… [01:28:39] like, CLS is the sequence embedding for a given sequence of tokens. [01:28:59] So, as we can see, tell me about Qualcomm. Was… this is the CLS token. This is a separator token. Since, uh, the entire sequence is just a singular sequence, sequence A, [01:29:10] That's why we use token type 0. [01:29:12] And we need to pay attention to all of these. Like, we are… there are no masked tokens here. This is the entirety of the sentence, has no mask tokens. That's why the attention mask is one. That is, all of the tokens need to be considered. [01:29:23] Uh, if you were masking this input to, let's say, 512, um, for a set output, you would then specify the mask for, uh, the attention for the mask tokens, [01:29:33] to be zero. But in our case, we don't have that, so that's why the tokens, or the attention mask for the tokens are all one. [01:29:41] Then, this is the output that we get from both the OR token output. So, there were 9 input tokens, we can see. [01:29:48] 3, 6, and 39. So, there is 9 outputs. Each output has an embedding size 768, so… [01:29:55] This is for CLS, this is for TEL, this is for me, and so on and so forth. [01:30:01] So this is the per-token output emitting. [01:30:05] And the CLS token output is passed through a special puller layer to get us the sentence embedding. [01:30:11] So this is the, uh, this is the context, or the sentence embedding. This… this corresponds to the meaning, or the contextual meaning of a sentence. [01:30:20] In a given vector space. So this can be used for various tasks such as spam detection, or sentiment analysis, [01:30:28] And so on and so forth. And these per-token emittings can be used for, uh, named and detailed information, uh, question answering, [01:30:34] And other switches. [01:30:37] So, this is about the base BERT model. [01:30:41] Um, we can then look at a few of the derived models, like Alberta, Robert, and DistilBot, and what were the issues, or what were the tasks that they were trying to [01:30:53] Um, actually help with. [01:31:08] So, let's take, for example, Albert Robotheim, distilbert as, uh, bird-based models. [01:31:15] So, let's take an example of Albert. [01:31:17] Albert stands for a light version of BERT. The BERT-based model has 12 layers, and the BERT large model has 24 layers, and [01:31:25] Each layer that we add falls back. [01:31:27] causes an exponential increase in the number of parameters of themodel. Like, [01:31:32] 12 to 24, it's just double the number of layers, but it's 3 times the number of parameters that we have. [01:31:37] So, to solve this, Albert uses, uh, paramedic sharing. [01:31:40] Uh, as the good idea. [01:31:43] So, in order to prevent this parameter explosion, because that exponentially increases the cost of inference and the cost of training, the moderate X as well. So, to solve this, Albert uses a cross-layer parameter sharing. [01:31:54] Instead of unique parameters for each layer, the parameters are learned for the first layer, and then they can be shared amongst [01:31:59] other layers. So, [01:32:02] Um, these… this sharing can be done either only for the feed-forward, or it can be only done for the self-attention, or it can be done for both. [01:32:10] the, uh, the paper here that I specified, uh, [01:32:13] explores all three of these, uh, parameter sharing methods. [01:32:17] And, uh, then it tries to benchmark what is the performance penalty, like, what is the accuracy loss that we face by sharing the parameters amongst various layers. [01:32:27] And the second core idea of Albert was to reduce the size of the embedding matrix. As we saw, the token embedding matrix. [01:32:32] or the, uh, the first layer of the model. [01:32:35] It is a huge insight, because for 30,522 tokens, each token has 768 dimensions. [01:32:42] The embedding for each token has 768 dimensions. So the memory required to store this 30,522 token embedents is quite huge. [01:32:51] Right? And this only releases as we increase the vocabulary size. So, to do that, uh, to reduce this memory, Albert [01:32:59] Oh, I came up with an idea of factorizing this entire matrix. So, instead of storing this 30,522 per 768 matrix directly in memory, it decomposed this into two smaller matrices. [01:33:09] 30,000 to 1, 522 cross 100, and 100 cross 768. [01:33:15] So, when we do a magma between these two factorized matrices, [01:33:19] that is when we get the actual embedding matrix out, right? So, I need not load the entire thing into memory, I just need to load this one into memory, and take the corresponding, um, [01:33:30] factorized token embedding from this one, and multiply it with this to get my, uh… [01:33:35] like, token embedding for Apple. For example, if I take Apple, [01:33:39] I just need to load this part into memory, and multiply it, uh, this… this is a matrix operation between the factorized, uh, [01:33:45] width and this vector. [01:33:48] And I get this vector as the output. [01:33:50] So, this significantly reduces the data size of the embedding methods. [01:33:57] So, for example, the data size for this, in full float 32 precision would be 90mg, while the data size to hold both of these in memory would only be 12MB, and we need not hold this in memory, right? We can only hold this one and load this as required, based on [01:34:11] which tokens we actually need. And this significantly reduced the amount of memory required. [01:34:16] for the embedding matrices. [01:34:21] And then, coming to parameter sharing, Albert, uh, the paper, right, in the paper, Lou et al. [01:34:26] Uh, the, uh, the paper where Albert was, uh, actually [01:34:32] Oh, Nick, or test it out, and propose. [01:34:35] So, in that, they benchmarked all the shared weights, which is shared attention, shared FFN, [01:34:41] And they benchmark only shared attention and only shared FFN, and no shared bids. So, if there were no shared weights, there were 108 million parameters. [01:34:49] And on the race benchmark, we got, uh, race is a reading comprehension, English reading comprehension benchmark. We got 68.2% with no share rates. [01:34:58] With shared FFN, it… [01:35:00] dropped significantly, but with shared attention, [01:35:03] like, it's only the attention bits are shared. We got 67.7%. [01:35:06] Which is quite close in performance. And, uh, this is 20% less parameters than [01:35:13] When none of the parameters are shared. [01:35:15] So, and another modification made by the Albert authors was [01:35:21] the ineffectiveness of, uh, [01:35:24] at most, or, like, NSP, uh, in the training process. Because, uh, mass language modeling and next sentence prediction as tasks are quite similar. They are training on almost the same task, and this was also realized by the authors of Roberta and Excelmik. [01:35:40] So, Alberta replaced, uh, so Albert, Alberta, and Roberta replaced this next sentence prediction task with sentence order protection instead. Sentence order prediction takes two consecutive parts of a document as a positive case and swaps them as a negative case. [01:35:53] So, let's look at this. For example, given this document X, sentence 1, Sentence 2, and a random document, uh, sentence 1, sentence 2, sentence, 3, sentence 4. If I take sentence 1 and sentence 2, [01:36:03] then they are next sentence, uh, prediction tasks, right? Like, is sentence 2? [01:36:08] succeeding sentence 1. Yes, it is an excellent prediction task, but does sentence 1 precede sentence 3 from this random document? Well, no. So, this is the next sentence protection pass. [01:36:18] While for sentence order protection, I just need to take sentence from sentence 2 and reverse their orders. So, if I take, I completed high school, then I joined undergrad, [01:36:25] This is the correct order, sentence order prediction. If I swap them out, then this is not the correct order, right? They joined, then I joined undergrad, I completed high school. This is not the correct sentence model. [01:36:35] So this was the task that, uh, Albert [01:36:38] Alberta and Roberta, or Trinkor, rather than, uh, Nickcenters prediction, which is quite close to mass language modeling. [01:36:49] And we can also look at proberta, which is robustly Optimized BERT approach. [01:36:54] Um, so what was the optimization? Well, they just changed the training process. There were 3 major changes that were made. [01:37:00] First one is that we have already looked at. Uh, next sentence prediction task was removed from the training process, and it was only trained on the mass language modeling task. [01:37:10] Because of the ineffectiveness of Nick Sanders' prediction as a training task. [01:37:12] Um, and then the second change that was made was direct masking. So, we have already looked at mass language modeling. We take 15% of the inputs and we mask them randomly. Well, um, if we mask them randomly, uh, it is still a static mask. [01:37:25] So, in order to, uh, like, because the same tokens have been masked for all of the epochs, in order to overcome this, [01:37:30] Word duplicated the data 10 times, and then masked it with different strategies, like, different random strategies, every single time, giving us various different, uh, masked data. [01:37:40] However, this caused a huge increase in the training time. [01:37:44] Robota introduced dynamic masking, so each time a particular data was fed into the model for training, [01:37:50] It was mass-driven differently. [01:37:53] So, robot I use dynamic masking, where a new masking pattern was generated each time a sequence was input to the model during the training process. [01:38:00] And the third change was in the… in the batch signs. So, bird was trained with a batch size of 256 for 1 million steps. [01:38:08] Robota instead used a dynamic strategy, where they used a sequence length of 2,000 sequences. [01:38:15] 2000 has a bad size for the sequences, and they used 31,000 stems to train on that. [01:38:23] So, we had, like, [01:38:24] Sorry, sorry. Basically, it trained for 125 steps on 2,000 sequences and 31,000 steps on 8,000 sequences. [01:38:33] instead of bird, which was just trained on 256 sequence length, uh, batch size, on… [01:38:37] For 1 million steps. So with larger batches, it gave greater perplexity for downstream tasks and MLM training objectives. [01:38:44] That is the reason that Roberta increases the batch sizes, rather than the training steps. [01:38:49] So, basically, increase their tech… the width of the training, rather than the depth of the training. [01:38:54] Um, and the other difference between Roberta and BERT was that Robota used a byte by encoding BPE tokenizer, rather than the WordPiece model that Bert used. [01:39:03] And this… this is basically, uh, for handling a different [01:39:08] like, a different vocabulary. So, [01:39:10] with, uh, Bert, what happened was that a lot of, uh, in a lot of real-life cases, uh, there were quite a few unknown tokens that were found when [01:39:20] tokenizing unknown, like, attacks that were not previously seen. With bytepair encoding, we get less of these unknown tokens, and it allows for a more efficient handling, effective handling of larger vocabularies. [01:39:31] for National Language Corpus. So, that is why Robota used BP rather than WP. That was another improvement that Robota made on the base BERT model. [01:39:40] And it used a lot more training data. Uh, it used 10 times the bird training data. It used the book purpose, the CC news, the open web text, and the stories dataset. [01:39:49] Right? So, it was a total of about, um, [01:39:53] something around 200 GB that it used. [01:39:56] to train on. [01:39:59] Any questions so far? [01:40:07] Okay. [01:40:09] So… [01:40:11] The next one that… the next word derivative that we look at is distilled bird. Distlebird solved a completely different, uh, [01:40:18] problem statement. So what Disputbird did was that it was pre-trained using knowledge distillation for faster inference. [01:40:24] So, this is, uh, aimed at edge devices, like phones, where, uh, [01:40:31] the on-device capabilities are quite limited. [01:40:34] Right? We cannot run an entire BERT model or an entire Roberta model. [01:40:38] Because we just don't have enough memory, we don't have, um, enough, uh, performance on the edge devices. [01:40:45] For it to load. So we need a smaller model that can run within, uh… [01:40:50] relatively fast, and it can give you a decent, uh… [01:40:54] user experience for any of the downstream tasks that it might be used for on edge devices. [01:40:59] So, for that, uh, for that purpose, knowledge distillation was used. So, that's why it's called distal board. Distilled here stands for knowledge distillation, but what is knowledge distillation? [01:41:07] It is a technique where the larger model acts as a teacher for a smaller model. The smaller model is trained, or it is trained to replicate the larger model's outputs on a given set of inputs. So, [01:41:20] This is called knowledge distillation, where it basically needs to mimic the teacher model from, uh, [01:41:25] like, the teacher model has actually learned all of the, uh, uh, all of the weights. [01:41:30] And, uh, we are actually trying to replicate a smaller set of weights, which can mimic as closely as possible the teacher model's intermediate activations and the model outputs. [01:41:41] So, it used a triple loss function. The, uh, basically, there is a language modeling loss, which is the mass language modeling loss using the next word prediction performance. Like, this is similar to what was used in BERT, the mass language modeling task of BERT is basically the language modeling laws in this still book. [01:41:56] The distillation loss is the distance or the loss between the distant model output, or the student model's output, [01:42:03] to the actual BERT model's outputs. So, what is… like, we need to mimic the BERT models output using the smaller distal BERT model. That is why this distillation loss is very important. [01:42:12] And the third loss is the cosine similarity loss, which is calculated between the sublayer activations. So for every layer output, like, [01:42:19] from the attention weight output, the SSN output, each of these layers outputs must closely align with the base, uh, BERT models outputs. [01:42:29] So, in order to achieve this, the cosine distance loss was used. [01:42:31] The effective loss was the sum of all of these three losses. [01:42:35] So after training, this little bird was able to mimic 97% of birds' accuracy. [01:42:41] with just 40%, uh, which is 60% of the parameters of board, which is 40% lesser. [01:42:46] parameters invert. And all of this allowed it to achieve 60% faster inference time than Word. So, for example, if BERT was taking 1 second, [01:42:54] It took just 0.4 seconds, 60% faster than work. [01:42:59] So, this is, uh, like, a bird's-eye view of how Disnetwork was made. [01:43:03] So you… you get the mass language modeling laws on the actual text corpus. You get the cosine similarity distance loss between each of these sublayer activations. So, these sublayer activations to sublayer activations from booster bird. [01:43:16] And, uh, then we have the displacement law of itself, which is the distance between the output from BERT and the output from distal input. And this is… this combined loss function is then used for back-propagating the errors on this text purpose for, uh, to train the distal board. [01:43:31] to mimic the outputs from both. [01:43:32] So, if we look at the number of parameters, Nesselbert had just 66 million parameters, whereas bird-based had 110 million, and Elmo had 180 million to have the same set of, uh, like, [01:43:44] similar performance on the blue dataset, on… this is the semantic textual similarity task of the, uh, of the general language understanding dataset. [01:43:53] So, to achieve this similar output, like, similar accuracy or similar performance on, uh, [01:43:59] these, uh, on this dataset, uh, DistilBird requires 66 million parameters, bird-based required 110, whereas Elmore required 180. And then, if we look at the inference time, [01:44:09] in seconds. Elmo requires 895 seconds to perform this task, whereas bird-based requires 668, while listed BERT requires 410. [01:44:17] So, it's much faster than bird-based. [01:44:20] And it… it has an accuracy which is 97% of that of birds' accuracy. So it's quite… [01:44:26] quite close, and much faster, allowing, uh, for various [01:44:30] edge use cases as well. [01:44:33] So, now let us look at these benchmarks. We have talked about benchmarks such as race and glue. [01:44:38] But what do these actually signify? What are these evaluation metrics? How do we decide whether a model is good in language understanding or not? [01:44:47] And, uh, how do we decide whether the model is good, uh, for [01:44:50] matching language tasks, downstream tasks such as semantic textual similarity, or, uh… [01:44:56] sentiment analysis or spam detection. Well, to quantify these model performances, we have these standard benchmarks. [01:45:03] So, one of these is Blue, the general language understanding evaluation. Another is SPORT, which is Stanford Question Answering Dataset. [01:45:09] Reading Comprehension in English. [01:45:13] callner, which is, uh, Conference on Natural Language Learning, named entity recognition, uh, task. [01:45:18] So, what is the glue benchmark? The glue benchmark has 9 datasets per different tasks. [01:45:23] One of it is, uh, natural language inference, another one is question-answer pairs, another one is, uh… [01:45:29] question natural language inference, and this is semantic textual similarity, which is how close or how similar to [01:45:36] two sentences are. Another one is paraphrasing corpus, where the model is expected to paraphrase the inputs. [01:45:42] Uh, after, like, given an input sequence, it must be able to paraphrase the next sequence. [01:45:47] of that. And, uh, then the recognizing textual entailment, the Winograd natural language inference, [01:45:55] And then there is a sentence classification task as well, which is a corpus of linguistic acceptability and Stanford Sentimentary power. [01:46:00] So, these are the 9, uh… [01:46:03] tasks that the Glue Benchmark includes. [01:46:05] To evaluate any model of the Blue Benchmark, it is, uh, first trained or fine-tuned over the Glue training dataset. [01:46:12] And then it is stored across the validation dataset. The final performance, or the final score, of the blue benchmark is the average of all of these 9 tasks. [01:46:20] So, for the bird base and the bird large model, the spores in percentages on each of these tasks are listed here. [01:46:27] And the average of the word base is 79.6, and bird large is 81.9. So, [01:46:33] this, uh, was the blue benchmark course that the BERT model achieved. These were state-of-the-art when, uh, BERT model was released, and, uh, these were the best performances, and it outperformed Elmo by quite a bit. [01:46:46] And then the second one is Stanford question answering dataset. [01:46:49] Uh, this contains, uh, one black question-answer pairs. [01:46:52] And, uh, the… the answer to each of these questions is basically a span of the question text itself, like we saw. Bert was asked to highlight the span, which was the answer to a question in a given… [01:47:05] in a given paragraph. So this is basically a comprehension, or a reading comprehension, uh, dataset, where you try to answer a question based on a given text. [01:47:16] So, all of these question answers are taken from Wikipedia articles, like, one lakh question answer pairs from Wikipedia articles. [01:47:22] Uh, bird-based and bird large scored 80.8 exact matches, and 84.1 exact matches, uh, percentage, again. [01:47:29] And the F1 score was 88.5 for bird pace, and 90.9 for bird large. [01:47:34] Though… [01:47:36] The next benchmark that we look at is race. This is a reading comprehension in English. This dataset is human-prepared. It consists of 28,000 passages and 1 lakh questions in each of these passages. [01:47:49] Uh, this was a dataset that was collected from English examinations for Chinese students, English proficiency examinations for Chinese students. So, this is reading comprehension in English. [01:47:57] Uh, that is the name of this benchmark. So, again, BERT was given the, uh, the questions, or the passages, and the questions to be answered by highlighting spans from those passages. [01:48:08] 28,000 passengers and 1 lakh spans. [01:48:11] So, the BERT large model achieved a score of 73.8% correct answers on race benchmark. [01:48:18] And the last one that we look at is the named entities recognition. As we have seen an example of previously. [01:48:24] So, this is an input sequence, and BERT model is expected to highlight the, uh, [01:48:31] the… the tokens that correspond to named entities, like Jim Henn is a person, [01:48:38] So, this is a named entity. Sun was a puppet who was… is not an entity, right? So, we should not be, uh, highlighting any of these as name-netes. [01:48:47] So, for this benchmark, the callner dataset, which is Conference on Natural Language Learnings, named Retail Recognition Dataset, is used. [01:48:55] It consists of about 2 lakh training words, and this has been annotated as person, organization, location, and miscellaneous. [01:49:03] Uh, which is not named entity. So, person, organization, and location are named entities, and miscellaneous or other are not named entities. [01:49:10] So, bird-based model achieved a F1 score of 96.4% here in BERT large model achieved an input score of 96.6% here. [01:49:20] So, this was about the benchmarks. [01:49:24] Oh, now we have the question… question and answer session. If, uh, do you have any questions? Any… [01:49:30] thing that I, uh, should go through again, or, uh, explain in a bit more detail. [01:49:38] Hello. [01:49:39] No. Yeah. [01:49:40] Uh, for that NEAR model, uh… [01:49:43] If the person… with the person name there is a degradable, for example, Loretz. [01:49:48] Dr. Jim Ham will need that. [01:49:54] Dr. B, also enable the clip? [01:49:57] Yeah, yeah, it would be. So, in the example that I had shown, so this is Mr. Trump, so all of these is person. [01:50:09] that doctor will also be exported as an email, right? [01:50:14] Yeah, it's the salutation along with the first name, middle name, last name, anything, would be your personal name negatively. [01:50:24] looking at Pfizer. [01:50:30] Any other questions, anyone? [01:50:39] Could you please explain, like, how is distilled bird different from bird again? [01:50:45] Yeah, sure. So, distal bird is trained using a technique called knowledge distillation. [01:50:50] Uh, in Norus legislation, what we want, uh, to do is we want to mimic a larger model. [01:50:55] output, uh, using a smaller model. Why this is done is, uh, because for… as I, for devices with a lower memory footprint, uh, a lower, um, [01:51:05] computational capacity, we really need, um… [01:51:09] smaller models that can run faster. So, [01:51:12] for example, on your phone, you would not want to wait a couple of seconds for a particular thing to run with, because that's not good user experience. So for such edge devices, [01:51:20] Uh, this knowledge distillation is a technique that's used to compress models, larger models, into smaller ones. So, the larger model acts as a teacher, and the smaller model is the student that needs to mimic the teacher's outputs using a smaller number of parameters. [01:51:34] So, this little bird is faint using this, uh, very same technique. It uses a triple loss function, where one of the losses is the mass language modeling loss, which is the same as the BERT-based model. [01:51:43] The second one is a distillation loss, where the distillation loss is calculated, uh, [01:51:49] based on the student model's output and the teacher model's output. Uh, the loss between the student model and the teacher model output is called… or basically the distal burden-bird model. So, when the distal birds output and the BERT models output are not close enough, then there is a large distillation loss. [01:52:05] When they are quite close, there's a lower distillation loss. And then the third one is the cosine distance loss. So, apart from the final output, we also want the intermediate tensors, or the intermediate sublayer activations, as we call them. [01:52:18] to be close as well, because if they are close, then the final output is quite close. So… [01:52:22] But in order to ensure that these sublayer activations are close enough, the third loss function, which is the cosine distance loss, is used. [01:52:29] The, uh, model is trained using a loss that is computed as a weighted sum of all of these three losses. [01:52:35] The language modeling laws, the translation laws, and the cosine distances. [01:52:39] Uh, using all of this, it is able to mimic [01:52:42] the BERT model's output using just about 60… 66 million parameters, rather than 110 million parameters, that the bird-based model uses. So that's quite a bit smaller this. [01:52:53] 40% smaller than bird bits, while still achieving 97% of the accuracy that we're achieving. [01:53:04] Uh, there's a question by Sachin on the chat, uh, Mr. Somni, if you could address that. [01:53:09] Yeah. [01:53:11] Um, let's see… [01:53:13] Can you please explain on how to improve performance of inference on edge devices? [01:53:18] Okay, uh, so performance improvement on edge devices depends on a variety of factors. [01:53:26] The first one is, uh, what is the amount of memory that is available to you? What is the performance of the flops that you have available on the device? [01:53:32] And the third one is, uh, what is a cap… like, you need to define a KPI. What is acceptable? How… how long? [01:53:38] Uh, does it need to wait in order to get? And also, KPI in terms of accuracy, like, how much accuracy loss is mostly? Because [01:53:44] Whenever we try to go faster, or make a model smaller, uh, we do encounter accuracy losses. Like, even in this little boat, we had approximately 3% loss rate, so… [01:53:55] you need to come up with KPIs for all of that. Once all of these KPIs are in place, there are a variety of techniques available. The first one is, uh, we try to make the model smaller. [01:54:03] Because most of these are transformer-related… any of the LMOs, transformer encoders, transformer decoders, encoder-decoder models, all of them are heavily memory-bandwidth limited. The CPU spends a lot of stall cycles. [01:54:14] It's not able to get enough data, a CPU, GPU, even NPU, anything that you might have. They are not able to get data, enough data fast enough in order to run, uh, in order to make it a compute-bound. [01:54:25] model. It's mostly memory-bound. [01:54:28] And in order to reduce this memory-bound nature of the model, you can use techniques such as quantization. [01:54:34] So, quantization is basically when you, uh, try to [01:54:39] mimic floating point numbers using integers. So, given, let's say, I have a weight that has a range of minus 0.1 to plus 0.1, I can represent these as integers instead in the range of minus 128 to plus 128, where, uh, minus 0.1 is minus 128, 0 is 0, and uh… [01:54:56] 0.1 is plus 128. Now, uh, these floating… each of these floating point numbers are 4 bytes, whereas the integer is just 1 byte. So, I have reducedthe model size by 1 fourth, by just quantizing these, uh, weights. [01:55:08] The next one is, uh, the infrareds itself, like, the, uh, the hardware itself is able to perform integer inference much faster. [01:55:16] Then, uh, the floating-point inference. So the ANU for integer units, or the integer ALU units, are much faster than the, um, [01:55:24] The floating point daily events. And lastly, it comes down to the hardware, uh, which is basically, you can reorder the data, uh, for your given hardware. For example, if I have a CPU that has the neon [01:55:36] dot product instruction available to me. I can use that, or if I have, for example, the name an IAT demo, which is an optional V8.7 extension, if I have that available to me, I can use… I can basically reorder my weights into submitted of 2 plus 8 in order to make use of that instruction. [01:55:51] Because that uses 2 plus 8 plus 2 submatrices, and gets a 2 plus 2 output, whereas the dot product will use, uh, 1 plus X and X plus 1 to be 1 plus 1 output. So, I am able to achieve, uh, 4 partial sums. [01:56:04] Instead of just one partition in the same number of cycles using that. [01:56:08] So, putting this all together, [01:56:11] is… is the way that you make inference faster, or any of the edge devices. [01:56:22] Yeah. [01:56:23] Okay, it is basically based on the requirement, what is the hardware limitation which we have on the edge. So, based on, like, okay. Thank you. [01:56:24] Yeah, yeah. [01:56:30] Any other questions, anyone? [01:56:41] Any questions? [01:56:47] Okay, if there are no further questions, then we may wrap up for today. [01:56:52] And we'll thank Mr. Sommei for his, uh… [01:56:56] insightful and interesting session. [01:56:58] And we'll continue with the other transformer models. [01:57:03] Uh, in the forthcoming sessions. [01:57:05] Okay, thank you all. Have a great day. Thank you. Bye-bye. [01:57:11] Thank you. [01:57:12] Thanks, Mr. Somani. You can feel free to leave the session. [01:57:15] Thank you.