# 09 2026-06-07 Introduction to RAG course: Module 4 — Generative AI & LLMs module: Module-4-Generative-AI-LLMs date: 2026-06-07 type: transcript video_url: https://personal-learn.armco.dev/files/_Recordings/Module-4-Generative-AI-LLMs/09_2026-06-07_Introduction_to_RAG.mp4 --- [00:21:13] Rati, are you there in the session? [00:22:12] Good morning, everyone. [00:22:22] I will stop. [00:22:52] Uh, hello everyone. I think, uh… [00:22:55] Uh, so I think, uh, is just joining the session. Let me just check. Pratik, are you there already? [00:23:01] Yes, ma'am. [00:23:02] Yeah, I'll request you to please enable your video. [00:23:11] Yeah. So, welcome, Sratif. Uh, we'll, uh… so, we'll, uh, be choosing a hands-on session on T5 today. [00:23:19] Uh, just a second, Katik, before we start, is everyone ready with the files? [00:23:26] Yeah, you can share your screen, no issue on that. [00:23:30] Is everyone ready with the files on your systems, beast? [00:23:34] I'll request the co-host to please share a link. [00:23:39] on the chart for the files, for those who are not yet ready. [00:23:47] Hugh Channel's co-host, are you there? [00:23:54] Insurance, can you please come in? [00:23:55] Yes, ma'am, I'm sharing it [00:23:56] Yeah, okay. [00:24:00] So, shall we begin? All of you, if you all are ready? [00:24:10] Yeah, I think people are ready. [00:24:12] So, then, over to you, Pratik. You can please start. We'll be discussing about the T5 model today. [00:24:19] And, uh, before that, uh… [00:24:21] So, with this, we will be completing with one logical portion, that is on text LLMs. [00:24:28] So, Pratik, over to you to take out one. [00:24:32] Okay, ma'am. [00:24:33] To take up the hands-on. [00:24:41] We did a brief session on GPT model, right? It was based on the task of [00:24:46] Question answering [00:24:48] to present you guys with something of a new task today. Like today we'll be dealing with mostly with the summarization task. And for that task we have chosen this Samsung [00:25:02] data set, right? So… [00:25:06] Like, before we load anything, let me just give you a quick overview, like, what this dataset is about. Samsung is a collection of 16,000 sought conversations that look like text message or messenger channels, right? Like multiple people are talking in between [00:25:22] Right? You can see, I'll give you some examples here. And from there, like, along with it, we have a summary of, like, how the dialogue went [00:25:31] So each conversation comes paired with a short summary written by a person. Our task for the whole tutorial is simple. [00:25:37] To state, given the conversation, produce a summary. [00:25:41] close to that human one, which the [00:25:44] like the humans have written [00:25:48] Now, a single example of like a data set would be having ID [00:25:54] dialogue and summary. ID is a unique identifier for the convention. We do not use it for modeling, it's just available. And dialogue is the conversation itself, right, written between like [00:26:06] Like, let's say something, suppose Ram and Sita were talking, so it's a conversation between them, and a summary is, like, sorts like [00:26:16] description of what the conversation was about, right? So, in this, like, summary will be the target [00:26:24] Okay? Now, [00:26:25] To give you a brief overview, like, how it was [00:26:29] like, the origin of Samsung dataset. Samsung dataset was basically created by research group at Samsung, released alongside 2019 academic paper, right? And, this, just to give you a brief, the computations were not scraped from the real private chats. Instead, linguists who are fluent in English wrote them by hand [00:26:48] And model kinds of messages they will send and every day, okay? Now, how it is structured, the conversation are spread fairly across four groups, right? Based on how many back and forth tons they contain, short ones with three [00:27:03] Right? To six tons, and then 7 to 12 tons, and then 30 to 80 tons. And the longer ones can be up to 19 to 30 tons, right? Most conversations, around 3 quarters of them [00:27:13] are just two people, right? While rest of them have 3 or more, okay? Now, the train data set, [00:27:20] Like the data set is going to speak into three parts this time, right? Unlike yesterday. This time we have a trained data set of trained training set of around 14,700 samples. Validation set of 800 examples around [00:27:36] And taste of around, again, 800, right? Training will be used to, like, [00:27:42] To help the model learn, right? Validation will be used just to check if the loss is decreasing or not, or… and to choose the best, right? And for testing, it will be used to measure like how well the bottle really does. Okay. [00:27:58] So… [00:27:59] We chose this dataset particularly because it is modestly clean, right? And it is very intuitive to understand and easy to work. Okay [00:28:14] Now, one honest note on usage, just be careful. Like, Samsung released for research and learning under [00:28:20] non-commercial license, okay? This was just to highlight that. Okay, so it's good for a tutorial, but it's good to know this. [00:28:28] Okay, so let me… your outputs [00:28:33] With the install session [00:28:46] So, yeah. Like always, here we are choosing the device that you will be working on. Since we are working on Colab, and we are… most of us are in future, we'll be choosing CPU, we can CPU. [00:28:57] Right? That's it. Then, in the next part, we are… we will download all the required libraries for this one. So, we are downloading transformers, data sets, evaluate, Rogue score, BERT score, sentence phase, acceler [00:29:13] 57ZR, okay? So, now what of these… each of these packages? Transformers are, like, basically for the hugging Face library, that gives us 35 model and tokenizer. Data sets is basically for one line loading of the Samsung data set from the Hugging Face Hub. Evaluate [00:29:28] will be used for the modern wrapper that runs metrics, right? Here we are using Rogue and all, okay? The rogues for the actual rogue implementation that you will get calls under the hood, and BERT score, our secondary metric, [00:29:39] which is used to see the semantic similarity metric right then sentence piece. It is required by the T5 tokenizer [00:29:48] And, like, just to give you a heads up, T5 fails to load without it, okay? Accelerated is required by the hardware trainer. We use it block three and pi 7za is Samsung like it is used for Samsung chips or like basically dot 7 [00:30:03] Is it archived, right? This lets data sets unpack it, okay? [00:30:08] So, let me quickly run this [00:30:19] Okay. Now, like the fourth, the imports… import part, like, we'll be importing torch. Torch, as we all know by now, is, like, it's for core deep learning library, powers the model, and for training. And from the data sets we are importing load data set [00:30:37] This one will be primarily used to call, this pull this Samsung dataset from Hugging Face Hub, right? Then, we import auto-tokenizer, auto model for sequence-to-sequence LM [00:30:49] from transformers right and [00:30:53] Then we import, evaluate [00:30:55] Now, this auto tokenizer lecturn stakes into token IDs. The model understands and auto model for sec to sequence to sequence alarm. This lowers the T5 in its text in text outs form [00:31:08] Now, just, like, [00:31:11] Heads up, yesterday we were using automotive causal LM. That was, because GPT was a, like, [00:31:19] A decoder model, right? So, for that, we had that, and its major task was next word prediction. So, for those kind of things, like, we used, auto model casual LM. Now, today, since we are dealing with sequence-to-sequence, right, we are expecting summary, and, [00:31:36] This T5 has an encoder part as well as a decoder part, so therefore that we are using sequence to sequence. Just to give you a heads up, it is from yesterday's tutorial [00:31:47] Right [00:31:49] Okay, now [00:31:50] This is, like, just a verification on which device we are, right? And to report the tox version [00:31:57] Now, let me quickly run this [00:32:13] Okay. [00:32:14] Now, yesterday, like, also we had used this node dataset. Like, I wanted to just add this because I wanted to show, like, how it is different, different than the, like, just pulling the local CSV, right? [00:32:29] Now, when we were using this PT and dot read under CV and we were putting the file name of the path where the fileis, right? Where like in that basically in [00:32:40] Like, our data was on hard disk, right? Nothing was downloaded. If we move to another computer, we had to bring that file with you. Now, when we are using this Hugging Face app, the data lives online. The first time you run the load, and, like, it is downloaded and cached on your [00:32:55] So later runs are fast, and you do not need to re-download, right? And you don't have to carry it, even if you change the machine. You just, like, if you just have this code, you, like, press enter, then it will be downloaded and cached in your computer [00:33:10] Okay? [00:33:11] Now, pd.read underscore CV would give you a pandas data frame, and now this load underscore dataset gives you a dataset dictionary, dataset date, right? Which, is a container holding multiple splits at once like just as we saw in the example above when we were discussing about the… when we are giving the introduction to the dataset [00:33:31] Right [00:33:32] It has training set, validation set, and the testing set. [00:33:36] In our case, it already contains this right with a local Csv. If we had a whole data set, we had to split it mostly right? If it is not split it already. [00:33:47] Now, we use this approach for the reproducibility and like it is easier. We don't have to carry on the data set, right? So no manual file handling is there [00:33:58] Welcome. [00:33:59] Now, this is how we use the load dataset, right? [00:34:03] From the load data set we have already imported it, and this is where the, how we can call the file name, and then it will be downloaded and by reference by dataset. [00:34:15] Okay [00:34:16] So, now if we… suppose if it was not splitted, right? So when we call it, it will be called for the whole data set, right? If you just needed the one split, what we can do is we can do data set same as before, and then we can just call this [00:34:32] trading split. Okay [00:34:35] This is what the difference is, and [00:34:39] Okay, I think now we are good to go. [00:34:43] So, like, here we'll be loading the dataset, and as discussed earlier, like, we are calling it from, like [00:34:50] K and Karthik, and slash Samsung, okay? Here, like, we have it on Hugging Face, and sometimes we are referencing it through dataset, okay? [00:35:00] Now here [00:35:02] Since we are on a collab, right, we have time constraint, so we cannot run everything on CPU, right? So, I have kept this as a toggle. Like, if you want to use the full data set, you can set it as false, and then you don't have to worry about all of this [00:35:18] Right? And then, whole dataset will be refused. But if you since we are running on this [00:35:24] CPU. I… we have trouble with that too. [00:35:28] So that we can do a quick CPU pass [00:35:31] For this, number of training samples we have taken 60. Number of valid samples is, like, for validation, we have taken 10, and from number of test samples, we have taken 20. [00:35:41] Okay? Now, for the IDA purpose, I wanted to show it to you, like, on the real dataset, rather than just showing you… showing it to you on 60 samples, right? Let's have a look. And once we are done with the IDA part [00:35:58] Like, we'll come… quickly come back and toggle this off again, right? And then, work on this. [00:36:05] So… [00:36:10] I'm setting this as false for now. I mean, we set this as false, the full data set will be used for now, right? [00:36:18] Okay. And, like, this is to take the split, right? So, whichever split, we want to select, how many of them if we select. So if this were into coming to picture, then this would be there, this should be used, okay? [00:36:32] Now, data sec, in DataSec, we have trained validation and test speed, right? From there, we are selecting and train, endvalid, and end test. These are all the numbers we have defined here [00:36:43] Okay. [00:36:45] Now, we'll print the dataset and just quickly look at, 5 examples of how a dataset is right [00:36:55] 5 years, we'll just quickly look at it [00:37:04] So [00:37:06] Here is it [00:37:09] It's okay. So, data set, like, as I told you, it is in this form, right? Then it has training set, it has validation set, it has test set, okay? The feature is across all of this set is same, ID, dialog, and summary. The number of rows [00:37:24] In this is 14,731. Validation is 818, and again, test is 819. [00:37:32] Okay [00:37:33] So in order to give you an example of how it is there [00:37:39] In each of the road, let's see. [00:37:42] So first of all, we have dialogue, right? [00:37:50] The input we feed to the model, right? This will be the dialogue. This is the summary, right? That we want to produce, okay, ideally. [00:38:00] So, let's see the first conversation. It's between Amanda and Jerry. Amanda says, like, I… because do you want some? Jerry says, sure. And Amanda says, I'll bring you tomorrow. [00:38:09] And the summary that the, like, the human has written already is like Amanda baked cookies and will bring Jerry some tomorrow. So it's just a summary of this between conversation between the two people, okay? We'll look at one more, and then we'll proceed. So like this is second dial [00:38:26] I think here, only sorry here, Olivia and Oliver is here, right? Olivia says, who are you voting for in this election? Oliver says liberals, as always. Then Olivia, me too, and Oliver says gay, right? The summary is of this conversation, Oliver [00:38:42] Sorry, Olivia and Oliver are voting for labor in this election. [00:38:46] Now, this is between team and team, and from this conversation, again, we had team may try to promote by team to get most of them, okay? So this is how the examples are not go to each or not those because it's just from there. So I'll just show it to you, just have a quick look [00:39:04] And then we proceed. [00:39:06] So here again, we have this advertising season love with Bella. Congratulations once airborne to open his door. Rachel is outside. So this is the conversation between [00:39:16] Like, Edward and Richard [00:39:19] And this is the conversation between Simon. [00:39:24] Right? And then we have a summary right here [00:39:26] Now you see this, dialog is quite long, right? It has many terms [00:39:32] turns by… by turns, I mean that, like, how many times it is going back and forth between the two people [00:39:41] Now [00:39:43] Now, I'll quickly move to IRA and processing. Here, we'll try to see how long the dialogues are, and the summaries are, in terms of tokens, okay? [00:39:53] Not, not the words like we were doing yesterday. Okay? So [00:39:58] For this, what we are doing is [00:40:02] We import NumPy as NP, right? Then we import auto tokenizer because we'll be dealing with tokens. We import that from transformers right model we are using is T5 soul. I'll give you a brief about this model [00:40:16] Below, like, when we start for, modeling, right? [00:40:20] For now, like, just understand that this is 60 million parameter T5. Okay [00:40:26] Now, tokenizer, we are calling Autotokenizer dot from pretend model name is referencing from your T5 small, and then it will be reference to tokenizer okay [00:40:36] Then, we try to pull up the, tokens for [00:40:41] dialogue and summary in the training spread. So, for that, what we have done is, for each of the rows in the dataset train, and from there, we are taking the dialogs [00:40:53] visa. We are passing it to the tokenizer, right? And from there we are getting its token IDs, right? And then getting the length for it. So that dialog length is a story here, okay, for each of the road and [00:41:06] Now, similarly, we do this for the summary row. Sorry, summary feature, right [00:41:13] That is, again, being pulled from the dataset train. We pass it to the tokenizer, right? And from there, we get the input ids, and then we take the length of it, okay? Then that's how we have a summary broken ones. And this is stored in formal list [00:41:29] Okay. Now we'll try to print this all, and try to have some statistical summary, right? Mainly consisting of mean, maximum, 95th percentile, 99th percentile [00:41:42] Okay? [00:41:43] And, like, we'll try to… this is the describe function that we created for this one, and we'll try to do this for dialogue and summary. So we are calling this [00:41:54] Okay. [00:41:57] Now, we'll try to oversee our reports. [00:42:17] So, now if you see in the dialog, we have approx, like, 149 tokens. Maximum tokens we have in dialogue is, 1153 tokens, 95, 95th percentile is somewhere around [00:42:32] And the 567, sorry, 367 tokens, and 99th percentile is around 512 tokens, right? For summary, the mean token is around 29 tokens, right? At maximum, we have is 94 tokens [00:42:47] 95th percentile would cover around 59 tokens, right? And 99th percentile [00:42:52] is covering around 73 tokens [00:42:55] Right? So, if you look at the distribution also, this will also justify it. So if you see [00:43:03] Right, [00:43:06] Somewhere around here, we have peak, right? For 100 parts, right? And here we have around, like, 20s we have the peak [00:43:17] Okay, it is reflected here also, closer to that [00:43:21] Okay. [00:43:22] So this is for dialogue length, and this is for the summer length. Same thing we have depicted in our histogram [00:43:31] Now, just to give you a brief summary, dialogues are far longer than summaries, about 149 tokens versus 29 on average. The summary is roughly the fifth of the dialogue, another sign of how much the model must compress, right? So, around, like, one-fifth of the dialogue is the summary. Like, we get this [00:43:49] It's a rough estimate, but below, we'll try to see the ratio also, how much it will serve you. So most dialogues are short, but few are very long, like mean… having mean 149 and maximum tokens of 1153. The percentiles matter more than mean because mean hides the square [00:44:06] We don't have the actual idea, like, how it is happening. So the percentile sets our length index, right? For dialogues [00:44:15] 95 percentile falls under 365 token. So, cap around 512 keeps all of them whole. For summaries, 95 percentile, like, falls under 59 tokens. So, like, we have, cover, like, a cap of around… a little bit more than that we have taken. We have taken 64 [00:44:33] Okay, for… to keep [00:44:35] So that it can cover the large majority, right, and trim the longest view, which gives the generation fast, okay? [00:44:43] Now, the caps are like, it's a grade, basically, if you want, you can keep the limit higher, right? But, like, it's a trade. So… [00:44:52] If, like, if the caps are taken very strictly, and it is pulled lower, then [00:44:58] The generation may not be so good because, like, suppose it may end before the actual summary was even generated, right? And if it is very large, that means we are simply wasting memory, right? It may not also need that much [00:45:14] So, we try to make, like, [00:45:17] We have some trade, so, between these two, right? We want to find a sweet spot, like, okay? So, these lands are tokens and not words, which is on purpose. The model works in tokens and later limits that we had taken here [00:45:34] will be, like, taken into consideration. [00:45:38] Okay? Now, just to have a brief, like, idea of how the… how much the, like, model is actually, syncing the dialogue, right, to have a summary [00:45:50] So, for that, we are importing NumPy. We are importing Matcotlib.pig.spreid, we'll be reforcing that, okay? Then [00:45:59] This dialogue lens and summary lens we are getting from the previous sale itself, right, that we took it out. [00:46:07] Okay, let me start to you, yeah, from there, right? It got reference from there, we are taking from there, and we are at this time converting into NumPyRL. [00:46:16] So that after this, once this is converted, we try to find the compression ratio, okay? That is summary lens by, lens by, dialogue lens, okay? So, this is one ratio, for example, we'll have it, okay? [00:46:32] Then we'll try to report the key numbers, mean, median, and we'll try to have an interpretation, and see in graph also how it looks [00:46:43] So, if you see mean compression ratio is somewhere around 0.26 and median compression ratio is 0.2 right? So typical summary is about 22%. The length of its dialogue. So like if a dialogue is say 100%, then 100 [00:47:00] Then it summaries of around 22% of it [00:47:04] Okay? So, just to give you a brief summary, like, a typical summary is about 22% the length of his dialog. Roughly four-fifths of the text is dropped, and 1 fifth of the text is skip, right? [00:47:18] Now we can, say this with confidence also. Earlier, we were estimating it, like, we are having a hunch, like, this is happening, now we can tell it for sure, okay? The mean is 0.266 above the median is pulled up by the sort dialogs [00:47:34] That I'm little, to compress, right? And the meeting is a safer number to test here [00:47:39] This sets an expectation. Good summary is solved. If the model later produces something nearly as long as the dialogue, it is [00:47:47] It is copying and not summarizing. So our task here is summarization, not to copy itself. Okay, so compression. We want compression, not paraphrase. Paraphrasing keeps the length and swaps words. Summarizing keeps only the point. A ratio near one [00:48:03] Fifth source, which one we are trailing for, okay? So this, all this we did, we got, we inferred that from there. Okay [00:48:13] So, now we'll do some quickly some data checks, right? If any of the [00:48:18] Data is empty, or, like, how many duplicates are there, right? So, for that, what we are doing is [00:48:25] We are… this is a helper, like, there is a bad text, we are creating this function. If the value is none, then it will be returned true, and if the value strip is, like, like, it's a blank, right? Then again, we return true, else we'll [00:48:39] Return it as false, right? So now to check the missing values. [00:48:46] So, for, for, the… from this, like, for each plate name in dataset, right, we take each split, right, and basically those are what? Training, test, and validation, right? Now we'll call this function [00:49:01] this [00:49:03] It's bad text for each of the X meaning row in the, like, dialogue, right? Okay, similarly, we do that for summary also, to have… to see how many bad summaries are there, and then we try to filter [00:49:18] Okay? Then we try to check the, duplicate checks. Like, how many of them are duplicates, right, from the total dialog, how many uniques are there [00:49:28] Okay [00:49:29] And, next we'll try to see [00:49:32] Like, from here, we are… what we are doing is, we are taking a set [00:49:38] training set, right? And taking a set of it, of the dialogues, so that all the… when we do set, it will have all the unique dialogues, right? And similarly, from test also, we are doing a set so that we have unique dialogues. And to see the leak [00:49:52] leak part, right, that is present in both training set and test set. Like, what we are doing is, we'll do this and preparation so that we can know, like, how many of the, common [00:50:04] Let's have a quick look. [00:50:08] So [00:50:11] Like, these are the outputs we got. Duplicate check, so basically there are bad dialogs are 0, bad summaries are zero, like, in all training validation and test set. In duplicate check, training split [00:50:23] Total dialogues were 14,731, right? And unique dialogues were 14,255. Duplicates were around, 476 [00:50:35] So now to see the leakage, dialogues appearing in both the train and test web 4 [00:50:42] So basically, if we have, like, the same things coming in, training and test, it will bump up the number, but we want our test to have [00:50:53] Like, like unseen things, right? So that's why we did this test. [00:50:59] Now, to give you a brief summary, no EMT or missing values were, like, were there anywhere, right? Nothing to clean there. 476, duplicate dialogue is in train, which is about 3%, right? Harmless [00:51:13] But, like, wasteful. Since the model would see them more than once, okay? The 42 dialogs appear in both training and test, but data lake, this is a data leakage, and like [00:51:24] That would matter. So, we would try to remove this from training set, so that, the results are not inflated [00:51:33] At about 5% of the test set, the leakage would not up slightly, not rigged them. We remove it on principles. That's what I said just now. The takeaway habit is, like, we check… so this is why I care this one. So, if a dataset is large, right, large enough, then we can check for this. If they have any empty values, duplicates [00:51:51] Right? Or any leakage before training, so that we can have a result that we can trust right rather than just hoping the result is right. [00:52:01] Okay. [00:52:04] Until here, doesanyone have any doubt [00:52:09] Yes, it does. [00:52:12] Today we are using a transformer sequence-to-sequence LM, and yesterday we have used the casual LM [00:52:20] Yes. [00:52:22] Since our, it is because today is our objective is to [00:52:27] produce a summary, and yesterday we were trying to auto-tune the GP22 model to… for question and answer format, or it is like 45 always we have to use a sequence to sequence LM [00:52:42] Yeah, for T5, always we have to use sequence to sequence LM, because yesterday we were using GPT-2, right? It was decoder only [00:52:51] Right. [00:52:55] Okay. [00:52:56] Right? So in GP22, we only have stacks of decoder, right? Whereas in T5, if you remember the architecture, we have encoder as well as a decoder [00:53:01] Right, so first a sequence, if we give it a sequence as input, it will take it to the encoder. Whatever context it gets, then it will be passed on to the decoder for the generation point, right? So we have encoder as a decoder. That's why for T5, right, we use sequence to sequence element. And for GPT we used [00:53:18] Let's schedule [00:53:21] Understood. So, let's say if for, okay, today's objective is to generate a summary out, but for, NT5, if I have to generate a question-answer based [00:53:32] stuff, sequence to sequence is sufficient, or… [00:53:36] Any other technique is also has to be used in conjunction. [00:53:39] Oh, sorry, come again, see once, I don't get to it. [00:53:47] Correct. [00:53:48] So today, as our use case objective is to generate a summary out of using dialogs and summary, but if I have to, instead of summary or conversation, I have to [00:53:54] train this model, foreign question-answer format, like yesterday. So, do we need to add any additional techniques along with sequence-to-sequence transformer? [00:54:04] That… see, see, once, don't relate this, with, task in hand. Instead, relate this with the architecture of the LLM that we are using. [00:54:13] Okay. So, Casual LM you used, because it was decoder only, right? So we only had it does not have encoder part right. So that's why casual Lm was in and for that task it was enough right here. T5. What is happening is we have like, if you remember the architecture, we have imported as well as a decoder. So it it comes as a sequence to sequence [00:54:37] So that's why we are [00:54:38] Understood. That's what I'm saying, in the bot, I can understand this was encoder only, and it was filling blanks for classification purpose. GPA22 is for decoder only, it's for generating continuous tests. [00:54:49] Correct. [00:54:52] Correct, correct. [00:54:55] Got it [00:54:56] But T5, we can do all those things, translating, summarizing, and answering questions. Today, we are focusing on summarizing a text from the dialogues datasets, but if I need to turn out into an answering and question format, 45, [00:55:04] Is there any additional techniques I have to include, or this transformers you are explaining is sufficient for them? [00:55:22] Yes, yes. [00:55:23] This LLM, it is sufficient, but you have to keep in mind, like, the task we are going for, right? So we have to form a desk study also, if you remember, we did some formatting in our input, right? How we wanted. So that part you have to take therefore rather apart from that, we are just using different kinds of elements, right, which were trained for different purposes. [00:55:32] Like, we are using for a specific use case. So, according to that, you do have to do some formatting, just remember that part. [00:55:40] Okay. [00:55:42] Understood. Second question is, yesterday we have used auto tokenizer, and today we are using NumPy for [00:55:50] No, no, auto tokenizer only we have lost. [00:55:53] Good organizer [00:55:54] Okay, and then we have used… along with that, we are using NumPy also for [00:56:00] Some other [00:56:01] Where? In which part [00:56:03] Yeah, sure. [00:56:04] No, I see, it's an auto-tokenizer on the list. [00:56:09] Okay, and you have used below NumPy also, right? [00:56:12] Dumpai was to convert into NumPy array so that we could do some operations that, NumPy supports. [00:56:21] Oh, okay. [00:56:24] Okay? [00:56:25] Okay, order. [00:56:28] Thank you. [00:56:33] So… where was I? [00:56:46] So, now we'll do the, like, some cleaning, okay? We'll try to remove the duplicates in the training data itself, so that, it does not see the same examples again and again, and like [00:56:58] Also, we remove the duplicates between the training set and the test, so the test is always able to see the, like, [00:57:07] Like, unseen data. Okay [00:57:10] So, here, like, like a design choice we did was, we, cleaned our training set rather than the test set, because, test set already has fewer examples, right? And train has a huge train has used examples, so we instead we remove [00:57:28] elements, sorry, those that are present in both be removed from training itself. [00:57:33] It's already going close. Okay. Now, like for this, what we do is we import data set from datasets, right? And then what we do is we record the train size before cleaning like so we can know how many of them were removed [00:57:49] Right [00:57:50] Similarly, now what we do is we try to build a set right out of the test set, dialogs, okay? [00:57:58] That's why we have used the state operation. [00:58:02] Now, what we are doing is [00:58:05] We walked through the training split once, and then decide which rows to kick. So, we are keeping this, keep indices. We are basically that will having, that will have the positions, like, where, in the training set, which we want to keep [00:58:19] Right? Then we run this loop, right? So that if there is dialogues present in the dialogue, from the test dialogue is present in the test dialogue, we skip it right [00:58:31] And, which is, like, leakage which is present in both. And if, like, the dialogue is already, [00:58:40] Like, was already passed, right? In the training side, and it is appearing again. Then also we skip it, right? If, like, once we are out of this, loop, then what we do is, we add those dialogues which are, failing both of these cases, right? That is good for us [00:58:55] And we keep those in business. Okay? Then what we do is, [00:59:00] Yes, this is where we keep the… whatever indices were stored here, like, we keep those, and then we have our new clean dataset [00:59:07] Right? Called by Clean Trail. Now we create a new dataset, right? [00:59:13] With the help of data set which was matching the earlier format we had like train in train we had clean train that we just clean validation and validation set and test set we changed nothing there. So it is as it is, okay? I'm [00:59:28] Let's just have a look [00:59:32] Okay. So, train rows were around 14,000 to 113 rows, remote were around 518, right? This had duplicate plus leak test dialogs. Valtation rows were around 818 unchanged [00:59:48] And test rules were around 819 unchanged. Now, just to give you a brief, like, of what has happened, 518 rows were removed from the trend, 476 duplicates, plus the 42 leaked test dialogs that were present in both training set and the test set. [01:00:04] And now, like, there is no overlap, okay? Only training site was clean, validation and test site were left under, so the test says the same, standard 819, and our final score remained comparable to the Samsung results. Leakage was fixed by removing and [01:00:21] offending Rose from train notice, like it's repeating. And next part is, like, train set is slightly smaller now, because we removed some rows, right? Which is the one. So [01:00:33] Okay, so it will not exactly match the Samsung results, right? But it will be closer to me. [01:00:40] Let's go to the modeling part [01:00:43] Okay, so before I tell you about the strategy we'll be using today, let me quickly tell you about the T5 model. Just a brush up. T5 basically stands for text-to-textromer, right? It was released by Google in 2019 [01:00:58] Since prefix package, like, yeah, okay. Now, every task in, its core idea is every task is in text-in, text out, translation, question answering, summarization, all frame the same way, right? So this, this model can do translation, question answering, summarization [01:01:15] Right? And also,a short prefix added to the inputs tells the model which task to do. So, before we do any task, suppose we are… in our case, we are doing summarization. So we add the prefix summarization in like before we pass the [01:01:30] Input for the model. Okay, so from this, the model will know that the task and hand is summarization [01:01:37] Now, why the prefix matters? Like, since one T5 model can do many tasks, it has to be told which one, right? Summarized [01:01:47] Right? So here we are doing summarization. So we are keeping this prefix summarize in front of the dialog, and from this, the model will understand that we want to summarize the dialogue, not translate it, right? [01:01:57] use this exact prefix every time we create the model, so the instruction says consistent, okay? The size we use is T5, small. It, although it comes in small base, large, XL, XXL, right? Bigger is more capable, but slower to tell. So that's why, 4 out of 12 purpose, we are using small [01:02:17] You can use bigger models for this, and like the results may be better like in terms of the evaluation matrix, but it will be a little slower to do. [01:02:27] Okay [01:02:29] tokenizer, like, always, like, it basically reads the like the model reason numbers, not text, so tokenizer basically converts between the two. It breaks the text in tokens, maps each to an ID and can turn IDs back into words [01:02:45] Every input passes through it first, and every task, every output passes back through it. [01:02:53] Okay, now, T5 tokenizer basically works is based on the sentence piece, right? That's why we had installed this earlier while installing in the installation cell. Okay [01:03:08] So, it basically splits the text into subword pieces rather than whole words. So arrayed word becomes more into several pieces. This lets it handle almost any text, including typos, slang names and common name, which is common in sensual terms [01:03:23] Okay. [01:03:25] Now, one link. This is why length analysis was done in tokens, not words. One word was only, does not mean it will be one token, it can be various tokens inside. Now, just to, just for CVANs, right, [01:03:38] Like, he did… like, he was asking, right? So this also answers, right? See, this T5 was, like, can do various tasks, right, here also. Now, for this [01:03:48] particular task, we are working with Summarize, and this model was trained to do various tasks. So if we just add this prefix, right, the model will have some idea that the dialogue [01:03:57] Like, it is meant for summarization. But it is not necessary that every model will work in the same way, that every model was trained in a slightly different way. So it's necessary to go understand, like, how, like, before we feed anything the model [01:04:13] How to format it, so that the model can understand, okay? [01:04:17] Got it. Thank you. [01:04:19] No, sir. [01:04:30] So, now what we do is, we load the T5 small and look at the zero sort summaries [01:04:36] Okay, so from here, we import auto model for sequence-to-sequence element. [01:04:42] From Transformers, right? And then we lowered the zero-sort model, and so what we're doing is basically auto model for sequence-to-sequence LM from pretend model. The model limit is referencing, we have already given T5 small. It's referencing from there [01:04:57] Then what we do is, we move the model onto our chosen device. We have chosen CPU, so you don't move to CPU. Then we put it in eval mode [01:05:05] This basically, turns off the training behaviors, like dropout. So, generation is deterministic. And also, we are not training here, we are just generating [01:05:15] Because, you know, this is zero salt [01:05:17] Like we had already discussed yesterday, we are just starting right with zero examples. We are just telling the model to summarize it, okay? And we'll try to have a look on it. [01:05:29] Now, this eval subset side is basically taking the length of the test data, and then from [01:05:36] Here we'll select… so right now, earlier I had during building the pipeline, I had kept this. Okay, so right now it will select everything because we have loaded the whole data set, so it will select everything from the dataset test. Okay, then we import torch [01:05:52] Right? And then, what we do is we create a function for summarization. So what we do is [01:05:58] We create a prompt first that is shared as input to the [01:06:03] 5 model. So, like we discussed earlier, we have kept this prefix [01:06:08] Right? Before, and then the text file, okay? The text will be feeding to function. This is actually now the inputs tokenizer prompt, tokenizer, we have, like, we give the prompt, right? Then the whatever outputs we have will be coming out as Python tensor [01:06:25] Right? Truncation is true, like, if it crosses max length of 512, right, it will truncate, right? And this is moved to a device that we selected, that is CPU. [01:06:37] Okay? [01:06:38] Then we go with Todd Snow Grant. This will basically disable the gradient tracking, right? Now, max new token, we have set it as 64. This is a little more than 95th percentile and [01:06:53] Number of beams is we have set it as for basically beam search will explore a few options and pick the best one. Okay. Rather than just 3D recording [01:07:03] So after that, we return the decoded output [01:07:09] So, there will have the output IDs, and that will be decoded with the help of tokenizer.requil. Okay? Now, we'll have few examples of how it looks [01:07:18] But right now, we have loaded everything, right? 7,000… sorry, there were roughly around 800 examples. That will take a lot of time. So I'll quickly go up, toggle it right where I had said to use the entire data set [01:07:34] And you just, sub-sample of it [01:07:38] And rerun this. Give me a minute [01:07:41] Since the IDA party is over the quarterly part [01:07:46] Okay, so this would be a faster room. I need to… [01:07:50] Restart session. [01:07:56] We'll put you around me, sir. Please, bear with me for a minute. [01:09:28] Okay, we are here. We had already discussed this, because if you're on it [01:09:35] Including the weights [01:09:47] So, let's have a look. [01:09:50] Okay? So, in example 0, it was a conversion between Haina and Avrinda, right? [01:09:56] Hannah says, do you… hey, do you have BT's number? Amanda said, let me check, right? There is some file exchange between them, right? Amanda says, sorry, I can't find it. [01:10:06] Amanda says, ask Larry, right? Then again, Amanda says he called her last time, you were at the park provider. Hannah says, I don't know him well. Again, some G first said, okay? Then Amandi said, don't be shy, he's very nice. Hana says, if you say so, I don't… and then again says, I would rather [01:10:22] You texted him. Amanda said, just text him, and she says, alright, Anna says bye, and Amanda says bye, ma'am, okay? Now, the true summary of this was Hana needs Betty's number, but [01:10:34] Amanda doesn't have it. She needs to contact Larry, and the, T5 producers, the zero-sort summary as, Amanda, let me check Hannah [01:10:43] Right? Sorry, I can't find it [01:10:46] I don't know him well. Don't, don't be shy, he's very [01:10:58] He's very nice. And then Hannah agrees to do so. So, right now, what we have seen, instead of [01:11:04] Writing a summary, what he has done is, what the model has done is, it has simply, you know, copied few parts of the dialogue, if you see, along with the names also [01:11:14] So it's… and in the true summer, if you see it was just like a summary between the [01:11:20] two people with our referencing the names, and it was in a structure right [01:11:25] Similarly, we have for this example one, where, like, let me go through one more, and then we'll just look at the summaries. [01:11:33] So Eric says machine, Rob says, that's all great. Eric says, I know MSO's how Americans see Russians, right? Rob says, and it's really funny. Eric says, I know I especially like the train part. Rob says, haha, no one talks to the machine like that. Is this only stand up? Rob says, I don't know. I'll check [01:11:51] Roxys turns out, no, there are some of his startup on YouTube. Eric is great, I'll watch him now. Robs is me too. Eric Machine, Rob says machine. Eric says, talk to you later. Like, question mark, and Rob says, sure. The true summary is Eric and Rob are going to WhatsApp stand up. [01:12:08] On YouTube. And the T50 Sword Summer is I know I especially like the train pirates, haha, no one talks to the machine like that. Roxas turned out to know there are some of his standoffs on mutual [01:12:23] Okay, so that's it. [01:12:26] So, it simply copied the dialogues from there, right? [01:12:30] Similarly, we have it here also. [01:12:34] Right? The summary is, Leni can [01:12:36] Lenny can't decide whether which trousers to buy, Bob will buys Leni on that topic. Lenny goes with Bob's advice to take the trousers that are of best quality [01:12:46] And the zero salt matter is what matters is the most will give, [01:12:50] But, you will… what matters you [01:12:54] What matters is what will give you the most outfit options [01:12:58] Boxes pick the best quality, then any says you are right, and boxes, like, no problem. [01:13:07] Okay? [01:13:08] So, what the model just doing, like, to give you a summary of, like, what is happening is T5 is… T5 small can already write something right? It is 0 sort, and we just get it prompt to summarize it, and it is already writing [01:13:22] output of the like [01:13:26] Model is English, right? Not exactly Gibberish, the model is not broken, but like this is something we, can… it's fluent, but it's not relevant at all, right? [01:13:37] So [01:13:40] But the thing to note about is, right, English is not the same as it summarizes his mail also. So if you see closely, like I had mentioned, we'll see that, like, it's copying a straight line, a line straight from the dialog, right? And grabbing it where the first speaker said [01:13:57] And then, [01:13:59] Like, the second speaker said in multiple cases, and instead of capturing the whole point, right? [01:14:04] And sometimes it is also missing who did that. Samsung Dialogs have several people talking and a good summary names them and what they are agreed on, right? Zero sort model tends to flatten all of that and keeps it in a single survey. This [01:14:19] And, like, this makes… because T5 was originally pointed at news article text, not to group sites, right? And we are asking it to do a job it was never specifically trained for. Okay, so the honest vertex is it is [01:14:34] Like, we can… it is possible, but it is not good, right? The gap is the whole reason we are. This tutorial exists, and we'll be going for fine tuning and just see like how [01:14:44] It, it works, right? [01:14:46] Now, we'll look at the three we just looked at the three examples and it gives us a brief idea. But let's also evaluate it and try to find like [01:14:59] How it is performing in ZeroSort, and how it will perform after fine-tuning, right? [01:15:06] So, the metric that we'll be using is, road score, okay? [01:15:12] Like [01:15:14] So give you, like, a heads up. [01:15:16] Yeah, matrix from Robsco, we are using Rogue 1, rogue to Rogue 1. I think you will, everyone would be familiar with this. Basically, Rogue 1, it compares single words, right? And in Rogue 2, it compares the word pairs. [01:15:31] Single words would take the vocabulary, rote would check the word base, and rogue L is basically longest matching word sequence in order [01:15:39] So this would, like, help us check if it is in right structure or void order, in comparison to the reference. Reference meaning the human summary. [01:15:50] Okay, that was provided before the gold [01:15:57] Now [01:15:58] Let's start with the evaluation part. [01:16:01] What we do is we import, evaluate, and from, like, Tikodium.auto, we import T equium. This is to see the progress bar, right? [01:16:11] Then, we load a road, right? With the help of ePilot.load road, right? Then, we create two lists, empty list right now, which has zero sort predictions and references. References will have, like, the original human summary [01:16:28] And zero-sort predictions will have the summaries with the human… Sorry, the model generated. Okay? [01:16:34] Now, what we'll do is we'll try to [01:16:37] Look at the eval set, right? And, feeling both in, like, and we'll try to, [01:16:45] Just compare it both. Like, this is how the rogues will work. It compares the… both of the summaries, right? No matter what the model predicted, and how the gold tooth was, tries to compute it. [01:16:58] Okay, so here again, it is being computed, right, with the help of these predictions, zero sort predictions, and then the in references, we are given differently because we have defined all the human references here [01:17:12] Now, let me quickly run this [01:17:15] Basically, we have a record [01:18:16] Now we have [01:18:19] Like, rogue short… like, on Zero Short with the help of 20 examples, we have some stores, right, on rogue stores. Rogue one is around 0.25, right? Rogue 2 is around 0.06, right? And Rogue oil is 0.2 [01:18:35] Right, so if you see row 2 is very less, and row girl, and Rogue 1 is doing [01:18:43] Like, like, it's slightly higher than the road, okay? Now we'll try to see [01:18:51] How it actually performs for this one. [01:18:55] And, full data set. So there, [01:18:58] We had Rogue 1 score of around 0.27 right this was just to give you a heads up, this was on the full test set. [01:19:08] So there we had road score of around 0.27. This means a little over a quarter of the world in the genitive summaries are also found in human summaries. Some overlap is present, but most of it is being missed [01:19:21] The road 2 is of 0.07, right? Because two word sequences are, like, what this metric said. Correct phrasing, word order are what it rewards. A score of a score this low that individual words are sometimes shared with the reference [01:19:36] But they are almost never stung together, right? Strung together. Basically, it may be there, but, like, maybe sometimes two words are not coming together, right? And this basically checks the word pace, right? The way our human could actually [01:19:51] Now, Rogue L, which is, like, longest… for longest moments of sequence, right? In order. So, that was found to be 0.21, the longest run of matching words is that, like, what is matrix effect? So the overall structure of the summary is what is being captured here. Some shape is being bought, but most of it is not [01:20:08] taken together, like, we can say that, scattered correct words are being produced, not coherent correct summaries. The distance between the right words and the right sentence is what these numbers are pointing at. [01:20:18] A fair, fixed aortic yardstick has now been set. These dispose measures on like a full set of 819 examples generated in exact way. The same measurement will be repeated after fine tuning, and then this number [01:20:34] Like, with this, with the result we had in XeroSort, we'll compare that with the, like, the fine-tuned one. [01:20:41] Now we move on to the fine-tuning section, and before we do that, like, we need to do some, [01:20:48] Pre-processing and like tokenize the data set properly for the fine-tuning, okay? [01:20:55] So here the main goal is to convert the raw text, which is dialogue plus Summer into token IDs the model can train on. Dialogs [01:21:05] Based on the dialogue, we expect the model to provide a good summary, right? Let it be closer to the ground term [01:21:13] Okay. [01:21:14] Now, like, 3 things will happen, like, dialogue will get a summarized prefix, okay, which tells the T5 will start to do. Dialogue will be tokenized as input, and human summary will be tokenized as label, or, like, or the target, right, that we [01:21:29] Padding will not be done here on purpose, because, [01:21:34] It should be done in the next cell, okay? [01:21:38] So, this is, like, maximum input length we have set on the basis of the EDA that we had taken 95%, a little more than that. So for maximum input length, that is for dialogs it is 512. And for the [01:21:52] output part, that is summary we have taken as 64, okay? [01:21:56] Prefix is summarized like we had kept in zero sort also, right? We have said that. Now we try to pre-process the examples. [01:22:06] So, for each of the dialogue, right in [01:22:10] Like, in the data, like, we'll add a prefix in the inputs, so that's what we are doing here. Then we tokenize the inputs, set the truncation as true, right? And this padding we are not doing here, that the collateral… we are using collater, right? It'll be handling it, it will handle it later [01:22:29] Okay? Similarly, we, tokenize the summaries, right? Again, truncation is said true, right? And max, target length that we are doing a set of 64,right? So that will be used [01:22:40] So now [01:22:42] What we do is, we attach the tokenized summary IDs as the labels field, right? So we got the labels, we'll get some labels, right? [01:22:51] Sorry, this summary after tokenization, we get some ids right? These are tokenids. So that is like that is being referred as input ids. So all of these sequences there, and then that will be [01:23:06] being kept here in model inputs labels, right? Then we are returning the model [01:23:12] So this was the function, okay? [01:23:17] Now, we confirm the result [01:23:19] So, what we do is, we apply it in one of the rows, right? And, try to see, like, [01:23:26] The first 12 token IDs of input token ID, and labor token ID, meaning, target token ID, right? Let's have a look. [01:23:38] So, if you see here, dataset has features, input IDs, attention mask, and labels, number of rows 60, because we have reduced it now, right, for this tutorial purpose. Validation is a number of rows is 10, and features is, like, input IDs, attention map label same [01:23:55] Number of those is 20, and features are again soon, like all [01:24:00] Okay? Now, in one training example, we'll have input id, attention mask and labels. First 12 token ID is like this is for the first [01:24:10] Input one, and this is for the [01:24:13] this output, which is the summary. [01:24:16] And this is the attention mask, okay? Now, [01:24:21] Right now, like, [01:24:23] If you see, all are one year, okay? Now I'll summarize everything. [01:24:30] So, basically, raw text volumes are gone. Now 3 numeric fills remain input IDs, attention mask, and logos. These are what the model actually read since the text is never fade directly. Now, every input begins with the same token ID 21603 and 10 [01:24:45] Okay, this is, that pair is for the summarize, this part [01:24:51] Prefix, so that, its presence at the start of each example confirms the prefixing or terms as intended [01:24:59] So, 621603 and 10. [01:25:01] Okay? The labels are the human summary turned into token IDs, and they end in 1, the final one is the end of sequence marker. This, which is how the model is taught, where the summary should stop. [01:25:14] Okay, so [01:25:15] Right? [01:25:17] The attention mask is the same, length as the input IDs right now [01:25:23] Every value is one. This is equated because no padding was added in this step. Since we don't use any padding here, everything is one, okay? [01:25:32] Because if there was padding, suppose, then whichever token corresponding to it was padding, it would have been 0. [01:25:41] Do all one's mask is not a finished picture. Once patches are padded in the next step, some positions will turn 0 to mark the filler tokens. [01:25:51] The change from all ones to some zeros is what the mass… is what makes the mass rock visible. Now the data collider thing that you are talking about in the previous will try to do that. So here, what we'll do is we'll try to set up 2 last pieces needed before the training [01:26:06] A data collater [01:26:09] That means each training badge pads the example in a batch equal to length marks padded label positions with minus 100, right? A phrase copy of T5 small, like, will train them. Now, why this [01:26:25] Okay, so example in like all the batch will have different lengths certain ones will be padded with the filler tokens to make it a [01:26:35] Neat example, but we must not, grade the model on a predicting filler. [01:26:41] The loss function ignores any label position equal to minus 100, so the collater swaps every label to confirm minus 100, right? So, basically, when we swap this, the model will understand that not to calculate loss here, right? It will simply ignore that [01:26:56] Okay? So, for now, let's come to the task in hand [01:27:01] From transformers, we import a data collater for sequence to sequence, right? And auto model for sequence-to-sequence LM. [01:27:11] Right? Now, like, we'll try to, load a fresh model, auto model for sequential sequence LM, like before, from pretend model name, and then it will be taken to device, right? And it will be referenced as fine-tuned model node [01:27:26] Now we are having a data collater. Basically, it will need a tokenizer, and it will need the model. Model, we, got it from here, and tokenizer, we have already defined above. Okay, so from that, data collater is being created [01:27:40] Okay? [01:27:42] So, at batch period time, like, batch bill… like, what happens is, like [01:27:48] These, inputs and labels to, like, it is padded to the longest item in that batch. This is called, let's say, like, there were 10 items, right, in that batch, right? So, it will [01:28:01] take the longest one in that, batch, and then, with the help of that, it'll, you know, pad it dynamically, for the other sequences. Let's say the longest was 10, and the shortest was, 5 [01:28:17] Right? So based on this longest time, it'll pad the, [01:28:20] The other file [01:28:22] So it will dynamically learn [01:28:24] Okay? Now we'll try to have a look at a sample batch right for that we have taken [01:28:31] Two rows, a zeroth and the first rows, and we'll try to see the batch level, batch attention, mass set labels, for example, labels, labels basically mean target, right? Attention mask, for example, 0 and [01:28:45] And then we'll try to dorsal inference. [01:28:56] Now, if you see [01:28:58] This batch level SAP right now is [01:29:01] Two, because two examples, and one is 14, right? And the other one is 13. [01:29:08] Now, for, example 0, these are the, tokens, right? For example, one, these are the tokens. These are for attention mask [01:29:17] Okay? If you say these are all ones, and here in example 1, we have one 0. [01:29:22] Now, we'll just try to draw some influence from this. [01:29:27] Okay? So, what the collater just did was, it made two neat rectangles out of the uneven examples, right? One was [01:29:35] Like, of two examples and 14, right? And the other was, 2 cross 30, okay? Two examples is padded to 30 input programs. [01:29:46] Okay? So the two sets differ, 14 versus 30, because labels and input are padded separately. Summaries are sought, dialogues are long, which is given. Now, each pads only the longest item in this batch, dynamic padding. So. [01:30:01] Not to some fixed 512. So, if… if in the batch this is not required, then it will not go on till 512. That was the whole idea. Okay? Now, the binary trick was that basically [01:30:14] In the levels part, if, so, example is, like, in the example zero levels end with 3 minus 110, so this minus 100, right? That means example 0 summary was sorted, so the tail got padded [01:30:28] So if you see her [01:30:31] Level 3 here [01:30:32] I would say minus 100. That means the example was sort of, right? And hence it got paddle. [01:30:38] This means examples, yeah. So minus 100 basically means, do not compute loss here. The model is not graded on this filler. Okay? And see in example one there are no minus 100. It means it was the longest summary. So nothing was fine. [01:30:52] The attention mass shows the same idea on the input side, right? Example zero, like, there is all ones [01:31:00] Right? And in example, one, the last one is 0. [01:31:04] Okay. [01:31:06] So basically, like, where it is zero, it says that the dialogue was shorter, so it water one pad token, right? Having zero means ignore this position. The model will pay no attention. Okay [01:31:19] The key insight is to separate ignore me signals. Minus 100 will hide the padding from the loss, and 0 hides the padding from the modal attention, so it does not read the filler [01:31:29] Okay, these are independent, and you not land on the same example. Here, example 0 has minus 100, sort summary, but example 1 has 0 in the mass, sort dialog, okay? Same concept, two jobs, one protects the loss, another protects the attention, and knows what to read and what not to read [01:31:46] Okay. [01:31:47] Now, for the training configuration [01:31:53] Before this, like, is there any question? Does anyone have? [01:32:02] Okay. So now we'll move on to, training configuration, sequence-to-sequence training arguments, plus sequence-to-sequence trainer [01:32:10] Okay, well, that is our goal here. So, main goal, like, we are… basically, we are trying to wire up everything needed to fine-tune, and [01:32:18] Basically, we are trying to have sequence-to-sequence trainer arguments, and then we'll have sequence-to-sequence trainer. So for that, what we are doing is we are importing sequence to sequence trainer arguments and sequence to sequence trainer from the transformers [01:32:33] Okay, now we are trying to pick up the season, right? T5 is known, like, this was, just a design choice, okay? [01:32:42] Some, like, T5 in, like, is known for unstable in FP16 sometimes, okay? And it was treated in, [01:32:50] This BF16. So [01:32:53] We never use FP16. Instead, if the GPU supports… if Sapone is working on GPU, and if it supports BF16, right, then… [01:33:03] Use BF16. [01:33:05] Right? Otherwise, like, you can switch it to FP16 also. But, like, better to use on DF60. [01:33:12] Okay. [01:33:14] So this is basically for the precision time, okay? Okay, and now the cleaning arguments, we have, like. [01:33:24] Basically, these are the hyperparameters that we are selecting, okay? All the outputs will be stored in T5 small, this iPhone Samsung, right? Seed via sitting as 42, so that, you know, we are able to produce [01:33:39] Same results across the ranks, even if it is sampled, you know randomly, okay? Number of epochs, we are taking it as 3. [01:33:47] 3 is the first because sometimes in one the model is not able to understand. And if we have three, at least we get to choose okay which of the [01:33:58] had produced a better result. Okay [01:34:01] Now, training that size per device you have selected as 8, right? And, for eval batch size also, we have selected as 8, okay? So, if you take the, larger, batch size, then that would require [01:34:16] your GPU memory or memory, right? And if we require it small, then it will, [01:34:22] Would you take less memory, but more time, okay? [01:34:26] So just in case if someone runs into like while running into their machine, you can halve it if you are running out of memory or something like that. [01:34:35] Right? Reduce it [01:34:37] The learning rate we are taking it as, 5 into 10th of power minus 5, right? This is our standard and reliable starting point for filing, right? Too high, if you take it too high, then basically will be, you know [01:34:51] Like, it will affect the, [01:34:55] T5 model rates negatively more than learning it, it will [01:34:59] Like, it'll hamper the model. [01:35:02] for the task, okay? [01:35:04] We are taking warmer pressure 0.1 because like, so what happens is like [01:35:11] Instead of starting the learning rate straight from 0.1, it will slowly ramp up the learning grid. [01:35:18] of the first up to the first 10% of the steps. [01:35:22] Right? So that we have a smooth start at the start of training, which generally improves the stability, okay? [01:35:31] Weight decay, we have kept it as, 0.01. This is the card we have kept for overfitting. Logging steps is, like, for each step, how many times, so during the training, how many… how many… at what steps you want to keep the logs? So that way I kept as [01:35:47] But if you do it for the full data set, you can increase it, okay? Like yesterday. [01:35:53] Since we had very few examples, so I had asked this lower number [01:35:58] Eval strategy, ebook and safe strategy epoch is basically used to [01:36:03] This will be done after each epoch, right? So, after each epoch, there, there'll be some evaluation that will help us, get the validation loss or something like that, that will help us see if the loss will decreasing or not. [01:36:16] And safe strategy is, like, it is safe [01:36:20] The base model, after the epoch-based computer [01:36:24] Okay? So, this, predict we generate, we have set it as, [01:36:29] false, because, [01:36:32] Like [01:36:33] This generation is slow, and rogue is already measured on the full data set in cell 4 right during training. We only want to watch all the eval loss. Okay, so generating here is a wasted one. Okay [01:36:44] And, this one, load-based model at the end is true, because, like, we want to have the base model when we go for the evaluation. Okay? And because we are running for 3 epochs, right? And during 3 epochs, we'll have 3 different [01:36:59] validation loss, and based on that, we can sell it. So, for how this metric that we are going to select with the help of which we are going to select the best model is email loss, okay? And greater is better is cost of also means we want it to have a, when the model is, [01:37:16] has lower loss, we want to select that [01:37:20] Okay. [01:37:22] So, this is for the same reason that I told for patient. And a safe limit, total limit is to, like, it means we are keeping two checkpoints, right? At any given time [01:37:36] Okay? So, and we are not producing these results to Hugging Face Hub, so this is kept as [01:37:43] Now, this was for the trainer arguments [01:37:47] Right? [01:37:48] Now we are going for train. [01:37:51] trainer, right? So now, like, will, for the sequential sequence trainer, we have model will be you that will be used is fine tune model. We have we have already loaded with T 5 small like in cell 6 [01:38:03] Arguments is training hours, training hours is basically just the things that we defined here. [01:38:08] Okay? Then eval dataset, sorry, terrain data set is trained, tokenized data set of train. Eval dataset is tokenized data set of validation. Data Collator, we already had defined earlier, right? Processing class is tokenizer [01:38:22] Right? [01:38:27] So, because what happened was, earlier when I was, [01:38:33] Doing tokenized tokenizer, it was giving some errors. This is because of the… some transformer version that's why I clicked this error [01:38:55] Now, we'll run the full fine-tuning, right? The main goal, actually, is, like, till now, what we have done, everything is [01:39:05] kept in place. Now he go for the, [01:39:09] Training part. It will run for 3 APOCs on full training set, but right now we have taken, some sample of it, right? And it will evaluate on the validation set once after repo, right? Now, what to watch while it runs? We'll watch at loss [01:39:24] Right of the training loss, that is, eval loss is basically based on the validation data, what loss they are getting, and the time. [01:39:33] So [01:39:34] We'll kick off the training. We'll call trainer.train, right? Then, the matrix trainer.result matrix will be stored in matrix, right? And then we'll save the fine-tuned model plus tokenizer to that this, okay? So that we can reload it later if we want to [01:39:51] Let me quickly run it. [01:40:03] So, this will take some time. In meanwhile, like, does anyone have any questions? If yes, then we can answer those. [01:40:10] If not, then let's wait for a couple of minutes. [01:40:21] I think, you know, take around 4 to 5 minutes [01:44:28] Okay, two epochs are fast, right? And if you see the… you know, it is decreasing. It was 2.7 knots for the testicle [01:44:40] The painting loss also decreased only 3.07. [01:46:44] Now in our case, luckily, you know, the training loss has decreased [01:46:51] Right? Because we are going… [01:46:54] And for validation laws, also, like, if you see, it will reduce from 2.71 to 2.59 right from 2.5 million import produced to 2.5 [01:47:07] So, now I guess this will be [01:47:10] Best moment, sir, right [01:47:13] I'm saving my mother [01:47:19] It's sold in this folder [01:47:24] This order is here. [01:47:30] My brother will result [01:47:35] So, like, in between, I was getting some morning, so just so that everyone does not get [01:47:43] like, started with it, so I had all this note. [01:47:52] Right? Now quickly moving on [01:47:59] So, we'll try to measure the, fine-tuned model, right, on the same thing, like, we have done before. We'll do it on HubSpot, right? [01:48:09] So, basically, we'll try to now produce the after number before fine-tuning number, we got from the zero sort model itself. Now we'll try to get it after the fine-tool model has been trained right [01:48:24] And, eval set will be saying that we had used before, right? And, we'll use the same summarizeOne function [01:48:32] Okay? And the same generation strings act as the zero shot that we did in cell 4 because only the models weights got changed. Any difference in the roles were installed by the fine telling me and nothing else [01:48:45] This was a heads up. Okay, so from we import TQDM auto. Sorry, we import TQDM from TQDM auto right? And then we run this model in eval mode [01:48:58] Because now we'll be using poison. The trainer we have already trained it. Now we just want to [01:49:05] evaluated, okay? We created a… we create an empty list, just like we did it in, Salesforce, right? We call it fine-tune predictions, where all the predictions will be stored [01:49:19] No [01:49:21] Like, for, from the eval set, like, we'll try to have production. This is summarize one example we had already created a function before, right? And then we pass the dialogue to it [01:49:36] And then the fine-tune model right [01:49:39] as arguments. and whatever results we get [01:49:43] We store it in [01:49:45] predictions, this prediction, okay? Then we append that in fine-tune predictions. [01:49:51] After this [01:49:53] No references was already built in cell phone, because there already we had created from the same event set, so we don't have to do it again, right? And, now to compute rogue score, what we do is, we use the fine-tuned prediction this time. Last time we had done zero sort predict [01:50:09] Right? And references will be the same [01:50:11] Okay, we have already defined this. Now let's quickly have a look at it. [01:50:39] I'll have a quick look [01:51:20] Okay, so, like, on the 20 test examples, room score was 0.26, row 2, 0.06, row was 0.1 right? So results like it is understandable that it is not so good [01:51:36] First of all, like, [01:51:39] We are working… we treat it around only 60 samples, right? And that is very less. We had around 14,000 samples to work with, and out of there, we do this, time constraint and all we train for the future purpose. But if you guys have time, like, if you take it off on home. [01:51:57] If you drain it, you can have better results. Now, I did train on the, like, a full data set, and I have [01:52:05] awesome results to share with you, so that, like, if anyone does not have the resources to run this by themselves, right? Robots are basically when I did this robots for robots from 0.27 to 0.45, right? [01:52:20] An increase of around 1.7 times. Far more of the right words are now present in the summaries, okay? Row 2 rose from 0.07 to 0.21. That is an increase of about 3 times, right? This is the largest jump of the 3, and it is the most element, because two word sequences are [01:52:38] at what it means is, right? This says the model is now stringing words together the way our human world, not just sharing the isolated words. Grow well also rose from 0.21 to 0.37, an increase of about 1.8 times [01:52:53] The overall structure of the summaries are now lined up [01:52:55] lines up far more closely than the human virgins [01:52:59] The pattern across the three scores is the real story. A metric that [01:53:05] Rewards correct phrasing improved the most, which means the largest gain came in from the hardest part of the task. Writing fluent, summary sentences rather than scattered correct words. So if you remember, like, earlier, we had sentences that directly took the [01:53:21] you know, copied the text from the, [01:53:25] this dialogues itself and pasted it, right? So now it is now when I had the results, it was more in the form all the summary events right now. These numbers were measured on the full set of around 819 examples using the same generation settings as a baseline [01:53:42] Only the model gets sent between the two measures, you know, so the improvement can be attributable to find any anode. [01:53:48] Now we try to do a close, before close, like, side-by-side comparison, for the before and after, right? Versus human [01:53:59] Okay, so for this, what we are doing, we are choosing 5 examples like [01:54:05] We run it 5 times, and we'll have, 5 sets of dialogues before is basically zero short predictions are very well, after fine-tuned predictions. [01:54:14] Okay, we'll try to print it. [01:54:18] Now, the results that I'm showing right now is from, like. [01:54:23] This our, [01:54:25] model that was trained on 20 examples, right? So and by the results also we know that it has [01:54:32] It is still not good, right? [01:54:36] Still, let's have a look. [01:54:39] So, like, for the first dialogue, just to… [01:54:45] So, dialogue was, Hannah, do you have Betty's number? Amanda, let me check. Hannah says. Amanda says, sorry, I can't find it, Amanda says, ask Larry. He called her last time, we, [01:54:56] operator. Anna says, I don't know him well. Hannah again says a joke, right? Amanda says, don't be sang, he's very nice if you say so [01:55:04] And after that, Hannah said, I'd rather you texted him. Amanda said jurisdiction, or Hannah says, all night, bye, and Amanda said [01:55:11] Now, the human summary that was intended was, Hannah needs Betty's number, but Amanda doesn't have it. She needs to contact Larry. Zero sort was this [01:55:21] Let me check, Hannah, right, Amanda, sorry, I can't find it. It's okay, right? [01:55:27] And [01:55:40] Great, and where is it? [01:55:46] And, now, the fine-tuned example is Amanda, sorry, can't find it, right? So it's not very good. It's a very evident, right? And it's still not following the structure also. So 20 examples where clearly not enough [01:55:59] Sorry, 60 examples will clearly not enough for the model to join. [01:56:04] Right [01:56:05] So, similarly, if you look at the other human summaries, right? [01:56:09] This is what human salary was [01:56:12] This is what the zero sort was [01:56:16] Right? And this is what the fine tone was. [01:56:21] So, in many cases, you can see [01:56:25] Like, it was [01:56:27] It's not that, there's no, like, no [01:56:32] improvement as such, right? For the 60 examples [01:56:38] Similarly, [01:56:42] And, yeah, also. [01:56:45] I think it's almost the same. [01:56:49] Right, between, before zero sort and after the fine bill, yeah. [01:56:53] Now, for the dialogue 3 [01:56:56] Again, like, 34 is also very similar [01:57:00] The work we had after [01:57:05] Now, like, this is understandable because, like, for 60 examples, like, we can say that do not watch a training happen, right? [01:57:16] Now, I'll try to show you some of the examples I got from during the after training it on the full data set. Let me just show it to you [01:57:25] So this was a dialogue. Human summary was now again, I'm just highlighting this was on the full data set. This one was on 60 samples, right? [01:57:36] So, this time the human summary is Hannah needs Betty's number, but Amanda doesn't have it, she needs to contact Larry. And before zero sort is Amanda, let me check Hana file, Jif, right? Amanda, sorry, can't find it, Hannah says, I don't know him well. And so, this is [01:57:52] Still in the after, it's much better. If you can see, the dialogue that it was simply copying in from, that has [01:58:03] Been reduced right now it is following some structure, right? So that is good. So, now it's… Amanda has Betty's number, she can't find it. Larry Potter last time [01:58:13] We were at the park [01:58:16] Great. So [01:58:18] It's much, much better than the zero shot, right? But it's still not up to the mark. Like, we were at the top. It says we right instead of saying they [01:58:29] Kind of [01:58:30] Right? And it also means the fact that she needs to contact Larry [01:58:37] Right? [01:58:40] So [01:58:42] That part is not there. So, we see that it has learned some structure, right? But the summarization part is still like it still needs some work like it's there, but it can be better [01:58:56] Now, let's see, for the example one [01:59:01] Is it Eric? So this example also we have really tones. So before was Rob. I know especially like the train part. No one talks about the machine like that. There are some of his stand-ups on YouTube. [01:59:12] Human summary was Eric and Rob are going to watch a stand-up on Youtube [01:59:16] And after fine-tuned, we had Rob and Eric are going to watch on Eric, stand up on YouTube. Now, they missed this part. It was not Eric's stand-up, right? They were going to, [01:59:28] So, there were, like, there are, machines, that's so great. I know, like, also American sea restaurants, right? So, this added… this is the extra part. [01:59:37] Like, and here it looks like, you know… [01:59:41] They are just going to watch a stand-up in some other stand-up, right? And it took us Eric's stand-up, because Eric was referring it [01:59:50] So we see that it is very close right? But still it needs some work. So this is also a problem. [01:59:56] We can also say from there. [01:59:58] And now [02:00:02] So here, it's about the genes part. So your WhatsApp, Bob says Henny says, babe, can you help me with something? Bob says your WhatsApp says which one should I pitch? Bob says send me photos, she sends she photos. Okay [02:00:16] And, boxes, I like the first one. I like first ones best, and Lenny says, but I already have purple trousers, does it make sense to buy… does it make sense to have two pairs? What says I have four black pairs, and Lenny says, yeah, but, shouldn't I pick a different color? Box says. [02:00:33] What you will give you the most outfit options. Then he says, so I'll… so I guess I'll buy the first or the third bed then. The box says, pick the best quality then, and then it says you're right, thanks. Box is no problem. [02:00:45] So the human summary was Leni can't decide which trousers to tie. Bob advised Lenny on that topic. Lenny goes with the Bob's advice to pick the trousers that are best quality, okay? And after fine tuning, we have results that Bob will help Lenny without the choices. Lenny will buy the first job [02:01:01] Okay [02:01:03] It is still able to, you know, get, some summary, but, you know, some… it can… little improvement room for environmental risk [02:01:12] Structure is that summarization has been done. The human main part has been removed and it is correctly summarizing it. So that is [02:01:26] Now, for the example 3, right, similarly, we have, now, I'll not read the, some of this dialogue once. Okay, I'll just see the summary. The human summary will be home soon, and she will know late will [02:01:40] Wait, what before is, like this one? After it is, Amal pickup will when she gets home. [02:01:48] Okay, so here, the context was missed, I guess. Right? Because here, the summary was, Emma will be home soon, and she will let well know. [02:01:57] And there it is written, and I will pick up well when she gets home. [02:02:01] So, so a little bit of context is missing. [02:02:05] Okay. [02:02:06] Now, similarly, for example. [02:02:10] Also, if you see Janice Varshall, and Oli and Jane has a party, Jane lost her calendar, they will get a lunch this week or on Friday, only accidentally called Jane and talk about whiskey. Jane canceled funds. They will [02:02:25] For a T-axis pointing it is gen and only had Warsaw for dinner on 16th and 18th. They have lunch this way, they will check on Friday [02:02:38] Right? So [02:02:41] It was able to get it, right? But again, there is some room for improvement. [02:02:45] Now, just, what did we learn from all of these examples, right? [02:02:51] So, format initially, like, we had in before after that got flipped [02:02:57] Right? Initially, we had, husband, like, Amanda, names were there in the summary, but after, like, the structure was there, right? [02:03:08] Okay [02:03:09] And this will… this is what helps in to get the good growth score, right? Later down, I have used bark score also has to do some comparison [02:03:18] And, if any of these examples, the fact was still imperfect, there was room for improvement, right? And somewhere it missed the context also. [02:03:27] bonus takeaway is fine-tuning is reliably not the, like, has reliably taught the style of good summary, concise thought person, abstractive, right? A dramatic physical win. It did not make a small model perfectly faithful, it still drops details, reverse roles, and sometimes even invent those hallucinations [02:03:44] Like, this is what we saw in the examples also, right? So that's why we need the examples and check metrics. [02:03:50] Right metrics things much better on average example better is not solved. Okay, so more data. So a workaround could be that right now we did return 60 million model, right? We could have taken a better [02:04:03] modernized model. Like yesterday for decoder, we had taken around because 124 million parameters, right? This one is only 60 million. So maybe if we go with [02:04:14] The performance can be better, but keeping at the same time, we don't want the model to overload it, right? So that should also be taken in consideration. [02:04:28] So, like, we did a quick bugs for test on this test set, as well as [02:04:36] Sorry, before this 0 said and the fine-tuning set also like to quickly run it. So basically what we did was we imported valid. We did a bug spot [02:04:48] So with the help of evaluate.load, also we loaded it, right? And then we did a zero sort to find the zero sort [02:04:57] Budspoon? [02:04:58] And similarly, we did it for fine-tuned. [02:05:01] Right? We computed the credit, with the help of fine-tune predictions and the references, just like we did for road score, right? Language is selected as English. [02:05:10] Okay, then we have [02:05:12] I'm looking around this [02:05:33] Till then, does anyone have any questions? [02:06:20] So, but so we got around 0.85 fine-tune was also no significant improvement, that is understandable, because we didn't get much. [02:06:27] Right? We saw that with the examples also. Now we'll try to visualize all the metrics that we had right? And we'll have graphs. Also, if you see it's everything is comparable only, and it was clearly evident from the examples also that we saw that 60 examples were clearly not enough [02:06:45] Now I did this, like, always, like, I have done this for the full data set, right? And there, what I saw was birthgo increased birth for iPhone increased from 0.86 to 91 again of about 0.05. [02:07:00] At first glance, it looks small, but, like, this is a good improvement, okay? [02:07:08] Now, like, now we are almost at the end of the, tutorial. Let me just quickly summarize of what we have done, right? We have taken a pre-trained T5 small, and, loaded it with Samsung [02:07:23] A dataset of a messenger-style dialogue paired with human written summaries. Before any training, the data was inspected and cleaned. Token lengths were measured to choose a sensible limits, which shows 95 percentile in our case. The fee which was 99 also, right? And, duplicates and dict examples you remove from the training site [02:07:41] So that the finance force could be trusted, okay? At zero sort baseline was recorded first, on purpose, so that there was a clear and honest baseline to compare it, compare the fine-tuning 5 also [02:07:54] So the model was later fine-tuned for three epochs on the full training site, validation loss was watched. And from there we [02:08:01] With the help of that, we found that it was no overfitting, and both training loss and the validation loss decreased as the number of epochs increased. [02:08:10] The loss was measured with three ways for water overlap and pod score for meaning. And also we saw side by side examples like we can take it as a human evaluation. We also saw all the examples side by side and just to see if it [02:08:25] How far is it from the gold answers. Now this is just a summary of what we saw, right? Before zero, this is on the full data set, right? And so it has 819 examples in this set. So before we had 0.27 of row [02:08:45] And after we had it, after fine-tuning, we had 0.45. It bumped off, around 1.7 times [02:08:52] to like was three times the zero sort, right? Rogal was around 1.8 times it increased right and birth score F1 [02:09:03] Like, it was a decent increase, okay? [02:09:08] Now, what we learned is like fine tune we saw the effect of fine-tuning, right? [02:09:14] The largest landed on the metric, that was rewarding the correct phrasing, right? With message that side-by-side output also. [02:09:21] This is the reason automatic metrics were paired with reading the outputs by hand. So, always remember, whenever you're working with LLMs [02:09:45] Just try to look at some of the examples also, right? Just to know, have an idea, just [02:09:52] Don't kindly trust the metrics, because sometimes, even though numbers are high, we do not know if, like, how far is it from the real truth, right? So it's a good practice to look at the output also. [02:10:07] Okay? [02:10:08] So, like, with this, also, one more thing, you could improve the results, with, let's say, using higher like borders with more parameters, right? Here we use only 16 million, right, which is [02:10:24] Fair enough, right? We saw some significant improvement from the 0. But there is still room for improvement. We can all agree on that, right? And also, if you have more data, right, to work with, that's even better [02:10:38] So I think with this, we can close. We are towards the end of the class. [02:10:45] If anyone has any doubt, we can address that [02:10:48] Does anyone have any doubt? [02:11:02] Okay, so I'll take that as a no. Thank you everyone for joining the class. I hope you guys had [02:11:09] hardboard learning session right so please when you when you sit at home, just look at the notebook. I hope [02:11:17] Right now, maybe you got the list. If you do it, try to do it once by yourself, just try going through the port, you'll definitely understand better, and I hope the comments that were, like, though it makes the code look very long, it… I hope it helps you guys to understand, especially for the beginners who are not so comfortable with both [02:11:36] So that the comments are helping them out, right? Apart from that, like, I also had a very great, like, good opportunity to, like, teach a wonderful badge like yours. Thank you so much, guys. [02:11:50] Yeah, with that, I think, we can close for today. [02:11:56] Best of luck. [02:12:01] Thank you. [02:12:02] Thank you. [02:12:03] Yeah, thanks, Pratik, for taking us through the session. Thank you all [02:12:09] Thank you. Bye-bye, we'll close the session for today.