# 08 2026-06-06 T5 Hands On

course: Module 4 — Generative AI & LLMs
module: Module-4-Generative-AI-LLMs
date: 2026-06-06
type: transcript
video_url: https://personal-learn.armco.dev/files/_Recordings/Module-4-Generative-AI-LLMs/08_2026-06-06_T5_Hands_On.mp4

---
[00:25:54] PRATIK: Hello, everyone.
[00:25:56] PRATIK: Let's wait for a couple of minutes, and we'll start.
[00:26:58] PRATIK: Can anyone please confirm if my screen is visible?
[00:27:05] Sacheen Adavinavar: this…
[00:27:50] PRATIK: Statistic.
[00:27:55] Deepak Katara: Yeah, I think we haven't got the material link today, the material link.
[00:28:03] Deepak Katara: Can anyone confirm?
[00:28:16] PRATIK: Okay, hi, Aviva. Very good evening. Today.
[00:28:20] PRATIK: Like, I think you guys must have studied GPT in your last class.
[00:28:24] PRATIK: Right? Last class, or last class. Like, so, based on that, today we are going to have a tutorial.
[00:28:32] PRATIK: And if you guys… Like, remember last time, in the last tutorial.
[00:28:39] PRATIK: We used BERT, and in that case, we had classification. Yes, someone is raising hand.
[00:28:46] Shivansh Sharma: As we not got the lab materials off today.
[00:28:50] PRATIK: It must be uploaded. I had already saved it with Matt.
[00:28:54] Shivansh Sharma: We took… I cross-checked on the LMS, I didn't find if there are any… anywhere else.
[00:28:58] GenAI Batch-2 Manager: The class materials have been uploaded already.
[00:29:02] GenAI Batch-2 Manager: Cool.
[00:29:02] Deepak Katara: We don't… if you could short the link.
[00:29:05] GenAI Batch-2 Manager: Okay, let me check.
[00:29:18] PRATIK: Okay, and you guys want to wait till the class materials are there, or is it good if I begin? Because we are… we have limited time, right? So, whichever suits you.
[00:29:30] Deepak Katara: I think for me, it's reflecting right now, so…
[00:29:33] Deepak Katara: I think others can also try this.
[00:29:35] PRATIK: Oh, it's reflective, right? So, are they good to begin, or should I wait?
[00:29:39] Sacheen Adavinavar: Is it the 3rd May, lab material?
[00:29:42] Deepak Katara: 66-2026. So, in Module 4, you'll see it in the last.
[00:29:50] Deepak Katara: I can share… I'm not getting the option to share a message to public.
[00:29:58] Deepak Katara: Boom.
[00:30:03] PRATIK: Thank you, guys.
[00:30:03] GenAI Batch-2 Manager: Hello?
[00:30:04] PRATIK: districts.
[00:30:05] GenAI Batch-2 Manager: Can you check under the Module 4?
[00:30:10] Deepak Katara: I, I, yeah, I can see it now.
[00:30:12] GenAI Batch-2 Manager: Okay, okay.
[00:30:17] Deepak Katara: Can you share…
[00:30:19] Neeraj Kumar: Oh, yes.
[00:30:20] Deepak Katara: Share the link.
[00:30:22] Neeraj Kumar: Please share the link.
[00:30:28] PRATIK: Unfortunately, I don't have access to that one.
[00:30:33] PRATIK: But, like, I guess, for Deepak, it was there, right? Deepak that are, like,
[00:30:39] Deepak Katara: Yeah, I could see it now.
[00:30:41] PRATIK: So, guys, if it is there for him, then most probably it should be there for you, because it's the same problem.
[00:30:46] Nirav Mehta: Module 4 in my LMS.
[00:30:48] PRATIK: Can you rephrase it maybe once?
[00:30:51] Deepak Katara: Yeah, you have to scroll to the last,
[00:30:58] Deepak Katara: In… under Official Material, tab.
[00:31:15] PRATIK: Is everyone open?
[00:31:17] PRATIK: We'll begin with.
[00:31:21] Shivansh Sharma: Yes.
[00:31:24] PRATIK: So, as I was telling.
[00:31:27] PRATIK: Like, last time, we had worked on, BERT, right? And the problem we dealt there was with classification problem.
[00:31:37] PRATIK: So, this time… this time around, we'll be dealing with squad dataset. It's basically… the problem, basically, if you see, it's a question-answering problem.
[00:31:47] PRATIK: Okay.
[00:31:49] PRATIK: So, we'll try with, like, we are basically trying to have different problem statements, right, as we are dealing with different elements, so that,
[00:32:01] PRATIK: Everyone can have a, you know…
[00:32:03] PRATIK: taste of, like, what the different problem statements are there in the NLP domain.
[00:32:10] PRATIK: Okay, so a brief on Squad dataset, quickly. Squad dataset is basically, like, the full form is standard, sorry, Stanford question answering dataset, right? And, like, each item gives the modal passage of text, right? Question about that passes, and the answer tool.
[00:32:28] PRATIK: Now, the difference here, like, from the previous tutorials, we are… we do not have specifically CSV file, right?
[00:32:38] PRATIK: for the data set. We are directly loading it, right, from Hugging Face, because it's a very popular dataset, and it's readily available there, so we'll be importing it from there.
[00:32:49] PRATIK: In tomorrow's tutorial, I'll give a little bit more brief about this low dataset, right? So think of it like, instead of, getting the files locally, we are directly getting it from the hugging Face, right? So, that's the difference.
[00:33:05] PRATIK: Now, this dataset was basically like, it was built from a set of Wikipedia articles.
[00:33:12] PRATIK: Right? The context was gathered from there, and the questions and the answers are basically written by human crowd workers, right? Many people have contributed to it.
[00:33:21] PRATIK: Okay?
[00:33:22] PRATIK: And, like, this is a very good fit for a beginner tutorial, because
[00:33:27] PRATIK: It's mostly clean and very documented. It's very realistic, because it has been taken from Wikipedia types, and I'm sure most of us have used Wikipedia.
[00:33:38] PRATIK: Right? And, like, it is widely used, right? And if you go through the research papers, like, SPOT is one of the… and if the… if the problem statement is around question answering, like, SPOT is very heavily used, for benchmarking.
[00:33:55] PRATIK: Right? The version that we are using here for Squad dataset is version 1.1, right? And just to give a quick example of, like, how the data, how one row would look is, like, first of all, we would have one ID.
[00:34:08] PRATIK: a unique identifier, for the example, then you would have title, the title of the Wikipedia article the passage is coming from.
[00:34:17] PRATIK: Then the context. Context is basically, it has the passage, or, like, suppose if you open a Wikipedia, it's article, right? So, we are basically taking paragraphs of that, and those are passed as context, okay? And based on that context, there are some questions.
[00:34:34] PRATIK: That, cloud workers have taken, and answers to it.
[00:34:39] PRATIK: Now, answer, if we take a closer look.
[00:34:43] PRATIK: Basically, it has, it is a small dictionary with two parallel lists. One is the text, right, the answer string, and one is the, like, answer start.
[00:34:54] PRATIK: Basically, this basically tells that where in the, you know, context, the, that ID, where the, or the character, where the answer starts.
[00:35:05] PRATIK: Right? So, in many of the models that were trained on Squad dataset, it's like, we have to find out the start position, where the answer starts, and the end position, where the answer ends.
[00:35:18] PRATIK: But here, we will be working a little bit differently, and I'll discuss it with her.
[00:35:23] PRATIK: Okay.
[00:35:24] PRATIK: Here, we'll be more, more… instead of predicting the tokens where the answer starts, we'll be more focused on predicting the… generating the answer. Okay, so that is the difference, and the details will follow up later on.
[00:35:38] PRATIK: Now, one more difference is that, like, in the training split, for, every, in the, all the training split, for all the rows, we have exactly one answer, right? But in the validation split, there are… there can be many, several answers, written by different auditors, right? And they are all correct.
[00:35:58] PRATIK: So I'll explain you how this will be working, when we'll be dealing with the evaluation and work.
[00:36:03] PRATIK: Okay.
[00:36:05] PRATIK: So, basically, Squad was originally designed for extractive question answering, right? And, where, like, just like I told before, it, we had to predict thestart and end of the answer inside the passage, or the context, okay?
[00:36:19] PRATIK: Now, we'll be doing something different. So, this is how it is. The context, the passive will be there, and then it is portion, right? Then, the portion will be there, like, like additional way, right? An answer, we have to find out now.
[00:36:32] PRATIK: well predicted.
[00:36:34] PRATIK: Okay.
[00:36:35] PRATIK: And here, the model's job will be basically to continue with text.
[00:36:38] PRATIK: Get down, sir.
[00:36:40] PRATIK: Okay, suppose context is given, question is given, then the gen… The model has to submit the answer.
[00:36:46] PRATIK: Okay. Now, traditionally, Squad dataset that we are importing from Hugging Fest, it has, training
[00:36:55] PRATIK: It has training split, right, and a validation split.
[00:36:59] PRATIK: Right? We are, breaking that, validation split.
[00:37:04] PRATIK: Right? And, we are, like, here, out of this 2 split, we are, converting, we are taking 3 splits, okay? We are, from those splits, we are also taking out a split.
[00:37:16] PRATIK: So that we have something that is completely untouched. And validation is basically, we'll use it while training only, just to see where the…
[00:37:25] PRATIK: You know, while training where the validation error was less, so that we can choose the best epoch.
[00:37:31] PRATIK: Right? Later double.
[00:37:34] PRATIK: Now, also, one thing to keep in mind is, answer of score dataset are very small. Like, this works in our favor because, because of the evaluation metrics also we are using. I'll discuss that below.
[00:37:47] PRATIK: Mostly, we are, in the evaluation, we are using, exact match and F1 store, okay? And I'll discuss those below, okay? And,
[00:37:56] PRATIK: One good thing is that the answer is always present in the passage.
[00:38:01] PRATIK: So, like, that helps, because sometimes if the answer is not present, even that can be a problem, right? But, like, in this dataset, answers are always present. So, models should be able to find the answer and not invoke it.
[00:38:14] PRATIK: Okay.
[00:38:16] PRATIK: So I think everyone is clear, right, on this part.
[00:38:20] PRATIK: So…
[00:38:21] PRATIK: Okay? Okay. So, let's move on. Just to confirm now, I'm beginning with this cool part. Is my screen visible to all? Is the size okay?
[00:38:38] PRATIK: Okay, I didn't take that as a yes. Okay, so…
[00:38:43] PRATIK: In the first part, in the first cell, as usual, we will be selecting the device, what we are working with. Since we are in Google Colab, we'll be selecting CPU here, right?
[00:38:54] PRATIK: So, the device, throughout the tutorial will be selected as CPU. It is set to CPU.
[00:39:00] PRATIK: Now, we'll be, installing some of the libraries that we'll be needing for the,
[00:39:07] PRATIK: this problem statement. One is transformers, data sets, accelerate, matplotlib, bundles.
[00:39:15] PRATIK: Okay, we'll, duplicate install for all of this.
[00:39:18] PRATIK: The book will play around this.
[00:39:27] PRATIK: Okay, installation is complete. Now, based on this installation, we'll be also using imports, right? So, we'll be majorly importing Torch, transformers, datasets, pandas as PD, right? Metro Clip.
[00:39:41] PRATIK: Now, let me quickly, tell you what all these are. Like, Torch, I think by now we are familiar. It's, basically for the PyTorch, and, like, we are be… we'll be using this for a deep learning framework, right? That will be used for GPT model and, model and training.
[00:39:58] PRATIK: Okay? Then the transformers for, importing that, model, tokenizer, right? And the trainer. Dataset, for, loading the squad dataset, okay? And Pandas, is… Pandas and NumPy are something that we are using regularly, right? NumPy is also used below.
[00:40:17] PRATIK: Like, it's there in the specific zone.
[00:40:20] PRATIK: So this one, Pandas, we are using basically for the EDA section, and Matplotlift for the, these graphs and all.
[00:40:27] PRATIK: Okay, so these are just the print statements to verify the versions and all. Okay, and then this, like, CUDA available, and all this is just to check if we have selected the device properly or not. Better to have error here than to have error at the
[00:40:43] PRATIK: When we are training, right? So we'll simply run this.
[00:40:53] PRATIK: Okay, so it looks like, in the import was done, nicely, okay? And…
[00:41:00] PRATIK: since we are using CPU, it is false, put our value is false, and we are working on SQL, which is correctly reflected. We are good to go.
[00:41:09] PRATIK: Now, some of the configuration that we'll be working on with the dataset, we have defined this very early. Okay, so…
[00:41:17] PRATIK: Like, due to the resource, constraint and the time constraint, we will not be using
[00:41:24] PRATIK: like,
[00:41:25] PRATIK: the full dataset here, although I have, worked with the full dataset, like, and I can confirm you that it takes a lot of time, because this dataset is huge, okay? And,
[00:41:39] PRATIK: It takes quite some time and memory to train it.
[00:41:43] PRATIK: Right? So we'll be working on a sample of
[00:41:46] PRATIK: Of the dataset. And, like, if you guys, just in case, if you guys want to, use it to, you know, like, train it on full dataset, you can just set it as true.
[00:42:00] PRATIK: Okay, so for right now, till the eta part, I'll set it as true.
[00:42:05] PRATIK: Okay, and, like, and after the EDA part is there, I'll quickly toggle it as false, and then rerun it, and then show you the tools and all, okay? The training part, I'll continue from there.
[00:42:17] PRATIK: So, once I toggle, it falls. The active will be, like, for training, but we are… we'll just be taking 20 samples, right? And for the…
[00:42:27] PRATIK: you know, this validation test split, right? We are taking, this validation, we'll be only taking 10, and from here itself, there will be validation and test split, okay?
[00:42:39] PRATIK: For that, we are taking total time.
[00:42:42] PRATIK: Okay, so the number of equals… for the full run, I took 3, okay?
[00:42:47] PRATIK: But here, I have only taken two, because it was taking a lot of time, and if I wait for it to train, I would not be able to complete two touring.
[00:42:57] PRATIK: Okay. Now, from this…
[00:42:59] PRATIK: validation subsample size, we are taking 0.7 as, we are… we have taken 0.7, right? So, 70% will be test, and 30% will be validation.
[00:43:10] PRATIK: Now, to have, every time the same samples,
[00:43:15] PRATIK: Right? So that even in random sub-sample, we have the same examples, those… just so that
[00:43:22] PRATIK: We have the same examples again while sowing, because if we have different results every time.
[00:43:28] PRATIK: So maybe, like, for example, if I had different examples here, right, it would be very difficult for me to, write any inference or draw any inference, right? Or to show it, because one time I may have different… little bit different reasons, and one time I may have…
[00:43:42] PRATIK: little different. So, for that, I'm setting seed as 42, okay? The model that we'll be using is GPT-2. It's a small model.
[00:43:51] PRATIK: of 124 million parameters, and I have told, like, more about this. There's a small marker, we can discuss it there.
[00:43:59] PRATIK: Now for, like…
[00:44:01] PRATIK: for every LLN, we have to give some prompt, right? So here, our prompt is in the format, like, first, the context. Context is basically the fastest thing that we had discussed earlier, and then the question part.
[00:44:13] PRATIK: whatever question we are passing. So these two will be passed out to the element, okay? This is the one template that we'll be reusing identically.
[00:44:23] PRATIK: Right? So, here, we have done two parts, right? First, we have taken
[00:44:30] PRATIK: like, GPT, where simply we start using whatever trend rates it has already… it was already initialized with that. Like, we'll start with that, we'll try to get some results, and then we'll try to do some fine-tuning and just see if fine-tuning really helps or not. And if it helps, like, how better were the results.
[00:44:49] PRATIK: So, that's what we aredoing here.
[00:44:52] PRATIK: So let me quickly run this.
[00:44:55] PRATIK: These are all queries checked on. So right now, we are using the full data set. Okay. Test fraction is 0.7 we have taken. Seed value is 42, and model name is GP2.
[00:45:06] PRATIK: Prompt template is in the format, context, then a question, then the answer, you have to get it.
[00:45:13] PRATIK: Okay.
[00:45:15] PRATIK: So, yeah, just as I said, full dataset is 2, then we are training on the whole dataset. Three epochs we had used for all the, like, for the main training on the whole
[00:45:28] PRATIK: data set, right? And, from the validation split, we took 0. I mean, 70% as test split and 30% for validation split, right? For validation purpose. Suite, we fixed as 42, right? Model, we took as GPT, which is, what, 24 million.
[00:45:43] PRATIK: Okay, and we define the target, though.
[00:45:46] PRATIK: Okay. Now, here, what we are doing is, we are loading the dataset.
[00:45:52] PRATIK: squad dataset, that was. So, let's see how we have done. Previously, we used to use pandas, right? So, it's a little different here. Let's… let's quickly have a look.
[00:46:02] PRATIK: So from datasets, we are importing… sorry, from,
[00:46:06] PRATIK: From datasets, we are importing load dataset, right? And then… Like…
[00:46:12] PRATIK: Here, with the low dataset, we are falling squad, right? And, keeping, in an object called SWOT, right?
[00:46:21] PRATIK: Okay, then, this has, train and validation split, right? Sorry, train and validation. It's like, it's stored, and I'll show you below how it is there. So, from there, we are calling for train, and from, again, and from in full valid… full spot validation, we are keeping the validation file of the dataset, okay?
[00:46:41] PRATIK: Now, if you use the full data set, then the, like, all the rows for the training will be saved in training data, and all the
[00:46:49] PRATIK: that, you know, validation will be kept in.
[00:46:54] PRATIK: squad validation. But since we are using samples, right, whatever we have defined before, so basically, we'll be taking… so for… in the else part, we are defining it, so that we don't have to, you know, write it again and again. We just have to toggle one button here.
[00:47:11] PRATIK: That is true or false.
[00:47:13] PRATIK: Right? And then we can quickly serve. So this was done so that, like, you know, when you guys try it at home, like, it is easier for you guys also to experiment with.
[00:47:24] PRATIK: Okay, so quickly here.
[00:47:27] PRATIK: So, here we are basically taking, like.
[00:47:31] PRATIK: how many subsamples we are considering, right? And we have taken minimum of the subsample size that we have defined above, right? And, full train. Whichever is minimum, we will take that, right? And…
[00:47:45] PRATIK: For, like, After that is done.
[00:47:49] PRATIK: here, the number is set, right? And now, here, basically, we are suffering, because if we don't do this, what'll happen is we'll take the, let's say, first 10 rows only, right? And that may have one kind of
[00:48:02] PRATIK: like, data, similar kind of data, right? And the model may not be able to run properly, and when it goes to test, like, it may not be able to perform at all.
[00:48:11] PRATIK: Right? So, for that, what we have done is, we select the number of rows that we defined, right? And then with the seed value, we are suffering the training data, right? And then selecting it.
[00:48:24] PRATIK: Similarly, we do for, validation also. Here, again, we select how many rows we need for validation, right? And then with the seed value, we are shuffling the validation data and then taking it, the number of rows, how many we require.
[00:48:40] PRATIK: Okay? Now, this is a print statement of how many it was taken, like, what was the speed and all.
[00:48:45] PRATIK: Okay? And here… Basically, We are splitting our validation data into validation data, into our plus test data.
[00:48:56] PRATIK: As I told you before, the spot dataset that we imported from Hugging Face has had only two parts, training, training set, and validation set. From the validation set, we are again splitting it, for,
[00:49:10] PRATIK: Validation and testing, okay?
[00:49:14] PRATIK: You could have done this with, training as well.
[00:49:19] PRATIK: But,
[00:49:20] PRATIK: the split of the validation could have taken as… from training also, but here I found that there were substantial number of, examples in, validation also. So, from the validation split only, I took, took the…
[00:49:33] PRATIK: You know, validation set, and the tested set.
[00:49:37] PRATIK: Okay.
[00:49:38] PRATIK: Now, let me quickly see, like, what all we have.
[00:49:44] PRATIK: Maybe this… It's difficult.
[00:49:47] PRATIK: This is just a small warning, because we are not using any token, right?
[00:49:52] PRATIK: So, if you are using Hugging Face, and if you are using any proprietary model, like, we usually require a token, right? Hugging Face token. But, since this is publicly available data set, we don't really require, so this is just,
[00:50:07] PRATIK: one, okay? Now, as I was telling you.
[00:50:10] PRATIK: the data, like, the raw data structure that we get is something like this. In the data set tick, right, we have training, right, and then we have variation.
[00:50:21] PRATIK: like, the features that we had discussed above is, like, ID, which has ID, title, context, question, answers. Similarly for validation data, it has ID,
[00:50:32] PRATIK: title, context, question, and answers. So, the number of rows in the original data set was 87,599 for the training one, and for the validation split, it was…
[00:50:45] PRATIK: 10,5… sorry, 10,517. Now, from this set, again, we are splitting, right, from the validation itself.
[00:50:53] PRATIK: We are splitting into validation data and test data rows, okay? So, since we took 70% as test and 50% as validation, this is the num… for validation data rows, 3,171 is taken, and for test data rows, 7,399 was taken.
[00:51:11] PRATIK: And, with both of the, like, all three having…
[00:51:16] PRATIK: ID, title, context, portion answers, as clear, does anyone have any doubt?
[00:51:27] Neeraj Kumar: Hello.
[00:51:28] PRATIK: Yep.
[00:51:29] Neeraj Kumar: Yeah, I have a query. While we are doing just, like, test, validation on validation data, we are just, again, we are splitting it into test and,
[00:51:38] PRATIK: Yeah, so, you guys, if you remember, even in the last tutorial, we did something very similar to this.
[00:51:45] PRATIK: But there, the difference was that we had taken the validation split from the training data itself.
[00:51:52] Neeraj Kumar: Yes. Great.
[00:51:53] PRATIK: Here, I, took it from the test, this valid test, because here, it's confusing a little bit because of the name.
[00:52:00] PRATIK: Right? Because usually we have train and test. Here, the name is given as validation. And from the validation only, I took test and validation data.
[00:52:09] PRATIK: Right? And the reason I did this for this particular case was because, you see, for validation itself, it has substantial number of, you know, examples there.
[00:52:19] PRATIK: So, if we take just 30% of it from there itself, it's, like, still, we have 7,000… approximately 7,400 samples, right?
[00:52:29] PRATIK: So, it was just a design choice. You could have taken this from training set also. We could have done that.
[00:52:36] PRATIK: Right? It was just my design. It was just my design choice.
[00:52:40] Neeraj Kumar: Yes, thank you.
[00:52:41] PRATIK: Now, just for a quick summary,
[00:52:46] PRATIK: like, it has two original splits, train and validation, hence the confusion, right? Mostly, we have train split and test split. Here, it is written as a validation split, and from this validation split, we are again splitting into validation, this
[00:53:00] PRATIK: Validation rule and the, sorry, validation and test rows, okay?
[00:53:06] PRATIK: And all of them have these, features, ID, title, context, person, answers, okay?
[00:53:12] PRATIK: And, when we set it as true, full dataset is taken. And for validation split, we have taken 30%. Fortest split, we have taken 70% of the… from the validation.
[00:53:22] PRATIK: This is showing here also. Originally, it has 10,570, and after the split, like.
[00:53:29] PRATIK: We have 3,171 in validation data rows, and 7,399 in test data rows.
[00:53:39] PRATIK: Okay. So basically, we are using this validation, data to choose the best repo. When we'll be doing, fine-tuning later down in the fine-tuning section, we'll be choosing the best
[00:53:53] PRATIK: base report, where the… from the validation data, we had the least validation blocks. So, this would be useful.
[00:54:02] PRATIK: Okay.
[00:54:04] PRATIK: So now we'll try to look at the data, how it, how it is looking, right?
[00:54:12] PRATIK: We'll try to look at 3 rows, so we have taken 3 examples, right? And from there, from trend data, like, we have called a loop, right?
[00:54:22] PRATIK: And, for, like, it'll be running 3 times, and from there, we'll be seeing the title, context, for sponsor, right?
[00:54:29] PRATIK: Those ID we are not specifically using, right? We'll just, see title, contest portion answers, and later down in the training, we are not even using title, we are just okay with context portion and answer. Even this was enough for us.
[00:54:45] PRATIK: So… Quick examples.
[00:54:48] PRATIK: Okay, so the title, like, one of the examples was University of Notre Dame, right? And context was this, right? Architecturally, the school has a Catholic character, and so on, right? It's a long context.
[00:55:02] PRATIK: Let's see. See? It's very long.
[00:55:06] PRATIK: Okay.
[00:55:07] PRATIK: And from there, like, a question has been asked.
[00:55:11] PRATIK: To whom did the, to whom did the Virgin Mary allegedly appear in 1858 in,
[00:55:16] PRATIK: to raise funds, or then, like, if I pronouncing it wrong, right? And from there, we have this as answered.
[00:55:22] PRATIK: Right?
[00:55:24] PRATIK: And also, like,
[00:55:26] PRATIK: like, in all the examples, we also have this. In answers token, as we discussed, right, we have answer as well as the character index, where the answer is actually there. So, we are not exactly into this. We are just looking at the dataset right now. But in our, like, exact
[00:55:45] PRATIK: Like, in this tutorial, we are not losing this. We are focused more on generating it, rather than predicting where the
[00:55:53] PRATIK: You know, the character, where the, answer actually starts from and ends at.
[00:56:00] PRATIK: Okay, so similarly, we have another example, okay?
[00:56:04] PRATIK: Now, you… you can also notice that, like, since it was pulled
[00:56:09] PRATIK: Right? Sequentially, what has happened is, The context is same.
[00:56:14] PRATIK: Right? Here also, even the title is the same, but the questions are a little different.
[00:56:18] PRATIK: Right? And the answer is also detail different. And, even the answer where, where the, like, character index, where the answer is, there's also different, right?
[00:56:29] PRATIK: Similarly for this one, this is also from the same title, same context, the question is different, and the answer is different.
[00:56:38] PRATIK: Okay. So this is how each dataset is looked, so each row in the dataset is…
[00:56:44] PRATIK: After giving the same as I do.
[00:56:47] PRATIK: So, to give you a brief summary, the row here is in dictionary format, right? Answers basically have the main answer and the character where the answer starts, right?
[00:57:00] PRATIK: Right? And, if you notice, the answers are very solid, right? Which is…
[00:57:06] PRATIK: nice for us, okay? It helps, you know, at the end.
[00:57:12] PRATIK: Okay, so we'll quickly do some EDAP.
[00:57:15] PRATIK: Right? And… let's see.
[00:57:24] PRATIK: So, first we'll, try to see, like, how many rows and all are there, in train and validation, right?
[00:57:33] PRATIK: And then we'll try to applaud a figure for it, so…
[00:57:40] PRATIK: So, if you see, in terms of bar graph, it's very easy progression. 3,171 for validation.
[00:57:47] PRATIK: This will be used along with training only, and testing is altogether different.
[00:57:52] PRATIK: Some 7,400 samples, I think.
[00:57:55] PRATIK: So, this is verified. Here, also, graph is also showing accurately. Here, it's, on the left-hand side, we have number of examples. On the right-hand side, we have the splits.
[00:58:07] PRATIK: So, like, as we have already seen, training has 87,599, and validation has 300… Sorry, 3,171 words.
[00:58:15] PRATIK: In the total capacity.
[00:58:17] PRATIK: now we'll try to see the length distribution of context, questions, and answers. Now, like, here we are doing with words, just to have a, like, understanding, but, like, later down, we'll be using… instead of words, we'll be using tokens, right?
[00:58:35] PRATIK: Let's say… this is because a single word, like, let's say, unpredictable, can… can be three…
[00:58:42] PRATIK: different tokens, right? So…
[00:58:45] PRATIK: This is just for, like, exploratory purpose, okay? Last time, someone had asked me a question on this. So, we are just looking at it, like, how the dataset is. We'll try to see the mean, medium, right? How it is for, the context questions and answers.
[00:59:00] PRATIK: Very quickly, sorry to put it up.
[00:59:04] PRATIK: Now for, like, each row in the training data, we are trying to…
[00:59:09] PRATIK: Get the length of context, questions, and answers, okay?
[00:59:15] PRATIK: Here, The text, and then in bracket 0 is written, because, like, there are two parts, right?
[00:59:22] PRATIK: The main answer, and the… in the second part, it is written that, where the…
[00:59:29] PRATIK: Character where the… sorry, character index where the answer starts.
[00:59:33] PRATIK: So here we are calling the zero index for this one.
[00:59:37] PRATIK: So from there, we are getting context lens, question lens, answer ones, right? And from there.
[00:59:43] PRATIK: Okay, now we'll try to quick, try to see some stats on it, like, mean, median, max, and minimum words. In, context, question, and answers.
[00:59:56] PRATIK: And then we are drawing a plot for it.
[01:00:01] PRATIK: People couldn't listen.
[01:00:16] PRATIK: Mr. C.
[01:00:17] PRATIK: Regardless of your service type.
[01:00:26] PRATIK: That's right, yeah, I forgot. I'm doing this for the whole data, so that's why it's done.
[01:00:30] PRATIK: Okay, now let's see.
[01:00:33] PRATIK: The median for context is around 110 words, mean is around 119, minimum 20 words are there in context, and maximum is 653 words.
[01:00:42] PRATIK: Okay. Question… in question, we have minimum 10 words, mean is also near to 10 only, minimum is 1 word, and maximum is 40 words.
[01:00:51] PRATIK: Answer, if you see, like, like we said before, we had also told this in the introductory part, introductory part, like, where the… we have told that answer is very short, right?
[01:00:59] PRATIK: One to… one or two words, mostly.
[01:01:02] PRATIK: So, median is 2, mean is also 3 words, minimum is 1, and maximum is 43.
[01:01:09] PRATIK: Right? So this is, like, an outlier, right? But if you see… if you look at mean and median, this will give you a brief idea, like, you know.
[01:01:17] PRATIK: Their answers are mostly 3, like, two to three words.
[01:01:25] PRATIK: Now, these are some of the… histogram, okay?
[01:01:29] PRATIK: Africa and Syria, also contextualize.
[01:01:32] PRATIK: Mostly, like, they are here, right?
[01:01:37] PRATIK: It's more, the peak is there around 110, which is also, we are seeing there.
[01:01:42] PRATIK: Right, similarity for words also, we are saying it's… the peak here is around 10, right? So, which is upgrading, right?
[01:01:49] PRATIK: Right? And for answers, the peak is near this 2 to 3 apart.
[01:01:54] PRATIK: 001, and 21, so somewhere around here. It decreases. So, histograms are also reflecting that only.
[01:02:03] PRATIK: Okay.
[01:02:04] PRATIK: Now, to give you a brief summary, contexts are very long, right? The median, context is around 110 words.
[01:02:10] PRATIK: Great.
[01:02:12] PRATIK: And, questions are medium and consistent, right? So questions are around 10 words, right?
[01:02:19] PRATIK: And, answers are very short, 2 to 3 words, like we said, and the longest is 43, which will be, like, an outlier.
[01:02:29] PRATIK: now, like, we'll try to see in questions what kind of questions are there exactly, right? So, there are what questions more, or the questions having who…
[01:02:41] PRATIK: as the first word more, when, where? So these are the questions, like, these are the words that will start a push move, right? What, who, when, where, why, how, which?
[01:02:51] PRATIK: Right? So we'll try to see, like, how many of the questions are actually starting with this, okay? Just, for informational basis, right?
[01:03:03] PRATIK: Okay. So, what we are doing for this is, we are, from collections, we are importing counter. Matplotlink was already important, but again, we have imported it here.
[01:03:11] PRATIK: Okay, now, I have created a list, right, where I have, kept all the words from where, like, user question starts, right? Then I have, like, we have created one, this function, right? Where we are trying to see if the first word, is here.
[01:03:30] PRATIK: Right, belongs to this list.
[01:03:34] PRATIK: Mojo, you want to say something? Sorry.
[01:03:44] PRATIK: Nope.
[01:03:49] PRATIK: Let me switch on his life was taken.
[01:03:52] PRATIK: Okay. So, like, we'll try to see these words, like, in how many of them are there. Sorry, I was here, okay. So, I'll try to, like, we have created one function, right? And there we'll try to, first of all, we'll, lowercase it, and then split it, right? And we'll check if the tokens
[01:04:11] PRATIK: If… if the token is not in this, we return it as other, right? And if, like, and for the first one, we are initializing with 0, right? So we are again checking.
[01:04:25] PRATIK: If, like, if it is there, if the first token belongs to this, right, then we'll keep it in first, else we'll keep it as other.
[01:04:37] PRATIK: Okay, so then we are trying to have,
[01:04:41] PRATIK: Now, we have created this function, now we'll try to apply this for the training data. So that's what we have done. For row in training data, first word category? This is the function that we have created, and we have the row portion.
[01:04:56] PRATIK: From the, like, we have kept the portion going. Sorry, portion column.
[01:05:02] PRATIK: Then, from this counter, we are trying to have counts.
[01:05:06] PRATIK: Okay, so, like, of what… how many watts are there, how many hulls are there, and all.
[01:05:10] PRATIK: Nope, and… And, like, after that, what we have done, we are trying to…
[01:05:16] PRATIK: order our, we are trying to take one graph below, and just showing it down, right? So, first we'll have the WH words, and then we'll have the other parts, right?
[01:05:27] PRATIK: the word… the question that does not start with all these words, okay? And then…
[01:05:33] PRATIK: like, we'll try to fetch, the values of it, okay? So, from all the…
[01:05:40] PRATIK: ordered labels, we'll try for label, for each level, we'll try to find. And this dot label
[01:05:47] PRATIK: Comma 0 is basically telling, if it is, say, if the count was 0, right, it did not… if it did not get anything, right, then it will simply take 0.
[01:05:58] PRATIK: default value will be 0, okay? So this is to avoid error.
[01:06:02] PRATIK: And then we'll plot the, graph. Let's see.
[01:06:15] PRATIK: So, we can see that what type, the questions that start with what is very dominating, right? There are around 37,593 questions that start.
[01:06:28] PRATIK: And who, when, where, why, how, which are, like.
[01:06:32] PRATIK: in decent numbers, and there are also questions that do not start with all the WHs, and those are populated in the other category.
[01:06:41] PRATIK: Okay. Now, just to have a quick summary of what we discussed, what, what questions are dominating? Like, all the rest of the WH questions, right?
[01:06:53] PRATIK: are in the middle tier, and if you notice, Y type of questions are the least, are the least ones, right? Because you usually have to explain, right, for,
[01:07:05] PRATIK: more wide type of portions, right? And if you have noticed that,
[01:07:09] PRATIK: Like, in our data set, answers are very short.
[01:07:13] PRATIK: So, this is also, like, this… there is also a good inference for this one. So, Y is also less.
[01:07:20] PRATIK: Open.
[01:07:24] PRATIK: So…
[01:07:26] PRATIK: Now, let's go to this one. This is very optional, you can do this, you cannot do it, it's not exactly helping anywhere, right? But it's just to have a good idea, like, what kind of words are being used, right, in the questions.
[01:07:40] PRATIK: Right? So what I have done is, like, for here, we have done, we have plotted two graphs. One graph, including the stop words, right? And the other graph, excluding the stop words, just to see how it is working. Okay, so here again, we are importing the…
[01:07:58] PRATIK: countered.
[01:08:00] PRATIK: From collections, you still have, to give additional, right? Where, for each of the words, we'll have columns and one.
[01:08:07] PRATIK: at the end, that will help us load the graph. So, we have defined some of the stop words, and in all of the data for, like, we have created one list in all of worlds, like, named as all worlds, and from, each row in training data, first we,
[01:08:23] PRATIK: lowercase it, split it, right? We drop the digits, punctuation, and basically, we only take, the characters which are, like.
[01:08:33] PRATIK: Alpha… alphanumeric, right?
[01:08:37] PRATIK: Or, sorry, alphabets, right? And then, if it is clean.
[01:08:41] PRATIK: So, and we keep it in clean, right? And if it is clean, then we append it in.
[01:08:47] PRATIK: onwards.
[01:08:48] PRATIK: So, from here, we try to… after that, what we do is, we do, use this counter to store the raw counts, right? And then again, we, like, we…
[01:08:58] PRATIK: we decided, right, first of all, we'll have the raw counts, and in the second one, in the second graph, we'll have the flitted counts. Flitter counts will be basically all the words without this raw counts, like, how it is. So we'll try to see the top 15 words, right? So, that's why we have to take, I mean, top underscore N is 15, right?
[01:09:16] PRATIK: And then…
[01:09:17] PRATIK: will, will… with this DOS, the dot most common, we'll be selecting the top 15.
[01:09:25] PRATIK: Right? Raw counts, and similarly, from the filter counts also, we'll be doing the same, right? And now, let's plot… this is just for plotting them.
[01:09:35] PRATIK: Hold on this.
[01:09:46] PRATIK: So, if you see, when we did not remove, right,
[01:09:51] PRATIK: Like, this one. Stop words. Most of them were swap words only, the county still.
[01:09:56] PRATIK: As soon as we remove the stop words, there are words like many, have first, name, type, used, right, as, new, most, city, doing. So there are now different words, apart from, and even we did not do valet, they were mostly stop words. So this kind of helped us to understand
[01:10:14] PRATIK: So, in the portions, like, we don't have many, like, how many…
[01:10:18] PRATIK: And on that kind of person. What year was it, kind of?
[01:10:23] PRATIK: When did it happen first? So, that kind of… so it picked up from there.
[01:10:30] PRATIK: So, like, a brief summary on this. A raw list is, mostly filler, like, of stop words and all, right? Filtered list is, more meaningful, right? And, many near the top inside the counting questions, right? How many, right? Which ask for numbers and all.
[01:10:48] PRATIK: And the other stop words are just for common things, like words like year, name, city, people, line up with the factual nature of support, right? How many people were there, which city was it, what was the name, kinds of questions.
[01:11:01] PRATIK: So, with this, we have completed the EDA section, right?
[01:11:05] PRATIK: And, till now, what we have done is, we have done, GPU selection, we discussed about, like, in brief, we discussed about zero-shift and, zero sort and fine-tuning, right?
[01:11:17] PRATIK: And we discussed about, like, how it has happened.
[01:11:20] PRATIK: Okay.
[01:11:21] PRATIK: And, key things we learned about was
[01:11:25] PRATIK: the, like, vortex was very long, portion was sort, and answer was even sorted, right? And spot dataset was mostly dominated by what?
[01:11:34] PRATIK: And answers are typically just few words, roughly speaking, two to three words, right? And, now, in, the next section, we'll be just, we'll be going on to,
[01:11:47] PRATIK: like, zero-sort question answering, right? And after that, we'll be going through, fine-tuning, right? Fine-tuning is basically, we are training it on our dataset, and then expecting the model to perform better, if not, worse, right? And then we will do up a comparison.
[01:12:05] PRATIK: Okay. So, till now, we had done, this full dataset.
[01:12:10] PRATIK: Right? But…
[01:12:12] PRATIK: like, as I've told, like, it will be very difficult to complete it. I don't know if installed it will be completed even by tomorrow, right? So let me quickly toggle it to false, and let me renun everything to here.
[01:12:26] PRATIK: Listen will take me at 2 minutes.
[01:12:28] PRATIK: This was just to show the data on,
[01:12:31] PRATIK: the actual data set, rather than just taking a sample, because 20 is very less, right? You would not get any idea.
[01:12:38] PRATIK: fungus, so that's why I did this.
[01:12:44] PRATIK: Okay, let me flip this as false.
[01:12:49] PRATIK: moment.
[01:12:53] PRATIK: Oops.
[01:12:56] PRATIK: You start session.
[01:12:58] PRATIK: Mr. Fleur of my partner now.
[01:13:02] PRATIK: Let's see…
[01:13:24] PRATIK: Just a second, please.
[01:13:25] PRATIK: Maybe 2 minutes.
[01:14:20] PRATIK: The training sample is shown here, why did I show you a different one has 3.
[01:14:28] PRATIK: Since now data is very nice, right? The graph is very sparse.
[01:14:38] PRATIK: It's still dominate over what questions, right? But these are close to zero, because we are picking very small sample size more.
[01:14:51] PRATIK: And these are, like…
[01:14:54] PRATIK: The most vulnerable work, then the stop, after removing stop work, these are the words.
[01:15:07] PRATIK: Oh, yeah.
[01:15:09] PRATIK: Sorry for the delay, but this was necessary because I really wanted you guys to see the data on actual dataset rather than the Senate code.
[01:15:16] PRATIK: Okay, now, let me quickly tell you more about… a little bit about the GP22 model. Like, you guys already know about it, right? This model was, this is, like, just a…
[01:15:28] PRATIK: I'll just touch it, because this has already been discussed with you, sir. So, GPT-2 was basically released by OpenAI, right? And it has been basically trained to, predict
[01:15:39] PRATIK: next word, given the sequence of her text, right? And they do this over and over, and, like, with this, they generate longer passages, one piece at a time.
[01:15:50] PRATIK: Now, it was, trained on large amount of text from the internet with no specific argument, just, like, next word prediction, more often. So, that's why it can produce fluent English, right?
[01:16:02] PRATIK: But, like, on its own, if you… you cannot expect it to do different tasks, like portion answering and all. So, we'll just see an example, right?
[01:16:12] PRATIK: of this.
[01:16:13] PRATIK: Now, we are using GPT-2.
[01:16:16] PRATIK: small, right? GPT-2, basically, it's a small version, it's 124 million parameters, it's small, right? But very capable, right? The larger versions would be a little slower and heavier, right? And, in our use case, GPT-2 small was working decently.
[01:16:34] PRATIK: Hence, GPT-2 small, like, was taken into consideration.
[01:16:39] PRATIK: Okay?
[01:16:40] PRATIK: For tokenizer, like, like any tokenizer, right? It breaks specs into units called tokens, right? And then each token is,
[01:16:49] PRATIK: map to an integer, that the model understands, right? And then, from the same ITs, we can convert it back to text, okay?
[01:16:56] PRATIK: Now, the tokenizer here, that we are using here is AutoTokenizer from pre-trend, model name, right? Model name with GPT Smart. So, we are, initializing our tokenizer from here, right? It's a byte pair… it's a tokenizer based on byte pair encoding, right?
[01:17:14] PRATIK: And one important thing that we have to take into consideration is
[01:17:21] PRATIK: Like, GP22 tokenizer has no padding token, right? And later on, we'll be needing, the padding is there. We are using some padding. So, what we are doing there is, like,
[01:17:33] PRATIK: Instead of, you know, like, having a separate pairing token, we are using the
[01:17:40] PRATIK: This end of text token.
[01:17:42] PRATIK: Right? And somehow, we are adapting that so that the model understands that after this, the sequence has entered. Right? We are using that again and again.
[01:17:54] PRATIK: So, like, this is a very common and safe choice, right, that we are doing it, like everyone usually does when working with GP2 and all.
[01:18:03] PRATIK: Okay. And now to give a quick example on what zero sort is. Zero sort is, basically.
[01:18:09] PRATIK: Zero sort means asking the modal to do tasks without giving it any training or examples for the task first.
[01:18:16] PRATIK: Zero refers to zero training examples, right?
[01:18:19] PRATIK: just like that, if we say about, one shot, that means one example is the one, right? So, similar to that. So…
[01:18:27] PRATIK: Here, what we are doing is,
[01:18:29] PRATIK: it has… this GPT was, not trained for, the squad.
[01:18:34] PRATIK: Right? So basically, it was generally… it was, trained to predict the next open.
[01:18:40] PRATIK: So I'm expecting, like, I was… when I was doing this, I was expecting it not to perform so good on the zero short, right? Because, it is just doing next-word prediction, and the answer, I mean, it does predict some, like, English, but it is not up to the mark, right? I'll just show you how it was happening.
[01:18:58] PRATIK: With the examples, also.
[01:19:02] PRATIK: So that's what I did, and we expect it to struggle, and that is fine. A weak zero-sort result is exactly what motivates fine-tuning in the next one. Okay? So, so what we do is, we import auto, model for CasualLM and auto tokenizer from Transformers.
[01:19:20] PRATIK: And in this cell, basically, we'll be loading the computer model and the tokenizer, right? So for… from… from transformers, we are importing these, right?
[01:19:29] PRATIK: And then, we are initializing our, tokenizer. Model name, we have already given ZPTool, right? And that, from that, auto tokenizer is named.
[01:19:39] PRATIK: based on our model, it is automatically detecting which model we are using, and then the tokenizer is being initialized. Similarly.
[01:19:48] PRATIK: For,
[01:19:51] PRATIK: Okay, before going there. Yeah, so as I said, for padding token also, we are using…
[01:19:59] PRATIK: since it does not have a dedicated padding token, right, we are using EOS token itself for this task, okay? And, like, for now, then, we are initializing our model also, that is done by auto model for casualLM.from pretend, and then we call model name, model name is already defined as GPT-2.
[01:20:19] PRATIK: This is the one block.
[01:20:21] PRATIK: Now… We are basically,
[01:20:25] PRATIK: doing this, tokenizer.pad, underscore token ID, and we are keeping, like, we are syncing this with modal config, so generation and the loss code can create the same ID, right?
[01:20:38] PRATIK: It basically tells the model which ID means padding, and keeping it consistent with the tokenizer. Then what we do, we move the model, to the chosen device, the GPU, in our case, it's CPU, right? Whatever we have chosen.
[01:20:53] PRATIK: Then…
[01:20:55] PRATIK: We, we put the model in evaluation mode, right? So when we put the model in evaluation mode, this will basically disable the trainingbehavior, like, dropout and giving
[01:21:07] PRATIK: And it will give the, stable and repeatable outputs.
[01:21:11] PRATIK: Now, just to see the number of parameters, how many, CP2 has, let's have a look.
[01:21:17] PRATIK: Review that was.
[01:21:22] PRATIK: So, like I said, like, it's around 1.24 million, right?
[01:21:27] PRATIK: And it is on CPU, and paired, and UAS token are taken as synth.
[01:21:33] PRATIK: An ID, like, if you… if you get to see some examples, this is the ID.
[01:21:39] PRATIK: Okay, just to give you a quick summary, GPT-2, till now, what we have done is GPT-2, small is loaded, right? I did it on GPU, so I wrote it in GPU. Actually, it's on CPU, right? And, padding, padding is said to be the,
[01:21:55] PRATIK: end of text token, because GPT does not specifically have that token, right? And that token is defined… that token ID is 50256.
[01:22:10] PRATIK: Any questions?
[01:22:19] PRATIK: Okay.
[01:22:21] PRATIK: Still here, does anyone have any questions?
[01:22:31] PRATIK: Okay, so I'll turn again.
[01:22:35] PRATIK: So, what we are doing here is, what we are doing here is, basically, we are creating an answer generation function. It's a reusable function we'll be using here itself, and later down, when we are fine-tuning it, they also will be using it.
[01:22:48] PRATIK: Okay, so for that, what we are doing is,
[01:22:51] PRATIK: we are, importing dots, right? And from there, we are creating, this function, generate unsett, right?
[01:22:59] PRATIK: Where context is… should be there, portions should be there, and maximum tokens is 30. This is, like, for the output, okay? Please don't get it confused,
[01:23:07] PRATIK: Okay, later on.
[01:23:09] PRATIK: So, like, this is for the output.
[01:23:12] PRATIK: Now, prompt, like we discussed before, it should be in, con… there should be… prompt, there should be a context question, and from this prompt, like, we are expected to… expected to have the answer again.
[01:23:25] PRATIK: Okay. Then…
[01:23:28] PRATIK: Like, after that, what we have is, like, in inputs, what we are doing is, we'll convert the, prompt text into token ID that the model understands, right, and return it, return it as, this
[01:23:42] PRATIK: PyTorot tensors, right? And then, move it to the device.
[01:23:48] PRATIK: Right? That we have selected above.
[01:23:50] PRATIK: And, So this…
[01:23:55] PRATIK: We are also… what we are doing is we are trying to remember the prompt length. Like, how many length
[01:24:01] PRATIK: like, what's the length of the form, so that we can later slide it up and keep only the newly generated answer, right? So what happens is, when you pass it on.
[01:24:10] PRATIK: when you pass the prompt to the LLM, sometimes what happens is, like, the…
[01:24:14] PRATIK: prompt is also echoed back, along with the answer. So with this, what we'll do is, we'll try to have the prompt length, and from the answer itself, we'll try to, you know, re-slice it off. Whatever the prompt length is, we'll slice the initial part from there, and just have the answer.
[01:24:31] PRATIK: update.
[01:24:32] PRATIK: Now for the sedation part.
[01:24:35] PRATIK: We'll start with TulschnowGrad. Toss no grad basically disables the gradient tracking, right? We are… since we are only doing… we are not training here, right? It's zero salt.
[01:24:44] PRATIK: So we give the tokenized prompt, right? Max new tokens is taken, right? Answers are very… answers are very short.
[01:24:54] PRATIK: Okay, and then do sample is across to false is basically we are taking,
[01:24:59] PRATIK: We are doing 3D generation, right? And,
[01:25:04] PRATIK: After that, we are, basically telling the model that, you know, use US inter-sequence as PAT token, right? This is for that.
[01:25:15] PRATIK: now, like, for, like, for generated IDs, we are only storing the output IDs, right? This is, like, from prompt length, we are slicing it off, right? From prompt length, wherever the prompt length, whatever prompt length, after that, whatever is there, it can be stored in the generated IDs, right? And,
[01:25:35] PRATIK: Then, whatever, generated IDs we got, right, based on that, we'll decode it back to the text, right? And then return the answer.
[01:25:44] PRATIK: Stripping the leading spaces and the, trailing spaces.
[01:25:48] PRATIK: Okay, so we'll just do a smoke test, right? If everything is working fine or not, we'll take the first row from the validation data, right, and try to generate answer based on it.
[01:26:00] PRATIK: Let's just see if everything is working now.
[01:26:16] PRATIK: Perfect.
[01:26:17] PRATIK: So, prompt integrated in groups, like, just we had discussed in cell 1 or cell 2 also, A for something like this. So, from there, we have first, we pass on the context, right, then we pass on the question, then we are expecting… we expect the answer to be generated by them.
[01:26:34] PRATIK: this GPT model. So, question would be…
[01:26:40] PRATIK: What does LGM stands for? The gold answer is basically telling, like, what the actual ground growth answer is. So, answer should be this.
[01:26:48] PRATIK: But right now, it is… You know, not accurate, it is.
[01:26:53] PRATIK: This is us doing a NEX for production, and it's not relevant at all.
[01:26:57] PRATIK: Okay.
[01:26:58] PRATIK: So this is… this is just one example we'll try to look at.
[01:27:01] PRATIK: other tests also, okay? So tutorial books, brief summary.
[01:27:07] PRATIK: the… like, what has happened is, we are using the prompt format, right? And we… and the cold answer is.
[01:27:15] PRATIK: like, short, right? Like we saw here, last, glacial mission, like, 3 volts. Very shortly. And,
[01:27:23] PRATIK: like, when generating the answer, the model is basically rambling, and it misses the actual answer, right? It just… it's just… if you look at it, it's, like, just doing the network prediction.
[01:27:34] PRATIK: Okay? But not getting to the point at all.
[01:27:37] PRATIK: Okay.
[01:27:38] PRATIK: So this is the behavior we expect from the model, right? Because it has been never trained for the question-answering task.
[01:27:45] PRATIK: And, the language, though we can notice that the language ability is… are still good, it's fluent only, but, for the task, it has not been able to generate the answer, which is very… which should be, like, ideally sought, right, and cleaned.
[01:28:03] PRATIK: So this whole gap will… is the whole motivation for fine-tailing.
[01:28:09] PRATIK: So… We'll do a quick… Like,
[01:28:16] PRATIK: like, test, from the test set also, we'll take few examples and just check, right? From there, we'll take 5 examples and just try to see how it is behaving, right? We'll have the context, question, answers, right? And, the golden, this gold prediction… sorry, gold truth. Sorry, ground truth also.
[01:28:35] PRATIK: So for that, what we have done is we have taken a number…
[01:28:39] PRATIK: number underscore demo as 5. We're basically choosing how many samples you want to look at, right? And from the taste data, we are selecting
[01:28:48] PRATIK: this, right? All the answers, zero-shot answers, will be stored in this,
[01:28:54] PRATIK: list, zero sort, zero underscore sort underscore answers, right? And then we loop over it.
[01:29:00] PRATIK: Right? For each of the, examples, we'll have context question and word answer.
[01:29:08] PRATIK: Right? And after that, we have the paid answer, using the function that we have created in the previous cell.
[01:29:15] PRATIK: Right? Then, whatever answer we get it, we'll append it into zero short answers, right? And then…
[01:29:22] PRATIK: For, since the contexts are very long, we'll only take the 160…
[01:29:28] PRATIK: Right? So that, it is small enough, and we can accurately see all the example settings.
[01:29:34] PRATIK: Now, there are some filling statements. Let's have a look.
[01:29:47] PRATIK: Make it load first.
[01:30:05] PRATIK: Okay, so if you see…
[01:30:08] PRATIK: Context review, we have, in India, private schools are called independent schools, and so on, right? Person is how many examination board exists in India? That should have been 30.
[01:30:18] PRATIK: Right? But if you look at the zero-sort answer.
[01:30:21] PRATIK: The number of examinations limited to the state level, the number of examinations linked… it's basically repeating the same thing again and again.
[01:30:28] PRATIK: Right? Now, let's see. Example one.
[01:30:32] PRATIK: here the context is the Writer Guild of America.
[01:30:36] PRATIK: The Writers Guild of America strike that halted production of network programs, and so on. Question is, who started rumors in 2008 that ABC would sell its 10 owned-operated stations, owned and operated stations? The cold answer would be Karish and who, right?
[01:30:53] PRATIK: GPT answer is the network is GPT is including, and it's, like, going on and on, where the answer, it just should have been this, or, like, it doesn't even have the answer file.
[01:31:03] PRATIK: Right?
[01:31:05] PRATIK: And look at the example, too, here. Here the context is private schooling in the United States has been limited, sorry, debated by educators, and so on.
[01:31:16] PRATIK: The question is, in what year did Massachusetts, Massachusetts, first require children to be educated in schools? The golden answer would be 1852, right? Now, in this case, we did have 1852, but after that, still, the generation continues, right? Because, this has not been trained for this past year.
[01:31:35] PRATIK: Example 3.
[01:31:38] PRATIK: Context, is CBS bought, CBS broadcast Super Bowl 50 in the US, and charged an average of $5 million, right, and so on. Answer is which network broadcast is the 50th Super Bowl game associated with CBS. Now, this time it has, CBS in the answer, but it is written CBS, NBC, and ABC.
[01:31:57] PRATIK: Right? Which network broadcast… so it repeated the question also, this time.
[01:32:02] PRATIK: Okay, and then we see that answer is, again, CVS NVC. It's repeating this.
[01:32:09] PRATIK: this bike.
[01:32:11] PRATIK: Okay. And last example, in the 1890s, the University of Chicago fearful that its vast resources enjoyed smaller schools, right, high drawing and so on. The question is, in 1890, who did the university decide to team up with? The answer is several regional colleges and universities.
[01:32:30] PRATIK: In the answer, we have University of Chicago was founded in 1891 by the late
[01:32:36] PRATIK: Charles, and so on, right? Not at all relevant.
[01:32:41] PRATIK: Okay.
[01:32:42] PRATIK: So, we observe that often, like, answers are often, like, off-topic, repetitive, right, in this case, right? Or, like, just continual updates. This is expected.
[01:32:52] PRATIK: Now, just to summarize what we just saw, like, right now, in zero sort results, the model is not giving, you know, sort answers, and in most of the cases, it is missing the answers entirely, right? And in few of the cases, it was repeating itself again and again, like here.
[01:33:11] PRATIK: Right? It was repeating again, and then the answer was repeating.
[01:33:14] PRATIK: Okay, and it continues with a pattern instead of answering, like,
[01:33:21] PRATIK: Right? And, like, in, like, example 3, the model says CVS, right? But then keeps on going and generates its own mixed question and answer, like you just see here.
[01:33:29] PRATIK: CBS, NBC, and ABC. Then again, it generates a question. And then again, it generates the answer.
[01:33:35] PRATIK: Right?
[01:33:36] PRATIK: And and it's sometimes drifts into invented detail also, right? In this example, University of Chicago, right? Where the answer was several regional colleges, right?
[01:33:50] PRATIK: So…
[01:33:53] PRATIK: like, the key takeaway is, basically, GPT-2 can produce fluent text and occasionally land a near-relevant fact, right?
[01:34:02] PRATIK: Basically, we are trying to say that it can… it is good with fluent English, but it does not understand the question, and it is answering randomly, right? It just rambles, repeats, invents, and does not know when to stop. It is going on and on, sometimes repeating the same things again and again.
[01:34:18] PRATIK: So, we are expecting that, you know, when we do fine-tuning, it will read the context and return a sort correct answer. This is our expectation, right? Before we do that, we'll just do a quick
[01:34:30] PRATIK: Evaluation of what has happened here, just to have a benchmark, right, of where we are in zero sort, and then where we land after.
[01:34:39] PRATIK: fine-tuning.
[01:34:41] PRATIK: Okay.
[01:34:42] PRATIK: So the key evaluation metrics that we are using here is exact match and F1, right? And I'm sure, like, you guys would have, like, some understanding of this already. Just to give you a brief how it is working.
[01:34:56] PRATIK: like, first, what happens if we are doing lowercasing everything, right? So, Paris and small Paris will be, like, lowercase parish will be same. We remove the punctuation, small words, like this,
[01:35:07] PRATIK: The stop words, like, A and these are removed, right? And extra spaces are collapsed. This way, like, when we do this, the Eiffel Tower, with all the, first words, capital, and the Eiffel Tower will be treated as, same answer, right?
[01:35:24] PRATIK: Now, let's look at the evaluation metric exact match. This is a very strict evaluation metric. It's basically, like, if the answer is same as the gold.
[01:35:34] PRATIK: answer, right? Then only it will be 1, else it will be 0. So, for example.
[01:35:39] PRATIK: like, if the gold answer was 1852, and model says 1852, right, then it will score 1, right? Now, if it says 18, 1852 instead.
[01:35:50] PRATIK: the answer will be 0, although 1852 was an answer. Just because of the presence of this in and the space, it will be treated as zero. So this is not an exact answer.
[01:36:00] PRATIK: Okay? Whereas F1 is a little bit, forgiving, it gives, like, you know, partial credit based on how many words, the two answers said.
[01:36:09] PRATIK: And based on that, like, we have precision recall, right? Precision, basically, of the world's model set, how many were correct. And recall is of the words in the bold answer, how many is the model recalled. And, based on this, we take the harmony mean of this, and then get the F1 score. Like, this is how we are, using the evaluation matrix, right?
[01:36:28] PRATIK: just… I just touched upon this, just so that you have some idea of how we are calculating it.
[01:36:33] PRATIK: Okay.
[01:36:35] PRATIK: Now, we are using both, because EM tells, how often the model is exactly right, and F1 tells… F1 tells us how close it is often, when the… when, like, it is not perfect, even the answer is not perfect, how close it is, right? Together, we have a, like.
[01:36:52] PRATIK: Better picture, like, you know, how the model is performing.
[01:36:58] PRATIK: Now, we'll try to evaluate, Like, our zero sort.
[01:37:04] PRATIK: like, predictions, okay? So for that, what we have done is we have imported, regular expression, string, we have imported counter from collections, TQDMauto, sorry, TQDM from TQDM.auto, this is basically to have the progress bar.
[01:37:19] PRATIK: Okay? Strings basically gives us a list of punctuation characters, and RE is basically for the regular expressions.
[01:37:26] PRATIK: Right?
[01:37:28] PRATIK: Now, we are basically taking the length of test data. It was basically already defined above, so we are taking that, right? Initially, I had, while making this notebook, I had capabilities, just so that it would run faster and all, because I was trying to create the pipeline first.
[01:37:45] PRATIK: Wait.
[01:37:47] PRATIK: Then, like what we do, like I discussed before, we do some text normalization, where we…
[01:37:54] PRATIK: lowercase everything, right? Keep only the non-punctuation characters, right? Remove the articles A and the, right? And,
[01:38:05] PRATIK: Split on white spaces, right? And again, join it. So if we have extra white spaces, right, those are collapsed. And then, finally, we return it.
[01:38:15] PRATIK: After that, we… to calculate, exact match code, basically, we first pass the prediction.
[01:38:22] PRATIK: To normalize text, and same we do for gold answer also. And then we try to see if both are same. If both are same, then, it returns,
[01:38:33] PRATIK: 1, right? Else, it'll be 0.
[01:38:36] PRATIK: And later on, it will be addressed. Similarly, for iPhone score, what we are doing is we normalize both prediction and goal tokens, right?
[01:38:44] PRATIK: And then what we do is,
[01:38:47] PRATIK: This is an age case where, both… if either of the prediction tokens or the gold answer… gold tokens is zero, then, will return as 1 only in that case.
[01:38:58] PRATIK: Okay? Just say that, like, it is like…
[01:39:01] PRATIK: Model is agreeing with the world code.
[01:39:03] PRATIK: Right? If any of them is zero.
[01:39:06] PRATIK: And, this is also to avoid errors, okay? This is an age case, okay? Now, what we are doing is we are trying to find the common words between trade tokens and the gold tokens, right? And then…
[01:39:19] PRATIK: We take the sum of it, right? And store it in num underscore same.
[01:39:25] PRATIK: So if num underscore same is 0, then we return 0, right?
[01:39:31] PRATIK: Right? And then, like,
[01:39:33] PRATIK: 0.0, basically, we are converting here integer to float, because here, the answer was returning at 1.0, okay?
[01:39:41] PRATIK: One, and then, for float operations, we are… we are doing this.
[01:39:46] PRATIK: So that we are ready for flood operations, okay? So, precision, we are calculating in this way. Whatever, common thing we got, divided by length of predicted tokens, and whatever, common words, sum we got.
[01:40:00] PRATIK: And then we divide it by length of the gold tokens. Then we take the harmonic mean of the prison recall, and hence we have the
[01:40:07] PRATIK: Like, Apple's work.
[01:40:10] PRATIK: now we, like, we have to do this for all… right now, this was a function, right? And now we have to do this for all the gold dancers, right? And take the best.
[01:40:21] PRATIK: So, basically, What you're doing is…
[01:40:25] PRATIK: Now, bold answers is basically, right now, list of
[01:40:29] PRATIK: Example, like, where are the answers in the first text part, right? We return the max EM and the max F1.
[01:40:36] PRATIK: Right? And, prediction appears against any single acceptable answer, right?
[01:40:42] PRATIK: So, that's what we are doing here.
[01:40:45] PRATIK: Okay, then we are trying to create a reusable evaluation loop over a dataset, right? So this can be used later on by the fine-tuning also.
[01:40:55] PRATIK: Right? So…
[01:40:57] PRATIK: Now we'll try to evaluate it, okay? And for that, what we have been initialized, EM total, NF1 total by 0, right? Then…
[01:41:06] PRATIK: And, we'll take…
[01:41:08] PRATIK: the, like, we… this minimum funds and is used so that we don't take any more rows than the dataset actually has, right? And then…
[01:41:17] PRATIK: We'll try to run the loop, and try to have our,
[01:41:21] PRATIK: EM total and F1 total calculating.
[01:41:24] PRATIK: Okay? And once those are calculated, we average it over the number of examples that it was calculated on.
[01:41:31] PRATIK: Right? And at the end, we'll just bring those answers.
[01:41:35] PRATIK: Okay, this is the function. Let me borrow it.
[01:41:39] PRATIK: This function that we had just defined.
[01:41:42] PRATIK: And then we get the results, and then we print it. You just layer on it.
[01:42:10] PRATIK: So, if you see here, the results are very poor, right?
[01:42:14] PRATIK: Even though we did it on many test samples, 7 test samples, the zero sort exact match is zero, given… we had already seen, this was very much expected, because in the examples we had already seen, right?
[01:42:28] PRATIK: a few of the cases. And 0Z1 is, also very less, it's 4.7%.
[01:42:34] PRATIK: These are our… these are, like.
[01:42:38] PRATIK: like, our before numbers, right? And, like, later on, we'll try to see, like, how, in fine-tuning, it, actually work… it is actually working.
[01:42:48] PRATIK: Known?
[01:42:50] PRATIK: Now, as you know, like, I did this for the whole dataset also, right?
[01:42:56] PRATIK: And, first of all, I did with 50, so I forgot to…
[01:43:03] PRATIK: Yes.
[01:43:04] PRATIK: Actually, I did it for the whole dataset, and with that, like, I had around 7,400 samples, right?
[01:43:13] PRATIK: And 3,000 for validation, right? So right now, I did 4 test samples, and, like, exact match was still 0%, right? F1 score was a little better, it was 8.5%, right?
[01:43:26] PRATIK: It's a slight bump from what we got for around 20 samples. Sorry, for around 27 samples, right?
[01:43:35] PRATIK: Okay, and just to, like, give you a brief overview of what has happened is, like.
[01:43:41] PRATIK: We, like, we also saw a few examples, right?
[01:43:45] PRATIK: You do.
[01:43:48] PRATIK: Yeah, and it is expected that exact match would be zero, because
[01:43:52] PRATIK: Right now, the model is giving long answers, that is expected, right? And many of the times, it is not even giving the answers.
[01:44:00] PRATIK: Right? So, when it was giving very long answers, and somehow it had few of the actual answers there.
[01:44:07] PRATIK: then, that contributed to some F1 score, right? And, for exact match, since we did not have the exact answer at any of the… like, if we… as you can see in the examples also, it's mostly zero.
[01:44:20] PRATIK: Although we have only taken 5 examples, sir.
[01:44:23] PRATIK: So, hence, we have exact match as zero.
[01:44:27] PRATIK: Now, what we'll do is we'll fine-tune the model. We'll move on the second half. What we'll be doing is we'll fine-tune it.
[01:44:36] PRATIK: And let's see, like, I'll try to give you the basic idea of how we are doing it. So, basically, we turn every squad example into single-line text in our standard format, so…
[01:44:49] PRATIK: We have context, question, and then we expect the answer, right?
[01:44:54] PRATIK: And then, at the end, we want end of text.
[01:44:57] PRATIK: Okay? So, the model trains by reading these lines and, learn to predict them. Over many examples, it will pick up pattern, right? After answer, we want a sort answer, right? It tries to pick, pick, like, it tried to learn from that.
[01:45:12] PRATIK: Right.
[01:45:13] PRATIK: And once the answer is, like, there, end of text, the text should stop, okay? So, how the model will actually learn is, basically, there are very few simple points, right?
[01:45:26] PRATIK: GPT works for predicting next token over and over. During training, it will basically compare its prediction to the real next token and measure how wrong it was. And then we'll have a loss. Those loss will help us adjust our model weights, and so that, like, in the next iteration, like, we'll have predictions closer to the real text.
[01:45:45] PRATIK: And once we do this across the whole dataset, the model will, gradually learn the task, okay?
[01:45:54] PRATIK: Now, one more thing to note is, we'll only store the answers, because, like, when the, like.
[01:46:02] PRATIK: Since it is,
[01:46:04] PRATIK: since the model is basically for, like, network prediction, right? We do not want the model to
[01:46:11] PRATIK: like, you know, waste the effort for learning to reproduce the long passes. Long passes includes, like, this contextual question. So even if it is giving the next word wrong, we do not waste… we want to waste our, like, effort there. We just want to see the correct answer, right? So what we'll do during training is we'll
[01:46:30] PRATIK: masters.
[01:46:31] PRATIK: Right? While calculating the loss, okay?
[01:46:38] PRATIK: So that the loss is only computed for the answers.
[01:46:42] PRATIK: Right? And, yeah.
[01:46:44] PRATIK: And, like, one epoch is basically full pass over the training data, and in the full data, we train for three epochs. Right now, we are doing it for 2 epochs, right?
[01:46:55] PRATIK: three splits we are doing train, validation, and test. Train, for the model where, like, train is basically where the model is running. Validation is basically… will be helpingus to choose the best epoch, right?
[01:47:08] PRATIK: Although the model… just remember that the model is never learned from it, right? There is no weight ablation from that, but it's just to see how it would perform on the test set.
[01:47:16] PRATIK: Right? So it just gives a… and test it, we keep it completely separate, and only… we'll use it only at the end.
[01:47:23] PRATIK: Thank you.
[01:47:25] PRATIK: So, like, before the fine-tuning, like, the motor was producing rambled, and, like, rarely gave the right answer. After fine-in, we expected to give the sort answer, direct answer, to match the format. And we'll try to see if this is actually happening or not, and
[01:47:43] PRATIK: Our evaluation metrics will also help us, you know.
[01:47:47] PRATIK: to see if… Like, if it is working or not.
[01:47:53] PRATIK: Now, tail here, does anyone have any questions?
[01:48:06] PRATIK: Okay.
[01:48:09] PRATIK: So, like, next what we'll do is, like, we'll…
[01:48:12] PRATIK: Format each example into a single training string.
[01:48:17] PRATIK: For that, what we have done is.
[01:48:19] PRATIK: we are creating this function, build training text example, right? Where we create the exact prompt used, like, everywhere else. Like, in ZeroSort also, we use a similar structure, right? Where we had question, right? Sorry, context, and then we had question.
[01:48:39] PRATIK: Okay? Now, from there, And the gold answers, we are storing it in this answer variable.
[01:48:47] PRATIK: Right?
[01:48:48] PRATIK: And what'll happen is,
[01:48:52] PRATIK: the full text will have, first, the prompt. Prompt is basically having this context and question, right? Then the answer, right? And at the end, we have, this tokenizer EOS token.us token, whatever, like, it's basically saying.
[01:49:08] PRATIK: That's the end of text.
[01:49:09] PRATIK: Okay.
[01:49:10] PRATIK: And then, whatever we get it at the end, we'll return it as a disturi.
[01:49:17] PRATIK: Great.
[01:49:19] PRATIK: Okay, so now let's see.
[01:49:21] PRATIK: Now, we'll apply this function over the training data and the validation data.
[01:49:27] PRATIK: Thanks.
[01:49:28] PRATIK: And then…
[01:49:30] PRATIK: And, like, this will be used for computing the validation loss, and this will be used for training later on.
[01:49:37] PRATIK: So, we have done the formatting. Now, technically, we run this.
[01:49:43] PRATIK: Okay, so example of one training, the model you learned from, like, context is this, right? You see?
[01:49:50] PRATIK: then… Yeah, somewhere really where she's gonna start.
[01:50:00] PRATIK: This isn't missed it. Oh.
[01:50:02] PRATIK: From here, the portion is starting.
[01:50:04] PRATIK: Great.
[01:50:05] PRATIK: All the way till here, and then till here we have…
[01:50:08] PRATIK: From here, like, answer starts, right? From here, 84% is the answer, right? And then, the end of its token.
[01:50:17] PRATIK: I doom.
[01:50:19] PRATIK: So this is, like, formatted.
[01:50:21] PRATIK: Like, as we wanted.
[01:50:25] PRATIK: So, like, just to give you a quick summary, now we have same format as our earlier forms. Answer is now included, just so that, like, the model can learn from you during training.
[01:50:37] PRATIK: And at the end, we have added the end of text token.
[01:50:40] PRATIK: Right? So that it tells the model where to… once the answer is emitted, stop generating.
[01:50:49] PRATIK: This isn't…
[01:51:08] PRATIK: Then, now, here we'll try to set the max.
[01:51:12] PRATIK: this max length, okay? Max length is basically…
[01:51:17] PRATIK: In the… if the formatted text
[01:51:20] PRATIK: Right? That will be used for training. Is more than a number, certain number, then we'll try to truncate it.
[01:51:27] PRATIK: Okay, and like always, like we did it last, last class also, we'll take the 99th percentile.
[01:51:36] PRATIK: Okay, now this whole thing, and this whole block is for that, and…
[01:51:42] PRATIK: So, first of all, we have imported microdep plotvplot as PAT. This is for plotting the histogram.
[01:51:49] PRATIK: Just to give some representation.
[01:51:52] PRATIK: Then we are, importing NumPy, okay?
[01:51:55] PRATIK: That will be used below, to convert into a NumPy tag, right?
[01:51:59] PRATIK: and numpy array, okay? And, like, now what we do is… We count tokens per formatted.
[01:52:08] PRATIK: like, train sequence.
[01:52:10] PRATIK: So, what we are doing here is, we take the… this text, right, and…
[01:52:16] PRATIK: like, we pass it to the tokenizer, and from there, whatever input IDs we get, we take that, right? And then take the length of it. This we do for each row in train formatting.
[01:52:28] PRATIK: Okay? So from there, like, we'll have token length, and then we convert into the NumPy array.
[01:52:35] PRATIK: So that, we can take out some stats.
[01:52:38] PRATIK: Okay, we have predefined… percentile, function there, right? So we'll try to do… do that.
[01:52:45] PRATIK: So we'll take out the 50th percentile, 95th percentile, 99th percentile, and the longest.
[01:52:50] PRATIK: Okay? And we'll try to print it, what it is.
[01:52:54] PRATIK: Now, from there, we will take the 99% type.
[01:52:57] PRATIK: Right? And, what we are doing is, we'll take a 99th percentile such that it's, and round it off to 16.
[01:53:06] PRATIK: Right? A multiple of 60.
[01:53:08] PRATIK: Also, it should not exceed 1024. 1024 is the maximum context length for GPT.
[01:53:16] PRATIK: Now, we are not taking maximum, because, the, We don't want to…
[01:53:23] PRATIK: like, add it unnecessarily, right? Like, we had this question last class also, so we are taking such that 99 percentile is covered.
[01:53:31] PRATIK: Dick?
[01:53:33] PRATIK: It saves us time, compute, and… Okay. And,
[01:53:39] PRATIK: So, like, after that is done, we'll try to, like, break, like, plot the frequency for this group.
[01:53:46] PRATIK: That's completely printed.
[01:53:53] PRATIK: So, 58th percentile is, 159, 95th percentile is 278 total, 99th percentile is 329, and the longest sequence is 342.
[01:54:03] PRATIK: V chose 336.
[01:54:06] PRATIK: Rounding 329 to the nearest multiple of 60.
[01:54:11] PRATIK: Okay? We take 16 because, you know.
[01:54:14] PRATIK: GPUs and most of the processors usually take, work well with a multiple of 16.
[01:54:22] PRATIK: Well, good.
[01:54:26] PRATIK: So… yeah, yeah, this is the graph. And…
[01:54:30] PRATIK: If you did it for the actual one, right, for the whole data set, I was getting it somewhere around 400. MaxLint token, I was taking 400. For this use case, it is 336. If you take the whole data, you would have it somewhere around 400.
[01:54:49] PRATIK: Now… now we'll, like, let's see what has happened.
[01:54:54] PRATIK: To give you a brief summary, most sequences are sort, median is around 170, right? And now to cover most of the
[01:55:02] PRATIK: Like, 99 percentile, we have taken, like, for… okay, this result is for the actual data.
[01:55:09] PRATIK: Right? That means the whole data. And this result here is for the subsample, or the sample of the data that we are working with.
[01:55:16] PRATIK: So I have intentionally kept this so that some of you may not have enough compute power to run this, right? So, here I have kept the…
[01:55:27] PRATIK: Like, the result for the actual one, which I got when I was running it on the full data set, okay? So, 90th percentile… 99th percentile was around 397, right? And almost every sequence fits around, about 400 tokens.
[01:55:42] PRATIK: There was one extreme outlier, the longest sequence was around 25,733 tokens, which is, like.
[01:55:48] PRATIK:far beyond everything else, this is like an outlier. We can safely ignore this one, right? And after, rounding of this to the nearest 16th multiple of 16, right? 400, max length was set as 400, and we saw that only 0.94 sequences were lost.
[01:56:09] PRATIK: Right? Sorry, 0.94 sequences were longer than the 400 tokens.
[01:56:15] PRATIK: Right? Race fit completely. 0.94% is very small. Given the number, I think we had around 87,000
[01:56:23] PRATIK: around approximately a data, right? So, that's very less.
[01:56:27] PRATIK: Right?
[01:56:28] PRATIK: The takeaway is that token counts run higher, and the word counts we saw earlier, and, like, sizing max length of 19% lets us keep almost every answer intact while avoiding the waste of padding everything to cover the red jive outliers, right? So, if we had just set it to this one.
[01:56:46] PRATIK: it was unnecessary, right? For just one of this, we had to keep the max bedding length as this one, right? So…
[01:56:56] PRATIK: That was waste of guiding, right?
[01:56:59] PRATIK: And here, we have taken as 3 touches was, sufficient for us.
[01:57:05] PRATIK: Okay. Now, for sale 13, like, what we have done is, we are trying to tokenize and build the labels for answers only, casual alumnos. So, as I told you, by training.
[01:57:18] PRATIK: We'll be masking… after, like, while calculating the loss, we'll be masking the…
[01:57:25] PRATIK: this prompt part, right? We'll be only calculating the loss for the answer part, right? So, let's begin for that one.
[01:57:36] PRATIK: So, now, like, a quick,
[01:57:38] PRATIK: hint of… how will we find the answer, where the answer starts, right? We build from… we build from the prompt prefix, right? We… we'll have… we have our prompt, right? Context, question, answer. For each example, we'll be counting its token, right? Then.
[01:57:55] PRATIK: we'll, like, whatever the prompt part was there, right, we'll mask it, and the remaining real tokens will be answered, plus the end of sequence token, that will be kept, right? And then we'll be calculating the loss.
[01:58:08] PRATIK: B?
[01:58:10] PRATIK: So, let's start. So, first of all, we start with creating one function, tokenize and mask, right? We will…
[01:58:17] PRATIK: basically tokenize the full formatted sequences, right? And we'll be padding it to a uniform maximum length, right?
[01:58:26] PRATIK: We'll be setting the truncation rule, that is, like, whatever, processed the maximum length, we'll be, truncating it there. Padding, we have kept as match length that we have already defined. For… in this case, we had around what? 336, I guess, for those of that
[01:58:42] PRATIK: 326, yes.
[01:58:44] PRATIK: Over the space equals 336.
[01:58:46] PRATIK: Okay, and then…
[01:58:49] PRATIK: What we're doing is, like, we'll build one label list, per example. For that, we have created this variable, empty list, okay? And then…
[01:58:58] PRATIK: Like, we'll process each example in the batch individually, right?
[01:59:02] PRATIK: From, like, this batch of length of text, because
[01:59:07] PRATIK: sorry, batch text, and this will be running for the length of the batch, right? And…
[01:59:13] PRATIK: So, now let's see what is happening now.
[01:59:16] PRATIK: So, first of all, We are, like…
[01:59:19] PRATIK: We are taking the token IDs, right? Then we are having the attention mask. Attention mask, basically for the real token, it'll be 1. For padding, it'll be 0. Wherever the padding is there, it'll be 0, okay?
[01:59:33] PRATIK: So that model knows, okay? So, from, we are building this prompt, like, this is everywhere in cell 1, cell 12, right? So, we'll be formatting this.
[01:59:44] PRATIK: having context and question, right? Then, based on this prompt, we'll have our prompt length, and then those… from that prompt.
[01:59:54] PRATIK: Basically, what you're doing is, from this prompt, we'll pass it to a tokenizer. From that tokenizer, once we have the input IDs, right, we'll find its length.
[02:00:02] PRATIK: Right? Then, hence, we have the, prompt length. Now, to avoid, you know, changing these input IDs, right, what we are doing is, we are creating a copy of it and storing it in levels.
[02:00:16] PRATIK: Okay, then what we do is, here the masking starts.
[02:00:21] PRATIK: we pro- we must the prompt part, right? Like,
[02:00:26] PRATIK: positions from prompt till the length of the prompt, right? So if it starts from zero, then it will be prompt to length minus 1, prompt length minus 1, right? So that the answer-only part is
[02:00:38] PRATIK: like, taken, right? And for the padding part, what we are doing is.
[02:00:43] PRATIK: Any position in the attention marks, like 0, that will be taken… that would be,
[02:00:47] PRATIK: Convert it to minus 100, so that the model knows not to compute loss over it.
[02:00:53] PRATIK: Okay, and then, we'll append,
[02:00:57] PRATIK: This, labels in the labels badge list.
[02:01:00] PRATIK: Okay? Then, what we are doing is, we'll attach the labels, we built as the new field alongside input IDs.
[02:01:07] PRATIK: Right? And return the model inputs.
[02:01:10] PRATIK: Okay, now we apply this to trading and validation.
[02:01:15] PRATIK: This is the function we created.
[02:01:17] PRATIK: Now, we'll be applying it to trend and, validation, right? So, what we do is, we map
[02:01:25] PRATIK: In trend tokenize, to have this trend tokenized, we map it to tokenize and must. This is a function, right? Bash2 is basically telling us to…
[02:01:34] PRATIK: It will not do it one by one, it will take it in batch, and, you know, get the answer.
[02:01:39] PRATIK: sorry, it will, take it in batch and apply the fun stuff, right? And whatever columns were there initially, it will be removed.
[02:01:47] PRATIK: Right? In trend format, right?
[02:01:49] PRATIK: Similarly, for, valid.
[02:01:52] PRATIK: like, we are doing a very similar thing, right? Everything is same, and we applied a valid
[02:01:57] PRATIK: formatted, okay? Now, we'll do some sanity checks. Basically, we'll take the first example from trend tokenized 0, right, and see what of the ID was kept, right? And, we'll try to see the number of IDs kept, right, and what were removed.
[02:02:13] PRATIK: Okay, so this is a, print statements for that.
[02:02:18] PRATIK: We can get out of this.
[02:02:31] PRATIK: So, like, total positions, we have 336, this is what we had decided as max length, right? Positions get, like, since answers were only 3, right? We had,
[02:02:41] PRATIK: 3 tokens, including answer plus end of sequence token.
[02:02:47] PRATIK: Answer was, now, if we decoded these tokens, we had 84%, and the end of sequence, or end of text token, right? This is what was decoded back. And, zero tokens were truncated, right? Now, when I was doing this for the full dataset, right?
[02:03:05] PRATIK: Like, I had 400 tokens, right? And, there, only 9 contributed to the loss.
[02:03:12] PRATIK: Right? Here we have 3. There, it was 9.
[02:03:17] PRATIK: Decode it, kept tokens and confirm the masking work, right?
[02:03:22] PRATIK: like, we have this end of text also, right? So that means the… this… everything worked, right? And,
[02:03:32] PRATIK: So, what happens is, sometimes, like, when you truncate it in the outlier cases, since we are considering only 99th percentile.
[02:03:41] PRATIK: suppose the answer was in… for some of the remaining question and answer, very truncated, right? Suppose it was very long, and it got truncated. And the answer is in the end, right?
[02:03:53] PRATIK: So, the answer will end up translating, right? So, we just took a quick count of how many examples we had in such case. So, in that case, it was 755,
[02:04:03] PRATIK: Right? 75… 755 answers had zero answer tokens. Answer tokens instead was… Truncated, right? And, like…
[02:04:16] PRATIK: And in this one, like, for these, all these examples, like,
[02:04:22] PRATIK: the loss would not be… could not be calculated only, because there was no answer token, right? So everything was removed.
[02:04:30] PRATIK: So, like…
[02:04:32] PRATIK: These are very few examples in respect to 87,000 approx examples that we trained it on, so it is safe to ignore this, and this will not anyway, like, you know, contribute to loss also. So it's okay.
[02:04:48] PRATIK: And if this number was, like, very high, right, then we would have to raise our maximum length taken.
[02:04:56] PRATIK: Okay, so this was our takeaway from here.
[02:04:59] PRATIK: Okay, and here, in this example, none of them were truncated. We are only working with 20 examples here, so luckily, in those examples, none of them were truncated anywhere. All the answers are there.
[02:05:13] PRATIK: Quoted.
[02:05:14] PRATIK: Till here, any questions? Anyone?
[02:05:22] PRATIK: Yeah, Sivant, please.
[02:05:25] Shivansh Sharma: Yeah, a couple of things I need to check with you.
[02:05:29] Shivansh Sharma: In one place, we are doing a padding.
[02:05:34] Shivansh Sharma: And considering end-off statements. So, is there any possibility to do so?
[02:05:40] Shivansh Sharma: Or we may land with, junk data or something like that.
[02:05:44] PRATIK: Sorry, come again, Simon, I didn't get the end part.
[02:05:47] PRATIK: Like, okay.
[02:05:48] Shivansh Sharma: I mean, when we use the end of string as in padding, like you have explained in the earlier session, the bot is having
[02:05:55] Shivansh Sharma: auto-padding concepts, but here we are using end-of statement as in padding criteria, right? So, while training, it may generate, give us,
[02:06:07] Shivansh Sharma: Dirty data, also.
[02:06:09] PRATIK: No, no, no. So what happens is, this end of text token has its ID, right? And, GPT model, like, there's an ID for this one. And whenever it comes across this ID, the model will know that, you know, not to calculate loss for that.
[02:06:24] PRATIK: So, till, for the loss part, only till here it is calculated. Whatever is after this one, right, whatever the, let's say, let's say there will be more end of text, end of text, end of text, right, to match that, padding length and all. Whatever is kept after that, it will contribute to nothing.
[02:06:42] PRATIK: So, like, that would not be taken into consideration itself.
[02:06:47] Shivansh Sharma: Okay, and one clarity I need is, is, when we are training, we are using three components, like context, question and answer, but do we need to provide a gold answer also to model to train and give the question-answer format? Because, as you mentioned, it is giving just a plain
[02:07:05] Shivansh Sharma: English language detailing, but if we need a response from the GPTO,
[02:07:10] Shivansh Sharma: From question-answer format, we need to provide a good answer also to them.
[02:07:15] PRATIK: So, that was, the plain answer format, right? Whatever you're saying is that, like, it is, just rambling. That part was in zero salt, where we had not trained the model, right? Decoder, this GPT model was initially trained to do just, next-word prediction, right? It was not trained on a specific task.
[02:07:34] PRATIK: Like, like question answering or summarization, nothing like that, right? It is just…
[02:07:38] Shivansh Sharma: Yeah.
[02:07:38] PRATIK: used to predict, like, next token. So… so, since it was not trained on it, and we were just giving the prompt, and, like, we were giving the context questions, right, it did know what to do then, right? So…
[02:07:53] PRATIK: like, it was giving, all the… it was all, like, the answers were all gibberish, right? Somewhere, it was just repeating and all, because it… it was not trained specifically for that task. So, hence, what we are trying to do is, we saw that example, we saw that how it is performing from zero sort.
[02:08:10] PRATIK: zero sort, right? And now, right now, we are doing all this fine-tuning, and just try to see… we are… after this, we'll be trying to see that, you know, like, if we give some examples, if we train on some of these examples, will the model learn?
[02:08:22] PRATIK: to give for this task, question answering task, that's what we are trying to see here. And hopefully, it'll be better than the zero salt. That's our assumption.
[02:08:34] Shivansh Sharma: Yes, so, I mean, for fine-tuning, you're saying is we need to rely on, other than context parameter, we need to provide a gold answer to the modeler to get trained himself.
[02:08:43] PRATIK: Yeah, initially, you need to… So that is basically, like, telling the model, you know, this is the context.
[02:08:50] PRATIK: This is the cushion, and based on this, this is the answer.
[02:08:54] PRATIK: It's basically telling, like, you know, this context was long, question was sorted, and answer was very sorter, right? And it is… so that answer… if you see later down, right, I don't want to spoil it for you, but later down, you'll see the answers are also, you know, sorted in length.
[02:09:10] PRATIK: So, modal, like, is slowly learning, right? Even if it is not giving the right answer, the answers are… answers have strength, right? And reliably, we have got better answers.
[02:09:23] PRATIK: Right? So, this is basically, initially.
[02:09:26] PRATIK: if he teaches, let's say, if he teach a small, baby, right, to do 2 plus 2, right? So, initially, we tell him, right, 2 plus 2 is 4, right? And after that, slowly he learns that 2 plus 2 is close to 4, and later on, if he gives 2 plus 3 also, he will be able to do that.
[02:09:45] PRATIK: So, that's the idea. That's what we're doing here. So, initially, we tell, like, this is the correct… this was the context. For that context, this is the correct answer. And we try to give many examples of this process, and expect… at the same time, we expect the model to learn this.
[02:10:01] PRATIK: Right? And so that when we give it on the… when we test it on the unseen data, we expect the model that, you know, since you have already, you know, learned on this much data, I'm expecting that, like, you know, you'll at least give some
[02:10:16] PRATIK: useful output. If not, then the model is not usable, right?
[02:10:21] Shivansh Sharma: Nope.
[02:10:22] Shivansh Sharma: Okay, and any other info… okay, go ahead, thank you. And the other quick check is, is we have used two library auto-tokenizer and AutoModel Cache LM from Transformer. So, is that sufficient for GPT-2, or we are… we need to… for fine-tune, we need to require other libraries also?
[02:10:40] PRATIK: like, if you, look at the name itself, right? So, Auto Tokenizer and, Auto Causal Model LM, something was there, right?
[02:10:48] PRATIK: So, the name Auto itself is there, right? So, it will automatically load the tokenizer.
[02:10:55] PRATIK: Right? Related to the model that we have given. So, that's why we are using it. This, like,
[02:11:02] PRATIK: This we have taken from Hugging Face, and if you know, Hugging Face is like a GitHub of different foundation models, right? So, since there were various models, so they came up with this auto-tokenizer, and, like, luckily, this supports for GPT-2 also.
[02:11:18] PRATIK: Right? So, for various models, this supports.
[02:11:22] Shivansh Sharma: to fine-tuning, we… we are, I mean, casual LM and tokenizer is sufficient for…
[02:11:29] PRATIK: Sufficient. That's sufficient. Okay. It's sufficient. From one, we can call modal, right, that we want to use, and from the other, we are calling the tokenizer that we want to, like, use on the text.
[02:11:43] Shivansh Sharma: Understood. Okay, thank you.
[02:11:45] PRATIK: Lovely.
[02:11:47] PRATIK: Okay, sorry guys, I think in the beginning, we had 10 to 15 minutes,
[02:11:53] PRATIK: confusion because of that, like, we are, we have reached the mark of 8, right? 8 PM. Give me 10 more minutes, I'll try to complete it as soon as possible, so that everyone can have dinner in time.
[02:12:05] PRATIK: I'm sorry for that part. Let me quickly continue.
[02:12:09] PRATIK: Okay? So…
[02:12:12] PRATIK: cell 14, like, we try to set up the trainer, right? And, here, the main goal is to do that, this fine-tuning, okay? Till now, we had, done all whatever we had, like, we were building up for this step.
[02:12:27] PRATIK: Okay, so now we have different various parameters, okay, just let's have a look at it.
[02:12:32] PRATIK: from… From transformers, we import trainer and trainer arguments, right? This is basically…
[02:12:40] PRATIK: Like, the training engine plus its config object, right? And then weare, like, whatever checkpoints is, or the logs will be created, that will be stored here, right?
[02:12:52] PRATIK: like, later on, we'll try to see… we'll see where it is being, right? Bath size we have taken as 8, right? And these are all the training… the training arguments.
[02:13:02] PRATIK: That's funny.
[02:13:04] PRATIK: Let me quickly, take you to this. Output directory, we have already set it here. Here, it will be stored, right?
[02:13:10] PRATIK: Like, and then, like,
[02:13:14] PRATIK: Number of trained epochs, we have set it as 2 right now. For the main data, please change it to 3, you'll have better results. Although, like, what I got was, in the second epoch, I had less validation loss, so I used that checkpoint, okay, for this evaluation.
[02:13:32] PRATIK: Okay? Now, like, train batch size, I have taken as 8 only. Eval batch size is also 8 only, because both of them are, like, I've initialized with batch size. Eval strategy is epoch, like, basically, we'll be running
[02:13:46] PRATIK: validation after each epoch, right? And also, a checkpoint will be saved after each epoch.
[02:13:54] PRATIK: And we'll be using the base model at the end. Base model will be calculating at the… with the help of the validation split that we had initially taken, validation, right? And
[02:14:05] PRATIK: A metric for the base model will be, evaluation loss, right?
[02:14:10] PRATIK: So for a loss, we are taking,
[02:14:15] PRATIK: Lower is better, not higher, okay? We want lower, okay? Logging strategy, we have done in… doing it in steps. We are… right now, we are doing it in 3 steps. When we… so when we are doing… right now, we have… we are doing this for only 20, rows, right? So it's very less.
[02:14:31] PRATIK: Right? So I did it, login steps as equals to 3. Now, if you do this for the entire, make sure you raise this.
[02:14:38] PRATIK: up to 50 or 100, maybe, right? Because we will be dealing with 87,000, and we don't want the logs to be
[02:14:46] PRATIK: like, very noisy. If it is very less, then there will be noise, like, the laws will be very noisy, right? So you increase it, because you'll be doing it for, 87,000 rows in the full data set, if you do it in home by yourself.
[02:15:00] PRATIK: Learning rate we are taking as 5 into 10th of a minus 5, weight decay, we took at, 0.01, warm appreciation was 0.05.
[02:15:07] PRATIK: Basically, the moreop pressure is, basically, we start with a lower running rate and gradually increase it.
[02:15:15] PRATIK: Great.
[02:15:16] PRATIK: How about the first 5% of the steps, okay?
[02:15:19] PRATIK: No?
[02:15:21] PRATIK: Report 2, we have said, like, we have kept it as none, because we are not sending logs to any external tools, right? Seed, we are keeping as 42 to have the reproducible output every time, and save total limit is basically one. Basically, we are telling the model that we only need one base checkpoint.
[02:15:39] PRATIK: There are, like, if you keep two, then you will be having two base checkpoints.
[02:15:42] PRATIK: Then, okay.
[02:15:44] PRATIK: Once the training arguments are done, we will…
[02:15:47] PRATIK: like, we'll, we'll use a trader that tries everything, right? Where modal is modal. Modal is basically GPT-2, we have already said. Training arguments is this, whatever we have kept here, okay? Train tokenized and valid tokenized is, like.
[02:16:03] PRATIK: Nick.
[02:16:04] PRATIK: with the data we had worked earlier, right, we had formatted properly, right? And, we had kept everything that is, kept here. And tokenizer, we use, we are calling it from the above cell, where, auto-tokenizer, for the GP22 was used, okay? So it's referencing from there. Now, let me quickly run this.
[02:16:25] PRATIK: So, epoch 2, batch size 8, trained rows 3, validation rows, sorry, trained rows 20, validation rows 3, base model.
[02:16:33] PRATIK: Like, we kept for the lowest eval loss, okay? Now we'll run the fine-tuning, let me run it, this will take some time. So let me run it right now, and then I'll start exploring.
[02:16:43] PRATIK: So, here we'll also try to see how much time it takes, okay? So we'll be importing time and store the starting time in the start.
[02:16:53] PRATIK: variable, right? And then…
[02:16:56] PRATIK: will, start the trainer, right? This will run the loops across all the books, and, like, until it is done. And whatever, summary is there, it will store it, a train.result. Whatever time elapses there, we'll get it from time.time from here.
[02:17:12] PRATIK: Whatever is the now time, and then minus start.
[02:17:16] PRATIK: the time we began with, so that we know how much time has gone. Okay?
[02:17:20] PRATIK: bust.
[02:17:21] PRATIK: So, like, we'll… these are quickly the pin statement to see the total time, final step, and training loss. Okay, I'll show it to you. And, validation loss.
[02:17:31] PRATIK: Like, we get it by the trainer.state.logist, what this… we are creating some logs here, right? I, I showed you.
[02:17:40] PRATIK: From there, we try to have the validation loss and all.
[02:17:43] PRATIK: Okay, it'll try to fetch it from there.
[02:17:47] PRATIK: Okay, and then,
[02:17:49] PRATIK: Similarly, at the end, we'll try to see the best metric, okay? The, like, best… best score by the… with the total attract, trainer track, like, that is the low eval loss.
[02:18:01] PRATIK: Okay, so right now it is training.
[02:18:03] PRATIK: It will take roughly around 5 to 6 minutes, I'm assuming.
[02:18:07] PRATIK: Okay, let's…
[02:18:10] PRATIK: If you guys are okay, can I go below? And, like, till… to save time, so that…
[02:18:17] PRATIK: Like, I can tell you what results I had for the full data, and then by the time this is done, I can tell you the results I got for this one.
[02:18:26] PRATIK: Is that okay?
[02:18:32] Gunjan Bhaiya: No, okay.
[02:18:35] PRATIK: Sure. So…
[02:18:37] PRATIK: like, now these are the trading results I got for the full dataset, right? I basically ran it for 3 full reports. It took around 3,000… sorry, 32,850 steps on full dataset, right? Validation initially increased from EPOC 1 to EPOC 2, and then from EPOC 2 to,
[02:18:56] PRATIK: sorry, validation decreased… sorry, my bad. Validation decreased from EPOC 1 to EPOC 2, and from EPOC 2 to Epoch 3, it increased slightly. Hence, we took the EPOC 2 as the best one there, right?
[02:19:09] PRATIK: And, from there, the base model was automatically loaded. Okay, so let's see, it's one more minute to go. This is the validation loss we got at the first depot.
[02:19:21] PRATIK: This is the training loss.
[02:19:32] PRATIK: And no… I got from the simple.
[02:19:36] PRATIK: Let's wait, I guess, guys.
[02:19:38] PRATIK: One minute.
[02:21:09] PRATIK: Meanwhile, does anybody have, does anybody have any questions? We have one minute.
[02:21:15] PRATIK: When you notice any other questions?
[02:22:40] PRATIK: Okay, so if you see here, in the first iteration, the loss was around 2.63, right? And in the second iteration, it was around 1.19. So we'll take this model.
[02:22:52] PRATIK: Now, when I had run this for 3 iterations, for the full dataset, like, what happened was, the third iteration loss slightly increased from here, right? So, I used the… in fact, I used the, you know, weights from the second epoch, in my case.
[02:23:09] PRATIK: Okay? And since we are only using two epochs, so I think this will be the learning goal.
[02:23:15] PRATIK: Let's see… okay, so training, it took around 6.6 minutes. Final step.
[02:23:22] PRATIK: There are a total of 6 steps, right?
[02:23:25] PRATIK: Training loss, average… over average, it was 2.82.
[02:23:30] PRATIK: And base model, is the one… will be the second one.
[02:23:34] PRATIK: Okay, now after fine-tuning, same prompt, same scoring, now we'll try to see, we'll calculate the scores, right, for the fine-tuning part, and then, do a quick comparison.
[02:23:47] PRATIK: So for that, what we do is, we again do it, run for the model.em.
[02:23:53] PRATIK: difficult.
[02:23:55] PRATIK: We'll do model node eval, right? And then, like, basically when we dothis, nick.
[02:24:05] PRATIK: It will move from training mode to eval mode, right, for stable output, and
[02:24:11] PRATIK: After that, what we are doing is, we create one variable, this fine-tune underscore answers, which is store the list of, like, answers, right? Correct answers.
[02:24:20] PRATIK: And then we loop over the same frozen demo examples we used in zero sort, cell 11, right? The four or five examples we had seen there, right? We'll loop over it, and then at the end, we'll compute, the, like, test scores that we had done in cell 11B after the zero sort, right?
[02:24:40] PRATIK: like, based on this, we have already created that font, so this time it's very easy, we don't have to write the whole code again, we just pass this one again, right? And try to calculate it. Okay, let me quickly run this.
[02:24:57] PRATIK: So, this is the example zero.
[02:25:03] PRATIK: Justice…
[02:25:15] PRATIK: This is the example. Now, if you see quickly, we have the same examples. Now, let's compare the gold answer. In gold answer, this is 30, right? And answer, like here, it's again.
[02:25:28] PRATIK: Yeah, giving gibberish, right? It just repeated the question again, right?
[02:25:32] PRATIK: In example 1. In example 2, it exactly matched the Second, this bold answer.
[02:25:39] PRATIK: Example 1, sorry. In example 2, like, it is again matching the gold answer, exactly. Example 3, it is again matching CBS and CBS, it's matching, right? Completely.
[02:25:51] PRATIK: Now, in example 4, again, it is giving, some, jibberish answer, right?
[02:25:56] PRATIK: So this… now, please note that this was trained on only 20 examples, right? Which is very less, very less for a model to be trained on. But still, even for 20 examples, we can see that, in few of the examples, the answers have become sorter and to the point, right? Like we had wanted.
[02:26:16] PRATIK: So, these are the results that we had, after this one. So, in… before zero sort, exact match was 0%, right? After, like, the exact match is around 42.9% for 7 test examples, right? It's very less data on both sides.
[02:26:34] PRATIK: Right? But still, we can see that for earlier, the results were very bad. Now, it is for, I think… I guess, 3 out of 5 answers are good.
[02:26:44] PRATIK: Right? And F1 score initially was 4.7400 sort, right? And after fine-tuning, we have 49.5%.
[02:26:54] PRATIK: Now, this, as I told you, and I'm repeating again, this is for, like, very few samples that we have taken, 20 examples, and
[02:27:02] PRATIK: This, 30 examples for training and 7 for, testing.
[02:27:06] PRATIK: Right? So these are very less. Now I'll try to show you the results I got for the whole data sync, okay?
[02:27:13] PRATIK: So that, like.
[02:27:14] PRATIK: Like, people who cannot run it on their machine, like, can have some reference. Now, when I was trading it, exact match rose from 0%. Earlier, I had also got 0% for 0 sort, right? It rose to 69.7%, which is a huge jump, right? From 0 to 69.7. Like, after fine-tuning it, I got this.
[02:27:34] PRATIK: F1 rose from 8.8.5% to 78%, 78.8, almost 79, right? So, we can see that both gains are large. EM improved by 69.7 points.
[02:27:47] PRATIK: and F1 by 70.3 points, right? These are huge numbers in terms of, like, increment, just based on fine-tuning.
[02:27:56] PRATIK: Right? Now, the key takeaway is, like, a small 1, 24 million parameters model went initially from Rambling, right? It was giving all gibberish outputs, scoring zero on exact match, answering most of the questions is now, like, answering most of the questions correctly and concisely, right?
[02:28:17] PRATIK: If not correct, somewhat it is towards the correct answer, right? We can conclude from this one also. 78 and a half is, like, almost 80%, like, 80 times out of 100, it is giving, like, decent results, right? And almost 70 times it is giving the
[02:28:34] PRATIK: Out of 170 times it is giving the almost, like, the exact result, which is a significant improvement.
[02:28:41] PRATIK: The numbers, like, we measured are on the held-out test set that we had separated it from the validation set, right? Using the exact same scoring as the baseline, so the comparison is fair, right? Part of the F1 gains come from the model, learning to be concise, not just more accurate.
[02:28:57] PRATIK: Right? Sorted answers, stop dragging the score down with extroverts. So if we had more answers, that's what I said in the beginning, right? If you have sorter answer to the point answer, it's easier to, you know, train the model, because
[02:29:10] PRATIK: The model that is not… Sorry, the answer… so, there are many words in the answer, right? Suppose some words, if we have long answers, some words may not be contributing to the real answer, right? So, while training on longer answer sequences, we may have a problem that, you know.
[02:29:29] PRATIK: those words, which do not even contribute to the answer, are also contributing to the loss. So training may take longer, training may not be as good, right? But in this case, since it was sorted, like, it also worked in our favor, right?
[02:29:45] PRATIK: So, that was there. So to give you a brief summary of what has happened, we trained a smart, like, GPT-2 model to answer questions, right? Which was initially trained to do, next… simply do network prediction. The English was fluent.
[02:30:01] PRATIK: initially, right, with zero-sort model, right? But it was not up to the task for question answering, right? And we saw that, exact match was zero, and effort score was around 4%.
[02:30:13] PRATIK: Now, once we did the fine-tuning, it sold out to 69.7% for exact math and… sorry, 78.8% to F1. The main lesson that we learned is a pretend model knows language, but not your specific task, right? Fine-tuning on a task, set example, is what teaches it to
[02:30:31] PRATIK: The format and the behavior you want.
[02:30:33] PRATIK: The before and after comparison on the held-out test data is how we prove the improvement is real.
[02:30:42] PRATIK: Now, there are some limitations, yeah. The metrics here are very strict, right? And, like, and you can see, there is still room for improvement.
[02:30:51] PRATIK: like, even when I did it for the whole dataset, I found some places where the model was hallucinating, right? It was not giving the right answers and giving some serious answers, right? This model is very small, 124 million. Like, we can expect larger models to do better in this task.
[02:31:09] PRATIK: Right? And,
[02:31:11] PRATIK: We made some simplifying choices also. We capped the sequence length and all, and 1% of the examples got truncated, right, because of that. So those were some of the design choices that we took. So, these were some of the limitations, right? You can, if you want, if you have the capability, and if you're,
[02:31:31] PRATIK: like, system is, well-powered. Maybe you can take longer sequences also, right? But, 99% covering, 99 percentile is, like, somewhat decent. Like, it covers most of the things.
[02:31:42] PRATIK: Right?
[02:31:44] PRATIK: So, the key takeaway is we took a general language model, and with a clear task format and modest amount of fine-tuning, turned it into a working question-answering system, then measured the improvement honestly on data it had never seen. The full loop, explore baseline, train, and compare is the poor workload, and can be reused for many other tasks, right? So, our aim here was basically
[02:32:08] PRATIK: to get you in touch, like I said in the beginning, also with different kinds of models, right?
[02:32:13] PRATIK: and, various tasks as well, right? BERT, with BERT, we saw classification task, right?
[02:32:19] PRATIK: Today, we saw, GPT, this was a question-answering task, and tomorrow, like, it'll be on TE5, and I'll try to come up with a new task tomorrow, okay? With this, I think, I already have lost some time.
[02:32:34] PRATIK: Please forgive me for that.
[02:32:36] PRATIK: Thank you so much, guys.
[02:32:37] PRATIK: Is there any question you would like to address? Or if not, then we can close the class today.
[02:32:42] PRATIK: So that everyone can have a dinner.
[02:32:51] PRATIK: Hello.
[02:32:52] PRATIK: Any questions?
[02:32:59] PRATIK: Okay, I take that as no. And with that, thank you guys, thank you everyone for attending the class. We really had some good questions today, thank you for those, and I think with those questions, many of the doubts that other people may, like, have had, like, even those got cleared.
[02:33:15] PRATIK: Thank you so much. I hope this was a very educational tutorial for you guys, and you got to learn some things, if not everything. And I promise you, if you go home and, like, do this tutorial and try to run it, everything yourself, try tweaking here and there.
[02:33:31] PRATIK: I'm very sure that you will get to learn a few more things, right? With this, I think we can close the class for today. Thank you so much, guys. Have a good night.
[02:33:40] Shivansh Sharma: Thank you.
[02:33:41] Deepak Bobade: Thank you, thank you, Prateek. I would like to ask one question to Jain AI Manager.
[02:33:48] GenAI Batch-2 Manager: Yes.
[02:33:49] Deepak Bobade: Yeah, so I, you said, recordings are available for Pathway to Success series, right?
[02:33:56] Deepak Bobade: So could you tell me how I can access it?
[02:34:00] GenAI Batch-2 Manager: It will be on Zoom.
[02:34:03] GenAI Batch-2 Manager: I think…
[02:34:04] Deepak Bobade: Yeah, I checked the cloud recordings, and I am not able to find a Pathway to Success series, any video related to that series.
[02:34:12] GenAI Batch-2 Manager: Okay, okay, you can raise a ticket regarding that, and I will just check with the tech team.
[02:34:18] Deepak Bobade: Okay, thank you.
[02:34:19] GenAI Batch-2 Manager: Okay, thank you.
[02:34:21] PRATIK: Thank you.
[02:34:21] PRATIK: I think with that, we can close it, right?
[02:34:25] PRATIK: Close the system's looking.
[02:34:29] PRATIK: Thank you.
[02:34:34] PRATIK: Good night. Good night, guys.