# 02 2026-05-16 Hands On Transformer

course: Module 4 — Generative AI & LLMs
module: Module-4-Generative-AI-LLMs
date: 2026-05-16
type: transcript
video_url: https://personal-learn.armco.dev/files/_Recordings/Module-4-Generative-AI-LLMs/02_2026-05-16_Hands_On_Transformer.mp4

---
[00:19:32] Hello. Good evening, guys.
[00:19:38] Give me a second, I'll quickly
[00:19:44] Say your ice cream
[00:20:09] Can anyone please confirm if
[00:20:11] Like,
[00:20:13] My screen is reserved
[00:20:14] Yes, it's visible, Pratik.
[00:20:18] Still sick.
[00:20:27] Let's wait for a couple of minutes, so that everyone is here, and then we'll start.
[00:22:42] Meanwhile, this, uh…
[00:22:46] Files were shared with all?
[00:22:51] Yeah
[00:22:52] For the practice, lab material, I'm not able to locate it.
[00:22:53] Yeah.
[00:22:55] Sorry, come again.
[00:22:56] I'm not able to look it in the portal.
[00:22:59] Is it not uploaded?
[00:23:01] At least I'm not able to see it.
[00:23:05] notes there.
[00:23:07] Uh, can you please navigate, please? Uh, who is saying it's there, at least?
[00:23:12] I've shared the link. I think they have added one more module, like, Module 4, so you'll find it over there.
[00:23:48] You're saying they added a module?
[00:23:51] 4, okay?
[00:23:52] Let me copy this link, Deepak.
[00:23:57] Yeah.
[00:25:04] Okay
[00:25:05] Let's begin. First of all, very good evening to everyone
[00:25:11] Like, today will be basically
[00:25:14] Working on… like, I hope everyone, like, is familiar with Transformers, especially given the, like, a class that's already happened on it
[00:25:23] In the previous tutorial, we did about we studied a little bit about LSDM attention, and we tried to work upon. It was little bit of
[00:25:30] Which is still left. Hopefully, if we get time today, we'll try to complete it. Today, like, we'll try to start with a new topic, since everyone is in fresh minds.
[00:25:40] Right? And then later on, like, if we get time, we'll continue with the previous left
[00:25:47] So, like, last time we studied tutorial for LSTM. Lstm without attention and Lstm with attention. This time we'll try to see like how transformer is actually like helping the
[00:26:02] Language translation task, right?
[00:26:05] So, just like last time, we have taken the same data set. The data set is basically
[00:26:13] bilingual satisfy dataset, right? And
[00:26:17] Like, it has English sentence and then magic French translation right next to it, okay?
[00:26:26] So, like
[00:26:27] These are the two columns, English and French, right? And the size of the data set just to like
[00:26:35] get everyone on the same page. Like, it's around 1,75,000 sentence pace, right?
[00:26:42] Okay, and for demonstration, this time we have taken only 200 rows
[00:26:48] 200 euros is, yes, pretty less, but, like, for demonstration, to complete the training and to show how it is working, we have taken this, right? If anyone wants to, like, get the best
[00:27:01] performance, right? Base goals
[00:27:05] Like, training it over the whole dataset, right? Instead of selecting 200. I'll exactly show you where I selected the 200 rows of the data, right? It is recommended that you train it on the whole data set. Okay? And apart from that, in this data set, sorry, in this tutorial
[00:27:23] What we have done is, like, we have trained for one epoch, right?
[00:27:28] Due to the time constraints and all
[00:27:31] So if you want to do it, like, you can stretch it to 3 maybe right? I have done it to 4, 3 epochs, and I have got some results which I'll be sharing at the end
[00:27:41] Right? So the results
[00:27:44] A decent, it's good, but, you know, it's not the, like, the best that we can get, okay?
[00:27:52] So, with that, let's start
[00:27:53] So this, like, in the first cell, what we are doing is, like, we are trying to select the, like, if anyone has GPU, if anyone has CPU, right? So for now, we'll be trying to select the CPU, because everyone needs Colab and like we'll try to work with it. So what you do is, like, you try
[00:28:11] CPU here and code it will select the device CPU. So later on, all the training that will happen is, like, will be on CPU only. So that's what we are doing here, okay?
[00:28:22] Now, like, in the next cell, we are basically installing some required libraries for this tutorial, right? Some of these are like transformers, sentence piece, secretly pandas
[00:28:36] Right? Okay, let me, like, quickly go over it. Transformers is basically, we have Hugging Face library, right? That lets us and download pre-trained AI models. Okay, we will use it for
[00:28:52] Like, in our tutorial, we have used Marian MT model, right? That has been already pre-trained
[00:28:59] Right? Between
[00:29:01] Languages, okay? So we don't have to train it from zero like we did last time. So we'll try to import it and then work on it, okay? How we have done, we'll discuss below.
[00:29:11] Then we have used will be importing will be downloading the sentence piece also. Okay, it's a helper tool that breaks sentences into smaller pieces called subwords. Okay, this is very important, like, because models understand small chunks better than the whole sentences, given the context
[00:29:29] issue, right? The context issue will be also discussed in the further classes, right? As you deep dive more into Llms and all like how context is very important for the models to understand and all
[00:29:43] Right
[00:29:44] Then we have used, we'll be downloading Secretary Blue. It's basically a text processing library used internally
[00:29:52] We used basically for blue score to calculate the blue score, okay?
[00:29:59] And then we have used saccharomosis, right? It's basically like a text processing library that will
[00:30:08] like, internally be used by Secret Blue
[00:30:12] Okay? It basically… what it does is, it will, consistently,
[00:30:17] You know, tokenize the translations before comparing them
[00:30:21] Right? And it is basically to get back the, you know, converting tokens back to the readable sentences and all
[00:30:29] Okay?
[00:30:31] Then pandas, I think everyone is familiar with by now, okay? And this is how we have called it. Now, a quick info. Like, we use exclamation mark here, right?
[00:30:42] Because we want the, you know, like it to be treated as a terminal command rather than a Python code. Okay? And with this we have installed will be installing transformer sentence do sector, Moses and pandas.
[00:30:55] Okay, let me quickly run it.
[00:31:04] Okay
[00:31:06] We have
[00:31:08] All… everything done now.
[00:31:10] Then, moving on to the imports, in cell 1, we'll import, like, basically import all the libraries that we'll be using throughout this tutorial. We'll be importing OS.
[00:31:21] It's basically a tool like all the it provides us all the tools that we need to work with folder path and files, right?
[00:31:31] Then
[00:31:32] Pandas, everyone knows, like, import torch, basically, we are doing to work with PyTorch, right? So, we have used in multiple scenarios. One we have used to, like, check if GPU is available or not, right? If, in the above
[00:31:47] Right, in zero B cells, 0V. Then we have used to, you know, train our model also, pre-trained model will be, you know, a little bit fine-tuning it later down. For that, also, we have used, okay?
[00:31:58] And from Transformers, we'll be importing Marian MT model and Marian tokenizer. Marian MT model is basically, it'll be used for translation, right? That model we are importing and tokenizer is basically now this model has its own tokenizer
[00:32:15] Right? So, like, for that, we are using Marian tokenizer
[00:32:21] It will be basically, you know, converting the sentence, you
[00:32:26] Two numbers and all, like we have discussed in the previous classes also
[00:32:31] Secretary Blue for calculating blue scores.
[00:32:34] So let's see. We'll try
[00:32:38] And see if everything is done. So, Pytor version is checked. Device selected was CPU, as we wrote there.
[00:32:43] Okay, now we'll try to load the data
[00:32:48] So, here we are loading it, everyone's… like, I hope everyone has it here, okay? From there, we'll call it.
[00:32:58] This is the data path
[00:32:59] Right? If anyone is confused how we do it, then, like, you take this copy path and then you paste it here. As simple as that
[00:33:07] So, like, after that, we will be, loading it into the DF frame right? The Csv file. And here we have selected number of rows right? If you remove this
[00:33:18] and coma, like, it will select the whole rows. Like, everything in the data. Like, that's how you work on the full data set.
[00:33:25] For now, due to compute restraints and all, and time restraint, like, we have selected this as 200, so that everything is completed within the shoot over.
[00:33:36] Okay? Then, just like last time, we have renamed it right to English and French.
[00:33:41] Right? Because last time also it was renamed to something different.
[00:33:46] This is the last tutorial file.
[00:33:49] We'll see.
[00:33:50] The header was like this, right? So this time for easy, ease and all, like, we have done that.
[00:33:57] Okay, now we'll, we'll be printing a dataset, say, columns, and the first few rows, just to have a look on it
[00:34:03] So, since we selected the 200th
[00:34:09] And there are two columns. So the shape is this, okay? Then the column names has been changed to English and French for ease. And these are the first three rows
[00:34:18] Let's see how it has been brought in.
[00:34:21] Now we'll try to do some basic IDA, just to have some recap. This we have already done in the last tutorial also, so we will be just briefly going through it.
[00:34:34] So, like, we try to, like, we'll… here we'll try to explore it, okay? So, we'll see the total number of rows, total number of columns, okay?
[00:34:43] Then we'll try to see if there are any missing values per column, and for that we do isdf dot is null, and then we take the sum of it, okay? Like, how many total values, like, how many rows are actually missing the, you know
[00:34:57] values for that.
[00:35:01] Then we, take for duplicate rows, if any, like, rows are duplicated or not, right? To avoid duplicates while training
[00:35:10] Okay, for that, we have done DFW duplicated, and then some, okay? Then we take five random English sentence phase, and just to have, like, again, look
[00:35:21] And this time, we have not taken index.
[00:35:23] Like last time we had taken place. And then what we'll do is we'll quickly
[00:35:31] calculate the sentence length for each row
[00:35:33] Okay, I forgot to ask, is my screen readable to all? Like, is there any issues with it?
[00:35:41] Can anyone please confirm?
[00:35:42] No, um, it's visible clear.
[00:35:48] So, like
[00:35:49] The text size is, good, right? I don't need to zoom in or anything. That is fine, right?
[00:35:58] Okay, thank you.
[00:35:59] So like now we'll be trying to calculate the sentence length for each row, just to see, like, you know, later on we can choose max length and all
[00:36:09] For… to, so that every sentence has the same length. Okay, so for that we'll try to have a look on it
[00:36:17] Okay, so for that, what you have done is, for English sentence, we have tried to take a length, and then we have added a new column for this one.
[00:36:24] Okay, then friends, we have taken the length, okay? Now, now these are two columns added
[00:36:32] Then, what we'll do is, like, we'll try to
[00:36:36] Like, we'll, use .describe function, right? And, it'll be routed to two digits, okay?
[00:36:44] And similarly, we'll do for friends. Now let's have a look. I'll quickly run it.
[00:36:51] So, the total rows are 200, as we took. Total columns now are 4, because, now there is English length and French length added right missing values is 0, right? Everything has some values
[00:37:04] And, like, 5 random pairs, sentence pairs, we took it. And this is how it looks now.
[00:37:10] English land or French land, right?
[00:37:14] And then, like, these are some, you know, statistics for English sentence length.
[00:37:20] Right? And French sentence length. So, count is 200, right? And mean for all the sentences is somewhere around 1.85 standard deviation, 25%, 50%, 75%, everything is there. Okay
[00:37:33] And similarly, there it is for French.
[00:37:36] Like, count 200, 2.57 is the mean, right? And then standard deviation, minimum 25%, 50%, and 75%. These are the quality range and all. Okay? And then maximum length is 6
[00:37:49] Here, the maximum length is 2.
[00:37:50] In the first 200 rows
[00:37:52] Now, if you take, like, you know, the whole dataset, the results might have been a little bit different, because if you have seen the dataset, like, the sort of sentences are above, and, like, as you go down the dataset, like, longer sentences are also there
[00:38:10] So here, just to clarify
[00:38:14] We are using two models, not one. By this, I don't mean to confuse you. It's just that, like, I just want to make a clear distinction, okay? So here we have used Marian Anti architecture, okay? Now, this is like
[00:38:31] Briefly tell you about this.
[00:38:34] Like, for now, the brief I have given below. For now, just understand that this is basically
[00:38:42] like a blueprint, right? Architecture was created, right? And this Helensky NLP, right, slash Opus MT English to French, right? Is the train model, like, it's basically you can think of it like
[00:38:56] They trained the translation from English to French
[00:39:01] Right? Using
[00:39:03] Some data, like, I'll also briefly discuss about the data, and then once it was trained, they have saved the weights and all, and, like, we are using it as a pre-trained model, and we'll be importing those rates, right? Think of it like that
[00:39:19] Then, like, this was basically done by the, you know, language translation, sorry, language technology research group of University of Alanski like it has been written here right
[00:39:32] So
[00:39:33] Like, this is how we have, you know, used the marine empty model.
[00:39:37] And this is basically like a task that it was trained for, right? It's calling it lag time
[00:39:44] So, as it… as I just told, it's basically marine empty turkey chicken. The weights of this Helensky NLP
[00:39:51] Like it is loaded, which has been trained for English to French translation.
[00:39:55] Now, more on the marinent this transformer, okay? It's basically, encoder-decoder architecture, like, that was, you know, designed specifically for machine translation
[00:40:09] Right? Now, various models are trained for various tasks, right?
[00:40:15] Like, and some of them are trained for specific purposes, right? Like, in this case, it has been trained for language translation right? The good thing and the like, you know the pro about this model is like it's very focused
[00:40:30] lightweight model, and you can even use CPU for this. Usually, when you go, like, if the model's parameters increases and all
[00:40:39] like, say, billion parameters and all. It's very difficult to work with CPU, if you have a task, right? It takes lots of time. Like, it takes days, weeks, and so on, right? So, even sometimes having GPU is not enough, like, you need, like
[00:40:54] state-of-the-art GPUs for such tasks to handle such large data and all, right? So, this one, like, luckily is, like, very lightweight model, and it can work on CPU
[00:41:07] Now, similar, like, to what we have studied in sequence-to-sequence architectures and all, this also has in photo decoder. Encoder basically reads the English sentence
[00:41:20] And on the basis of that, it tries to have some context, right? By context, like you can also take it as understanding or try to get its meaning and all. And after that, like, then it is
[00:41:33] passed on… that context is passed on to the decoder for generating and understanding the French translation word by word, okay?
[00:41:42] Now, basically, like, that we discussed before, it was trained on OPAS dataset. This is a dataset. You can look it up on Google and all. It's a huge data set right and basically the dataset was made by
[00:41:56] collecting translated text from the web, scraping it, right? And including, like, it can also have… it also had this subtitles, books, news articles, right? And it was trained, like, right now we are working on 1, like, 75,000. This one was trained on, like, you know
[00:42:13] millions of sentences. Okay
[00:42:16] So
[00:42:19] Like, this model basically has 75 million parameters, right? And as I told you before, it works on CPU, right? It is freely available on Hugging face
[00:42:28] Usually for gated models, like,
[00:42:32] It is still available for free, but you have to do some authentication of sort. Like, you have to get some permission, right? But this one is freely available, you don't need any Hugging Face token or anything, it works freely, right? And it was specifically built for English to French translation, right?
[00:42:48] And, like, despite its small size, on a pre-trained model, like, we achieve, blue score, a good blue score, right? Ignore this for now, I'll let you know more on the results side later at all
[00:43:02] So, first step, what we do is, like,
[00:43:06] We
[00:43:07] load the pre-tokenizer, pre-train tokenizer on the MT model, right? And the process will look… and the higher level, it will look something like this. Like, this is a sentence, right? The tokenizer encodes it, and then it is sent to a model
[00:43:23] Right? Model does the translation, right? And the translation modal outputs are used numbers and all right from there, tokenizer has to decode it back to get the frames text. Okay?
[00:43:39] And as I told before, Mariana Tokenizer is the specific tokenizer
[00:43:44] That was, designed to work with marine empty family of process, because this is not only
[00:43:50] This is not only the only model that, like, they have, right? They have multiple models also. And this one they have used for the Marine MT.
[00:44:00] Now, here we have defined the model name
[00:44:04] Right
[00:44:05] Helensky NLP is basically a research group that trained it, right? So that's why they have this name
[00:44:13] And this was their task
[00:44:15] Opacity and French, whatever.
[00:44:18] Now, we download the tokenizer. It's about, like, 1 to 2 MB, okay? And this is how we do it. Marine tokenizer from… .from pretend, and then the modal name, modal name is, like, it takes from here
[00:44:30] Okay, it's only given
[00:44:33] like, we'll quickly do a sanity check if it is working or not, right? A sample is taken and then we pass it to the tokenizer.
[00:44:45] And, like
[00:44:47] Now, whatever comes from there, like, we'll try to encode it
[00:44:53] Okay
[00:44:55] And one more thing, this tokenizer will automatically add, like,
[00:45:00] Like, some, like, like, for example, it will add zero, right? This is, like, basically, end of EOS token, like, we have studied last time also right
[00:45:14] Then, now we'll print the sample, the tokens, the token IDs, and the vocabulary size.
[00:45:19] Now, let's see
[00:45:26] Oh, this is done.
[00:45:28] Let's this
[00:45:33] So,
[00:45:35] Hello, how are you, was the sentence, right? The tokens are created as, like, underscore hello
[00:45:42] Underscore how? Underscore R, underscore U. Now, this is one of the tokenization scheme, right?
[00:45:49] this has been not taught, there are many different types of tokenizers, right? And like different models have generally different tokenizers. And some of the popular tokenizing techniques will be discussed in the class
[00:46:05] So, right now, like, I can show you one quick observation. Like every word is starting with an underscore, if you see
[00:46:12] Right?
[00:46:14] And sometimes, like, if the world is bigger, right, like
[00:46:20] In other tokenizers also, like, what happens is, suppose they say it's understanding, right? I'm just giving example, it may not happen exactly like this. So what will happen is like underscore will be there, then maybe understand comes in one token, right? Or under comes in one token
[00:46:36] Then the next subword, right? It'll break into small, small tokens, right? So what will happen is, next subword, maybe, let's say, stand, and this time the stand will not start with an underscore, right? And,
[00:46:49] And after that, it will be in, right? So that's how it will be like first converted into tokens.
[00:46:55] So, it is basically, we are doing it to break it into small numbers, small sub words, right? If the word is, like, a huge word, or a long word. Okay, having many characters.
[00:47:09] And, like, after that, it is converted into token IDs. Like, these are the numbers, right? And then we quickly see the vocabulary space
[00:47:19] Now, like
[00:47:21] In the previous session, like, we had gone through multiple normalizing techniques and pre-processing techniques, sorry, preprocessing techniques that had normalizing, lowercasing, right? Removing extra spaces and all right? So here we have not done that
[00:47:38] This has been taken care, by the, you know, model itself. Since we are using a pre-trained transformer, most of the heavy takes processing, like, it has been automatically handled by the tokenizer itself, right?
[00:47:52] like, you know, punctuation splitting, right? Your coma was splitted, right? And the, this question mark
[00:47:58] Then subword tokenization
[00:48:00] Avoids are broken into smaller pieces called soft words. Okay
[00:48:05] like, use of underscore and all, and special tokens.
[00:48:11] like, having zero at the end, which is basically saying end of sentence token, right?
[00:48:17] And, like, converting takes two numbers and all. Last time, we did this ourselves, okay?
[00:48:24] Now
[00:48:25] We do need to be… this doesn't mean that, like, we can be completely carefree. Like, we still have to be careful about few things. Having missing values, extremely long sentences
[00:48:38] presence of extremely long sentences, and, like, there may be some encoding issues if, like, CSV is corrupted and all
[00:48:46] So in the extremely long sentences part, like what can happen is, like, the model, like, the model that we are dealing with can handle, like, around 512 tokens at a time, right? So sentences longer than this, like it can be cut off
[00:49:01] and lose. But we don't have to worry about this, because in our dataset, like, the longest sentence we have is only 44 words. So we are well within the limits, right? And this suits our like
[00:49:17] Sorry? Are you sick? You're saying something?
[00:49:21] Maybe
[00:49:24] Quickly mute him
[00:49:27] So, like, missing values we have already checked, we don't have any, so we, we can be carefree on this part that we have already checked it right in cell 3
[00:49:38] Okay, now
[00:49:42] So for this one, we are basically skipping the pre-processing and cleaning step because we have checked for missing values and all length of data set is well within 512 tokens, right? And CSV is not corrupt. We already checked the outputs and all. It's looking fine
[00:49:57] So now, we'll try to load the model
[00:50:00] Okay?
[00:50:02] And, like, so let's see what happens.
[00:50:05] So, till now, what we have done is, like, in Salesforce, we had the tokenizer, it converts text to numbers, right? And those numbers will be passed on to the model, and then the translation will happen, right? So
[00:50:20] So, now what we do is, tokenizer, we have already discussed, okay? So yeah. Now we'll download the weights and all of the model, right? And the model name was already like given above in the previous list. So that's how we are calling it Marian
[00:50:37] empty model . And then the model name. Okay. Now.
[00:50:43] Like, whatever device we have selected in the first zero B, like that is taken here. Then the all the model weights
[00:50:53] It is, you know, transferred to this one
[00:51:00] Okay. Then, what we do is,
[00:51:04] Like, we already know, like, from our experience that a neural network can behave differently during training and inference, right? In training mode
[00:51:14] Some neurons are dropped out to prevent model right from memorizing and all
[00:51:19] We also use many normalizing techniques like batch normalization, right? And then in evaluation, suddenly like what happens is we like
[00:51:31] We turn off the dropout and all, right? So, and there, all neurons are active, because
[00:51:36] Their dropout was specifically used, or any regularization techniques were specifically used so that you know model does not like memorize it or overlearn it, right?
[00:51:48] So for that,just to check, like, what's the… how it will act on the, you know test data like in evaluation mode, all these things are turned off
[00:51:58] Okay?
[00:51:59] Now, since we are only using this model to translate, right, right now, so what we have done is, like,
[00:52:08] We have, done two things here, right? First, we have used, you
[00:52:16] pre-trained tokenizer, you… and then, with the help of that, we have generated translation without doing any training, just using the pre-trained weights, right? And with the help of that, we have tried to, you know, translate
[00:52:29] And the next thing what we have done is we have taken a small data part of a data, part of data from the data set. And with that, we have, you know, trained it, trained the very NMD model
[00:52:43] And try to see if the performance increases for our use case, right? So, we have done that, right? So, right now, if you don't use, right, you know, like, if you don't do any fine-tuning, you can… you are good to go from
[00:52:59] model.eval, right? So, now we'll quickly do the sanity checks and all
[00:53:04] If the model is loaded safely, and model is running on, and the number of parameters, okay?
[00:53:09] Let me see if I run the previous…
[00:53:13] That is done
[00:53:19] Let's see. So it is loading the weights right now, as you can see, right?
[00:53:24] And, like, it is running on CPU, and, like…
[00:53:30] of parameters here. Okay, it is slightly more than what I had told
[00:53:33] I had told around 75 million. It's slightly more than that, sorry for the
[00:53:41] Okay.
[00:53:44] I think the… this one is for the English to French translation, right, for that task, this is the marine empty architecture by itself has… okay, I'll just confirm on this, okay? I'll have a look on it again.
[00:53:59] Now, like,
[00:54:03] Basically, first, in the first part, we'll try to fine-tune the pre-trained model, our dataset, right? And in the second part, what we'll do is, like.
[00:54:12] We'll use the pre-trained model itself, right? We'll have our respective blue scores for each of the model, and then we'll try to see the difference and all. And also we'll be comparing with our
[00:54:25] You know, tutorials from our last class, like how it had behaved for those. Okay?
[00:54:33] So, like, just to give a brief about fine tuning
[00:54:37] What happens is, like, in fine-tuning, what we do is, like, we already take a model that has already been trained, right, you know, train it with our data
[00:54:46] Right? And then, like.
[00:54:50] What we do is
[00:54:51] Like
[00:54:53] After that
[00:54:54] Like, since it has seen our data also, right, apart from the fine tuning it has done
[00:55:00] Like, we can have better results. And in lesser time, with lesser data. Now, let me quickly explain you how this is happening. If we had to, you know, find, you know, train this network ourselves, right?
[00:55:14] It would be very difficult. First of all, the sheer amount of data that this model has been trained on is huge. We would have huge computation issues and all
[00:55:23] Hence, like, there is a, like, you know
[00:55:27] We are
[00:55:29] Like, it's… it's we are getting some help from already,
[00:55:35] fine-tuned, sorry, pre-trained data, pre-trained model, okay? Because that has already been trained on some data. So there are already some training has already happened
[00:55:44] Okay, so with that, what we can have is, like
[00:55:48] We can train our, like, you know, train for our,
[00:55:53] results
[00:55:56] like, very quickly, for our problem statement, right? It can be done in the fraction of time of what it would take to, you know, train the whole network from scratch
[00:56:06] Right? And the other thing is, like, it would
[00:56:10] Also, save us some computation power, right? We don't have to load that used data in our system and all. It is just a 1, 75,000 is very less, right? And since it has already been trained, we don't have to worry about training it from scratch
[00:56:26] I believe you would also be finding this code simpler than what we had for like LSTMs, right? So that is also there, because we are already using fine tuned
[00:56:39] Sorry, this pre-trained model like we don't have many of the things
[00:56:44] Like, are already done for us, okay?
[00:56:48] Now
[00:56:50] For fine-tuning, what we'll do is, moving on, we'll, like, what we'll do is, like, we'll
[00:56:56] distribute into two sets, training set and validation set, right? In training set, we'll have 80% data, and in validation set, we'll be, having 20%, right? Now, in this, validation one, like, the sentences, these sentences will enable you… be, like, seen by
[00:57:12] You know.
[00:57:15] our model, okay?
[00:57:21] That's what we have done here. In fine-tuning step one
[00:57:26] Right, first of all, we have imported this, trend test plate from sklnmodel.selection, right?
[00:57:33] We have selected,
[00:57:35] Test side at 0.2, which basically means 20%, right? And the remaining 80% will go for the training.
[00:57:43] Right now, I have used 200 rows. I was experimenting this
[00:57:48] Take it as an example.
[00:57:49] Right? So, 160 rows will be used for training set, and 40 rows will be using as, like, validation, right?
[00:57:58] Then we perform the split
[00:58:00] Okay.
[00:58:02] So, we have training list, well English, train French, and VAL French contract split
[00:58:08] Okay, using the English sentences and the French sentences.
[00:58:12] Okay, and those will be converted to,
[00:58:15] you know, list, and then, like,
[00:58:19] This side we have already given from above. A random set is, like, to have reproducibility to some extent. Okay, we have set it as 42.
[00:58:29] Now, we'll… like, these are some checks that we are doing in the print statement will look below
[00:58:36] number of total sentences, training sentences, and the validation sentences. We'll just have a look at it
[00:58:44] Okay? Then,
[00:58:45] We'll have some like we'll try to see some of the English and the French pairs okay how they are corresponding to each other.
[00:58:53] Like, could you run this
[00:58:57] Okay. So, as we have foreseen, like, total sentences were 200, training sentences are no
[00:59:01] 160, right? 80%, and validation sentences are, like, 40, 20%, right? And some of the, sample training this way.
[00:59:13] Now, what we'll do is, we'll try to create a custom dataset class. Basically
[00:59:21] PyTorch needs to have data in a specific format, right? And a dataset class is basically like a small container that will hold our sentence space. And
[00:59:33] It can also tell us, like, you know.
[00:59:36] How many pairs it has, right? And it can handle one pair at a time when asked. Okay
[00:59:42] So this is, like, an example that I have given, yeah, for a date. So basically, it's like, holding a deck, like, you know
[00:59:51] Like, it's having a deck cough that holds all the cards, and then you have to see, like, you can check how many cards are there in a deck, basically asking for length. And then, if someone asks, like, give me card number 5, it can simply get it for you.
[01:00:07] So, this will have, look at the class that we have defined below, so I'll show you.
[01:00:12] Okay, and now, like,
[01:00:17] the fringe tokens, basically, for the target has basically become labels. Labels as in the tokens and all right that token IDs and all that the model is trying to learn and produce. Okay
[01:00:37] Okay. So let's quickly go. This I'll tell you why, like, in the example itself.
[01:00:45] So, like, as we said, we have three, two things here, right? One, we can have the length, and then we can see, like, if you have to get a particular item from the dataset class, we can do that.
[01:00:59] So
[01:01:01] Just like we discussed before, we import dataset from Torch, torch.utils.data, right?
[01:01:09] Then we create a class
[01:01:11] Right?
[01:01:12] Then we have 3 methods here.
[01:01:15] one in it that all the classes,like, as soon as the object is like created, it will inherit that right
[01:01:23] So, basically, we are, like, saving the tokenizer if it needs later, and the max length, okay?
[01:01:31] Now, what we'll do is, we'll tokenize the English sentences
[01:01:36] With the help of this. So, let's have a look
[01:01:39] So basically here, the tokenizer will basically
[01:01:44] convert a string to tensors of numbers, right? And padding length… padding is the max length that we have decided, okay? And
[01:01:56] Here, I think we have taken 128
[01:01:59] By default.
[01:02:01] Okay?
[01:02:04] So
[01:02:06] So, 128, and, like…
[01:02:11] Truncation, we have kept it as true, like, if it has, you know, 128 more than it, more than 128, we cut it off, right? To prevent, like, so that all the, you know, sentences have the same tokens
[01:02:26] Right
[01:02:27] And then we return, like, this is, like, PT is basically, like, calling PyTorch tenses, and we return it as Python tenses. So that it can be directly passed to the model.
[01:02:43] Okay, so this is, like, what the tokenizer is doing for English sentences, okay? So we have, in the tokenizer, we have passed English sentences, padding length
[01:02:54] Padding is like 128. We have set up a… right, truncation is true, right, max length is, like, like what we discussed. Like, that is the max it can have, right? And return transfer is like PT, like PyTorch transfer, okay? That's how it will be returned
[01:03:10] Provided to the models. Similarly, we do for French language, right? French sentences. So, here also, we call the tokenizer on the French sentences, having padding, truncation, true, and max length. Okay, and this will also be returned as a
[01:03:26] Sorry, Pytos tensors
[01:03:30] Now, like, what we'll do is, we'll replace the padding token IDs
[01:03:35] With, like, minus 100, or something like that. Let's see here
[01:03:41] Like
[01:03:43] If padding id is one right like initially it is given one. We have
[01:03:49] Like, suppose if the sentence is lesser than 128, right? So we'll have some padding done at the end, right? And for that, so that, like, the model understands it is all padding tokens, whatever you do is, like, suppose, for example, this is 1, right? Then what will happen is
[01:04:06] Those ones, after, like, you know, masking it, it'll be converted to, like, minus 100, right? And
[01:04:15] So that, like, this is basically it is done, so that the model knows, like, these are the, padding tokens, right?
[01:04:22] And now, what we have done is, like, we are cloning our, like, targets, so that
[01:04:28] We have an independent copy of the tensor, and if there is any modification done right and suppose it goes, it is not right, right? We can fetch our labels again.
[01:04:42] Right? So for that, we have cloned the input IDs
[01:04:50] Then, what we have done is, like,
[01:04:53] Like, the trailing underscore, like, whatever we have, right? In what we are seeing. We'll, we'll have in place operation for those things
[01:05:04] Yeah, right? For that, we have done, yeah.
[01:05:07] Minus 100, whatever it is
[01:05:11] like, token ID, whenever it is padding is done, it'll simply replace with minus on here, right? Now, this is the function that we have
[01:05:19] method sorry this is the method that we have for
[01:05:22] Having length, like, self-input, input IDs, and .safezero will give
[01:05:29] The length
[01:05:30] And this is the class we had discussed before, where, like, if we give the ID right, then it will the like dataset class can give us the sentence by that we are referring to.
[01:05:41] Just, like, what we are saying in the example before, like, if we ask for 5, then the dataset will give, you know the fifth sentence pair.
[01:05:54] Okay, so a quick example, if you call for data set 0, right? Then it will give input ID attention mask and level
[01:06:06] So, and at the end, we'll have… this will have returning us input IDs, attention mask, and labels, or whatever we have done.
[01:06:20] Now, we create the actual train and validation dataset objects
[01:06:25] From training list, train French and the tokenizer right and this will be our train data cell. Similarly, we'll create a translation data set with well English Val French
[01:06:39] And, like, you know, this one. Tokenizer
[01:06:42] Now we'll do some sanity checks if the data is created correctly or not, right? And try to retrieve the zeroth item
[01:06:51] Right, and to see, like, if it's there or not
[01:06:55] The previous one is run or not
[01:06:59] Okay, Mr. Chuan, so let me check this one.
[01:07:06] Now, the training dataset size has
[01:07:09] 160 sentence space, validation data size is up to 40 sentence space, right? Structure of one data set
[01:07:15] item is, like, input is
[01:07:18] It will have keys, input IDs say, attention mask safe and label safe.
[01:07:24] Right? This, we asked for it, and these are the keys
[01:07:29] Correct, like, in a dictionary, how we have, keys and values, okay?
[01:07:32] And for the first time, the input ID is… will be, like, 10 side… like, in the tensor form, it is… it was returning, right?
[01:07:40] So this is how it is
[01:07:42] Potential phone. Notice, like, after end of, sequence, 0,
[01:07:48] The remaining padding has been replaced by minus 100
[01:07:53] Okay.
[01:07:54] Now we'll try to create a data set loader, okay
[01:07:59] Here, what we were doing was, we were
[01:08:02] See
[01:08:05] We were trying to create a data set class. Basically, it's like having cards and all, right? Now, with that
[01:08:14] we are trying to have our dataset
[01:08:17] loader, like, to load all the,
[01:08:21] It's basically like, how should I explain? Okay. It's basically like having a small container that will have our, like, English sentence and the French sentence pair, and, like.
[01:08:34] So, basically, before, it was passed out one pair at a time, right? Now here, we can call it in batches. We don't have to call it one by one. We can simply take, say, 16 sentences at a time, 32 sentences at a time, and
[01:08:49] The batch size will depend on how much memory we have, right? If we have more memory, let's say we can take 32 sentences at a time. But let's say if we have less
[01:09:00] memory, right? And then if we call for more, the batch size is more, then what will happen is there's a risk of going will having an issue called out of memory, right? You'll be running out of memory, and then
[01:09:12] The training will stop there itself. So you have to find the correct batch size
[01:09:17] For your training also. That is also very important.
[01:09:21] Okay
[01:09:23] So, like, as I said, a dataset
[01:09:27] The loader will basically wrap our data set
[01:09:30] And automatically group it into batches of each size. Batches like whatever we have defined.
[01:09:36] If it's 16, then we are calling for 16 sentences right? And then also one of the advantages like it will be shuffling the training data. So basically it will not it will. It restricts the model by doing this. It basically is
[01:09:53] He's checking the model to, you know memorize it. Like, if… if, say, we already know that Sam is coming after RAM, right? Every time, then instead of learning what Sam is doing or what RAM is doing, like, it will memorize the positions and try to predict it
[01:10:11] So in training, we usually have some kind of shuffling happening, okay? So this dataset loader does that for us. Okay
[01:10:20] And in the background, one suppose, let's say, a 16 sentences are, like, have gone for, you know
[01:10:27] in one epoch, right? Sorry, for training. Let's let's not say that. So if 16 sentences have gone right, then in the background, the next 16 sentences will be prepared, so that you know it is very efficient and less time is
[01:10:43] In loading one batch to another. Okay, so this also takes care of that
[01:10:50] Now, like,
[01:10:52] We'll have two data loaders. One is for training, and one is for validation. We already know why, kind of because train loader we'll be using for training, right? And well loader will use it for validation purpose
[01:11:08] Now, one of the differences in training
[01:11:10] The train loader will have some kind of shuffling, right? And in validation, it will not suffer, because it's like testing on the real world, right? So we don't need to do that.
[01:11:21] So in next step, like, as we just discussed below, above, we will try to create the data loader. Okay, so we'll be importing data loader from torch.utils.data
[01:11:35] Okay, we have defined the patch size as 16, right? This worked out for me, right? So I have kept it as 16 in Google Collaborate is running at around 5 min
[01:11:45] For 200 sentences, right? So it's fine for me, and I'm not running out of movement, so it's a good news, right?
[01:11:54] Then
[01:11:55] Like, we have created one train loader, right?
[01:12:00] Like, how we have done is, like, we have data set loader. Sorry, data loader, that object that we have
[01:12:07] From there, right, we import a data loader, okay? And then for that, we have passed train dataset that we had created in the above cell, right? Batch size is 16, we have given that. Suffol is true is basically we are shuffling the order in every epoch
[01:12:22] Right? And, number of workers, we have set it as 2 is basically
[01:12:28] We get to decide how many, you know, work, like, CPU, like, how many of the workers can be used for faster loading and all, right?
[01:12:39] Similar thing, I think we have in Random Forest also, okay? So…
[01:12:44] There, maybe you have some list. Now we'll be similarly creating a validation data loader right? And similarly, instead of this one
[01:12:55] training, like, the differences in the training data loader, we had passed the training data set, now we are passing the well data set. Batch size is exactly the same we have given as false, because in validation, we don't need it, right? And the number of workers is set
[01:13:14] Okay? Now, we'll do quick sanity checks to confirm the data loaders. If it is looking good or not, okay? So for that, we'll have… try to see the length
[01:13:24] Right? If it is matching or not, okay? And, how many batches we'll have, okay? So that also. What we can do is number of the sentences, let's say 200, right? And if we have 16, then how many matches we will have
[01:13:37] for
[01:13:38] So, let's say, if we had 1600 sentences, right?
[01:13:42] Then
[01:13:44] The training will have, basically, 100 batches, right? And in validation, we have 400 sentences, so it'll have 25 batches.
[01:13:54] Okay?
[01:13:55] Now
[01:13:57] We, try to peak in one of the batch, and confirm if it looks structurally correct or not, okay?
[01:14:04] So
[01:14:06] Let's quickly run this.
[01:14:09] Okay
[01:14:15] Okay. So, the batch size is basically, like, is exactly what we have said, 16, right? And the training batches, let's see. So, 160 sentences we had, and by 16, if we do, then that will have 10 batches, right? And in validation
[01:14:31] Like, we had 40 sentences, and batch size we had set is 16, so that will have 3 batches, okay? Now, if we try to have a look at one of the batches, how the structure looks, then this is how we are looking at it. So, 16 is the number of batches
[01:14:48] And 128, because we had said the maximum length as 128, right? So, in one sentence, like, let's say, like, one sentence in a batch
[01:14:58] Right? It would have, 128,
[01:15:02] Tokens, right? So, then, like, if it is less than that, it'll be, you know, padded. If it is more than that, then it will be truncated at that. So that's why, like, input is 128, and to match the order, like, everything will have the same.
[01:15:17] So this is this looks good, okay? Now we are good to go for the next step.
[01:15:22] Now we'll start our training for the fine-tuning
[01:15:26] So, like, in training, we already know what happens, like, there is one training loop
[01:15:32] And in that we have a 4 hour pass, where you know the model looks at the English sentences and tries to predict the French translations, and then it will compare it to the correct French levels and calculate a loss. Okay, and in the backward pass, what happens is
[01:15:49] Model calculates, like, how far it is from the, you know, the actual ones, and then it will try to calculate the error, and then,
[01:15:58] Once, like, this… like, doing this is called back propagation, right? I think you guys have studied that already. And then, at the final step, the weight updation happens, right, based
[01:16:12] Like, with the help of optimizer and all
[01:16:17] So
[01:16:19] Like, in the validation loop, what will happen is, like,
[01:16:22] Here, this will not happen.
[01:16:26] The backward pass will not happen, right? We are just trying to see how far is the model from the actual truth. So let's see
[01:16:37] And in the validation, we also see if the loss is going down every epoch or not. So that is also good idea. And if the loss is going up while the training goes down, that is very bad. That is one of
[01:16:51] key identifiers for like overfitting. Basically, which means, like, it's memorizing the training data
[01:17:00] And,
[01:17:02] The optimizer, we have used AdamW, running rate, we have used 5 into 10 to the power of minus 5 right and epoch
[01:17:13] I had trained for the whole dataset in 3 epochs. Right now, due to, like, time constraint, I have taken it as
[01:17:22] Okay, so…
[01:17:25] This will take around 5 minutes
[01:17:30] Let's
[01:17:32] Let me run this
[01:17:34] Okay, so training has started. I am doing this to save time, okay? Now let's have a look.
[01:17:41] In this city
[01:17:49] So here, what we are doing is
[01:17:52] Like, just like we discussed, we'll be having our training and the validation, okay? So, the optimizer, we have used Adam, right? W basically means, wait, it is, like, for weight decay, right? TQDM, like we said,
[01:18:09] In the early tutorial also, it is used to display live progress file like this.
[01:18:13] I completely show it to you.
[01:18:18] Play progress bar. And the good thing about is having this is, like, you can also see, like, you know, for this epoch, how many times… how many minutes have already passed in real time
[01:18:28] like, this is the remaining time, and this is the like number of like like what time has passed already.
[01:18:35] portraying, right?
[01:18:36] And here we have bachelor's and all, let's see how that will happen
[01:18:52] Before we move on, like, does anyone have any questions?
[01:18:55] Right? Before I move on to explaining this.
[01:19:00] Anyone?
[01:19:04] Sure. Okay, so I'll quickly move on to, like, explaining this.
[01:19:08] Okay, so number of you folks, as I said, like I have taken it as one. Okay, learning rate, I have taken this 520 per minus 5 right? And then we have set up the
[01:19:21] The optimizer
[01:19:23] Thank you. Notice that, like, I have taken the learning rate very low, because I don't want it to, you know, drastically move here and there, right?
[01:19:33] So
[01:19:38] Okay, now we have set up the optimizer, like AdamW, right?
[01:19:42] learning rate we have already fixed, weight decay 0.01. As I said, W is for that weight decay
[01:19:48] And this is how it is
[01:19:51] Now, we'll switch to training mode, right? So for that, what we do is we call model.train
[01:19:58] So, when we do this dropout will activate right? And then, like in during training, what will happen is some of the neurons will be randomly selected, and it will be turned off. It will be deactivated, basically. So while so it will directly pass
[01:20:13] Pass through it, right? And
[01:20:17] This is done. Basically, we don't want the model to over memorize it, right? So this helps in this acts as like a regularizer regularizer.
[01:20:29] Now, for storage of tracking loss across epochs, we have to
[01:20:34] We have created two variables, train losses and well losses, and after every epoch, we'll try to have the average loss over epoch, because for each sentence, maybe we may be having some
[01:20:47] loss, right? And, like, since we are taking a batch, it will be calculated across the epoch.
[01:20:57] Okay. Now, then we move on to the training loop
[01:21:01] Right? For epoch in range of norm of epochs we are basically number of epochs, here it will be one only, right? So it will run for one epoch.
[01:21:09] Then, we have, for training this model.train that we did, right? Training loss is initially set as 0. Total training loss. And then for each batch
[01:21:19] Right? First, we move the batch data to the correct device, right? Whatever we had selected, ZPU or CPU in the first zero B, right? Then in here, by doing this optimizer.zero grad, basically we are zeroing all the gradients
[01:21:36] from the previous patch right and giving first start for calculation, right? And then we do the forward pass to compute the loss, right?
[01:21:46] So
[01:21:48] Like, for that, input IDs will, like, we have already defined this, right? Input IDs, attention mask, and labels, right? All these things we have decided already like done
[01:22:01] And from there it will fetch it, and then we'll calculate the loss. Okay? And
[01:22:06] at the, like, third step, like, after the forward pass, what we'll do is, like, we do backward pass. This is basically computing the gradients, how far it is from the truth value, right? So for that we'll do loss.back word right
[01:22:20] And, like we discussed in the last tutorial also, it was used there also. So we are tipping the gradient. So if a gradient is too large, right? So that may drastically affect the training of the model, and it may suppose that maybe not intended, right? Sometimes
[01:22:38] there are errors, and we don't model to, like, you know.
[01:22:42] would be unstable. So what we do is, even if it has more, right, we clip it at,
[01:22:49] At some point. So for here, like, in our use case, we have clipped it at 1.0, right?
[01:22:56] So this basically will help us to prevent any, drastic update
[01:23:00] After that, the weight ablation stage is there. For that, we have used this optimizer.state.
[01:23:08] And
[01:23:09] For this batch, we'll be calculating loss
[01:23:12] right? And that is like training loss. Total training loss is equal to total training loss plus loss.item for each item that is calculated across the batch. Okay? And then, like, this is for the
[01:23:25] This TQDM path. Okay, for that we are doing this
[01:23:29] And, now we'll be calculating the average training loss, right? And the
[01:23:35] whatever, like, losses we'll have, we'll try to have a list for it, so that we can see below. I'll equally sorry to you at the end
[01:23:44] And,
[01:23:46] Here, average loss is basically total training loss by the length of the train loader, like, whatever
[01:23:52] Like we did right for the
[01:23:56] train loader… train loader
[01:24:00] And each trainloader, like, is loaded with 10 items right now
[01:24:06] Okay
[01:24:07] And we'll print it, we'll check, I will have a look at that, you know. Now, for validation, like, after, each epoch, there is a validation also happening, just to see, like, calculate the loss and all.
[01:24:18] So here the dropout scene after we call model.eval, like we discussed before, it will
[01:24:26] The dropouts and all will be turned off, right?
[01:24:29] And, no, this bypropagation and weight updation will happen.
[01:24:35] Just, we'll just try to see the loss accumulated. So for that now we have set this variable as 0 total validation loss. Okay? And we start with with basically it is telling them ptorch to not to track gradients and all, because we'll not be doing
[01:24:53] There's bank propagation. Okay, so this saves memory and will speed things up for us
[01:25:01] Like, and
[01:25:02] Belbar is basically, like, a bar that we have… that TQDM bar is for that. It is just for sewing. And, for each batch, right?
[01:25:13] We'll again like similar to what we had done in training, we'll move it to the correct device, do forward pass on it, and this time we'll not do any backward pass and weight updation that we did before. We'll calculate the average validation loss for this epoch. Okay
[01:25:27] Like, wait for turning order. So, here in total, for validation, we'll have total valossal loader.
[01:25:36] Okay, and here, like, it is for 3 items each, right? And average, validation loss will be calculated. So we'll just have to… we'll try to see the summary across all the epochs at the end.
[01:25:50] Let's see what has happened. We have trained for one epoch, right? Batches, as we have seen, it is for 3 batches
[01:25:57] And then,
[01:25:59] Sorry, batches is 10 batches, like, we had calculated this, because batch size, we have taken as 16, and 160 were the, like, total sentences in the training, right? So 16… 160 divided by 16 is 10, so that's how
[01:26:14] And at the end, validation, like, after training loss, like, we try to calculate the validation, right? And this is the
[01:26:23] ticket and bar, right, for validation
[01:26:26] Okay, this is the bar that we had for validation, okay? And now
[01:26:31] like, this is for 3, right? This also, we had seen how the calculation has come
[01:26:36] Now, this is like a quick print statement of like what has happened in one epoch. Train loss was 3.64, and band loss was 2.8629, right? Right now, I've trained for one epoch only. Please note that, guys, you can do this for
[01:26:52] You know, 3 box, and taking the whole dataset, and you'll have some good results, right? One epoch is very less for the model to learn anything. Right now, due to time constraint and all, we have done this
[01:27:06] So, like, what we do is, now, we'll save the fine-tuned model, right? And, because training it takes time, right? So… and if you have to repeat this process, or if you have to use that fine-tuned model for some
[01:27:22] Like, let's say some other dataset, right? So then what we can do is, you can simply save the model, and saving the model will basically… what will happen is like
[01:27:34] like even like we can simply get the weights from that and continue the training from there itself, or you can use it for some other task right
[01:27:48] So…
[01:27:49] When we save the model, basically two things get saved. Modal weights, right? And, like
[01:27:57] Motor weights, and also the tokenizer that was used, okay? So, the modal weights is basically the millions of numbers, the weights, like, it learns during training, and, like, during the process of fine-tuning
[01:28:12] Like, it'll be updated slightly, right? And when we call this model.save. underscore pretend, it'll save those updated weights. Okay
[01:28:23] And, the next thing is, like, vocabulary and the rules converting for text and the numbers, text to numbers and back. Like, first it is converted to text to numbers, and then numbers to text, right? At the end, during the decoding stage. So
[01:28:38] That is saved by the
[01:28:40] tokenize tokenizer dot save underscore pre-trained. Okay
[01:28:46] So, right now, 4Es, what I have done is, whatever data path has been defined, there the, you know, like, this
[01:28:56] It will be saved
[01:28:57] the model will be saved, okay? So, right now in cell 2, we have defined
[01:29:05] this path right where our data is, that will only be available
[01:29:13] So, like,
[01:29:16] like, here, as we discussed, we'll save it, right? For that, what you have done is we have imported OS here once again, just to
[01:29:24] So, okay, where it is being used. Okay, and the data path right now is this, right?
[01:29:29] Okay, so if we… like, this is the whole
[01:29:33] data path for the this one.
[01:29:37] our CSV file. So it'll take the directory name right? That is only slash content. So it will fetch that, and then the safe path will be slash content slash fine tune model
[01:29:50] Okay.
[01:29:52] So, for this, like, let's try to save.
[01:29:55] Okay, so, say, path we have defined as,
[01:30:00] Like, data path, whatever it took from here, right? And then it will be saved under a fine-tuned model. Okay
[01:30:11] Now we'll do some print statements just to see where it is saved and all.
[01:30:17] So, right now, we are trying to see if the file folder actually exists or not. Like, that slash content or not.
[01:30:25] So that, like, we can ignore the… like, if it does not, then we can work on it. Just confirming that, and once we have confirmed that, we'll save it to the date, like, whatever save path we have created, yeah?
[01:30:38] Okay, and
[01:30:41] We'll be also saving the tokenizer, right? So, for saving the
[01:30:47] Modal weights, we have done all the, like, saving the model, what you have done is model.save underscore pretend, like we discussed, and for the tokenizer, we have done tokenizer.save, underscore pretrain, okay?
[01:30:58] So we are doing some checks like where it is saved or not, if it is saved or not. Okay
[01:31:04] And, we'll try to see the size of it also.
[01:31:09] And from that safe path, we'll try to reload the model later if we need it, okay? Let's see if we can run this.
[01:31:25] Okay? So
[01:31:27] The dataset location is, like, slash content, right? Moodle is, like, save to slash content slash, fine-tuned model, right? And,
[01:31:38] folder was created, right? And, like, first of all, the model weights are stored, right? Here it is, it is stored. Then the tokenizer is stored, okay?
[01:31:51] And these are the 7 files in total that were saved.
[01:31:56] Now, to reload, like, what we can do is from transformers import this, marine empty model, marine tokenizer, and model we can what we had done earlier right from pretend we had that Helen's key NLP slash NLP
[01:32:11] Sorry, hyphen NLP, and slash for that English to French translation thing we had, right? Now, see? For the pretend thing, we can also call our model also from there.
[01:32:23] Whatever we had sale. Similar for the tokenizer.
[01:32:28] So, this is one important learning that if you want to use your model later on, right, to maybe do some other tasks, right
[01:32:38] Maybe the dataset may be different instead of training everything from scratch. If the problem is very similar, you can load those right and you know further train it. If the task is very similar, like in this case.
[01:32:49] Okay. And this is how we have used it, okay?
[01:32:55] Now quickly do so, if it's there. See, the fine-tuned model is there, right? And the same files. 1, 2, 3, 4, 5,
[01:33:05] 6, 7, right?
[01:33:07] So all these are there, what is enlisted here
[01:33:12] Now, like, once we have understood this, we'll try to, you know, translate a single sentence
[01:33:19] to friends, just to see if everything is working or not, right? So ideally, what should happen is, the sentence raw English text would be there, then you would tokenize it, right? Moodle reads those tokens, right? And, generate some output
[01:33:35] And then after that, what will happen is
[01:33:38] Decoding will start based on those generated outputs. And after that, at the end, we'll have the French text. Okay
[01:33:48] So let's see a quick example of how this is working. I love learning new things every day. Is the sample sentence like we took before, right?
[01:33:57] We first of all, in step 1, we'll tokenize the input sentence. We are just trying to create the pipeline. These codes we have already done before also, but right now for sample sentence, we are trying to do this. The above one was done for the whole training
[01:34:10] Right? Training set
[01:34:13] So first of all, we tokenized it and saved these as input, right? Input sentence. Then
[01:34:21] Like we'll from there, we'll try to have our input IDs. Okay
[01:34:28] And, like
[01:34:30] From actual dataset also, we can see, we'll try to see some examples, okay? So for that, what we have done is we have sampled three rows from our dataset. Random set 42 we have set so that it is, you know, every time we can
[01:34:45] I have the same 3 samples
[01:34:48] That is randomly chosen.
[01:34:51] So
[01:34:54] like, here what we try to… it'll basically try to show us the subword pieces as the text was split, okay? And then
[01:35:02] Like, after the encoding has happened, we can see how 0 was added automatically at the end. Okay, this will see in the twin statements and all. Then we'll try to see the generated translated outputs
[01:35:17] From there
[01:35:20] So just give like how it is working is, like, encoder will read the English token, build a rich numerical representation. And this here we are using Marian MT tokenizer and all right? And then once it has
[01:35:36] Like
[01:35:37] learned everything, based on the tokens, right, context and all, then it'll be used by the decoder to then generate one French token at a time, right? And, like, once the translation has ended, it'll produce an end-of-sequence token, and from there, the translation will stop
[01:35:57] Okay.
[01:35:59] So, we'll try to
[01:36:01] Generate the translated tokens, okay, from here
[01:36:07] And basically this import is basically like, if you remember, we had input IDs, attention mask, and label, I think or yeah, here I think only two things input IDs and attention masks
[01:36:23] These are, like, like, keys, right? So values will be you on like it will basically unpack
[01:36:28] the dictionary and the values of those will come here, right? And based on based on those, the tokens will be translated
[01:36:38] And,
[01:36:40] This is, like
[01:36:41] From here, the tokens will have, like, and we'll have that, and then we'll, from those tokens, we'll try to decode it back to the text that was translated, okay? And just, like, we'll try to print those sample sentence, and the translation text
[01:36:58] Translated text, and just have a look, okay?
[01:37:04] I'll quickly run this
[01:37:08] So
[01:37:10] Input sequence as we saw right
[01:37:13] I love learning new things every day. The tokenized input ID is 4717793655 and so on. And at the end of the sentence, right? 0 is added.
[01:37:27] So, just see, 1, 2, 3, 4, 5, 6, 7, 1, 2, 3, 4, 5, 6, 7…
[01:37:34] Sorry, 1, 2, 3, 4, 1, 2, 3, 4, 5, 6, 7
[01:37:40] 1, 2, 3, 4, 5, 6, 7… and then, like, at the end, 8, and then 0 at the end, okay? So here are some, like, some word is breaking into two, that's why we have one more
[01:37:53] And then
[01:37:56] Let's see how it is
[01:37:57] Like some examples from the dataset
[01:38:01] For callers, you know, underscore call, then underscore us, right? And the token IDs are this one
[01:38:09] 67033683 and 0.0 basically is saying, like, end of
[01:38:15] this sentence, right?
[01:38:18] And then simply go on also, underscore go, underscore on, then punctuation, whatever you have. That is also being taken as token here. And you can see that for both the punctuations
[01:38:29] 3 has been here
[01:38:32] This is also, like… so it's okay, right? If you have to verify and all. So, and then get up
[01:38:39] Right? The last one
[01:38:44] Last one is, like, again
[01:38:45] Sorry, I mistook it. 3 is for the full stop.
[01:38:52] Okay.
[01:38:54] And then the translation token IDs produced are like these, right?
[01:38:59] And, the English sentence was this, and then translated sentence is this one. Now, also, I wanted to show this, like, 59513, what is this, okay? Let's have, like, I had written something about it
[01:39:13] Yeah.
[01:39:15] So, whatever, like, token ID, we had, right.
[01:39:21] So we, like, for that, 59513 is being used.
[01:39:28] Okay, for, like, you
[01:39:31] See? The first one.
[01:39:33] Here, like, in this case, the first token that comes is 59513.
[01:39:38] Life. It's basically like saying, you know, start, start of sequence. It's basically like telling that, you know, you can start, generating
[01:39:48] the translation, right? So, like, we had that SOS taken, right? Start of sequence. So, 59513 is acting like the, that for… in mer and architecture
[01:40:03] So
[01:40:04] Yeah.
[01:40:09] Yeah. So in Nextcel, like, we'll basically, right now, we did all this for a single sentence, right? Now we'll do this for the
[01:40:21] We'll try to create a function, so that the translation can happen for all the sentences, training sentences that we have, okay?
[01:40:31] So we'll try to create a function for that
[01:40:35] Now, like.
[01:40:37] In DevTranslate, we are basically passing on the text, and it will return a list of translated fence strings, okay?
[01:40:47] At the end. So now let's see what we are doing here. First of all, like
[01:40:55] His instance, this text underscore string will check whether the input is a single string, right? If it is wrap, if it is, we wrap into a list, right? So let's say a hello, like if it is hello, right, it will be
[01:41:12] This in under a list, right? If it is hello, hi, then it is like it stays like that.
[01:41:17] It's already a list. If it is single word, then that will be also changed to a list, okay? So that's what we are doing here.
[01:41:27] So next, what you do is, we'll tokenize the input, right? Just like we discussed before, truncation is true, padding is true, right? And the tensors will be returned as Python tensors, which will be needed by the model
[01:41:41] Okay, so, yeah.
[01:41:43] PT for return chances, like Python tenses, padding is true, truncation is true, like, 128 we have set. If it is more than 128, it'll be truncated, right? And all this is taken to, like, our device that is selected CPU, right?
[01:41:58] Now
[01:41:59] With that, we'll try to translate the token IDs. Okay? Now,
[01:42:06] Let's see. So
[01:42:10] Like in… like, what we can do is, like.
[01:42:13] This, as we discussed before, this will unpack a dictionary, right? Inputs and inputs is basically like having those input and then
[01:42:23] And there was
[01:42:24] At, this one, ma'am.
[01:42:26] Great, I'll show it to you.
[01:42:32] Like, it was having those, input IDs, attention masks, right? So those will be unpacked from the dictionary, and straight away it will have those values here.
[01:42:44] Okay? Yeah. Okay. Then…
[01:42:47] And the translated tokens will be kept in this variable, translated underscore tokens, right? After it is generated
[01:42:55] Like, once we have those tokens, we have to, like, decode it, right, to the French tokens. So whatever translated tokens we have got, we'll
[01:43:07] Pass it on to tokenizer.batch_decode, right? And if there is any special tokens, like end of sequence or whatever, like, is there, I will skip it for the decoding part
[01:43:20] And at the end, we'll return the translation, whatever we have.
[01:43:24] So, we'll do a sanitary check on it, test sentence, I love learning new things every day, and result we should have a translation
[01:43:33] like, translated sentence.
[01:43:36] So, let me run that, let me see if I have run the previous… okay, let's run
[01:43:43] So, English sentences, sentences having, I love learning new things every day, and the French is this one, okay? I'm not pronounce this for my lack of knowledge in French.
[01:43:53] Okay? Now, like, once this every… we have created a function, we have tested a single sentence, if it… if that works on that function, right? Now we'll try to translate a real sentence from the dataset
[01:44:06] Okay
[01:44:07] So, this time we have taken a sample of 10 rows, right, with a random state of 42, and we'll be dropping the index. Index is basically, like, 0, 1, 2, that is there in the data set.
[01:44:19] So that one we are dropping, and we'll be keeping the rest.
[01:44:22] In the sample DF
[01:44:26] Now, we'll try to extract the English sentences as a plain Python list, okay?
[01:44:32] Like, two days will convert it to Python
[01:44:34] Python list. Okay
[01:44:36] And then, like, we'll translate all the 10 sentences in one batch.
[01:44:42] So, all the English sentences that was loaded is passed on to the translate function, right? At the end, we whatever output it comes, it is stored in the predicted underscore French. Okay?
[01:44:54] Now.
[01:44:55] We… whatever like predictions we have, we'll have like we'll try to have a separate predicted column to compare it later on. Okay? And at the end, we'll display the formatted table
[01:45:07] So, for this, we are trying to basically see English, and then predicted French, and then the actual French, just to see how it is
[01:45:15] I'll quickly run this
[01:45:22] So, this is the English sentence, right? And this is the predicted French.
[01:45:28] Okay, so you see this is exactly the same. This is somewhat little different. This is
[01:45:34] like little different right and
[01:45:39] This one is exactly the same. This one is exactly the same
[01:45:43] This one is exactly the same. And here and there, we have few, difference
[01:45:48] From the predicted and the actual
[01:45:51] Now, like, while I was doing this for the whole dataset, like, I found that, like, there were 3 scenarios, okay?
[01:46:00] Like, there was… there was perfect match, matches in most of the cases, like, you know, because both the models, like,
[01:46:09] like sorry the model that we fine-tuned, and the ground truth are basically very trained on the you know
[01:46:18] This model was trained for that only. So, like, somehow it has been able to capture it perfectly, right? So in those cases, we have perfect matches. In some of the cases, what has happened is, like,
[01:46:31] The sentence translated is correct, but it has slightly different,
[01:46:36] Wording, okay? For an example, like, for the English sentence, I'm not scared to die, right? The model has given this. It's like, not afraid to die, is the translated text. And in the truth, we have do not fear dying
[01:46:52] Right? So, like.
[01:46:56] context is somewhat the same, little bit wording here and there is a little different, right?
[01:47:05] So another example, how did the audition go?
[01:47:08] like, here, if you see
[01:47:11] It's almost the same here.
[01:47:13] Here, in one part it is used as a D-roll, and I may not be pronouncing it right and here in the other part it has been used. In the truth, it has been used as a passive, something like this, okay
[01:47:25] So, basically.
[01:47:28] These are very similar words, right? And they can be interchanged
[01:47:35] And, like, here is one other example. I really like this shirt, right? Here, also, it's, like, very similar
[01:47:44] Just that instead of this word
[01:47:47] Sorry, instead of this word, this one was predicted
[01:47:51] And, like, both of them have some meaning of, like, really, like, right?
[01:47:56] So this is one of the like good observations, right?
[01:48:00] Like, that we can make. And then
[01:48:03] In
[01:48:06] And sometimes, like, what is happening is, like, let's say if most of the people understand Hindi, right? Like, sentences like that are also like
[01:48:21] there, right? But if you see word-to-word, those are different
[01:48:25] Right?
[01:48:26] But the meaning is same, but the formal, like if we have elder, we say like and if we have someone to like very small, or we are talking rudely, then it's like two words
[01:48:42] And so it's similar to that.
[01:48:46] Okay.
[01:48:47] So this is one other example
[01:48:50] Right, and in some of the cases, like, model was better than the ground truth also, because the model was trained on very huge corpus, right? So, in some of the cases, we found that sometimes model was predicting better
[01:49:06] It was very closer than the ground truth also. Okay.
[01:49:10] Means both are right, but the prediction was somewhat, like, closer
[01:49:18] Now, like, while doing the evaluation, as I said, these things will, you know, matter a lot, because, usually what happens is we do word-to-void matching, right? And with that, what happens is, like, if you don't know
[01:49:34] Friends exactly like we not be able to correctly identify, oh, like, this is this sentence and the other predictive sentence, ground truth and the predicted sentence are basically having the same meaning
[01:49:46] But, like, somewhere, like, these scenarios can come, like formal, informal, right? And correct, but like different wordings
[01:49:56] Right? So
[01:49:58] These are the things, like, that may not be accounted in the evaluation. Like, especially in the blue sport that we are doing, right? Because they… we try to see for, like, how many words are, like, are matching in some context, okay? Not exactly that, but in some context.
[01:50:16] So, basically, key takeaways that we can, like, see from all those examples are
[01:50:22] Most differences are about, word choice, and not actual errors, okay? And in some cases, the modal output was better also. We also found that. And in some cases, naturally, the ground truth is better, right? And while computing blue score also, we saw that, you know, like, later on we like in blue score what happens is
[01:50:44] Valid alternative sentence translations will be
[01:50:49] penalized because maybe the voids did not match exactly right? So
[01:50:56] having a proper, like, very good, evaluation is also, like, a tough thing, right?
[01:51:04] So
[01:51:07] Okay, so this is for that sentences, I repeated it by mistake
[01:51:11] So you know this one
[01:51:13] So now we'll try to compute the blue score for the loaded data set, whatever we had. And then we'll try to compute the compute the blue score for the pre-trained one, like without using any fine tuning without
[01:51:30] Doing anything, we'll just pass our sentences to the pre-trained model and try to see how the evaluations are, and then try to evaluate on the blue score, okay?
[01:51:40] metric. So, like, here what we have, like, this is what we are exactly doing
[01:51:48] Okay, so we have completed blue scroll twice, right? One for the original pre-trained model, like that from scratch we had got, and one after fine-tuning
[01:51:56] Okay, and
[01:51:58] So
[01:52:00] So, I was experimenting with different sizes so that I was trying to take as many numbers as I can for tutorial, but like after experimenting a bit, like, I found 200 was able to, you know I was able to do that training within the
[01:52:17] window. So, like, pardon this mistake typo here
[01:52:22] Right? It's basically we had done for, we have done for 200 sentences and one epoch, right?
[01:52:29] I did this for 10,000, 2,000, 10,000
[01:52:33] And the full data set, right? And I'll also tell you the results that I got for the full data set, because that is more significant
[01:52:41] Okay.
[01:52:43] So, right now we are doing it for 200. So the results may not be, you know, quite accurate to what we would have for
[01:52:54] You know, the actual data. So, please consider that as well.
[01:53:01] Okay
[01:53:03] Now, in the earlier sale, we also saw that, right, you know, the actual, in the actual blue score evaluation, like, some words may… because the sum words are a little different, but having the same meaning, blue score would still penalize those valid translations.
[01:53:17] Today
[01:53:20] So, in reality, the true quality from fine-tuning is little better, right, than the numbers alone would suggest here, in our case
[01:53:32] Okay, so let's begin.
[01:53:36] So, like, first of all, what we do is we take that validation set that we had, because that has not been touched by the fine-tuned model. And on the basis of that, like we'll try to add some ground truth versus predicted truth
[01:53:51] predicted, sentence comparison and all, okay?
[01:53:56] So first of all, we'll have our good old progress bar right batch size
[01:54:02] I have taken as 32 right now, because I wanted to get done with it first, because we have very few sentences also, and this was working out for me, okay? And then we build a data frame from the validation set for easy handling, right
[01:54:17] And this is how we do it.
[01:54:19] Like data frame we had already, we already have, right? From there
[01:54:23] We'll have this, English and French sentences, right? And then, like, this is basically a simple print statement for, like, when we run this, we'll have a look on it.
[01:54:34] Okay, I run the previous cell
[01:54:38] Okay, I did run it.
[01:54:44] Now, like, we'll start for the blue score calculation for the fine-tuned model. Okay
[01:54:50] So again, we are calling for model.eval, right? I hear the dropout will be turned off, right? And the fine-tuned predictions will be all stored in this variable, right?
[01:55:03] Okay, and
[01:55:05] Number of batches and all, we'll try to see how many we have at the end, okay? For each batch, what we'll be doing is,
[01:55:14] We'll slice one batch of English sentence from the validation set. We'll try to, you know, use the fine-tuned model to translate it, yeah.
[01:55:23] Okay? And then, you know, add it at the end.
[01:55:27] So predictions will be like one after the other.
[01:55:31] Now, all the references basically
[01:55:36] The difference was, like, for… like, remember we had imported this, Sacred Blue that will help us to calculate the blue score. For that, like, we need to have reference translation also, right? So, reference translation are basically, like, you can take it as
[01:55:55] ground truth, right? So we'll need a ground truth right? And then we'll also need fine tuned predictions also right? So that is being passed here to calculate the blue
[01:56:08] Score for the fine-tuned one. So, fine-tuned prediction is passed, or apart from the, all… also with the ground truth
[01:56:15] Right
[01:56:17] Okay, and tokenizer, like, this one we have used, INTL, because, like, I found that this was able to better tokenize the words, like, for, French and all right because French has all those
[01:56:34] You know, that E having some like comma thing like in the above, right?
[01:56:39] So, for those, like, this works better.
[01:56:44] Okay?
[01:56:45] And then we'll have a fine-tuned model score. And similarly, we'll do for
[01:56:50] pre-trained model, okay?
[01:56:54] So, we'll be reloading the pre-trained model for the baseline comparison. Pre-trained model, basically, we are saying that we'll directly take from the Marian MT, right, whatever they have, whatever the mates they had, we'll take those, right?
[01:57:10] And from there
[01:57:11] Like, you see here, like, we have done that. Modal name, we had already defined above. We had, importing that model, right? It is taken to
[01:57:20] like, device, whatever we have said, CPU in our case, then we start the evaluation process, okay? Totally well
[01:57:27] Now, we have created a small function for this, like we had done before, right? If it's a single sentence, it is converted into list as well, right? And
[01:57:39] Sorry, not a sentence. If it is a single word, right, then it is converted to list of like having word
[01:57:49] Right? Then the inputs are passed right where like tokenizer in the tokenizer we pass and this text right? And the output will be returned as Python tensor padding is goes true truncation is goes true like
[01:58:04] We did before, exactly like we did before, okay? And, we are, since we'll not be updating any weights, alright, we'll be just using whatever weights the model has, and directly generate
[01:58:18] Right? That translated tokens
[01:58:21] So, translated tokens will have those generated in these tokens.
[01:58:28] And then, from there, like, we'll try to decode those translation, translated tokens, right? And then return the translation. Now we have created a function for this one
[01:58:40] Then we'll start for the
[01:58:43] You know.
[01:58:45] like this computing of blue score
[01:58:49] So, all the original predictions are right now, this was a fun, sir. So right now we are initializing with an empty list. Then we'll again do the same thing. Like we'll, you know, slice the English sentences from the you
[01:59:04] Sorry, from the dataset, and then keep it in one batch. Send those batch towards
[01:59:13] This model, then from there, we'll have our translated
[01:59:17] There's
[01:59:19] like, our translated variable, right, where everything, our tokens and all are stored. Then what will happen
[01:59:26] like, once it is decoded and all, it'll be simply added to the translated
[01:59:33] This translated again, so that it, like, it'll keep on stacking for all those translation
[01:59:41] Okay. Now, after that, we'll compute the blue score, okay? So, original predictions
[01:59:49] all references, all references is basically like ground truth. Original prediction is basically the prediction we had from the model, okay? And then
[02:00:03] original blue score calculated. Original blue score is basically having
[02:00:10] Like, all the models, whatever our pre-trained model had predicted, okay?
[02:00:15] And then we'll do the side-by-side comparison, okay? So here we'll have a blue fine-tuned score minus blue original score. Basically, we are trying to compute the difference between those. Okay
[02:00:28] And then
[02:00:30] We'll try to have it more than zero, naturally, because we would expect our fine-tuned model to, you know, perform better, because that has been specifically trained for our use case, okay?
[02:00:43] So, like, now these are some of the
[02:00:46] like, a helpful note that I have added. Right now, we are using 200 sentences and, one epoch only, right?
[02:00:53] So that's what I, like, I tried. So
[02:00:57] You can do this for the whole data set, as I told, and you can train it to more than, like, you know, 3 books, 3 epochs, or 3 epochs was fine for me, right? You can do that. And if the results are not good, then you can slightly play with the learning rate as well. You can try to reduce it a lower
[02:01:17] Or maybe try to increase it a bit, and then see what results you have, right? If it works out for you, then you know that, you know, increasing the learning rate helps. If it didn't, then you know that, you know, increasing is not helping. We either have to stay with that one, or we go lower
[02:01:36] Now, after this, we'll do the side-by-side translation.
[02:01:41] And then, let's see. Now, these are, like, print statements that basically will just have a look after a couple of seconds. And what we do is, like, from the validation set, we'll try to have 5 random sentences, and then see the translation photos
[02:01:57] Okay, so that's what we are doing there
[02:02:00] Let's go quickly
[02:02:03] Notice
[02:02:07] So this is the bar that I was talking about.
[02:02:15] the… right now, this is, like, for the pre-trained model, right?
[02:02:19] So, it's importing again and all, the weights and all
[02:02:26] So, for the fine-tuned model, we had the score of around 24.72, right?
[02:02:32] And,
[02:02:34] like, from the, like, pre-trained, this actual one, we have 24.73. Now, the result is not significant here. It is right
[02:02:46] Because we, like, trained it on, like
[02:02:51] very, you know, less data, right? Now, like.
[02:02:58] Let me show it to you, the results that I got, okay?
[02:03:01] So when I did this for,
[02:03:04] 75,000 data, right? I found out that, you know, the results increased dramatically from
[02:03:16] pre-trained model itself, like, I was able to reach a blue score of around
[02:03:23] You know, 51, if I'm not wrong. Let me see.
[02:03:27] Yeah, right, 5351. Right? 53, right? And from Marian empty fine-tuned, like, whatever we fine-tuned, the result was, 57.60. So
[02:03:41] 5 blue points worth of improvement, just using the, like, our dataset, right? And if you see the code was, like, not so complicated,
[02:03:53] Like, if you see… if you have to code a transformer, like, by yourself, it's a huge task, okay? And it, like, there is a lot of debugging. The code is a little bit complex, right? Even if you compare it to the last tutorial of, like, having
[02:04:08] There'sms, right? So even there, we had, like, lengthy portion of
[02:04:14] codes and all. But here, if we just have to use the pre-trained model, like, it's very easy, right? What we did, we just imported the weights, and we created a function right where the input
[02:04:28] statements, sorry, input sentences were passed, and by the end, we had the translated
[02:04:34] like decoded sentences, right? It was that easy, very easy. For fine-tune, yes, we had to do something, we had to code, we had to train the model first. Based on that, we had some results at the end, right? And
[02:04:49] It was complicated, but with very less data, we got good results. Now, like.
[02:04:57] I'll try to conclude this.
[02:05:01] Okay.
[02:05:02] See?
[02:05:03] comparison
[02:05:06] The common, like, fine-tuning, like, is very close, right? This results is, like, 24.73, 24, because 200 data set is, like, 200 sentences is, like, nothing, like, close to no fine-tuning only. But,
[02:05:18] If you talk on the real data, because at least 2,000, if you had trend, then you had got, like, more better results. But right now is 200, so very similar results, right? Even fine tuned is working as
[02:05:33] Like good as pre-trained model. But
[02:05:37] Like, let's see
[02:05:39] Some examples that we can, okay? So the English sentence for Vitra is
[02:05:45] This one
[02:05:48] Original model gave this, right? Fine-tuned model gave this, and the ground truth is this. Okay
[02:05:56] Now you see this one.
[02:05:58] Let's see the
[02:06:00] This one, join us
[02:06:02] Okay? Original model gave this
[02:06:06] Now, this one fine-tuned model gave this, and ground truth is this.
[02:06:11] Right? So without fine tuning also you see
[02:06:16] It's very close, right?
[02:06:19] To the ground truth
[02:06:24] So like let's see another
[02:06:26] go
[02:06:28] is there. Then another this one
[02:06:32] And the ground truth is this one. Right? Now you see this one
[02:06:37] Like, if you see just the exact word matching, like, meaning maybe slightly different, but, you know, exact word matching is also quite good.
[02:06:45] Okay.
[02:06:46] Now, if you, like, do this for more number of sentences, naturally, fine-tuned model will be better than the original model, because the fine-tuned model actually learned from our data set. Okay, so this is
[02:07:02] Like, what I also observed.
[02:07:06] Now, if we remember the last class, last tutorial, right, I was not able to show the results. But if you had gone little down, then like if I did Lstm with attention, right? And if I did for 30 epochs in that in the last tutorial
[02:07:22] If I did it for 30 walks, and if I had taken the huge, like, or the whole data set 1, like, 75,000, then I was able to reach
[02:07:31] The blue score of 21.17 right
[02:07:34] And, without attention was somewhere around 18, right? So just by using attention mechanism, we were able to improve 3 to 4 blue points. Okay, then
[02:07:46] Like, in this tutorial, we saw that, you know, just using the Marian MT, without doing anything. Just importing the model weights, right, and passing the sentences, like, whatever we have in our dataset to it, I was able to get a blue score of around 53.72.
[02:08:02] Right? And after fine-tuning it, right?
[02:08:07] like I found, just by doing it, I was able to improve around 5 blue points, approximately, right? So, like, this is one of the major takeaways that we can take from our class, right? That, transform
[02:08:23] Because of their.
[02:08:25] architecture and, the amount of data it has been trained on, the performance increases very drastically. Like, if you see, it's almost 3 times
[02:08:36] Right? Then what we have we had with Lstm with attention. So I really wanted to show you this results because it it will help you have a grasp of why transformers are being talked about now like and like after that, LLMs have been talked about
[02:08:54] Right
[02:08:55] So this is why, because
[02:08:58] Like, even after doing the same task, right, we are able to do like a better performance in most of the cases, right?
[02:09:10] So with this, almost I think we are good to, like, end our class also
[02:09:18] I hope you guys had a good learning session.
[02:09:21] And, in most… in this tutorial, I've added comments also, like, so that, you know, when you go home, when you sit alone, you can also, like, just having a look at it, you know, you can just read it and maybe understand it, because code may be complicated, but if you get the intuition
[02:09:40] Right? You may be able to understand the concepts, and whatever concepts you have learned in class, you can relate to it, so that there is a better learning experience for all of you.
[02:09:51] Okay.
[02:09:52] Is there anything
[02:09:59] How was it
[02:10:01] If, like, this made sense from whatever you have studied in class
[02:10:14] Okay.
[02:10:17] Thank you so much, then. We have, like, almost, like, we can, leave.
[02:10:28] We can end the class for today.
[02:10:29] Unfortunately, I could not complete the
[02:10:32] the remaining tutorial
[02:10:34] Let's see if I get chance in another class
[02:10:39] But I hope, like, you guys are a good learning experience today.
[02:10:53] Thank you
[02:11:13] Okay, everyone can drop in. Okay, let's stop. Thank you.