# 04 2026-05-23 BERT Hands On

course: Module 4 — Generative AI & LLMs
module: Module-4-Generative-AI-LLMs
date: 2026-05-23
type: transcript
video_url: https://personal-learn.armco.dev/files/_Recordings/Module-4-Generative-AI-LLMs/04_2026-05-23_BERT_Hands_On.mp4

---
[00:22:49] PRATIK: Hello, good evening, everyone.
[00:22:53] PRATIK: Let's wait for a few minutes.
[00:22:57] PRATIK: So that everyone is able to zoom.
[00:26:48] PRATIK: Let me quickly share my screen.
[00:26:50] PRATIK: I think most of us are here already.
[00:26:53] PRATIK: Let's see…
[00:27:14] PRATIK: Hello, is my screen visible?
[00:27:18] PRATIK: Hey everyone, can anyone please confirm?
[00:27:22] Ankit Sood: Yes, it is visible.
[00:27:27] PRATIK: Okay, and is the text also visible? Like, is it readable?
[00:27:32] Ankit Sood: Yes, it is.
[00:27:34] PRATIK: So, hello, good evening, everyone. I would like to quickly start the session for today.
[00:27:39] PRATIK: So… You must have studied GERT in the last class, right?
[00:27:46] PRATIK: like, it's a LLM, and today, our problem statement will be revolving around that, and we'll be using BERT to solve.
[00:27:56] PRATIK: for our problem today. Okay, so the data set we are considering today is AG News classification.
[00:28:03] PRATIK: Okay? And AG News is, like, basically,
[00:28:08] PRATIK: like, it's a data set from, that is readily available in Kaggle.
[00:28:12] PRATIK: And it is a most widely used benchmark dataset that can be used for this text classification in NLP.
[00:28:19] PRATIK: And the, you know, like, it has been gathered from various sources, right? Or from around thousands of new sources, and it is very commonly used to evaluate and compare machine learning models and sort of text classification.
[00:28:36] PRATIK: Right?
[00:28:38] PRATIK: Antonio… Antonio Gulley was, like, the…
[00:28:42] PRATIK: person who compiled the dataset originally, right? And it was derived from the AG's corpus of news articles.
[00:28:50] PRATIK: Right? And, like, from around 2,000 news sources, right? And from around, like, containing, which contains around 1 million news articles. The structure is such that, first of all, we have class index, then we have title, right? Then we have, description.
[00:29:07] PRATIK: Okay? So, class index is basically telling, like, which field the news belongs to. Title is, like, a headline, you can imagine it like that. And description is the news, okay?
[00:29:18] PRATIK: So…
[00:29:20] PRATIK: like, like that, it is mentioned here also. The headline of the news article is stored in the title field, and the summary of the article is, like, in the description, right? And then, again, a numerical label is there, from 1 to 4, assigned for the class index, okay? Now, the class index is such that, like, the…
[00:29:37] PRATIK: there are four class indexes, right, belonging to world, right, world category, sports category, business category, and science intake category.
[00:29:46] PRATIK: Thanks.
[00:29:48] PRATIK: So, like, this time the dataset is, like, usually what we… till… in the last 2-3 tutorials, what we have been dealing is, like, we have a dataset, and we ourselves used to, like, you know, split in the training and test.
[00:30:01] PRATIK: data from a single data source, right? Now, this time, like, if you have noticed, like, there are… like, a zip folder has been set, right? Inside, there are two files, train CSV and, test CSV.
[00:30:14] PRATIK: Right? So, you have to import it, like, 2 times, okay? And then… Nope.
[00:30:22] PRATIK: Just a second.
[00:30:25] PRATIK: It's fun.
[00:30:27] PRATIK: So, like, in that, training, dataset, like, there are around 1,000… 1,20,000
[00:30:35] PRATIK: samples, right? And, like, if you see, like, we have done, this EDA also, then what we have noticed is, like, there are around 30,000 samples.
[00:30:46] PRATIK: each class, okay? And for test, there are around 7,600 samples, right? Roughly around 1,900 samples allocated per class.
[00:30:57] PRATIK: Both splits are perfectly balanced, meaning every class is represented equally across both states.
[00:31:04] PRATIK: Now, this dataset is usually recommended for beginners because
[00:31:10] PRATIK: As you saw, 30,000 samples is…
[00:31:13] PRATIK: You know, roughly, like, for each of the categories.
[00:31:18] PRATIK: Right? For, world, sports, business, Science and Tech.
[00:31:23] PRATIK: Right? And so you don't have to deal with, this one, class imbalance problem, right? Which is a major problem in, classification, that, like,
[00:31:35] PRATIK: Data science, person, data scientist, like, has to deal with, right?
[00:31:40] PRATIK: Known?
[00:31:41] PRATIK: And since the test, this… all the text is, like, very spo… Like, relatively short… of short length, it is very easy to have a look and understand it, right?
[00:31:53] PRATIK: And it is large enough to, like, work with LLMs also. So, this is one of the, like, easiest data sets that we can, like, deal with, okay? So, it's beginner-friendly.
[00:32:05] PRATIK: Okay? Now, as always, let's start with selecting our device, right?
[00:32:11] PRATIK: So… let me… Brickly start.
[00:32:16] PRATIK: Let's… Restart session…
[00:32:25] PRATIK: It will start.
[00:32:32] PRATIK: So, in the first cell, we will be basically selecting, like, CPU, or if you have GPU available, right, you can select which of the GPU you want, right? So, since we are dealing with,
[00:32:45] PRATIK: dual collab, and everybody is in free tier. Let's go with CPU right now.
[00:32:50] PRATIK: Right? As it is saying, device set to CPU.
[00:32:54] PRATIK: Now we'll, go to our imports.
[00:32:57] PRATIK: Like, some of the major imports that we have done is import OS, which is basically used for file path operations and reading environment variables, then import random for, like, random number generation, right? And then there is, we, as always, we have imported
[00:33:13] PRATIK: NumPy and Pandas, okay? As NP and PD. So whenever we call NP, it will be, like, like, using NumPy only, right?
[00:33:20] PRATIK: Then, for, for deep learning, like, like, a deep learning framework, or, like, using whenever PyTorch is involved, like, we, like, here we have imported Torch, import Torch, right? And then.
[00:33:34] PRATIK: we have imported toss.nn as nn, right? And then we have imported dataset and data loader from Tots.utils.data, right?
[00:33:48] PRATIK: then, like…
[00:33:50] PRATIK: from the… for the… this pre-trained language model, what we have done is, we have imported BERT tokenizer and BERT model from Transformers, right? This BERT tokenizer, like, like we had, Marianempty tokenizer last time, it basically converts raw text
[00:34:06] PRATIK: two token IDs that BERT understands. And, like, we also discussed, right, Last time that,
[00:34:13] PRATIK: Usually, like, a family of LLMs, they say, this BERT has its tokenizer, different model has its, different kind of tokenizer, right? So, for BERT, we are using BERT tokenizer. There is one more tokenizer, BERT tokenizer first, we'll be discussing it below.
[00:34:29] PRATIK: Okay, I'll just quickly… down below, I have said the difference between them also. And in this tutorial, we have used both, okay?
[00:34:38] PRATIK: For evaluation metrics, we have imported, accuracy score classification report F1 score from sklun.matrix.
[00:34:46] PRATIK: Okay? And for progress tracking, we have imported TQDM from TQDM.
[00:34:52] PRATIK: Okay, let's quickly run this.
[00:35:11] PRATIK: Okay, so all imports are successfully. If it, just in case, if, like, if anyone is running on their own, like, if any of them is not there, then you can simply do keep install, okay? Whatever module is not there.
[00:35:23] PRATIK: Now, whatever device we have selected, let's do a sanity check.
[00:35:27] PRATIK: Right, if we have selected CPU, GPU, whatever it is, this is just a sanity check to confirm if the PyTOS or the TOS that we have used below will see the correct device. So, it's right, it's showing CPU.
[00:35:41] PRATIK: Great.
[00:35:42] PRATIK: Then, in part one, we have done, like, some data exploration and preparation.
[00:35:48] PRATIK: Right? And, like, as we know, like, before any model is trained, and this data is new, like, previous to… in previous tutorial, we were using dataset that was already discussed in the previous, so, like, very less of data exploration was done there, but this time it's new, so we'll be having a look.
[00:36:05] PRATIK: At it, okay? So, how it is and what.
[00:36:07] PRATIK: So, first of all, we'll check the class distribution, just to verify if it's everything
[00:36:12] PRATIK: As we had discussed before, in the first markdown, right?
[00:36:18] PRATIK: We'll inspect a few samples, just to see if it's okay or not, and we'll confirm that the input that we'll be using to the model is clean and consistent.
[00:36:29] PRATIK: Right? Now, the AG News dataset contains news articles across four categories. World, sports, business, and scitech, like we discussed before. Each sample will be provided, each sample is provided with a title and description, right? Now.
[00:36:44] PRATIK: like, I'll show you how the title and description is, like, there in the, like, of the… a preview of the dataset, right? But it's, like, there is one title and description, so people may get confused at which one to use.
[00:37:00] PRATIK: Right? So here, what we are doing is, instead of choosing one, we are using both, okay? We are using both the title and the description, right? With the separation token in between.
[00:37:11] PRATIK: Okay, and…
[00:37:14] PRATIK: So, basically, this will be… how will the model know what is the, this description and the title? So, like, there is something called segment ID, right? If it is using title, it'll be, showing… it'll be… the segment ID for that will be 0, right? And if it is the title, right, one will be, used.
[00:37:38] PRATIK: Right? That will be the segment ID.
[00:37:41] PRATIK: Okay, so…
[00:37:42] PRATIK: what we have done here is, like, we have done trend test split, right? We have loaded trend and test, sorry, we have loaded the trend and test CSV, checked the save, like, looked at the class distribution to see the class balance or not, right? Inspected some class rows.
[00:37:58] PRATIK: Right? And, kept the title and description separate for now, and later on, it is passed to the tokenizer, right? Joined internally with the same token.
[00:38:09] PRATIK: Okay?
[00:38:10] PRATIK: Now, let's have a look.
[00:38:14] PRATIK: So here, like, like I said, like, this time we have,
[00:38:19] PRATIK: reading two different CSVs. One is the trained CSV, one is the test CSV. Now, this header is equal to 0 is basically telling that the first row is basically the names of the columns, okay? So, please don't get confused, because in the last two or three tutorials, we didn't mean this, but this time we are using this, okay?
[00:38:36] PRATIK: And, like, first of all, like, we are standardizing the volume names, right?
[00:38:43] PRATIK: So, the first one is the label, right? Label would be the category, and instead of using the names there, it has been labeled, like, 1234, right? And in title, there are different titles, like, as we have discussed, and then the description. Similarly, we have done for test.
[00:39:01] PRATIK: Data also.
[00:39:03] PRATIK: Now, these are some sanity checks, where we'll be, checking the, train dataset, like, data frame save, and the test data frame save, okay?
[00:39:14] PRATIK: And…
[00:39:15] PRATIK: This one, like, I was checking with Anne Rose is request to end it, that's why I have written this comment. Please ignore this for now. Right now, we are considering the whole of the,
[00:39:23] PRATIK: test data frame, okay? And then also we'll be looking at the class distribution, right, of train data frame, right? And let's see.
[00:39:37] PRATIK: So, the train size is, like we had discussed in the description of the dataset, it's 1,20,000, and three columns, right? Labels, titles, and description. Similarly, for test size, we have 7,600 samples, right? And with 3 columns. Label, test, and…
[00:39:56] PRATIK: sorry, label, title, and description, okay? And I'll show it to you also, a preview of train data, data frame and the test data frame, okay? So the class distribution for each of the label is, 30,000, 30,000, 30,000, like, 30,000 each, right? And these are some sample rows, okay?
[00:40:16] PRATIK: Let me also show it here. So, right? So, label is here. This is, like, basically, this row number, right? This is the label, this is the title.
[00:40:27] PRATIK: Okay, Wall Street, wall is sending us claw backed into the black.
[00:40:33] PRATIK: Right? And then again, there is a description.
[00:40:37] PRATIK: So here we can see also 1, like, 20,000 rows into 3 columns.
[00:40:42] PRATIK: Now, again, let's look at the test DF also.
[00:40:47] PRATIK: So, label is there, right, title is there, and description is very similar to what we have in the trend, and, like, it's expected behavior.
[00:40:57] PRATIK: Now, let's also confirm the, you know, distribution, like, how it is for the test.
[00:41:02] PRATIK: Right? So each of the labels has around 1,900 samples each.
[00:41:07] PRATIK: Okay.
[00:41:08] PRATIK: It's confirmed here also.
[00:41:10] PRATIK: Now, basically, the labels that we have here is, like, from 1 to 4, right? So, we'll be converting into, like, this to 023.
[00:41:20] PRATIK: Great.
[00:41:21] PRATIK: So that it is easier for us when we are using cross-entropy laws, right? And, birds classifier head basically expects zero-based levels, right? So that's why we'll be doing this. Also, remember this, like,
[00:41:36] PRATIK: to run this sale only once, if you run this twice, then it… the range of the label may change. It will start from minus 1. So, this is,
[00:41:45] PRATIK: This cell is destructive in nature, so be very careful just to run this once, okay?
[00:41:50] PRATIK: So, here we have done that, right? From train, like, label index, like, we have… from each of the columns, we have reduced one, right? And similarly, from test label also, we have done it.
[00:42:05] PRATIK: Now, like, we'll quickly verify the label range is correct or not. It's, the expected should be 0.
[00:42:11] PRATIK: should be minimum, right? And maximum should be 3, right? If it is, like, some, say, minus 1 to 2, then it is… it means that it has been run twice, right? And it is silently breaking the training, right? So if… if this happens by mistake, what you can do is you can restart the kernel and start running the code again.
[00:42:28] PRATIK: Okay? Now we'll do some sanity checks. We'll check the sample title, sample description, and label of the first row.
[00:42:36] PRATIK: Okay, and, with iLog… zero, right? What title, full description, full level.
[00:42:44] PRATIK: Let's see. So, the label range is 0 to 3, okay? And sample title is this, right? Sample description is this, and label is 2.
[00:42:55] PRATIK: Okay, so if you see here…
[00:42:58] PRATIK: Alright, this was it, right? So this was 3 initially, it got changed to 2, and this is that… this was the title there, and this was the description.
[00:43:06] PRATIK: So, this was the description.
[00:43:11] PRATIK: Okay? Exactly that.
[00:43:15] PRATIK: Now, now we'll move on to check the distribution plot.
[00:43:19] PRATIK: Right? Of how it is.
[00:43:21] PRATIK: Right? So this time you'll use, matplotlib, okay? And then, like, what we are doing is, for readability, we are creating, like, a dictionary first. Like, the zero index will be mapped to world category, right? One first index will be,
[00:43:38] PRATIK: map to sports, second to business, and third to, sighting, okay? Now, we'll apply, like, what we are doing, we'll map this label to the trend EF, okay? And add a new column called label name, okay? Now, after that, we'll count samples per class. For that, what we have done is, this, whatever,
[00:43:57] PRATIK: this column we have created just now. From there, we'll do .value counts, which will calculate the value of each of the labels, right? And then sort it with index, right? So that
[00:44:10] PRATIK: It is sorted by count, okay? It's,
[00:44:14] PRATIK: Like, sorry, it would not be sorted by count, but it would be sorted by order, alphabetically.
[00:44:19] PRATIK: Okay.
[00:44:20] PRATIK: Now, like, this is the code to plot this graph.
[00:44:25] PRATIK: Great.
[00:44:26] PRATIK: Where we have given that it's classCons.index, classCons.values, right?
[00:44:31] PRATIK: And these are some of the, like, narrower bar.
[00:44:35] PRATIK: Right?
[00:44:37] PRATIK: To improve readability and all. Let's quickly see it here.
[00:44:45] PRATIK: So, let me zoom out a little.
[00:44:50] PRATIK: So, this is the class distribution in training set, right? Business, scitech, sports, world, right? And each of them have 30,000 samples each.
[00:45:00] PRATIK: Let me quickly zoom in.
[00:45:05] PRATIK: So… like, a model setup, now we'll do model setup, okay? Now…
[00:45:13] PRATIK: like, last time also, we imported, like, a pre-trained model, MarinMT, right? This time also, we are importing a pre-trained model, right? Basically, in Hugging Face Transformers library, there are some pre-trained models, right? And tokenizers also at the same time, corresponding to the pre-trained model.
[00:45:32] PRATIK: Right? And we are loading it with the help of from pre-trained method, right? We have… Okay.
[00:45:39] PRATIK: And the modal name is passed, right? In this case, we are using BERT base uncached in this tutorial today, right? And the corresponding weights and the configuration are downloaded automatically on the first one. Next time, if you… once we have used it, next time the files will be used, it will be cached locally. So, basically, this will avoid repeated downloads in, like, if you use it again and again, right?
[00:46:02] PRATIK: And last time, we had imported BERT tokenizer only, right? This time, like, there is one more tokenizer. This is BERT tokenizer first. Basically, it, like, it also does the same thing as, BERT tokenizer only.
[00:46:16] PRATIK: But this is a little faster, okay? So, like, how it works is, basically, it converts the, like, a normal tokenizer, it converts the raw text into BERT,
[00:46:25] PRATIK: like a format that a bird can process, right? Its input string is broken down into sub-word units called word pieces tokens, right? Special tokens are added automatically, like, we have two special tokens, right? CLS and SEP token, right?
[00:46:41] PRATIK: So, CLS token is at the beginning of the sequence, right? And after that, after every sequence, right, if there is, like, let's say a two sequence, right, two sentences, then CLS will be at the beginning, then after each sentence ends, there will be a SEP token.
[00:46:58] PRATIK: Now, apart from this, our tension mask is also produced, which tells, like, you know, where the real tokens are, and, like, it also helps the model to understand, like, which of the tokens are the real tokens, and which of them are the padding tokens. So, padding tokens, like, are usually avoided by the model, so that it can
[00:47:18] PRATIK: Not impact the training, okay?
[00:47:22] PRATIK: Now, as we had discussed, there's a fast and slow tokenizer, right? So, there is one BERT tokenizer fast also. So, basically, BERT tokenizer fast is backed by Rust.
[00:47:32] PRATIK: Right? It is, very… like, since it is…
[00:47:36] PRATIK: backed by Rust, it is a little faster, and the other one is written in pure Python.
[00:47:41] PRATIK: Okay? So, like, we have used both, but right now, like, we have… like, I have used past tokenizer, right, and a normal tokenizer in this, so we'll use both.
[00:47:53] PRATIK: Okay? And both bersena will produce identical autos, so there is no such issue, right? And the… so, since they'll produce
[00:48:01] PRATIK: identical outputs, the model VA will also be relatively same.
[00:48:06] PRATIK: Okay?
[00:48:08] PRATIK: Now, let's understand what bird-based uncased means. Now, there are various bird-based, bird bottles, right? So, there is bird-based uncased and bird-based cased, okay? So, bird-based basically refers to the two standard bird sizes, right?
[00:48:25] PRATIK: 12 transformer layers, 768 hidden dimensions, with approximately 110 million parameters. There is one more bulb, right? It has 24 transformer layers, and then there are more parameters in it.
[00:48:40] PRATIK: Uncased basically means that all the input text
[00:48:44] PRATIK: will be lowered, before that organization. So, words like this, Paris with a capital P, will be treated same as Paris with a small p, right?
[00:48:55] PRATIK: And, this variant is generally preferred for tasks where capitalization does not carry meaningful information, which is, like, in the most, case for classifications.
[00:49:06] PRATIK: Right? We are basically more interested in the context, okay?
[00:49:13] PRATIK: So, now let's do some, text length analysis, okay?
[00:49:18] PRATIK: Where we'll try to see, like.
[00:49:22] PRATIK: How, after combining this title and description together, like, are we crossing the…
[00:49:30] PRATIK: birds, context limit or not. So, birds,
[00:49:34] PRATIK: context limit is somewhere around, like, 512 tokens. We'll just try to have a look at it. Okay? So, for this, what we have done is, we have loaded a fast tokenizer.
[00:49:43] PRATIK: Right? For Ida, right? And from, so for that, what we have done is, from transformers import BERT tokenizer first, which is very similar. There, it was just BERT tokenizer in the slower one, okay?
[00:49:56] PRATIK: So, like, now we have imported it.
[00:50:00] PRATIK: from pre-trend, right? Bird base, uncased, and, like.
[00:50:06] PRATIK: Then what we have done is, we have created a, like, a copy of training data and saved it as… and we'll be reflecting as EDARDF, right? So that the main data is not
[00:50:19] PRATIK: we are not, you know, corrupting it in any way, or, you know, polluting it any way, like, whatever you can say. So…
[00:50:28] PRATIK: After that, what we are doing is, we are combining the title, Right? And the description.
[00:50:34] PRATIK: Right? With a space.
[00:50:36] PRATIK: Right? Then we'll try to see the word count for it, with the help of this.
[00:50:41] PRATIK: Right? Dot apply lambda X, then for each, we'll do the X… we'll take the X or split, right, and then check the length, right?
[00:50:51] PRATIK: This will basically do the word count based on white space, okay? Now for… after that, we are doing character count. So whatever word was found out, right? Whatever word was taken from there, again, we are doing…
[00:51:04] PRATIK: we'll do length for each of the words, right? And based on this, for each of the rows, we'll have a raw character count.
[00:51:11] PRATIK: Okay.
[00:51:12] PRATIK: Now, we'll try to see.
[00:51:15] PRATIK: the token count, okay? Right, before this, we were saying word count and character count. Now we'll try to see, this token count. For this, what we have done is.
[00:51:27] PRATIK: on this EDA underscore DF, like, the copy of trade data, we are applying this lambda function, where tokenizer, EDA is being used upon this title and description of our
[00:51:41] PRATIK: data.
[00:51:42] PRATIK: Okay.
[00:51:44] PRATIK: And… Like, with XS equals to 1, and, like, input IDs are basically telling, like, what…
[00:51:52] PRATIK: what it is, right? Which of the index it is, somewhere… it is giving that idea, okay? Now, we'll simply print
[00:52:00] PRATIK: Right? Some of the statistics, it is for that, like, whatever word count, CAD count, and token count we found out, we'll try to see it.
[00:52:09] PRATIK: Okay, this is for that. These are some print statements. We'll be looking at the output below.
[00:52:13] PRATIK: Okay, and if any of them exceed the token limit, right, if any of the rows exceed the token limit, those are, like, those are counted here, and taken some, right?
[00:52:24] PRATIK: Similarly, if there are over 128… basically, in this tutorial, we are taking 128 hours, hour.
[00:52:30] PRATIK: Limit.
[00:52:35] PRATIK: Just a second. Someone has… Raised…
[00:52:39] Ankit Sood: Yeah, I have a question. So, I mean, what are we trying to do by combining this sentence… by making this sentence and actually
[00:52:48] Ankit Sood: identifying… The number of characters and number of.
[00:52:53] PRATIK: Yeah, so let me…
[00:52:54] Ankit Sood: return.
[00:52:54] PRATIK: You can answer that.
[00:52:56] PRATIK: Just a second.
[00:52:59] PRATIK: So, we had, in the dataset, we have this title and description, right? Initially, we sawthat. So, what we are doing… what we are doing is, we are trying to use this both at the same time.
[00:53:11] PRATIK: Okay.
[00:53:12] Ankit Sood: Okay.
[00:53:12] PRATIK: So that the model is able to see both the title and description.
[00:53:16] Ankit Sood: Okay?
[00:53:18] PRATIK: But at the same time, what we are… there is a context limit for BERT, right? It can only.
[00:53:22] Ankit Sood: Yeah.
[00:53:23] PRATIK: It's 512 tokens at a time.
[00:53:25] PRATIK: So, once we combine this, we are also trying to see if that is exceeding the BERT, this context limit or not.
[00:53:32] PRATIK: So, if I.
[00:53:33] Ankit Sood: Goodbye.
[00:53:33] PRATIK: Sorry, come.
[00:53:35] Ankit Sood: We are taking character then, because token is…
[00:53:39] Ankit Sood: more or less a word, right? Or maybe a string of characters. But we are taking character size also.
[00:53:47] PRATIK: Sorry, sorry, I didn't get you. Come again.
[00:53:49] Ankit Sood: So, the example that you shared, right, there was… we were taking the word size by splitting it via white spaces, and we were also taking the character size also. But how many characters are present in this? So, how character is going to help us, because we have.
[00:54:03] PRATIK: No, no, no. We are just seeing, like, that was done for exploratory purpose, right? We are just trying to explore. Like, we're just trying to see, like, what all is there. Like, what is the nature of the data, right? Because from words only, tokens are being created and all, right? So we are just trying to understand that.
[00:54:19] PRATIK: It is not directly helping you, we are just trying to understand. We are going step by step.
[00:54:23] Rashmi Ranjan Singh: Okay, cool.
[00:54:23] PRATIK: We are trying to see the data, and then we are, you know, trying to see, like, if this is… because later on, those words itself will be converted into tokens, right?
[00:54:33] PRATIK: So…
[00:54:33] Ankit Sood: Word spoken, I understand. I was a little confused about the characters that… what advantage characters are…
[00:54:39] PRATIK: Explore entry purpose, yeah.
[00:54:40] Ankit Sood: Oak.
[00:54:41] PRATIK: Nothing as such, okay? We're just trying to see.
[00:54:45] Ankit Sood: Thank you.
[00:54:47] PRATIK: Any other question?
[00:54:58] PRATIK: Missouri.
[00:55:06] PRATIK: So, yeah, here we will try to see, like, how many of the tokens, means how many of the rows actually crossed the 512 token mark, right? And how many crossed 128. In this, we have taken 128.
[00:55:19] PRATIK: Right? In our tutorial, okay? We'll try to see.
[00:55:23] PRATIK: And then… What is the actual length of the data?
[00:55:28] PRATIK: Right? Data set. So, we'll… now, this is a plot, okay, simply to reflect all of this. Let's quickly have a look.
[00:55:50] PRATIK: So here it is tokenizing also, right? So that's why it is taking some time. Let's wait for a couple of seconds.
[00:56:25] PRATIK: So, basically, if you see, we are looking at 1, like, 20,000 samples, right? Around 37 words are there in average, right? And 236 character in average, and token count is somewhere around 54.16 in average, okay?
[00:56:41] PRATIK: So that's average. Now, we want to consider as many as we can, so that, like, we cover maximum of the rows, right? We don't want to miss out on any context. So for that, we have done this. Now, this here is the 512 mark, so we can see that none of the sentence have crossed that.
[00:57:00] PRATIK: Right? Now, for the… for the tutorial purpose, we have taken 128 token here, right?
[00:57:05] PRATIK: So, we see that some of the, you know, sentences have crossed that. To be precise, 956 rows.
[00:57:12] PRATIK: I crossed it, right? That's about 0.80, right? So, we can see, simply, even if we keep 128, it won't affect much. Like, that's a trade-off that we'll be accepting.
[00:57:26] PRATIK: Okay.
[00:57:31] PRATIK: Now, we'll.
[00:57:32] Ankit Sood: We've got some…
[00:57:32] PRATIK: samples? Oh, sorry, yeah.
[00:57:34] Ankit Sood: Question is, but it can still fit our token limit size, right? Then why we are actually ignoring those?
[00:57:39] PRATIK: We didn't know this, right?
[00:57:41] Ankit Sood: After knowing it only we came to know, right?
[00:57:43] PRATIK: Initially, we didn't know it would cross 512 hours.
[00:57:46] Ankit Sood: Yup.
[00:57:47] PRATIK: Now, we have taken 128 because
[00:57:49] PRATIK: Because for smaller, there are many smaller sentences also, right? So if we… so if we take a huge, this token, this,
[00:57:59] PRATIK: Like, consider more tokens, right? Let's say there is a smaller token, a smaller sentence, okay? When it is converted, it may not reach 128, then we have to do padding. So we have to take unnecessary.
[00:58:10] Ankit Sood: Okay.
[00:58:11] PRATIK: happens every time, right? So, it is.
[00:58:13] Ankit Sood: Excuse me.
[00:58:13] PRATIK: basically, take unnecessary time, also. So this is a trade-off that we are accepting, that, you know, 128 is okay. If you want, you can take it more than that.
[00:58:22] Ankit Sood: Got it.
[00:58:25] PRATIK: Okay.
[00:58:25] Ankit Sood: And it… it can be, like, 130, 130, I mean, it can be any number, but…
[00:58:30] PRATIK: It's been 5 minutes.
[00:58:31] Ankit Sood: 112, right? It doesn't have to be multiple of 2, or power of 2.
[00:58:35] PRATIK: No, no, nothing.
[00:58:35] Ankit Sood: Nothing like that. Yeah. Okay.
[00:58:43] PRATIK: So, next, we'll do, like, sample inspection, right, per classes. We'll just try to see how many samples are there, right?
[00:58:51] PRATIK: Sorry, we'll try to look at some of the samples, right, from each of the category, just to have a look at it, okay? So right now, we'll be looking at two samples from each of the category, right?
[00:59:04] PRATIK: And, from this, we had, created one label map, right, which was defined in cell 2A, where it, there was a dictionary setup, like, 0 for World, 1 for sports, 2 for business, right? And 3 for SciTech, right? From there, we are doing dot items and trying to access, right?
[00:59:23] PRATIK: So, from there, we are trying to, like, what we'll do is, like, from the trend DF, we'll use that label index, right, and then try to sample.
[00:59:32] PRATIK: Okay,
[00:59:33] PRATIK: This number of samples is set to here, so we'll be, taking two samples each. With a seed, we have… a random state we have taken as 42, right? So that it is reproducible, okay? And we'll be considering this title and, description. We are only keeping two. And this value is to basically convert into a NumPy array.
[00:59:53] PRATIK: Okay, now we'll, display this title description, right? And this title and description here, we have, segregated with the help of this bar, okay? Now we'll print it. Let's see.
[01:00:09] PRATIK: So, from World, we have two samples here, right? Explosion Rocks, Baghdad neighborhood, right? BBC reporters, blog, right? And these, this much is the title, and after that is the description, okay?
[01:00:23] PRATIK: This is just for, you know, just to have a look at it, how each of the samples are for each category, right? So this is for world, this is for sports, right?
[01:00:35] PRATIK: This is the title, and this is the description.
[01:00:38] PRATIK: This is the title.
[01:00:39] PRATIK: This is a restriction. Similarly for business. U.S. house sales fall in July, right? And then this is a restriction. And similarly for SciTech also. Gartner optimistic about cheap numbers, and then there is a description for it.
[01:00:53] PRATIK: Similarly, fossil indicates groundways went south, right? And then the this description.
[01:01:00] PRATIK: Now, what we have… now, this is very important, okay?
[01:01:04] PRATIK: So this is important because if you want to run this by your own, in home, it's for that. Right now, we have used a very small
[01:01:14] PRATIK: data, right? Or a sample from our original data set, due to…
[01:01:19] PRATIK: this time constraint, because we cannot do this. It will take lots of hours on CPU, so if you have CPU also, it's fine, but it will take a lot of time, right?
[01:01:30] PRATIK: So, I'll just tell you about the parameters that I have considered.
[01:01:34] PRATIK: So, if you want to train it on the full dataset, and full training set, you would have to do nothing, but you have to change this true, okay? Once you change this as true, everything is adjusted in the code, like, basically, it'll use the full data set.
[01:01:49] PRATIK: trained subset will use the train dataset. Now, if this is false, right, then whatever samples you have suggested here, right, that will be taken. So.
[01:01:59] PRATIK: since we don't want class imbalance, right? So we cannot directly take first 100 or first 200, right? So there, like, we don't know if they are appearing properly in, like, in a proper distribution or not.
[01:02:12] PRATIK: So, to avoid that, what we have done is, we have taken, like, 25 samples each, right, from each category, right, for a class.
[01:02:20] PRATIK: So… and this is for similarly from the test set. We have not taken 1,900 test samples, right? From there, we'll… we are selecting only 25 per class, and then evaluating it upon. So, this is right now due to the time constraint and all.
[01:02:33] PRATIK: later on, you can, if you want to do it on full, right, you can do it true. Even if you don't want to do it on true, you're finding it's too long, then what you can do is you can just change the sample per class and test sample per class. Let's say if you take 500 here.
[01:02:47] PRATIK: Right? Then you'll have, like, 2,000 subset, right? And then if you want to test it on, let's say.
[01:02:55] PRATIK: 100 per class.
[01:02:57] PRATIK: test samples, right? Then you'll have 500 test samples.
[01:03:01] PRATIK: Right.
[01:03:02] PRATIK: So, this is per class, just remember that.
[01:03:06] PRATIK: Okay, so like I said, if we… if full dataset is true, then we'll make a train
[01:03:13] PRATIK: this trend dataset copy and, like, reference it by, trend subset, and you'll use it by trend subset, copy of it that is there, okay?
[01:03:20] PRATIK: And… a lake.
[01:03:23] PRATIK: If not, then we'll basically sample… use, we'll sample with the help of this lambda function, right? And reset the index, right? So that grouped indexes
[01:03:35] PRATIK: Flattened and all, right?
[01:03:39] PRATIK: Okay? Similarly, we are doing for test subset.
[01:03:42] PRATIK: If full dataset is used, then you will be using the full test set. Sorry, yeah, if you are using the full dataset, then you'll be using the full test set. If not, then whatever samples you have defined here.
[01:03:53] PRATIK: We'll be sampling that from the test set and using it.
[01:03:57] PRATIK: Okay.
[01:03:58] PRATIK: Now, from the 10, this,
[01:04:01] PRATIK: this trendset, right? We are using trend test split and splitting into trend subset and validation subset, okay? So, what you have done is.
[01:04:12] PRATIK: In trend test split, we have given input trend subset. Test site, we have taken around 10%, right? Random state 42, and stratified train subset label. That means, basically, we are telling that
[01:04:23] PRATIK: preserve the, this balance in both the halves. We are telling them, you know, the trend is split, that, you know, when you are splitting it, make sure that each of the halves are relatively balanced.
[01:04:34] PRATIK: Okay? And we'll be verifying that below also.
[01:04:39] PRATIK: Okay? So, like, after that, we'll be resetting the index, okay?
[01:04:45] PRATIK: So that it is cleanly zero index, right?
[01:04:50] PRATIK: Okay, and then we'll be verifying the train size, validation size, and test size. Okay.
[01:04:56] PRATIK: And then we'll also see the class distribution in trend, like, for each, right? In class distribution in validation, in-class distribution in tests. Let's quickly run this, I'll just quickly check if, yeah, I run this too short one.
[01:05:11] PRATIK: So, 25 samples per class we have used, right? So that makes around 100 total rows for training, 25 samples per class for testing, so that makes around 100 total rows, right, for testing.
[01:05:22] PRATIK: train size is somewhere around 90, right? Because 100… because 10 is given to validation, right?
[01:05:30] PRATIK: And test size is around 100, pink.
[01:05:36] PRATIK: Okay, now, class distribution. Class distribution, if you see, for 0th category, first category, second category, or third category, you have 22, 23, 22, 23, like, they have tried to, and whatever is remaining from the 25, that has been given to the validation, 3, 2, 3, 2.
[01:05:52] PRATIK: Right? For each of the categories.
[01:05:55] PRATIK: And similarly, if you see in,
[01:05:58] PRATIK: class distribution in test, each of the category, each of the class has 25, yeah. Because we, in four tests, we don't have to do validation split, right? That's fine.
[01:06:06] PRATIK: Everything is considered.
[01:06:11] PRATIK: Now, here we'll be loading the bulk organizer.
[01:06:15] PRATIK: So, here we are not using the… a fast one, okay? We are just using normal. I just wanted to show, like, you know, both are working similarly.
[01:06:27] PRATIK: Now, like, from pre-trend, we have, like, using boiled-based encased, just like before.
[01:06:34] PRATIK: Right?
[01:06:36] PRATIK: And, we'll be using this on a sample text.
[01:06:39] PRATIK: Just to see how it is working.
[01:06:40] PRATIK: So, let's say sample text is, BERT is a powerful language model, and sample text is passed.
[01:06:46] PRATIK: And where, tensors are returned as PyTorch, we don't want numpy errors here, right? Because we'll be dealing with training, and we are using, PyTorch for that, right? And padding is max length. Max length, we are, taking 128.
[01:07:01] PRATIK: Right? And, but here we have defined 10 specifically just to see, because, to preview this, it'll be hard, right? 128 tokens, so we have just considered 10 here, okay?
[01:07:13] PRATIK: Now, we'll just see the token IDs, attention mask, and tokens.
[01:07:17] PRATIK: Just to have a look how Tokenizer is working exactly.
[01:07:20] PRATIK: Now, if you look here… The sentence was this, right?
[01:07:27] PRATIK: This one. BERT is a powerful language model.
[01:07:30] PRATIK: So…
[01:07:32] PRATIK: Yeah, one more thing I forgot to tell you. Okay, so C, for ID for CLS is 101, and for, SAP token is 102, right? CLS is, basically prepended, it's in the beginning of, start of the each sentence, right? And,
[01:07:49] PRATIK: shape is appended after a sentence ends, or boundary between the two sentences, okay?
[01:07:56] PRATIK: And attention mask basically, tells us which one is the real token we have to attend to. And 0 basically tells it is a padding token, which it can be ignored.
[01:08:06] PRATIK: Right? So we are trying to see that. And after that, we'll be trying to, you know, like, we'll try to convert it back to
[01:08:14] PRATIK: Human-readable subwords.
[01:08:16] PRATIK: Okay.
[01:08:17] PRATIK: So, now let's see. So, 101 for CLS token, 102 for SIP token.
[01:08:24] PRATIK: 0 for padding, right?
[01:08:26] PRATIK: Now, let's see here.
[01:08:28] PRATIK: Like, 111 is basically telling that these are all the relevant ones. 0, because at the end, we have this padding token, okay?
[01:08:37] PRATIK: Now, similarly, if we, you know, convert it back to human-readable subwords, then you can also see that in front, it was added, CLS was added. After the end of the sentence, save was added. SAPE token was added. And once, to match the padding length, this maximum length, 10.
[01:08:54] PRATIK: Right? A padding token was added. So this is how it works. Now, if you take a huge
[01:09:00] PRATIK: let's say context, right? And max length we took it around, let's say, 512. Then there will be so many padding tokens at the end, right?
[01:09:08] PRATIK: So, that's a computation that it needs some… even for this, it has to know, right? Like…
[01:09:15] PRATIK: After this, it is there. So there is some competition involved.
[01:09:18] PRATIK: Right? So, we try to take the appropriate, this maximum length.
[01:09:24] PRATIK: Now, like, moving on, we'll be trying to define the custom, this PyTorse data, dataset class, okay? Now, basically, in PyTorse, we require two methods. One is length, and another is get item. Length is basically telling, tells us to
[01:09:41] PRATIK: When we use length method, it is, like, basically returning the total number of the samples in the dataset, and
[01:09:47] PRATIK: gate item basically returns a single tokenized sample by index, okay?
[01:09:52] PRATIK: So for this, we have defined, this class.
[01:09:56] PRATIK: AG News dataset, okay? Then we have defined this init method, where
[01:10:02] PRATIK: We are, first of all, resetting index. Okay.
[01:10:06] PRATIK: to… Basically.
[01:10:09] PRATIK: So that we can use this .log index, okay, with a clean zero-based index, integer index, because it may have, index from before also, right? It is taking from the trend data set and all, right? And then we are initializing the tokenizer and the max length, right, per sample.
[01:10:29] PRATIK: Okay? Then…
[01:10:30] PRATIK: For the length of the data, we have done this length.lin, and then save.data, right? How many batches to generate per epoch.
[01:10:41] PRATIK: then there is gate item, if you have to pull, like, a data
[01:10:45] PRATIK: From it, so… so it'll have the, this title, description, and label. Okay.
[01:10:53] PRATIK: So, like, this is retrieved separately, so that we can insert the saved token between the
[01:11:00] PRATIK: correct token type IDs, right?
[01:11:05] PRATIK: then… What we do is, we'll try to encode this.
[01:11:10] PRATIK: Right? Right now, it's into words and all, so we'll try to encode this. For this, what we have done is, we have passed the title and description together, right? And then, max length is considered, right? Padding…
[01:11:24] PRATIK: is again taken by max length only. Truncation only second, basically, it means that if we have to truncate, if it just in case it crosses the 128 mark, then we'll not truncate the title, we'll truncate the description.
[01:11:37] PRATIK: Because, this is a design choice, because, like, if we read the title, even in newspapers, we have the title, right? By title, we understand more, like, somewhat we have some idea, right? So, instead of, truncating title, we are truncating description first, okay?
[01:11:56] PRATIK: And then we are returning is the,
[01:11:58] PRATIK: this Python, PyTorse tensors.
[01:12:03] PRATIK: Okay? Then at the end, these input IDs, attention marks, and token type IDs and label is returned. Now, like, token type IDs is basically, like a segment ID, right? Where that we had discussed in the beginning, right? It is basically, since we are using,
[01:12:21] PRATIK: this one, both title and description, right? So… it basically…
[01:12:28] PRATIK: helps both distinguish, right, which of them is, from the title segment, which of the token is from this title segment, and which of them is from the description segment. So here, title segment for token type ID, for, title segment is
[01:12:45] PRATIK: Like, we have taken as 0, right? And for the description segment, we have taken as 1.
[01:12:50] PRATIK: Right? And this is how Boyd was also trained initially, right?
[01:12:57] PRATIK: Now, we'll do some smoke test if, like, everything is okay or not, right? So, we are basically creating a sample dataset, right? Passing the train subset and the tokenizer, right? And then we'll try to look at the first instance of it.
[01:13:12] PRATIK: Right? Then we'll look at, input ID shape, attention mask shape, token ID shape, and label.
[01:13:18] PRATIK: Right? And then for each of the token type ID, we'll inspect the pattern, like, how it is.
[01:13:24] PRATIK: Okay? And then we'll try to decode the first sample virtually, verifying if a CLS token is there or not, safe token is there or not. Okay, so I'll quickly run this one.
[01:13:37] PRATIK: So, input ID shape is 128, because we took 128, okay? Then, attention shape is 128, token type ID is 128, okay? And then, level is 0, okay?
[01:13:48] PRATIK: Then?
[01:13:49] PRATIK: Segment A tokens. Basically, this will include title plus padding, right?
[01:13:56] PRATIK: And segment B tokens is basically the description.
[01:13:59] PRATIK: Okay. Now, if we, now, right now, what we have done is we have decoded the first 120 characters, right? So, it's like this. CLS, the China Post.
[01:14:09] PRATIK: Right? CLS is basically the sequence is starting, the sign-up post is the…
[01:14:14] PRATIK: like, this title, right? And then.
[01:14:19] PRATIK: Like, after that, the description is starting.
[01:14:22] PRATIK: Okay, so since we have only considered the first 120, tokens, so basically, the first 21 characters, so it's only coming to here, if you remove this.
[01:14:32] PRATIK: Then we'll have… The full thing.
[01:14:37] PRATIK: See?
[01:14:41] PRATIK: Right.
[01:14:42] PRATIK: This is, like, sentence ends, and after that, padding is there, like this.
[01:14:46] PRATIK: So, I'll keep it 121.
[01:14:52] PRATIK: Okay.
[01:14:53] PRATIK: Now, from the…
[01:14:54] PRATIK: Now, we have already created a data loader. Now, we'll be, sorry, we have created this dataset class, okay? Now, then we'll be going towards the, data loaders.
[01:15:05] PRATIK: Okay? Now for that, our, basically main goal is here to
[01:15:12] PRATIK: wrap each dataset in Data Loader for this training, validation, and testing, right? And Data Loader will basically help us
[01:15:21] PRATIK: for batching, shuffling, and loading, okay? Now, there is one, like, this one I have used because… okay, I'll explain this. So, basically, when data loaders are used, right, and then we have one parameter called number of workers there, which,
[01:15:38] PRATIK: Causes, like, forces… right now, in our case, what is happening is, it is forcing parallelism off.
[01:15:45] PRATIK: Right? Because Hugging Face also has,
[01:15:49] PRATIK: this Rust-backed thread pool internally, okay, which also helps in parallelism, so it's… it's somewhat conflicting there, okay? So we have… what we have done is…
[01:15:58] PRATIK: We'll be trying to, use, close that, okay? So for… force… Okay.
[01:16:04] PRATIK: So let's see.
[01:16:06] PRATIK: So for that, what you have done is, we are doing import OS, right? And then, tokenized parallelism, we are selecting this and setting it as false, okay? Then batch size, we have considered 32. If it is too much, right, for the GPU that you are running, or the CPU you are running, you can reduce it.
[01:16:23] PRATIK: The training time may be a little longer, but at the same time, you'll be using less of the memory, so it's okay.
[01:16:31] PRATIK: So, from here, we instant, like, we'll be… we'll be instantiating the… this dataset objects. So, first, we are using for train subset. Second, we are doing for validation subset.
[01:16:43] PRATIK: Right? And then third, we are doing for test subset. And then tokenizer is same for all these three, so we are passing it to the train dataset, validation dataset, and test dataset.
[01:16:51] PRATIK: Now I'll be wrapping this in… data loader
[01:16:55] PRATIK: Okay, so for, train loader, validation loader, and test loader, we have passed train dataset.
[01:17:02] PRATIK: in the train one, for validation, we have passed the validation dataset, this well dataset, and then for the testing one, we have passed the test dataset, right? Bath size is already defined 32. Suffol is true for training, because we don't want the model to memorize the sample order, right?
[01:17:19] PRATIK: And number of workers.
[01:17:21] PRATIK: We have set it as 0, because to avoid that parallelism and all, okay, that we just discussed, okay?
[01:17:28] PRATIK: So… and for validation and test, we are setting the sample for false, because order will not matter much here, right?
[01:17:37] PRATIK: Okay.
[01:17:39] PRATIK: Now, we have a defined number of workers for all, this zero for all loaders, right? Because if we are using more than that, this is spawning separate, like, call gate item in parallel. And this Hugging Face tokenizer disables their internal thread pool when forked, right? So, to avoid this.
[01:17:58] PRATIK: Which means, worker… when we are doing this, worker would actually run slower than the main process.
[01:18:03] PRATIK: So, for this, we are setting num workers equal to zero, okay? Now, we'll quickly verify all the batch counts.
[01:18:13] PRATIK: So, training batch is 3, right? Val batch will be 1, test batch is 4, right? And batch size is 32.
[01:18:22] PRATIK: Like we had done. Now, finally, we are on to our pretend BUT model. Now, I'll tell you quickly what, like, the architecture that we have taken into consideration here.
[01:18:34] PRATIK: We have done basically two experiments. In the first experiment, what we have done is, we have taken the pretend bug, we have done… taken it as it is, right? And then we are putting a classifier on top of the CLS token there, right? And, since we just want to classify, right? So…
[01:18:53] PRATIK: We are… the… the…
[01:18:55] PRATIK: So, we are training a small classification layer just on top of the CLS token to have, to the classification task. So, the number of parameters that will be trained here is very less, right? So…
[01:19:11] PRATIK: And when we are doing this, it's basically we do this, in cases where we have very less data.
[01:19:19] PRATIK: Right? And with small data, maybe the model may not learn also properly, right? And the…
[01:19:27] PRATIK: Modal weights, right, that the pre-trend BERT already had.
[01:19:32] PRATIK: May get disturbed also.
[01:19:34] PRATIK: So, in that case, what we do is, it hurts more than you know, benefiting. So…
[01:19:41] PRATIK: So just to see, like, how pretend buy it.
[01:19:45] PRATIK: with a classification layer on top of it is working. And in the second experiment, what we have done is,
[01:19:52] PRATIK: Like, we have trained the whole network.
[01:19:55] PRATIK: from classifi- we have trained BERT also, and then the classification layer on top of it also.
[01:20:01] PRATIK: And… and try to see the performance of it. And at the end, we have done the comparative study. Now let's have a… quickly have a look at it, okay?
[01:20:09] PRATIK: no… Basically, pretend board means weights have been already learned on a large surface, and we are
[01:20:16] PRATIK: The modal weights are saved, right, in logging phase, and we are simply importing it.
[01:20:20] PRATIK: Right? So, now…
[01:20:23] PRATIK: we already know there is CLS token and all, right? And CLS token basically tries to capture the essence of the… all the sentence, right? We are hoping for that, and let's see if it's useful for our, like.
[01:20:38] PRATIK: Case or no?
[01:20:40] PRATIK: Okay, so this is what I just told you.
[01:20:43] PRATIK: we have used bird-based uncased, right? It's a 2L transformer layer with,
[01:20:49] PRATIK: 12 attention heads, and the hidden size at the end is of 768 dimension. Approximately, it has about 110 million parameters, right? And the uncased,
[01:21:01] PRATIK: the encased, like, it basically means that text is lowercase before the tokenization. So, BERT and the small BERT will be treated as same. Two special tokens are used, the CLS token and SAPE token. CLS at the start, SAPE at the end, okay? The example is this.
[01:21:19] PRATIK: So, if we say 52 tokens are passed, right, then the output save will be around 52 cross 768, meaning each of the tokens will be represented in 768 dimension, right? So, hence this output.
[01:21:34] PRATIK: Okay. Now, to access the CLS token, what we have done is, we have taken this from the output last written state.
[01:21:42] PRATIK: when we do this, we have access to the, this CLS token.
[01:21:46] PRATIK: Okay? So the output will be whatever bad size we have considered, cross 768.
[01:21:52] PRATIK: Okay, and then the linear classifier head that we have used is somewhat like this. Okay, 768, because it is taking the… its input will be the hidden step from the BERT, right, which is of 768 dimension, and output, we have 4 categories here, that's why 4.
[01:22:09] PRATIK: Okay.
[01:22:11] PRATIK: So, let's see the flow ones. Input text will be passed to the tokenizer. From tokenizer, we'll have CLS input IDs, including the SAP token, right? Plus the attention mask, which will be passed to the BERT encoder.
[01:22:24] PRATIK: right, of 12, layers, transformer layers, right? From there, we'll have the CLS token vector, which is of 768 dimension, which will be passed to the linear classifier layer, right? And, which will do… which will help us
[01:22:39] PRATIK: For the classification of the four, categories.
[01:22:43] PRATIK: Okay, world, sports, business, and sightseeing.
[01:22:46] PRATIK: Now, in Frozen BERT, basically, encoder weights are fixed, right? Only the linear classifier is trend, and in the fine-tune, everything, end-to-end is trend, right? 110 million parameters from… plus the 3076 parameters from the linear classifier, right?
[01:23:02] PRATIK: Okay.
[01:23:03] PRATIK: So, firstly, we'll define the frozen bird classifier model, okay? So, this is a…
[01:23:11] PRATIK: class we have defined, which is inheriting the, inheriting this nn.module, right?
[01:23:16] PRATIK: We have defined the method.
[01:23:19] PRATIK: And this is basically, apart from its init, it is also, it will also initialize whatever the init, NN.module has had before, right? And from there, we'll load the… after that, we'll be loading the ART encoder, right, from .
[01:23:37] PRATIK: Byrdmodel.from, pre-trained, byt-based encased, okay? Like, just like before.
[01:23:43] PRATIK: After that, we'll be freezing all the parameters. Since here, we'll be only dealing with the linear classified, we only want the way it's updated for it, right? For that, we have done this. Parameter requires grad is equals to false, right?
[01:23:58] PRATIK: Now, this is the classification head on top, right? For that, linear classification layer is there.
[01:24:06] PRATIK: like, which has hidden size, it's from the, like, board, 768 dimension, right? And number of classes.
[01:24:15] PRATIK: like, that we have is 4, okay? Now, one more thing, if this was a bird large, right, then the hidden size would be 1024.
[01:24:26] PRATIK: Okay? In our case, it's 760. Just, just to give you some idea.
[01:24:32] PRATIK: Okay? Now, in forward, like, now we'll have some, this forward method, right? What we are doing is…
[01:24:40] PRATIK: We'll be passing the input IDs, attention marks, and token type IDs. Token type ID is that segment ID, which I said, 0 for title token, and one for description, right?
[01:24:51] PRATIK: output, okay? And then from those output, we take the last written state, right? And from there,
[01:24:59] PRATIK: we take the last region, which is… and for the CLS token, right? And then pass it to the linear classifier layer, right?
[01:25:06] PRATIK: self.classify and CLS output is passed, and whatever logic we get, we return it.
[01:25:11] PRATIK: So this is, like, what we have defined as a class.
[01:25:15] PRATIK: Okay? Now, we'll be creating one object from this, where we are setting this number of classes equal to 2, and whatever device.
[01:25:23] PRATIK: We are using, it'll be…
[01:25:25] PRATIK: you know, all the parameters will move to that device, okay? Now, quickly, we'll see the, trainable parameters.
[01:25:33] PRATIK: Okay.
[01:25:35] PRATIK: we have for here. For BERT, how many parameters are there? And for, like, how many are we actually training in this?
[01:25:43] PRATIK: For this experiment.
[01:25:45] PRATIK: Okay, so this is just a warning, you can safely ignore this one.
[01:25:50] PRATIK: Here, total parameters for BERT is, yeah, approx… This 110 million. Trainable parameters are 3076, right? Fosen parameters out of that are… sorry, sorry, this one is for BERT, right? 307… 3076 is for the linear classifier, and these are the total parameters.
[01:26:09] PRATIK: Okay, and here, only this will be trend, which is corresponding to the linear classifier heading.
[01:26:14] PRATIK: Okay? Now we'll quickly define the evaluate function.
[01:26:20] PRATIK: Right? And, here will be in, like, in this cell.
[01:26:26] PRATIK: we are trying to see… we'll try to evaluate on the accuracy score, precision, recall, F1 score, right? So, we have imported all those, and then classification report also from sklun.metrix, okay? Now, since we'll be evaluating for this
[01:26:41] PRATIK: trained one also, okay? So, for that, what you have done is, instead of writing this again and again, we have created a, evaluate function for this one.
[01:26:50] PRATIK: So, let's see how it is working. First of all.
[01:26:53] PRATIK: like, a model is passed, loader is passed, and a device, like, whichever device you'll be using, that is passed to this function, okay? Then we are using model.eval. Basically, this tells that gradients will not be updated.
[01:27:07] PRATIK: Right? Then, whatever predicted… prediction is happening here will be here, and, true class labels will be automated in this all labels. Okay.
[01:27:18] PRATIK: Then, like, here, we are, disabling the, this.
[01:27:24] PRATIK: gradient computation with torch… with torch.no, no grad, okay? Then, this progress bar is basically… we want to see some progress, right? How fast is there, how slow it is? Okay, let me quickly run this also.
[01:27:39] PRATIK: By that time… oh, sorry.
[01:27:41] PRATIK: It's smart.
[01:27:44] PRATIK: Okay.
[01:27:45] PRATIK: So, like, here we'll try to see the progress bar also, so…
[01:27:49] PRATIK: So, like, everything, we are, like…
[01:27:53] PRATIK: Like, taking it to the device that we want.
[01:27:56] PRATIK: this…
[01:27:57] PRATIK: input IDs, attention mask, token type IDs, label for each batch, okay? Then we do the forward pass, right?
[01:28:05] PRATIK: And then, you know, whatever prediction is happening, we take the last dimension of it, right? And then.
[01:28:13] PRATIK: Like, we extend it, right? Means we are adding it to the,
[01:28:19] PRATIK: We are keeping it in the all parades and all levels, okay?
[01:28:23] PRATIK: Now, for the computation of all the metrics, we have calculated accuracy score based on all levels and all parades. Similarly, precision score, recall score, and F1 score has also been calculated.
[01:28:36] PRATIK: Now, average, we have taken weighted average, okay? And zero division is basically, initially, sometimes what happens is, the model is not able to do some prediction. So, this is, usually absorbed in, like, let's say first typo kernel, so you will see a lot of warning signs. It's basically telling that,
[01:28:54] PRATIK: You know, whenever 0x0 instances are there, so that is basically
[01:28:58] PRATIK: undefined, right? So whenever cases like that happen, simply take that as 0, instead of taking it as 0 by 0, or a NAN value, or something like that.
[01:29:07] PRATIK: So, it's basically telling that.
[01:29:09] PRATIK: Right? And,
[01:29:11] PRATIK: After that, we'll be having classification report based on the all levels, all levels, all prints, right? And,
[01:29:19] PRATIK: These are the targets, which are mapped from 0 to 3.
[01:29:23] PRATIK: Right, and then we'll be returning accuracy, precision, Recall, F1, report, and all labels and all prints.
[01:29:29] PRATIK: Okay, we have already done this. Now, a quick question, okay? Bird has around 110 million parameters in total in this setup. Only classified head is being trained.
[01:29:39] PRATIK: Which is roughly around 3076 parameters, okay?
[01:29:44] PRATIK: So, what do you think?
[01:29:46] PRATIK: Is it enough to produce the meaningful classification result?
[01:29:50] PRATIK: Right, and the next question is, like, CLS token representation produced by Boyd was not exactly trained for news classification, right? It was trained on marked language modeling and next sequence prediction. Do you think this representation will still carry enough information to separate news into four categories?
[01:30:09] PRATIK: Right? So what are your, thoughts on this? Anyone, like, if interested?
[01:30:21] PRATIK: Okay, looks like no one is interested. So, what we'll do is, like, we'll quickly…
[01:30:25] PRATIK: Run this and try to answer these, okay?
[01:30:30] PRATIK: So, first of all, like, we are importing this TQDM from TQDM, okay?
[01:30:36] PRATIK: we are taking 3 epochs. Learning rate we are taking as 2E minus 3, which is, like, a little bit higher. Like, usually we take 2E to the power minus 5, or some… somewhat lower, right? So, this time we are taken, because we are taking as 3 epochs, and, it will not be affecting our training also, right? So…
[01:30:56] PRATIK: training for the BERT, right?
[01:30:59] PRATIK: bolt weights. So, we'll be training the…
[01:31:02] PRATIK: classifier head, so it's okay if we take that. We may take it as small also, doesn't matter.
[01:31:07] PRATIK: In this case.
[01:31:09] PRATIK: Okay? Then for loss, we have taken, loss, we have taken cross entropy loss, and optimizer, we have taken, this item W, where there is some weight decay of 0.01. Learning rate, we have taken from over.
[01:31:22] PRATIK: 2E minus 3.
[01:31:24] PRATIK: Right?
[01:31:25] PRATIK: And here, we are telling the optimizer that only the frozen… this classifier will be trained, only the linear head, right? Not the bite encoder.
[01:31:36] PRATIK: Now, we are quickly defining the… This training book.
[01:31:41] PRATIK: Right? For the training loop. Okay, we are switching it to modal.train. In model.train, what happens is, dropout will be enabled, right? And, batch normalization will be done.
[01:31:54] PRATIK: Okay.
[01:31:55] PRATIK: And, after that, like,
[01:31:58] PRATIK: like, we also have seen that model.evil, right? In that…
[01:32:03] PRATIK: we have to… if we use that, we had to set it explicitly to use dropout, batch form, and all, right? Here, it happens by its own for dropout and all. So, first of all, we have set total loss as 0, and total correct predictions across all batches as 0, okay? Now…
[01:32:21] PRATIK: This is for the progress bar, okay?
[01:32:24] PRATIK: Which, like, which has the input, loader, Right? And,
[01:32:29] PRATIK: leave false is basically… it helps to remove the bar, right, after EPOC is, EPOC completes, so that,
[01:32:37] PRATIK: Unnecessarily does not stack up in the output.
[01:32:41] PRATIK: Okay, now for each batching progress bar, we'll… We'll try to…
[01:32:46] PRATIK: You know, first of all, move…
[01:32:48] PRATIK: all the input IDs, attention mask, token type IDs, and labels to the device, whatever we selected, CPU in our case, right? Then, what we do is we zero out all the gradients.
[01:32:59] PRATIK: Right? Sometimes it is already accumulated by default, right? So we want to zero it out, we clear it, then we do the forward pass, right? We pass input IDs, attention mask, token type IDs to model, and whatever is the output is stored in logits.
[01:33:16] PRATIK: Okay, and then we compute the, loss based on the…
[01:33:20] PRATIK: These logits and the two levels.
[01:33:22] PRATIK: Okay? Then, based on that, we do, backward pass, right? Where, like, gradients are computed.
[01:33:33] PRATIK: Right? With respect to all the parameters. And if some gradients are, like, more than 1.0, right, it will be clipped.
[01:33:41] PRATIK: Right? So, we are doing some gradient clipping here.
[01:33:45] PRATIK: And after that, weights are updated.
[01:33:48] PRATIK: Okay.
[01:33:50] PRATIK: And now, after that, all of this, Like…
[01:33:55] PRATIK: It is this optimizer… sorry. Okay. Now, after that, what happens is we'll be calculating the total loss, right?
[01:34:05] PRATIK: And then, like, we'll update this, progress bar also.
[01:34:10] PRATIK: Okay.
[01:34:12] PRATIK: So, let's try to look at it.
[01:34:16] PRATIK: Are we running it here?
[01:34:19] PRATIK: Okay, it's a training loop, right? Okay, I'll quickly run this also, so that we don't have to wait at the end. So that's how our progress bar comes initially, okay? Just to show it to you.
[01:34:29] PRATIK: Okay, we're here.
[01:34:31] PRATIK: So average loss is calculated based on the total loss and the length of the loader, right? How many instances were there in a batch? And after that, accuracy is calculated.
[01:34:41] PRATIK: And average loss and accuracy is return.
[01:34:43] PRATIK: Okay.
[01:34:44] PRATIK: no…
[01:34:46] PRATIK: year the training loop is run, or epoch, in whatever epoch we have selected. We have selected 3 epochs right now, so…
[01:34:54] PRATIK: We train the epoch. This train epoch function, we just defined it, right? Frozen model is passed, train loader is passed.
[01:35:01] PRATIK: And frozen model is passed, train loader is passed, optimizer, criterion, device, we have all already said this, right? And here.
[01:35:09] PRATIK: we'll be only using the validation accuracy, other metrics we'll not be using, so these are all the placeholders, we'll not be using those, okay? And at the end, we'll be printing it.
[01:35:19] PRATIK: So, let's see…
[01:35:25] PRATIK: So, as we're… this validation accuracy is used, right? So, that's how it came.
[01:35:36] PRATIK: It will take some time, I guess, 2 minutes, approx.
[01:35:47] PRATIK: Meanwhile, if anyone wants to discuss on the questions.
[01:35:50] PRATIK: These questions, if anyone is interested.
[01:35:57] PRATIK: Anyone?
[01:36:07] PRATIK: Okay.
[01:36:14] PRATIK: We are on the last epoch. It'll be completed soon.
[01:36:51] PRATIK: Okay.
[01:36:57] PRATIK: So, the training accuracy and the validation accuracy was noted. Training accuracy was done on the training set, and validation accuracy is on the test set.
[01:37:06] PRATIK: Right?
[01:37:07] PRATIK: Yeah, sorry, validation set, not test set. Test set we'll be doing at the end, right? Now, we'll be quickly… now, if you had taken, you know.
[01:37:16] PRATIK: a large number… more number of data sets, then usually what we observe is, this accuracy increases, right? Which happened in the first step to second step, but since the data was very less, right, initially.
[01:37:28] PRATIK: So, like… Initially, it increased.
[01:37:32] PRATIK: Right? Validation increase, accuracy increased, then again it decreases. Right?
[01:37:37] PRATIK: And meanwhile, if you see the trending accuracy, it is increasing, which is a good sign, okay, for now.
[01:37:47] PRATIK: Now, like, quickly to evaluate on the frozen bird.
[01:37:52] PRATIK: For the frozen bird on the test site, okay?
[01:37:55] PRATIK: So we'll be trying to see the frozen accuracy, frozen precision, recall, F1, and, like, and for that, we'll evaluate, okay? So, for that, what we have done is…
[01:38:05] PRATIK: We have passed, this frozen model test loader device, in the eval.function that we had defined above.
[01:38:12] PRATIK: Okay, then…
[01:38:14] PRATIK: We'll try to print all the scalar metrics, okay, and then we'll also try to see it per class report, like, how it was… how was the performance for each of the classes.
[01:38:23] PRATIK: Let's see Okay.
[01:38:41] PRATIK: seconds. Till then, let's quickly discuss
[01:38:47] PRATIK: Like this, okay? It's happening.
[01:38:51] PRATIK: So, like, BERT has around 110, just a quick question recap, okay? BERT has 110 million parameters in total, and in the setup, only classifier head is being trained, which amounts to roughly around 30… I mean, 6 parameters. Do you think this is enough to produce meaningful classification results?
[01:39:09] PRATIK: So, right now…
[01:39:12] PRATIK: only 3076 parameters are being trained, right? Okay, we have the report, let's quickly… So, accuracy is 0.54, precision is 0.60, recall it's 0.54, and F1 score is 0.47, right? Now.
[01:39:27] PRATIK: the results are, like, not that perfect, right? Or not, like, what we would want to be, given that we have only used 100 test, 100 training samples, right?
[01:39:39] PRATIK: So, even for 3076 parameters, that is very less. The modus was not able to learn properly, okay? Now…
[01:39:47] PRATIK: in the later stage, we'll try to use the whole data and fine-tune upon the data, and then see if it's any better or not. Now, I have also done this experiment for the full dataset and subset of 2,000 datasets, okay? And I'll show you the results at the end.
[01:40:03] PRATIK: Don't worry. So, quickly to answer this, basically, CLS token is carrying the summary of the entire sequence, right? And…
[01:40:15] PRATIK: A simple linear layer is, right now, seems to be sufficient to separate categories. It's… it's not because this classification head is working very nicely, this 3076 parameters is doing wonders. It's not like that. It's because BERT has already been trained on
[01:40:32] PRATIK: like, use data, right? And somehow it has, learned
[01:40:39] PRATIK: from different kinds of data, and based on that, and the linear head, we are able to produce, you know, relatively good results, right? Right now, for this one, it's not that good.
[01:40:54] PRATIK: But if we had run this for the whole dataset, I'll show you the results below, okay? How… how it is.
[01:41:00] PRATIK: Now, the next question was the CLS token representation produced by BERT was not trained for news classification. It was trained on mass language modeling and next sequence prediction, right? Do you think the presentation will carry strong enough information to separate news into categories, right?
[01:41:17] PRATIK: So, the answer to this is…
[01:41:19] PRATIK: You know, pre-training is basically, like.
[01:41:23] PRATIK: We are using a pre-turned bot, right? It has already been…
[01:41:27] PRATIK: trained on various, like, you know, data source. Now.
[01:41:34] PRATIK: Basically, this general… it was on a general language task, right? And it was produced, somewhat useful, this one.
[01:41:45] PRATIK: It was producing somewhat, useful,
[01:41:49] PRATIK: representations, right? And since it has learned on so many, like, you know,
[01:41:56] PRATIK: data. Like, the expectation is that
[01:42:00] PRATIK: CLS, after going through all those data, has some general, you know.
[01:42:08] PRATIK: Has done some generalization and is able to perform well for some of the general things, like classification being one of it.
[01:42:16] PRATIK: Right? So whenever we pass two sentences, or any sentence, it is somewhat able to capture the essence of what the sentence is trying to say. And after that, based on the CLS token, we are able to capture the, you know, meaning.
[01:42:28] PRATIK: Of it. And which later, with the help of linear head and all, we are able to… trying to… we are able to classify it.
[01:42:36] PRATIK: So… Okay.
[01:42:39] PRATIK: So that was a quick answer. Now we'll quickly move on to the fine-tuned board.
[01:42:44] PRATIK: It is very similar to what we had done in frozen bird, right, or pre-trend bird. Here, what we are doing is, in contrast to the previous
[01:42:55] PRATIK: A method, what we are doing this, we are trying to train everything.
[01:42:59] PRATIK: in doing.
[01:43:01] PRATIK: Okay, so it'll be all the parameters, plus the 30… all the parameters of BERT, which is approximately 110 million, and then extra, 3076 parameters, okay?
[01:43:14] PRATIK: So, similarly, what we have done is we have defined the fine-tuned, bird classifier, right, which has inherited all from, inherited from the nn.module also, right? Then, for this init method, it is also…
[01:43:30] PRATIK: Apart from its own init method, it is also…
[01:43:34] PRATIK: initializing all the init methods from the NN module, right? And then, similarly, like we had done for Frozen BERT, like, we have defined this, like, where we are importing, from list. From pre-trend, we are using this BERT-based encased, right?
[01:43:51] PRATIK: This time, the difference is we are not freezing anything. We'll be, updating everything here, so the freezing part is not here, and
[01:44:01] PRATIK: like, the classification head remains as it is, right? The hidden state from the world is passed on to the linear classification layer, right?
[01:44:13] PRATIK: And then… In the forward… now, the forward method is defined, where, like,
[01:44:20] PRATIK: Like, we have passed the input IDs, attention mark… attention mask, and token type IDs, and for the… to get the respective output.
[01:44:28] PRATIK: Okay? Then, with the help of,
[01:44:32] PRATIK: that CLS token. We take that CLS token, and, like, with the help of that, we try to classify the,
[01:44:40] PRATIK: We pass it to the linear layer, right? And then try to classify it.
[01:44:46] PRATIK: See?
[01:44:47] PRATIK: In our case, we are doing, like, it's, like, there are four categories, that's why four.
[01:44:52] PRATIK: Okay, and then we return the logic.
[01:44:55] PRATIK: So this is the class that was defined. Now we have simply created object of it, right? Number of classes we have defined as 4, and then we have transferred this to, like, this has been moved to device. This is CPU in our case. We quickly try to see the parameters that has been used by the total number of parameters, right? Model.
[01:45:15] PRATIK: And then the total trainable parameters.
[01:45:18] PRATIK: complete with our Nick.
[01:45:22] PRATIK: Okay, so if you see… Total parameters is this one.
[01:45:28] PRATIK: Right? And total trainable parameters is also this one, because we are training end-to-end everything.
[01:45:34] PRATIK: Okay? Now, a quick question. Like, before, we were using
[01:45:39] PRATIK: The learning rate of,
[01:45:43] PRATIK: 2 to the power 10 minus 3, right? And for fine-tuning, we are using, in below, if you see, 2 to the power 10 to the power minus 5, right? Why do you think, like, this is there? Like, why do you think, specifically, we have reduced it?
[01:45:57] PRATIK: So this is one question, and another one is, like.
[01:46:00] PRATIK: We are using these bolt weights that are meaningful because of pre-training, but the classifier head is randomly initialized, right, in the few training steps.
[01:46:10] PRATIK: gradient flow from this random classifier all the way back through BERT, what problem could this cause, and how does the choice of a very small learning rate help avoid it?
[01:46:20] PRATIK: Okay, so we'll try to, after we run this, we'll try to answer this.
[01:46:25] PRATIK: Okay.
[01:46:27] PRATIK: So… Now, for the training loop, we have done something very similar.
[01:46:33] PRATIK: Epochs, we have taken as 3. Learning rate we have taken as 2 to the power… 2… 2 into 10 to the power minus 5.
[01:46:39] PRATIK: Right? We have reduced the learning rate, because we don't want it to, you know, because since the bird has already been trained on large number of data, right? It has already learned some weights, and we don't want, like.
[01:46:55] PRATIK: our initial data to, you know, disturb all those weights, right? Because this time it is training end-to-end, and the linear classifier head at the top is randomly initialized, right? So initially, it will change, if we keep it very high, the weights will be changed.
[01:47:15] PRATIK: Right? Very drastically, right? So we want a lower learning rate, so that the weights do not, you know, cause use disturbance. It is, updating the weights smoothly. That's why we do that.
[01:47:29] PRATIK: Okay? Then loss criterion we have taken as,
[01:47:34] PRATIK: cross entropy loss, and optimizer, we have taken at MW, just like before, right?
[01:47:41] PRATIK: And…
[01:47:44] PRATIK: We used a train epoch from the previous cell, it was already defined in the function, we are using that, right? That function was already there. Okay, and now we'll try to run the full training loop.
[01:47:55] PRATIK: Okay, I'll quickly run this also.
[01:47:57] PRATIK: We'll take something.
[01:47:59] PRATIK: Now, for each, we are trying to see the training loss and training accuracy.
[01:48:05] PRATIK: Right? And, this one, validation accuracy also.
[01:48:09] PRATIK: So, these are all the placeholders. We are just curious for the validation accuracy. That's why we are evaluating.
[01:48:15] PRATIK: I'm just trying to see it here.
[01:48:17] PRATIK: Oh, it is performing the one.
[01:48:23] PRATIK: Let's wait a couple of minutes.
[01:48:42] PRATIK: Meanwhile, if anyone has any question.
[01:49:19] Sacheen Adavinavar: Pratik, I have one question.
[01:49:21] PRATIK: Yeah, please.
[01:49:22] Sacheen Adavinavar: So here, the train size is, I mean, compared to the train size, right, we have a test size which is, less than 10%. Why, I mean, is there any reason? Because earlier, whatever, till now,
[01:49:36] Sacheen Adavinavar: I mean, the model which we have tried out.
[01:49:39] Sacheen Adavinavar: So, we had at least the test set as, I mean, nearly 20% or 15%, like that.
[01:49:46] PRATIK: No, no, this is not a test size, this is just a validation.
[01:49:50] PRATIK: Test size, so, while doing training, we are also checking out validation.
[01:49:56] Sacheen Adavinavar: Oh, okay.
[01:49:57] PRATIK: Training set also, we have… what we have done is, we have splitted it into training set and the validation set.
[01:50:03] Sacheen Adavinavar: Okay.
[01:50:03] PRATIK: Just so that… so this is done, basically, to check if we are… just in case, as the… as we are going through epochs.
[01:50:11] PRATIK: Is the, means model over-learning, over-memorizing?
[01:50:16] PRATIK: Right? Because it's going through the same training samples, right?
[01:50:21] PRATIK: To judge that, we are using this validation set. So after that, after the weights… after it has learned the weights from the first epoch, right?
[01:50:30] PRATIK: It takes… it takes those weights, right, and then goes through the validation set.
[01:50:35] PRATIK: And then comes up with the validation accuracy. So from there, we get some idea that, you know, after training, it is improving, the training is actually helping or not, right? Whether it is overfitting, whether it is underfitting.
[01:50:49] PRATIK: So we get all those kinds of ideas. And test set, we are touching at the end only, during the final evaluation.
[01:50:55] Sacheen Adavinavar: Okay.
[01:50:56] Sacheen Adavinavar: It is like a cross-validation.
[01:51:00] PRATIK: You can say it, but not exactly that. For cross-validation, there are different techniques also, so you can take this as that. But there are various other techniques as well.
[01:51:11] Sacheen Adavinavar: Yeah.
[01:51:24] PRATIK: This is because this time it is taking more time, because we are going through all the training from end to end, right?
[01:51:32] Sacheen Adavinavar: E.
[01:51:32] PRATIK: from one end to another end. That's why it is taking a little bit more time.
[01:51:37] PRATIK: Do you want me to continue to save time, or should I wait for this to complete?
[01:51:46] PRATIK: Anyone?
[01:51:48] Pallavi Chakravarty: You can continue.
[01:51:51] PRATIK: Okay.
[01:51:52] PRATIK: So, like, after this, we'll, like, once this, we have the training done, we can do the evaluation, right?
[01:52:00] PRATIK: So for that, we have… we are doing pretty much the same what we had done for the frozen part, right? Frozen part meaning where we are just training the, this linear classifier, and rest of the weights are frozen.
[01:52:12] PRATIK: Right? So the evaluated function from before, is being used, right? And, this time it is using fine-tune model, test loader, right?
[01:52:23] PRATIK: And device. So, as you can see, Sachin, like, we are using test loader here, right? Which was based on the test set. Okay. So, yeah. So, earlier part, we were using on the validation set.
[01:52:36] PRATIK: It's well loaded, see.
[01:52:38] PRATIK: Okay.
[01:52:40] PRATIK: Okay. Now, after that is done, we'll quickly have a look at accuracy, precision, recall, and F1 score, and also the per-class report, meaning, like, for each… how… for each of the class, how it is performing.
[01:52:54] PRATIK: Let's see, almost done.
[01:54:07] PRATIK: So, as you can see, Sachin, like, here we get some idea, right, from the train accuracy, if the train accuracy is increasing or not, right? And with the validation accuracy, we have some idea how it is performing on the validation… the unseen data, right?
[01:54:23] PRATIK: So, you see that in this case, it has stayed constant, but in… if you do it for a large number of epochs, you will see that validation accuracy, let's say, in the initial stages are less, slowly it is increasing, and as you increase the number of epochs, you can see it again decreasing also.
[01:54:41] PRATIK: So that's how you know that after this number of epochs, it is underfitting, right? Sorry, overfitting. So then you can…
[01:54:49] PRATIK: You know, work. And if it is not decreasing at all, if it is increasing again and again, that means it is underfitted, kind of, right? So, for this, we are doing more.
[01:55:04] Sacheen Adavinavar: Thanks.
[01:55:13] PRATIK: just to reiterate, okay? If this increases, right, we know, and if this is also increasing, this is also increasing, then that means the model is learning, right? If this… the difference between them is more, right?
[01:55:30] PRATIK: then we can say some overfitting or underfitting is happening. I thought I said opposite, that's why.
[01:55:37] PRATIK: Nope.
[01:55:39] PRATIK: So, now…
[01:55:44] PRATIK: To do the evaluation. Let's see, again, to take time.
[01:55:48] PRATIK: Okay. Till then, let's quickly discuss on the answers for this one. So, like, as we discussed, like, I had complete, given some hint already. So, what happens is, since we are training for both of the questions, answer is very similar, right? So, basically what happens is, since we are doing end-to-end training, and it already has some weights.
[01:56:08] PRATIK: Right? That it has learned on the large corpus.
[01:56:11] PRATIK: We don't want… and we have a linear classifier head on the top, right? So it will be initialized with some random widths initially. So…
[01:56:19] PRATIK: There'll be, some huge…
[01:56:22] PRATIK: of, since there'll be noisy gradients initially, right? It will try to up those weights, and if we keep learning rate high, then what'll happen is, like, it'll change the weights drastically, and whatever it,
[01:56:37] PRATIK: the pre-trained model had learned those important myths. Weights also could have been, you know, disturbed.
[01:56:43] PRATIK: So, what we do is, we keep the learning rate less initially, right, for this one, because we don't want those, untrained weights in the classifier head, because of that, learning, the…
[01:56:56] PRATIK: the, you know, the good weights that the pre-trained classifier already has to disturb that very largely, right? We want small changes to it, not huge changes to it.
[01:57:08] PRATIK: Great.
[01:57:09] PRATIK: So that's why we do this.
[01:57:12] PRATIK: Okay, so let's quickly have a… You know?
[01:57:17] PRATIK: look at it. So, F1 score, this time is, 0.75 for void. Sports, it is 0.63. Business, 0.61, site tech, it is 0.51, right? And accuracy, efficiency.
[01:57:30] PRATIK: Overall accuracy is 0.63, precision is 0.75, recall is 0.63, and iPhone score is 0.62, right?
[01:57:39] PRATIK: We have also resulted comparison between, you know, the pre-trained and the frozen bird. Let's quickly have a look at it.
[01:57:46] PRATIK: So for that, what we have done is whatever, you know, print, whatever results were there before, out of that, we have created a data frame, and based on that.
[01:57:57] PRATIK: We have created one graph also. Let me move quickly.
[01:58:01] PRATIK: Renick?
[01:58:03] PRATIK: graph is below. So first, let's have a look at this.
[01:58:07] PRATIK: So, frozen bird accuracy was 0.54, right? Fine-tuned bird is 0.63, like, improvement, right? So if you multiply this with 100, it's around 54% and 63%, like, a huge improvement.
[01:58:21] PRATIK: Preseason was?
[01:58:23] PRATIK: For frozen bird, 60% approx, right? And then for fine-tuned bird, it is 75% approx. Recall, 0.554 is 54%, and 63%. F1 score increased from 47%, approximately 48% to 62%, right? That's a huge jump.
[01:58:42] PRATIK: Vague.
[01:58:43] PRATIK: and if we had used more data, the result would be even better.
[01:58:50] PRATIK: Right, and now let's have a quick look at, this per-class F1 comparison.
[01:58:56] PRATIK: So, this 4-wall category, it increased by 8% when we used fine-tuned, right? For sports,
[01:59:06] PRATIK: it increased by 3%. For business, it increased by 5%, right? And SciTech, it was almost, like, 6%. It did not perform so well, right? For a fine-tuned one.
[01:59:17] PRATIK: on-site.
[01:59:18] PRATIK: But, like, if you see… if we had used the whole data, I'll show you the results, it has improved for each of the class, right? Performance is very good. Now, we'll quickly…
[01:59:31] PRATIK: visualize whatever is here with the help of matplotlib.
[01:59:36] PRATIK: Great.
[01:59:40] PRATIK: So, like, the same results, this time it is in, this.
[01:59:48] PRATIK: Rough form, right? So, here, if you see.
[01:59:51] PRATIK: everything is improving in general, right? This is the overall metrics for accuracy, precision, recall, and F1 score.
[01:59:58] PRATIK: Blue one is the frozen bird, and
[02:00:00] PRATIK: orange one is the fine-tuned part, right? And in per class, it… in general, it is increasing for all these three, except the SciTech. SciTech, it was not so good. It did not perform so well.
[02:00:12] PRATIK: Okay. Now… To summarize, like, how it was, everything was done.
[02:00:19] PRATIK: Okay. Till now, if anybody has any questions, then I'll quickly move on to Samurai and the results that I had for the full data set and a subset of 2000 datasets.
[02:00:28] PRATIK: Anyone, any question?
[02:00:36] PRATIK: Okay? So, what we have done till now, okay? Basically, we created a…
[02:00:42] PRATIK: Jupyter Notebook, right? We compared two BERT-based models, right? One which was frozen, BERT-based was frozen, and then only the linear classifier I had on top of was trained. And on the other hand, we used a BERT model.
[02:00:57] PRATIK: Linear classifier on top, and this time, everything was trained.
[02:01:02] PRATIK: From end to end, right? Both of them were used for a classification task.
[02:01:06] PRATIK: classification task involved, classifying each of the rows into four classes, one of the four classes, okay? So first, for that, we did some data exploration and preparation, right? And then we, then we used the
[02:01:23] PRATIK: pre-trend bird.
[02:01:24] PRATIK: Right?
[02:01:25] PRATIK: with learning rate 2 into 10 to the power minus 3, right? And next, we use the fine template with learning rate 2 into 10 to the power minus 5.
[02:01:34] PRATIK: Okay. Now, result comparison, basically, both the models were evaluated and compared using accuracy, precision recall, and F1 score. Also, we did, per class, F1 score comparison also, and we saw the graphs also, side by side.
[02:01:51] PRATIK: Now, like, as I said, I had done some experiments on the 2,000 sample subset, right? And the full dataset, comprising of 1,20,000 sample, right? Both evaluated on 7,600 row test set.
[02:02:06] PRATIK: Okay.
[02:02:08] PRATIK: Now, if you see.
[02:02:10] PRATIK: And in this notebook, I have used only 100 sample of training subset and 100 sample of certified test set.
[02:02:18] PRATIK: Right?
[02:02:19] PRATIK: Now, let's see the results. Over 2,000 training samples.
[02:02:23] PRATIK: Right? Now, just have a look, like…
[02:02:26] PRATIK: 2,000 training samples on which the frozen, like, frozen and pre-trained, frozen and the fine-tuned boat was trained, and it was tested on 7,600, test samples, which is very large than… more than triple, right? But still, see, the results were, much, much better.
[02:02:44] PRATIK: Like, than what we have. So, if we see 4 frozen birds, it's around 85% accuracy.
[02:02:50] PRATIK: With having around 86% precision.
[02:02:53] PRATIK: 86, around 86% or 85% recall, and, having iPhone score of around 85%.
[02:03:01] PRATIK: F1 score, right? Whereas, in fine-tuned board.
[02:03:05] PRATIK: like, I, like, I was able to get around 89% of accuracy.
[02:03:11] PRATIK: 90% of precision, 89% of recall, and 89% of F1 score. Now, this model, being trained on 2,000 training sample, this is, like, giving very good results, right?
[02:03:24] PRATIK: It means that it is able to do tasks very nicely.
[02:03:30] PRATIK: now, like, if you see the per-class F1, comparison, right?
[02:03:35] PRATIK: Pull void?
[02:03:36] PRATIK: the difference, for a boiled category, the difference between the frozen boiled F1 and the fine-tuned boiled F1, the difference is around 3%, right?
[02:03:47] PRATIK: For sports, it's around 3.4%. For business, it's around 4.4%. For scitech, it improved around, 4%, 4.3%, right? So, there's notable…
[02:04:00] PRATIK: like, improve… improvement in performance when we used fine-tuned BERT, right? Now, let's see the, like, how the performance was for the full data set, of… comprising of around 1,20,000 sample, right? And, test sample, again, of 7,600. Like, this is the maximum we have.
[02:04:18] PRATIK: Okay? So, this time, the accuracy for the frozen bird was…
[02:04:23] PRATIK: 90.21, precision was 90.22, recall was 90.21, and F1 score was 90.16. All touching 90, right? Very good performance.
[02:04:34] PRATIK: And, like…
[02:04:36] PRATIK: For fine-tuned Bird, like, the accuracy was 94.30, precision was 94.31, recall was 94.30, and F1 score was around 94.30 again. So, almost there is an increase of 4%, right, when we use Fine Tuned Bird.
[02:04:54] PRATIK: Right? Now, even if we had not fine-tuned it, 90%
[02:04:59] PRATIK: Having 99%, like, performance across all these metrics is a very good like, performance, right?
[02:05:07] PRATIK: Like, given by our pre-trained bot also. So we must highlight that as well, right? Now, quickly, to see about the per-class F1 comparison, right?
[02:05:19] PRATIK: The frozen bolt was around, 90.74, right?
[02:05:25] PRATIK: This is the iPhone score, okay? And for… This one.
[02:05:30] PRATIK: fine-tuned, it was 95.74. Difference was around 4.99%.
[02:05:35] PRATIK: Right? Touching almost 5%. Sports, it increased by around 1%, 1.7%. Business, it increased by around 5.4%. SciTech, it increased around, like, 4.38%, right? Very good performance.
[02:05:50] PRATIK: Then… Like, these are, like, these were some of the key findings, okay?
[02:05:56] PRATIK: So, no, like, the largest gains, like, that we got was around 4.46% for business, right, on the 2,000 samples.
[02:06:04] PRATIK: Right? And the SciTech, 4.37%, right?
[02:06:08] PRATIK: Business was the weakest class, right? 81.20 and 85.66, okay? And on the full data set, if you see, in contrast to, like.
[02:06:20] PRATIK: frozen… Dataset, like, everything was…
[02:06:26] PRATIK: well above 90%. There was almost a 4% difference here also. And even in the subset of 2,000 samples, the gains were around 4% for, like, for, like, overall metrics, right?
[02:06:41] PRATIK: in F1 score. So, we can see that pre-trained, bird is also, like.
[02:06:50] PRATIK: Working very nice if we have Enough samples.
[02:06:54] PRATIK: like, in the initial case, in our case, we may say that, like, the performance was not so good, because maybe the 3076 parameters were also not, well trained, because the model was able to see very few examples. 100 samples is very less.
[02:07:11] PRATIK: Right? We can confidently say that. And as you increase the number of samples, we see… we have seen better performance in our case, right?
[02:07:21] PRATIK: Like, we… we can say that with confidence with this example. Like, we had just in with 2,000 training samples, right? Which was tested on 7,600 test samples.
[02:07:32] PRATIK: And the performance was pretty good, right? 85% F1 score on Frozen Bird, and 89% of this on fine-tuned Bird is a very good score.
[02:07:43] PRATIK: So, and also we saw that, like, as we increase the number of
[02:07:47] PRATIK: this data, right? During training, like, the performance is also increasing.
[02:07:55] PRATIK: Okay.
[02:07:56] PRATIK: So, these are some of the key takeaways. I hope you like
[02:08:00] PRATIK: And enjoyed today's lecture. You had… I hope, like, you had some great learning today.
[02:08:06] PRATIK: If there is any question, we can address it.
[02:08:10] PRATIK: Anyone?
[02:08:23] PRATIK: Any question, anyone?
[02:08:28] PRATIK: Okay. I think, then we can close the class.
[02:08:32] PRATIK: Thank you so much, guys.
[02:08:35] PRATIK: I hope I was able to answer your questions properly, and I hope you had a great learning session today. Thank you so much. We can close for today. Thank you.
[02:08:44] PRATIK: Good night.