# 11 2026-06-20 Gen AI LLMs Session

course: Module 4 — Generative AI & LLMs
module: Module-4-Generative-AI-LLMs
date: 2026-06-20
type: transcript
video_url: https://personal-learn.armco.dev/files/_Recordings/Module-4-Generative-AI-LLMs/11_2026-06-20_Gen_AI_LLMs_Session.mp4

---
[00:16:06] Anirban Paul: Good evening, guys.
[00:16:27] Anirban Paul: Am I audible?
[00:16:32] MuniPrakash Ganji: Yes.
[00:16:33] Nirav Mehta: It's…
[00:16:35] Anirban Paul: Okay.
[00:16:41] Anirban Paul: Good evening, yes.
[00:16:43] Anirban Paul: Okay, just a moment…
[00:16:48] Anirban Paul: Still not a lot of people are yet to join.
[00:16:53] Anirban Paul: Okay, anyways, we will continue.
[00:16:57] Anirban Paul: On top of what we were discussing yesterday.
[00:17:07] Anirban Paul: about…
[00:17:32] Anirban Paul: Okay, what we were discussing yesterday about diffusion models. That was the last topic that we had left, and we saw a very, high-level understanding about diffusion model, and just the forward diffusion process that we have seen.
[00:17:45] Anirban Paul: Okay, so just a quick recap will do, and we will head on to other parts of diffusion models as well. Okay, so…
[00:17:54] Anirban Paul: And then we'll also use, one of the diffusion models from Hugging Face and
[00:18:01] Anirban Paul: We will try with some experiments as well. Okay. So, first of all is, in forward diffusion, what did I… I told you that there was a paper known as Deep Unsupervised learning using, non-equilibrium thermodynamics, right?
[00:18:22] Anirban Paul: That is the paper from where this concept of division models got inspired in the year 2025, but the popularity came later on as the compute power increased.
[00:18:33] Anirban Paul: Okay, and that is how diffusion model, you know, gained popularity. Before that, the popular models were GANs, and before that, it was VAE, and nowadays, it is all diffusion models only. So the concept of
[00:18:47] Anirban Paul: forward diffusion was? That, if you have an image.
[00:18:57] Anirban Paul: Like this, you keep on… Adding noise at every step.
[00:19:04] Anirban Paul: Pair the noise added at
[00:19:07] Anirban Paul: at every step is just a little bit more compared to the previous one, so it is… it follows a Markov chain
[00:19:12] Anirban Paul: rule. Basically, the state of the current, the current state of the image is basically closest to the most recent stage, and it is just a modification of the most recent stage. Okay, so you update, you start from this position.
[00:19:28] Anirban Paul: And then, like that, you keep on updating.
[00:19:35] Anirban Paul: Keep on upgrading, and let's take T is equals to 100,000.
[00:19:44] Anirban Paul: By here, your image will be filled with TV static.
[00:19:48] Anirban Paul: That it will be unrecognizable that even a human existed over there.
[00:19:52] Anirban Paul: Okay, it'll be so unrecognizable that
[00:19:55] Anirban Paul: you will not be able to say that whether a human existed or not. So, it's like this, okay?
[00:20:00] Anirban Paul: So, we add this Gaussian noise at every, image.
[00:20:05] Anirban Paul: And we slowly, slowly, slowly, slowly, slowly, slowly make it very unrecognizable, and by the level we reach X1T1000, it is completely unrecognizable. This process was known as forward diffusion, and there are a few things that you… we understood from this process was adding noise.
[00:20:22] Anirban Paul: Add noise. Then we also had time steps concept.
[00:20:28] Anirban Paul: Then we also had Markov chain.
[00:20:35] Anirban Paul: Okay.
[00:20:36] Anirban Paul: And then, I told you that from here.
[00:20:41] Anirban Paul: This is the initial process. This is a very, very, you know, process for you to actually create the data, okay? Create the data, because this reverse diffusion is the actual learning process. So, let's understand if you're reverse diffusion. This is the actual
[00:20:57] Anirban Paul: Process of diffusion is the reverse diffusion process.
[00:21:00] Anirban Paul: So…
[00:21:02] Anirban Paul: In reverse diffusion process, you actually go the opposite way. You go from here to here, here to here, here to here, like that, you go and reach here.
[00:21:10] Anirban Paul: Okay, so in reverse diffusion process, the concept is that you are starting from a position where you have complete noise.
[00:21:22] Anirban Paul: Okay.
[00:21:24] Anirban Paul: You have complete noise.
[00:21:26] Anirban Paul: And then, from there, you are basically going to predict
[00:21:31] Anirban Paul: From here, you are actually going to predict That what was the noise
[00:21:38] Anirban Paul: One step before this, so this is T1000. Before, one step before this, let's say T… at T1… T999.
[00:21:47] Anirban Paul: how did the image look like? What was the noise that was there? So, let's say you predict. This is your predicted…
[00:21:53] Anirban Paul: Value?
[00:21:55] Anirban Paul: at T999, these were the NAS that got added, let's say.
[00:22:00] Anirban Paul: These were the noise that got added.
[00:22:03] Anirban Paul: Okay, so basically you do a minus of this.
[00:22:06] Anirban Paul: If you do a minus of full noise from this noise that got predicted in the T999 step, then you will get X
[00:22:18] Anirban Paul: T999, or whatever was your time step minus 1, so XT minus 1, you will get as an output which looks like this.
[00:22:32] Anirban Paul: An output which is devoid of this noise that got predicted in the 9999 step. So this…
[00:22:37] Anirban Paul: That T999 step that you have, that image that you have, that serves as a data.
[00:22:49] Anirban Paul: Got it?
[00:22:51] Anirban Paul: So this is noise predicted.
[00:22:53] Anirban Paul: This is random noise.
[00:22:58] Anirban Paul: or XT1000, you can call it as, or you can also refer… for generic purpose, you can refer it as XT only, and this make it like XT-1.
[00:23:08] Anirban Paul: Okay, and like that, you will keep on predicting, predicting, and then from this step, you will come to the next step, where you will predict which noise was there. So again, you will predict a noise. Let's say this noise got added at this step, so you will minus that from here, and you will get an image which is free of that noise.
[00:23:25] Anirban Paul: So you will get an image which is free of that noise.
[00:23:28] Anirban Paul: In between.
[00:23:30] Anirban Paul: Okay, so like that, you will keep on repeating these steps until a point comes. It is not you, as in, like, the stable diffusion, the diffusion models. A point will come where you will have images where, let's say, slowly, slowly, a human will start appearing.
[00:23:48] Anirban Paul: A human will start appearing, and…
[00:23:51] Anirban Paul: You will last probably predict the last few prediction, that these are the things that you will have to remove, and then you will have a very clear image.
[00:23:59] Anirban Paul: This is the concept of reverse diffusion.
[00:24:02] Anirban Paul: Okay, now, this purely obeys the concept of equilibrium and reverse equilibrium, okay? So, like I told you yesterday, you have a beaker on which a red dye was added, or a blue dye was added. The dye got spreaded into the old beaker.
[00:24:18] Anirban Paul: Taylor states are equilibrium reached, where every particle has spread across.
[00:24:24] Anirban Paul: Okay, now you have to reverse diffusion that person, and we are doing that alone.
[00:24:27] Anirban Paul: Instead of red dye, you have noise over here.
[00:24:32] Anirban Paul: Got it? So, it also use… it is also using a Markov chain formula. So, it is mixing that Markov chain capabilities with
[00:24:40] Anirban Paul: With your image generation capabilities. Now.
[00:24:45] Anirban Paul: This process of using reverse convolution, you can say this was a convolution thing, this was process… this process is called as a unit architecture.
[00:24:55] Anirban Paul: a unit CNN architecture. So, as you are going a U-turn, U-turn, you are using a U-turn of neural networks, okay, U-neural networks. Okay, that's why that architecture which got popular in diffusion model was UNET CNN architecture. So, using UNET, you do
[00:25:11] Anirban Paul: Basically,
[00:25:15] Anirban Paul: a stable diffusion model, or, sorry, a reverse diffusion process. Okay. Now, what is the loss function? In this case, the MSC is a loss function. At every step, you are basically
[00:25:26] Anirban Paul: doing this minus the actual one that you had. Okay, sorry, this minus the actual one you had, and like that, you find out, like, how much you are able to predict the actual image as well, at every step.
[00:25:41] Anirban Paul: So, this is the most plain version of
[00:25:45] Anirban Paul: reverse diffusion process. This is known as unconditional reverse diffusion process.
[00:25:50] Anirban Paul: Any question, guys, in this topic?
[00:25:53] Anirban Paul: As of now?
[00:26:12] Anirban Paul: Okay, so are you all understanding, guys?
[00:26:15] Anirban Paul: Can you give a yes in the chat, if you're understanding?
[00:26:19] Deepan Kanagaraj: Sir, one question. Let's say in the DALI model, I'm working on a project where the farmer has to sow the seeds, and then he needs to… that's activity, let's say, right? So we are actually giving that as a context to the DALI model, and then DALI actually generates the farmer, with the background…
[00:26:37] Anirban Paul: you're talking about, not Dali 3.
[00:26:38] Deepan Kanagaraj: DALI3, yeah, DALI 3. And then it is actually generating that image. So now you also told DALI3 actually using the diffusion model to do that, right? So now, this is my prompt, and then I'm getting the farmer image. From where, actually, this is coming to play?
[00:26:52] Anirban Paul: I'll come, I'll come. That is known as conditional diffusion process, conditional reverse diffusion process, CRD. Now, we are doing unconditional diffusion process. There is no condition. We are just adding noise, and from that noise, again, we are getting back that image. Now, I have chain by chain,
[00:27:09] Anirban Paul: image that got created while doing the forward diffusion, so that acts as a helper for me, that this is how the initial image will look. That's how MSE is happening at every step. Got it? Now.
[00:27:24] Anirban Paul: Yeah.
[00:27:26] Anirban Paul: I will explain conditional as well, but now, first, also understand, at a very pixel level, this you are understanding at an image level. So, actually, what is happening at a pixel level? So, what are images? Images are pixels of colors, right? Colors of RGB.
[00:27:41] Anirban Paul: Okay. So, let's say if I put them inside a coordinate system, then if This is…
[00:27:50] Anirban Paul: If this is R, this is G, this is B.
[00:27:54] Anirban Paul: Okay. And, let's say I have the color pixels plotted it out, so I have
[00:28:00] Anirban Paul: Let's say one here.
[00:28:02] Anirban Paul: One here, one pixel. Various pixels I have plotted it out. Okay. One not here,
[00:28:10] Anirban Paul: I have one here, one here.
[00:28:13] Anirban Paul: Okay, so what is happening at a pixel level is… so, let's say this pixel.
[00:28:18] Anirban Paul: So, at this pixel, you can call it as 255, because it's on the R side. So, 25500, because it's on the pixel. So, let's say this is, like, 255.00.
[00:28:29] Anirban Paul: Okay, that I have placed, okay? For example, Now…
[00:28:33] Anirban Paul: I start adding a noise onto this. Let's say I add a noise of… Minus 2.
[00:28:40] Anirban Paul: Zero. Zero.
[00:28:42] Anirban Paul: If I do this, then the value will become 253.
[00:28:47] Anirban Paul: 2… Zero, zero.
[00:28:50] Anirban Paul: So, can I say that this has moved away from here, like this?
[00:28:56] Anirban Paul: This point has moved away from here like this. Can I see it like that?
[00:29:00] Anirban Paul: Everybody agrees?
[00:29:03] gunjan bhaiya: Yeah, because the red values… Decreased.
[00:29:07] Anirban Paul: Yeah, so it has kind of moved…
[00:29:08] Muni Prakash Ganji: 3 minus, correct?
[00:29:10] Anirban Paul: Yeah, it is, it is, it is minus.
[00:29:12] Anirban Paul: Okay.
[00:29:13] Anirban Paul: So, it has moved away, let's say. Similarly, at every pixels, move away from here. So, it starts from an initial point, it moves away, it moves away, it moves away, it moves away, and until a point will come that it will be completely unrecognizable.
[00:29:28] Anirban Paul: So, these pixels were the initial portions of the pixels, and they got just moved away. This is the concept of reverse diffusion process… forward diffusion process. Now, we are adding a reverse diffusion process to bring it back to the original thing. This is only… at a pixel level, this is what is happening.
[00:29:43] Anirban Paul: First, I showed you at image level, now I'm telling you at a pixel level, that is what is happening.
[00:29:48] Anirban Paul: Okay.
[00:29:49] Anirban Paul: So, now, coming to the next part.
[00:29:52] Anirban Paul: Peace?
[00:29:54] Anirban Paul: This was unconditional diffusion
[00:29:58] Anirban Paul: process. This was initially how the division model came into existence. Now, the question is, how does prompt to image happen? So this is, like, image-to-image use of, okay? Image to how new images are getting created, and now the point is, like, how…
[00:30:20] Anirban Paul: some AI tools. Deep on that… I'll talk about those, you know, just… let me just cover this, okay? So…
[00:30:28] Anirban Paul: Now the point is about unconditional
[00:30:31] Anirban Paul: Unconditional, the point is about conditional diffusion process.
[00:30:39] Anirban Paul: diffusion process. Reverse diffusion. Conditional reverse diffusion. So, in conditional reverse diffusion, what happens is.
[00:30:47] Anirban Paul: along with the noise. So, what was happening? You are having an image of noise, completely a data of noise. From there, you are coming to a clear image.
[00:31:00] Anirban Paul: clear image.
[00:31:02] Anirban Paul: Now… Now what happens is… In conditional reverse diffusion process.
[00:31:10] Anirban Paul: You are conditioning this process with the help of a text.
[00:31:15] Anirban Paul: So first, you have a text, which is your prompt.
[00:31:19] Anirban Paul: You have a text, which is basically prompt.
[00:31:22] Anirban Paul: That is first converted to… Huh?
[00:31:27] Anirban Paul: What? What? Take a guess.
[00:31:31] Anirban Paul: What, guys?
[00:31:35] gunjan bhaiya: The instructions?
[00:31:37] gunjan bhaiya: collector.
[00:31:40] Anirban Paul: Anything else?
[00:31:42] ranjani siva: Random?
[00:31:43] ranjani siva: Random dots, or, making a random structure.
[00:31:48] Sacheen Adavinavar: encoder.
[00:31:48] Anirban Paul: Vector, vector, X vector.
[00:31:51] Anirban Paul: a text vector out of it. So, ultimately.
[00:31:55] Anirban Paul: Your noise, which is also a matrix.
[00:31:58] Anirban Paul: Okay.
[00:32:00] Anirban Paul: And this text vector is also a matrix. So both have bought… have been bought to the same level. Then only things can start, okay? So this is automatically done, like, you don't have to do it. But, you know, in the training process, like, when the division creators, they first take the text, convert it into a text vector. So now, it is not only noise, it is noise
[00:32:18] Anirban Paul: Which is also a matrix, noise matrix. Noise matrix.
[00:32:23] Anirban Paul: Plus, text vector, which is also a matrix.
[00:32:26] Anirban Paul: This is given as an input now.
[00:32:28] Anirban Paul: Okay, and from here, you come to a clear image. Now, How do you…
[00:32:36] Anirban Paul: this… this is the reverse diffusion process, but how did the forward diffusion process happen in this case? In this case, you actually… you had the… in the training purpose, during the training, you had images.
[00:32:52] Anirban Paul: you had… I image like this. Okay.
[00:32:56] Anirban Paul: Plus, you had…
[00:32:58] Anirban Paul: a text also like this. The text representing it. So you have the pairs. You have the pairs of images and their text. Text is this text.
[00:33:07] Anirban Paul: Okay, you added noise onto this, and that is your final noise.
[00:33:11] Anirban Paul: Okay.
[00:33:12] Anirban Paul: Then, from that noise, you are going one step.
[00:33:15] Anirban Paul: by, one step, like, that is how you're doing conditional reverse diffusion process. So, previously, you just had image.
[00:33:22] Anirban Paul: And your image got converted to numbers, and those numbers, like, in that numbers, new noise numbers was added.
[00:33:31] Anirban Paul: Now, you have image, those will get converted numbers. Along with that, you also have a pair of text vector as well. That is also number. So first, you have this concept.
[00:33:41] Anirban Paul: of bringing both of them to a same embedding space is known as shared embedding space. This concept is popular in diffusion models. This concept got popular, this
[00:33:51] Anirban Paul: this space, once you convert them both into same numbers, okay, both into a similar number format, then this thing becomes known as shared embedding space. So that is the point from where the process begins. Now, did you get it, Deepan?
[00:34:09] Deepan Kanagaraj: Yes, yes sir, I can understand now.
[00:34:11] Anirban Paul: Yeah, so first, you just had image, you didn't have any condition. Now you have image, which are converted to numbers. Along with that, you have numbers of the text vectors as well. So, that is your input now, while training the process. And you add noise, noise, noise, noise, noise, it becomes unrecognizable. Then, from that unrecognizable noise, again, you come back and create back that image.
[00:34:30] Anirban Paul: So, that reverse process, reverse diffusion process is known as conditional reverse diffusion process, and your training images… training pairs are now images, and they're text pairs.
[00:34:40] Anirban Paul: That is how it is done. So, CRD is what is our, basically, all these stable diffusion models, if you have heard of. Like, previously it was known as diffusion models, then, stable diffusion is a class of model that this company known as RunwayML came up with. So…
[00:34:58] Anirban Paul: So yeah, if I go back to the PPT.
[00:35:01] vinit shah: Sir, I have one doubt, with regards to this, Anirban, so… What you're telling is.
[00:35:07] vinit shah: I'm going to train the model by saying that, say, for example, I take hand as an example, okay? So I'm going to train the model that this is hand with 5 fingers, and I'll.
[00:35:16] Anirban Paul: Yeah, not you, not you, just to take that thing up. Yeah, yeah. The researcher of, say.
[00:35:21] vinit shah: Whoever has done it, or wherever that stock has come, right? Now, what is happening is…
[00:35:29] vinit shah: if I'm putting across 20, 30 different… the researcher is putting across hundreds of different hands, the model then understand that
[00:35:38] vinit shah: This is what a hand is, or this is what a, expression of a hand is, or this is what an action is happening, based on the text.
[00:35:47] vinit shah: So this comes in a pair, is it, when the model is getting trained?
[00:35:51] Anirban Paul: When the model is creating that data of noise.
[00:35:55] Anirban Paul: Along with that, it comes as a pair that noise plus text.
[00:35:59] Anirban Paul: Okay. Put that in numbers, it's a mix of numbers, like, two… like, you had, let's say, for example, you had, four-dimensional matrix, now you have, let's say, 6-dimensional matrix. Two-dimensional added by the text. Okay.
[00:36:13] vinit shah: Okay, okay. So, so, but this is… this combination is always there, right? Like, it's not like… say the images are trained separately, and then the, descriptions of it is, I don't know, trained separately, it's not like that.
[00:36:25] Anirban Paul: No, no, no, no, no. If it is a text-to-image model, it has to be there. If it is a image-to-image model, okay, then it's a different goal, okay? So sometimes image-to-image models, even image-to-image models are also conditioned. Image, along with that, text is given. And you say, like, improve this image.
[00:36:43] Anirban Paul: So, that is how it happens. So, even that kind of pairs are also trained.
[00:36:48] Anirban Paul: Okay, that is also text-to-image models only.
[00:36:51] vinit shah: Okay.
[00:36:52] Anirban Paul: It's… it's just that.
[00:36:54] Anirban Paul: When you use image-to-image models.
[00:36:57] Anirban Paul: The kind of input that you take is an image plus text. When you use text-to-image models, you just give the text.
[00:37:04] Anirban Paul: Okay, so you just do the reverse diffusion process. When you give tech image-to-image models, you give… you do the entire process.
[00:37:12] vinit shah: Okay.
[00:37:13] Anirban Paul: So that is… so image-to-image models are still now not so pop…
[00:37:18] Anirban Paul: like, it is fine, like, Nano and all are doing a good job. Okay, it is both the image-to-image model as well as text-to-image, it can identify both. Okay.
[00:37:27] vinit shah: Yeah.
[00:37:28] Anirban Paul: But still now, text-to-image has improved really well. Like, that was struggling for… since 2015 to 2022, it was struggling. In 2022, the concept of stable division came by a company known as RunwayML.
[00:37:40] Anirban Paul: Okay, RunwayML, just like OpenAI was, you know, doing pretty good with their GPTs and all, so RunwayML was very good with their diffusion architecture, because OpenAI was still stuck in their GANs and all those things, okay? That was their approach.
[00:37:54] Anirban Paul: It is RunwayML that actually did all this thing. RunwayML also has a video model as well nowadays. Now they have fully focused into video. Okay. Okay. And that is how it is, yeah.
[00:38:04] vinit shah: Okay, so just one more question with regards to this, right? So, say, for example, the GPDs and these AI tools and all, they have the options of
[00:38:10] vinit shah: we can put a snip of an image, right? Like, forget the straightforward text ones, but if I'm creating a program.
[00:38:16] vinit shah: the output of it can be put as a snip and said that, okay, I want to change this to so-and-so and all that. So, is the diffusion concept even coming up over there, or that's a different concept altogether? Like, how is it being able to read that this is the output and, you know, I need to change from red to black, or I don't know, small box to big box, and…
[00:38:33] Anirban Paul: No, see, if it is the image generation, then there is a diffusion process.
[00:38:36] vinit shah: Okay.
[00:38:37] Anirban Paul: If it is image through text, then it is not a division model. Okay. Then it… then it is plain, the concept of images being converted to their image vectors. Okay. And you… and there is a shared embedding space.
[00:38:49] Anirban Paul: Shared embedding space means… shared embedding space concept is there for everything, okay? Shared embedding space means you have a mix of image vectors and sound vectors or audio… or text vectors, okay? Okay. So, let's say…
[00:39:02] Anirban Paul: you have image vectors, which are made to look like a text vector. So…
[00:39:08] vinit shah: Okay.
[00:39:08] Anirban Paul: Let's say, if I explain you in a form of I have your image.
[00:39:14] vinit shah: Okay.
[00:39:14] Anirban Paul: you in a form of numbers, from your image, and with your behavior that you have towards me, I explain you in a form of numbers. And this explanation that I have about you in a form of numbers, from both of your behavior as well as image, this explanation, this representation that I have about you, should be similar.
[00:39:34] Anirban Paul: Got it. So, you… Vinit coming from image to… image to numbers, and Vinit coming from text to numbers.
[00:39:42] vinit shah: Okay.
[00:39:42] Anirban Paul: so similar that… that you can look up against each other. You can find similarity against each other.
[00:39:49] Anirban Paul: So… let's say I have, for example, I will, to answer your question, Let's say…
[00:40:00] Anirban Paul: I want to do a basic lookup. I'm just explaining a different concept of shared embedding space. If you understand this concept, you will understand everything around
[00:40:10] Anirban Paul: this… this mode to that mode. So I have a picture of…
[00:40:14] Anirban Paul: Let's say Vinit. I have a description about Vinit, okay? I have a description about Vinit. Vinit is, let's say, for example, Vinit is tall.
[00:40:23] Anirban Paul: Okay, let's say… Vineet… codes… Always.
[00:40:33] Anirban Paul: Always… where…
[00:40:39] Anirban Paul: a white t-shirt. So this is just to explain the concept, like, this, you will have thousands of explanation about you.
[00:40:45] vinit shah: Got it, yeah.
[00:40:45] Anirban Paul: Let's convert this into vectors.
[00:40:49] Anirban Paul: Okay, so let's say this is the vector meaning of this.
[00:40:56] Anirban Paul: This is the vector meaning of whatever I have told you.
[00:40:58] Anirban Paul: Okay, or I have 3-4 vectors like this. Okay, explain what… the entire description about you.
[00:41:04] Anirban Paul: Now, I give a picture of you.
[00:41:07] vinit shah: Okay.
[00:41:08] Anirban Paul: Okay, that picture…
[00:41:10] Anirban Paul: of you, let's say you coding on computers, okay? So, something, something like that, you are coding on a laptop, okay, so that picture is there. So, the goal of shared embedding space is such that that this embedding
[00:41:25] Anirban Paul: this embedding and this embedding, the image embeddings of you, and the text embedding of you, should be very similar to each other. Then only you can do a lookup.
[00:41:34] vinit shah: Got it.
[00:41:35] Anirban Paul: So then only I can find If, let's say.
[00:41:39] Anirban Paul: This is Vinith. Then only I can find this is Vinith. So if this is Vinith, then this is my nearest text vector I have.
[00:41:46] Anirban Paul: Okay, if from the image… I… image also has a vector?
[00:41:50] vinit shah: Got it, yeah.
[00:41:51] Anirban Paul: And that vector should be as close as each other.
[00:41:54] vinit shah: Got it.
[00:41:54] Anirban Paul: So, when these models are trained, if you have a shared mode along each other, then there
[00:42:00] Anirban Paul: Common understanding. What is the common understanding? Vectors. Their vectors' representations should be very close to each other. Then only I can compare image with text, and text with image. Now, I can do a lookup. From here, I can find, okay, Vinit is coding, let's say. Vinith codes always. This I found.
[00:42:16] Anirban Paul: So now… Once I have found this.
[00:42:19] Anirban Paul: Okay, I can create different versions of this, I can create a LLM, I can use a LLM to create an auto…
[00:42:26] Anirban Paul: Using the autoregressive model, I can create any other version of it. We need this coding here.
[00:42:31] Anirban Paul: Okay.
[00:42:32] vinit shah: Got it.
[00:42:33] Anirban Paul: Got it. So that is the thing. The concept of shared embedding is that if you have multiple, various types of modality in your data, then your embedding space of explaining those modalities in such a way that they
[00:42:46] Anirban Paul: That even from emails, you can create text, and from text, you can create emails.
[00:42:50] vinit shah: Okay.
[00:42:52] Anirban Paul: Got it.
[00:42:53] vinit shah: Got it, got it.
[00:42:56] Anirban Paul: So yeah, so do, guys, like, whenever you are starting, like, you all can refer around shared embedding space to understand
[00:43:03] Anirban Paul: these concepts a bit more, but this is the entire concept about shared embedding, that when you have that lookup kind of thing, like, what is this guy doing and all, so there is a text, there is a numerical explanation of this image.
[00:43:17] Anirban Paul: And that numerical explanation has to match with similar kind of numerical explanation you have about the text also.
[00:43:23] Anirban Paul: If you can match that, then you can create these lookups. Let's say I have tens of thousands of explanation about what this person is doing. So, which one to show?
[00:43:32] Anirban Paul: which word to show next. And then you can apply a causal LM, and… or a, you know, predict the next word kind of ILM, and just predict… keep on predicting, keep on predicting, like that.
[00:43:42] Anirban Paul: Okay, so anyways, so that is about CHI's reverse diffusion and conditional reverse diffusion process. So, we will now go to our practical, of, you know, exploring these models. I will share you one of the notebooks,
[00:43:59] Anirban Paul: over here.
[00:44:02] Anirban Paul: And I will share you an experimentation notebook that is for you all to play around with these models, okay? This one, every one of you, try along with me. You'll play it. It might take some time to play, I just now played it. Okay, so…
[00:44:16] Anirban Paul: In this model, what we are basically doing, first open the collab, first open the collab notebook, everyone.
[00:44:38] Anirban Paul: So, Vineet, you understood that the vector has to be of same language.
[00:44:43] vinit shah: Yeah, I understand that, right? Just one, one counter question to this is this, right? Say, for example, when text to images, I'm using the diffusion model, correct? Where you said that the model learns, by, diffusion and reverse diffusion. So, over a period of time, millions of photos, you know, like, I was just taking hand as an example, because
[00:45:01] vinit shah: That was one thing which never used to come properly whenever I give prompt, like, sometimes there are six fingers, sometimes the fingers look, weird. Till now, I have always seen that as a problem.
[00:45:10] vinit shah: Right? So that part, I understand that. Over a period of time, say, millions of photos, it knows, and when I say waving hand, it would have understood that. This is how I create a waving hand. That is the diffusion model, which help it understand. My question was, like, now the claud is there, right? I wrote a code with it, saying that, ki, create a page which says, hello world.
[00:45:28] vinit shah: Now, Hello World came up, I took a snip of that Hello World, put it in Claude, and told it, I want the Hello World to be looking red in color, or, I don't know, say, purple in color.
[00:45:38] Anirban Paul: This is not the fusion process.
[00:45:39] vinit shah: So, what does that come up? Like, how does it understand this particular, way of reading the images? So, is this also trained on certain data that is available, or what is the.
[00:45:52] Anirban Paul: No, no, no, it is definitely, it is definitely trained on our data, so I will tell you. So, how… if you wouldn't have taken a photo, how you'd have, how you would have done it? Otherwise, you would have copied, you would have probably copied that area and would have sent it.
[00:46:04] vinit shah: Yeah.
[00:46:05] Anirban Paul: Okay, if you would have copied that code.
[00:46:08] Anirban Paul: what it would have internally done. It would have converted to a text vector.
[00:46:12] vinit shah: Correct.
[00:46:13] Anirban Paul: Right? And that text vector, from that text vector, it would have found out an answer.
[00:46:17] vinit shah: Correct.
[00:46:18] Anirban Paul: similar thing it does. Imagine, from that image vector, it is converting that to a number vector, okay, image vector, and that image vector has a similar
[00:46:27] Anirban Paul: kind of…
[00:46:29] Anirban Paul: Appearance, just like how your text vector would have looked for the same thing if you had copied it.
[00:46:35] vinit shah: Carted.
[00:46:37] Anirban Paul: It is just like a translator in sitting in between. So your vector is just like a common translator who visits, with the Prime Minister to different countries.
[00:46:46] vinit shah: Okay.
[00:46:47] Anirban Paul: Depends upon the language he talks.
[00:46:50] vinit shah: So basically, there is a mechanism which allows it to interpret the picture or a snippet.
[00:46:56] Anirban Paul: In a formal learning with it.
[00:46:57] vinit shah: language, and then it then translates my requirement, and again, puts it back and learns it. So that is how it is.
[00:47:03] Anirban Paul: And that common language is known as shared embedding space.
[00:47:06] vinit shah: Oh, okay.
[00:47:07] Anirban Paul: So, shared embedding space means whether you are coming from image, audio, text, anything, video, your representation should be similar to each other.
[00:47:16] vinit shah: Got it.
[00:47:17] Anirban Paul: Imagine you have not given the image only, you have given the text only, as if.
[00:47:22] vinit shah: Huh, huh.
[00:47:23] Anirban Paul: And that is where the improvement of the image model has happened in the last few years, that your images, repairs, also from repair to text should be so good that, as if you have copied the text only.
[00:47:34] vinit shah: Got your point. So basically, if I would have written it that I got an output, which is Hello World in black color, but now make it as red color, it is at the center of the screen. I want it at the top right. So the same thing, it is able to even interpret it, looking at the picture that I have presented it, and that is where the concept of share embedding comes into picture.
[00:47:54] Anirban Paul: Shared embedding space, yes.
[00:47:55] vinit shah: So the integrity…
[00:47:57] Anirban Paul: Yeah, the embedding space is such a way, it is not text vector separately, image vector separately, and all these things. Now that these models, who does this common embedding.
[00:48:08] Anirban Paul: Okay, they are trained in such a way that their embedding should be near to each other. Got it. They have a common region towards their embedding, okay? Whether you create image, that image should also look like
[00:48:18] Anirban Paul: If I write, we need this coding versus actually Vinitis coding of image, both should have a same representation.
[00:48:26] vinit shah: Got your point.
[00:48:27] Anirban Paul: Similar, or similar.
[00:48:28] Anirban Paul: Because image will also come with pixels, so some white color, black color, all those things will also come. Okay, so those are added additional things, but the core part should be, like, your numerical representation should be similar, and that is the concept of shared image.
[00:48:42] vinit shah: Okay. Yeah, I'll read more around this, and if there is doubt…
[00:48:44] Anirban Paul: What's shared embedding.
[00:48:46] vinit shah: Yeah, next…
[00:48:46] Anirban Paul: embedding space, how sh… explain me the concept of shared embedding, with respect to image-to-text model and text-to-image models.
[00:48:56] vinit shah: Got it. I'll do that. Thanks a lot.
[00:48:57] Anirban Paul: Yeah, yeah, yeah.
[00:48:59] Anirban Paul: Okay, now guys, I have shared this, I hope everybody is open to… able to open this. Everybody open? Collab Notebook?
[00:49:07] vinit shah: Yes.
[00:49:12] Anirban Paul: Okay, so now, first, we install your torch and diffusers and all these things, so, and these are some… our basic requirements for installation. We also will be, using Matplotlib because we want to plot the image.
[00:49:28] Anirban Paul: And, yeah, we'll go ahead and do these installation. I have already done it. Okay, so then we will import
[00:49:34] Anirban Paul: from Diffuser's import stable diffusion pipeline.
[00:49:37] Anirban Paul: Okay, so this is the pipeline. So this might trigger your hugging face, it might ask you, like.
[00:49:44] Anirban Paul: whether you Hugging Face access is there or not, but I don't think I have removed all the dependency, but let's see, let's see.
[00:49:53] Anirban Paul: Now, we take the stable division model, we first download the model, you pass the model name, RunwayML, Stable Division 1.5. Version 1 to 5, this 1-5 means 1.5. Okay, so the model will… is getting downloaded, and after that.
[00:50:08] Anirban Paul: We use the… since this is an offline model we are using, this is an open source offline model, we are not using any API call to use this model. So, obviously, it will use your GPU or your CPU availability of your collab. So, first, we check if the GPU is there or not. Device equals to CUDA, you all know that it checks for your GPU. If
[00:50:27] Anirban Paul: the CUDA is available, then use… else use CPU, okay? So whatever is your device, on that, please shift your pipeline onto that. Shift your diffusion pipeline that you have created, shift in diffusion model, move it to CPU or GPU, depending upon your availability. Okay.
[00:50:44] vinit shah: One question, sorry, Anderman. So, this one, right, where you're doing it via this device, could dive, blah blah, is there a way to do it at our system level? Like, whenever we are doing something so that it, by default, takes it up rather than checking it, or that's not possible.
[00:50:59] Anirban Paul: Like, you want to do this in your local system.
[00:51:02] vinit shah: Madhlab, I have a GPU under CPU, imagine. So, is there a way to say that I want all this computation to happen on GPU at a system level, rather than, you know, having these checks or something like that?
[00:51:12] Anirban Paul: Yeah, then you always make your device equals to CUDA only.
[00:51:15] vinit shah: Oh, God.
[00:51:15] Anirban Paul: We don't have to put this if If condition.
[00:51:18] Anirban Paul: Okay, so if you look at the logic, it is CUDA if CUDA is available. So CUDA means your GPU is available. Else CPU. So if you remove this, then this becomes permanently CUDA.
[00:51:29] vinit shah: Kuda, okay.
[00:51:30] Anirban Paul: Okay, so now we have created some prompts, a serene sunset over calm lake and a bustling cityscape at night.
[00:51:37] Anirban Paul: we will take one of these… each of these prompts, and we'll try to create an image. So we'll run a for loop, we will read one prompt at a time, and we'll just enumerate over the image prompts list that we have.
[00:51:48] Anirban Paul: We will take…
[00:51:49] Anirban Paul: One image at a time, and we'll try to create that image, and then whatever image gets created. Okay, this is how you create it. Pipeline, pass the prompt, dot images. If you do this, image will be created, and whatever image gets created is also being saved, as well as we are also appending that image into a list.
[00:52:08] Anirban Paul: Such that I can show it later. So, this is using Matplotlib, we are showing, displaying the images. Okay, so 3 images, 3 different matprodLib.
[00:52:16] Anirban Paul: Okay, so everybody run till here, and tell me if it is running for you. Okay, I've already ran it once.
[00:52:21] Anirban Paul: So, it might take some time for you to run. So, when you are running, no, you might see like this, that this 50 by 50.
[00:52:29] Anirban Paul: So this 5550 is the number of time steps. It is known as inference steps, for reverse diffusion. So this is, like, how many reverse diffusion processes will be there to remove the noise. So lesser, if you do, you will see more bloody image, more noisy image. If you… this is not a noisy image, by the way, guys. This is a…
[00:52:48] Anirban Paul: That is a cityscape, so as you wrote cityscape, it draws… it drew a painting kind of a thing.
[00:52:53] Anirban Paul: Okay. So, if you… if you want to change this value, I will show you those experiments as well. Like, I have a different notebook file for you all to play around with this, and all of this creation, image creation, takes a lot of time, okay? It might take you 10 minutes to create this image. First, try out, guys.
[00:53:09] Anirban Paul: Everyone try it out.
[00:53:50] vinit shah: I got the first photo different, Anirvan, compared to what.
[00:53:53] Anirban Paul: Yeah, you will get a different photo only, because it is generation model, no? So there is…
[00:53:58] vinit shah: No, like, the… the second and… the third photo is, like, similar. Second is almost similar, but the first one is totally different.
[00:54:05] Anirban Paul: Yeah, it'll be different, because that's where the creativity and generation comes, or else it becomes a static model, no?
[00:54:10] vinit shah: Got it.
[00:54:11] Anirban Paul: So now… What about the rest, guys? Everyone created?
[00:54:17] Anirban Paul: Step one?
[00:54:17] Deepan Kanagaraj: One question, one question. In the pipeline, one of the parameters you have given was, tarch underscore DD type as float 16. Yeah. Yeah. If I readjust this float parameter, will that have any impact on the model or image?
[00:54:34] Anirban Paul: If you readjust it to, what, 32?
[00:54:36] Deepan Kanagaraj: 32 or 64.
[00:54:37] Anirban Paul: You can try 32. I have personally researched and I have seen, usually, on GPU, 16 is more safer.
[00:54:45] Anirban Paul: Okay, see, if you make it bigger, your GPU weightage will become increasingly high. Okay, the tensors. So this is, like, how your tensors will be represented, in a form of 62… 16-bit or 32-bit. If you make your GPU weightage higher,since we are using a free GPU, which
[00:55:01] Anirban Paul: collab gives us. So, I have seen it becomes a little unstable. You can try with 32.
[00:55:06] Deepan Kanagaraj: Okay.
[00:55:07] Anirban Paul: But if you are using CPU, you can try 32, no problem.
[00:55:10] Anirban Paul: But CPU will work out or not, that is a different issue.
[00:55:13] Anirban Paul: But you can play around, there's… these options are there. You can… you can change this to 32 and try it out.
[00:55:19] Deepan Kanagaraj: Okay, sure.
[00:55:58] vinit shah: So, Nirvan, again, one more question. So, chances are an even of us might get the same photo, or that is very low?
[00:56:04] Anirban Paul: Very, very similar, but not…
[00:56:06] vinit shah: Okay.
[00:56:06] Anirban Paul: Exact same is very rare. Okay, similar photo you will get.
[00:56:10] Anirban Paul: Okay.
[00:56:10] vinit shah: Oh, God.
[00:56:11] Anirban Paul: So, while everybody is done, then I can show you the next part. This is for you all to try it out, okay, because this is going to take a lot of time for you all to experiment. So, I have some 6-7 experiments for you all to try it out. Okay, I'm sharing you this experiments file, and I'll explain you what you will have to do.
[00:56:28] Anirban Paul: Okay, I have very well detailed it out, okay, for you all to consume it. So this is like, you know, you…
[00:56:34] Anirban Paul: trying out with different, different prompts. Okay, so this is like that. So I have shared you this…
[00:56:39] Anirban Paul: file. Can you all access this file?
[00:56:52] Deepak Bobade: Yes.
[00:56:53] Anirban Paul: Yeah. Don't play anything over here, okay?
[00:56:57] Anirban Paul: Don't have to play anything over here. I had played it, and it is still now running it. Okay, so don't have to play anything over here.
[00:57:03] Anirban Paul: You do one thing.
[00:57:05] Anirban Paul: is there are some experiments we are doing over here. So, first experiment is… I will explain the first one, and the second, third one was… it's very well written. The first experiment
[00:57:16] Anirban Paul: is… Everybody, are you hearing?
[00:57:20] vinit shah: Yes.
[00:57:21] Anirban Paul: Yes, you have to be very much with me, because you have to try these out.
[00:57:25] Anirban Paul: Okay.
[00:57:26] Anirban Paul: So… First experiment is, if you want to play around with these models, You have a prompt variation.
[00:57:35] Anirban Paul: Okay, so take one other prompt, a futuristic sports car. Then take a futuristic sports car, but with cinematic lighting.
[00:57:41] Anirban Paul: Also, take a futuristic sports car with cinematic lighting and ultra-realistic.
[00:57:46] Anirban Paul: A futuristic sports car, cinematic lighting, ultra-realistic, 8K, highly detailed. A futuristic sports car, cyberpunk city background, neon lights, cinematic, highly detailed.
[00:57:55] Anirban Paul: Just take this prompt variations instead of the prompt variation you saw here. So, you have played this, okay?
[00:58:01] Anirban Paul: Now, you can remove this, and take those, and paste it here.
[00:58:06] Anirban Paul: Okay, that… 8.
[00:58:09] Anirban Paul: And paste it here.
[00:58:11] Anirban Paul: Okay, and from here, Image prompts. And from here, you try out how your pictures are coming in.
[00:58:20] Anirban Paul: That is the first experiment. Got it?
[00:58:24] Anirban Paul: Everybody understood what is the experiment that you'll have to do.
[00:58:28] Deepak Bobade: Yes.
[00:58:30] Anirban Paul: Okay, that is the first thing, to play around. This is to play around that with the same prompt, and with just little, little addition on top of that, how much it is changing. You will get an idea about stable division if you get a hang of stable division, like, you know, whether you can get the similar image or not, and all those things will be very much clearer now. Got it, Vinit?
[00:58:48] vinit shah: Sure, sure, trying it out.
[00:58:49] Anirban Paul: Yeah, only the first experiment is a little spreaded out, that's why I explained it, because this is the entire code of the first experiment. I was trying it out. Here, it never… it never plays, okay? It never plays, because of the different versions I was trying to load. Okay, no worries. So, you copy this and try it there. Now comes to the second experiment onwards. It is very well written now.
[00:59:08] Anirban Paul: Experiment number 2 is…
[00:59:11] Anirban Paul: See, there is a code that you had used like this, pipeline prompt image. You had used this code, right?
[00:59:17] Anirban Paul: This… Change this from here, change this to this.
[00:59:23] Anirban Paul: Where you are giving the prompt, you're also giving number of inference steps. What was number of inference steps, guys?
[00:59:32] vinit shah: 50, I think.
[00:59:34] Anirban Paul: Yeah, 50. The denoising step. So, by default, it is 50. You change it to 30 and try it out.
[00:59:40] Anirban Paul: This is one experiment. Okay, you can try it out with 10 as well, you can try it out with 50 as well. 50 was already there, so you can try it out with 10 and 30. So 50, 30 will see balanced, 50 will be… 10 will be faster, less detailed image, and 50 will be low… slower. This is the idea.
[00:59:56] Anirban Paul: Okay, now…
[00:59:58] Anirban Paul: One more thing. Along with that, there is another parameter, hyperparameter, known as guidance scale. Guidance scale is how much
[01:00:06] Anirban Paul: Your prompt should be followed by your model.
[01:00:09] Anirban Paul: So, guidance scale also ranges between 1 to 20 is the actual value we start. To use style, you try this with guidance scale equals to 3. So, change this line, basically, to this line. Add this guidance scale as a prompt over
[01:00:27] Anirban Paul: Over here, instead of this, comma, add this.
[01:00:38] Anirban Paul: Over here, just add this, and try generating an image, and try with different diet and skills values.
[01:00:43] Anirban Paul: So these are experiments that you are going to do as a… as a kind of, you know, as a task that you are going to do. Okay, do try this, like, post your class, because this is going to eat a lot of time, okay? This entire thing might take you half an hour or one hour to do, the entire thing, okay? Trying out so many values, okay? So…
[01:01:03] Anirban Paul: Now, One more thing, one more experiment. This is experiment number 3.
[01:01:11] Anirban Paul: Sorry, experiment number 4, because 3 was this. Experiment number 4. Instead of changing the prompts, keep one prompt and compare the guidance. Keep one prompt.
[01:01:22] Anirban Paul: Just change the guidance value.
[01:01:24] Anirban Paul: Try with different, different guidance. See how it is changing.
[01:01:28] Anirban Paul: Okay? That you do. That you can do in the last one only, so same kind of experiment only.
[01:01:33] Anirban Paul: Then… Same thing you also do with different, different inference steps also. Same thing.
[01:01:40] Anirban Paul: Now comes to the next one, that is adding a seed. Do you all know what is a seed in Python?
[01:01:47] Deepak Bobade: Yeah.
[01:01:49] Anirban Paul: What is a seat? What does it help you to do?
[01:01:51] Deepak Bobade: So, it kind of, fixes the randomizer, to give same output.
[01:01:58] Anirban Paul: Yes, so…
[01:02:00] Anirban Paul: You have a random number generator. So, from there, you… if you fix a seed, so this is how you fix a seed.
[01:02:07] Anirban Paul: So, instead of changing the above line, you had this line to create an image, change it to this entire line, where you provide a manual seed like this, you can provide a seed of, let's say, 42 over here, which you have taken over here. Okay.
[01:02:22] Anirban Paul: And you can put a seed value over here, 42 or whatever it is, and try generating. We need… this will create
[01:02:28] Anirban Paul: Your… this will answer your question, can we create the same image at all?
[01:02:33] vinit shah: Okay.
[01:02:35] Anirban Paul: With the same prompt, you try out. Okay, again and again. The same seed, you try out.
[01:02:42] Anirban Paul: Okay, so this is your one of the experiments.
[01:02:45] Anirban Paul: And then… there is… Experiment of same prompt, different seeds. That also you try out.
[01:02:53] Anirban Paul: Okay, change the value of different seat with the same prompt, see how it is changing.
[01:02:58] Anirban Paul: Okay, and the last thing is…
[01:03:01] Anirban Paul: There is another image generation, object that you can create, which is a parameter that is negative prompt.
[01:03:09] Anirban Paul: So, negative prompt is, your LLM will… your diffusion models will try to avoid these things. It will try to avoid bloody, low quality, distorted, bad anatomy, ugly and deformed images.
[01:03:22] Anirban Paul: Okay, your LLM will try to avoid… avoid this. So, see how it is obeying the negative prompt as well. Along with prompt, you pass negative prompt as well. So, it is like…
[01:03:31] Anirban Paul: That this is what it should avoid at any cost.
[01:03:36] Anirban Paul: Got it? Everybody? These are the experiments that you all have to try with these diffusion models in your part-time and you're playing around with it.
[01:03:45] Anirban Paul: Okay? Clear?
[01:03:56] Anirban Paul: Okay.
[01:03:58] Anirban Paul: Okay, guys. So, guys, today, understood all the topics.
[01:04:05] Anirban Paul: Till now.
[01:04:09] Anirban Paul: Anything went above your head, diffusion process, or anything, reverse diffusion, conditional diffusion?
[01:04:13] Anirban Paul: All these things, like.
[01:04:16] Anirban Paul: Because with that, we will wind up stable diffusion and image generation models. So this… these are your… some of the image generation models that you might be seeing on public.
[01:04:27] Anirban Paul: Mostly these, because these are gone. Style dance and all these things hardly we use. So…
[01:04:32] Anirban Paul: you will see DALI 3, Midjourney, Stable Diffusion, which is by RunwayML, then you see Nano Banana, Flux, then OpenAI is using Sora video generation.
[01:04:42] Anirban Paul: That also we are seeing from the diffusion models. So, these are the applications that you are seeing. So diffusion models are heavily… heavily used to create, how would
[01:04:53] Anirban Paul: A particular type of cancer cell would look like on
[01:04:59] Anirban Paul: this tissue of your body. So, let's say you… you have lots of tissue pictures of… of humans, and you would give, let's say, how would this cancer cell look like on this tissue?
[01:05:11] Anirban Paul: Yes, I will share this VPT. Just a moment.
[01:05:28] Anirban Paul: Yeah, so… so, these are some of the models that are popular, and one of the good applications about deficient models, I am seeing, like, with one of my, colleagues, he's also working on one of the
[01:05:41] Anirban Paul: research healthcare company, where they are using a lot of diffusion models to create how cancer cells
[01:05:49] Anirban Paul: would look like with prompting only, with different type of prompting, different type of experiments, these experiments I told you, with different types of parameters, experimentations, and with different advanced models, like, not only Stable Division 1.5, maybe more advanced models which can support in your GPU, they're also trying to create
[01:06:07] Anirban Paul: some of the, you know, these cancer cells, how they will look like on an image, on a tissue, on a skin, on all these things. So they are trying to detect, like, after six, seven stage, how it will look like. So, yeah, so those are some of the applications I have seen.
[01:06:23] Anirban Paul: You know, around me. Apart from that, there are a lot of startups as well. There is a startup known as…
[01:06:28] Anirban Paul: At creative.io.
[01:06:31] Anirban Paul: So, ad creative, it has been renamed to adcreative.ai. Now, ad creative actually creates images for marketing. So, if you…
[01:06:41] Anirban Paul: if you give a prompt, okay, or if you give an image and just give some prompt, it will create a very product-ready image, or some product shoot images, like this. This is for what?
[01:06:55] Anirban Paul: There's some sort of a nail polish or something, maybe.
[01:06:58] Anirban Paul: Okay, yeah. So, then for headphones, this is how it looks like. So, this is also one of the startups that is really doing really well. So, this is, like, Valentine's Day's images. All these things are done using these diffusion models and AI generations. Okay, these are some of the startups which are really excelling. Okay, they have also done, like, in 18 months, I think they reached
[01:07:22] Anirban Paul: Around 40 million ARR, okay, within 18 months. So this guy is a Turkish guy who was the founder, Tufan Gawk, his name is. So he did that. Anyways, so these are some of the applications that I have seen around myself, marketing and healthcare, okay?
[01:07:42] Anirban Paul: Diffusion models, image generation as a whole, I believe it has a lot of
[01:07:46] Anirban Paul: ways to go from here onwards, like, it has a lot of improvement to make. Diffusion models are, guys, slowly, slowly coming towards text generation as well. Okay, there are a few reasons to this, so that is the last thing I will tell you before we go.
[01:08:01] Anirban Paul: To agentic and LLMs and RAGs. So, one thing is.
[01:08:09] Anirban Paul: One thing about diffusion models, autoregressive versus diffusion is, so all your LLMs are autoregressive in nature. Okay, they are autoregressive in nature. And you all said that you all know about autoregressive, right? You all know autoregressive, right?
[01:08:29] Anirban Paul: Everyone knows.
[01:08:37] Anirban Paul: Deeper, deeper.
[01:08:39] Anirban Paul: Auto…
[01:08:40] Deepak Bobade: Autoregressive, as in, which has an encoder, decoder, both, or.
[01:08:45] Anirban Paul: No, no, no, it is… it is just the next word prediction.
[01:08:49] Deepak Bobade: Oh, okay, yeah.
[01:08:50] Anirban Paul: Okay, I will explain that when I come to LLMs, but now I'm just trying to say that your LLMs, many, slowly, slowly, division models are coming to text as well, text generation as well. Division models are coming to text generation as well, and they are, you know, they are…
[01:09:05] Anirban Paul: being looked upon by a choice as an alternate to autoregressive models, okay, diffusion models are. The problem with diffusion models are they have… they have a high requirement of high compute.
[01:09:18] Anirban Paul: Okay, autoregressive doesn't require high compute compared to division models. That's why division models took so much time to generate, no? Okay, they have high compute requirement, usually. And autoregressive models has a very low compute requirement.
[01:09:31] Anirban Paul: Because… There… the pros and… there are some pros and cons. In auto-regressive model.
[01:09:38] Anirban Paul: What happens is, you are trying to predict the next word at a time.
[01:09:43] Anirban Paul: Okay.
[01:09:44] Anirban Paul: In diffusion models, that is not the case. In diffusion models, it doesn't predict the next steps. It predicts all the possibilities that could happen.
[01:09:54] Anirban Paul: all the possibilities. So, from the current state, where it can go, it predicts… it can predict all the possibilities. Okay, diffusion model. So, that is there. Like, this is… that is a diffusion model of text, okay? So, but the problem with diffusion models is that
[01:10:09] Anirban Paul: It is… it has worked good for images. Why? Because, see, guys, if you… if you have a blue dog.
[01:10:17] Anirban Paul: Okay, for example, if you have a blue color dock.
[01:10:20] Anirban Paul: Okay, I know, I know it's not possible, but let's say, hypothetically, if you have a blue-colored dog, Even…
[01:10:27] Anirban Paul: Even, let's say, if you add a little distortion to it, or little color here and there, or little, you know, fading, little, you know, filtering effect if you do.
[01:10:37] Anirban Paul: A blue is still a blue.
[01:10:40] Anirban Paul: Right?
[01:10:42] Anirban Paul: A blue is still a blue.
[01:10:44] Anirban Paul: Right?
[01:10:46] Deepak Bobade: Yes, yes.
[01:10:47] Anirban Paul: Are you… are you getting the thought process, guys? So, if I take… if you take my photo, and apply a little bit of noise onto it, little bit of noise.
[01:10:54] Anirban Paul: So, your photo is still… it is an advanced photo only, okay, with the noise added.
[01:10:59] Anirban Paul: So, I'm not talking about the diffusion, reverse diffusion process and all this. I'm telling, in general, if your model predicts, instead of Anirban, it predicts almost something like Anirban. It is also fair to
[01:11:10] Anirban Paul: For a division model, it is fair.
[01:11:12] Anirban Paul: But in text, this might not work out. Because, let's say if you have cat, If you have cat.
[01:11:20] Anirban Paul: By mistake, if you make it chat, the complete meaning changes.
[01:11:25] Anirban Paul: So that's where your diffusion models are a little struggling. So that's where autoregressive is still better.
[01:11:30] Anirban Paul: Okay, division models can create a very drastically different version
[01:11:35] Anirban Paul: If it… if you look with respect to text.
[01:11:38] Anirban Paul: Okay, but with respect to image, it is not there. Okay. A lake?
[01:11:43] Anirban Paul: Which is…
[01:11:45] Anirban Paul: Green, lake which is blue in color, you add a little bit of noise, it might be a little light blue, but it is still blue.
[01:11:50] Anirban Paul: Okay, so that is why diffusion models are working fine with images. With text, it is trying to come, it is slowly, slowly, they are seeing that diffusion models is being also being explored upon on text, but it has not achieved that much success. But there are some pros and cons.
[01:12:03] Anirban Paul: The pros are, like I told you, Division Models has more possibilities of creativeness. Okay, it can come up with more random predictions, like.
[01:12:11] Anirban Paul: like Vineet was saying, like, every time you're generating something different. But in autoregressive, guys.
[01:12:18] Anirban Paul: you will not drastically generate everything different, okay?
[01:12:22] Anirban Paul: the cat sat on a mat. If you write, if you ask it to predict, the next time also it can predict mat. But the image, exact image cannot be created. Very rarely it creates.
[01:12:33] Anirban Paul: some pixels here and there, it will happen. But still, that image still indicates the same
[01:12:39] Anirban Paul: Similar explanation. But in your text, if you do that, the complete meaning changes.
[01:12:44] Anirban Paul: So, these are the pros and cons about diffusion models. It is very high compute, but there are… and also some negativity about diffusion models, but it is looked upon as more for creativity, better for creativity, whereas autoregressive are lesser compute.
[01:12:57] Anirban Paul: But it only outputs within a limited amount of, you know, output. That is good also, it is a little deterministic compared to diffusion models.
[01:13:07] Anirban Paul: Okay.
[01:13:09] Anirban Paul: So, yeah. So… so that is the current thing. Diffusion model is still now in research. On the video side, heavy research are going on. Okay, I'm not saying that they have achieved the pinnacle of
[01:13:20] Anirban Paul: image generation yet. Okay, I think still it will improve. Okay. But, yeah, but some level they have achieved, that's why they have moved on to video generation. Heavily into video generation.
[01:13:30] Anirban Paul: Okay.
[01:13:31] Anirban Paul: So, that is all, guys, about VAE, intro to VAE to GANs, which are your older models, to Diffusion, which is the current state of the art model. Okay, now, I will share this thing also with all of you.
[01:13:49] Deepak Bobade: Could you also ask our, manager to upload it on the LMS?
[01:13:57] Anirban Paul: Oh, yes.
[01:14:05] Anirban Paul: Just a moment.
[01:14:25] Anirban Paul: Akansha, are you there?
[01:14:28] Anirban Paul: Oh, okay.
[01:14:29] GenAI Batch-2 Manager: Yes.
[01:14:30] Anirban Paul: So I'm…
[01:14:30] GenAI Batch-2 Manager: True.
[01:14:31] Anirban Paul: Yeah, I've given a lot of links, so I've given one link now, and I've given two links before this, which is of collab, to download a notebook file and just upload.
[01:14:42] GenAI Batch-2 Manager: Okay.
[01:14:44] Anirban Paul: Okay, now guys, something very interesting we are moving towards.
[01:14:49] Anirban Paul: Which is… Agentic bot. Okay. So…
[01:14:54] Anirban Paul: Next, I would say, just a moment, I would check your… Toc a bit.
[01:15:02] Anirban Paul: From my notes.
[01:15:04] Anirban Paul: So, next…
[01:15:07] Anirban Paul: Few sessions is only about AI agents. AI agents, AI agents development, then advanced agenting AI, then developing domain-specific agents, developing agents that automate productivity and planning.
[01:15:19] Anirban Paul: Developing agents that automate productivity and planning again.
[01:15:23] Anirban Paul: data research, and EDA agents, Okay.
[01:15:27] Anirban Paul: So, all these things are there. So, it's completely about agenting. So, I have a… like…
[01:15:34] Anirban Paul: we will follow this, but not, like, your topics might be here and there, so let's say
[01:15:41] Anirban Paul: In the next week only, we have, let's say, for example, MCP as a topic for
[01:15:53] Anirban Paul: Okay.
[01:15:55] Anirban Paul: So, next few weeks, as a topic-wise, let's say next week only we have MCP,
[01:16:01] Anirban Paul: As a topic. Okay, but MCP is supposed to be there towards the end, after you have learned agentic architecture. Then MCP makes more sense, because when you create the agents, then only as a MCP you can expose it.
[01:16:15] Anirban Paul: Okay, so we will shuffle the orders, but we will follow this TOC, but the orders will be a little different. Like, only the MCP will change, other orders are fine, and maybe the,
[01:16:27] Anirban Paul: Yeah, only the MCP will change.
[01:16:29] Anirban Paul: Okay, so that is how it is going to go, just letting you know this.
[01:16:36] Anirban Paul: Can you share the collab link? Yes, I will share it again.
[01:16:39] Anirban Paul: So this is the experimental link.
[01:16:42] Anirban Paul: That you will have to use and play with your experiment in the other notebook, which is the full notebook.
[01:16:50] Anirban Paul: This is the experimentation link.
[01:17:00] Anirban Paul: And this is the main link.
[01:17:04] Anirban Paul: Okay, so guys, you all understood the expectation, everybody?
[01:17:09] Anirban Paul: Yes, I will start with RAG. I will start… Pawan, today's topic is about RAG only, so we will… but I'm just letting you the expectation that there is MCP just next week, we will do that topic name is MCP, but after your agent is complete, then only we'll be able to do MCP. Otherwise, other order will be followed as it is. Okay.
[01:17:28] Anirban Paul: But it's just that, someday, You know, like today.
[01:17:33] Anirban Paul: Actually, this LLM agent's concept is supposed to start from next.
[01:17:37] Anirban Paul: But… You all… we ha- we all have done…
[01:17:41] Anirban Paul: the major topic of diffusion and guns, so it starts from today itself, okay? So, that is a thing. So, though, due to that, some orders might go here and there, okay? But the main thing is to follow the TOC. Okay, so now, first thing is, how do you
[01:17:56] Anirban Paul: build your first agentic bot. Okay.
[01:17:59] Anirban Paul: So don't go by the name. Agent is going to start way later, but this is a storyline that we have to starting our first concept of RAG. Okay, this is RAG only. Okay.
[01:18:11] Anirban Paul: We just now have talked about
[01:18:13] Anirban Paul: LLMs and division models. Now, slowly, slowly, let's relate to that. So, as you know, this is the space of Gen AI now. By now, you all have done with machine learning, you all have done deep learning, and then slowly, slowly, we… you are moving towards Gen AI.
[01:18:28] Anirban Paul: And you are going to build a lot of use cases around NLP, natural language processing. This is the entire domain that we all have, and
[01:18:37] Anirban Paul: You all know the use cases of Gen AI being content creation, image generation, port generation, fraud detection, idea generation, text generation, predictive analytics, although this predictive analytics, fraud detection, data analytics, this thing is still… ML is preferred. Classical ML and deep learning processes are still preferred.
[01:18:55] Anirban Paul: But all the other generation, wherever you are seeing generation, task management, all these things are still
[01:19:00] Anirban Paul: are being ruled, like, you know, it is very booming with generative AI.
[01:19:06] Anirban Paul: Now, what are large language models?
[01:19:16] Anirban Paul: What are large language models, guys? Question to you.
[01:19:22] Deepak Bobade: Models trained with a lot of numbers of, parameters.
[01:19:28] Anirban Paul: What was Bert?
[01:19:30] Deepak Bobade: It was a model, yeah.
[01:19:32] Anirban Paul: What is a, what is a… was it a large language model?
[01:19:36] Deepak Bobade: But… No, no, it wasn't. It wasn't, no.
[01:19:42] Anirban Paul: Okay, Gunjan?
[01:19:42] gunjan bhaiya: Yeah, but it was like a SLM model, you can say that it works on a very small scale of data with a limited
[01:19:50] gunjan bhaiya: Token size and everything.
[01:19:52] Anirban Paul: Yeah, you're right, but, you know, saying BERT as SLM might be too low for its fan, one of… the fan being me, so the actual name for BERT is a language model. BERT was just called as a language model, because that time, large and small, this concept didn't come up.
[01:20:08] Anirban Paul: Why language model? Because BERT was trained on 2,500 million dataset of Wikipedia, and 800 million unpublished books.
[01:20:17] Anirban Paul: Using that, BERT was trained. Using that two process of mask language model and next sentence prediction. You know this process? Mask language model and next sentence prediction. This process, you all know about BERT?
[01:20:31] Deepak Bobade: Yes, actually, BERT has the capability to, look into past as well as future, right? To predict the next token.
[01:20:40] Anirban Paul: Yes. And, there was a two training process that was used. One is masked language model, and one is next sentence prediction. Using that, it has become a very good encoder model, right?
[01:20:50] Anirban Paul: So…
[01:20:52] Anirban Paul: Now, I told you, from 2017, guys, you have to keep timeline into account, like, how these timelines of events have happened. So, from 2017 to 2020, 2021, I told you that BERT was ruling, and language… encoder model was doing very well. Okay, so your transformers was…
[01:21:09] Anirban Paul: going to become the future. And encoder was not improving anymore. So your bird was the best, and after that, no encoder good model has come up.
[01:21:18] Anirban Paul: That is at the level of BERT. A lot of decoder improvement has happened, okay? Decoder, imagine, like, GPTs. All your GPTs are a collection of decoders. Decoders means which take the encodings, embeddings, and predicts next word
[01:21:31] Anirban Paul: at any point of time, it just predicts the next word, okay? Given all the input words still now, what is the prediction of the next word? That is your autoregressive model. Okay. So, now…
[01:21:42] Anirban Paul: your GPTs started improving, and there came a time period, which is around 2022, where your GPT came up with large language model concepts. Why large? Because it is trained on a huge corpus of a dataset. It is not limited to 200 million Wikipedia and 800 million unpublished book. It is just trained on the entire internet.
[01:21:59] Anirban Paul: Okay, all data, all forums, all public data, all, like, you know, different type of forums, discussion forums, news forums, using all of them it got trained on. And plus, the large is not only because of that, the large name also comes because of the parameters. It is a huge amount of model.
[01:22:19] Anirban Paul: It is a huge size model, so that's why you hear the name 24B, 5B, 7B, that is the number of parameters that you have, 7 billion parameters. 24 billion parameters. So that B is for the number of parameters or weights that you have… that your model got trained on.
[01:22:35] Anirban Paul: Okay, so that's… that's where the large thing comes in. So large defines that. And language model, because it got trained on English language and, you know, human language. Not only English language, other languages as well, but works with human language. And it is an algorithm that trained… that got trained and got trained to do something based on our data, so we can call it as a model.
[01:22:53] Anirban Paul: Okay, so it… it has the… it can understand context, it can understand semantics in language, it can generate coherent text.
[01:23:01] Anirban Paul: That's why you have a very…
[01:23:03] Anirban Paul: You know, very predictable kind of output if you keep your temperature low.
[01:23:08] Anirban Paul: And…
[01:23:10] Anirban Paul: perform diverse tasks, like translation, summarization, question-answering. So that was the whole concept of LLMs, that it came onto the market.
[01:23:18] Anirban Paul: Okay.
[01:23:19] Anirban Paul: Now, some of the limitations of LLMs are, if you look at the underlying LLMs, the underlying LLMs doesn't
[01:23:27] Anirban Paul: have.
[01:23:28] Anirban Paul: memory with it. Doesn't have the concept of memory. We will learn about memory, but it doesn't have the concept of memory as a whole, so it… there is no concept of memory. Your LLMs are
[01:23:39] Anirban Paul: underlying LLM, if you go and ask, if you say, my name is Anitwan, and if you say the next question in the next call, that what is my name, it will not be able to sell you.
[01:23:47] Anirban Paul: Because in the context, there is no memory.
[01:23:52] Anirban Paul: In the context, if you want to give memory, you have to store all the previous chats between you and the AI. You have to store it somewhere. Okay, maybe in a DB, or maybe in some sort of a, you know, data storage. You have to store it.
[01:24:06] Anirban Paul: And when you are having that session of conversation, I have activated one session with my chatbot. In that session, you provide that as a context. So in the context, you are giving that memory.
[01:24:17] Anirban Paul: Along with that, you're giving the next question. That's how memory comes in. So memory, inherently is not there on LLM. LLM doesn't have the concept of memory. Underlying LLM doesn't have it.
[01:24:27] Anirban Paul: Okay?
[01:24:28] Anirban Paul: So… One more thing, these underlying LLMs also requires training very often. They get strained again and again.
[01:24:37] Anirban Paul: And, to keep up with the newer data.
[01:24:44] Anirban Paul: Akansha, it is for yesterday as well.
[01:24:57] Anirban Paul: Okay, so…
[01:24:58] Anirban Paul: they often need to retrain the model again and again, so that the memory… so that there is a… you know, so that more recency of data comes in. So now, if you go to ChatGPT and if you ask, you will see probably 2025 August till which it is trained on. So there is that cutoff, that issue is there.
[01:25:16] Anirban Paul: So that's why, if you ask about the current war, Middle Eastern war, and all these things, it will generate the answer. But it is generating the answer because of its capability to access the internet. So it has the capability to access the internet and come back with an answer.
[01:25:31] Anirban Paul: So, all those capability you provide
[01:25:34] Anirban Paul: to the LLM when you're creating the chatbot. It is not that what's your LLM has it. Yes, Aditya?
[01:25:41] Aditya Banda: Yeah, just one basic question, Aniphan. When you say the model is trained on knowledge, I understand it's trained on the language-specific semantics and all of that, so that you can talk like a human. But then, why do you think it should be trained on the knowledge itself, the data itself in the internet?
[01:25:59] Anirban Paul: Because… because or else, you know, sometime, if it is not trained on the data, sometime actual information can be missed out.
[01:26:08] Anirban Paul: Got it.
[01:26:09] Aditya Banda: That information is anyway, persisted in the vector space somewhere, right? So when we chat with the chatbot, it can convert our utterance into vector.
[01:26:19] Anirban Paul: vector space is with regards to your chatbot that you're making. But a LLM, which is solving a bigger purpose, what is your goal of using the LLM?
[01:26:28] Anirban Paul: That it can answer also from its own behalf, right, Aditya?
[01:26:32] Aditya Banda: Hmm.
[01:26:33] Anirban Paul: Okay, so if it doesn't know the recency of information, how will it answer from its own behalf?
[01:26:38] Anirban Paul: Okay, vector space is you creating a chatbot, that is you requiring it, and it is creating it. But I'm talking about the LLM that you're seeing on ChatGPT,
[01:26:46] Anirban Paul: Over there.
[01:26:47] Anirban Paul: They want to serve the entire world, and they are solving the vector space, they are using vector space, till the next training happens. So from 2025 August till June 2026, how are they answering from the questions that we are asking? That is about more recency. Who is the president? It will… if you say, don't access the internet, it will not be able to tell.
[01:27:06] Anirban Paul: Okay, okay, or that time the president of US was chosen. Okay, but let's say it got trained before that. So, it'll not be able to answer if it doesn't access the internet. So, till the next training is done.
[01:27:18] Anirban Paul: Vector store, it is used.
[01:27:21] Anirban Paul: But training a model is the most idealistic process, guys. Training a LLM is the most idealistic process. If you can do it, if you can achieve what you want to achieve, if it can be done, if you have the GPU access, everything, then training nothing is better than training.
[01:27:36] Anirban Paul: Okay.
[01:27:38] Anirban Paul: And… If you can't train it, then rag and vector store is a…
[01:27:44] Anirban Paul: It is more of a Jugada.
[01:27:46] Aditya Banda: Hmm.
[01:27:47] Anirban Paul: Got it, Aditya. All these things are Jugar. So why… why more updation of a RAG is coming now? A RAG is also becoming, vector-less RAG concept is coming now, in the market. Okay, it has not gained popularity yet. Okay, few of the people have started discussing around it. Okay.
[01:28:02] Anirban Paul: Vectorless rack. These updates are coming. Why? Because rack is also not very reliable.
[01:28:08] Anirban Paul: Got it.
[01:28:09] Anirban Paul: So, idealistic situation would have been to train a model.
[01:28:13] Anirban Paul: Till you can do your next training, you can mix up with your vector store concept.
[01:28:17] Aditya Banda: Okay, the delta, maybe you can mix with the vector store, is what you're saying.
[01:28:21] Anirban Paul: Yes.
[01:28:22] Aditya Banda: Yes. We do training, yes.
[01:28:24] Anirban Paul: So, for your own setup, for your own chatbot, fine, vector stored is good. But, if I'm creating Oh.
[01:28:32] Anirban Paul: See, I'm… my goal is to sell LLMs, okay? Then I cannot manage it with Vector Store, and just train the LLM ones. OpenAI just cannot forever run with Vector Store. Got it.
[01:28:44] Anirban Paul: If it is going to sell LLM, okay, OpenAI is selling the software license to Azure and to AWS to use for other enterprise clients to go and buy it from there.
[01:28:54] Aditya Banda: Yeah, so your persona is somebody who's building those LLMs, not the channel.
[01:28:57] Anirban Paul: Yes, yes, you are the creator of LLM. Vector store is the person who is consuming that and making it.
[01:29:03] Anirban Paul: You are also one of the consumer. You are means not you, like, I'm talking about the OpenAI has created ChatGPT. ChatGPT is one of the consumption of the LLM.
[01:29:10] Anirban Paul: ChatGPT is an agentic architecture that has vector stored, that has capability to search the internet, that has the capability to draw an image, or, like, call the diffusion model and get an image. Okay, all those, those are, like, agentic architectures.
[01:29:23] Aditya Banda: Right.
[01:29:24] Anirban Paul: But the main point that I'm coming back to, your LLM doesn't have those capabilities.
[01:29:30] Anirban Paul: Okay? If you go to ChatGPT, it says, I am ChatGPT. If you go to a LLM call and ask, what are you, it will say, I am a model trained by this, this, this, this, this. It will not say I'm ChatGPT. So, don't think your LLM and your ChatGPTs are safe.
[01:29:45] Anirban Paul: That is my point.
[01:29:47] Aditya Banda: Got it.
[01:29:50] Anirban Paul: Okay.
[01:29:51] Anirban Paul: So, now, we will go to the next part. Before that, we'll go for a break. Okay, so this is where the rag part starts in. Okay, so we'll come back after the break of around…
[01:30:02] Anirban Paul: 7 minutes, and we'll do the rest. Guys, are you all enjoying till now? Excited and enjoying?
[01:30:09] Anirban Paul: Yes, good day.
[01:30:11] Deepak Bobade: S.
[01:30:12] Anirban Paul: Okay, things will go up like this now, because, see, stable division, all these things, if you can take lightly, it's fine, because we hardly use. And while using, we learn, and we use it. Okay, but from here onwards is actual the things that you are going to build, because your next few
[01:30:26] Anirban Paul: Two months is all about this only. Agents, agents, agents. So this is the actual application you are going to learn. Okay.
[01:30:33] Anirban Paul: Okay, so today we'll do the intro to RAC, and we will start building one part of RAC. Next day, we'll continue on top of that, okay? So, see you in 7 minutes.
[01:38:17] Anirban Paul: Okay, so… guys…
[01:38:45] Anirban Paul: Okay, are you able to see the screen, Guy?
[01:38:50] Anirban Paul: automating.
[01:38:52] Sacheen Adavinavar: No, it says.
[01:38:56] Anirban Paul: Okay. Now?
[01:38:58] Sacheen Adavinavar: Nice.
[01:38:59] Anirban Paul: Okay.
[01:39:00] Anirban Paul: So, now my question is, can you build a chatbot that answers the question, what is the capital of France?
[01:39:07] Anirban Paul: Can you build it?
[01:39:09] Anirban Paul: Just answering these questions. Everyone?
[01:39:13] vinit shah: Yes.
[01:39:14] Aditya Banda: Yeah.
[01:39:15] Anirban Paul: Fine, like, this is a very general knowledge from which it can answer.
[01:39:21] Anirban Paul: Now…
[01:39:23] Anirban Paul: Can you answer… can you… can your chatbot also answer this question? I live in Paris. Should I carry umbrella if I'm going out today?
[01:39:32] Anirban Paul: Can your chat would answer that?
[01:39:34] Aditya Banda: Nope.
[01:39:36] Anirban Paul: Inherently, I'm talking about chatbot by chatbot, I mean is your LLM? Can your LLM answer that?
[01:39:41] Ravindra Singh: No, right?
[01:39:43] Anirban Paul: Because it doesn't have access to…
[01:39:45] Anirban Paul: Any of your recent information, or any extra reference data that can be required to answer this question.
[01:39:51] Anirban Paul: Now, let's take it from the enterprise-level point of view.
[01:39:55] Anirban Paul: Okay, so… Let's say… You are trying to build a… Chatbot. Okay.
[01:40:03] Anirban Paul: proprietary chatbot for your… for your, company. Okay, let's say it's an HR policy chatbot you are building. Now, for building that HR policy chatbot, let's say there are some leave policy you have, okay, about your company, that your chatbot… that your HR knows, and there are, like.
[01:40:22] Anirban Paul: Lot of PDFs, lot of, you know, PDFs carrying this information. Every time the HR has to give this answer, they will have to refer to this PDF and come up with an answer.
[01:40:34] Anirban Paul: Nope!
[01:40:35] Anirban Paul: What if you can build chatbot
[01:40:38] Anirban Paul: or a chatbot powered by LLM, which is on… which sits on top of this HR document.
[01:40:46] Anirban Paul: Okay, you have a HR document. On top of that, your LLM sits.
[01:40:51] Anirban Paul: And your LLM can actually answer based on that, right? What if you can build something like that?
[01:40:59] Anirban Paul: Right? So, this is the concept that we are going to talk about today. Okay. So, this is the problem that we are going to solve today. So, agents will help you achieve this. Okay. So.
[01:41:11] Anirban Paul: Now, let's first understand what are agents. We will resume agents again after completing RAG. So first, just to give you an intro, RAG is also a special type of agent only. Okay, so…
[01:41:24] Anirban Paul: Agents, Are something thathas these 3 important things that you should remember.
[01:41:31] Anirban Paul: One is known as planning.
[01:41:34] Anirban Paul: Memory and tools.
[01:41:36] Anirban Paul: I remember in a different way, memory, action, and planning.
[01:41:40] Anirban Paul: Action is tools only. So tools means what are the accessible information, like, how you can access the information. It could be a DB, it could be an internet, it could be a weather app, anything. The tools access that you have. So those are the action it can take.
[01:41:53] Anirban Paul: Memory is like a persistent kind of a memory, which the agent shares for the entire flow. It is like to remember what is the flow that is happening, because, see, your LLM doesn't have any memory.
[01:42:05] Anirban Paul: If you ask Anirban, if you again ask what is Anirban, okay, who is Anirban or, it will not be able to answer.
[01:42:12] Anirban Paul: So, you need to give a persistent kind of a memory, which is shared across your agentic architecture, so that it remembers what has happened till now. It's like a state. You're giving a state to your agents. Okay, a state to remember.
[01:42:26] Anirban Paul: Okay, and then come planning. Planning is basically orchestration, which agents to call, what to call now. From here, how will it go to there? Like, this is what is
[01:42:37] Anirban Paul: what we are going to build today. What we are going to build in Agent Tech Architecture. Okay, so this is the core concept of agents, that there are three major things. Memory planning tools, or memory action and planning. Map, or MPT.
[01:42:50] Anirban Paul: However want… however… whichever way you want to remember. Okay, but this is the core concept of our agents. Agents have nothing else. Okay, now, many people, you know, consider the input as one component, like input to the agent, which is known as perceiver… perceiving an agent.
[01:43:05] Anirban Paul: Okay, how you perceive the information. So that is also there, input to the agent, output from the agent.
[01:43:10] Anirban Paul: Those are also considered as core comp… The components of agents, but those are not the core. The codes are these three. That is what the differentiating factor. See, input is there in LLM also, output is there in LLM also. But in agents, you have these core things, which makes it different, and which makes it more powerful than just a plain LLM call.
[01:43:27] Anirban Paul: Okay. So, obviously, an agent can understand a task, it can… a task described in natural language because of LLM. It can decide what actions to take, which tools or API to call. It can execute everything.
[01:43:40] Anirban Paul: Together. Okay, so this is the concept of agents. Now, we will take
[01:43:47] Anirban Paul: aside, and we will come towards RAG now.
[01:43:50] Anirban Paul: So…
[01:43:51] Anirban Paul: Few more things, when to use an agent, when you have a complex task, when you have lots of tool integration, when you have a lot of dynamic, you know.
[01:43:59] Anirban Paul: not a straightforward thinking, not a very linear way, but you have some dynamic thinking to do. Okay, dynamic thinking, you know, quality analysis you need to do before you go to the next steps. Those are the situations when you use agents.
[01:44:12] Anirban Paul: If you have to build a customer bot agent for your company, how will the agent know about your customer's basic information to handle the queries? This is the question I was raising.
[01:44:21] Anirban Paul: Now, we will slowly come towards RAC.
[01:44:24] Anirban Paul: So, what is RAG? You can call RAG as a component A special component
[01:44:32] Anirban Paul: In an agent, or you can call the rag is also agent in an agentic architecture. So, agentic architecture is multiple agents.
[01:44:40] Anirban Paul: one, every agent doing certain tasks. RAG is doing one task of… out of that.
[01:44:45] Anirban Paul: So, RAG is a special type of agent, which is a special type of architecture, I would say, that
[01:44:50] Anirban Paul: One of the agent of your agentic architecture uses in order to achieve something.
[01:44:54] Anirban Paul: So first, we will build the rag without an agentic architecture. We'll just build a plain RAG. We will see multiple versions of RAG as well. Then, slowly, slowly, we will come to agentic Rag as well. AgentIt RAG is there almost, like, after one month of learning agents. Then once you, I think.
[01:45:11] Anirban Paul: Your agent degrad, is there in 16th of August.
[01:45:16] Anirban Paul: As a mini-project, okay.
[01:45:19] Anirban Paul: So…
[01:45:20] Anirban Paul: Okay, this is that… that is there as a mini-project, okay. So anyways, so now we will learn… we will take a site, and we will learn React the most, you know, basic to advanced way, and
[01:45:32] Anirban Paul: And then slowly, slowly, you will start integrating towards Agent, okay?
[01:45:37] Anirban Paul: Till now, any question, guys, with regards to agents or anything? Because this is a very high-level overview of agents. Agents, again, we will land on to agents later. Now, we are starting with RAG.
[01:45:46] Anirban Paul: Any question regarding agents?
[01:45:52] vinit shah: Nope.
[01:45:56] Anirban Paul: Okay, so this, whatever I have told you, is just to give you the intro so that RAG can be introduced. Now, the main thing about RAG is RAG is like, now I'll share another PPT with you. So this is another PPT that is to be…
[01:46:11] Anirban Paul: Now, this I will share later, because there's a lot of other things as well. So, as of now, I will share you a rack PPT.
[01:46:34] Anirban Paul: So this is the RAG PBT that I'm sharing.
[01:46:38] Anirban Paul: This is the core.
[01:46:40] Anirban Paul: rack people. You just need this for this week's LMS.
[01:46:44] Anirban Paul: The other one I'll share later.
[01:46:50] Anirban Paul: Okay.
[01:46:51] Anirban Paul: Okay. So now, guys, with your concepts, with your understanding about RAG, what all do you understand by RAG?
[01:46:59] Anirban Paul: What is the thing that you know about RAG?
[01:47:10] vinit shah: At a high level, we have access to certain artifacts about that particular,
[01:47:15] vinit shah: item, right? Like, say, for example, documents or live updating details and stuff like that, which gives me a bit more up-to-date data.
[01:47:26] Anirban Paul: Hmm. Have you seen an architecture of RAG, Vinit, ever?
[01:47:30] vinit shah: At a very high level.
[01:47:33] Muni Prakash Ganji: Yeah, RAG is specifically which the LLM is not aware of it, so it will go like a query.
[01:47:40] Muni Prakash Ganji: to the RAC, that means the personal data of companies, suppose the HR have a different policy, which the LLM is not trying.
[01:47:48] Muni Prakash Ganji: So, it will take those details, and again, this both will go into the LLM as a query, and it will be tuned with the help of LLM again, and the output will come into a required format, like how you want it.
[01:48:03] Anirban Paul: Okay.
[01:48:04] Anirban Paul: Fair. Anything else? Aditya, you raised your hand.
[01:48:07] Aditya Banda: Yeah, yeah, I like to think about it this way. Llm is a library, general library.
[01:48:13] Aditya Banda: But, I want information from my personal notes or journal, so I will ask LLM to only refer to my books.
[01:48:22] Aditya Banda: Whenever I ask any question.
[01:48:24] Anirban Paul: Hmm.
[01:48:25] Anirban Paul: Okay.
[01:48:28] Anirban Paul: Anything extra fancy you want to add, or make it…
[01:48:33] HITESH SURYAWANSHI: Good morning.
[01:48:34] Sonam Manwal: On internet, I had read one example.
[01:48:38] Sonam Manwal: What they were saying, like, you can think of, like, you are hiring a new software engineer in your company, so you can think of him as a LLM, but he must be knowing all the technologies, but he's not aware how your company specific, processes specific to your company are working.
[01:48:53] Sonam Manwal: So, by doing the trainings and all, we make him know our company's specific policies know. Same way, we apply the RAG on top of LLM so that it can give answer to specific to our, requirements, or whatever the data knowledge we have shared with it.
[01:49:13] Anirban Paul: Got it.
[01:49:15] Anirban Paul: Okay.
[01:49:16] HITESH SURYAWANSHI: So, basically, in the RAG, I mean, we can store our documents in the form of
[01:49:23] HITESH SURYAWANSHI: vectors into the vector DB,
[01:49:26] HITESH SURYAWANSHI: And, while querying on that, particular database, so we, encode our query, into the vector form and do the, cosine similarities with it, and get the.
[01:49:40] Anirban Paul: similarity or Euclidean, let's say. Euclidean, yeah.
[01:49:43] HITESH SURYAWANSHI: And then, we get the context, like, we give that context to the LLM, and the LLM gives the result based on that.
[01:49:53] Anirban Paul: Got it.
[01:49:54] Anirban Paul: Okay. So, let me tell you in this way, guys.
[01:49:58] Anirban Paul: RAG is like a… one of the examples I've given over your years. Here is… it's like a… it's like an open book exam. Have you ever sat for an open book exam, guys?
[01:50:11] Deepak Bobade: Yes.
[01:50:13] Anirban Paul: Okay, so in an open book exam, do you prepare one night before, yes or no?
[01:50:19] Deepak Bobade: Yes. But we also carry the book to the exam.
[01:50:23] Anirban Paul: Yes, you do carry the book with your exam, but still you prepare it. So, in RAC, also has the similar concept, same concept. So, in RAC, so today we'll just do the introduction to RAC. So in RAC, what happens is, guys, you have
[01:50:40] Anirban Paul: In an open book exam, one night before, you do…
[01:50:43] Anirban Paul: What is? You go through the book. You go through the book.
[01:50:47] Anirban Paul: You do a… Like, like, a very rough reading through the books.
[01:50:55] Anirban Paul: You go through it, so that… Next morning, during the exam, You can save time.
[01:51:02] Anirban Paul: During the exam, you can save time.
[01:51:05] Anirban Paul: You can save time.
[01:51:08] Anirban Paul: Okay. So, similar concept.
[01:51:10] Anirban Paul: happens over here. Okay, in LA… in a RAG concept, first is…
[01:51:18] Anirban Paul: I… your bottom line LLM doesn't have the concept off…
[01:51:26] Anirban Paul: the questions that you are asking, okay? So there could be questions that you are asking about your company's policy, about your HR policy, and all these things, and there could be scenario that you are not able to get a good answer, okay? From your LLM, obviously, you'll not get… your LLM doesn't know about your company, it could be a startup that has come up just now.
[01:51:44] Anirban Paul: Okay, it doesn't know, and your HR policy might not be public also.
[01:51:48] Anirban Paul: Okay, so how will we solve this? So, there comes the concept of retrieval. Retrieval.
[01:51:56] Anirban Paul: augmented
[01:52:01] Anirban Paul: Generation.
[01:52:04] Anirban Paul: So what happens is, every word of here has a meaning, okay?
[01:52:09] Anirban Paul: What happens is, let's say I have a PDF.
[01:52:14] Anirban Paul: I have a PDF. I have multiple PDFs like this. These are my PDFs.
[01:52:20] Anirban Paul: Okay. So, these PDFs have a lot of… paragraphs like this.
[01:52:29] Anirban Paul: Now… what I will do is, first.
[01:52:33] Anirban Paul: I will convert them into a representation which machine understands. So, I will use an encoder.
[01:52:41] Anirban Paul: encoder model.
[01:52:44] Anirban Paul: Let's say, but… To convert these things into
[01:52:49] Anirban Paul: I will pass them through encoder model.
[01:52:52] Anirban Paul: Okay? And convert them into something known as embeddings. I will come to chunking and all these things later, okay? First, let's look at the most vanilla explanation of RAG. Okay, I will convert them into embeddings.
[01:53:06] Anirban Paul: And… I will keep these embeddings with me.
[01:53:10] Anirban Paul: So, embeddings of this page… Page 1.
[01:53:14] Anirban Paul: Page 2… Page 3, like that, I will convert for every page.
[01:53:19] Anirban Paul: And I will keep this embeddings, embeddings 1, embeddings 2, embeddings 3. Now, if somebody asks me a question.
[01:53:26] Anirban Paul: First, I will go ahead and convert this also using an encoder model, this also I'll convert into an embedding.
[01:53:33] Anirban Paul: Because, like I told you, when you can compare things, when they're…
[01:53:37] Anirban Paul: similarity, when they speak the same language. So I'll take this query, convert this into embeddings.
[01:53:42] Anirban Paul: These embeddings and these embeddings?
[01:53:45] Anirban Paul: will be searched. So, embedding of page 1, embeddings of page 2, embeddings of page 3 will be searched.
[01:53:51] Anirban Paul: So once I have searched this.
[01:53:53] Anirban Paul: Let's say my question is there in page 2.
[01:53:56] Anirban Paul: Now.
[01:53:59] Anirban Paul: I don't need page 1, I don't need page 2. I can take just, sorry, I don't need page 3. I can just take page 2,
[01:54:06] Anirban Paul: Give the query, along with page 2, to my LLM.
[01:54:13] Anirban Paul: And I can write a prompt also, let's say, based on the query of the user, please answer a question from page 2, and I can give that to the LLM.
[01:54:21] Anirban Paul: And the LRM will start generating answer.
[01:54:24] Anirban Paul: For this question from here.
[01:54:26] Anirban Paul: So there is a hyper-relevancy that you have got just added to your concept, to your, context, by giving your context, without even training your model.
[01:54:37] Anirban Paul: This is the overall high-level concept of our RAT. Now, we'll discuss more things. Yes.
[01:54:42] Anirban Paul: Tell me.
[01:54:44] Anirban Paul: Yogan, right?
[01:54:48] Anirban Paul: How's your day pronounced?
[01:54:50] Anirban Paul: Hello.
[01:54:54] Anirban Paul: You're not audible.
[01:55:00] Anirban Paul: You're not audible.
[01:55:03] yogan naik: Can you hear me now?
[01:55:04] Anirban Paul: Yeah, yeah, I know.
[01:55:05] yogan naik: Okay. So when it does this…
[01:55:08] yogan naik: Now, if the same question has been asked tomorrow, will it do the same exercise again, or will it remember or under, you know, learn it?
[01:55:16] Anirban Paul: No, no, no, it will not learn it, okay? It will not learn it, so that is how you train, like, use your model, but ideally, like, at any stage, it will not learn it. That is, like, when you are using LLM, when you get the LLM subscription, no, there is an option that LLM asks you, that whether you want to use your data for training or not.
[01:55:35] Anirban Paul: That is not with respect to this PDF only, even if you are saying, I am Anirban, and I…
[01:55:42] Anirban Paul: I'm a very good cricketer, played first-class cricket in, let's say, in my state, and if you continuously give this, and if you sell the LRM to learn it, if you give the permission to learn it, then your data will be used for learning.
[01:55:55] Anirban Paul: Got it. So, that data could be that PDF data also, that the query that you give also. Okay. But that learning happens when the next training happens.
[01:56:05] yogan naik: Okay.
[01:56:06] Anirban Paul: So, that is a different thing.
[01:56:08] Anirban Paul: Okay, but usually on enterprise-level LLMs, usually they don't give the permission to use your data to train their model.
[01:56:16] yogan naik: Okay, but there is an option. They can, enable it to learn it so that it can respond.
[01:56:20] Anirban Paul: Yeah, yeah, but that training will happen probably…
[01:56:25] Anirban Paul: let's say a year after… you see ChatGPT hasn't been trained for the last one year.
[01:56:30] Anirban Paul: Got it.
[01:56:31] Anirban Paul: So… so by that time, your policy will again change again.
[01:56:35] Anirban Paul: Okay.
[01:56:37] Anirban Paul: So… so yeah, so instantly it will not learn.
[01:56:41] yogan naik: Okay.
[01:56:42] Anirban Paul: Yeah. Here's G2.
[01:56:48] Jithu Tagore: Am I audible?
[01:56:49] Anirban Paul: Yes, yes, you do, yes.
[01:56:50] Jithu Tagore: If I am using RAG, then the test generation, doing it multiple times, will the generated response will be same all the time? Like, I need to have the same response all the time, that is my out… outcome. Is it possible?
[01:57:08] Anirban Paul: No, that has nothing to do with RAG. That is… that is your normal LLM call only. That depends on your temperature. You're ideally…
[01:57:16] Anirban Paul: it will not be very exactly similar, but very similar. So that, like, how you choose your temperature value, how you choose your,
[01:57:25] Anirban Paul: let's say, top key, top K, but that has nothing to do with RAG itself.
[01:57:29] Jithu Tagore: Yeah, for example, I will give an example. I have developed a, like, code for CV tailoring. So, with the job description, I was tailoring the summary of my receipt with the job description.
[01:57:42] Jithu Tagore: For that, if I use the job description plus the, resume, to generate the summary again, if I, do the process multiple times, will it hallucinate, or…
[01:57:56] Jithu Tagore: making, to make the… to generate the same response again, more accurately, what should I do?
[01:58:03] Anirban Paul: If you… if you want to generate the more response accurately, then your temperature value should be as low as possible. You may have to make the model very deterministic in nature.
[01:58:13] Anirban Paul: a temperature value, if it is low, then only very confident output only will come out of it. And you have to put a top P parameters, you have to put top K parameters, then…
[01:58:24] Anirban Paul: You'll also have to use some few short examples. Okay, that is your prompt engineering concepts also will come in. But this.
[01:58:30] Jithu Tagore: Oh, besides…
[01:58:31] Anirban Paul: It has nothing to do with RAG, it's what I'm saying. RAG is just a way to add extra knowledge. You're augmenting. What is augmentation? Adding.
[01:58:40] Anirban Paul: Augmentation is artificially adding or adding or increasing what you have, right?
[01:58:45] Jithu Tagore: Okay. Yeah.
[01:58:46] Anirban Paul: So, you're augmenting, like, you learned about data augmentation on CV, right? Computer vision. You augmented the data. Let's say if you have one picture of cat, you flip it, you change the color, you change the filter, and you create more images. So, same concept, DRAC does.
[01:59:02] Anirban Paul: that instead of answering from its own knowledge, what if I can provide you the most relevant context
[01:59:09] Anirban Paul: And you answered from there.
[01:59:12] Jithu Tagore: Okay.
[01:59:12] Anirban Paul: Okay, that is the concept of RAG. It has nothing to do with you generating repeatedly same answer. That is… that is, again, coming back to your temperature and hyperparameters of your LLMs.
[01:59:22] Jithu Tagore: Okay, okay, I got it. Yeah, thank you.
[01:59:25] Anirban Paul: So, if we give the text in the document through LLM through prompt, instead of using RAG.
[01:59:34] Anirban Paul: Will it work differently? No, ideally, it will work the same way. Okay, it will work same way, only if you can manually give it that. But the point is, you cannot scale that.
[01:59:45] Anirban Paul: Okay, let's say if you want to make it for 10,000 PDFs, or 1,000 PDFs, or 100 PDFs only, you cannot manually do this.
[01:59:53] Anirban Paul: The result is same. Query, and instead of giving the page 2 through RAG, you are giving… copying and pasting the page 2 content.
[02:00:03] Anirban Paul: Yes, Ravindra.
[02:00:05] Ravindra Singh: So I have a question, like, here we have a multiple page, so… so let's say we have a 100 page.
[02:00:12] Ravindra Singh: And whenever I will do any query, so, in the vector database, so, like, do we have some kind of, you know, like, indexing from where it should search, or, like, it will scan entire, database or course, basically?
[02:00:27] Anirban Paul: No, no, no. It is… this method is known as indexing only, so… In…
[02:00:33] Anirban Paul: Normally, when you search, your embeddings will be searched against all the PDFs.
[02:00:37] Anirban Paul: Okay. All the pages. Okay. Based on the distance, it will give you the lowest one, and it will give you… it is the most vanilla rag. Now, there are more other rags, like, you know, people do something known as ANN, Artificial, Approximate Nearest Neighbor Search. So, where, instead of doing
[02:00:55] Anirban Paul: the… search from all the pages. Okay, let's say you have 6 pages.
[02:01:02] Anirban Paul: Okay. These two are talking about one topic.
[02:01:06] Anirban Paul: this…
[02:01:07] Anirban Paul: this, and this is talking about one topic, and this is talking… talking about another topic. So what? First, you search the distance from your query with only
[02:01:17] Anirban Paul: with only… The centroid of the… these… these chunks, these pages.
[02:01:24] Anirban Paul: Okay, so these two, instead of searching both the pages, let's say you had 20 pages. In this topic, you had 20 pages.
[02:01:31] Anirban Paul: So, you will have the centroid of this… embeddings of these 20 pages? Centroid.
[02:01:36] Ravindra Singh: Okay.
[02:01:37] Anirban Paul: You search a distance from the send right to your query.
[02:01:40] Anirban Paul: Okay, from there.
[02:01:43] Anirban Paul: Like that, you do for all the centroids, okay, of all the trees. Whichever is the nearest, lowest.
[02:01:50] Anirban Paul: That topic, you go deeper.
[02:01:53] Anirban Paul: Let's say you found one cluster, which is page 1 and page 2 together, that is the lowest. So then you further go into page 1, check their distance from this topic.
[02:02:03] Anirban Paul: go to page 2, check the distance from this topic. Instead of doing everything, you do approximate nearest neighbor.
[02:02:08] Anirban Paul: So those are also there.
[02:02:10] Anirban Paul: Those are different forms of rags.
[02:02:13] Ravindra Singh: Oh, okay.
[02:02:14] Anirban Paul: Okay. Not the vanilla rag. Yes, the Rakesh.
[02:02:17] Dwarakesh T P: Yeah, so, for example, if, if the document that I'm providing has a different kind of information, different from whatever I'm asking, so will it still give the closest match, which may not be relevant to the question?
[02:02:32] Anirban Paul: depends on your… that depends on your prompt. That depends on few of the things that we will learn next week about hyperparameters of rags. Okay, hyperparameters of rags, which are basically, your threshold. You will use a threshold technique. Like, how many… how much distance will you measure before you come to the answer?
[02:02:49] Anirban Paul: Okay, so you have a thresholding technique. You have K, you have the K… how many K you want to… like, when you are searching how many PDFs it should show. Will it show all the PDFs?
[02:03:02] Anirban Paul: Like, and their distance, or will it show only top 2 pages only?
[02:03:08] Anirban Paul: Got it, Dura Kesh. So, let's say you have 100 pages. When you're searching, all the 100 pages, there'll be some distance, right? Maybe negative distance, but there will be some distance.
[02:03:19] Anirban Paul: Okay, but do you want to consider all of them?
[02:03:21] Anirban Paul: How much you want to consider? Those will depend on your K, and if you want to consider this much also, then also how much minimum threshold distance you are looking at. That depends on threshold.
[02:03:32] Anirban Paul: Okay, so using all these things, you will build the rack, and that time you'll be able to filter this out.
[02:03:38] Anirban Paul: this information out. So anything which is not relevant, you can write in your prompt that I do not want… if you do not know… if the answer is not from the context, please don't answer. If the answer is not from the context, you can answer from your knowledge. You can write prompts like that.
[02:03:53] Dwarakesh T P: Got it, okay.
[02:03:54] Anirban Paul: That is how you design your rack. See, RAC is a very open-ended architecture. What you do, you have n number of combinations. We will use the most
[02:04:02] Anirban Paul: enterprise-level combination. Now, there are n number of way racks can be built. There is vector-less rackalso that is coming to the market. It is not yet popular. It is not, yes, cracked enterprise level yet. Few… few people are doing, few researchers are exploding. But that is also there.
[02:04:20] Anirban Paul: Okay, you have lots of LLM access, then you can use vector-less rack, but it uses a lot of LLM, by the way.
[02:04:26] Anirban Paul: So, RAG reduce the need of LLM.
[02:04:30] Anirban Paul: Vectorless RAG is increasing the need of LLM.
[02:04:32] Anirban Paul: Okay, again, but apparently it is…
[02:04:36] Anirban Paul: supposed to do a better job than a traditional RAC, but it has not gained the traction yet so much to comment it on about. It might fade away also. So this kind of
[02:04:46] Anirban Paul: additional concepts of RAG you will see many times, because this is an architecture, guys. There are probably 10, 15 versions of BERT, do you know that? Because BERT is, ultimately, it's a transformer architecture model.
[02:04:57] Anirban Paul: Okay, so from there, people created their own BERT. Legal BERT is there, Digital Robota is there, sorry, Digital BERT is there, Robota is there. So, so many different BERT models are also there. Okay.
[02:05:10] Anirban Paul: So, do we do all of them? No. We use the main one, which we find it from sentence transformer.
[02:05:16] Anirban Paul: So that is a thing. So there will be multiple versions of RAG. You can create lots of
[02:05:21] Anirban Paul: you know, your own strategy inside BERT, inside RAG, and make fancy, fancy, you know, setup, for RAG. Okay, this is how my rag works. This is how this rag works. So, that is your skill.
[02:05:35] Anirban Paul: Okay.
[02:05:36] Anirban Paul: But anyways, this is a very, very high-level overview about RAC. Pawan, do you understand? You, since you asked the first time, please introduce RAC. This is a very high introduction about RAC. We'll get into the depth of RAC by next day, where we will build
[02:05:50] Anirban Paul: the first data ingestion pipeline. This was the data ingestion pipeline. First, we'll build the data ingestion pipeline, then we will learn about Langchain, and then we'll build an entire Langchenic rag. Then we will build a non-langchenic rag also, because with Langchain, you don't get a lot of options.
[02:06:04] Anirban Paul: So we'll do a lot of 2-3 versions of rags.
[02:06:10] Anirban Paul: Yeah, Pawan, like I was asking, did you understand?
[02:06:14] Pawan Misra: Yes, yes, it was a good explanation today.
[02:06:18] Anirban Paul: Yeah. So till now, guys, we haven't even scratched the surface. We have just learned, you can say, Probably…
[02:06:25] Anirban Paul: 10% of RAG. We have done 10% of RAG here. Okay.
[02:06:30] Anirban Paul: Now, we'll go into the deeper. Slowly, slowly, we will break, we'll understand the concept of chunking, like, do we give the page, entire page, or we chunkify it into multiple chunks?
[02:06:41] Anirban Paul: then we will also see how LangChain is used, how this threshold K is used, how… what are the different libraries that is there.
[02:06:49] Anirban Paul: For creating the VectorDB. There is FAS, there's Chroma. Okay, I will do it with FAS. That is the most fastest, and the most earliest thing that has come to the market.
[02:07:00] Anirban Paul: You will try out with Chroma as well, from yourself, okay? Same thing, just change fast to Chroma, it works.
[02:07:06] Anirban Paul: Okay. Anyways, so that is the concept. Next day, we will get into the depth of this, and I think probably
[02:07:14] Anirban Paul: One… two classes will be required.
[02:07:18] Anirban Paul: So next day, both the days will be about RAG, and then we will go into other agentic things.
[02:07:24] Jithu Tagore: Yeah, hello? Yes, GP.
[02:07:27] Jithu Tagore: higher word, another one term also, it's, called hybrid, like VectorDB plus some other DB included, like, for that.
[02:07:35] Anirban Paul: Yeah, yeah, yeah. So, those are different versions. So, hybrid is very open-ended. You can create your own hybrid version. Okay, so hybrid is like, you're not happy with VectorDB, you mix it VectorDB with some other kind of search. Let's say exact text match.
[02:07:48] Anirban Paul: What a G2? I will look for exact text, then only I will get the answer, not using VectorDB.
[02:07:55] Anirban Paul: So, people use that also. So…
[02:07:57] Jithu Tagore: Yeah, for example, for, I've worked on e-commerce, DB, just a small project. For that, to some queries, like, how many, what are the, how many orders are there? Like, the count-based answer, the vector DB will not fetch it properly, right?
[02:08:15] Jithu Tagore: VectorDB solution architecture is not good for, like, in my understanding, I don't know whether it's right or not.
[02:08:22] Anirban Paul: No, no, no, you're correct, you're correct. See, that's what I'm saying, like, every… see, these are not… these are, like, very specific, specific with respect to your project. In your project, this didn't work out. Okay, in my project, there was a different reason for using hybrid. Okay, so those are different, different versions.
[02:08:38] Anirban Paul: What did you do?
[02:08:40] Jithu Tagore: Yeah, okay, okay.
[02:08:41] Anirban Paul: Yeah. But anyways…
[02:08:43] Jithu Tagore: I'm not hybrid, I, like, I am expecting to learn from this course.
[02:08:47] Anirban Paul: Yeah, we will talk about one hybrid, RAG as well, but that will come way later. First, we'll have to process… proceed ahead with RAG, and plain rags, and few different versions of RAG, and then go to agents, and then we will have agentic rags, and all those things we'll learn about hybrid, I think.
[02:09:04] Jithu Tagore: Yeah, okay, thank you.
[02:09:07] Anirban Paul: Okay, thank you guys.
[02:09:10] Anirban Paul: Bye, everyone.
[02:09:12] Anirban Paul: See you next week.
[02:09:14] vinit shah: Thank you. Do you want to be the same time management, 5 to 7?
[02:09:17] Anirban Paul: 5 to 7.
[02:09:19] vinit shah: Saturday, Sunday, both.
[02:09:20] Anirban Paul: Yeah.
[02:09:21] vinit shah: Okay.
[02:09:22] Anirban Paul: These guys do give the feedback, guys, that's where… that's what I depend on, and we also depend on. So, yeah.
[02:09:30] vinit shah: Okay.
[02:09:31] Anirban Paul: Thank you.
[02:09:31] yogan naik: Thank you.
[02:09:32] Nirav Mehta: Thank you.