# 02 2026-06-28 Exploring LLM and AI Agents

course: Module 5 — AI Agents & Agentic Frameworks
module: Module-5-AI-Agents-Agentic-Frameworks
date: 2026-06-28
type: transcript
video_url: https://personal-learn.armco.dev/files/_Recordings/Module-5-AI-Agents-Agentic-Frameworks/02_2026-06-28_Exploring_LLM_and_AI_Agents.mp4

---
[00:19:56] Okay, so…
[00:19:58] Yeah. Good evening to everyone. Uh, welcome back, welcome back to…
[00:20:05] the session on Agent EKI nowadays. Uh, from here onwards. So…
[00:20:09] We already gave a very brief
[00:20:12] Intro, port, uh…
[00:20:15] the concept of agent, a very, very high-level overview, just to enter into RAG, because RAG is also a part of the agent.
[00:20:22] Uh, that we do. And, uh, and then we…
[00:20:27] We just gave you a very analogical explanation about
[00:20:31] Uh, what is RAC, and how it works. So, in the last one week, did you all, uh…
[00:20:39] like, acquire any knowledge out of it, or…
[00:20:42] Like, anybody who wants to add anything extra from last one week.
[00:20:47] Also, recap what RAG is, and…
[00:20:49] what you feel is that…
[00:20:52] These are the things that you have seen anybody out of this
[00:20:55] what, 35 people.
[00:20:57] Uh, yeah, let me put that.
[00:20:59] So, we are using REC very frequently from last two years, so…
[00:21:02] If I look at this in this perspective,
[00:21:04] Uh, this… all these models which we have, like LLM models, right? They're typically trained on certain duration. Cutoff is there, right? And this is trained on general data which is available across the globe, which is public data, I would say.
[00:21:17] Each and every enterprise need their…
[00:21:19] propriety data, right? The confidential data is there, right? And we need that. Example, policies…
[00:21:26] financial data, ethical, or whatever, like, legal documents and everything, right?
[00:21:31] Hmm.
[00:21:32] So, if we have to leverage LLM on top of that document,
[00:21:35] Right? This whole concept, in a short term, you can say RAG. RAG means that
[00:21:39] we are using LLM model.
[00:21:42] But the source of data is not going to be populated, it is going to be, like, our provided data. It can be formed up.
[00:21:48] structured, unstructured, SQL, non-SQL,
[00:21:51] documentation, anything can be done.
[00:21:53] Right? So that is more about RAC, right? And it is used
[00:21:58] In the enterprise level, I would say, the corporate levels, right?
[00:22:01] been for public users, it might not be available.
[00:22:03] Right? In fact, in the, like, a normal GPT models, what we're at, right? If you upload into that some document,
[00:22:09] Right? And tell them to summarize. It's more like a RAG concept only, right? We are, like, example,
[00:22:15] Uh, we are putting in a document, we want to summarize the document, any kind of example.
[00:22:19] Right? So, it's kind of a RAG model, where
[00:22:21] The source of information will become our provided data.
[00:22:25] So that is a rack, I would say that. That is a…
[00:22:29] Huh. Huh.
[00:22:30] high-level, right? Now we can go in deep.
[00:22:32] Uh, okay, so Gunyan, you said that you have been using it for 2 years, you have been using it as in your enterprise, or something like you have been using?
[00:22:39] interpret. Enterprise.
[00:22:40] We update it, we have built it.
[00:22:41] Yes?
[00:22:43] Oh, okay, that's great. So you have built it using Langchain?
[00:22:46] Langchin, line graphs,
[00:22:48] And all, right?
[00:22:49] Huh. Minecraft is agentic. Agentic rack.
[00:22:50] Yeah, if you don't think about that, we are using it for multiple orchestration by Google, like, instead of sequential language, like, sequential, right? But yeah, we link chain and all.
[00:23:00] Hmm, hmm. Okay, that's great.
[00:23:03] Okay, uh… yeah, anybody else? Any… any difference?
[00:23:08] to that, or, uh…
[00:23:11] Any con… any conceptual clarity you want to get out of that point, which Gunjan made?
[00:23:28] Okay, anyways, so, if nothing is there, we'll get into the deepest perspective of RAG, and we'll do the implementation as well.
[00:23:36] Today. Uh, uh, by the way, I had given you all one experimentation file for diffusion models, and I told you that from here,
[00:23:44] You read these experimentation, experiment 1, 2, 3, 4, like that, and you all try it out. Did you all try it out?
[00:23:55] Not all, I would say one or two experiments I have done it, but not for all. Yeah, just…
[00:23:59] Okay. But you understood the purpose, right? That you have to…
[00:24:02] Uh, kind of, I would say, but you're saying that it is, like, some are, like, not directly getting used, so yeah, but…
[00:24:08] We have to do more digive, like some do up and downs, right? Some changes and all.
[00:24:13] Yeah, some changes, please.
[00:24:14] And it will be more well. Then we're going to get more values around that, how it can get impacted, what is going to there, but for that, we need some time
[00:24:21] Hmm.
[00:24:22] to spend on that, and only able to learn that, yeah.
[00:24:26] Okay, so…
[00:24:28] Now, uh, we will, uh, you know, do a dissection of
[00:24:33] Iraq today. And, uh,
[00:24:36] And we will find some…
[00:24:39] you know, clarity with…
[00:24:42] Storm and sharing my screen.
[00:24:47] Yeah, can you all see the screen?
[00:24:54] Yeah
[00:24:55] Yeah. Okay, so last day, I told you that RAG is like a open book exam.
[00:25:03] Where you are bypassing the need of
[00:25:08] LLM answering from its general knowledge, because you have a proprietary data from where you want the answer out of.
[00:25:16] So that is why you are building something known as a retrieval augmented generation, which is like an open book exam.
[00:25:22] Where you are carrying a book, and based on the question, you are querying the book first, and then you are… you know
[00:25:29] you have a fair idea.
[00:25:31] Where the actual answer is, and from there,
[00:25:34] Uh, you get that relevant part of that answer, and then along with that query, you try to
[00:25:39] answer. So, this is the same thing happens in…
[00:25:42] in RAG as well. So, uh… so what was the last part you heard about RAG was chunking, right?
[00:25:49] So, that is the last part we discussed about RAC. I just recap a few things, uh, before we come into the, uh, you know,
[00:25:56] One by one part of executing it. So today will be a mixture of RAG,
[00:26:01] And a little bit about Langchain as well, and also understanding
[00:26:05] Uh, how vector embeddings and all these things becomes very vital in this case.
[00:26:10] Okay, so, uh, the first thing, uh, that, uh, we do…
[00:26:15] is… just a moment…
[00:26:18] Yeah, so the first thing that we all
[00:26:34] Yeah. So, uh, what, uh, what actually is, uh, RAG, I told you, in RAG, there are a few components to RAG.
[00:26:42] Okay, so in RAG, there are few components to it. So, first of all, you have something known
[00:26:48] as data ingestion.
[00:26:52] ingestion.
[00:26:55] So, you have something known as data ingestion. So now, since this is zoomed in, that's why you are seeing a little blurry.
[00:27:00] Okay, as at the end, you will see it zoomed out, you will see it very clearly.
[00:27:05] Okay, so data ingestion.
[00:27:07] Uh, uh, first of all, you have is data ingestion, that is,
[00:27:11] If I relate it with the open book exam, this is like a one night before activity that you are doing, data ingestion.
[00:27:17] So, data ingestion involves few of the very critical, critical steps.
[00:27:23] Okay, so first of all, uh,
[00:27:26] you know, you need… the data can be in various forms of documents. It can be a document, it can be…
[00:27:32] HTML document, it can be a web page itself,
[00:27:35] It can be a PDF document, it can be a SQL connection, a SQL connection to your DB.
[00:27:42] As well, many things can happen.
[00:27:44] Uh, so, as per… as per, you know, your data can be in various formats, okay? The most other thing is
[00:27:50] Uh, most of the RAG applications is on a structured data, unstructured data like, you know, PDF documents and web pages and text document, and docs document.
[00:27:59] So, in data ingestion, first of all, you have
[00:28:02] Document Loading.
[00:28:05] Okay, that is the first step.
[00:28:07] that you load the document.
[00:28:09] Okay, so first step is document loading.
[00:28:13] That is, you load the document. So, there could be many… various types of document, I told you. It could be PDFs,
[00:28:20] Uh, it could be even web page,
[00:28:27] It could be even text document.
[00:28:31] It could be docs as well.
[00:28:33] And various types like this, there could be n number of possibilities. Okay. Then, after you have loaded the document,
[00:28:41] your document is usually divided into pages. Okay, so you, first,
[00:28:46] chunk the page, so you have chunking.
[00:28:51] Okay, chunking of the document. So, this is the first data ingestion part. After
[00:28:56] your chunking is done. Chunking is breaking down the pages into few paragraphs or few set of paragraphs. That is what you do.
[00:29:03] And then, after chunking, you also have
[00:29:07] embedding creation.
[00:29:12] embedding creation.
[00:29:15] Okay, so you take… you ingest the data, load the document, you're chunking, you do chunking, and then you do embedding creation.
[00:29:23] So, these are some of the things that is there in your rag.
[00:29:27] Okay, uh, in your data ingestion pipeline, this is like one night before activity or one-time activity. Now,
[00:29:32] Every time new document come, then again, this…
[00:29:35] process will repeat.
[00:29:37] And, uh, this process doesn't take much time. It is just… even if you're doing for the first time, it might take you half a day.
[00:29:44] Okay, uh, since you're learning now and doing it, so it might take you weeks, okay, but if you're doing it for the first time, after learning it, then it might take you half a day, and then…
[00:29:53] Every then onwards, you might have a, you know, pipeline just doing this entire thing. You just…
[00:29:58] Give the document inside this folder, just run your code, it will create the document.
[00:30:03] And it will create the embedding and everything. All these things will be done.
[00:30:07] And once this embeddings are done, you've…
[00:30:10] These embeddings are created, they are captured in the form of a vector database. So, this is your ingestion part.
[00:30:15] Next comes… is your…
[00:30:17] Retrieving part.
[00:30:21] Okay, so retrieving part.
[00:30:23] involves that you
[00:30:27] Yeah, somebody was…
[00:30:28] Nirvan, I just want doubt before you go for the next part. So, in the data ingestion, the data loading part, what we are talking about, let's say, for example, I downloaded, say, 10 PDFs, okay? Tomorrow, in that 10 PDF, one of the PDFs just got modified and became 200, say, 210 pages.
[00:30:50] No, no.
[00:30:51] So, if I again reupload, is it going to hit the entire PDF, or it is smart enough to only, 10 more pages?
[00:30:52] No, no, no, no. You have to read the entire thing, because your vector database will get misaligned otherwise.
[00:30:57] Okay, okay, fine, thank you.
[00:30:59] Yeah, so it is better to replace that PDF and just read that extra thing.
[00:31:04] Okay.
[00:31:05] Okay
[00:31:07] Read the entire thing again, as the extra thing, along with the entire thing, okay?
[00:31:12] Got it, like a new, new, document
[00:31:16] Okay.
[00:31:17] Yeah, like a new, new entire thing only. So, usually, this data ingestion is like a script.
[00:31:19] How much is it?
[00:31:20] Okay, you just run that script, everything is done.
[00:31:21] Okay, got it.
[00:31:22] Okay? It is like one sequential flow you do, and everything is created.
[00:31:26] Okay. Okay.
[00:31:39] Okay, so next is, uh, so this is, this is, like, in form of a retrieving.
[00:31:43] Okay, and uh… then, uh, after retrieving, you also have the LLM part.
[00:31:49] So, in retrieving, you have a lot of things like, you know, how many… how many documents you want to consider,
[00:31:55] How many, you know, paragraphs or chunks and all these things you want to consider before you come to the final answer.
[00:32:01] And uh… and then, uh, you know, retrieving also involves, let's say, I'm not happy with whatever I'm getting from my embedding.
[00:32:09] Okay, can I also use a full text?
[00:32:12] text match kind of a DB. So…
[00:32:15] If you are retrieving quality is poor, okay, you're not happy with it, you want to hybrid
[00:32:20] You want to make a hybrid rag. Hybrid rag as in you're not happy with the plain rag.
[00:32:24] Uh, you can, you can go ahead and, you know, change your, uh, not only have embedding, but also have a…
[00:32:29] database which… which just stores the actual text, and you can do a text match onto it. So there could be n number of algorithms that you can use, n number of processes, like Lucene search,
[00:32:37] And full-text search, plain full-text search, you can do open search, elastic search, although…
[00:32:42] Now, OpenSarge Elasticsearch has a fork of each other, okay, from one version onwards, one became open source and one became paid.
[00:32:49] So, you can use either of them, and there are a lot of these text search capabilities also you can use.
[00:32:55] Okay, apart from…
[00:32:57] So those are searching capabilities. Those have nothing to do with rack speciality. That text search was there before also.
[00:33:03] Uh, even… even RAG content… RAG is for contextual search.
[00:33:08] Okay, RAG, uh, especially VectorDB is for contextual search. So, the concept of VectorDB
[00:33:13] is especially when you have semantic and contextual search. But if you are not happy with the semantic search, you can add these things.
[00:33:18] These are like customization that you can do in order to, you know, satisfy your rank.
[00:33:23] satisfy and make other, like, RAG as per your need.
[00:33:27] Okay, so you're retrieving a lot controls what you do… what you have done till now, and you can make changes.
[00:33:34] Uh, in the chunking and in the document loading and in the embedding part of it as well.
[00:33:37] Okay, uh, so after that, you have LLM, where you have your prompt engineering and… or your…
[00:33:43] Uh, writing how effective prompt you can write, so that, you know, your answers that come out is very effective out of that retrieved text that you have had.
[00:33:52] Okay, so, uh, so in retrieving, you have mostly, uh, two things. That is, one is known as K, and the other thing is, uh…
[00:34:00] known as thresholding technique.
[00:34:03] Okay, so I'm just, you know, giving you the idea as of now, because we'll do all these things slowly, slowly.
[00:34:08] And, uh, key and thresholding technique, which controls, like, how many responses you are willing to consider.
[00:34:13] Okay, and then in LLM, in LLM, you will have temperature and top P,
[00:34:18] top K temperature, top P,
[00:34:21] And then you also have topK. Topk is not used so much, but, uh, top P and temperature are used.
[00:34:27] Uh, so to get… to play around with the creativity. So, during this rag, you are going to, you know, involve in all these concepts, uh, this entire journey.
[00:34:37] Apart from this, uh, now let's also underst- any question, guys, as of now?
[00:34:44] Yeah, I have one question. So, as you are saying, like, in case of, uh…
[00:34:49] data ingestion part, so if we are not happy with the retrieving, so maybe we could have
[00:34:55] You know, like, full search text.
[00:34:58] So, in that case, like, we are not, uh…
[00:35:01] going, you know, like, we are not doing embedded, so just, we will keep our data as a…
[00:35:07] As it is, as it is, yeah.
[00:35:08] text, but in that case, like, do you think, like, we would have…
[00:35:11] require host database, I mean, like, without embedded. We need more space, I mean.
[00:35:18] Yeah, yeah, yeah, definitely, a full-text search will have more space. Even your embedding model will also have a more…
[00:35:24] space, because when you create an embedding, your actual text is also stored. Or else, how will retain the chunk?
[00:35:30] Yeah, but… yeah, but it will store as a form of vector. I mean, like, uh…
[00:35:35] No, no, no, no, no. The vector is stored along with the form of a pickle file. The pickle file is the actual text.
[00:35:40] Okay, so…
[00:35:41] So, it is not only stored as a form of a vector.
[00:35:44] So, if you store it in a form of vector, how will… how you will get back the extra? So you have a pickle file also that is stored.
[00:35:49] or you… or if you do not want to store the pickle file, you will have to store the actual text also.
[00:35:54] Okay.
[00:35:55] Okay, so in vector database also, like, we are storing the text.
[00:35:57] Yes, yes, yes, yes, yes.
[00:35:59] Okay, okay.
[00:36:00] So that is also… that is also there, okay? But these are very small space. Space is never a problem, you know?
[00:36:07] You know, there was a model known as, uh…
[00:36:12] Yeah.
[00:36:13] Bloom. Bloom. Bloom was one of the first open source
[00:36:15] LLM model came in 2022. Okay, before that, open source,
[00:36:19] LLM development was not, uh, evolving as a… at a time, okay? And nobody was sharing what, uh…
[00:36:26] how I will design my architecture. And also, Bloom came up, okay? So, uh, Blooms…
[00:36:33] was 176 billion parameter model. It was a very big… it was bigger than ChatGPT, uh, GPT 3.5.
[00:36:40] Bloom's entire data was 1.8TB of data.
[00:36:44] And it was supposed to be the entire internet data.
[00:36:47] So, 1.8 terabytes…
[00:36:50] of data. That is entire internet data. So…
[00:36:54] With the kind of data that we deal in, okay, it is not going to cross even few MBs also.
[00:37:00] Text datas are so lighter in nature.
[00:37:03] The moment it goes to image data, then complexity comes in.
[00:37:07] Got it.
[00:37:08] Okay.
[00:37:09] Yeah, so text data space has never been of a problem, although optimizing, like, getting the right
[00:37:17] thing, like…
[00:37:20] like, in terms of fasteness, that has been a problem, but space is not a big problem.
[00:37:25] Okay, because the hosting model…
[00:37:28] The hosting model is way bigger than the actual space of the text.
[00:37:32] Okay, yes, Neera's telling me.
[00:37:38] Yeah, yeah, I said, can we use the JSN data also in, uh, data ingestion?
[00:37:40] Yeah, yeah, yeah, any data, like, there'll be n number of document loader, okay, if you search
[00:37:46] JSON document loader. You will get
[00:37:48] Lot of options. Okay, so…
[00:37:51] a lot of things are there. Okay, it can be JSON, it can be XML, it can be HTML, anything.
[00:37:57] Okay, so as these days are passing, more options are…
[00:38:00] coming. These are the basic options, XML, JSON, and all these things.
[00:38:05] But as the data's applying, people are now coming up with Figma loaders also. Like, you can load the Figma, entire Figma, that…
[00:38:11] What is that called in Figma that… what UI people uses it?
[00:38:14] Uh, wireframes or something. That is the terminology. Even that also you can load, apparently.
[00:38:20] Okay.
[00:38:22] Hi, everyone. Good question.
[00:38:24] So this RAG can be used to store data continuously, like streaming of data, for example.
[00:38:32] Um, a user behavior, it has known.
[00:38:37] Um, features, but this continuously evolving.
[00:38:40] So, a user can log in at one time,
[00:38:43] He is usage pattern of internet.
[00:38:46] Or what websites he generally access, such kind of…
[00:38:49] Yeah, your data ingestion can be a cron job. You can make it like a… see, these are very, you know, very specific needs.
[00:38:56] Okay.
[00:38:57] Okay, see, in the topics, you will not have DC, because these are very specific to your need, but I'm letting you know that if you have a data ingestion, like a Airflow pipeline, you know Apache Airflow, have you heard of it?
[00:39:09] Uh, not really.
[00:39:10] Okay, Apache Airflow helps you to run
[00:39:15] Okay.
[00:39:16] Uh, background jobs, okay? It is not part of Agent DK or anything, it is a normal thing.
[00:39:18] in your normal…
[00:39:22] coding back in your backend job only.
[00:39:24] You run Apache Airflow Pipelines to…
[00:39:28] let's say load data every day. So, let's say one of the job is doing the data ingestion for you.
[00:39:33] Okay, and data… data ingestion is streaming data. So every day, new user behavior data gets loaded,
[00:39:34] Okay.
[00:39:40] And Apache Airflow just loads that into
[00:39:42] Uh, into… into, uh, let's say, a DB or some sort of a file where it loads, and then automatically this data ingestion runs for everyday morning, uh, at 1AM, it runs. It's like that.
[00:39:55] Yeah, it can be done.
[00:39:56] Okay. Yeah, sure. Thanks.
[00:39:59] I… actually, how to handle that dreamage, but uh…
[00:40:04] So we can able to do the OCR per image to extract the text, but in the driving, I just need the reference of, uh…
[00:40:11] image, right? So whenever I'm asking a few questions,
[00:40:14] let's say how to enable some features. So, our PDF contains…
[00:40:20] lot of steps involving images. So, along with the answer text, I need the image as well, how to handle these kind of things.
[00:40:28] Okay, uh, so, uh, so there is…
[00:40:32] something known as, uh, like, you have document loading.
[00:40:35] Uh, you have most of the popular document loaders are PDFs, web page, and text, and docs.
[00:40:41] The… the point is, these are good with textual data.
[00:40:46] But the moment your document has images and text,
[00:40:51] a mixture of all these things, then these loaders are not enough. Okay, that time, the image…
[00:40:58] that time, the kind of loader you require is something known as Unstructured Document Loader. Usually,
[00:41:04] unstructured Document Loader requires a very high-performing system. Okay, which loads
[00:41:10] Everything, your entire document, like a HTML page.
[00:41:13] Okay, so in HTML page, uh, if you, if you over, ever open a HTML page and there are HTML document, you will see TD. Do you… have you ever seen TD, TR, like that?
[00:41:23] Yeah, yeah.
[00:41:24] entries? Yeah, so… so it will divide your entire page into a table… table.
[00:41:30] And everything will be like a table… a TD and table rows, okay? So one of the images will be a table rows, starting point of the table row, ending point of the table row, it's like that.
[00:41:38] So, your pages are converted to in the form of a HTML, and then they are loaded using unstructured loader.
[00:41:44] After they are loaded using that, then this images are converted to image vectors.
[00:41:50] Okay, and the texts are converted to text vectors.
[00:41:53] Mm-hmm, mm-hmm.
[00:41:54] Got it. So… so now, when you are searching,
[00:41:58] It is searched from both the places, from both text and both from images. And the answers come from both text and images.
[00:42:05] No.
[00:42:06] Got it. And when you are using the LLM, when you are using this LLM, make sure
[00:42:11] You are using
[00:42:13] For the image part, or from the text… for the text plus image plus, you are using an image model.
[00:42:19] Okay, which can read both text and image. It can pass both text and image together. It's like that.
[00:42:25] Oh, okay.
[00:42:26] Got it? So it'slike, analyze this image, uh, and based off, uh, based on…
[00:42:31] Uh, that, please answer the user query, followed by your image.
[00:42:35] No.
[00:42:36] Okay. Yeah, that is how it is. That is known as multimodal RAG.
[00:42:39] Okay, it is very difficult in terms of…
[00:42:42] It is… it is not very difficult in terms of building it, but it is very difficult in terms of maintaining it, because
[00:42:47] Unstructured, it's very different to load, and it's very difficult to load. Unstructured doesn't usually work in our systems, okay? I have tried many times in my local system, in my…
[00:42:57] office system which was given. Uh, it didn't work out, unstructured, so…
[00:43:02] Unstructured Loader usually requires a GPU.
[00:43:05] Uh, uh, Gunjani, you have worked on unstructured?
[00:43:07] Multimodal?
[00:43:08] Um, multimodal, we use unstructured.
[00:43:12] Uh, but…
[00:43:14] When you say multimodal, it means what exactly? How are you looking in that?
[00:43:17] Uh, so, your PDF can have image plus…
[00:43:20] Image as in not OCRable image. Uh, I'm talking about rough image. It can be a figure diagram or anything.
[00:43:26] It can be…
[00:43:27] No, in that manner, we have the images, like graphs and everything in the documents, we have some, uh, that, we have…
[00:43:34] Um, seen that in the past.
[00:43:35] But not, like, a separate image. I'm not able to recall, yeah.
[00:43:39] So, mostly documentations we have, like,
[00:43:41] the legal document is a contractuals, financials,
[00:43:46] I got it. So did you extract things out of there?
[00:43:47] And apologies.
[00:43:48] Yeah?
[00:43:50] You do extract things out of that… those graphs and all?
[00:43:52] Uh, yes, but it was happen…
[00:43:57] As… as he is convinced.
[00:44:00] Yes, it is.
[00:44:01] Partially, it was not accurate, so the accuracy was not correct, so I would say that it was adapted very well.
[00:44:04] Yes, yes. So, yeah, that answers actually the question. So, multimodal RAG is a very, uh…
[00:44:10] Uh, it's not a very…
[00:44:13] Easy thing to do it, because it doesn't maintain the order.
[00:44:16] Okay, so you are creating an image vector, you are creating a text vector, but that order is not maintained. Let's say you're…
[00:44:24] your text vector, your third chunk, is talking about…
[00:44:28] The third page, second paragraph. But, in your image vector, it might be the first
[00:44:34] Vector is talking about that image, because you don't… you might not have similar number of image, the amount of text you have, right?
[00:44:40] So that orders and all those things gets lost.
[00:44:43] Okay, so sometimes the continuation gets lost.
[00:44:45] So maybe, in… from the image vector, you are retrieving
[00:44:50] Some content from the 10th page, and from the text vector, you are achieving some content for the third page. Okay, so those things are also very difficult.
[00:44:56] To build. Okay, it is still striving, so there are a lot of these libraries that is coming out known as layout LM, okay, Layout Language Model, which can, you know, understand these unstructured data.
[00:45:06] And build it. But, uh, yeah, multimodal rag is a very difficult thing. If it is a very glory-glory type of a, you know, data where you have…
[00:45:15] in a very expected, predictable structure across all the pages, then it works out very well, it works out. But it is not very partially… it is partially not very good.
[00:45:24] Like, it is very, like…
[00:45:27] unpredictable.
[00:45:28] No, okay.
[00:45:29] Yeah.
[00:45:30] Another question. So, when, uh, during data ingestion, right,
[00:45:34] Let's say I'm ingesting… I'm building some knowledge base, uploading some PDFs, documentation.
[00:45:39] So, multiple users are uploading some documents about the same features, right?
[00:45:44] So, for example, at certain point of time, so let's say, uh, we are building… we are ingestion… ingestion during a 6 months of span, okay? So, same feature containing
[00:45:53] and different documents, having a different version of it.
[00:45:58] Okay? Okay.
[00:45:59] So… so while redriving, right? So…
[00:46:02] Uh, let's say, uh, we are returning the top K chance of, uh, documents.
[00:46:07] Okay, so we are explaining the question, and we are returning the top key agents.
[00:46:10] But it doesn't… are… what is a recent version of that feature, right?
[00:46:15] Yeah, yeah.
[00:46:16] new transition itself, we should arrest it.
[00:46:19] So…
[00:46:20] Yeah, so that… your rag can't help it, that is your architecture capability. So, what we used to do is, we used to ingest Confluence pages.
[00:46:27] Okay, so Confluence pages used to be converted to PDFs first, so in Conflas, there export to PDF.
[00:46:33] So, by, uh, programmatically, only we used to do. We used to take Confluence pages, convert them into PDF,
[00:46:38] And then from PDF, we used to create all these things, okay? Now, once you convert conference page to PDF, now you get some metadata along with that. That is like Confluence capability, nothing to do with RAG.
[00:46:49] Conference used to give it version number.
[00:46:51] If some changes have been in, the version number will be different. So, we used to monitor
[00:46:56] Mm-hmm.
[00:46:57] that if this
[00:46:58] a row. This… this Confluence page has an entry in the table, and
[00:47:03] If the version number is updated, then only take
[00:47:07] a newer version and replace with the older version.
[00:47:13] He is.
[00:47:14] Yeah, here we are, uh, the data is coming from the conflates only, let's say I'm uploading, uh, a document of same feature through PDF as well.
[00:47:17] In confines also, I have the same thing. If we are maintained, get as coming from confiance,
[00:47:21] I have the version number, these kind of capabilities, right?
[00:47:24] Hm.
[00:47:25] So, if I'm uploading, uh, same feature in the PDF or doc or something like that, right?
[00:47:31] Huh.
[00:47:32] I just need to… I can't… right now, I don't have, uh, ability to find it, right? Because we don't have a version number, these kind of capabilities. Because, uh, to find that conflict happens, our existing documents, we need to search the rack only, then only we can able to identify it, right?
[00:47:47] Yeah, so in that case, what you can do is you can find… you can take the text from both the versions of the PDF,
[00:47:55] And you can find out that which is more recent, which is more recent in terms of
[00:48:02] Uh, let's say you can either apply a formula like, you know, is that number of tokens more recent? Uh, like, if there are more number of tokens in the newer version, you can take that as a more reviewer version.
[00:48:13] If you're not happy, okay, maybe… maybe in the newer version, my number of tokens can be less. Okay, in that… in that way, you will have to…
[00:48:20] Uh, you know, when you read the document, na,
[00:48:24] Uh, your PDF loader, when it loads, it comes up with few of the metadata. And over there, there are creation dates as well. Okay.
[00:48:31] So, you can compare the creation dates with the previous creation date also. So, these are things that, in hand, that
[00:48:38] that comes. That, like, in hand, I have faced it. And using those, you can do it. Okay. But…
[00:48:44] If you are creating a data ingestion pipeline, you should have a separate pipeline to track your…
[00:48:49] BDFs. Okay, that this is an updated version of your PDF. That has nothing… that is what I'm trying to say, that has nothing to do with RAG in it.
[00:48:55] React cannot solve that problem. React has no capability to solve that problem.
[00:49:00] That is you having an ingestion pipeline or having a quality check kind of a pipeline, or a version control kind of a pipeline, to check that.
[00:49:08] Okay, and those are these methods. So either you, while loading the document, there is a metadata, you can use that metadata.
[00:49:14] Okay, uh, but sometime that metadata might not be very accurate. Okay.
[00:49:19] Uh, it might not give you the actual date of the PDF creation. Okay, so then it might not help you. Other things you can do is you can actually
[00:49:27] see if more number of text is there, so you take that one. Prefer that one. That one you can do. So, these are some of the things that is pre-rag.
[00:49:35] Okay, so it is like your version controlling system, nothing to do with RAC.
[00:49:39] Okay. So…
[00:49:41] No.
[00:49:42] So, yeah, this more than you, this is your… the data analyst or data an… uh…
[00:49:48] Data collector's job, actually, that… to ensure that this is sent to you.
[00:49:53] Mm-hmm, mm-hmm. No.
[00:49:56] Yes.
[00:49:57] Yeah, thank you.
[00:49:58] Yeah, so guys, everybody, whoever is asking these questions, these are things to learn about, okay? These are, you know, expected
[00:50:05] Uh, out of… when you are building rags. So, you know, I hope this is, you know, adding a value to everyone.
[00:50:11] Okay, yes, yes, Sivanch and Deepak.
[00:50:14] One of you, please.
[00:50:19] Okay.
[00:50:20] Another type of problem statement, just trying to understand is how these components could help or build that capability. So for now, let's forget about image. It's only a text
[00:50:34] Something like, what is the correlation between localized microstructure deformation in that mine
[00:50:41] Blade route of Gen 3 turbofan added a 50 degree exceedance is a design limited stage turbine temperature during the hot end
[00:50:52] high takeoff phase.
[00:50:54] Yeah, so when such type of input is an ERAC model receive, it's a kind of a problem like it's a jargon heavy
[00:51:02] Company-specific knowledge
[00:51:05] And also, external reference matters here to
[00:51:14] Hmm.
[00:51:15] get out the answer, right, correct answer for the chatbot so how these subcomponents should function or we should require any pre or post after this
[00:51:20] Rag modeling
[00:51:22] No, see, uh, Sivan, to answer this, so your model's capability. So, you have an embedding model.
[00:51:28] It is purely at the hand of the embedding model, okay? So, see, you as a developer of these
[00:51:33] RAG engine. Your…
[00:51:36] capability is… see, like, let's say if I ask you to build
[00:51:41] Uh… space algorithm.
[00:51:45] Which maintains…
[00:51:47] The orbital movement of a…
[00:51:50] of a, let's say,
[00:51:52] of our space body inside
[00:51:56] Within a proper way, and there is no human involvement involved.
[00:51:59] probably will not be able to do this at this stage, because I had studied about pace, space, that's why I'm talking about it. So, you will not be able to do it, because you're limited by the technology of it.
[00:52:08] Okay, so still people are solving it, and we… me, as my company, is also solving this.
[00:52:13] problem. Okay, uh, uh, with less human intervention, how can you maintain that body in this way? So, this is the…
[00:52:20] you know, limitation we are limited by, because of the technology. Maybe we are speaking about it too high of its time, okay?
[00:52:25] So, same thing for… to answer your question.
[00:52:28] your… your capability of your model searching capability, the searching capability by… it's… it's only the hand of your embedding model that you take.
[00:52:38] Okay, if you're embedding model is able to do that search,
[00:52:41] Okay, uh, then only you will be able to answer or get the actual chunk, which has this answer.
[00:52:47] Or else, it is… you have to try out with different embedding models. You have to try out different combination of embedding models.
[00:52:54] So, this trying out embedding models is why, you know, your AI projects are very empirical in nature.
[00:53:00] Okay, so you have to try out different, different embedding models. Maybe, you know, go for a more granular chunking value. Okay, so that you…
[00:53:09] If that text which you are searching is present inside
[00:53:11] that chunk, it will be able to search, but if contextually, also, it can understand these jargons,
[00:53:18] Then also it can search. Okay. Otherwise…
[00:53:21] If that is not present,
[00:53:23] then your threshold value will be very low, your threshold value will be low, your score value that, uh, that there is no actual relevant chunk from…
[00:53:31] from where I can find the answer, it will be very low, and there'll be no answer coming to it, coming out of it.
[00:53:37] Got it, Sivash.
[00:53:46] Yes.
[00:53:47] Yeah, I understood, so, I mean, we have to… yeah, as you are trying to say is we have to more rely on the embedding technique, but how is the end-to-end workflow or pipeline which you are saying, so when chatbot or user are such type of domain specific question, jargon specific question
[00:53:57] So what should be our supervised
[00:54:00] Learning plan should be, how should we classify or decode should be looks like
[00:54:03] No, there is no classifier, nothing, nothing over here, so I'd, uh, maybe you…
[00:54:13] Hacho.
[00:54:14] I'm just putting what I'm looking for, the answer. I'm not maybe the right words I'm placing, but what the workflow or the pipeline component should be to suffice such needs is, okay, maybe sometime for that domain
[00:54:25] Like, I have asked for aerospace, you are mentioning about orbit things. We don't have that much domain knowledge binded in machine learning. But what should be the workflow
[00:54:35] See, whatever the purpose of RAG is to solve the same question which you are asking, which is,
[00:54:42] Uh,
[00:54:46] my normal LLM is not able to solve this.
[00:54:48] I… my… there are some proprietary terminologies, nomenclatures, my company follows.
[00:54:53] My normal LLM is not able to answer that. That's why I am building
[00:54:58] Exactly.
[00:54:59] RAG, or VectorDB, on top of my company documents, for which
[00:55:03] It can answer those questions. So, you are building RAC for that. The RAG entire workflow…
[00:55:08] about RAD is done.
[00:55:09] Let me ask you more specifically, let's say nowadays, in a cloud, we have a concept of any skills. So, when we have something not surprised with a regular thing, or we need to specialize, we put some skills or MD files into it. So
[00:55:23] In consideration of RAG, when we are saying Islam is not so far is a case, and we are ingestion with our domain-specific knowledge, what should be our pipeline look like, or what should… how the workload should be looked like? And what are those specific name of those components?
[00:55:37] These are the general forms embedding data ingestation, but what we are knowing in the real world with the themes
[00:55:44] Uh, so, see, uh, this is the most smallest unit of building a rag.
[00:55:51] There can be no…
[00:55:54] Ultimately, whatever you do, you will have to come to a chunking, you will have to come to an embedding only. Whatever you do.
[00:56:01] You might build a rack without embedding, you might build a rack purely on, uh, textual search.
[00:56:06] whatever some new, new algorithms that keep on coming up, like BM25 and all those things.
[00:56:13] But, what my point is?
[00:56:14] This is the most…
[00:56:17] In our smallest, simplest,
[00:56:19] thing that you will have to come towards. It is like linear regression in terms of classical ML. So, it is just like that. So, whatever you do, your workflow will be this only.
[00:56:28] You might have a wrapper on top of this and build a more complex architecture on top of this, but this
[00:56:34] is where you will eventually come towards. So even, let's say, if your thresholding answer is not correct,
[00:56:41] Okay, so let's say you're ha- you're not happy with the embedding answer. Then the other thing is you will go for hybrid rack.
[00:56:47] But there can be no bottom… there can be no other versions of that.
[00:56:52] Got it, Sivanch.So, it's like that. See, skills, that is not comparable, that is an agentic architecture.
[00:56:59] That is not at all comparable with this. This is not an agent-tick architecture. This is one of the…
[00:57:05] architecture of an agentic architecture.
[00:57:08] So you have reached to the bottom of the, uh, you know, tree.
[00:57:12] Where you are now operating. RAG is the bottom of the entire agentic architecture, that this is the lowest you can come to.
[00:57:18] Okay, uh, so this is this what builds an agent.
[00:57:23] Okay, you… skills, if you look at it, if the concept of skills, that MD files and sub-agents and all those things are very, very much at the top, because
[00:57:32] Those are operated very much by no-code peoples.
[00:57:34] And very much by people who are not into coding so much. Okay, even people who are into coding also, like, who are doing full-stack, they also operate.
[00:57:42] But the point is, that is at a very superficial level. Then,
[00:57:46] Slowly, slowly, slowly, slowly, what is the most bottom-line thing about skills? Prompt engineering.
[00:57:51] So you… you are coming to that most bottom-line thing, that is prompt engineering, RAG, fine-tuning, all these things we are learning.
[00:57:58] courtesy Vange. So…
[00:58:01] This is the most plain amount of architecture that you will come to, even if you don't get. So,
[00:58:06] Even if you don't get the answer, your flow will be this only. You're not happy with the answer, then you will go to a full-text search. If you're not happy with the answer, you will probably change your chunking strategy.
[00:58:17] Got it. Sivanj.
[00:58:19] Okay.
[00:58:20] Yeah, so this is the most lowest thing, like, there is nothing that can come out after this also.
[00:58:27] Okay, now you can make versions out of this, let's say, uh…
[00:58:31] I… somehow I… my rag answers are not good.
[00:58:36] So, many people are building vector-less rack also nowadays, which I told, I think, last week that Vectorless RAG are also becoming popular.
[00:58:42] Where people, uh, are…
[00:58:45] that is LLM heavy, where instead of storing the embedding, you store a summarized version of your chunks.
[00:58:51] Okay, in a, in a library known as PageIndex, and you use a LLM only to do the searching. So, there's a… it's very LLM-heavy.
[00:58:59] Okay, so maybe you are not happy with your normal rag. You go for vector-less rag.
[00:59:04] Okay, for a more… better answer. So, these are different, different architecture you can make changes to.
[00:59:10] make tune into, okay, to come to a final answer.
[00:59:16] Okay.
[00:59:20] Yeah. Any questions?
[00:59:22] Yeah
[00:59:23] Yes, debug.
[00:59:25] Right. So, I have a use case. For an example, right, where, you know, our source for ingestion pipeline is a GitHub code wrapper, which is
[00:59:35] Like, which has thousands of files, you know, because we… in Enterprise, the code gets bigger and bigger. For an example, we might have two use case. One is
[00:59:46] Suppose we want to fix some bug, like, for an example, right? So that require, some sort of indexing and understanding of the code, which will sort of go through the code and understand which part of the code has to be fixed, something like that, right?
[01:00:01] So, in that way, suppose if we take this use case
[01:00:05] How do we approach in a way that it's more optimized? For an example, right, if I keep indexing the code rep again and again, then there is a cost to it
[01:00:18] Yeah.
[01:00:19] How do I make sure that, you know, the cost is optimized, and I don't heavily rely on the LLM for this job
[01:00:27] Because
[01:00:28] Otherwise, it'll defeat the purpose. I can't really do that, you know. Probably not… if not now, tomorrow, I'll have some restrictions
[01:00:38] Hmm. Okay.
[01:00:40] Okay. So, see, see, this is completely…
[01:00:44] an off-topic question, why I'll tell you? This has something to do with coding agents. Okay.
[01:00:49] So, what you are telling is how coding agents and clot code internal working. So…
[01:00:56] they… this clot code and coding agents, they have…
[01:00:59] Uh, this will actually diverge the topic after today's topic.
[01:01:03] Uh, so, Deepak, it is… I'm just giving you a high-level overview. So, there is something known as
[01:01:08] agent harness, which all your cursor, anti-gravity, clot code, open code, command code, everybody uses.
[01:01:16] agent harness. So, in that agent harness,
[01:01:19] Uh, you can define rules as to how many times it will traverse through a particular piece of code,
[01:01:24] How many times it will traverse through your GitHub,
[01:01:27] How many, uh, like, in the agentic mode, what it… all it will do.
[01:01:31] Okay, so you can…
[01:01:34] you know, make changes in your agent hardness over there. So, agent harness is like a, uh, it's like a set of rules that you set, okay? So…
[01:01:41] over there, you can make those changes.
[01:01:44] when you are designing your coding agents. Uh, but it has nothing to do with RAG as such, and, uh, neither…
[01:01:52] you know, any of our topics as such, Deepak. So…
[01:01:54] Because I'm seeing from angle where we are creating a knowledge base, on top of LLM, where, you know, it has a context of our code repos. I'm seeing from that angle.
[01:02:07] Okay
[01:02:08] Yeah, I get it, I get it, but there are a lot of things in a coding agents, okay, a coaching, uh, how many coding agents are there in the world?
[01:02:12] probably 10 to 15. So, you can… you can understand that the complexity that comes with coding a generic ag.
[01:02:20] we all build. We all, as an AI developer, we all build. Probably 74 of you, all 74 of you will build it. Okay.
[01:02:27] RAG. But a coding agent is…
[01:02:30] is… your question is referring to a coding agent. Coding agent is something of…
[01:02:34] enterprise-level SaaS application.
[01:02:36] Okay, so for that, you have to understand agent hardness, you also have to understand how fine-tuning works as well.
[01:02:43] A lot of things, okay? So, uh…
[01:02:48] Over there, also, knowledge base options is definitely there, but that is way more than that, actually. There are skills, like, you know, Sivanj was telling, there are skills
[01:02:57] their, uh, skills is something that's… if you're not
[01:03:01] Happy with the answer, then you can utilize this skill to get the answer. Skill is somebody, let's say your agent,
[01:03:06] Python skills. You activate that skills, or you activate an agent, uh, sub-agent, which will go and use the Python skills capability to come up with optimized Python code, for example. Okay, so these things all gets connected through that agent harness.
[01:03:22] So, it's that agent harness that sits on these coding agents. But coding agent itself is a…
[01:03:28] Uh, it's a separate topic. It's a… it's a completely different topic.
[01:03:31] So what you're saying is coding agent is way more complex, so can't
[01:03:35] Yes, yes, yes, definitely, definitely.
[01:03:37] Sure.
[01:03:38] That is why… that is why we don't create coding agents, we use coding agents.
[01:03:42] And it is like that. Yeah, coding agents are definitely very, very complex thing.
[01:03:46] Got it. So, if I change this use case from according to, for an example, a wiki pages
[01:03:53] Yeah. Yeah, makes sense.
[01:03:56] So… so you are trying to say that, uh, uh, to answer your question, you are trying to, uh, say that
[01:04:05] For a particular prompt, let's say the answer
[01:04:06] is already there with me. How…
[01:04:09] How can I, you know, uh…
[01:04:12] relate to, uh, like, instead of doing the search, uh, like, not search, what did you say?
[01:04:18] Instead of doing the LLM prompting technique again, can we…
[01:04:23] How can we do this?
[01:04:24] Wiki knowledge base, we can sort of
[01:04:26] Yeah, that I got it, but what is then after that? After that, what is your question?
[01:04:30] You can figure it out, like, because if we have already solved it before, so we would understand from the wiki knowledge base.
[01:04:37] Where and what to be fixed, something like that
[01:04:40] So, there is something known as, uh…
[01:04:43] prompt, uh, caching, uh, which all these LLMs are giving options when you are using via their enterprise level.
[01:04:51] So, prompt caching is something where your previous calculation about a particular prompt or similar-looking prompt is kept.
[01:04:59] Okay, as a side, in the cache memory. Okay, so next time, if that similar kind of search is there, that prompt caching
[01:05:06] Technique can be used to answer the previous type of similar type of question. So, that is there, that saves a lot of tokens, actually.
[01:05:13] So, that, uh, can be used. So, when I was
[01:05:17] using GCP recently, I saw a prompt… enable prompt caching. Okay, so that they come up with. So, apparently, they save…
[01:05:25] of these LLMs tool, uh, they save the previous calculation that you had while coming up to this answer.
[01:05:31] in a cache memory. Next time, if somebody, they try to relate, if a similar question is asked, they try to relate to that first.
[01:05:38] If that is solved already, those calculations are used, and your tokens cost go low over there. And from there, that same answer comes out.
[01:05:45] So that is one technique that is there.
[01:05:48] Okay, so this cache is for the user-specific, or for the whole enterprise, like, how do…
[01:05:52] Uh, no, you have to enable it. As an API, when you get that API from… as an enterprise,
[01:05:58] You'll have to enable it, and every API provider, every LLM service provider doesn't give that option.
[01:06:03] Got it. So, in both of the things, we don't see the use case for RAG, right? That's what you're trying to do.
[01:06:10] No, no, no, this has nothing to do with RAG, okay? In RAG, there are a few optimization steps that we will discuss, that is chunking and embedding and all these things, uh, and uh… and let's say,
[01:06:21] There is also, like, when we see other versions of RAC, uh…
[01:06:26] Uh, then we will see, like, as we go towards more agentic purpose of it, so we will see other versions of Agile tryouts with some more versions of air. I'll give you some activity also.
[01:06:34] to try out some experiments on this. And you will see about approximate nearest neighbors, so that also, you know, optimizes your searching capability as well.
[01:06:42] Okay.
[01:06:43] Thank you.
[01:06:47] Uh, sorry, Ariman, so here, like, you are saying, like, we cannot implement this, say, functionality in Rec, but here also, like, let's say if I'm asking same question,
[01:06:57] You know, like, again and again, so maybe I can put in the casset.
[01:07:01] Yeah.
[01:07:02] And it lifted up, you know, like, going, uh, maybe our vector database, so just… I don't have to do, you know, like, again, embedded and all.
[01:07:09] And again, I have to search in vector database, so instead of…
[01:07:12] Now, embedded, uh, see, so problem is, when you are using this embedding model, there is no cost involved.
[01:07:18] Okay, Rhendra. So…
[01:07:19] Yeah, but there is a call, I mean, like, I have to go there, and I have to search, but instead of going and searching from vector database,
[01:07:26] So, just I can take it from Kese, and then…
[01:07:29] Yeah, that… that you can do, but that has, again, that has nothing to do with RAGI. That is your cache memory management and
[01:07:34] Yeah.
[01:07:35] All those things. Yeah. So, that's what I'm seeing, like, in terms of cost manipulation, you can do…
[01:07:40] a lot of, uh, things, uh, there are lots of…
[01:07:45] you know, things can be done, okay? Uh, so, you know, one of you, Prakash, is also saying there is something known as batching you can do.
[01:07:52] And, uh,
[01:07:55] What is the first one, Prakash, you mentioned?
[01:08:00] Pay attention
[01:08:01] page it. Acha, page it, page it addiction mechanism.
[01:08:03] Yes.
[01:08:05] Yeah, so there are n number of things that can be, you know, like this that can be done. Okay, that…
[01:08:09] You can utilize it to, you know, uh…
[01:08:12] get a faster query response to that, yeah.
[01:08:19] Yeah. But anyways, uh, so coming back,
[01:08:22] Uh, to our topic. So, these are the major, major things that we will be doing. So, now, to achieve this, uh, there is something that we'll be exploring, definitely, that is
[01:08:34] LangChain. Okay. So, LangChain is a library.
[01:08:38] to build all of these, uh, you know, components and various components. Previous to Langchain, so I will show you one…
[01:08:46] Partial langchain usage, uh, rag, and I will show you one lang… purely LangChain usage rack.
[01:08:52] With Langchain, uh…
[01:08:55] you don't get to build
[01:08:58] everything, every component doesn't have a lot of freedom to it.
[01:09:02] If you are building, uh, things that Langchain already, if those options, if you are exploring at Langchain already has those options,
[01:09:09] You can… you can… you can build a rag as it is. But if…
[01:09:13] let's say Langchain don't have the options, uh, of…
[01:09:17] of… of doing this, uh, of making certain changes. So, let's say if I want…
[01:09:23] Uh, let's say I don't want to search,
[01:09:26] every, uh, text of it.
[01:09:28] Okay, I want to search only particular text of it. I don't… I want to reduce my timing on that, uh, search capability. Then I will use approximate nearest neighbor. So,
[01:09:36] Langtans, by default, doesn't have that option, so I will go non-lanch anyway, a partial way. So,
[01:09:42] Why am I telling you? Because a lot of the rags are not
[01:09:46] built entirely using LinkedIn, because a lot have been built before
[01:09:49] Before Langchen came up with a lot of options. So now, what has happened is, Langchain also came in the year 2022.
[01:09:55] Okay, uh, with ChatGPT, it also came in the year 2022, and Langchain also…
[01:10:01] had its fair share of struggles, so Langchen previously was built only to build something known as
[01:10:07] Sequential agents.
[01:10:10] Okay, sequential agents.
[01:10:13] So, let's say I have a prompt.
[01:10:17] Okay, and that prompt will be given to LLM.
[01:10:20] LLM will give some answer.
[01:10:22] And that answer will go to…
[01:10:24] That part of that answer will go to some another prompt.
[01:10:27] And that will go to another LLM.
[01:10:30] And like this. So, when you have sequential agent like this, sequential agent like this,
[01:10:35] you… Langchen was famous for that. Okay, as the name says, it's a chain of language
[01:10:41] LLM-powered application. Okay, chain of LLM-powered application. Initially, Langchain came with this concept only.
[01:10:47] Okay, but it didn't get popularity. Parallel, other, you know, things were also coming up in the market, known as Creview AI.
[01:10:57] Uh, which was coming in the market, and which is… which was also exploding the possibility of
[01:11:01] non-sequential agents, where our hierarchical agents, where there is no sequential steps that could be… it can go anywhere, okay? And there also came the concept, uh,
[01:11:12] Uh, of, uh, RAG as well. So, in RAG, you have lots of things. You have Data Loader, you have…
[01:11:20] data, uh, you know, embedding creation and all these things, and then finally a RAG, and then, uh, finally a vector database, and then you have the LLM call, and all these things. So, there are lots of components that came in.
[01:11:30] So, what Langchen did is Langchen started becoming an ecosystem, more of an ecosystem.
[01:11:36] LinkedIn has its initial inherent process of building a chain of LL-empowered application, but it also shifted to something known as Umbrella.
[01:11:44] An umbrella of library. So, it started creating some of the sub-libraries under it, so known as Langchain, Community,
[01:11:54] Langchain community.
[01:12:00] Okay, then also it came up withLangchain Core.
[01:12:03] Which has the code LangChain, uh…
[01:12:06] concepts, then you have, uh, Langchain Classic as well, same thing.
[01:12:11] Uh, same thing, but classic.
[01:12:13] Over here. Uh, then you have, uh…
[01:12:17] Langchain, uh…
[01:12:19] document loader as well. So all those things became part of Lang Jungle. So, Langchin started creating other libraries,
[01:12:26] and made these options, like Document Loader. Document Loader… what document loader was not there before.
[01:12:32] PDF document loader was there before also. So why did I… why did I need a Langchin community for that?
[01:12:37] Langchin started making an umbrella where you will not have to go to anywhere
[01:12:44] You can use just Langchain to build anything you want.
[01:12:47] So, Langchen started becoming like that. It started including document loader under itself.
[01:12:52] It… all the same document loader, which is there before also, it started including something under itself. Then, there are lots of ways to build VectorDBs. So, there is Chroma, there is…
[01:13:02] that is fast.
[01:13:04] There is… uh, there is Pinecone,
[01:13:07] And there are…
[01:13:09] others as well, like Quadrant and all these things. So, FAS became
[01:13:12] the most, you know, fastest, uh, document, uh, you know, searching
[01:13:17] Uh, VectorDB, you know, searching capability is fast. Facebook AI similarity search. So, FAS was individually available, but Langchain started making available under them. Okay, even all our embedding models, embedding models, the BERT models that we'll be using,
[01:13:31] That BERT models was individually available from Sentence Transformer, you can easily download it, but Langchen started making it available.
[01:13:38] So, Langchain started making everything under them. Okay, so that, if you want this, you can import from Langchains this
[01:13:45] sub-library only. So, Langchain became an ecosystem for everything.
[01:13:49] Then, slowly, slowly,
[01:13:51] As Crew AI grew, Crew.ai grew, Gru.ai was by Andrew NG, okay? Andrew NG…
[01:13:57] started Preview AI. It is one of the first
[01:14:00] agent-tic architecture. If you… if you keep LanChain aside, Langchain was sequential agents.
[01:14:05] So if you keep that aside, like, Andrew NG came up with Creeview AI, which was… which was
[01:14:10] agentic architecture as a whole, like, the true nature of agentic architecture came out of it. Very simple, very…
[01:14:18] very less learning curve is there, so people who wants to build rapid prototypes,
[01:14:22] Agents using rapid prototyping, preview AI is the way to go. So, Creview AI came in, so Langchain faced competition, Langchain came up with something known as Langraph.
[01:14:31] Okay, so LandGraph is…
[01:14:34] an agentic architecture.
[01:14:36] Powered by Langchen. This is where LangChain actually kicked off.
[01:14:41] Langchin truly found
[01:14:43] It's, uh, strength with Langraph, because Langraph…
[01:14:48] give… first of all, Langraph is complex compared to Cree AI.AI is very, very used for rapid prototyping and all. Langraph is way complex.
[01:14:56] learning curve is higher.
[01:14:57] Uh, in order to learn Langraph, and…
[01:15:00] And of, uh, then…
[01:15:03] While learning… while creating Langraph, Langraph gave you more freedom.
[01:15:08] Cree UI was not giving a lot of freedom, a lot of controllable, uh, things were not there in Query BI. Langraph gave you that. Langraph
[01:15:16] can help you to make a lot of
[01:15:20] you know, conditional edging, and all these things. You will, once you go to Langraph, you will understand, but as of now, I'm, you know, stopping here.
[01:15:26] Uh, as we continue to rag, uh, that Langraph is up
[01:15:31] These are Lang Chains Agentic architecture, and that, actually, what Enterprise also started using it.
[01:15:37] Okay, ClearUI is good for travel planning and all those things, or itinerary planner, all those kind of use cases, TrioA is very good.
[01:15:44] Okay, but if you want to…
[01:15:47] do something more of… more than that, then Langgraph is the way to do it.
[01:15:51] Okay, so that is how the overall language into Langraph. So, if you go to Langchain,
[01:15:56] Uh, so we will see more about Langchain in the practical
[01:16:00] thing. Uh, so if you go to Langchain,
[01:16:07] Okay, so this is the entire website, they keep on changing it, okay? So, if you go to products, you will see a lot of…
[01:16:15] Langraph and, uh, Langchain, uh, as a, as a product itself also. You have a lot of documentation,
[01:16:20] That comes with Langchain. And you can explore this, but usually when I require anything, I just search document
[01:16:29] a loader slang chain, and I go here, let's say I want to use any document loader, I can go here, I can read about… see, there are different types of loaders. Aquium loader, there are different types of connectors. These are connectors.
[01:16:42] for you to load documents from. Okay, so there is HTML, there is web page, there are…
[01:16:47] lot of… there is ChatGPT Loader,
[01:16:50] Uh, where you can load conversation from exported ChatGPT as well, given your API key. Most of them will be, like, based on an API key.
[01:16:56] Uh, so there are a lot of these loaders, and you can see them, you can, you know, explore a lot of Figma files. This is a Figma file loader where you have to…
[01:17:05] pass the access token, the ID and the key, like this.
[01:17:09] So, there is GitLoader, there is all of these things, Langchen started including so many options were not there in the first time when I was learning about it.
[01:17:16] Okay, so, so many options. There is…
[01:17:18] There is Trello Loader, and every tool, every microSaaS, or every SaaS application that we use,
[01:17:23] Uh, you will find… you'll find a loader being, you know,
[01:17:27] is coming up with LangChain.
[01:17:29] Okay, so Langchain, this is not only for loaders. You also will have
[01:17:34] Langchain for chunking as well, so if you search text splitter langchain, so you will see some
[01:17:41] Text filters as well. Uh…
[01:17:44] So, you will see some HTML header splitters, spaCyText splitter, all these things, but the most popular is a recursive character text blitter that we'll be using.
[01:17:52] Uh, so there is Python code text printer. So different, different parsing techniques are there for different, different loading capabilities that is there.
[01:18:00] Okay, so like this Langdin, all these capabilities, Langchen started including. These were there initially also.
[01:18:06] Okay, but nowadays, I'm seeing even FAS is also underlying chain, everything is under Langchain. Okay, so, fast, you separately don't have to import, you can import using
[01:18:16] Separated import process.
[01:18:17] They're going in
[01:18:19] Yeah, sorry?
[01:18:25] Yeah, Shubam, you were saying something?
[01:18:27] No, I guess somebody was on mute, Anirban
[01:18:31] I'm sorry, I'll do that right now
[01:18:35] Okay, so, uh, so as you, you know,
[01:18:39] See Langchen, you will see more and more options. So, as we go towards the practical, we will see that. So, next, we will go towards the practical as well.
[01:18:48] Okay, uh, so we'll go for a break, and then we'll come back, we'll do the practical. We'll first do the VectorDB part, the ingestion part, the data ingestion part.
[01:18:58] How detention looks like, how…
[01:19:00] you know, how usually when you load a document, how their structure looks like, all those things.
[01:19:05] And, uh, then we will move on to a full-length chain-based rack.
[01:19:08] Okay? Fine, guys.
[01:19:14] Okay.
[01:19:15] Yeah. Okay, guys. So, yeah, guys, uh, we'll go for a break for now, and then we'll come back and we'll do this.
[01:19:23] Okay. Approximately a 7 minutes break, and then we'll come back and do it.
[01:19:28] Thank you.
[01:28:11] Yeah. Okay, guys, so…
[01:28:14] Back, uh…
[01:28:16] All of you?
[01:28:27] Okay, so I am sharing one notebook with
[01:28:31] You all?
[01:29:03] Are you able to access it?
[01:29:14] Still loading
[01:29:15] Okay, fine.
[01:29:20] Yeah, we can
[01:29:26] Okay.
[01:29:33] Okay, so, uh…
[01:29:36] Now, the thing is, uh, the concept of…
[01:29:39] you know, vectored…
[01:29:41] DB. Uh, and, you know, how embeddings works, how the searching retrieval part of
[01:29:47] these LLMs work. That is the first thing that we are going to learn today. Okay.
[01:29:51] Uh, so there is a lot of langchain involvement over here, and whatever I have told you till now, this theoretical concept that I've told you, is explained very well over here as well.
[01:30:00] Uh, so, anyways, we will directly go into the implementation of it now. So, here are some of the libraries that we are installing.
[01:30:09] So, first of all, we'll be using our FAS VectorDB.
[01:30:13] Uh, for creating our VectorDB, and then…
[01:30:18] you will be using sentence transformers for creating the… for using the embedding model.
[01:30:27] embedding model.
[01:30:28] Okay. Uh, then, uh, we'll also be using the concept of
[01:30:34] Obviously, Langchain will be using Langchain, Community will be using, okay, for document loaders and all those things.
[01:30:40] And, uh, from Langchain Order Loader, we'll also be using the embeddings. Okay.
[01:30:44] So, so that we can use the sentence transformer, all the, uh, sentence transformer BERT models. So, these are the libraries that will be required.
[01:30:51] Then, after that, I'll give you a task to change this to ChromaDB as well.
[01:30:55] And, uh, you know, deploy Chroma, even try to use ChromaDB.
[01:31:13] Okay, so let's install this first.
[01:31:15] And, uh, in the meantime, uh, as this is installing, uh, so…
[01:31:20] We will, you know, understand what we are doing over here.
[01:31:25] So, this is one of the BERT
[01:31:27] In our models. It is one of the mini-BERT models that we have, so if you go to send…
[01:31:33] tens, transformers model list.
[01:31:36] If you go over here, pre-trained models,
[01:31:38] You will see a lot of these model options. Okay, you have all, mini,
[01:31:43] Uh, all MPNet-based, we do this, one of…
[01:31:46] uh, allegedly, like, this is one of the best models.
[01:31:48] Uh, and, uh, we are using is this all-mini L6V2.
[01:31:54] Okay, you can try with other models as well. Some of them are good for question answering, so QA, MPNet you are seeing. There is Digital Robata, which is a distilled version
[01:32:02] of Roberta model. Roberta was created by Facebook at some point of time. It is a… it is BERT only, but
[01:32:09] Facebook's version of it, and they took that and made some changes and created it. So, you can try with different model.
[01:32:16] You can try the Mini LM L2, L12V2 also. There's a different version from L6V2, it's a more…
[01:32:22] Uh, newer version.
[01:32:24] So, these are different models to try out, okay? So this, we do when?
[01:32:28] So, when we are not happy with the models, uh, we are not getting good answers, uh, the…
[01:32:34] Brekka, so these, uh, the, you know,
[01:32:37] the required answer, then these are the options that we are willing to consider.
[01:32:43] So, over there, if we change this…
[01:32:45] part, uh, from L6V2 to that model, uh, we can actually change the models.
[01:32:50] Okay, so now, uh, LanChain provides you
[01:32:54] embeddings to be used from sentence transformer directly, we are using from Sentence Transformer directly. Okay, that's why…
[01:33:01] Uh, we just have to mention the model name.
[01:33:03] Also, it provides you to install models from Hugging Face also. So, Hugging Face also has this sentence transformer model, same model,
[01:33:11] It's just that when you are using from Hugging Face, you will have to write even this
[01:33:14] these things also. You have to write the entire thing. Okay.
[01:33:18] So… so Hugging phase is also one of the model listing tools, uh, that we have, where
[01:33:23] People just list open source models, and uh… and so that way, also, we can do it.
[01:33:29] But I will show you another version where we'll be using Hugging Face as well.
[01:33:32] Where we'll show the completed rag, where
[01:33:34] The same model will do via AgingPace.
[01:33:36] Okay, so now this is the model. We are creating an embedding object first.
[01:33:40] So, let's download this model. So, this download model will get downloaded in our
[01:33:45] session. At this point.
[01:33:48] And… let's see.
[01:34:07] Okay, this might take some down time to get downloaded.
[01:34:11] Yeah.
[01:34:32] Okay, so we have downloaded the model. Now, see, if you run this in your
[01:34:39] own system, then again and again, you will not have to do, but since we are running in the collab,
[01:34:43] So, our sessions, once we restart the session, again, this downloading is required, but if you do it in your own system or in the VM space where you will be doing it, or you will be coding out, this model will be downloaded and saved already.
[01:34:56] Okay, so that will not be required. Now comes, we are creating some sample documents.
[01:35:03] So, the purpose of this sample document is to understand the
[01:35:08] capability of…
[01:35:10] document loader. Okay, first.
[01:35:13] So, when you load a document,
[01:35:15] Okay, that loaded documents…
[01:35:18] will be…
[01:35:22] present across multiple periods, or maybe in one PDF only.
[01:35:26] So, when you load a document, there are two things that we'll see with the loaded document. One is…
[01:35:32] page content. The other thing is metadata.
[01:35:36] So, page content is… especially has
[01:35:39] Uh, the thing, uh, the entire text of the page.
[01:35:44] Okay? And metadata will have some extra things, like source of the page.
[01:35:50] version of the page. Okay, uh, then author of the PDF,
[01:35:54] Then, uh, creator, uh, any… if you… if you have used a special,
[01:35:59] tool to create that PDF, so that
[01:36:01] tool will also be there, mentioned, let's say, FPDF. FPDF is used to create PDFs on Python.
[01:36:07] Okay, so Creator is FPDF.
[01:36:09] So all those things will be mentioned. So, any extra knowledge, extra informations are there on metadata. You can have even versions also.
[01:36:16] mentioned over there. But that depends. That version might not be the most accurate, but it also comes.
[01:36:22] Okay, so these are extra information. You have the page content,
[01:36:26] And extra information are present inside metadata.
[01:36:29] metadata. And this metadata is like a dictionary. You can add more key-value pairs if you want, you can add website,
[01:36:36] of the content. Okay, so you can have a column known as website. So, like that, you can add more things over there.
[01:36:43] So, when you load a PDF, there are two things. Page content and PDF. So, imagine this is, like, one
[01:36:49] page content of a PDF. Although this is just…
[01:36:52] Uh, 7-8 words, okay? But imagine it is as if we have read a PDF.
[01:36:59] Imagine, this is like one PDF, this is like…
[01:37:01] The same PDF, because the source name is same. So, source name is still mammalPets.doc.
[01:37:07] So, this… this…
[01:37:09] And this is, like, one PDF.
[01:37:13] Okay, having 3 different pages, imagine, like that. And, uh, then this is a different PDF, this is a different PDF, this is a different PDF, this is a different PDF.
[01:37:20] Okay, this is how your PDFs will look like. Okay.
[01:37:25] your PDF pages will look like. This is, like, one page of a PDF. Once you chunk this, let's say if I chunk this, if I chunk this, or if I break this,
[01:37:33] This will be further divided into two parts. So, let's say, till here will be one page content.
[01:37:38] Okay, one entry like this.
[01:37:41] Okay, till here, till here, and then there'll be another entry,
[01:37:44] Which will comprise the rest of the text. So if you chunk it,
[01:37:48] you will have further splits of this.
[01:37:50] But, uh, but when you load it, you will have
[01:37:54] One page content, one metadata per page of the PDF, or per page of a web page.
[01:37:59] It's like that. Everybody clear?
[01:38:05] Document is the PDF, and page contents are your pages, sir. Can we think of it in that way?
[01:38:09] Yeah, you can think of it as that way.
[01:38:11] Okay. And this method
[01:38:13] But that is applicable for PDF and…
[01:38:15] And that is applicable for PDF and Docs.
[01:38:18] Got it, got it.
[01:38:19] Okay, for… for text file, it can be different, because text file doesn't have page concept, nah? In one file only.
[01:38:24] Challenge again, huh?
[01:38:25] So everything will be one page container. You can further chunk it. If you chunk it, then you can have multiple base contracts.
[01:38:32] Okay.
[01:38:33] Okay, and what's the use of this metadata, exactly?
[01:38:35] So this metadata serves as some sort of a, like, one of you was asking, like, I forgot.
[01:38:41] Uh, the person who was asking, uh, was also asking about versions and, you know, extra information, like,
[01:38:47] Uh, let's say… let's say for over your source, I change this source to the resume name. Let's say this is…
[01:38:53] This is about a resume about you, Vinit, okay?
[01:38:55] Okay.
[01:39:00] Okay.
[01:39:01] Source can be your resume name itself. Let's say Vinit underscore resume underscore latest dot PDF.
[01:39:03] Okay.
[01:39:04] So, this helps with extra additional information, additional context.
[01:39:07] Okay.
[01:39:08] You can… you can also…
[01:39:10] add more key-value pairs, let's say…
[01:39:13] domain. Let's say I deliberately add domain.
[01:39:16] Okay, I deliberately, after reading all these things, I take one of the documents and just deliberately add domains. So your domain is, let's say, uh… for example, backend.
[01:39:27] Okay.
[01:39:28] Okay, this can also help me, this can help me to, you know, cluster things.
[01:39:32] So that is…
[01:39:33] Okay. So basically, if you cluster it or chunk it into separate sections, sir.
[01:39:39] Okay.
[01:39:41] Okay.
[01:39:42] Yeah, separate sections, separate subsections, it can help you with that. And sometimes this metadata also helps you with extra additional information. Okay, that…
[01:39:44] Okay.
[01:39:45] This is from where the actual information is coming out.
[01:39:48] Okay. Okay.
[01:39:49] Okay, let's… let's say when you search on Google, Google show you the…
[01:39:52] few write-ups, and uh…
[01:39:54] Correct.
[01:39:55] when you search on Google,
[01:40:00] Correct, correct.
[01:40:01] So, Google… see, this is the text, this is the page content. Imagine. This is like… this is like a metadata. This… this thing is like a metadata example.
[01:40:05] Okay, so let's say heading.
[01:40:07] And the source. This is the source. Imagine this is, like,
[01:40:11] how it is happening internally.
[01:40:13] Got it, thanks.
[01:40:14] Yeah, so as I show you more…
[01:40:17] you know, capability we'll see, more metadata, and then you will see, okay, now I get it. Okay.
[01:40:23] Okay, so these are like, uh, imagine we have a few, you know,
[01:40:28] page content of the page, a few pages of the PDF. We don't need to chunk it now. Okay, chunking will do for a bigger page.
[01:40:34] So, we don't need to chunk it now. So, uh, now we have just read the PDF as… imagine we have read it, and now we are just doing to do our vector store creation.
[01:40:43] And just searching from that vector store.
[01:40:45] Okay, so for that, we use langchaincommunity.vectorStore. Again, I told you, under Langchain community, all the extra things.
[01:40:52] like, fast, ChromaDB, uh, document loaders, all these things are added.
[01:40:57] So, from Langchen Community, we import this fast.
[01:41:00] And we just do fast from documents.
[01:41:03] Okay, so this stores…
[01:41:06] Uh, over here, what will we pass? We pass the document.
[01:41:08] that 7 items document. See, if I hover around it, you will see 7-item documents, 7 items that you have.
[01:41:14] And the embedding model that you have created. So you have created an embedding model over here.
[01:41:18] So this embedding model and this document
[01:41:22] is what you are passing. So let's pass this.
[01:41:26] This is creating my vector store. My vector store got created, so this is your data ingestion part, guys.
[01:41:31] Your data ingestion got done. Now, this doesn't end here. You might not like the answer.
[01:41:37] That… that… so this is a… this is a very empirical process. It's like a machine learning model. You have created it.
[01:41:42] Now, you might not like the accuracy. You might go back, you might change it, you might change some parameters, you might come back, you might change this model also.
[01:41:50] Uh, you might go for a bigger model also. Uh, okay, you're not happy with the answer. Those things are still open.
[01:41:55] Okay, those things can… you can do. You can go back, you can do a lot of experimentation goes in this phase.
[01:42:03] Okay, once you have finalized that this is the…
[01:42:04] VectorDB that you are going to choose, then you are final.
[01:42:07] Okay, so usually this experiment takes almost, like,
[01:42:11] 2-3 days to try it out, like, uh, especially when the choosing of the models. Okay, based on research.
[01:42:19] Now, based on your user
[01:42:22] response. Let's say your user is not liking it. You have created a product which your user is also consuming, and they are giving feedback, okay?
[01:42:29] they are not happy with the answer. Uh…
[01:42:31] So, that time, you can go for…
[01:42:34] Uh, different, different, uh…
[01:42:36] you know, models as well, different strategies as well. Maybe you can change from FAST to Chroma as well, you're not liking it.
[01:42:43] So FAS is supposed to be the fastest, okay? It is the… one of the most
[01:42:47] uh, you know, simplest vector DB
[01:42:51] library that is present, ever.
[01:42:53] Okay, Krumah is a little, little heavier compared to FAS, okay, because apparently Chroma internally, Chroma, Chroma is a library.
[01:43:01] Chroma apparently stores the text also.
[01:43:04] Chroma doesn't only stores the vector DB.
[01:43:06] In FAS, the concept of text is not there. It only stores the vector DB.
[01:43:12] It doesn't store the text. The text
[01:43:15] storage capability is provided by Langchain. Langchain is helping with that. Okay, FAS doesn't store, so internally, FAS doesn't store.
[01:43:23] So, when Langchain has made FAS available, they have given the capability to store the text also.
[01:43:28] But internally, when you are searching,
[01:43:30] Using your search,
[01:43:32] It is not carrying the baggage of storing the text. It is only storing the embedding vector.
[01:43:38] only storing the vector DB.
[01:43:40] Okay? That is it. But…
[01:43:43] Langchain has provided that extra feature that you can store the text also. So, when you save this, if you ever save this model locally,
[01:43:51] You can save both the text as well as the vector as well.
[01:43:55] Okay, but inherently fast doesn't have that capability.
[01:44:00] Okay, so anyways, so this is how you create the VectorDivia. VectorDBs are created.
[01:44:04] without chunking anything, we have not done chunking, we've just created the data ingestion. Okay.
[01:44:09] Now comes to similarity search. So, your vector store, which
[01:44:13] you have, which you have created this object.
[01:44:16] So, Langchain has provided some options, like similarity search. Okay.
[01:44:19] So, this similarity search uses similarity search only. So, it's internally all using Langchains capability only, so…
[01:44:27] It just named it a similarity search. Langgen has a different name. Okay, uh, to this.
[01:44:32] index.search, it's kind of like that. So, similarity search, if you use this, if you… right, cats are an amazing pet.
[01:44:38] case goes to 3 if you do. K… can you all take a guess? What is K by now?
[01:44:45] Okay.
[01:44:47] Yeah, top 3. So, top 3 responses will be shown.
[01:44:52] So, like this, we are getting our top 3 responses.
[01:44:56] Okay. Now, guys, a question to you. Things to ponder upon.
[01:45:01] Uh, so…
[01:45:03] Do you think these answers are relevant?
[01:45:10] with the search, but maybe if I look at it from a
[01:45:14] Overall thing it is talking about these animals or mammals or something.
[01:45:18] Okay? Anybody else?
[01:45:31] No, first one seems to be good, but, uh, rest of the answers are not good. I mean, they are not related to cats at all.
[01:45:32] Huh. Okay.
[01:45:33] Okay, so guys, do you all know about cosine similarity?
[01:45:38] Euclidean distance, all these things you all know.
[01:45:42] Everybody… is there anybody who doesn't know? Can you… can you give a down thumbs up if you don't know at all?
[01:45:53] Down symbol, down reaction, if you don't know.
[01:45:57] I'm not aware on your list.
[01:45:59] Okay, very good. Okay, fine. Let's… let's talk about this. Just a moment.
[01:46:12] Yeah.
[01:46:15] So,
[01:46:18] What happens is, others, do you all know cosine similarity?
[01:46:22] Uh, was it…
[01:46:24] No, not
[01:46:25] Yeah, you're gonna give thumbs-down reaction if you all don't know.
[01:46:28] Okay, so any, any topic?
[01:46:30] Guys, considering all the people who are very basic to this,
[01:46:33] Uh, considering them, you all tell me. Like, there are people who already know about this, so they might say they know it, so…
[01:46:40] Uh, but I have to consider the person who doesn't know at all, so…
[01:46:45] Till now, do you have any doubts?
[01:46:47] From here, guys.
[01:46:49] From this topic. Till now, any step you have not understood.
[01:46:53] Till now, except this similarity search.
[01:47:07] Yeah, hi's one of them.
[01:47:08] Yeah, oh, just one thing, like, uh, the vector store which we have created,
[01:47:12] It is an in-memory, correct? Like, we are not using any persistent DBEs, vector DBAs.
[01:47:18] Uh, so we are using FAS VectorDB, like, the library we are using, but uh… but, yeah, it is currently in memory, so if you don't save it, it will not be…
[01:47:26] seen in your files over here. You will not be able to see the files. You will just…
[01:47:30] be able to see the… wait…
[01:47:34] If you load the model, you will just see the model, you will not be able to see the… any memory used for that vector DB yet.
[01:47:40] It is just like when you train a model, it is in memory, so it is like that, it is in memory as of now.
[01:47:49] So, saving and all those things also, I'll show how you can save, you can load,
[01:47:52] And all those things, and…
[01:47:54] All those things I'll show you, I don't know why it is not loading.
[01:47:57] Uh, yeah, it came up. So yeah, there is no space that has taken.
[01:48:01] for the vector DB, so you can't see anything over here.
[01:48:08] Okay. Got it.
[01:48:12] Yeah. So…
[01:48:14] Now, the concept of…
[01:48:16] Cosine similarity and Euclidean distance. Very important concept. So…
[01:48:20] Remember, your embedding models
[01:48:24] Your embedding models does what? It creates matrices like this.
[01:48:32] This is what your all…
[01:48:34] Mini…
[01:48:36] LM…
[01:48:38] doing. This is creating this…
[01:48:46] this kind of a model. All mini-LM, L6V2, okay?
[01:48:49] L6V2.
[01:48:51] This is what this is doing, okay? Uh, it is actually taking your text and converting them into numbers, which your system understands.
[01:49:00] Which is your system understand. So, according to your system, this is what cats are, and whatever
[01:49:05] Cats are this, that, whatever it is, it was written. So, this is your system understanding.
[01:49:11] of your…
[01:49:13] input that you have given.
[01:49:16] Okay, so now…
[01:49:19] When you want to do, let's say I… these things are considered as different dimensions.
[01:49:25] Okay, so let's say this is, like, 4 dimension. 4 dimension. You have 4 different entries, so you have 4 different dimensions.
[01:49:32] Okay. Now, what does cosine similarity does is cosine similarity is basically, it tries to
[01:49:39] measure, actually,
[01:49:41] It tries to see that
[01:49:44] Between two vectors.
[01:49:46] These are known as vectors. When you have array of numbers, these are known as vectors.
[01:49:50] So, let's say you have one vector like this, one other vector over here.
[01:49:53] So, let's say I am trying to check if this one and this one are talking about the same topic or not.
[01:49:58] So, I can do two things. Either I can do cosine similarity. Cosine similarity is basically this, if this is A, a vector, this is B vector.
[01:50:06] It is a dot product of A
[01:50:10] AB. It is a dot product of AB.
[01:50:15] And…
[01:50:17] Normalization.
[01:50:19] Using length of A and length of B.
[01:50:22] So this is your
[01:50:24] Cosine similarity. This gives you a value between minus 1
[01:50:29] 2 plus 1.
[01:50:32] Okay, where and 0, also. Obviously, zero in between.
[01:50:36] Okay, so value ranges between minus 1 to plus 1. Minus 1 means…
[01:50:40] Put the things at top talking exactly opposite. So, let's say if this is talking about hot milk,
[01:50:45] Okay, this is talking about cold milk, for example. For example, okay?
[01:50:50] See, although both of them are talking about milk, okay, so there will… they will not be exactly minus 1.
[01:50:55] But, since it is hot and cold is there, to explain that part, I'm explaining, let's say, hot and cold, if both are talking exactly opposite, then it will be minus 1.
[01:51:04] Okay, if it is talking something relatable, let's say this is talking about
[01:51:08] Uh… mmm…
[01:51:13] Cold milk, or milk, and this is talking about fermented milk, for example.
[01:51:18] So, both of them are talking about milk only. Uh, so both of them will be greater than zero and closer towards 1.
[01:51:25] If they are unrelated, let's say if this is talking about milk and this is talking about coal.
[01:51:30] coal. Okay, then there will be zero. They are not completely opposite, but they are unrelated.
[01:51:37] Okay, so this is the concept of…
[01:51:40] cosine similarity. So, this gives you an idea that are all your dimensions?
[01:51:44] I'll order your dimensions of your embedding vector are talking about the same topic, or they are talking about a different topic.
[01:51:51] Okay? Very useful in using for RAC, because then only you can find out
[01:51:56] that the query that you are making
[01:51:59] Is that query matching?
[01:52:01] The topic about which the chunks are written about.
[01:52:06] Got it? Vinit? Understood?
[01:52:08] Yes, sir.
[01:52:10] Okay, uh, there was another person who didn't understand…
[01:52:14] who didn't know.
[01:52:18] Uh… Narcimado.
[01:52:20] Did you understand?
[01:52:29] Okay, so this is with respect to
[01:52:31] Cosine similarity. There is another metric that can be used.
[01:52:35] Okay, uh, that is Euclidean. So these are also, guys, your hyperparameters. These are your choices that you have, like…
[01:52:41] If you want to… if you're not happy with your, uh, you know, method, uh, if you're not happy with your answers,
[01:52:47] Uh, from the rack, then you can change these metrics also. Usually, this metric change doesn't happen
[01:52:53] Like, this doesn't benefit so much. Usually, cosine similarity will be same as Euclidean only.
[01:52:58] So, co-equiladian is like a distance metric. So, it is like how to measure distance between these two.
[01:53:03] So, if you want to… let's say this is X, this is Y, this is Z, this is A.
[01:53:08] Okay, similarly, we'll have XYZA. Okay, so it is like X…
[01:53:12] One, this is…
[01:53:14] X1, Y1,
[01:53:17] Similarly, you have X2, Y2, Z2, and A2.
[01:53:21] So this is like X1 minus X2.
[01:53:24] Whole square plus Y1 minus Y2.
[01:53:28] Whole square. Like that, Z1 minus Z2.
[01:53:32] whole square, plus…
[01:53:34] A1 minus A2.
[01:53:38] All squared.
[01:53:40] So, just root over of that will give you a distance between this matrix and this matrix. Imagine if this is a 3D space. Okay, I know this is 4D, this is 4 dimension, I cannot draw 4 dimensions, so I can draw 3 dimensions, so…
[01:53:53] Imagine a space, uh, uh, a…
[01:53:56] a point over here, and a point over here. You're measuring the distance between them.
[01:54:00] Okay, using this formula. So that is Euclidean distance. So, if you have anything where you want to measure distances between two words or two topics that is being talked about, you use Euclidean. If you want to get
[01:54:11] direction-based understanding, like, they are talking about that same thing or not, then you use cosine similarity.
[01:54:16] Uh, both of them kind of does the same job, okay? The end that is there, the end that you get, the…
[01:54:23] ending of using these both.
[01:54:26] are kind of same. I don't see… I have seen… I have not seen them changing a lot.
[01:54:31] Okay, uh, so… but if you keep RAG aside, if you keep RAG aside, let's say if you want to do clustering-related activity, distance-related activity, I have seen people preferring Equilidian distance more.
[01:54:42] And if… if you have RAG and all these things, people
[01:54:47] usually have preferred about cosine a little more than Euclidean.
[01:54:51] Okay. Uh, people have…
[01:54:53] This one, one question
[01:54:55] USB panel.
[01:54:56] This is on the vector representation. So let's say the query which I am making is nothing but so I love cat. But whereas the embeddings, what I already have right now has a full amount of sentence. Let's say cat is a domestic animal and then it eat milk and then
[01:55:13] quite a lot of information right now. From vector representation standpoint, here I'm talking about only ILOCATs. So now, this… it has just the three words, assume that, okay, by splitting it, it may get into 6 tokens, and the six tokens will be having some
[01:55:28] vector dimension range, right? And similarly, I have a full para which has multiple occurrences of cat, along with a lot more other teams, and that is also will get represented in the same dimension of vector space. So now, how the number… is it… is
[01:55:43] Some numbers in the vector space will have a match on the other vector space.
[01:55:46] No, no, no, it is not one-on-one mapping, one-on-one mapping. So, Deepan, I understood your question.
[01:55:51] So, I will tell you, so if that clears. So, first of all, your embedding models is trained to take text input,
[01:55:59] Thank you.
[01:56:00] Okay, and convert it into a limited number of vector size. That will be, in this case, it is a mini
[01:56:05] Many LM we are using, so it will be 386 vector, okay? So, mini LMs are 386, it is a half of the full model. Full model is around…
[01:56:15] 768. Okay, the full model. So, these are your… so if you go to Sentence Transformer list,
[01:56:22] Uh, so these are your full models. So, it'll be a little higher in size. So, these are your full model, MPNet, base, V2, these are full models.
[01:56:29] So they have a little higher size. If you take all Mini LM, they are a little lower in size. See, ATMB.
[01:56:35] Okay.
[01:56:36] Okay, so this is the approx size. Uh, it depends on your system where you're downloading also, so…
[01:56:39] So anyways, the point is, you're embedding models takes in input, and they convert into
[01:56:45] a limited number of vector size, dimension size. Okay. So…
[01:56:49] It is like, let's say, you are a embedding model,
[01:56:53] You take whatever input people give you, you will summarize them in your local language. Local language is your embedding vector over here.
[01:57:02] And you are summarizing them into local language within 10 words. Whatever people might give you.
[01:57:07] As input. Okay.
[01:57:08] Okay
[01:57:09] So, now you tell me, which one will be better? People
[01:57:13] giving everything together, and you summarizing them into 10…
[01:57:17] Uh, words are people giving everything but in…
[01:57:22] phase manner, and instead of 110 vector, you are getting multiple 10 vectors. Which one is better?
[01:57:28] Second option
[01:57:29] Yes, and why?
[01:57:31] Because I can represent the data in a better dimension.
[01:57:36] Yes, you are not losing information.
[01:57:37] Rather than compressing everything.
[01:57:40] Right.
[01:57:41] You're not losing information, you are capturing everything. So, same concept over here. So…
[01:57:43] you… now, that… that is where the chunk size come is. So, when I come to chunk size, that is the thing that you'll have to…
[01:57:50] you know, workout, as… if your models… if you answer, uh, if you respond you, then you reduce your chunk size. Maybe you are going for 2,000 characters first.
[01:58:02] per chunk. Now, you will not consider in 2000 chunk at 2,000 characters at all. You're considering 200.
[01:58:09] 400, 500, like that.
[01:58:10] Got it.
[01:58:11] So that is the thing.
[01:58:13] Okay, got it.
[01:58:15] Yeah.
[01:58:17] So, so when you are doing a search over here, this similarity search actually is doing a distance match with this
[01:58:26] With all these chunks, chunk 1, uh, sorry, all these, uh, page content.
[01:58:31] Page content 1, 2,
[01:58:34] 3…
[01:58:35] 4, like this, it is doing with all of them. Okay. And from there, okay, embedding model got trained on a language
[01:58:43] model only. It is a language model, so it got trained on a language data. So, somewhere,
[01:58:47] It fee, it obviously… cats will come. Cats will have a higher weightage. But somewhere, it felt…
[01:58:53] that dogs is also relatable.
[01:58:55] Why do you think, guys, dog is relatable to cat compared to others?
[01:59:03] The supplemented data.
[01:59:04] Yes, because when the model was trained, when I embedding model was trained,
[01:59:09] You will see dogs and cats are appearing together multiple times.
[01:59:13] Because both are mammals, both are pet animals,
[01:59:17] Both can be, you know,
[01:59:20] can be, you know, used at your homegrown pet's animals, so you can actually, you know, pet them, and they have lots of similar characteristics.
[01:59:28] So that is why, when you are searching for cats, your second option you're getting as dogs and not getting goldfish,
[01:59:34] not getting parrots, and not getting Lang Chin and LLM. That is why dogs is coming, because
[01:59:39] In some space, in some, you know, dimension, they are a similar thing also. They are pets. Let's in the pets dimension,they are same thing.
[01:59:48] Okay, they are a companion as, like, like a dog, but not like a…
[01:59:52] parrot or a rabbit. Got it. So that is why you are seeing that second response as, uh, when you are searching, you are seeing the second response as dogs, and the third response as rabbits.
[02:00:03] Okay, somewhere there is a relatability, but this is not enough.
[02:00:07] So, just printing case goes to 3 and printing top 3 is not enough.
[02:00:11] Okay, you will give this to your LLM, no?
[02:00:15] whatever answer you got, you will give this to your LLM, and your LLM will formulate a prompt
[02:00:19] that based on this query, formulate an answer. Your LLM will give… give like that.
[02:00:25] But your LLM will see unnecessary tokens like this. See, everything you cannot avoid. Something will go.
[02:00:30] Definitely. But at least you could have avoided rabbits, let's say, if not dogs, then at least rabbits you could have avoided, you could have avoided so much of tokens.
[02:00:38] Okay, because this will go to your LLM.
[02:00:40] So, you can avoid so much, so much of token.
[02:00:43] So that is why we use something known as similarity search with relevance Scores.
[02:00:48] So, not only Similarity Search, let's do Similarity Search with relevant score. So now, there is lots of transparency, we'll see. If you print this,
[02:00:55] See, the score is coming. This is not a distance, this is a score. So, FAS has, you know, uh, sorry, Langchain has, you know, uh, taken Eclidan distance, and they have added, they have converted that to a…
[02:01:06] Inversely proportional score. So if the distance is less, the score is high. Okay. So…
[02:01:12] So they have taken… so that's why you are seeing that cats,
[02:01:15] And this chunk ad has a higher score of 0.42.
[02:01:18] But now, if you see with dogs, it has…
[02:01:20] almost 10 points lesser. Now, if you see with rabbits,
[02:01:24] See, guys, when the moment you start printing the score,
[02:01:27] That is where the real thing comes. It is in minus.
[02:01:30] Just because you gave KZ goes to 3, it gave you.
[02:01:34] But the moment you start printing,
[02:01:36] It shows it's in minus, guys.
[02:01:41] Got it?
[02:01:46] Yes
[02:01:48] So,
[02:01:53] Yeah. So, that is why you are…
[02:01:56] getting this…
[02:01:59] answer. That is why you are getting rabbit. So now, it is more visible in front of you that should I consider dogs?
[02:02:05] Should I consider cats? What should I consider? Or should I consider rabbit? So, now you can use the
[02:02:10] You can use a thresholding technique that I… if the threshold is above 0.35, or let's say…
[02:02:17] Sometimes, it will not be so idealistic, like, you know, you will get to know that what is the threshold for dogs.
[02:02:22] Let's say if the threshold is above 0.30, then only you will pin the answer. So, if you do that, you're getting two answers, at least
[02:02:29] one unnecessary answer got waved off.
[02:02:32] So, this is what thresholding I was talking about. I told you, nah, in retrieving, you have two things. One is K, and one is threshold. This is that threshold thing.
[02:02:39] Okay,
[02:02:42] And now, how do you decide on this threshold?
[02:02:45] How do you… guys, like, this is a question to you all.
[02:02:48] Like, how do you think you can decide on this session? I decided like this, but uh… let's say…
[02:02:54] proper for a, let's say,
[02:02:56] for a use case, how will you decide this threshold?
[02:03:02] How we will optimize on this area.
[02:03:18] Anyone?
[02:03:28] We get more response, probably I'll go with a higher number
[02:03:31] Okay, okay.
[02:03:35] Okay? Anybody, anybody, Gunjan, what about you? If you are there in the session?
[02:03:55] I think I will just try to be on the positive side, like, uh, more than zero.
[02:03:56] Okay.
[02:03:57] So, there comes in, let me tell you,
[02:03:58] There comes in a concept of golden dataset.
[02:04:00] So what if you, let's say your company,
[02:04:03] will give you a testing scenario, right? If, for these kind of queries, it is working or not.
[02:04:09] So, it'll give you, like, 10-15 queries.
[02:04:10] Okay, or let's say 200 queries. It gives you 200 queries.
[02:04:14] For all the 200 queries, you know the expected answer.
[02:04:17] You… you… it will also give you the expected answer, that if this is the query, this is the answer that should come.
[02:04:22] So, you…
[02:04:25] create a table in your Excel file, or whatever it is, you create a table,
[02:04:32] You create a table,
[02:04:35] a table, which has…
[02:04:43] Wait, guys, wait, wait, wait.
[02:04:47] You create a table,
[02:04:51] Which has query,
[02:04:55] I don't know, my thing is all of a sudden not working. Just a moment.
[02:05:02] You create a table where you have query,
[02:05:05] You have the expected response.
[02:05:09] expected chunk, you can say, or expected page, or whatever is expected page content.
[02:05:13] And you have your result.
[02:05:16] You have your result, number 1.
[02:05:19] Result number 2, because you are returning
[02:05:22] 3 responses, no. Result number…
[02:05:25] Uh, 2, and result number 3.
[02:05:29] So, you collect these things.
[02:05:32] You collect these things for 200 queries like this, for 200 queries, or 100 queries, whatever.
[02:05:36] Wherever you feel you are happy. It can be 200. For my use case, which I was building,
[02:05:41] Uh, first time, it was 50 queries. For 50 queries,
[02:05:45] is my expected response? Is my expected response is coming
[02:05:50] within top one.
[02:05:52] top one result, top two result.
[02:05:55] So let's say it's not coming over here, but it is coming here. So within top 2, it is coming.
[02:05:59] And within top 3, it is coming. Sometimes, it might not come. So, in top 3,
[02:06:04] definitely higher chance of coming. So, top 3, you will get accuracy, let's say, for example, 95%, for example.
[02:06:10] Now, a top 2, maybe you will get to 80.
[02:06:14] Top 1, you will get to 70.
[02:06:16] So, this is how you measure.
[02:06:20] Now, you…
[02:06:22] are not getting the response in your top 3 also, okay?
[02:06:26] Uh, let's say you are not getting. So, your top 3 will be naturally low, so then you change your threshold. You try out, you play around with different, different threshold value,
[02:06:34] to, you know, to manage this number, or to optimize this number.
[02:06:40] Got it, guys?
[02:06:44] Got it?
[02:06:45] Okay, look, please?
[02:06:46] Yeah, I will just repeat once, because the session is almost coming to an end. Then we'll repeat this again tomorrow, but I will just repeat once.
[02:06:54] So, we need, you create a table of 200 or 50 or 100 queries.
[02:06:59] You create a column of expected answers.
[02:07:02] And you create… you also collect all the answers, top 3 answers that you are getting.
[02:07:08] And then, you try to see whether your answers
[02:07:12] is actually…
[02:07:14] Coming in top 3, top 2, or top 1 or not, depending on the threshold.
[02:07:18] And then, if you're not happy, change the threshold.
[02:07:21] Got it.
[02:07:23] Got it.
[02:07:24] So, this is the empirical process I was talking about. This is the hyperparameter that… this is the rag hyperparameter, that you change
[02:07:30] And you come up with a better, better, better, better, more accurate answer.
[02:07:34] So, this is how you do it. Now,
[02:07:35] So this is more of a model for manually evaluating it, right?
[02:07:39] Yeah, this is, like, usually, this is…
[02:07:42] Like, like, manual as in, like,
[02:07:44] You… no, this thing can be purely done through Excel.
[02:07:49] So much, yeah, but I mean to say that validating it will be via my
[02:07:53] Oh, yeah, yeah, yeah, yeah, yeah. No, that way, you know, how you can do is, uh…
[02:07:59] Instead of manually seeing it, you can just…
[02:08:01] Uh, you know, check the source metadata.
[02:08:08] Correct.
[02:08:09] you will get a metadata for this response, and you have an expected metadata.
[02:08:10] If that expected metadata and this matches, then you are fine. We never used to do manual validating.
[02:08:15] Makes sense.
[02:08:16] So that is where metadata is used.
[02:08:17] Okay, so rather than concentrating on the sentence output, the metadata comes into picture, so that at least the accuracy is closer to the threshold that we might be looking at, rather than being 100% accurate.
[02:08:29] Yeah, so it will not be 100%.
[02:08:32] In top 3, you can still get 100%, because top 3 are top 4, you will definitely get it, but I don't think, like, it is very difficult to get 100% in top 1.
[02:08:44] She got it.
[02:08:45] Okay, because these are contextual search, these are not…
[02:08:46] Like, user will always give exact text, and exactly you will get the search.
[02:08:50] Got it, got it.
[02:08:51] In… yeah, so contextual search is more like meaning… meaning search, okay?
[02:08:57] Yeah, yeah. Semantic meeting, yes.
[02:08:58] More like semantic, semantic closeness kind of a thing, right? Okay.
[02:09:01] Okay, anyways…
[02:09:02] Okay, thanks, everyone.
[02:09:03] Yeah, anyways, guys, so I will take, uh, uh…
[02:09:07] you know, I'll wrap it up here. So, this is what we had about the vector DB ingestion part.
[02:09:13] Tomorrow, we will do the complete rag, and we will do the rest of the part of the LLM part, and
[02:09:19] How to create a chain in our langchain, and also accessing different open source models as well. Uh, you know…
[02:09:25] open source API service provider as well, all those things will do tomorrow.
[02:09:28] Okay, so…
[02:09:31] So, guys, I will, you know, wrap it up here. Uh, you know, uh…
[02:09:36] Any questions you all have?
[02:09:46] This vector DVD data will remain only for this session
[02:09:50] No, no, no, I will show you how to save it also, so Shari.
[02:09:53] Okay.
[02:09:54] And any kind of traditional TV also you can use, right? SQL Server
[02:10:01] Post-race
[02:10:02] Mmm… see, RDBMS is not specialized in searching semantic search.
[02:10:05] So, Oracle actually created a VectorDB,
[02:10:09] From there, Oracle SQL TB.
[02:10:13] Okay.
[02:10:19] Yeah.
[02:10:20] Okay, so traditional DB cannot be used. Traditional search like this, where, and you use a big SQL statement, right? Over here, that thing is not there.
[02:10:22] Okay.
[02:10:23] Okay, so that thing is not there.
[02:10:25] That is why. You know,
[02:10:27] FAS is already providing you how to search the distance. It is based on a distance search, not on…
[02:10:32] structured query language.
[02:10:34] Okay.
[02:10:36] Yeah.
[02:10:37] So, FAS is not going to remain, beyond the session, right? There would be some other database for that.
[02:10:40] No, no. No, no, no, no, no. Fast only will remain after the session. I'll show you how to save it as well.
[02:10:45] Okay, okay.
[02:10:46] You can say fast. You can save the model.
[02:10:49] Okay. Okay.
[02:10:52] Yeah.
[02:10:53] Okay. Yeah.
[02:10:54] Thank you.
[02:10:55] Okay, thank you guys, uh, see you tomorrow, guys, uh…
[02:10:59] A wonderful session, wonderful audience.
[02:11:02] Uh, I hope I have added value, uh, all people who are newcomers, if you have still have doubts, no worries, we'll take it tomorrow.
[02:11:17] Yeah, thank you. Yeah, Sadish, it will be like that. It actually depends on the mentor timings. So, since, uh, from the day I have been assigned, so that is where…
[02:11:26] Both have been evenings.
[02:11:31] Okay, thank you. Bye, everyone.