# 04 2026-07-12 AI Agents Contd

course: Module 5 — AI Agents & Agentic Frameworks
module: Module-5-AI-Agents-Agentic-Frameworks
date: 2026-07-12
type: transcript
video_url: https://personal-learn.armco.dev/files/_Recordings/Module-5-AI-Agents-Agentic-Frameworks/04_2026-07-12_AI_Agents_Contd.mp4

---
[01:40:02] First of all.
[01:40:10] Okay, so, first of all, in RAG, guys, as you build, so this is like how you learn RAG, you have learned about
[01:40:19] You know, different two ways of RAG. One is
[01:40:22] Like when you are learning about else over there, you had one part of RAG.
[01:40:27] Okay, you had the component of rag as one part of the sequence.
[01:40:32] Other than that, you also have just learned RAG itself also. So giving the LSL aside.
[01:40:39] Okay, now, are there different versions of RAGs? Yes, there are. There are lots of changes that keeps on coming because, see, guys, this is the architecture.
[01:40:47] Architecture, there is no standard rule. This is the most standard version that you have run.
[01:40:51] Now, there are different things that you can change.
[01:40:54] Okay, so few things that, you know, that is coming at the top of my mind that you can change.
[01:41:00] is definitely your…
[01:41:02] Uh, you know, embedding model can change.
[01:41:05] Okay, your embedding model can change, your embedding model, now we are using open source embedding model from Sentence Transformer, we are using a mini BERT.
[01:41:12] You can use a full-size bird. Okay, if you want to go to transcendence transformer model list, I had shown you last week.
[01:41:17] And you can change the model over there. You can go for some paid models as well. You can use the OpenAI keys.
[01:41:24] You can use the Google Gemini keys and get Google Gemini embeddings as well. Their embeddings are also there. Their embeddings are bigger in sizes as well.
[01:41:32] But that will come with a cost, okay? Rags is done to save a lot of cost, okay?
[01:41:37] But that will, might be counterintuitive for certain people if the cost is a big concern for you.
[01:41:42] So, embedding models you can change, you can try out with other embedding models.
[01:41:46] Other things, you know, other components that you can change is obviously chunk size.
[01:41:51] Your chunk size and
[01:41:53] Your overlap and overlap you can change.
[01:41:58] Apart from that, uh
[01:42:00] Apart from embedding model and chunk size overlap, you can, uh,
[01:42:04] You know, uh, you can also change.
[01:42:07] The type of document loader you are using. So I had shown you that if you go to Langchain.
[01:42:11] Document loaders documentation.
[01:42:15] You will see a lot of document loaders are there. We have used PDF, we have used web-based, but there are N number. There's Figma document loader, there's archive document loader, there's unstructured, there are unstructured email document loader.
[01:42:26] Lots of document loaders are there.
[01:42:28] So let's say you have a directory.
[01:42:31] Inside directory, you have PDF, you have docs, you have text files, you have HTML files, lots of different, different files.
[01:42:38] So you can run a for loop inside that directory blob.
[01:42:42] Inside that directory blob,
[01:42:45] Okay, directory blob.
[01:42:47] And you can, based on the extension of that.
[01:42:50] Thing, let's say this extension is PDF.
[01:42:53] This extension is docs.
[01:42:56] This is HTML. So based on that.
[01:42:59] Running a for loop using an if-else condition, you can read, you can use different, different loader.
[01:43:04] for reading them separately.
[01:43:06] Then, you know, combined them.
[01:43:09] Some way, you know, combine those loaders in some way, uh, combine those loaders and create a chunk out of them.
[01:43:14] And Chang Sai is an overlap also you can change.
[01:43:17] Then, uh, apart from that, so these are the varieties that you will see as you matured in this field or.
[01:43:24] As you, you know, start more development.
[01:43:27] Ah, different RAG will have a different journey.
[01:43:29] Uh, so…
[01:43:32] Apart from that, you can have a different kind of a splitter as well. We are using recursive character text splitter.
[01:43:36] And 90 to 95% use cases I have seen recursive character detector only.
[01:43:40] Okay, that is not changed so much, but you have other text printers as well. You have latest explanators as well, that also you can use.
[01:43:46] So those things are there. Then definitely the LLM models also you can change. Then the top K method, like how much to use topk, top case value, what is the method that you want to use?
[01:43:56] So, lot of things you can modify and you will get an exposure as you come into the industry.
[01:44:02] You will get a lot of exposure about changing how will the, you know, your rack change. But overall concept is this, the fundamentals is this.
[01:44:09] Then this changes this custom version of rags.
[01:44:12] You know, can appear, ok.
[01:44:14] So now my point is that how do you decide?
[01:44:18] On what are the parameters and what is the value that you want to change using evaluation of rags, ok, you evaluate your RAG, ok, based on the actual.
[01:44:27] answers based on the actual context you evaluate.
[01:44:31] And based on that, you get the answer. So evaluation will do after.
[01:44:34] Little after, uh, maybe, you know, next week.
[01:44:37] uh, will do about evaluation.
[01:44:40] But, uh…
[01:44:42] The other thing about, uh
[01:44:46] That you need to understand is
[01:44:47] That there is a concept of ranking of your answers.
[01:44:51] So in RAG, when you are searching,
[01:44:55] Okay, let's say you have
[01:44:57] hundreds of PDFs.
[01:45:00] 100 PDFs, from there you have created.
[01:45:03] Uh, per PDF 10 pages, so 10,000.
[01:45:07] Chunks.
[01:45:09] Okay, uh, sorry, per PDF 10 chunks, so 10 1,000 chunks.
[01:45:14] So, you are running a retrieving engine, you are running a retrieval pipeline.
[01:45:19] Okay, a retrieval pipeline from there.
[01:45:22] At any point of time, let us say you are retrieving
[01:45:25] Top.
[01:45:27] See, if you are getting top three, it's not a problem, but let's
[01:45:30] Your answers could be present in top 10 answers also because your documents are too closer to each other.
[01:45:37] So you are willing to look at top 10 documents.
[01:45:41] Now, from this top 10 documents, sometimes there comes a need.
[01:45:46] that you will give
[01:45:50] from this top 10 also, you will give only top 3.
[01:45:52] or top five. Okay.
[01:45:55] Rough 3 to 125. So,
[01:45:58] Based on this top 10 documents.
[01:46:00] There is a ranking that comes, okay, rank one, rank two, rank 3, rank 4.
[01:46:05] Based on your engine, based on your embedding model that you use, you get a ranking based on their score, you get a ranking.
[01:46:14] Right. Now, why will you trust this ranking? Is this ranking enough for you?
[01:46:18] Okay, for that, we need to learn something known as re-ranking strategy.
[01:46:24] Okay, so that is the next part that we are going to do next.
[01:46:29] That is re-ranking of.
[01:46:31] uh, your answers that you get from.
[01:46:34] rack. So, there are
[01:46:36] Now, what is this AE ranking? So, one of the most generic re-ranking technique.
[01:46:41] is using something like, like if I leave the need of LLM and everything.
[01:46:47] Okay, because why do you need to re-rank? Because
[01:46:49] I told you, from top 10 you want to rank only top 3.
[01:46:53] So then, you know, you need to re-rank, you need to consider if this rank is fine or not.
[01:46:59] Okay, because that is ultimately going to the LLM.
[01:47:03] Okay, so there are many ways you can re-rank. R-rank strategies, there are n number of re-ranking strategies.
[01:47:09] But when I came to the field, ok, there was very less pre-ranking strategies and 1 of them.
[01:47:16] is one of the most standard re-ranking technique.
[01:47:19] is… this is like plain if you are using only semantic search.
[01:47:23] There is a rankings technique known as
[01:47:26] Just a moment.
[01:47:30] Yes, there is a rankings technique known as cross-encoder ranking technique.
[01:47:35] So, what is this cross-encoder ranking technique? Let me share this.
[01:47:57] Okay.
[01:48:01] So, there is this cross encoder ranking technique. In cross encoder ranking technique.
[01:48:06] What happens is, see, first of all,
[01:48:09] You using an embedding model, all mini LM.
[01:48:17] Using this embedding model, you got some embedding.
[01:48:20] Okay, uh…
[01:48:25] You got some embedding using this embedding model.
[01:48:28] Okay, uh…
[01:48:32] AND.
[01:48:33] You also pass the query through this embedding, you also pass the query, query embedding.
[01:48:38] You also have the query embedded.
[01:48:43] Yeah, query embedding. Now, what do you do usually?
[01:48:48] You have an embedding, you have a query embedding, you find a distance, so you find the distance.
[01:48:54] Between query and
[01:48:57] Chang Guan, Chang Tu.
[01:48:59] Chunk 3, chunk 4.
[01:49:06] Chang 4, Chung 5.
[01:49:10] Like this, you keep on finding the distance.
[01:49:13] This is what you do. Now from there, let's say you found this, this, this as your top relevant chunks.
[01:49:19] This is what you do in bi-encoding format. This format is known as bi-encoder format. By encoding format. There is something known as cross-encoding format as well.
[01:49:27] So, cross-encoding format, forget, forget drag for now. Forget retrieval for now, forget anything.
[01:49:34] Anyways, there is a concept, this technique called Mini LM is a bi-encoder.
[01:49:39] model. All mini LM.
[01:49:41] is a bi-encoder model.
[01:49:43] by encoder. In this case, you calculate.
[01:49:46] The encodings for the document, and you calculate the encoding for the query, and then you do a distance search.
[01:49:52] You do them independently.
[01:49:54] Okay, so you keep it…
[01:49:56] Keep the encoding safe somewhere.
[01:49:58] And then when the query comes, you do
[01:50:01] You do a distance match, okay, distance search.
[01:50:04] Now, there is something known as cross encoder. So there is a famous data set by Microsoft known as MS Marco data set, very famous data set.
[01:50:12] It was purely based on question answering dataset that Microsoft had created.
[01:50:16] Okay, and BERT was trained on that data set.
[01:50:20] Okay, and that dataset model, that model name is MS Marco Mini LM ERGC because that also we are using Mini. So this also let's keep it to mini. Then only we can compare.
[01:50:30] Mini-LM. Okay, this also has a full version of it, not the mini, the full version is also there.
[01:50:35] But let's say we take the mini version of it, MS Marco Mini LM.
[01:50:40] Okay, what this does, this is a cross-encoder model. So, what cross-encoder works is
[01:50:45] In cross-encoder model,
[01:50:49] You actually pass a pair of text, so text 1.
[01:50:53] Text to forget drag, okay?
[01:50:55] Text1 and text two. This internally is…
[01:50:59] you know, kept like this class.
[01:51:02] Class token, followed by text one.
[01:51:06] Followed by a separator.
[01:51:08] Followed by text to.
[01:51:10] Followed by another separator. Separator means your text has ended, your sentence has ended.
[01:51:15] So, when you pass a text1 and text2 pairs into Marco Mini LM,
[01:51:21] It gets this…
[01:51:23] 2 sentences clubbed together.
[01:51:26] Separated by a separator and this class token actually gives you the score how much this text 1 and text 2 are interrelated with each other if they are placed.
[01:51:35] side by side. Let's say if I say
[01:51:39] the sentence that I love watching football, for example, and you have created an embedding out of this. Like this, you have created for N number of PDFs.
[01:51:46] And if somebody search what do you love watching, then probably I love watching football.
[01:51:51] Is the chunk that will come.
[01:51:54] So there is no direct relationship. You have counted that independently in a bioencoder model.
[01:52:00] And you are counting this, the query independently, and you are just doing an embedding distance calculation.
[01:52:06] But in this cross encoder,
[01:52:08] They both are placed together side by side, as if.
[01:52:12] As if they… these sentences are somewhere related.
[01:52:15] Okay, and that related or not, that will be judged by your class token. So if you pass into MSMarko a pair of sentences.
[01:52:22] Okay, separated by separators. You don't have to set the separators, it automatically done.
[01:52:26] It will give you a score, MS Marco LM will give you a cross encoder score.
[01:52:31] And that score actually understands the relationship between.
[01:52:35] two sentences. Now,
[01:52:38] Usually, it has been seen that MS Marco is very good, like cross encoder models are better than BI encoders smaller.
[01:52:45] model in terms of ranking.
[01:52:47] In terms of ranking of your text and their like text pairs, it is very good, okay.
[01:52:53] Instead of doing a distance search, if you do this, this is more efficient.
[01:52:57] Sorry, this is not efficient, this is more, you know, accurate.
[01:53:01] But the problem is there is an efficiency problem.
[01:53:05] cross-encoder doesn't store any vector DB or anything. On spot, it takes the pair.
[01:53:09] And it tries to calculate the rank and it gives you.
[01:53:12] Okay, so there is an involvement of time at every element.
[01:53:16] If you give, let's say, 10 documents,
[01:53:19] 10 text like this, so with all the
[01:53:23] Uh, queries, uh, with all the text.
[01:53:25] It will formulate a query. So it will be like query.
[01:53:28] Chang Guan.
[01:53:30] Query Chunk 2.
[01:53:32] So like that, with all the pairs, query and all the rest of the pairs.
[01:53:37] You will be feeded to the MS Marco model.
[01:53:42] So, on spot, it will do all the calculation, it will give you a score. So, that is why it takes time.
[01:53:45] So that is why you cannot apply MS marko for a normal RAG retrieval engine.
[01:53:49] You apply a bioencoder, get the top 10 responses, top 50 responses.
[01:53:55] And from there, you do a re-ranking.
[01:53:57] Of those 50 responses with the query.
[01:54:00] And then, once you have done that re-ranking, you can give that to the LLM.
[01:54:04] That based on this.
[01:54:06] MS Marco, uh, were based on this cross encoder, this is your ranking, this you will get, and then you can feed that to the LLM.
[01:54:13] Okay, so this is one of the most early.
[01:54:17] re-ranking strategies.
[01:54:19] that was there in the market.
[01:54:21] Okay, it is good for semantic
[01:54:23] like search like this. So for this only, I have shared a notebook with all of you.
[01:54:29] Okay, are you able to access and see it yourself once?
[01:55:03] So, this notebook which I have shared with you.
[01:55:07] It has the difference between cross encoders and bi-encoders. So first, we have the query.
[01:55:12] Then we have the collection of chunks. So, imagine these are the chunks, like, I have 5 different chunks.
[01:55:16] Okay, there is no LangChain, nothing, like plain Python.
[01:55:20] Okay, so first…
[01:55:22] V over here we are creating by encoder cross encoder objects. Okay, these are our two model MS Marko mini LM and all mini LM L6V2 same, same versioning.
[01:55:31] Okay, now…
[01:55:34] We will beusing this for
[01:55:37] Now we are doing a query.
[01:55:41] And we are encoding the query also.
[01:55:45] This is the query embedding and I have the document embeddings. So I have the query and the document.
[01:55:49] I will just do a search using cosine similarity.
[01:55:53] Using cosine civility and get the top k.
[01:55:55] Okay.
[01:55:58] So, cosine similarity will give you, like, if they are relatable or not.
[01:56:01] Positive means 0 to cosine similarity ranges between minus 1 to plus 1.
[01:56:05] Okay, uh…
[01:56:07] If they are related, they are similar content, they are talking about the same thing, then similar thing, then it will be more than zero, closer towards 1.
[01:56:13] And if it is completely opposite, then minus 1.
[01:56:17] closer towards minus 1. If it is unrelated, then around 0 it will be, like, 0.01, 0.1, maximum 0.1.
[01:56:24] or minus 0.1 also, so around that, it will range.
[01:56:28] Okay, just a moment. Let the…
[01:56:34] Both the model read itself.
[01:56:46] Okay, so this is a bi-encoder score. So see, as per this.
[01:56:49] This is the bi-encoder score. So, let's say.
[01:56:52] If I get top 5,
[01:56:55] Uh, sorry, top 3. In AA, this is the score that I am getting.
[01:56:58] Now, drowsiness, I have asked a query, which is the medicine not cause drowsiness? It has given drowsiness is a common side effect of several allergic medicine.
[01:57:07] And it has also given us a top answer. This medicine commonly causes drowsiness.
[01:57:12] This medicine commonly causes drowsiness. This is the second one.
[01:57:16] And the last one is this is a non-drowsy medicine. So this is what it is giving. Okay.
[01:57:21] Now, if you look at MS Marco, if you apply the cross encoder.
[01:57:25] Where you take every query and the
[01:57:28] Every chunk, like, and you
[01:57:31] put it in a pairs.
[01:57:33] You make a pair and pass it to crossencoder.predict.
[01:57:36] If you do that, you will get a cross-encoder score.
[01:57:40] And over here, guys, if you sort by the cross encoder score C, drowsiness has gone to the 2nd ranking.
[01:57:46] So this is what it is, guys. You first, if you have lots of search to do.
[01:57:51] Okay, then first use, guys. First, what do you use is?
[01:57:55] Use the, uh, first you use by encoder.
[01:57:58] Okay, it will give you the semantic search.
[01:58:00] Then, from that 100, top 100 or top 50, you apply a
[01:58:04] cross encoder. So, if you apply the cross encoder, you are seeing that drowsiness has come to second one and this medicine commonly causes drowsiness.
[01:58:14] This medicine commonly causes drowsiness should be taken at night is preferred.
[01:58:19] Okay.
[01:58:22] This is
[01:58:27] Yeah, should be taken at night. So you can't see there is a huge differences.
[01:58:33] In terms of ranking.
[01:58:35] So, this is what I am…
[01:58:38] what I'm trying to say, that these are the ranking strategy that we use. This is one of the ranking strategies we will discuss more.
[01:58:44] tomorrow, but this is like re-ranking strategy. See, guys, the formula and all these things are not important. You can create your own formula.
[01:58:52] 90%, the concept is more important, that you are understanding re-ranking over here.
[01:58:57] Okay, the score, why is this, why is that? That is an empirical process. You will apply lots of these formulas like this.
[01:59:04] And you will finally come to a ranking strategy. And what is your ranking strategy is your own call.
[01:59:08] Okay, so I will explain one of the ranking strategy. I had done it, and I think that is one of the most industry across.
[01:59:14] used as a hybrid ranking strategy.
[01:59:16] That we'll be using, because I had a hybrid RAG.
[01:59:19] Okay, so over there we'll discuss about hybrid RAG as well.
[01:59:23] Okay.
[01:59:27] Got it guys?
[01:59:30] Guys understood the concept of ranking.
[01:59:33] R-ranking, like, if you are not happy with your rank, see, usually.
[01:59:38] You know, in 40% of the cases.
[01:59:41] People are happy with their rank, but in 60, 70% of the cases, ranks, they are not happy with.
[01:59:46] And usually, de-ranking comes when you want to work with top 50, top 100 responses like that. Not for top 3.
[01:59:53] We are doing for top 3 for the class, but…
[01:59:55] Usually you will do it with top 50, top 100 like that.
[02:00:08] Okay.
[02:00:11] Got it, everybody clear?
[02:00:20] Yeah, please explain this encoder once again. It is more likely a query such like that
[02:00:26] See, bye later, I understood, whatever the relevant documents, it will pull into the ranking 1 to 3.
[02:00:32] Then, how this cross encoder will validate actually in terms of semantic, it will retrieve the matched documents to the first layer, correct?
[02:00:43] No, see, there is a pattern.
[02:00:46] of the sentences.
[02:00:48] Prakash, let's say.
[02:00:50] If I say, for example,
[02:00:51] Yeah
[02:00:52] You know, actually cross encoder works in this way.
[02:00:58] That if two sentences are very much interrelated with each other,
[02:01:03] There is a chance that their score will be high.
[02:01:06] Okay, so that is why what they do is on spot they put two sentences side by side and they try to calculate the interrelationship between each other.
[02:01:15] So, let's say, I say
[02:01:18] Prakash is a good cricketer.
[02:01:21] First sentence. Second sentence is
[02:01:23] He plays for Bangalore.
[02:01:26] For example, so this is a natural sentence.
[02:01:31] Got it, yeah. Go ahead
[02:01:32] Okay, in this kind of case, MS Marco will give you a higher score.
[02:01:37] In a more natural relationship.
[02:01:39] Sentences. But let's say if I say Prakash is a good cricketer.
[02:01:47] And this, and the query is
[02:01:50] Let us say.
[02:01:53] Prakash is what?
[02:01:55] So this is not a natural sentence relationship. Do you think that these sentence can come together?
[02:02:03] No.
[02:02:04] No, right? The query is completely different. Like if you look at the meaning of it.
[02:02:06] A query is completely different.
[02:02:09] So, in this kind of case, MS Marco will give you a lower score.
[02:02:12] But sentence encoder, by encoder can still give you a higher score.
[02:02:16] Because you're talking about…
[02:02:18] and cricket and only.
[02:02:20] Got it.
[02:02:22] Yeah.
[02:02:23] So the point is like you can.
[02:02:26] Use MS Marko as one of the ranking strategies, but there are others as well, but this is just to give you the idea about ranking strategy exist.
[02:02:33] That once you get the answer, you don't use it as it is. You can use some re-ranking strategy, and one of them I'm showing you is MSMarko.
[02:02:41] Based cross-encoder technique. There could be others as well.
[02:02:44] Okay, so tomorrow we will discuss about
[02:02:47] Uh, one more, which is the most popular one.
[02:02:49] Okay, that we will discuss tomorrow, which is for hybrid ranking.
[02:02:53] Okay. Anyways, so guys with that guys you know today's session comes to an end. I think guys you all learned something new.
[02:03:01] Streamlit to, you know,
[02:03:04] Non-L-Cell-based rags and then
[02:03:06] to de-ranking overview about re-ranking.
[02:03:09] You'll learn something new.
[02:03:19] So, was I slow enough?
[02:03:22] For all of you guys or things was fast.
[02:03:30] Guys, if you can give me the feedback, guys.
[02:03:32] Because I have like another two, three minutes left.
[02:03:38] Okay, so tomorrow we will get into the depth of this, we will look at some others as well, and then
[02:03:45] at the… and tomorrow we'll also go into tool calling as well. We'll also see some LangChain agents as well.
[02:03:51] And all those things.
[02:03:53] At some point, we'll do evaluation of rags also when we matured with other things and when we start integrating RAG to other things.
[02:04:00] We'll see that as well.
[02:04:02] Okay, so anyways.
[02:04:08] Okay, Deepan, I will try, I'll try, like
[02:04:11] You know, I'll try as much as possible.
[02:04:13] These tools I will try to give you.
[02:04:17] Okay. Okay. Thank you guys. Great having you today.
[02:04:21] Bye everyone, see you tomorrow and I will try to solve some the timing issue around August mid.
[02:04:27] Ah, because by that time I will not have any sessions around that time, so I will move it to 9 to 11.
[02:04:36] Okay, fine, thank you.
[02:04:37] Thank you, thank you.
[02:04:40] Thank you.