# 07 2026-07-25 Live Session Agents

course: Module 5 — AI Agents & Agentic Frameworks
module: Module-5-AI-Agents-Agentic-Frameworks
date: 2026-07-25
type: transcript
video_url: https://personal-learn.armco.dev/files/_Recordings/Module-5-AI-Agents-Agentic-Frameworks/07_2026-07-25_Live_Session_Agents.mp4

---
[00:15:05] So actually we had some questions last week, and just want to follow up that where she is right now on that
[00:15:14] Okay, I think Gulin is there.
[00:15:20] Hello, Colin.
[00:15:25] Wait, being…
[00:15:42] This week, I knew that she will not be there.
[00:15:47] Okay, okay.
[00:15:48] In the meantime, let's say in the break, let's say if I get some time.
[00:15:52] I'll call her, okay?
[00:15:54] Sure, sure, sure, sure. But, but one thing, like, whatever the document you were sharing during these sessions, right, while going through this video, right, recordings, right, we are not getting that one, so that is only the concern right now
[00:16:07] For the LMS, it's not there, yeah.
[00:16:11] Okay, great. I'll share, I'll share.
[00:16:12] No, it is not there, and if we'll be getting any way to share that one over
[00:16:14] No, no, I'm shading, I'm shading now only. I have this thing created for you all.
[00:16:21] I, uh…
[00:16:22] Hi, good evening.
[00:16:23] Yeah, I'm shit.
[00:16:24] oh yeah is the manager available?
[00:16:27] Even I'm also pinging, I think they are usually shared across multiple sessions, so they're not replying.
[00:16:33] Actually, I created a ticket for some information.
[00:16:37] It was very important and I created it 1 month ago.
[00:16:40] Every week I'm dropping a reminder for them to update it, but they are not, they're simply saying we'll check.
[00:16:46] And it is like embarrassing for me to give a reminder every week.
[00:16:52] Okay, I will ask, uh, there is a break in between Richik, I will, from my side, I will.
[00:16:57] Call her and her manager is there. Yes.
[00:16:58] Yes, yes sir, I'm here, what happened?
[00:17:02] Yeah, the chick is having some…
[00:17:04] Actually, there was someone before me who had a doubt, so maybe after that I'll ask.
[00:17:09] Yes.
[00:17:12] Yes
[00:17:13] Yeah, Chandrashikhar, you can go ahead. I have shared one doc.
[00:17:14] Yeah, yeah, yeah
[00:17:15] You can go through this. This is, you know.
[00:17:16] Yeah, so we
[00:17:17] Fine for now.
[00:17:18] Well, all the notebooks are here, till now.
[00:17:20] And I'll be adding over here only, ok.
[00:17:23] Okay, okay.
[00:17:24] What's…
[00:17:26] Yeah, it is, now you can go ahead, manager is there.
[00:17:28] Hello, ma'am, uh, actually I had created a ticket… I had created several tickets on Tixie.
[00:17:34] So…
[00:17:36] Um, my ticket number is 5030.
[00:17:39] And this was, uh…
[00:17:40] What was the concern?
[00:17:41] Yeah, it was about attendance. I had sent this query 1 month ago and, uh,
[00:17:46] I think… yeah, I think I have replied you yesterday, will you please check your email?
[00:17:50] No, no, I have not received a reply, otherwise it would have been there in the ticket.
[00:17:56] Okay, let me recheck.
[00:17:59] Anything else?
[00:18:01] Uh, I created it 1 month ago, and almost every week I'm dropping a reminder.
[00:18:09] So, kindly resolve it at the earliest.
[00:18:13] Okay.
[00:18:16] And also, one more thing, whenever we are sending an email, right, at least we are looking into or we are following up, no such communications we are getting. It is simply we are sending mail and we are waiting for the response.
[00:18:28] It is,
[00:18:29] On which email ID… on which email ID you're sending those emails?
[00:18:31] Yeah, yeah, yeah, yes.
[00:18:32] ID, ID, Jupur, sorry, ID rookie.
[00:18:38] Hello, am I audible?
[00:18:40] Yeah, yeah, you're doing.
[00:18:41] I'm asking, can… on which email you're sending?
[00:18:46] IT Rudki, delivery ID
[00:18:49] Yes, yes, I do Ruki, yes
[00:18:51] Okay.
[00:18:56] Actually, we are expecting some kind of response that we are looking into it, or we are following up, because
[00:19:01] Yes, I understand. I understand that it should be there. Actually, Akasha was the one who was working, now she is here
[00:19:07] I'll be taking care of this program for now. So, I had checked your ticket, it's not there. It is closed.
[00:19:15] Uh, are you asking, are you referring to my ticket?
[00:19:18] Yes, yes.
[00:19:19] So my ticket number is 5030.
[00:19:20] There is no open
[00:19:23] Yeah, and I still have the option to close the ticket, so I see that it is still in progress.
[00:19:34] But in my hotel, there is no 530 is showing
[00:19:39] Okay, can you please write again?
[00:19:44] Yeah.
[00:19:45] But then, do I have to wait one more month to get a response?
[00:19:46] No, no, I'll reply during your session only.
[00:19:49] So I should just create a new ticket, or should I send an email somewhere?
[00:19:54] It is up to you. You can apply only in that ticket as well, ma'am
[00:20:02] They might, it will be on my portal
[00:20:03] Yeah, but you're saying you are not able to view their ticket, so there's no use in this.
[00:20:06] Do it once, let's see if it will
[00:20:08] What is your name?
[00:20:10] It's really…
[00:20:11] Good lane. Okay, I'll just type now and let me know if you have received it.
[00:20:29] Yeah, so my comment is updated in my ticket.
[00:20:39] It is rated.
[00:20:41] Yeah. Okay.
[00:20:48] Okay, sir, I'm reviewing his attendance.
[00:20:51] Yeah, I've sent all the required documents and the attachments, so maybe…
[00:20:56] Yes, yes, I have luck.
[00:20:57] Thank you. Thank you, sir.
[00:21:00] Thank you, Annabar, for letting me.
[00:21:04] Yeah, yeah, welcome. Thanks.
[00:21:06] Okay, so guys, how are you all?
[00:21:12] Yeah, but
[00:21:14] How's your meeting and all? Today, you had right
[00:21:15] Okay. Yes, yes, yes, today I just came. I'm in Delhi only for that.
[00:21:23] Hmm. So…
[00:21:24] Nice, like, you know, the only problem was it was not directly for space.
[00:21:29] It was like, you know, most of the companies that has come are around AI only. So there were a lot of.
[00:21:37] AI powered this, AI powered that, AI powered.
[00:21:41] Marketing agency, AI-powered healthcare provider, AI powered everywhere, AI powered.
[00:21:46] was there. Okay, it's it's like I'm like company something else, but.
[00:21:51] Everywhere it is listed as AI powered.
[00:21:54] So, we got some drone startup, we got some family offices connection.
[00:21:59] Investors connections.
[00:22:01] As a talking point for our next few calls that we can make.
[00:22:07] So, we got one drone startup through which we can do some collaboration.
[00:22:11] Uh, because, uh, that is…
[00:22:14] Because that is a space our product is also moving towards drone plus defense plus space.
[00:22:20] So… so yeah, that is one of the good connections, but rest of the things were all.
[00:22:26] Uh, in a very generic. Lots of e-commerce company came, chocolates and all these related companies, so we didn't find any, you know.
[00:22:32] Collaboration on that site.
[00:22:34] So, so yeah, that was the thing, uh.
[00:22:38] It was little generic expo, not very specific to space and defense, which we were, which we have on September 1 more.
[00:22:46] Which is organized by IDEXX only.
[00:22:47] Hydex is India's defense.
[00:22:49] Uh, you know, uh, council which actually
[00:22:52] All these fundings and all that all these defense company gets is through IDEXX.
[00:22:57] So they make sure it happens, so…
[00:23:00] There is an expo by that, so.
[00:23:02] That will be more specific and the expectations will be very clear in front of us. This is like generic.
[00:23:07] All type of startup came. Niti Ayok, president also came.
[00:23:08] Goodbye.
[00:23:10] So all very generic kind of expo.
[00:23:14] So…
[00:23:15] I think the drone expo.
[00:23:17] Uh, no, we, we came for Bharat Preneur.
[00:23:19] No, the upcoming one.
[00:23:20] Okay.
[00:23:22] Uh, that is space and defense expo, so drones will be there, definitely.
[00:23:26] defense drones, not the commercial drones.
[00:23:27] Okay.
[00:23:28] Drones will be there.
[00:23:30] That in September.
[00:23:32] So that also in Delhi only.
[00:23:35] So yeah.
[00:23:38] So anyways, uh.
[00:23:39] That is what, uh, you know, for this we created the product and
[00:23:44] All those things happened, that is a good thing that has happened.
[00:23:48] But, you know, the kind of people came.
[00:23:53] Lacked that understanding except that drone person.
[00:23:55] Wow. So…
[00:23:58] Anyways, it got us good for the next call. We can make with these family offices and all.
[00:24:03] For investments and all these things. Let's see how it goes.
[00:24:09] Yeah, so that is what.
[00:24:11] Uh…
[00:24:15] So yeah, so guys, last day, I, at the end, I think I… we started with.
[00:24:21] The LangChain agents concepts.
[00:24:24] So, I had told you that we will do Langchain agents and then we will do the memory part.
[00:24:29] And then we will move on to prompt J and then LangGraph again.
[00:24:32] Sorry, other agentic architecture, maybe Creview AI first, and then LandGraph.
[00:24:38] Like that. So this is what I had told you.
[00:24:40] Now, my question is, like, do you all remember just
[00:24:51] Yeah, so my point is, guys, do you…
[00:24:54] Or remember last days long chain.
[00:24:59] topic, like, did we…
[00:25:01] Use the weather app and all those things, did we?
[00:25:04] See that, do you all remember that?
[00:25:11] Yeah.
[00:25:12] Hazat, can you all just can you just remind me till…
[00:25:16] What point did we start the history part, or did we not?
[00:25:20] No, not the spot that we have done till the weather one, and
[00:25:26] A startup API was also there, right?
[00:25:29] Hmm.
[00:25:30] Yeah, yeah, yes. API was there, and this one. Weather Agent Memory version
[00:25:36] That one was there
[00:25:37] Aha, so, and the topic about history, chat history was just started, right? I told you.
[00:25:43] Yes, yes, yes.
[00:25:44] You all can go through it and next day we'll discuss.
[00:25:45] Yes.
[00:25:46] So whether one was completed, right, whether and…
[00:25:49] Yeah, yeah, right.
[00:25:50] Oh yeah, yeah, I remember, you all were straying with different, different agents, then you all were trying out.
[00:25:54] Correct, yes.
[00:25:57] some addition to that.
[00:26:01] Yeah
[00:26:02] Okay, got it.
[00:26:03] Thank you. Yeah. So…
[00:26:07] So yeah, so this is what we had last week.
[00:26:11] And we have…
[00:26:12] Could you please share this link, Anirvan? Or it is there in that file wherever
[00:26:16] Yeah, yeah, yeah, the last one, last one, yeah, it is there.
[00:26:19] Okay, okay.
[00:26:29] Okay. So today we will complete this history part first and then we'll move on to prompt engineering.
[00:26:37] basics of prompt engineering will do. I don't know, like, how much you all know about prompt engineering. I heard last day in the class also, like, many… none of you know actually.
[00:26:45] Uh, so, once you see prompt engineering, you'll feel like, like, this is…
[00:26:49] Of thing that we have been doing it.
[00:26:51] Okay, so prompt engineering is very simple also.
[00:26:55] uh, you know,
[00:26:57] 1R should be good, one, one and a half hours should be good for you to understand.
[00:27:02] Okay, so anyways, so this is what first we will do. I'll just restart once this.
[00:27:13] Can you share the link? Somehow I missed it.
[00:27:17] Okay, okay, share it.
[00:27:20] Do you have this link? This link has all the links.
[00:27:23] till now, whatever we have discussed.
[00:27:25] Okay, probably can share that as well
[00:27:26] Yeah, the last, the last one will be that.
[00:27:29] Link.
[00:27:30] No, the last one is not that file, yeah.
[00:27:36] Oh, it is not wait. Yeah, this is what it is. Memory variations.
[00:27:41] Oh
[00:27:45] I'm seeing weather rack, yeah, maybe it's not updated.
[00:27:52] Hey, same thing.
[00:27:56] We have the same thing, we have the history as well.
[00:28:00] Okay, okay, sorry.
[00:28:12] Okay, so first we give the SERP API search open router. Now we give the
[00:28:18] So BPI and then we give…
[00:28:22] Open weathers.
[00:28:25] Here is open weather.
[00:28:30] Okay, then we do the import, then we build the rag.
[00:28:35] Just…
[00:28:38] And then tools…
[00:28:40] Sorry, can you share the link explicitly? Because I don't
[00:28:44] See that on the document
[00:28:46] This, you don't see, this is the last one.
[00:28:49] This…
[00:28:56] No, this one is…
[00:28:57] Yeah, I tried, like, but the last two links are same, like, search plus weather plus rack, not the memory one.
[00:29:00] This one is without memory.
[00:29:03] And you see?
[00:29:05] Yeah, I do have this one, but whatever you're opening, right, on the memory
[00:29:11] answer.
[00:29:12] I don't have that, yeah.
[00:29:17] Here it is.
[00:29:21] Thank you.
[00:29:30] Okay, so it will ask for the file.
[00:29:53] We'll do the…
[00:29:56] database building, all these things are done.
[00:29:59] Now we have the open.
[00:30:02] Open Router LLM.
[00:30:11] We have the agent call, yeah.
[00:30:13] So this is where till last day we have done and I had told you that we will be discussing about chat history.
[00:30:21] Which is which like people feel little.
[00:30:26] You know, little overwhelmed by hearing it because they feel like, you know, what?
[00:30:30] What is this chat history? Because LLM
[00:30:32] Underlying LLM doesn't have any memory concept. They are
[00:30:35] Completely, they are a thing which has learned.
[00:30:39] And they understand human language, they do autoregressive.
[00:30:43] you know, prediction, and they do not have any past.
[00:30:47] You know.
[00:30:50] Think about whatever, whatever you have talked till now.
[00:30:53] Okay, all those things are not there. So you have to create a system that can.
[00:30:57] do that thing that can explicitly manage.
[00:31:01] uh, that kind of a history.
[00:31:02] That kind of a management it can do.
[00:31:04] Okay, so that is what we will be doing from here onwards.
[00:31:08] Okay, it's very simple, guys. History is one of the very, very simple thing, yet people feel like it's a very big thing.
[00:31:15] If you see it, it is pure, pure Python.
[00:31:18] Okay, now there are like, you know, using the frameworks also you can do many people.
[00:31:22] do purely using LangChain. You know, LangChain also provides
[00:31:27] You know, these frame in Langchen also there are options where you can manage history.
[00:31:30] Okay, same thing only, it's just that it's instead of direct syntax, you use Langchain syntax, okay?
[00:31:36] I prefer using without LangChain, okay, because I have been doing it before LangChain came up with history.
[00:31:42] Okay, so it is very simple without LangChain. If you see the frameworks, it's easy to remember also.
[00:31:47] It's very simple. Okay.
[00:31:50] So the concept is LangChain is involved but LangChain's history concept we are not taking, we are
[00:31:55] taken directly Python list only. So we have created a list. So there are three, four variations, okay?
[00:32:02] Of managing history, so one can be, like, one is known as full history.
[00:32:05] So what we are doing is known as full chat history.
[00:32:08] In this case, first we are creating an empty list.
[00:32:12] Okay. And…
[00:32:13] Then we are making the empty list global variable.
[00:32:17] You all know what is global variable, what is local variable, right? You all know, right?
[00:32:18] There we go.
[00:32:23] If you, if you design a variable within
[00:32:24] Yes.
[00:32:26] Yeah, you all know it, right? Okay, so we have made it global so that, you know, it is accessible even outside this thing also and it can be used later also.
[00:32:34] And all these things. So, you know, full chat history.
[00:32:37] We have made it global so that the value you set over here, it is accessible even.
[00:32:42] In the other function also it is accessible.
[00:32:45] Same values can be accessible. Okay, so first we make this one as global full chat history.
[00:32:52] Then, see, first this is empty. Nothing is there.
[00:32:56] Ask agent full chat history is a function that takes in.
[00:32:59] Questions till now, whatever, whatever question you will ask.
[00:33:03] It will take that question.
[00:33:05] It will take the previous chat till now.
[00:33:07] Okay, which is empty as of now. Okay, in the first go, it is empty.
[00:33:11] Then, whatever is the user question will be attached to
[00:33:15] uh, dictionary like this, you know.
[00:33:17] You know, our chats are managed like this, role user content is equals to
[00:33:21] The question.
[00:33:23] And you will invoke, you will invoke your agent. So what is your agent? Let's go back. This is our agent.
[00:33:29] We know our agent will be invoked with this user question.
[00:33:33] With this user question, with this message.
[00:33:36] It will be invoked. So
[00:33:37] If you do that, the result.
[00:33:39] The result, since you are using LangChain over here, the result.
[00:33:45] Will be.
[00:33:47] Role user, content,
[00:33:49] As well as the response as well.
[00:33:52] So when you do result messages, when you do result messages.
[00:33:56] You will get the user question.
[00:33:59] As well as the answer as well.
[00:34:03] So that you can keep.
[00:34:06] as a history till now.
[00:34:09] So, again,
[00:34:10] You, I will just repeat again. So first,
[00:34:16] You have full chat history.
[00:34:20] Okay, imagine this is the first question you are asking.
[00:34:24] Okay, so full chat history is empty.
[00:34:27] You are creating, uh,
[00:34:31] Another.
[00:34:34] Which is…
[00:34:37] Messages is also a list only.
[00:34:39] So, it is a list of this FCH, I am writing it as FCH.
[00:34:43] full chat history. Plus another list.
[00:34:47] Which is a dictionary that has a role.
[00:34:52] And the
[00:34:54] Question Content.
[00:34:58] Role is user, I should have written this in a bigger way. Wait, wait, wait.
[00:35:04] Yeah, role.
[00:35:07] User and
[00:35:11] Content.
[00:35:14] Question.
[00:35:17] See, till, this is till…
[00:35:19] you have given your question till here, till this step.
[00:35:23] Then you are using agent.invoke. This is LangChain's capability as you know.
[00:35:28] You have created an agent, you can just invoke using that.
[00:35:30] Message that you have created at the top.
[00:35:33] So, if you just pass this message over here.
[00:35:36] Okay, if you just pass…
[00:35:39] Messages equals to message, let me write this in a short form.
[00:35:45] Regards to message till now, so if you pass it.
[00:35:50] Now, LangChain has a capability of
[00:35:54] Managing prompts and responses in its own way, you know that you have that, you know.
[00:36:00] role user and the question, and then followed by the assistant response that LangChain has it.
[00:36:06] So once you do an agent.invoke,
[00:36:09] Your agent will answer something and that answer will be part of result no.
[00:36:14] So result inside result.
[00:36:17] Inside result, there is a message.
[00:36:19] There is this message.
[00:36:21] Which is responsible by your LLM itself.
[00:36:26] and maintained by your LangChain.
[00:36:28] This message.
[00:36:31] Okay, this will have…
[00:36:34] This kind of a format, this will have this kind of format. It is a list.
[00:36:40] Of dictionary that will have a role.
[00:36:45] User.
[00:36:49] Followed by content.
[00:36:51] User question.
[00:37:00] This one is finished, one dictionary done, then another.
[00:37:04] Dictionary which is role.
[00:37:06] Can you take a guess? What will be it?
[00:37:13] What will be it, guides?
[00:37:18] agent.
[00:37:19] Yes, node agent, assistant.
[00:37:23] Okay, agent is also fine, like if you are answering it as agent, it's fine.
[00:37:27] Just a moment, guys, I'm getting up.
[00:37:55] Hello? Yeah. Okay, so…
[00:38:00] role, assistant, and then you'll have again the content, which is the response of AI.
[00:38:07] Which is the AI response.
[00:38:09] Okay, this is the AI response.
[00:38:11] Okay, about the user question. So this will be part of your full chat history. This is your FCH.
[00:38:18] After your result is already out.
[00:38:21] Now, if you just want to get FCH says answer.
[00:38:27] You can do full chat history minus one if you do, then you will get
[00:38:30] Only this patch. Minus one is this, no? Last content.
[00:38:34] Okay, so that is stored as an answer over here, so that's why.
[00:38:38] It is just being shown over here, user question, whatever the question has asked.
[00:38:42] An assistant dancer, whatever is the answer. So, this is just a
[00:38:46] you know, fancy way of printing. Okay.
[00:38:49] Now we also have clear full chat history. If you do this, again, this global full chat history.
[00:38:54] is just made forcefully, they are made an empty list, that's all. Your chat history gets cleared. So, this function you can call.
[00:39:00] Whenever you want to clear the chat history.
[00:39:03] So that is the
[00:39:05] is the plan. Okay.
[00:39:08] Now, guys, my runtime got disconnected.
[00:39:10] uh I might have to rerun this.
[00:39:15] Okay, in the meantime, you already start running it, ok.
[00:39:23] Oh, it's connected only, fine.
[00:39:29] Okay, this is done.
[00:39:31] Now, let me ask the first question.
[00:39:35] Okay, what is the current weather in?
[00:39:39] Let's say Delhi.
[00:39:42] Okay.
[00:39:46] See weather tool called
[00:39:48] Okay, what is the question is, what is the current weather in Delhi?
[00:39:54] The assistant gives the current radar in Delhi is as follows, in this way.
[00:39:56] I used weather tool to get this information.
[00:39:58] The current weather, this is the final answer. Again, it is being printed.
[00:40:02] Okay. As in the agent that we have built,
[00:40:10] We are printing it, so…
[00:40:12] That's why you are getting an answer like this, ok.
[00:40:16] Now,
[00:40:18] Now see guys, last time somebody is on unmute.
[00:40:23] And…
[00:40:27] Yeah, so now guys see.
[00:40:28] Last time, when we had this question, and if you ask a follow-up question, it wasn't giving answer.
[00:40:34] But let's say if you write over here,
[00:40:36] Based on that, should I carry an umbrella or not?
[00:40:43] Based on that, should I carry an umbrella or not? Since the current weather in Dili is overcast, there is a possibility of rain, although it is not explicitly stated that it is raining.
[00:40:50] It might be good to carry an umbrella just in case.
[00:40:53] This is the power of history, guys.
[00:40:55] When you asked this question.
[00:40:58] It already has…
[00:41:00] Can you ask the 2nd question?
[00:41:07] It already had that role.
[00:41:10] user
[00:41:14] Content, first question.
[00:41:19] Last question.
[00:41:22] Then, the second question is role.
[00:41:25] User.
[00:41:27] Or the answer, oh, sorry, assistant.
[00:41:30] Assistant.
[00:41:34] Content is answer.
[00:41:36] First AI answer.
[00:41:40] And when you ask the 2nd question.
[00:41:43] It got added over here.
[00:41:45] A rule.
[00:41:49] User.
[00:41:52] Second question.
[00:41:55] And that's why you are getting a second answer.
[00:41:58] You are getting a second answer now.
[00:42:09] Second answer, which is content.
[00:42:14] Okay, let me go to this and take this to, uh…
[00:42:18] White page, you will be able to see it.
[00:42:23] Awful.
[00:42:27] C. Okay, so if you want to see this history, let's…
[00:42:34] Just try to let's see if I can print and show it to you.
[00:42:41] Yeah, so here is the full history.
[00:42:43] Here is a full history. See, human message, what is the current weather?
[00:42:48] Now, next time. What is the current weather in Delhi?
[00:42:51] Okay, AI message.
[00:42:54] Now, first we'll not have any AI message.
[00:42:58] Uh, because in the first structure, you have only the human message, and then the AI message gets attached to it.
[00:43:02] If you go below, oh, there is a tool message also that will be there. Current weather in Delhi, so.
[00:43:07] There is another role attached to it. So by the way, this is also I wanted to cover. So whenever you are using tools, guys.
[00:43:13] So, there are three roles attached to it. It is not only AI.
[00:43:17] Uh, it is not only the AI history.
[00:43:19] Okay, AI and human history. You will also have a tool message also. There is another role that will be there that is known as the tool message.
[00:43:26] tool message, so you have human.
[00:43:29] You have AI.
[00:43:32] Which is triggering which tool to use.
[00:43:34] And then you have tool which
[00:43:40] Just to say
[00:43:42] You have a tool, the tool will have some response, so you are having a tool over here, which is giving some response, current weather in Delhi is this, this, this, this, this.
[00:43:50] temperature and all these things, and then.
[00:43:53] The AI will formulate this response and give it in a proper way.
[00:43:57] Okay.
[00:44:00] Got it. So, from the AI you do not have a reply, you do not have a content in the first.
[00:44:05] thing, but you have a tool call over here, which is the weather tool.
[00:44:09] See, in the first time, in the AI, you do not have a content coming out, but you will have a tool call instead.
[00:44:15] Okay, this is the response. So you have human message.
[00:44:19] AI message, tool call.
[00:44:20] AI message, then.
[00:44:22] the next human message that goes in.
[00:44:25] Then the next human message goes in. Based on that, should I carry an umbrella or not?
[00:44:29] Then, AI says,
[00:44:32] Since the current weather in Delhi is overcast.
[00:44:35] There is a possibility of rain. So
[00:44:38] What did happen then?
[00:44:40] If you look at it,
[00:44:42] User.
[00:44:45] Followed by
[00:44:49] EA, AI just gave you tools to be used.
[00:44:51] Tools to be used, not content.
[00:44:53] Tools to be used.
[00:44:56] Then followed by tools.
[00:44:58] Tools is also a role, by the way, okay? Tool message will be a role.
[00:45:02] Okay, so that will be like weather tool.
[00:45:06] Whether, with our tool.
[00:45:09] Okay, and then you have a AI response.
[00:45:13] So, this is the first chat.
[00:45:16] Then you have an AI response.
[00:45:19] Responds to…
[00:45:21] user.
[00:45:23] And then followed by that.
[00:45:26] You have the next question.
[00:45:29] First, this is the first.
[00:45:33] Iteration.
[00:45:37] The second iteration.
[00:45:45] Okay, where you have the next question and then next time you do not have a tool call, you just have the AI response.
[00:45:51] This is the concept of history, guys.
[00:45:53] This is known as full chat history.
[00:45:56] You are not having any window or anything.
[00:45:59] Got it, guys? All good guys if you are good, can you give a thumbs up?
[00:46:04] Symbol only, chat history.
[00:46:09] Is it get persisted in the workspace
[00:46:12] Yes, it is only in the session.
[00:46:17] Okay.
[00:46:18] Sivanj, currently it is using your…
[00:46:20] collabs RAM, collapse memory, to store it.
[00:46:22] Okay.
[00:46:23] But let's say if you want to store this somewhere.
[00:46:26] You can take this full chart history and store it in a DB.
[00:46:29] Got it.
[00:46:30] Like, you have initialized the agent, so let's, maybe tomorrow when I'm, again, initialize the agent, so you're able to get all those histories of my today's work
[00:46:38] That… that will… that will involve a DB. You will need to involve a DB over there.
[00:46:45] Okay.
[00:46:46] Got it, you just have to store this in a SQL DB or SQLite DB.
[00:46:49] or a SQL DB. This entire full chat history, you can just…
[00:46:52] Insert in a DB with an ID.
[00:46:55] Okay, and every time of initialization, he will read the DBA and reload all the context.
[00:47:00] Yes, if you give, so, let's say you make an option that, uh
[00:47:04] You want to load the chat history from before.
[00:47:07] Then you say yes.
[00:47:10] Uh, and then please provide me the ID. You can give an ID, let's say GUID, you have a GUID to store it.
[00:47:16] Okay, you provide the GUID, it will reload the chat.
[00:47:20] How in ChatGPT happens?
[00:47:22] Yeah, it's the same concept of, like, cloud memory, or we… something
[00:47:26] Different
[00:47:27] See, Claude memory is something that I haven't used. I have used ChatGPT.
[00:47:32] uh, memory.
[00:47:33] Yeah, ChatGPT also having a memory concept, so where they have
[00:47:34] Yeah, same thing. If it is the same thing, then yeah, then…
[00:47:38] If your cloud memory is same as your ChatGPT memory, yeah, that is the same thing.
[00:47:42] Where once you go and you click your older chats,
[00:47:51] Yeah.
[00:47:52] In that older chats is saved as a GUID. And on the UI, you are showing last 10, 20, 30, 40 chats.
[00:47:53] Okay, that you have done. So that, imagine you are being shown that as a user.
[00:47:57] In the backend, it is stored like a GUID.
[00:48:00] And this full chat history is stored as a history till now.
[00:48:04] Okay, so you can have a column.
[00:48:05] Okay, and can we create any sort of metadata, like, here is you're capturing three attribute
[00:48:12] So we can define any number of
[00:48:15] metadata attributes…
[00:48:17] Uh, like, for example, what you are referring to metadata over here, you are talking about.
[00:48:21] Ah, so…
[00:48:22] Role, user, agent, 3 information error
[00:48:24] So in this case, usually full chat history, you can save it as it is. Like you can have a column only.
[00:48:32] Which is full chat history, okay?
[00:48:34] And you can… you can…
[00:48:36] You can just keep the entire full chart history entry as a one column only, and
[00:48:40] Whenever you are retrieving, you retrieve this entire full chat history.
[00:48:44] And then, if this full chat history is like a list object,
[00:48:47] on your Python, you can retrieve using that list.
[00:48:50] Blitz retrieval technique. Like, you can use a slicing concept to get the last chat, previous chat, last 10 chats, 20 chats, 30 chats.
[00:49:00] No.
[00:49:01] Got it. So I would say store it as it is, stored like a list only.
[00:49:05] Got it.
[00:49:06] Go ahead.
[00:49:08] So maybe when we are doing the mini-project, okay.
[00:49:12] When we are doing a mini-project, we will
[00:49:14] You know, that time we will create a DB and we will try to store last two, three.
[00:49:19] people's chat.
[00:49:20] Sure, sure.
[00:49:21] That time will do it. Okay.
[00:49:23] Okay.
[00:49:24] So since DBs are not covered in the AI.
[00:49:29] Batches, so that for that I don't have a ready-made content ready.
[00:49:33] In the mini project, we are a little flexible about doing things.
[00:49:36] So over there, we'll try to show some like SQLite. Maybe I'll take a.
[00:49:41] SQLiteDB created on Colab itself, and over there only I show it to you.
[00:49:46] Okay.
[00:49:54] Yeah. In agent.invoke, invoke is an internal function, right? What else do we have as a
[00:50:00] function. So, invoke all these agents, uh, if you go to
[00:50:06] These object, uh, where is it?
[00:50:13] Just a moment.
[00:50:18] If you go to these objects, you will
[00:50:28] See, these are the options.
[00:50:31] Different, different options that we have. Sugar.
[00:50:35] That you can look forward to, there are lots of things.
[00:50:38] Okay, there is C schema as well, and
[00:50:41] That will, that is more like a function, it's more like an attribute.
[00:50:45] Okay, there are lots of things like you can get get name and getToolsName and get input schema and all these things.
[00:50:52] Okay.
[00:51:00] Okay, one more. Yes, Neeraj, you had raised your hand.
[00:51:04] Uh, sorry to reach you, sir. Sir, as we are storing all the char- uh, chat session.
[00:51:10] Chatter, all the chat session in the one variable is a limitation of the chat session?
[00:51:17] The only limitation.
[00:51:20] Chat session is more of a memory thing. It is storing in your memory, in your RAM.
[00:51:21] Yes. Okay, yes, yes, that's a message about the memory things, yes, yes.
[00:51:24] Yeah, so I think one of you asked this question last week also with that time, theoretically when we were discussing, you asked.
[00:51:31] So, that time I had told you that your memory is the limitation, like whatever is your memory. So, you can show, like,
[00:51:37] You can store like probably thousands of chats also in this full chat history.
[00:51:41] Okay, the point is…
[00:51:43] Even if you'd stored thousands of chats.
[00:51:46] Maybe in the thousandth chat, let's say, not thousandth,
[00:51:47] Yes.
[00:51:50] But let's say 500th chat, when you give it to the LLM.
[00:51:53] The LLM will not take all of them as input. It will max, it will cap
[00:51:58] It will, you know, in between it will.
[00:52:02] Uh, you know, cap when it is taking as an input because
[00:52:04] Because LLM's limitation is there.
[00:52:06] Coordinated. Your chat is stored.
[00:52:09] In full chat history. But when you are giving to the LLM,
[00:52:12] In the LLM will cap out few last, uh, you know, last few chats will be capped out.
[00:52:17] Previous chats will be captured, because LLM can only take up to its context side limits, na.
[00:52:18] Okay.
[00:52:21] Limit. What are Ni Raj.
[00:52:22] Yeah, actually, why I'm asking in a programming concept, when we're storing large amount of memory in any variable.
[00:52:30] It gives you sometimes error that's like a auto buffer memory error error.
[00:52:36] That's why I.
[00:52:37] Yeah, so in… see, again, AI agent is a separate concept. I'm talking about
[00:52:41] Like AI agent has.
[00:52:43] The concept of I had told you before also like hardness and
[00:52:48] They have lots of tools to manage.
[00:52:51] those charts, but I am talking about, let's say, when you are building a chat bot.
[00:52:54] Okay, that is you are using a tool and in that you are asking this question.
[00:52:55] Yes.
[00:52:58] Okay, in that tool, you face this question. I'm talking about when you are building a chat bot.
[00:53:02] In your session, you can load whatever, however size you want. That depends on the size of your RAM and memory, and how much it can take.
[00:53:11] Okay, but your LLM is limited. Your LLM will obviously be limited.
[00:53:16] Okay, so that time…
[00:53:18] You can give an option to the user that your memory is 70% full.
[00:53:23] Okay, your memory is 80% full. That time, that means you.
[00:53:28] that you are almost going to touch your LLM's limit.
[00:53:31] That's all.
[00:53:33] So, let's say, let's say, let's say you have seen.
[00:53:37] You have counted all the tokens that is present in your full chat history, all the tokens of user, assistant, user, assistant, everything.
[00:53:38] Yeah, got it. Thanks.
[00:53:45] You have counted. And you have saw.
[00:53:49] that your cloud limit is, let's say, 2 million.
[00:53:51] And you have almost touched.
[00:53:54] 1.8 million already.
[00:53:56] Bye now. Okay, by all the history,
[00:54:00] By all, by taking all the history till now, by all the tool calls and everything.
[00:54:05] If you have taken, you have counted almost 1.8 million, you have touched it.
[00:54:08] Okay. Next chat, if you give, it is going to cross 2 million.
[00:54:13] So that time you provoke the user.
[00:54:16] that your…
[00:54:18] Chat memory is full, ok, or your chat memory is about to get full. Similar thing happens to your tool as well. When the tool asks you.
[00:54:25] that allow… that your memory is going to… like, become full.
[00:54:30] Okay, so they also touch that memory thing only.
[00:54:34] Got it.
[00:54:37] Yeah.
[00:54:38] Yes, sir. Yes, yes, yes, sir. Thanks, sir.
[00:54:41] So, see, guys, if you are relating
[00:54:44] If you are relating a lot of things to your AI agent.
[00:54:47] It is fine. It is fine for you to understand. But guys, when you are learning AI.
[00:54:54] You know, bring yourself out of the AI agent part, I would say, because many people's
[00:55:00] EI limitation is only up to the AI agent. Yesterday I was
[00:55:04] Having a mentoring session and with a student and student.
[00:55:10] Uh, is also going through AI course.
[00:55:13] And they are asking, like, you know,
[00:55:17] How do I build, I have been building.
[00:55:19] agent throughout my life using plot code.
[00:55:22] Okay, uh…
[00:55:24] I told why they told because, you know, on LinkedIn I see a lot of people using plot code to build this agentic application, that agentic application.
[00:55:33] I told them, I told her that in those cases, I have mostly.
[00:55:38] Mostly I have seen the person who is building is mostly a software dev.
[00:55:43] And who, who is excited in using that AI agent, AI coding agent while making that application, but.
[00:55:50] They have no idea about how.
[00:55:52] agentic framework, a framework as in not.
[00:55:55] Like you are building the entire application, I'm talking about framework as in Langchain, LangGraph, Autogen, Crew,
[00:56:01] All these things function. That person is just know how to write prompts.
[00:56:05] And just hitting the prompt and making changes, probably giving the access of.
[00:56:10] Claude code to your GitHub as well.
[00:56:12] Making changes in your GitHub as well.
[00:56:14] And maybe activating skills, uh, skills, various, various skills to get a job done. All those things they are doing. So they know Claude Code very well.
[00:56:22] They know how to use skills in skills, there are prompt engineering that also they know.
[00:56:26] Okay, but they don't know how an agent function and all these things. So, if there is a little bit of change,
[00:56:32] In the architecture framework or anything, or if you are asked to lead.
[00:56:36] a team, or team as in, like, a project, not a team, team, that's a managerial thing, but I'm talking about, let's say, if you are a solution architect.
[00:56:42] of an AI product, you will not never be able to guide your team or take a decision.
[00:56:47] Because your decision is limited to AI agent, AI coding tool.
[00:56:51] Okay, so use them as a tool.
[00:56:54] has a boilerplate code creation.
[00:56:58] Okay, but don't be super reliable on that, that you have built so many agents, but you're not you don't know what is LangChain.
[00:57:02] So that person came on the call.
[00:57:04] For the one hour, almost like 45 minutes almost.
[00:57:08] She learned Crew AI, LangChain, LangGraph.
[00:57:12] All these things, she's heard these terms for the first time. But before that, she has built almost 3, 4 SaaS application already, using all agentic.
[00:57:19] Using Claude Code. She has some subscription. She continuously hits with prompts.
[00:57:24] And she has done it. And prompt, also, she doesn't do prompt engineering.
[00:57:27] She just rides on the go, like how we ride normally on ChatGPT because it's free, no.
[00:57:32] So that's why it is done. She has doing like that.
[00:57:35] But when you go to enterprise, okay, like Gunjan can relate as well, because he building it.
[00:57:41] So over there, the usage of these tools are
[00:57:43] Very limited to fixing your problems, debugging.
[00:57:46] Building boilerplate code, small code, like that.
[00:57:50] Okay, but it is not about, like, you build the entire agent and you don't have an answer for a question.
[00:57:55] Okay, because that is where the AI engineering, then what is AI engineering? Then, then you are just hitting with prompts.
[00:58:03] With you, you are just.
[00:58:06] You know, a little bit of
[00:58:09] Glorified AI engineer, like, I would say, like
[00:58:11] You know, software engineering.
[00:58:14] And you know how to use Claude tools.
[00:58:16] So, if you are…
[00:58:18] or everything is related to that.
[00:58:21] I would say move away from that a bit. Okay, when you are learning this course, after that, you start using it, no problem. But at least.
[00:58:28] Away from that, that way you will…
[00:58:30] You know, understand these LangChain, LangGraph.
[00:58:34] models will be a little hesitant towards doing that.
[00:58:36] So remember this, this is very important because I have seen this with my co-founder as well.
[00:58:41] My co-founder also, you know, used these tools a lot. So he knows these superb, like, and he keeps on telling me like.
[00:58:48] Hey, if you don't use this, you will get left behind. But.
[00:58:52] You know, sometimes he only asked me how a RAG works, but he has built RAG applications already.
[00:58:56] So that's the point I'm trying to make. Learn the core of it.
[00:59:00] AI engineering. Focus should be on AI engineering more.
[00:59:03] These tools, even non-AI people can also learn. That person was non-AI only, no.
[00:59:08] My co-founder is also non-AI only, software developer.
[00:59:11] Okay, so anyways, so low code has a different industry, has a different crowd.
[00:59:17] And code like us has a different.
[00:59:20] Market. Okay, we are more on the AI development side, AI developer.
[00:59:24] Okay, they are more of a software developer using AI tools. It's like that. So there's a difference.
[00:59:30] Okay, so anyways, so guys, this is what full chat history is.
[00:59:35] Now, there are applications where you will not use full chart history. I will show you variations, ok.
[00:59:39] You, you're.
[00:59:41] Application.
[00:59:44] Might use, you know, last few chats, last 10 chats.
[00:59:48] Okay, in that case, guys, simple slicing concept will help you.
[00:59:53] If you just take your chat history.
[00:59:56] Okay. And just limited only to, let's say, last 15 chats, last 10 chats, last 20 chats.
[01:00:03] Okay, before that, I will just limit it off.
[01:00:06] Okay, usually…
[01:00:08] The kind of chatbots we built.
[01:00:10] Okay, it is usually last 10 chats, last 12 chats, okay?
[01:00:14] But if you're building, obviously, if you are ever, if you ever build an AI agent, an AI chatbot, let's say.
[01:00:21] AI coding tool ever. Okay, I know that's a very mammoth thing.
[01:00:25] Okay, if you ever build it, then probably.
[01:00:29] Windowed chat history might not help, but prompting the user help.
[01:00:32] Okay, prompting the user, hey, we are almost going to run out of memory.
[01:00:37] So, ah.
[01:00:39] Do you want to move, allow to remove few chats? So over there you use the same concept again.
[01:00:44] Okay, cut few things, okay, cut few older chats like that.
[01:00:50] Okay, so anyways, so this is what ask
[01:00:53] Last 10 questions, uh, ask agent last 10.
[01:00:57] Okay, what this does is, instead of…
[01:01:00] Uh, previously your full chat history.
[01:01:02] was an empty thing, and that started and that, uh, then on that we kept on attaching new, new questions.
[01:01:10] Now, you have full chart history is still an empty thing, and the first thing it is starting.
[01:01:14] But it is going to take the last 10 messages. Okay, so if I mention 10 over here, so minus 10 will be the starting point.
[01:01:21] Okay, so you know how minus 10 works guys, you all know right how minus 10 will work, so if you have
[01:01:28] It's a 0, 1, 2, 3.
[01:01:29] 4, 5, 6.
[01:01:32] 7, 8, 9, 10…
[01:01:35] 11… 11…
[01:01:37] Then 12, so minus 10 will be
[01:01:40] Around here, this will be minus 10.
[01:01:43] Because this is minus 1.
[01:01:45] N as 2.
[01:01:56] Yeah, minus 10 will be here.
[01:01:57] So, this will be the starting point. So from here, only the last
[01:02:01] These many charts will be kept. Okay, all these charts will be cut off.
[01:02:04] When you are doing this.
[01:02:06] Okay, so you're starting from minus 10.
[01:02:08] Minus 10 is your starting point. So max matches is minus 10th messages.
[01:02:12] Uh, onwards messages will be kept. So that is what.
[01:02:16] Wind.chat history is.
[01:02:19] Okay, so what we are doing over here.
[01:02:21] is first we take IMT list again like full chat history, we have.
[01:02:25] empty list, instead of the…
[01:02:28] In the function,
[01:02:31] Instead of taking window chat history as it is,
[01:02:34] We are taking last 10 chats, okay? That will become your recent messages.
[01:02:39] Now, recent messages plus same format again, user question.
[01:02:43] Messages agent.invoke will send this message.
[01:02:46] We'll get an answer.
[01:02:48] And whatever answer you get.
[01:02:51] will again increase the chat, right?
[01:02:53] Let's say minus 10 till
[01:02:56] 0, you had the message.
[01:02:58] This is going to increase the amount of chat, right? After you get a response.
[01:03:02] So again, you need to store only minus 10.
[01:03:06] Because the response from the AI.
[01:03:08] is going to make it minus 11, like…
[01:03:10] The limitation will be like, it will explode, right.
[01:03:13] So, that's why we are again reinforcing, we are again taking minus 10 even after generating the answer.
[01:03:19] So even after generating the result message result message will have the entire thing.
[01:03:23] All the 10 chats. Okay, plus the 11th answer.
[01:03:28] Okay, so we are again.
[01:03:30] You know, taking from minus 10 and making sure that the
[01:03:35] Third chat, which was there, 0.
[01:03:37] In our case, 0, 1, 2,
[01:03:41] 1, 2, 3. From here, minus 10 started, right?
[01:03:44] And, like, this it started, right? Now, when it becomes 11, then this is from where minus 10 will start. So, even this will go out.
[01:03:52] After the answer is generated. So that is what we are doing. So we are…
[01:03:56] Basically, taking this Windows chat history,
[01:03:59] From minus 10 again onwards, but again, uh.
[01:04:03] After the answer is generated.
[01:04:05] Okay.
[01:04:10] Yeah, let's do this.
[01:04:11] So, so it means that, uh, when the street increased, so it will, like…
[01:04:16] automatically remove that.
[01:04:20] Adolf. Yes, cut off.
[01:04:21] Yep.
[01:04:22] Uh, yeah. So, in many chatbots, uh, you know, you will… once you build in the industry, you will.
[01:04:27] See that many chatbots.
[01:04:30] your older chats is not valuable anymore.
[01:04:34] Okay, in ManyChat see.
[01:04:38] ChatGPT is a very generic chatbot. In that case, anything can happen, ok.
[01:04:42] Even older chats might also be important.
[01:04:45] Okay.
[01:04:47] But many chatbots that in enterprise we build.
[01:04:52] Where the conversation doesn't goes on forever.
[01:04:55] Let's say I'm building a chatbot.
[01:04:58] Since it is so niche.
[01:05:00] It is so specific chatbot about HR.
[01:05:02] A person who is using it,
[01:05:04] Hardly we'll ask two, three question.
[01:05:07] Maximum follow-up, 4 to 5 questions.
[01:05:10] So that's why this minus 10 is a very…
[01:05:12] Important part, like
[01:05:14] As you can see, like, it doesn't cross minus 10. Many people use minus 5 also.
[01:05:19] What is Aditya?
[01:05:23] Hello.
[01:05:30] Okay, understood, right?
[01:05:33] Yeah, yeah, got it, but I was in…
[01:05:34] Oh, sorry, you're on mute. Uh-huh.
[01:05:37] So basically, yeah, there are chatbots where you will see.
[01:05:41] That the value
[01:05:43] of older chats is not there after 5 or 6 chats. There are chatbots like that.
[01:05:48] Okay, uh…
[01:05:51] Since when we are discussing.
[01:05:54] Probably cloud perplexity and ChatGPT is running in your mind, so you are not able to think that, because those are generally chatbots in the same chat.
[01:06:01] Only you keep on asking questions.
[01:06:03] Okay, you keep on asking a new question, different topic questions.
[01:06:07] It keeps on flowing, even after 10 chats also, it keeps on going.
[01:06:10] Okay, but ah, but when you are going for, you know, chat bot.
[01:06:14] let's say for a healthcare app. Okay, now we recently.
[01:06:18] are signing one healthcare client, okay.
[01:06:21] So over there also they will build some chatbot for a particular clients profile, all the health history. It's about a U.S. app.
[01:06:28] Okay, you were in US healthcare is something serious, okay, you all know it, so there is like 1 database for all your.
[01:06:36] previous healthcare history and everything. So you can ask questions and all those things.
[01:06:41] You, let's say if I build a chatbot there.
[01:06:46] They will not have conversation more than, you know, 10, 15.
[01:06:48] Conversation. After that, it will not be there.
[01:06:51] Okay, they will not ask you question in the healthcare chat, but they will not not ask you question about.
[01:06:55] Ah, you know, HR, okay? They will not ask you generic question. They will ask question only about that.
[01:06:59] Okay, and if they even ask questions about that, there will be guardrails to prevent that, that.
[01:07:04] You know, please ask relevant question like that. There will be guardians which will prevent.
[01:07:08] So the chat will not go for a long.
[01:07:11] Okay, usually. So usually niche chatbots are very limited to
[01:07:15] 10, 15, maybe 20.
[01:07:17] Okay, it's like that. But when you're making a generic chatbot, then the game is something different. Then you need something known as a summary also.
[01:07:26] Okay, so that we will do next. So, over here, as you can see,
[01:07:29] We are taking last 10 chats again.
[01:07:32] And same way we are showing the output.
[01:07:35] Okay, and…
[01:07:37] you can see the output as well. So, you'll probably have to trigger this.
[01:07:41] Last change are cleared.
[01:07:44] Yeah, so as you print this, you'll be able to see if you need
[01:07:48] So when you ask the first question, what is the current weather in Kolkata?
[01:07:52] Okay, the current weather in Kolkata is as follows, so it will give you the information.
[01:07:56] And it also gives you this if you need more information, ask currently message stored is 4.
[01:08:02] Can you tell me why 4 messages stored?
[01:08:04] Guys?
[01:08:29] You'll find it out and tell me why 4 messages wrote. No problem, you'll find it out. Use your
[01:08:34] Capability and finder, while 4 messages stored. Next one, after asking the next question, it is 6 message.
[01:08:39] Then asking the next question, it is 8 message.
[01:08:41] So, it is increasing by 2, but why in the first case 4 messages stored.
[01:08:54] That might be saving the user details, maybe, that users
[01:08:55] And…
[01:08:59] Not so, but I'm just guessing it
[01:09:01] No problem. That is not the right answer, but you…
[01:09:05] Try, try finding it out.
[01:09:08] Maybe tools
[01:09:12] Very good.
[01:09:14] User roles and go away suggestion.
[01:09:17] User AI, AI recommending you the tools.
[01:09:22] Then tools, then AI again.
[01:09:26] Got it?
[01:09:29] Got it guys, so since the weather tool is called, so first user.
[01:09:33] Then the AI telling you to call the weather tool.
[01:09:36] Then the tool, which is called and you get the answer from the tool.
[01:09:40] And then the AI again, which is framing in a good manner like this.
[01:09:45] That's why you are seeing 4 messages for the 1st time.
[01:09:48] Okay. So, then 6, then 8 and
[01:09:52] If you keep on hitting with more questions, you will see that this message will vanish off after a point.
[01:09:58] Okay, after two more questions, it'll start vanishing off.
[01:10:02] Okay, it will not store more than 10.
[01:10:05] Okay, but anyways, the most…
[01:10:07] The most exciting and the most used by.
[01:10:12] Many companies, is this summarization, so all your
[01:10:16] EI agent and coding tools, which you all.
[01:10:19] use very often the Claude code and all these things.
[01:10:23] All they use is this concept of summarization.
[01:10:27] Okay, before me going through it, you all just go through it. I'll just take…
[01:10:32] of 7 minutes break. You all go through it once, okay.
[01:10:35] This is the last point that we'll discuss, and then we'll move on to prompt engineering guys.
[01:10:40] This is the last point that is left for memory. I'll just be back in 7 to 8 minutes, okay? Just go through this and tell me the flow after this. We'll discuss the flow, but you all will have to tell me, because I have.
[01:10:49] told you two chat variations already. Shoot chat history variations. Now, we all will tell me.
[01:10:54] Okay, after we come back.
[01:10:56] Okay, just go through it once.
[01:11:01] You are going to play, play and go through it, okay? Don't take guess and all. Just go through it, what is happening.
[01:11:07] Uh, you know, tell me the flow. This is a bigger flow, this is way bigger, this is not simple like the last one.
[01:11:11] Okay.
[01:18:03] Okay.
[01:18:08] Did you all guys go through it?
[01:18:10] Did you understand what is happening?
[01:18:27] If you are done, can you give a thumbs up? Or if you are still going through it?
[01:19:28] Azat, Chandrashikarth, Deepak.
[01:19:33] Not seeing Deepan.
[01:19:41] Torragesh, yes.
[01:19:44] Okay, done. Can one of you explain what is happening? So why are we using message to text and then.
[01:19:51] What is this?
[01:19:54] uh where we using a prompt over here.
[01:19:57] Okay, what is the need of it? So, then in the ask, the main, the main part where we are managing this.
[01:20:04] over here first we do a check.
[01:20:09] And then…
[01:20:10] We keep the older message.
[01:20:14] Till 10.
[01:20:16] over here.
[01:20:20] And…
[01:20:23] And then from 10 onwards, we keep it in a here, so keep it over here.
[01:20:27] So…
[01:20:29] So why are we creating an older message? What is the use of this, first of all?
[01:20:34] And…
[01:20:38] After that,
[01:20:40] What we are doing is.
[01:20:42] We are
[01:20:45] We are creating a prompt, some system prompt. In the system prompt, we are sending this summary text.
[01:20:52] Okay, summary text that we have.
[01:20:55] And then…
[01:20:57] We are doing this message.extend.
[01:20:59] So why are we doing all of these things? So can you, can one of you tell me?
[01:21:03] The entire flow, what has, what is happening and then again we are doing a check over here.
[01:21:10] Okay, it is written also, like, all of these things are well written as well, like, it's
[01:21:14] It's very self-intuitive. Okay.
[01:21:18] Let's explained me this much, like
[01:21:20] I want you all also to, you know, think in the class a bit more.
[01:21:26] So, in summary, it's pretty much very similar to when we use any LLM tool, right? And when we use any of the chat for a longer time
[01:21:37] Then it shows the message that they are compacting the conversations
[01:21:42] Huh.
[01:21:43] Every tour has a different messaging, but the idea is, rather than sending the whole messages or the chat history that we are having
[01:21:50] Hmm, hmm.
[01:21:52] If you send everything, every time, then it's a problem, right?
[01:21:57] They have to process a lot. So, I think that's the core idea
[01:21:59] Here. Yeah, the core idea is that only. So basically this concept is used by all.
[01:22:05] This who asked me, Neeraj asked me, right, about Claude.
[01:22:09] And, uh…
[01:22:11] Or some agent tool has asked me that.
[01:22:15] After a certain point, it says 70% used already.
[01:22:18] So, during that time, it gives you an option. Do you want to summarize the chat till now? And that is where you use this.
[01:22:24] Even, like, many of them don't even ask this. Like, in a normal chatbots don't ask, like, ChatGPT don't ask you this thing. ChatGPT already summarizes it.
[01:22:32] For you. And they often summarize your important data.
[01:22:35] Yes, Dwaragesh.
[01:22:38] I just had one question, comparing this with the actual way this, how this cloud or ChatGPT works, right? So, so here we are keeping into the window, like, we have a 6- window of 10, but,
[01:23:04] Yes.
[01:23:05] the actual message that we have in the window might be shorter or longer, that depends on the context, right? But there, when we use ChatGPT, it mostly works on a number of tokens, right? So when the token buffer is like 70% is filled, then it automatically kind of summarizes
[01:23:10] Yeah, correct, correct, correct. You are correct about it.
[01:23:14] So over here, we are having messages.
[01:23:16] In that, in the other one, you can convert this into token concept.
[01:23:20] Okay, where you can take like 4 to 5 words as one token kind of a thing, okay?
[01:23:26] Or let's say, uh, six words as one token.
[01:23:29] And the moment you see that.
[01:23:31] That 6 words, how many 6 words you have, let's say.
[01:23:35] Let's say you have 506 words, okay? The moment it crosses that.
[01:23:39] You start summarizing. That also you can do.
[01:23:41] Dwarakesh.
[01:23:44] So for that, do we again make use of the recursive character text splitter to keep
[01:23:48] For recursive character text, by the way, recursive character text splitter is not token based, it is character-based, by the way, that 500 chunk size you write, it is character.
[01:23:56] Oh, got it, yeah, okay, okay, yes, okay.
[01:23:57] Yeah, so either you can use some sort of a tokenizer.
[01:24:03] Okay, dwarakir, ah.
[01:24:04] Okay.
[01:24:06] Which is a more tougher job, like, tougher job as in.
[01:24:09] You have to have the tokenizer downloaded of GPT.
[01:24:13] And you can use the actual tokenizer.
[01:24:15] Okay
[01:24:16] And measured it that way, or you can take an approximation. Usually in the industry, I've seen people take approximation.
[01:24:22] Okay, they take the token.
[01:24:24] It's like size of 4 to 5 words as one token or Rakesh.
[01:24:28] Okay, and the moment they see, like,
[01:24:31] 54 to 5 words crossed, they start summarizing, let's say.
[01:24:36] Because after 500 tokens, you are summarizing.
[01:24:38] Okay.
[01:24:39] Okay, is that one of the reasons that they usually keep, like, 70% to 80% buffer, and they don't go very close to the actual window
[01:24:47] Yeah, they… they don't go, but it depends, some certain tools are even going till 90 as well, but
[01:24:52] When you are using a coding agent, since coding agent, the next token, the next question that you are asking.
[01:24:58] can trigger, can actually use…
[01:25:01] The rest of the 30% token as well.
[01:25:04] Okay.
[01:25:05] What is Dwarrakesh? See, that's why they give you a buffer.
[01:25:07] That you have already done 70%.
[01:25:10] Hmm, okay.
[01:25:11] Got it. So the point I'm trying to make that the next call we will make.
[01:25:15] Since all these coding agents are thinking.
[01:25:18] And they are, they are working in build or plan mode, whatever.
[01:25:23] mode you are working in since they have access to so many coding files.
[01:25:26] So all those things will come to the context, right?
[01:25:30] Okay.
[01:25:31] So maybe your next question might be.
[01:25:32] Analyze all the codes and find out where… why is this error coming.
[01:25:36] So, if you give that kind of a prompt, it can actually go over that 100% limit as well.
[01:25:42] Got it.
[01:25:44] Okay.
[01:25:45] So that is why they give a buffer.
[01:25:46] Okay, but in a normal chatbot, it is not given.
[01:25:49] Because normal chatbots, so much tokens is not used, no? In a coding agent, more tokens are used.
[01:25:56] Yeah.
[01:25:57] Got it. But what really happens, like, like, so you're already at 70%, but your next prompt is actually going to blow up the window. In that case, what will happen? Like
[01:26:06] So… so usually it asks you, you know, that, uh, when it is 70%, it asks you that.
[01:26:12] You know, do you want to summarize older chats? So, that is what we… I have been doing. Like, I just press that.
[01:26:19] Only. But if you don't summarize.
[01:26:22] window, then.
[01:26:24] I, like, as per my understanding that older chats are capped.
[01:26:28] Like they automatically summarize it. Older chats.
[01:26:31] Okay.
[01:26:32] Got it. So after a capping limit?
[01:26:33] Okay, before making the alarm call, it'll automatically make sure that it is within the window, right?
[01:26:38] Sorry, I… when I was telling you, told something, I didn't hear.
[01:26:48] Mmm hmm.
[01:26:49] So my question was, like, like, before we actually make the LLM call, it will make sure that it's within that window. It is not exceeding the
[01:26:51] No, no, no, yeah.
[01:26:53] So these are mostly seen in coding agents. That's the point I'm trying to make, because coding agents.
[01:27:00] uses a lot of tokens. Okay, but the chatbot that you will design.
[01:27:04] Okay, normally will not have so much tokens, usually.
[01:27:08] Okay, usually they will not have chat. If you were making again a generic chatbot, then it's a different issue.
[01:27:13] Okay, let's say if you are making an all-in-one chatbot for your companies,
[01:27:16] All kind of information.
[01:27:18] Okay, then there.
[01:27:21] Could be some generic.
[01:27:23] Uh, you know, chats that is being asked by user and that time.
[01:27:27] One chatbot session can keep on going, keep on going, keep on going.
[01:27:32] But again, these are utility tools inside the company, okay? People will not keep on hitting your chatbot. They will again go back to ChatGPT only.
[01:27:40] They will just ask specific questions.
[01:27:42] Right? You only think, Daurakesh, let's say you have a…
[01:27:44] Yeah.
[01:27:46] You have a chatbot in your company. Let's say you work at
[01:27:49] For me to refer to a
[01:27:51] company, let's say you are working at Atlassian.
[01:27:54] In Atlassian, you are, let's say, building.
[01:27:57] Confluence, confluence in confluence.
[01:28:00] Your department only you work, you make fixes over there, you, you know, you are trying to do some AI over there.
[01:28:06] in that team. And then, there are certain errors, there are certain bugs about confluence.
[01:28:11] Which are again present inside in another Confluence document. So, you are working in Atlassian, in the Confluence team,
[01:28:17] And in that also there are some errors.
[01:28:21] For certain codes and that also there is a Confluence page.
[01:28:24] From there, it is generating some answers for you.
[01:28:27] Will you keep on asking for a long, long time?
[01:28:30] Do you think? Because your goal is to come to the right confluence page and just
[01:28:34] Just visit there. After that, you'll get most of your answer by reading that conference document.
[01:28:40] Right.
[01:28:41] Okay, you'll not keep on hitting, like
[01:28:42] Okay, this is fine now.
[01:28:44] Okay, now do me at Excel also create an Excel. So you will not do a unnecessary task.
[01:28:49] So it is not an open-ended chatbot like ChatGPT where you do everything, you generate image also with chatbot chat.
[01:28:54] You do Excel also, you generate proposal also, document also, code also, everything.
[01:29:00] So that time the chat keeps on growing, growing, growing, but it's a free tool also.
[01:29:06] So, so that is the point I'm trying to make. Usually.
[01:29:10] In this kind of specific kind of chat, we
[01:29:13] You know, limited, uh…
[01:29:16] 10, 15 chats are good, then after that, you can summarize it.
[01:29:20] Okay
[01:29:21] Oh, I need one, just, uh, maybe a little bit of a digression on this. So, this 90% buffer and all you are telling, right?
[01:29:26] There have been, uh, many cases that I have sometimes seen this, okay? Like, it'll show that 90% has been used, another 10% are pending.
[01:29:34] And my next question will kind of consume that 10%, and it say that key next Char punch, let's visit after the next.
[01:29:41] Cooldown window of 4 or 5 hours, okay? I'm talking more from, uh…
[01:29:45] over a period of time that I've been using it.
[01:29:47] So I was just thinking something around it.
[01:29:51] Is there a way, like, if I'm creating something which says that, okay, if I use… if you are using the current model,
[01:29:56] This 90%, whatever 10% buffer is there, may not be that, uh, great, right? Like, go to a lower model,
[01:30:02] Which uses lesser token or…
[01:30:06] I don't know, uh, intelligently, can it do such switches? Is that possible?
[01:30:11] Uh, so, there is a concept known as intent engineering.
[01:30:18] Okay.
[01:30:19] Okay. So there is a concept known as internet engineering that happens inside your model.
[01:30:23] Okay, which basically detects?
[01:30:24] Uh-huh.
[01:30:25] uh, what the user wants.
[01:30:26] Okay.
[01:30:27] Okay, there are two concepts. When you write, when you do prompt engineering.
[01:30:31] Uh huh.
[01:30:32] I mean, so you, as a user, you do prompt engineering. Inside the model, there are two things that happens. One is intent engineering.
[01:30:39] One is context engineering.
[01:30:40] Okay. Okay.
[01:30:41] Engineering is what is to be done.
[01:30:44] Okay, uh-huh.
[01:30:45] Uh, what the user wants, and the next context engineering is how to get it done.
[01:30:49] Okay.
[01:30:50] What are the tools I can use to get it done?
[01:30:52] Okay.
[01:30:53] Every LLM.
[01:30:55] Inside every LLM architecture has this kind of an architecture, ok.
[01:30:58] Okay.
[01:30:59] So every LLM call that you make will have this kind of an architecture, so
[01:31:03] When the intent
[01:31:04] Maha.
[01:31:05] In the intent, if you, through guardrails or through your prompt, if you can design.
[01:31:09] Uh-huh.
[01:31:10] A system where even if your model is asking, let's say HTML related question, which will unnecessarily cause.
[01:31:16] A lot of token to be used if I start using Claude.
[01:31:18] Okay.
[01:31:19] So, I will write in the intents LLM prompts only that.
[01:31:24] When you are detecting an intent, if the intent is about user creating a HTML/CSS kind of a page, which has lots of lines written, usually.
[01:31:31] HTML is a lot of things written.
[01:31:34] HTML lines are way longer.
[01:31:35] Mm-hmm. Correct.
[01:31:36] So you can
[01:31:38] That time, do a model routing, model switching.
[01:31:42] Okay, so internally, like if the intent is about
[01:31:46] Uh, HTML, go to Gemini Pro, uh, Gemini Flash.
[01:31:47] Okay. Okay.
[01:31:50] Because, no, not Gemini Flash, Gemini Pro.
[01:31:53] Okay, Gemini Flash is good for basic question answering, like, how is the weather?
[01:31:57] How is this, what is HTML? This kind of question Gemini flash.
[01:32:02] From that intent engineering of the user,
[01:32:05] You first detect.
[01:32:07] Whether the user wants a simple thing.
[01:32:10] Medium thing or tough thing?
[01:32:12] Okay, let's say you have an agent.
[01:32:14] Okay, that agent has access to
[01:32:15] Okay. Okay.
[01:32:17] Claude only, let's say Claude is the LLM, Claude.
[01:32:21] Using that LLM, you have given a big prompt that your job is to define.
[01:32:23] Uh-huh.
[01:32:25] Whenever the user comes a question, you are the first agent.
[01:32:28] And your job is to define what is the intent. Is it simple intent? I'm just making it very simple for you. Simple intent.
[01:32:32] Thomas gay huh
[01:32:33] Easy. Simple, medium, or hard. If it is hard.
[01:32:36] Okay.
[01:32:37] Then your model selection, so the, your tool, your tool response, like you are a tool kind of a thing, you are.
[01:32:44] Or your response should be use Claude Opus.
[01:32:46] If you are, if your answer is medium.
[01:32:50] Then use all the Gemini Pro versions.
[01:32:53] Okay. If your answer is…
[01:32:54] Okay.
[01:32:55] Easy. If it is an easy question, like, what is, what is basic question, like, what is this, what is that, then use Gemini flash.
[01:33:02] So this, I internally, I usually do in my…
[01:33:06] when I'm using anti-gravity or clot.
[01:33:08] Uh, well.
[01:33:10] I just, whenever I'm asking basic question, no.
[01:33:13] I use Gemini Flash.
[01:33:14] Okay. Okay.
[01:33:15] Okay. And whenever asking complex question, I go to Claude Opus.
[01:33:20] Okay.
[01:33:21] No, no, understood. So, you are doing this from your end, right? Like, Madlab, you are selecting it. Ha, so what…
[01:33:24] I'm doing manually. You can do this through an intent. You intended identification.
[01:33:30] Okay, so this can happen only when I am building a
[01:33:34] Application of my own.
[01:33:37] Or can it be done in the existing Clause or our ChatGPTs as well?
[01:33:41] Oh, yeah, they have auto options.
[01:33:43] Auto option, is it?
[01:33:44] The yeah auto option, auto model selections is there. Cursor provides auto model selections.
[01:33:49] Anti-gravity, ah, I have always choose chosen a model and done it.
[01:33:54] Okay.
[01:33:55] Uh, wait, wait, wait. Let me see auto is there. On anti-gravity, there is no auto.
[01:33:59] I am seeing clots on it.
[01:34:00] Welcome.
[01:34:01] Yeah, in anti-gravity, I'm not seeing, so see if you can see my anti-gravity screen just a moment.
[01:34:12] Yeah.
[01:34:13] Could assist having that auto mode, so am I thinking anti-gravity also will have
[01:34:19] No, I am seeing, like…
[01:34:23] Man also.
[01:34:24] options, manual selection only.
[01:34:28] And that's HLA
[01:34:29] But we had this…
[01:34:30] Germany CLI is having an option of an auto mode
[01:34:31] Ortonia, right?
[01:34:32] Like
[01:34:34] Yeah, it could be there, over there. Again, maybe if I sign in again, it will come, because recently I have recharged my
[01:34:42] geo. So, with Geo, you get 18 months mini pro-free.
[01:34:45] Do you have any free, yeah, exactly.
[01:34:47] Yeah, I had told you.
[01:34:50] I had told you this thing already, uh, maybe in the 1st class.
[01:34:54] Yes, and 5TB of data for us.
[01:34:58] Yeah, yeah, yeah.
[01:35:00] Okay.
[01:35:01] So, wait, just a moment, let's see if my geo is recharged, if this thing has come up or not.
[01:35:09] Maybe with authentication it will come back, but
[01:35:11] Let me see once. In clawed cursor it is there I know for sure.
[01:35:16] Yes, yes. Yeah, we use that.
[01:35:17] It is there. No, actually.
[01:35:19] Yeah, I…
[01:35:20] Yeah, we can do the ask plan and coding always. Yes
[01:35:24] Yeah, so over here, actually, there is model selection manually.
[01:35:29] Okay, maybe in a different option.
[01:35:32] We'll have to, but always I have done manually only. Like, if I'm…
[01:35:36] generating, you know.
[01:35:38] HTML, this thing I always switch to Pro. And if I am asking like.
[01:35:44] What is this? Explain me this. Let's say, explain me skills.
[01:35:48] Okay, then I will go to Flash.
[01:35:50] Take that.
[01:35:53] Okay, so what initially when you asked this question, Vineet, what I thought is you are asking like, how do we design it?
[01:36:00] So, that's why I told you, like,
[01:36:01] Both, but, both. So, I'm asking it in this way, right? One is,
[01:36:07] Yeah, so designing you got, you understood, right?
[01:36:08] How do I design it if I'm doing something simultaneously in the existing one? Is there, I got it. Yeah, yeah.
[01:36:10] So, yeah, every, every tool.
[01:36:15] Yeah, every tool will have…
[01:36:21] Of, uh, intent identification.
[01:36:22] and context engineering.
[01:36:25] Okay.
[01:36:26] So in your intent identification, first you identify the intent based on that, you provide that context. And your context could be.
[01:36:31] tools, all these things, so.
[01:36:32] Over here also it is kind of doing internet identification only first.
[01:36:37] Your intent identifying the intent. That's why it is calling the weather tool.
[01:36:40] So, in that case, we will call a model.
[01:36:41] Correct.
[01:36:42] change tool kind of a thing.
[01:36:44] Now, in cursor, it is already there, but I don't have the paid cursor, so…
[01:36:50] I anyways do not have a model option, model selection option, I'm just seeing.
[01:36:53] And it comes as auto only.
[01:36:56] whatever model is given over there, that only option I get usually. Let me see.
[01:37:04] There is water.
[01:37:05] Okay, there was one more thing that I wanted to ask, so, uh, I was just…
[01:37:08] Yeah, I'll just share, I'll just share.
[01:37:10] Yeah, you keep on asking, Vinit.
[01:37:12] Yeah, so there was one more thing that I was just exploring the other day, where it says that, you know, whenever using the, uh, agents that are
[01:37:18] Provided by this, uh, at least with regards to Claude.
[01:37:22] It allows me to select
[01:37:24] 2 separate thinking minds, right? Like, it says that I can have a
[01:37:28] Combination of, say, uh, Haiku and, uh, Opus.
[01:37:33] Okay, so it is…
[01:37:34] So, the
[01:37:35] Ah, so what is that concept like?
[01:37:36] Yeah, so there is a concept of sub-agents that has come up.
[01:37:41] Okay, where in Cloud Code it is there. I don't know if anti-gravity is also included that or not nowadays.
[01:37:47] So, sub-agents, what they do, by the way, guys, this is a little different topic, like we are just discussing since we have come towards summarization, we have come… we are discussing, but this has nothing to do with.
[01:37:58] today's topic, just few people are raising the doubt, that's why I'm telling, okay? So, my point is.
[01:38:03] Subagent concept is, let's say for a particular task.
[01:38:07] Instead of using all the agents as Claudopus and causing a lot of lot of token.
[01:38:14] What if I can have multiple sub-agents doing different, different tasks?
[01:38:18] Two subagents will be Haikus.
[01:38:21] Two haikus and one could be, let's say, a Sonnet model.
[01:38:24] Got it?
[01:38:26] Got it. Let's say.
[01:38:29] Let's say we need to achieve one task, I'm giving a very basic example.
[01:38:33] You want to create a UI, you want to create.
[01:38:37] a backend, and you want to create a middle layer.
[01:38:42] Okay, middleware. So maybe the middleware and the backend will be done by Sonnet and Haiku will do the HTML page it like that. It's two sub-agents. So this is the concept of sub-agents.
[01:38:51] Where each agents take different, different models while coming up with an answer.
[01:38:56] Okay, so this sub-agents concept has come up in Glot Code now.
[01:38:57] Yes, yes.
[01:39:01] Okay, got it, Aniruban.
[01:39:03] Yeah, yeah.
[01:39:05] Okay, so over here,
[01:39:11] Uh-huh, just now I had…
[01:39:13] option over here.
[01:39:14] And, uh, over here just now I had auto option, ok, it just now refreshed and it went away.
[01:39:20] Okay, so…
[01:39:23] If you, if you, uh, if you are doing it in your thing.
[01:39:27] Let me see if I do this, it comes or not.
[01:39:34] Just now over here, I had a auto option.
[01:39:37] Okay, so if anybody is using paid cursor, you can see it.
[01:39:42] That, uh, over here it is using Composer 2.5 fast. Over there.
[01:39:45] Just note this entire list opened.
[01:39:47] Okay, and over there I saw auto option. So if you select that auto option, it automatically chooses a model. Even
[01:39:52] When I use Amazon Q in my in one of our clients and environment, I see auto option.
[01:39:59] Nowadays, I have removed the auto option, and I am forcing the usage of Claude saw.
[01:40:02] Opus.
[01:40:05] But anti-gravity is not showing auto option. Anti-gravity, since I have it and over there only I'm not seeing the auto option.
[01:40:18] It is always referring me to a model, maybe in the settings change somewhere will be there.
[01:40:23] But yeah, but anyways, that is how I use it, guys. You know, sometime I don't even use these things, you know, sometimes.
[01:40:32] For example, you'll know the concept of skills, you'll know it, right?
[01:40:37] Yes, Mini
[01:40:39] Yeah, maybe you'll, okay.
[01:40:40] Indeed bank and skillset was the same thing. Like, future uses the memory bank and Casal uses the skill to save the context of the chat we are doing, or some
[01:40:50] Let's including a drink, right? So, like, you want to set that context, or some finding with you, right? Some planning we did there, so that you can save as a skill or the memory bank. That, like, once we go, like, doing some code changes for the same
[01:41:04] model you can say. So, like, the we go and first check from the skill and the memory bank, then it will go and, like, set the code
[01:41:10] Yeah, yeah, so skills is something on demand, right? Instruction on demand.
[01:41:13] Yes, yes.
[01:41:14] Okay, so many people create skills. You can get.
[01:41:19] Anthropic Skills, do you know that? Anthropic.
[01:41:23] Skills. There is a page, do you know that?
[01:41:26] Yes, yes.
[01:41:27] Public repository for Anthropic skills, so over here you will have lots of
[01:41:33] Skills and you can.
[01:41:35] Visit. But anyways, you can create skills.
[01:41:37] Okay, uh, see, there are different skills. One is for algorithmic art, one is for canvas design, one is for Claude API.
[01:41:44] PPD, PPDS skills. So, when you create a PPD, have you.
[01:41:49] I don't know if you have created a PPT on Claude or not.
[01:41:51] Uh, when you create a PPD, you will see this name.
[01:41:54] skills dot going to skills.md, then pptx.
[01:41:58] It will show you. So this is that page. It goes to this repository, and in this repository.
[01:42:03] It will be written, like, use this skill anytime to time.pptx or .potx is involved in any way.
[01:42:10] Okay, so skills are basically instruction. It's a MD file.
[01:42:14] Where you have these prompts kinds of instruction.
[01:42:17] And since we'll talk about prompt engineering, you'll be able to delete to this.
[01:42:22] Okay, so where you have these kinds of prompts, which…
[01:42:25] will be used on demand. Let's see, want to create.
[01:42:27] A PPD, then this skill will be used, because this has all the
[01:42:31] Things required to create a PPD. Okay.
[01:42:33] This might internally call more Pythons as well over here, that's why you will see internal lots of Python things as well.
[01:42:39] So, this will call, but it will do the things for you. So these are like instructions that will be played on demand.
[01:42:45] Based on your requirement, so.
[01:42:47] You can create these kind of skills as well for your work, okay, for the project that you are building.
[01:42:52] So these skills, many people use the coding tool only to define the skills.
[01:42:56] I would say, why are you wasting prompt over there?
[01:42:58] Go to Claude Free. Free Claude.
[01:43:01] Okay, and creator skills over there only.
[01:43:03] And download that skills and then put it on that folder. So many time you
[01:43:08] You know, I…
[01:43:10] That's why I told you, like, I have
[01:43:12] very less use the auto version.
[01:43:15] I mostly used like this way, like manually.
[01:43:18] Going to Claude, getting the skills downloaded, because that will save me a cost, no.
[01:43:22] It will not use my memory, chat, token, anything. My subscription will be reduced.
[01:43:27] Because that is for free.
[01:43:29] Because skills usually uses a lot of tokens.
[01:43:32] If I want to create a big file like this, it is going to be a lot of tokens. So, create it for free.
[01:43:37] From Claude, download it and put it over there. I can just watch a video how to put the skills file in store.
[01:43:43] into the skills folder. That will do it.
[01:43:46] So anyways, this is like how you optimize this usage of these tools. Always don't use auto.
[01:43:51] Okay, sometime you change this thing. See, if we get it from client and client gives us a
[01:43:56] freedom, like, you can do whatever you want. That time I used as it is. I don't think about all these things, I use the auto only.
[01:44:02] Okay, I will use that only to generate skills, but let's say if I'm using out of my own pocket, then I will do these things.
[01:44:09] Okay. Okay, now guys coming to the summarization, let's discuss the summarization. So guys, so first of all tell me why am I doing message to text?
[01:44:19] You all have understood, no. Hazat, did you understand what is message to text?
[01:44:22] What are we doing?
[01:44:26] Yeah, actually, like
[01:44:28] Others kind of health requirements, actually, like, that we use to, like, save the message context, like, message object we can say, okay? And
[01:44:38] You can see that rule content and all in the same format, you can save it.
[01:44:43] Yeah, so basically, all the Langchain.
[01:44:46] Memory message objects, they are in a form of type, they will be either in a form of type, or they'll be either in a form of role.
[01:44:53] And they will have some content with it.
[01:44:56] Okay, so in…
[01:44:59] And they will have a lot of metadata, additional, all those things. So in order to get rid of all those things,
[01:45:04] and just store the role and the message, role and the message.
[01:45:08] Okay, that is why we are using this get attribute, this get ATTR like this.
[01:45:15] If you do this, it will give you the actual content.
[01:45:19] Which is saved as content, and it will give you
[01:45:23] The type, if the type, sometimes type is as human, like, instead of role, sometimes it is type also.
[01:45:28] Certain things will have it stored as type, certain things will have it stored as role.
[01:45:32] So that's what we are doing is over here we are cleaning everything.
[01:45:35] other than role and content. Later on, the role and content will just be saved like this. Let's say the role is user, and the
[01:45:42] actual value. So instead of storing it as
[01:45:47] Instead of making it stored as role,
[01:45:49] Then followed by user.
[01:45:51] Then followed by content.
[01:45:54] Then followed by the actual content, then additional.
[01:45:59] Additional…
[01:46:01] Keywords.
[01:46:03] Keyword arguments.
[01:46:05] Then you have metadata.
[01:46:09] Then you will have some ID.
[01:46:11] You know, for a particular message.
[01:46:13] Instead of storing everything, we are just cleaning out
[01:46:17] All these things.
[01:46:18] And at the end, we are just returning in a form of user and the content. User.
[01:46:25] And whatever the user said, this message.
[01:46:28] Then AI, whatever.
[01:46:30] The AI said.
[01:46:33] Like this. Then tools.
[01:46:36] Whatever is the tool's message, like that.
[01:46:39] This is what message to text is doing, because
[01:46:42] Because we will be sending a prompt to the LLM, and that prompt.
[01:46:47] These unnecessary metadata and all these things is going to weigh in a lot of into your context.
[01:46:52] So that is why we are doing this. Okay, so we are keeping this message to context, to make it a very clean kind of a readable text.
[01:47:00] Okay, instead of storing like this…
[01:47:01] See, just 4 message, so much of token usage. If we send everything.
[01:47:08] Okay, so first we do that. So we have a function to do that.
[01:47:13] And then we have conversation, update conversation history.
[01:47:15] So what we are doing is we are creating a summary text. This summary text is basically going to create the summary.
[01:47:21] Previous summary, sorry, summary and what we are doing is
[01:47:26] Whatever is the message to text has done, we are
[01:47:29] Clubbing them together, joining them via backslash n.
[01:47:32] and creating a conversation text over here.
[01:47:34] So all your messages.
[01:47:37] User message backslash N.
[01:47:40] User AI message backslash n like that.
[01:47:43] All the messages will be stored in conversation text, and we are giving that as input to the
[01:47:47] prompt. So, you maintain a conversation history.
[01:47:51] Followed by the existing summary till now.
[01:47:52] You might have no summary also.
[01:47:55] Followed by existing summary till now.
[01:47:57] And the conversation text is
[01:48:00] The text that you have joined like this.
[01:48:02] Create an updated summary. So, the moment you add a new text.
[01:48:05] So, let's say the summary is already created.
[01:48:09] But you might have some new text coming in.
[01:48:10] So that's message to summarize. We'll have that new text.
[01:48:15] So, existing summary plus the new text.
[01:48:19] Add it together, an updated summary will be created. So this is for that.
[01:48:22] Updating your summary, continuously updating your summary.
[01:48:25] Okay. And…
[01:48:28] You are invoking a LLM to do that summary. So summary text is the invoked LLM summary.
[01:48:33] that you have. Okay, now comes to the next part summary chat history and summary text.
[01:48:39] This is where the main…
[01:48:42] Elephant in the room, that is, summary chat history. Summary chat history is basically
[01:48:46] At any point, the last
[01:48:49] 10 chats. Okay.
[01:48:51] at any point, like, that is the goal, like, to keep the last 10 chats, because the window is 10 over here.
[01:48:56] Anything greater than 10.
[01:48:59] When the summary window becomes greater than 10,
[01:49:02] Okay, so anything before that.
[01:49:05] Anything till 10, anything till 10.
[01:49:09] Okay, will be part of the older message. So let's say if you have 15 chats, for example, if you have 15 chats,
[01:49:14] Okay. From 15th till 10.
[01:49:18] We'll go to older message.
[01:49:20] Okay. And then from 10 onwards you will keep it in your summary chat history.
[01:49:26] Okay, so from 10 onwards, you are keeping it in your summary chart history. Everything goes to your older message.
[01:49:31] So, then you are just calling this update conversation summary, it will go to update conversation summary, and it will
[01:49:37] keep summarizing. Now, again,
[01:49:39] After the 10 chat, you will
[01:49:42] Probably generate another answer from the AI.
[01:49:45] And that is going to add up more things.
[01:49:47] Okay, so again, your 10th chat will become your 11th chat now.
[01:49:50] So that extra one chat will again be sent. Next time when you are again triggering the agent,
[01:49:56] The 11 chat is again going to go and generate the summary like that, it will continuously generate updated summary, updated summary, updated summary.
[01:50:06] So this is a check that we do to ensure that, and this check is done once before you hit the LLM.
[01:50:12] And once after you hit that.
[01:50:14] Because after you hit the LLM, what happens? Why do you need to do after you hit the LLM?
[01:50:20] We get the response, right, so to save that response might be
[01:50:22] Yes. So, when you get the response again, your chat increases.
[01:50:27] More than 10. So that's why you do it once after the LLM is hit also.
[01:50:32] If the summary text is there, then you send that as a system prompt.
[01:50:35] to your… this thing, conversation summary from earlier messages, and you send the system prompt.
[01:50:39] Okay. And you put this, this is like a temporary thing, okay? Usually we delete it after this.
[01:50:45] Okay. And in the message till now, which was empty list, you just extend the summary chat window, all the 10 chats.
[01:50:52] And the new question you attach, so this is the 11 chat, guys.
[01:50:56] 10 chat is becoming the 11 chat. Or let's say 9 chat is becoming 10 chat. Okay, let's say even it is 9 chat.
[01:51:02] It is becoming 10 chat.
[01:51:04] You are invoking the LLM. So all your nine chats plus 10 chat is going along with that, your summary is also going.
[01:51:11] Okay, the summary is going now.
[01:51:14] This is going to generate another response.
[01:51:16] And your chat is going to increase.
[01:51:18] So you would remove the temporary
[01:51:20] System prompt first, you remove it.
[01:51:23] First of all, and
[01:51:26] Again, you do perform this. Summarize immediately after.
[01:51:30] This if this interaction pushed the history over the limit.
[01:51:34] So this interaction will push the history over the limit, right? So, you need to summarize… you need to again do this.
[01:51:40] Okay, so…
[01:51:42] If the summary text exists, and if there is a return message.
[01:51:46] Then return message should be everything other than the
[01:51:49] Prompt system prompt.
[01:51:53] So that's why you're taking from 1 onwards, because this one is hidden.
[01:51:56] I removed off, ok.
[01:51:58] Because again, you are going to use it, no, next time.
[01:52:02] Apart from that, whatever the user
[01:52:05] Whatever the AI generated the answer, that is going to
[01:52:07] you know, push the limit over.
[01:52:09] So that is why you, you know, again, you should hit the LLM.
[01:52:14] So that is the overall thing about summarization. Okay.
[01:52:18] Now, guys, see, this is the most standard structure of summarizing.
[01:52:22] Now, there are different ways as well, like.
[01:52:27] uh, you can, you know.
[01:52:29] uh, maybe do it in a more different way, you can, instead of doing, like, like Dwarrakesh was discussing.
[01:52:35] You can do it character-based as well. You can do it token-based as well. You can keep approximation of 4 to 5 tokens.
[01:52:41] As your token size, and do it. And why am I telling, guys, approximation of token size? Because.
[01:52:47] See guys, do you know that token size is very.
[01:52:53] It is not agnostic to LLMs, it is very specific to LLM. Do you know that?
[01:52:57] Do you know that every LLM will have a different token calculation?
[01:53:01] Do you know that?
[01:53:02] Yes.
[01:53:03] Yes, yes.
[01:53:05] Yes.
[01:53:09] Yes.
[01:53:10] And that's what switching from one token to one LLM to another is always a panic. Sometimes the results getting changed.
[01:53:11] Completely.
[01:53:12] Yes, because every token mapping of everything is different. So if you ever go to tick tokenizer.
[01:53:20] If you ever go to tick tokenizer and let's say, if you ever go here,
[01:53:24] So there is an option over here.
[01:53:26] where you can check like what are like, let's say you write I love.
[01:53:30] Football.
[01:53:32] So these are the tokens. But if you change the models to a completely different model.
[01:53:37] Okay.
[01:53:40] This, see, within I love only 3 tokens got hit.
[01:53:43] Okay, maybe because of I love space, because I have given I love and space, so that's why.
[01:53:48] But this is…
[01:53:50] Complete, this will be different from different, different tokens.
[01:53:53] different, different LLMs.
[01:53:55] a 5733 at the end.
[01:54:08] See, it was 5733, I love football.
[01:54:11] Now, it is 1, 3, 3, 2.
[01:54:14] So, every token, the way it is calculated is completely different.
[01:54:19] So, that's why there is approximation, how much this will be. So, this is 3 token.
[01:54:23] Uh, the moment you go to ilafootball dot, it becomes 4 token. So, that's why keep an approximation while doing the token calculation. If you do it… if you do this token-based.
[01:54:32] Okay, let's say instead of doing message-based, you can do this token-based. Usually, in coding tools, it is token based.
[01:54:38] Okay, in the chatbot that we designed, usually it is message-based.
[01:54:41] Okay, but it depends. If you are building something critical, very critical, where
[01:54:46] You know, lots of tokens are used usually.
[01:54:50] Uh, then you can again use token-based Azure, but it depends on LLM to LLM.
[01:54:53] So, if you…
[01:54:55] If the best is if you can host.
[01:54:58] the LLM's tokenizer.
[01:55:00] And use that tokenizer to like this is a tokenizer that is actually measuring how much token.
[01:55:06] So if you can use that particular LLMs,
[01:55:09] Tokenizer to found out how much token is this actually.
[01:55:14] till now, and then go, then this will be more accurate.
[01:55:16] Okay, or else you can take approximation.
[01:55:19] Okay, like 4 to 5 words is one token.
[01:55:21] Our 4 to 6 was 1 token. So, you can take a random number, let's say sometimes you take 6 word as 1 token, sometime you take
[01:55:28] You know.
[01:55:31] For, sorry, 42.
[01:55:33] What did I say?
[01:55:35] 4 to 5 characters, yeah. So if I… if I just, uh, you know, sometimes you can take a random number, let's say.
[01:55:43] Sometimes 6 characters is 1 token, sometimes 4 characters is 1 row, like that.
[01:55:47] Okay.
[01:55:49] So that is it guys, that is like if you want to go with token-based. So, anyways, this is what I have wanted to discuss about memory variations.
[01:55:56] Uh, okay, so this is an entire chat of 4, 5, you know, messages, so you can see how the message is going.
[01:56:03] So we asked about current weather in Kolkata first and then.
[01:56:07] It gives the weather, so.
[01:56:09] First, obviously, the tool call is done, whether tool call, and then
[01:56:13] Uh, this message is generated and then
[01:56:18] See, recently stored messages 4 again, same concept again, because weather tool is called and
[01:56:23] AI responds with weather.
[01:56:26] Then again, what is the humidity in Kolkata? It gives 61%. This is a follow-up question because it got 61 over here.
[01:56:32] And then I use the weather tool to
[01:56:35] Obtain this information, summary available false. Summary available becomes true when you.
[01:56:40] Start hitting this. See, this is the fifth question.
[01:56:42] Will I get drenched if I go outside?
[01:56:44] Somebody available is true.
[01:56:47] And summary update is summary is, the user asked about whether current weather in Kolkata.
[01:56:53] So, this is what?
[01:56:55] Okay guys, understood guys, very much clear about history, memory.
[01:57:01] All these things, guys, I don't think you all, like, everybody had… see, one kind of idea you have when you use these coding agents.
[01:57:08] But how do you build it? Do you get that idea now?
[01:57:16] Dorakesh, what about you? You understood?
[01:57:23] Great. So, uh, guys, after this, we will move on to prompt engineering.
[01:57:28] Uh, so…
[01:57:31] Very simple, prompt engineering is very simple, like these, these are techniques that we do at our industry level, so
[01:57:38] Just to give you all an idea.
[01:57:43] And then, one question. So whatever the superization that you did is, like, we are doing ourselves, right? So, you also mentioned LangChain also supports it, right? So, it also has
[01:57:54] Yeah, yeah, Langchen also I have a code for that.
[01:57:58] Okay.
[01:57:59] Okay, LangChain also supports that. That I will share. You can read it.
[01:58:01] Yourself also, but LangChain one is.
[01:58:05] Like, it is… it will move you out of Python a little bit more, because Langchain has its own syntax.
[01:58:12] Okay, so if you want, I can…
[01:58:13] But what I wanted to understand is, like, LangChain also does it in the same way, like, however you expect, or it's, like, the token.
[01:58:17] Same, just little, little extra lines of, uh, syntaxes of lang chains.
[01:58:24] Okay.
[01:58:25] Now, over here, we are easily managing through lists and dictionaries.
[01:58:27] Sorry.
[01:58:28] Okay. And as a Python developer, we all understand list.
[01:58:32] More than this, LangChain's new frameworks.
[01:58:34] So, so it becomes very simple with…
[01:58:38] You know, us building it than Langchain. LangChain memory also I have.
[01:58:42] Okay.
[01:58:43] I'll share that also.
[01:58:45] Sure, thanks.
[01:58:48] Yeah, so that I'll share for you, since you all are asking, like,
[01:58:52] Uh, I'll share it for you.
[01:59:00] Um, just a moment…
[01:59:03] Yeah, this is the LangChain one I have. Same thing, like again, summary.
[01:59:09] Okay, I am checking over here if…
[01:59:11] You know, uh, so, yeah, so LangChain does this using this.
[01:59:16] So, there is something known as in Langchen, known as chat message history.
[01:59:19] Okay, it is a class of LangChain.
[01:59:21] Using that, by inheriting that, you can create another function.
[01:59:27] Okay, known as summary memory, and…
[01:59:29] All your calculation will be inside summary memory.
[01:59:32] Okay, and you might have to call some.
[01:59:35] You know, uh, parent class init function.
[01:59:38] Parent class, add message function. When you are, you are adding a new message, no.
[01:59:42] You'll have to call this and a new message gets added.
[01:59:45] Okay, and this is how you are invoking the LLM.
[01:59:48] So you are getting it, like, it becomes little LangChain focused.
[01:59:54] Whoever asked this question, if, is it Dwarakesh who asked?
[01:59:57] Yeah, yeah, yes.
[01:59:58] Yeah, so, so, uh, same thing we have, we have the full memory.
[02:00:02] Okay, this is the full memory. Again, full memory is this.
[02:00:05] Okay, uh, this is the full memory concept we have, so.
[02:00:10] You build a concept, you ask three more questions, three, four, five questions, LLM gives you an answer.
[02:00:15] Here's how it is saved. This is the full memory. This is the windowed memory, okay, where you mentioned in a K how much window you want. Let's say window is one.
[02:00:22] Windows 10. So, all your operation happens. So, in LangChain, there is something known as chat message history. Inside chat message history.
[02:00:28] You have window memory, you have summary memory.
[02:00:31] And you have…
[02:00:33] Full memory is not there, anything other than that is full memory only, so you don't need to do full memory, so.
[02:00:41] Yeah, you have window memory, you have summary memory, and you have token memory.
[02:00:44] So, in this case, you just check like how many words are there.
[02:00:48] This is not actually token, this is just checking how many words are there. It is not doing that approximation of 4 to 5.
[02:00:54] Characters for 1 token.
[02:00:55] It is just checking how many words are there. If that word is crossed.
[02:01:00] Okay, the moment see that the sum of the length of the content dot split, so if you split.
[02:01:04] Everything, then you will get the words. So if the moment it crosses 30, you summarize, you basically.
[02:01:10] You know, limit it. Okay.
[02:01:13] So this is using a simple approximation.
[02:01:15] Okay, simple approximation is the number of words is number of tokens. That's how.
[02:01:19] So this is there, this is in LangChain.
[02:01:21] Okay.
[02:01:23] Same thing. Okay. But…
[02:01:26] Okay.
[02:01:27] Way more simple is this. This is the most simple one. LangChain is unnecessary including the concept of chat message history, which is their function, which is their class.
[02:01:35] You call that to implement all these things. Got it, Daurakesh.
[02:01:39] Yeah, yeah, got it.
[02:01:40] Yeah.
[02:01:43] So, unnecessary, a lot of supernate to call the parent class function, that's why.
[02:01:50] You know, I have seen like, you know, people also find it little easy to understand the Pythonic way.
[02:01:56] So that's why I had created a Pythonic way. So this also I'm sharing as an extra file.
[02:02:00] Okay.
[02:02:04] This is additional.
[02:02:08] Okay. Anyways, guys, uh, so tomorrow we'll do the prompt engineering part.
[02:02:15] And once…
[02:02:21] Additional for reading is this.
[02:02:23] Okay. Everything additional for reading, if I add anything, I will refer it to as AFR from here onwards. Okay.
[02:02:31] So, this is what you have, okay?
[02:02:35] Okay, so guys, thank you guys. See you all tomorrow.
[02:02:39] Uh, and, uh, let's discuss on the prompt engineering part as well.
[02:02:44] Okay?
[02:02:47] Okay. How did your… this thing go? Is it today or tomorrow?
[02:02:50] It was done, it was done, Vineet.
[02:02:53] It's finally, like, it was a little generic.
[02:02:55] Okay.
[02:02:56] It was a space defense specific.
[02:03:01] Okay.
[02:03:02] It was only people coming from all ends, like e-commerce people were there and every company is AI powered.
[02:03:04] Okay.
[02:03:05] So AI-powered digital marketer, AI powered this, AI powered that.
[02:03:08] Okay, so…
[02:03:10] Like, I found two, three family offices, I'd found.
[02:03:13] One drone startup as well.
[02:03:14] Mute. I was in mute the time.
[02:03:16] But I couldn't find any space specific people.
[02:03:17] Mm-hmm. Okay.
[02:03:21] Okay, so it was very generic. I found an ETIOSCA president as well, so that is also good.
[02:03:27] For government grants and scheme.
[02:03:28] Okay. Okay.
[02:03:29] Okay. Let's see how it happens. Next week we are planning some connects with these people, so…
[02:03:32] It was a good networking event, I would say.
[02:03:35] But nothing speaks to the idea.
[02:03:36] Okay, got it.
[02:03:38] Yeah.
[02:03:39] Bye bye have a great weekend.
[02:03:41] Okay. Thank you guys, bye.