# 01 2025-11-08 Introduction to AI course: Module 1 — Foundations of AI & ML module: Module-1-Foundations-AI-ML date: 2025-11-08 type: transcript video_url: https://personal-learn.armco.dev/files/_Recordings/Module-1-Foundations-AI-ML/01_2025-11-08_Introduction_to_AI.mp4 --- [00:18:45] intelligence, uh, that humans exhibit. [00:18:48] Uh, with focus on how, uh, human nervous system [00:18:53] represents and processes. [00:18:55] information and transforms it. [00:18:59] And, uh, AI is actually, uh… [00:19:01] a part of cognitive science, generally. [00:19:05] If we talk about AI, uh… [00:19:08] There can be so many ways to describe what AI is all about. [00:19:13] It could be about making computers to think, [00:19:17] or automating some activities, uh, which are associated with human and human thinking. [00:19:24] Like, decision-making and so on and so forth. [00:19:27] Or it could be creating machines that perform [00:19:30] functions that require human-like intelligence. [00:19:35] Or we can think of it as a field that tries to emulate human intelligence. [00:19:41] And, uh, automation. [00:19:44] So, finally, we could actually define artificial and intelligence as two words. [00:19:50] Artificial, which is something that has been created. [00:19:53] Uh, by, uh, some explicit human effort. [00:19:56] It's not, uh, occurring naturally. [00:19:59] And intelligence, of course, uh, describes the ability to acquire knowledge. [00:20:05] And utilize it. [00:20:07] And so, AI actually deals with that part of computer science, that is… [00:20:14] Uh, focuses on design of computer systems that exhibit [00:20:17] human-like intelligence, and… [00:20:19] multiple ways. [00:20:23] So, uh, this means that when we talk about AI, we are primarily talking about two things. [00:20:31] Uh, one, uh, to study the intelligence concerned with the humans. [00:20:35] And then representing it. [00:20:37] with techniques or actions, uh… [00:20:42] that utilize computers. [00:20:46] So, artificial intelligence in itself is also multidisciplinary. [00:20:51] It draws upon different fields, like computer science, statistics, mathematics, sociology. [00:20:57] And, uh, so and so forth. [00:21:02] Uh, if we talk about the relationship between AI, machine learning, [00:21:07] deep learning, NLP, [00:21:09] And, uh, generative AI and agentic AI, this diagram… [00:21:14] uh… shows that. [00:21:16] So, the superset is artificial intelligence. [00:21:20] And within it, we have a subset, which is machine learning. [00:21:24] And in machine learning, we have deep learning. [00:21:27] And then generative AI is a subset of deep learning. [00:21:31] And then agentic AI is actually a subset of generative AI. [00:21:38] NLP is a subset of AI. Uh, it draws upon concepts from machine learning, deep learning, [00:21:43] Generative AI, agentic AI and all, so it has a… [00:21:47] intersection with, uh, all of these. [00:21:50] So, often, many times, people get confused with what is a relationship [00:21:54] in all of these, uh, terms. [00:21:57] So, I've included this diagram here. [00:22:01] So, if we talk about tasks in AI, uh, there are, uh… [00:22:05] Uh, different types of tasks associated. [00:22:09] with AI, such as knowledge representation and reasoning, learning, [00:22:14] Natural language processing, searching, path planning, and so on and so forth. [00:22:21] So, we'll be primarily talking about or discussing learning at length, and of course, [00:22:25] Uh, this will involve concepts of NLP, that is natural language processing. [00:22:31] export systems and all. [00:22:34] So, since, uh, machine learning involves learning, so our focus would be on learning. [00:22:41] And, uh, if we have a system that actually is, uh, able to… [00:22:47] Uh, understand the situation and change its course of action. [00:22:51] based on the experience. [00:22:54] Uh, then we can say that it may be utilizing learning. [00:22:58] So, learning would, uh, require generating new facts from old ones. [00:23:04] Generating new concepts completely, or distinguishing amongst different types of environments and situations. [00:23:13] And, uh, of course, machine learning would involve making machines intelligent in [00:23:20] Somebody. [00:23:24] So, if we talk about major AI techniques, we have supervised learning, [00:23:28] In which we have classification and prediction. [00:23:31] We have unsupervised learning. [00:23:33] Uh, that, uh, involves clustering. [00:23:37] Then we have regression, and of course, generative and agentic AI have not included because it's different and doesn't, uh… [00:23:44] fall in this categorization. [00:23:47] So, we'll be talking, uh, today, I'll be just introducing, uh, these terms. [00:23:53] Uh, technically. And then we'll have, uh… [00:23:57] sessions covering, uh, each and every concept. [00:24:02] Uh, many of you might be using supervised, unsupervised learning. [00:24:06] or regression, or all of them. [00:24:08] Uh, and you might just use this as a recap. [00:24:13] So, if we talk about supervised learning, we are primarily talking about classification. [00:24:20] And, uh, which is based on discriminative models. [00:24:24] So, uh, what is classification? [00:24:27] So, um… so in classification, [00:24:30] We have, uh, we are always given… [00:24:33] Our special data, called a training dataset. [00:24:37] Uh, so I'll be using, uh, the terminology, like, records. [00:24:42] or tuples, or instances. [00:24:46] They all mean the same thing. [00:24:48] And what… what does it mean? It means… [00:24:51] rows in a table. [00:24:53] So, if we talk about a tabular data, then the rows would represent the records, [00:24:59] or instances. And, uh, the fields or the columns would represent the attributes. [00:25:05] So, uh, in case of classification, we are actually given a set of attribute, uh, um, uh, [00:25:12] training data, which has a set of attributes, and one particular attribute is special, and is called a classifying attribute, or the class label. [00:25:24] And this is given to us. So, we have… [00:25:26] records that have got, uh, different attribute values, and also the class label given to us. [00:25:31] In the training dataset. [00:25:34] And, uh, the tasks of classification involves [00:25:39] mapping, um… [00:25:42] input to the output. The output is in form of the class label. [00:25:47] And the input is the set of attributes. So, we want to find out a model [00:25:52] That takes as input the attributes. [00:25:55] Uh, and, uh, gives, uh, or, uh, gives the class label. [00:26:00] So, uh, this means that we have something like Y is equal to function of X, [00:26:07] Where X is the set of input or input attributes, and Y will be the class label. [00:26:14] And the task of classification is to find out the function F. [00:26:19] So, uh, this is what is classification all about. [00:26:23] Once we build them, uh, we, uh, build a model using the training dataset, [00:26:30] Then, what we do is that we… [00:26:32] Test the performance of the model on something called the test data. [00:26:36] The test data is also special. [00:26:39] Because it has attributes and the class label given to us. [00:26:44] And once we see… [00:26:47] that the performance of the classification model [00:26:50] is, um, is fine. Uh, it is acceptable on the test data. [00:26:56] Then we can use it to predict the class label of unseen record. [00:27:00] So, the primary goal of, uh… [00:27:03] Making a classifier, or… [00:27:05] Uh, learning, uh, a classifier. [00:27:08] From the training data is… [00:27:11] to make predictions on unseen data in future. [00:27:16] Unseen data is a dataset. [00:27:17] Whose class label is not known to us, and we want to find it out. [00:27:21] Using the classifier that we have. [00:27:24] built. So, uh, this is a brief on what is, uh, [00:27:30] classification, and we'll be taking it in much more details as we go along. [00:27:37] So, there are a huge amount of applications of classification. I've just mentioned few of them. [00:27:43] Uh, so, for example, it could be used for targeting the consumers which are likely to buy a new product. [00:27:51] So, uh, here, what we want to do is that we want to look at the set of the customers that we are having. [00:27:58] And we could actually classify them into [00:28:02] Uh, those customers who will buy the product, or who will not buy the product. [00:28:06] Based on their attribute values. The attribute values could include the salary of the person, [00:28:14] The lifestyle demographics, [00:28:15] Uh, and so many other things. [00:28:18] Uh, related to the customer. So, in this way, we could actually use [00:28:23] Uh, classification for predicting whether [00:28:27] The person is going to buy or not buy, uh, the… [00:28:31] product. [00:28:34] Another important application of classification is in trot, uh, identification. [00:28:40] of fraud detection. So, for example, uh, if we want to identify [00:28:44] Fraudulent transactions on credit card, [00:28:48] Uh, then also we can use classification. [00:28:50] So, uh, what the classifier will do is that [00:28:55] Uh, it's going to learn the previously available training dataset, [00:29:00] And identify, uh, what are the cases, or how it is going to learn [00:29:05] Uh, how the fraudulent cases can be, uh, identified. [00:29:10] So, for example, a fraudulent transaction could be one that involves a… [00:29:15] Uh, very high transaction amount. [00:29:18] Which is much higher than the average spent on the card, on the credit card. [00:29:22] And I think this is the way the… [00:29:25] Financial institutes also work. [00:29:27] Uh, I think we all often must have encountered situations where… [00:29:31] Uh, we have… we get a call from our, uh, credit card company, [00:29:37] That is just the transaction that you have made, or somebody else has made? [00:29:41] If the transaction is made [00:29:44] Either it is made on a foreign currency, [00:29:48] or the purchase is being done from a website which has not been used earlier, [00:29:53] bias, or if the amount that is spent is very high. [00:29:58] So, all these can be learned by the classifier. [00:30:01] And can be used to make predictions, uh, for, uh, fraud identification. [00:30:08] Um, so before I proceed, uh, I'll just, uh… [00:30:13] Ask if there are any questions, then you can please let me know. [00:30:18] Uh, then we can proceed further. [00:30:22] Any questions, anyone? [00:30:24] So, uh, one point, ma'am, is there? [00:30:29] Mm-hmm, mm-hmm. [00:30:30] Regarding classification. [00:30:31] Yeah. [00:30:32] So, can you see that classification? [00:30:33] is a mechanism where we are distributing data across, uh, [00:30:36] various, uh… [00:30:39] Uh, section. [00:30:43] Various sections. [00:30:44] It can be a multi- like a multiple example is… it might be, like, uh, you mentioned the previous one, a marketing one, right? Which one is the fair one, right? Which is not the… [00:30:51] target. So it can be multi-classification also, right? Not only two. [00:30:56] Yeah, yeah, yes. So, classification is of two types. [00:31:01] In fact, more than two types. One is binary classification, [00:31:04] Mm-hmm. [00:31:05] Where there are only two class labels, like yes or no, buy, don't buy, like that. [00:31:10] It could also be multi-class classification. [00:31:14] Where there are some five… [00:31:16] categories, or five class labels. For example, uh, you know, let's say that if we are talking about some kind of [00:31:23] fraudulent active… activity, and instead of just identifying fraud or non-fraud, [00:31:29] We are identifying the type of fraudulent activity, so there could be 5 types. [00:31:34] So that is a multi-class problem. [00:31:35] Mm-hmm. [00:31:36] There's another way to classify, which is called, uh… [00:31:39] One versus the rest. So, in this case, what we do is, we have a multi-class problem. [00:31:45] That is, let's say we have some 5 class labels, but we are interested only in one particular class label. [00:31:52] Then what we'll do, we'll, uh, identify that class label, [00:31:56] And all the rest we will put in, uh, one particular label. [00:32:01] So, that is called one versus the rest. [00:32:04] That is, we are classifying one. [00:32:06] And the rest, we are all putting together. [00:32:10] Okay? [00:32:11] Okay, yeah. Thanks. [00:32:14] Any other questions, anyone? [00:32:18] Okay, so we'll just proceed. [00:32:24] Uh, so we have something called discriminative AI models. [00:32:29] Uh, discriminative models are those that, uh, actually, uh, [00:32:34] Learn a decision boundary, [00:32:36] And this decision boundary helps to categorize the different classes in the given data. [00:32:42] Uh, and uh… the idea is only to distinguish or identify [00:32:49] The different classes or categories. [00:32:52] Uh, in the given data. [00:32:54] So, uh, what do the discriminative models will do? [00:33:00] These will actually map the input attributes. [00:33:02] And, uh, that is, they are going to model the probability [00:33:07] Uh, that… [00:33:09] What will be the class label for the given data? [00:33:12] So, such models are called discriminative models. Discriminative models… [00:33:17] Primarily, uh, try to discriminate or distinguish [00:33:21] Between the different classes that are available in the data. [00:33:27] So, uh, the key characteristics of discriminative models is that the goal of the model will always [00:33:33] B to classify or to predict the class labels from the input data. [00:33:38] And, uh, how they do it, they do it by building a decision boundary that separates the different classes. [00:33:46] In the data space, and the output will always be in the probability for the given class label. [00:33:52] Uh, say, for example, uh, it is going to… [00:33:56] Use a decision boundary, and then… [00:33:59] Find out, uh, the probability if we have two classes, yes or no, then what will be the probability of occurrence? [00:34:05] of the class label yes, or the class label no. [00:34:08] But this is just a general term, discriminative models. [00:34:12] And, uh, there are different kinds of discriminative models. There are also called classifiers. [00:34:19] For example, logistic regression, support vector machines, [00:34:23] Decision, trees, random forest. [00:34:26] Knive-based classifier, and so on and so forth. We'll be talking about [00:34:30] Each one in more details as we go along. This is just to mention. [00:34:34] There are different kinds of discriminative models. [00:34:38] Then, the next kind of, uh… [00:34:41] Learning is unsupervised learning, okay? [00:34:44] So, previously, I talked about supervised learning. [00:34:48] Which was classification and, uh, building the classifier model to make a prediction. [00:34:54] So, uh, classification is called supervised learning. [00:34:59] Because of availability of class labels, [00:35:02] Uh, in the given training data. [00:35:04] So, in that data, as I told you, we have attribute values and the class labels given to us. [00:35:10] And hence, we call it supervised learning. [00:35:13] unsupervised learning is one where [00:35:16] We have no information about the class labels, or we have no group information. [00:35:22] about the data, and we arrive at it at the end of the process. [00:35:28] So, unsupervised learning is also called clustering. [00:35:33] So, um… [00:35:35] In clustering, we are given a set of points. [00:35:39] Uh, having some attributes, set of attributes, and we use some similarity measure. [00:35:44] to identify groupings in the given data, which are also called clusters. [00:35:50] So, uh, the idea here is that… [00:35:54] the points, the data points that belong to one particular cluster, [00:35:57] are much more similar with respect to each other, rather than [00:36:02] Those data points that belong to… [00:36:03] different, uh, separate clusters. [00:36:06] So, points inside a cluster will be very much similar [00:36:09] to each other, versus those. [00:36:12] points that are, uh, in some other cluster. [00:36:15] Uh, they'll be less similar. [00:36:18] So, uh, and then there are many similarity measures, for example, Euclidean distance. [00:36:24] And so on and so forth. We'll talk about, uh, most of the important similarity measures. [00:36:30] Uh, so, the broad idea here is to identify the groupings [00:36:36] The natural groupings that is available in the data. [00:36:39] In form of clusters. [00:36:41] Now, the thing is that the word… [00:36:45] The English word classification [00:36:48] or the English word clustering. [00:36:49] Both mean grouping of the data. [00:36:52] However, classification is very different from clustering. [00:36:56] Because, uh… [00:36:58] Uh, in classification, we have the group information in form of class label given to us. [00:37:06] Whereas in clustering, we don't have any group information, [00:37:09] We, uh, in fact, we form the groups and arrive at the group information. [00:37:15] At the end of the process. [00:37:19] So this is a diagrammatic representation, for example. [00:37:24] Uh, you can see the three-dimensional data space in the… on the slide. [00:37:28] It has some natural groupings of the data. [00:37:31] Uh, that you can see very easily that the data points are crowded. [00:37:36] Uh, in the, uh, in form of three regions. [00:37:40] One which is shown in the light blue color, dark blue color, and a red color. [00:37:45] And the whole idea is to actually… [00:37:47] maximize the inter-cluster distances, that is the distances across [00:37:54] clusters and minimize the intra-cluster distances. [00:37:58] Uh, this is the whole idea. [00:38:05] So, an example of clustering, I thought it will be interesting. [00:38:07] is to see the histograms which form the clusters, which I did on the [00:38:12] basis of first names of the people, of all of you in this class. [00:38:17] It's interesting to note that the cluster, or the histogram, [00:38:21] With the alphabet S for the first name shows highest frequency, [00:38:25] Followed by, uh, A. [00:38:28] Followed by M, followed by V, [00:38:31] And then R and followed by R and P. [00:38:35] And D, and so on and so forth. [00:38:37] So, these are the clusters that are done on the alphabet. [00:38:41] Uh, of the first name. [00:38:43] of all of you. Of course, these clusters I've shown in form of histograms. [00:38:48] Uh, so just to add, uh, this, uh, insight. [00:38:54] Earlier, I had shown clusters of all of you based on your locations on the… [00:39:00] map of India. [00:39:02] Say, for example, uh, and uh… if we talk about, uh, [00:39:07] You know, some case study where… [00:39:09] Uh, how we can use this concept. [00:39:13] So, we can use it, like, for example, we have location clusters, then [00:39:18] We can, uh, look at, uh, the given location. [00:39:23] And then use it to identify what will be the number of participants, given any particular location. [00:39:28] or given a course, what will be? [00:39:31] The probable number of participants at the different locations. [00:39:35] And so and so forth. [00:39:38] So, uh, clustering is often used, uh, for market segmentation. [00:39:43] Where we want to divide the set of customers into subsets. [00:39:48] And, for example, there could be… [00:39:52] Budget buyers, uh, that spend, uh, economically. [00:39:57] There could be high spending, uh… [00:40:00] set of customers who are… who buy niche products, [00:40:04] Uh, which are costly, and they are not meant for the common person. [00:40:09] And so on and so forth, so… [00:40:12] Uh, market segmentation is, uh… [00:40:13] an application of clustering, where we want to identify the different [00:40:18] types of customers based on their buying habits. [00:40:22] It can also be based on geographical location or lifestyle, or many other things. [00:40:29] So, we want to identify customers which are similar in some way. [00:40:32] That could be location of the customer, that could be spend habits, spending habit. [00:40:37] there could be… that could be anything. [00:40:42] Another application of clustering is in, uh, document, uh… [00:40:47] grouping or document clustering. [00:40:49] Where we want to, uh, uh, find out groups or clusters from the given set of documents. [00:40:56] Uh, in such a way that, uh, the topics of, uh, similar documents [00:41:02] come together in a cluster. So, we want to identify [00:41:07] groups from the given set of documents, and that is based on [00:41:11] What is the topic on, uh, on which the document has been aligned? [00:41:16] Say, for example, uh, we have a large set of documents. It will be very difficult to each time see [00:41:24] what the document is about. [00:41:26] So, a very simple way is to, let's say we have thousands of documents, instead of every time seeing the document and identifying, [00:41:35] what topic the document is covering. It will be much better to form [00:41:40] clusters out of the given set of documents, and then, uh… [00:41:45] See, uh, whenever there's a new document, we can see to which cluster it belongs. [00:41:50] And accordingly, uh, it will be based on some particular topic. [00:41:54] So, like this, uh, we can form clusters from different, uh, documents, different sets of documents. [00:42:02] And this can help in a lot of ways. It can help us to quickly access documents belonging to some particular topic or a set of topics. [00:42:10] And so on and so forth. And uh… this is a very common thing we all also do in our common, uh, in our day-to-day life, for example. [00:42:19] We might keep all the documents in some particular folder. [00:42:24] based on, uh… [00:42:25] Uh, you know, some particular [00:42:27] thing the document is covering. Say, for example, [00:42:31] we might keep all the news articles, or we might… [00:42:35] have some particular documents on AI, then we group them together. [00:42:39] And so on and so forth. So… [00:42:41] Document clustering is very common. We've been doing it ourselves also in our daily life. [00:42:46] However, it can be done on, uh, technically on documents. [00:42:50] Uh, using clustering. [00:42:54] So, uh, another application of clustering would be… [00:42:59] to identify clusters, uh… [00:43:02] of behavior on stock movements. For example, there could be a cluster. [00:43:07] Uh, that, uh, has all, uh, the stocks that may be going down. [00:43:12] or going up, and so on and so forth. [00:43:17] So, having talked about, uh… [00:43:21] Uh, supervised and unsupervised learning. [00:43:23] We'll next, uh, our next discuss in brief what is regression and forecasting. [00:43:30] So, many times you might have heard the term forecasting, uh, which is also very common in predictive analytics. [00:43:38] So, uh, when we talk about [00:43:41] forecasting of values. So, whenever we are predicting values, [00:43:46] Then, what we have is forecasting. [00:43:49] So, forecasting always talks about value prediction. [00:43:53] And regression can be used for forecasting values. [00:43:58] And, uh, in such a case… [00:44:00] We can call, uh, the regression to be performing the task of forecasting. [00:44:04] Forecasting is very important in a lot of applications. For example, weather forecast. [00:44:11] forecasting zone, or forecasting the stock price. [00:44:14] or forecasting any kind of other values, forecasting the performance, and so on and so forth. [00:44:21] And predictive analytics actually tries to [00:44:24] uh, analyze the current and the historical data. [00:44:28] And then use it [00:44:31] Uh, to find out the predictions about [00:44:34] future, uh, data. [00:44:37] And, uh, we call such a data as unseen data. [00:44:42] Uh, and, uh, this kind of predictive analytics has so many applications, about which I already told you. [00:44:48] Some of them. So, prediction is used pretty much in anything. Prediction is also used by doctors. [00:44:56] Uh, to identify, uh, the healthy patient. [00:45:00] And so on and so forth. [00:45:03] So, we talked about forecasting. [00:45:07] Uh, which is used in predictive analytics, where we talk about a value that is getting predicted. [00:45:13] Uh, then, uh, we have anomaly detection. [00:45:17] So, what is anomaly detection? [00:45:19] Anything that is very different from the normal behavior, [00:45:24] is anomaly detection. [00:45:26] So, uh, and, uh… [00:45:28] in the… with the help of machine learning and AI, [00:45:31] We can identify abnormal behavior. [00:45:34] Abnormal behavior is called an anomaly or outlier. [00:45:38] And it is very important in certain applications to identify [00:45:43] Abnormal behavior rather than the normal one. [00:45:46] For example, in medical, um… [00:45:48] applications, it may be of a lot of importance to identify the abdominal [00:45:53] behavior, rather than the normal one. [00:45:57] Uh, and, uh, we could use supervised or unsupervised methods to identify anomalies from the given data. [00:46:05] Uh, say, for example, if we use unsupervised learning, then we are talking about clustering. [00:46:12] So, if we try to cluster the given data, [00:46:14] or form groups out of the given data. [00:46:17] Then, uh, anomalies would be lying very far away, [00:46:21] from the normal groups. And therefore, uh, we can identify anomalies using [00:46:28] unsupervised learning. Similarly, we can also identify anomalous behavior. [00:46:32] using, uh, supervised learning. [00:46:37] Any questions, anyone, so far? [00:46:50] Yes, I'll be sharing the slides with you. [00:46:51] Yeah, we'll get this material in email. [00:46:53] Uh, post the class, yeah. [00:46:54] Yeah, thanks. [00:46:57] And madam, can you, uh, just explain a little more about that intra node, and then intranode clustering? [00:47:02] intra-cluster? Yeah. Okay. [00:47:06] Oh. [00:47:07] So then, uh, allow me to… [00:47:08] Yeah, yeah. [00:47:11] Just to… I'll just do a stop share so I can actually connect my… [00:47:16] Ben, just another meal. [00:48:29] I'll just share the slides and, uh… [00:48:31] We'll discuss it more. [00:48:39] Okay, so let me just… [00:48:42] Go back. [00:48:56] Yeah. [00:49:05] So, what I was telling you was that here, [00:49:08] In case of trustring, these are the clusters that we want to identify. [00:49:13] And to do that, actually, [00:49:15] We want to minimize… [00:49:17] The distances within the points in a cluster. [00:49:21] So, this is called inter-cluster [00:49:24] distance. That is distance within the, uh, in the points. [00:49:30] I've just… auto boot. [00:49:33] Just a sec. [00:49:35] So, we are also mapping that also. We are checking that points also. [00:49:41] the distance between them. [00:49:42] The rest of the points? [00:49:44] Yeah, so… [00:49:45] Yeah. Means, uh, okay. [00:49:46] These clusters, uh, which you have right now, that there are… that are gray. [00:49:52] Mm-hmm. [00:49:53] We are also mapping the distance between them also. [00:49:57] Yeah, so what we do is that… [00:49:58] In clustering, we try to identify, uh, groupings of the points. So, what we see is that [00:50:05] Those points, which are very near to each other, [00:50:09] They form a cluster, and therefore we say the inter-cluster distance needs to be [00:50:14] Sorry, uh, this is intra-cluster, I'm sorry. [00:50:18] So, this is intra-cluster distance that needs to be [00:50:21] minimized, that is the distances within the points. [00:50:25] Whereas distances across the clusters need to be maximized, so this needs to be maximized. [00:50:31] So, we want to identify [00:50:33] Points which are very close together. And in this process, what will happen is we'll actually have to [00:50:39] Find out the distances of all the points with respect to each other. So, I'll be taking it in details as we go along. [00:50:46] Suppose we are having n number of points, right? So what we will have to do is, let's say we have a matrix, [00:50:53] Like, 1, 2, 3, till n. [00:50:56] And these are the columns, and we also have rows like 1, 2, 3. [00:50:59] Till n. So, what we do is we find out the distance of… [00:51:03] Uh, all the points with respect to all others, like this pairwise distance. [00:51:08] And then, based on this distance, we put all the points which are very close to each other in [00:51:14] A single cluster. [00:51:16] Whereas points which are across cluster are very far away. This is how we do. [00:51:21] And to find out this distance, we'll have to use a similarity measure. [00:51:26] So, I'll be discussing these in details as we go along. [00:51:28] Is it okay? The doubt is clarified? [00:51:31] Yes. [00:51:34] Okay. [00:51:41] Yes? [00:51:42] Hello, ma'am. What do you think one question around clustering or deterministic model. [00:51:47] Yeah, sure, go ahead. [00:51:49] Uh, so unlike, uh, supervised learning, it has a bunch of unlabeled data. [00:51:56] So, how, uh… Would we explain, uh, with the training data what clustering actually. [00:52:03] Uh, does whether… I mean, what are you really asking Motor to learn about? [00:52:10] In clustering. [00:52:11] In the absence of those attributes. [00:52:12] Interesting, right? So, in… yeah, so in clustering, [00:52:17] What data we have is… [00:52:19] Uh, the data could be point P1. [00:52:20] in classroom. [00:52:22] P2, like this, and we have attributes. [00:52:27] Let me write it here, that may be better. We have 0.1, 0.2, and so on and so forth. [00:52:31] And for each of these points, we are given attribute 1, attribute 2, [00:52:36] Like this. Maybe some attributes are given. [00:52:40] But what we don't have here… [00:52:42] is the group identification. [00:52:44] We don't have any group information in the given data. [00:52:50] Yes. [00:52:51] Okay, so this is not given. So, what are given are these points, let's say these are some values, I'm just randomly putting some things here. [00:52:58] Maybe this may be 1, 0, something, something. [00:53:01] So, by these values, that is, these are the records, so by these values, we want to identify which points are close to each other. [00:53:10] And those points which are close to each other are put in a cluster. [00:53:14] But this information is not given to us. [00:53:18] But by this process of clustering, we are actually able to identify what are the groups that are there in the data, because we are [00:53:25] Finding out points that are similar to each other. [00:53:28] In classification, this group information is given in the form of [00:53:33] Classifying attribute or class label. [00:53:37] So, uh, in case of classification, [00:53:39] So, when we are talking about classification, [00:53:43] Then what do we do? We have the input data X. [00:53:46] Let's say these are the set of attributes. [00:53:49] Like, A1, A2, A3 till AA. [00:53:52] Okay? And why is the class labeled? [00:53:56] That is what we predict after having built the model. [00:54:00] And f is a function. So we try to find out this function F in such a way that when we give these attributes as the input, [00:54:08] What we get is a correct class label. [00:54:11] But this is the case in… [00:54:14] classification, and in the training data, we have the class label given to us. It's a special kind of data. [00:54:20] So that our model can be built. [00:54:23] But in case of clustering, no class label is given, we just have the attribute values. [00:54:28] And we use, uh, these attributes to find out what points are similar to each other, and then [00:54:35] Uh, we put them in a group. [00:54:37] Is that okay? [00:54:40] Yeah, understood. So that means, uh, in the deterministic model, uh, there isn't two-way approach. First, he has to identify the class label. [00:54:48] And then segment it out, or group it out. [00:54:52] So, if we are doing classification, [00:54:53] those entities. [00:54:55] Then, definitely, we'll need the class label in advance. [00:55:00] But suppose there is a new data, let's say you collect some data on your own, [00:55:04] And you really don't know what are the classes in the given data. [00:55:08] Then what you could do is, you can first cluster the data, [00:55:12] Based on the similarity between the points, and then once the clusters are formed, [00:55:17] Then you can use the clusters [00:55:19] Uh, you can just look at the cluster, study them, and see what is similar. [00:55:25] So, based on those similar things, you can actually give them class labels. [00:55:30] So, classification is often used [00:55:34] When, uh, in, uh, sorry, clustering is often used, [00:55:37] Along with classification, [00:55:39] When the class information is not available, [00:55:42] But to build a classifier, we need the class label, then how do we get it? [00:55:46] So, what we do is we perform clustering. [00:55:49] And then in clustering, groups are formed. [00:55:52] of similar points, then we try to study… [00:55:54] Uh, you know, those points that come within a group or a cluster, [00:55:59] And, uh, according to the, um, the attributes or the behavior of those points, [00:56:06] We can label them. We can identify labels for each of the class labels. [00:56:14] Yep. Yeah, yeah. [00:56:17] Anyone else? [00:56:18] Got it. Thank you so much. [00:56:22] Oh, yeah, hi, ma'am. Uh, I have a question. [00:56:24] Yeah. Mm-hmm. [00:56:25] why it is important to have the maximum distance between the… [00:56:30] intra-cluster. [00:56:33] Okay. [00:56:36] Here. So you want, uh, you're talking about this distance, why it needs to be maximized? [00:56:43] Yeah, yeah, sorry, inter… Mrs. [00:56:44] Yeah. So, the inter-cluster distance needs to be maximized because… [00:56:51] We want similar… [00:56:53] See, the idea of clustering is to put similar points together. [00:56:57] If we are putting, uh… if we are having a large amount of distances within the cluster, then it is not a cluster. [00:57:05] For example, if… [00:57:08] I had a cluster which was like this. [00:57:11] Okay? So, in that case, what would happen? Some points are very close to each other, like these. [00:57:17] Whereas, if we talk about these two points, [00:57:21] They are far apart. [00:57:23] So, uh, see, the idea of clustering is to identify [00:57:28] The naturally occurring, uh… [00:57:31] groupings in the data, not to create clusters just for the sake of it. [00:57:36] We want to identify… [00:57:38] Uh, how the points are crowded, what are the regions in the data space where points are there? [00:57:44] So, that can be correctly done if and only if we find out regions which are full of points. So, when there is a region that will be full of points, [00:57:55] Then the points will be closely packed. [00:57:57] Therefore, the intracluster distance, that is the distance amongst the points in that region, will be minimized. [00:58:03] And we want to maximize the inter-cluster distance so that dissimilar points are not put together. [00:58:11] A point that is dissimilar to some other point will be far away. [00:58:17] That is the whole idea. Points that are similar will be close by. [00:58:23] Points that are dissimilar will be far away. For example, [00:58:25] Uh, if we have, let's say, uh, a 2D space, we have X and we have Y. [00:58:31] And we have a point like, uh, let's say… [00:58:36] So I have a point like… [00:58:39] Just a second. So, this is a point, like, one… [00:58:42] Comma 1, and then I also have a point like… [00:58:46] 2, then I have a point here somewhere. [00:58:51] Which is, like, 10 comma 10. [00:58:55] So, if I have these 3 points, [00:58:58] Then, definitely, [00:59:00] Which two will be similar? [00:59:03] Let's say I call them A, B, and C. [00:59:06] Definitely, A and B are more similar as compared to A and C or beyond C. [00:59:13] Why? Because the distance between A and B is very less. [00:59:15] Whereas, the distance between B and C is very large, [00:59:20] ANC is also very large. [00:59:22] So, uh, so the… [00:59:24] idea here, or the intuition, is that similar points are close to each other. [00:59:29] So, points within a cluster must be similar and must be close to each other. [00:59:35] And points which are across clusters must be dissimilar. [00:59:39] And therefore, the distance between them should be high. [00:59:44] Is it okay? [00:59:49] Mm-hmm. [00:59:50] I'm, uh… I have one doubt, right? So, I mean, uh, how do we set the cutoff for the distance? For example, if we have a scenario, right? I mean. [00:59:55] The distance between the two entries is almost equal to the one that is, you know, in a different cluster. [01:00:03] Is there a way that we do it, or we might be covering this in the later lectures? I'm not sure, yeah. [01:00:08] Okay. Okay. [01:00:09] Yes, yes, we'll be covering it in the later lecture, because today we will not go into details of clustering. [01:00:14] There are a lot of ways to identify [01:00:18] How to cluster the points, and then how to see [01:00:19] Mm-hmm. [01:00:20] Points are similar or not similar, okay? I'll be covering them. [01:00:25] Yeah, thanks. [01:00:27] Okay. Wow. [01:00:28] Oh, sure, ma'am, thank you. [01:00:30] So, I think there is some comment by Meet that, uh… [01:00:34] Centroids need to be found out. [01:00:37] Uh, actually, centroids can… so there are different ways, or different algorithms that perform, uh, that help to perform clustering. [01:00:46] Uh, and in one of these techniques, centroids are used. [01:00:51] to form the cluster, but not all techniques are, uh… [01:00:56] are using centroids. There are other techniques also, which we'll talk about. [01:01:02] As we go along. Uh, and hyperparameter tuning, uh, actually… [01:01:07] Uh, is different from… [01:01:09] Finding out, uh, similar, similarity in points, and [01:01:15] grouping them into clusters. So that's actually very different from hyperparameter tuning. [01:01:23] Any other questions? [01:01:28] Mm-hmm. [01:01:29] Oh, ma'am, I just have one question. So, uh… The given data, for example, may not readily form the clusters. So, in that case, do we apply any sort of tuning on the input data. [01:01:38] To make sure we are able to get some clusters out of it. [01:01:42] Yes, sometimes what happens is that, uh, the data is, uh, actually… [01:01:49] organized in such a way [01:01:51] that it may, uh, in the raw form, it may be difficult to… [01:01:56] form the clusters. [01:01:59] For example, we have a very high-dimensional data, [01:02:03] I'll talk about high-dimensional data slightly later. You can think of data. [01:02:08] Uh, you know, uh, that is… there are points which are very far away. Generally very far away. [01:02:15] Then, how do we form clusters? Because all of them look dissimilar to each other. [01:02:18] So, in such cases, we performed some transformations on the data. [01:02:23] And then we use the data for clustering. [01:02:26] So, uh, there are different techniques. We will talk about them as we go along. Today is just an introductory class, so I'm… [01:02:34] introducing the different concepts to all of you, okay? [01:02:38] Yep, got it, thanks. [01:02:48] Okay, so let's, uh, talk about deep learning. [01:02:52] So, uh, deep learning definitely is a subset of machine learning, which in turn is a subset of [01:02:57] artificial intelligence. [01:02:59] In which we use neural networks. [01:03:02] Uh, and, uh, we think of neural networks to being organized in form of layers. [01:03:08] And these layers make the model deep. [01:03:11] Uh, so, deep learning is very useful when we have very complex data with us, and it's very difficult [01:03:19] to very easily identify the patterns or the relationships in the data. [01:03:25] In such cases, we use deep learning. In deep learning, [01:03:29] We talk about neural networks. [01:03:32] And neural network, uh, actually is comprised of neurons. [01:03:36] And neurons are based on human brain. [01:03:40] So, just like in human brain, we have… [01:03:44] Uh, neurons. [01:03:45] Similarly, here, we have neural networks. [01:03:48] that have notes and, uh, like in brain, the neurons help to… [01:03:54] understand or identify, or do the learning. [01:03:59] The neural networks in deep learning also help to perform the learning. [01:04:02] Especially in cases of very complex patterns and… [01:04:07] Uh, data. [01:04:09] from data which is very… [01:04:13] Uh, otherwise difficult to handle, such as text data, images, [01:04:17] audio, and so on and so forth. [01:04:21] And, uh, the special thing about neural networks is that in machine learning, [01:04:26] We actually try to manually extract, uh… [01:04:29] features from the given data. [01:04:32] But in case of neural networks, the feature extraction is not, uh… [01:04:37] handcrafted, but it is done automatically. [01:04:41] So, we'll be talking a lot of… a lot more in details about deep learning. There'll be [01:04:46] several sessions on deep learning as we go along, where I'll discuss… [01:04:52] Uh, the different techniques of deep learning. [01:04:54] And how they are used. [01:04:56] So right now, I'm just introducing the term. [01:04:59] The applications of deep learning could be immense. For example, in natural language processing, computer vision, [01:05:06] Speech identification, healthcare, finance, forecasting of weather, stock price, and so on and so forth. [01:05:13] So, computer vision means we want to identify, with the help of some images. [01:05:18] So, that is useful in a lot of fields, like, uh… [01:05:23] Nowadays, autonomous vehicles are very, very… [01:05:26] popular. And, uh, if not fully autonomous, partially or semi-autonomous vehicles are commonly available. [01:05:34] For example, we have the… [01:05:36] Uh, vehicles which are having some assistance. [01:05:40] Uh, like, uh… [01:05:42] Uh, then we have cameras deployed on vehicles, on cars, in different directions to give us a… [01:05:49] 360-degree view, or different kinds of views of the surrounding [01:05:54] vehicles, and so on and so forth. [01:05:57] And then there could be automatic sleep identification, driver's sleep identification, and alarm. [01:06:04] There could be automatic, uh… [01:06:06] breaking, and so on and so forth. So, all this can be done with the help of images, which in… which are covered in computer vision. [01:06:16] Natural language processing means [01:06:18] Uh, the, uh… [01:06:19] the techniques that handle [01:06:23] language as it is spoken. It is called natural because, uh… [01:06:27] We as humans speak some language. [01:06:30] That comes… that is a natural thing that we use. [01:06:34] So, natural language processing… [01:06:36] uh… deals with handling this kind of language. [01:06:40] Uh, that is the language spoken [01:06:44] by humans. Uh, it could be English, it could be Hindi, it could be any other regional language, and so on and so forth. [01:06:51] Then, speech recognition means it helps us to identify [01:06:55] Uh, speech data. [01:06:57] to derive features, and this is very much useful in voice assistants like Alexa. [01:07:03] And, uh, and of course, transcription tools. [01:07:07] Uh, and AI summarization, uh, like, AI notes. [01:07:12] Uh, that is used often with meetings and talks and all that. [01:07:16] Then, uh, deep learning has a lot of application in healthcare. [01:07:20] It can help to identify certain… [01:07:23] situations in the, uh, persons, or… [01:07:28] It can also help in drug discovery. [01:07:31] And, uh, recommendations for treatments or plans, or… and so and so forth. [01:07:37] Then, uh, finance also sees a lot of applications of deep learning for, for example, in algorithmic trading. [01:07:45] fraud identification, credit risk modeling, and so on and so forth. [01:07:52] So, some domains in which AI is used have already [01:07:58] covered some of them, like vision, social media, data analytics, NLP. [01:08:03] And primarily, we, uh, I wanted to mention generative AI and agentic AI, which are… [01:08:09] The primary focus of this particular certification course. [01:08:15] So, uh, just an introduction, uh, to generative AI. So, generative AI, uh… [01:08:22] Uh, is a category of algorithms or AI models that help to generate new content. [01:08:29] that new content could be in any form. It could be text, it could be images, it could be code. [01:08:35] It could be, uh… [01:08:37] some audio, it could be anything. [01:08:40] Uh, that is very much similar to real-world [01:08:44] examples. So, um… [01:08:47] Generative AI can help to understand [01:08:50] the context of the given data, [01:08:53] Whether it is image or it is text, or whatever it is, [01:08:56] And then use, uh, the context. [01:09:00] The given style and structure, and then… [01:09:04] Use all of them to produce content. [01:09:05] Which looks like a natural or a humanly generated content. [01:09:10] And, uh, uh… [01:09:13] Primarily, the traditional AI systems are more discriminative in nature, [01:09:18] That is, they are meant to classify or to predict. [01:09:21] However, uh, generative AI is not discriminative, rather it is generative in nature, that is, [01:09:28] It helps to generate. [01:09:30] And, of course, the learning here is also based on patterns that is learned [01:09:37] From vast amount of data, [01:09:38] And the concept is to use these parts, learn these patterns, and use them to generate new [01:09:45] completely new kinds of outputs. [01:09:47] Then there are a large number of generative AI models that you'll be also [01:09:52] Learning in this course, for example, [01:09:55] A bot, GPT… [01:09:58] then, um… [01:09:59] You can also talk… discuss T5, GANs, LLMs, and so and so forth. [01:10:10] So, I'll just skip this then. [01:10:13] Then comes, finally, Agentic AI. So… [01:10:17] So, uh, we are just advancing ahead, more and more ahead. [01:10:22] As we go along. So, initially, uh, when we… if we talk about, uh, the sequence, [01:10:28] Initially, we had simple, traditional machine learning algorithms. [01:10:33] These got diversified. [01:10:35] into deep learning techniques? [01:10:38] And then the deep learning techniques further got diversifi… the diversified. [01:10:45] into, uh, generative AI techniques, which are not just discriminative, they are generative. [01:10:52] And then, uh… [01:10:54] evolved the agentic AI, where… [01:10:56] Uh, it's not just generating new concepts, [01:11:00] We have, uh, agentic AI systems that behave like, uh, you know, autonomous… [01:11:06] Partially or fully autonomous agents that [01:11:09] that not just generate, but they are able to [01:11:12] big decisions, learn the situation, [01:11:15] Then take feedback, and accordingly adapt. [01:11:19] It can help to plan [01:11:21] And, you know, act in dynamic environments. [01:11:25] And it doesn't actually involve the human intervention all the time. [01:11:30] So, for example, uh, you know, you might have seen in gaming, [01:11:34] Uh, that there are, uh, you know, different, uh… [01:11:38] roles that the people, or avatars that are taken up by the players, [01:11:43] And, uh, they, uh, the player needs to decide [01:11:47] Uh, what action [01:11:49] Uh, he or she is going to take, and based on that, the environment is going to dynamically [01:11:56] change, and whatever action… [01:11:58] The person is going to take [01:12:00] There'll be reward or penalty, something will be associated. So, we have agents. [01:12:04] Agents behave in such a way [01:12:07] That, for their actions, they either earn [01:12:10] Uh, rewards or penalties. Here, reward or penalty could be… [01:12:15] Improvement in performance or degradation in performance, and so on and so forth. [01:12:18] And then, based on that, the agents could take corrective action, [01:12:23] could plan, uh, the task. [01:12:26] or set certain goals. [01:12:28] And adapt to dynamic environments. [01:12:32] So, uh, this is about agentic AI. [01:12:36] And, uh, they actually work in different kinds of environments, and so on and so forth. [01:12:41] And use different kinds of data. [01:12:45] Uh, for their decision-making. [01:12:47] So, I think, uh, then applications of agentic AI are immense. Actually, [01:12:53] Uh, in the state-of-the-art agentic AI techniques are pretty much useful in a lot of [01:13:01] Um, real-world scenarios. For example, autonomous robotics is one particular area in which [01:13:08] Uh, agentic AI can be used. [01:13:11] Many times, there are safety-critical systems [01:13:14] For example, uh, systems, uh, [01:13:17] Safety-critical situations where it may not be, um… [01:13:22] good for a human to actually go. [01:13:26] So, there are autonomous robots [01:13:28] that actually are able to, uh… [01:13:31] go at that particular location, that may be critical, or… [01:13:35] risky for the human. [01:13:37] And then they are successfully able to perform the task and return back. [01:13:41] Uh, autonomous robotics is being currently utilized in a lot of areas, including manufacturing. [01:13:49] Cleaning of, uh… cleaning off some hazardous materials. [01:13:54] Uh, and so on and so forth. [01:13:56] Uh, then Agent Kai could be used in software development. [01:14:01] It could be used as personal AI assistance, like the virtual agents that, uh, [01:14:07] manage calendars, send emails, do reports. [01:14:11] do bookings, and so on and so forth, and adapt to preferences. [01:14:16] And, uh, other things. [01:14:18] Then, we also have, uh, agentic AI being used in scientific discoveries for running experiments, for deciding [01:14:27] Uh, how many, uh, runs to, uh, you know, go for. [01:14:32] Then, after deciding the number of runs, uh… [01:14:35] Going through all those runs of experiments, interpreting the data, [01:14:41] And then, uh, based on that interpretation, refining the experiments and running them again, and so on and so forth, or… [01:14:49] making inferences. [01:14:52] So, with this, I will, uh, wrap up, uh, on, uh, this part, and we'll begin a new part. [01:14:58] Uh, that is data preprocessing. I'll start today itself. [01:15:03] But before that, if there are any questions on anything that I've covered so far… [01:15:08] Uh, please ask. [01:15:12] Yes, ma'am. I absolutely said. So, ma'am, actually, my basic question is, so, when we talk about the agents in AI, so what is the fundamental difference between the agent [01:15:25] And then, uh, LLM model. So, [01:15:27] we can use the LLM model in one of the scripts, which can be… which can be targeted to do this on a specific task. [01:15:34] Hmm. [01:15:35] But the same kind of instructions as your prompt and using an LM model we are giving into the… [01:15:41] Yeah. [01:15:42] agent could also. So, then, what is the difference, actually? [01:15:44] So, actually, LLMs can be used, uh, in agent TK directly. [01:15:49] They can be used as agents also. [01:15:52] So, if you're talking specifically about LLMs, [01:15:56] The… they just not only… [01:15:59] You know, do the required, uh, [01:16:03] processing of NLP and all that thing, but they can be used as agents also. [01:16:09] So, there are, uh, LLNs that can be used in Agentic AI as agents. [01:16:15] And without being… [01:16:16] That's why… that's why you have that question. [01:16:19] Okay, okay, my next question is, so without LM also, we can have the agendas, what you're saying? [01:16:24] Yes, of course. [01:16:26] Okay. [01:16:31] Any other question? That you wanted to ask? [01:16:35] Uh, ma'am, one question is there. I always get confused on that. [01:16:37] Mm-hmm. [01:16:39] Mm-hmm. [01:16:40] Um, just want to understand, there's a thin line between AI agent versus authentic AI. [01:16:45] Right? So, what actually makes identify different? Because. It's, again, like, we think about this, like, um, some process. [01:16:54] Which execute on demand, or, like, schedule-based, which is perform, which have a AI capability to perform certain tasks. [01:17:00] Right? And even identity AI also needs some kind of trigger to be happened, right? From… it might be a human trigger, it might be an event triggered, right? So how… what is the difference in that? [01:17:14] So, actually, if we talk about agentic AI, [01:17:15] So I always get confusion with these two. [01:17:18] In this case, such agents are actually able to, uh… [01:17:24] You know, they may or may not rely on any trigger. They are like, [01:17:29] kind of autonomously working over a task. [01:17:33] They can ex… they can plan the task, they can execute the task, then they can [01:17:37] Refine the task, stop it, set goals. [01:17:42] And dynamically change, and all that it does. So, that is what is… [01:17:45] Uh, agentic AI all about? [01:17:48] Now, there are, uh, traditionally, in AI, we also talk about AI agents. [01:17:55] Hmm. [01:17:56] So those agents are not smart agents. For example, we might have agents to do… [01:18:01] ticket booking, or to decide, you know, [01:18:04] Uh, let's say, if you're talking about ticketing, there might be agent, software agents, [01:18:09] or AI agents that can help us to do the booking of a train, or booking of a… [01:18:15] bus or taxi or something. [01:18:17] But, uh, if I talk about agentic AI, and I give it a task that I have to travel [01:18:24] From point A to point B. [01:18:26] And for that, my criteria is I want to travel in minimum amount of time. [01:18:30] So, the agentic AI is going to take this as a task, [01:18:34] Look at different modes of transport available. [01:18:38] Look at that, and then I can also specify [01:18:39] that I want to optimize the cost. [01:18:42] or whatever my criteria is. Then the Agentech AI not only takes up the task, [01:18:48] It decides what are the different modes, then it looks at the permutations, [01:18:54] It might then decide which one is taking, uh, least amount of time, and also, uh, [01:19:00] is optimal over cost, and take an optimal combination of these two constraints. [01:19:06] I gave. So, these all things, and then give us a decision. [01:19:10] that, okay, this should be your way, the… suppose I'm going from A to B to C. [01:19:15] So, to go over this route, [01:19:18] What all, uh, modes should be used? [01:19:21] to mete out less time, and… [01:19:24] optimal cost. So, this is what Agentic AI can do, because it is autonomous. [01:19:30] The simple, uh, agents, the simple AI agents, they are meant to do a specific task, they do it, and that's all. [01:19:38] For example, as I said, there could be an agent… [01:19:41] Which is assigned to, uh, you know, do a… [01:19:46] Uh, booking for a particular journey. [01:19:47] to a particular mode of transport, it will do it and finish, that's all. [01:19:53] But they don't work autonomously. [01:19:54] They don't take decisions like the way Agent TKI does. [01:20:00] Okay. [01:20:01] So… so, ma'am, uh, besides this set. So, basically, you're saying that, um… But a rule-based, simple agent. [01:20:09] is not using the MLM model, because. Yes, we have utilized some rules, and because of that, we have created a. [01:20:15] Yeah. [01:20:16] But an AI agent with a LLM model. can take other decisions also. That is not going to be rule based. [01:20:25] Yeah, yeah. [01:20:27] That's correct. [01:20:29] Yeah, thanks. Uh, who's next? I think Pallavi? [01:20:30] Um. Yeah, um, on the lines, when you said that agentic AI does not, like, may not always include LLM. [01:20:41] So, in generative AI, what we're basically asking is you have some pattern and you generate new. [01:20:49] Mm-hmm. [01:20:50] Patterns, like, and in agenting AI. Uh, like, how is it without an LLM? If you could give an example? [01:20:59] So, uh, actually… [01:21:00] On that, I mean… [01:21:02] I… I was more referring to… [01:21:04] The language model, but there could be vision models also. [01:21:08] There could be vision-based AI also, and then you have agentic AI for those vision models also. [01:21:09] Oh. [01:21:16] Yes. [01:21:17] So, essentially, we're using generative AI to. Right? Okay. [01:21:18] There could be no agentic AI without generative AI. [01:21:22] Yeah. [01:21:23] But the… I mean, the prompt or the instruction is already prefixed based on some task, and… [01:21:30] Yeah, so definitely it will be, uh… [01:21:35] Fixed based on some tasks, but it could be dynamic… I mean, it could be dynamically changing also. [01:21:44] Yeah, thanks. [01:21:45] Okay, yeah, got it. [01:21:46] Ma'am, uh, can you give me one example of, like. [01:21:50] Uh, I see that, uh, um, use LLMs or VLMs. [01:21:56] So, apart from that, what model you can use, and then, uh. [01:22:00] Or do a work. [01:22:04] So, you will be reading this in the entire course, right? So, I think you'll be studying it at length. [01:22:10] So, this will be a little early, right? How they work and all, you'll be [01:22:14] reading all about generative and agentic AI, so you'll have to have a little bit of a patience. [01:22:20] Okay? Because this is the course all about. It cannot be covered or taught in a day. [01:22:26] But if there's a specific question, you can ask me. [01:22:30] pertaining to that. [01:22:35] Is it okay, Sachin? [01:22:37] I think it was 13 only, yeah. [01:22:42] Any other questions, anyone? [01:22:43] Yeah. Yeah, yeah. [01:22:50] So if there are no further questions, then I will, uh, share… [01:22:56] The next slide deck, allow me… [01:22:59] Uh, moment, please. [01:23:18] So, you were supposed to start it tomorrow, but I'm actually starting it today only. [01:23:45] So, we'll be talking about data preprocessing today. [01:23:59] So, uh, any particular data… [01:24:03] Uh, is of, uh… [01:24:06] not of much use till and until, uh, it is in a… [01:24:11] It is in a condition. [01:24:13] Uh, that it is useful for us. [01:24:17] And, uh… [01:24:19] Uh, to make it useful, we have to bring it to, uh… [01:24:23] You know, to the farm. [01:24:26] or to a state that it is rendered useful. [01:24:30] So, let me just show you some things. [01:24:35] About the data. [01:24:41] So, basically, which means unstructured data to structured data. [01:24:46] Not necessarily, it will be unstructured to structure. [01:24:49] Data which is not in good shape. [01:24:50] Mm-hmm. [01:24:53] That we want to pre-process and bring it [01:24:56] up to the mark. So, I'll tell you, uh, what it means, just, uh… [01:25:01] Uh, allow me to explain this to you in the subsequent slides. [01:25:07] So, let's talk about, uh… [01:25:10] data preprocessing, and why it is useful. [01:25:14] So, pre-processing the data is very important, because to have quality results from the data, [01:25:20] We want the data… [01:25:22] to be of good quality. [01:25:24] And it must be up to Mark. [01:25:27] So, there are some, uh, important tasks that form part of data preprocessing. Of course, there are many others also. [01:25:34] These are some important ones that we always include when we talk about data preprocessing. [01:25:40] So, first of all is data cleaning. [01:25:42] That includes filling missing values, [01:25:45] Uh, smoothing of noisy data, resolving inconsistencies and all. [01:25:50] Then, data integration. So, the data may be coming from multiple sources. [01:25:55] And we need to integrate into a single… [01:25:58] Uh, warehouse or a single structure. [01:26:02] Uh, so that it can be useful. [01:26:04] Then, data transformation may involve certain… [01:26:08] operations that may be performed on the data, so that, uh… [01:26:13] Uh, so that it is brought up to the mark that, uh, or, uh, it is brought up to the quality that we want it to be. [01:26:20] I'll explain each one of these as we go along. [01:26:24] Then we have data reduction. [01:26:26] Uh, so data reduction is useful when we have a huge quantity of data. [01:26:31] And it's really difficult to handle this because of compute or some other constraints. [01:26:37] Then we have data discretization. [01:26:40] Uh, that is, uh… [01:26:42] putting the data into some VINs is what is data discretization about. [01:26:49] So, let's talk about, uh… [01:26:50] Each one of these in more details as we go along. [01:26:54] So, first of all is data cleaning. [01:26:56] So, data cleaning, the first thing that we do [01:26:59] in-data cleaning is filling in of missing values. [01:27:04] Okay. So, let's say we have some data. You can see here. [01:27:09] As you can see in this table, and we have, uh… [01:27:13] There are certain records that belong to one class, which is the majority class, that is. [01:27:20] We have 6 samples belonging to Class C2. [01:27:23] And we have 3 samples belonging to class C1. [01:27:28] Which is shown here. So, there are 6 samples of C2 class and 3 samples of C1 class. [01:27:33] And there are total 10, um… [01:27:36] Uh, for records. [01:27:39] And for one record, we don't know the class label. Therefore, [01:27:42] We are talking about 6 and 3 here. [01:27:46] And, uh, what we can see here is that… [01:27:51] We have, like, 3 cells which do not contain any attributes. [01:27:56] So, uh, in this particular case, uh, [01:27:59] We are having about 5% of the data having… [01:28:02] missing values in it. [01:28:04] Now, the question is, how can we handle these missing values? [01:28:09] So, any suggestions or any ideas how we can handle missing values? [01:28:15] Uh, in the given data. [01:28:17] take, uh… take, like, for example, A2 column. We can take a mean or average. [01:28:24] And put the data, because that is going to be nearby value. [01:28:28] It's not going to be discrete much. [01:28:33] Okay, so you want to take the average of what? [01:28:37] For example, A2 column, right? So, 2, 4, 2, 2, and all, so we take an average. [01:28:42] Okay. [01:28:43] Right, so… what the value will come, we can put into R9. [01:28:47] Okay. So, that is average over a column. [01:28:51] That you want to say? [01:28:54] Okay. [01:28:55] Correct. Yeah, so same for E3, same for these two, and uh… But I assume that X is… Coming from the… derived value from A1, A2, A45, right? So, we can… Put a formula, once it's value here, value, then we can take on a… [01:29:09] Um, clustering, if you want, like, okay, example 1, 2, 3, the first row. [01:29:16] Mm-hmm. [01:29:17] Right, some certain value is this. We can form a clusters, okay, C1 category belongs. If the value is above. [01:29:19] let's say, plus is, like, example, 10 or less than 10. [01:29:23] Okay. [01:29:24] 15, right? Similarly. T2 is above that, because I'm able to see some… Some value has, uh, higher values, higher numbers, right? So, like. [01:29:31] So, you're talking about some rule-based… [01:29:33] Class prediction, right? [01:29:34] So, by that way, we can… [01:29:35] That's what you said. [01:29:36] Yes. Yeah, at least to, uh, handle missing values. [01:29:41] Okay. Okay, anybody else? Anything? One answer we got is replace. [01:29:44] And then, right, so… [01:29:46] Can we… can we mark them to 0 or minus 1? [01:29:47] Your column number A3. They are spread across a range of 1 to 6, so probably we can consider taking the median of it. [01:29:55] And, uh, for column X, since it's a class label, we can probably take the mode of it. [01:30:01] I see. [01:30:02] The one that is repeating most number of times. [01:30:04] Okay? [01:30:06] fitting most. [01:30:07] Anyone else? Any other solution? [01:30:09] Yeah, we can replace it with most frequent value. [01:30:14] Okay, so you want to talk about mold? [01:30:18] Okay? [01:30:20] Okay, we got some answers, uh, to, uh, the problem of handling missing values. [01:30:27] One taking mean of the column, and for the class label, having some rule-based [01:30:33] inference on the class label. [01:30:37] Then, using median to, uh… [01:30:40] Fill in the missing value for the attributes and mode for the class label.