# RAG LangChain FAISS

course: Module 4 โ€” Generative AI & LLMs
module: Module-4-Generative-AI-LLMs
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-4-Generative-AI-LLMs/General/21_June/RAG_LangChain_FAISS.pdf
pages: 7

---
[page 1]
LECTURE SERIES ยท AI SYSTEMS
Retrieval-Augmented
Generation
with LangChain & FAISS
LangChain FAISS Vector Search

[page 2]
RAG = Open Book Exam
An LLM using RAG is like a student in an open-book exam
Question
Asked
Exam question
= User's query
Open Book
= Knowledge Base
Notes & textbooks
= Vector database
Smart Answer
Generated
Student synthesizes
= LLM generates reply
๐Ÿ’ก
  Closed-book (vanilla LLM) relies on memory. RAG lets the model consult sources before answering.

[page 3]
RAG Pipeline
1
Documents
PDFs, Web, DBs
2
Chunk & Embed
Split โ†’ Vectors
3
FAISS Index
Vector Store
4
Query
User asks โ†’
Embedded
5
Retrieve
Top-K similar
chunks
6
LLM Answer
Context +
Prompt โ†’ Reply
INDEXING PHASE (offline) RETRIEVAL & GENERATION PHASE (runtime)
# LangChain + FAISS quickstart 
from langchain.vectorstores import FAISS 
from langchain.embeddings import OpenAIEmbeddings 
db = FAISS.from_documents(docs, OpenAIEmbeddings()) 
retriever = db.as_retriever(search_kwargs= {"k": 4} )

[page 4]
How FAISS Finds Similar Vectors
Vector Space
Doc A
Doc B
Doc C
Doc D
Doc E
Doc F Doc G
Query
Top-K
neighbours
Embedding
Text โ†’ float vector (e.g. 1536-dim OpenAI ada)
Cosine Similarity
Measures angle between vectors; closer = more similar
IVF Index
FAISS clusters space for fast Approximate NN search
Top-K Retrieval
Returns the K most similar chunks to the query

[page 5]
LangChain RAG Chain
Question
 Retriever
(FAISS)
Relevant
Chunks
Prompt
Template
LLM
(GPT / Claude)
 Answer
# Full RAG chain in LangChain 
from langchain.chains import RetrievalQA 
from langchain.chat_models import ChatOpenAI 
from langchain.prompts import PromptTemplate 
llm = ChatOpenAI(model= "gpt-4o", temperature= 0)
chain = RetrievalQA.from_chain_type( 
    llm=llm, retriever=retriever, chain_type= "stuff"
    return_source_documents= True)
result = chain.invoke( {"query": "What is RAG?"} )

[page 6]
RAG vs Vanilla LLM
๐Ÿค–
  Vanilla LLM
โŒ
  Knowledge frozen at training cut-off
โŒ
  Hallucinations with no source
โŒ
  Can't access private documents
โŒ
  No citations possible
โœ…
  Fast (no retrieval step)
๐Ÿ”
  RAG
โœ…
  Fresh knowledge โ€” any external source
โœ…
  Grounded answers with source context
โœ…
  Works on private/proprietary docs
โœ…
  Citable, verifiable responses
โš 
  Slightly higher latency

[page 7]
Key Takeaways
01 RAG = Retrieve relevant docs โ†’ Augment the prompt โ†’ Generate grounded answer
02 FAISS converts text to vectors and finds semantically similar passages in milliseconds
03 LangChain wires everything together: loaders, splitters, embeddings, retrievers & chains
04 RAG beats vanilla LLM on freshness, factuality & private-document Q&A
05 Think open-book exam: the LLM consults your knowledge base before answering