# GPT Summary-30-05-2026

course: Module 4 — Generative AI & LLMs
module: Module-4-Generative-AI-LLMs
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-4-Generative-AI-LLMs/General/GPT_Summary-30-05-2026.pdf
pages: 40

---
[page 1]
Transformer Models for NLP
GPT-2
Generative Pre-trained Transformer

[page 2]
Recap: BERT
• BERT (Devlin et al., 2019), Bi-directional Encoder Representations from Transformers, makes use of a 
Transformer Encoder to learn contextual relationships between words in any given text input.
• The input to BERT is a sequence of tokens (obtained from the tokenization of the input text).
• The output from BERT is a set of token embeddings.
• BERT representations are jointly conditioned on both past and future context simultaneously using bi-
directional self attention

[page 3]
Causal Language Modelling
• In essence, it is just next word prediction based on an input
sequence.

[page 4]
Causal Language Modelling
• We can think of the model as this black box:

[page 5]
What is GPT?
• GPT, Generative Pre-trained Transformer:
• Generative: Language Model that generates the next word in a sequence
• Pre-trained: Trained on the Language Modelling task
• Transformer: Uses a TransformerArchitecture
Version Title Parameters
GPT-1 Improving Language Understanding 
by Generative Pre-Training
117M
GPT-2 Language Models are Unsupervised 
Multitask Learners
1.5B

[page 6]
What is GPT?
• GPT uses Transformer Decoder blocks.
• It is an Autoregressive model.
• An autoregressive (AR) model is a type of statistical model used for analyzing
and forecasting time series data. It is based on the idea that the current value
of a time series can be explained by its previous values.
• It outputs one token at a time. After each token is produced, it is
added to the input sequence. This combined sequence becomes the
input to the model at the next iteration.

[page 7]
What is GPT?

[page 8]
GPT vs BERT
Model GPT BERT
Transformer Decoder Encoder
Type Autoregressive Autoencoding
Training Objective Causal Language Modelling Masked Language Modelling
Evaluation Tasks GLUE, Perplexity GLUE, SQuAD, NER etc.

[page 9]
GPT vs BERT

[page 10]
GPT: Training Task
• Objective: Causal Language Modelling
• Training Task: Unsupervised Pre-training
• Given an unlabelled (unsupervised) corpus of tokens 𝑈 = {𝑢1, … , 𝑢𝑛}, the objective is to maximise the likelihood:
𝐿 𝑈 = ෍
𝑖
log(𝑃 𝑢𝑖 | 𝑢𝑖−𝑘, … , 𝑢𝑖−1; 𝜃 )
• Where, 𝑘 = 𝑠𝑖𝑧𝑒 𝑜𝑓 𝑡ℎ𝑒 𝑚𝑜𝑑𝑒𝑙′𝑠 𝑐𝑜𝑛𝑡𝑒𝑥𝑡
𝑃 = 𝑐𝑜𝑛𝑑𝑖𝑡𝑖𝑜𝑛𝑎𝑙 𝑝𝑟𝑜𝑏𝑎𝑏𝑖𝑙𝑖𝑡𝑦 𝑚𝑜𝑑𝑒𝑙𝑙𝑒𝑑 𝑤𝑖𝑡ℎ 𝑛𝑒𝑡𝑤𝑜𝑟𝑘 𝑝𝑎𝑟𝑎𝑚𝑒𝑡𝑒𝑟𝑠 𝜃
• This is achieved by using a transformer decoder model with 𝑛 layers as:
ℎ0 = 𝐺𝑎𝑡ℎ𝑒𝑟 𝑈′, 𝑊𝑒 + 𝑊𝑝
ℎ𝑙,𝑖 = 𝑡𝑟𝑎𝑛𝑠𝑓𝑜𝑟𝑚𝑒𝑟𝑖 ℎ𝑙 −1 ∀ 𝑖 ∈ [1, 𝑛]
𝑃𝑈′ = 𝑠𝑜𝑓𝑡𝑚𝑎𝑥 ℎ𝑙,𝑛𝑊𝑒𝑇
• Where, 𝑊𝑒 = 𝑣𝑜𝑐𝑎𝑏𝑢𝑙𝑎𝑟𝑦 𝑡𝑜𝑘𝑒𝑛 𝑒𝑚𝑏𝑒𝑑𝑑𝑖𝑛𝑔𝑠
𝑊𝑝 = 𝑝𝑜𝑠𝑖𝑡𝑖𝑜𝑛𝑎𝑙 𝑒𝑚𝑏𝑒𝑑𝑑𝑖𝑛𝑔𝑠
𝑈′ = 𝑐𝑜𝑛𝑡𝑒𝑥𝑡 𝑣𝑒𝑐𝑡𝑜𝑟 𝑜𝑓 𝑡𝑜𝑘𝑒𝑛𝑠

[page 11]
GPT: Training Task
• Objective: Discriminative Fine Tuning
• Training Task: For a labelled (supervised) downstream task, maximise the log probability for each
pair of inputs (𝑥, 𝑦)
• After pre-training the model as above, the parameters are fine-tuned to a supervised downstream task given a
labelled dataset 𝐶. The dataset consists of a sequence of input tokens and a label 𝑥1, … , 𝑥𝑚 ∶ 𝑦.
• Each sequence of tokens in 𝐶 is passed through the transformer model and the final transformer outputs, ℎ𝑙,𝑛 , are
fed into an added FFN output layer (having parameters 𝑊) to predict 𝑦:
𝑃 𝑦 𝑥1, … , 𝑥𝑚 = 𝑠𝑜𝑓𝑡𝑚𝑎𝑥(ℎ𝑙,𝑛𝑊)
• The objective function for the fine tuning is thus:
𝐿′ 𝐶 = ෍
𝑥,𝑦 ∈𝐶
log(𝑃 𝑦 𝑥1, … , 𝑥𝑚 )
• The loss is also backpropagated into the transformer model using a parameter 𝜆 as:
𝐿𝑓 𝐶 = 𝐿′ 𝐶 + 𝜆 ∗ 𝐿(𝐶)

[page 12]
GPT: Training Task
• Input structures for various fine-tuning tasks

[page 13]
GPT: Datasets and Details
• The datasets used for training were:
• WebText: Around 40GB of text from over 8 million web pages. These web pages were selected using outbound links from highly voted
(posts that had received at least 3 karma points prior to December 2017) Reddit posts.
• Duplicate links and all Wikipedia articles were removed from the dataset (since these would lead to overfitting).
• GPT-2 models were available in 4 sizes:
Model Parameters Decoder Layers Embedding Dimension
GPT-2 Small 117M 12 768
GPT-2 Medium 345M 24 1024
GPT-2 Large 762M 36 1280
GPT-2 XL 1.5B 48 1600

[page 14]
GPT vs BERT: Architecture
BERT GPT2

[page 15]
GPT: Input Tokenization
• GPT uses a Byte Pair Encoding (BPE) model as its tokenizer.
• BPE starts with a small vocabulary of tokens including the special tokens and the initial alphabet.
• It starts by computing the set of unique words and then builds the initial vocabulary using all the 
symbols used to write those words.
• After the base vocabulary is ready, it adds new tokens until the desired vocabulary size is reached 
by learning the merges i.e., the rules to merge two existing tokens to form a new one.
• The merges are learnt by using the most frequent pair of existing tokens:
𝑠𝑐𝑜𝑟𝑒 = 𝑓𝑟𝑒𝑞_𝑜𝑓_𝑐𝑢𝑟𝑟𝑒𝑛𝑡_𝑝𝑎𝑖𝑟

[page 16]
GPT: Input Tokenization
• Consider the following corpus:
("hug", 10), ("pug", 5), ("pun", 12), ("bun", 4), ("hugs", 5)
• The splits will be:
("h" "u" "g", 10), ("p" "u" "g", 5), ("p" "u" "n", 12), ("b" "u" "n", 4), ("h" "u" "g" "s", 5)
Vocabulary: ["b", "g", "h", "n", "p", "s", "u“]
• Let’s assume a vocabulary size of 10 for the current example. So, we need to learn 3 merges since the current vocabulary is 7.
• The most frequent pair is ("u", "g"), which is present 20 times in the vocabulary. So, the first merge learned is ("u", "g") -> ("ug").
Vocabulary: ["b", "g", "h", "n", "p", "s", "u", "ug"]
Corpus: ("h" "ug", 10), ("p" "ug", 5), ("p" "u" "n", 12), ("b" "u" "n", 4), ("h" "ug" "s", 5)
• The most frequent pair at this stage is ("u", "n"), present 16 times in the corpus, so the second merge rule learned is ("u", "n") -> ("un").
Vocabulary: ["b", "g", "h", "n", "p", "s", "u", "ug", "un"]
Corpus: ("h" "ug", 10), ("p" "ug", 5), ("p" "un", 12), ("b" "un", 4), ("h" "ug" "s", 5)
• Now, the most frequent pair is ("h", "ug"), so we learn the merge rule ("h", "ug") -> ("hug").
Vocabulary: ["b", "g", "h", "n", "p", "s", "u", "ug", "un", "hug"]
Corpus: ("hug", 10), ("p" "ug", 5), ("p" "un", 12), ("b" "un", 4), ("hug" "s", 5)
• The final tokenizer is now given as:
Vocabulary:  ["b", "g", "h", "n", "p", "s", "u", "ug", "un", "hug"]
Merge Rules: [("u", "g") -> ("ug"), ("u", "n") -> ("un"), ("h", "ug") -> ("hug")]

[page 17]
GPT: Input Tokenization
• Using the tokenizer from the previous slide:
Vocabulary:  ["b", "g", "h", "n", "p", "s", "u", "ug", "un", "hug"]
Merge Rules: [("u", "g") -> ("ug"), ("u", "n") -> ("un"), ("h", "ug") -> ("hug")]
• Tokenization of the word "bug" is done as:
"bug" -> ["b", "u", "g"] -> ["b", "ug"]
• Tokenization of the word "mug" is done as:
"mug" -> ["[UNK]", "u", "g"] -> ["[UNK]", "ug"]
• Tokenization of the word "thug" is done as:
"thug" -> ["[UNK]", "h", "u", "g"] -> ["[UNK]", "h", "ug"] -> ["[UNK]", "hug"]

[page 18]
GPT: Input Representation
• The input to GPT is tokenized using the BPE tokenizer giving us a list of tokens.
• The tokens are then converted into token embeddings. For every token in the BPE model’s 
vocabulary, GPT learns token embedding representations. For GPT, the vocabulary size is 50,257 
tokens. The token embeddings are stored as a lookup table indexed via the token-id.
• To help GPT express the positions of words within the sentence, position embeddings are used. 
This allows the model to capture the sequence/order of information. 
• GPT uses absolute position embeddings. These are learnt during model training. (This is different 
from BERT which uses sinusoidal position embeddings)

[page 19]
Recap: Sinusoidal Positional Embeddings
• BERT uses Sinusoidal Positional Embeddings to represent sequence order.
• These embeddings are obtained as:
𝑝𝑖 =
sin 𝑖
10000
2∗1
𝑑
𝑐𝑜𝑠 𝑖
10000
2∗1
𝑑
⋮
sin 𝑖
10000
2∗𝑑
2
𝑑
𝑐𝑜𝑠 𝑖
10000
2∗𝑑
2
𝑑
Where, 𝑖 = position of the token ∈ 0, 511
                              𝑑 = embedding dimension of the model
• For any fixed offset 𝑘, 𝑝𝑖+𝑘 can be easily represented as a linear transformation of 𝑝𝑖. This allows the model 
to learn to attend to relative positions.

[page 20]
GPT: Input Representation

[page 21]
GPT: Self Attention
• Once we feed the input to the GPT decoder, it understands the 
context of each word using the Masked/Causal Multi-Head Self 
Attention mechanism.
• It relates each word in the sentence to all the other words preceding 
it in sentence order and learns the causal relationships and meanings 
of the words.

[page 22]
GPT: Self Attention
• From the input embeddings matrix (X), we create three matrices as:
• Query Matix (Q): 𝑄 = 𝑋𝑊𝑄
• Key Matrix (K): 𝐾 = 𝑋𝑊𝐾
• Value Matrix (V): 𝑉 = 𝑋𝑊𝑉
• Next, we compute the self-attention. This has four steps:
• First, we compute the dot products of each of the queries and keys. This is implemented as a matrix multiplication as: 𝑄𝐾𝑇
• Next, we scale the dot products by the square root of the dimension of the vectors:
𝑄𝐾𝑇
𝑑
• The scaled dot products are then masked to prevent attention scores of a word to be affected by words that appear later than it in 
the sentence order:
𝑄𝐾𝑇
𝑑 + 𝑚𝑎𝑠𝑘
• Then, the softmax function is applied to normalize all the scores: 𝑠𝑜𝑓𝑡𝑚𝑎𝑥
𝑄𝐾𝑇
𝑑 + 𝑚𝑎𝑠𝑘
• Lastly, we obtain the attention scores by multiplying the score matrix (post softmax) with the value matrix:
𝑎𝑡𝑡𝑒𝑛𝑡𝑖𝑜𝑛 𝑄, 𝐾, 𝑉 = 𝑠𝑜𝑓𝑡𝑚𝑎𝑥 𝑄𝐾𝑇
𝑑
+ 𝑚𝑎𝑠𝑘 𝑉

[page 23]
Recap: Bi-directional Self Attention
Bi-directional Scaled Dot-
Product Attention
“The animal didn’t cross the street because it was too tired”

[page 24]
GPT: Self Attention
Masked/Causal Scaled Dot-
Product Attention
“The animal didn’t cross the street because it was too tired”

[page 25]
GPT: Self Attention Masking
• Take an input consisting of 4 tokens, “robot must obey orders”. Assume that we are observing the 2nd token 
at the current decoding step. The last two tokens must be masked to prevent them from affecting the 
attention scores (since they appear in the future of the current decoding step). 
• Essentially, all the future tokens should be scored as 0 in the attention matrix.

[page 26]
GPT: Self Attention Masking
• When the model processes the first input from the dataset (row 1), which contains only one word (“robot”), 
100% of its attention will be on that word.
• When the model processes the second input from the dataset (row 2), which contains the words (“robot”,  
“must”), 48% of its attention will be on “robot”, and 52% of its attention will be on “must”. No attention will 
be paid to any tokens appearing later i.e., (“obey”, “orders”).

[page 27]
GPT: Self Attention
𝑎𝑡𝑡𝑒𝑛𝑡𝑖𝑜𝑛 𝑄, 𝐾, 𝑉 = 𝑠𝑜𝑓𝑡𝑚𝑎𝑥 𝑄𝐾𝑇
𝑑
+ 𝑚𝑎𝑠𝑘 𝑉
𝑀𝑢𝑙𝑡𝑖𝐻𝑒𝑎𝑑 𝑄, 𝐾, 𝑉 = 𝑐𝑜𝑛𝑐𝑎𝑡(ℎ𝑒𝑎𝑑1, … , ℎ𝑒𝑎𝑑ℎ)𝑊𝑂
where, ℎ𝑒𝑎𝑑𝑖 = 𝑎𝑡𝑡𝑒𝑛𝑡𝑖𝑜𝑛 𝑄𝑊𝑖
𝑄, 𝐾𝑊𝑖
𝐾, 𝑉𝑊𝑖
𝑉

[page 28]
GPT: Self Attention
Multi-Head Attention “The animal didn’t cross the street because it was too tired”

[page 29]
GPT: Model Output
• The output from the final decoder layer consists of a set of vectors. Each input token has a corresponding output vector. 
For the GPT-2 Small model, the size of this vector is 768 (for each token).
• This is then fed into a “language modelling head”, which is an FFN (with softmax activation) trained to predict the next 
token probabilities from the decoder output.
• This FFN gives an output vector of size 50,257 (vocabulary size of GPT-2 model) for each input token which predicts the 
next token probabilities.
GPT-2
[BOS]
Keep

[page 30]
GPT: KV Caching
• The GPT model is autoregressive i.e., it essentially uses the past to predict the future. 
• Given a sequence of input tokens (𝑥1, … , 𝑥𝑛), the self-attention block computes the vectors (𝑞1, … , 𝑞𝑛),
𝑘1, … , 𝑘𝑛  and 𝑣1, … , 𝑣𝑛  in order to predict 𝑥𝑛+1.
• In the next step, given the sequence of input tokens (𝑥1, … , 𝑥𝑛+1), the self-attention block computes the 
vectors (𝑞1, … , 𝑞𝑛+1), 𝑘1, … , 𝑘𝑛+1  and 𝑣1, … , 𝑣𝑛+1  in order to predict 𝑥𝑛+2 . Here, the vectors
𝑘1, … , 𝑘𝑛  and 𝑣1, … , 𝑣𝑛  are being recomputed.
• In-fact, all previous Q, K & V vectors are recomputed each time a new token is taken into consideration. 
• A key observation is that the K & V set of vectors contain the sequence context, while the query vector is 
only needed for the last token.
• The dot product between last query vector & K corresponds to computing the attention scores between the 
last token and all the previous ones.

[page 31]
GPT: KV Caching
• Some observations during the 
process:
• During the sequence generation 
one token at a time, K & V sets of 
vectors and do not change. We just 
add another vector to each of the K 
& V sets when we process a new 
token.
• Once we have computed the K & V 
embeddings for the new token, it’s 
not going to change, no matter how 
many more tokens we generate.

[page 32]
GPT: KV Caching
• We can essentially cache the K & V 
vectors for generating the future tokens. 
• This way we can avoid recomputing the 
entire set of vectors 𝑘1, … , 𝑘𝑛  and
𝑣1, … , 𝑣𝑛  in order to predict 𝑥𝑛+1.
• This leads to a significant speedup in 
generation of new token.

[page 33]
GPT: KV Caching
Without KV Caching
With KV Caching

[page 34]
GPT: Fine Tuning
• The output from the GPT model can be fine-tuned to a specific downstream task (rather than just language 
modelling i.e., next word prediction). 
• This can be done by training a different final modelling head (FFN + Softmax). Let’s  take the summarization 
task for example. Given a text passage, the decoder model must generate a summary for it. 
• The training dataset could be Wikipedia articles with their summaries:

[page 35]
GPT: Fine Tuning
GPT-2

[page 36]
GPT: In-Context Learning
• Thus far, we have seen that the models learn via the following mechanisms:
• Unsupervised Pre-Training
• Fine-tuning to task specific domains
• Large Language Models also seem to perform some kind of learning directly from 
examples provided to them without any SGD steps.
• This kind of in-context example based learning is referred to as in-context 
learning. For example:
• Input:    thanks -> merci, hello -> bonjour, mint -> menthe, otter -> ?
• Output: loutre
• GPT is a prime example of this kind of learning.

[page 37]
GPT: In-Context Learning

[page 38]
GPT: Perplexity
• Perplexity is a widely-used metric for evaluating the performance of language models. It 
measures the uncertainty of a model's predictions, specifically how well it predicts the next word 
in a sequence. A lower perplexity indicates higher confidence and better performance, while a 
higher perplexity suggests more uncertainty.
• It is mathematically defined as:
𝑃𝑒𝑟𝑝𝑙𝑒𝑥𝑖𝑡𝑦 =  𝑒−1
𝑁 σ ln(𝑃 𝑥𝑖 | 𝑥1, … , 𝑥𝑖−1 )
where, 𝑃 𝑥𝑖 | 𝑥1,  … , 𝑥𝑖−1 = 𝑐𝑜𝑛𝑑𝑖𝑡𝑖𝑜𝑛𝑎𝑙 𝑝𝑟𝑜𝑏𝑎𝑏𝑖𝑙𝑖𝑡𝑦 𝑜𝑓𝑥𝑖 𝑔𝑖𝑣𝑒𝑛 𝑡ℎ𝑒 𝑝𝑟𝑒𝑐𝑒𝑒𝑑𝑖𝑛𝑔 𝑤𝑜𝑟𝑑𝑠
             𝑁 = 𝑇ℎ𝑒 𝑡𝑜𝑡𝑎𝑙 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑤𝑜𝑟𝑑𝑠 𝑖𝑛 𝑡ℎ𝑒 𝑚𝑜𝑑𝑒𝑙 𝑣𝑜𝑐𝑎𝑏𝑢𝑙𝑎𝑟𝑦
• Equivalently, it can also be defined as:
𝑃𝑒𝑟𝑝𝑙𝑒𝑥𝑖𝑡𝑦 =
𝑁
ෑ 1
𝑃 𝑥𝑖 | 𝑥1,  … , 𝑥𝑖−1
where, 𝑃 𝑥𝑖 | 𝑥1,  … , 𝑥𝑖−1 = 𝑐𝑜𝑛𝑑𝑖𝑡𝑖𝑜𝑛𝑎𝑙 𝑝𝑟𝑜𝑏𝑎𝑏𝑖𝑙𝑖𝑡𝑦 𝑜𝑓𝑥𝑖 𝑔𝑖𝑣𝑒𝑛 𝑡ℎ𝑒 𝑝𝑟𝑒𝑐𝑒𝑒𝑑𝑖𝑛𝑔 𝑤𝑜𝑟𝑑𝑠
             𝑁 = 𝑇ℎ𝑒 𝑡𝑜𝑡𝑎𝑙 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑤𝑜𝑟𝑑𝑠 𝑖𝑛 𝑡ℎ𝑒 𝑚𝑜𝑑𝑒𝑙 𝑣𝑜𝑐𝑎𝑏𝑢𝑙𝑎𝑟𝑦

[page 39]
GPT: Perplexity
If we weren’t limited by a model’s context size, we would evaluate the model’s perplexity by autoregressively factorizing a sequence and 
conditioning on the entire preceding subsequence at each step:
However, we have a constraint on the maximum number of tokens the model can process. For GPT-2 this is 1024 (maximum context length). 
So, we must break the sequence down into smaller chunks. However, breaking the text down into segments equal to the model’s max context 
length and calculating the likelihoods of each segment independently is suboptimal. This is because, it gives the model very little context to use 
for prediction at the beginning of each individual segment.
The optimal approach is to instead employ a sliding window strategy, where the context window is continually moved across the sequence, 
allowing the model to take advantage of the available context.