# Summary Slides T5-31-05-2026
course: Module 4 — Generative AI & LLMs
module: Module-4-Generative-AI-LLMs
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-4-Generative-AI-LLMs/General/Summary_Slides_T5-31-05-2026.pdf
pages: 33
---
[page 1]
Transformer Models for NLP
T5
Text-to-Text Transfer Transformer
[page 2]
Recap: Transfer Learning
• Transfer learning refers to pre-training a model on a large unlabeled
dataset with a self supervising task such as language modelling, and
then fine-tuning the model on a smaller labelled dataset for a specific
downstream task.
• We have looked at GPT, BERT, ALBERT & RoBERTa all of which used
the same training methodology of pre-training followed by task
specific fine-tuning.
[page 3]
What is T5?
• T5 (Raffel et al., 2019) is an encoder-decoder transformer model. It
stands for Text-to-Text TransferTransformer.
• Text-to-Text: It takes text input and produces text output
• Transfer: It used the transfer learning methodology for training
• Transformer: It is based on the transformer architecture
Model Title Parameters
T5-Base Exploring the Limits of Transfer
Learning with a Unified Text-to-Text
Transformer
220M
[page 4]
What is T5?
• Pose every NLP task as a text-to-text task (McCann et al., 2018).
• Neural Machine Translation
Translate English to German: English text -> German text
• Linguistic Acceptability
CoLA sentence: text -> acceptable | not acceptable
• Semantic Textual Similarity
STS sentence: A sentence: B -> x.y
• Text Summarization
Summarize: text -> summary
• Pretrain the network on a large unlabelled dataset.
• Fine-tune for each downstream task.
[page 5]
What is T5?
[page 6]
GPT vs BERT vs T5
Model GPT BERT T5
Transformer Decoder Encoder Encoder-Decoder
Type Autoregressive Autoencoding Autoencoding &
Autoregressive
Training Objective Causal Language
Modelling
Masked Language
Modelling
Span Corruption
Evaluation Tasks GLUE, Perplexity GLUE, SQuAD, NER etc. GLUE, SQuAD, SGLUE etc.
[page 7]
Recap: GPT vs BERT Architecture
[page 8]
T5 Architecture Overview
T5
Decoder Stack
Encoder Stack
’ ‘ ’
[page 9]
T5: Training Task
• Unsupervised Objective: Span Corruption
• Input sequence:
• Randomly sample and dropout ~15% of the tokens in the input sequence.
• All consecutively dropped out tokens are considered spans. Each span is replaced by a special sentinel token.
• These sentinel tokens are special tokens added to the model vocabulary.
• Target sequence:
• All the dropped out spans of tokens delimited by their corresponding sentinel token in the input sequence.
• An extra sentinel token is added at the end of the target sequence.
• The model is trained to predict the target sequence given the span corrupted input sequence.
• The loss function used is cross-entropy loss:
𝐻 𝑝, 𝑞 = −
𝑖
𝑝 𝑖 log(𝑞 𝑖 )
where, 𝑝 𝑖 = 𝑇ℎ𝑒 𝑡𝑟𝑢𝑒 𝑝𝑟𝑜𝑏𝑎𝑏𝑖𝑙𝑖𝑡𝑦 𝑜𝑓 𝑡ℎ𝑒 𝑡𝑎𝑟𝑔𝑒𝑡 𝑡𝑜𝑘𝑒𝑛 𝑖. 𝑒. , 1
𝑞 𝑖 = 𝑇ℎ𝑒 𝑝𝑟𝑒𝑑𝑖𝑐𝑡𝑒𝑑 𝑝𝑟𝑜𝑏𝑎𝑏𝑖𝑙𝑖𝑡𝑦 𝑜𝑓 𝑡ℎ𝑒 𝑡𝑎𝑟𝑔𝑒𝑡 𝑡𝑜𝑘𝑒𝑛
[page 10]
T5: Training Task
[page 11]
T5: Training Task
• The input tokens are passed through the
encoder to output the encoder hidden
states or encoder embeddings.
• The encoder’s output embeddings are
passed to the decoder, which then predicts
the masked token spans.
• The decoder’s output hidden states are
passed to a language modeling head
(classifier), which performs a linear
transformation to convert these hidden
states into logits.
• The cross-entropy loss is then calculated
between the logits and the target token.
[page 12]
T5: Training Task
• Objective: Supervised Fine-Tuning
• The T5 model was fine-tuned using the same cross-entropy loss function for various downstream
tasks such as:
• Text Classification: GLUE
• Text Summarization: CNN, Daily Mail datasets
• Question Answering: SQuAD
• Neural Machine Translation: WMT English to French, German and Romanian
• The fine-tuning for each task was done over 262,144 (218) steps and a checkpoint was stored
every 5000 steps. The checkpoint with the highest benchmark on the validation datasets was
used for reporting results.
[page 13]
GPT: Datasets and Details
• The datasets used for training were:
• Colossal Clean Crawled Corpus (C4): A cleaned version of Common Crawl Web Extracted Text dataset. It is approximately 745GB.
• The cleaning process involved deduplication of content, discarding incomplete sentences, and removing offensive or noisy content.
• The T5 Base Encoder and T5 Base Decoder details are as follows:
Model Parameters Layers Embedding Dimension
T5 Base Encoder 110M 12 768
T5 Base Decoder 110M 12 768
[page 14]
T5: Training Overview
[page 15]
Recap: GPT vs BERT
BERT GPT2
[page 16]
T5: Architecture
[page 17]
T5: Input Tokenization
• T5 uses Unigram algorithm in combination with SentencePiece (Kudo et al., 2018) as its tokenizer.
• SentencePiece addresses the fact that not all languages use spaces to separate words. Instead, it treats
the input as a raw stream which includes spaces in the set of characters to use. It then uses the
Unigram algorithm to construct the appropriate vocabulary.
• This is done by computing the set of unique words. The initial vocabulary consists of all the symbols
used to write those words. Compared to BPE and WordPiece, Unigram works in the other direction: it
starts from a large initial vocabulary and removes tokens from it until it reaches the desired vocabulary
size.
• At each step of the training, the Unigram algorithm computes a loss over the corpus by tokenizing
every word in the corpus, using the current vocabulary.
• Then, for each symbol in the vocabulary, the algorithm computes how much the overall loss would
increase if the symbol was removed.
• The symbols that have a lower effect (least increase) on the overall loss over the corpus are the best
candidates for removal. A subset of these symbols (controlled by a hyperparameter p) are removed.
[page 18]
T5: Input Tokenization
• Loss Function:
𝑃 𝑡𝑜𝑘 = 𝑓𝑟𝑒𝑞(𝑡𝑜𝑘)
σ𝑡∈𝑣𝑜𝑐𝑎𝑏 𝑓𝑟𝑒𝑞(𝑡)
𝑃 𝑡1, … , 𝑡𝑛 = ෑ
𝑡=1
𝑛
𝑃(𝑡𝑡)
𝐿𝑜𝑠𝑠 =
𝑤∈𝐶
𝑓req 𝑤 ∗ −log 𝑃 𝑤
where, 𝑤 = 𝑡1, … , 𝑡𝑛 𝑤𝑜𝑟𝑑 𝑡𝑜𝑘𝑒𝑛𝑖𝑧𝑎𝑡𝑖𝑜𝑛
𝐶 = 𝐶𝑜𝑟𝑝𝑢𝑠 𝑜𝑓 𝑤𝑜𝑟𝑑𝑠
[page 19]
T5: Input Tokenization
• Consider the following corpus:
("hug", 10), ("pug", 5), ("pun", 12), ("bun", 4), ("hugs", 5)
• The splits will be:
Vocabulary: ["h", "u", "g", "hu", "ug", "p", "pu", "n", "un", "b", "bu", "s", "hug", "gs", "ugs"]
Frequencies: [("h", 15) ("u", 36) ("g", 20) ("hu", 15) ("ug", 20) ("p", 17) ("pu", 17) ("n", 16), ("un", 16) ("b", 4) ("bu", 4) ("s", 5)
("hug", 15) ("gs", 5) ("ugs", 5)]
• Consider the possible tokenization "pug" for with this vocabulary:
"pug" = ["p", "u", "g"] | ["pu", "g"] | ["p", "ug"]
The score for each of these tokenization would be:
["p", "u", "g"] = P("p") * P("u") * P("g") =
17
210 ∗
36
210 ∗
20
210 = 0.0003
["pu", "g"] = P("pu") * P("g") =
17
210 ∗
20
210 = 0.007
["p", "ug"] = P("p") * P("ug") =
17
210 ∗
20
210 = 0.007
The tokenization with the highest score is selected. Here, since there is a tie, we select the tokenization that occurred first i.e., ["pu", "g"] as the tokenization for this word.
• Tokenizing all the words in the corpus similarly, we get:
"hug": ["hug"] (score 0.0714)
"pug": ["pu", "g"] (score 0.0077)
"pun": ["pu", "n"] (score 0.0061)
"bun": ["bu", "n"] (score 0.0014)
"hugs": ["hug", "s"] (score 0.0017)
[page 20]
T5: Input Tokenization
• The loss can now be computed as:
"hug": ["hug"] (score 0.0714)
"pug": ["pu", "g"] (score 0.0077)
"pun": ["pu", "n"] (score 0.0061)
"bun": ["bu", "n"] (score 0.0014)
"hugs": ["hug", "s"] (score 0.0017)
Loss = 10 * (-log(0.0714)) + 5 * (-log(0.0077)) + 12 * (-log(0.0061)) + 4 * (-log(0.0014)) + 5 * (-log(0.0017)) = 169.8
• Now we need to compute how removing each token affects the loss.
• For example, "pug" could be tokenized ["p", "ug"] with the same score. Thus, removing the "pu" token from the vocabulary will give the exact same loss i.e.,
∆= 0.
• On the other hand, consider removing "hug" from the vocabulary will change the tokenizations and loss as:
"hug": ["hu", "g"] (score 0.0068)
"hugs": ["hu", "gs"] (score 0.0017)
Loss = 10 * (-log(0.0068)) + 5 * (-log(0.0077)) + 12 * (-log(0.0061)) + 4 * (-log(0.0014)) + 5 * (-log(0.0017)) = 193.3
Now, ∆= 23.5
• So, at this step of training the token "pu" will be dropped from the vocabulary giving us:
Vocabulary: ["h", "u", "g", "hu", "ug", "p", "n", "un", "b", "bu", "s", "hug", "gs", "ugs"]
Frequencies: [("h", 15) ("u", 36) ("g", 20) ("hu", 15) ("ug", 20) ("p", 17) ("n", 16), ("un", 16) ("b", 4) ("bu", 4) ("s", 5) ("hug", 15)
("gs", 5) ("ugs", 5)]
[page 21]
T5: Input Representation
• The input to T5 is tokenized using the BPE tokenizer giving us a list of tokens.
• The tokens are then converted into token embeddings. For T5, the vocabulary size is 32,128
tokens. The token embeddings are stored as a lookup table indexed via the token-id.
• T5 does not add any positional encoding to the token embeddings unlike GPT which uses absolute
position embeddings and BERT which uses sinusoidal position embeddings.
• T5 instead uses relative positional embeddings during the self-attention computation (Shaw et al.
2018).
[page 22]
T5: Relative Position Self-Attention
Absolute Position
Relative Position
[page 23]
T5: Relative Position Self-Attention
• The relative position is defined as, 𝑟𝑒𝑙𝑝𝑜𝑠 = 𝑘𝑒𝑦𝑝𝑜𝑠 − 𝑞𝑢𝑒𝑟𝑦𝑝𝑜𝑠 i.e., the distance between the query and
key tokens.
• These relative distances are then split into two categories, close and long distances. This is done
based on a model parameter 𝑛𝑏𝑢𝑐𝑘𝑒𝑡𝑠.
• Any relative distance ቂ0, ቃ
𝑛𝑏𝑢𝑐𝑘𝑒𝑡𝑠
2 is called a close distance. The model learns a relative positional
embedding (per head) for each close distance.
• Any relative distance ቂ ቃ
𝑛𝑏𝑢𝑐𝑘𝑒𝑡𝑠
2 , 𝑚𝑎𝑥𝑑𝑖𝑠𝑡𝑎𝑛𝑐𝑒 is called a long distance. These long distances are binned
using a logarithmic function into the remaining
𝑛𝑏𝑢𝑐𝑘𝑒𝑡𝑠
2 buckets. The model learns a relative
positional embedding (per head) for each of these long distance buckets. All distances > 𝑚𝑎𝑥𝑑𝑖𝑠𝑡𝑎𝑛𝑐𝑒
are binned into the same bucket.
𝑓(𝑟𝑒𝑙𝑝𝑜𝑠) = 𝑛𝑏𝑢𝑐𝑘𝑒𝑡𝑠
2 ∗ 1 +
log 𝑟𝑒𝑙𝑝𝑜𝑠 ∗ 2
𝑛𝑏𝑢𝑐𝑘𝑒𝑡𝑠
𝑙𝑜𝑔 𝑚𝑎𝑥𝑑𝑖𝑠𝑡𝑎𝑛𝑐𝑒 ∗ 2
𝑛𝑏𝑢𝑐𝑘𝑒𝑡𝑠
[page 24]
T5: Relative Position Self-Attention
• For the T5 model, 𝑛𝑏𝑢𝑐𝑘𝑒𝑡𝑠 = 32 and 𝑚𝑎𝑥𝑑𝑖𝑠𝑡𝑎𝑛𝑐𝑒 = 128:
• Consider the T5 Encoder. This uses a bidirectional self-attention block similar to BERT. Here, relative positional
embeddings need to be learnt for both directions i.e., left to right and right to left.
• Thus, the total number of buckets, 𝑛𝑏𝑢𝑐𝑘𝑒𝑡𝑠, are split into two. One for left to right relative positional embeddings
and one for right to left relative positional embeddings. Thus, the T5 Encoder learns 16 left to right relative positional
embeddings and 16 right to left relative positional embeddings.
• Each of these sets of 16 embeddings is split into close and large distance groups. Thus, for the T5 encoder, any
relative distance < 8 is considered a close distance and has a unique relative positional embedding. Relative distances
> 8 are binned using the logarithmic function. There are 8 such bins each of which share a relative positional
embedding.
[page 25]
T5: Relative Position Self-Attention
• For the T5 model, 𝑛𝑏𝑢𝑐𝑘𝑒𝑡𝑠 = 32 and 𝑚𝑎𝑥𝑑𝑖𝑠𝑡𝑎𝑛𝑐𝑒 = 128:
• Consider the T5 Decoder. Here, relative positional embeddings need to be learnt for the left to right direction since
this uses a causal attention block.
• Thus, the total number of buckets, 𝑛𝑏𝑢𝑐𝑘𝑒𝑡𝑠 are used to learn left to right relative positional embeddings.
• Now, for the T5 encoder, any relative distance < 16 is considered a close distance and has a unique relative positional
embedding. Relative distances > 16 are binned using the logarithmic function. There are 16 such bins each of which
share a relative positional embedding.
[page 26]
T5: Relative Position Self-Attention
𝑠𝑜𝑓𝑡𝑚𝑎𝑥 𝑄𝐾𝑇 + 𝛼
𝑑
+ 𝑚𝑎𝑠𝑘 𝑉
𝑚𝑎𝑠𝑘
[page 27]
T5: Encoder Self-Attention
Bi-directional Scaled Dot-
Product Attention
“The animal didn’t cross the street because it was too tired”
[page 28]
T5: Decoder Self-Attention
Masked/Causal Scaled Dot-
Product Attention
“Das Tier überquerte die Straße nicht, weil es zu müde war”
[page 29]
T5: Encoder-Decoder Cross-Attention
[page 30]
T5: Encoder-Decoder Cross-Attention
• From the input embeddings matrix (X) and the encoder hidden states (Y), we create three matrices as:
• Query Matix (Q): 𝑄 = 𝑋𝑊𝑄
• Key Matrix (K): 𝐾 = 𝑌𝑊𝐾
• Value Matrix (V): 𝑉 = 𝑌𝑊𝑉
• Next, we compute the cross-attention. This has four steps for each head:
• First, we compute the dot products of each of the queries and keys. This is implemented as a matrix multiplication as: 𝑄𝐾𝑇
• Next, we scale the dot products by the square root of the dimension of the vectors:
𝑄𝐾𝑇
𝑑
• Then, the softmax function is applied to normalize all the scores: 𝑠𝑜𝑓𝑡𝑚𝑎𝑥
𝑄𝐾𝑇
𝑑
• Lastly, we obtain the attention scores by multiplying the score matrix (post softmax) with the value matrix:
𝑎𝑡𝑡𝑒𝑛𝑡𝑖𝑜𝑛 𝑄, 𝐾, 𝑉 = 𝑠𝑜𝑓𝑡𝑚𝑎𝑥 𝑄𝐾𝑇
𝑑
𝑉
• Note that the T5 Cross Attention block does not use any relative positional embedding. They are used only
during the self attention blocks in the encoder/decoder.
[page 31]
T5: Encoder-Decoder Cross-Attention
“The animal didn’t cross the street because it was too tired”
“Das Tier überquerte die Straße nicht, weil es zu müde war”
[page 32]
T5: Benchmarks
• The below table provides the comparison of T5 model’s performance against BERT Base across
the following GLUE Benchmark tasks:
• Sentence Pair Tasks:
• MNLI: Multi Genre Natural Language Inference
• STS-B: Semantic Textual Similarity Benchmark
• MRPC: Microsoft Research Paraphrase Corpus
• RTE: Recognizing Textual Entailment
• Single Sentence Classification:
• SST-2: Stanford Sentiment Treebank
Model MNLI SST-2 STS-B MRPC RTE GLUE Average
BERT Base 84.6 93.5 85.8 88.9 66.4 79.6
T5 Base 86.2 95.2 85.0 90.7 70.1 82.7
[page 33]
T5: Benchmarks
• The T5 model also outperforms BERT Base on the SQuAD benchmark:
Model F1 Score
BERT Base 88.5
T5 Base 92.1