# Transformer Models for NLP-17-05-2026

course: Module 4 — Generative AI & LLMs
module: Module-4-Generative-AI-LLMs
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-4-Generative-AI-LLMs/General/Transformer_Models_for_NLP-17-05-2026.pdf
pages: 47

---
[page 1]
Transformer Models for NLP
BERT
Bi-directional Encoder Representations from Transformers

[page 2]
Outline
• Background
• Introduction to BERT
• Text Processing in BERT
• Bi-Directional Self Attention
• Other BERT Based Models
• Benchmarks for BERT Models

[page 3]
Background: ELMo
• ELMo, Embeddings from Language Models, (Peters et al., 2018) is used to produce word embeddings from text.
• Instead of using a fixed embedding for each word, it looks at the entire sentence before assigning each word an embedding vector.

[page 4]
Background: ELMo
• ELMo consists of L layers of language models (biLMs).
• Input tokens are processed by a character level CNN.
• Every layer of ELMo captures different information form the source
sentence. The final embeddings (per token) are computed as a
weighted sum across the layers.

[page 5]
Background: ELMo
• ELMo consists of L layers of language models (biLMs).
• Forward LM:   𝑝 𝑡1, 𝑡2, … , 𝑡𝑁 = ς𝑘 = 1
𝑁 𝑝(𝑡𝑘|𝑡1, … , 𝑡𝑘−1)
• Backward LM: 𝑝 𝑡1, 𝑡2, … , 𝑡𝑁 = ς𝑘 = 1
𝑁 𝑝(𝑡𝑘|𝑡𝑘+1, … , 𝑡𝑁)
• Input tokens are processed by a character level CNN.
• Every layer of ELMo captures different information form the source
sentence. The final embeddings (per token) are computed as a
weighted sum across the layers.

[page 6]
Background: ELMo

[page 7]
What is BERT?
• BERT (Devlin et al., 2019), Bi-directional Encoder Representations
from Transformers, makes use of a Transformer Encoder to learn
contextual relationships between words in any given text input.
• Bi-directional: Allows to represent context from both past and future
• Encoder: Uses a TransformerEncoder
• Representation: Uses bi-directional self attention
• The input to BERT is a sequence of tokens (obtained from the
tokenization of the input text).
• The output from BERT is a set of token embeddings.

[page 8]
What is BERT?
• BERT has a maximum context size of 512 tokens, i.e., it can represent
inputs with a maximum sequence length of 512 tokens.
• Any sequence less than 512 tokens has to be padded to a length of
512.

[page 9]
What is BERT?
• ELMo uses a concatenation of forward and backward LMs to generate contextual token embeddings.
• BERT representations are jointly conditioned on both past and future context simultaneously using bi-
directional self attention

[page 10]
BERT: Training Task
• Masked Language Modelling.
• For every input sequence,
randomly select 15% of the
input tokens and:
• For 80% of them, replace
them with the [MASK] token.
• For 10% of them, replace
them with a random token.
• Keep the remaining 10%
unchanged.
 BERT

[page 11]
BERT: Training Task
• Next Sentence Prediction. 
• For every input sequence, 
randomly select a split:
• Store split A.
• For 50% of the splits, replace 
split B with another sentence’s 
split B.
• For the remaining 50% use the 
original split B.
BERT

[page 12]
BERT: Datasets and Details
• BERT uses the sum of the mean masked LM likelihood and mean next sentence prediction likelihood as it 
training loss function.
• The datasets used for training were:
• BooksCorpus: 800M words from around 7,000 unique unpublished books from various genres.
• English Wikipedia: 2,500M words
• BERT-Base:
• Number of transformer blocks: 6
• Embedding dimension: 768
• Maximum Sequence Length: 512
• Total Parameters: 110M

[page 13]
BERT: Visualization

[page 14]
BERT: Input Tokenization
Wu et al., 2016
• BERT uses a WordPiece (Wu et al., 2016) model as its tokenizer.
• The WordPiece tokenizer starts from a small vocabulary including the special tokens used by the model and 
the initial alphabet. 
• It identifies subwords by adding a prefix (like ## for BERT), each word is initially split by adding that prefix to 
all the characters inside the word. So, for instance, "word" is split as:
w ##o ##r ##d
• Next, the tokenizer learns the merge rules i.e., rules to merge these initial tokens to form larger tokens. The 
merges are based on a score computed as:
𝑠𝑐𝑜𝑟𝑒 = 𝑓𝑟𝑒𝑞_𝑜𝑓_𝑝𝑎𝑖𝑟
(𝑓𝑟𝑒𝑞_𝑜𝑓_𝑓𝑖𝑟𝑠𝑡 ∗ 𝑓𝑟𝑒𝑞_𝑜𝑓_𝑠𝑒𝑐𝑜𝑛𝑑)
• This score allows the algorithm to ensure merging of pairs of tokens where the individual parts are less 
frequent in the vocabulary.

[page 15]
BERT: Input Tokenization
• Consider the following corpus:
("hug", 10), ("pug", 5), ("pun", 12), ("bun", 4), ("hugs", 5)
• The splits will be:
("h" "##u" "##g", 10), ("p" "##u" "##g", 5), ("p" "##u" "##n", 12), ("b" "##u" "##n", 4), ("h" "##u" "##g" "##s", 5)
• Consequently, the initial vocabulary would be:
["b", "h", "p", "##g", "##n", "##s", "##u"]
• The most frequent pair is ("##u", "##g") (present 20 times), but the individual frequency of "##u" is very high, so its score is not the highest (it’s 1 / 36). So, the 
best score goes to the pair ("##g", "##s") at 1 / 20, and the first merge learned is ("##g", "##s") -> ("##gs"). The new vocabulary is now:
["b", "h", "p", "##g", "##n", "##s", "##u", "##gs"]
Splits: ("h" "##u" "##g", 10), ("p" "##u" "##g", 5), ("p" "##u" "##n", 12), ("b" "##u" "##n", 4), ("h" "##u" "##gs", 5)
• We continue merging tokens like this till we reach the desired vocabulary size (let’s assume 10 for the current example). The final vocabulary would be:
["b", "h", "p", "##g", "##n", "##s", "##u", "##gs", "hu", "hug"]
• The tokenizer only saves the final vocabulary (the merge rules are not saved). During tokenization, the longest subword (starting at the word to be tokenized) 
present in the vocabulary is found and split on. This continues till the word has been tokenized. Consider the following examples using the vocabulary above:
hugs -> [“hug”, “##s”]
bugs -> [“b”, “##ugs”] -> [“b”, “##u”, “##gs”]
• If during tokenization of a word, no subword can be found in the vocabulary to proceed, the whole word is tokenized using a special token [“[UNK]”]

[page 16]
BERT: Input Representation
• The input to BERT is tokenized using the WordPiece tokenizer giving us a list of tokens.
• The tokens are then converted into token embeddings. For every token in the WordPiece model’s 
vocabulary, BERT learns token embedding representations. For BERT Base, the vocabulary size is 30,522 
tokens. The token embeddings are stored as a lookup table indexed via the token-id.
• BERT also learns segment embeddings. These are embeddings for the first and second sentences (separated 
by the [SEP] token). It helps the model distinguish between the sentence pairs for certain tasks such as next 
sentence prediction.
• Lastly, to help BERT express the positions of words within the sentence, position embeddings are used. This 
allows the model to capture the sequence/order of information. BERT uses sinusoidal position embeddings.

[page 17]
BERT: Input Representation

[page 18]
BERT: Positional Embeddings
• BERT uses Sinusoidal Positional Embeddings to represent sequence order.
• These embeddings are obtained as:
𝑝𝑖 =
sin 𝑖
10000
2∗1
𝑑
𝑐𝑜𝑠 𝑖
10000
2∗1
𝑑
⋮
sin 𝑖
10000
2∗𝑑
2
𝑑
𝑐𝑜𝑠 𝑖
10000
2∗𝑑
2
𝑑
Where, 𝑖 = position of the token ∈ 0, 511
                              𝑑 = embedding dimension of the model
• For any fixed offset 𝑘, 𝑝𝑖+𝑘 can be easily represented as a linear transformation of 𝑝𝑖. This allows the model 
to learn to attend to relative positions.

[page 19]
BERT: Self Attention
• Once we feed the input to the BERT encoder, it understands the 
context of each word using the Multi-Head Self Attention 
mechanism.
• It relates each word in the sentence to all the other words and learns 
the relationships and contextual meanings of the words.
• The attention mechanism produces the contextual representation of 
each word in the sentence.

[page 20]
BERT: Self Attention
attention
• From the input embeddings matrix (X), we create three matrices as:
• Query Matix (Q): 𝑄 = 𝑋𝑊𝑄
• Key Matrix (K): 𝐾 = 𝑋𝑊𝐾
• Value Matrix (V): 𝑉 = 𝑋𝑊𝑉
• Next, we compute the self-attention. This has four steps:
• First, we compute the dot products of each of the queries and keys. This is implemented as a matrix multiplication as: 𝑄𝐾𝑇
• Next, we scale the dot products by the square root of the dimension of the vectors:
𝑄𝐾𝑇
𝑑
• Then, the softmax function is applied to normalize all the scores: 𝑠𝑜𝑓𝑡𝑚𝑎𝑥
𝑄𝐾𝑇
𝑑
• Lastly, we obtain the attention scores by multiplying the score matrix (post softmax) with the value matrix:
𝑎𝑡𝑡𝑒𝑛𝑡𝑖𝑜𝑛 𝑄, 𝐾, 𝑉 = 𝑠𝑜𝑓𝑡𝑚𝑎𝑥 𝑄𝐾𝑇
𝑑
𝑉

[page 21]
BERT: Self Attention
Bi-directional Scaled Dot-
Product Attention
“The animal didn’t cross the street because it was too tired”

[page 22]
BERT: Self Attention
• The core idea behind using multiple heads in parallel instead of a 
single head is that each head can attend to different contextual 
relationships between words in the sentence.
• For example, one head could learn the subject-verb relationship while 
another could focus on the actor-verb relationships.

[page 23]
BERT: Self Attention
𝑎𝑡𝑡𝑒𝑛𝑡𝑖𝑜𝑛 𝑄, 𝐾, 𝑉 = 𝑠𝑜𝑓𝑡𝑚𝑎𝑥 𝑄𝐾𝑇
𝑑
𝑉
𝑀𝑢𝑙𝑡𝑖𝐻𝑒𝑎𝑑 𝑄, 𝐾, 𝑉 = 𝑐𝑜𝑛𝑐𝑎𝑡(ℎ𝑒𝑎𝑑1, … , ℎ𝑒𝑎𝑑ℎ)𝑊𝑂
where, ℎ𝑒𝑎𝑑𝑖 = 𝑎𝑡𝑡𝑒𝑛𝑡𝑖𝑜𝑛 𝑄𝑊𝑖
𝑄, 𝐾𝑊𝑖
𝐾, 𝑉𝑊𝑖
𝑉

[page 24]
BERT: Self Attention
Multi-Head Attention “The animal didn’t cross the street because it was too tired”

[page 25]
Why Self-Attention?
• Self-Attention has a lower computation complexity than RNNs making it more efficient.
• RNNs are sequential and require “n” steps to process “n” input tokens. Self-Attention achieves 
this in a single operation.
Note: n = sequence length & d = embedding representation dimension
Layer Type Complexity Per Layer Sequential Operations
Self-Attention O(n2d) O(1)
Recurrent Neural Network O(nd2) O(n)

[page 26]
BERT: Model Output
• The output from BERT is a set of embedding vectors. Each input token has a corresponding output 
embedding vector.
• The size of the vector is the “hidden_size” of the model. It is 768 for BERT Base.

[page 27]
BERT: Model Output
• The output embedding vectors can be used for a variety of downstream tasks. The embedding 
corresponding to the special [CLS] token provides the context for the entire sentence.
• For example, the embedding vector corresponding to the [CLS] token can be used as an input to a classifier. 
This classifier can be trained for various tasks such as sentiment analysis of the sentence (classifying the 
embedding as positive, negative or neutral), spam detection (classifying the embedding as spam or not 
spam) etc.
• The individual tokens’ output embedding vectors can be used for tasks such as Named Entity Recognition, 
Question Answering etc. The individual output embedding vectors can be provided as input to a classifier to 
classify them as a named entity (person, place etc.), start/end of an answer etc.

[page 28]
BERT: Downstream Tasks
• Spam Detection (/ Sentiment Analysis / Other Sentence Classification Tasks):
BERT

[page 29]
BERT: Downstream Tasks
• Named Entity Recognition:
BERT
Input Token Class
[CLS] O
Mr B-per
. B-per
Trump I-per
’ O
s O
tweets O
began O
just O
moments O
after O
a O
Fox B-org
News I-org
report O
by O
Mike B-per
Tobin I-per
, O
a O
reporter O
for O
the O
network O
, O
about O
protests O
in O
Minnesota B-geo
and O
elsewhere O
. O
[SEP] O

[page 30]
BERT: Downstream Tasks
• Question Answering: To feed a QA task into BERT, we pack both the question and the reference text into the 
input.

[page 31]
BERT: Downstream Tasks
• Question Answering:
• BERT needs to highlight a "span" of text containing the answer. This is represented as simply predicting which token marks the start 
of the answer, and which token marks the end.
• For every token in the text, we feed its final embedding into the start token classifier. Whichever word has the highest probability of 
being the start token is the one that we pick as the start of the answer.
BERT

[page 32]
BERT: Downstream Tasks
• Question Answering:
• We repeat this process for the end token. For every token in the text, we feed its final embedding into the end token classifier. 
Whichever word has the highest probability of being the end token is the one that we pick as the end of the answer.
BERT
{'answer': '340M', 'start': 27, 'end’: 30}

[page 33]
BERT Based Models
• Various variants of BERT have been developed to cater to different 
types of NLP based workloads. Here are a few variants of BERT:
• ALBERT
• RoBERTa
• DistilBERT

[page 34]
ALBERT: A Lite version of BERT
• The BERT Base model has 12 layers, while the BERT Large model has 24 layers. Adding layers increases the 
parameters exponentially.
• To solve this, ALBERT (Lan et al., 2020) uses cross-layer parameter sharing. Instead of learning unique 
parameters for each layer, the parameters are learnt only for the first layer and shared amongst the layers. 
The parameters can be shared either only for the FFN layer, only for the Self Attention layer or both. 
• ALBERT also uses factorized embedding i.e., it factorized the input token embeddings to make them much 
smaller.
Version Number of Layers Parameters
BERT-Base 12 110M
BERT-Large 24 340M

[page 35]
ALBERT: A Lite version of BERT
• Factorized Embeddings:
Data Size
90 MB
Data Size
12 MB

[page 36]
ALBERT: A Lite version of BERT
• The evaluation on the Reading Comprehension (RACE) benchmark yields the following results:
• RoBERTa (Liu et al., 2019) & XLNet (Yang et al., 2020) shed light on the ineffectiveness of the Next Sentence 
Prediction (NSP) task since NSP and Masked Language Modeling (MLM) are quite similar. ALBERTA replaces 
the NSP task with Sentence Order Prediction (SOP). SOP Takes two consecutive parts of a document as a 
positive case. Swapping the order yields a negative case.
Model Parameter Sharing Number of Parameters RACE
ALBERT Base
(Embedding Size = 768)
All Shared 31M 63.3%
Shared Attention 83M 67.7%
Shared FFN 57M 62.6%
Not Shared 108M 68.2%

[page 37]
ALBERT: A Lite version of BERT
NSP SOP

[page 38]
RoBERTa: Robustly Optimized BERT Approach
• RoBERTa (Liu et al., 2019) improved on BERT with the following changes to the training process:
• Dynamic masking instead of static in the MLM task. BERT masked the data once during the preprocessing stage, leading to a single 
static mask. To allow for different masks, BERT duplicated the training data 10 times and masked it with different strategies. RoBERTa 
used dynamic masking where a new masking pattern was generated each time a sequence was input to the model.
• The Next Sentence Prediction (NSP) task was removed from the training process. RoBERTa was trained only on the MLM task. It was 
found that this led to improved results on the downstream tasks.
• BERT was trained on a batch size of 256 sequences for 1M steps. RoBERTa instead used larger batch sizes. It was trained for 125 steps 
with a batch size of 2,000 sequences and for 31,000 steps with a batch size of 8,000 sequences. Larger batches improved perplexity 
on the MLM training objective. 
• RoBERTa uses a Byte-Pair Encoding (BPE) tokenizer with a vocab size of 50,265 tokens instead of BERT’s 
WordPiece (WPM) tokenizer with a vocab size of 30,522 tokens. This allows for effective handling of the 
extensive vocabularies typical in various natural language corpora.
• It was trained on 10x the BERT training data. The datasets used are:
• BookCorpus (16 GB)
• CC-NEWS (76 GB)
• OpenWebText (38 GB)
• Stories (31 GB)

[page 39]
DistilBERT
• DistilBERT (Sanh et al., 2020) was pretrained using knowledge distillation to create a smaller 
model which allows for faster inference. 
• Knowledge Distillation is a model compression technique where a larger model acts as a 
“teacher” for a smaller model. The smaller model tries to replicate the larger model’s layer 
activations and outputs for a given set of inputs.
• DistilBERT uses a similar encoder architecture as the BERT Base model. However, it has half the 
number of encoder layers. 
• It used a triple loss function:
• Language Modelling Loss: Calculated using next word prediction performance.
• Distillation Loss: Calculated as the loss between the student/teacher model’s predictions
• Cosine-Distance Loss: To align the sublayer activations (hidden states) of the student model with the teacher model.
• Using these combined losses, DistilBERT was able to mimic (97%) of BERT’s performance with 40% 
fewer parameters than BERT. 
• It also achieved 60% faster inference than BERT.

[page 40]
DistilBERT
Text Corpus
Combined 
Loss
Backpropagation
DistilBERT
BERT
Model Number of Parameters Inference Time (s) (GLUE Task: STS-B)
ELMo 180 895
BERT-Base 110 668
DistilBERT 66 410

[page 41]
BERT: Benchmarks
• BERT model performance and accuracy has been continuously 
evaluated over different datasets for various NLP tasks. 
• This is done to quantify the performance/accuracy of the BERT 
models. 
• Some of these benchmarks are:
• GLUE: General Language Understanding Evaluation
• SQuAD: Stanford Question Answering Dataset
• RACE: Reading Comprehension
• CoNLL NER: Conference on Natural Language Learning Named Entity 
Recognition

[page 42]
BERT: Benchmarks GLUE
• GLUE, General Language Understanding Evaluation is a collection of datasets that can be used to 
train, evaluate and analyze NLP models.
• The GLUE benchmark includes 9 diverse task datasets:
• Sentence Pair Tasks:
• MNLI: Multi Genre Natural Language Inference
• QQP: Quora Question Pairs
• QNLI: Question Natural Language Inference
• STS-B: Semantic Textual Similarity Benchmark
• MRPC: Microsoft Research Paraphrase Corpus
• RTE: Recognizing Textual Entailment
• WNLI: Winograd Natural Language Inference
• Single Sentence Classification:
• SST-2: Stanford Sentiment Treebank
• CoLA: Corpus of Linguistic Acceptability 
• To evaluate a model, it is first trained over the training dataset provided by GLUE and then scored 
across all the tasks. The final performance score is the average of all the 9 tasks.

[page 43]
BERT: Benchmarks GLUE
Model MNLI QQP QNLI SST-2 CoLA STS-B MRPC RTE Average
BERT Base 84.6 71.2 90.1 93.5 52.1 85.8 88.9 66.4 79.6
BERT Large 86.7 72.1 91.1 94.9 60.5 86.5 89.3 70.1 81.9
The table above denotes an example of a GLUE benchmark run over two models: BERT Base and 
BERT Large.

[page 44]
BERT: Benchmarks SQuAD
• The SQuAD, Stanford Question Answering Dataset, is a reading comprehension dataset consisting of 
question asked on a set of Wikipedia articles. The answer to each of the questions is a text segment from the 
passage. SQuAD has 100,000 question/answer pairs.
Model Dev
Exact Match (EM) F1 Score
BERT Base 80.8 88.5
BERT Large 84.1 90.9

[page 45]
BERT: Benchmarks RACE
• RACE is a large-scale reading comprehension dataset from examinations. 
• The RACE dataset is used to evaluate models on a reading comprehension task. This dataset was collected 
from English examinations of Chinese students. It consists of nearly 28,000 passages and 100,000 questions 
generated by human experts. The number of questions is much larger in RACE as compared to other 
benchmark datasets.
• The BERT large model achieves a score of 73.8% on the RACE benchmark.

[page 46]
BERT: Benchmarks CoNLL NER
• The CoNLL 2003 Named Entity Recognition (NER) dataset consists of 200,000 training words which have 
been annotated as Person, Organization, Location, Miscellaneous, or Other (non-named entity).
Model Dev F1 Test F1
BERT Base 96.4 92.4
BERT Large 96.6 92.8