# Transformer Summary-10-05-2026

course: Module 4 — Generative AI & LLMs
module: Module-4-Generative-AI-LLMs
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-4-Generative-AI-LLMs/General/Transformer_Summary-10-05-2026.pdf
pages: 35

---
[page 1]
Transformer Neural 
Network Models

[page 2]
Encoder 
Word Embedding
Positional Encoding
Self Attention
Residual Connections
Decoder
Word Embedding
Positional Encoding
Self Attention
Encoder-Decoder Attention
Obtaining Translated Tokens using NN and Softmax
Transformer NN Model

[page 3]
Transformer NN Model

[page 4]
Transformer NN Model
Word embeddings capture word identity
• Each word (or token) has a vector representation, its embedding, that encodes its meaning 
and relationships with other words (e.g., “cat” and “dog” are close).
•
On its own, this embedding says nothing about position — “cat” at the start or end of a 
sentence has the same embedding.
Positional encoding adds order information
• Transformers need a way to represent where each word occurs in the sequence.
•
They achieve this by adding (not replacing) a positional encoding vector:
𝑥𝑖 = word_embedding 𝑤𝑖 + positional_encoding 𝑖
• The addition is elementwise. 
• The positional encoding shifts the embedding slightly to inject order information, but doesn’t 
change its identity.

[page 5]
Transformer NN Model
Addition ≠ overwriting
Why doesn’t adding positional encoding turn a word into “some other word”?
Because:
• The word embedding space and the positional encoding space occupy the same dimension, 
but they carry different kinds of information.
• The embedding matrix learned during training is huge (e.g., 30k–50k tokens). The chance that 
the sum of embedding + position vector happens to exactly match another token’s embedding is 
astronomically small.
• Moreover, there’s no lookup after addition — the sum isn’t reinterpreted as a new “word”. It’s just 
a richer vector representation passed to the next layer.
• Positional encodings augment the word embeddings with order information — they don’t 
replace or distort them enough to change their identity.
•
The model learns to interpret this combination to understand both what the word is and where it 
is in the sequence.

[page 6]
Transformer NN Model
What the model actually “sees”
Each token’s input vector now encodes:
• Semantic meaning (from embedding)
• Position in sequence (from positional encoding)
For example:
• The representation for “dog” in position 2 and position 10 differ, but they are still recognizably “dog”.
• The Transformer can now attend differently depending on where the word is.
• Think of positional encoding like adding a small tag to each word’s meaning:
“dog” + “position 2” → “dog at start”
“dog” + “position 10” → “dog at end”
• Neither becomes “cat” or “apple”; they’re just distinct versions of “dog”.

[page 7]
Transformer NN Model

[page 8]
Transformer NN Model

[page 9]
Encoder 
Word Embedding
Positional Encoding
Self Attention
Residual Connections
Decoder
Word Embedding
Positional Encoding
Self Attention
Encoder-Decoder Attention
Obtaining Translated Tokens using NN and Softmax
Transformer NN Model

[page 10]
Transformer NN Model

[page 11]
Transformer NN Model

[page 12]
Transformer NN Model

[page 13]
Transformer NN Model

[page 14]
Transformer NN Model

[page 15]
Transformer NN Model

[page 16]
Transformer NN Model

[page 18]
Transformer NN Model

[page 19]
Transformer NN Model

[page 20]
Transformer NN Model

[page 21]
Encoder 
Word Embedding
Positional Encoding
Self Attention
Residual Connections
Decoder
Word Embedding
Positional Encoding
Self Attention
Encoder-Decoder Attention
Obtaining Translated Tokens using NN and Softmax
Transformer NN Model

[page 22]
Transformer NN Model
How do Residual Connections Help ?
They Help Gradients Flow During Backpropagation
Transformers are deep networks — they can have dozens or even hundreds of layers. 
Without residual (skip) connections, gradients can vanish or explode, making training unstable 
or causing convergence to fail.
Residual path: 𝑥𝑙+1 = 𝑥𝑙 + 𝑓 𝑥𝑙
The gradient can flow directly through the identity path 𝑥𝑙 reducing vanishing effects. This 
allows very deep models (like GPT or BERT) to train efficiently.

[page 23]
Transformer NN Model
How do Residual Connections Help ?
Preserve Information Across Layers
• Each Transformer layer transforms its input representation. Without a skip connection, deeper 
layers might overwrite or distort important features learned earlier.
• The residual connection ensures the original information (the input to the layer) is retained and 
combined with the new features.
• This helps the model maintain context and coherence over long sequences.
Encourage Incremental Learning
• Rather than forcing each layer to learn a full transformation from scratch, residual connections let 
each layer refine the previous representation.
• The layer learns a residual function — a correction or enhancement — on top of the input.
• This makes optimization smoother and helps avoid overfitting or divergence.

[page 24]
Transformer NN Model
How do Residual Connections Help ?
Preserve Information Across Layers
• Each Transformer layer transforms its input representation. Without a skip connection, deeper 
layers might overwrite or distort important features learned earlier.
• The residual connection ensures the original information (the input to the layer) is retained and 
combined with the new features.
• This helps the model maintain context and coherence over long sequences.
Encourage Incremental Learning
• Rather than forcing each layer to learn a full transformation from scratch, residual connections let 
each layer refine the previous representation.
• The layer learns a residual function — a correction or enhancement — on top of the input.
• This makes optimization smoother and helps avoid overfitting or divergence.

[page 25]
Transformer NN Model

[page 26]
Transformer NN Model

[page 27]
Transformer NN Model

[page 28]
Transformer NN Model

[page 29]
Transformer NN Model

[page 30]
Transformer NN Model

[page 31]
Transformer NN Model

[page 32]
Transformer NN Model
What is Multi-Head Attention in Decoder
Multi-head attention lets a model look at different parts of the input 
sequence from multiple perspectives simultaneously.
•“Attention” = how much focus a word should give to other words
•“Multi-head” = multiple attention computations run in parallel, each 
learning different relationships (semantics, long-range 
dependencies, etc.)

[page 33]
Transformer NN Model
How is it implemented?
We use a triangular mask matrix that blocks attention scores to future 
positions by setting them to −∞ before softmax.
Example mask for 5 tokens:

[page 34]
Transformer NN Model

[page 35]
Transformer NN Model