# Transformer Summary-10-05-2026 course: Module 4 — Generative AI & LLMs module: Module-4-Generative-AI-LLMs type: pdf source_url: https://personal-learn.armco.dev/files/Module-4-Generative-AI-LLMs/General/Transformer_Summary-10-05-2026.pdf pages: 35 --- [page 1] Transformer Neural Network Models [page 2] Encoder Word Embedding Positional Encoding Self Attention Residual Connections Decoder Word Embedding Positional Encoding Self Attention Encoder-Decoder Attention Obtaining Translated Tokens using NN and Softmax Transformer NN Model [page 3] Transformer NN Model [page 4] Transformer NN Model Word embeddings capture word identity • Each word (or token) has a vector representation, its embedding, that encodes its meaning and relationships with other words (e.g., “cat” and “dog” are close). • On its own, this embedding says nothing about position — “cat” at the start or end of a sentence has the same embedding. Positional encoding adds order information • Transformers need a way to represent where each word occurs in the sequence. • They achieve this by adding (not replacing) a positional encoding vector: 𝑥𝑖 = word_embedding 𝑤𝑖 + positional_encoding 𝑖 • The addition is elementwise. • The positional encoding shifts the embedding slightly to inject order information, but doesn’t change its identity. [page 5] Transformer NN Model Addition ≠ overwriting Why doesn’t adding positional encoding turn a word into “some other word”? Because: • The word embedding space and the positional encoding space occupy the same dimension, but they carry different kinds of information. • The embedding matrix learned during training is huge (e.g., 30k–50k tokens). The chance that the sum of embedding + position vector happens to exactly match another token’s embedding is astronomically small. • Moreover, there’s no lookup after addition — the sum isn’t reinterpreted as a new “word”. It’s just a richer vector representation passed to the next layer. • Positional encodings augment the word embeddings with order information — they don’t replace or distort them enough to change their identity. • The model learns to interpret this combination to understand both what the word is and where it is in the sequence. [page 6] Transformer NN Model What the model actually “sees” Each token’s input vector now encodes: • Semantic meaning (from embedding) • Position in sequence (from positional encoding) For example: • The representation for “dog” in position 2 and position 10 differ, but they are still recognizably “dog”. • The Transformer can now attend differently depending on where the word is. • Think of positional encoding like adding a small tag to each word’s meaning: “dog” + “position 2” → “dog at start” “dog” + “position 10” → “dog at end” • Neither becomes “cat” or “apple”; they’re just distinct versions of “dog”. [page 7] Transformer NN Model [page 8] Transformer NN Model [page 9] Encoder Word Embedding Positional Encoding Self Attention Residual Connections Decoder Word Embedding Positional Encoding Self Attention Encoder-Decoder Attention Obtaining Translated Tokens using NN and Softmax Transformer NN Model [page 10] Transformer NN Model [page 11] Transformer NN Model [page 12] Transformer NN Model [page 13] Transformer NN Model [page 14] Transformer NN Model [page 15] Transformer NN Model [page 16] Transformer NN Model [page 18] Transformer NN Model [page 19] Transformer NN Model [page 20] Transformer NN Model [page 21] Encoder Word Embedding Positional Encoding Self Attention Residual Connections Decoder Word Embedding Positional Encoding Self Attention Encoder-Decoder Attention Obtaining Translated Tokens using NN and Softmax Transformer NN Model [page 22] Transformer NN Model How do Residual Connections Help ? They Help Gradients Flow During Backpropagation Transformers are deep networks — they can have dozens or even hundreds of layers. Without residual (skip) connections, gradients can vanish or explode, making training unstable or causing convergence to fail. Residual path: 𝑥𝑙+1 = 𝑥𝑙 + 𝑓 𝑥𝑙 The gradient can flow directly through the identity path 𝑥𝑙 reducing vanishing effects. This allows very deep models (like GPT or BERT) to train efficiently. [page 23] Transformer NN Model How do Residual Connections Help ? Preserve Information Across Layers • Each Transformer layer transforms its input representation. Without a skip connection, deeper layers might overwrite or distort important features learned earlier. • The residual connection ensures the original information (the input to the layer) is retained and combined with the new features. • This helps the model maintain context and coherence over long sequences. Encourage Incremental Learning • Rather than forcing each layer to learn a full transformation from scratch, residual connections let each layer refine the previous representation. • The layer learns a residual function — a correction or enhancement — on top of the input. • This makes optimization smoother and helps avoid overfitting or divergence. [page 24] Transformer NN Model How do Residual Connections Help ? Preserve Information Across Layers • Each Transformer layer transforms its input representation. Without a skip connection, deeper layers might overwrite or distort important features learned earlier. • The residual connection ensures the original information (the input to the layer) is retained and combined with the new features. • This helps the model maintain context and coherence over long sequences. Encourage Incremental Learning • Rather than forcing each layer to learn a full transformation from scratch, residual connections let each layer refine the previous representation. • The layer learns a residual function — a correction or enhancement — on top of the input. • This makes optimization smoother and helps avoid overfitting or divergence. [page 25] Transformer NN Model [page 26] Transformer NN Model [page 27] Transformer NN Model [page 28] Transformer NN Model [page 29] Transformer NN Model [page 30] Transformer NN Model [page 31] Transformer NN Model [page 32] Transformer NN Model What is Multi-Head Attention in Decoder Multi-head attention lets a model look at different parts of the input sequence from multiple perspectives simultaneously. •“Attention” = how much focus a word should give to other words •“Multi-head” = multiple attention computations run in parallel, each learning different relationships (semantics, long-range dependencies, etc.) [page 33] Transformer NN Model How is it implemented? We use a triangular mask matrix that blocks attention scores to future positions by setting them to −∞ before softmax. Example mask for 5 tokens: [page 34] Transformer NN Model [page 35] Transformer NN Model