# Class Materials- Attention course: Module 3 — Deep Learning & NLP module: Module-3-Deep-Learning-NLP type: pdf source_url: https://personal-learn.armco.dev/files/Module-3-Deep-Learning-NLP/General/Class_Materials-_Attention.pdf pages: 44 --- [page 1] Transformers [page 2] LLMs are built out of transformers Transformer: a specific kind of network architecture, like a fancier feedforward network, but based on attention [page 3] A very approximate timeline 1990 Static Word Embeddings 2003 Neural Language Model 2008 Multi-Task Learning 2015 Attention 2017 Transformer 2018 Contextual Word Embeddings and Pretraining 2019 Prompting [page 4] Transformers Attention [page 5] Instead of starting with the big picture Stacked Transformer Blocks So long and thanks for long and thanks forNext token all … … … U Input tokens x1 x2 Language Modeling Head x3 x4 x5 Input Encoding E 1+ E 2+ E 3+ E 4+ E 5+ … … ……… U U U U … logits logits logits logits logits Stacked Transformer Blocks So long and thanks for long and thanks forNext token all … … … U Input tokens x1 x2 Language Modeling Head x3 x4 x5 Input Encoding E 1+ E 2+ E 3+ E 4+ E 5+ … … ……… U U U U … logits logits logits logits logits Let's consider the embeddings for an individual word from a particular layer [page 6] Problem with static embeddings (word2vec) They are static! The embedding for a word doesn't reflect how its meaning changes in context. The chicken didn't cross the road because it was too tired What is the meaning represented in the static embedding for "it"? [page 7] Contextual Embeddings • Intuition: a representation of meaning of a word should be different in different contexts! • Contextual Embedding: each word has a different vector that expresses different meanings depending on the surrounding words • How to compute contextual embeddings? • Attention [page 8] Contextual Embeddings The chicken didn't cross the road because it What should be the properties of "it"? The chicken didn't cross the road because it was too tired The chicken didn't cross the road because it was too wide At this point in the sentence, it's probably referring to either the chicken or the street But the static meaning of "it" can't represent that, it just means " non-human pronoun". [page 9] Intuition of attention Build up the contextual embedding from a word by selectively integrating information from all the neighboring words We say that a word "attends to" some neighboring words more than others [page 10] Intuition of attention: test [page 11] Attention (general mechanism) Idea: A model doesn’t treat all input information equally. It pays attention to the most relevant parts when making a decision. How it works: • You have a query (what you’re focusing on), • A set of keys (possible places to look), and • Values (the information linked to each key). Attention computes how much each key is relevant to the query (via similarity), then takes a weighted sum of the values. [page 12] Attention definition A mechanism to help compute the embedding for a token by selectively attending to and integrating information from surrounding tokens (at the previous layer). More formally: a method for doing a weighted sum of vectors. [page 13] Self-Attention (special case of attention) Idea: The query, keys, and values all come from the same sequence. Why: This allows each element (e.g., each word in a sentence) to look at other elements and gather context. How it works: Every word projects into query, key, and value vectors. Each word’s query compares with all other words’ keys to decide how much attention to pay. Then it aggregates the weighted values. [page 14] Attention is left-to-right attentionattentionSelf-Attention Layer attentionattentionattention a1 a2 a3 a4 a5 x3 x4 x5x1 x2 [page 15] Attention Versus Self-Attention Attention = general mechanism of focusing on relevant information (Q, K, V can be from anywhere). Self-attention = queries, keys, values are from the same sequence, so each token learns from all other tokens in its context. [page 16] Self-Attention (special case of attention) Why Self-Attention Matters It lets every word directly interact with every other word in the sequence. Captures long-range dependencies better than RNNs (no sequential bottleneck). Scales nicely with parallelization (matrix multiplications). [page 17] Simplified version of attention: a sum of prior words weighted by their similarity with the current word Given a sequence of token embeddings: x1 x2 x3 x4 x5 x6 x7 xi Produce: ai = a weighted sum of x1 through x7 (and xi) Weighted by their similarity to xi 10.1 • THE TRANSFORMER:A SELF-ATTENTION NETWORK 5 Self-Attention Layer x1 a1 x2 a2 a3 a4 a5 x3 x4 x5 Figure10.2 Informa tionflow in a ca us a l(or ma s ke d)s e lf-a tte ntionmode l.In proce s s ing e a che le me ntof thes e que nce ,themode la tte ndsto a lltheinputsup to, a ndincluding,the curre ntone .Unlike RNNs ,thecomputa tionsat e a chtimes te pa reinde pe nde ntof a llthe othe rs te psa ndthe re foreca nbepe rforme din pa ra lle l. 10.1.3 S e lf-atte ntionmoreformally We’vegiventheintuitionof s e lf-a tte ntion(a sawaytocomputere pre s e nta tionsof a wordata givenla ye rby integra tinginforma tionfromwordsatthepreviousla ye r) a ndwe ’ve de fine dcontext asa llthepriorwordsin theinput. Le t’s now introduce thes e lf-a tte ntioncomputa tionits e lf. Thecoreintuitionof a tte ntionis theide aof com paringanite mof inte re s tto a colle ctionof othe rite msin awaytha treve a lsthe irre leva ncein thecurre ntcontext. In theca s eof s e lf-a tte ntionfor la ngua ge ,thes e tof compa ris onsa reto othe rwords (ortoke ns )withinagivens e que nce .There s ultof the s ecompa ris onsisthe nus e dto computeanoutputs e que ncefor thecurre ntinputs e que nce .For exa mple ,re turning to Fig.10.2, thecomputa tionof a3 is ba s e don a s e tof compa ris onsbe twe e nthe inputx3 a nditspre ce dinge le me ntsx1 a ndx2, a ndtox3 its e lf. How s ha llwe compa rewordsto othe rwords ?S inceour re pre s e nta tionsfor wordsa reve ctors ,we ’ll ma ke us eof ourold frie ndthedot producttha tweus e d for computingwords imila rityin Cha pte r6, a nda ls opla ye da rolein a tte ntionin Cha pte r9. Le t’s re fe rto there s ultof this compa ris onbe twe e nwordsi a ndj asa s core(we ’ll beupda ting thise qua tionto a dda tte ntionto thecomputa tionof this s core ): Vers on1: s core(xi,xj)= xi ·xj (10.4) There s ultof a dotproductis a s ca la rva luera ngingfrom − • to • , thela rge r theva luethemores imila rtheve ctorstha ta rebe ingcompa re d.Continuingwithour exa mple ,thefirs ts te pin computingy3 wouldbeto computethre es core s :x3 ·x1, x3 ·x2 a ndx3 ·x3. The nto ma keeffe ctiveus eof the s es core s ,we ’ll norma lizethe m with a s oftma xto cre a tea ve ctorof we ights ,aij , tha tindica te stheproportiona l re leva nceof e a chinputtotheinpute le me nti tha tis thecurre ntfocusof a tte ntion. aij = s oftma x(s core(xi,x j)) 8 j i (10.5) = exp(s core(xi,xj))P i k=1exp(s core(xi,xk)) 8 j i (10.6) Ofcours e ,thes oftma xwe ightwill like lybehighe s tforthecurre ntfocuse le me nt i, s inceve cxi is ve rys imila rto its e lf,re s ultingin a highdot product.But othe r context wordsma ya ls obes imila rtoi, a ndthes oftma xwill a ls oa s s igns omewe ight tothos ewords . Giventheproportiona ls core sin a, wege ne ra teanoutputva lueai bys umming 10.1 • THE TRANSFORMER:A SELF-ATTENTION NETWORK 5 Self-Attention Layer x1 a1 x2 a2 a3 a4 a5 x3 x4 x5 Figure10.2 Informa tionflow in a ca us a l(or ma s ke d)s e lf-a tte ntionmode l.In proce s s ing e a che le me ntof thes e que nce ,themode la tte ndsto a lltheinputsup to, a ndincluding,the curre ntone .Unlike RNNs ,thecomputa tionsat e a chtime s te pa reinde pe nde ntof a llthe othe rs te psa ndthe re foreca nbepe rforme din pa ra lle l. 10.1.3 S e lf-atte ntionmoreformally We’vegiventheintuitionof s e lf-a tte ntion(a sawayto computere pre s e nta tionsof a wordata givenla ye rby integra tinginforma tionfrom wordsat thepreviousla ye r) a ndwe ’ve de fine dcontext asa lltheprior wordsin theinput. Le t’s now introduce thes e lf-a tte ntioncomputa tionits e lf. Thecoreintuitionof a tte ntionis theide aof com paringanite mof inte re s tto a colle ctionof othe rite msin awaytha treve a lsthe irre leva ncein thecurre ntcontext. In theca s eof s e lf-a tte ntionfor la ngua ge ,thes e tof compa ris onsa reto othe rwords (ortoke ns )withinagivens e que nce .There s ultof the s ecompa ris onsis the nus e dto computeanoutputs e que ncefor thecurre ntinputs e que nce .For exa mple ,re turning to Fig. 10.2, thecomputa tionof a3 is ba s e don a s e tof compa ris onsbe twe e nthe inputx3 a nditspre ce dinge le me ntsx1 a ndx2, a ndto x3 its e lf. How s ha llwe compa rewordsto othe rwords ?S inceour re pre s e nta tionsfor wordsa reve ctors ,we ’ll ma ke us eof ourold frie ndthedot producttha tweus e d for computingwords imila rityin Cha pte r6, a nda ls opla ye da rolein a tte ntionin Cha pte r9. Le t’s re fe rto there s ultof this compa ris onbe twe e nwordsi a ndj asa s core(we ’ll beupda ting this e qua tionto a dda tte ntionto thecomputa tionof this s core ): Vers on1: s core(xi,x j)= xi ·x j (10.4) There s ultof a dotproductis a s ca la rva luera ngingfrom − • to • , thela rge r theva luethemores imila rtheve ctorstha ta rebe ingcompa re d.Continuingwith our exa mple ,thefirs ts te pin computingy3 wouldbeto computethre es core s :x3 ·x1, x3 ·x2 a ndx3 ·x3. The nto ma keeffe ctiveus eof the s es core s ,we ’ll norma lizethe m with a s oftma xto cre a tea ve ctorof we ights ,aij , tha tindica te sthe proportiona l re leva nceof e a chinputto theinpute le me nti tha tis thecurre ntfocusof a tte ntion. aij = s oftma x(s core(xi,x j)) 8 j i (10.5) = exp(s core(xi,x j))P i k= 1exp(s core(xi,xk)) 8 j i (10.6) Of cours e ,thes oftma xwe ightwill like lybehighe s tforthecurre ntfocuse le me nt i, s inceve cxi is ve rys imila rto its e lf,re s ultingin a high dot product. But othe r context wordsma ya ls obes imila rto i, a ndthes oftma xwill a ls oa s s igns omewe ight to thos ewords . Giventheproportiona ls core sin a, wege ne ra teanoutputva lueai by s umming [page 18] Intuition of attention: test x1 x2 x3 x4 x5 x6 x7 xi [page 19] An Actual Attention Head: slightly more complicated High-level idea: instead of using vectors (like xi and x4) directly, we'll represent 3 separate roles each vector xi plays: • query: As the current element being compared to the preceding inputs. • key: as a preceding input that is being compared to the current element to determine a similarity • value: a value of a preceding element that gets weighted and summed [page 20] Attention intuition x1 x2 x3 x4 x5 x6 x7 xi query values [page 21] Intuition of attention: x1 x2 x3 x4 x5 x6 x7 xi query values k v k v k v k v k v k v k v keys k v [page 22] Attention Attention = general mechanism of focusing on relevant information (Q, K, V can be from anywhere). Self-attention = queries, keys, values are from the same sequence, so each token learns from all other tokens in its context. In Attention, every input token (like a word embedding) is projected into three different spaces using learned weight matrices: • Queries: 𝑄 = 𝑋𝑊𝑄 • Keys: 𝐾 = 𝑋𝑊𝐾 • Values: 𝑉 = 𝑋𝑊𝑉 The values 𝑉are just linear transformations of the input embeddings via 𝑊𝑉 [page 23] Attention In Attention, every input token (like a word embedding) is projected into three different spaces using learned weight matrices: • Queries: 𝑄 = 𝑋𝑊𝑄 • Keys: 𝐾 = 𝑋𝑊𝐾 • Values: 𝑉 = 𝑋𝑊𝑉 The values 𝑉are just linear transformations of the input embeddings via 𝑊𝑉 Analogy Think of a library search: • Query (Q): your search phrase. • Keys (K): book titles in the catalog (used to measure similarity with the query). • Values (V): the actual content of the books. You don’t return the book title (the key) — you use the key only to decide which book is relevant. What you retrieve is the book’s content (the value). [page 24] Attention Step 1: Attention formula Attention 𝑄 𝐾 𝑉 = softmax 𝑄𝐾𝑇 𝑑𝑘 𝑉 • 𝑄𝐾𝑇 →similarity scores (how much each token should attend to others). • Softmax → turns scores into weights (attention distribution). • Multiplying by 𝑉→ uses those weights to blend value vectors into one output vector. [page 25] Attention Step 2: What does “blend” mean? Each value vector 𝑣𝑗 from each token contains features about that token. • The softmax gives weights 𝛼𝑗 between 0 and 1, sum to 1. • The new representation for a token = weighted sum of other tokens’ values: output𝑖 = 𝑗 𝛼𝑖𝑗 𝑣𝑗 That’s the “combining” — we average the content vectors (values) according to relevance weights. [page 26] Attention Tiny numeric example Suppose we have 2 tokens with value vectors: •𝑣1 = 1 0 maybe “cat” •𝑣2 = 0 1 maybe “mat” For the query “sat,” attention scores (after softmax) might give weights: 𝛼 = 0.7 0.3 Then the output vector is: output = 0.7 ⋅ 𝑣1 + 0.3 ⋅ 𝑣2 = 0.7 1 0 + 0.3 0 1 = 0.7 0.3 This new embedding for “sat” blends in information from “cat” and “mat.” [page 27] Attention Keys 𝐾: used only to compute similarity scores with the query (how relevant one token is to another). Values 𝑉: contain the actual information that gets mixed together and passed on to the next layer. After we compute attention weights (via Query–Key dot products → softmax), we use those weights to take a weighted average of the Value vectors. Mathematically: Attention 𝑄 𝐾 𝑉 = softmax 𝑄𝐾𝑇 𝑑𝑘 𝑉 The term inside softmax gives attention weights. Multiplying by 𝑉means we are literally picking and combining the “content” vectors (values). [page 28] An Actual Attention Head: slightly more complicated We'll use matrices to project each vector xi into a representation of its role as query, key, value: • query: WQ • key: WK • value: WV [page 29] An Actual Attention Head: slightly more complicated Given these 3 representation of xi To compute similarity of current element xi with some prior element xj We’ll use dot product between qi and kj. And instead of summing up xj , we'll sum up vj [page 30] Final equations for one attention head [page 31] Calculating the value of a3 6. Sum the weighted value vectors 4. Turn into i,j weights via softmax a3 1. Generate key, query, value vectors 2. Compare x3’s query with the keys for x1, x2, and x3 Output of self-attention Wk Wv Wq x1 k q v x3 k q vx2 k q v × × Wk Wk Wq Wq WvWv 5. Weigh each value vector ÷ √dk 3. Divide score by √dk ÷ √dk ÷ √dk 3,1 3,2 3,3 [page 32] Actual Attention: slightly more complicated • Instead of one attention head, we'll have lots of them! • Intuition: each head might be attending to the context for different purposes • Different linguistic relationships or patterns in the context [page 33] Multi-head attention [page 34] Model Dimension What is Model Dimension (d)? • It is the size of the vector representation for each token at every layer of the Transformer. • In other words: every word (after tokenization and embedding) is represented by a vector of length d. • This dimension stays the same across all layers of the Transformer (so residual connections can be applied easily). Example Suppose d = 512 in a Transformer: ◦ Each token (word piece) is represented as a 512-dimensional vector. ◦ If your sequence length is N = 20, the input to the model is a matrix of shape 20 × 512 (20 tokens, each a 512-dim vector). [page 35] Model Dimension How it connects with Attention When we split into h heads, the model dimension gets divided: 𝑑𝑘 = 𝑑𝑣 = 𝑑 ℎ ◦ Example: if d = 512 and h = 8, then each head works on 64-dimensional vectors since 512 ÷ 8 = 64. ◦ After attention is computed across heads, we concatenate back to 512, and project again to stay at dimension d. [page 36] Model Dimension Why keep dimension fixed? • Consistency: Every layer outputs vectors of dimension d so that they can be added to the residual stream. • Flexibility: Inside the layer, we can branch out into multiple heads, or use feedforward layers with larger hidden sizes, but the main "stream" always comes back to d. Summary: • The model dimension (d) is the length of the embedding vector that represents each token throughout the Transformer. • It’s the “width” of the model — bigger d means richer representations but also higher computational cost. [page 37] Model Dimension Why divide by √dₖ in Attention? • In scaled dot-product attention, we compute similarity scores as: • Here dₖ is the dimensionality of the key vectors. • Without the division, the dot products grow large in magnitude when dₖ is big → leading to very large values before softmax, which then produces very small gradients (bad for training). • Dividing by √dₖ normalizes the scale of the dot products, making training stable. In short: √dₖ is just a scaling factor to control the variance of dot products so softmax works well [page 38] Model Dimension Versus Number of Parameters How the model dimension (d) affects the number of parameters in a Transformer. 1. Embedding Layer • V ocabulary size = ∣ 𝑉 ∣ • Model dimension = 𝑑 • Embedding matrix size = ∣ 𝑉 ∣× 𝑑 • Example: If ∣ 𝑉 ∣= 30,000 and 𝑑 = 512: 30,000 × 512 = 15.36 million parameters [page 39] Model Dimension Versus Number of Parameters How the model dimension (d) affects the number of parameters in a Transformer. 2. Attention Layer Each attention head has three learned projection matrices: 𝑊𝑄, 𝑊𝐾, 𝑊𝑉 , each of size 𝑑 × 𝑑𝑘 Typically, 𝑑𝑘 = 𝑑/ℎ , and we have ℎ heads. So total size: 𝑑 × 𝑑 for all heads combined And one more projection 𝑊𝑂 output, also 𝑑 × 𝑑. So, attention block ≈ 4 × (d²) parameters from 𝑊𝑄, 𝑊𝐾, 𝑊,𝑉 𝑊𝑂 Example: if 𝑑 = 512, that’s about 1 million parameters just for attention. [page 40] Model Dimension Versus Number of Parameters How the model dimension (d) affects the number of parameters in a Transformer. Key Takeaways • d (model dimension) is the single biggest factor in scaling model size. • Parameters grow as O(d²) because of all the matrix multiplications. • Larger d = richer token representations, but also much more compute/memory. Summary: Embeddings: scale with ∣ 𝑉 ∣× 𝑑. Attention: scales as 4𝑑2. Feedforward: scales as 8𝑑2. Per layer ≈ 12𝑑2. Full model ≈ 𝐿 × 12𝑑2 +∣ 𝑉 ∣× 𝑑. [page 41] Unembedding / Language Modeling Head • Instead of one attention head, we'll have lots of them! • Intuition: each head might be attending to the context for different purposes • Different linguistic relationships or patterns in the context [page 42] Unembedding / Language Modeling Head What is W₀? • After the transformer layers, we have a hidden vector of size 1 × d (for one token). • To convert this into vocabulary logits (size 1 × |V|, where |V| is vocab size), we multiply by a matrix W₀. • W₀ is the unembedding weight matrix (same as the embedding matrix Eᵀ in weight tying). • Embedding: maps vocab index → d-dimensional vector. • Unembedding (W₀): maps d-dimensional hidden vector → vocab logits. • In short: W₀ is the weight matrix used to project the hidden state back into vocabulary space [page 43] Multi-head attention output projection From 1 × h· dᵥ → 1 × d • This refers to multi-head attention output projection. • Each attention head produces an output vector of size 1 × dᵥ (value dimension). • If there are h heads, concatenating them gives size 1 × (h·dᵥ). • But the model needs the output to remain in the same dimension d (for residual connections). • So we multiply by a learned projection matrix Wᴼ of shape (h·dᵥ × d) to map it back. • In short: Concatenated head outputs (1 × h·dᵥ) are linearly projected by Wᴼ into size 1 × d, preserving the model dimension [page 44] Summary Attention is a method for enriching the representation of a token by incorporating contextual information The result: the embedding for each word will be different in different contexts! Contextual embeddings: a representation of word meaning in its context.