# Class Materials- Attention

course: Module 3 — Deep Learning & NLP
module: Module-3-Deep-Learning-NLP
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-3-Deep-Learning-NLP/General/Class_Materials-_Attention.pdf
pages: 44

---
[page 1]
Transformers

[page 2]
LLMs are built out of transformers
Transformer: a specific kind of network architecture, like a 
fancier feedforward network, but based on attention

[page 3]
A very approximate timeline
1990 Static Word Embeddings
2003 Neural Language Model
2008 Multi-Task Learning
2015 Attention
2017 Transformer
2018 Contextual Word Embeddings and Pretraining
2019 Prompting

[page 4]
Transformers
Attention

[page 5]
Instead of starting with the big picture
Stacked
Transformer
Blocks
So long and thanks for
long and thanks forNext token all
…
…
…
U
Input tokens
x1 x2
Language
Modeling
Head
x3 x4 x5
Input
Encoding
 E
1+
E
2+
E
3+
E
4+
E
5+
…
… ………
U
 U
 U
 U
…
logits logits logits logits logits
Stacked
Transformer
Blocks
So long and thanks for
long and thanks forNext token all
…
…
…
U
Input tokens
x1 x2
Language
Modeling
Head
x3 x4 x5
Input
Encoding
 E
1+
E
2+
E
3+
E
4+
E
5+
…
… ………
U
 U
 U
 U
…
logits logits logits logits logits
Let's consider the embeddings for an individual word from a particular layer

[page 6]
Problem with static embeddings (word2vec)
They are static!  The embedding for a word doesn't reflect how its 
meaning changes in context.
 The chicken didn't cross the road because it was too tired
What is the meaning represented in the static embedding for "it"?

[page 7]
Contextual Embeddings
• Intuition: a representation of meaning of a word 
should be different in different contexts!
• Contextual Embedding: each word has a different 
vector that expresses different meanings 
depending on the surrounding words
• How to compute contextual embeddings?
• Attention

[page 8]
Contextual Embeddings
The chicken didn't cross the road because it
What should be the properties of "it"?
The chicken didn't cross the road because it was too tired
The chicken didn't cross the road because it was too wide
At this point in the sentence, it's probably referring to either the chicken or the street
But the static meaning of "it" can't represent that, it just means " non-human pronoun".

[page 9]
Intuition of attention
Build up the contextual embedding from a word by 
selectively integrating information from all the 
neighboring words
We say that a word "attends to" some neighboring 
words more than others

[page 10]
Intuition of attention: 
test

[page 11]
Attention (general mechanism)
Idea: A model doesn’t treat all input information equally. It pays attention to the 
most relevant parts when making a decision.
How it works:
• You have a query (what you’re focusing on),
• A set of keys (possible places to look), and
• Values (the information linked to each key).
Attention computes how much each key is relevant to the query (via similarity), then 
takes a weighted sum of the values.

[page 12]
Attention definition
A mechanism to help compute the embedding for a 
token by selectively attending to and integrating 
information from surrounding tokens (at the previous 
layer).
More formally: a method for doing a weighted sum of 
vectors.

[page 13]
Self-Attention (special case of attention)
Idea: The query, keys, and values all come from the same sequence.
Why: This allows each element (e.g., each word in a sentence) to look at other 
elements and gather context.
How it works:
Every word projects into query, key, and value vectors.
Each word’s query compares with all other words’ keys to decide how much 
attention to pay.
Then it aggregates the weighted values.

[page 14]
Attention is left-to-right
attentionattentionSelf-Attention
Layer
attentionattentionattention
a1 a2 a3 a4 a5
x3 x4 x5x1 x2

[page 15]
Attention Versus Self-Attention  
Attention = general mechanism of focusing on relevant information 
(Q, K, V can be from anywhere).
Self-attention = queries, keys, values are from the same sequence, 
so each token learns from all other tokens in its context.

[page 16]
Self-Attention (special case of attention)
Why Self-Attention Matters
It lets every word directly interact with every other word in the 
sequence.
Captures long-range dependencies better than RNNs (no sequential 
bottleneck).
Scales nicely with parallelization (matrix multiplications).

[page 17]
Simplified version of attention: a sum of prior words 
weighted by their similarity with the current word
Given a sequence of token embeddings:
 x1 x2   x3   x4   x5   x6   x7   xi
Produce: ai = a weighted sum of x1 through x7 (and xi)
Weighted by their similarity to xi
10.1 • THE TRANSFORMER:A SELF-ATTENTION NETWORK 5
Self-Attention
Layer
x1
a1
x2
a2 a3 a4 a5
x3 x4 x5
Figure10.2 Informa tionflow in a ca us a l(or ma s ke d)s e lf-a tte ntionmode l.In proce s s ing
e a che le me ntof thes e que nce ,themode la tte ndsto a lltheinputsup to, a ndincluding,the
curre ntone .Unlike RNNs ,thecomputa tionsat e a chtimes te pa reinde pe nde ntof a llthe
othe rs te psa ndthe re foreca nbepe rforme din pa ra lle l.
10.1.3 S e lf-atte ntionmoreformally
We’vegiventheintuitionof s e lf-a tte ntion(a sawaytocomputere pre s e nta tionsof a
wordata givenla ye rby integra tinginforma tionfromwordsatthepreviousla ye r)
a ndwe ’ve de fine dcontext asa llthepriorwordsin theinput. Le t’s now introduce
thes e lf-a tte ntioncomputa tionits e lf.
Thecoreintuitionof a tte ntionis theide aof com paringanite mof inte re s tto a
colle ctionof othe rite msin awaytha treve a lsthe irre leva ncein thecurre ntcontext.
In theca s eof s e lf-a tte ntionfor la ngua ge ,thes e tof compa ris onsa reto othe rwords
(ortoke ns )withinagivens e que nce .There s ultof the s ecompa ris onsisthe nus e dto
computeanoutputs e que ncefor thecurre ntinputs e que nce .For exa mple ,re turning
to Fig.10.2, thecomputa tionof a3 is ba s e don a s e tof compa ris onsbe twe e nthe
inputx3 a nditspre ce dinge le me ntsx1 a ndx2, a ndtox3 its e lf.
How s ha llwe compa rewordsto othe rwords ?S inceour re pre s e nta tionsfor
wordsa reve ctors ,we ’ll ma ke us eof ourold frie ndthedot producttha tweus e d
for computingwords imila rityin Cha pte r6, a nda ls opla ye da rolein a tte ntionin
Cha pte r9. Le t’s re fe rto there s ultof this compa ris onbe twe e nwordsi a ndj asa
s core(we ’ll beupda ting thise qua tionto a dda tte ntionto thecomputa tionof this
s core ):
Vers on1: s core(xi,xj)= xi ·xj (10.4)
There s ultof a dotproductis a s ca la rva luera ngingfrom − • to • , thela rge r
theva luethemores imila rtheve ctorstha ta rebe ingcompa re d.Continuingwithour
exa mple ,thefirs ts te pin computingy3 wouldbeto computethre es core s :x3 ·x1,
x3 ·x2 a ndx3 ·x3. The nto ma keeffe ctiveus eof the s es core s ,we ’ll norma lizethe m
with a s oftma xto cre a tea ve ctorof we ights ,aij , tha tindica te stheproportiona l
re leva nceof e a chinputtotheinpute le me nti tha tis thecurre ntfocusof a tte ntion.
aij = s oftma x(s core(xi,x j)) 8 j i (10.5)
= exp(s core(xi,xj))P i
k=1exp(s core(xi,xk))
8 j i (10.6)
Ofcours e ,thes oftma xwe ightwill like lybehighe s tforthecurre ntfocuse le me nt
i, s inceve cxi is ve rys imila rto its e lf,re s ultingin a highdot product.But othe r
context wordsma ya ls obes imila rtoi, a ndthes oftma xwill a ls oa s s igns omewe ight
tothos ewords .
Giventheproportiona ls core sin a, wege ne ra teanoutputva lueai bys umming
10.1 • THE TRANSFORMER:A SELF-ATTENTION NETWORK 5
Self-Attention
Layer
x1
a1
x2
a2 a3 a4 a5
x3 x4 x5
Figure10.2 Informa tionflow in a ca us a l(or ma s ke d)s e lf-a tte ntionmode l.In proce s s ing
e a che le me ntof thes e que nce ,themode la tte ndsto a lltheinputsup to, a ndincluding,the
curre ntone .Unlike RNNs ,thecomputa tionsat e a chtime s te pa reinde pe nde ntof a llthe
othe rs te psa ndthe re foreca nbepe rforme din pa ra lle l.
10.1.3 S e lf-atte ntionmoreformally
We’vegiventheintuitionof s e lf-a tte ntion(a sawayto computere pre s e nta tionsof a
wordata givenla ye rby integra tinginforma tionfrom wordsat thepreviousla ye r)
a ndwe ’ve de fine dcontext asa lltheprior wordsin theinput. Le t’s now introduce
thes e lf-a tte ntioncomputa tionits e lf.
Thecoreintuitionof a tte ntionis theide aof com paringanite mof inte re s tto a
colle ctionof othe rite msin awaytha treve a lsthe irre leva ncein thecurre ntcontext.
In theca s eof s e lf-a tte ntionfor la ngua ge ,thes e tof compa ris onsa reto othe rwords
(ortoke ns )withinagivens e que nce .There s ultof the s ecompa ris onsis the nus e dto
computeanoutputs e que ncefor thecurre ntinputs e que nce .For exa mple ,re turning
to Fig. 10.2, thecomputa tionof a3 is ba s e don a s e tof compa ris onsbe twe e nthe
inputx3 a nditspre ce dinge le me ntsx1 a ndx2, a ndto x3 its e lf.
How s ha llwe compa rewordsto othe rwords ?S inceour re pre s e nta tionsfor
wordsa reve ctors ,we ’ll ma ke us eof ourold frie ndthedot producttha tweus e d
for computingwords imila rityin Cha pte r6, a nda ls opla ye da rolein a tte ntionin
Cha pte r9. Le t’s re fe rto there s ultof this compa ris onbe twe e nwordsi a ndj asa
s core(we ’ll beupda ting this e qua tionto a dda tte ntionto thecomputa tionof this
s core ):
Vers on1: s core(xi,x j)= xi ·x j (10.4)
There s ultof a dotproductis a s ca la rva luera ngingfrom − • to • , thela rge r
theva luethemores imila rtheve ctorstha ta rebe ingcompa re d.Continuingwith our
exa mple ,thefirs ts te pin computingy3 wouldbeto computethre es core s :x3 ·x1,
x3 ·x2 a ndx3 ·x3. The nto ma keeffe ctiveus eof the s es core s ,we ’ll norma lizethe m
with a s oftma xto cre a tea ve ctorof we ights ,aij , tha tindica te sthe proportiona l
re leva nceof e a chinputto theinpute le me nti tha tis thecurre ntfocusof a tte ntion.
aij = s oftma x(s core(xi,x j)) 8 j i (10.5)
= exp(s core(xi,x j))P i
k= 1exp(s core(xi,xk))
8 j i (10.6)
Of cours e ,thes oftma xwe ightwill like lybehighe s tforthecurre ntfocuse le me nt
i, s inceve cxi is ve rys imila rto its e lf,re s ultingin a high dot product. But othe r
context wordsma ya ls obes imila rto i, a ndthes oftma xwill a ls oa s s igns omewe ight
to thos ewords .
Giventheproportiona ls core sin a, wege ne ra teanoutputva lueai by s umming

[page 18]
Intuition of attention: 
test
x1  x2  x3  x4  x5  x6  x7   xi

[page 19]
An Actual Attention Head: slightly more complicated
High-level idea: instead of using vectors (like xi and x4) 
directly, we'll represent 3 separate roles each vector xi plays:
• query: As the current element being compared to the 
preceding inputs. 
• key: as a preceding input that is being compared to the 
current element to determine a similarity
• value: a value of a preceding element that gets weighted 
and summed

[page 20]
Attention intuition
x1  x2  x3  x4  x5  x6  x7   xi
query
values

[page 21]
Intuition of attention: 
x1  x2  x3  x4  x5  x6  x7  xi
query
values
k
v
k
v
k
v
k
v
k
v
k
v
k
v
keys k
v

[page 22]
Attention
Attention = general mechanism of focusing on relevant information 
(Q, K, V can be from anywhere).
Self-attention = queries, keys, values are from the same sequence, 
so each token learns from all other tokens in its context.
In Attention, every input token (like a word embedding) is projected into 
three different spaces using learned weight matrices:
• Queries: 𝑄 = 𝑋𝑊𝑄
• Keys: 𝐾 = 𝑋𝑊𝐾
• Values: 𝑉 = 𝑋𝑊𝑉
The values 𝑉are just linear transformations of the input embeddings via 
𝑊𝑉

[page 23]
Attention
In Attention, every input token (like a word embedding) is projected into three different 
spaces using learned weight matrices:
• Queries: 𝑄 = 𝑋𝑊𝑄
• Keys: 𝐾 = 𝑋𝑊𝐾
• Values: 𝑉 = 𝑋𝑊𝑉
The values 𝑉are just linear transformations of the input embeddings via 𝑊𝑉
Analogy
Think of a library search:
• Query (Q): your search phrase.
• Keys (K): book titles in the catalog (used to measure similarity with the query).
• Values (V): the actual content of the books.
You don’t return the book title (the key) — you use the key only to decide which book 
is relevant. What you retrieve is the book’s content (the value).

[page 24]
Attention
Step 1: Attention formula
Attention 𝑄 𝐾 𝑉 = softmax  𝑄𝐾𝑇
𝑑𝑘
𝑉
• 𝑄𝐾𝑇 →similarity scores (how much each token should attend to others).
• Softmax → turns scores into weights (attention distribution).
• Multiplying by 𝑉→ uses those weights to blend value vectors into one output 
vector.

[page 25]
Attention
Step 2: What does “blend” mean?
Each value vector 𝑣𝑗 from each token contains features about that token.
• The softmax gives weights 𝛼𝑗 between 0 and 1, sum to 1.
• The new representation for a token = weighted sum of other tokens’ values:
output𝑖 = ෍
𝑗
𝛼𝑖𝑗 𝑣𝑗
That’s the “combining” — we average the content vectors (values) according 
to relevance weights.

[page 26]
Attention
Tiny numeric example
Suppose we have 2 tokens with value vectors:
•𝑣1 = 1 0  maybe “cat”
•𝑣2 = 0 1  maybe “mat”
For the query “sat,” attention scores (after softmax) might give weights:
𝛼 = 0.7 0.3
Then the output vector is:
output = 0.7 ⋅ 𝑣1 + 0.3 ⋅ 𝑣2 = 0.7 1 0 + 0.3 0 1 = 0.7 0.3
This new embedding for “sat” blends in information from “cat” and “mat.”

[page 27]
Attention  
Keys 𝐾: used only to compute similarity scores with the query (how 
relevant one token is to another).
Values 𝑉: contain the actual information that gets mixed together and 
passed on to the next layer.
After we compute attention weights (via Query–Key dot products → 
softmax), we use those weights to take a weighted average of the Value 
vectors.
Mathematically:
Attention 𝑄 𝐾 𝑉 = softmax  𝑄𝐾𝑇
𝑑𝑘
𝑉
The term inside softmax gives attention weights.
Multiplying by 𝑉means we are literally picking and combining the 
“content” vectors (values).

[page 28]
An Actual Attention Head: slightly more complicated
We'll use matrices to project each vector xi into a 
representation of its role as query, key, value:
• query: WQ
• key: WK
• value: WV

[page 29]
An Actual Attention Head: slightly more complicated
Given these 3 representation of xi
To compute  similarity of current element xi with 
some prior element xj
We’ll use dot product between  qi and kj. 
And instead of summing up xj ,  we'll sum up vj

[page 30]
Final equations for one attention head

[page 31]
Calculating the value of a3
6. Sum the weighted 
value vectors
4. Turn into  i,j weights via softmax
a3
1. Generate 
key, query, value 
vectors
2. Compare x3’s query with
the keys for x1, x2, and x3
Output of self-attention
Wk
Wv
Wq
x1
k
q
v x3
k
q
vx2
k
q
v
×
×
Wk Wk
Wq Wq
WvWv
5. Weigh each value vector
÷
√dk
3. Divide score by √dk
÷
√dk
÷
√dk
 3,1  3,2  3,3

[page 32]
Actual Attention: slightly more complicated
• Instead of one attention head, we'll have lots of them!
• Intuition: each head might be attending to the context for different purposes
• Different linguistic relationships or patterns in the context

[page 33]
Multi-head attention

[page 34]
Model Dimension
What is Model Dimension (d)?
• It is the size of the vector representation for each token at every layer of the Transformer.
• In other words: every word (after tokenization and embedding) is represented by a vector of 
length d.
• This dimension stays the same across all layers of the Transformer (so residual connections 
can be applied easily).
Example
Suppose d = 512 in a Transformer:
◦ Each token (word piece) is represented as a 512-dimensional vector.
◦ If your sequence length is N = 20, the input to the model is a matrix of shape 
20 × 512 (20 tokens, each a 512-dim vector).

[page 35]
Model Dimension
How it connects with Attention
When we split into h heads, the model dimension gets divided:
𝑑𝑘 = 𝑑𝑣 = 𝑑
ℎ
◦ Example: if d = 512 and h = 8, then each head works on 64-dimensional 
vectors since 512 ÷ 8 = 64.
◦ After attention is computed across heads, we concatenate back to 512, and 
project again to stay at dimension d.

[page 36]
Model Dimension
Why keep dimension fixed?
• Consistency: Every layer outputs vectors of dimension d so that they 
can be added to the residual stream.
• Flexibility: Inside the layer, we can branch out into multiple heads, or 
use feedforward layers with larger hidden sizes, but the main "stream" 
always comes back to d.
Summary:
• The model dimension (d) is the length of the embedding vector that 
represents each token throughout the Transformer. 
• It’s the “width” of the model — bigger d means richer representations 
but also higher computational cost.

[page 37]
Model Dimension
Why divide by √dₖ in Attention?
• In scaled dot-product attention, we compute similarity scores as:
• Here dₖ is the dimensionality of the key vectors.
• Without the division, the dot products grow large in magnitude when 
dₖ is big → leading to very large values before softmax, which then 
produces very small gradients (bad for training).
• Dividing by √dₖ normalizes the scale of the dot products, making 
training stable.
In short: √dₖ is just a scaling factor to control the variance of dot 
products so softmax works well

[page 38]
Model Dimension Versus Number of Parameters
How the model dimension (d) affects the number of parameters in a 
Transformer.
1. Embedding Layer
• V ocabulary size = ∣ 𝑉 ∣
• Model dimension = 𝑑
• Embedding matrix size = ∣ 𝑉 ∣× 𝑑
• Example: If ∣ 𝑉 ∣= 30,000  and 𝑑 = 512:
30,000 × 512 = 15.36 million parameters

[page 39]
Model Dimension Versus Number of Parameters
How the model dimension (d) affects the number of parameters in a 
Transformer.
2. Attention Layer
Each attention head has three learned projection matrices:
𝑊𝑄, 𝑊𝐾, 𝑊𝑉 , each of size 𝑑 × 𝑑𝑘
Typically, 𝑑𝑘 = 𝑑/ℎ , and we have ℎ heads.
So total size:
𝑑 × 𝑑 for all heads combined
And one more projection 𝑊𝑂 output, also 𝑑 × 𝑑.
So, attention block ≈ 4 × (d²) parameters from 𝑊𝑄, 𝑊𝐾, 𝑊,𝑉 𝑊𝑂
Example: if 𝑑 = 512, that’s about 1 million parameters just for 
attention.

[page 40]
Model Dimension Versus Number of Parameters
How the model dimension (d) affects the number of parameters in a 
Transformer.
Key Takeaways
•  d (model dimension) is the single biggest factor in scaling model size.
•  Parameters grow as O(d²) because of all the matrix multiplications.
•  Larger d = richer token representations, but also much more 
compute/memory.
Summary:
Embeddings: scale with ∣ 𝑉 ∣× 𝑑.
Attention: scales as 4𝑑2.
Feedforward: scales as 8𝑑2.
Per layer ≈ 12𝑑2.
Full model ≈ 𝐿 × 12𝑑2 +∣ 𝑉 ∣× 𝑑.

[page 41]
Unembedding / Language Modeling Head
• Instead of one attention head, we'll have lots of them!
• Intuition: each head might be attending to the context for different purposes
• Different linguistic relationships or patterns in the context

[page 42]
Unembedding / Language Modeling Head
What is W₀?
• After the transformer layers, we have a hidden vector of size 1 × d (for one 
token).
• To convert this into vocabulary logits (size 1 × |V|, where |V| is vocab size), we 
multiply by a matrix W₀.
• W₀ is the unembedding weight matrix (same as the embedding matrix Eᵀ in 
weight tying).
• Embedding: maps vocab index → d-dimensional vector.
• Unembedding (W₀): maps d-dimensional hidden vector → vocab logits.
• In short: W₀ is the weight matrix used to project the hidden state back into 
vocabulary space

[page 43]
Multi-head attention output projection
From 1 × h· dᵥ → 1 × d
• This refers to multi-head attention output projection.
• Each attention head produces an output vector of size 1 × dᵥ (value 
dimension).
• If there are h heads, concatenating them gives size 1 × (h·dᵥ).
• But the model needs the output to remain in the same dimension d 
(for residual connections).
• So we multiply by a learned projection matrix Wᴼ of shape (h·dᵥ × 
d) to map it back.
• In short: Concatenated head outputs (1 × h·dᵥ) are linearly 
projected by Wᴼ into size 1 × d, preserving the model dimension

[page 44]
Summary
Attention is a method for enriching the representation of a token by 
incorporating contextual information
The result: the embedding for each word will be different in different 
contexts!
Contextual embeddings: a representation of word meaning in its 
context.