# LSTM Summary Slides

course: Module 3 — Deep Learning & NLP
module: Module-3-Deep-Learning-NLP
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-3-Deep-Learning-NLP/General/LSTM_Summary_Slides.pdf
pages: 28

---
[page 1]
LSTM

[page 2]
Long Short Term Memory (LSTM)

[page 3]
Long Short Term Memory (LSTM)

[page 4]
Long Short Term Memory (LSTM)

[page 5]
Long Short Term Memory (LSTM)

[page 6]
Long Short Term Memory (LSTM)

[page 7]
Long Short Term Memory (LSTM)

[page 8]
Long Short Term Memory (LSTM)

[page 9]
Long Short Term Memory (LSTM)

[page 10]
Long Short Term Memory (LSTM)

[page 11]
Long Short Term Memory (LSTM)

[page 12]
Long Short Term Memory (LSTM)

[page 13]
Long Short Term Memory (LSTM)

[page 14]
Long Short Term Memory (LSTM)

[page 15]
Long Short Term Memory (LSTM)

[page 16]
Long Short Term Memory (LSTM)

[page 17]
Long Short Term Memory (LSTM)
Think of LSTM like a smart memory box 
with 3 switches.
Meaning of Symbols
xₜ: Input at time step t
hₜ₋₁: Previous hidden state
Cₜ₋₁: Previous cell state (long-
term memory)
U: Weight matrix for input
W: Weight matrix for hidden state
b: Bias term
σ (sigmoid): Outputs values 
between 0 and 1
tanh: Outputs values between -1 
and 1
⊙: Element-wise multiplication

[page 18]
Long Short Term Memory (LSTM)
Forget Gate — "What should I erase?"
𝒇𝒕 = 𝝈(𝒙𝒕𝑼𝒇 + 𝒉𝒕−𝟏𝑾𝒇 + 𝒃𝒇)
• Output between 0 and 1 
• Multiplies old memory 𝐶𝑡−1
Cell State Update: Cₜ = fₜ ⊙ Cₜ₋₁ + iₜ ⊙ C̃ₜ
If:
• 𝑓𝑡 = 1→ keep everything 
• 𝑓𝑡 = 0→ forget everything 
• 𝑓𝑡 = 0.5→ keep half 
This prevents useless old information from staying 
forever.
Meaning of Symbols
xₜ: Input at time step t
hₜ₋₁: Previous hidden state
Cₜ₋₁: Previous cell state (long-term 
memory)
U: Weight matrix for input
W: Weight matrix for hidden state
b: Bias term
σ (sigmoid): Outputs values between 
0 and 1
tanh: Outputs values between -1 and 1
⊙: Element-wise multiplication

[page 19]
Long Short Term Memory (LSTM)
Meaning of Symbols
xₜ: Input at time step t
hₜ₋₁: Previous hidden state
Cₜ₋₁: Previous cell state (long-term 
memory)
U: Weight matrix for input
W: Weight matrix for hidden state
b: Bias term
σ (sigmoid): Outputs values between 0 
and 1
tanh: Outputs values between -1 and 1
⊙: Element-wise multiplication
Input Gate — "What new information 
should I write?"
Two parts:
(a) Gate decision:
𝒊𝒕 = 𝝈(𝒙𝒕𝑼𝒊 + 𝒉𝒕−𝟏𝑾𝒊 + 𝒃𝒊)
(b) Candidate memory:
෩𝑪𝒕 = 𝐭𝐚𝐧𝐡(𝒙𝒕𝑼𝒈 + 𝒉𝒕−𝟏𝑾𝒈 + 𝒃𝒈)
• 𝑖𝑡 decides how much to write 
• ሚ𝐶𝑡contains new candidate information 
Together they add useful new knowledge.

[page 20]
Long Short Term Memory (LSTM)
Meaning of Symbols
xₜ: Input at time step t
hₜ₋₁: Previous hidden state
Cₜ₋₁: Previous cell state (long-term 
memory)
U: Weight matrix for input
W: Weight matrix for hidden state
b: Bias term
σ (sigmoid): Outputs values between 0 
and 1
tanh: Outputs values between -1 and 1
⊙: Element-wise multiplication
Cell State Update — "Combine 
old + new“
𝑪𝒕 = 𝒇𝒕 ⊙ 𝑪𝒕−𝟏 + 𝒊𝒕 ⊙ ෩𝑪𝒕
This is the most important equation.
First term → kept old memory 
Second term → new memory added 
This addition (not multiplication) is why 
LSTM avoids vanishing gradients.

[page 21]
Long Short Term Memory (LSTM)
Meaning of Symbols
xₜ: Input at time step t
hₜ₋₁: Previous hidden state
Cₜ₋₁: Previous cell state (long-term 
memory)
U: Weight matrix for input
W: Weight matrix for hidden state
b: Bias term
σ (sigmoid): Outputs values between 0 
and 1
tanh: Outputs values between -1 and 1
⊙: Element-wise multiplication
Output Gate — "What should I 
reveal?"
𝒐𝒕 = 𝝈(𝒙𝒕𝑼𝒐 + 𝒉𝒕−𝟏𝑾𝒐 + 𝒃𝒐)
𝒉𝒕 = 𝒐𝒕 ⊙ 𝒕𝒂𝒏𝒉(𝑪𝒕)
Controls what part of memory 
becomes output 
Hidden state ℎ𝑡is what goes to: 
Next time step 
Final prediction layer

[page 22]
Long Short Term Memory (LSTM)
How “Open” vs. “Closed” Helps Gradients
Why LSTM Solves Vanishing Gradient
Normal RNN:
ℎ𝑡 = tanh(𝑊ℎ𝑡−1)
→ repeated multiplication → gradients vanish
LSTM:
𝐶𝑡 = 𝑓𝑡𝐶𝑡−1+. . .
→ additive update → gradients flow easily
That’s why LSTM remembers long-term dependencies.
LSTM is mainly designed to solve the vanishing gradient problem, not completely 
eliminate exploding gradients.
But it helps control exploding gradients indirectly.

[page 23]
Long Short Term Memory (LSTM)
What is “Internal Representation”?
Inside the LSTM at each time step, it stores:
𝐶𝑡 →long-term memory 
ℎ𝑡 →short-term/output memory 
These are vectors, not single numbers.
Each element of these vectors is:
A learned feature 
A learned abstract pattern 
A learned signal detector 
This vector is what we call the internal representation.

[page 24]
Long Short Term Memory (LSTM)
What is Hidden Size?
Hidden size = number of LSTM memory units (neurons) in one LSTM layer.
Intuition: What Does Each Hidden Unit Represent?
Each hidden unit (one dimension of ℎ𝑡 learns something like:
“Is the sentence talking about sports?” 
“Are we inside quotation marks?” , “Is the subject singular or plural?” “Is the sentiment positive?” 
We don’t manually define these — the network learns them.
So if hidden size = 64: The LSTM is learning 64 different internal signals 
simultaneously.
Why Is Hidden Size Independent of Input Size?
Input size depends on data.     Example: Word embedding = 300 dimensions → 𝑛𝑥 = 300
But hidden size is a design choice.

[page 25]
Long Short Term Memory (LSTM)
What is Hidden Size?
Hidden size = number of LSTM memory units (neurons) in one LSTM layer.
Intuition: What Does Each Hidden Unit Represent?
Why Bigger Hidden Size = More Power?
Because:
More hidden units = more memory slots 
More parameters in weight matrices 
More expressive representation 
But:
Too small → underfitting 
Too large → overfitting 
Too large → slower training

[page 26]
Long Short Term Memory (LSTM)
Example  what do we mean by - “4 units of LSTM connected in a NN” 
It could mean either:
• 4 time steps unrolled (same LSTM cell repeated 4 times over a sequence) - most 
common 
• hidden size = 4 (the LSTM has 4 memory units inside one time step)

[page 27]
Long Short Term Memory (LSTM)
4 time steps unrolled (same LSTM cell repeated 4 times over a sequence) - most common

[page 28]
Long Short Term 
Memory (LSTM)
hidden size = 4
(the LSTM has 4 
memory units inside 
one time step)