# LSTM Summary Slides course: Module 3 — Deep Learning & NLP module: Module-3-Deep-Learning-NLP type: pdf source_url: https://personal-learn.armco.dev/files/Module-3-Deep-Learning-NLP/General/LSTM_Summary_Slides.pdf pages: 28 --- [page 1] LSTM [page 2] Long Short Term Memory (LSTM) [page 3] Long Short Term Memory (LSTM) [page 4] Long Short Term Memory (LSTM) [page 5] Long Short Term Memory (LSTM) [page 6] Long Short Term Memory (LSTM) [page 7] Long Short Term Memory (LSTM) [page 8] Long Short Term Memory (LSTM) [page 9] Long Short Term Memory (LSTM) [page 10] Long Short Term Memory (LSTM) [page 11] Long Short Term Memory (LSTM) [page 12] Long Short Term Memory (LSTM) [page 13] Long Short Term Memory (LSTM) [page 14] Long Short Term Memory (LSTM) [page 15] Long Short Term Memory (LSTM) [page 16] Long Short Term Memory (LSTM) [page 17] Long Short Term Memory (LSTM) Think of LSTM like a smart memory box with 3 switches. Meaning of Symbols xₜ: Input at time step t hₜ₋₁: Previous hidden state Cₜ₋₁: Previous cell state (long- term memory) U: Weight matrix for input W: Weight matrix for hidden state b: Bias term σ (sigmoid): Outputs values between 0 and 1 tanh: Outputs values between -1 and 1 ⊙: Element-wise multiplication [page 18] Long Short Term Memory (LSTM) Forget Gate — "What should I erase?" 𝒇𝒕 = 𝝈(𝒙𝒕𝑼𝒇 + 𝒉𝒕−𝟏𝑾𝒇 + 𝒃𝒇) • Output between 0 and 1 • Multiplies old memory 𝐶𝑡−1 Cell State Update: Cₜ = fₜ ⊙ Cₜ₋₁ + iₜ ⊙ C̃ₜ If: • 𝑓𝑡 = 1→ keep everything • 𝑓𝑡 = 0→ forget everything • 𝑓𝑡 = 0.5→ keep half This prevents useless old information from staying forever. Meaning of Symbols xₜ: Input at time step t hₜ₋₁: Previous hidden state Cₜ₋₁: Previous cell state (long-term memory) U: Weight matrix for input W: Weight matrix for hidden state b: Bias term σ (sigmoid): Outputs values between 0 and 1 tanh: Outputs values between -1 and 1 ⊙: Element-wise multiplication [page 19] Long Short Term Memory (LSTM) Meaning of Symbols xₜ: Input at time step t hₜ₋₁: Previous hidden state Cₜ₋₁: Previous cell state (long-term memory) U: Weight matrix for input W: Weight matrix for hidden state b: Bias term σ (sigmoid): Outputs values between 0 and 1 tanh: Outputs values between -1 and 1 ⊙: Element-wise multiplication Input Gate — "What new information should I write?" Two parts: (a) Gate decision: 𝒊𝒕 = 𝝈(𝒙𝒕𝑼𝒊 + 𝒉𝒕−𝟏𝑾𝒊 + 𝒃𝒊) (b) Candidate memory: ෩𝑪𝒕 = 𝐭𝐚𝐧𝐡(𝒙𝒕𝑼𝒈 + 𝒉𝒕−𝟏𝑾𝒈 + 𝒃𝒈) • 𝑖𝑡 decides how much to write • ሚ𝐶𝑡contains new candidate information Together they add useful new knowledge. [page 20] Long Short Term Memory (LSTM) Meaning of Symbols xₜ: Input at time step t hₜ₋₁: Previous hidden state Cₜ₋₁: Previous cell state (long-term memory) U: Weight matrix for input W: Weight matrix for hidden state b: Bias term σ (sigmoid): Outputs values between 0 and 1 tanh: Outputs values between -1 and 1 ⊙: Element-wise multiplication Cell State Update — "Combine old + new“ 𝑪𝒕 = 𝒇𝒕 ⊙ 𝑪𝒕−𝟏 + 𝒊𝒕 ⊙ ෩𝑪𝒕 This is the most important equation. First term → kept old memory Second term → new memory added This addition (not multiplication) is why LSTM avoids vanishing gradients. [page 21] Long Short Term Memory (LSTM) Meaning of Symbols xₜ: Input at time step t hₜ₋₁: Previous hidden state Cₜ₋₁: Previous cell state (long-term memory) U: Weight matrix for input W: Weight matrix for hidden state b: Bias term σ (sigmoid): Outputs values between 0 and 1 tanh: Outputs values between -1 and 1 ⊙: Element-wise multiplication Output Gate — "What should I reveal?" 𝒐𝒕 = 𝝈(𝒙𝒕𝑼𝒐 + 𝒉𝒕−𝟏𝑾𝒐 + 𝒃𝒐) 𝒉𝒕 = 𝒐𝒕 ⊙ 𝒕𝒂𝒏𝒉(𝑪𝒕) Controls what part of memory becomes output Hidden state ℎ𝑡is what goes to: Next time step Final prediction layer [page 22] Long Short Term Memory (LSTM) How “Open” vs. “Closed” Helps Gradients Why LSTM Solves Vanishing Gradient Normal RNN: ℎ𝑡 = tanh(𝑊ℎ𝑡−1) → repeated multiplication → gradients vanish LSTM: 𝐶𝑡 = 𝑓𝑡𝐶𝑡−1+. . . → additive update → gradients flow easily That’s why LSTM remembers long-term dependencies. LSTM is mainly designed to solve the vanishing gradient problem, not completely eliminate exploding gradients. But it helps control exploding gradients indirectly. [page 23] Long Short Term Memory (LSTM) What is “Internal Representation”? Inside the LSTM at each time step, it stores: 𝐶𝑡 →long-term memory ℎ𝑡 →short-term/output memory These are vectors, not single numbers. Each element of these vectors is: A learned feature A learned abstract pattern A learned signal detector This vector is what we call the internal representation. [page 24] Long Short Term Memory (LSTM) What is Hidden Size? Hidden size = number of LSTM memory units (neurons) in one LSTM layer. Intuition: What Does Each Hidden Unit Represent? Each hidden unit (one dimension of ℎ𝑡 learns something like: “Is the sentence talking about sports?” “Are we inside quotation marks?” , “Is the subject singular or plural?” “Is the sentiment positive?” We don’t manually define these — the network learns them. So if hidden size = 64: The LSTM is learning 64 different internal signals simultaneously. Why Is Hidden Size Independent of Input Size? Input size depends on data. Example: Word embedding = 300 dimensions → 𝑛𝑥 = 300 But hidden size is a design choice. [page 25] Long Short Term Memory (LSTM) What is Hidden Size? Hidden size = number of LSTM memory units (neurons) in one LSTM layer. Intuition: What Does Each Hidden Unit Represent? Why Bigger Hidden Size = More Power? Because: More hidden units = more memory slots More parameters in weight matrices More expressive representation But: Too small → underfitting Too large → overfitting Too large → slower training [page 26] Long Short Term Memory (LSTM) Example what do we mean by - “4 units of LSTM connected in a NN” It could mean either: • 4 time steps unrolled (same LSTM cell repeated 4 times over a sequence) - most common • hidden size = 4 (the LSTM has 4 memory units inside one time step) [page 27] Long Short Term Memory (LSTM) 4 time steps unrolled (same LSTM cell repeated 4 times over a sequence) - most common [page 28] Long Short Term Memory (LSTM) hidden size = 4 (the LSTM has 4 memory units inside one time step)