# Seq2seq Encoder-Decoder 2 May course: Module 4 — Generative AI & LLMs module: Module-4-Generative-AI-LLMs type: pdf source_url: https://personal-learn.armco.dev/files/Module-4-Generative-AI-LLMs/General/Seq2seq_Encoder-Decoder_2_May.pdf pages: 41 --- [page 1] Seq2seq and Encoder Decoder Neural Network Models [page 2] What is Encoder–Decoder? • It’s a sequence-to-sequence (seq2seq) architecture. • The encoder processes an input sequence (e.g., a sentence in French) into a hidden representation. • The decoder generates an output sequence (e.g., the same sentence in English) from that hidden representation. Seq2Seq and Encoder Decoder NN Models [page 3] Seq2Seq and Encoder Decoder NN Models The Encoder • Usually an RNN or LSTM (rarely GRU). • Reads the input sequence one token at a time: 𝑥1, 𝑥2, … , 𝑥𝑇 Produces hidden states ℎ𝑡. • Final hidden state (or cell state in LSTM) becomes the context vector. • Example: after “Je suis étudiant”, the encoder condenses the meaning into one vector. [page 4] Seq2Seq and Encoder Decoder NN Models The Decoder • Another RNN / LSTM. • Takes the context vector from the encoder as its initial hidden state. • Generates the output sequence token by token: 𝑦1, 𝑦2, … , 𝑦𝑇 • At each step, it predicts the next token based on: • Its current hidden state • The previous token it generated [page 5] Seq2Seq and Encoder Decoder NN Models Training During training, we often use teacher forcing: • Instead of feeding the decoder its own previous prediction, we give it the true previous word. • This speeds up learning. Example Task: Translate “I am a student” into French. 1.Encoder reads tokens one by one → builds context vector = “meaning of sentence.” 2.Decoder starts with that context vector → generates: 1. First word: “Je” 2. Second word: “suis” 3. Third word: “étudiant” [page 6] Seq2Seq and Encoder Decoder NN Models Summary: Encoder: Reads and compresses the input sequence → context vector. Decoder: Expands that vector into an output sequence, step by step. Together: Encoder–Decoder = the foundation of seq2seq models. [page 7] Seq2Seq and Encoder Decoder NN Models The Problem In a sequence model (like an LSTM decoder or Transformer decoder), each new word is generated based on the previous word. • At inference time, the model doesn’t know the correct next word — it must rely on its own previous prediction. • But during training, we do know the correct sequence (ground truth). • If we always feed the model’s own predictions during training, errors would accumulate quickly (a mistake at step 1 messes up steps 2, 3, …). [page 8] Seq2Seq and Encoder Decoder NN Models What Teacher Forcing Does Instead of feeding the decoder its own previous output, we feed it the true previous token from the dataset. So at timestep 𝑡, input to the decoder is 𝑦𝑡−1 true rather than 𝑦𝑡−1 predicted. This makes training much more stable and faster. [page 9] Seq2Seq and Encoder Decoder NN Models Trade-off Pros: Speeds up training, helps the model learn correct dependencies. Cons: Creates a train–test mismatch during inference, the model has to use its own predictions, which it didn’t practice much during training. Researchers sometimes use scheduled sampling — gradually replacing ground truth with model predictions during training to make it more robust. [page 10] Seq2Seq and Encoder Decoder NN Models [page 11] Seq2Seq and Encoder Decoder NN Models Pro Problem 1 : To convert a sentence in English to Spanish Problem 2 : To convert amino acid sequences into 3D structures like alpha - helices Both are [page 12] Seq2Seq and Encoder Decoder NN Models Pro say, we are interested to convert [page 13] Seq2Seq and Encoder Decoder NN Models Pro [page 14] Seq2Seq and Encoder Decoder NN Models ProWhat an LSTM Unit Is • An LSTM unit is the recurrent cell that processes one timestep of input. • It has gates (input, forget, output) and internal memory. • It takes input 𝑥𝑡at time 𝑡, plus the hidden state ℎ𝑡−1and cell state 𝑐𝑡−1from the previous step, and outputs ℎ𝑡, 𝑐𝑡. [page 15] Pro What “Unrolling” Means • “Unrolling” does not mean we are adding new LSTM cells. • It means: showing how the same LSTM cell, with the same parameters is applied repeatedly across time steps. • For example, if the sequence has 5 words: 𝑥1, 𝑥2, 𝑥3, 𝑥4, 𝑥5 • Then the LSTM will be unrolled 5 times, one for each timestep: LSTM 𝑥1 → LSTM 𝑥2 → ⋯ → LSTM 𝑥5 • Each “copy” in such a diagram is the same LSTM unit re-used with the same weights. • What changes are the inputs (𝑥𝑡 (and the hidden/cell states flowing through. • Unrolling is just shown for visualization: to make clear how hidden states flow across timesteps. • “Unrolling an LSTM” means representing the repeated application of the same LSTM unit across timesteps of a sequence. It doesn’t mean adding new units; the parameters are shared across all steps. [page 16] Seq2Seq and Encoder Decoder NN Models Pro [page 17] Seq2Seq and Encoder Decoder NN Models Pro [page 18] Seq2Seq and Encoder Decoder NN Models Pro [page 19] Seq2Seq and Encoder Decoder NN Models Pro [page 20] Seq2Seq and Encoder Decoder NN Models Pro [page 21] Seq2Seq and Encoder Decoder NN Models Pro [page 22] Seq2Seq and Encoder Decoder NN Models Pro [page 23] Seq2Seq and Encoder Decoder NN Models Pro Its also called Start of Sentence <SOS> sometimes [page 24] Seq2Seq and Encoder Decoder NN Models Pro [page 25] Seq2Seq and Encoder Decoder NN Models Pro [page 26] Seq2Seq and Encoder Decoder NN Models Pro [page 27] Seq2Seq and Encoder Decoder NN Models Pro [page 28] Seq2Seq and Encoder Decoder NN Models Pro [page 29] Seq2Seq and Encoder Decoder NN Models Pro [page 30] Seq2Seq and Encoder Decoder NN Models Pro [page 31] Seq2Seq and Encoder Decoder NN Models Pro [page 32] Seq2Seq and Encoder Decoder NN Models Pro [page 33] Seq2Seq and Encoder Decoder NN Models Pro [page 34] Seq2Seq and Encoder Decoder NN Models Pro [page 35] Seq2Seq and Encoder Decoder NN Models Pro [page 36] Seq2Seq and Encoder Decoder NN Models Pro [page 37] Seq2Seq and Encoder Decoder NN Models Pro [page 38] Seq2Seq and Encoder Decoder NN Models Pro [page 39] Seq2Seq and Encoder Decoder NN Models Pro [page 40] Seq2Seq and Encoder Decoder NN Models Pro [page 41] Seq2Seq and Encoder Decoder NN Models Pro