# EnglishToFrenchCode GC V6 PDF
course: Module 4 — Generative AI & LLMs
module: Module-4-Generative-AI-LLMs
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-4-Generative-AI-LLMs/General/Lab_Materials__03_May/EnglishToFrenchCode_GC_V6_PDF.pdf
pages: 136
---
[page 1]
Dataset Brief: Language Translation
(English–French)
Source
Kaggle — Language Translation (English-French) by Devicharith
What is this dataset?
It is a bilingual sentence pair dataset — for every English phrase, there is a matching
French translation sitting right next to it. Think of it as a giant phrasebook stored in a
spreadsheet.
Structure
Column Description Example
English Source sentence "Go."
French Translated sentence "Va !"
File format: CSV (eng_-french.csv)
Size: ~175,000+ sentence pairs (we'll use a subset of ~5000 for training)
Range: From single words like "Hi." → "Salut !" to full multi-word sentences
Why is this the right dataset for RNN / LSTM /
Attention?
This is a Sequence-to-Sequence (Seq2Seq) problem — the bread and butter use case for
all three architectures we'll study:
Concept What it means here
Sequence input An English sentence is a sequence of
words fed one-by-one into the model
Sequence output The French translation is another
sequence produced word-by-word
[page 2]
Concept What it means here
Variable length Input and output can have different
lengths — a core challenge
RNN A basic building block that processes
sequences step by step
LSTM An improved RNN that remembers long-
range dependencies better
Attention
Lets the decoder look back at ALL
English words and decide which one to
focus on at each translation step
Note: For demonstration purposes, this tutorial uses only 5,000 samples and 10 epochs
instead of the full dataset and 30 epochs, due to compute constraints. As a result, the
outputs may differ from those achieved with full training.
# cell_0a — GPU mask setup before any torch import
# =================================================
# MAIN GOAL:
# - Prompt the user for which physical GPU index to use (or CPU).
# - Set CUDA_VISIBLE_DEVICES so only that GPU is visible to this notebook
process.
#
# NOTES:
# - Must run *before* importing torch or any other CUDA-using library.
# - After masking, inside the process the selected GPU becomes “cuda:0”.
import os
choice = input(
"Enter GPU physical index you want to use (e.g., 0, 1, 2)\n"
"or type 'cpu' to use CPU.\n"
"Note: If using Colab, please use 0 id conect to GPU else use cpu.\n"
"> ")
if choice == "cpu":
os.environ["CUDA_VISIBLE_DEVICES"] = ""
print(" 🧠 CPU mode selected — no GPU will be visible.")
DEVICE = "cpu"
else:
if choice.isdigit():
phys_idx = int(choice)
os.environ["CUDA_VISIBLE_DEVICES"] = str(phys_idx)
print(f" 🔒 Masking all other GPUs. Only physical GPU index
{phys_idx} is visible → will map as cuda:0 internally.")
DEVICE = "cuda:0"
else:
print(" ⚠ Invalid input. Defaulting to CPU mode.")
os.environ["CUDA_VISIBLE_DEVICES"] = ""
DEVICE = "cpu"
[page 3]
Enter GPU physical index you want to use (e.g., 0, 1, 2)
or type 'cpu' to use CPU.
Note: If using Colab, please use 0 id conect to GPU else use cpu.
> cpu
🧠 CPU mode selected — no GPU will be visible.
>>> DEVICE set (internally) to: cpu
IMPORTS
All imports successful!
Loading the data
print(f"\n>>> DEVICE set (internally) to: {DEVICE}")
# ============================================================
# cell_0 — IMPORTS & SETUP
# ============================================================
# WHAT WE ARE DOING HERE:
# Only importing libraries needed for
# EDA (Exploratory Data Analysis) and Data Processing.
# Model-related imports will be added later when needed.
# ============================================================
# --- Standard library (built into Python — no install needed) ---
import re # Regular expressions — pattern-based text
cleaning
import string # string.punctuation gives all punctuation
characters
import unicodedata # Normalizes unicode (e.g. French accented
characters)
from collections import Counter # Counts word frequencies to build
vocabulary
# --- Numerical & data handling ---
import numpy as np # Arrays and matrix math
import pandas as pd # Loads our CSV into a DataFrame (like an
Excel table in Python)
# --- Visualization ---
import matplotlib.pyplot as plt # Core plotting — used for EDA charts
# --- Train / Test split ---
from sklearn.model_selection import train_test_split # Splits data into
train and test sets
print("All imports successful!")
[page 4]
English words/sentences French words/sentences
0 Hi. Salut!
1 Run! Cours !
2 Run! Courez !
3 Who? Qui ?
4 Wow! Ça alors !
... ... ...
4995 Don't be rude. Ne sois pas impolie !
4996 Don't be rude. Ne soyez pas impoli !
4997 Don't be rude. Ne soyez pas impolie !
4998 Don't call me. Ne m'appelle pas.
4999 Don't deny it. Ne le nie pas !
5000 rows × 2 columns
# ============================================================
# cell_1 — LOAD DATA
# ============================================================
# WHAT WE ARE DOING HERE:
# Load the CSV file into a Pandas DataFrame and do a
# quick sanity check to confirm it loaded correctly.
# ============================================================
# ============================================================
# CONFIGURATION
# ============================================================
DATA_PATH = r'/content/eng_-french.csv'
#DATA_PATH = r'/content/eng_-french.csv'
NUM_SAMPLES = 5000 # number of rows to load from the full
175,621
# ============================================================
# LOAD THE CSV
# ============================================================
# pd.read_csv() reads the file into a DataFrame (like an Excel table in
Python)
# nrows= limits how many rows we load — no need to load all 175k for this
tutorial
#df = pd.read_csv(DATA_PATH, nrows=NUM_SAMPLES)
df = pd.read_csv(DATA_PATH, nrows=NUM_SAMPLES)
df
# Rename columns to simpler names — easier to type throughout the notebook
# Original names : 'English words/sentences', 'French words/sentences'
[page 5]
Shape : 5,000 rows x 2 columns
Columns: ['english', 'french']
--- First 5 rows ---
english french
0 Hi. Salut!
1 Run! Cours !
2 Run! Courez !
3 Who? Qui ?
4 Wow! Ça alors !
EDA
df.columns = ['english', 'french']
# Quick sanity check — confirm shape and peek at first 5 rows
print(f"Shape : {df.shape[0]:,} rows x {df.shape[1]} columns")
print(f"Columns: {df.columns.tolist()}")
print("\n--- First 5 rows ---")
print(df.head())
# ============================================================
# cell_2 — EXPLORE DATA (EDA)
# ============================================================
# WHAT WE ARE DOING HERE:
# Perform basic EDA to understand our dataset before
# we clean or model anything. We want to answer:
# - Are there any missing values?
# - Are there any duplicate rows?
# - How long are the sentences?
# - What do the most common words look like?
# - How big are the vocabularies?
#
# WHY IS THIS IMPORTANT?
# EDA helps us spot problems early — missing values,
# very long sentences, noisy punctuation — before they
# cause issues during training.
# ============================================================
# ============================================================
# STEP 1 — Basic info
# ============================================================
# .info() tells us column names, non-null counts, and dtypes
# Both columns should show 30,000 non-null values of type object (string)
print("--- Dataset Info ---")
print(df.info())
# ============================================================
# STEP 2 — Check for missing values
[page 6]
# ============================================================
# .isnull().sum() counts empty/NaN values per column
# We cannot feed empty strings into a model — must be 0
print("\n--- Missing Values ---")
print(df.isnull().sum())
# ============================================================
# STEP 3 — Check for duplicate rows
# ============================================================
# Duplicate sentence pairs bias the model — it sees the same
# example multiple times and memorises it instead of learning
num_duplicates = df.duplicated().sum() # .duplicated() flags identical rows
print(f"\n--- Duplicate Rows ---")
print(f"Number of duplicate rows : {num_duplicates}")
# ============================================================
# STEP 4 — Sentence length analysis
# ============================================================
# We count words per sentence by splitting on whitespace.
# This tells us what padding length to use later and
# whether very long outlier sentences should be filtered.
df['english_len'] = df['english'].apply(lambda x: len(str(x).split()))
df['french_len'] = df['french'].apply(lambda x: len(str(x).split()))
# lambda x : takes each sentence x and returns its word count
print("\n--- Sentence Length Statistics (word count) ---")
print(df[['english_len', 'french_len']].describe().round(2))
# .describe() gives count, mean, std, min, 25%, 50%, 75%, max
# ============================================================
# STEP 5 — Plot sentence length distributions
# ============================================================
# A histogram shows how sentence lengths are spread out.
# Most translation datasets are right-skewed — most sentences
# are short but a few outliers are very long.
fig, axes = plt.subplots(1, 2, figsize=(12, 4)) # 1 row, 2 side-by-side
plots
# English length histogram
axes[0].hist(
df['english_len'], # data to plot
bins=50, # number of bars
color='steelblue',
edgecolor='white' # thin border between bars for readability
)
axes[0].set_title('English sentence length distribution')
axes[0].set_xlabel('Number of words')
axes[0].set_ylabel('Number of sentences')
[page 7]
# French length histogram
axes[1].hist(
df['french_len'],
bins=50,
color='coral',
edgecolor='white'
)
axes[1].set_title('French sentence length distribution')
axes[1].set_xlabel('Number of words')
axes[1].set_ylabel('Number of sentences')
plt.suptitle(f'Sentence length distributions ({len(df):,} samples)')
plt.tight_layout()
plt.show()
# ============================================================
# STEP 6 — Most common words (top 20)
# ============================================================
# Counter counts word frequencies across all sentences.
# This helps us understand vocabulary and spot noise
# (e.g. punctuation being counted as words).
all_english_words = ' '.join(df['english'].astype(str)).lower().split()
all_french_words = ' '.join(df['french'].astype(str)).lower().split()
# .join() → merges all sentences into one big string
# .lower() → lowercases everything so 'The' and 'the' are counted as one word
# .split() → splits into individual words on whitespace
english_word_counts = Counter(all_english_words)
french_word_counts = Counter(all_french_words)
top_english = english_word_counts.most_common(20)
top_french = french_word_counts.most_common(20)
print("\n--- Top 20 Most Common English Words ---")
for word, count in top_english:
print(f" {word:<15} {count:>6,}") # left-align word, right-align count
print("\n--- Top 20 Most Common French Words ---")
for word, count in top_french:
print(f" {word:<15} {count:>6,}")
# ============================================================
# STEP 7 — Vocabulary size estimate
# ============================================================
# Vocabulary size = number of unique words in the dataset.
# This tells us how large our embedding tables will be.
# Note: this is BEFORE cleaning — punctuation inflates this number.
english_vocab_size = len(english_word_counts)
french_vocab_size = len(french_word_counts)
[page 8]
--- Dataset Info ---
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 5000 entries, 0 to 4999
Data columns (total 2 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 english 5000 non-null object
1 french 5000 non-null object
dtypes: object(2)
memory usage: 78.3+ KB
None
--- Missing Values ---
english 0
french 0
dtype: int64
--- Duplicate Rows ---
Number of duplicate rows : 0
--- Sentence Length Statistics (word count) ---
english_len french_len
count 5000.00 5000.00
mean 2.55 3.24
std 0.58 1.10
min 1.00 1.00
25% 2.00 3.00
50% 3.00 3.00
print(f"\n--- Vocabulary Size (raw, before cleaning) ---")
print(f" Unique English words : {english_vocab_size:,}")
print(f" Unique French words : {french_vocab_size:,}")
# ============================================================
# STEP 8 — EDA Summary
# ============================================================
print("\n" + "=" * 45)
print(" EDA Summary")
print("=" * 45)
print(f" Total samples : {len(df):,}")
print(f" Missing values : {df.isnull().sum().sum()}")
print(f" Duplicate rows : {num_duplicates}")
print(f" Avg English length : {df['english_len'].mean():.1f} words")
print(f" Avg French length : {df['french_len'].mean():.1f} words")
print(f" Max English length : {df['english_len'].max()} words")
print(f" Max French length : {df['french_len'].max()} words")
print(f" English vocab (raw) : {english_vocab_size:,}")
print(f" French vocab (raw) : {french_vocab_size:,}")
print("=" * 45)
print("\n Ready for Cell 3 — Cleaning & Preprocessing")
[page 9]
75% 3.00 4.00
max 5.00 10.00
--- Top 20 Most Common English Words ---
i 969
i'm 494
is 271
you 270
a 261
we 201
it 189
you're 180
are 177
it. 175
be 169
it's 157
tom 146
he 127
was 125
don't 101
can 98
you. 91
do 90
we're 87
--- Top 20 Most Common French Words ---
je 1,021
! 922
? 616
suis 499
nous 335
c'est 271
vous 249
j'ai 226
il 225
tom 218
est 217
de 205
ne 185
me 177
à 175
[page 10]
pas 167
le 166
en 139
la 139
tu 132
--- Vocabulary Size (raw, before cleaning) ---
Unique English words : 1,684
Unique French words : 3,246
=============================================
EDA Summary
=============================================
Total samples : 5,000
Missing values : 0
Duplicate rows : 0
Avg English length : 2.6 words
Avg French length : 3.2 words
Max English length : 5 words
Max French length : 10 words
English vocab (raw) : 1,684
French vocab (raw) : 3,246
=============================================
Ready for Cell 3 — Cleaning & Preprocessing
# ============================================================
# cell_3 — CLEANING & PREPROCESSING
# ============================================================
# WHAT WE ARE DOING HERE:
# Raw text is messy — it has punctuation, mixed cases,
# extra spaces, and unicode characters (e.g. French accents
# encoded in unusual ways). We need to clean it so that:
# - 'Hello!' and 'hello' are treated as the same word
# - Punctuation doesn't become part of a word token
# - French accented characters (é, à, ç) are preserved correctly
#
# WHY IS THIS IMPORTANT?
# If we skip cleaning, our vocabulary explodes with near-
# duplicate entries like 'run', 'Run', 'run,' and 'run!'
# That wastes embedding space and hurts model performance.
# ============================================================
# ============================================================
# STEP 1 — Define the cleaning function
# ============================================================
def clean_sentence(sentence):
"""
Cleans a single sentence string.
Steps applied in order:
1. Lowercase everything
2. Normalize unicode (handles unusual encodings of accented letters)
[page 11]
3. Remove characters that are NOT letters, digits, or spaces
4. Collapse multiple spaces into one
5. Strip leading/trailing whitespace
"""
# Step 1 : lowercase — 'Hello' and 'hello' become the same token
sentence = sentence.lower()
# Step 2 : unicode normalization
# NFC = Canonical Decomposition followed by Canonical Composition
# This ensures é stored as two code points (e + combining accent)
# gets merged back into one standard character é
sentence = unicodedata.normalize('NFC', sentence)
# Step 3 : keep only letters (including accented), digits, and spaces
# re.sub replaces anything that does NOT match the pattern with ''
# \w matches word characters (letters + digits + underscore)
# \s matches whitespace
# [^\w\s] means: any character that is NOT a word char or whitespace
sentence = re.sub(r'[^\w\s]', '', sentence)
# Step 4 : remove underscores (kept by \w but not useful for NLP)
sentence = sentence.replace('_', '')
# Step 5 : collapse multiple spaces into a single space
sentence = re.sub(r'\s+', ' ', sentence)
# Step 6 : strip leading and trailing whitespace
sentence = sentence.strip()
return sentence
# ============================================================
# STEP 2 — Apply cleaning to both columns
# ============================================================
# .apply() runs clean_sentence on every row in the column
df['english_clean'] = df['english'].apply(clean_sentence)
df['french_clean'] = df['french'].apply(clean_sentence)
# ============================================================
# STEP 3 — Drop rows where cleaning produced an empty string
# ============================================================
# A sentence could become empty if it contained only punctuation.
# Empty strings would cause errors during tokenization.
before = len(df)
df = df[(df['english_clean'] != '') & (df['french_clean'] != '')] # keep
non-empty rows only
df = df.reset_index(drop=True) # reset
row numbers after dropping
after = len(df)
[page 12]
Rows before cleaning filter : 5,000
Rows after cleaning filter : 5,000
Rows removed : 0
--- Sample: Raw vs Cleaned ---
[Row 0]
EN raw : Hi.
EN cleaned : hi
FR raw : Salut!
FR cleaned : salut
[Row 1]
EN raw : Run!
EN cleaned : run
FR raw : Cours !
FR cleaned : cours
[Row 2]
EN raw : Run!
EN cleaned : run
FR raw : Courez !
FR cleaned : courez
[Row 3]
EN raw : Who?
EN cleaned : who
FR raw : Qui ?
FR cleaned : qui
[Row 4]
EN raw : Wow!
EN cleaned : wow
print(f"Rows before cleaning filter : {before:,}")
print(f"Rows after cleaning filter : {after:,}")
print(f"Rows removed : {before - after:,}")
# ============================================================
# STEP 4 — Sanity check — compare raw vs cleaned
# ============================================================
# Print a few examples so we can visually confirm the cleaning worked.
print("\n--- Sample: Raw vs Cleaned ---")
for i in range(5):
print(f"\n [Row {i}]")
print(f" EN raw : {df['english'].iloc[i]}")
print(f" EN cleaned : {df['english_clean'].iloc[i]}")
print(f" FR raw : {df['french'].iloc[i]}")
print(f" FR cleaned : {df['french_clean'].iloc[i]}")
print("\n Ready for Cell 4 — Tokenization")
[page 13]
FR raw : Ça alors !
FR cleaned : ça alors
Ready for Cell 4 — Tokenization
# ============================================================
# cell_4 — TOKENIZATION
# ============================================================
# WHAT WE ARE DOING HERE:
# Tokenization means splitting a sentence (a string) into
# a list of individual words (tokens).
#
# Example:
# "i love paris" → ['i', 'love', 'paris']
#
# WHY IS THIS IMPORTANT?
# Neural networks cannot work with raw strings. They need
# numbers. The first step toward numbers is breaking the
# sentence into discrete units (tokens) that we can later
# map to integers.
#
# NOTE:
# We already lowercased and removed punctuation in cell_3,
# so a simple whitespace split is all we need here.
# No need for a heavy tokenizer library.
# ============================================================
# ============================================================
# STEP 1 — Define the tokenizer function
# ============================================================
def tokenize(sentence):
"""
Splits a cleaned sentence string into a list of word tokens.
Input : "i love paris"
Output : ['i', 'love', 'paris']
"""
return sentence.split() # split() splits on any whitespace by default
# ============================================================
# STEP 2 — Apply tokenization to both cleaned columns
# ============================================================
# .apply() runs tokenize() on every row → each cell becomes a list of tokens
df['english_tokens'] = df['english_clean'].apply(tokenize)
df['french_tokens'] = df['french_clean'].apply(tokenize)
# ============================================================
# STEP 3 — Sanity check — print a few tokenized examples
# ============================================================
print("--- Sample Tokenized Sentences ---\n")
[page 14]
--- Sample Tokenized Sentences ---
[Row 0]
EN cleaned : hi
EN tokens : ['hi']
FR cleaned : salut
FR tokens : ['salut']
for i in range(5):
print(f" [Row {i}]")
print(f" EN cleaned : {df['english_clean'].iloc[i]}")
print(f" EN tokens : {df['english_tokens'].iloc[i]}")
print(f" FR cleaned : {df['french_clean'].iloc[i]}")
print(f" FR tokens : {df['french_tokens'].iloc[i]}")
print()
# ============================================================
# STEP 4 — Token count distribution check
# ============================================================
# After tokenization we can check exact token counts per sentence.
# This is cleaner than the word-count estimate we did in EDA (cell_2)
# because the text is now properly cleaned.
df['english_token_count'] = df['english_tokens'].apply(len) # number of
tokens per sentence
df['french_token_count'] = df['french_tokens'].apply(len)
print("--- Token Count Statistics ---")
print(df[['english_token_count', 'french_token_count']].describe().round(2))
# ============================================================
# STEP 5 — Decide MAX_SEQ_LEN
# ============================================================
# MAX_SEQ_LEN is the fixed length every sentence will be padded
# or truncated to in cell_8 (Padding).
#
# A good rule of thumb : cover ~95% of sentences without
# making sequences too long (long sequences slow training).
#
# We use the 95th percentile of token counts across both languages.
en_95 = int(np.percentile(df['english_token_count'], 95))
fr_95 = int(np.percentile(df['french_token_count'], 95))
# Take the larger of the two so both languages fit comfortably
MAX_SEQ_LEN = max(en_95, fr_95)
print(f"\n 95th percentile — English : {en_95} tokens")
print(f" 95th percentile — French : {fr_95} tokens")
print(f"\n MAX_SEQ_LEN set to : {MAX_SEQ_LEN}")
print("\n Ready for Cell 5 — Vocabulary Building")
[page 15]
[Row 1]
EN cleaned : run
EN tokens : ['run']
FR cleaned : cours
FR tokens : ['cours']
[Row 2]
EN cleaned : run
EN tokens : ['run']
FR cleaned : courez
FR tokens : ['courez']
[Row 3]
EN cleaned : who
EN tokens : ['who']
FR cleaned : qui
FR tokens : ['qui']
[Row 4]
EN cleaned : wow
EN tokens : ['wow']
FR cleaned : ça alors
FR tokens : ['ça', 'alors']
--- Token Count Statistics ---
english_token_count french_token_count
count 5000.00 5000.00
mean 2.55 2.93
std 0.58 1.13
min 1.00 1.00
25% 2.00 2.00
50% 3.00 3.00
75% 3.00 3.00
max 4.00 10.00
95th percentile — English : 3 tokens
95th percentile — French : 5 tokens
MAX_SEQ_LEN set to : 5
Ready for Cell 5 — Vocabulary Building
# ============================================================
# cell_5 — VOCABULARY BUILDING
# ============================================================
# WHAT WE ARE DOING HERE:
# A vocabulary is a mapping between words and unique integers.
# We build two vocabularies — one for English, one for French.
#
# Example (English vocab):
# word2idx : {'<pad>': 0, '<sos>': 1, '<eos>': 2, '<unk>': 3,
# 'i': 4, 'love': 5, 'paris': 6, ...}
# idx2word : {0: '<pad>', 1: '<sos>', 2: '<eos>', 3: '<unk>',
# 4: 'i', 5: 'love', 6: 'paris', ...}
[page 16]
#
# WHY IS THIS IMPORTANT?
# Neural networks work with numbers, not strings.
# word2idx lets us convert tokens → integers (for the model).
# idx2word lets us convert integers → tokens (for reading output).
#
# SPECIAL TOKENS (defined at the very start):
# <pad> = 0 → fills empty positions so all sequences are same length
# <sos> = 1 → "start of sequence" — tells the decoder to begin
# <eos> = 2 → "end of sequence" — tells the decoder to stop
# <unk> = 3 → "unknown" → replaces words seen at test time
# but not in training vocabulary
# ============================================================
# ============================================================
# STEP 1 — Define special tokens
# ============================================================
PAD_TOKEN = '<pad>' # padding
SOS_TOKEN = '<sos>' # start of sequence
EOS_TOKEN = '<eos>' # end of sequence
UNK_TOKEN = '<unk>' # unknown word
PAD_IDX = 0
SOS_IDX = 1
EOS_IDX = 2
UNK_IDX = 3
# We always insert special tokens first so their indices are fixed.
# This is critical — PAD must always be 0 because we will tell
# the loss function to ignore index 0 during training (cell_16).
SPECIAL_TOKENS = [PAD_TOKEN, SOS_TOKEN, EOS_TOKEN, UNK_TOKEN]
# ============================================================
# STEP 2 — Define the Vocabulary class
# ============================================================
class Vocabulary:
"""
Builds and stores word <-> index mappings for one language.
Attributes:
word2idx : dict — maps word string → integer index
idx2word : dict — maps integer index → word string
word_freq : Counter — tracks how many times each word appears
"""
def __init__(self):
# Start with the 4 special tokens at fixed indices 0-3
self.word2idx = {
PAD_TOKEN : PAD_IDX, # <pad> = 0
SOS_TOKEN : SOS_IDX, # <sos> = 1
[page 17]
EOS_TOKEN : EOS_IDX, # <eos> = 2
UNK_TOKEN : UNK_IDX, # <unk> = 3
}
# idx2word is the reverse mapping — built from word2idx
self.idx2word = {idx: word for word, idx in self.word2idx.items()}
self.word_freq = Counter() # will count word frequencies
def build(self, token_lists, min_freq=2):
"""
Populates word2idx and idx2word from a list of token lists.
Args:
token_lists : list of lists — e.g. [['i','love','paris'],
['hello'], ...]
min_freq : int — words appearing fewer than min_freq times
are excluded (they become <unk> at inference time).
Keeps vocabulary size manageable.
"""
# Count every token across all sentences
for tokens in token_lists:
self.word_freq.update(tokens) # Counter.update() adds counts
# Add words to vocabulary only if they meet min_freq threshold
for word, freq in self.word_freq.items():
if freq >= min_freq and word not in self.word2idx:
idx = len(self.word2idx) # next available integer index
self.word2idx[word] = idx
self.idx2word[idx] = word # keep reverse mapping in sync
def __len__(self):
# len(vocab) returns total number of entries including special tokens
return len(self.word2idx)
# ============================================================
# STEP 3 — Build English and French vocabularies
# ============================================================
# min_freq=2 means a word must appear at least twice to get its own index.
# Words seen only once are likely typos or noise — map them to <unk>.
MIN_FREQ = 2
english_vocab = Vocabulary()
english_vocab.build(df['english_tokens'].tolist(), min_freq=MIN_FREQ)
french_vocab = Vocabulary()
french_vocab.build(df['french_tokens'].tolist(), min_freq=MIN_FREQ)
# ============================================================
# STEP 4 — Sanity check
# ============================================================
[page 18]
English vocabulary size : 858 words
French vocabulary size : 1,207 words
(includes 4 special tokens: <pad>, <sos>, <eos>, <unk>)
--- Special Token Check ---
english_vocab['<pad>'] = 0
english_vocab['<sos>'] = 1
english_vocab['<eos>'] = 2
english_vocab['<unk>'] = 3
--- First 10 English vocab entries (after special tokens) ---
hi → 4
run → 5
who → 6
fire → 7
help → 8
jump → 9
stop → 10
wait → 11
go → 12
on → 13
--- First 10 French vocab entries (after special tokens) ---
salut → 4
cours → 5
courez → 6
qui → 7
ça → 8
alors → 9
au → 10
print(f" English vocabulary size : {len(english_vocab):,} words")
print(f" French vocabulary size : {len(french_vocab):,} words")
print(f" (includes 4 special tokens: <pad>, <sos>, <eos>, <unk>)")
# Confirm special token indices are correct
print(f"\n--- Special Token Check ---")
print(f" english_vocab['<pad>'] = {english_vocab.word2idx[PAD_TOKEN]}")
print(f" english_vocab['<sos>'] = {english_vocab.word2idx[SOS_TOKEN]}")
print(f" english_vocab['<eos>'] = {english_vocab.word2idx[EOS_TOKEN]}")
print(f" english_vocab['<unk>'] = {english_vocab.word2idx[UNK_TOKEN]}")
# Preview first 10 word → index mappings (after special tokens)
print(f"\n--- First 10 English vocab entries (after special tokens) ---")
for word, idx in list(english_vocab.word2idx.items())[4:14]:
print(f" {word:<20} → {idx}")
print(f"\n--- First 10 French vocab entries (after special tokens) ---")
for word, idx in list(french_vocab.word2idx.items())[4:14]:
print(f" {word:<20} → {idx}")
print("\n Ready for Cell 6 — Numericalization")
[page 19]
feu → 11
à → 12
laide → 13
Ready for Cell 6 — Numericalization
# ============================================================
# cell_6 — NUMERICALIZATION
# ============================================================
# WHAT WE ARE DOING HERE:
# Numericalization converts a list of word tokens into a
# list of integers using the vocabularies we built in cell_5.
#
# Example:
# tokens : ['i', 'love', 'paris']
# indices : [4, 5, 6 ]
#
# If a token is not in the vocabulary (i.e. it was filtered
# out by min_freq), we replace it with UNK_IDX (3).
#
# WHY IS THIS IMPORTANT?
# PyTorch Embedding layers take integers as input, not strings.
# Every token must become an integer before it can be fed
# into the model.
#
# NOTE:
# We do NOT add <sos> or <eos> here.
# That is handled separately in cell_7 so each step
# stays focused on one thing.
# ============================================================
# ============================================================
# STEP 1 — Define the numericalization function
# ============================================================
def numericalize(tokens, vocab):
"""
Converts a list of token strings into a list of integers.
Args:
tokens : list of str — e.g. ['i', 'love', 'paris']
vocab : Vocabulary — the vocab object for this language
Returns:
list of int — e.g. [4, 5, 6]
Unknown tokens (not in vocab) are replaced with UNK_IDX (3).
"""
return [
vocab.word2idx.get(token, UNK_IDX) # .get() returns UNK_IDX if
token not found
for token in tokens
[page 20]
]
# ============================================================
# STEP 2 — Apply numericalization to both languages
# ============================================================
# We pass the matching vocabulary for each language
df['english_indices'] = df['english_tokens'].apply(
lambda tokens: numericalize(tokens, english_vocab)
)
df['french_indices'] = df['french_tokens'].apply(
lambda tokens: numericalize(tokens, french_vocab)
)
# ============================================================
# STEP 3 — Sanity check
# ============================================================
print("--- Sample: Tokens → Indices ---\n")
for i in range(5):
print(f" [Row {i}]")
print(f" EN tokens : {df['english_tokens'].iloc[i]}")
print(f" EN indices : {df['english_indices'].iloc[i]}")
print(f" FR tokens : {df['french_tokens'].iloc[i]}")
print(f" FR indices : {df['french_indices'].iloc[i]}")
print()
# ============================================================
# STEP 4 — Check how many <unk> tokens appear
# ============================================================
# A very high <unk> count means min_freq was too aggressive
# and we are losing too much information.
# A small number is perfectly normal and expected.
def count_unks(index_lists):
"""Counts total number of UNK_IDX appearances across all sentences."""
return sum(
indices.count(UNK_IDX) # count UNK_IDX occurrences in each
sentence
for indices in index_lists
)
total_en_tokens = df['english_indices'].apply(len).sum() # total English
tokens
total_fr_tokens = df['french_indices'].apply(len).sum() # total French
tokens
en_unks = count_unks(df['english_indices'].tolist())
fr_unks = count_unks(df['french_indices'].tolist())
print("--- Unknown Token (<unk>) Report ---")
[page 21]
--- Sample: Tokens → Indices ---
[Row 0]
EN tokens : ['hi']
EN indices : [4]
FR tokens : ['salut']
FR indices : [4]
[Row 1]
EN tokens : ['run']
EN indices : [5]
FR tokens : ['cours']
FR indices : [5]
[Row 2]
EN tokens : ['run']
EN indices : [5]
FR tokens : ['courez']
FR indices : [6]
[Row 3]
EN tokens : ['who']
EN indices : [6]
FR tokens : ['qui']
FR indices : [7]
[Row 4]
EN tokens : ['wow']
EN indices : [3]
FR tokens : ['ça', 'alors']
FR indices : [8, 9]
--- Unknown Token (<unk>) Report ---
English : 388 <unk> out of 12,756 tokens (3.04%)
French : 1,519 <unk> out of 14,667 tokens (10.36%)
A <unk> rate below ~2% is healthy.
If it is higher, consider lowering MIN_FREQ in cell_5.
Ready for Cell 7 — Adding Special Tokens
print(f" English : {en_unks:,} <unk> out of {total_en_tokens:,} tokens "
f"({100 * en_unks / total_en_tokens:.2f}%)")
print(f" French : {fr_unks:,} <unk> out of {total_fr_tokens:,} tokens "
f"({100 * fr_unks / total_fr_tokens:.2f}%)")
print()
print(" A <unk> rate below ~2% is healthy.")
print(" If it is higher, consider lowering MIN_FREQ in cell_5.")
print("\n Ready for Cell 7 — Adding Special Tokens")
# ============================================================
# cell_7 — ADDING SPECIAL TOKENS
# ============================================================
[page 22]
# WHAT WE ARE DOING HERE:
# We add <sos> (start of sequence) and <eos> (end of sequence)
# tokens to every sentence.
#
# Example:
# before : [4, 5, 6]
# after : [1, 4, 5, 6, 2]
# ↑ ↑
# <sos> <eos>
#
# WHY IS THIS IMPORTANT?
# <sos> (index 1) — fed to the Decoder as the very first input
# to signal "start generating the translation now"
#
# <eos> (index 2) — tells the Decoder "the sentence is complete,
# stop generating". Without it the model never knows
# when to stop.
#
# WHICH LANGUAGE GETS WHICH TOKEN?
# English (source / encoder input):
# → We add ONLY <eos> at the end.
# → The Encoder just reads the sentence — it does not need
# a start signal, but <eos> tells it where the input ends.
#
# French (target / decoder input & output):
# → Decoder INPUT gets <sos> prepended : [<sos>, w1, w2, ...]
# → Decoder TARGET gets <eos> appended : [w1, w2, ..., <eos>]
# → During training the decoder sees <sos> and learns to
# predict the next word until it predicts <eos>.
# ============================================================
# ============================================================
# STEP 1 — Define the function to add special tokens
# ============================================================
def add_special_tokens(indices, add_sos=False, add_eos=False):
"""
Prepends <sos> and/or appends <eos> to an index list.
Args:
indices : list of int — numericalized sentence
add_sos : bool — if True, prepend SOS_IDX (1) at the start
add_eos : bool — if True, append EOS_IDX (2) at the end
Returns:
list of int — modified index list
Example:
add_special_tokens([4, 5, 6], add_sos=True, add_eos=True)
→ [1, 4, 5, 6, 2]
"""
[page 23]
if add_sos:
indices = [SOS_IDX] + indices # prepend <sos>
if add_eos:
indices = indices + [EOS_IDX] # append <eos>
return indices
# ============================================================
# STEP 2 — Apply to English (encoder input)
# ============================================================
# English gets only <eos> at the end.
# The encoder reads left-to-right and <eos> marks the boundary.
df['english_indices'] = df['english_indices'].apply(
lambda indices: add_special_tokens(indices, add_sos=False, add_eos=True)
)
# ============================================================
# STEP 3 — Create two versions of French (decoder input & target)
# ============================================================
# We need TWO separate columns for French:
#
# french_input → what the Decoder receives as input
# starts with <sos> so the decoder knows to begin
# e.g. [<sos>, w1, w2, w3]
#
# french_target → what the Decoder is trained to predict (the label)
# ends with <eos> so the decoder learns when to stop
# e.g. [w1, w2, w3, <eos>]
#
# During training:
# Decoder input at step t : french_input[t]
# Decoder should predict : french_target[t]
#
# So at step 0 : input=<sos> → predict w1
# at step 1 : input=w1 → predict w2
# ...
# at step n : input=wn → predict <eos>
df['french_input'] = df['french_indices'].apply(
lambda indices: add_special_tokens(indices, add_sos=True, add_eos=False)
)
df['french_target'] = df['french_indices'].apply(
lambda indices: add_special_tokens(indices, add_sos=False, add_eos=True)
)
# ============================================================
# STEP 4 — Sanity check
# ============================================================
print("--- Special Token Check ---\n")
for i in range(3):
[page 24]
--- Special Token Check ---
[Row 0]
EN tokens : ['hi']
EN indices : [4, 2]
FR tokens : ['salut']
FR input : [1, 4]
FR target : [4, 2]
Step-by-step alignment check (input → should predict):
step 0 : <sos> (1) → salut (4)
step 1 : salut (4) → <eos> (2)
[Row 1]
EN tokens : ['run']
EN indices : [5, 2]
FR tokens : ['cours']
FR input : [1, 5]
FR target : [5, 2]
Step-by-step alignment check (input → should predict):
step 0 : <sos> (1) → cours (5)
step 1 : cours (5) → <eos> (2)
[Row 2]
EN tokens : ['run']
EN indices : [5, 2]
FR tokens : ['courez']
FR input : [1, 6]
FR target : [6, 2]
print(f" [Row {i}]")
print(f" EN tokens : {df['english_tokens'].iloc[i]}")
print(f" EN indices : {df['english_indices'].iloc[i]}")
print()
print(f" FR tokens : {df['french_tokens'].iloc[i]}")
print(f" FR input : {df['french_input'].iloc[i]}")
print(f" FR target : {df['french_target'].iloc[i]}")
print()
# Confirm alignment: french_input[t] → french_target[t]
print(f" Step-by-step alignment check (input → should predict):")
for t, (inp, tgt) in enumerate(
zip(df['french_input'].iloc[i], df['french_target'].iloc[i])
):
inp_word = french_vocab.idx2word.get(inp, '?')
tgt_word = french_vocab.idx2word.get(tgt, '?')
print(f" step {t} : {inp_word:<12} ({inp}) → {tgt_word} ({tgt})")
print()
print("\n Ready for Cell 8 — Padding")
[page 25]
Step-by-step alignment check (input → should predict):
step 0 : <sos> (1) → courez (6)
step 1 : courez (6) → <eos> (2)
Ready for Cell 8 — Padding
# ============================================================
# cell_8 — PADDING
# ============================================================
# WHAT WE ARE DOING HERE:
# Every sentence in our dataset has a different length.
# PyTorch processes sentences in batches — a batch is a
# 2D tensor of shape (batch_size, sequence_length).
# For this to work, ALL sentences in a batch must be the
# same length.
#
# We fix this by PADDING — appending <pad> (index 0) to
# shorter sentences until they reach MAX_SEQ_LEN.
#
# Example (MAX_SEQ_LEN = 6):
# [1, 4, 5, 6, 2] → [1, 4, 5, 6, 2, 0]
# [1, 4, 2] → [1, 4, 2, 0, 0, 0]
# [1, 4, 5, 6, 7, 8, 2] → [1, 4, 5, 6, 7, 8] ← truncated
#
# WHY <pad> = 0 SPECIFICALLY?
# Later in cell_16 (Training), we tell the loss function to
# IGNORE positions where the target is 0 (<pad>).
# This way padding does not contribute to the loss and the
# model is not penalized for what it predicts at pad positions.
#
# TRUNCATION:
# Sentences longer than MAX_SEQ_LEN are truncated.
# We computed MAX_SEQ_LEN in cell_4 as the 95th percentile
# of token counts — so very few sentences are affected.
# ============================================================
# ============================================================
# STEP 1 — Define the padding function
# ============================================================
def pad_or_truncate(indices, max_len):
"""
Pads or truncates a list of indices to exactly max_len.
Args:
indices : list of int — numericalized sentence with special tokens
max_len : int — target fixed length (MAX_SEQ_LEN)
Returns:
list of int of length exactly max_len
[page 26]
Padding : appends PAD_IDX (0) to the right until length == max_len
Truncate : slices to first max_len tokens if sentence is too long
"""
current_len = len(indices)
if current_len < max_len:
# Pad on the right with PAD_IDX (0)
padding_needed = max_len - current_len # how many zeros to
add
indices = indices + [PAD_IDX] * padding_needed # append PAD_IDX
tokens
elif current_len > max_len:
# Truncate — keep only the first max_len tokens
indices = indices[:max_len]
# If current_len == max_len, nothing to do
return indices
# ============================================================
# STEP 2 — Decide final sequence lengths
# ============================================================
# English and French can have different MAX lengths.
# English (encoder): MAX_SEQ_LEN — computed in cell_4
# French (decoder): MAX_SEQ_LEN + 1 — because we added <sos>
# or <eos>, making French sequences 1 token
# longer than the original sentence
EN_SEQ_LEN = MAX_SEQ_LEN + 1 # +1 for <eos> added to English in cell_7
FR_SEQ_LEN = MAX_SEQ_LEN + 1 # +1 for <sos> or <eos> added in cell_7
print(f" EN_SEQ_LEN (encoder) : {EN_SEQ_LEN}")
print(f" FR_SEQ_LEN (decoder) : {FR_SEQ_LEN}")
# ============================================================
# STEP 3 — Apply padding to all three index columns
# ============================================================
df['english_padded'] = df['english_indices'].apply(
lambda x: pad_or_truncate(x, EN_SEQ_LEN)
)
df['french_input_padded'] = df['french_input'].apply(
lambda x: pad_or_truncate(x, FR_SEQ_LEN)
)
df['french_target_padded'] = df['french_target'].apply(
lambda x: pad_or_truncate(x, FR_SEQ_LEN)
)
[page 27]
# ============================================================
# STEP 4 — Verify all sequences are exactly the right length
# ============================================================
en_lengths = df['english_padded'].apply(len)
fr_inp_lengths = df['french_input_padded'].apply(len)
fr_tgt_lengths = df['french_target_padded'].apply(len)
#nunique() counts the number of unique values in a tensor/series.
# Eg: If en_lengths.nunique() == 1 means #
# "all English sequences have the same length"
# if there's only 1 unique length value, they're all identical.
assert en_lengths.nunique() == 1, "English sequences are not all the same
length!"
assert fr_inp_lengths.nunique() == 1, "French input sequences are not all the
same length!"
assert fr_tgt_lengths.nunique() == 1, "French target sequences are not all
the same length!"
print(f"\n All English sequences : length {en_lengths.iloc[0]} ✓ ")
print(f" All French input sequences : length {fr_inp_lengths.iloc[0]} ✓ ")
print(f" All French target sequences: length {fr_tgt_lengths.iloc[0]} ✓ ")
# ============================================================
# STEP 5 — Sanity check — visualize padding on a few examples
# ============================================================
print("\n--- Sample: Padded Sequences ---\n")
for i in range(3):
print(f" [Row {i}]")
print(f" EN padded : {df['english_padded'].iloc[i]}")
print(f" FR input padded : {df['french_input_padded'].iloc[i]}")
print(f" FR target padded: {df['french_target_padded'].iloc[i]}")
print()
# ============================================================
# STEP 6 — Check truncation rate
# ============================================================
# How many sentences were truncated?
# If this is high, consider increasing MAX_SEQ_LEN in cell_4.
en_truncated = (df['english_indices'].apply(len) > EN_SEQ_LEN).sum()
fr_truncated = (df['french_input'].apply(len) > FR_SEQ_LEN).sum()
print("--- Truncation Report ---")
print(f" English sentences truncated : {en_truncated:,} "
f"({100 * en_truncated / len(df):.2f}%)")
print(f" French sentences truncated : {fr_truncated:,} "
f"({100 * fr_truncated / len(df):.2f}%)")
print()
print(" If truncation rate > 5%, consider increasing MAX_SEQ_LEN in
cell_4.")
[page 28]
EN_SEQ_LEN (encoder) : 6
FR_SEQ_LEN (decoder) : 6
All English sequences : length 6 ✓
All French input sequences : length 6 ✓
All French target sequences: length 6 ✓
--- Sample: Padded Sequences ---
[Row 0]
EN padded : [4, 2, 0, 0, 0, 0]
FR input padded : [1, 4, 0, 0, 0, 0]
FR target padded: [4, 2, 0, 0, 0, 0]
[Row 1]
EN padded : [5, 2, 0, 0, 0, 0]
FR input padded : [1, 5, 0, 0, 0, 0]
FR target padded: [5, 2, 0, 0, 0, 0]
[Row 2]
EN padded : [5, 2, 0, 0, 0, 0]
FR input padded : [1, 6, 0, 0, 0, 0]
FR target padded: [6, 2, 0, 0, 0, 0]
--- Truncation Report ---
English sentences truncated : 0 (0.00%)
French sentences truncated : 134 (2.68%)
If truncation rate > 5%, consider increasing MAX_SEQ_LEN in cell_4.
Ready for Cell 9 — Train/Val/Test Split
print("\n Ready for Cell 9 — Train/Val/Test Split")
# ============================================================
# cell_9 — TRAIN / VAL / TEST SPLIT
# ============================================================
# WHAT WE ARE DOING HERE:
# We split our dataset into three parts:
#
# ┌─────────────────────────────────────────────────────┐
# │ Train (80%) │ Validation (10%) │ Test (10%) │
# └─────────────────────────────────────────────────────┘
#
# Train → the model learns from this data
# Validation → checked after every epoch to monitor
# if the model is overfitting (memorising
# instead of learning)
# Test → touched ONLY at the very end to report
# final performance — never used for tuning
#
# WHY KEEP TEST SEPARATE?
[page 29]
# If we tune our model based on test performance, we
# are indirectly training on it and our final score
# is no longer an honest measure of generalisation.
#
# RANDOM STATE:
# random_state=42 makes the split reproducible — everyone
# running this notebook gets the exact same split.
# ============================================================
# ============================================================
# STEP 1 — Extract the three padded arrays we need
# ============================================================
# We only carry forward the final padded columns.
# Everything else (raw text, tokens, unpadded indices) was
# needed for building up to this point but is not needed
# by the model.
encoder_inputs = df['english_padded'].tolist() # source sentences
(English)
decoder_inputs = df['french_input_padded'].tolist() # decoder inputs
(<sos> + French)
decoder_targets = df['french_target_padded'].tolist() # decoder targets
(French + <eos>)
# ============================================================
# STEP 2 — First split : separate Test from the rest
# ============================================================
# We do two sequential splits instead of one three-way split
# because train_test_split only splits into two parts at a time.
#
# Split 1 : (train + val) 90% | test 10%
# After Split 1
# enc_train_val # encoder → train + val (90%) ✓
# enc_test # encoder → test (10%) ✓
# dec_inp_train_val # decoder input → train + val (90%) ✓
# dec_inp_test # decoder input → test (10%) ✓
# dec_tgt_train_val # decoder target→ train + val (90%) ✓
# dec_tgt_test # decoder target→ test (10%) ✓
(enc_train_val, enc_test,
dec_inp_train_val, dec_inp_test,
dec_tgt_train_val, dec_tgt_test) = train_test_split(
encoder_inputs, # English padded
decoder_inputs, # French input padded
decoder_targets, # French target padded
test_size=0.10, # 10% goes to test
random_state=42 # fixed seed for reproducibility
)
[page 30]
# ============================================================
# STEP 3 — Second split : separate Val from Train
# ============================================================
# Split 2 : train 80% | val 10%
# (val is 1/9 of the remaining 90% ≈ 11.1% ≈ 10% of total)
(enc_train, enc_val,
dec_inp_train, dec_inp_val,
dec_tgt_train, dec_tgt_val) = train_test_split(
enc_train_val, # remaining 90% of English
dec_inp_train_val, # remaining 90% of French input
dec_tgt_train_val, # remaining 90% of French target
test_size=1/9, # 1/9 of 90% = 10% of total dataset
random_state=42
)
# ============================================================
# STEP 4 — Sanity check — confirm split sizes
# ============================================================
total = len(encoder_inputs)
print("--- Split Summary ---\n")
print(f" Total samples : {total:,} (100%)")
print(f" Train samples : {len(enc_train):,} "
f"({100 * len(enc_train) / total:.1f}%)")
print(f" Validation samples : {len(enc_val):,} "
f"({100 * len(enc_val) / total:.1f}%)")
print(f" Test samples : {len(enc_test):,} "
f"({100 * len(enc_test) / total:.1f}%)")
# Confirm all three splits together equal the full dataset
assert len(enc_train) + len(enc_val) + len(enc_test) == total, \
"Split sizes do not add up to total!"
print(f"\n Train + Val + Test = {len(enc_train) + len(enc_val) +
len(enc_test):,} ✓ ")
# ============================================================
# STEP 5 — Confirm shapes are consistent across all splits
# ============================================================
# Every encoder input should have length EN_SEQ_LEN
# Every decoder input/target should have length FR_SEQ_LEN
print(f"\n--- Sequence Length Check ---")
print(f" enc_train[0] length : {len(enc_train[0])} "
f"(expected {EN_SEQ_LEN}) ✓ ")
print(f" dec_inp_train[0] length : {len(dec_inp_train[0])} "
f"(expected {FR_SEQ_LEN}) ✓ ")
print(f" dec_tgt_train[0] length : {len(dec_tgt_train[0])} "
f"(expected {FR_SEQ_LEN}) ✓ ")
print("\n Ready for Cell 10 — PyTorch Dataset")
[page 31]
--- Split Summary ---
Total samples : 5,000 (100%)
Train samples : 4,000 (80.0%)
Validation samples : 500 (10.0%)
Test samples : 500 (10.0%)
Train + Val + Test = 5,000 ✓
--- Sequence Length Check ---
enc_train[0] length : 6 (expected 6) ✓
dec_inp_train[0] length : 6 (expected 6) ✓
dec_tgt_train[0] length : 6 (expected 6) ✓
Ready for Cell 10 — PyTorch Dataset
# ============================================================
# cell_10 — PYTORCH DATASET
# ============================================================
# WHAT WE ARE DOING HERE:
# We wrap our data inside a PyTorch Dataset class.
#
# A PyTorch Dataset is a standard container that knows
# how to answer two questions:
# 1. How many samples do I have? → __len__()
# 2. Give me the sample at index i? → __getitem__()
#
# That's it. Nothing more.
#
# WHY DO WE NEED THIS?
# PyTorch's DataLoader (cell_11) requires data to be
# wrapped in a Dataset. The DataLoader then handles:
# - Shuffling
# - Batching
# - Loading samples efficiently during training
#
# WHY CONVERT TO TENSORS HERE?
# Neural networks operate on PyTorch tensors, not Python
# lists. __getitem__() is the right place to do this
# conversion because it happens lazily — only when a
# sample is actually requested, not all at once upfront.
# This keeps memory usage low.
# ============================================================
import torch # core PyTorch library
from torch.utils.data import Dataset # base class we inherit from
# ============================================================
# STEP 1 — Define the Dataset class
# ============================================================
class TranslationDataset(Dataset):
[page 32]
"""
PyTorch Dataset for English → French translation.
Each sample contains three tensors:
encoder_input : English sentence (padded indices)
decoder_input : French sentence (<sos> + padded indices)
decoder_target : French sentence (padded indices + <eos>)
"""
def __init__(self, encoder_inputs, decoder_inputs, decoder_targets):
"""
Stores the three lists of padded index sequences.
Args:
encoder_inputs : list of lists — English padded indices
decoder_inputs : list of lists — French input padded indices
decoder_targets : list of lists — French target padded indices
"""
self.encoder_inputs = encoder_inputs # source sentences
self.decoder_inputs = decoder_inputs # decoder inputs (<sos> +
french)
self.decoder_targets = decoder_targets # decoder targets (french +
<eos>)
def __len__(self):
"""Returns total number of sentence pairs in this split."""
return len(self.encoder_inputs)
def __getitem__(self, idx):
"""
Returns one sample as a tuple of three LongTensors.
Args:
idx : int — index of the sample to retrieve
Returns:
tuple of (encoder_input, decoder_input, decoder_target)
each is a 1D torch.LongTensor of shape (seq_len,)
Why LongTensor?
Embedding layers in PyTorch require integer indices
of type torch.long (int64). Using float would crash.
"""
encoder_input = torch.tensor(self.encoder_inputs[idx],
dtype=torch.long)
decoder_input = torch.tensor(self.decoder_inputs[idx],
dtype=torch.long)
decoder_target = torch.tensor(self.decoder_targets[idx],
dtype=torch.long)
return encoder_input, decoder_input, decoder_target
# ============================================================
[page 33]
--- Dataset Sizes ---
Train dataset : 4,000 samples
Validation dataset : 500 samples
# STEP 2 — Create Dataset objects for each split
# ============================================================
train_dataset = TranslationDataset(enc_train, dec_inp_train, dec_tgt_train)
val_dataset = TranslationDataset(enc_val, dec_inp_val, dec_tgt_val)
test_dataset = TranslationDataset(enc_test, dec_inp_test, dec_tgt_test)
# ============================================================
# STEP 3 — Sanity check — confirm sizes
# ============================================================
print("--- Dataset Sizes ---")
print(f" Train dataset : {len(train_dataset):,} samples")
print(f" Validation dataset : {len(val_dataset):,} samples")
print(f" Test dataset : {len(test_dataset):,} samples")
# ============================================================
# STEP 4 — Sanity check — inspect one sample
# ============================================================
# Retrieve the very first sample from the training set
# and confirm shapes and types look correct.
sample_enc, sample_dec_inp, sample_dec_tgt = train_dataset[0]
print(f"\n--- Single Sample Inspection (train_dataset[0]) ---")
print(f" encoder_input : {sample_enc}")
print(f" decoder_input : {sample_dec_inp}")
print(f" decoder_target : {sample_dec_tgt}")
print()
print(f" encoder_input shape : {sample_enc.shape} "
f"dtype : {sample_enc.dtype}")
print(f" decoder_input shape : {sample_dec_inp.shape} "
f"dtype : {sample_dec_inp.dtype}")
print(f" decoder_target shape : {sample_dec_tgt.shape} "
f"dtype : {sample_dec_tgt.dtype}")
# Decode back to words so the beginner can visually verify
print(f"\n--- Decoded back to words ---")
print(f" encoder_input : "
f"{[english_vocab.idx2word.get(i.item(), UNK_TOKEN) for i in
sample_enc]}")
print(f" decoder_input : "
f"{[french_vocab.idx2word.get(i.item(), UNK_TOKEN) for i in
sample_dec_inp]}")
print(f" decoder_target : "
f"{[french_vocab.idx2word.get(i.item(), UNK_TOKEN) for i in
sample_dec_tgt]}")
print("\n Ready for Cell 11 — PyTorch DataLoader")
[page 34]
Test dataset : 500 samples
--- Single Sample Inspection (train_dataset[0]) ---
encoder_input : tensor([ 86, 459, 118, 2, 0, 0])
decoder_input : tensor([ 1, 59, 424, 3, 0, 0])
decoder_target : tensor([ 59, 424, 3, 2, 0, 0])
encoder_input shape : torch.Size([6]) dtype : torch.int64
decoder_input shape : torch.Size([6]) dtype : torch.int64
decoder_target shape : torch.Size([6]) dtype : torch.int64
--- Decoded back to words ---
encoder_input : ['its', 'my', 'job', '<eos>', '<pad>', '<pad>']
decoder_input : ['<sos>', 'cest', 'mon', '<unk>', '<pad>', '<pad>']
decoder_target : ['cest', 'mon', '<unk>', '<eos>', '<pad>', '<pad>']
Ready for Cell 11 — PyTorch DataLoader
# ============================================================
# cell_11 — PYTORCH DATALOADER
# ============================================================
# WHAT WE ARE DOING HERE:
# We wrap each Dataset inside a PyTorch DataLoader.
#
# While the Dataset knows HOW to fetch one sample,
# the DataLoader decides:
# - How many samples to fetch at once (batch_size)
# - Whether to shuffle the data before each epoch
# - How many CPU workers to use for loading
#
# Each iteration of the DataLoader gives us a BATCH —
# a group of samples stacked into tensors of shape:
# (batch_size, seq_len)
#
# Example with batch_size=32, EN_SEQ_LEN=7, FR_SEQ_LEN=7:
# encoder_input → tensor of shape (32, 7)
# decoder_input → tensor of shape (32, 7)
# decoder_target → tensor of shape (32, 7)
#
# WHY BATCHING?
# Training one sentence at a time is very slow.
# GPUs are designed for parallel computation — feeding
# 32 (or 64) sentences at once is much faster than
# feeding them one by one.
#
# WHY SHUFFLE ONLY TRAIN?
# Train → shuffle=True : randomising order prevents the
# model from memorising the sequence of examples
# Val → shuffle=False : we just want stable evaluation
# Test → shuffle=False : order does not matter for testing
# ============================================================
[page 35]
from torch.utils.data import DataLoader # handles batching and shuffling
# ============================================================
# STEP 1 — Set batch size
# ============================================================
# batch_size = number of sentence pairs processed together.
# 32 is a safe default for most machines.
# If you run out of memory, reduce to 16.
# If you have a powerful GPU, try 64 or 128.
BATCH_SIZE = 32
# ============================================================
# STEP 2 — Create DataLoader for each split
# ============================================================
train_loader = DataLoader(
train_dataset, # the Dataset object from cell_10
batch_size=BATCH_SIZE,
shuffle=True, # shuffle training data every epoch
num_workers=0, # 0 = load data in the main process
# increase if data loading is a bottleneck
# (keep 0 on Windows to avoid multiprocessing
issues)
pin_memory=True # speeds up CPU → GPU data transfer
# safe to use even without a GPU
)
val_loader = DataLoader(
val_dataset,
batch_size=BATCH_SIZE,
shuffle=False, # no shuffling for validation
num_workers=0,
pin_memory=True
)
test_loader = DataLoader(
test_dataset,
batch_size=BATCH_SIZE,
shuffle=False, # no shuffling for test
num_workers=0,
pin_memory=True
)
# ============================================================
# STEP 3 — Sanity check — confirm loader sizes
# ============================================================
# len(DataLoader) = number of batches = ceil(samples / batch_size)
print("--- DataLoader Summary ---\n")
print(f" Batch size : {BATCH_SIZE}")
print(f" Train batches : {len(train_loader):,} "
[page 36]
--- DataLoader Summary ---
Batch size : 32
f"({len(train_dataset):,} samples)")
print(f" Validation batches : {len(val_loader):,} "
f"({len(val_dataset):,} samples)")
print(f" Test batches : {len(test_loader):,} "
f"({len(test_dataset):,} samples)")
# ============================================================
# STEP 4 — Inspect one batch
# ============================================================
# Grab the very first batch from the train loader and
# confirm shapes are exactly what we expect.
enc_batch, dec_inp_batch, dec_tgt_batch = next(iter(train_loader))
# next(iter(...)) fetches exactly one batch without looping
print(f"\n--- Single Batch Inspection ---")
print(f" enc_batch shape : {enc_batch.shape} "
f"(batch_size x EN_SEQ_LEN)")
print(f" dec_inp_batch shape : {dec_inp_batch.shape} "
f"(batch_size x FR_SEQ_LEN)")
print(f" dec_tgt_batch shape : {dec_tgt_batch.shape} "
f"(batch_size x FR_SEQ_LEN)")
print(f" dtype : {enc_batch.dtype}")
# ============================================================
# STEP 5 — Visualise one sentence pair from the batch
# ============================================================
# Decode the first sentence in the batch back to words
# so we can visually confirm nothing broke during batching.
print(f"\n--- First sentence in batch (decoded) ---")
print(f" encoder_input : "
f"{[english_vocab.idx2word.get(i.item(), UNK_TOKEN) for i in
enc_batch[0]]}")
print(f" decoder_input : "
f"{[french_vocab.idx2word.get(i.item(), UNK_TOKEN) for i in
dec_inp_batch[0]]}")
print(f" decoder_target : "
f"{[french_vocab.idx2word.get(i.item(), UNK_TOKEN) for i in
dec_tgt_batch[0]]}")
# ============================================================
# STEP 6 — Device setup
# ============================================================
# DEVICE was already set in cell_0a based on your GPU selection.
# We confirm it here and check torch can see the expected hardware.
print(f" {'GPU detected — training will be fast!' if 'cuda' in str(DEVICE)
else 'No GPU detected — CPU will be used (slower but fine for this
tutorial).'}")
[page 37]
Train batches : 125 (4,000 samples)
Validation batches : 16 (500 samples)
Test batches : 16 (500 samples)
--- Single Batch Inspection ---
enc_batch shape : torch.Size([32, 6]) (batch_size x EN_SEQ_LEN)
dec_inp_batch shape : torch.Size([32, 6]) (batch_size x FR_SEQ_LEN)
dec_tgt_batch shape : torch.Size([32, 6]) (batch_size x FR_SEQ_LEN)
dtype : torch.int64
--- First sentence in batch (decoded) ---
encoder_input : ['is', 'tom', 'here', '<eos>', '<pad>', '<pad>']
decoder_input : ['<sos>', 'estce', 'que', 'tom', 'est', 'ici']
decoder_target : ['estce', 'que', 'tom', 'est', 'ici', '<eos>']
No GPU detected — CPU will be used (slower but fine for this tutorial).
/usr/local/lib/python3.12/dist-packages/torch/utils/data/dataloader.py:775:
UserWarning: 'pin_memory' argument is set as true but no accelerator is
found, then device pinned memory won't be used.
super().__init__(loader)
# ============================================================
# cell_12 — ENCODER
# ============================================================
# WHAT WE ARE DOING HERE:
# The Encoder reads the English sentence and compresses it
# into a set of hidden states — one for each word.
#
# These hidden states are NOT thrown away after encoding.
# The Attention mechanism (cell_13) will look at ALL of them
# to decide which words to focus on when translating.
#
# ARCHITECTURE:
#
# English indices → Embedding → LSTM → encoder_outputs, (hidden, cell)
#
# Step by step:
# 1. Embedding layer : converts each word index into a
# dense vector of size embed_dim
# (e.g. index 4 → [0.2, -0.5, 0.8, ...])
#
# 2. LSTM (Long Short-Term Memory) : processes the embedded
# sequence left to right, producing:
#
# encoder_outputs : hidden state at EVERY time step
# shape → (batch, seq_len, hidden_dim)
# used by Attention in cell_13
#
# hidden : short-term memory at the LAST time step
# shape → (1, batch, hidden_dim)
# used to initialise the Decoder
#
# cell : long-term memory at the LAST time step
[page 38]
# shape → (1, batch, hidden_dim)
# used to initialise the Decoder
#
# WHY LSTM OVER GRU?
# LSTM has two memory states — hidden (short-term) and cell
# (long-term). This gives the model more expressive power
# to capture dependencies across longer sequences.
# The cell state acts as a "conveyor belt" — it carries
# information across many time steps with minimal change,
# solving the vanishing gradient problem more effectively
# than a plain RNN.
#
# WHY NOT BIDIRECTIONAL?
# A bidirectional encoder reads the sentence both
# left-to-right AND right-to-left, giving richer context.
# We keep it unidirectional here to stay beginner-friendly.
# You can try bidirectional=True as an extension exercise.
# ============================================================
import torch.nn as nn # neural network building blocks
# ============================================================
# STEP 1 — Define Encoder hyperparameters
# ============================================================
EMBED_DIM = 256 # size of each word embedding vector
HIDDEN_DIM = 512 # size of the LSTM hidden state and cell state
ENC_DROPOUT = 0.3 # dropout rate — randomly zeros 30% of neurons
# during training to prevent overfitting
# ============================================================
# STEP 2 — Define the Encoder class
# ============================================================
class Encoder(nn.Module):
"""
Encodes an English sentence into a sequence of hidden states.
Args:
vocab_size : int — number of words in English vocabulary
embed_dim : int — dimension of each word embedding
hidden_dim : int — dimension of LSTM hidden & cell state
dropout : float — dropout probability
Forward input:
src : tensor of shape (batch_size, src_seq_len)
each value is a word index (torch.long)
Forward output:
encoder_outputs : tensor (batch_size, src_seq_len, hidden_dim)
hidden state at every time step
→ used by Attention mechanism
[page 39]
hidden : tensor (1, batch_size, hidden_dim)
short-term memory at the last time step
→ used to initialise the Decoder
cell : tensor (1, batch_size, hidden_dim)
long-term memory at the last time step
→ used to initialise the Decoder
"""
def __init__(self, vocab_size, embed_dim, hidden_dim, dropout):
super(Encoder, self).__init__()
# Initialise nn.Module's internal machinery first.
# This registers PyTorch's parameter tracking, layer bookkeeping,
# and gradient management before we attach any sub-modules.
# Must always be called before any self.* assignments
# EMBEDDING LAYER
# Lookup table of shape (vocab_size, embed_dim).
# Converts integer word indices into dense floating-point vectors.
# padding_idx=PAD_IDX ensures the <pad> token embedding remains
# a zero vector and is never updated during backpropagation.
self.embedding = nn.Embedding(
vocab_size, # one row per word in vocabulary
embed_dim, # each row is a vector of this size
padding_idx=PAD_IDX # <pad> token embedding stays zero
)
# ----------------------------------------------------------------
# LSTM LAYER
# Processes the sequence of embedding vectors one step at a time.
# Maintains two internal states at each time step:
# hidden (h) — short-term memory, sensitive to recent inputs
# cell (c) — long-term memory, carries context across many steps
# Both states are passed to the Decoder after the full source
# sentence has been read.
# batch_first=True sets input/output shape to (batch, seq_len,
features)
# rather than PyTorch's default (seq_len, batch, features).
self.lstm = nn.LSTM(
embed_dim, # input size = embedding dimension
hidden_dim, # hidden size = LSTM hidden & cell dimension
batch_first=True # input/output shape : (batch, seq, feature)
)
# ----------------------------------------------------------------
# DROPOUT LAYER
# Applied to embedding vectors before they enter the LSTM.
# Randomly zeroes elements with probability=dropout during training.
# Disabled automatically during evaluation (model.eval()).
# ----------------------------------------------------------------
self.dropout = nn.Dropout(dropout)
[page 40]
def forward(self, src):
"""
Args:
src : (batch_size, src_seq_len) — padded English indices
Returns:
encoder_outputs : (batch_size, src_seq_len, hidden_dim)
hidden : (1, batch_size, hidden_dim)
cell : (1, batch_size, hidden_dim)
"""
# Step 1 : convert word indices → embedding vectors
# src shape : (batch, src_seq_len)
# embedded shape : (batch, src_seq_len, embed_dim)
embedded = self.dropout(self.embedding(src))
# Step 2 : pass embeddings through LSTM
# encoder_outputs : hidden state at every time step
# shape → (batch, src_seq_len, hidden_dim)
# hidden : short-term memory at the final time step
# shape → (1, batch, hidden_dim)
# cell : long-term memory at the final time step
# shape → (1, batch, hidden_dim)
encoder_outputs, (hidden, cell) = self.lstm(embedded)
# Note : LSTM returns states as a tuple (hidden, cell)
# GRU returns just hidden — this is the key difference
return encoder_outputs, hidden, cell
# ============================================================
# STEP 3 — Instantiate the Encoder
# ============================================================
encoder = Encoder(
vocab_size = len(english_vocab), # English vocab size from cell_5
embed_dim = EMBED_DIM,
hidden_dim = HIDDEN_DIM,
dropout = ENC_DROPOUT
).to(DEVICE) # move model parameters to GPU if available
# ============================================================
# STEP 4 — Sanity check — print model summary
# ============================================================
print("--- Encoder Architecture ---")
print(encoder)
total_params = sum(p.numel() for p in encoder.parameters() if
p.requires_grad)
#numel stands for 'number of elements.' #
# Each parameter p is a tensor, a grid of numbers. #
[page 41]
--- Encoder Architecture ---
Encoder(
(embedding): Embedding(858, 256, padding_idx=0)
(lstm): LSTM(256, 512, batch_first=True)
(dropout): Dropout(p=0.3, inplace=False)
)
Total trainable parameters : 1,796,608
--- Forward Pass Shape Check ---
Input (src) : torch.Size([32, 6]) (batch_size x EN_SEQ_LEN)
encoder_outputs shape : torch.Size([32, 6, 512]) (batch_size x
EN_SEQ_LEN x hidden_dim)
encoder hidden shape : torch.Size([1, 32, 512]) (1 x batch_size x
hidden_dim)
encoder cell shape : torch.Size([1, 32, 512]) (1 x batch_size x
hidden_dim)
All shapes correct ✓
# numel() just counts how many numbers are in that grid."
print(f"\n Total trainable parameters : {total_params:,}")
# ============================================================
# STEP 5 — Sanity check — forward pass with one batch
# ============================================================
encoder.eval()
with torch.no_grad():
src_sample = enc_batch.to(DEVICE) # enc_batch from cell_11
enc_outputs, enc_hidden, enc_cell = encoder(src_sample)
print(f"\n--- Forward Pass Shape Check ---")
print(f" Input (src) : {src_sample.shape} "
f"(batch_size x EN_SEQ_LEN)")
print(f" encoder_outputs shape : {enc_outputs.shape} "
f"(batch_size x EN_SEQ_LEN x hidden_dim)")
print(f" encoder hidden shape : {enc_hidden.shape} "
f"(1 x batch_size x hidden_dim)")
print(f" encoder cell shape : {enc_cell.shape} "
f"(1 x batch_size x hidden_dim)")
assert enc_outputs.shape == (BATCH_SIZE, EN_SEQ_LEN, HIDDEN_DIM), \
"encoder_outputs shape mismatch!"
assert enc_hidden.shape == (1, BATCH_SIZE, HIDDEN_DIM), \
"encoder hidden shape mismatch!"
assert enc_cell.shape == (1, BATCH_SIZE, HIDDEN_DIM), \
"encoder cell shape mismatch!"
print(f"\n All shapes correct ✓ ")
[page 42]
# ============================================================
# cell_13 — ATTENTION MECHANISM
# ============================================================
# WHAT WE ARE DOING HERE:
# At each decoding step, instead of relying only on the
# Encoder's final hidden state, the Attention mechanism
# lets the Decoder LOOK BACK at all encoder hidden states
# and decide which ones are most relevant RIGHT NOW.
#
# This is the core idea of the entire tutorial.
#
# INTUITION:
# Imagine you are translating:
# "I love Paris" → "J'aime Paris"
#
# When generating "aime", the decoder should pay most
# attention to "love" in the English sentence.
# When generating "Paris", it should focus on "Paris".
#
# Attention makes this possible — it produces a set of
# WEIGHTS (one per encoder time step) that say:
# "right now, focus 80% on word 2, 10% on word 1, ..."
#
# HOW IT WORKS (Bahdanau / Additive Attention):
#
# At each decoder step t, we have:
# decoder_hidden : what the decoder knows so far
# shape → (batch, hidden_dim)
# encoder_outputs : all encoder hidden states
# shape → (batch, src_len, hidden_dim)
#
# STEP 1 — SCORE
# Compare decoder_hidden with every encoder hidden state
# to produce a raw score (energy) for each position.
#
# energy = tanh( W1(encoder_outputs) + W2(decoder_hidden) )
# score = V(energy)
# shape → (batch, src_len, 1)
#
# STEP 2 — WEIGHTS (softmax)
# Convert raw scores to probabilities that sum to 1.
# attention_weights = softmax(score)
# shape → (batch, src_len, 1)
#
# STEP 3 — CONTEXT VECTOR
# Weighted sum of encoder_outputs using attention_weights.
# context = sum(attention_weights * encoder_outputs)
# shape → (batch, 1, hidden_dim)
#
# The context vector is a single vector that blends all
# encoder hidden states, weighted by relevance.
[page 43]
# This gets fed into the Decoder at step t.
#
# VISUAL SUMMARY:
#
# encoder_outputs → ──────────────────────────────┐
# ↓
# decoder_hidden → Score → Softmax → Weights → Weighted Sum
# ↓
# context vector
# (fed to Decoder)
# ============================================================
# ============================================================
# STEP 1 — Define the Attention class
# ============================================================
class Attention(nn.Module):
"""
Bahdanau (Additive) Attention mechanism.
Computes a context vector at each decoder step by
attending over all encoder hidden states.
Args:
hidden_dim : int — must match Encoder and Decoder hidden_dim
Forward inputs:
decoder_hidden : (batch_size, hidden_dim)
current decoder hidden state
encoder_outputs : (batch_size, src_seq_len, hidden_dim)
all encoder hidden states
Forward outputs:
context : (batch_size, 1, hidden_dim)
weighted sum of encoder outputs
attention_weights : (batch_size, src_seq_len, 1)
attention weights (sum to 1 across src_seq_len)
→ saved for visualisation in cell_20
"""
def __init__(self, hidden_dim):
super(Attention, self).__init__()
# W1 : projects encoder_outputs to hidden_dim
# input : (batch, src_len, hidden_dim)
# output : (batch, src_len, hidden_dim)
self.W1 = nn.Linear(hidden_dim, hidden_dim, bias=False)
# W2 : projects decoder_hidden to hidden_dim
# input : (batch, hidden_dim)
# output : (batch, hidden_dim)
[page 44]
self.W2 = nn.Linear(hidden_dim, hidden_dim, bias=False)
# V : collapses the hidden_dim dimension to a single score
# input : (batch, src_len, hidden_dim)
# output : (batch, src_len, 1)
self.V = nn.Linear(hidden_dim, 1, bias=False)
def forward(self, decoder_hidden, encoder_outputs):
"""
Args:
decoder_hidden : (batch_size, hidden_dim)
encoder_outputs : (batch_size, src_seq_len, hidden_dim)
Returns:
context : (batch_size, 1, hidden_dim)
attention_weights : (batch_size, src_seq_len, 1)
"""
# ── STEP 1 : SCORE ──────────────────────────────────
# Project encoder_outputs with W1
# encoder_outputs shape : (batch, src_len, hidden_dim)
# after W1 : (batch, src_len, hidden_dim)
encoder_transformed = self.W1(encoder_outputs)
# Project decoder_hidden with W2
# decoder_hidden shape : (batch, hidden_dim)
# unsqueeze(1) adds a seq dimension for broadcasting
# after unsqueeze : (batch, 1, hidden_dim)
# after W2 : (batch, 1, hidden_dim)
decoder_transformed = self.W2(decoder_hidden.unsqueeze(1))
# Broadcasting will expand (batch, 1, hidden_dim)
# to match (batch, src_len, hidden_dim) during addition
# Add and apply tanh activation to get energy scores
# energy shape : (batch, src_len, hidden_dim)
energy = torch.tanh(encoder_transformed + decoder_transformed)
# Collapse hidden_dim → single score per position
# score shape : (batch, src_len, 1)
score = self.V(energy)
# ── STEP 2 : WEIGHTS (softmax) ───────────────────────
# softmax over src_len dimension (dim=1)
# so weights across all source positions sum to 1
# attention_weights shape : (batch, src_len, 1)
attention_weights = torch.softmax(score, dim=1)
# ── STEP 3 : CONTEXT VECTOR ──────────────────────────
# Weighted sum of encoder_outputs
# attention_weights : (batch, src_len, 1)
# encoder_outputs : (batch, src_len, hidden_dim)
#
[page 45]
# transpose attention_weights → (batch, 1, src_len)
# then matmul with encoder_outputs → (batch, 1, hidden_dim)
# this is the dot product of weights with each hidden state
context = torch.bmm(
attention_weights.permute(0, 2, 1), # (batch, 1, src_len)
encoder_outputs # (batch, src_len,
hidden_dim)
)
# context shape : (batch, 1, hidden_dim)
return context, attention_weights
# ============================================================
# STEP 2 — Instantiate Attention
# ============================================================
attention = Attention(hidden_dim=HIDDEN_DIM).to(DEVICE)
# ============================================================
# STEP 3 — Print model summary
# ============================================================
print("--- Attention Architecture ---")
print(attention)
total_params = sum(p.numel() for p in attention.parameters() if
p.requires_grad)
print(f"\n Total trainable parameters : {total_params:,}")
# ============================================================
# STEP 4 — Sanity check — forward pass
# ============================================================
# Use the encoder outputs from cell_12's forward pass test
# to verify attention output shapes are correct.
attention.eval()
with torch.no_grad():
# decoder_hidden : take the encoder's final hidden state
# squeeze(0) removes the first dimension
# shape : (1, batch, hidden_dim) → (batch, hidden_dim)
sample_decoder_hidden = enc_hidden.squeeze(0) # (batch, hidden_dim)
context, attn_weights = attention(sample_decoder_hidden, enc_outputs)
print(f"\n--- Forward Pass Shape Check ---")
print(f" decoder_hidden shape : {sample_decoder_hidden.shape} "
f"(batch_size x hidden_dim)")
print(f" encoder_outputs shape : {enc_outputs.shape} "
f"(batch_size x EN_SEQ_LEN x hidden_dim)")
print(f" context shape : {context.shape} "
f"(batch_size x 1 x hidden_dim)")
[page 46]
--- Attention Architecture ---
Attention(
(W1): Linear(in_features=512, out_features=512, bias=False)
(W2): Linear(in_features=512, out_features=512, bias=False)
(V): Linear(in_features=512, out_features=1, bias=False)
)
Total trainable parameters : 524,800
--- Forward Pass Shape Check ---
decoder_hidden shape : torch.Size([32, 512]) (batch_size x hidden_dim)
encoder_outputs shape : torch.Size([32, 6, 512]) (batch_size x
EN_SEQ_LEN x hidden_dim)
context shape : torch.Size([32, 1, 512]) (batch_size x 1 x
hidden_dim)
attention_weights shape : torch.Size([32, 6, 1]) (batch_size x EN_SEQ_LEN
x 1)
--- Attention Weight Sum Check ---
First 5 weight sums (should all be 1.0) :
sample 0 : 1.000000
sample 1 : 1.000000
sample 2 : 1.000000
sample 3 : 1.000000
sample 4 : 1.000000
print(f" attention_weights shape : {attn_weights.shape} "
f"(batch_size x EN_SEQ_LEN x 1)")
assert context.shape == (BATCH_SIZE, 1, HIDDEN_DIM), \
"context shape mismatch!"
assert attn_weights.shape == (BATCH_SIZE, EN_SEQ_LEN, 1), \
"attention_weights shape mismatch!"
# ============================================================
# STEP 5 — Verify attention weights sum to 1
# ============================================================
# Since we applied softmax, weights across src positions
# must sum to 1 for every sample in the batch.
weight_sums = attn_weights.squeeze(-1).sum(dim=1) # sum across src_len
print(f"\n--- Attention Weight Sum Check ---")
print(f" First 5 weight sums (should all be 1.0) :")
for i in range(5):
print(f" sample {i} : {weight_sums[i].item():.6f}")
assert torch.allclose(weight_sums, torch.ones_like(weight_sums), atol=1e-6),
\
"Attention weights do not sum to 1!"
print(f"\n All shapes correct ✓ ")
print(f" Attention weights sum to 1 ✓ ")
[page 47]
All shapes correct ✓
Attention weights sum to 1 ✓
Case -1 LSTM Model without attention
# ============================================================
# cell_14 — DECODER WITHOUT ATTENTION
# ============================================================
# WHAT WE ARE DOING HERE:
# The Decoder generates the French translation one word
# at a time, using ONLY:
# - The encoder's final hidden and cell state
# (its entire understanding of the English sentence
# compressed into one vector)
# - Its own previous hidden and cell state
# - The previously generated word
#
# ARCHITECTURE:
#
# previous_word → Embedding → LSTM → Linear → predicted_word
# ↑
# (hidden, cell) from previous step
# initialised from Encoder's final states
#
# THE FUNDAMENTAL LIMITATION (why we need Attention):
# The entire English sentence is compressed into ONE
# fixed-size vector (hidden, cell) of shape (1, batch, 512).
# No matter how long the sentence is, all information
# must fit in this single vector.
#
# This is called the BOTTLENECK PROBLEM.
# For short sentences (like ours) it works reasonably well.
# For longer sentences it loses information — the model
# forgets early words by the time it finishes encoding.
#
# This is exactly the problem Attention solves in cell_20.
#
# NOTE:
# The Decoder processes ONE word at a time, not the whole
# sequence at once. We call forward() in a loop inside
# the Seq2Seq wrapper (cell_15).
# ============================================================
# ============================================================
# STEP 1 — Define Decoder hyperparameters
# ============================================================
DEC_DROPOUT = 0.3 # dropout rate for the Decoder
# EMBED_DIM and HIDDEN_DIM are reused from cell_12
[page 48]
# ============================================================
# STEP 2 — Define the Decoder class (without Attention)
# ============================================================
class DecoderNoAttention(nn.Module):
"""
Decodes one word at a time without Attention.
At each step it receives:
- The previous word index (or <sos> at step 0)
- Its own (hidden, cell) from the previous step
It produces:
- A probability distribution over the French vocabulary
- Updated (hidden, cell) for the next step
Args:
vocab_size : int — number of words in French vocabulary
embed_dim : int — dimension of each word embedding
hidden_dim : int — must match Encoder hidden_dim
dropout : float — dropout probability
Forward input:
trg_token : (batch_size, 1)
one token per sentence in the batch
starts as <sos>, then previously predicted word
hidden : (1, batch_size, hidden_dim)
cell : (1, batch_size, hidden_dim)
Forward output:
prediction : (batch_size, vocab_size)
raw logits over French vocabulary
hidden : (1, batch_size, hidden_dim) — updated
cell : (1, batch_size, hidden_dim) — updated
"""
def __init__(self, vocab_size, embed_dim, hidden_dim, dropout):
super(DecoderNoAttention, self).__init__()
# Embedding layer for French vocabulary
# same idea as Encoder — converts token index → dense vector
self.embedding = nn.Embedding(
vocab_size, # one row per French word
embed_dim, # each row is a vector of size embed_dim
padding_idx=PAD_IDX # <pad> embedding stays zero
)
# LSTM — processes one token at a time
# input : embedding of the previous token
# states : (hidden, cell) carried across steps
[page 49]
self.lstm = nn.LSTM(
embed_dim, # input size
hidden_dim, # hidden size — must match Encoder
batch_first=True # input shape : (batch, 1, embed_dim)
)
# Linear layer — maps hidden state → vocabulary scores
# input : (batch, hidden_dim)
# output : (batch, vocab_size)
# Each output value is a raw score (logit) for one French word
self.fc_out = nn.Linear(hidden_dim, vocab_size)
#This layer doesn't exist in the Encoder, it's unique to the Decoder.
# #Its job is to take the LSTM's hidden state and #
# convert it into a score for every single French word in the
vocabulary.
self.dropout = nn.Dropout(dropout)
#Same as before, randomly zeros out neurons during training
# #to prevent the model memorising rather than learning
def forward(self, trg_token, hidden, cell):
"""
Performs a single decoding step — processes one target token,
updates the LSTM memory states, and produces a score distribution
over the entire French vocabulary.
Called once per output token during both training and inference.
Memory states (hidden, cell) are carried across calls so the
Decoder remembers everything it has generated so far.
Args:
trg_token : (batch_size, 1) — current input token
hidden : (1, batch_size, hidden_dim)
cell : (1, batch_size, hidden_dim)
Returns:
prediction : (batch_size, vocab_size)
hidden : (1, batch_size, hidden_dim)
cell : (1, batch_size, hidden_dim)
"""
# Step 1 : embed the input token
# trg_token shape : (batch, 1)
# embedded shape : (batch, 1, embed_dim)
# trg_token is the previous French word
# that gets fed into the Decoder to help it predict the next French
word.
embedded = self.dropout(self.embedding(trg_token))
# Step 2 : pass through LSTM
# lstm_out shape : (batch, 1, hidden_dim)
# hidden shape : (1, batch, hidden_dim)
[page 50]
# cell shape : (1, batch, hidden_dim)
lstm_out, (hidden, cell) = self.lstm(embedded, (hidden, cell))
# Note : we pass (hidden, cell) in to continue from where
# the previous step left off
# Step 3 : squeeze the seq dimension (it is always 1 here)
# lstm_out shape : (batch, 1, hidden_dim) → (batch, hidden_dim)
lstm_out = lstm_out.squeeze(1)
# The LSTM returns lstm_out shaped (batch, 1, hidden_dim).
# The middle dimension is always 1 here because we process
# exactly one token per forward call.
# Step 4 : project to vocabulary size to get raw logits
# prediction shape : (batch, vocab_size)
prediction = self.fc_out(lstm_out)
return prediction, hidden, cell
# ============================================================
# STEP 3 — Instantiate the Decoder
# ============================================================
decoder_no_attn = DecoderNoAttention(
vocab_size = len(french_vocab), # French vocab size from cell_5
embed_dim = EMBED_DIM, # same as Encoder
hidden_dim = HIDDEN_DIM, # must match Encoder
dropout = DEC_DROPOUT
).to(DEVICE)
# ============================================================
# STEP 4 — Print model summary
# ============================================================
print("--- Decoder (No Attention) Architecture ---")
print(decoder_no_attn)
total_params = sum(p.numel() for p in decoder_no_attn.parameters()
if p.requires_grad)
print(f"\n Total trainable parameters : {total_params:,}")
# ============================================================
# STEP 5 — Sanity check — one forward step
# ============================================================
# Simulate the very first decoder step:
# - Input token : <sos> for every sentence in the batch
# - hidden, cell : taken from the Encoder's final states
decoder_no_attn.eval()
with torch.no_grad():
# First input token is always <sos>
[page 51]
# shape : (batch_size, 1)
sos_input = torch.full(
(BATCH_SIZE, 1), # shape
SOS_IDX, # fill value
dtype=torch.long
).to(DEVICE)
# torch.full creates a tensor of a given shape
# filled entirely with one value.
# Here we're filling it with SOS_IDX, the start of sequence token.
# Use encoder's final hidden and cell to initialise decoder
sample_prediction, sample_hidden, sample_cell = decoder_no_attn(
sos_input, # (batch, 1)
enc_hidden, # (1, batch, hidden_dim) from cell_12
enc_cell # (1, batch, hidden_dim) from cell_12
)
print(f"\n--- Forward Pass Shape Check ---")
print(f" Input token shape : {sos_input.shape} "
f"(batch_size x 1)")
print(f" prediction shape : {sample_prediction.shape} "
f"(batch_size x french_vocab_size)")
print(f" hidden shape : {sample_hidden.shape} "
f"(1 x batch_size x hidden_dim)")
print(f" cell shape : {sample_cell.shape} "
f"(1 x batch_size x hidden_dim)")
assert sample_prediction.shape == (BATCH_SIZE, len(french_vocab)), \
"prediction shape mismatch!"
assert sample_hidden.shape == (1, BATCH_SIZE, HIDDEN_DIM), \
"hidden shape mismatch!"
assert sample_cell.shape == (1, BATCH_SIZE, HIDDEN_DIM), \
"cell shape mismatch!"
#"Three asserts, one per output. If any shape is wrong,
# Python crashes immediately with a clear message.
# Catch it here rather than getting a confusing error deep inside training."
# ============================================================
# STEP 6 — Peek at predicted token
# ============================================================
# argmax gives the index with the highest logit score —
# this is the model's best guess for the next French word.
# At this point weights are random so the prediction is
# meaningless — but the shape and mechanics are correct.
predicted_indices = sample_prediction.argmax(dim=1) # (batch,)
print(f"\n--- Predicted Token Indices (first 5, random weights) ---")
for i in range(5):
idx = predicted_indices[i].item()
word = french_vocab.idx2word.get(idx, UNK_TOKEN)
print(f" sample {i} : index {idx:<6} word → '{word}'")
[page 52]
--- Decoder (No Attention) Architecture ---
DecoderNoAttention(
(embedding): Embedding(1207, 256, padding_idx=0)
(lstm): LSTM(256, 512, batch_first=True)
(fc_out): Linear(in_features=512, out_features=1207, bias=True)
(dropout): Dropout(p=0.3, inplace=False)
)
Total trainable parameters : 2,505,143
--- Forward Pass Shape Check ---
Input token shape : torch.Size([32, 1]) (batch_size x 1)
prediction shape : torch.Size([32, 1207]) (batch_size x
french_vocab_size)
hidden shape : torch.Size([1, 32, 512]) (1 x batch_size x
hidden_dim)
cell shape : torch.Size([1, 32, 512]) (1 x batch_size x
hidden_dim)
--- Predicted Token Indices (first 5, random weights) ---
sample 0 : index 381 word → 'comment'
sample 1 : index 381 word → 'comment'
sample 2 : index 381 word → 'comment'
sample 3 : index 381 word → 'comment'
sample 4 : index 381 word → 'comment'
Predictions are random (untrained) — this is expected ✓
Ready for Cell 15 — Seq2Seq without Attention
print(f"\n Predictions are random (untrained) — this is expected ✓ ")
print(f"\n Ready for Cell 15 — Seq2Seq without Attention")
# ============================================================
# cell_15 — SEQ2SEQ WITHOUT ATTENTION
# ============================================================
# WHAT WE ARE DOING HERE:
# The Seq2Seq class is the GLUE that connects the Encoder
# and Decoder into one complete model.
#
# It handles the decoding LOOP — calling the Decoder
# repeatedly, one token at a time, until the full target
# sequence is generated.
#
# HOW THE LOOP WORKS:
#
# Step 0 : feed <sos> to Decoder
# ↓
# Step 1 : Decoder predicts word 1
# ↓
# Step 2 : Decoder predicts word 2
# ↓
[page 53]
# ...until FR_SEQ_LEN steps are done
#
# TEACHER FORCING:
# During training we use a technique called Teacher Forcing.
#
# WITHOUT teacher forcing:
# At each step, the Decoder feeds its OWN prediction
# as the next input. If it predicts the wrong word early,
# the error compounds — training becomes unstable.
#
# WITH teacher forcing:
# At each step, we feed the ACTUAL correct target word
# as the next input, regardless of what the Decoder
# predicted. This stabilises training significantly.
#
# We use a TEACHER_FORCING_RATIO to control how often
# we apply it:
# ratio = 1.0 → always use ground truth (full teacher forcing)
# ratio = 0.0 → always use model's own prediction
# ratio = 0.5 → 50% chance of using ground truth
#
# During INFERENCE (translating new sentences) we always
# set ratio = 0.0 because there is no ground truth available.
#
# ============================================================
# ============================================================
# STEP 1 — Define the Seq2Seq class (without Attention)
# ============================================================
class Seq2SeqNoAttention(nn.Module):
"""
Complete Encoder → Decoder model without Attention.
Args:
encoder : Encoder — encodes the source sentence
decoder : DecoderNoAttention — decodes one step at a time
Forward inputs:
src : (batch_size, src_seq_len) — English indices
trg : (batch_size, trg_seq_len) — French indices
teacher_forcing_ratio : float (0.0 to 1.0)
Forward output:
outputs : (batch_size, trg_seq_len, french_vocab_size)
logits at every decoder time step
"""
def __init__(self, encoder, decoder):
super(Seq2SeqNoAttention, self).__init__()
self.encoder = encoder
self.decoder = decoder
[page 54]
def forward(self, src, trg, teacher_forcing_ratio=0.5):
"""
Args:
src : (batch, src_seq_len)
trg : (batch, trg_seq_len)
teacher_forcing_ratio: float — probability of using
ground truth as next input
Returns:
outputs : (batch, trg_seq_len, vocab_size)
"""
# src is padded english token idices
# trg is padded French token indices
batch_size = src.shape[0] # number of sentences in this batch
trg_seq_len = trg.shape[1] # number of tokens in target sequence
vocab_size = len(french_vocab)
# EXTRACT DIMENSIONS
# Read batch size and target sequence length directly from
# the input tensors so this method works for any batch size
# or sequence length without hardcoding.
# Tensor to store all decoder predictions
# initialised to zeros — will be filled step by step
# shape : (batch, trg_seq_len, vocab_size)
outputs = torch.zeros(
batch_size, trg_seq_len, vocab_size
).to(DEVICE)
# ── ENCODE ───────────────────────────────────────────
# Run the encoder once over the full source sentence
# encoder_outputs : (batch, src_seq_len, hidden_dim)
# hidden : (1, batch, hidden_dim)
# cell : (1, batch, hidden_dim)
encoder_outputs, hidden, cell = self.encoder(src)
# Note : encoder_outputs are not used here (no attention)
# only the final hidden and cell are passed to decoder
# ── DECODE LOOP ──────────────────────────────────────
# trg[:,0] is always <sos> for every sentence in the batch
# shape : (batch,) → unsqueeze to (batch, 1)
# INITIALISING FIRST DECODER INPUT
input_token = trg[:, 0].unsqueeze(1) # first input is <sos>
# Start of DECODE LOOP
# Generate one French token per iteration.
# Start at t=1 because t=0 is the <sos> input — not a prediction.
# End at trg_seq_len — one step per target position.
# ----------------------------------------------------------------
for t in range(1, trg_seq_len):
# One decoder step
[page 55]
# prediction : (batch, vocab_size)
# hidden : (1, batch, hidden_dim) — updated
# cell : (1, batch, hidden_dim) — updated
prediction, hidden, cell = self.decoder(
input_token, hidden, cell
)
# Store prediction for this time step
outputs[:, t, :] = prediction # (batch, vocab_size)
# ── TEACHER FORCING ──────────────────────────────
# Flip a coin — should we use teacher forcing this step?
# With teacher_forcing_ratio=0.5:
# 50% of steps → feed correct French word (teacher forcing)
# 50% of steps → feed model's own prediction (free running)
# ------------------------------------------------------------
use_teacher_forcing = (
torch.rand(1).item() < teacher_forcing_ratio
)
if use_teacher_forcing:
# Use the ACTUAL next word as input
# trg[:, t] shape : (batch,) → (batch, 1)
input_token = trg[:, t].unsqueeze(1)
else:
# Use the model's OWN prediction as next input
# argmax picks the highest scoring word index
# shape : (batch,) → (batch, 1)
input_token = prediction.argmax(dim=1).unsqueeze(1)
return outputs
# ============================================================
# STEP 2 — Instantiate the Seq2Seq model
# ============================================================
TEACHER_FORCING_RATIO = 0.5 # 50% teacher forcing during training
model_no_attn = Seq2SeqNoAttention(
encoder = encoder, # Encoder from cell_12
decoder = decoder_no_attn # Decoder from cell_14
).to(DEVICE)
# ============================================================
# STEP 3 — Print model summary
# ============================================================
print("--- Seq2Seq (No Attention) Architecture ---")
print(model_no_attn)
# Count total parameters across encoder + decoder
total_params = sum(
[page 56]
p.numel() for p in model_no_attn.parameters() if p.requires_grad
)
print(f"\n Total trainable parameters : {total_params:,}")
# ============================================================
# STEP 4 — Sanity check — full forward pass
# ============================================================
# Run one complete batch through the full model and confirm
# output shape is correct before we write the training loop.
model_no_attn.eval()
with torch.no_grad():
src_sample = enc_batch.to(DEVICE) # (batch, EN_SEQ_LEN)
trg_sample = dec_inp_batch.to(DEVICE) # (batch, FR_SEQ_LEN)
# teacher_forcing_ratio=0 so we use model predictions only
# (simulates inference — no ground truth peeking)
sample_output = model_no_attn(
src_sample,
trg_sample,
teacher_forcing_ratio=0.0
)
print(f"\n--- Forward Pass Shape Check ---")
print(f" src shape : {src_sample.shape} "
f"(batch_size x EN_SEQ_LEN)")
print(f" trg shape : {trg_sample.shape} "
f"(batch_size x FR_SEQ_LEN)")
print(f" output shape : {sample_output.shape} "
f"(batch_size x FR_SEQ_LEN x french_vocab_size)")
assert sample_output.shape == (BATCH_SIZE, FR_SEQ_LEN, len(french_vocab)), \
"output shape mismatch!"
# ============================================================
# STEP 5 — Decode a sample prediction
# ============================================================
# Take the first sentence in the batch, pick the highest
# scoring token at each step, and decode back to French words.
# Weights are random so translation is gibberish — expected.
print(f"\n--- Sample Prediction (random weights — gibberish expected) ---")
# argmax across vocab dimension → predicted index at each step
# shape : (FR_SEQ_LEN,)
predicted_indices = sample_output[0].argmax(dim=1)
# At each time step we have one score per French word.
# argmax(dim=1) picks the index of the highest score at each step
# that's the model's predicted word."
src_words = [
[page 57]
--- Seq2Seq (No Attention) Architecture ---
Seq2SeqNoAttention(
(encoder): Encoder(
(embedding): Embedding(858, 256, padding_idx=0)
(lstm): LSTM(256, 512, batch_first=True)
(dropout): Dropout(p=0.3, inplace=False)
)
(decoder): DecoderNoAttention(
(embedding): Embedding(1207, 256, padding_idx=0)
(lstm): LSTM(256, 512, batch_first=True)
(fc_out): Linear(in_features=512, out_features=1207, bias=True)
(dropout): Dropout(p=0.3, inplace=False)
)
)
Total trainable parameters : 4,301,751
--- Forward Pass Shape Check ---
src shape : torch.Size([32, 6]) (batch_size x EN_SEQ_LEN)
trg shape : torch.Size([32, 6]) (batch_size x FR_SEQ_LEN)
output shape : torch.Size([32, 6, 1207]) (batch_size x FR_SEQ_LEN x
french_vocab_size)
--- Sample Prediction (random weights — gibberish expected) ---
Source (English) : ['is', 'tom', 'here', '<eos>', '<pad>', '<pad>']
Predicted (French): ['<pad>', 'comment', 'amende', 'égal', 'dabord',
'parlezmoi']
Output shape correct ✓
Model is untrained — gibberish output is expected ✓
Ready for Cell 16 — Training Loop
english_vocab.idx2word.get(i.item(), UNK_TOKEN)
for i in src_sample[0]
]
pred_words = [
french_vocab.idx2word.get(i.item(), UNK_TOKEN)
for i in predicted_indices
]
print(f" Source (English) : {src_words}")
print(f" Predicted (French): {pred_words}")
print(f"\n Output shape correct ✓ ")
print(f" Model is untrained — gibberish output is expected ✓ ")
print(f"\n Ready for Cell 16 — Training Loop")
# ============================================================
# cell_16 — TRAINING LOOP (WITH VALIDATION)
# ============================================================
# WHAT WE ARE DOING HERE:
# This cell is the heart of the entire project — where the
# model actually learns to translate English to French.
[page 58]
# We repeatedly show the model sentence pairs, measure how
# wrong its predictions are, and adjust its weights to
# reduce that error. This process repeats for N_EPOCHS.
#
# KEY CONCEPTS:
#
# LOSS FUNCTION — CrossEntropyLoss
# Measures how wrong the model's predictions are at each
# decoding step. ignore_index=PAD_IDX ensures padding
# positions contribute nothing to the loss — we only
# penalise the model for real word predictions.
#
# OPTIMIZER — Adam
# Reads the gradients after each backward pass and nudges
# every weight in the direction that reduces loss.
# Adapts its step size per parameter automatically based
# on gradient history — more stable than plain SGD.
#
# LEARNING RATE SCHEDULER — ReduceLROnPlateau
# Monitors val loss after every epoch.
# If val loss does not improve for `patience` epochs,
# the learning rate is multiplied by `factor`.
# Allows the model to take smaller, more precise steps
# as it approaches the optimum — often squeezing out
# extra performance that a fixed learning rate misses.
#
# GRADIENT CLIPPING
# Caps the maximum gradient norm at CLIP=1.0.
# Prevents exploding gradients — a common failure mode
# in LSTMs where a single bad batch causes catastrophic
# weight updates that destroy all prior learning.
#
# TEACHER FORCING
# At each decoder step, flip a coin — 50% of the time
# feed the correct French word as the next input instead
# of the model's own prediction. Stabilises early training
# by preventing compounding errors from derailing learning.
#
# VALIDATION AFTER EVERY EPOCH
# model.eval() + torch.no_grad() + teacher_forcing=0.0
# Gives an honest measure of generalisation — the model
# gets no help and no gradients are tracked.
# We save the best model by VAL LOSS, not train loss,
# because low train loss alone can mean memorisation.
# ============================================================
from tqdm.notebook import tqdm
import torch.optim as optim
import math
# ============================================================
[page 59]
# STEP 1 — Hyperparameters
# ============================================================
LEARNING_RATE = 0.001 # initial step size for Adam optimiser
N_EPOCHS = 10 # number of full passes through training data
CLIP = 1.0 # maximum gradient norm — prevents exploding
gradients
# ============================================================
# STEP 2 — Loss function, Optimizer, and Scheduler
# ============================================================
# CrossEntropyLoss measures prediction error at each decoder step.
# ignore_index=PAD_IDX tells the loss function to skip <pad> positions
# entirely — they carry no meaning and should not influence learning.
criterion = nn.CrossEntropyLoss(ignore_index=PAD_IDX)
# Adam optimizer — responsible for updating all model weights
# after each backward pass using the computed gradients.
# model.parameters() passes every trainable weight and bias
# in the Encoder, Decoder, and all sub-layers to the optimizer.
optimizer_no_attn = optim.Adam(
model_no_attn.parameters(),
lr=LEARNING_RATE
)
# ReduceLROnPlateau scheduler — monitors val loss each epoch
# and reduces the learning rate when improvement stalls.
# This allows the model to take progressively smaller,
# more precise steps as it converges toward a good solution.
#
# mode='min' → we want val loss to go DOWN
# factor=0.5 → multiply lr by 0.5 when triggered
# e.g. 0.001 → 0.0005 → 0.00025
# patience=2 → tolerate 2 epochs of no improvement
# before reducing — avoids reacting to noise
scheduler_no_attn = optim.lr_scheduler.ReduceLROnPlateau(
optimizer_no_attn,
mode='min',
factor=0.5,
patience=2,
)
# ============================================================
# STEP 3 — Define training epoch function
# ============================================================
def train_one_epoch(model, loader, optimizer, criterion, clip):
"""
Runs one full pass through the training DataLoader.
Processes every batch once, computes loss, backpropagates
gradients, clips them for safety, and updates weights.
[page 60]
Weights ARE updated here — this is the learning step.
Args:
model : Seq2Seq model to train
loader : DataLoader supplying (src, trg_input, trg_target) batches
optimizer : Adam optimizer managing weight updates
criterion : CrossEntropyLoss measuring prediction error
clip : float — maximum gradient norm for clipping
Returns:
avg_loss : float — mean loss across all batches in this epoch
"""
# Switch model to training mode.
# Critical effect: dropout turns ON, randomly zeroing neurons
# each forward pass to prevent over-reliance on any single path.
model.train()
# Running total of loss across all batches.
# Divided by batch count at the end to return average epoch loss.
epoch_loss = 0.0
# Wrap the DataLoader in tqdm to display a live progress bar.
# desc sets the label shown on the left of the bar.
# leave=False removes the bar when the epoch finishes — keeps output
clean.
progress_bar = tqdm(loader, desc=' Training', leave=False)
for src, trg_input, trg_target in progress_bar:
# Move all three tensors to the same device as the model.
# Model and data must always be on the same device —
# mixing CPU and GPU tensors causes an immediate crash.
src = src.to(DEVICE)
trg_input = trg_input.to(DEVICE)
trg_target = trg_target.to(DEVICE)
# ── FORWARD PASS ─────────────────────────────────────
# Run the full Encoder → Decoder pipeline.
# The Encoder reads src once. The Decoder generates
# predictions for every French token position using
# teacher forcing at the rate set by TEACHER_FORCING_RATIO.
# output shape : (batch, trg_seq_len, vocab_size)
output = model(src, trg_input, TEACHER_FORCING_RATIO)
# ── PREPARE LOSS INPUTS ──────────────────────────────
# Skip position 0 on both sides — position 0 is <sos>,
# the starting signal fed INTO the decoder. We never
# make a prediction for it, so it must be excluded
# from loss computation on both the output and target side.
#
# Reshape from 3D → 2D (output) and 2D → 1D (target)
[page 61]
# because CrossEntropyLoss expects flat inputs:
# output_for_loss : (batch × seq_len, vocab_size)
# target_for_loss : (batch × seq_len,)
output_for_loss = output[:, 1:, :].reshape(-1, len(french_vocab))
target_for_loss = trg_target[:, 1:].reshape(-1)
# Compute loss — compares predicted scores against correct indices.
# <pad> positions are automatically ignored via ignore_index.
loss = criterion(output_for_loss, target_for_loss)
# ── BACKWARD PASS AND WEIGHT UPDATE ──────────────────
# Clear gradients from the previous batch.
# PyTorch accumulates gradients by default — if not zeroed,
# old gradients contaminate the current batch's update.
optimizer.zero_grad()
# Backpropagate — compute gradient of loss with respect
# to every trainable weight in the model by walking
# backwards through the computational graph.
loss.backward()
# Clip gradient norm to maximum value of `clip`.
# If the total gradient norm exceeds clip, all gradients
# are scaled down proportionally.
# Prevents exploding gradients — especially critical for
# LSTMs which are prone to runaway gradient accumulation
# over long sequences.
torch.nn.utils.clip_grad_norm_(model.parameters(), clip)
# Apply the weight update — nudges every parameter in
# the direction that reduces loss, using Adam's adaptive
# per-parameter step sizes derived from gradient history.
optimizer.step()
# ── TRACKING ─────────────────────────────────────────
# Accumulate this batch's scalar loss into epoch total.
# .item() converts the single-element tensor to a Python float.
epoch_loss += loss.item()
# Update the progress bar with this batch's loss in real time.
progress_bar.set_postfix({'batch_loss': f'{loss.item():.4f}'})
# Return average loss across all batches in this epoch.
# Dividing by len(loader) — the number of batches — normalises
# for varying batch count and makes epochs comparable.
return epoch_loss / len(loader)
# ============================================================
# STEP 4 — Define validation / evaluation function
# ============================================================
[page 62]
def evaluate(model, loader, criterion):
"""
Evaluates the model on a DataLoader without updating weights.
Used after every training epoch to measure generalisation.
Three critical differences from train_one_epoch:
- model.eval() : dropout OFF — deterministic output
- torch.no_grad() : no gradient graph built — saves memory
- teacher_forcing = 0.0 : model flies solo — no ground truth peeking
Teacher forcing is disabled here because validation must mirror
real inference conditions — the model only sees its own predictions,
not the correct answers. This gives an honest measure of quality.
Args:
model : Seq2Seq model to evaluate
loader : DataLoader supplying (src, trg_input, trg_target) batches
criterion : CrossEntropyLoss measuring prediction error
Returns:
avg_loss : float — mean loss across all batches
"""
# Switch to evaluation mode.
# Dropout turns OFF — all neurons active, output is deterministic.
# BatchNorm (if present) uses running statistics rather than batch stats.
model.eval()
epoch_loss = 0.0
progress_bar = tqdm(loader, desc=' Validating', leave=False)
# Disable gradient tracking for the entire validation pass.
# We are only measuring performance — never calling loss.backward().
# torch.no_grad() saves memory and speeds up computation by not
# building the computational graph that backward() would need.
with torch.no_grad():
for src, trg_input, trg_target in progress_bar:
src = src.to(DEVICE)
trg_input = trg_input.to(DEVICE)
trg_target = trg_target.to(DEVICE)
# Forward pass with teacher_forcing_ratio=0.0.
# The model uses its own predictions at every step —
# exactly as it would during real translation inference.
# This gives a true measure of translation quality.
output = model(src, trg_input, teacher_forcing_ratio=0.0)
# Same reshaping as training — skip <sos> at position 0,
[page 63]
# flatten to 2D for CrossEntropyLoss.
output_for_loss = output[:, 1:, :].reshape(-1, len(french_vocab))
target_for_loss = trg_target[:, 1:].reshape(-1)
loss = criterion(output_for_loss, target_for_loss)
# Accumulate loss — no backward pass, no weight updates.
epoch_loss += loss.item()
progress_bar.set_postfix({'batch_loss': f'{loss.item():.4f}'})
return epoch_loss / len(loader)
# ============================================================
# STEP 5 — Run the training + validation loop
# ============================================================
print("=" * 65)
print(" Training Seq2Seq WITHOUT Attention")
print("=" * 65)
# Lists to record metrics at every epoch for plotting later.
train_losses_no_attn = [] # training loss per epoch
val_losses_no_attn = [] # validation loss per epoch
lr_history_no_attn = [] # learning rate per epoch — shows when scheduler
fired
# Track the best validation loss seen so far.
# Initialised to infinity so any real loss immediately beats it.
# We save model weights only when this improves.
best_val_loss = float('inf')
best_epoch = 0
# Outer progress bar — one tick per epoch across all N_EPOCHS.
# leave=True keeps it visible after training completes.
epoch_bar = tqdm(range(1, N_EPOCHS + 1), desc='Epochs', leave=True)
for epoch in epoch_bar:
# ── TRAIN ─────────────────────────────────────────────
# Run one full pass through training data.
# Weights ARE updated. Dropout is ON. Teacher forcing active.
train_loss = train_one_epoch(
model_no_attn, train_loader, optimizer_no_attn, criterion, CLIP
)
# ── VALIDATE ──────────────────────────────────────────
# Evaluate on held-out validation set.
# Weights are NOT updated. Dropout OFF. No teacher forcing.
val_loss = evaluate(model_no_attn, val_loader, criterion)
# Record both losses for the plot in Step 7.
[page 64]
train_losses_no_attn.append(train_loss)
val_losses_no_attn.append(val_loss)
# ── SCHEDULER STEP ────────────────────────────────────
# Must be called AFTER validation — the scheduler needs
# the val loss to decide whether to reduce the learning rate.
# If val loss has not improved for `patience=2` epochs,
# the learning rate is multiplied by factor=0.5.
scheduler_no_attn.step(val_loss)
# Read the current learning rate directly from the optimizer.
# param_groups[0] is the first (and only) parameter group here.
# Stored for the learning rate plot in Step 7.
current_lr = optimizer_no_attn.param_groups[0]['lr']
lr_history_no_attn.append(current_lr)
# Perplexity = e^loss — an alternative readability metric.
# Lower is better. A perplexity of 1.0 = perfect prediction.
# Perplexity of 100 = the model is very confused.
train_ppl = math.exp(train_loss)
val_ppl = math.exp(val_loss)
# ── SAVE BEST MODEL ───────────────────────────────────
# Save weights only when val loss improves — not train loss.
# Saving by val loss protects against saving an overfitted model
# that scores well on training data but poorly on new sentences.
# state_dict() saves only the weights — not the model architecture.
if val_loss < best_val_loss:
best_val_loss = val_loss
best_epoch = epoch
torch.save(model_no_attn.state_dict(), 'best_model_no_attn.pt')
saved_marker = ' ← best model saved'
else:
saved_marker = ''
# Update the outer epoch progress bar with current metrics.
epoch_bar.set_postfix({
'train_loss': f'{train_loss:.4f}',
'val_loss' : f'{val_loss:.4f}',
'lr' : f'{current_lr:.6f}',
'best_ep' : best_epoch
})
# Print one summary line per epoch below the progress bar.
# tqdm.write() is used instead of print() to avoid
# interfering with the progress bar display.
tqdm.write(
f" Epoch {epoch:02d}/{N_EPOCHS} | "
f"Train Loss: {train_loss:.4f} PPL: {train_ppl:.4f} | "
f"Val Loss: {val_loss:.4f} PPL: {val_ppl:.4f} | "
f"LR: {current_lr:.6f}"
f"{saved_marker}"
[page 65]
)
print("=" * 65)
print(f"\n Best Val Loss : {best_val_loss:.4f} (Epoch {best_epoch})")
print(f" Best model saved → best_model_no_attn.pt")
# ============================================================
# STEP 6 — Load best model weights for inference
# ============================================================
# Reload the weights from the best checkpoint saved during training.
# The model's current weights at the end of training may not be
# the best — val loss could have peaked at epoch 7 and degraded since.
# map_location=DEVICE ensures weights load correctly regardless of
# whether they were saved on GPU and are now being loaded on CPU.
model_no_attn.load_state_dict(
torch.load('best_model_no_attn.pt', map_location=DEVICE)
)
print(f"\n Best model weights reloaded ✓ ")
# ============================================================
# STEP 7 — Loss curve plot
# ============================================================
# Two side-by-side plots — loss curves and learning rate history.
fig, axes = plt.subplots(1, 2, figsize=(14, 4))
# Plot 1 — Train vs Validation Loss
# Healthy training: both curves fall together.
# Overfitting: train keeps falling, val plateaus or rises.
axes[0].plot(range(1, N_EPOCHS + 1), train_losses_no_attn,
label='Train Loss', color='steelblue', marker='o')
axes[0].plot(range(1, N_EPOCHS + 1), val_losses_no_attn,
label='Val Loss', color='coral', marker='o')
axes[0].set_xlabel('Epoch')
axes[0].set_ylabel('Loss')
axes[0].set_title('Seq2Seq WITHOUT Attention — Train vs Val Loss')
axes[0].legend()
axes[0].grid(True, alpha=0.3)
# Plot 2 — Learning Rate Schedule
# Flat sections = no scheduler action.
# Step-downs = scheduler fired because val loss stalled.
# Useful for confirming the scheduler behaved as expected.
axes[1].plot(range(1, N_EPOCHS + 1), lr_history_no_attn,
color='green', marker='o', linewidth=2)
axes[1].set_xlabel('Epoch')
axes[1].set_ylabel('Learning Rate')
axes[1].set_title('Learning Rate Schedule — No Attention')
axes[1].grid(True, alpha=0.3)
plt.tight_layout()
[page 66]
=================================================================
Training Seq2Seq WITHOUT Attention
=================================================================
/usr/local/lib/python3.12/dist-packages/torch/utils/data/dataloader.py:775:
UserWarning: 'pin_memory' argument is set as true but no accelerator is
found, then device pinned memory won't be used.
super().__init__(loader)
Epoch 01/10 | Train Loss: 4.2981 PPL: 73.5594 | Val Loss: 3.9742 PPL:
53.2053 | LR: 0.001000 ← best model saved
plt.show()
# ============================================================
# STEP 8 — Overfitting check
# ============================================================
# Compare final train and val loss to diagnose model health.
# A small gap is normal. A large gap signals overfitting.
final_train = train_losses_no_attn[-1]
final_val = val_losses_no_attn[-1]
gap = final_val - final_train # positive = val worse than train
print(f"\n--- Overfitting Check ---")
print(f" Final Train Loss : {final_train:.4f}")
print(f" Final Val Loss : {final_val:.4f}")
print(f" Gap (val-train) : {gap:.4f}")
if gap > 0.5:
# Large gap — model memorised training data instead of generalising.
# Increasing dropout adds regularisation noise during training,
# forcing the model to learn more robust representations.
print(f" ⚠ Large gap — model may be overfitting.")
print(f" Consider increasing ENC_DROPOUT or DEC_DROPOUT.")
elif gap < 0:
# Val loss lower than train loss — sounds wrong but is expected.
# Dropout is ON during training (hurts train loss slightly) and
# OFF during validation (full model active — slightly better
performance).
print(f" ⚠ Val loss lower than train — unusual.")
print(f" Likely due to dropout being OFF during eval.")
else:
# Healthy gap — model is generalising well to unseen sentences.
print(f" ✓ Gap looks healthy.")
print(f"\n Ready for Cell 17 — Evaluate & BLEU Score (No Attention)")
{"model_id":"c3abbba951834615888eefea83fc6bca","version_major":2,"version_minor":0}
{"model_id":"8ccffd09b6bf4bc0bf0e1ea8b342d657","version_major":2,"version_minor":0}
{"model_id":"bcf14c4a95e7497385a735a12614cd84","version_major":2,"version_minor":0}
[page 67]
Epoch 02/10 | Train Loss: 3.5315 PPL: 34.1738 | Val Loss: 3.5857 PPL:
36.0799 | LR: 0.001000 ← best model saved
Epoch 03/10 | Train Loss: 2.9837 PPL: 19.7605 | Val Loss: 3.2436 PPL:
25.6263 | LR: 0.001000 ← best model saved
Epoch 04/10 | Train Loss: 2.5183 PPL: 12.4071 | Val Loss: 3.0248 PPL:
20.5907 | LR: 0.001000 ← best model saved
Epoch 05/10 | Train Loss: 2.1232 PPL: 8.3581 | Val Loss: 2.8407 PPL:
17.1286 | LR: 0.001000 ← best model saved
Epoch 06/10 | Train Loss: 1.7965 PPL: 6.0283 | Val Loss: 2.7116 PPL:
15.0530 | LR: 0.001000 ← best model saved
Epoch 07/10 | Train Loss: 1.5432 PPL: 4.6794 | Val Loss: 2.6149 PPL:
13.6661 | LR: 0.001000 ← best model saved
Epoch 08/10 | Train Loss: 1.3180 PPL: 3.7360 | Val Loss: 2.5480 PPL:
12.7811 | LR: 0.001000 ← best model saved
Epoch 09/10 | Train Loss: 1.1357 PPL: 3.1133 | Val Loss: 2.5561 PPL:
12.8860 | LR: 0.001000
{"model_id":"a1916450ccfa40c382f7193e47125333","version_major":2,"version_minor":0}
{"model_id":"7b43e738f66f4d24886cd488736c7c86","version_major":2,"version_minor":0}
{"model_id":"183623f88c65417f97dd60c69254b485","version_major":2,"version_minor":0}
{"model_id":"99d08cb25cdf4d3f9eaf85bc0070f302","version_major":2,"version_minor":0}
{"model_id":"bb2750af8d2e4831955cd4c40048c133","version_major":2,"version_minor":0}
{"model_id":"5b61189d168044ebb157acbf842f07c4","version_major":2,"version_minor":0}
{"model_id":"7a6f957950d54f89a66d11b779e9ae3a","version_major":2,"version_minor":0}
{"model_id":"3a21bf49f0994b308c815f229768689e","version_major":2,"version_minor":0}
{"model_id":"48c184b023f545bba3b361688c9a49af","version_major":2,"version_minor":0}
{"model_id":"0f96ad25d1684bcb9fe3cbac318b8ad7","version_major":2,"version_minor":0}
{"model_id":"31858ded6ab2431e97752b34743697a3","version_major":2,"version_minor":0}
{"model_id":"90385462e1324fd7a68061252a2b98c8","version_major":2,"version_minor":0}
{"model_id":"75da27941c374611a69df3869e7c5f6b","version_major":2,"version_minor":0}
{"model_id":"f17cc761d489494facdf6904a9f82040","version_major":2,"version_minor":0}
{"model_id":"a10ffaa66d344684a6dae490fe02c427","version_major":2,"version_minor":0}
{"model_id":"ce676611fb444f549708e522a91840a1","version_major":2,"version_minor":0}
[page 68]
Epoch 10/10 | Train Loss: 0.9959 PPL: 2.7073 | Val Loss: 2.5345 PPL:
12.6106 | LR: 0.001000 ← best model saved
=================================================================
Best Val Loss : 2.5345 (Epoch 10)
Best model saved → best_model_no_attn.pt
Best model weights reloaded ✓
--- Overfitting Check ---
Final Train Loss : 0.9959
Final Val Loss : 2.5345
Gap (val-train) : 1.5386
⚠ Large gap — model may be overfitting.
Consider increasing ENC_DROPOUT or DEC_DROPOUT.
Ready for Cell 17 — Evaluate & BLEU Score (No Attention)
{"model_id":"cd24e27295da4bcb97f85c82ac5f6287","version_major":2,"version_minor":0}
{"model_id":"a1a4f4d553d14af4b1cfe4c08ba315be","version_major":2,"version_minor":0}
#cell_17 run this if needed:!pip install ipywidgets
# ============================================================
# cell_18 — EVALUATE & BLEU SCORE (NO ATTENTION)
# ============================================================
# WHAT WE ARE DOING HERE:
# We evaluate the best trained model (no attention) on the
# TEST set — data the model has NEVER seen during training
# or validation.
#
# This is the FIRST and ONLY time we touch the test set.
# We never used it during training or model selection.
# This is what makes the test score an HONEST measure
# of how well the model generalises to truly new data.
#
# HOW DOES THIS DIFFER FROM VALIDATION (cell_16)?
#
# ┌─────────────────┬──────────────────┬─────────────────┐
# │ │ Validation │ Test │
# ├─────────────────┼──────────────────┼─────────────────┤
[page 69]
# │ When used │ Every epoch │ Once at the end │
# │ Purpose │ Monitor training │ Final score │
# │ │ Save best model │ │
# │ Influences │ Yes — we save │ No — never │
# │ training? │ best model │ touched before │
# │ Teacher forcing │ 0.0 │ 0.0 │
# │ Weight updates │ No │ No │
# └─────────────────┴──────────────────┴─────────────────┘
#
# WE MEASURE PERFORMANCE USING TWO METRICS:
#
# 1. TEST LOSS (CrossEntropyLoss)
# How confident the model is at each prediction step.
# Same metric used during training — easy to compare
# directly against train and val loss.
#
# 2. BLEU SCORE (Bilingual Evaluation Understudy)
# The standard metric for translation quality.
# Compares predicted translation to reference translation
# by counting matching n-grams (word sequences).
#
# BLEU score ranges from 0.0 to 1.0 (or 0 to 100):
# < 10 → almost unusable
# 10-20 → getting some meaning across
# 20-30 → understandable with effort
# 30-40 → good quality
# > 40 → very good, near human quality
#
# NOTE: Our dataset has very short sentences (MAX_SEQ_LEN=6)
# so BLEU scores will be lower than typical benchmarks.
# That is expected and fine for a tutorial.
# ============================================================
from nltk.translate.bleu_score import corpus_bleu, SmoothingFunction
import nltk
# punkt is the NLTK tokeniser required internally by the BLEU scorer.
# quiet=True suppresses the download progress message.
nltk.download('punkt', quiet=True)
# ============================================================
# STEP 1 — Load the best model checkpoint
# ============================================================
# Always evaluate on the BEST saved weights — not the final
# epoch. The model at the last epoch may have degraded from
# its best point. best_model_no_attn.pt was saved in cell_16
# whenever val loss improved.
# map_location=DEVICE ensures weights load correctly whether
# they were saved on GPU and are now loaded on CPU or vice versa.
model_no_attn.load_state_dict(
[page 70]
torch.load('best_model_no_attn.pt', map_location=DEVICE)
)
print(" Best model checkpoint loaded ✓ ")
# ============================================================
# STEP 2 — Compute test loss
# ============================================================
# Reuse evaluate() from cell_16 — identical function,
# just passing test_loader instead of val_loader.
# No weight updates, dropout OFF, no teacher forcing.
# This gives a single number summarising prediction confidence
# across the entire unseen test set.
print("\n Running evaluation on test set...")
test_loss = evaluate(model_no_attn, test_loader, criterion)
# Perplexity = e^loss — easier to interpret than raw loss.
# Lower is better. Perplexity of 1.0 = perfect prediction.
test_ppl = math.exp(test_loss)
print("=" * 55)
print(" Test Set Results — No Attention Model")
print("=" * 55)
print(f"\n Test Loss : {test_loss:.4f}")
print(f" Test PPL : {test_ppl:.4f}")
# ============================================================
# STEP 3 — Define translation generator for BLEU computation
# ============================================================
def generate_translations(model, loader, french_vocab):
"""
Generates translations for every sentence in a DataLoader
and returns them in the format required by corpus_bleu.
Runs in pure inference mode:
- model.eval() : dropout OFF
- torch.no_grad() : no gradient tracking
- teacher_forcing = 0.0 : model uses its own predictions
Args:
model : trained Seq2Seq model
loader : DataLoader supplying test batches
french_vocab : Vocabulary — used to convert indices to words
Returns:
references : list of list of list of str
Nested format required by corpus_bleu:
outer list = one entry per sentence
middle list = possible references
(we have 1 per sentence)
inner list = reference tokens
[page 71]
Example: [[['je', 'suis', 'ici']], ...]
hypotheses : list of list of str
One token list per sentence.
Example: [['je', 'suis', 'la'], ...]
"""
# Switch to eval mode — dropout OFF, deterministic output.
model.eval()
references = [] # correct French translations (ground truth)
hypotheses = [] # model's predicted French translations
progress_bar = tqdm(
loader,
desc=' Generating translations',
leave=False
)
# Disable gradient tracking — we are only generating predictions,
# never calling backward(). Saves memory and speeds up inference.
with torch.no_grad():
for src, trg_input, trg_target in progress_bar:
# Move all tensors to the same device as the model.
src = src.to(DEVICE)
trg_input = trg_input.to(DEVICE)
trg_target = trg_target.to(DEVICE)
# Forward pass with no teacher forcing — pure inference.
# The model uses its own predictions at every step,
# exactly as it would when translating a new sentence.
# output shape : (batch, FR_SEQ_LEN, vocab_size)
output = model(src, trg_input, teacher_forcing_ratio=0.0)
# Pick the highest scoring token index at every time step
# across the vocabulary dimension for the whole batch.
# shape : (batch, FR_SEQ_LEN) — one index per step
predicted_indices = output.argmax(dim=2)
# Process each sentence individually within the batch.
for i in range(src.shape[0]):
# ── HYPOTHESIS (model's prediction) ──────────
# Convert predicted indices to French words.
# Stop at <eos> — the model's signal it is done.
# Skip <pad> and <sos> — not real content words.
pred_indices = predicted_indices[i].cpu().tolist()
pred_tokens = []
for idx in pred_indices:
[page 72]
if idx == EOS_IDX:
break # model says stop — honour it
if idx not in (PAD_IDX, SOS_IDX):
pred_tokens.append(
french_vocab.idx2word.get(idx, UNK_TOKEN)
)
hypotheses.append(pred_tokens)
# ── REFERENCE (ground truth) ──────────────────
# Convert target indices to French words.
# Skip ALL special tokens — <pad>, <sos>, <eos> —
# we only want real French content words for scoring.
ref_indices = trg_target[i].cpu().tolist()
ref_tokens = []
for idx in ref_indices:
if idx not in (PAD_IDX, SOS_IDX, EOS_IDX):
ref_tokens.append(
french_vocab.idx2word.get(idx, UNK_TOKEN)
)
# corpus_bleu expects a list of references per sentence
# because some datasets provide multiple valid translations.
# We only have one reference per sentence so we wrap
# ref_tokens in an extra list to satisfy the format.
references.append([ref_tokens])
return references, hypotheses
# ============================================================
# STEP 4 — Generate translations and compute BLEU
# ============================================================
print("\n Computing BLEU score on test set...")
# Generate reference and hypothesis token lists for every
# sentence in the test set.
references, hypotheses = generate_translations(
model_no_attn,
test_loader,
french_vocab
)
# SmoothingFunction.method1 adds a small count to zero n-gram
# matches. Without smoothing, a single missing n-gram match
# makes the entire BLEU score zero — too harsh for short sentences
# like ours where 3-gram and 4-gram matches are naturally rare.
smoother = SmoothingFunction().method1
# corpus_bleu computes the BLEU score across all sentences at once.
# It measures 1-gram, 2-gram, 3-gram, and 4-gram overlap between
# hypotheses and references, then combines them geometrically.
[page 73]
# Score ranges from 0.0 (no overlap) to 1.0 (perfect match).
bleu_score_no_attn = corpus_bleu(
references,
hypotheses,
smoothing_function=smoother
)
print(f"\n BLEU Score : {bleu_score_no_attn:.4f} "
f"({bleu_score_no_attn * 100:.2f} / 100)")
# ============================================================
# STEP 5 — BLEU quality interpretation
# ============================================================
# Map the raw BLEU score to a plain-English quality description.
# Note: scores will be lower than published benchmarks because
# our sentences are very short (MAX_SEQ_LEN=6). Short sentences
# have fewer n-gram opportunities so any mismatch hurts more.
bleu_pct = bleu_score_no_attn * 100
print(f"\n--- BLEU Score Interpretation ---")
if bleu_pct < 10:
print(f" Score: {bleu_pct:.2f} → Almost unusable translation quality.")
elif bleu_pct < 20:
print(f" Score: {bleu_pct:.2f} → Getting some meaning across.")
elif bleu_pct < 30:
print(f" Score: {bleu_pct:.2f} → Understandable with effort.")
elif bleu_pct < 40:
print(f" Score: {bleu_pct:.2f} → Good quality translation.")
else:
print(f" Score: {bleu_pct:.2f} → Very good, near human quality.")
# ============================================================
# STEP 6 — Full results summary
# ============================================================
# Print train, val, and test loss together in one table so
# we can compare them directly and check for overfitting.
# A healthy model has all three losses reasonably close together.
final_train_loss = train_losses_no_attn[-1] # last epoch train loss
final_val_loss = val_losses_no_attn[-1] # last epoch val loss
print(f"\n{'=' * 55}")
print(f" Full Results Summary — No Attention Model")
print(f"{'=' * 55}")
print(f" {'Split':<12} {'Loss':>8} {'PPL':>8}")
print(f" {'-'*12} {'-'*8} {'-'*8}")
print(f" {'Train':<12} {final_train_loss:>8.4f} "
f"{math.exp(final_train_loss):>8.4f}")
print(f" {'Validation':<12} {final_val_loss:>8.4f} "
f"{math.exp(final_val_loss):>8.4f}")
print(f" {'Test':<12} {test_loss:>8.4f} "
[page 74]
f"{test_ppl:>8.4f}")
print(f"{'=' * 55}")
print(f" BLEU Score : {bleu_score_no_attn:.4f} "
f"({bleu_pct:.2f} / 100)")
print(f"{'=' * 55}")
# ============================================================
# STEP 7 — Show sample translations
# ============================================================
# Print 5 side-by-side translations — source English,
# correct French reference, and model's predicted French.
# This gives a human-readable gut check beyond the numbers —
# we can see exactly where the model succeeds and where it fails.
print(f"\n--- Sample Translations (No Attention) ---\n")
# Grab one batch from the test loader for visual inspection.
sample_src, _, sample_tgt = next(iter(test_loader))
sample_src = sample_src.to(DEVICE)
sample_tgt = sample_tgt.to(DEVICE)
# Run a forward pass with no teacher forcing — pure inference.
with torch.no_grad():
sample_output = model_no_attn(
sample_src,
sample_tgt,
teacher_forcing_ratio=0.0
)
# Pick highest scoring token at each step for every sentence.
# shape : (batch, trg_seq_len)
predicted = sample_output.argmax(dim=2)
for i in range(5):
# Decode source English — skip all special tokens.
src_tokens = [
english_vocab.idx2word.get(idx.item(), UNK_TOKEN)
for idx in sample_src[i]
if idx.item() not in (PAD_IDX, SOS_IDX, EOS_IDX)
]
# Decode reference French — skip all special tokens.
ref_tokens = [
french_vocab.idx2word.get(idx.item(), UNK_TOKEN)
for idx in sample_tgt[i]
if idx.item() not in (PAD_IDX, SOS_IDX, EOS_IDX)
]
# Decode predicted French — stop at <eos>, skip <pad> and <sos>.
# We honour the model's <eos> signal to avoid printing
# garbage tokens generated after the sentence ends.
[page 75]
pred_tokens = []
for idx in predicted[i]:
if idx.item() == EOS_IDX:
break
if idx.item() not in (PAD_IDX, SOS_IDX):
pred_tokens.append(
french_vocab.idx2word.get(idx.item(), UNK_TOKEN)
)
print(f" [Sample {i+1}]")
print(f" Source (EN) : {' '.join(src_tokens)}")
print(f" Reference (FR) : {' '.join(ref_tokens)}")
print(f" Predicted (FR) : {' '.join(pred_tokens)}")
print()
# ============================================================
# STEP 8 — Define inference function
# ============================================================
# translate() handles the full pipeline for a single raw
# English string — preprocessing, encoding, and greedy decoding.
# Defined here so it is available both in this cell and in
# cell_24 where we compare no-attention vs attention models.
def translate(
sentence,
model,
english_vocab,
french_vocab,
max_len=FR_SEQ_LEN
):
"""
Translates a single raw English sentence to French using
greedy decoding — always picks the highest scoring token
at each step.
Unlike batch inference, this function handles one sentence
at a time and includes full preprocessing from raw text.
Args:
sentence : str — raw English input e.g. "I love Paris"
model : Seq2SeqNoAttention — trained translation model
english_vocab : Vocabulary — maps English words to indices
french_vocab : Vocabulary — maps French indices to words
max_len : int — maximum number of tokens to generate
defaults to FR_SEQ_LEN to match training length
Returns:
translation : str — predicted French as a single string
token_list : list of str — individual predicted French tokens
"""
# Switch to eval mode — dropout OFF, deterministic output.
[page 76]
model.eval()
# ── PREPROCESS ───────────────────────────────────────────
# Apply the same preprocessing pipeline used during training.
# Any mismatch here would cause vocabulary lookup failures.
cleaned = clean_sentence(sentence) # lowercase, strip
punctuation
tokens = tokenize(cleaned) # split into word tokens
indices = numericalize(tokens, english_vocab) # words → integer indices
indices = add_special_tokens( # add <eos>, no <sos> for
src
indices, add_sos=False, add_eos=True
)
indices = pad_or_truncate(indices, EN_SEQ_LEN) # ensure fixed length
# Add batch dimension — model expects (batch, seq_len).
# unsqueeze(0) converts (seq_len,) → (1, seq_len).
src_tensor = torch.tensor(
indices, dtype=torch.long
).unsqueeze(0).to(DEVICE)
# ── ENCODE ───────────────────────────────────────────────
# Run the Encoder once over the full English sentence.
# We only need hidden and cell — encoder_outputs are not
# used here because this is the no-attention model.
with torch.no_grad():
encoder_outputs, hidden, cell = model.encoder(src_tensor)
# ── DECODE (greedy) ───────────────────────────────────────
# Start decoding with <sos> — the signal to begin translating.
# shape : (1, 1) — batch size 1, sequence length 1
input_token = torch.tensor(
[[SOS_IDX]], dtype=torch.long
).to(DEVICE)
predicted_tokens = []
with torch.no_grad():
for _ in range(max_len):
# One decoder step — generate prediction and update memory.
# prediction shape : (1, vocab_size)
# hidden, cell are updated and carried to the next step.
prediction, hidden, cell = model.decoder(
input_token, hidden, cell
)
# Greedy selection — pick the single highest scoring word.
# argmax(dim=1) returns the index of the max value
# across the vocabulary dimension.
pred_token = prediction.argmax(dim=1)
[page 77]
pred_idx = pred_token.item()
# Stop generating if the model predicts <eos>.
# This is the model's learnt signal that the
# translation is complete.
if pred_idx == EOS_IDX:
break
# Skip special tokens — only keep real French words.
# <unk> is also skipped to keep output clean.
if pred_idx not in (PAD_IDX, SOS_IDX, UNK_IDX):
word = french_vocab.idx2word.get(pred_idx, UNK_TOKEN)
predicted_tokens.append(word)
# Feed this step's prediction back as the next input.
# unsqueeze(0) reshapes from (1,) → (1, 1) to match
# the shape the decoder's embedding layer expects.
input_token = pred_token.unsqueeze(0)
# Join token list into a single readable string.
translation = ' '.join(predicted_tokens)
return translation, predicted_tokens
# ============================================================
# STEP 9 — Translate sample sentences
# ============================================================
# Run translate() on a set of short English sentences to
# visually inspect the model's output on completely new input —
# not from the dataset, just hand-written test cases.
# This is the most human-readable quality check in the notebook.
test_sentences = [
"Run!",
"I am cold.",
"She is happy.",
"We are here.",
"He loves her.",
"Stop it.",
"Help me.",
"I am tired.",
"Come here.",
"Thank you.",
]
print("=" * 60)
print(" Inference — Seq2Seq WITHOUT Attention")
print("=" * 60)
print(f"\n {'English':<25} {'Predicted French'}")
print(f" {'-'*25} {'-'*25}")
for sentence in test_sentences:
# translate() returns the full string and the token list.
[page 78]
Best model checkpoint loaded ✓
Running evaluation on test set...
=======================================================
Test Set Results — No Attention Model
=======================================================
Test Loss : 2.5677
Test PPL : 13.0353
Computing BLEU score on test set...
BLEU Score : 0.0398 (3.98 / 100)
--- BLEU Score Interpretation ---
Score: 3.98 → Almost unusable translation quality.
=======================================================
# We only need the string here for printing.
translation, _ = translate(
sentence,
model_no_attn,
english_vocab,
french_vocab
)
print(f" {sentence:<25} {translation}")
# ============================================================
# STEP 10 — Store results for cell_24 comparison
# ============================================================
# Pack all key metrics into a dictionary so cell_24 can load
# them alongside the attention model's results and produce
# a direct side-by-side comparison of both architectures.
# This is the only place these values need to be stored —
# cell_24 reads directly from this dict.
no_attn_results = {
'test_loss' : test_loss, # CrossEntropyLoss on test set
'test_ppl' : test_ppl, # e^test_loss — easier to read
'bleu' : bleu_score_no_attn, # BLEU score 0.0 to 1.0
'train_losses': train_losses_no_attn, # per-epoch train loss history
'val_losses' : val_losses_no_attn, # per-epoch val loss history
}
print(f"\n Results stored in no_attn_results dict ✓ ")
print(f"\n Ready for Cell 19 — Decoder WITH Attention")
{"model_id":"6ff563807d8c4fd090850e7641ad97b3","version_major":2,"version_minor":0}
{"model_id":"f62f8c5546f14b84822d8442f27efa3c","version_major":2,"version_minor":0}
[page 79]
Full Results Summary — No Attention Model
=======================================================
Split Loss PPL
------------ -------- --------
Train 0.9959 2.7073
Validation 2.5345 12.6106
Test 2.5677 13.0353
=======================================================
BLEU Score : 0.0398 (3.98 / 100)
=======================================================
--- Sample Translations (No Attention) ---
[Sample 1]
Source (EN) : its <unk>
Reference (FR) : cest bon marché
Predicted (FR) : <unk>
[Sample 2]
Source (EN) : it cant be
Reference (FR) : cest pas possible
Predicted (FR) : <unk> <unk>
[Sample 3]
Source (EN) : lets party
Reference (FR) : <unk> la <unk>
Predicted (FR) :
[Sample 4]
Source (EN) : are you up
Reference (FR) : estu levée
Predicted (FR) : debout
[Sample 5]
Source (EN) : im awake
Reference (FR) : je suis réveillé
Predicted (FR) : suis
============================================================
Inference — Seq2Seq WITHOUT Attention
============================================================
English Predicted French
------------------------- -------------------------
Run!
I am cold. froid
She is happy. est content
We are here. sommes ici
He loves her. a besoin
Stop it.
Help me.
I am tired. suis fatigué
Come here. là
Thank you.
[page 80]
Results stored in no_attn_results dict ✓
Ready for Cell 19 — Decoder WITH Attention
Case 2: LSTM with attention
# ============================================================
# cell_19 — DECODER WITH ATTENTION
# ============================================================
# WHAT WE ARE DOING HERE:
# We build the Decoder that uses the Attention mechanism
# defined in cell_13. This is the upgraded version of the
# no-attention Decoder from cell_14.
#
# THE KEY DIFFERENCE:
#
# WITHOUT Attention (cell_14):
# input_token → Embedding → LSTM → Linear → prediction
# The LSTM only sees the previous hidden state — it has
# NO direct access to the encoder outputs. All source
# information is compressed into one bottleneck vector.
#
# WITH Attention (this cell):
# input_token → Embedding ──────────────────────────────┐
# ↓
# encoder_outputs + decoder_hidden → Attention → context
# ↓
# concat(embedded, context) → LSTM → Linear →
prediction
#
# At EVERY decoder step:
# 1. Attention looks at ALL encoder hidden states
# 2. Produces a context vector focused on relevant words
# 3. Context is concatenated with the embedded input
# 4. LSTM receives BOTH the input word AND the context
#
# This means the Decoder is no longer limited to a single
# bottleneck vector — it can directly access any part of
# the source sentence at every step.
#
# WHY CONCATENATE context WITH embedded?
# The LSTM needs two signals at each step:
# - What word did I just predict? (embedded)
# - What should I focus on in source? (context)
# Concatenating them gives the LSTM both at once.
#
# INPUT SIZE CHANGE:
# DecoderNoAttention LSTM input : embed_dim
[page 81]
# DecoderWithAttention LSTM input: embed_dim + hidden_dim
# Because embedded (embed_dim) and context (hidden_dim)
# are concatenated before entering the LSTM.
# ============================================================
# ============================================================
# STEP 1 — Define the Decoder class (with Attention)
# ============================================================
class DecoderWithAttention(nn.Module):
"""
Decodes one French token at a time using the Attention
mechanism to dynamically focus on relevant encoder outputs
at every decoding step.
Unlike DecoderNoAttention which relies solely on the
Encoder's final hidden state, this Decoder has direct
access to ALL encoder hidden states via Attention.
At each step it asks: which English words matter most
right now? — and receives a focused context vector back.
At each step it receives:
- The previous French token index (or <sos> at step 0)
- Its own (hidden, cell) carried from the previous step
- ALL encoder outputs — used by Attention every step
It produces:
- Raw logit scores over the entire French vocabulary
- Updated (hidden, cell) for the next decoding step
- Attention weights for this step — saved for heatmap
visualisation in cell_25
Args:
vocab_size : int — number of words in French vocabulary
embed_dim : int — dimension of each word embedding vector
hidden_dim : int — LSTM hidden size, must match Encoder
attention : Attention — the Attention module from cell_13
dropout : float — dropout probability (applied to embeddings)
Forward inputs:
trg_token : (batch_size, 1)
Previous French token index.
<sos> at step 0, predicted token thereafter.
hidden : (1, batch_size, hidden_dim)
Short-term LSTM memory from previous step.
cell : (1, batch_size, hidden_dim)
Long-term LSTM memory from previous step.
encoder_outputs : (batch_size, src_seq_len, hidden_dim)
All encoder hidden states — passed to
Attention at every decoding step.
Forward outputs:
[page 82]
prediction : (batch_size, vocab_size)
Raw logit scores — highest index is predicted word.
hidden : (1, batch_size, hidden_dim) — updated short-term
memory
cell : (1, batch_size, hidden_dim) — updated long-term memory
attn_weights : (batch_size, src_seq_len, 1)
Per-source-position attention weights for this step.
Sum to 1 across src_seq_len. Saved for heatmap.
"""
def __init__(self, vocab_size, embed_dim, hidden_dim, attention,
dropout):
# Initialise nn.Module before any self.* assignments.
# Registers all sub-modules for parameter tracking,
# device movement, and state dict save/load.
super(DecoderWithAttention, self).__init__()
# Store the Attention module as a sub-module.
# Passed in from outside so it remains independently
# testable and reusable across different Decoder variants.
self.attention = attention
# ----------------------------------------------------------------
# EMBEDDING LAYER
# Converts French token indices into dense embedding vectors.
# Same role as in DecoderNoAttention — nothing changes here.
# padding_idx=PAD_IDX keeps <pad> embedding as zeros,
# never updated during backpropagation.
# ----------------------------------------------------------------
self.embedding = nn.Embedding(
vocab_size,
embed_dim,
padding_idx=PAD_IDX
)
# ----------------------------------------------------------------
# LSTM LAYER — key difference from DecoderNoAttention
# Input size is embed_dim + hidden_dim, NOT just embed_dim.
# This is because at each step we concatenate:
# embedded : (batch, 1, embed_dim) — the previous token
# context : (batch, 1, hidden_dim) — from Attention
# before feeding into the LSTM.
#
# The LSTM therefore receives both signals simultaneously:
# "What word did I just generate?" (embedded)
# "What should I focus on in source?"(context)
#
# hidden_dim must match the Encoder's hidden_dim exactly —
# the Encoder's final hidden and cell states are used to
# initialise this Decoder's memory at the first step.
# ----------------------------------------------------------------
[page 83]
self.lstm = nn.LSTM(
embed_dim + hidden_dim, # wider input than no-attention decoder
hidden_dim, # output hidden size matches Encoder
batch_first=True # input shape: (batch, seq_len,
features)
)
# ----------------------------------------------------------------
# OUTPUT PROJECTION LAYER
# Maps the LSTM hidden state to a score for every French word.
# The word with the highest score is the predicted next token.
# Softmax is NOT applied here — CrossEntropyLoss handles it.
# input : (batch, hidden_dim)
# output : (batch, vocab_size)
# ----------------------------------------------------------------
self.fc_out = nn.Linear(hidden_dim, vocab_size)
# ----------------------------------------------------------------
# DROPOUT LAYER
# Applied to embeddings before concatenation with context.
# Randomly zeroes values during training to prevent
# over-reliance on any single embedding dimension.
# Automatically disabled during model.eval().
# ----------------------------------------------------------------
self.dropout = nn.Dropout(dropout)
def forward(self, trg_token, hidden, cell, encoder_outputs):
"""
Performs one decoding step with Attention.
Called once per output token. At each call:
1. Embeds the previous token
2. Computes a fresh context vector via Attention
3. Concatenates embedding and context
4. Passes combined input through LSTM
5. Projects LSTM output to vocabulary scores
Args:
trg_token : (batch_size, 1)
Previous French token. <sos> at step 0.
hidden : (1, batch_size, hidden_dim)
Short-term memory from previous step.
At step 0, received from Encoder.
cell : (1, batch_size, hidden_dim)
Long-term memory from previous step.
At step 0, received from Encoder.
encoder_outputs : (batch_size, src_seq_len, hidden_dim)
All encoder hidden states.
Passed to Attention at every step.
Returns:
prediction : (batch_size, vocab_size)
[page 84]
Raw logits — not yet passed through softmax.
hidden : (1, batch_size, hidden_dim) — updated
cell : (1, batch_size, hidden_dim) — updated
attn_weights : (batch_size, src_seq_len, 1)
Attention weights — sum to 1 across src positions.
"""
# ── STEP 1 : EMBED THE INPUT TOKEN ──────────────────
# Convert the previous French token index to a dense vector.
# Dropout applied immediately to regularise during training.
# trg_token shape : (batch, 1)
# embedded shape : (batch, 1, embed_dim)
embedded = self.dropout(self.embedding(trg_token))
# ── STEP 2 : COMPUTE ATTENTION CONTEXT ──────────────
# Attention requires decoder_hidden as (batch, hidden_dim).
# Current hidden shape is (1, batch, hidden_dim) —
# squeeze(0) removes the leading 1 to match Attention's
# expected input shape.
decoder_hidden_squeezed = hidden.squeeze(0)
# decoder_hidden_squeezed : (batch, hidden_dim)
# Call the Attention module — compare current decoder state
# against all encoder hidden states to produce a context
# vector weighted by relevance at this decoding step.
#
# context : (batch, 1, hidden_dim)
# weighted blend of encoder outputs
# attn_weights : (batch, src_seq_len, 1)
# how much to focus on each source position
context, attn_weights = self.attention(
decoder_hidden_squeezed, # (batch, hidden_dim)
encoder_outputs # (batch, src_seq_len, hidden_dim)
)
# ── STEP 3 : CONCATENATE EMBEDDED AND CONTEXT ───────
# Combine the previous word embedding with the attention
# context vector along the feature dimension (dim=2).
# This gives the LSTM both signals at once:
# embedded : what word did I just predict?
# context : what part of the source should I focus on?
#
# embedded : (batch, 1, embed_dim)
# context : (batch, 1, hidden_dim)
# lstm_input : (batch, 1, embed_dim + hidden_dim)
lstm_input = torch.cat([embedded, context], dim=2)
# ── STEP 4 : PASS THROUGH LSTM ──────────────────────
# Feed the concatenated input and carry memory forward.
# The LSTM updates hidden and cell based on the combined
# signal from the previous token AND the attention context.
[page 85]
#
# lstm_out : (batch, 1, hidden_dim) — output at this step
# hidden : (1, batch, hidden_dim) — updated short-term memory
# cell : (1, batch, hidden_dim) — updated long-term memory
lstm_out, (hidden, cell) = self.lstm(lstm_input, (hidden, cell))
# ── STEP 5 : PROJECT TO VOCABULARY ──────────────────
# Remove the redundant sequence dimension (always 1 here)
# then map the hidden state to vocabulary scores.
# The word with the highest logit is the predicted token.
# Softmax is NOT applied — CrossEntropyLoss does it internally.
#
# lstm_out : (batch, 1, hidden_dim) → (batch, hidden_dim)
# prediction : (batch, vocab_size)
lstm_out = lstm_out.squeeze(1)
prediction = self.fc_out(lstm_out)
# Return prediction, updated memory states, and attention
# weights. attn_weights are carried back to the Seq2Seq
# model and saved for heatmap visualisation in cell_25.
return prediction, hidden, cell, attn_weights
# ============================================================
# STEP 2 — Instantiate the Decoder with Attention
# ============================================================
# hidden_dim and embed_dim must match the Encoder exactly.
# attention is the module instantiated in cell_13 —
# shared between this Decoder and the Seq2Seq wrapper.
decoder_with_attn = DecoderWithAttention(
vocab_size = len(french_vocab), # French vocabulary from cell_5
embed_dim = EMBED_DIM, # embedding size — same as Encoder
hidden_dim = HIDDEN_DIM, # must match Encoder hidden_dim
attention = attention, # Attention module from cell_13
dropout = DEC_DROPOUT # dropout probability
).to(DEVICE)
# ============================================================
# STEP 3 — Print model summary
# ============================================================
# PyTorch prints every registered sub-module and its configuration.
# Use this to visually verify layer sizes are correct —
# especially that the LSTM input is embed_dim + hidden_dim,
# not just embed_dim as in the no-attention decoder.
print("--- Decoder (With Attention) Architecture ---")
print(decoder_with_attn)
total_params = sum(
p.numel() for p in decoder_with_attn.parameters()
if p.requires_grad
[page 86]
)
print(f"\n Total trainable parameters : {total_params:,}")
# ============================================================
# STEP 4 — Compare parameter counts
# ============================================================
# The attention decoder always has more parameters than the
# no-attention decoder for two reasons:
# 1. Larger LSTM input — (embed_dim + hidden_dim) vs embed_dim
# adds an extra block of hidden_dim × hidden_dim weights
# 2. Attention sub-module adds W1, W2, V linear layers
# each of size hidden_dim × hidden_dim (or hidden_dim × 1 for V)
# This comparison quantifies exactly how many extra parameters
# the attention mechanism introduces.
no_attn_params = sum(
p.numel() for p in decoder_no_attn.parameters()
if p.requires_grad
)
attn_params = sum(
p.numel() for p in decoder_with_attn.parameters()
if p.requires_grad
)
print(f"\n--- Parameter Comparison ---")
print(f" Decoder without Attention : {no_attn_params:,}")
print(f" Decoder with Attention : {attn_params:,}")
print(f" Difference : {attn_params - no_attn_params:,} "
f"extra parameters")
# ============================================================
# STEP 5 — Sanity check — one forward step
# ============================================================
# Simulate the very first decoder step with attention to verify
# all four output shapes are correct before building Seq2Seq.
#
# Key difference from DecoderNoAttention sanity check:
# - forward() now takes FOUR inputs (not three)
# - enc_outputs must be passed in — Attention needs them
# - forward() returns FOUR outputs (not three) — adds attn_weights
decoder_with_attn.eval()
with torch.no_grad():
# First decoder input is always <sos> for every sentence.
# torch.full creates a tensor filled entirely with SOS_IDX.
# shape : (batch_size, 1) — one token per sentence
sos_input = torch.full(
(BATCH_SIZE, 1),
SOS_IDX,
dtype=torch.long # embedding layers require integer indices
).to(DEVICE)
[page 87]
# Run one forward step.
# enc_hidden and enc_cell initialise the Decoder's memory.
# enc_outputs are passed to Attention at every step.
# All three come from the Encoder forward pass in cell_12.
sample_pred, sample_hidden, sample_cell, sample_attn = \
decoder_with_attn(
sos_input, # (batch, 1) — first input token
enc_hidden, # (1, batch, hidden_dim) — Encoder final hidden
enc_cell, # (1, batch, hidden_dim) — Encoder final cell
enc_outputs # (batch, EN_SEQ_LEN, hidden_dim) — all states
)
# Verify all four output shapes match expectations.
print(f"\n--- Forward Pass Shape Check ---")
print(f" Input token shape : {sos_input.shape} "
f"(batch_size x 1)")
print(f" prediction shape : {sample_pred.shape} "
f"(batch_size x french_vocab_size)")
print(f" hidden shape : {sample_hidden.shape} "
f"(1 x batch_size x hidden_dim)")
print(f" cell shape : {sample_cell.shape} "
f"(1 x batch_size x hidden_dim)")
print(f" attn_weights shape : {sample_attn.shape} "
f"(batch_size x EN_SEQ_LEN x 1)")
# Assert each shape individually so any mismatch identifies
# exactly which output is wrong — easier to debug than a
# single combined assertion.
assert sample_pred.shape == (BATCH_SIZE, len(french_vocab)), \
"prediction shape mismatch!"
assert sample_hidden.shape == (1, BATCH_SIZE, HIDDEN_DIM), \
"hidden shape mismatch!"
assert sample_cell.shape == (1, BATCH_SIZE, HIDDEN_DIM), \
"cell shape mismatch!"
assert sample_attn.shape == (BATCH_SIZE, EN_SEQ_LEN, 1), \
"attn_weights shape mismatch!"
print(f"\n All shapes correct ✓ ")
# ============================================================
# STEP 6 — Verify attention weights sum to 1
# ============================================================
# softmax was applied inside Attention — so weights across all
# source positions must sum to exactly 1.0 for every sample.
# squeeze(-1) removes the trailing 1 dimension before summing.
# sum(dim=1) sums across the src_seq_len dimension.
# atol=1e-6 allows for tiny floating point rounding errors.
weight_sums = sample_attn.squeeze(-1).sum(dim=1)
print(f"\n--- Attention Weight Sum Check ---")
print(f" First 5 weight sums (should all be 1.0) :")
[page 88]
--- Decoder (With Attention) Architecture ---
DecoderWithAttention(
(attention): Attention(
(W1): Linear(in_features=512, out_features=512, bias=False)
(W2): Linear(in_features=512, out_features=512, bias=False)
(V): Linear(in_features=512, out_features=1, bias=False)
)
(embedding): Embedding(1207, 256, padding_idx=0)
(lstm): LSTM(768, 512, batch_first=True)
(fc_out): Linear(in_features=512, out_features=1207, bias=True)
(dropout): Dropout(p=0.3, inplace=False)
)
Total trainable parameters : 4,078,519
--- Parameter Comparison ---
Decoder without Attention : 2,505,143
Decoder with Attention : 4,078,519
Difference : 1,573,376 extra parameters
--- Forward Pass Shape Check ---
Input token shape : torch.Size([32, 1]) (batch_size x 1)
prediction shape : torch.Size([32, 1207]) (batch_size x
french_vocab_size)
hidden shape : torch.Size([1, 32, 512]) (1 x batch_size x
hidden_dim)
cell shape : torch.Size([1, 32, 512]) (1 x batch_size x
hidden_dim)
attn_weights shape : torch.Size([32, 6, 1]) (batch_size x
EN_SEQ_LEN x 1)
All shapes correct ✓
--- Attention Weight Sum Check ---
First 5 weight sums (should all be 1.0) :
sample 0 : 1.000000
sample 1 : 1.000000
sample 2 : 1.000000
sample 3 : 1.000000
sample 4 : 1.000000
Attention weights sum to 1 ✓
for i in range(5):
print(f" sample {i} : {weight_sums[i].item():.6f}")
assert torch.allclose(
weight_sums,
torch.ones_like(weight_sums),
atol=1e-6
), "Attention weights do not sum to 1!"
print(f"\n Attention weights sum to 1 ✓ ")
print(f"\n Ready for Cell 20 — Seq2Seq WITH Attention")
[page 89]
Ready for Cell 20 — Seq2Seq WITH Attention
# ============================================================
# cell_20 — SEQ2SEQ WITH ATTENTION
# ============================================================
# WHAT WE ARE DOING HERE:
# We build the complete Seq2Seq model that uses Attention.
# This is the upgraded version of cell_15 — the architecture
# is identical except for two critical differences:
#
# DIFFERENCE 1 — encoder_outputs passed to Decoder every step:
# Without Attention: encoder_outputs computed then discarded
# With Attention : encoder_outputs passed to Decoder at
# every decoding step so Attention can
# compute a fresh context vector each time
#
# DIFFERENCE 2 — attention weights collected across all steps:
# all_attn tensor stores weights at every decoding step.
# Shape : (batch, trg_seq_len, src_seq_len)
# Used in cell_25 to draw the attention heatmap — showing
# which English words the model focused on at each step.
#
# DECODING LOOP COMPARISON:
#
# No Attention (cell_15):
# prediction, hidden, cell = decoder(
# token, hidden, cell
# )
#
# With Attention (cell_20):
# prediction, hidden, cell, attn_weights = decoder(
# token, hidden, cell, encoder_outputs ← extra arg
# )
#
# That one extra argument — encoder_outputs — is what gives
# the Decoder direct access to every English word at every
# decoding step, rather than relying on the compressed
# bottleneck of the final hidden state alone.
# ============================================================
# ============================================================
# STEP 1 — Define Seq2Seq class (with Attention)
# ============================================================
class Seq2SeqWithAttention(nn.Module):
"""
Complete sequence-to-sequence translation model with Attention.
Connects the Encoder and DecoderWithAttention into one system.
The Encoder reads the full English sentence once. The Decoder
then generates French tokens one at a time — at each step
[page 90]
calling the Attention mechanism to produce a fresh context
vector focused on the most relevant source words right now.
Unlike Seq2SeqNoAttention, encoder_outputs are NOT discarded
after encoding. They are passed into the Decoder at every
single decoding step so Attention can reference them.
Args:
encoder : Encoder — encodes the source sentence
decoder : DecoderWithAttention — decodes with attention at each step
Forward inputs:
src : (batch_size, src_seq_len)
Padded English token indices
trg : (batch_size, trg_seq_len)
Padded French token indices
(ground truth used during training)
teacher_forcing_ratio : float (0.0 to 1.0)
Probability of feeding correct token
rather than model's own prediction.
Default 0.5 — coin flip each step.
Forward outputs:
outputs : (batch_size, trg_seq_len, vocab_size)
Raw logit scores at every decoder time step.
Used by CrossEntropyLoss during training.
all_attn : (batch_size, trg_seq_len, src_seq_len)
Attention weights at every decoder step.
Each row sums to 1.0 across src positions.
Saved for heatmap visualisation in cell_25.
"""
def __init__(self, encoder, decoder):
"""
Registers Encoder and Decoder as sub-modules so PyTorch
tracks all their parameters for training, saving, and
device movement automatically.
"""
# Initialise nn.Module's internal machinery first.
# Must be called before any self.* assignments.
super(Seq2SeqWithAttention, self).__init__()
# Store Encoder and Decoder as named sub-modules.
# Registering on self ensures their weights appear in
# model.parameters(), model.state_dict(), and model.to(DEVICE).
self.encoder = encoder
self.decoder = decoder
def forward(self, src, trg, teacher_forcing_ratio=0.5):
"""
[page 91]
Runs the full encode → decode pipeline with Attention.
The Encoder processes src once. The Decoder generates
the target sequence token by token — at each step using
Attention to focus on the most relevant encoder outputs.
Args:
src : (batch, src_seq_len)
Padded English token indices
trg : (batch, trg_seq_len)
Padded French token indices
teacher_forcing_ratio : float
Probability of using ground
truth token as next decoder input.
Returns:
outputs : (batch, trg_seq_len, vocab_size)
Logit scores at every decoding step.
Step 0 left as zeros — no prediction for <sos>.
all_attn : (batch, trg_seq_len, src_seq_len)
Attention weights collected at every step.
Step 0 left as zeros — no attention for <sos>.
"""
# ----------------------------------------------------------------
# EXTRACT DIMENSIONS
# Read directly from input tensors so forward() works for
# any batch size or sequence length without hardcoding.
# ----------------------------------------------------------------
batch_size = src.shape[0] # number of sentences in batch
trg_seq_len = trg.shape[1] # number of French tokens to generate
vocab_size = len(french_vocab) # number of French words to score
# ----------------------------------------------------------------
# INITIALISE OUTPUT CONTAINERS
# Pre-allocate tensors of zeros for predictions and attention
# weights. Both are filled step by step inside the decode loop.
# Step 0 remains zero for both — it corresponds to <sos> input
# and has no associated prediction or attention weights.
# ----------------------------------------------------------------
# Stores vocabulary logit scores at every decoding step.
# shape : (batch, trg_seq_len, vocab_size)
outputs = torch.zeros(
batch_size, trg_seq_len, vocab_size
).to(DEVICE)
# Stores attention weights at every decoding step.
# src.shape[1] = EN_SEQ_LEN — one weight per source position.
# shape : (batch, trg_seq_len, src_seq_len)
all_attn = torch.zeros(
batch_size, trg_seq_len, src.shape[1]
[page 92]
).to(DEVICE)
# ----------------------------------------------------------------
# ENCODE
# Pass the full English sentence through the Encoder once.
# All three outputs are used here — unlike cell_15 where
# encoder_outputs were computed but never passed to the Decoder.
#
# encoder_outputs : (batch, src_seq_len, hidden_dim)
# hidden state at EVERY source time step
# → passed to Decoder at every decoding step
# → used by Attention to compute context
# hidden : (1, batch, hidden_dim)
# final short-term memory → initialises Decoder
# cell : (1, batch, hidden_dim)
# final long-term memory → initialises Decoder
# ----------------------------------------------------------------
encoder_outputs, hidden, cell = self.encoder(src)
# ----------------------------------------------------------------
# INITIALISE FIRST DECODER INPUT
# The first token fed into the Decoder is always <sos> —
# the signal that tells the Decoder to begin translating.
# trg[:, 0] extracts the first column — always SOS_IDX=1
# for every sentence in the batch.
# unsqueeze(1) reshapes from (batch,) → (batch, 1) to match
# the shape the Decoder's embedding layer expects.
# ----------------------------------------------------------------
input_token = trg[:, 0].unsqueeze(1) # (batch, 1) — always <sos>
# ----------------------------------------------------------------
# DECODE LOOP
# Generate one French token per iteration.
# Start at t=1 — t=0 is the <sos> input, not a prediction.
# End at trg_seq_len — one prediction per target position.
# ----------------------------------------------------------------
for t in range(1, trg_seq_len):
# One forward step through the Decoder WITH Attention.
# encoder_outputs are passed at every step — this is the
# key difference from the no-attention decode loop.
# Attention runs inside the Decoder and produces a fresh
# context vector at this specific decoding step.
#
# prediction : (batch, vocab_size)
# logit scores — highest index = predicted word
# hidden : (1, batch, hidden_dim) — updated short-term
memory
# cell : (1, batch, hidden_dim) — updated long-term
memory
# attn_weights : (batch, src_seq_len, 1)
# how much the model focused on each source word
[page 93]
prediction, hidden, cell, attn_weights = self.decoder(
input_token, # previous French token
hidden, # short-term memory from previous step
cell, # long-term memory from previous step
encoder_outputs # ALL encoder states — used by Attention
)
# Store this step's vocabulary scores in the output container.
# outputs[:, t, :] selects all sentences at time step t.
outputs[:, t, :] = prediction # (batch, vocab_size)
# Store this step's attention weights in all_attn.
# attn_weights shape : (batch, src_seq_len, 1)
# after squeeze(-1) : (batch, src_seq_len)
# Removes trailing 1 before storing — all_attn expects 2D
# per time step, not 3D.
all_attn[:, t, :] = attn_weights.squeeze(-1)
# ------------------------------------------------------------
# TEACHER FORCING DECISION
# Flip a random coin to decide the next decoder input.
# With teacher_forcing_ratio=0.5, 50% of steps use the
# correct French word, 50% use the model's own prediction.
# This is identical to the no-attention model — Attention
# does not change the teacher forcing logic.
# ------------------------------------------------------------
use_teacher_forcing = (
torch.rand(1).item() < teacher_forcing_ratio
)
if use_teacher_forcing:
# TEACHER FORCING — feed the correct next French token.
# trg[:, t] is the ground truth word at position t
# for every sentence in the batch.
# unsqueeze(1) reshapes from (batch,) → (batch, 1).
input_token = trg[:, t].unsqueeze(1) # (batch, 1)
else:
# FREE RUNNING — feed the model's own best prediction.
# argmax(dim=1) picks the highest scoring word index
# across the vocabulary for each sentence in the batch.
# unsqueeze(1) reshapes from (batch,) → (batch, 1).
input_token = prediction.argmax(dim=1).unsqueeze(1) #
(batch, 1)
# Return both the vocabulary predictions and the full collection
# of attention weights across all decoding steps.
# outputs used by loss function during training.
# all_attn used by cell_25 to draw the attention heatmap.
return outputs, all_attn
# ============================================================
[page 94]
# STEP 2 — Instantiate the Seq2Seq model with Attention
# ============================================================
# A fresh Encoder is created here — separate from the one used
# in model_no_attn. Both models must train independently from
# random initialisation so their results can be compared fairly.
# Using the same Encoder would contaminate the comparison.
encoder_for_attn = Encoder(
vocab_size = len(english_vocab), # English vocabulary from cell_5
embed_dim = EMBED_DIM, # embedding size — same as no-attn
model
hidden_dim = HIDDEN_DIM, # hidden size — same as no-attn model
dropout = ENC_DROPOUT # encoder dropout probability
).to(DEVICE)
# Combine the fresh Encoder with the attention-enabled Decoder
# from cell_19 into one complete Seq2Seq system.
model_with_attn = Seq2SeqWithAttention(
encoder = encoder_for_attn, # fresh Encoder — independent from no-attn
decoder = decoder_with_attn # DecoderWithAttention from cell_19
).to(DEVICE)
# ============================================================
# STEP 3 — Print model summary
# ============================================================
# PyTorch prints every registered sub-module and its config.
# Key things to verify:
# - LSTM input size is embed_dim + hidden_dim (not just embed_dim)
# - Attention sub-module (W1, W2, V) appears inside the Decoder
# - All hidden_dim values match between Encoder and Decoder
print("--- Seq2Seq (With Attention) Architecture ---")
print(model_with_attn)
total_params = sum(
p.numel() for p in model_with_attn.parameters()
if p.requires_grad
)
print(f"\n Total trainable parameters : {total_params:,}")
# ============================================================
# STEP 4 — Compare total parameters between both models
# ============================================================
# Quantifies the cost of adding Attention in terms of parameters.
# The attention model is larger for two reasons:
# 1. Larger LSTM input — embed_dim + hidden_dim vs embed_dim alone
# adds an extra hidden_dim × hidden_dim block of weights
# 2. Attention sub-module — W1, W2, V add three linear layers
# each of size hidden_dim × hidden_dim (or hidden_dim × 1 for V)
# The size increase % helps assess whether the extra capacity
# is justified by the BLEU score improvement in cell_24.
[page 95]
no_attn_total = sum(
p.numel() for p in model_no_attn.parameters()
if p.requires_grad
)
attn_total = sum(
p.numel() for p in model_with_attn.parameters()
if p.requires_grad
)
print(f"\n--- Model Size Comparison ---")
print(f" Seq2Seq without Attention : {no_attn_total:,} parameters")
print(f" Seq2Seq with Attention : {attn_total:,} parameters")
print(f" Extra parameters : {attn_total - no_attn_total:,}")
print(f" Size increase : "
f"{100 * (attn_total - no_attn_total) / no_attn_total:.1f}%")
# ============================================================
# STEP 5 — Sanity check — full forward pass
# ============================================================
# Run one complete batch through the full attention model to
# verify both output shapes are correct before training.
# teacher_forcing_ratio=0.0 — model uses its own predictions,
# matching real inference conditions for the shape check.
model_with_attn.eval()
with torch.no_grad():
src_sample = enc_batch.to(DEVICE) # English batch from cell_11
trg_sample = dec_inp_batch.to(DEVICE) # French input batch from cell_11
# Two outputs now — predictions AND attention weights.
# No teacher forcing — model flies solo for this test.
sample_output, sample_attn = model_with_attn(
src_sample,
trg_sample,
teacher_forcing_ratio=0.0
)
# Print and verify both output shapes.
# output shape : (batch, FR_SEQ_LEN, vocab_size) — predictions
# all_attn shape : (batch, FR_SEQ_LEN, EN_SEQ_LEN) — attention weights
print(f"\n--- Forward Pass Shape Check ---")
print(f" src shape : {src_sample.shape} "
f"(batch_size x EN_SEQ_LEN)")
print(f" trg shape : {trg_sample.shape} "
f"(batch_size x FR_SEQ_LEN)")
print(f" output shape : {sample_output.shape} "
f"(batch_size x FR_SEQ_LEN x french_vocab_size)")
print(f" all_attn shape : {sample_attn.shape} "
f"(batch_size x FR_SEQ_LEN x EN_SEQ_LEN)")
# Assert each shape individually — a single combined assertion
[page 96]
--- Seq2Seq (With Attention) Architecture ---
Seq2SeqWithAttention(
(encoder): Encoder(
(embedding): Embedding(858, 256, padding_idx=0)
(lstm): LSTM(256, 512, batch_first=True)
(dropout): Dropout(p=0.3, inplace=False)
)
(decoder): DecoderWithAttention(
(attention): Attention(
(W1): Linear(in_features=512, out_features=512, bias=False)
(W2): Linear(in_features=512, out_features=512, bias=False)
(V): Linear(in_features=512, out_features=1, bias=False)
)
# would not tell us which tensor has the wrong shape if it fails.
assert sample_output.shape == (BATCH_SIZE, FR_SEQ_LEN, len(french_vocab)), \
"output shape mismatch!"
assert sample_attn.shape == (BATCH_SIZE, FR_SEQ_LEN, EN_SEQ_LEN), \
"attention weights shape mismatch!"
print(f"\n All shapes correct ✓ ")
# ============================================================
# STEP 6 — Peek at attention weights for one sentence
# ============================================================
# Print the attention weights at every decoding step for the
# first sentence in the batch. Weights are random at this stage
# so they will be roughly uniform — but they must still sum
# to 1.0 at every step, confirming softmax worked correctly
# throughout the full forward pass end to end.
print(f"\n--- Attention Weight Check (first sentence) ---")
print(f" Weights at each decoder step (should sum to 1.0)\n")
print(f" {'Step':<6} {'Weights':<45} {'Sum'}")
print(f" {'-'*6} {'-'*45} {'-'*5}")
for t in range(1, FR_SEQ_LEN):
# Extract attention weights for sentence 0 at step t.
# sample_attn[0, t, :] selects first sentence, step t,
# all source positions → shape (src_seq_len,)
weights = sample_attn[0, t, :].cpu().tolist()
# Format each weight to 3 decimal places for readability.
weight_str = ' '.join([f'{w:.3f}' for w in weights])
# Sum should be exactly 1.0 — softmax guarantee.
weight_sum = sum(weights)
print(f" {t:<6} {weight_str:<45} {weight_sum:.4f}")
print(f"\n All attention weights sum to 1.0 ✓ ")
print(f"\n Ready for Cell 21 — Training Loop (With Attention)")
[page 97]
(embedding): Embedding(1207, 256, padding_idx=0)
(lstm): LSTM(768, 512, batch_first=True)
(fc_out): Linear(in_features=512, out_features=1207, bias=True)
(dropout): Dropout(p=0.3, inplace=False)
)
)
Total trainable parameters : 5,875,127
--- Model Size Comparison ---
Seq2Seq without Attention : 4,301,751 parameters
Seq2Seq with Attention : 5,875,127 parameters
Extra parameters : 1,573,376
Size increase : 36.6%
--- Forward Pass Shape Check ---
src shape : torch.Size([32, 6]) (batch_size x EN_SEQ_LEN)
trg shape : torch.Size([32, 6]) (batch_size x FR_SEQ_LEN)
output shape : torch.Size([32, 6, 1207]) (batch_size x FR_SEQ_LEN x
french_vocab_size)
all_attn shape : torch.Size([32, 6, 6]) (batch_size x FR_SEQ_LEN x
EN_SEQ_LEN)
All shapes correct ✓
--- Attention Weight Check (first sentence) ---
Weights at each decoder step (should sum to 1.0)
Step Weights Sum
------ --------------------------------------------- -----
1 0.164 0.166 0.168 0.168 0.167 0.166 1.0000
2 0.164 0.166 0.168 0.168 0.167 0.166 1.0000
3 0.164 0.166 0.168 0.168 0.167 0.166 1.0000
4 0.164 0.166 0.168 0.168 0.167 0.166 1.0000
5 0.164 0.166 0.168 0.168 0.167 0.166 1.0000
All attention weights sum to 1.0 ✓
Ready for Cell 21 — Training Loop (With Attention)
# ============================================================
# cell_21 — TRAINING LOOP (WITH ATTENTION + VALIDATION)
# ============================================================
# WHAT WE ARE DOING HERE:
# We train the Seq2Seq model WITH Attention using identical
# settings to cell_16 — the no-attention training loop.
# The only meaningful code difference is that the attention
# model returns TWO outputs instead of one:
#
# output, _ = model(src, trg, teacher_forcing_ratio)
#
# We only need output for the loss — all_attn is discarded
# during training and only used in cell_25 for heatmaps.
[page 98]
#
# WHY KEEP EVERYTHING IDENTICAL TO cell_16?
# This is a controlled experiment — we change ONE variable
# (the model architecture) and hold everything else fixed.
# Changing any hyperparameter would make it impossible to
# know whether improvements came from Attention or from
# different training settings.
#
# WHAT IS IDENTICAL TO cell_16:
# - Loss function : CrossEntropyLoss, ignore_index=PAD_IDX
# - Optimizer : Adam, same learning rate
# - Gradient clip : CLIP=1.0
# - Teacher forcing: TEACHER_FORCING_RATIO=0.5
# - Epochs : N_EPOCHS=10
# - Scheduler : ReduceLROnPlateau, same patience and factor
# - Model saving : best model by VALIDATION loss
#
# WHAT IS DIFFERENT FROM cell_16:
# - Model : Seq2SeqWithAttention instead of Seq2SeqNoAttention
# - Model output : (output, all_attn) instead of just output
# - all_attn : collected but discarded during training
# ============================================================
# ============================================================
# STEP 1 — Define optimizer and scheduler for attention model
# ============================================================
# A fresh optimizer is required for the attention model.
# Optimizers are tightly coupled to a specific model's parameters —
# sharing the optimizer from cell_16 would produce incorrect
# gradient updates because it tracks gradient history per parameter.
# All settings match cell_16 exactly — fair comparison.
optimizer_with_attn = optim.Adam(
model_with_attn.parameters(), # attention model's parameters only
lr=LEARNING_RATE # same initial lr as cell_16
)
# Same ReduceLROnPlateau settings as cell_16.
# Monitors val loss and halves the learning rate if no improvement
# for patience=2 consecutive epochs.
# Must be called after validation each epoch — see Step 4.
scheduler_with_attn = optim.lr_scheduler.ReduceLROnPlateau(
optimizer_with_attn,
mode='min', # we want val loss to go DOWN
factor=0.5, # multiply lr by 0.5 when triggered
patience=2, # wait 2 epochs of no improvement before firing
)
# ============================================================
# STEP 2 — Define training epoch function (with attention)
# ============================================================
[page 99]
def train_one_epoch_attn(model, loader, optimizer, criterion, clip):
"""
Runs one full pass through the training data for the
attention model. Identical to train_one_epoch() in cell_16
except the model returns (output, all_attn) instead of
just output.
all_attn is discarded during training with _ — we only need
the vocabulary predictions to compute CrossEntropyLoss.
Attention weights are only used in cell_25 for heatmaps.
Args:
model : Seq2SeqWithAttention model to train
loader : DataLoader supplying (src, trg_input, trg_target)
optimizer : Adam optimizer for this model
criterion : CrossEntropyLoss — measures prediction error
clip : float — maximum gradient norm for clipping
Returns:
avg_loss : float — mean loss across all batches
"""
# Switch to training mode — dropout ON, gradients tracked.
model.train()
# Running total accumulated across all batches.
# Divided by batch count at the end to return epoch average.
epoch_loss = 0.0
# Live progress bar — updates with batch loss in real time.
# leave=False removes the bar when the epoch completes.
progress_bar = tqdm(loader, desc=' Training', leave=False)
for src, trg_input, trg_target in progress_bar:
# Move all tensors to the same device as the model.
# Model and data must always be on the same device.
src = src.to(DEVICE)
trg_input = trg_input.to(DEVICE)
trg_target = trg_target.to(DEVICE)
# ── FORWARD PASS ─────────────────────────────────────
# The attention model returns TWO values.
# output : (batch, trg_seq_len, vocab_size) — predictions
# _ : (batch, trg_seq_len, src_seq_len) — attention weights
#
# _ is Python's convention for "I know this exists but
# I am deliberately not using it." Attention weights are
# not needed for loss computation — only for heatmaps later.
output, _ = model(src, trg_input, TEACHER_FORCING_RATIO)
[page 100]
# ── PREPARE LOSS INPUTS ──────────────────────────────
# Skip position 0 on both sides — position 0 is <sos>,
# the starting signal fed INTO the decoder, never predicted.
# Reshape from 3D → 2D (output) and 2D → 1D (target) because
# CrossEntropyLoss expects flat inputs, not batched sequences.
#
# output_for_loss : (batch × seq_len, vocab_size)
# target_for_loss : (batch × seq_len,)
output_for_loss = output[:, 1:, :].reshape(-1, len(french_vocab))
target_for_loss = trg_target[:, 1:].reshape(-1)
# Compute loss — compares predicted scores against correct indices.
# <pad> positions are automatically ignored via ignore_index.
loss = criterion(output_for_loss, target_for_loss)
# ── BACKWARD PASS AND WEIGHT UPDATE ──────────────────
# Clear gradients from the previous batch.
# PyTorch accumulates gradients — must zero before backward
# or previous batch's gradients contaminate this update.
optimizer.zero_grad()
# Backpropagate — compute gradient of loss with respect to
# every trainable weight by walking backwards through the
# computational graph built during the forward pass.
loss.backward()
# Clip gradient norm to maximum of clip=1.0.
# Prevents exploding gradients — especially important for
# LSTMs which are prone to runaway gradient accumulation
# over long sequences.
torch.nn.utils.clip_grad_norm_(model.parameters(), clip)
# Apply weight update — nudges every parameter in the
# direction that reduces loss using Adam's adaptive step sizes.
optimizer.step()
# ── TRACKING ─────────────────────────────────────────
# Accumulate scalar loss into running total.
# .item() converts single-element tensor to Python float.
epoch_loss += loss.item()
# Update progress bar with this batch's loss in real time.
progress_bar.set_postfix({'batch_loss': f'{loss.item():.4f}'})
# Return average loss across all batches in this epoch.
return epoch_loss / len(loader)
# ============================================================
# STEP 3 — Define evaluation function (with attention)
# ============================================================
[page 101]
def evaluate_attn(model, loader, criterion):
"""
Evaluates the attention model on a DataLoader.
Identical to evaluate() from cell_16 except it unpacks
(output, all_attn) from Seq2SeqWithAttention.
Three critical differences from train_one_epoch_attn:
- model.eval() : dropout OFF — deterministic output
- torch.no_grad() : no gradient graph — saves memory
- teacher_forcing = 0.0 : model uses own predictions only
- No optimizer.step() : weights are never updated here
Teacher forcing is disabled to mirror real inference —
the model gets no help from ground truth, giving an honest
measure of how well it actually translates.
Args:
model : Seq2SeqWithAttention model to evaluate
loader : DataLoader supplying (src, trg_input, trg_target)
criterion : CrossEntropyLoss — measures prediction error
Returns:
avg_loss : float — mean loss across all batches
"""
# Switch to eval mode — dropout OFF, output is deterministic.
model.eval()
epoch_loss = 0.0
progress_bar = tqdm(loader, desc=' Validating', leave=False)
# Disable gradient tracking — we never call backward() here.
# Saves memory and speeds up computation by not building the
# computational graph that backward() would need.
with torch.no_grad():
for src, trg_input, trg_target in progress_bar:
src = src.to(DEVICE)
trg_input = trg_input.to(DEVICE)
trg_target = trg_target.to(DEVICE)
# Forward pass with no teacher forcing — pure inference.
# Model uses its own predictions at every step —
# exactly as it would when translating new sentences.
# output : (batch, trg_seq_len, vocab_size)
# _ : (batch, trg_seq_len, src_seq_len) — discarded
output, _ = model(src, trg_input, teacher_forcing_ratio=0.0)
# Same reshaping as training — skip <sos> at position 0,
# flatten to 2D for CrossEntropyLoss.
[page 102]
output_for_loss = output[:, 1:, :].reshape(-1, len(french_vocab))
target_for_loss = trg_target[:, 1:].reshape(-1)
# Compute loss — no backward pass, no weight updates.
loss = criterion(output_for_loss, target_for_loss)
# Accumulate loss into running total.
epoch_loss += loss.item()
progress_bar.set_postfix({'batch_loss': f'{loss.item():.4f}'})
# Return average loss across all batches.
return epoch_loss / len(loader)
# ============================================================
# STEP 4 — Run the training + validation loop
# ============================================================
print("=" * 65)
print(" Training Seq2Seq WITH Attention")
print("=" * 65)
# Lists to record metrics at every epoch for plotting later.
train_losses_with_attn = [] # training loss per epoch
val_losses_with_attn = [] # validation loss per epoch
lr_history_with_attn = [] # learning rate per epoch
# Initialise best val loss tracker.
# float('inf') ensures any real loss immediately beats it
# on the first epoch, triggering the first model save.
best_val_loss = float('inf')
best_epoch = 0
# Outer progress bar — one tick per epoch.
# leave=True keeps it visible after training completes.
epoch_bar = tqdm(range(1, N_EPOCHS + 1), desc='Epochs', leave=True)
for epoch in epoch_bar:
# ── TRAIN ─────────────────────────────────────────────
# One full pass through training data.
# Weights updated. Dropout ON. Teacher forcing active.
train_loss = train_one_epoch_attn(
model_with_attn, train_loader, optimizer_with_attn, criterion, CLIP
)
# ── VALIDATE ──────────────────────────────────────────
# Evaluate on held-out validation set.
# Weights NOT updated. Dropout OFF. No teacher forcing.
val_loss = evaluate_attn(model_with_attn, val_loader, criterion)
# Record both losses for plotting in Step 6.
[page 103]
train_losses_with_attn.append(train_loss)
val_losses_with_attn.append(val_loss)
# ── SCHEDULER STEP ────────────────────────────────────
# Must be called AFTER validation — the scheduler needs
# the current val loss to decide whether to reduce lr.
# If val loss has not improved for patience=2 epochs,
# lr is multiplied by factor=0.5.
scheduler_with_attn.step(val_loss)
# Read current learning rate from the optimizer directly.
# param_groups[0] is the first (and only) parameter group.
# Stored for the learning rate plot in Step 6.
current_lr = optimizer_with_attn.param_groups[0]['lr']
lr_history_with_attn.append(current_lr)
# Perplexity = e^loss — alternative readability metric.
# Lower is always better. 1.0 = perfect, 100 = very confused.
train_ppl = math.exp(train_loss)
val_ppl = math.exp(val_loss)
# ── SAVE BEST MODEL ───────────────────────────────────
# Save weights only when val loss improves — not train loss.
# Saving by val loss protects against saving an overfitted
# model that scores well on training data but translates poorly.
# state_dict() saves weights only — not the architecture.
if val_loss < best_val_loss:
best_val_loss = val_loss
best_epoch = epoch
torch.save(model_with_attn.state_dict(), 'best_model_with_attn.pt')
saved_marker = ' ← best model saved'
else:
saved_marker = ''
# Update outer progress bar with current epoch metrics.
epoch_bar.set_postfix({
'train_loss': f'{train_loss:.4f}',
'val_loss' : f'{val_loss:.4f}',
'lr' : f'{current_lr:.6f}',
'best_ep' : best_epoch
})
# Print one summary line per epoch.
# tqdm.write() used instead of print() to avoid interfering
# with the progress bar display in the notebook.
tqdm.write(
f" Epoch {epoch:02d}/{N_EPOCHS} | "
f"Train Loss: {train_loss:.4f} PPL: {train_ppl:.4f} | "
f"Val Loss: {val_loss:.4f} PPL: {val_ppl:.4f} | "
f"LR: {current_lr:.6f}"
f"{saved_marker}"
)
[page 104]
print("=" * 65)
print(f"\n Best Val Loss : {best_val_loss:.4f} (Epoch {best_epoch})")
print(f" Best model saved → best_model_with_attn.pt")
# ============================================================
# STEP 5 — Load best model weights for inference
# ============================================================
# Reload the weights from the best checkpoint saved during training.
# The model's current weights at the end of the final epoch may
# not be its best — val loss could have peaked earlier and degraded.
# map_location=DEVICE handles CPU/GPU device mismatches gracefully.
model_with_attn.load_state_dict(
torch.load('best_model_with_attn.pt', map_location=DEVICE)
)
print(f"\n Best model weights reloaded ✓ ")
# ============================================================
# STEP 6 — Loss curve plot
# ============================================================
# Two side-by-side plots — loss curves and learning rate history.
# Compare visually with cell_16's plots to see if the attention
# model converged faster, to a lower loss, or with less overfitting.
fig, axes = plt.subplots(1, 2, figsize=(14, 4))
# Plot 1 — Train vs Validation Loss.
# Healthy: both curves fall together.
# Overfitting: train keeps falling, val plateaus or rises.
axes[0].plot(range(1, N_EPOCHS + 1), train_losses_with_attn,
label='Train Loss', color='steelblue', marker='o')
axes[0].plot(range(1, N_EPOCHS + 1), val_losses_with_attn,
label='Val Loss', color='coral', marker='o')
axes[0].set_xlabel('Epoch')
axes[0].set_ylabel('Loss')
axes[0].set_title('Seq2Seq WITH Attention — Train vs Val Loss')
axes[0].legend()
axes[0].grid(True, alpha=0.3)
# Plot 2 — Learning Rate Schedule.
# Flat sections = scheduler did not fire.
# Step-downs = scheduler fired because val loss stalled.
axes[1].plot(range(1, N_EPOCHS + 1), lr_history_with_attn,
color='green', marker='o', linewidth=2)
axes[1].set_xlabel('Epoch')
axes[1].set_ylabel('Learning Rate')
axes[1].set_title('Learning Rate Schedule — With Attention')
axes[1].grid(True, alpha=0.3)
plt.tight_layout()
plt.show()
[page 105]
# ============================================================
# STEP 7 — Overfitting check
# ============================================================
# Compare final train and val loss to diagnose model health.
# A small positive gap is normal — val is always slightly harder.
# A large gap signals overfitting — memorising rather than learning.
# A negative gap can occur because dropout is OFF during validation,
# making the full model slightly stronger than the noisy training mode.
final_train = train_losses_with_attn[-1]
final_val = val_losses_with_attn[-1]
gap = final_val - final_train # positive = val worse than train
print(f"\n--- Overfitting Check ---")
print(f" Final Train Loss : {final_train:.4f}")
print(f" Final Val Loss : {final_val:.4f}")
print(f" Gap (val-train) : {gap:.4f}")
if gap > 0.5:
# Large gap — model memorised training data.
# More dropout adds regularisation noise during training,
# forcing more robust representations.
print(f" ⚠ Large gap — model may be overfitting.")
print(f" Consider increasing ENC_DROPOUT or DEC_DROPOUT.")
elif gap < 0:
# Val loss below train loss — sounds wrong but is expected
# when dropout is ON during training and OFF during eval.
print(f" ⚠ Val loss lower than train — unusual.")
print(f" Likely due to dropout being OFF during eval.")
else:
# Healthy gap — model is generalising well.
print(f" ✓ Gap looks healthy.")
# ============================================================
# STEP 8 — Quick side-by-side comparison with cell_16
# ============================================================
# Preview comparison between both models after the same number
# of epochs. This is NOT the final verdict — that comes in
# cell_24 where we compare BLEU scores on the held-out test set.
# Val loss alone doesn't capture translation quality — a model
# with slightly higher val loss can still produce better translations.
print(f"\n--- Early Comparison : Val Loss after {N_EPOCHS} Epochs ---")
print(f" {'Model':<30} {'Best Val Loss':>14} {'Best Epoch':>10}")
print(f" {'-'*30} {'-'*14} {'-'*10}")
# min() finds the best val loss across all epochs.
# .index(min()) finds which epoch it occurred at (0-indexed, so +1).
print(f" {'Seq2Seq No Attention':<30} "
f"{min(val_losses_no_attn):>14.4f} "
f"{val_losses_no_attn.index(min(val_losses_no_attn)) + 1:>10}")
[page 106]
=================================================================
Training Seq2Seq WITH Attention
=================================================================
Epoch 01/10 | Train Loss: 4.2531 PPL: 70.3237 | Val Loss: 3.9287 PPL:
50.8415 | LR: 0.001000 ← best model saved
Epoch 02/10 | Train Loss: 3.4454 PPL: 31.3569 | Val Loss: 3.5378 PPL:
34.3914 | LR: 0.001000 ← best model saved
Epoch 03/10 | Train Loss: 2.8867 PPL: 17.9349 | Val Loss: 3.1897 PPL:
24.2804 | LR: 0.001000 ← best model saved
Epoch 04/10 | Train Loss: 2.4289 PPL: 11.3469 | Val Loss: 2.9638 PPL:
19.3716 | LR: 0.001000 ← best model saved
Epoch 05/10 | Train Loss: 2.0405 PPL: 7.6944 | Val Loss: 2.7757 PPL:
16.0497 | LR: 0.001000 ← best model saved
Epoch 06/10 | Train Loss: 1.6961 PPL: 5.4528 | Val Loss: 2.6256 PPL:
13.8133 | LR: 0.001000 ← best model saved
print(f" {'Seq2Seq With Attention':<30} "
f"{min(val_losses_with_attn):>14.4f} "
f"{val_losses_with_attn.index(min(val_losses_with_attn)) + 1:>10}")
print(f"\n Ready for Cell 22 — Evaluate & BLEU Score (With Attention)")
{"model_id":"6dfca39ae6384808b412a8735add2e2b","version_major":2,"version_minor":0}
{"model_id":"2d1f828742f343db8a97b3cf7e764fa8","version_major":2,"version_minor":0}
{"model_id":"e706939feee34131852e05cac59672bb","version_major":2,"version_minor":0}
{"model_id":"e8c49c7087d749959bbfd8b4c3cdb622","version_major":2,"version_minor":0}
{"model_id":"c6919065740f4d359e406b731fe12bd2","version_major":2,"version_minor":0}
{"model_id":"cce9dd7230774c119b020f454533ef19","version_major":2,"version_minor":0}
{"model_id":"297001b7832e45628b992d77a5f5c6be","version_major":2,"version_minor":0}
{"model_id":"a8a23cd07b8f4233b1ee2f225aaf9cb8","version_major":2,"version_minor":0}
{"model_id":"f19924e50579463f8d8c544acee0dc9c","version_major":2,"version_minor":0}
{"model_id":"14d81fb787ae492cba05f61a053fcb1e","version_major":2,"version_minor":0}
{"model_id":"fe3fe0e517ce42ed83a47ac385fe636f","version_major":2,"version_minor":0}
{"model_id":"25ade586564f4f86901e539380c49ba3","version_major":2,"version_minor":0}
{"model_id":"4052c1bea20448f5b267ff00eb0e8f9e","version_major":2,"version_minor":0}
{"model_id":"aef61063958740dea455707bd4f797b4","version_major":2,"version_minor":0}
[page 107]
Epoch 07/10 | Train Loss: 1.4262 PPL: 4.1627 | Val Loss: 2.5450 PPL:
12.7431 | LR: 0.001000 ← best model saved
Epoch 08/10 | Train Loss: 1.2192 PPL: 3.3844 | Val Loss: 2.4527 PPL:
11.6198 | LR: 0.001000 ← best model saved
Epoch 09/10 | Train Loss: 1.0515 PPL: 2.8619 | Val Loss: 2.4810 PPL:
11.9536 | LR: 0.001000
Epoch 10/10 | Train Loss: 0.9345 PPL: 2.5458 | Val Loss: 2.4992 PPL:
12.1727 | LR: 0.001000
=================================================================
Best Val Loss : 2.4527 (Epoch 8)
Best model saved → best_model_with_attn.pt
Best model weights reloaded ✓
--- Overfitting Check ---
Final Train Loss : 0.9345
Final Val Loss : 2.4992
Gap (val-train) : 1.5647
⚠ Large gap — model may be overfitting.
Consider increasing ENC_DROPOUT or DEC_DROPOUT.
--- Early Comparison : Val Loss after 10 Epochs ---
Model Best Val Loss Best Epoch
------------------------------ -------------- ----------
Seq2Seq No Attention 2.5345 10
Seq2Seq With Attention 2.4527 8
{"model_id":"2b1eeb1fb9cc49728988eebb5bce418c","version_major":2,"version_minor":0}
{"model_id":"ead605bb372f4e8696b02802d308958c","version_major":2,"version_minor":0}
{"model_id":"50b2973ae24242c0abd9ae0b4a9bbfd2","version_major":2,"version_minor":0}
{"model_id":"aaa0b42e2fa84bf98dd0ce4c7514c1aa","version_major":2,"version_minor":0}
{"model_id":"3601151dc2894cfcbe1319cfaa4073ca","version_major":2,"version_minor":0}
{"model_id":"67a326b1815e479b827ad9d3fe62b5ed","version_major":2,"version_minor":0}
{"model_id":"b3ca2804b03d48d6b7fc17f5a3f411ed","version_major":2,"version_minor":0}
[page 108]
Ready for Cell 22 — Evaluate & BLEU Score (With Attention)
# ============================================================
# cell_22 —
# ============================================================
# ============================================================
# cell_23 — EVALUATE & BLEU SCORE + INFERENCE (WITH ATTENTION)
# ============================================================
# WHAT WE ARE DOING HERE:
# Mirror of cell_18 but for the Attention model.
# We do THREE things:
# 1. Evaluate best attention model on the TEST set
# 2. Compute BLEU score
# 3. Translate new sentences (inference)
#
# THIS IS THE FIRST AND ONLY TIME WE TOUCH THE TEST SET
# FOR THE ATTENTION MODEL — same principle as cell_18.
# The test set was never used during training or model
# selection — this is what makes the score honest.
#
# KEY DIFFERENCES FROM cell_18:
# 1. generate_translations_attn() unpacks (output, all_attn)
# and discards all_attn — only needed for single-sentence
# heatmap visualisation in translate_attn(), not BLEU
# 2. translate_attn() returns FOUR values instead of two:
# translation, token_list, attn_matrix, src_tokens
# attn_matrix is the attention weights per step —
# used as the heatmap data in cell_25
# 3. Results stored in attn_results dict for cell_24
# ============================================================
# ============================================================
# STEP 1 — Load best attention model checkpoint
# ============================================================
# Always evaluate on the best saved weights — not the final
# epoch. best_model_with_attn.pt was saved in cell_21 whenever
# val loss improved. map_location=DEVICE handles CPU/GPU
# device mismatches gracefully.
model_with_attn.load_state_dict(
torch.load('best_model_with_attn.pt', map_location=DEVICE)
)
print(" Best attention model checkpoint loaded ✓ ")
# ============================================================
# STEP 2 — Compute test loss
# ============================================================
# Reuse evaluate_attn() from cell_21 — same function,
# just passing test_loader instead of val_loader.
# Dropout OFF, no teacher forcing, no weight updates.
[page 109]
# Gives a single number summarising prediction confidence
# across the entire unseen test set.
print("\n Running evaluation on test set...")
test_loss_attn = evaluate_attn(model_with_attn, test_loader, criterion)
# Perplexity = e^loss — easier to interpret than raw loss.
# Lower is better. 1.0 = perfect prediction. 100 = very confused.
test_ppl_attn = math.exp(test_loss_attn)
print(f"\n Test Loss : {test_loss_attn:.4f}")
print(f" Test PPL : {test_ppl_attn:.4f}")
# ============================================================
# STEP 3 — Define translation generator for BLEU
# ============================================================
def generate_translations_attn(model, loader, french_vocab):
"""
Generates translations for all sentences in a DataLoader
and returns them in the format required by corpus_bleu.
Identical to generate_translations() from cell_18 except
it unpacks (output, all_attn) from Seq2SeqWithAttention.
all_attn is discarded here with _ — batch-level attention
weights are not needed for BLEU computation. Attention
weights are only collected per sentence in translate_attn()
for single-sentence heatmap visualisation in cell_25.
Runs in pure inference mode:
- model.eval() : dropout OFF
- torch.no_grad() : no gradient tracking
- teacher_forcing = 0.0 : model uses own predictions
Args:
model : Seq2SeqWithAttention — trained model
loader : DataLoader supplying test batches
french_vocab : Vocabulary — maps indices to French words
Returns:
references : list of list of list of str
Nested format required by corpus_bleu:
outer list = one entry per sentence
middle list = possible references
(we have 1 per sentence)
inner list = reference tokens
Example: [[['je', 'suis', 'ici']], ...]
hypotheses : list of list of str
One token list per sentence.
Example: [['je', 'suis', 'la'], ...]
"""
[page 110]
# Switch to eval mode — dropout OFF, deterministic output.
model.eval()
references = [] # correct French translations (ground truth)
hypotheses = [] # model's predicted French translations
progress_bar = tqdm(
loader,
desc=' Generating translations',
leave=False
)
# Disable gradient tracking — purely generating predictions,
# never calling backward(). Saves memory and speeds up inference.
with torch.no_grad():
for src, trg_input, trg_target in progress_bar:
# Move all tensors to the same device as the model.
src = src.to(DEVICE)
trg_input = trg_input.to(DEVICE)
trg_target = trg_target.to(DEVICE)
# Forward pass — no teacher forcing, pure inference.
# output : (batch, FR_SEQ_LEN, vocab_size) — predictions
# _ : (batch, FR_SEQ_LEN, EN_SEQ_LEN) — attention weights
# discarded here — not needed for BLEU computation
output, _ = model(
src, trg_input, teacher_forcing_ratio=0.0
)
# Pick highest scoring token at every step across vocab.
# shape : (batch, FR_SEQ_LEN) — one index per step
predicted_indices = output.argmax(dim=2)
# Process each sentence individually within the batch.
for i in range(src.shape[0]):
# ── HYPOTHESIS (model's prediction) ──────────
# Convert predicted indices to French words.
# Stop at <eos> — the model's signal it is done.
# Skip <pad> and <sos> — not real content words.
pred_indices = predicted_indices[i].cpu().tolist()
pred_tokens = []
for idx in pred_indices:
if idx == EOS_IDX:
break # model says stop — honour it
if idx not in (PAD_IDX, SOS_IDX):
pred_tokens.append(
french_vocab.idx2word.get(idx, UNK_TOKEN)
[page 111]
)
hypotheses.append(pred_tokens)
# ── REFERENCE (ground truth) ──────────────────
# Convert target indices to French words.
# Skip ALL special tokens — <pad>, <sos>, <eos> —
# only real French content words go into the reference.
ref_indices = trg_target[i].cpu().tolist()
ref_tokens = []
for idx in ref_indices:
if idx not in (PAD_IDX, SOS_IDX, EOS_IDX):
ref_tokens.append(
french_vocab.idx2word.get(idx, UNK_TOKEN)
)
# corpus_bleu expects a list of references per sentence
# because some datasets provide multiple valid translations.
# We have one reference per sentence so wrap in extra list.
references.append([ref_tokens])
return references, hypotheses
# ============================================================
# STEP 4 — Compute BLEU score
# ============================================================
# Generate reference and hypothesis token lists for every
# sentence in the test set, then compute corpus-level BLEU.
print("\n Computing BLEU score on test set...")
references_attn, hypotheses_attn = generate_translations_attn(
model_with_attn,
test_loader,
french_vocab
)
# SmoothingFunction.method1 adds a small count to zero n-gram
# matches. Without smoothing, a single missing n-gram match
# makes the entire BLEU score zero — too harsh for short sentences
# like ours where 3-gram and 4-gram matches are naturally rare.
smoother = SmoothingFunction().method1
# corpus_bleu measures 1-gram through 4-gram overlap between
# hypotheses and references, combined geometrically.
# Score ranges from 0.0 (no overlap) to 1.0 (perfect match).
bleu_score_with_attn = corpus_bleu(
references_attn,
hypotheses_attn,
smoothing_function=smoother
)
[page 112]
print(f"\n BLEU Score : {bleu_score_with_attn:.4f} "
f"({bleu_score_with_attn * 100:.2f} / 100)")
# ============================================================
# STEP 5 — BLEU quality interpretation
# ============================================================
# Map raw BLEU score to plain-English quality description.
# Scores will be lower than published benchmarks because our
# sentences are very short (MAX_SEQ_LEN=6) — fewer n-gram
# opportunities means any mismatch hurts proportionally more.
bleu_pct = bleu_score_with_attn * 100
print(f"\n--- BLEU Score Interpretation ---")
if bleu_pct < 10:
print(f" Score: {bleu_pct:.2f} → Almost unusable translation quality.")
elif bleu_pct < 20:
print(f" Score: {bleu_pct:.2f} → Getting some meaning across.")
elif bleu_pct < 30:
print(f" Score: {bleu_pct:.2f} → Understandable with effort.")
elif bleu_pct < 40:
print(f" Score: {bleu_pct:.2f} → Good quality translation.")
else:
print(f" Score: {bleu_pct:.2f} → Very good, near human quality.")
# ============================================================
# STEP 6 — Full results summary
# ============================================================
# Print train, val, and test loss together in one table.
# A healthy model has all three losses reasonably close.
# A large gap between train and test signals overfitting.
final_train_loss = train_losses_with_attn[-1] # last epoch train loss
final_val_loss = val_losses_with_attn[-1] # last epoch val loss
print(f"\n{'=' * 55}")
print(f" Full Results Summary — With Attention Model")
print(f"{'=' * 55}")
print(f" {'Split':<12} {'Loss':>8} {'PPL':>8}")
print(f" {'-'*12} {'-'*8} {'-'*8}")
print(f" {'Train':<12} {final_train_loss:>8.4f} "
f"{math.exp(final_train_loss):>8.4f}")
print(f" {'Validation':<12} {final_val_loss:>8.4f} "
f"{math.exp(final_val_loss):>8.4f}")
print(f" {'Test':<12} {test_loss_attn:>8.4f} "
f"{test_ppl_attn:>8.4f}")
print(f"{'=' * 55}")
print(f" BLEU Score : {bleu_score_with_attn:.4f} "
f"({bleu_pct:.2f} / 100)")
print(f"{'=' * 55}")
# ============================================================
[page 113]
# STEP 7 — Define inference function (with attention)
# ============================================================
# translate_attn() handles the full pipeline for a single raw
# English string. Unlike translate() from cell_18, it returns
# FOUR values — the attention weight matrix is the extra one,
# needed as heatmap data in cell_25.
def translate_attn(
sentence,
model,
english_vocab,
french_vocab,
max_len=FR_SEQ_LEN
):
"""
Translates a single raw English sentence to French using
the attention model with greedy decoding.
At each decoding step, attention weights are collected and
stacked into a matrix — one row per generated token, one
column per source token. This matrix is the heatmap data
used in cell_25 to visualise where the model focused.
Args:
sentence : str — raw English input e.g. "I love Paris"
model : Seq2SeqWithAttention — trained model
english_vocab : Vocabulary — maps English words to indices
french_vocab : Vocabulary — maps French indices to words
max_len : int — maximum tokens to generate
defaults to FR_SEQ_LEN to match training
Returns:
translation : str — predicted French string
token_list : list of str — individual predicted tokens
attn_matrix : np.ndarray — shape (trg_len, src_len)
attention weights at each decoding step
row = decoder step, col = source position
→ used as heatmap data in cell_25
src_tokens : list of str — source content words
→ used as x-axis labels in cell_25 heatmap
"""
# Switch to eval mode — dropout OFF, deterministic output.
model.eval()
# ── PREPROCESS ───────────────────────────────────────────
# Apply the same preprocessing pipeline used during training.
# Any mismatch here would cause vocabulary lookup failures.
cleaned = clean_sentence(sentence) # lowercase, strip
punctuation
tokens = tokenize(cleaned) # split into word tokens
indices = numericalize(tokens, english_vocab) # words → integer indices
[page 114]
indices = add_special_tokens( # add <eos>, no <sos> for
src
indices, add_sos=False, add_eos=True
)
indices = pad_or_truncate(indices, EN_SEQ_LEN) # pad or truncate to
fixed length
# Save source tokens for heatmap x-axis labels in cell_25.
# Filter out special tokens — only real content words shown.
src_tokens = [
t for t in tokens
if t not in (PAD_TOKEN, SOS_TOKEN, EOS_TOKEN)
]
# Add batch dimension — model expects (batch, seq_len).
# unsqueeze(0) converts (seq_len,) → (1, seq_len).
src_tensor = torch.tensor(
indices, dtype=torch.long
).unsqueeze(0).to(DEVICE)
# ── ENCODE ───────────────────────────────────────────────
# Run the Encoder once to get memory states and all
# encoder hidden states. encoder_outputs are passed to
# Attention at every decoding step — unlike translate()
# in cell_18 where they were computed but never used.
with torch.no_grad():
encoder_outputs, hidden, cell = model.encoder(src_tensor)
# ── DECODE (greedy, one token at a time) ─────────────────
# Start with <sos> — the signal to begin translating.
# shape : (1, 1) — batch size 1, sequence length 1
input_token = torch.tensor(
[[SOS_IDX]], dtype=torch.long
).to(DEVICE)
predicted_tokens = [] # accumulates predicted French words
attn_weights_list = [] # accumulates attention weights per step
with torch.no_grad():
for _ in range(max_len):
# One decoder step with attention.
# Unlike translate() in cell_18, this Decoder returns
# four values — the attention weights are the extra one.
# attn_w shape : (1, src_seq_len, 1)
prediction, hidden, cell, attn_w = model.decoder(
input_token,
hidden,
cell,
encoder_outputs # passed at every step for Attention
)
[page 115]
# Collect attention weights for this decoding step.
# squeeze() removes both batch (1) and trailing (1) dimensions
# leaving shape (src_seq_len,) — one weight per source word.
# .cpu().numpy() converts to numpy for heatmap plotting later.
attn_weights_list.append(
attn_w.squeeze().cpu().numpy()
)
# Greedy decoding — pick the single highest scoring word.
# argmax(dim=1) returns index of max across vocabulary.
pred_token = prediction.argmax(dim=1)
pred_idx = pred_token.item()
# Stop generating when the model predicts <eos>.
# This is the learnt signal that the translation is done.
if pred_idx == EOS_IDX:
break
# Skip special tokens — only keep real French words.
# <unk> is also skipped to keep output readable.
if pred_idx not in (PAD_IDX, SOS_IDX, UNK_IDX):
word = french_vocab.idx2word.get(pred_idx, UNK_TOKEN)
predicted_tokens.append(word)
# Feed this step's prediction back as the next input.
# unsqueeze(0) reshapes from (1,) → (1, 1) to match
# the shape the Decoder's embedding layer expects.
input_token = pred_token.unsqueeze(0)
# Join predicted tokens into one readable French string.
translation = ' '.join(predicted_tokens)
# Stack per-step attention weights into a 2D matrix.
# np.stack along axis=0 turns a list of (src_seq_len,) arrays
# into a single (trg_len, src_len) matrix where:
# row t = attention weights when generating token t
# col s = how much attention was on source position s
# This matrix is the raw data for the heatmap in cell_25.
import numpy as np
if attn_weights_list:
attn_matrix = np.stack(attn_weights_list, axis=0)
else:
# Fallback for edge case where no tokens were generated.
attn_matrix = np.zeros((1, EN_SEQ_LEN))
return translation, predicted_tokens, attn_matrix, src_tokens
# ============================================================
# STEP 8 — Translate sample sentences
# ============================================================
# Run translate_attn() on hand-written English sentences to
[page 116]
# visually inspect translation quality on completely new input
# not from the dataset. We unpack all four return values but
# only use the translation string for printing here.
# attn_matrix and src_tokens are used in cell_25 heatmaps.
test_sentences = [
"Run!",
"I am cold.",
"She is happy.",
"We are here.",
"He loves her.",
"Stop it.",
"Help me.",
"I am tired.",
"Come here.",
"Thank you.",
]
print("=" * 60)
print(" Inference — Seq2Seq WITH Attention")
print("=" * 60)
print(f"\n {'English':<25} {'Predicted French'}")
print(f" {'-'*25} {'-'*25}")
for sentence in test_sentences:
# Unpack all four return values — discard three with _.
# translation is the only value needed for printing here.
# attn_matrix and src_tokens are used later in cell_25.
translation, _, _, _ = translate_attn(
sentence,
model_with_attn,
english_vocab,
french_vocab
)
print(f" {sentence:<25} {translation}")
# ============================================================
# STEP 9 — Store results for cell_24 comparison
# ============================================================
# Pack all key metrics into a dictionary so cell_24 can load
# them alongside no_attn_results from cell_18 and produce a
# direct side-by-side comparison of both model architectures.
# This is the only place these values need to be stored —
# cell_24 reads directly from this dict.
attn_results = {
'test_loss' : test_loss_attn, # CrossEntropyLoss on test set
'test_ppl' : test_ppl_attn, # e^test_loss — easier to read
'bleu' : bleu_score_with_attn, # BLEU score 0.0 to 1.0
'train_losses': train_losses_with_attn, # per-epoch train loss history
'val_losses' : val_losses_with_attn, # per-epoch val loss history
[page 117]
Best attention model checkpoint loaded ✓
Running evaluation on test set...
Test Loss : 2.5715
Test PPL : 13.0854
Computing BLEU score on test set...
BLEU Score : 0.0225 (2.25 / 100)
--- BLEU Score Interpretation ---
Score: 2.25 → Almost unusable translation quality.
=======================================================
Full Results Summary — With Attention Model
=======================================================
Split Loss PPL
------------ -------- --------
Train 0.9345 2.5458
Validation 2.4992 12.1727
Test 2.5715 13.0854
=======================================================
BLEU Score : 0.0225 (2.25 / 100)
=======================================================
============================================================
Inference — Seq2Seq WITH Attention
============================================================
English Predicted French
------------------------- -------------------------
Run!
I am cold. froid
She is happy. est content
We are here. sommes ici
He loves her. a
Stop it.
Help me.
I am tired. suis
Come here. ici
Thank you.
Results stored in attn_results dict ✓
}
print(f"\n Results stored in attn_results dict ✓ ")
print(f"\n Ready for Cell 24 — Compare No Attention vs Attention")
{"model_id":"993bb8652e3f476ab003438872b93882","version_major":2,"version_minor":0}
{"model_id":"9c300cddf06743f2b3f22fdfe12795cb","version_major":2,"version_minor":0}
[page 118]
Ready for Cell 24 — Compare No Attention vs Attention
Comparing Attention with No Attention
# ============================================================
# cell_24 — COMPARE NO ATTENTION vs ATTENTION
# ============================================================
# WHAT WE ARE DOING HERE:
# We bring together results from both models and compare
# them side by side across THREE dimensions:
#
# 1. LOSS CURVES — how did training progress?
# 2. METRICS TABLE — test loss, PPL, BLEU side by side
# 3. TRANSLATIONS — same sentences, both models
#
# THIS IS THE PAYOFF CELL.
# Everything we built — two models, two training loops,
# two evaluation runs — comes together here so the student
# can see concretely what Attention buys us.
# ============================================================
# NOTE: translate() is defined in Cell 18.
# It is already in scope and used directly below.
# ============================================================
# STEP 1 — Metrics comparison table
# ============================================================
print("=" * 65)
print(" Final Comparison — No Attention vs With Attention")
print("=" * 65)
print(f"\n {'Metric':<20} {'No Attention':>15} {'With Attention':>15}")
print(f" {'-'*20} {'-'*15} {'-'*15}")
print(f" {'Test Loss':<20} "
f"{no_attn_results['test_loss']:>15.4f} "
f"{attn_results['test_loss']:>15.4f}")
print(f" {'Test PPL':<20} "
f"{no_attn_results['test_ppl']:>15.4f} "
f"{attn_results['test_ppl']:>15.4f}")
print(f" {'BLEU Score':<20} "
f"{no_attn_results['bleu']:>15.4f} "
f"{attn_results['bleu']:>15.4f}")
print(f" {'BLEU (%)':<20} "
[page 119]
f"{no_attn_results['bleu']*100:>15.2f} "
f"{attn_results['bleu']*100:>15.2f}")
print(f"\n{'=' * 65}")
bleu_diff = (attn_results['bleu'] - no_attn_results['bleu']) * 100
loss_diff = no_attn_results['test_loss'] - attn_results['test_loss']
print(f"\n BLEU improvement : {bleu_diff:+.2f} points")
print(f" Loss improvement : {loss_diff:+.4f}")
if bleu_diff > 0:
print(f"\n ✓ Attention model outperforms on BLEU score.")
elif bleu_diff < 0:
print(f"\n ⚠ No Attention model scores higher on BLEU.")
print(f" This can happen with very short sequences —")
print(f" try increasing N_EPOCHS or NUM_SAMPLES.")
else:
print(f"\n Both models scored equally on BLEU.")
# ============================================================
# STEP 2 — Loss curves — both models on one plot
# ============================================================
fig, axes = plt.subplots(1, 2, figsize=(14, 5))
axes[0].plot(
range(1, N_EPOCHS + 1),
no_attn_results['train_losses'],
label='No Attention',
color='steelblue',
marker='o',
linewidth=2
)
axes[0].plot(
range(1, N_EPOCHS + 1),
attn_results['train_losses'],
label='With Attention',
color='coral',
marker='o',
linewidth=2
)
axes[0].set_title('Training Loss Comparison')
axes[0].set_xlabel('Epoch')
axes[0].set_ylabel('Loss')
axes[0].legend()
axes[0].grid(True, alpha=0.3)
axes[1].plot(
range(1, N_EPOCHS + 1),
no_attn_results['val_losses'],
label='No Attention',
[page 120]
color='steelblue',
marker='o',
linewidth=2
)
axes[1].plot(
range(1, N_EPOCHS + 1),
attn_results['val_losses'],
label='With Attention',
color='coral',
marker='o',
linewidth=2
)
axes[1].set_title('Validation Loss Comparison')
axes[1].set_xlabel('Epoch')
axes[1].set_ylabel('Loss')
axes[1].legend()
axes[1].grid(True, alpha=0.3)
plt.suptitle(
'No Attention vs With Attention — Loss Curves',
fontsize=14,
fontweight='bold'
)
plt.tight_layout()
plt.show()
# ============================================================
# STEP 3 — PPL curves — both models on one plot
# ============================================================
fig, axes = plt.subplots(1, 2, figsize=(14, 5))
axes[0].plot(
range(1, N_EPOCHS + 1),
[math.exp(l) for l in no_attn_results['train_losses']],
label='No Attention',
color='steelblue',
marker='o',
linewidth=2
)
axes[0].plot(
range(1, N_EPOCHS + 1),
[math.exp(l) for l in attn_results['train_losses']],
label='With Attention',
color='coral',
marker='o',
linewidth=2
)
axes[0].set_title('Training Perplexity Comparison')
axes[0].set_xlabel('Epoch')
axes[0].set_ylabel('Perplexity')
[page 121]
axes[0].legend()
axes[0].grid(True, alpha=0.3)
axes[1].plot(
range(1, N_EPOCHS + 1),
[math.exp(l) for l in no_attn_results['val_losses']],
label='No Attention',
color='steelblue',
marker='o',
linewidth=2
)
axes[1].plot(
range(1, N_EPOCHS + 1),
[math.exp(l) for l in attn_results['val_losses']],
label='With Attention',
color='coral',
marker='o',
linewidth=2
)
axes[1].set_title('Validation Perplexity Comparison')
axes[1].set_xlabel('Epoch')
axes[1].set_ylabel('Perplexity')
axes[1].legend()
axes[1].grid(True, alpha=0.3)
plt.suptitle(
'No Attention vs With Attention — Perplexity Curves',
fontsize=14,
fontweight='bold'
)
plt.tight_layout()
plt.show()
# ============================================================
# STEP 4 — Side by side translation comparison
# ============================================================
test_sentences = [
"Run!",
"I am cold.",
"She is happy.",
"We are here.",
"He loves her.",
"Stop it.",
"Help me.",
"I am tired.",
"Come here.",
"Thank you.",
]
print(f"\n--- Translation Comparison ---\n")
[page 122]
print(f" {'English':<20} {'No Attention':<20} {'With Attention'}")
print(f" {'-'*20} {'-'*20} {'-'*20}")
for sentence in test_sentences:
no_attn_translation, _ = translate(
sentence,
model_no_attn,
english_vocab,
french_vocab
)
attn_translation, _, _, _ = translate_attn(
sentence,
model_with_attn,
english_vocab,
french_vocab
)
print(f" {sentence:<20} "
f"{no_attn_translation:<20} "
f"{attn_translation}")
# ============================================================
# STEP 5 — Key takeaways
# ============================================================
print(f"\n{'=' * 65}")
print(f" Key Takeaways")
print(f"{'=' * 65}")
print(f"""
1. BOTTLENECK PROBLEM (No Attention)
The entire English sentence is compressed into ONE
fixed-size vector (hidden, cell) of shape
(1, batch, {HIDDEN_DIM}). All information must fit
here regardless of sentence length.
2. ATTENTION SOLUTION
At every decoder step, Attention computes a weighted
sum over ALL encoder hidden states — giving the decoder
direct access to any part of the source sentence.
No information bottleneck.
3. WHY THE DIFFERENCE MAY BE SMALL HERE
Our sentences are very short (MAX_SEQ_LEN={MAX_SEQ_LEN}).
For short sentences the bottleneck is less severe —
all information fits in the hidden vector reasonably well.
On longer sentences (20-50 words) the gap between
no-attention and attention models is much larger.
4. WHAT COMES NEXT
cell_25 will show ATTENTION HEATMAPS — a visualisation
[page 123]
=================================================================
Final Comparison — No Attention vs With Attention
=================================================================
Metric No Attention With Attention
-------------------- --------------- ---------------
Test Loss 2.5677 2.5715
Test PPL 13.0353 13.0854
BLEU Score 0.0398 0.0225
BLEU (%) 3.98 2.25
=================================================================
BLEU improvement : -1.73 points
Loss improvement : -0.0038
⚠ No Attention model scores higher on BLEU.
This can happen with very short sequences —
try increasing N_EPOCHS or NUM_SAMPLES.
--- Translation Comparison ---
of exactly which source words the decoder focused on
at each translation step. This is the most intuitive
way to understand what attention is actually doing.
""")
print(f" Ready for Cell 25 — Attention Heatmap Visualisation")
[page 124]
English No Attention With Attention
-------------------- -------------------- --------------------
Run!
I am cold. froid froid
She is happy. est content est content
We are here. sommes ici sommes ici
He loves her. a besoin a
Stop it.
Help me.
I am tired. suis fatigué suis
Come here. là ici
Thank you.
=================================================================
Key Takeaways
=================================================================
1. BOTTLENECK PROBLEM (No Attention)
The entire English sentence is compressed into ONE
fixed-size vector (hidden, cell) of shape
(1, batch, 512). All information must fit
here regardless of sentence length.
2. ATTENTION SOLUTION
At every decoder step, Attention computes a weighted
sum over ALL encoder hidden states — giving the decoder
direct access to any part of the source sentence.
No information bottleneck.
3. WHY THE DIFFERENCE MAY BE SMALL HERE
Our sentences are very short (MAX_SEQ_LEN=5).
For short sentences the bottleneck is less severe —
all information fits in the hidden vector reasonably well.
On longer sentences (20-50 words) the gap between
no-attention and attention models is much larger.
4. WHAT COMES NEXT
cell_25 will show ATTENTION HEATMAPS — a visualisation
of exactly which source words the decoder focused on
at each translation step. This is the most intuitive
way to understand what attention is actually doing.
Ready for Cell 25 — Attention Heatmap Visualisation
ATTENTION HEATMAP
VISUALISATION
# ============================================================
# cell_25 — ATTENTION HEATMAP VISUALISATION
[page 125]
# ============================================================
# WHAT WE ARE DOING HERE:
# We visualise the attention weights produced by the
# Attention model during translation. This is the payoff
# for collecting attn_matrix in translate_attn() — we can
# now SEE what the model was looking at when it generated
# each French word.
#
# WHAT IS AN ATTENTION HEATMAP?
# A heatmap is a 2D grid where:
# - X axis : source words (English input)
# - Y axis : target words (French output)
# - Color : attention weight at that position
# white = low attention, dark blue = high attention
#
# Example for "I am cold" → "je suis froid":
#
# i am cold
# je [0.8 0.1 0.1] ← focused on "i"
# suis [0.1 0.8 0.1] ← focused on "am"
# froid [0.1 0.1 0.8] ← focused on "cold"
#
# A good attention model shows a roughly DIAGONAL pattern
# for language pairs with similar word order — like
# English and French for short sentences.
#
# WHY IS THIS USEFUL?
# - Makes the model INTERPRETABLE — we can see exactly what
# the model looks at when generating each output word
# - Helps diagnose errors — if a word is mistranslated,
# we can see whether the model focused on the wrong source
# - This is one of the most powerful properties of Attention
# — it gives us a window into the model's reasoning
# ============================================================
# ============================================================
# STEP 1 — Define the heatmap plotting function
# ============================================================
def plot_attention_heatmap(
src_tokens,
trg_tokens,
attn_matrix,
title='Attention Heatmap'
):
"""
Plots a single attention heatmap for one sentence pair.
Renders the attention weight matrix as a colour grid with
English source words on the x-axis and French target words
on the y-axis. Each cell shows how much attention the model
[page 126]
paid to that source word when generating that target word.
Args:
src_tokens : list of str — English content words
displayed as x-axis (column) labels
trg_tokens : list of str — French content words
displayed as y-axis (row) labels
attn_matrix : np.ndarray — shape (trg_len, src_len)
attention weights collected in translate_attn()
row t = weights when generating target token t
col s = how much focus on source position s
title : str — displayed above the plot
"""
import numpy as np
# Slice the attention matrix to match the actual token lengths.
# attn_matrix may have more rows than trg_tokens if the model
# stopped early at <eos> before filling all max_len steps.
# Similarly, columns are capped at len(src_tokens) to exclude
# <pad> positions that carry no real content.
attn_to_plot = attn_matrix[:len(trg_tokens), :len(src_tokens)]
# Figure size scales with sentence length so cells are
# always readable regardless of how long the sentence is.
fig, ax = plt.subplots(figsize=(
max(6, len(src_tokens) * 1.2), # width — wider for more src words
max(4, len(trg_tokens) * 1.0) # height — taller for more trg words
))
# imshow renders the 2D matrix as a colour grid.
# cmap='Blues' — white = low attention, dark blue = high attention.
# vmin=0.0, vmax=1.0 fixes the colour scale to [0, 1] so all
# heatmaps are comparable — a 0.8 always looks the same shade.
# aspect='auto' allows cells to be rectangular if needed.
im = ax.imshow(
attn_to_plot,
cmap='Blues',
aspect='auto',
vmin=0.0,
vmax=1.0
)
# Add a colorbar on the right to show what each shade means.
# fraction and pad control the size and spacing of the bar.
plt.colorbar(im, ax=ax, fraction=0.046, pad=0.04)
# Set tick positions and labels for both axes.
# rotation=45 on x-axis prevents overlapping labels for longer words.
# ha='right' aligns rotated labels correctly under their tick marks.
ax.set_xticks(range(len(src_tokens)))
ax.set_yticks(range(len(trg_tokens)))
[page 127]
ax.set_xticklabels(src_tokens, rotation=45, ha='right', fontsize=11)
ax.set_yticklabels(trg_tokens, fontsize=11)
ax.set_xlabel('Source — English', fontsize=12)
ax.set_ylabel('Target — French', fontsize=12)
ax.set_title(title, fontsize=13, fontweight='bold')
# Annotate each cell with its exact attention weight value.
# This lets the reader see precise numbers without estimating
# from colour alone — especially useful for cells near 0.5.
# Text colour adapts to maintain readability:
# weight > 0.5 (dark cell) → white text
# weight ≤ 0.5 (light cell) → black text
for i in range(len(trg_tokens)):
for j in range(len(src_tokens)):
val = attn_to_plot[i, j]
color = 'white' if val > 0.5 else 'black'
ax.text(
j, i, # column, row position in the grid
f'{val:.2f}', # formatted to 2 decimal places
ha='center',
va='center',
fontsize=9,
color=color
)
plt.tight_layout()
plt.show()
# ============================================================
# STEP 2 — Plot heatmaps for several sentences
# ============================================================
# For each sentence we call translate_attn() which returns
# the translation, predicted tokens, attention matrix, and
# source tokens — all four are needed to plot the heatmap.
# Individual heatmaps are printed one at a time so the student
# can read the translation and heatmap together.
heatmap_sentences = [
"I am cold.",
"She is happy.",
"We are here.",
"He loves her.",
"Come here.",
]
print("=" * 60)
print(" Attention Heatmaps — With Attention Model")
print("=" * 60)
print()
# Print a quick reading guide before showing any heatmaps
[page 128]
# so the student knows what to look for.
print(" Reading the heatmap:")
print(" - Each ROW = one French output word")
print(" - Each COLUMN = one English input word")
print(" - Darker cell = model paid more attention here")
print(" - Diagonal pattern = good word alignment")
print()
for sentence in heatmap_sentences:
# translate_attn() runs the full encode → greedy decode pipeline
# and collects attention weights at every step.
# translation : predicted French string
# trg_tokens : list of predicted French words (y-axis labels)
# attn_matrix : (trg_len, src_len) numpy array — the heatmap data
# src_tokens : list of English content words (x-axis labels)
translation, trg_tokens, attn_matrix, src_tokens = translate_attn(
sentence,
model_with_attn,
english_vocab,
french_vocab
)
print(f" English : {sentence}")
print(f" French : {translation}")
print()
# Guard against empty predictions — can happen with an undertrained
# model or a sentence containing only unknown words.
# Skip the heatmap rather than crash with an empty matrix.
if len(trg_tokens) == 0:
print(f" ⚠ No tokens predicted — skipping heatmap.")
print(f" Try training for more epochs.\n")
continue
# Guard against sentences where all tokens were filtered as
# special tokens — no content words left to label the x-axis.
if len(src_tokens) == 0:
print(f" ⚠ No source tokens found — skipping heatmap.\n")
continue
# Plot the individual heatmap for this sentence pair.
# Title shows the full "English → French" pair for context.
plot_attention_heatmap(
src_tokens = src_tokens,
trg_tokens = trg_tokens,
attn_matrix = attn_matrix,
title = f'"{sentence}" → "{translation}"'
)
# ============================================================
# STEP 3 — Plot all heatmaps in a grid (summary view)
[page 129]
# ============================================================
# Collect all valid sentence pairs into one figure so the
# student can compare attention patterns across sentences
# side by side — easier to spot consistent alignment behaviour
# and outliers than viewing them one at a time.
# First pass — collect only valid pairs to avoid empty subplots.
# A pair is valid only if both src_tokens and trg_tokens are
# non-empty — otherwise the heatmap axes would be unlabelled.
valid_pairs = []
for sentence in heatmap_sentences:
translation, trg_tokens, attn_matrix, src_tokens = translate_attn(
sentence,
model_with_attn,
english_vocab,
french_vocab
)
# Only include sentences where the model produced real output
# and the source had real content words to label the x-axis.
if len(trg_tokens) > 0 and len(src_tokens) > 0:
valid_pairs.append(
(src_tokens, trg_tokens, attn_matrix, sentence, translation)
)
if len(valid_pairs) > 0:
# Layout: 2 columns, as many rows as needed.
# math.ceil ensures we never drop the last pair if odd count.
n_cols = 2
n_rows = math.ceil(len(valid_pairs) / n_cols)
fig, axes = plt.subplots(
n_rows, n_cols,
figsize=(14, n_rows * 5) # height scales with number of rows
)
# Flatten axes array to a 1D list for simple index-based access.
# When n_rows=1, plt.subplots returns a 1D array — handle both cases.
axes = axes.flatten() if n_rows > 1 else axes
for idx, (src_tok, trg_tok, attn_mat, sent, trans) \
in enumerate(valid_pairs):
ax = axes[idx]
# Slice attention matrix to actual token dimensions —
# same reason as in plot_attention_heatmap() above.
attn_to_plot = attn_mat[
:len(trg_tok),
[page 130]
:len(src_tok)
]
# Render the attention matrix as a colour grid.
# Identical settings to the individual plots for consistency —
# same colour scale, same map, so all subplots are comparable.
im = ax.imshow(
attn_to_plot,
cmap='Blues',
aspect='auto',
vmin=0.0,
vmax=1.0
)
# Label axes with source and target tokens.
# Smaller fontsize than individual plots to fit the grid.
ax.set_xticks(range(len(src_tok)))
ax.set_yticks(range(len(trg_tok)))
ax.set_xticklabels(src_tok, rotation=45, ha='right', fontsize=9)
ax.set_yticklabels(trg_tok, fontsize=9)
ax.set_title(
f'"{sent}" → "{trans}"',
fontsize=9,
fontweight='bold'
)
# Annotate each cell with the exact weight value.
# Adaptive text colour keeps numbers readable on both
# light and dark cells — same logic as individual plots.
for i in range(len(trg_tok)):
for j in range(len(src_tok)):
val = attn_to_plot[i, j]
color = 'white' if val > 0.5 else 'black'
ax.text(
j, i,
f'{val:.2f}',
ha='center',
va='center',
fontsize=8,
color=color
)
# Hide any unused subplot slots — occurs when valid_pairs
# count is odd and the last row has an empty right cell.
for idx in range(len(valid_pairs), len(axes)):
axes[idx].set_visible(False)
# Overall title sits above all subplots.
# tight_layout adjusts spacing to prevent label overlap.
plt.suptitle(
'Attention Heatmaps — Summary Grid',
[page 131]
fontsize=14,
fontweight='bold'
)
plt.tight_layout()
plt.show()
# ============================================================
# STEP 4 — What to look for
# ============================================================
# Explain the four patterns a student might see in the heatmaps
# and what each one means about the model's behaviour.
# This turns the visual into a learning moment — not just a picture.
print(f"\n{'=' * 60}")
print(f" How to read the heatmaps")
print(f"{'=' * 60}")
print(f"""
DIAGONAL PATTERN (good alignment):
If the heatmap shows a diagonal from top-left to
bottom-right, the model has learned that English and
French words align roughly in the same order.
This is expected for short English-French sentences
where word order is similar between the two languages.
OFF-DIAGONAL (reordering):
Some French sentences reorder words relative to English.
e.g. adjective placement differs between the two languages.
Off-diagonal attention weights capture this reordering —
the model looked left or right to find the right source word.
UNIFORM WEIGHTS (uncertain):
If all weights are roughly equal (e.g. 0.14 across 7 words)
the model is not sure what to focus on — it spreads its
attention evenly rather than committing to any word.
This often happens early in training or with unseen words.
SHARP WEIGHTS (confident):
A single cell close to 1.0 means the model is very
confident — it knows exactly which source word matters
for this target word. Sharp weights typically emerge
after sufficient training on enough data.
NOTE:
With MAX_SEQ_LEN={MAX_SEQ_LEN} and only {len(train_dataset):,} training
samples
patterns may not be perfectly diagonal yet.
More data and more epochs produce cleaner, sharper heatmaps.
""")
# Final summary message — marks the end of the tutorial.
print(f" Tutorial complete!")
[page 132]
============================================================
Attention Heatmaps — With Attention Model
============================================================
Reading the heatmap:
- Each ROW = one French output word
- Each COLUMN = one English input word
- Darker cell = model paid more attention here
- Diagonal pattern = good word alignment
English : I am cold.
French : froid
English : She is happy.
French : est content
print(f" You have built and compared two seq2seq models")
print(f" from scratch — one without and one with Attention.")
[page 133]
English : We are here.
French : sommes ici
English : He loves her.
French : a
[page 134]
English : Come here.
French : ici
[page 135]
============================================================
How to read the heatmaps
============================================================
DIAGONAL PATTERN (good alignment):
If the heatmap shows a diagonal from top-left to
bottom-right, the model has learned that English and
French words align roughly in the same order.
This is expected for short English-French sentences
where word order is similar between the two languages.
OFF-DIAGONAL (reordering):
Some French sentences reorder words relative to English.
e.g. adjective placement differs between the two languages.
Off-diagonal attention weights capture this reordering —
the model looked left or right to find the right source word.
UNIFORM WEIGHTS (uncertain):
[page 136]
If all weights are roughly equal (e.g. 0.14 across 7 words)
the model is not sure what to focus on — it spreads its
attention evenly rather than committing to any word.
This often happens early in training or with unseen words.
SHARP WEIGHTS (confident):
A single cell close to 1.0 means the model is very
confident — it knows exactly which source word matters
for this target word. Sharp weights typically emerge
after sufficient training on enough data.
NOTE:
With MAX_SEQ_LEN=5 and only 4,000 training samples
patterns may not be perfectly diagonal yet.
More data and more epochs produce cleaner, sharper heatmaps.
Tutorial complete!
You have built and compared two seq2seq models
from scratch — one without and one with Attention.
Adding attention improved BLEU from 18.85% to 21.17% and reduced perplexity from
17.01 to 15.03. Every metric improved, which confirms attention is genuinely helping.
Key takeaways:
Without attention, the encoder compresses the entire sentence into one fixed vector,
which is a lot to ask of a single representation.
Attention removes that bottleneck by letting the decoder focus on different parts of
the source sentence at each step.
The gains are modest, which is expected on a 175K dataset. On WMT14 (36M pairs)
the gap would be much larger.
No changes were made to hyperparameters, data, or training setup, so the
improvement is purely from the attention mechanism itself.