# EngToFre usingTransformer GC v3

course: Module 4 — Generative AI & LLMs
module: Module-4-Generative-AI-LLMs
type: notebook
source_url: https://personal-learn.armco.dev/files/Module-4-Generative-AI-LLMs/General/Lab_Materials_-16-05-2026/EngToFre_usingTransformer_GC_v3.ipynb

---
[cell 1 markdown]
# Dataset Brief: Language Translation (English–French)

## Source
[Kaggle — Language Translation (English-French) by Devicharith](https://www.kaggle.com/datasets/devicharith/language-translation-englishfrench)

---

## What is this dataset?

It is a **bilingual sentence pair dataset** — for every English phrase, there is a matching French translation sitting right next to it. Think of it as a giant phrasebook stored in a spreadsheet.

---

## Structure

| Column | Description | Example |
|---|---|---|
| `English` | Source sentence | `"Go."` |
| `French` | Translated sentence | `"Va !"` |

- **File format:** CSV (`eng_-french.csv`)
- **Size:** ~175,000+ sentence pairs (we'll use a subset of ~1000–5000 for training)
- **Range:** From single words like `"Hi." → "Salut !"` to full multi-word sentences

---

Note: For demonstration purpose, we have selected only 200 rows of data.

[cell 2 code]
# cell_0a — GPU masking BEFORE any torch import
# =================================================
# MAIN GOAL:
#   Decide which physical GPU (if any) this notebook is allowed to see, and
#   hide all other GPUs from the process. This MUST happen before `import torch`
#   because torch reads CUDA_VISIBLE_DEVICES exactly once at import time —
#   changing it later has no effect.
#
# WHY THIS MATTERS:
#   On a shared machine with multiple GPUs, you don't want this notebook to
#   accidentally allocate memory on a GPU someone else is using. Setting
#   CUDA_VISIBLE_DEVICES="2" (for example) makes physical GPU 2 the *only*
#   GPU torch can see — and inside this process it will be re-numbered as
#   cuda:0 (because it's the only visible device).
#
# FLOW:
#   1. Ask the user for a GPU index (e.g. "0", "1", "2") or "cpu".
#   2. Validate the input is one of those forms — if not, STOP loudly instead
#      of silently falling back (silent fallback to CPU could waste hours).
#   3. Set the env var and remember the logical device string for later cells.

import os  # used to read/write process environment variables

# Prompt the user. .strip() removes accidental spaces, .lower() normalizes case.
choice = input(
    "Enter GPU physical index to use for cuda (e.g., 0, 1, 2) or type 'cpu' for CPU mode: "
).strip().lower()

if choice == "cpu":
    # Empty string for CUDA_VISIBLE_DEVICES = no GPUs visible to this process.
    os.environ["CUDA_VISIBLE_DEVICES"] = ""
    DEVICE = "cpu"
    print("CPU mode selected — no GPU will be visible to this process.")

elif choice.isdigit():
    # User gave a number like "0" or "2". We mask every GPU except this one.
    phys_idx = int(choice)
    os.environ["CUDA_VISIBLE_DEVICES"] = str(phys_idx)
    # Inside the process, the selected physical GPU is always re-indexed as 0,
    # because it is the *only* visible CUDA device.
    DEVICE = "cuda:0"
    print(
        f"Masking other GPUs. Physical GPU {phys_idx} is now visible "
        f"and will appear inside this process as cuda:0."
    )

else:
    # Anything else (typos like 'gpu0', 'cuda', empty string) -> raise loudly.
    # We deliberately do NOT silently fall back to CPU, because that can lead
    # to hours of slow training before the user notices.
    raise ValueError(
        f"Invalid choice {choice!r}. Please re-run this cell and enter "
        f"either a digit (e.g. '0') or 'cpu'."
    )

print(f"\n>>> DEVICE (logical, used in later cells): {DEVICE}")
print(f">>> CUDA_VISIBLE_DEVICES (env): {os.environ.get('CUDA_VISIBLE_DEVICES')!r}")

[cell 3 markdown]
# Installing required Libraries

[cell 4 code]
# cell_0
# ============================================================
# STEP 1: Install all the required libraries for this tutorial
# ============================================================
# Before we can build our translator, we need to install some Python libraries.
# Think of libraries as "toolboxes". Each one gives us ready-made tools so we
# don't have to build everything from scratch.

# transformers:
#    Hugging Face's library that lets us download and use pretrained AI models.
#    We'll use it to load "MarianMT", a model already trained to translate
#    between languages, so we don't have to train one from zero (which would
#    take days and a lot of computing power).

# sentencepiece:
#    A helper tool that breaks sentences into smaller pieces called "subwords".
#    Example: "unhappiness" becomes ["un", "happi", "ness"]
#    This is important because models understand small chunks better than
#    whole words, especially for rare or new words.

# sacrebleu:
#    A tool to measure how good our translations are.
#    It compares our model's output to a correct human translation and gives
#    a score called "BLEU" (higher means better translation).
#    Example: If the model says "Bonjour le monde" and the correct answer is
#    "Bonjour monde", sacrebleu tells us how close they are.

# sacremoses:
#    A text cleaning helper. It properly handles punctuation, spaces, and
#    capitalization when converting tokens back into readable sentences.
#    Example: ["Hello", ",", "world", "!"] becomes "Hello, world!"

# pandas:
#    A library to load and work with tables of data (like Excel in Python).
#    We'll use it to read our CSV file with English-French sentence pairs.

# The "!" at the start tells Jupyter Notebook: "Run this as a terminal command,
# not as Python code." "pip install" is how we install Python libraries.


!pip install transformers sentencepiece sacrebleu sacremoses pandas

[cell 5 markdown]
# Imports

[cell 6 code]
# cell_1
# ============================================================
# STEP 2: Import all the libraries we will use throughout this tutorial
# ============================================================
# "Importing" a library means loading it into our current Python session
# so we can use its tools. Think of it like opening a toolbox before
# starting work. We only need to do this once at the beginning.

import os                        # provides tools to work with file and folder paths
                                 # Example: checking if a file exists, joining paths

import pandas as pd              # loads and manipulates tabular data (like a spreadsheet)
                                 # "as pd" means we can write "pd" instead of "pandas"
                                 # Example: pd.read_csv("data.csv") loads a CSV file

import torch                     # PyTorch, the deep learning framework
                                 # we use it to check if a GPU is available
                                 # and to handle "tensors" (the number arrays that
                                 # the model reads and produces)

from transformers import (
    MarianMTModel,               # the pretrained model for translation
                                 # "MarianMT" is a family of models trained by
                                 # the University of Helsinki on millions of
                                 # sentence pairs. We load the English->French version.

    MarianTokenizer              # the tokenizer that pairs with MarianMTModel.
                                 # A tokenizer converts raw text into numbers
                                 # (which the model understands) and back to text.
                                 # Example: "Hello" -> [4, 156, 23] -> "Bonjour"
)

import sacrebleu                 # used later to evaluate translation quality
                                 # with a standard metric called BLEU score
                                 # (Bilingual Evaluation Understudy)

# ---------------------------------------------------------------
# Confirm setup: print PyTorch version and the device being used
# ---------------------------------------------------------------
# This is a quick sanity check to make sure everything imported correctly.

# torch.__version__ tells us which version of PyTorch is installed.
print(f"PyTorch version : {torch.__version__}")

# DEVICE tells us whether we are using a GPU (faster) or CPU (slower).
# It was defined in the previous cell (cell_0a).
# GPU example output  -> "Device selected : cuda"
# CPU example output  -> "Device selected : cpu"
print(f"Device selected : {DEVICE}")  # DEVICE was set in cell_0a

[cell 7 markdown]
# Data Loading

[cell 8 code]
# cell_2
# Goal: Load the English-French dataset from the local path into a pandas DataFrame

# Full path to the dataset CSV file
DATA_PATH ="/content/eng_-french.csv"

# Load the CSV file into a DataFrame
#df = pd.read_csv(DATA_PATH)
df = pd.read_csv(DATA_PATH, nrows=200)  # load only the first 200 rows

# Rename columns to simpler names for easier access throughout the notebook
df.columns = ["english", "french"]

# Confirm the dataset loaded correctly
print(f"Dataset shape : {df.shape}")
print(f"Columns       : {list(df.columns)}")
print(f"\nFirst 3 rows:")
print(df.head(3))

[cell 9 markdown]
# Basic EDA

[cell 10 code]
# cell_3
# ============================================================
# STEP 4: Explore the dataset to understand its structure and content
# ============================================================
# Before training or using any model, it is good practice to explore
# the data. This helps us catch problems early (like missing values
# or duplicates) and understand what kind of sentences we are working with.

# ---------------------------------------------------------------
# Basic info: how many rows and columns do we have?
# ---------------------------------------------------------------
# df.shape[0] gives the number of rows (sentence pairs).
# df.shape[1] gives the number of columns (should be 2: english, french).
print(f"Total rows    : {df.shape[0]}")
print(f"Total columns : {df.shape[1]}")

# ---------------------------------------------------------------
# Check for missing values
# ---------------------------------------------------------------
# A missing value means a cell in the table is empty (no text at all).
# df.isnull() returns True for every empty cell, and .sum() counts them
# per column. Ideally both columns should show 0.
# Example output:
#   english    0
#   french     0
# If we see a non-zero number, we would need to remove those rows
# before feeding the data to our model.
print(f"\nMissing values per column:")
print(df.isnull().sum())

# ---------------------------------------------------------------
# Check for duplicate rows
# ---------------------------------------------------------------
# A duplicate row means the exact same English-French pair appears
# more than once. Duplicates can make the model over-learn certain
# sentences, so it is good to know how many there are.
# df.duplicated() marks each repeated row as True, and .sum() counts them.
# Example output: "Duplicate rows: 32"
print(f"\nDuplicate rows: {df.duplicated().sum()}")

# ---------------------------------------------------------------
# Sample 5 random English-French sentence pairs
# ---------------------------------------------------------------
# df.sample(5) picks 5 random rows from the dataset so we can visually
# inspect what the sentence pairs look like.
# random_state=42 is a fixed "seed" so we get the same 5 rows every
# time we run this cell (makes results reproducible).
# .to_string(index=False) prints the table without the row numbers.
# Example output:
#   english                  french
#   I am tired.              Je suis fatigue.
#   She likes cats.          Elle aime les chats.
print(f"\n5 random sentence pairs:")
print(df.sample(5, random_state=42).to_string(index=False))

# ---------------------------------------------------------------
# Calculate sentence length (in words) for each row
# ---------------------------------------------------------------
# We want to know how long the sentences are on average.
# This helps us choose a good "max_length" when tokenizing later.
#
# How it works:
#   str(x).split() splits a sentence into a list of words by whitespace.
#   Example: "I love Paris".split() -> ["I", "love", "Paris"]
#   len(...) then counts how many words are in that list.
#   lambda x: ... is a small one-line function applied to every row.
#
# We store the word counts in two new columns: "english_len" and "french_len".
df["english_len"] = df["english"].apply(lambda x: len(str(x).split()))
df["french_len"]  = df["french"].apply(lambda x: len(str(x).split()))

# ---------------------------------------------------------------
# Print summary statistics for sentence lengths
# ---------------------------------------------------------------
# .describe() gives us useful statistics for a numeric column:
#   count  - how many rows were counted
#   mean   - average sentence length
#   std    - how much lengths vary (standard deviation)
#   min    - shortest sentence (in words)
#   25%    - 25% of sentences are shorter than this length
#   50%    - median sentence length
#   75%    - 75% of sentences are shorter than this length
#   max    - longest sentence (in words)
# .round(2) rounds all numbers to 2 decimal places for cleaner output.

print(f"\nEnglish sentence length (words):")
print(df["english_len"].describe().round(2))

print(f"\nFrench sentence length (words):")
print(df["french_len"].describe().round(2))

[cell 11 markdown]
## The Model We Are Using

### One model, not two
Before we begin, one important clarification. You will see two names used
throughout this notebook -- MarianMT and Helsinki-NLP/opus-mt-en-fr. These
are not two separate models. They refer to the same thing from two different
angles.

- **MarianMT** is the architecture -- the design and structure of the transformer.
  Think of it as the blueprint.
- **Helsinki-NLP/opus-mt-en-fr** is the trained model -- a specific instance of
  the MarianMT architecture that has been trained on millions of English-French
  sentence pairs by the Language Technology Research Group at the University of
  Helsinki. Think of it as the finished product built from that blueprint.

So when we write:

    model = MarianMTModel.from_pretrained("Helsinki-NLP/opus-mt-en-fr")

We are saying -- use the MarianMT architecture and load the weights that
Helsinki-NLP trained specifically for English to French translation.

### What is MarianMT?
MarianMT is a transformer based encoder-decoder architecture designed
specifically for machine translation. Unlike large general purpose language
models, MarianMT is a focused, lightweight model that does one thing well --
translating text from one language to another.

It works in two stages:
- The **encoder** reads the English sentence and builds a rich understanding
  of its meaning
- The **decoder** takes that understanding and generates the French translation
  word by word

### How was Helsinki-NLP/opus-mt-en-fr trained?
The model was trained on the OPUS dataset, which is a large collection of
translated texts gathered from the web, including subtitles, books, news
articles, and more. It was trained on millions of English-French sentence
pairs, which is why it performs so well even on sentences it has never
seen before.

### Why are we using it?
- It is small (75 million parameters) and fast -- works well even on CPU
- It is freely available on Hugging Face with no authentication required
- It is purpose built for English to French translation
- It is easy to use with just a few lines of code through the Hugging Face
  transformers library
- Despite its small size, it achieves a BLEU score of 51.16 on our dataset

[cell 12 markdown]
# Loading Tokenizer

[cell 13 code]
# cell_4
# ============================================================
# STEP 5: Load the pretrained tokenizer for the MarianMT model
# ============================================================
# Before the model can translate text, it needs the text to be converted
# into numbers. This is what a "tokenizer" does.
#
# The process looks like this:
#   "Hello, how are you?"
#        |
#        v  (tokenizer encodes)
#   [10537, 2, 541, 52, 55, 54, 0]   <-- numbers the model reads
#        |
#        v  (model translates)
#   [34, 78, 120, 56, 0]             <-- numbers the model outputs
#        |
#        v  (tokenizer decodes)
#   "Bonjour, comment allez-vous?"   <-- readable French text
#
# MarianTokenizer is the specific tokenizer designed to work with
# the MarianMT family of translation models.

# ---------------------------------------------------------------
# Define the model name
# ---------------------------------------------------------------
# This is the official name of the pretrained model on Hugging Face.
# "Helsinki-NLP" is the research group that trained it.
# "opus-mt-en-fr" means: trained on OPUS data, MarianMT, English to French.
# We define it as a variable here so we can reuse the same name in cell_5
# when we load the model itself (both the tokenizer and the model must match).
MODEL_NAME = "Helsinki-NLP/opus-mt-en-fr"

# ---------------------------------------------------------------
# Download and load the tokenizer
# ---------------------------------------------------------------
# from_pretrained() downloads the tokenizer files from Hugging Face
# on the first run (about 1-2 MB) and caches them locally.
# On subsequent runs it loads from the local cache (no internet needed).
tokenizer = MarianTokenizer.from_pretrained(MODEL_NAME)

# ---------------------------------------------------------------
# Quick sanity check: tokenize a sample sentence
# ---------------------------------------------------------------
# We run a simple English sentence through the tokenizer to see
# what it produces, just to confirm everything is working correctly.
sample = "Hello, how are you?"

# tokenizer.tokenize() splits the sentence into "subword pieces".
# Instead of splitting by whole words, it uses smaller units.
# This helps handle rare or unknown words.
# Example: "unhappiness" might become ["un", "happi", "ness"]
# The "▁" symbol marks the beginning of a new word (comes from SentencePiece).
# Example output: ['▁Hello', ',', '▁how', '▁are', '▁you', '?']
# Note: tokenize() does NOT add any special tokens, it just splits text.
tokens = tokenizer.tokenize(sample)

# tokenizer.encode() does two things in one step:
#   1. Splits the text into subword pieces (same as tokenize())
#   2. Converts each piece into its integer ID from the vocabulary
#      and appends a special EOS (End-Of-Sentence) token at the end.
# EOS token (ID = 0) tells the model "this is where the input ends".
# Example output: [10537, 2, 541, 52, 55, 54, 0]
#                                                ^ this 0 is the EOS token
# Notice encode() returns 7 numbers while tokenize() returned 6 pieces,
# because encode() adds the extra EOS token at the end.
token_ids = tokenizer.encode(sample)

# ---------------------------------------------------------------
# Print the results of the sanity check
# ---------------------------------------------------------------
print(f"Sample sentence : {sample}")

# The subword pieces the tokenizer split the sentence into.
# Example output: ['▁Hello', ',', '▁how', '▁are', '▁you', '?']
print(f"Tokens          : {tokens}")

# The integer IDs corresponding to each subword piece, plus EOS at the end.
# Example output: [10537, 2, 541, 52, 55, 54, 0]
print(f"Token IDs       : {token_ids}")

# The total number of unique tokens this tokenizer knows about.
# This is the size of its "vocabulary" (dictionary of all subword pieces).
# Example output: "Vocabulary size : 65001"
print(f"Vocabulary size : {tokenizer.vocab_size}")

[cell 14 markdown]
## Do we need to preprocess or clean the text?

Since we are using a pretrained transformer (MarianMT), most of the heavy text
processing is handled automatically by the tokenizer. Here is a clear breakdown:

### What the MarianMT tokenizer handles for you

- **Punctuation splitting** -- punctuation is separated from words automatically.
  We saw this in cell_4 where "Hello, how are you?" was split into:
  ['▁Hello', ',', '▁how', '▁are', '▁you', '?']
  Notice how the comma is a separate token.

- **Subword tokenization** -- words are broken into smaller pieces called subwords.
  We saw this in cell_4 where the ▁ prefix marks the start of each new word.
  This allows the tokenizer to handle words it has never seen before.

- **Special tokens** -- the tokenizer automatically adds an end of sentence marker.
  We saw this in cell_4 where the last token ID was 0, which is the end of sentence token.
  This tells the model where the input sentence ends.

- **Text to numbers** -- the tokenizer converts words into numbers the model can read.
  We saw this in cell_4 where "Hello" became 10537, "how" became 541, and so on.
  The model only works with these numbers, never with raw text directly.

### What you still need to be careful about

- **Missing values** -- passing an empty or null value to the tokenizer will crash.
  If a row in your dataset has no text, tokenizing it will throw an error.
  We already confirmed there are no missing values in cell_3, so we are safe.

- **Extremely long sentences** -- MarianMT can only handle up to 512 tokens at a time.
  Sentences longer than this get cut off and lose information at the end.
  Our longest sentence is only 44 words, so we are well within limits.

- **Encoding issues** -- unusual or special characters can sometimes confuse the tokenizer.
  This is common in corrupted CSV files or text copied from PDFs or web pages.
  Our dataset contains clean, well-formed text so this is not a concern here.

### Bottom line

For this specific dataset and this specific model, we can skip any preprocessing or
cleaning step entirely. The dataset is clean and the sentences are short and well-formed.

Note: If you were working with a raw, noisy dataset (for example, social media text,
OCR output, or user reviews), a cleaning step would be necessary before tokenization.

[cell 15 markdown]
# Loading Model

[cell 16 code]
# cell_5
# ============================================================
# STEP 6: Load the pretrained MarianMT model
# ============================================================
# Now that we have the tokenizer (cell_4), we load the actual translation
# model. The tokenizer and model always come as a pair:
#   - tokenizer : converts text <-> numbers
#   - model     : takes the numbers, does the translation
#
# Think of the tokenizer as a "translator's dictionary" and the model
# as the "translator's brain". Both must be from the same pretrained
# checkpoint to work correctly together.

# ---------------------------------------------------------------
# Download and load the model weights
# ---------------------------------------------------------------
# from_pretrained() downloads the model from Hugging Face on the first
# run (about 300 MB) and caches it locally for future runs.
# "Weights" are the millions of numbers inside the model that were learned
# during training on millions of English-French sentence pairs.
# We are loading those already-learned weights, so no training is needed.
# We reuse MODEL_NAME from cell_4 ("Helsinki-NLP/opus-mt-en-fr") to make
# sure the model and tokenizer are perfectly matched.
model = MarianMTModel.from_pretrained(MODEL_NAME)

# ---------------------------------------------------------------
# Move the model to the correct device (GPU or CPU)
# ---------------------------------------------------------------
# Deep learning models run much faster on a GPU than on a CPU.
# .to(DEVICE) moves all the model weights to the device we selected
# in cell_0a. If a GPU is available, DEVICE = "cuda", otherwise "cpu".
# It is important that both the model and the input data are on the
# same device, otherwise PyTorch will throw an error.
model = model.to(DEVICE)

# ---------------------------------------------------------------
# Set the model to evaluation mode
# ---------------------------------------------------------------
# A neural network behaves differently during training vs inference:
#
#   Training mode  : some neurons are randomly "dropped out" to prevent
#                    the model from memorizing data (called "dropout").
#                    Batch normalization uses running statistics.
#
#   Evaluation mode: dropout is turned off, all neurons are active.
#                    The model gives consistent, deterministic outputs.
#
# Since we are only using this model to translate (not to train it),
# we must call model.eval() to switch it to evaluation mode.
# If we forget this step, translations may randomly vary between runs.
model.eval()

# ---------------------------------------------------------------
# Quick sanity check: confirm the model loaded correctly
# ---------------------------------------------------------------
print(f"Model loaded successfully")

# next(model.parameters()) gets the first set of weights in the model.
# .device tells us which device those weights are stored on.
# Expected output: "cuda:0" (if GPU) or "cpu" (if no GPU available).
print(f"Model is running on : {next(model.parameters()).device}")

# This counts the total number of individual numbers (parameters) in the model.
# p.numel() returns the count of numbers in one layer's weight matrix.
# We sum these across all layers to get the total.
# The {:,} format adds commas for readability.
# Example output: "Number of parameters: 77,376,768"
# That is about 77 million numbers that were learned during training!
print(f"Number of parameters: {sum(p.numel() for p in model.parameters()):,}")

[cell 17 markdown]
# Fine-Tuning the Pretrained Model on Our Dataset

## What is Fine-Tuning?

So far we have used the MarianMT model exactly as it was downloaded from Hugging Face.
This is called **inference** -- we are just asking the model to translate without
teaching it anything new.

**Fine-tuning** means we take that already-trained model and train it a little more,
but this time specifically on OUR dataset. Think of it like this:

- The pretrained model went to university and learned to translate millions of sentences.
- Fine-tuning is like giving that graduate a short internship focused specifically
  on our style of sentences.

After fine-tuning, the model should perform better on sentences similar to those
in our dataset.

## Step 1 — Split the Data into Training and Validation Sets

Before we can fine-tune, we need to divide our dataset into two parts:

- **Training set (80%)** -- the sentences the model will learn from.
- **Validation set (20%)** -- sentences the model will NEVER see during training.
  We use these at the end of each epoch to check if the model is actually improving
  or just memorizing the training data.

Think of it like studying for an exam:
- Training set = your study material
- Validation set = a practice test with questions you haven't seen before

[cell 18 code]
# cell_5a
# ============================================================
# FINE-TUNING STEP 1: Split data into Train and Validation sets
# ============================================================
# Before we can fine-tune the model, we need to divide our dataset
# into two parts:
#   - Training set   : the model LEARNS from these sentences
#   - Validation set : we TEST the model on these after each epoch
#                      to check if it is improving
#
# We use sklearn's train_test_split() for this.
# It randomly shuffles and splits the data for us automatically.

from sklearn.model_selection import train_test_split

# ---------------------------------------------------------------
# Define the split ratio
# ---------------------------------------------------------------
# TEST_SIZE = 0.2 means:
#   20% of the data goes to validation
#   80% of the data goes to training
#
# Example with 2000 rows:
#   Training set   -> 1600 rows (80%)
#   Validation set ->  400 rows (20%)
TEST_SIZE = 0.2

# ---------------------------------------------------------------
# Perform the split
# ---------------------------------------------------------------
# train_test_split() takes two lists (English and French sentences)
# and splits both at the same time, keeping pairs matched correctly.
#
# Arguments:
#   df["english"].tolist() : all English sentences as a Python list
#   df["french"].tolist()  : all French sentences as a Python list
#   test_size=TEST_SIZE    : fraction to use for validation (0.2 = 20%)
#   random_state=42        : fixes the random shuffle so we get the
#                            same split every time we run this cell.
#                            Without this, the split would be different
#                            each run, making results hard to compare.
train_english, val_english, train_french, val_french = train_test_split(
    df["english"].tolist(),   # source sentences (input)
    df["french"].tolist(),    # target sentences (what we want the model to output)
    test_size=TEST_SIZE,
    random_state=42
)

# ---------------------------------------------------------------
# Confirm the split worked correctly
# ---------------------------------------------------------------
# We print the sizes to double-check the split looks right.
# Expected output for 2000 rows:
#   Total sentences    : 2000
#   Training sentences : 1600
#   Validation sentences: 400
print(f"Total sentences      : {len(df)}")
print(f"Training sentences   : {len(train_english)}  ({100 - int(TEST_SIZE*100)}%)")
print(f"Validation sentences : {len(val_english)}  ({int(TEST_SIZE*100)}%)")

# ---------------------------------------------------------------
# Peek at a few training pairs to confirm they are still matched
# ---------------------------------------------------------------
# It is important that English[i] and French[i] still correspond
# to each other after the split. Let's visually check 3 pairs.
print(f"\n--- Sample training pairs (English -> French) ---")
for i in range(3):
    print(f"  EN: {train_english[i]}")
    print(f"  FR: {train_french[i]}")
    print()

[cell 19 markdown]
# Step 2 — Create a Custom PyTorch Dataset Class

## What is a Dataset class?

PyTorch needs data to be served to the model in a very specific format.
A **Dataset class** is like a smart container that:
- Holds all our sentence pairs (English + French)
- Knows how many pairs it has
- Can hand out one pair at a time when asked

Think of it like a deck of cards:
- The deck holds all the cards (sentences)
- You can ask "how many cards are in the deck?" (length)
- You can ask "give me card number 5" (get one item)

PyTorch will use this class later to automatically feed batches of
sentences to the model during training.

## What is Tokenization of the Target (French)?

In cell_4 we only tokenized the English (input) side.
Now we also need to tokenize the French (target/output) side.

The French token IDs become the **labels** -- the correct answers
the model is trying to learn to produce.

One special thing: we replace the padding token ID with -100 in the labels.
This tells PyTorch "ignore this position when calculating the loss".
We do not want the model to be penalized for not predicting padding tokens
-- those are not real words, just fillers to make sentence lengths match.

[cell 20 code]
# cell_5b
# ============================================================
# FINE-TUNING STEP 2: Create a Custom PyTorch Dataset Class
# ============================================================
# PyTorch requires data to be wrapped in a Dataset class.
# This class acts as a smart container for our sentence pairs.
# It must implement exactly THREE methods:
#
#   __init__()  : called once when we create the dataset object.
#                 This is where we store and tokenize all the data.
#
#   __len__()   : returns the total number of sentence pairs.
#                 PyTorch calls this to know how many items exist.
#
#   __getitem__(): returns ONE tokenized sentence pair by index.
#                 PyTorch calls this repeatedly to build batches.
#
# Example:
#   dataset = TranslationDataset(train_english, train_french, tokenizer)
#   len(dataset)      -> 1600   (total pairs)
#   dataset[0]        -> {"input_ids": ..., "labels": ...}  (first pair)

from torch.utils.data import Dataset  # base class we inherit from; enforces the 3-method contract

class TranslationDataset(Dataset):
    """
    A custom Dataset that holds English-French sentence pairs
    and tokenizes them so PyTorch can feed them to the model.
    """

    def __init__(self, english_sentences, french_sentences, tokenizer, max_length=128):
        """
        Called once when we create the dataset.
        Tokenizes ALL sentences upfront and stores them.

        Args:
            english_sentences : list of English strings (the inputs)
            french_sentences  : list of French strings  (the targets)
            tokenizer         : the MarianTokenizer from cell_4
            max_length        : maximum number of tokens per sentence.
                                Sentences longer than this get cut off.
                                128 is safe for our dataset (max was 44 words).
        """

        self.tokenizer  = tokenizer   # save tokenizer so __getitem__ can access it later if needed
        self.max_length = max_length  # save max_length for reference

        # ---------------------------------------------------------------
        # Tokenize the English sentences (inputs)
        # ---------------------------------------------------------------
        # tokenizer() converts a list of raw strings into tensors of numbers.
        # Think of it as: ["Hello", "I am"] -> [[34, 56, 1, 1], [78, 99, 12, 1]]
        #
        # padding="max_length" : every sentence is padded to exactly 128 tokens.
        #                        e.g. "Hi" -> [416, 1, 1, 1, ..., 1]  (126 padding tokens added)
        #                        This is required because PyTorch batches must have
        #                        uniform shape -- all rows must be the same length.
        #
        # truncation=True      : if a sentence exceeds 128 tokens, cut it off.
        #                        Prevents crashes on unexpectedly long inputs.
        #
        # return_tensors="pt"  : return PyTorch tensors (not plain lists or numpy arrays).
        #                        Required so tensors can be directly passed to the model.
        #
        # Result: self.inputs is a dict with two keys:
        #   "input_ids"      : shape (N, 128) — one row per sentence, each row is 128 token IDs
        #   "attention_mask" : shape (N, 128) — 1 where there's a real token, 0 where padding
        self.inputs = tokenizer(
            english_sentences,
            padding="max_length",
            truncation=True,
            max_length=max_length,
            return_tensors="pt"
        )

        # ---------------------------------------------------------------
        # Tokenize the French sentences (targets / labels)
        # ---------------------------------------------------------------
        # Same tokenization settings as English for consistency.
        # These French token IDs are the "correct answers" the model must learn
        # to predict given an English input.
        #
        # Result: self.targets is a dict with the same two keys:
        #   "input_ids"      : shape (N, 128) — French token IDs
        #   "attention_mask" : shape (N, 128) — 1 for real tokens, 0 for padding
        self.targets = tokenizer(
            french_sentences,
            padding="max_length",
            truncation=True,
            max_length=max_length,
            return_tensors="pt"
        )

        # ---------------------------------------------------------------
        # Replace padding token IDs in labels with -100
        # ---------------------------------------------------------------
        # During training, PyTorch computes "loss" — how wrong the model's
        # predictions are compared to the correct French tokens (labels).
        #
        # Problem: padding tokens (value = tokenizer.pad_token_id, typically 1)
        # are NOT real words. We don't want the model penalized for getting
        # padding positions wrong — those positions are meaningless fillers.
        #
        # Solution: PyTorch's loss function (CrossEntropyLoss) has a special
        # convention: any label position with value -100 is completely ignored
        # in the loss calculation.
        #
        # So we replace every pad_token_id in the labels with -100.
        #
        # Concrete example (max_length=7):
        #   French sentence "Je t'aime" tokenizes to: [34, 78, 120, 56,  1,    1,    1  ]
        #                                             real tokens ^       ^ padding (ID=1)
        #   After masking:                            [34, 78, 120, 56, -100, -100, -100 ]
        #   Loss is now only computed on positions 0-3, not 4-6.

        self.labels = self.targets["input_ids"].clone()
        # .clone() makes a full independent copy of the tensor.
        # Without clone(), modifying self.labels would also modify self.targets["input_ids"]
        # because they'd point to the same memory.

        # .eq(pad_token_id)  -> boolean tensor: True where value == pad_token_id
        # .masked_fill_(mask, -100) -> in-place: replace True positions with -100
        # The trailing underscore _ means "in-place" operation (modifies self.labels directly)
        self.labels = self.labels.masked_fill(
            self.labels == tokenizer.pad_token_id, -100
        )

    def __len__(self):
        """
        Returns the total number of sentence pairs in this dataset.
        PyTorch calls this internally (e.g. to decide how many batches to create).

        Example:
            len(train_dataset) -> 1600
        """
        # .shape[0] gives the first dimension of the tensor, which equals the number of sentences.
        # e.g. if input_ids is shape (1600, 128), shape[0] = 1600
        return self.inputs["input_ids"].shape[0]

    def __getitem__(self, idx):
        """
        Returns ONE tokenized sentence pair by index.
        The DataLoader calls this repeatedly to assemble mini-batches during training.
        For example, to build a batch of 32 sentences, it calls this 32 times
        with different idx values and stacks the results.

        Args:
            idx : integer index of the sentence pair to retrieve (0-based)

        Returns:
            A dictionary with three tensors, each of shape (128,):
                "input_ids"      : English token IDs  e.g. [416, 2164, 123, ..., 1, 1]
                "attention_mask" : 1 for real tokens, 0 for padding  e.g. [1, 1, 1, ..., 0, 0]
                "labels"         : French token IDs with -100 at padding  e.g. [34, 78, ..., -100]

        Example:
            dataset[0] -> {
                "input_ids"      : tensor([416, 2164, 123, ..., 1, 1]),
                "attention_mask" : tensor([1,   1,    1,   ..., 0, 0]),
                "labels"         : tensor([34,  78,   56,  ..., -100, -100])
            }
        """
        # Index into the pre-tokenized tensors to get row `idx`.
        # e.g. self.inputs["input_ids"][3] returns the 4th English sentence's token IDs as a 1D tensor.
        return {
            "input_ids"      : self.inputs["input_ids"][idx],       # English tokens for this sentence
            "attention_mask" : self.inputs["attention_mask"][idx],  # mask: 1=real word, 0=padding
            "labels"         : self.labels[idx]                     # French tokens (-100 at padding)
        }


# ---------------------------------------------------------------
# Create the actual train and validation dataset objects
# ---------------------------------------------------------------
# Now we instantiate TranslationDataset twice:
#   train_dataset : wraps the 160 training sentence pairs
#   val_dataset   : wraps the 40 validation sentence pairs
#
# Both use the same tokenizer so the vocabulary (token IDs) is consistent.
# Tokenization happens here, inside __init__, not lazily per batch.
train_dataset = TranslationDataset(train_english, train_french, tokenizer)
val_dataset   = TranslationDataset(val_english,   val_french,   tokenizer)

# ---------------------------------------------------------------
# Sanity check: confirm the datasets were created correctly
# ---------------------------------------------------------------
print(f"Training dataset size   : {len(train_dataset)} sentence pairs")
print(f"Validation dataset size : {len(val_dataset)} sentence pairs")

# Retrieve the first item to verify the structure is as expected
sample_item = train_dataset[0]
print(f"\n--- Structure of one dataset item ---")
print(f"Keys                    : {list(sample_item.keys())}")
# All three tensors should be 1D with length 128 (our max_length)->set above in the method __init__ of TranslationDataset
print(f"input_ids shape         : {sample_item['input_ids'].shape}")        # expected: torch.Size([128])
print(f"attention_mask shape    : {sample_item['attention_mask'].shape}")   # expected: torch.Size([128])
print(f"labels shape            : {sample_item['labels'].shape}")           # expected: torch.Size([128])
# Preview the first 10 tokens of each tensor to spot-check values
print(f"\ninput_ids  (first 10)   : {sample_item['input_ids'][:10]}")   # English token IDs
print(f"labels     (first 10)   : {sample_item['labels'][:10]}")       # French token IDs (no -100 expected in first tokens)


#note:  59513 is a special token reserved for padding

[cell 21 markdown]
# Step 3 — Create DataLoaders

## What is a DataLoader?

In cell_5b we created a Dataset -- a smart container that holds all our
sentence pairs and can hand out one pair at a time.

But during training, we do not want to feed the model one sentence at a time.
That would be very slow. Instead we feed it a **batch** of sentences at once
(e.g. 16 sentences at a time).

A **DataLoader** wraps our Dataset and automatically:
- Groups sentences into batches of a fixed size
- Shuffles the training data before each epoch (so the model does not
  memorize the order of sentences)
- Loads the next batch in the background while the model is processing
  the current one (this speeds things up significantly)

Think of it like a conveyor belt in a factory:
- The Dataset is the warehouse of all raw materials (sentences)
- The DataLoader is the conveyor belt that delivers them in neat
  batches to the worker (the model) at a steady pace

## Training vs Validation DataLoader

We create TWO DataLoaders:
- **train_loader** : shuffles data each epoch (so the model sees sentences
  in a different order every time, which helps it learn better)
- **val_loader**   : does NOT shuffle (order does not matter for evaluation,
  and keeping it consistent makes results easier to interpret)

[cell 22 code]
# cell_5c
# ============================================================
# FINE-TUNING STEP 3: Create DataLoaders
# ============================================================
# A DataLoader wraps our Dataset and delivers data to the model
# in batches during training. It handles:
#   - Batching     : groups N sentences together for parallel processing
#   - Shuffling    : randomizes order each epoch (training only)
#   - Loading      : fetches the next batch while the current one is
#                    being processed (num_workers controls this)

from torch.utils.data import DataLoader  # PyTorch's built-in DataLoader class

# ---------------------------------------------------------------
# Define the batch size
# ---------------------------------------------------------------
# BATCH_SIZE controls how many sentence pairs the model sees at once.
#
# Larger batch = faster training BUT needs more GPU memory.
# Smaller batch = slower training BUT uses less GPU memory.
#
# 16 is a safe choice for fine-tuning on a dataset of our size.
# If you get an out-of-memory error, reduce this to 8.
# If you have a powerful GPU with lots of memory, you could try 32.
BATCH_SIZE = 16

# ---------------------------------------------------------------
# Create the Training DataLoader
# ---------------------------------------------------------------
# shuffle=True : randomly reorders the sentences before each epoch.
#   Why? Because if the model always sees sentences in the same order,
#   it might start to "memorize" the order rather than learning the
#   actual patterns. Shuffling prevents this.
#   Example: Epoch 1 order -> [5, 2, 8, 1, ...]
#            Epoch 2 order -> [3, 7, 1, 9, ...]  (different each time)
#
# num_workers=2 : uses 2 background processes to load the next batch
#   while the model is still processing the current one.
#   This keeps the GPU busy instead of waiting for data.
#   Set to 0 if you get any errors related to multiprocessing.
train_loader = DataLoader(
    train_dataset,       # the training dataset we created in cell_5b
    batch_size=BATCH_SIZE,
    shuffle=True,        # shuffle order every epoch
    num_workers=2        # background workers for faster loading
)

# ---------------------------------------------------------------
# Create the Validation DataLoader
# ---------------------------------------------------------------
# shuffle=False : we do NOT shuffle validation data.
#   Why? Because during validation we only care about the loss value,
#   not the order. Keeping it consistent also makes it easier to
#   debug if something goes wrong.
val_loader = DataLoader(
    val_dataset,         # the validation dataset we created in cell_5b
    batch_size=BATCH_SIZE,
    shuffle=False,       # no shuffling for validation
    num_workers=2        # same background loading for speed
)

# ---------------------------------------------------------------
# Sanity check: confirm the DataLoaders look correct
# ---------------------------------------------------------------
# The number of batches = total sentences / batch size (rounded up).
# Example with BATCH_SIZE=16:
#   Training   : 1600 sentences / 16 = 100 batches
#   Validation :  400 sentences / 16 =  25 batches
print(f"Batch size              : {BATCH_SIZE}")
print(f"Training batches        : {len(train_loader)}  "
      f"({len(train_dataset)} sentences / {BATCH_SIZE} per batch)")
print(f"Validation batches      : {len(val_loader)}  "
      f"({len(val_dataset)} sentences / {BATCH_SIZE} per batch)")

# ---------------------------------------------------------------
# Peek at one batch to confirm the structure looks right
# ---------------------------------------------------------------
# next(iter(train_loader)) fetches the very first batch.
#   iter()  : converts the DataLoader into an iterator
#   next()  : asks it for the first item (one batch)
#
# Each batch is a dictionary with three keys (same as our Dataset):
#   "input_ids"      : shape (BATCH_SIZE, 128) -- token IDs for English
#   "attention_mask" : shape (BATCH_SIZE, 128) -- 1s for real, 0s for padding
#   "labels"         : shape (BATCH_SIZE, 128) -- token IDs for French
sample_batch = next(iter(train_loader))

print(f"\n--- Structure of one batch ---")
print(f"input_ids shape      : {sample_batch['input_ids'].shape}")
      # Expected: torch.Size([16, 128])
print(f"attention_mask shape : {sample_batch['attention_mask'].shape}")
      # Expected: torch.Size([16, 128])
print(f"labels shape         : {sample_batch['labels'].shape}")
      # Expected: torch.Size([16, 128])

[cell 23 markdown]
# Step 4 — Training and Validation Loop

## What happens in a Training Loop?

This is the heart of fine-tuning. In each **epoch** (one full pass
through the training data), the model does the following for every batch:

1. **Forward pass** -- the model looks at the English sentences and
   tries to predict the French translations. It compares its predictions
   to the correct French labels and calculates a **loss** (how wrong it was).

2. **Backward pass** -- the model works backwards through its own
   calculations to figure out which internal numbers (weights) caused
   the error. This is called **backpropagation**.

3. **Update weights** -- the optimizer nudges the weights slightly in
   the direction that reduces the loss. Over many batches, the model
   gradually gets better.

Think of it like learning to throw darts:
- You throw (forward pass)
- You see how far off you were (loss)
- You adjust your technique (backward pass + weight update)
- You throw again, slightly better each time

## What is the Validation Loop?

After each epoch of training, we run the model on the validation set
(sentences it has never trained on) to check:
- Is the loss going DOWN? Good -- the model is learning.
- Is the validation loss going UP while training loss goes down? Bad --
  the model is memorizing training data (overfitting).

## Key concepts in this cell

- **AdamW optimizer** -- the algorithm that updates the model weights.
  "W" stands for weight decay, a technique that prevents the model from
  making any single weight too large, which helps avoid overfitting.

- **Learning rate (5e-5)** -- how big each weight update step is.
  Too large = the model overshoots and forgets what it already learned.
  Too small = training takes forever.
  5e-5 (0.00005) is the standard safe choice for fine-tuning transformers.

- **Epoch** -- one complete pass through all training data.
  We train for 3 epochs, meaning the model sees each sentence 3 times.

[cell 24 code]
# cell_5d
# ============================================================
# FINE-TUNING STEP 4: Training and Validation Loop
# ============================================================
# This cell fine-tunes the pretrained MarianMT model on our
# English-French dataset. It runs for a fixed number of epochs.
# Each epoch has two phases:
#   1. Training phase   : model learns from training batches
#   2. Validation phase : model is evaluated on unseen data

from torch.optim import AdamW  # the optimizer that updates model weights
from tqdm import tqdm          # displays a live progress bar

# ---------------------------------------------------------------
# Hyperparameters
# ---------------------------------------------------------------
# These are the settings that control how training behaves.
# They are called "hyperparameters" because we set them manually
# (the model does not learn them automatically).

# Number of times we go through the entire training dataset.
# 3 epochs is a good starting point for fine-tuning.
# More epochs = more learning BUT risk of overfitting.
NUM_EPOCHS = 1

# How big each weight update step is.
# 5e-5 means 0.00005 -- a very small step, which is intentional.
# Fine-tuning needs small steps so we don't destroy what the
# pretrained model already learned. This is called avoiding
# "catastrophic forgetting".
LEARNING_RATE = 5e-5

# ---------------------------------------------------------------
# Set up the optimizer
# ---------------------------------------------------------------
# AdamW is the standard optimizer for fine-tuning transformers.
# It takes all the model's weights (model.parameters()) and
# updates them slightly after each batch based on the loss.
#
# weight_decay=0.01 is a regularization technique.
# It gently penalizes very large weights, which helps prevent
# the model from overfitting to the training data.
optimizer = AdamW(model.parameters(), lr=LEARNING_RATE, weight_decay=0.01)

# ---------------------------------------------------------------
# Switch model to TRAINING mode
# ---------------------------------------------------------------
# Remember in cell_5 we called model.eval() to turn OFF dropout.
# Now we call model.train() to turn dropout back ON.
# During training, dropout randomly deactivates some neurons each
# step, which forces the model to learn more robust patterns
# instead of relying too heavily on any single neuron.
model.train()

# ---------------------------------------------------------------
# Storage for tracking loss across epochs
# ---------------------------------------------------------------
# We store the average loss for each epoch so we can print a
# summary at the end and see if training is going in the right direction.
# Loss should decrease over epochs if the model is learning correctly.
train_losses = []   # average training loss per epoch
val_losses   = []   # average validation loss per epoch

# ---------------------------------------------------------------
# Main training loop
# ---------------------------------------------------------------
# We loop NUM_EPOCHS times. Each iteration is one full pass
# through the entire training dataset.
for epoch in range(NUM_EPOCHS):

    print(f"\n{'='*60}")
    print(f"EPOCH {epoch + 1} of {NUM_EPOCHS}")
    print(f"{'='*60}")

    # ===========================================================
    # PHASE 1: TRAINING
    # ===========================================================
    model.train()          # ensure model is in training mode
    total_train_loss = 0   # accumulate loss across all batches

    # tqdm wraps train_loader to show a live progress bar.
    # desc= sets the label shown next to the bar.
    train_bar = tqdm(train_loader, desc=f"  Training  ")

    for batch in train_bar:

        # -------------------------------------------------------
        # Move batch data to the correct device (GPU or CPU)
        # -------------------------------------------------------
        # Every tensor must be on the same device as the model.
        # We move all three tensors in the batch to DEVICE.
        input_ids      = batch["input_ids"].to(DEVICE)
        attention_mask = batch["attention_mask"].to(DEVICE)
        labels         = batch["labels"].to(DEVICE)

        # -------------------------------------------------------
        # Step 1: Zero out gradients from the previous batch
        # -------------------------------------------------------
        # PyTorch accumulates gradients by default -- it adds new
        # gradients ON TOP of old ones. We must reset them to zero
        # before each batch, otherwise the updates will be wrong.
        # Think of it like wiping the whiteboard clean before
        # solving a new math problem.
        optimizer.zero_grad()

        # -------------------------------------------------------
        # Step 2: Forward pass -- compute the loss
        # -------------------------------------------------------
        # We pass the English tokens (input_ids, attention_mask)
        # AND the correct French tokens (labels) to the model.
        #
        # When labels are provided, MarianMT automatically:
        #   1. Runs the encoder on the English input
        #   2. Runs the decoder to predict French tokens one by one
        #   3. Compares predictions to labels
        #   4. Computes and returns the cross-entropy loss
        #
        # Cross-entropy loss measures how wrong the predictions are.
        # Lower loss = better predictions.
        outputs = model(
            input_ids=input_ids,
            attention_mask=attention_mask,
            labels=labels
        )
        loss = outputs.loss  # the scalar loss value for this batch

        # -------------------------------------------------------
        # Step 3: Backward pass -- compute gradients
        # -------------------------------------------------------
        # loss.backward() tells PyTorch to work backwards through
        # all the calculations and compute how much each weight
        # in the model contributed to the loss.
        # These are called "gradients" -- they tell the optimizer
        # which direction to nudge each weight.
        loss.backward()

        # -------------------------------------------------------
        # Step 4: Clip gradients
        # -------------------------------------------------------
        # Sometimes gradients can become extremely large, causing
        # the optimizer to make a huge update that destabilizes
        # training. This is called "exploding gradients".
        # clip_grad_norm_() caps the total gradient size at 1.0,
        # preventing any single update from being too drastic.
        # This is standard practice when fine-tuning transformers.
        torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)

        # -------------------------------------------------------
        # Step 5: Update the weights
        # -------------------------------------------------------
        # optimizer.step() uses the gradients computed in backward()
        # to nudge each weight slightly in the direction that reduces
        # the loss. The size of each nudge is controlled by LEARNING_RATE.
        optimizer.step()

        # -------------------------------------------------------
        # Track the loss for this batch
        # -------------------------------------------------------
        # loss.item() converts the tensor loss to a plain Python float.
        # We add it to our running total so we can compute the average
        # loss across all batches at the end of the epoch.
        total_train_loss += loss.item()

        # Update the progress bar to show the current batch loss
        # .4f formats the number to 4 decimal places
        train_bar.set_postfix({"batch_loss": f"{loss.item():.4f}"})

    # Compute average training loss for this epoch
    # We divide by the number of batches (len(train_loader))
    avg_train_loss = total_train_loss / len(train_loader)
    train_losses.append(avg_train_loss)
    print(f"\n  Avg Training Loss   : {avg_train_loss:.4f}")

    # ===========================================================
    # PHASE 2: VALIDATION
    # ===========================================================
    # After each epoch of training, we evaluate the model on the
    # validation set. We do NOT update any weights here -- we only
    # measure how well the model performs on unseen data.

    model.eval()          # switch to evaluation mode (turns off dropout)
    total_val_loss = 0    # accumulate validation loss

    # torch.no_grad() tells PyTorch not to track gradients.
    # We don't need gradients during validation (no weight updates),
    # so this saves memory and speeds things up.
    with torch.no_grad():

        val_bar = tqdm(val_loader, desc=f"  Validating")

        for batch in val_bar:

            # Move batch to the correct device
            input_ids      = batch["input_ids"].to(DEVICE)
            attention_mask = batch["attention_mask"].to(DEVICE)
            labels         = batch["labels"].to(DEVICE)

            # Forward pass only -- no backward pass, no weight update
            outputs = model(
                input_ids=input_ids,
                attention_mask=attention_mask,
                labels=labels
            )
            loss = outputs.loss
            total_val_loss += loss.item()

            val_bar.set_postfix({"batch_loss": f"{loss.item():.4f}"})

    # Compute average validation loss for this epoch
    avg_val_loss = total_val_loss / len(val_loader)
    val_losses.append(avg_val_loss)
    print(f"  Avg Validation Loss : {avg_val_loss:.4f}")

# ---------------------------------------------------------------
# Print final summary across all epochs
# ---------------------------------------------------------------
# A good sign: both training and validation loss should decrease
# across epochs. If validation loss starts increasing while
# training loss keeps decreasing, the model is overfitting.
print(f"\n{'='*60}")
print(f"TRAINING COMPLETE")
print(f"{'='*60}")
print(f"{'Epoch':<10} {'Train Loss':<20} {'Val Loss':<20}")
print(f"{'-'*50}")
for i, (tl, vl) in enumerate(zip(train_losses, val_losses)):
    print(f"{i+1:<10} {tl:<20.4f} {vl:<20.4f}")

[cell 25 markdown]
# Step 5 — Save the Fine-Tuned Model

## Why Save the Model?

Training takes time and computing power. If we close the notebook
or the runtime crashes, all that fine-tuning work is lost forever
unless we save it to disk.

## What gets saved?

Two things need to be saved together -- they are a matched pair
and both are needed to translate later:

- **Model weights** -- the millions of numbers inside the model
  that were updated during fine-tuning. Saved by model.save_pretrained()

- **Tokenizer** -- the vocabulary and rules for converting text to
  numbers and back. Saved by tokenizer.save_pretrained()

## Where does it get saved?

We reuse the DATA_PATH variable from cell_2 (where our CSV dataset
is located) and save the model in the SAME folder.

This means:
- No hardcoded paths
- No manual changes needed
- Works automatically on both Google Colab and any local PC
- Everything stays neatly in one place -- dataset and model together

[cell 26 code]
# cell_5e  (updated -- saves to the same folder as the dataset)
# ============================================================
# FINE-TUNING STEP 5: Save the Fine-Tuned Model
# ============================================================
# We save both the model weights and the tokenizer to the SAME
# folder where our dataset CSV file lives.
#
# We reuse DATA_PATH from cell_2 so there is no hardcoding --
# this works automatically on both Google Colab and local PC
# without any manual changes.

import os

# ---------------------------------------------------------------
# Derive the save path from DATA_PATH (defined in cell_2)
# ---------------------------------------------------------------
# os.path.dirname() extracts the folder part from a full file path.
#
# Example:
#   DATA_PATH = "/content/eng_-french.csv"
#   os.path.dirname(DATA_PATH) -> "/content"
#   SAVE_PATH                  -> "/content/finetuned_model"
#
# Another example on a local PC:
#   DATA_PATH = "C:/Users/you/project/eng_-french.csv"
#   os.path.dirname(DATA_PATH) -> "C:/Users/you/project"
#   SAVE_PATH                  -> "C:/Users/you/project/finetuned_model"
#
# This means the model is ALWAYS saved next to the CSV file,
# regardless of which machine or environment you are running on.
SAVE_PATH = os.path.join(os.path.dirname(DATA_PATH), "finetuned_model")

print(f"Dataset location : {os.path.dirname(DATA_PATH)}")
print(f"Model will be saved to : {SAVE_PATH}")

# ---------------------------------------------------------------
# Create the save folder if it does not already exist
# ---------------------------------------------------------------
# exist_ok=True means: do not raise an error if the folder
# already exists (e.g. if we run this cell a second time).
os.makedirs(SAVE_PATH, exist_ok=True)
print(f"Folder created (or already exists)")

# ---------------------------------------------------------------
# Save the fine-tuned model weights
# ---------------------------------------------------------------
# save_pretrained() saves the model architecture config AND
# all the weight values that were updated during fine-tuning.
# This is the large file (~300 MB) that contains everything
# the model learned during our fine-tuning in cell_5d.
model.save_pretrained(SAVE_PATH)
print(f"\nModel saved      : {SAVE_PATH}")

# ---------------------------------------------------------------
# Save the tokenizer
# ---------------------------------------------------------------
# The tokenizer must always be saved alongside the model.
# They are a matched pair -- using a different tokenizer with
# this model would produce garbage translations.
tokenizer.save_pretrained(SAVE_PATH)
print(f"Tokenizer saved  : {SAVE_PATH}")

# ---------------------------------------------------------------
# Confirm what was saved by listing the files
# ---------------------------------------------------------------
# We list every file in the save folder along with its size
# so we can visually confirm everything was saved correctly.
saved_files = os.listdir(SAVE_PATH)
print(f"\nFiles saved ({len(saved_files)} total):")
for f in sorted(saved_files):
    size_mb = os.path.getsize(os.path.join(SAVE_PATH, f)) / (1024 * 1024)
    print(f"  {f:<40} {size_mb:.2f} MB")

# ---------------------------------------------------------------
# Show how to reload the model later
# ---------------------------------------------------------------
# We print the exact code the user needs to reload this model
# in any future notebook session.
print(f"""
To reload your fine-tuned model in any future session:

    from transformers import MarianMTModel, MarianTokenizer

    model     = MarianMTModel.from_pretrained(r"{SAVE_PATH}")
    tokenizer = MarianTokenizer.from_pretrained(r"{SAVE_PATH}")
    model     = model.to(DEVICE)
    model.eval()
""")

[cell 27 markdown]
# Translation

[cell 28 code]
# cell_6
# ============================================================
# STEP 7: Translate a single English sentence to French
# ============================================================
# This cell walks through the full translation pipeline step by step.
# We do it manually here (instead of using a function) so we can see
# exactly what happens at each stage. Later we will wrap this into
# a reusable function.
#
# The full pipeline looks like this:
#   "I love learning new things every day."   (raw English text)
#            |
#            v  Step 1: Tokenize
#   tensor([[416, 2164, 2onal, ..., 0]])       (numbers the model reads)
#            |
#            v  Step 2: Model generates output
#   tensor([[34, 78, 120, 56, 0]])             (numbers the model outputs)
#            |
#            v  Step 3: Decode
#   "J'adore apprendre de nouvelles choses."  (readable French text)

sample_sentence = "I love learning new things every day."

# ---------------------------------------------------------------
# Step 1: Tokenize the input sentence
# ---------------------------------------------------------------
# The tokenizer converts the raw English text into a PyTorch tensor
# of integer IDs that the model can process.
#
# Arguments explained:
#   return_tensors="pt" : return the result as a PyTorch tensor.
#                         "pt" stands for PyTorch. The alternative is
#                         "np" for NumPy arrays, but the model needs "pt".
#
#   padding=True        : if we pass multiple sentences, they may have
#                         different lengths. Padding adds extra zeros to
#                         shorter sentences so all inputs are the same length.
#                         Not strictly needed for a single sentence, but good
#                         practice for when we process batches later.
#
#   .to(DEVICE)         : moves the tensor to the same device (GPU or CPU)
#                         as the model. Both must be on the same device or
#                         PyTorch will throw an error.
inputs = tokenizer(
    sample_sentence,
    return_tensors="pt",   # return as PyTorch tensors
    padding=True           # pad the input if needed
).to(DEVICE)

print(f"Input sentence     : {sample_sentence}")

# inputs is a dictionary with key "input_ids" holding the token ID tensor.
# Example output: tensor([[416, 2164, 123, 456, 78, 90, 12, 0]])
#                                                              ^ EOS token
print(f"Tokenized input IDs: {inputs['input_ids']}")

# ---------------------------------------------------------------
# Show tokenization examples from the actual dataset
# ---------------------------------------------------------------
# Looking at real sentences from our dataset (not just the sample above)
# helps us understand how the tokenizer handles everyday language.
# We pick 3 random rows using the same random_state as before for
# reproducibility.
print(f"\n--- Tokenization examples from the dataset ---")
print(f"{'English':<35} {'Tokens':<55} {'Token IDs'}")
print("-" * 120)

for _, row in df.sample(3, random_state=42).iterrows():
    # tokenize() shows us the subword pieces the text was split into
    tokens_ex    = tokenizer.tokenize(row["english"])
    # encode() shows us the integer IDs, including the final EOS token (0)
    token_ids_ex = tokenizer.encode(row["english"])
    print(f"{row['english']:<35} {str(tokens_ex):<55} {token_ids_ex}")

# ---------------------------------------------------------------
# Step 2: Generate the translated output tokens
# ---------------------------------------------------------------
# model.generate() runs the actual translation. It takes the input
# token IDs and produces a sequence of output token IDs in French.
#
# How it works internally (simplified):
#   - The encoder reads the English token IDs and builds a rich
#     numerical representation of the meaning of the sentence.
#   - The decoder then generates French token IDs one at a time,
#     each time looking at both the encoder output and the French
#     tokens it has already generated, until it produces the EOS token.
#
# torch.no_grad() is a context manager that tells PyTorch:
#   "We are not training, so do not store gradient information."
# During training, PyTorch tracks every calculation to compute gradients
# for updating the model weights. During inference we do not need this,
# so turning it off saves memory and speeds up the translation.
#
# **inputs unpacks the dictionary {"input_ids": ..., "attention_mask": ...}
# and passes each item as a separate argument to model.generate().
with torch.no_grad():
    translated_tokens = model.generate(**inputs)

# The output is a tensor of integer IDs representing the French translation.
# Example output: tensor([[38, 200, 456, 789, 23, 0]])
#                                                  ^ EOS token
print(f"\nTranslated token IDs: {translated_tokens}")

# ---------------------------------------------------------------
# Step 3: Decode the output tokens back into readable French text
# ---------------------------------------------------------------
# tokenizer.decode() converts the integer IDs back into a readable string.
#
# translated_tokens[0] selects the first (and only) translation in the batch.
# If we had passed a batch of 3 sentences, we would have [0], [1], [2].
#
# skip_special_tokens=True removes special tokens like the EOS token (0)
# from the output. Without this, the output might look like:
#   "J'adore apprendre de nouvelles choses.</s>"
# With it, we get clean readable text:
#   "J'adore apprendre de nouvelles choses."
translated_text = tokenizer.decode(translated_tokens[0], skip_special_tokens=True)

# Print the final result: the original English and its French translation
print(f"\nEnglish : {sample_sentence}")
print(f"French  : {translated_text}")

#Note: n MarianTokenizer, 59513 serves double duty:
#Context Middle/end of a sequence       :   Meaning -> Padding — ignore me
#Context First token of decoder input   :   Meaning -> Start of sequence — begin generating
print(tokenizer.pad_token_id)          # 59513
print(model.config.decoder_start_token_id)  # also 59513 ← same token!

[cell 29 code]
# cell_7
# ============================================================
# STEP 8: Wrap the translation logic into a reusable function
# ============================================================
# In cell_6 we translated one sentence by writing out every step manually.
# That works, but if we want to translate many sentences throughout the
# notebook, we would have to copy and paste the same code every time.
#
# Instead, we define a function called translate() once here, and then
# call it with a single line anywhere in the notebook.
#
# The function also handles two input formats automatically:
#   - A single string : translate("Hello")
#   - A list of strings: translate(["Hello", "How are you?"])
# In both cases it always returns a list of French translations.

def translate(texts):
    """
    Translates a single sentence or a list of sentences from English to French.

    Args:
        texts: a single string or a list of strings in English

    Returns:
        a list of translated French strings
    """

    # ---------------------------------------------------------------
    # Handle both single string and list inputs
    # ---------------------------------------------------------------
    # isinstance(texts, str) checks whether the input is a single string.
    # If it is, we wrap it in a list so the rest of the function always
    # works with a list. This avoids writing two separate code paths.
    # Example:
    #   "Hello"           becomes ["Hello"]        (single string wrapped)
    #   ["Hello", "Hi"]   stays   ["Hello", "Hi"]  (already a list, unchanged)
    if isinstance(texts, str):
        texts = [texts]

    # ---------------------------------------------------------------
    # Step 1: Tokenize the input
    # ---------------------------------------------------------------
    # Same tokenization as in cell_6, but now with one extra argument:
    #
    #   truncation=True : the model has a maximum input length of 512 tokens.
    #                     If a sentence is longer than that, truncation=True
    #                     cuts it off at 512 tokens instead of throwing an error.
    #                     Example: a very long paragraph gets cut to 512 tokens.
    #
    #   padding=True    : when translating a batch (list) of sentences, they
    #                     may have different lengths. Padding adds a special
    #                     padding token (ID = 1) to the end of shorter sentences
    #                     so all sentences in the batch have the same length.
    #                     Example:
    #                       "Hi."         -> [4, 0, 1, 1]   (padded with 1s)
    #                       "How are you?"-> [6, 52, 55, 0] (no padding needed)
    #
    #   return_tensors="pt" : return as PyTorch tensors (needed by the model).
    #
    #   .to(DEVICE)     : move tensors to the same device as the model.
    inputs = tokenizer(
        texts,
        return_tensors="pt",
        padding=True,
        truncation=True
    ).to(DEVICE)

    # ---------------------------------------------------------------
    # Step 2: Generate translated token IDs
    # ---------------------------------------------------------------
    # model.generate() translates all sentences in the batch at once.
    # Processing multiple sentences together (batching) is much faster
    # than translating them one by one, especially on a GPU.
    # torch.no_grad() disables gradient tracking to save memory and time,
    # since we are doing inference (not training).
    with torch.no_grad():
        translated_tokens = model.generate(**inputs)

    # ---------------------------------------------------------------
    # Step 3: Decode all outputs back into readable French text
    # ---------------------------------------------------------------
    # In cell_6 we used tokenizer.decode() for a single sentence.
    # Here we use tokenizer.batch_decode() which does the same thing
    # but for a whole list of translated token sequences at once.
    # It returns a list of French strings, one per input sentence.
    # skip_special_tokens=True removes the EOS token from each output.
    # Example output: ["J'adore apprendre de nouvelles choses chaque jour."]
    translations = tokenizer.batch_decode(
        translated_tokens,
        skip_special_tokens=True
    )

    # Return the list of French translations to the caller.
    # Even if the input was a single string, we always return a list,
    # so the caller can always access the result as result[0].
    return translations


# ---------------------------------------------------------------
# Quick test: confirm the function works correctly
# ---------------------------------------------------------------
# We test with the same sentence from cell_6 so we can compare outputs
# and confirm the function produces the same result as the manual steps.
test_sentence = "I love learning new things every day."
result = translate(test_sentence)

# result is a list, so we access the first (and only) translation with [0].
# Expected output:
#   English : I love learning new things every day.
#   French  : J'adore apprendre de nouvelles choses chaque jour.
print(f"English : {test_sentence}")
print(f"French  : {result[0]}")

[cell 30 code]
# cell_8
# ============================================================
# STEP 9: Translate a batch of real sentences from the dataset
# ============================================================
# In cell_7 we tested our translate() function on a single hand-picked
# sentence. Now we test it on real sentences from our dataset to see
# how well the model performs on actual data.
#
# We also display the model's predicted French translation alongside
# the actual (ground truth) French translation from the dataset.
# This lets us visually judge translation quality before we compute
# a formal score in the next cell.

# ---------------------------------------------------------------
# Sample 10 rows from the dataset
# ---------------------------------------------------------------
# df.sample(10) picks 10 random rows from our dataset.
# random_state=42 ensures we get the same 10 rows every time we run
# this cell (reproducibility). You can change the number to any value
# to see different sentences.
#
# reset_index(drop=True) resets the row numbers to 0-9.
# Without this, the row numbers would be random (e.g., 234, 1892, 7103...)
# because they came from random positions in the original dataset.
# drop=True means we discard the old row numbers entirely.
sample_df = df.sample(10, random_state=42).reset_index(drop=True)

# ---------------------------------------------------------------
# Extract the English sentences as a plain Python list
# ---------------------------------------------------------------
# Our translate() function expects a list of strings as input.
# .tolist() converts the pandas column (a Series) into a plain Python list.
# Example:
#   ["I am happy.", "She likes cats.", "We are learning.", ...]
english_sentences = sample_df["english"].tolist()

# ---------------------------------------------------------------
# Translate all 10 sentences in one batch
# ---------------------------------------------------------------
# We pass the entire list to translate() at once (batch translation).
# This is more efficient than calling translate() 10 times in a loop,
# because the model processes all sentences in parallel on the GPU.
# The result is a list of 10 French strings in the same order as the input.
# Example:
#   ["Je suis heureux.", "Elle aime les chats.", "Nous apprenons.", ...]
predicted_french = translate(english_sentences)

# ---------------------------------------------------------------
# Add the predictions as a new column in the DataFrame
# ---------------------------------------------------------------
# We store the model's translations back into sample_df so we can
# easily compare them with the actual French translations side by side.
# After this line, sample_df has three columns:
#   "english"          : the original English sentence
#   "french"           : the correct French translation (from the dataset)
#   "predicted_french" : the model's French translation
sample_df["predicted_french"] = predicted_french

# ---------------------------------------------------------------
# Display the results in a formatted table
# ---------------------------------------------------------------
# We print three columns side by side so we can visually compare:
#   - the original English sentence
#   - what the model predicted in French
#   - what the correct French translation actually is
#
# The :<45 format specifier left-aligns each value in a field of 45
# characters wide, so all three columns line up neatly regardless of
# the actual length of each sentence.
print(f"{'English':<45} {'Predicted French':<45} {'Actual French':<45}")
print("-" * 135)

# iterrows() loops through each row of the DataFrame one at a time.
# The underscore "_" is used for the row index, which we do not need here.
# For each row we print the three columns side by side.
# Example output line:
#   I am happy.                                  Je suis heureux.                             Je suis heureux.
for _, row in sample_df.iterrows():
    print(f"{row['english']:<45} {row['predicted_french']:<45} {row['french']:<45}")

[cell 31 markdown]
## Observations from Cell 8 -- Fine-Tuned Model Predictions vs Ground Truth

### Perfect Matches
- "I wish Tom was here" -> identical: "J'aimerais que Tom soit là"
- "The clock has stopped" -> identical: "L'horloge s'est arrêtée"

### Correct but Different Wording
- "I'm not scared to die"
  - Model : "Je n'ai pas peur de mourir" (not afraid to die)
  - Truth : "Je ne crains pas de mourir" (do not fear dying)
  - Both mean the same thing, just different verb choice

- "How did the audition go?"
  - Model : "Comment s'est déroulé l'audition?"
  - Truth : "Comment s'est passée l'audition?"
  - Both are natural French, dérouler and passer are interchangeable here

- "I really like this skirt. Can I try it on?"
  - Model : "J'aime vraiment cette jupe. Puis-je l'essayer?"
  - Truth : "J'aime beaucoup cette jupe, puis-je l'essayer?"
  - "vraiment" and "beaucoup" are both valid ways to say "really like"

### Formal vs Informal Register
- "Take a seat"
  - Model : "Assieds-toi" (informal tu form)
  - Truth : "Prends place!" (more formal)
  - The original English does not specify formality so both are correct

- "You'd better make sure that it is true"
  - Model : "Vous feriez mieux..." (formal vous form)
  - Truth : "Tu ferais bien..." (informal tu form)
  - Again, English is ambiguous here so both are valid

### Model Better than Ground Truth
- "I've no friend to talk to about my problems"
  - Model : "Je n'ai pas d'ami à qui parler de mes problèmes"
  - Truth : "Je n'ai pas d'ami avec lequel je puisse m'entretenir..."
  - The model's version is actually more natural everyday French
  - The ground truth uses an unnecessarily formal construction

### Genuine Error
- "Take any two cards you like"
  - Model : "Prends n'importe quelle carte que tu aimes"
  - Truth : "Prends deux cartes de ton choix"
  - The model dropped the word "two" which changed the meaning
  - This shows that models can miss specific numbers in longer sentences

### Key Takeaways
- Most differences are about register or word choice, not actual errors
- In some cases the model output is more natural than the ground truth
- BLEU score penalizes valid alternative translations, so our score
  likely understates the true quality of the model
- The only genuine error was dropping the word "two" in one sentence

[cell 32 markdown]
# BLEU score for the loaded dataset

[cell 33 markdown]
# BLEU Score Comparison: Pretrained vs Fine-Tuned Model

## What are we doing here?

So far we have:
1. Used the pretrained MarianMT model to translate (before fine-tuning)
2. Fine-tuned the model on our English-French dataset

Now we want to answer the key question:
**Did fine-tuning actually improve the translations?**

We answer this by computing the BLEU score TWICE:
- Once using the **original pretrained model** (reloaded from Hugging Face)
- Once using the **fine-tuned model** (saved in cell_5e)

Then we compare the two scores side by side.

## What should we expect?

With only 2,000 sentences and 3 epochs, the improvement may be small.
But even a small improvement confirms that fine-tuning is working correctly.

A large improvement would require:
- More data (tens of thousands of sentences)
- More epochs
- A dataset with a very specific style or vocabulary that differs
  from the original training data

## Important reminder about BLEU score limitations

As we saw in cell_8, BLEU score penalizes valid translations that use
different but correct wording. So the true quality improvement from
fine-tuning is likely BETTER than what the numbers alone suggest.

[cell 34 code]
# cell_9b  (updated)
# ============================================================
# BLEU SCORE COMPARISON: Pretrained vs Fine-Tuned Model
# ============================================================
# We compute BLEU scores for both models and compare them.
# This tells us whether fine-tuning improved translation quality.
#
# Steps:
#   1. Translate the VALIDATION SET using the FINE-TUNED model
#      (already loaded as `model` -- updated during cell_5d training)
#   2. Reload the ORIGINAL pretrained model from Hugging Face
#   3. Translate the VALIDATION SET using the ORIGINAL model
#   4. Compute BLEU scores for both and compare
#
# IMPORTANT: We use only the validation set (val_english, val_french)
# created in cell_5a. The fine-tuned model has NEVER seen these 40
# sentences during training, so the comparison is fair.

from tqdm import tqdm  # progress bar

BATCH_SIZE = 32  # same batch size as before for consistency

# ---------------------------------------------------------------
# Build a DataFrame from the validation set for easy handling
# ---------------------------------------------------------------
# val_english and val_french were created in cell_5a.
# We combine them into a DataFrame so we can slice batches easily,
# just like we did with df in the training cells.
val_df = pd.DataFrame({"english": val_english, "french": val_french}).reset_index(drop=True)

print(f"Evaluating on {len(val_df)} held-out validation sentences (never seen during training)")

# ===========================================================
# PART 1: BLEU score for the FINE-TUNED model
# ===========================================================
# The `model` variable currently holds our fine-tuned model
# (it was updated in-place during the training loop in cell_5d).
# We run batch translation on the validation set only.

print("\nComputing BLEU score for FINE-TUNED model...")
print("-" * 50)

# Make sure model is in evaluation mode (no dropout)
model.eval()

finetuned_predictions = []  # will hold all fine-tuned translations
total = len(val_df)
num_batches = (total + BATCH_SIZE - 1) // BATCH_SIZE

for i in tqdm(range(0, total, BATCH_SIZE), total=num_batches, desc="Fine-tuned model"):

    # Slice one batch of English sentences from the validation set
    batch = val_df["english"][i : i + BATCH_SIZE].tolist()

    # Translate using the fine-tuned model
    # translate() defined in cell_7 uses the global `model` variable
    # which is currently our fine-tuned model
    translated = translate(batch)
    finetuned_predictions.extend(translated)

# The references must be wrapped in a list because sacrebleu supports
# multiple reference translations per sentence. We only have one here.
all_references = [val_df["french"].tolist()]

bleu_finetuned = sacrebleu.corpus_bleu(
    finetuned_predictions,
    all_references,
    tokenize='intl'
)
print(f"Fine-Tuned Model BLEU Score : {bleu_finetuned.score:.2f} / 100")


# ===========================================================
# PART 2: BLEU score for the ORIGINAL pretrained model
# ===========================================================
# We reload the original pretrained model fresh from Hugging Face
# so that it has the ORIGINAL weights (before our fine-tuning).
# This gives us a fair baseline to compare against.
#
# We use a separate variable name `original_model` so we do NOT
# overwrite our fine-tuned `model` variable.

print(f"\nReloading original pretrained model for baseline comparison...")
print("-" * 50)

# Load the original pretrained model into a NEW variable
# This does NOT affect our fine-tuned `model` variable
original_model = MarianMTModel.from_pretrained(MODEL_NAME)
original_model = original_model.to(DEVICE)
original_model.eval()
print(f"Original pretrained model reloaded successfully")

# ---------------------------------------------------------------
# Define a temporary translate function for the original model
# ---------------------------------------------------------------
# Our existing translate() function in cell_7 uses the global
# `model` variable (fine-tuned). We define a separate function
# here that uses `original_model` instead, so we can fairly
# compare both models without any confusion.
def translate_original(texts):
    """
    Same as translate() in cell_7 but uses the ORIGINAL
    pretrained model instead of the fine-tuned one.
    """
    if isinstance(texts, str):
        texts = [texts]

    inputs = tokenizer(
        texts,
        return_tensors="pt",
        padding=True,
        truncation=True
    ).to(DEVICE)

    with torch.no_grad():
        translated_tokens = original_model.generate(**inputs)

    translations = tokenizer.batch_decode(
        translated_tokens,
        skip_special_tokens=True
    )
    return translations

# Translate validation set using the original pretrained model
print(f"\nComputing BLEU score for ORIGINAL pretrained model...")

original_predictions = []

for i in tqdm(range(0, total, BATCH_SIZE), total=num_batches, desc="Original model "):

    # Slice the same validation batches for a fair comparison
    batch = val_df["english"][i : i + BATCH_SIZE].tolist()
    translated = translate_original(batch)
    original_predictions.extend(translated)

# Compute BLEU score for original model
bleu_original = sacrebleu.corpus_bleu(
    original_predictions,
    all_references,
    tokenize='intl'
)
print(f"Original Pretrained Model BLEU Score : {bleu_original.score:.2f} / 100")


# ===========================================================
# PART 3: Side-by-side comparison
# ===========================================================

improvement = bleu_finetuned.score - bleu_original.score

print(f"\n{'='*55}")
print(f"{'BLEU SCORE COMPARISON SUMMARY':^55}")
print(f"{'='*55}")
print(f"{'Model':<35} {'BLEU Score':>10}")
print(f"{'-'*55}")
print(f"{'Original Pretrained Model':<35} {bleu_original.score:>10.2f}")
print(f"{'Fine-Tuned Model':<35} {bleu_finetuned.score:>10.2f}")
print(f"{'-'*55}")

# Show whether fine-tuning helped, hurt, or had no effect
if improvement > 0:
    print(f"{'Improvement after Fine-Tuning':<35} {improvement:>+10.2f}  ✓ Better")
elif improvement < 0:
    print(f"{'Change after Fine-Tuning':<35} {improvement:>+10.2f}  (see note below)")
else:
    print(f"{'Change after Fine-Tuning':<35} {improvement:>+10.2f}  No change")

print(f"{'='*55}")

# ---------------------------------------------------------------
# Helpful note if fine-tuning did not improve the score
# ---------------------------------------------------------------
if improvement <= 0:
    print(f"""
Note: Fine-tuning did not improve the BLEU score this time.
This is common when:
  - The dataset is small (we used only 200 sentences)
  - The number of epochs is low (we used 1 epoch)
  - The pretrained model was already trained on very similar data

To improve results, try:
  - Using more data (10,000+ sentence pairs)
  - Training for more epochs (5-10)
  - Reducing the learning rate slightly (e.g. 2e-5)
""")

# ===========================================================
# PART 4: Side-by-side translation examples
# ===========================================================
# Numbers alone don't tell the full story. Let's look at actual
# translation examples from both models side by side on 5 random
# validation sentences to visually judge quality differences.

print(f"\n--- Side-by-side translation examples (validation set only) ---\n")
print(f"{'English':<35} {'Original Model':<35} {'Fine-Tuned Model':<35} {'Ground Truth':<35}")
print("-" * 140)

# Sample 5 random sentences from the validation set
sample_df = val_df.sample(5, random_state=42).reset_index(drop=True)

for _, row in sample_df.iterrows():
    orig_translation      = translate_original(row["english"])[0]
    finetuned_translation = translate(row["english"])[0]
    print(
        f"{row['english']:<35} "
        f"{orig_translation:<35} "
        f"{finetuned_translation:<35} "
        f"{row['french']:<35}"
    )

[cell 35 markdown]
## Summary

In this notebook we built a complete English to French translation pipeline
using a pretrained transformer model and then fine-tuned it on our own dataset.
Here is a recap of everything we did and what we learned along the way.

### What we did

1. Loaded the dataset -- we worked with a real English-French translation
   dataset. We explored its structure, checked for missing values, duplicates,
   and understood the sentence length distribution.

2. Loaded a pretrained model -- we used Helsinki-NLP/opus-mt-en-fr, a MarianMT
   transformer model trained specifically for English to French translation.

3. Understood the tokenizer -- we saw how raw English text gets converted into
   numbers (token IDs) that the model can read, and how the output numbers get
   converted back into readable French text.

4. Translated sentences -- we translated a single sentence step by step first,
   then wrapped the logic into a reusable function, and finally translated the
   entire dataset in batches.

5. Evaluated with BLEU score -- we measured translation quality using the BLEU
   metric on the pretrained model and got a score of 53.68.

6. Fine-tuned the model -- we split our data into training and validation sets,
   created a custom PyTorch Dataset and DataLoader, and ran a training loop
   for 3 epochs to adapt the pretrained model to our specific dataset.

7. Compared results -- we computed BLEU scores for both the original pretrained
   model and our fine-tuned model and compared them side by side.

### What we learned

- Pretrained transformers handle most text processing automatically through
  the tokenizer. For clean datasets like this one, no manual preprocessing
  is needed.

- Fine-tuning on even a small dataset can
  meaningfully improve translation quality.

- BLEU score has limitations. It penalizes valid translations that use different
  but correct wording. Looking at actual translations in cell_8, many differences
  were about register or word choice rather than genuine errors.

- In some cases the fine-tuned model produced more natural French than the
  ground truth itself, which further shows that BLEU score alone does not
  tell the full story.

### Key results

| Model                   | BLEU Score |
|-------------------------|------------|
| LSTM with Attention     | 21.17      |
| MarianMT (pretrained)   | 53.72      |
| MarianMT (fine-tuned)   | 57.60     |

A few things stand out from this table:

- MarianMT pretrained already achieves more than double the BLEU score of an
  LSTM with attention, without any training on our part.

- Fine-tuning pushed the score further from 53.72 to 57.60.

- This shows that pretrained transformers are a strong starting point, and
  fine-tuning on your own data can make them even better with relatively
  little effort.