# EngToFre usingTransformer GC v3 PDF
course: Module 4 — Generative AI & LLMs
module: Module-4-Generative-AI-LLMs
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-4-Generative-AI-LLMs/General/Lab_Materials_-16-05-2026/EngToFre_usingTransformer_GC_v3_PDF.pdf
pages: 51
---
[page 1]
Dataset Brief: Language Translation
(English–French)
Source
Kaggle — Language Translation (English-French) by Devicharith
What is this dataset?
It is a bilingual sentence pair dataset — for every English phrase, there is a matching
French translation sitting right next to it. Think of it as a giant phrasebook stored in a
spreadsheet.
Structure
Column Description Example
English Source sentence "Go."
French Translated sentence "Va !"
File format: CSV (eng_-french.csv)
Size: ~175,000+ sentence pairs (we'll use a subset of ~1000–5000 for training)
Range: From single words like "Hi." → "Salut !" to full multi-word sentences
Note: For demonstration purpose, we have selected only 200 rows of data.
# cell_0a — GPU masking BEFORE any torch import
# =================================================
# MAIN GOAL:
# Decide which physical GPU (if any) this notebook is allowed to see, and
# hide all other GPUs from the process. This MUST happen before `import
torch`
# because torch reads CUDA_VISIBLE_DEVICES exactly once at import time —
# changing it later has no effect.
#
# WHY THIS MATTERS:
# On a shared machine with multiple GPUs, you don't want this notebook to
# accidentally allocate memory on a GPU someone else is using. Setting
# CUDA_VISIBLE_DEVICES="2" (for example) makes physical GPU 2 the *only*
[page 2]
# GPU torch can see — and inside this process it will be re-numbered as
# cuda:0 (because it's the only visible device).
#
# FLOW:
# 1. Ask the user for a GPU index (e.g. "0", "1", "2") or "cpu".
# 2. Validate the input is one of those forms — if not, STOP loudly instead
# of silently falling back (silent fallback to CPU could waste hours).
# 3. Set the env var and remember the logical device string for later
cells.
import os # used to read/write process environment variables
# Prompt the user. .strip() removes accidental spaces, .lower() normalizes
case.
choice = input(
"Enter GPU physical index to use for cuda (e.g., 0, 1, 2) or type 'cpu'
for CPU mode: "
).strip().lower()
if choice == "cpu":
# Empty string for CUDA_VISIBLE_DEVICES = no GPUs visible to this
process.
os.environ["CUDA_VISIBLE_DEVICES"] = ""
DEVICE = "cpu"
print("CPU mode selected — no GPU will be visible to this process.")
elif choice.isdigit():
# User gave a number like "0" or "2". We mask every GPU except this one.
phys_idx = int(choice)
os.environ["CUDA_VISIBLE_DEVICES"] = str(phys_idx)
# Inside the process, the selected physical GPU is always re-indexed as
0,
# because it is the *only* visible CUDA device.
DEVICE = "cuda:0"
print(
f"Masking other GPUs. Physical GPU {phys_idx} is now visible "
f"and will appear inside this process as cuda:0."
)
else:
# Anything else (typos like 'gpu0', 'cuda', empty string) -> raise
loudly.
# We deliberately do NOT silently fall back to CPU, because that can lead
# to hours of slow training before the user notices.
raise ValueError(
f"Invalid choice {choice!r}. Please re-run this cell and enter "
f"either a digit (e.g. '0') or 'cpu'."
)
print(f"\n>>> DEVICE (logical, used in later cells): {DEVICE}")
print(f">>> CUDA_VISIBLE_DEVICES (env):
{os.environ.get('CUDA_VISIBLE_DEVICES')!r}")
[page 3]
Masking other GPUs. Physical GPU 2 is now visible and will appear inside this
process as cuda:0.
>>> DEVICE (logical, used in later cells): cuda:0
>>> CUDA_VISIBLE_DEVICES (env): '2'
Installing required Libraries
# cell_0
# ============================================================
# STEP 1: Install all the required libraries for this tutorial
# ============================================================
# Before we can build our translator, we need to install some Python
libraries.
# Think of libraries as "toolboxes". Each one gives us ready-made tools so we
# don't have to build everything from scratch.
# transformers:
# Hugging Face's library that lets us download and use pretrained AI
models.
# We'll use it to load "MarianMT", a model already trained to translate
# between languages, so we don't have to train one from zero (which would
# take days and a lot of computing power).
# sentencepiece:
# A helper tool that breaks sentences into smaller pieces called
"subwords".
# Example: "unhappiness" becomes ["un", "happi", "ness"]
# This is important because models understand small chunks better than
# whole words, especially for rare or new words.
# sacrebleu:
# A tool to measure how good our translations are.
# It compares our model's output to a correct human translation and gives
# a score called "BLEU" (higher means better translation).
# Example: If the model says "Bonjour le monde" and the correct answer is
# "Bonjour monde", sacrebleu tells us how close they are.
# sacremoses:
# A text cleaning helper. It properly handles punctuation, spaces, and
# capitalization when converting tokens back into readable sentences.
# Example: ["Hello", ",", "world", "!"] becomes "Hello, world!"
# pandas:
# A library to load and work with tables of data (like Excel in Python).
# We'll use it to read our CSV file with English-French sentence pairs.
# The "!" at the start tells Jupyter Notebook: "Run this as a terminal
command,
# not as Python code." "pip install" is how we install Python libraries.
[page 4]
Requirement already satisfied: transformers in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (4.56.1)
Requirement already satisfied: sentencepiece in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (0.2.1)
Requirement already satisfied: sacrebleu in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (2.6.0)
Requirement already satisfied: sacremoses in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (0.1.1)
Requirement already satisfied: pandas in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (2.3.2)
Requirement already satisfied: filelock in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from transformers) (3.19.1)
Requirement already satisfied: huggingface-hub<1.0,>=0.34.0 in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from transformers) (0.34.4)
Requirement already satisfied: numpy>=1.17 in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from transformers) (1.26.4)
Requirement already satisfied: packaging>=20.0 in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from transformers) (25.0)
Requirement already satisfied: pyyaml>=5.1 in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from transformers) (6.0.2)
Requirement already satisfied: regex!=2019.12.17 in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from transformers) (2025.9.1)
Requirement already satisfied: requests in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from transformers) (2.32.5)
Requirement already satisfied: tokenizers<=0.23.0,>=0.22.0 in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from transformers) (0.22.0)
Requirement already satisfied: safetensors>=0.4.3 in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from transformers) (0.6.2)
Requirement already satisfied: tqdm>=4.27 in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from transformers) (4.67.1)
Requirement already satisfied: fsspec>=2023.5.0 in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from huggingface-hub<1.0,>=0.34.0->transformers) (2025.3.0)
Requirement already satisfied: typing-extensions>=3.7.4.3 in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
!pip install transformers sentencepiece sacrebleu sacremoses pandas
[page 5]
packages (from huggingface-hub<1.0,>=0.34.0->transformers) (4.15.0)
Requirement already satisfied: hf-xet<2.0.0,>=1.1.3 in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from huggingface-hub<1.0,>=0.34.0->transformers) (1.1.9)
Requirement already satisfied: portalocker in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from sacrebleu) (3.2.0)
Requirement already satisfied: tabulate>=0.8.9 in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from sacrebleu) (0.10.0)
Requirement already satisfied: colorama in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from sacrebleu) (0.4.6)
Requirement already satisfied: lxml in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from sacrebleu) (6.1.0)
Requirement already satisfied: click in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from sacremoses) (8.2.1)
Requirement already satisfied: joblib in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from sacremoses) (1.5.2)
Requirement already satisfied: python-dateutil>=2.8.2 in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from pandas) (2.9.0.post0)
Requirement already satisfied: pytz>=2020.1 in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from pandas) (2025.2)
Requirement already satisfied: tzdata>=2022.7 in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from pandas) (2025.2)
Requirement already satisfied: six>=1.5 in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from python-dateutil>=2.8.2->pandas) (1.17.0)
Requirement already satisfied: charset_normalizer<4,>=2 in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from requests->transformers) (3.4.3)
Requirement already satisfied: idna<4,>=2.5 in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from requests->transformers) (3.10)
Requirement already satisfied: urllib3<3,>=1.21.1 in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from requests->transformers) (2.5.0)
Requirement already satisfied: certifi>=2017.4.17 in
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages (from requests->transformers) (2025.8.3)
Imports
# cell_1
# ============================================================
[page 6]
# STEP 2: Import all the libraries we will use throughout this tutorial
# ============================================================
# "Importing" a library means loading it into our current Python session
# so we can use its tools. Think of it like opening a toolbox before
# starting work. We only need to do this once at the beginning.
import os # provides tools to work with file and
folder paths
# Example: checking if a file exists,
joining paths
import pandas as pd # loads and manipulates tabular data (like a
spreadsheet)
# "as pd" means we can write "pd" instead of
"pandas"
# Example: pd.read_csv("data.csv") loads a
CSV file
import torch # PyTorch, the deep learning framework
# we use it to check if a GPU is available
# and to handle "tensors" (the number arrays
that
# the model reads and produces)
from transformers import (
MarianMTModel, # the pretrained model for translation
# "MarianMT" is a family of models trained
by
# the University of Helsinki on millions of
# sentence pairs. We load the English-
>French version.
MarianTokenizer # the tokenizer that pairs with
MarianMTModel.
# A tokenizer converts raw text into numbers
# (which the model understands) and back to
text.
# Example: "Hello" -> [4, 156, 23] ->
"Bonjour"
)
import sacrebleu # used later to evaluate translation quality
# with a standard metric called BLEU score
# (Bilingual Evaluation Understudy)
# ---------------------------------------------------------------
# Confirm setup: print PyTorch version and the device being used
# ---------------------------------------------------------------
# This is a quick sanity check to make sure everything imported correctly.
# torch.__version__ tells us which version of PyTorch is installed.
print(f"PyTorch version : {torch.__version__}")
# DEVICE tells us whether we are using a GPU (faster) or CPU (slower).
[page 7]
PyTorch version : 2.6.0+cu118
Device selected : cuda:0
Data Loading
Dataset shape : (200, 2)
Columns : ['english', 'french']
First 3 rows:
english french
0 Hi. Salut!
1 Run! Cours !
2 Run! Courez !
Basic EDA
# It was defined in the previous cell (cell_0a).
# GPU example output -> "Device selected : cuda"
# CPU example output -> "Device selected : cpu"
print(f"Device selected : {DEVICE}") # DEVICE was set in cell_0a
# cell_2
# Goal: Load the English-French dataset from the local path into a pandas
DataFrame
# Full path to the dataset CSV file
DATA_PATH ="/content/eng_-french.csv"
# Load the CSV file into a DataFrame
#df = pd.read_csv(DATA_PATH)
df = pd.read_csv(DATA_PATH, nrows=200) # load only the first 200 rows
# Rename columns to simpler names for easier access throughout the notebook
df.columns = ["english", "french"]
# Confirm the dataset loaded correctly
print(f"Dataset shape : {df.shape}")
print(f"Columns : {list(df.columns)}")
print(f"\nFirst 3 rows:")
print(df.head(3))
# cell_3
# ============================================================
# STEP 4: Explore the dataset to understand its structure and content
# ============================================================
# Before training or using any model, it is good practice to explore
# the data. This helps us catch problems early (like missing values
# or duplicates) and understand what kind of sentences we are working with.
[page 8]
# ---------------------------------------------------------------
# Basic info: how many rows and columns do we have?
# ---------------------------------------------------------------
# df.shape[0] gives the number of rows (sentence pairs).
# df.shape[1] gives the number of columns (should be 2: english, french).
print(f"Total rows : {df.shape[0]}")
print(f"Total columns : {df.shape[1]}")
# ---------------------------------------------------------------
# Check for missing values
# ---------------------------------------------------------------
# A missing value means a cell in the table is empty (no text at all).
# df.isnull() returns True for every empty cell, and .sum() counts them
# per column. Ideally both columns should show 0.
# Example output:
# english 0
# french 0
# If we see a non-zero number, we would need to remove those rows
# before feeding the data to our model.
print(f"\nMissing values per column:")
print(df.isnull().sum())
# ---------------------------------------------------------------
# Check for duplicate rows
# ---------------------------------------------------------------
# A duplicate row means the exact same English-French pair appears
# more than once. Duplicates can make the model over-learn certain
# sentences, so it is good to know how many there are.
# df.duplicated() marks each repeated row as True, and .sum() counts them.
# Example output: "Duplicate rows: 32"
print(f"\nDuplicate rows: {df.duplicated().sum()}")
# ---------------------------------------------------------------
# Sample 5 random English-French sentence pairs
# ---------------------------------------------------------------
# df.sample(5) picks 5 random rows from the dataset so we can visually
# inspect what the sentence pairs look like.
# random_state=42 is a fixed "seed" so we get the same 5 rows every
# time we run this cell (makes results reproducible).
# .to_string(index=False) prints the table without the row numbers.
# Example output:
# english french
# I am tired. Je suis fatigue.
# She likes cats. Elle aime les chats.
print(f"\n5 random sentence pairs:")
print(df.sample(5, random_state=42).to_string(index=False))
# ---------------------------------------------------------------
# Calculate sentence length (in words) for each row
# ---------------------------------------------------------------
[page 9]
Total rows : 200
Total columns : 2
Missing values per column:
english 0
french 0
dtype: int64
Duplicate rows: 0
5 random sentence pairs:
english french
Call us. Appelle-nous !
Go on. Poursuivez.
Get up. Lève-toi.
I swore. J’ai promis.
Go home. Rentrez chez vous.
English sentence length (words):
# We want to know how long the sentences are on average.
# This helps us choose a good "max_length" when tokenizing later.
#
# How it works:
# str(x).split() splits a sentence into a list of words by whitespace.
# Example: "I love Paris".split() -> ["I", "love", "Paris"]
# len(...) then counts how many words are in that list.
# lambda x: ... is a small one-line function applied to every row.
#
# We store the word counts in two new columns: "english_len" and
"french_len".
df["english_len"] = df["english"].apply(lambda x: len(str(x).split()))
df["french_len"] = df["french"].apply(lambda x: len(str(x).split()))
# ---------------------------------------------------------------
# Print summary statistics for sentence lengths
# ---------------------------------------------------------------
# .describe() gives us useful statistics for a numeric column:
# count - how many rows were counted
# mean - average sentence length
# std - how much lengths vary (standard deviation)
# min - shortest sentence (in words)
# 25% - 25% of sentences are shorter than this length
# 50% - median sentence length
# 75% - 75% of sentences are shorter than this length
# max - longest sentence (in words)
# .round(2) rounds all numbers to 2 decimal places for cleaner output.
print(f"\nEnglish sentence length (words):")
print(df["english_len"].describe().round(2))
print(f"\nFrench sentence length (words):")
print(df["french_len"].describe().round(2))
[page 10]
count 200.00
mean 1.85
std 0.36
min 1.00
25% 2.00
50% 2.00
75% 2.00
max 2.00
Name: english_len, dtype: float64
French sentence length (words):
count 200.00
mean 2.57
std 0.97
min 1.00
25% 2.00
50% 2.00
75% 3.00
max 6.00
Name: french_len, dtype: float64
The Model We Are Using
One model, not two
Before we begin, one important clarification. You will see two names used throughout
this notebook -- MarianMT and Helsinki-NLP/opus-mt-en-fr. These are not two separate
models. They refer to the same thing from two different angles.
MarianMT is the architecture -- the design and structure of the transformer. Think of
it as the blueprint.
Helsinki-NLP/opus-mt-en-fr is the trained model -- a specific instance of the
MarianMT architecture that has been trained on millions of English-French sentence
pairs by the Language Technology Research Group at the University of Helsinki.
Think of it as the finished product built from that blueprint.
So when we write:
model = MarianMTModel.from_pretrained("Helsinki-NLP/opus-mt-en-fr")
We are saying -- use the MarianMT architecture and load the weights that Helsinki-NLP
trained specifically for English to French translation.
What is MarianMT?
MarianMT is a transformer based encoder-decoder architecture designed specifically for
machine translation. Unlike large general purpose language models, MarianMT is a
focused, lightweight model that does one thing well -- translating text from one language
to another.
It works in two stages:
[page 11]
The encoder reads the English sentence and builds a rich understanding of its
meaning
The decoder takes that understanding and generates the French translation word by
word
How was Helsinki-NLP/opus-mt-en-fr trained?
The model was trained on the OPUS dataset, which is a large collection of translated
texts gathered from the web, including subtitles, books, news articles, and more. It was
trained on millions of English-French sentence pairs, which is why it performs so well
even on sentences it has never seen before.
Why are we using it?
It is small (75 million parameters) and fast -- works well even on CPU
It is freely available on Hugging Face with no authentication required
It is purpose built for English to French translation
It is easy to use with just a few lines of code through the Hugging Face transformers
library
Despite its small size, it achieves a BLEU score of 51.16 on our dataset
Loading Tokenizer
# cell_4
# ============================================================
# STEP 5: Load the pretrained tokenizer for the MarianMT model
# ============================================================
# Before the model can translate text, it needs the text to be converted
# into numbers. This is what a "tokenizer" does.
#
# The process looks like this:
# "Hello, how are you?"
# |
# v (tokenizer encodes)
# [10537, 2, 541, 52, 55, 54, 0] <-- numbers the model reads
# |
# v (model translates)
# [34, 78, 120, 56, 0] <-- numbers the model outputs
# |
# v (tokenizer decodes)
# "Bonjour, comment allez-vous?" <-- readable French text
#
# MarianTokenizer is the specific tokenizer designed to work with
# the MarianMT family of translation models.
# ---------------------------------------------------------------
# Define the model name
# ---------------------------------------------------------------
# This is the official name of the pretrained model on Hugging Face.
[page 12]
# "Helsinki-NLP" is the research group that trained it.
# "opus-mt-en-fr" means: trained on OPUS data, MarianMT, English to French.
# We define it as a variable here so we can reuse the same name in cell_5
# when we load the model itself (both the tokenizer and the model must
match).
MODEL_NAME = "Helsinki-NLP/opus-mt-en-fr"
# ---------------------------------------------------------------
# Download and load the tokenizer
# ---------------------------------------------------------------
# from_pretrained() downloads the tokenizer files from Hugging Face
# on the first run (about 1-2 MB) and caches them locally.
# On subsequent runs it loads from the local cache (no internet needed).
tokenizer = MarianTokenizer.from_pretrained(MODEL_NAME)
# ---------------------------------------------------------------
# Quick sanity check: tokenize a sample sentence
# ---------------------------------------------------------------
# We run a simple English sentence through the tokenizer to see
# what it produces, just to confirm everything is working correctly.
sample = "Hello, how are you?"
# tokenizer.tokenize() splits the sentence into "subword pieces".
# Instead of splitting by whole words, it uses smaller units.
# This helps handle rare or unknown words.
# Example: "unhappiness" might become ["un", "happi", "ness"]
# The " ▁ " symbol marks the beginning of a new word (comes from
SentencePiece).
# Example output: [' ▁ Hello', ',', ' ▁ how', ' ▁ are', ' ▁ you', '?']
# Note: tokenize() does NOT add any special tokens, it just splits text.
tokens = tokenizer.tokenize(sample)
# tokenizer.encode() does two things in one step:
# 1. Splits the text into subword pieces (same as tokenize())
# 2. Converts each piece into its integer ID from the vocabulary
# and appends a special EOS (End-Of-Sentence) token at the end.
# EOS token (ID = 0) tells the model "this is where the input ends".
# Example output: [10537, 2, 541, 52, 55, 54, 0]
# ^ this 0 is the EOS token
# Notice encode() returns 7 numbers while tokenize() returned 6 pieces,
# because encode() adds the extra EOS token at the end.
token_ids = tokenizer.encode(sample)
# ---------------------------------------------------------------
# Print the results of the sanity check
# ---------------------------------------------------------------
print(f"Sample sentence : {sample}")
# The subword pieces the tokenizer split the sentence into.
# Example output: [' ▁ Hello', ',', ' ▁ how', ' ▁ are', ' ▁ you', '?']
print(f"Tokens : {tokens}")
[page 13]
Sample sentence : Hello, how are you?
Tokens : [' ▁ Hello', ',', ' ▁ how', ' ▁ are', ' ▁ you', '?']
Token IDs : [10537, 2, 541, 52, 55, 54, 0]
Vocabulary size : 59514
Do we need to preprocess or clean the text?
Since we are using a pretrained transformer (MarianMT), most of the heavy text
processing is handled automatically by the tokenizer. Here is a clear breakdown:
What the MarianMT tokenizer handles for you
Punctuation splitting -- punctuation is separated from words automatically. We saw
this in cell_4 where "Hello, how are you?" was split into: [' ▁ Hello', ',', ' ▁ how',
' ▁ are', ' ▁ you', '?'] Notice how the comma is a separate token.
Subword tokenization -- words are broken into smaller pieces called subwords. We
saw this in cell_4 where the ▁ prefix marks the start of each new word. This allows
the tokenizer to handle words it has never seen before.
Special tokens -- the tokenizer automatically adds an end of sentence marker. We
saw this in cell_4 where the last token ID was 0, which is the end of sentence token.
This tells the model where the input sentence ends.
Text to numbers -- the tokenizer converts words into numbers the model can read.
We saw this in cell_4 where "Hello" became 10537, "how" became 541, and so on.
The model only works with these numbers, never with raw text directly.
What you still need to be careful about
Missing values -- passing an empty or null value to the tokenizer will crash. If a row
in your dataset has no text, tokenizing it will throw an error. We already confirmed
there are no missing values in cell_3, so we are safe.
Extremely long sentences -- MarianMT can only handle up to 512 tokens at a time.
Sentences longer than this get cut off and lose information at the end. Our longest
sentence is only 44 words, so we are well within limits.
Encoding issues -- unusual or special characters can sometimes confuse the
tokenizer. This is common in corrupted CSV files or text copied from PDFs or web
pages. Our dataset contains clean, well-formed text so this is not a concern here.
# The integer IDs corresponding to each subword piece, plus EOS at the end.
# Example output: [10537, 2, 541, 52, 55, 54, 0]
print(f"Token IDs : {token_ids}")
# The total number of unique tokens this tokenizer knows about.
# This is the size of its "vocabulary" (dictionary of all subword pieces).
# Example output: "Vocabulary size : 65001"
print(f"Vocabulary size : {tokenizer.vocab_size}")
[page 14]
Bottom line
For this specific dataset and this specific model, we can skip any preprocessing or
cleaning step entirely. The dataset is clean and the sentences are short and well-formed.
Note: If you were working with a raw, noisy dataset (for example, social media text, OCR
output, or user reviews), a cleaning step would be necessary before tokenization.
Loading Model
# cell_5
# ============================================================
# STEP 6: Load the pretrained MarianMT model
# ============================================================
# Now that we have the tokenizer (cell_4), we load the actual translation
# model. The tokenizer and model always come as a pair:
# - tokenizer : converts text <-> numbers
# - model : takes the numbers, does the translation
#
# Think of the tokenizer as a "translator's dictionary" and the model
# as the "translator's brain". Both must be from the same pretrained
# checkpoint to work correctly together.
# ---------------------------------------------------------------
# Download and load the model weights
# ---------------------------------------------------------------
# from_pretrained() downloads the model from Hugging Face on the first
# run (about 300 MB) and caches it locally for future runs.
# "Weights" are the millions of numbers inside the model that were learned
# during training on millions of English-French sentence pairs.
# We are loading those already-learned weights, so no training is needed.
# We reuse MODEL_NAME from cell_4 ("Helsinki-NLP/opus-mt-en-fr") to make
# sure the model and tokenizer are perfectly matched.
model = MarianMTModel.from_pretrained(MODEL_NAME)
# ---------------------------------------------------------------
# Move the model to the correct device (GPU or CPU)
# ---------------------------------------------------------------
# Deep learning models run much faster on a GPU than on a CPU.
# .to(DEVICE) moves all the model weights to the device we selected
# in cell_0a. If a GPU is available, DEVICE = "cuda", otherwise "cpu".
# It is important that both the model and the input data are on the
# same device, otherwise PyTorch will throw an error.
model = model.to(DEVICE)
# ---------------------------------------------------------------
# Set the model to evaluation mode
# ---------------------------------------------------------------
# A neural network behaves differently during training vs inference:
[page 15]
Model loaded successfully
Model is running on : cuda:0
Number of parameters: 75,133,952
Fine-Tuning the Pretrained Model on Our
Dataset
What is Fine-Tuning?
So far we have used the MarianMT model exactly as it was downloaded from Hugging
Face. This is called inference -- we are just asking the model to translate without
teaching it anything new.
Fine-tuning means we take that already-trained model and train it a little more, but this
time specifically on OUR dataset. Think of it like this:
#
# Training mode : some neurons are randomly "dropped out" to prevent
# the model from memorizing data (called "dropout").
# Batch normalization uses running statistics.
#
# Evaluation mode: dropout is turned off, all neurons are active.
# The model gives consistent, deterministic outputs.
#
# Since we are only using this model to translate (not to train it),
# we must call model.eval() to switch it to evaluation mode.
# If we forget this step, translations may randomly vary between runs.
model.eval()
# ---------------------------------------------------------------
# Quick sanity check: confirm the model loaded correctly
# ---------------------------------------------------------------
print(f"Model loaded successfully")
# next(model.parameters()) gets the first set of weights in the model.
# .device tells us which device those weights are stored on.
# Expected output: "cuda:0" (if GPU) or "cpu" (if no GPU available).
print(f"Model is running on : {next(model.parameters()).device}")
# This counts the total number of individual numbers (parameters) in the
model.
# p.numel() returns the count of numbers in one layer's weight matrix.
# We sum these across all layers to get the total.
# The {:,} format adds commas for readability.
# Example output: "Number of parameters: 77,376,768"
# That is about 77 million numbers that were learned during training!
print(f"Number of parameters: {sum(p.numel() for p in
model.parameters()):,}")
[page 16]
The pretrained model went to university and learned to translate millions of
sentences.
Fine-tuning is like giving that graduate a short internship focused specifically on our
style of sentences.
After fine-tuning, the model should perform better on sentences similar to those in our
dataset.
Step 1 — Split the Data into Training and Validation
Sets
Before we can fine-tune, we need to divide our dataset into two parts:
Training set (80%) -- the sentences the model will learn from.
Validation set (20%) -- sentences the model will NEVER see during training. We
use these at the end of each epoch to check if the model is actually improving or just
memorizing the training data.
Think of it like studying for an exam:
Training set = your study material
Validation set = a practice test with questions you haven't seen before
# cell_5a
# ============================================================
# FINE-TUNING STEP 1: Split data into Train and Validation sets
# ============================================================
# Before we can fine-tune the model, we need to divide our dataset
# into two parts:
# - Training set : the model LEARNS from these sentences
# - Validation set : we TEST the model on these after each epoch
# to check if it is improving
#
# We use sklearn's train_test_split() for this.
# It randomly shuffles and splits the data for us automatically.
from sklearn.model_selection import train_test_split
# ---------------------------------------------------------------
# Define the split ratio
# ---------------------------------------------------------------
# TEST_SIZE = 0.2 means:
# 20% of the data goes to validation
# 80% of the data goes to training
#
# Example with 2000 rows:
# Training set -> 1600 rows (80%)
# Validation set -> 400 rows (20%)
TEST_SIZE = 0.2
# ---------------------------------------------------------------
[page 17]
Total sentences : 200
Training sentences : 160 (80%)
Validation sentences : 40 (20%)
--- Sample training pairs (English -> French) ---
EN: Be fair.
FR: Sois juste !
# Perform the split
# ---------------------------------------------------------------
# train_test_split() takes two lists (English and French sentences)
# and splits both at the same time, keeping pairs matched correctly.
#
# Arguments:
# df["english"].tolist() : all English sentences as a Python list
# df["french"].tolist() : all French sentences as a Python list
# test_size=TEST_SIZE : fraction to use for validation (0.2 = 20%)
# random_state=42 : fixes the random shuffle so we get the
# same split every time we run this cell.
# Without this, the split would be different
# each run, making results hard to compare.
train_english, val_english, train_french, val_french = train_test_split(
df["english"].tolist(), # source sentences (input)
df["french"].tolist(), # target sentences (what we want the model to
output)
test_size=TEST_SIZE,
random_state=42
)
# ---------------------------------------------------------------
# Confirm the split worked correctly
# ---------------------------------------------------------------
# We print the sizes to double-check the split looks right.
# Expected output for 2000 rows:
# Total sentences : 2000
# Training sentences : 1600
# Validation sentences: 400
print(f"Total sentences : {len(df)}")
print(f"Training sentences : {len(train_english)} ({100 -
int(TEST_SIZE*100)}%)")
print(f"Validation sentences : {len(val_english)} ({int(TEST_SIZE*100)}%)")
# ---------------------------------------------------------------
# Peek at a few training pairs to confirm they are still matched
# ---------------------------------------------------------------
# It is important that English[i] and French[i] still correspond
# to each other after the split. Let's visually check 3 pairs.
print(f"\n--- Sample training pairs (English -> French) ---")
for i in range(3):
print(f" EN: {train_english[i]}")
print(f" FR: {train_french[i]}")
print()
[page 18]
EN: Shut up!
FR: La ferme !
EN: Got it?
FR: T'as capté ?
Step 2 — Create a Custom PyTorch
Dataset Class
What is a Dataset class?
PyTorch needs data to be served to the model in a very specific format. A Dataset class is
like a smart container that:
Holds all our sentence pairs (English + French)
Knows how many pairs it has
Can hand out one pair at a time when asked
Think of it like a deck of cards:
The deck holds all the cards (sentences)
You can ask "how many cards are in the deck?" (length)
You can ask "give me card number 5" (get one item)
PyTorch will use this class later to automatically feed batches of sentences to the model
during training.
What is Tokenization of the Target (French)?
In cell_4 we only tokenized the English (input) side. Now we also need to tokenize the
French (target/output) side.
The French token IDs become the labels -- the correct answers the model is trying to
learn to produce.
One special thing: we replace the padding token ID with -100 in the labels. This tells
PyTorch "ignore this position when calculating the loss". We do not want the model to be
penalized for not predicting padding tokens -- those are not real words, just fillers to
make sentence lengths match.
# cell_5b
# ============================================================
# FINE-TUNING STEP 2: Create a Custom PyTorch Dataset Class
# ============================================================
# PyTorch requires data to be wrapped in a Dataset class.
# This class acts as a smart container for our sentence pairs.
[page 19]
# It must implement exactly THREE methods:
#
# __init__() : called once when we create the dataset object.
# This is where we store and tokenize all the data.
#
# __len__() : returns the total number of sentence pairs.
# PyTorch calls this to know how many items exist.
#
# __getitem__(): returns ONE tokenized sentence pair by index.
# PyTorch calls this repeatedly to build batches.
#
# Example:
# dataset = TranslationDataset(train_english, train_french, tokenizer)
# len(dataset) -> 1600 (total pairs)
# dataset[0] -> {"input_ids": ..., "labels": ...} (first pair)
from torch.utils.data import Dataset # base class we inherit from; enforces
the 3-method contract
class TranslationDataset(Dataset):
"""
A custom Dataset that holds English-French sentence pairs
and tokenizes them so PyTorch can feed them to the model.
"""
def __init__(self, english_sentences, french_sentences, tokenizer,
max_length=128):
"""
Called once when we create the dataset.
Tokenizes ALL sentences upfront and stores them.
Args:
english_sentences : list of English strings (the inputs)
french_sentences : list of French strings (the targets)
tokenizer : the MarianTokenizer from cell_4
max_length : maximum number of tokens per sentence.
Sentences longer than this get cut off.
128 is safe for our dataset (max was 44
words).
"""
self.tokenizer = tokenizer # save tokenizer so __getitem__ can
access it later if needed
self.max_length = max_length # save max_length for reference
# ---------------------------------------------------------------
# Tokenize the English sentences (inputs)
# ---------------------------------------------------------------
# tokenizer() converts a list of raw strings into tensors of numbers.
# Think of it as: ["Hello", "I am"] -> [[34, 56, 1, 1], [78, 99, 12,
1]]
#
[page 20]
# padding="max_length" : every sentence is padded to exactly 128
tokens.
# e.g. "Hi" -> [416, 1, 1, 1, ..., 1] (126
padding tokens added)
# This is required because PyTorch batches
must have
# uniform shape -- all rows must be the same
length.
#
# truncation=True : if a sentence exceeds 128 tokens, cut it
off.
# Prevents crashes on unexpectedly long
inputs.
#
# return_tensors="pt" : return PyTorch tensors (not plain lists or
numpy arrays).
# Required so tensors can be directly passed
to the model.
#
# Result: self.inputs is a dict with two keys:
# "input_ids" : shape (N, 128) — one row per sentence, each
row is 128 token IDs
# "attention_mask" : shape (N, 128) — 1 where there's a real token,
0 where padding
self.inputs = tokenizer(
english_sentences,
padding="max_length",
truncation=True,
max_length=max_length,
return_tensors="pt"
)
# ---------------------------------------------------------------
# Tokenize the French sentences (targets / labels)
# ---------------------------------------------------------------
# Same tokenization settings as English for consistency.
# These French token IDs are the "correct answers" the model must
learn
# to predict given an English input.
#
# Result: self.targets is a dict with the same two keys:
# "input_ids" : shape (N, 128) — French token IDs
# "attention_mask" : shape (N, 128) — 1 for real tokens, 0 for
padding
self.targets = tokenizer(
french_sentences,
padding="max_length",
truncation=True,
max_length=max_length,
return_tensors="pt"
)
# ---------------------------------------------------------------
# Replace padding token IDs in labels with -100
[page 21]
# ---------------------------------------------------------------
# During training, PyTorch computes "loss" — how wrong the model's
# predictions are compared to the correct French tokens (labels).
#
# Problem: padding tokens (value = tokenizer.pad_token_id, typically
1)
# are NOT real words. We don't want the model penalized for getting
# padding positions wrong — those positions are meaningless fillers.
#
# Solution: PyTorch's loss function (CrossEntropyLoss) has a special
# convention: any label position with value -100 is completely
ignored
# in the loss calculation.
#
# So we replace every pad_token_id in the labels with -100.
#
# Concrete example (max_length=7):
# French sentence "Je t'aime" tokenizes to: [34, 78, 120, 56, 1,
1, 1 ]
# real tokens ^ ^
padding (ID=1)
# After masking: [34, 78, 120, 56, -100,
-100, -100 ]
# Loss is now only computed on positions 0-3, not 4-6.
self.labels = self.targets["input_ids"].clone()
# .clone() makes a full independent copy of the tensor.
# Without clone(), modifying self.labels would also modify
self.targets["input_ids"]
# because they'd point to the same memory.
# .eq(pad_token_id) -> boolean tensor: True where value ==
pad_token_id
# .masked_fill_(mask, -100) -> in-place: replace True positions with
-100
# The trailing underscore _ means "in-place" operation (modifies
self.labels directly)
self.labels = self.labels.masked_fill(
self.labels == tokenizer.pad_token_id, -100
)
def __len__(self):
"""
Returns the total number of sentence pairs in this dataset.
PyTorch calls this internally (e.g. to decide how many batches to
create).
Example:
len(train_dataset) -> 1600
"""
# .shape[0] gives the first dimension of the tensor, which equals the
number of sentences.
# e.g. if input_ids is shape (1600, 128), shape[0] = 1600
return self.inputs["input_ids"].shape[0]
[page 22]
def __getitem__(self, idx):
"""
Returns ONE tokenized sentence pair by index.
The DataLoader calls this repeatedly to assemble mini-batches during
training.
For example, to build a batch of 32 sentences, it calls this 32 times
with different idx values and stacks the results.
Args:
idx : integer index of the sentence pair to retrieve (0-based)
Returns:
A dictionary with three tensors, each of shape (128,):
"input_ids" : English token IDs e.g. [416, 2164, 123,
..., 1, 1]
"attention_mask" : 1 for real tokens, 0 for padding e.g. [1,
1, 1, ..., 0, 0]
"labels" : French token IDs with -100 at padding
e.g. [34, 78, ..., -100]
Example:
dataset[0] -> {
"input_ids" : tensor([416, 2164, 123, ..., 1, 1]),
"attention_mask" : tensor([1, 1, 1, ..., 0, 0]),
"labels" : tensor([34, 78, 56, ..., -100, -100])
}
"""
# Index into the pre-tokenized tensors to get row `idx`.
# e.g. self.inputs["input_ids"][3] returns the 4th English sentence's
token IDs as a 1D tensor.
return {
"input_ids" : self.inputs["input_ids"][idx], # English
tokens for this sentence
"attention_mask" : self.inputs["attention_mask"][idx], # mask:
1=real word, 0=padding
"labels" : self.labels[idx] # French
tokens (-100 at padding)
}
# ---------------------------------------------------------------
# Create the actual train and validation dataset objects
# ---------------------------------------------------------------
# Now we instantiate TranslationDataset twice:
# train_dataset : wraps the 160 training sentence pairs
# val_dataset : wraps the 40 validation sentence pairs
#
# Both use the same tokenizer so the vocabulary (token IDs) is consistent.
# Tokenization happens here, inside __init__, not lazily per batch.
train_dataset = TranslationDataset(train_english, train_french, tokenizer)
val_dataset = TranslationDataset(val_english, val_french, tokenizer)
[page 23]
Training dataset size : 160 sentence pairs
Validation dataset size : 40 sentence pairs
--- Structure of one dataset item ---
Keys : ['input_ids', 'attention_mask', 'labels']
input_ids shape : torch.Size([128])
attention_mask shape : torch.Size([128])
labels shape : torch.Size([128])
input_ids (first 10) : tensor([ 2476, 3921, 3, 0, 59513, 59513,
59513, 59513, 59513, 59513])
labels (first 10) : tensor([ 752, 900, 424, 51, 291, 0, -100,
-100, -100, -100])
Step 3 — Create DataLoaders
What is a DataLoader?
In cell_5b we created a Dataset -- a smart container that holds all our sentence pairs and
can hand out one pair at a time.
But during training, we do not want to feed the model one sentence at a time. That would
be very slow. Instead we feed it a batch of sentences at once (e.g. 16 sentences at a time).
# ---------------------------------------------------------------
# Sanity check: confirm the datasets were created correctly
# ---------------------------------------------------------------
print(f"Training dataset size : {len(train_dataset)} sentence pairs")
print(f"Validation dataset size : {len(val_dataset)} sentence pairs")
# Retrieve the first item to verify the structure is as expected
sample_item = train_dataset[0]
print(f"\n--- Structure of one dataset item ---")
print(f"Keys : {list(sample_item.keys())}")
# All three tensors should be 1D with length 128 (our max_length)->set above
in the method __init__ of TranslationDataset
print(f"input_ids shape : {sample_item['input_ids'].shape}") #
expected: torch.Size([128])
print(f"attention_mask shape : {sample_item['attention_mask'].shape}") #
expected: torch.Size([128])
print(f"labels shape : {sample_item['labels'].shape}") #
expected: torch.Size([128])
# Preview the first 10 tokens of each tensor to spot-check values
print(f"\ninput_ids (first 10) : {sample_item['input_ids'][:10]}") #
English token IDs
print(f"labels (first 10) : {sample_item['labels'][:10]}") #
French token IDs (no -100 expected in first tokens)
#note: 59513 is a special token reserved for padding
[page 24]
A DataLoader wraps our Dataset and automatically:
Groups sentences into batches of a fixed size
Shuffles the training data before each epoch (so the model does not memorize the
order of sentences)
Loads the next batch in the background while the model is processing the current one
(this speeds things up significantly)
Think of it like a conveyor belt in a factory:
The Dataset is the warehouse of all raw materials (sentences)
The DataLoader is the conveyor belt that delivers them in neat batches to the worker
(the model) at a steady pace
Training vs Validation DataLoader
We create TWO DataLoaders:
train_loader : shuffles data each epoch (so the model sees sentences in a different
order every time, which helps it learn better)
val_loader : does NOT shuffle (order does not matter for evaluation, and keeping it
consistent makes results easier to interpret)
# cell_5c
# ============================================================
# FINE-TUNING STEP 3: Create DataLoaders
# ============================================================
# A DataLoader wraps our Dataset and delivers data to the model
# in batches during training. It handles:
# - Batching : groups N sentences together for parallel processing
# - Shuffling : randomizes order each epoch (training only)
# - Loading : fetches the next batch while the current one is
# being processed (num_workers controls this)
from torch.utils.data import DataLoader # PyTorch's built-in DataLoader
class
# ---------------------------------------------------------------
# Define the batch size
# ---------------------------------------------------------------
# BATCH_SIZE controls how many sentence pairs the model sees at once.
#
# Larger batch = faster training BUT needs more GPU memory.
# Smaller batch = slower training BUT uses less GPU memory.
#
# 16 is a safe choice for fine-tuning on a dataset of our size.
# If you get an out-of-memory error, reduce this to 8.
# If you have a powerful GPU with lots of memory, you could try 32.
BATCH_SIZE = 16
# ---------------------------------------------------------------
# Create the Training DataLoader
[page 25]
# ---------------------------------------------------------------
# shuffle=True : randomly reorders the sentences before each epoch.
# Why? Because if the model always sees sentences in the same order,
# it might start to "memorize" the order rather than learning the
# actual patterns. Shuffling prevents this.
# Example: Epoch 1 order -> [5, 2, 8, 1, ...]
# Epoch 2 order -> [3, 7, 1, 9, ...] (different each time)
#
# num_workers=2 : uses 2 background processes to load the next batch
# while the model is still processing the current one.
# This keeps the GPU busy instead of waiting for data.
# Set to 0 if you get any errors related to multiprocessing.
train_loader = DataLoader(
train_dataset, # the training dataset we created in cell_5b
batch_size=BATCH_SIZE,
shuffle=True, # shuffle order every epoch
num_workers=2 # background workers for faster loading
)
# ---------------------------------------------------------------
# Create the Validation DataLoader
# ---------------------------------------------------------------
# shuffle=False : we do NOT shuffle validation data.
# Why? Because during validation we only care about the loss value,
# not the order. Keeping it consistent also makes it easier to
# debug if something goes wrong.
val_loader = DataLoader(
val_dataset, # the validation dataset we created in cell_5b
batch_size=BATCH_SIZE,
shuffle=False, # no shuffling for validation
num_workers=2 # same background loading for speed
)
# ---------------------------------------------------------------
# Sanity check: confirm the DataLoaders look correct
# ---------------------------------------------------------------
# The number of batches = total sentences / batch size (rounded up).
# Example with BATCH_SIZE=16:
# Training : 1600 sentences / 16 = 100 batches
# Validation : 400 sentences / 16 = 25 batches
print(f"Batch size : {BATCH_SIZE}")
print(f"Training batches : {len(train_loader)} "
f"({len(train_dataset)} sentences / {BATCH_SIZE} per batch)")
print(f"Validation batches : {len(val_loader)} "
f"({len(val_dataset)} sentences / {BATCH_SIZE} per batch)")
# ---------------------------------------------------------------
# Peek at one batch to confirm the structure looks right
# ---------------------------------------------------------------
# next(iter(train_loader)) fetches the very first batch.
# iter() : converts the DataLoader into an iterator
[page 26]
Batch size : 16
Training batches : 10 (160 sentences / 16 per batch)
Validation batches : 3 (40 sentences / 16 per batch)
--- Structure of one batch ---
input_ids shape : torch.Size([16, 128])
attention_mask shape : torch.Size([16, 128])
labels shape : torch.Size([16, 128])
Step 4 — Training and Validation Loop
What happens in a Training Loop?
This is the heart of fine-tuning. In each epoch (one full pass through the training data),
the model does the following for every batch:
1. Forward pass -- the model looks at the English sentences and tries to predict the
French translations. It compares its predictions to the correct French labels and
calculates a loss (how wrong it was).
2. Backward pass -- the model works backwards through its own calculations to figure
out which internal numbers (weights) caused the error. This is called
backpropagation.
3. Update weights -- the optimizer nudges the weights slightly in the direction that
reduces the loss. Over many batches, the model gradually gets better.
Think of it like learning to throw darts:
You throw (forward pass)
You see how far off you were (loss)
You adjust your technique (backward pass + weight update)
You throw again, slightly better each time
# next() : asks it for the first item (one batch)
#
# Each batch is a dictionary with three keys (same as our Dataset):
# "input_ids" : shape (BATCH_SIZE, 128) -- token IDs for English
# "attention_mask" : shape (BATCH_SIZE, 128) -- 1s for real, 0s for padding
# "labels" : shape (BATCH_SIZE, 128) -- token IDs for French
sample_batch = next(iter(train_loader))
print(f"\n--- Structure of one batch ---")
print(f"input_ids shape : {sample_batch['input_ids'].shape}")
# Expected: torch.Size([16, 128])
print(f"attention_mask shape : {sample_batch['attention_mask'].shape}")
# Expected: torch.Size([16, 128])
print(f"labels shape : {sample_batch['labels'].shape}")
# Expected: torch.Size([16, 128])
[page 27]
What is the Validation Loop?
After each epoch of training, we run the model on the validation set (sentences it has
never trained on) to check:
Is the loss going DOWN? Good -- the model is learning.
Is the validation loss going UP while training loss goes down? Bad -- the model is
memorizing training data (overfitting).
Key concepts in this cell
AdamW optimizer -- the algorithm that updates the model weights. "W" stands for
weight decay, a technique that prevents the model from making any single weight too
large, which helps avoid overfitting.
Learning rate (5e-5) -- how big each weight update step is. Too large = the model
overshoots and forgets what it already learned. Too small = training takes forever. 5e-
5 (0.00005) is the standard safe choice for fine-tuning transformers.
Epoch -- one complete pass through all training data. We train for 3 epochs, meaning
the model sees each sentence 3 times.
# cell_5d
# ============================================================
# FINE-TUNING STEP 4: Training and Validation Loop
# ============================================================
# This cell fine-tunes the pretrained MarianMT model on our
# English-French dataset. It runs for a fixed number of epochs.
# Each epoch has two phases:
# 1. Training phase : model learns from training batches
# 2. Validation phase : model is evaluated on unseen data
from torch.optim import AdamW # the optimizer that updates model weights
from tqdm import tqdm # displays a live progress bar
# ---------------------------------------------------------------
# Hyperparameters
# ---------------------------------------------------------------
# These are the settings that control how training behaves.
# They are called "hyperparameters" because we set them manually
# (the model does not learn them automatically).
# Number of times we go through the entire training dataset.
# 3 epochs is a good starting point for fine-tuning.
# More epochs = more learning BUT risk of overfitting.
NUM_EPOCHS = 1
# How big each weight update step is.
# 5e-5 means 0.00005 -- a very small step, which is intentional.
# Fine-tuning needs small steps so we don't destroy what the
# pretrained model already learned. This is called avoiding
[page 28]
# "catastrophic forgetting".
LEARNING_RATE = 5e-5
# ---------------------------------------------------------------
# Set up the optimizer
# ---------------------------------------------------------------
# AdamW is the standard optimizer for fine-tuning transformers.
# It takes all the model's weights (model.parameters()) and
# updates them slightly after each batch based on the loss.
#
# weight_decay=0.01 is a regularization technique.
# It gently penalizes very large weights, which helps prevent
# the model from overfitting to the training data.
optimizer = AdamW(model.parameters(), lr=LEARNING_RATE, weight_decay=0.01)
# ---------------------------------------------------------------
# Switch model to TRAINING mode
# ---------------------------------------------------------------
# Remember in cell_5 we called model.eval() to turn OFF dropout.
# Now we call model.train() to turn dropout back ON.
# During training, dropout randomly deactivates some neurons each
# step, which forces the model to learn more robust patterns
# instead of relying too heavily on any single neuron.
model.train()
# ---------------------------------------------------------------
# Storage for tracking loss across epochs
# ---------------------------------------------------------------
# We store the average loss for each epoch so we can print a
# summary at the end and see if training is going in the right direction.
# Loss should decrease over epochs if the model is learning correctly.
train_losses = [] # average training loss per epoch
val_losses = [] # average validation loss per epoch
# ---------------------------------------------------------------
# Main training loop
# ---------------------------------------------------------------
# We loop NUM_EPOCHS times. Each iteration is one full pass
# through the entire training dataset.
for epoch in range(NUM_EPOCHS):
print(f"\n{'='*60}")
print(f"EPOCH {epoch + 1} of {NUM_EPOCHS}")
print(f"{'='*60}")
# ===========================================================
# PHASE 1: TRAINING
# ===========================================================
model.train() # ensure model is in training mode
total_train_loss = 0 # accumulate loss across all batches
# tqdm wraps train_loader to show a live progress bar.
[page 29]
# desc= sets the label shown next to the bar.
train_bar = tqdm(train_loader, desc=f" Training ")
for batch in train_bar:
# -------------------------------------------------------
# Move batch data to the correct device (GPU or CPU)
# -------------------------------------------------------
# Every tensor must be on the same device as the model.
# We move all three tensors in the batch to DEVICE.
input_ids = batch["input_ids"].to(DEVICE)
attention_mask = batch["attention_mask"].to(DEVICE)
labels = batch["labels"].to(DEVICE)
# -------------------------------------------------------
# Step 1: Zero out gradients from the previous batch
# -------------------------------------------------------
# PyTorch accumulates gradients by default -- it adds new
# gradients ON TOP of old ones. We must reset them to zero
# before each batch, otherwise the updates will be wrong.
# Think of it like wiping the whiteboard clean before
# solving a new math problem.
optimizer.zero_grad()
# -------------------------------------------------------
# Step 2: Forward pass -- compute the loss
# -------------------------------------------------------
# We pass the English tokens (input_ids, attention_mask)
# AND the correct French tokens (labels) to the model.
#
# When labels are provided, MarianMT automatically:
# 1. Runs the encoder on the English input
# 2. Runs the decoder to predict French tokens one by one
# 3. Compares predictions to labels
# 4. Computes and returns the cross-entropy loss
#
# Cross-entropy loss measures how wrong the predictions are.
# Lower loss = better predictions.
outputs = model(
input_ids=input_ids,
attention_mask=attention_mask,
labels=labels
)
loss = outputs.loss # the scalar loss value for this batch
# -------------------------------------------------------
# Step 3: Backward pass -- compute gradients
# -------------------------------------------------------
# loss.backward() tells PyTorch to work backwards through
# all the calculations and compute how much each weight
# in the model contributed to the loss.
[page 30]
# These are called "gradients" -- they tell the optimizer
# which direction to nudge each weight.
loss.backward()
# -------------------------------------------------------
# Step 4: Clip gradients
# -------------------------------------------------------
# Sometimes gradients can become extremely large, causing
# the optimizer to make a huge update that destabilizes
# training. This is called "exploding gradients".
# clip_grad_norm_() caps the total gradient size at 1.0,
# preventing any single update from being too drastic.
# This is standard practice when fine-tuning transformers.
torch.nn.utils.clip_grad_norm_(model.parameters(), max_norm=1.0)
# -------------------------------------------------------
# Step 5: Update the weights
# -------------------------------------------------------
# optimizer.step() uses the gradients computed in backward()
# to nudge each weight slightly in the direction that reduces
# the loss. The size of each nudge is controlled by LEARNING_RATE.
optimizer.step()
# -------------------------------------------------------
# Track the loss for this batch
# -------------------------------------------------------
# loss.item() converts the tensor loss to a plain Python float.
# We add it to our running total so we can compute the average
# loss across all batches at the end of the epoch.
total_train_loss += loss.item()
# Update the progress bar to show the current batch loss
# .4f formats the number to 4 decimal places
train_bar.set_postfix({"batch_loss": f"{loss.item():.4f}"})
# Compute average training loss for this epoch
# We divide by the number of batches (len(train_loader))
avg_train_loss = total_train_loss / len(train_loader)
train_losses.append(avg_train_loss)
print(f"\n Avg Training Loss : {avg_train_loss:.4f}")
# ===========================================================
# PHASE 2: VALIDATION
# ===========================================================
# After each epoch of training, we evaluate the model on the
# validation set. We do NOT update any weights here -- we only
# measure how well the model performs on unseen data.
model.eval() # switch to evaluation mode (turns off dropout)
total_val_loss = 0 # accumulate validation loss
# torch.no_grad() tells PyTorch not to track gradients.
[page 31]
============================================================
EPOCH 1 of 1
============================================================
Training : 100%|██████████| 10/10 [00:00<00:00, 16.19it/s,
batch_loss=2.9536]
# We don't need gradients during validation (no weight updates),
# so this saves memory and speeds things up.
with torch.no_grad():
val_bar = tqdm(val_loader, desc=f" Validating")
for batch in val_bar:
# Move batch to the correct device
input_ids = batch["input_ids"].to(DEVICE)
attention_mask = batch["attention_mask"].to(DEVICE)
labels = batch["labels"].to(DEVICE)
# Forward pass only -- no backward pass, no weight update
outputs = model(
input_ids=input_ids,
attention_mask=attention_mask,
labels=labels
)
loss = outputs.loss
total_val_loss += loss.item()
val_bar.set_postfix({"batch_loss": f"{loss.item():.4f}"})
# Compute average validation loss for this epoch
avg_val_loss = total_val_loss / len(val_loader)
val_losses.append(avg_val_loss)
print(f" Avg Validation Loss : {avg_val_loss:.4f}")
# ---------------------------------------------------------------
# Print final summary across all epochs
# ---------------------------------------------------------------
# A good sign: both training and validation loss should decrease
# across epochs. If validation loss starts increasing while
# training loss keeps decreasing, the model is overfitting.
print(f"\n{'='*60}")
print(f"TRAINING COMPLETE")
print(f"{'='*60}")
print(f"{'Epoch':<10} {'Train Loss':<20} {'Val Loss':<20}")
print(f"{'-'*50}")
for i, (tl, vl) in enumerate(zip(train_losses, val_losses)):
print(f"{i+1:<10} {tl:<20.4f} {vl:<20.4f}")
[page 32]
Avg Training Loss : 3.6104
Validating: 100%|██████████| 3/3 [00:00<00:00, 20.38it/s,
batch_loss=2.6098]
Avg Validation Loss : 2.8498
============================================================
TRAINING COMPLETE
============================================================
Epoch Train Loss Val Loss
--------------------------------------------------
1 3.6104 2.8498
Step 5 — Save the Fine-Tuned Model
Why Save the Model?
Training takes time and computing power. If we close the notebook or the runtime
crashes, all that fine-tuning work is lost forever unless we save it to disk.
What gets saved?
Two things need to be saved together -- they are a matched pair and both are needed to
translate later:
Model weights -- the millions of numbers inside the model that were updated during
fine-tuning. Saved by model.save_pretrained()
Tokenizer -- the vocabulary and rules for converting text to numbers and back.
Saved by tokenizer.save_pretrained()
Where does it get saved?
We reuse the DATA_PATH variable from cell_2 (where our CSV dataset is located) and
save the model in the SAME folder.
This means:
No hardcoded paths
No manual changes needed
Works automatically on both Google Colab and any local PC
Everything stays neatly in one place -- dataset and model together
[page 33]
# cell_5e (updated -- saves to the same folder as the dataset)
# ============================================================
# FINE-TUNING STEP 5: Save the Fine-Tuned Model
# ============================================================
# We save both the model weights and the tokenizer to the SAME
# folder where our dataset CSV file lives.
#
# We reuse DATA_PATH from cell_2 so there is no hardcoding --
# this works automatically on both Google Colab and local PC
# without any manual changes.
import os
# ---------------------------------------------------------------
# Derive the save path from DATA_PATH (defined in cell_2)
# ---------------------------------------------------------------
# os.path.dirname() extracts the folder part from a full file path.
#
# Example:
# DATA_PATH = "/content/eng_-french.csv"
# os.path.dirname(DATA_PATH) -> "/content"
# SAVE_PATH -> "/content/finetuned_model"
#
# Another example on a local PC:
# DATA_PATH = "C:/Users/you/project/eng_-french.csv"
# os.path.dirname(DATA_PATH) -> "C:/Users/you/project"
# SAVE_PATH -> "C:/Users/you/project/finetuned_model"
#
# This means the model is ALWAYS saved next to the CSV file,
# regardless of which machine or environment you are running on.
SAVE_PATH = os.path.join(os.path.dirname(DATA_PATH), "finetuned_model")
print(f"Dataset location : {os.path.dirname(DATA_PATH)}")
print(f"Model will be saved to : {SAVE_PATH}")
# ---------------------------------------------------------------
# Create the save folder if it does not already exist
# ---------------------------------------------------------------
# exist_ok=True means: do not raise an error if the folder
# already exists (e.g. if we run this cell a second time).
os.makedirs(SAVE_PATH, exist_ok=True)
print(f"Folder created (or already exists)")
# ---------------------------------------------------------------
# Save the fine-tuned model weights
# ---------------------------------------------------------------
# save_pretrained() saves the model architecture config AND
# all the weight values that were updated during fine-tuning.
# This is the large file (~300 MB) that contains everything
# the model learned during our fine-tuning in cell_5d.
model.save_pretrained(SAVE_PATH)
[page 34]
/mnt/home2/home/pratik_cs/.conda/envs/indirides_sentiment_analysis/lib/python
packages/transformers/modeling_utils.py:4034: UserWarning: Moving the
following attributes in the config to the generation config: {'max_length':
512, 'num_beams': 4, 'bad_words_ids': [[59513]]}. You are seeing this warning
because you've set generation parameters in the model config, as opposed to
in the generation config.
warnings.warn(
Dataset location :
/mnt/home2/home/pratik_cs/classwork/Attention_mechanism/English–French
dataset/data
Model will be saved to :
/mnt/home2/home/pratik_cs/classwork/Attention_mechanism/English–French
dataset/data/finetuned_model
Folder created (or already exists)
print(f"\nModel saved : {SAVE_PATH}")
# ---------------------------------------------------------------
# Save the tokenizer
# ---------------------------------------------------------------
# The tokenizer must always be saved alongside the model.
# They are a matched pair -- using a different tokenizer with
# this model would produce garbage translations.
tokenizer.save_pretrained(SAVE_PATH)
print(f"Tokenizer saved : {SAVE_PATH}")
# ---------------------------------------------------------------
# Confirm what was saved by listing the files
# ---------------------------------------------------------------
# We list every file in the save folder along with its size
# so we can visually confirm everything was saved correctly.
saved_files = os.listdir(SAVE_PATH)
print(f"\nFiles saved ({len(saved_files)} total):")
for f in sorted(saved_files):
size_mb = os.path.getsize(os.path.join(SAVE_PATH, f)) / (1024 * 1024)
print(f" {f:<40} {size_mb:.2f} MB")
# ---------------------------------------------------------------
# Show how to reload the model later
# ---------------------------------------------------------------
# We print the exact code the user needs to reload this model
# in any future notebook session.
print(f"""
To reload your fine-tuned model in any future session:
from transformers import MarianMTModel, MarianTokenizer
model = MarianMTModel.from_pretrained(r"{SAVE_PATH}")
tokenizer = MarianTokenizer.from_pretrained(r"{SAVE_PATH}")
model = model.to(DEVICE)
model.eval()
""")
[page 35]
Model saved :
/mnt/home2/home/pratik_cs/classwork/Attention_mechanism/English–French
dataset/data/finetuned_model
Tokenizer saved :
/mnt/home2/home/pratik_cs/classwork/Attention_mechanism/English–French
dataset/data/finetuned_model
Files saved (8 total):
config.json 0.00 MB
generation_config.json 0.00 MB
model.safetensors 284.87 MB
source.spm 0.74 MB
special_tokens_map.json 0.00 MB
target.spm 0.77 MB
tokenizer_config.json 0.00 MB
vocab.json 1.39 MB
To reload your fine-tuned model in any future session:
from transformers import MarianMTModel, MarianTokenizer
model =
MarianMTModel.from_pretrained(r"/mnt/home2/home/pratik_cs/classwork/Attention_
French dataset/data/finetuned_model")
tokenizer =
MarianTokenizer.from_pretrained(r"/mnt/home2/home/pratik_cs/classwork/Attenti
French dataset/data/finetuned_model")
model = model.to(DEVICE)
model.eval()
Translation
# cell_6
# ============================================================
# STEP 7: Translate a single English sentence to French
# ============================================================
# This cell walks through the full translation pipeline step by step.
# We do it manually here (instead of using a function) so we can see
# exactly what happens at each stage. Later we will wrap this into
# a reusable function.
#
# The full pipeline looks like this:
# "I love learning new things every day." (raw English text)
# |
# v Step 1: Tokenize
# tensor([[416, 2164, 2onal, ..., 0]]) (numbers the model reads)
# |
# v Step 2: Model generates output
[page 36]
# tensor([[34, 78, 120, 56, 0]]) (numbers the model outputs)
# |
# v Step 3: Decode
# "J'adore apprendre de nouvelles choses." (readable French text)
sample_sentence = "I love learning new things every day."
# ---------------------------------------------------------------
# Step 1: Tokenize the input sentence
# ---------------------------------------------------------------
# The tokenizer converts the raw English text into a PyTorch tensor
# of integer IDs that the model can process.
#
# Arguments explained:
# return_tensors="pt" : return the result as a PyTorch tensor.
# "pt" stands for PyTorch. The alternative is
# "np" for NumPy arrays, but the model needs "pt".
#
# padding=True : if we pass multiple sentences, they may have
# different lengths. Padding adds extra zeros to
# shorter sentences so all inputs are the same
length.
# Not strictly needed for a single sentence, but good
# practice for when we process batches later.
#
# .to(DEVICE) : moves the tensor to the same device (GPU or CPU)
# as the model. Both must be on the same device or
# PyTorch will throw an error.
inputs = tokenizer(
sample_sentence,
return_tensors="pt", # return as PyTorch tensors
padding=True # pad the input if needed
).to(DEVICE)
print(f"Input sentence : {sample_sentence}")
# inputs is a dictionary with key "input_ids" holding the token ID tensor.
# Example output: tensor([[416, 2164, 123, 456, 78, 90, 12, 0]])
# ^ EOS token
print(f"Tokenized input IDs: {inputs['input_ids']}")
# ---------------------------------------------------------------
# Show tokenization examples from the actual dataset
# ---------------------------------------------------------------
# Looking at real sentences from our dataset (not just the sample above)
# helps us understand how the tokenizer handles everyday language.
# We pick 3 random rows using the same random_state as before for
# reproducibility.
print(f"\n--- Tokenization examples from the dataset ---")
print(f"{'English':<35} {'Tokens':<55} {'Token IDs'}")
print("-" * 120)
[page 37]
for _, row in df.sample(3, random_state=42).iterrows():
# tokenize() shows us the subword pieces the text was split into
tokens_ex = tokenizer.tokenize(row["english"])
# encode() shows us the integer IDs, including the final EOS token (0)
token_ids_ex = tokenizer.encode(row["english"])
print(f"{row['english']:<35} {str(tokens_ex):<55} {token_ids_ex}")
# ---------------------------------------------------------------
# Step 2: Generate the translated output tokens
# ---------------------------------------------------------------
# model.generate() runs the actual translation. It takes the input
# token IDs and produces a sequence of output token IDs in French.
#
# How it works internally (simplified):
# - The encoder reads the English token IDs and builds a rich
# numerical representation of the meaning of the sentence.
# - The decoder then generates French token IDs one at a time,
# each time looking at both the encoder output and the French
# tokens it has already generated, until it produces the EOS token.
#
# torch.no_grad() is a context manager that tells PyTorch:
# "We are not training, so do not store gradient information."
# During training, PyTorch tracks every calculation to compute gradients
# for updating the model weights. During inference we do not need this,
# so turning it off saves memory and speeds up the translation.
#
# **inputs unpacks the dictionary {"input_ids": ..., "attention_mask": ...}
# and passes each item as a separate argument to model.generate().
with torch.no_grad():
translated_tokens = model.generate(**inputs)
# The output is a tensor of integer IDs representing the French translation.
# Example output: tensor([[38, 200, 456, 789, 23, 0]])
# ^ EOS token
print(f"\nTranslated token IDs: {translated_tokens}")
# ---------------------------------------------------------------
# Step 3: Decode the output tokens back into readable French text
# ---------------------------------------------------------------
# tokenizer.decode() converts the integer IDs back into a readable string.
#
# translated_tokens[0] selects the first (and only) translation in the batch.
# If we had passed a batch of 3 sentences, we would have [0], [1], [2].
#
# skip_special_tokens=True removes special tokens like the EOS token (0)
# from the output. Without this, the output might look like:
# "J'adore apprendre de nouvelles choses.</s>"
# With it, we get clean readable text:
# "J'adore apprendre de nouvelles choses."
translated_text = tokenizer.decode(translated_tokens[0],
skip_special_tokens=True)
[page 38]
Input sentence : I love learning new things every day.
Tokenized input IDs: tensor([[ 47, 1779, 3655, 191, 1609, 963, 613, 3,
0]],
device='cuda:0')
--- Tokenization examples from the dataset ---
English Tokens
Token IDs
-----------------------------------------------------------------------------
-------------------------------------------
Call us. [' ▁ Call', ' ▁ us', '.']
[6703, 368, 3, 0]
Go on. [' ▁ Go', ' ▁ on', '.']
[2120, 30, 3, 0]
Get up. [' ▁ Get', ' ▁ up', '.']
[4203, 205, 3, 0]
Translated token IDs: tensor([[59513, 234, 6, 4247, 7497, 5,
828, 2141, 182, 16,
735, 3, 0]], device='cuda:0')
English : I love learning new things every day.
French : J'aime apprendre de nouvelles choses tous les jours.
59513
59513
# Print the final result: the original English and its French translation
print(f"\nEnglish : {sample_sentence}")
print(f"French : {translated_text}")
#Note: n MarianTokenizer, 59513 serves double duty:
#Context Middle/end of a sequence : Meaning -> Padding — ignore me
#Context First token of decoder input : Meaning -> Start of sequence —
begin generating
print(tokenizer.pad_token_id) # 59513
print(model.config.decoder_start_token_id) # also 59513 ← same token!
# cell_7
# ============================================================
# STEP 8: Wrap the translation logic into a reusable function
# ============================================================
# In cell_6 we translated one sentence by writing out every step manually.
# That works, but if we want to translate many sentences throughout the
# notebook, we would have to copy and paste the same code every time.
#
# Instead, we define a function called translate() once here, and then
# call it with a single line anywhere in the notebook.
#
# The function also handles two input formats automatically:
# - A single string : translate("Hello")
# - A list of strings: translate(["Hello", "How are you?"])
# In both cases it always returns a list of French translations.
[page 39]
def translate(texts):
"""
Translates a single sentence or a list of sentences from English to
French.
Args:
texts: a single string or a list of strings in English
Returns:
a list of translated French strings
"""
# ---------------------------------------------------------------
# Handle both single string and list inputs
# ---------------------------------------------------------------
# isinstance(texts, str) checks whether the input is a single string.
# If it is, we wrap it in a list so the rest of the function always
# works with a list. This avoids writing two separate code paths.
# Example:
# "Hello" becomes ["Hello"] (single string wrapped)
# ["Hello", "Hi"] stays ["Hello", "Hi"] (already a list,
unchanged)
if isinstance(texts, str):
texts = [texts]
# ---------------------------------------------------------------
# Step 1: Tokenize the input
# ---------------------------------------------------------------
# Same tokenization as in cell_6, but now with one extra argument:
#
# truncation=True : the model has a maximum input length of 512 tokens.
# If a sentence is longer than that, truncation=True
# cuts it off at 512 tokens instead of throwing an
error.
# Example: a very long paragraph gets cut to 512
tokens.
#
# padding=True : when translating a batch (list) of sentences, they
# may have different lengths. Padding adds a special
# padding token (ID = 1) to the end of shorter
sentences
# so all sentences in the batch have the same length.
# Example:
# "Hi." -> [4, 0, 1, 1] (padded with 1s)
# "How are you?"-> [6, 52, 55, 0] (no padding
needed)
#
# return_tensors="pt" : return as PyTorch tensors (needed by the
model).
#
# .to(DEVICE) : move tensors to the same device as the model.
inputs = tokenizer(
texts,
[page 40]
return_tensors="pt",
padding=True,
truncation=True
).to(DEVICE)
# ---------------------------------------------------------------
# Step 2: Generate translated token IDs
# ---------------------------------------------------------------
# model.generate() translates all sentences in the batch at once.
# Processing multiple sentences together (batching) is much faster
# than translating them one by one, especially on a GPU.
# torch.no_grad() disables gradient tracking to save memory and time,
# since we are doing inference (not training).
with torch.no_grad():
translated_tokens = model.generate(**inputs)
# ---------------------------------------------------------------
# Step 3: Decode all outputs back into readable French text
# ---------------------------------------------------------------
# In cell_6 we used tokenizer.decode() for a single sentence.
# Here we use tokenizer.batch_decode() which does the same thing
# but for a whole list of translated token sequences at once.
# It returns a list of French strings, one per input sentence.
# skip_special_tokens=True removes the EOS token from each output.
# Example output: ["J'adore apprendre de nouvelles choses chaque jour."]
translations = tokenizer.batch_decode(
translated_tokens,
skip_special_tokens=True
)
# Return the list of French translations to the caller.
# Even if the input was a single string, we always return a list,
# so the caller can always access the result as result[0].
return translations
# ---------------------------------------------------------------
# Quick test: confirm the function works correctly
# ---------------------------------------------------------------
# We test with the same sentence from cell_6 so we can compare outputs
# and confirm the function produces the same result as the manual steps.
test_sentence = "I love learning new things every day."
result = translate(test_sentence)
# result is a list, so we access the first (and only) translation with [0].
# Expected output:
# English : I love learning new things every day.
# French : J'adore apprendre de nouvelles choses chaque jour.
print(f"English : {test_sentence}")
print(f"French : {result[0]}")
[page 41]
English : I love learning new things every day.
French : J'aime apprendre de nouvelles choses tous les jours.
# cell_8
# ============================================================
# STEP 9: Translate a batch of real sentences from the dataset
# ============================================================
# In cell_7 we tested our translate() function on a single hand-picked
# sentence. Now we test it on real sentences from our dataset to see
# how well the model performs on actual data.
#
# We also display the model's predicted French translation alongside
# the actual (ground truth) French translation from the dataset.
# This lets us visually judge translation quality before we compute
# a formal score in the next cell.
# ---------------------------------------------------------------
# Sample 10 rows from the dataset
# ---------------------------------------------------------------
# df.sample(10) picks 10 random rows from our dataset.
# random_state=42 ensures we get the same 10 rows every time we run
# this cell (reproducibility). You can change the number to any value
# to see different sentences.
#
# reset_index(drop=True) resets the row numbers to 0-9.
# Without this, the row numbers would be random (e.g., 234, 1892, 7103...)
# because they came from random positions in the original dataset.
# drop=True means we discard the old row numbers entirely.
sample_df = df.sample(10, random_state=42).reset_index(drop=True)
# ---------------------------------------------------------------
# Extract the English sentences as a plain Python list
# ---------------------------------------------------------------
# Our translate() function expects a list of strings as input.
# .tolist() converts the pandas column (a Series) into a plain Python list.
# Example:
# ["I am happy.", "She likes cats.", "We are learning.", ...]
english_sentences = sample_df["english"].tolist()
# ---------------------------------------------------------------
# Translate all 10 sentences in one batch
# ---------------------------------------------------------------
# We pass the entire list to translate() at once (batch translation).
# This is more efficient than calling translate() 10 times in a loop,
# because the model processes all sentences in parallel on the GPU.
# The result is a list of 10 French strings in the same order as the input.
# Example:
# ["Je suis heureux.", "Elle aime les chats.", "Nous apprenons.", ...]
predicted_french = translate(english_sentences)
# ---------------------------------------------------------------
# Add the predictions as a new column in the DataFrame
[page 42]
English Predicted French
Actual French
-----------------------------------------------------------------------------
----------------------------------------------------------
Call us. Appelez-nous.
Appelle-nous !
Go on. Vas-y.
Poursuivez.
Get up. Lève-toi !
Lève-toi.
I swore. J'ai juré.
J’ai promis.
Go home. Rentrez à la maison.
Rentrez chez vous.
Go away! Pars !
Dégage !
We won. On a gagné.
Nous avons gagné.
I'm hit! Je suis frappé !
Je suis touchée !
# ---------------------------------------------------------------
# We store the model's translations back into sample_df so we can
# easily compare them with the actual French translations side by side.
# After this line, sample_df has three columns:
# "english" : the original English sentence
# "french" : the correct French translation (from the dataset)
# "predicted_french" : the model's French translation
sample_df["predicted_french"] = predicted_french
# ---------------------------------------------------------------
# Display the results in a formatted table
# ---------------------------------------------------------------
# We print three columns side by side so we can visually compare:
# - the original English sentence
# - what the model predicted in French
# - what the correct French translation actually is
#
# The :<45 format specifier left-aligns each value in a field of 45
# characters wide, so all three columns line up neatly regardless of
# the actual length of each sentence.
print(f"{'English':<45} {'Predicted French':<45} {'Actual French':<45}")
print("-" * 135)
# iterrows() loops through each row of the DataFrame one at a time.
# The underscore "_" is used for the row index, which we do not need here.
# For each row we print the three columns side by side.
# Example output line:
# I am happy. Je suis heureux.
Je suis heureux.
for _, row in sample_df.iterrows():
print(f"{row['english']:<45} {row['predicted_french']:<45}
{row['french']:<45}")
[page 43]
I'm wet. Je suis mouillée.
Je suis mouillé.
I know. Je sais.
Je sais.
Observations from Cell 8 -- Fine-Tuned Model
Predictions vs Ground Truth
Perfect Matches
"I wish Tom was here" -> identical: "J'aimerais que Tom soit là"
"The clock has stopped" -> identical: "L'horloge s'est arrêtée"
Correct but Different Wording
"I'm not scared to die"
Model : "Je n'ai pas peur de mourir" (not afraid to die)
Truth : "Je ne crains pas de mourir" (do not fear dying)
Both mean the same thing, just different verb choice
"How did the audition go?"
Model : "Comment s'est déroulé l'audition?"
Truth : "Comment s'est passée l'audition?"
Both are natural French, dérouler and passer are interchangeable here
"I really like this skirt. Can I try it on?"
Model : "J'aime vraiment cette jupe. Puis-je l'essayer?"
Truth : "J'aime beaucoup cette jupe, puis-je l'essayer?"
"vraiment" and "beaucoup" are both valid ways to say "really like"
Formal vs Informal Register
"Take a seat"
Model : "Assieds-toi" (informal tu form)
Truth : "Prends place!" (more formal)
The original English does not specify formality so both are correct
"You'd better make sure that it is true"
Model : "Vous feriez mieux..." (formal vous form)
Truth : "Tu ferais bien..." (informal tu form)
Again, English is ambiguous here so both are valid
Model Better than Ground Truth
"I've no friend to talk to about my problems"
Model : "Je n'ai pas d'ami à qui parler de mes problèmes"
Truth : "Je n'ai pas d'ami avec lequel je puisse m'entretenir..."
The model's version is actually more natural everyday French
The ground truth uses an unnecessarily formal construction
[page 44]
Genuine Error
"Take any two cards you like"
Model : "Prends n'importe quelle carte que tu aimes"
Truth : "Prends deux cartes de ton choix"
The model dropped the word "two" which changed the meaning
This shows that models can miss specific numbers in longer sentences
Key Takeaways
Most differences are about register or word choice, not actual errors
In some cases the model output is more natural than the ground truth
BLEU score penalizes valid alternative translations, so our score likely understates
the true quality of the model
The only genuine error was dropping the word "two" in one sentence
BLEU score for the loaded dataset
BLEU Score Comparison: Pretrained vs
Fine-Tuned Model
What are we doing here?
So far we have:
1. Used the pretrained MarianMT model to translate (before fine-tuning)
2. Fine-tuned the model on our English-French dataset
Now we want to answer the key question: Did fine-tuning actually improve the
translations?
We answer this by computing the BLEU score TWICE:
Once using the original pretrained model (reloaded from Hugging Face)
Once using the fine-tuned model (saved in cell_5e)
Then we compare the two scores side by side.
What should we expect?
With only 2,000 sentences and 3 epochs, the improvement may be small. But even a
small improvement confirms that fine-tuning is working correctly.
A large improvement would require:
[page 45]
More data (tens of thousands of sentences)
More epochs
A dataset with a very specific style or vocabulary that differs from the original
training data
Important reminder about BLEU score limitations
As we saw in cell_8, BLEU score penalizes valid translations that use different but
correct wording. So the true quality improvement from fine-tuning is likely BETTER
than what the numbers alone suggest.
# cell_9b (updated)
# ============================================================
# BLEU SCORE COMPARISON: Pretrained vs Fine-Tuned Model
# ============================================================
# We compute BLEU scores for both models and compare them.
# This tells us whether fine-tuning improved translation quality.
#
# Steps:
# 1. Translate the VALIDATION SET using the FINE-TUNED model
# (already loaded as `model` -- updated during cell_5d training)
# 2. Reload the ORIGINAL pretrained model from Hugging Face
# 3. Translate the VALIDATION SET using the ORIGINAL model
# 4. Compute BLEU scores for both and compare
#
# IMPORTANT: We use only the validation set (val_english, val_french)
# created in cell_5a. The fine-tuned model has NEVER seen these 40
# sentences during training, so the comparison is fair.
from tqdm import tqdm # progress bar
BATCH_SIZE = 32 # same batch size as before for consistency
# ---------------------------------------------------------------
# Build a DataFrame from the validation set for easy handling
# ---------------------------------------------------------------
# val_english and val_french were created in cell_5a.
# We combine them into a DataFrame so we can slice batches easily,
# just like we did with df in the training cells.
val_df = pd.DataFrame({"english": val_english, "french":
val_french}).reset_index(drop=True)
print(f"Evaluating on {len(val_df)} held-out validation sentences (never seen
during training)")
# ===========================================================
# PART 1: BLEU score for the FINE-TUNED model
# ===========================================================
# The `model` variable currently holds our fine-tuned model
# (it was updated in-place during the training loop in cell_5d).
# We run batch translation on the validation set only.
[page 46]
print("\nComputing BLEU score for FINE-TUNED model...")
print("-" * 50)
# Make sure model is in evaluation mode (no dropout)
model.eval()
finetuned_predictions = [] # will hold all fine-tuned translations
total = len(val_df)
num_batches = (total + BATCH_SIZE - 1) // BATCH_SIZE
for i in tqdm(range(0, total, BATCH_SIZE), total=num_batches, desc="Fine-
tuned model"):
# Slice one batch of English sentences from the validation set
batch = val_df["english"][i : i + BATCH_SIZE].tolist()
# Translate using the fine-tuned model
# translate() defined in cell_7 uses the global `model` variable
# which is currently our fine-tuned model
translated = translate(batch)
finetuned_predictions.extend(translated)
# The references must be wrapped in a list because sacrebleu supports
# multiple reference translations per sentence. We only have one here.
all_references = [val_df["french"].tolist()]
bleu_finetuned = sacrebleu.corpus_bleu(
finetuned_predictions,
all_references,
tokenize='intl'
)
print(f"Fine-Tuned Model BLEU Score : {bleu_finetuned.score:.2f} / 100")
# ===========================================================
# PART 2: BLEU score for the ORIGINAL pretrained model
# ===========================================================
# We reload the original pretrained model fresh from Hugging Face
# so that it has the ORIGINAL weights (before our fine-tuning).
# This gives us a fair baseline to compare against.
#
# We use a separate variable name `original_model` so we do NOT
# overwrite our fine-tuned `model` variable.
print(f"\nReloading original pretrained model for baseline comparison...")
print("-" * 50)
# Load the original pretrained model into a NEW variable
# This does NOT affect our fine-tuned `model` variable
original_model = MarianMTModel.from_pretrained(MODEL_NAME)
original_model = original_model.to(DEVICE)
[page 47]
original_model.eval()
print(f"Original pretrained model reloaded successfully")
# ---------------------------------------------------------------
# Define a temporary translate function for the original model
# ---------------------------------------------------------------
# Our existing translate() function in cell_7 uses the global
# `model` variable (fine-tuned). We define a separate function
# here that uses `original_model` instead, so we can fairly
# compare both models without any confusion.
def translate_original(texts):
"""
Same as translate() in cell_7 but uses the ORIGINAL
pretrained model instead of the fine-tuned one.
"""
if isinstance(texts, str):
texts = [texts]
inputs = tokenizer(
texts,
return_tensors="pt",
padding=True,
truncation=True
).to(DEVICE)
with torch.no_grad():
translated_tokens = original_model.generate(**inputs)
translations = tokenizer.batch_decode(
translated_tokens,
skip_special_tokens=True
)
return translations
# Translate validation set using the original pretrained model
print(f"\nComputing BLEU score for ORIGINAL pretrained model...")
original_predictions = []
for i in tqdm(range(0, total, BATCH_SIZE), total=num_batches, desc="Original
model "):
# Slice the same validation batches for a fair comparison
batch = val_df["english"][i : i + BATCH_SIZE].tolist()
translated = translate_original(batch)
original_predictions.extend(translated)
# Compute BLEU score for original model
bleu_original = sacrebleu.corpus_bleu(
original_predictions,
all_references,
tokenize='intl'
[page 48]
)
print(f"Original Pretrained Model BLEU Score : {bleu_original.score:.2f} /
100")
# ===========================================================
# PART 3: Side-by-side comparison
# ===========================================================
improvement = bleu_finetuned.score - bleu_original.score
print(f"\n{'='*55}")
print(f"{'BLEU SCORE COMPARISON SUMMARY':^55}")
print(f"{'='*55}")
print(f"{'Model':<35} {'BLEU Score':>10}")
print(f"{'-'*55}")
print(f"{'Original Pretrained Model':<35} {bleu_original.score:>10.2f}")
print(f"{'Fine-Tuned Model':<35} {bleu_finetuned.score:>10.2f}")
print(f"{'-'*55}")
# Show whether fine-tuning helped, hurt, or had no effect
if improvement > 0:
print(f"{'Improvement after Fine-Tuning':<35} {improvement:>+10.2f} ✓
Better")
elif improvement < 0:
print(f"{'Change after Fine-Tuning':<35} {improvement:>+10.2f} (see note
below)")
else:
print(f"{'Change after Fine-Tuning':<35} {improvement:>+10.2f} No
change")
print(f"{'='*55}")
# ---------------------------------------------------------------
# Helpful note if fine-tuning did not improve the score
# ---------------------------------------------------------------
if improvement <= 0:
print(f"""
Note: Fine-tuning did not improve the BLEU score this time.
This is common when:
- The dataset is small (we used only 200 sentences)
- The number of epochs is low (we used 1 epoch)
- The pretrained model was already trained on very similar data
To improve results, try:
- Using more data (10,000+ sentence pairs)
- Training for more epochs (5-10)
- Reducing the learning rate slightly (e.g. 2e-5)
""")
# ===========================================================
# PART 4: Side-by-side translation examples
# ===========================================================
[page 49]
Evaluating on 40 held-out validation sentences (never seen during training)
Computing BLEU score for FINE-TUNED model...
--------------------------------------------------
Fine-tuned model: 100%|██████████| 2/2 [00:00<00:00, 23.49it/s]
Fine-Tuned Model BLEU Score : 26.46 / 100
Reloading original pretrained model for baseline comparison...
--------------------------------------------------
Original pretrained model reloaded successfully
Computing BLEU score for ORIGINAL pretrained model...
Original model : 100%|██████████| 2/2 [00:00<00:00, 22.91it/s]
Original Pretrained Model BLEU Score : 24.73 / 100
=======================================================
BLEU SCORE COMPARISON SUMMARY
=======================================================
Model BLEU Score
-------------------------------------------------------
Original Pretrained Model 24.73
Fine-Tuned Model 26.46
-------------------------------------------------------
Improvement after Fine-Tuning +1.73 ✓ Better
=======================================================
# Numbers alone don't tell the full story. Let's look at actual
# translation examples from both models side by side on 5 random
# validation sentences to visually judge quality differences.
print(f"\n--- Side-by-side translation examples (validation set only) ---\n")
print(f"{'English':<35} {'Original Model':<35} {'Fine-Tuned Model':<35}
{'Ground Truth':<35}")
print("-" * 140)
# Sample 5 random sentences from the validation set
sample_df = val_df.sample(5, random_state=42).reset_index(drop=True)
for _, row in sample_df.iterrows():
orig_translation = translate_original(row["english"])[0]
finetuned_translation = translate(row["english"])[0]
print(
f"{row['english']:<35} "
f"{orig_translation:<35} "
f"{finetuned_translation:<35} "
f"{row['french']:<35}"
)
[page 50]
--- Side-by-side translation examples (validation set only) ---
English Original Model Fine-
Tuned Model Ground Truth
-----------------------------------------------------------------------------
---------------------------------------------------------------
We try. Nous essayons. Nous
essayons. On essaye.
No way! C'est pas vrai ! C'est
pas possible ! Impossible !
Join us. Joignez-vous à nous.
Joignez-vous à nous. Joignez-vous.
Be fair. Sois juste. Soyez
justes. Soyez équitables !
Go home. Rentre chez toi.
Rentrez à la maison. Rentrez chez vous.
Summary
In this notebook we built a complete English to French translation pipeline using a
pretrained transformer model and then fine-tuned it on our own dataset. Here is a recap of
everything we did and what we learned along the way.
What we did
1. Loaded the dataset -- we worked with a real English-French translation dataset. We
explored its structure, checked for missing values, duplicates, and understood the
sentence length distribution.
2. Loaded a pretrained model -- we used Helsinki-NLP/opus-mt-en-fr, a MarianMT
transformer model trained specifically for English to French translation.
3. Understood the tokenizer -- we saw how raw English text gets converted into
numbers (token IDs) that the model can read, and how the output numbers get
converted back into readable French text.
4. Translated sentences -- we translated a single sentence step by step first, then
wrapped the logic into a reusable function, and finally translated the entire dataset in
batches.
5. Evaluated with BLEU score -- we measured translation quality using the BLEU
metric on the pretrained model and got a score of 53.68.
6. Fine-tuned the model -- we split our data into training and validation sets, created a
custom PyTorch Dataset and DataLoader, and ran a training loop for 3 epochs to
adapt the pretrained model to our specific dataset.
7. Compared results -- we computed BLEU scores for both the original pretrained
model and our fine-tuned model and compared them side by side.
[page 51]
What we learned
Pretrained transformers handle most text processing automatically through the
tokenizer. For clean datasets like this one, no manual preprocessing is needed.
Fine-tuning on even a small dataset can meaningfully improve translation quality.
BLEU score has limitations. It penalizes valid translations that use different but
correct wording. Looking at actual translations in cell_8, many differences were
about register or word choice rather than genuine errors.
In some cases the fine-tuned model produced more natural French than the ground
truth itself, which further shows that BLEU score alone does not tell the full story.
Key results
Model BLEU Score
LSTM with Attention 21.17
MarianMT (pretrained) 53.72
MarianMT (fine-tuned) 57.60
A few things stand out from this table:
MarianMT pretrained already achieves more than double the BLEU score of an
LSTM with attention, without any training on our part.
Fine-tuning pushed the score further from 53.72 to 57.60.
This shows that pretrained transformers are a strong starting point, and fine-tuning on
your own data can make them even better with relatively little effort.