# sqaud gpt qa v4 GC
course: Module 4 — Generative AI & LLMs
module: Module-4-Generative-AI-LLMs
type: notebook
source_url: https://personal-learn.armco.dev/files/Module-4-Generative-AI-LLMs/General/Materials-06-06-2026/sqaud_gpt_qa_v4_GC.ipynb
---
[cell 1 markdown]
## The SQuAD Dataset
### What it is
SQuAD (the Stanford Question Answering Dataset) is a reading comprehension dataset. Each item gives the model a passage of text and a question about that passage. The answer to the question is always a span of text taken directly from the passage.
The version we use here is SQuAD v1.1, loaded with `load_dataset("squad")` from the Hugging Face `datasets` library. It was built from a set of Wikipedia articles, with questions and answers written by human crowdworkers.
### Why we chose it
A few reasons make SQuAD a good fit for a beginner tutorial:
1. It is clean and well documented. The structure is simple and consistent, so we spend time learning the task instead of cleaning messy data.
2. It is realistic. The passages are real Wikipedia text and the questions are written by people, so the task resembles something you might actually build.
3. It is widely used. SQuAD is a standard benchmark, so the metrics and methods we learn here carry over to other work.
### What one example looks like
Every row in the dataset has the same fields:
- `id`: a unique identifier for the example.
- `title`: the title of the Wikipedia article the passage came from.
- `context`: the passage of text. This is what the model reads to find the answer.
- `question`: the question asked about the context.
- `answers`: the correct answer or answers.
The `answers` field is worth a closer look because its shape can be confusing at first. It is not a single string. It is a small dictionary with two parallel lists:
- `text`: the answer string or strings.
- `answer_start`: the character position in the context where each answer begins.
For example, a single row might look like this:
context: "The Amazon rainforest covers most of the Amazon basin
in South America. The basin spans nine countries."
question: "How many countries does the Amazon basin span?"
answers: {"text": ["nine"], "answer_start": [88]}
The answer "nine" appears in the context, and `answer_start` tells you the exact character index where it starts.
### One answer or several
The number of answers depends on the split:
- In the training split, each question has exactly one answer.
- In the validation split, each question often has several answers, written by different annotators. They are all considered correct.
This matters when we score the model. A fair score checks the prediction against every acceptable answer and keeps the best match, so we do not mark a correct answer wrong just because it matched a different annotator.
### How we use SQuAD in this tutorial
SQuAD was originally designed for extractive question answering, where the model points to the start and end of the answer inside the passage. We are doing something different. We treat the task as text generation: the model reads the passage and the question, then writes the answer as text.
To do this we put every example into one consistent prompt format:
Context: <the passage>
Question: <the question>
Answer:
The model's job is to continue this text with the answer. We use the same format everywhere, both when we test the untrained model and after fine-tuning, so the before and after comparison is fair.
### Splits we work with
SQuAD ships with two splits, a large training split and a smaller validation split. It does not include a public test split, because the original test set was kept private for an online leaderboard.
To match the usual train, validation, and test setup, we create our own three splits:
- `train`: used to fine-tune the model.
- `validation`: a portion of SQuAD's validation split, used to check the model after each training epoch and pick the best one.
- `test`: the remaining portion, held out and used only for the final before and after numbers.
### What to keep in mind
A couple of points are useful to remember as we go:
- Answers in SQuAD are short, usually just a few words. The passages, by contrast, are full paragraphs. This is why we ask the model to read a long input but produce a short output.
- Because the answer is always present in the passage, the task rewards careful reading rather than outside knowledge. The model should find the answer in the text, not invent it.
[cell 2 code]
# cell_0a — GPU mask setup before any torch import
# MAIN GOAL: Let the user pick which physical GPU to use (or CPU),
# then hide all other GPUs from this process via CUDA_VISIBLE_DEVICES.
# Must run before importing torch or any CUDA-using library.
import os
choice = input("Enter physical GPU index (e.g. 0, 1, 2) or 'cpu': ").strip().lower()
if choice == "cpu":
os.environ["CUDA_VISIBLE_DEVICES"] = ""
DEVICE = "cpu"
elif choice.isdigit():
os.environ["CUDA_VISIBLE_DEVICES"] = choice # hide every other GPU
DEVICE = "cuda:0" # selected GPU maps to cuda:0 internally
else:
print("Invalid input. Defaulting to CPU.")
os.environ["CUDA_VISIBLE_DEVICES"] = ""
DEVICE = "cpu"
print(f"Device set to: {DEVICE}")
[cell 3 code]
# cell_0b — install pinned libraries
# MAIN GOAL: Install the exact library versions this tutorial was built and tested against.
# Pinning versions is a deliberate safeguard: a surprise upgrade to transformers/datasets
# can change APIs mid-session and break the notebook for a beginner audience.
# This cell only INSTALLS packages — it does not import them. Importing torch happens
# later, AFTER the GPU mask in cell_0a has already been set.
# NOTE: on Colab, if pip reports that a package was updated, you may need to use
# Runtime > Restart session once, then re-run from cell_0a. Locally this is usually not needed.
# The leading "!" runs this line as a shell command from inside the notebook.
# -q = quiet, fewer log lines so beginners aren't overwhelmed.
# package==X = pin to an exact version for reproducibility.
!pip install -q \
transformers==4.44.2 \
datasets==2.21.0 \
accelerate==0.34.2 \
matplotlib==3.9.2 \
pandas==2.2.2
# Quick confirmation line so the user gets visible feedback that the cell finished.
print("Install step complete. If on Colab and prompted, restart the session, then re-run from cell_0a.")
[cell 4 code]
# cell_0c — imports and device verification
# MAIN GOAL: Import torch (for the first time, now that cell_0a has already hidden the
# unwanted GPUs) plus the core libraries we'll use throughout, then CONFIRM that the
# device we asked for in cell_0a is actually the device PyTorch sees and will use.
# This catches "I asked for GPU but it silently fell back to CPU" problems early,
# before we waste time loading models or data.
import torch # deep learning framework; backs the GPT-2 model and training
import transformers # Hugging Face: model + tokenizer + Trainer
import datasets # Hugging Face: loads SQuAD
import pandas as pd # tabular handling for the EDA section
import matplotlib # plotting for the EDA section
# Print versions so that if anything misbehaves later, we can see exactly what's installed.
# This should match the pins from cell_0b.
print("torch :", torch.__version__)
print("transformers:", transformers.__version__)
print("datasets :", datasets.__version__)
print("pandas :", pd.__version__)
print("matplotlib :", matplotlib.__version__)
print("-" * 40)
# Did PyTorch actually find a CUDA GPU in this process?
# Remember: because of cell_0a, AT MOST ONE GPU is visible here (or none, if 'cpu' was chosen).
cuda_available = torch.cuda.is_available()
print("CUDA available to this process:", cuda_available)
# Safety reconciliation between what we ASKED for (DEVICE from cell_0a) and what we GOT.
# If we asked for cuda but CUDA isn't available, fall back to CPU loudly rather than crashing later.
if DEVICE.startswith("cuda") and not cuda_available:
print("WARNING: cuda was requested in cell_0a, but no CUDA GPU is visible. Falling back to CPU.")
DEVICE = "cpu"
# If we do have a GPU, print its name so the user sees which card they actually landed on.
if cuda_available and DEVICE.startswith("cuda"):
# Index 0 here is the *visible* GPU, i.e. the one we selected in cell_0a (it was remapped to 0).
print("Active GPU:", torch.cuda.get_device_name(0))
print("Final device for this notebook:", DEVICE)
[cell 5 code]
# cell_1 — central configuration (subsample toggle lives here)
# MAIN GOAL: Put every knob the user might want to turn in ONE place, near the top,
# so the rest of the notebook just reads these variables and never needs editing.
# The headline controls are the SUBSAMPLE toggle (how much data) and NUM_EPOCHS (how long
# we train). Everything downstream (splitting, EDA, fine-tuning, evaluation) reads from here.
# --- Subsample toggle ---------------------------------------------------------
# USE_FULL_DATASET:
# True -> use all of SQuAD (slow to fine-tune; this is our "real run" for proper numbers)
# False -> use only SUBSAMPLE_SIZE rows from the training split (fast, for time-tuning later)
# We start with the FULL dataset to get trustworthy results first; we can flip this to
# False afterwards to tune for speed without changing any other code.
USE_FULL_DATASET = False # flip to False later to do a fast subsampled run
# SUBSAMPLE_SIZE:
# How many TRAIN rows to keep when USE_FULL_DATASET is False.
# A few thousand keeps fine-tuning quick. Ignored entirely when USE_FULL_DATASET is True.
SUBSAMPLE_SIZE = 20
# VALID_SUBSAMPLE_SIZE:
# How many rows to keep (BEFORE the val/test split below) when subsampling.
# Ignored when USE_FULL_DATASET is True.
VALID_SUBSAMPLE_SIZE = 10
# --- Train epochs -------------------------------------------------------------
# NUM_EPOCHS:
# How many full passes over the training data during fine-tuning.
# We use more than 1 so that "pick the best epoch by validation loss" is meaningful:
# with only 1 epoch there would be nothing to choose between. The Trainer will evaluate
# validation loss after each epoch and we'll keep the best checkpoint.
#NUM_EPOCHS = 3 #-> for full run on full datasets
NUM_EPOCHS = 2
# --- Validation / Test split --------------------------------------------------
# SQuAD ships with only "train" and "validation" (no public "test"). To match the usual
# train/val/test setup, we split SQuAD's validation split into our OWN val and test sets.
# TEST_FRACTION = portion of SQuAD-validation that becomes our TEST set (the rest is val).
# We give TEST the larger share because test carries the headline before/after numbers
# we report at the end, while our validation set only does light per-epoch model picking.
TEST_FRACTION = 0.70 # 70% test, 30% validation
# --- Reproducibility ----------------------------------------------------------
# A fixed seed makes the "random" subsample AND the val/test split the same every run,
# so results and demo examples are stable. Predictability matters when teaching.
SEED = 42
# --- Model choice -------------------------------------------------------------
# "gpt2" is GPT-2 small (124M), our primary model.
# Swap to "distilgpt2" if a later fast run needs to be lighter (the brief's fallback).
MODEL_NAME = "gpt2"
# --- Prompt template ----------------------------------------------------------
# ONE template, reused identically in BOTH the zero-shot and post-fine-tuning sections
# so the before/after comparison is apples-to-apples. Defining it here guarantees
# we never accidentally use two slightly different prompts.
# {context} and {question} get filled in per example later.
PROMPT_TEMPLATE = "Context: {context}\nQuestion: {question}\nAnswer:"
# --- Quick echo so the user can SEE the active configuration ------------------
print("Configuration for this run")
print("-" * 40)
print("USE_FULL_DATASET :", USE_FULL_DATASET)
print("SUBSAMPLE_SIZE :", SUBSAMPLE_SIZE, "(ignored if USE_FULL_DATASET=True)")
print("VALID_SUBSAMPLE_SIZE:", VALID_SUBSAMPLE_SIZE, "(ignored if USE_FULL_DATASET=True)")
print("NUM_EPOCHS :", NUM_EPOCHS)
print("TEST_FRACTION :", TEST_FRACTION, "(share of SQuAD-validation used as TEST)")
print("SEED :", SEED)
print("MODEL_NAME :", MODEL_NAME)
print("PROMPT_TEMPLATE :")
print(PROMPT_TEMPLATE)
[cell 6 markdown]
### Reading the configuration
This run is set up as follows:
- **Full dataset.** `USE_FULL_DATASET` is `True`, so we train on all of SQuAD. The two subsample sizes are shown but ignored here. They only take effect when `USE_FULL_DATASET` is `False`, which is the faster mode we can switch to later.
- **Three training epochs.** `NUM_EPOCHS` is 3. We use more than one epoch so that picking the best epoch by validation loss is meaningful. With a single epoch there would be nothing to compare.
- **Test split of 70 percent.** `TEST_FRACTION` is 0.7, so 70 percent of SQuAD's validation split becomes our test set and the remaining 30 percent becomes our validation set. Test gets the larger share because it carries the final before and after numbers, while validation only does light per-epoch model selection.
- **Fixed seed.** `SEED` is 42. This keeps the random parts of the run repeatable, the data split, any shuffling, and training, so the results and the demo examples are the same every time.
- **Model.** `MODEL_NAME` is `gpt2`, the 124M parameter GPT-2 small model.
- **Prompt template.** Every example is formatted the same way, a context, then a question, then `Answer:` for the model to continue. We reuse this exact format before and after fine-tuning so the comparison stays fair.
Keeping all of these settings in one place means the rest of the notebook just reads from them. To change the run, you change this cell and nothing else.
[cell 7 code]
# cell_2 — load SQuAD, apply the subsample toggle, and build train/val/test splits
# MAIN GOAL: Load SQuAD and produce THREE splits the rest of the notebook will use:
# train_data -> fine-tuning (and EDA)
# validation_data -> per-epoch model selection during fine-tuning (pick best epoch)
# test_data -> final, held-out before/after numbers we report at the end
# SQuAD only ships "train" and "validation" (its real test set is private), so we create
# our own val/test by splitting SQuAD's validation split using TEST_FRACTION from cell_1.
# We also honor the subsample toggle. This cell only LOADS/SIZES/SPLITS — no tokenizing.
from datasets import load_dataset
# Download (first run) or load-from-cache (later runs) the SQuAD dataset.
# Returns a DatasetDict with two keys: "train" (~87k rows) and "validation" (~10k rows).
# Subsequent runs are instant because HuggingFace caches to ~/.cache/huggingface/datasets.
print("Loading SQuAD … (first run downloads it; later runs use the local cache)")
squad = load_dataset("squad")
print("\nRaw dataset structure (note: only 'train' and 'validation' exist):")
print(squad)
# Store references to the two original SQuAD splits before any resizing.
# We keep "full_" prefixes so it's always clear we haven't touched these yet.
full_train = squad["train"]
full_squad_validation = squad["validation"] # we will CUT this into our val + test
# --- Apply the subsample toggle from cell_1 ----------------------------------
if USE_FULL_DATASET:
# Keep everything (our "real run" for trustworthy numbers).
# No copying happens here — these are just variable aliases pointing at the same data.
train_data = full_train
squad_validation = full_squad_validation
print("\nUSE_FULL_DATASET=True -> keeping the entire dataset.")
else:
# Subsample for speed (used during development / quick iteration).
# min() guards against requesting more rows than the split actually contains.
n_train = min(SUBSAMPLE_SIZE, len(full_train))
n_valid = min(VALID_SUBSAMPLE_SIZE, len(full_squad_validation))
# Shuffle BEFORE selecting so we don't grab a non-representative block (SQuAD is
# ordered by article, so the first N rows would cover only a handful of topics).
# The fixed SEED makes the subsample identical across runs for reproducibility.
train_data = full_train.shuffle(seed=SEED).select(range(n_train))
squad_validation = full_squad_validation.shuffle(seed=SEED).select(range(n_valid))
print(f"\nUSE_FULL_DATASET=False -> subsampled to {n_train} train / "
f"{n_valid} (pre-split) validation rows.")
# --- Split SQuAD's validation into OUR validation_data + test_data -----------
# train_test_split is a built-in Dataset method. Despite the name, here it just divides
# squad_validation into two parts. test_size=TEST_FRACTION sends that share to "test".
# seed=SEED makes the partition identical every run (stable test numbers + demo examples).
# (Internally it shuffles before splitting, so both halves are representative.)
split = squad_validation.train_test_split(test_size=TEST_FRACTION, seed=SEED)
# Naming note: HuggingFace always returns keys "train" and "test" from this method,
# regardless of what the two parts will actually be used for in your project.
# Here the "train" portion becomes our validation set (the smaller slice we check
# each epoch) and the "test" portion becomes our held-out test set.
validation_data = split["train"] # the SMALLER part (1 - TEST_FRACTION) -> our validation
test_data = split["test"] # the LARGER part (TEST_FRACTION) -> our test
# --- Final readout of what we'll actually work with --------------------------
# Printing row counts here is a quick sanity check: if the numbers look wildly off
# (e.g. test_data is empty) it usually means TEST_FRACTION was set incorrectly in cell_1.
print("-" * 40)
print("train_data rows :", len(train_data), "(fine-tuning + EDA)")
print("validation_data rows :", len(validation_data), "(per-epoch model selection)")
print("test_data rows :", len(test_data), "(final before/after numbers)")
print("columns :", train_data.column_names)
[cell 8 markdown]
### Reading the dataset output
This is what the loading and splitting step produced:
- **SQuAD has two original splits.** The raw structure shows a `train` split with 87,599 rows and a `validation` split with 10,570 rows. There is no `test` split, which is why we make our own.
- **Each row has the same five fields.** Every example contains `id`, `title`, `context`, `question`, and `answers`. These are the same across all splits.
- **The full dataset is kept.** Because `USE_FULL_DATASET` is `True`, the train split stays at its full 87,599 rows for fine-tuning and EDA.
- **Validation was split into our own validation and test sets.** SQuAD's 10,570 validation rows were divided using `TEST_FRACTION` of 0.7. This gives 7,399 rows for test (the final before and after numbers) and 3,171 rows for validation (per-epoch model selection). The two add up to the original 10,570.
From here on, the notebook works with three clear splits: `train_data` for training, `validation_data` for picking the best epoch, and `test_data` for the final held-out results.
[cell 9 code]
# cell_3 — meet the data: inspect structure and print full examples
# MAIN GOAL: Before any charts, make the dataset CONCRETE by looking at a few complete
# rows. A beginner should walk away knowing exactly what one SQuAD example contains:
# - context : a paragraph of text (the passage to read)
# - question: a question whose answer is found in that context
# - answers : the answer text(s), plus where in the context they start
# We print 2-3 full examples in a readable layout. No statistics yet — that's the next cell.
# First, look at ONE raw example exactly as the dataset stores it (a Python dict).
# Seeing the raw dict demystifies the column names we printed in cell_2.
print("One raw example (as stored):")
print(train_data[0])
print("=" * 70)
# The raw dict is a bit noisy, so now print a FEW examples in a clean, labeled format.
# We loop over the first 3 rows and pretty-print each field.
NUM_EXAMPLES_TO_SHOW = 3 # small number; just enough to see the pattern
for i in range(NUM_EXAMPLES_TO_SHOW):
example = train_data[i] # grab the i-th row as a dict
# The 'answers' field is itself a dict with two parallel lists:
# 'text' -> the answer string(s)
# 'answer_start' -> the character index in the context where each answer begins
# In the SQuAD TRAIN split there is normally exactly one answer per question.
answer_texts = example["answers"]["text"]
answer_starts = example["answers"]["answer_start"]
print(f"EXAMPLE {i}")
print(f" Title : {example['title']}") # the Wikipedia article it came from
print(f" Context : {example['context']}") # the full passage
print(f" Question : {example['question']}") # the question to answer
print(f" Answer(s): {answer_texts}") # the gold answer text(s)
print(f" Starts at: {answer_starts} (character index into the context)")
print("-" * 70)
# A plain-language takeaway printed for the audience, reinforcing the key insight
# that motivates the whole tutorial: answers are SHORT spans, but we'll be GENERATING
# them as text with GPT rather than extracting their positions.
print("Takeaway: each row = a context paragraph + a question + a short answer.")
print("We will teach GPT to GENERATE that answer as text, not to point at its location.")
[cell 10 markdown]
### Reading the example output
These first few rows make the structure of one SQuAD example concrete:
- **The raw row is a dictionary.** Printed as stored, one example is a Python dict with the five fields `id`, `title`, `context`, `question`, and `answers`. The cleaned-up view below it shows the same information in a more readable layout.
- **Answers are short and taken from the context.** Each answer is just a few words, like "Saint Bernadette Soubirous" or "the Main Building", and each one appears word for word in the passage.
- **`answer_start` is a character index.** The "Starts at" value, such as 515, is the position in the context string where the answer begins. It is measured in characters, not words.
The key takeaway is the shape of the task. Every row is a long context, a question, and a short answer. We will teach the model to generate that answer as text, rather than to point at where it sits in the passage.
[cell 11 markdown]
# EDA
[cell 12 code]
# cell_4 — EDA 1: train vs validation counts
# MAIN GOAL: The gentlest possible first chart — just how many examples are in each
# split. The point is comfort, not insight: get the audience used to the pattern of
# "compute a number with pandas/plain Python, then draw it with matplotlib" before we
# move on to the busier length-histogram cells.
# NOTE: these counts reflect whatever the subsample toggle (cell_1) produced, so if you
# subsampled, you'll see your subsample sizes here, not full SQuAD's ~88k/~10k.
import matplotlib.pyplot as plt # the standard plotting interface
# Count rows in each split. len() on a Hugging Face dataset gives its number of rows.
split_names = ["train", "validation"]
split_counts = [len(train_data), len(validation_data)]
# Print the raw numbers too, so the chart is backed by something explicit.
print("Row counts:")
for name, count in zip(split_names, split_counts):
print(f" {name:10s}: {count}")
# --- Draw the bar chart -------------------------------------------------------
plt.figure(figsize=(5, 4)) # a small, uncluttered canvas
bars = plt.bar(split_names, split_counts) # one bar per split
# Write the exact count on top of each bar so the chart is readable at a glance
# without having to eyeball the y-axis.
for bar, count in zip(bars, split_counts):
plt.text(
bar.get_x() + bar.get_width() / 2, # x: horizontal center of the bar
bar.get_height(), # y: top of the bar
str(count), # the label text (the count)
ha="center", va="bottom" # center horizontally, sit just above the bar
)
plt.title("Number of examples per split") # plain, descriptive title
plt.ylabel("number of examples") # label the axis that carries meaning
plt.tight_layout() # keep labels from getting clipped
plt.show() # render it
[cell 13 markdown]
### Reading the split sizes
This first chart simply counts the rows in each split:
- **Train has 87,599 rows.** This is the full SQuAD training split, used for fine-tuning and EDA.
- **Validation has 3,171 rows.** This is our own validation set, the 30 percent slice of SQuAD's validation split that we use to pick the best epoch.
The test split is not shown here because this chart only compares the two sets used during training. The counts reflect whatever the configuration in the earlier cell produced, so if you switch to a subsampled run these numbers will change.
[cell 14 code]
# cell_5 — EDA 2: length distributions of context, question, and answer
# MAIN GOAL: Show how long contexts, questions, and answers typically are. The memorable
# surprise for beginners is that ANSWERS ARE VERY SHORT (usually just a few words), while
# contexts are long paragraphs. Seeing this builds intuition for why we only ask GPT to
# generate a short answer at the end of a long prompt.
# We measure length in WORDS (split on spaces) rather than characters or tokens, because
# "words" is the most intuitive unit for a beginner. (Tokenization comes later, separately.)
import matplotlib.pyplot as plt
# --- Compute lengths ----------------------------------------------------------
# For each row we take the relevant text and count words via simple .split().
# We use a plain Python list comprehension so the operation is easy to read.
# Note: answers["text"] is a LIST; SQuAD train has one answer, so we take element [0].
context_lengths = [len(row["context"].split()) for row in train_data]
question_lengths = [len(row["question"].split()) for row in train_data]
answer_lengths = [len(row["answers"]["text"][0].split()) for row in train_data]
# --- Print quick summary stats so the chart is backed by numbers --------------
# We use pandas just for its convenient .describe()-style summary via a small helper.
import pandas as pd
def quick_stats(name, values):
s = pd.Series(values)
# Show the typical (median), the average (mean), and the extremes (min/max).
print(f"{name:9s} -> median: {s.median():.0f}, mean: {s.mean():.1f}, "
f"min: {s.min()}, max: {s.max()}")
print("Length in words:")
quick_stats("context", context_lengths)
quick_stats("question", question_lengths)
quick_stats("answer", answer_lengths)
print("Notice how small the answer numbers are compared to the context.")
# --- Draw three histograms side by side ---------------------------------------
# One row, three columns: each subplot is the distribution of one field's lengths.
fig, axes = plt.subplots(1, 3, figsize=(14, 4))
# Context lengths (long; wide spread)
axes[0].hist(context_lengths, bins=40)
axes[0].set_title("Context length (words)")
axes[0].set_xlabel("words")
axes[0].set_ylabel("number of examples")
# Question lengths (medium; fairly tight)
axes[1].hist(question_lengths, bins=40)
axes[1].set_title("Question length (words)")
axes[1].set_xlabel("words")
# Answer lengths (short; clustered near the low end — the surprise)
axes[2].hist(answer_lengths, bins=40)
axes[2].set_title("Answer length (words)")
axes[2].set_xlabel("words")
plt.tight_layout() # space the three plots so titles/labels don't overlap
plt.show()
[cell 15 markdown]
### Reading the length distributions
These numbers measure length in words for the training data:
- **Contexts are long.** The median context is 110 words, with an average near 120 and some passages running past 600 words. These are full paragraphs.
- **Questions are medium and consistent.** The median question is 10 words and the average is about the same, so most questions are short and fairly uniform.
- **Answers are very short.** The median answer is just 2 words and the average is around 3. Even the longest answer, at 43 words, is small next to the contexts.
This is the main thing to take away. The model reads a long passage and a short question, then produces a very short answer. That imbalance, a long input and a tiny output, is exactly the shape of the task we are training for. The long tail in each distribution, such as the 653-word context or the 43-word answer, comes from a small number of outliers and is worth remembering when we later set a maximum sequence length.
[cell 16 code]
# cell_6 — EDA 3: question types by first word (the memorable chart)
# MAIN GOAL: Categorize each question by its FIRST WORD (what / who / when / where /
# why / how, plus an "other" bucket) and show how many of each there are. This is the
# most memorable EDA chart because it reveals the SHAPE of the dataset at a glance:
# SQuAD is dominated by "what" questions, with the other types trailing. It also gives
# the audience a mental model of the kinds of questions GPT will be asked to answer.
import matplotlib.pyplot as plt
from collections import Counter # convenient tally tool: counts occurrences
# The standard "wh" question words we want to track explicitly.
# Anything not in this list will be lumped into an "other" bucket.
# "which" is included even though it's rarer — omitting it would silently inflate "other".
WH_WORDS = ["what", "who", "when", "where", "why", "how", "which"]
def first_word_category(question):
# Lowercase so "What" and "what" count as the same thing,
# then take the first whitespace-separated token.
# .split() with no argument also strips extra whitespace and handles tabs/newlines.
# .split() on an empty/odd string could yield nothing, so guard with a fallback.
tokens = question.lower().split()
if not tokens:
return "other"
first = tokens[0]
# If the first word is one of our tracked wh-words, use it as the category;
# otherwise bucket it as "other" (e.g. questions starting with "in", "the", a name…).
# Using a set lookup here would be O(1) vs O(n) for a list, but WH_WORDS is so
# small (7 items) that the difference is negligible — readability wins.
return first if first in WH_WORDS else "other"
# Apply the classifier to every question in one pass using a list comprehension.
# train_data is a HuggingFace Dataset, so iterating over it yields one dict per row,
# each with keys: "id", "title", "context", "question", "answers".
categories = [first_word_category(row["question"]) for row in train_data]
# Counter produces a dict-like object: {"what": 23k, "who": 5k, ...}.
# Missing keys return 0 (via .get() below), so categories with zero questions are safe.
counts = Counter(categories)
# Decide a sensible, STABLE display order: our wh-words in their listed order,
# then "other" at the end. This keeps the chart consistent run-to-run instead of
# ordering bars by whatever count happens to be largest.
# .get(label, 0) handles the edge case where a wh-word never appears in the data.
ordered_labels = WH_WORDS + ["other"]
ordered_values = [counts.get(label, 0) for label in ordered_labels] # 0 if a type never appears
# Print the tally so the numbers behind the bars are explicit.
print("Question counts by first word:")
for label, value in zip(ordered_labels, ordered_values):
print(f" {label:6s}: {value}")
# :6s left-pads the label to 6 characters so the counts align in a column.
# --- Draw the bar chart -------------------------------------------------------
plt.figure(figsize=(8, 4))
bars = plt.bar(ordered_labels, ordered_values)
# plt.bar returns a list of Rectangle objects — one per bar — which we iterate
# below to position the count labels. The order matches ordered_labels exactly.
# Put the exact count on top of each bar so readers don't have to squint at the y-axis.
# bar.get_x() + bar.get_width()/2 finds the horizontal centre of the bar.
# bar.get_height() is the bar's top edge, which equals the count value.
# va="bottom" places the text just above that edge rather than hanging below it.
for bar, value in zip(bars, ordered_values):
plt.text(
bar.get_x() + bar.get_width() / 2,
bar.get_height(),
str(value),
ha="center", va="bottom"
)
plt.title("Question types by first word")
plt.xlabel("first word of question")
plt.ylabel("number of questions")
plt.tight_layout() # prevents axis labels from being clipped at the figure edge
plt.show()
# Plain-language takeaway reinforcing the insight.
print("Takeaway: SQuAD is heavily 'what'-driven, with who/when/where/how making up most of the rest.")
[cell 17 markdown]
### Reading the question types
Grouping each question by its first word shows the shape of the dataset:
- **"What" dominates.** With 37,593 questions, "what" alone makes up far more than any other type, close to half of the training data.
- **A middle tier of common types.** "Who" (8,150), "how" (8,124), "when" (5,459), "which" (4,159), and "where" (3,291) each appear in meaningful numbers.
- **"Why" is rare.** At 1,201 questions, "why" is the least common of the standard question words, which fits the fact that SQuAD answers are short factual spans rather than explanations.
- **A large "other" bucket.** 19,622 questions do not start with one of the tracked words. These begin with things like names, prepositions, or articles, for example "In what year..." or "The team that...".
The takeaway is that SQuAD is heavily focused on "what" questions, with who, how, when, where, and which making up most of the rest. This gives a sense of the kinds of questions the model will be asked to answer.
[cell 18 code]
# cell_7 — EDA 4 (optional): most common words in questions
# MAIN GOAL: Look at which words appear most often across all questions. Done naively,
# the top of the list is boring filler ("the", "of", "in", "is"…) — which is itself a
# useful lesson about natural language. So we show TWO lists side by side:
# (a) raw most-common words (dominated by filler/stopwords)
# (b) most-common words AFTER removing a small stopword list (more meaningful)
# This teaches, in one cell, why people bother filtering stopwords before drawing
# conclusions from word counts.
import matplotlib.pyplot as plt
from collections import Counter
# A SMALL, hand-written stopword list. We keep it short and visible (rather than pulling
# in a big NLP library) so a beginner can see exactly what is being filtered and why.
# These are high-frequency function words that carry little topical meaning.
STOPWORDS = {
"the", "a", "an", "of", "to", "in", "on", "for", "and", "or", "is", "are",
"was", "were", "did", "does", "do", "what", "which", "who", "when", "where",
"why", "how", "that", "this", "with", "by", "at", "as", "from", "be", "it",
}
# Tokenize every question into lowercase words.
# We strip basic punctuation by keeping only alphabetic characters per token, so
# "city?" becomes "city" and doesn't count as its own separate word.
all_words = []
for row in train_data:
for token in row["question"].lower().split():
cleaned = "".join(ch for ch in token if ch.isalpha()) # drop digits/punctuation
if cleaned: # skip tokens that became empty
all_words.append(cleaned)
# (a) Raw counts — includes stopwords.
raw_counts = Counter(all_words)
# (b) Filtered counts — same words, minus anything in our STOPWORDS set.
filtered_counts = Counter(w for w in all_words if w not in STOPWORDS)
# How many top words to display in each list/chart.
TOP_N = 15
# Pull the top-N for each version. Counter.most_common returns (word, count) pairs
# already sorted from most to least frequent.
raw_top = raw_counts.most_common(TOP_N)
filtered_top = filtered_counts.most_common(TOP_N)
# Print both lists so the contrast is explicit in text, not just in the chart.
print(f"Top {TOP_N} words (RAW, includes stopwords):")
print(" " + ", ".join(f"{w}({c})" for w, c in raw_top))
print(f"\nTop {TOP_N} words (FILTERED, stopwords removed):")
print(" " + ", ".join(f"{w}({c})" for w, c in filtered_top))
# --- Draw the two as horizontal bar charts side by side -----------------------
# Horizontal bars (barh) are easier to read when the labels are words.
fig, axes = plt.subplots(1, 2, figsize=(14, 6))
# Helper to plot one ranked list into one subplot.
def plot_top(ax, pairs, title):
words = [w for w, c in pairs]
counts = [c for w, c in pairs]
# Reverse so the MOST frequent word lands at the TOP of the horizontal chart.
ax.barh(words[::-1], counts[::-1])
ax.set_title(title)
ax.set_xlabel("count")
plot_top(axes[0], raw_top, "Most common (raw)")
plot_top(axes[1], filtered_top, "Most common (stopwords removed)")
plt.tight_layout()
plt.show()
# Takeaway tying it together.
print("\nTakeaway: raw counts are mostly filler words; removing stopwords reveals the")
print("topical words questions actually ask about (e.g. 'name', 'year', 'city', 'first').")
[cell 19 markdown]
### Reading the common words
Comparing the two word lists shows why stopword removal matters:
- **The raw list is mostly filler.** The most frequent words are "the", "what", "of", "in", "to", and similar function words. They appear constantly in English but say nothing about what the questions are actually about.
- **The filtered list is more meaningful.** After removing stopwords, the top words become "many", "year", "first", "name", "type", "city", and "people". These point to the real content of the questions.
- **"Many" near the top hints at counting questions.** Its high rank reflects the many "how many" questions in the dataset, which ask for numbers.
- **The other top words suggest common themes.** Words like "year", "name", "city", and "people" line up with the factual nature of SQuAD, where questions often ask about dates, names, places, and who was involved.
The takeaway is that raw counts are dominated by filler words, and removing a small list of stopwords reveals the topical words that questions are really built around.
[cell 20 code]
print("Done so far:")
print(" - Setup: GPU selection, pinned installs, imports, device check")
print(" - Concepts: language models predict text; zero-shot vs fine-tuning")
print(" - EDA: split sizes, length distributions, question types, common words")
print()
print("Key things we learned about the data:")
print(" - Each example = long CONTEXT + short QUESTION + very short ANSWER")
print(" - SQuAD is dominated by 'what' questions")
print(" - Answers are typically just a few words long")
print()
print("Coming up next:")
print(" 1. Zero-shot QA — ask GPT-2 questions with NO training, watch it struggle")
print(" 2. Fine-tuning — train GPT-2 on our SQuAD subsample (1 epoch)")
print(" 3. Before/after — re-run the SAME prompts and compare")
print()
print("Reminder: everything below uses the configuration from cell_1")
print(f" model = {MODEL_NAME} | full dataset = {USE_FULL_DATASET} | "
f"train rows = {len(train_data)}")
print("=" * 70)
[cell 21 markdown]
## The GPT-2 Model and Its Tokenizer
### About GPT-2
GPT-2 is a language model released by OpenAI. At its core it does one simple thing: given a sequence of text, it predicts the next word. By doing this over and over, it can generate longer passages one piece at a time.
It was trained on a large amount of text from the internet, with no specific task in mind. This is why it can produce fluent English but does not, on its own, know how to follow a particular format like question answering. Teaching it that format is what fine-tuning does later.
GPT-2 comes in several sizes. We use the smallest, often called GPT-2 small, which has about 124 million parameters. The larger versions are more capable but slower and heavier. The small model is a good fit for a tutorial because it trains quickly and still clearly shows the before and after effect of fine-tuning.
### What a tokenizer does
A model does not read raw text. It reads numbers. The tokenizer is the piece that converts between the two:
- It breaks text into units called tokens.
- It maps each token to an integer id the model understands.
- It can also turn those ids back into text, which is how we read the model's output.
Tokens are not the same as words. The tokenizer uses subword pieces, so a common word might be a single token while a rarer word gets split into several. As a rough guide, token counts run a little higher than word counts for English text. This is why we measured token lengths separately before choosing a maximum sequence length.
### The tokenizer we use
We load the tokenizer that was built for GPT-2, using `AutoTokenizer.from_pretrained(MODEL_NAME)`. Using the model's own matching tokenizer is important, because the model only understands the exact token ids it was trained with. A mismatched tokenizer would feed it meaningless numbers.
GPT-2's tokenizer is a byte-level BPE (byte pair encoding) tokenizer. Two practical details are worth knowing:
- It encodes spaces as part of the token. A word at the start of a sentence and the same word in the middle can be different tokens, because one carries a leading space and the other does not. This is why we are careful about spacing when we format our prompts and answers.
- It has a special end-of-text token that marks where a piece of text ends.
### One adjustment we make
By default, GPT-2's tokenizer has no padding token. Padding is needed later so that sequences of different lengths can be batched together into a uniform shape. The standard fix, which we use, is to set the padding token to be the same as the end-of-text token.
This is a common and safe choice for GPT-2. We just have to remember that padding and the real end-of-text marker share the same id, so later we rely on the attention mask to tell genuine content apart from padding.
[cell 22 markdown]
## What is Zero-Shot?
Zero-shot means asking the model to do a task without giving it any training or examples for that task first. "Zero" refers to zero training examples.
- We take GPT-2 exactly as it comes, already trained on general text, and simply ask it our questions.
- It has never seen SQuAD and has never been shown what a good answer looks like in our format.
- We just hand it the prompt and see what it produces.
The point is to get a baseline. This shows us what the model can do on its own, before any fine-tuning. Whatever it scores here is the starting line we will try to beat.
We expect it to struggle, and that is fine. A weak zero-shot result is exactly what motivates fine-tuning in the next step.
[cell 23 markdown]
# Zero-shot
[cell 24 code]
# cell_9 — load GPT-2 model and tokenizer (still untrained / zero-shot)
# MAIN GOAL: Load the pretrained GPT-2 (the exact MODEL_NAME from cell_1) and its
# matching tokenizer, move the model onto the DEVICE we picked in cell_0a/0c, and put it
# in evaluation mode. This is the SAME pretrained model the brief calls "zero-shot": it
# has general language ability from its original training, but has seen NONE of our SQuAD
# data yet. We will ask it questions as-is in the next cell.
#
# One important GPT-2 quirk handled here: GPT-2 ships WITHOUT a padding token. We need a
# pad token later (for batching during generation/fine-tuning), so we set the pad token
# to be the same as the end-of-sequence (eos) token — the standard fix for GPT-2.
from transformers import AutoModelForCausalLM, AutoTokenizer
# --- Tokenizer ----------------------------------------------------------------
# The tokenizer converts text <-> token IDs. AutoTokenizer picks the right one for
# MODEL_NAME automatically. "Causal LM" = standard left-to-right GPT-style generation.
print(f"Loading tokenizer for '{MODEL_NAME}' …")
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
# GPT-2 has no dedicated pad token. Reuse the end-of-sequence token as padding.
# Without this, batched generation and the Trainer's collation can error out.
tokenizer.pad_token = tokenizer.eos_token
# --- Model --------------------------------------------------------------------
print(f"Loading model '{MODEL_NAME}' …")
model = AutoModelForCausalLM.from_pretrained(MODEL_NAME)
# Tell the model which token id means "padding", keeping it consistent with the tokenizer.
# Sync model config with tokenizer so generation and loss code read the same id.
# (Some generation/loss code reads this from the model config.)
model.config.pad_token_id = tokenizer.pad_token_id
# Move the model's weights onto our chosen device (GPU if available, else CPU).
model.to(DEVICE)
# Put the model in EVALUATION mode. This disables training-only behaviors like dropout,
# giving stable, repeatable outputs — appropriate for the zero-shot demo.
model.eval()
# --- Confirmation readout -----------------------------------------------------
# Parameter count gives a tangible sense of model size (~124M for gpt2 small).
num_params = sum(p.numel() for p in model.parameters())
print("-" * 40)
print("Model loaded.")
print(f" model name : {MODEL_NAME}")
print(f" parameters : {num_params:,}")
print(f" on device : {next(model.parameters()).device}")
print(f" pad == eos token: '{tokenizer.pad_token}' (id {tokenizer.pad_token_id})")
[cell 25 markdown]
### Reading the model load output
This confirms the model and tokenizer are ready:
- **GPT-2 small is loaded.** The model name is `gpt2` and it has 124,439,808 parameters, about 124 million, which matches the small version we chose.
- **It is on the GPU.** The device shows `cuda:0`, meaning the model's weights are on the selected GPU and training and generation will run there.
- **Padding is set to the end-of-text token.** Both share the token `<|endoftext|>` with id 50256. This is the fix for GPT-2 having no padding token by default.
One detail to carry forward: because padding and the real end-of-text marker are the same token, we cannot tell them apart by id alone. Later we rely on the attention mask to know which positions are genuine content and which are padding.
[cell 26 code]
# cell_10 — the answer-generation function (reused before AND after fine-tuning)
# MAIN GOAL: Define ONE function that takes a context + question, fills the shared
# PROMPT_TEMPLATE from cell_1, runs the model's text generation, and returns only the
# generated ANSWER (the text after "Answer:"). Building this as a single reusable function
# is what guarantees the before/after comparison is apples-to-apples: the zero-shot demo
# and the post-fine-tuning demo both call this exact same code with the exact same prompt.
#
# Important framing: right now `model` is the untrained GPT-2, so answers will likely be
# rambling/repetitive. That weakness is the intended lesson, not a bug.
import torch
def generate_answer(context, question, max_new_tokens=30):
# --- 1. Build the prompt ---------------------------------------------------
# Fill the shared template so every call uses the identical format the model will
# later be fine-tuned on. {context} and {question} are substituted; the string ends
# with "Answer:" so the model's job is to continue it with the answer.
prompt = PROMPT_TEMPLATE.format(context=context, question=question)
# --- 2. Tokenize -----------------------------------------------------------
# Convert the prompt text into token IDs the model understands.
# return_tensors="pt" gives PyTorch tensors; .to(DEVICE) puts them on the same
# device as the model so they can be processed together.
inputs = tokenizer(prompt, return_tensors="pt").to(DEVICE)
# Remember how many tokens the PROMPT was, so we can later slice it off and keep
# ONLY the newly generated answer tokens (not the echoed-back prompt).
prompt_length = inputs["input_ids"].shape[1]
# --- 3. Generate -----------------------------------------------------------
# torch.no_grad() disables gradient tracking: we're only doing inference here,
# so this saves memory and time.
with torch.no_grad():
output_ids = model.generate(
**inputs, # the tokenized prompt
max_new_tokens=max_new_tokens, # cap how much new text to produce (answers are short)
do_sample=False, # greedy/deterministic: same prompt -> same answer
pad_token_id=tokenizer.pad_token_id, # silences a warning; uses our eos-as-pad token
)
# --- 4. Keep only the NEW tokens ------------------------------------------
# output_ids contains prompt tokens + generated tokens. Slice off the prompt portion
# using prompt_length so we're left with just what the model added.
generated_ids = output_ids[0][prompt_length:]
# --- 5. Decode back to text ------------------------------------------------
# Turn the generated token IDs back into a human-readable string.
# skip_special_tokens=True drops things like the eos token from the output.
answer = tokenizer.decode(generated_ids, skip_special_tokens=True)
# .strip() trims leading/trailing whitespace/newlines for a clean result.
return answer.strip()
# --- Tiny smoke test so we see the function works (and how poor zero-shot is) --
# We borrow the first validation example just to exercise the function once.
_demo = validation_data[0]
_demo_answer = generate_answer(_demo["context"], _demo["question"])
print("PROMPT TEMPLATE IN USE:")
print(PROMPT_TEMPLATE)
print("=" * 70)
print("Question :", _demo["question"])
print("Gold answer :", _demo["answers"]["text"][0]) # the true answer, for comparison
print("GPT-2 says :", repr(_demo_answer)) # repr() so we can see stray newlines/spaces
[cell 27 markdown]
### Reading the zero-shot test
This is a first look at the untrained model answering one question:
- **The prompt format is in use.** The context, question, and `Answer:` template is filled in and passed to the model, exactly as it will be everywhere else.
- **The gold answer is short.** The correct answer is simply "Last Glacial Maximum".
- **The model rambles and misses.** GPT-2 picks up the abbreviation "LGM" and even echoes the prompt, but it never says what it stands for. Instead it drifts into invented text about vegetation loss in the Amazon basin and trails off mid-sentence.
This is the behavior we expect from a model that has never been trained on this task. It has the language ability to continue the text fluently, but it does not know that the answer should be a short, clean span, and here it does not even surface the right fact. That gap is the whole motivation for fine-tuning. Note that the model is using greedy decoding, so the same prompt always produces the same output, which keeps this demo stable.
[cell 28 code]
# cell_11 — zero-shot QA demo on a FIXED set of examples (from the TEST set)
# MAIN GOAL: Run the untrained GPT-2 (via generate_answer from cell_10) over a small,
# FIXED set of TEST examples and display question / gold answer / model answer together.
# Two things matter here:
# 1. We FREEZE the chosen examples into demo_examples now, so the after-fine-tuning cell
# reuses the EXACT same ones — apples-to-apples.
# 2. We SAVE these zero-shot answers (zero_shot_answers) so the final before/after cell
# can place "before" and "after" side by side without re-running this.
# We draw from test_data so the qualitative demo and the quantitative scores (cell_11b)
# are on the SAME held-out set. Expect weak/rambling answers — that IS the motivation
# for fine-tuning.
# How many examples to showcase. Small, because we'll read each one with the group.
NUM_DEMO = 5
# Freeze the demo set: take the first NUM_DEMO rows of test_data.
# test_data was produced by a seeded split in cell_2, so these are fixed, reproducible,
# and representative — the same every run, and a subset of what cell_11b scores.
demo_examples = [test_data[i] for i in range(NUM_DEMO)]
# We'll collect the model's zero-shot answers here, aligned with demo_examples by index.
zero_shot_answers = []
print("ZERO-SHOT QA (GPT-2 has NOT been trained on SQuAD)")
print("=" * 70)
# Loop over the frozen demo examples and ask the untrained model each question.
for i, ex in enumerate(demo_examples):
context = ex["context"]
question = ex["question"]
gold = ex["answers"]["text"][0] # first gold answer, for a readable display
# Generate the model's answer using the shared function (same one we'll reuse later).
predicted = generate_answer(context, question)
# Store it so the before/after cell can reuse it without recomputing.
zero_shot_answers.append(predicted)
# Short context preview (full contexts are long); enough to recall the passage.
context_preview = context[:160] + ("…" if len(context) > 160 else "")
print(f"[Example {i}]")
print(f" Context (preview): {context_preview}")
print(f" Question : {question}")
print(f" Gold answer : {gold}")
print(f" GPT-2 (zero-shot): {repr(predicted)}") # repr() exposes stray newlines/repetition
print("-" * 70)
print("\nObserve: answers are often off-topic, repetitive, or just continue the text.")
print("This is expected. Fine-tuning (next) is what teaches GPT-2 the QA *format* and task.")
[cell 29 markdown]
### Reading the zero-shot results
Running the untrained model on five test questions shows a clear pattern of failure:
- **It does not give short answers.** In every case the model keeps writing full sentences instead of the short span the question calls for, such as "30" or "CBS".
- **It often misses the answer entirely.** For the examination boards question the gold answer is "30", but the model never gives a number and just loops a vague sentence. For the ABC rumors question it invents a name, "Bob Odenkirk", instead of the correct "Caris & Co.".
- **It repeats itself.** Example 0 loops the same sentence about examinations again and again. This is a common failure of greedy generation with an untrained model.
- **It continues the pattern instead of answering.** In example 3 the model does say "CBS", but then keeps going and generates its own next "Question:" and "Answer:", showing it has learned the shape of the text without learning to stop.
- **It sometimes drifts into invented detail.** The University of Chicago answer is fluent and confident but factually made up, which is a preview of the hallucination problem we discuss later.
The takeaway is that GPT-2 can produce fluent text and occasionally lands near a relevant fact, but it does not understand the question answering task or format. It rambles, repeats, invents, and does not know when to stop. Fine-tuning is what teaches it to read the context and return a short, correct answer.
[cell 30 markdown]
## How We Measure the Answers
We use the two standard SQuAD metrics: Exact Match and F1. Both compare the model's answer to the correct (gold) answer, but they are strict in different ways.
### First, we clean up the text
Before comparing, we normalize both answers the same way, so trivial differences do not count as mistakes:
- Lowercase everything, so "Paris" and "paris" match.
- Remove punctuation.
- Remove the small words "a", "an", and "the".
- Collapse extra spaces.
This way "The Eiffel Tower" and "eiffel tower" are treated as the same answer.
### Exact Match (EM)
- The strict one. It asks: after cleanup, is the answer exactly equal to the gold answer?
- It scores 1 if they match completely, and 0 otherwise. There is no partial credit.
- Example: gold is "1852". If the model says "1852" it scores 1. If it says "in 1852" it scores 0, because it is not an exact match.
### F1
- The forgiving one. It gives partial credit based on how many words the two answers share.
- It balances two things:
- Precision: of the words the model said, how many were correct?
- Recall: of the words in the gold answer, how many did the model recover?
- F1 combines these into a single score, high only when both are high.
- Example: gold is "copper statue of Christ". If the model says "a copper statue", it shares several words, so EM would be 0 but F1 would still be fairly high.
- This is why F1 is fairer for a model that generates text, since it rewards being close even when not word-perfect.
### Handling multiple correct answers
- In the test data, a question can have several acceptable answers from different annotators.
- For each metric, we score the model against every gold answer and keep the best one.
- This avoids marking a correct answer wrong just because it matched a different annotator's wording.
### Why we use both
- EM tells us how often the model is exactly right.
- F1 tells us how close it is even when not perfect.
- Together they give a fuller picture. A model can have low EM but decent F1, which means it is on the right track but not precise.
[cell 31 code]
# cell_11b — quantitative zero-shot evaluation (Exact Match + F1) on the TEST set
# MAIN GOAL: Put a NUMBER on how good (bad) zero-shot GPT-2 is, computed on our held-out
# TEST set, so the before/after later is "X% -> Y%" on data the model never trained on.
# We compute the two standard SQuAD metrics:
# - Exact Match (EM): did the prediction, after light cleanup, exactly equal a gold answer?
# - F1: token-overlap score giving PARTIAL credit (fairer for generated text).
# IMPORTANT FIX vs. a naive version: SQuAD questions can have SEVERAL acceptable gold
# answers (different annotators). The official metric scores against ALL of them and takes
# the BEST (max). We do that here, so we don't unfairly mark a correct answer wrong just
# because it matched annotator #2 instead of #1.
# We build these as REUSABLE functions and store results (zero_shot_em / zero_shot_f1)
# for the final comparison cell. Expect LOW numbers — that's the motivation for fine-tuning.
import re # regular expressions, used to strip out articles
import string # gives us the list of punctuation characters
from collections import Counter # counts word occurrences, used for token overlap in F1
from tqdm.auto import tqdm # progress bar; .auto picks notebook vs terminal style
# How many TEST examples to score. Each is its own generate() call, so the full test set
# takes a while — the progress bar below shows count, percent, and estimated time left.
######################
######################
#NUM_EVAL = 50 # small number for quick development runs
NUM_EVAL = len(test_data) # set to full test set for the final run
# --- Text normalization (the official SQuAD way) ------------------------------
# Clean prediction and gold the SAME way so trivial differences (case, punctuation,
# articles, extra spaces) don't count as wrong.
def normalize_text(s):
s = s.lower() # case-insensitive: "Paris" == "paris"
s = "".join(ch for ch in s if ch not in string.punctuation) # keep only non-punctuation characters
s = re.sub(r"\b(a|an|the)\b", " ", s) # remove the articles a / an / the
s = " ".join(s.split()) # split on whitespace and rejoin: collapses extra spaces
return s # cleaned string, ready to compare
# --- EM / F1 for a single (prediction, ONE gold) pair -------------------------
def exact_match_score(prediction, gold):
# Normalize both sides, then check if they are identical.
# float(...) turns True/False into 1.0/0.0 so we can average it later.
return float(normalize_text(prediction) == normalize_text(gold))
def f1_score_single(prediction, gold):
# F1 = harmonic mean of precision and recall over shared WORDS (tokens).
# First, normalize each side and split into a list of word tokens.
pred_tokens = normalize_text(prediction).split()
gold_tokens = normalize_text(gold).split()
# Edge case: if either side is empty after normalization, there are no words to overlap.
# We say F1 is 1 only if BOTH are empty (they "agree" on emptiness), otherwise 0.
if len(pred_tokens) == 0 or len(gold_tokens) == 0:
return float(pred_tokens == gold_tokens)
# Count how many tokens the two answers share.
# Counter(...) & Counter(...) keeps the MINIMUM count of each common word,
# so duplicates are handled correctly (e.g. "the the" vs "the" shares only one "the").
common = Counter(pred_tokens) & Counter(gold_tokens)
num_same = sum(common.values()) # total number of overlapping word occurrences
# No shared words at all means no overlap, so F1 is 0.
if num_same == 0:
return 0.0
# Precision: of the words the MODEL produced, what fraction were correct?
precision = num_same / len(pred_tokens)
# Recall: of the words in the GOLD answer, what fraction did the model recover?
recall = num_same / len(gold_tokens)
# F1 combines the two; it is high only when BOTH precision and recall are high.
return 2 * precision * recall / (precision + recall)
# --- Score against ALL gold answers, take the best (official SQuAD behavior) --
# gold_answers is the LIST ex["answers"]["text"]; we return the max EM and max F1
# the prediction achieves against any single acceptable answer.
def best_em_f1(prediction, gold_answers):
# Guard: if somehow there are no gold answers, treat as a single empty string
# so the code below doesn't crash on an empty list.
if len(gold_answers) == 0:
gold_answers = [""]
# Score the prediction against EACH acceptable answer, then keep the best result.
# This gives the model credit if it matches ANY annotator's wording.
em = max(exact_match_score(prediction, g) for g in gold_answers)
f1 = max(f1_score_single(prediction, g) for g in gold_answers)
return em, f1
# --- Reusable evaluation loop over a dataset ----------------------------------
# Returns average EM and average F1 (percentages) over the first `n` rows.
# We pass `answer_fn` so the SAME loop scores zero-shot now and the fine-tuned model
# later — apples-to-apples, just like our shared prompt template.
def evaluate(dataset, n, answer_fn):
em_total = 0.0 # running sum of EM scores
f1_total = 0.0 # running sum of F1 scores
n = min(n, len(dataset)) # don't ask for more rows than the dataset has
# tqdm wraps the loop to show a live progress bar (count, %, and ETA), so a long
# full-test-set run isn't a silent black box. desc= labels the bar.
for i in tqdm(range(n), desc="Scoring"):
ex = dataset[i] # one example (context, question, answers, ...)
gold_answers = ex["answers"]["text"] # the FULL list of acceptable answers
pred = answer_fn(ex["context"], ex["question"]) # generate the model's answer
em, f1 = best_em_f1(pred, gold_answers) # best score over all gold answers
em_total += em # accumulate
f1_total += f1
# Divide by the number of examples to get the average, then scale to a percentage.
return 100.0 * em_total / n, 100.0 * f1_total / n
# --- Run the zero-shot evaluation on the TEST set -----------------------------
print(f"Scoring zero-shot GPT-2 on {NUM_EVAL} TEST examples … (one gen call each)")
# generate_answer is our shared answer function (cell_10); here it uses the UNTRAINED model.
zero_shot_em, zero_shot_f1 = evaluate(test_data, NUM_EVAL, generate_answer)
print("-" * 40)
print(f"ZERO-SHOT Exact Match : {zero_shot_em:.1f}%")
print(f"ZERO-SHOT F1 : {zero_shot_f1:.1f}%")
print("-" * 40)
print("These are our BEFORE numbers (on held-out test data). We'll compute the AFTER")
print("numbers with the identical evaluate() function once fine-tuning is done.")
[cell 32 markdown]
### Reading the zero-shot scores
These are the baseline numbers for the untrained model, measured on 50 held-out test examples:
- **Exact Match is 0 percent.** The model never produced an answer that exactly matched a gold answer after cleanup. Given how much it rambles, this is expected.
- **F1 is 8.5 percent.** F1 gives partial credit for overlapping words, so this small nonzero value reflects the occasional case where the model's output happens to share a word or two with the correct answer.
Together these are our "before" numbers. They are deliberately low, and that is the point. They give us a clear baseline to improve on. After fine-tuning we will run the exact same scoring function on the same test set, so the "after" numbers are directly comparable.
One honest note: part of the low F1 comes from the model being long winded. All those extra words drag the score down on top of the answers being wrong. After fine-tuning the model should become both more accurate and more concise, and both of those will push the score up.
[cell 33 markdown]
# Fine tuning
[cell 34 markdown]
## How We Fine-Tune the Model
### The basic idea
Fine-tuning means taking the already-trained GPT-2 and continuing its training, this time on our own data. The model already knows English. We are teaching it one specific skill: read a context and a question, then write a short answer.
### What the model learns from
We turn every SQuAD example into a single line of text in our standard format:
Context: <the passage>
Question: <the question>
Answer: <the correct answer><end-of-text>
The model trains by reading these lines and learning to predict them. Over many examples, it picks up the pattern: after "Answer:", a short answer should follow, and then the text should stop.
### How the model actually learns
A few simple points cover the mechanics:
- GPT-2 works by predicting the next token, over and over.
- During training, it compares its prediction to the real next token and measures how wrong it was. This number is the loss.
- Training nudges the model's weights to make the loss smaller, so its predictions get closer to the real text.
- Do this across the whole dataset and the model gradually learns the task.
### We only score the answer
This is an important choice in our setup:
- We do not want the model to waste effort learning to reproduce the long passage.
- So we only count the answer part when measuring the loss.
- The context and question are shown to the model as input, but they are ignored when grading its predictions.
- We also include the end-of-text marker in the answer, which teaches the model to stop instead of rambling.
### Training in epochs
- One epoch is one full pass over the training data.
- We train for three epochs so the model sees the data more than once.
- After each epoch, we check the model's loss on the validation set.
- We keep the version from the best epoch, the one with the lowest validation loss, rather than just the last one.
### Why three splits matter here
- The train set is what the model learns from.
- The validation set is used only to pick the best epoch. The model never learns from it.
- The test set is kept completely separate and used only at the end, to fairly measure how much the model improved.
### What to expect
- Before fine-tuning, the model rambled and rarely gave the right answer.
- After fine-tuning, it should give short, direct answers that match the format.
- We will prove this by running the exact same questions and the same scoring as before, and comparing the numbers.
[cell 35 code]
# cell_12 — format each example into a single training string
# MAIN GOAL: Convert each SQuAD row into ONE flat piece of text the model learns from.
# For causal-LM fine-tuning, training data is just text and the model learns to predict
# each next token. So we build the SAME prompt as the zero-shot demo, but now APPEND the
# gold answer after "Answer:" (during training we WANT the model to see the correct
# completion). We also append the end-of-sequence (eos) token so the model learns where an
# answer should STOP — this directly fights the rambling we measured in zero-shot.
#
# The prefix here is identical to PROMPT_TEMPLATE (cell_1), so what the model trains on
# matches what we prompt it with later. We format train AND validation (validation text is
# used to compute per-epoch validation loss for picking the best epoch). We do NOT format
# test_data — test is only ever scored via generation in evaluate(), never trained on.
# This cell only builds TEXT strings; tokenization happens next.
def build_training_text(example):
# Recreate the exact prompt prefix used everywhere else (no answer yet).
# .format() fills the {context} and {question} slots in PROMPT_TEMPLATE for this row.
prompt = PROMPT_TEMPLATE.format(
context=example["context"],
question=example["question"],
)
# The gold answer string. answers["text"] is a LIST; training uses a SINGLE target
# answer, so take element [0].
# (Multiple gold answers only matter for SCORING, which we do on val/test, not here.)
answer = example["answers"]["text"][0]
# Full training sequence = prompt + space + answer + eos token.
# - The space after "Answer:" matches natural text spacing (and gives a clean token
# boundary we rely on when masking in the next cell).
# - tokenizer.eos_token is the end-of-text marker; appending it teaches the model to
# STOP after a short answer instead of continuing forever.
full_text = f"{prompt} {answer}{tokenizer.eos_token}"
# Return a dict with a new "text" field. .map() (below) adds this as a new column.
return {"text": full_text}
# Apply the formatter to train and validation.
# .map() runs build_training_text on each row and adds the new "text" field to every row.
train_formatted = train_data.map(build_training_text)
valid_formatted = validation_data.map(build_training_text) # used for validation LOSS later
# --- Show what a finished training string looks like --------------------------
# Printing one example makes the abstract "format into a sequence" idea concrete.
print("Example of ONE training string the model will learn from:")
print("=" * 70)
# repr() shows the string with its escape characters visible, so we can actually SEE
# the newlines (\n) and the eos token rather than having them render invisibly.
print(repr(train_formatted[0]["text"]))
print("=" * 70)
print("\nNotice:")
print(" - It starts with the SAME 'Context/Question/Answer:' format as our prompts.")
print(" - The correct answer is included right after 'Answer:'.")
print(f" - It ends with the eos token {repr(tokenizer.eos_token)}, marking where to stop.")
[cell 36 markdown]
### Reading the formatted training string
This shows what one finished training example looks like:
- **Same format as our prompts.** It starts with the familiar `Context:` then `Question:` then `Answer:` layout, the same one we used for zero-shot testing.
- **The answer is now included.** Unlike the prompt we feed at test time, the training string has the correct answer, "Saint Bernadette Soubirous", written right after `Answer:`. This is the target the model learns to produce.
- **It ends with the end-of-text token.** The `<|endoftext|>` marker at the end tells the model where the answer stops. Training on this is what teaches the model to give a short answer and then stop, instead of rambling.
So each training example is the prompt plus the correct answer plus a stop signal, all as one piece of text. The model learns to fill in and end the `Answer:` section, which is exactly the behavior we want at test time.
[cell 37 code]
# cell_12b — token-length EDA on the formatted text, then set MAX_LENGTH
# MAIN GOAL: Earlier (cell_5) we measured lengths in WORDS to build intuition. But the model
# works in TOKENS, and tokens ≠ words (subword splitting makes tokens ~1.3–1.5x word count).
# Now that we have a tokenizer (cell_9) AND the final formatted training strings (cell_12),
# we measure the ACTUAL token length of the full "Context/Question/Answer: answer<eos>"
# sequences we'll train on, and use that to choose MAX_LENGTH honestly — instead of guessing.
# We target the 99th percentile so MAX_LENGTH covers almost every sequence with minimal
# truncation (we especially don't want to clip the answer at the end).
import matplotlib.pyplot as plt # plotting the histogram
import numpy as np # percentiles and array math
# Count tokens per formatted TRAIN sequence.
# We tokenize WITHOUT padding/truncation here so we measure each sequence's TRUE length.
# (This is measurement only; the real tokenize-for-training happens in cell_13.)
print("Measuring token lengths of formatted training sequences … (one pass over train)")
token_lengths = [
# tokenizer(...)["input_ids"] is the list of token ids for one sequence;
# len(...) of that list is how many tokens the sequence is.
len(tokenizer(row["text"])["input_ids"]) # raw token count, no padding/truncation
for row in train_formatted # do this for every formatted training row
]
token_lengths = np.array(token_lengths) # convert to a NumPy array for easy stats
# --- Percentile summary -------------------------------------------------------
# Percentiles tell us "X% of sequences are at most this many tokens."
# The 99th percentile is our candidate cap: it covers 99% of data with no truncation.
p50 = int(np.percentile(token_lengths, 50)) # median: half the sequences are shorter than this
p95 = int(np.percentile(token_lengths, 95)) # 95% of sequences are at most this long
p99 = int(np.percentile(token_lengths, 99)) # 99% of sequences are at most this long
longest = int(token_lengths.max()) # the single longest sequence (often an outlier)
print("Token-length summary (formatted train sequences):")
print(f" median (50th pct): {p50}")
print(f" 95th percentile : {p95}")
print(f" 99th percentile : {p99}")
print(f" longest sequence : {longest}")
# --- Choose MAX_LENGTH from the 99th percentile -------------------------------
# We round the 99th percentile UP to a "nice" multiple of 16. Rounding to a multiple of 8/16
# is a common habit because it can be slightly friendlier for GPU tensor operations.
# We also clamp to GPT-2's hard limit of 1024 tokens, just in case.
def round_up_to(value, multiple):
# np.ceil rounds up to the next whole number of "multiples", then we scale back.
# e.g. round_up_to(397, 16) -> 400.
return int(np.ceil(value / multiple) * multiple)
# min(..., 1024) makes sure we never exceed GPT-2's maximum context length.
MAX_LENGTH = min(round_up_to(p99, 16), 1024)
print("-" * 40)
print(f"Chosen MAX_LENGTH = {MAX_LENGTH} (99th pct {p99}, rounded up to a multiple of 16,")
print(f" capped at GPT-2's 1024-token limit)")
# What fraction of sequences will be truncated at this MAX_LENGTH? (Honesty check.)
# (token_lengths > MAX_LENGTH) gives a True/False array; .mean() is the fraction that are True.
truncated_frac = float((token_lengths > MAX_LENGTH).mean()) * 100
print(f"Sequences longer than MAX_LENGTH (will be truncated): {truncated_frac:.2f}%")
# --- Histogram of token lengths with the chosen cap marked --------------------
# Most sequences are short, but a few rare outliers (one is ~25,000 tokens) would
# stretch the x-axis and crush all the real data into a sliver. So we zoom the view
# to a sensible range instead of plotting the full spread.
# We cap the x-axis a bit past MAX_LENGTH so the cap line and the bulk of data are clear.
x_view_max = int(MAX_LENGTH * 1.5) # show a little beyond MAX_LENGTH for context
plt.figure(figsize=(8, 4))
# range=(0, x_view_max) keeps bins within the zoomed window; outliers beyond it are
# left out of the DRAWING only (they're still counted in the stats above).
# bins=50 splits that window into 50 bars.
plt.hist(token_lengths, bins=50, range=(0, x_view_max))
# axvline draws a vertical line at MAX_LENGTH so we can see where the cap falls in the data.
plt.axvline(MAX_LENGTH, color="red", linestyle="--", label=f"MAX_LENGTH = {MAX_LENGTH}")
plt.title("Token length of formatted training sequences (zoomed)")
plt.xlabel("tokens per sequence")
plt.ylabel("number of examples")
plt.legend() # show the MAX_LENGTH label
plt.tight_layout() # keep labels from getting clipped
plt.show()
print("\nTakeaway: tokens run longer than the word counts from cell_5 (subword splitting).")
print("We size MAX_LENGTH to the 99th percentile so we cover almost everything without")
print("clipping answers, while keeping sequences as short as possible for speed/memory.")
[cell 38 markdown]
### Reading the token-length analysis
This step measures how long our formatted sequences are in tokens, then sets the maximum length from that:
- **Most sequences are short.** The median is 170 tokens and 95 percent are at or below 305 tokens. The bulk of the data sits in a fairly narrow range.
- **The 99th percentile is 397.** Almost every sequence fits within about 400 tokens.
- **There is one extreme outlier.** The longest sequence is 25,733 tokens, far beyond everything else. This is a single unusual example and does not reflect the rest of the data.
- **MAX_LENGTH is set to 400.** We take the 99th percentile of 397 and round it up to a multiple of 16, then keep it under GPT-2's 1024-token limit. This covers nearly all sequences while keeping them as short as possible for speed and memory.
- **Very little is lost.** Only 0.94 percent of sequences are longer than 400 tokens and will be truncated. The rest fit completely.
The takeaway is that token counts run higher than the word counts we saw earlier, and sizing MAX_LENGTH to the 99th percentile lets us keep almost every answer intact while avoiding the waste of padding everything out to cover rare giant outliers.
[cell 39 code]
# cell_13 — tokenize and build labels for ANSWER-ONLY causal-LM loss
# MAIN GOAL: Turn the formatted "text" (cell_12) into model inputs (input_ids,
# attention_mask) AND build "labels" so that loss is computed ONLY over the answer tokens
# (plus the eos that teaches the model to stop). Everything else — the context/question
# prompt, and the padding — is set to -100, the special value that cross-entropy IGNORES.
# This is the "answer-only loss" approach: the model is graded on producing the answer,
# not on re-predicting the passage, which trains a sharper QA model and makes the
# validation loss actually reflect answer quality.
#
# How we find where the answer starts: we rebuild the prompt prefix ("Context:…Question:…
# Answer:") for each example and count ITS tokens. The first that-many positions are the
# prompt → masked. The remaining real tokens are " answer…<eos>" → kept. (cell_12 puts a
# space before the answer, so the answer's first token is "Ġanswer", giving a clean split.)
# MAX_LENGTH comes from cell_12b (chosen from the 99th-percentile token length).
def tokenize_and_mask(batch):
# Tokenize the FULL formatted sequences, padded/truncated to a uniform MAX_LENGTH.
# truncation=True -> cut anything longer than MAX_LENGTH
# padding="max_length" -> pad shorter sequences up to MAX_LENGTH (uniform shape)
# padding uses the eos token id (pad_token was set to eos in cell_9).
model_inputs = tokenizer(
batch["text"],
truncation=True,
padding="max_length",
max_length=MAX_LENGTH,
)
labels_batch = [] # we'll build one label list per example
# Process each example in the batch individually, because the prompt length (and thus
# where the answer begins) differs from row to row.
for i in range(len(batch["text"])):
input_ids = model_inputs["input_ids"][i] # token ids for this sequence
attention_mask = model_inputs["attention_mask"][i] # 1 = real token, 0 = padding
# Rebuild THIS example's prompt prefix from the same template (cell_1/cell_12).
# context and question are still present as columns in train_formatted.
prompt = PROMPT_TEMPLATE.format(
context=batch["context"][i],
question=batch["question"][i],
)
# Count the prompt's tokens (no padding/truncation here — we just need the length).
# Because cell_12 used this exact prompt as the prefix, the full sequence's first
# `prompt_len` tokens ARE the prompt; everything after is the answer (+eos).
prompt_len = len(tokenizer(prompt)["input_ids"])
# Start labels as a COPY of input_ids (so we don't modify the inputs themselves),
# then overwrite the parts that shouldn't contribute to the loss with -100.
labels = input_ids.copy()
# (a) Mask the PROMPT prefix: positions 0 .. prompt_len-1 -> -100 (ignored).
# This is what makes the loss "answer-only" — the model isn't graded on
# re-predicting the context/question.
# min(...) guards against a prompt longer than the (truncated) sequence.
for j in range(min(prompt_len, len(labels))):
labels[j] = -100
# (b) Mask PADDING: any position the attention_mask marks as 0 -> -100.
# Note: padding tokens ARE eos tokens (pad==eos), but attention_mask=0 tells us
# they're padding, NOT the real end-of-answer eos. The real eos has mask=1 and
# sits after the answer, so it stays UNMASKED — that's how the model learns to stop.
for j in range(len(labels)):
if attention_mask[j] == 0:
labels[j] = -100
labels_batch.append(labels) # save this example's labels
# Attach the labels we built as a new field alongside input_ids / attention_mask.
model_inputs["labels"] = labels_batch
return model_inputs
# Apply to train and validation.
# batched=True -> tokenize_and_mask receives a batch of rows at once (faster)
# remove_columns=... -> drop all original string columns so the Trainer only sees
# input_ids / attention_mask / labels (leftover strings would
# break tensor creation).
train_tokenized = train_formatted.map(
tokenize_and_mask,
batched=True,
remove_columns=train_formatted.column_names,
)
valid_tokenized = valid_formatted.map(
tokenize_and_mask,
batched=True,
remove_columns=valid_formatted.column_names,
)
# --- Sanity check: confirm masking kept ONLY the answer (+eos) ----------------
# We take one example and show which tokens are "kept" (label != -100) vs masked.
ex = train_tokenized[0]
# Pair each token id with its label; keep only tokens whose label is NOT -100 (i.e. the
# positions that actually contribute to the loss — should be just the answer + eos).
kept_ids = [tok for tok, lab in zip(ex["input_ids"], ex["labels"]) if lab != -100]
print("Masking sanity check on one training example:")
print(" total positions :", len(ex["labels"])) # = MAX_LENGTH
num_kept = sum(1 for lab in ex["labels"] if lab != -100) # count of unmasked positions
print(" positions kept (loss):", num_kept, "(should be just the answer + eos)")
# Decode just the kept tokens back to text to visually confirm it's the answer.
print(" decoded KEPT tokens :", repr(tokenizer.decode(kept_ids)))
print(" ^ should read as the answer text, ending in the eos token")
# --- Diagnostic: how many examples ended up with NO answer tokens kept? --------
# This can happen for very long contexts where truncation cut the answer off. Should be
# ~0 because MAX_LENGTH was sized to the 99th percentile, but we check honestly.
no_answer = sum(
1 for row in train_tokenized
if all(lab == -100 for lab in row["labels"]) # True when EVERY label is masked
)
print("-" * 40)
print(f"Examples with zero answer tokens (answer truncated away): {no_answer}")
if no_answer > 0:
print(" (These contribute nothing to loss. If this number is large, raise MAX_LENGTH.)")
[cell 40 code]
# cell_14 — set up the Trainer (training config + per-epoch validation)
# MAIN GOAL: Configure fine-tuning.
# - Train for NUM_EPOCHS (cell_1).
# - Check validation loss after EACH epoch.
# - Keep the BEST epoch's model automatically (lowest val loss).
from transformers import Trainer, TrainingArguments # the training engine + its config object
# Where checkpoints + logs get saved (a folder created on disk).
OUTPUT_DIR = "gpt2-squad-finetuned"
# Batch size = how many sequences the model processes per training step.
# Bigger = faster but more GPU memory. Lower this if you hit out-of-memory on the GPU.
BATCH_SIZE = 8
# TrainingArguments holds every setting that controls HOW training runs.
training_args = TrainingArguments(
output_dir=OUTPUT_DIR, # folder for checkpoints and logs
num_train_epochs=NUM_EPOCHS, # full passes over train data
per_device_train_batch_size=BATCH_SIZE, # train batch size (per GPU/CPU)
per_device_eval_batch_size=BATCH_SIZE, # eval batch size (per GPU/CPU)
eval_strategy="epoch", # run validation once per epoch
save_strategy="epoch", # save a checkpoint once per epoch
load_best_model_at_end=True, # restore best checkpoint when done
metric_for_best_model="eval_loss", # "best" = lowest validation loss
greater_is_better=False, # for loss, LOWER is better (not higher)
logging_strategy="steps", # log training loss periodically
logging_steps=3, # ...every 3 steps (frequent, so we see progress)
#bump this to 50 on full runs so logs are less noisy; keep it low for development/debugging
learning_rate=5e-5, # how big each weight update is; standard GPT-2 fine-tuning LR
weight_decay=0.01, # mild regularization, discourages over-large weights
warmup_ratio=0.05, # ramp LR up gradually over the first 5% of steps
report_to="none", # don't send logs to external tools (keep it simple)
seed=SEED, # reproducible training (same run every time)
save_total_limit=1, # only keep the single best checkpoint (saves disk)
)
# Trainer ties together: model, args, the two datasets, and the tokenizer.
# Labels are already in the dataset (cell_13), so the model computes the loss automatically.
trainer = Trainer(
model=model, # the GPT-2 loaded in cell_9
args=training_args, # the config we just built
train_dataset=train_tokenized, # data the model learns from (answer-only loss)
eval_dataset=valid_tokenized, # data used only to measure per-epoch validation loss
tokenizer=tokenizer, # so the Trainer knows how to pad/collate batches
)
print("Trainer ready.")
print(f" epochs : {NUM_EPOCHS}")
print(f" batch size : {BATCH_SIZE}")
print(f" train rows : {len(train_tokenized)}")
print(f" val rows : {len(valid_tokenized)}")
print(f" best model : kept automatically (lowest eval_loss)")
[cell 41 code]
# cell_15 — run fine-tuning
# MAIN GOAL: Actually train the model.
# - Runs NUM_EPOCHS passes over the training data.
# - Prints training loss periodically and validation loss each epoch.
# - When done, `model` holds the BEST epoch (lowest val loss), per cell_14.
# This is the SLOW cell. On full SQuAD it can take a while — narrate while it runs.
import time # used to measure how long training takes
print("Starting fine-tuning…")
print("Watch the table: 'Training Loss' should fall; 'Validation Loss' is how we pick the best epoch.\n")
start = time.time() # record the start time (seconds since epoch)
# Kick off training. This runs the whole loop (all epochs) and blocks until done.
# It returns summary stats about the run.
train_result = trainer.train()
elapsed = time.time() - start # total wall-clock seconds the run took
print("\nFine-tuning complete.")
print(f" total time : {elapsed/60:.1f} min") # convert seconds to minutes
print(f" final step : {train_result.global_step}") # total optimizer steps taken
print(f" train loss : {train_result.training_loss:.4f} (avg over run)") # mean training loss
# Show the validation loss recorded at each epoch, pulled from the logged history.
# trainer.state.log_history is a list of dicts; the Trainer appended one entry each time
# it logged or evaluated. We pick out the evaluation entries.
# This makes it explicit WHICH epoch was best (the one load_best_model_at_end restored).
print("\nValidation loss by epoch:")
for entry in trainer.state.log_history:
if "eval_loss" in entry: # only evaluation entries carry eval_loss
epoch = entry.get("epoch", "?") # which epoch this eval was for
print(f" epoch {epoch}: eval_loss = {entry['eval_loss']:.4f}")
# The best score Trainer tracked (lowest eval_loss), and which checkpoint it came from.
# best_metric is None if no evaluation ran; we guard against that.
if trainer.state.best_metric is not None:
print(f"\nBest validation loss: {trainer.state.best_metric:.4f}")
# Because load_best_model_at_end=True (cell_14), `model` now holds this best version,
# not necessarily the last epoch's version.
print("This best checkpoint is now loaded into `model`.")
[cell 42 markdown]
### Reading the training results
Fine-tuning finished, and the per-epoch numbers tell a clear story:
- **Training ran for three full epochs.** It took 32,850 steps on the full dataset. The average training loss over the run was 0.4480.
- **Validation loss improved, then leveled off.** It fell from 0.5368 after epoch 1 to 0.5045 after epoch 2, then rose slightly to 0.5155 after epoch 3.
- **Epoch 2 was the best.** Its validation loss of 0.5045 was the lowest, so that is the checkpoint we keep.
- **The best model is loaded automatically.** Thanks to the best-checkpoint setting, `model` now holds the epoch 2 version, not the final epoch 3 version.
This is exactly why we trained for more than one epoch and watched the validation loss. The small rise at epoch 3 is an early sign of overfitting, where the model starts fitting the training data a little too closely and stops improving on unseen data. By picking the epoch with the lowest validation loss rather than just the last one, we keep the version that should generalize best. The test numbers we compute next will measure how well that choice paid off.
[cell 43 code]
# cell_16 — after fine-tuning: same prompts, same scoring, before vs after
# MAIN GOAL: Re-run the EXACT same demo and test scoring as the zero-shot section, now
# with the fine-tuned model, and compare.
# - Same demo_examples (cell_11), same generate_answer (cell_10), same evaluate (cell_11b).
# - Only the model's weights changed → a true apples-to-apples before/after.
# Training left the model in train mode; switch back to eval for stable output (no dropout).
# generate_answer uses the global `model`, which now holds the fine-tuned best epoch.
model.eval()
# --- Qualitative: same demo questions, now answered by the fine-tuned model ---
print("AFTER FINE-TUNING — same questions as the zero-shot demo")
print("=" * 70)
finetuned_answers = [] # collect answers, aligned with demo_examples by index
# Loop over the SAME frozen demo examples we used zero-shot (cell_11), in the same order.
for i, ex in enumerate(demo_examples):
gold = ex["answers"]["text"][0] # first gold answer (for display only)
pred = generate_answer(ex["context"], ex["question"]) # fine-tuned model's answer
finetuned_answers.append(pred) # save it
print(f"[Example {i}]")
print(f" Question : {ex['question']}")
print(f" Gold answer : {gold}")
print(f" BEFORE (zero) : {repr(zero_shot_answers[i])}") # saved in cell_11, same index
print(f" AFTER (tuned) : {repr(pred)}") # what we just generated
print("-" * 70)
# --- Quantitative: same test scoring as cell_11b -----------------------------
# Same evaluate(), same test_data, same NUM_EVAL → directly comparable to the zero-shot run.
# The ONLY thing that changed since the zero-shot run is the model's weights.
print(f"\nScoring fine-tuned model on {NUM_EVAL} TEST examples…")
finetuned_em, finetuned_f1 = evaluate(test_data, NUM_EVAL, generate_answer)
# --- Before/after summary -----------------------------------------------------
# Print a small aligned table. The format specifiers control column width and decimals:
# {:<14} = left-aligned in 14 chars, {:>10} = right-aligned in 10 chars,
# {:>9.1f} = right-aligned, 1 decimal place (the % sign is printed separately).
print("\n" + "=" * 40)
print("RESULTS (held-out test set)")
print("=" * 40)
print(f"{'Metric':<14}{'Before':>10}{'After':>10}")
print("-" * 40)
print(f"{'Exact Match':<14}{zero_shot_em:>9.1f}%{finetuned_em:>9.1f}%")
print(f"{'F1':<14}{zero_shot_f1:>9.1f}%{finetuned_f1:>9.1f}%")
print("=" * 40)
# The gain is simply after minus before, for each metric.
print(f"EM gain: +{finetuned_em - zero_shot_em:.1f} F1 gain: +{finetuned_f1 - zero_shot_f1:.1f}")
[cell 44 markdown]
### Reading the final results
These are the headline numbers, measured on the full held-out test set:
- **Exact Match rose from 0.0 percent to 69.7 percent.** The untrained model never produced an exact answer. After fine-tuning, it matches the gold answer outright on about seven of every ten questions.
- **F1 rose from 8.5 percent to 78.8 percent.** Counting partial overlap, the fine-tuned model recovers most of the correct answer text on the large majority of questions.
- **Both gains are large.** EM improved by 69.7 points and F1 by 70.3 points, on data the model never saw during training.
The takeaway is that fine-tuning clearly worked. The key points:
- A small 124 million parameter model went from rambling and scoring zero on exact match to answering most questions correctly and concisely.
- The numbers are measured on the held-out test set, using the exact same scoring as the baseline, so the comparison is fair.
- Part of the F1 gain comes from the model learning to be concise, not just more accurate. Shorter answers stop dragging the score down with extra words.
[cell 45 markdown]
## Summary and Wrap-Up
### What we did
We taught a small GPT-2 model to answer questions, and we can measure that it worked.
- We started with the SQuAD dataset and explored it: long contexts, short questions, and very short answers, mostly "what" questions.
- We treated question answering as text generation, using one consistent prompt format throughout.
- We tested the untrained model and saw it ramble, repeat, and invent answers, scoring 0 percent exact match.
- We fine-tuned the model on the data, training for three epochs and keeping the best one by validation loss.
- We re-ran the exact same questions and scoring, and saw a clear jump: 0 to 69.7 percent exact match, and 8.5 to 78.8 percent F1.
### The main lesson
- A pretrained model knows language but not your specific task.
- Fine-tuning on task-shaped examples is what teaches it the format and the behavior you want.
- The before and after comparison, on held-out test data, is how we prove the improvement is real.
### Limitations to keep in mind
- **The metrics are strict.** One model answer named the specific colleges instead of the phrase "several regional colleges and universities". It was arguably correct, but scored as wrong because it did not match the gold text word for word.
- **Models can still hallucinate.** The untrained model confidently invented a fake founder for the University of Chicago. Even fine-tuned models can produce fluent, wrong answers, so outputs should not be blindly trusted.
- **This is a small model.** GPT-2 small has 124 million parameters. It does well here, but it has a ceiling, and harder questions will expose it.
- **We made simplifying choices.** We capped sequence length, which truncated about 1 percent of examples, and we trained on a single answer per question. These keep the tutorial simple but are not the most thorough setup.
### The takeaway
We took a general language model and, with a clear task format and a modest amount of fine-tuning, turned it into a working question answering system, then measured the improvement honestly on data it had never seen. That full loop, explore, baseline, train, and compare, is the core workflow you can reuse for many other tasks.