# sqaud gpt qa v4 GC PDF
course: Module 4 — Generative AI & LLMs
module: Module-4-Generative-AI-LLMs
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-4-Generative-AI-LLMs/General/Materials-06-06-2026/sqaud_gpt_qa_v4_GC_PDF.pdf
pages: 55
---
[page 1]
The SQuAD Dataset
What it is
SQuAD (the Stanford Question Answering Dataset) is a reading comprehension dataset.
Each item gives the model a passage of text and a question about that passage. The
answer to the question is always a span of text taken directly from the passage.
The version we use here is SQuAD v1.1, loaded with load_dataset("squad") from the
Hugging Face datasets library. It was built from a set of Wikipedia articles, with
questions and answers written by human crowdworkers.
Why we chose it
A few reasons make SQuAD a good fit for a beginner tutorial:
1. It is clean and well documented. The structure is simple and consistent, so we spend
time learning the task instead of cleaning messy data.
2. It is realistic. The passages are real Wikipedia text and the questions are written by
people, so the task resembles something you might actually build.
3. It is widely used. SQuAD is a standard benchmark, so the metrics and methods we
learn here carry over to other work.
What one example looks like
Every row in the dataset has the same fields:
id: a unique identifier for the example.
title: the title of the Wikipedia article the passage came from.
context: the passage of text. This is what the model reads to find the answer.
question: the question asked about the context.
answers: the correct answer or answers.
The answers field is worth a closer look because its shape can be confusing at first. It is
not a single string. It is a small dictionary with two parallel lists:
text: the answer string or strings.
answer_start: the character position in the context where each answer begins.
For example, a single row might look like this:
context: "The Amazon rainforest covers most of the Amazon basin
in South America. The basin spans nine countries."
question: "How many countries does the Amazon basin span?"
answers: {"text": ["nine"], "answer_start": [88]}
The answer "nine" appears in the context, and answer_start tells you the exact character
index where it starts.
[page 2]
One answer or several
The number of answers depends on the split:
In the training split, each question has exactly one answer.
In the validation split, each question often has several answers, written by different
annotators. They are all considered correct.
This matters when we score the model. A fair score checks the prediction against every
acceptable answer and keeps the best match, so we do not mark a correct answer wrong
just because it matched a different annotator.
How we use SQuAD in this tutorial
SQuAD was originally designed for extractive question answering, where the model
points to the start and end of the answer inside the passage. We are doing something
different. We treat the task as text generation: the model reads the passage and the
question, then writes the answer as text.
To do this we put every example into one consistent prompt format:
Context: <the passage>
Question: <the question>
Answer:
The model's job is to continue this text with the answer. We use the same format
everywhere, both when we test the untrained model and after fine-tuning, so the before
and after comparison is fair.
Splits we work with
SQuAD ships with two splits, a large training split and a smaller validation split. It does
not include a public test split, because the original test set was kept private for an online
leaderboard.
To match the usual train, validation, and test setup, we create our own three splits:
train: used to fine-tune the model.
validation: a portion of SQuAD's validation split, used to check the model after
each training epoch and pick the best one.
test: the remaining portion, held out and used only for the final before and after
numbers.
What to keep in mind
A couple of points are useful to remember as we go:
Answers in SQuAD are short, usually just a few words. The passages, by contrast,
are full paragraphs. This is why we ask the model to read a long input but produce a
short output.
Because the answer is always present in the passage, the task rewards careful reading
rather than outside knowledge. The model should find the answer in the text, not
invent it.
[page 3]
Enter physical GPU index (e.g. 0, 1, 2) or 'cpu': cpu
Device set to: cpu
# cell_0a — GPU mask setup before any torch import
# MAIN GOAL: Let the user pick which physical GPU to use (or CPU),
# then hide all other GPUs from this process via CUDA_VISIBLE_DEVICES.
# Must run before importing torch or any CUDA-using library.
import os
choice = input("Enter physical GPU index (e.g. 0, 1, 2) or 'cpu':
").strip().lower()
if choice == "cpu":
os.environ["CUDA_VISIBLE_DEVICES"] = ""
DEVICE = "cpu"
elif choice.isdigit():
os.environ["CUDA_VISIBLE_DEVICES"] = choice # hide every other GPU
DEVICE = "cuda:0" # selected GPU maps to
cuda:0 internally
else:
print("Invalid input. Defaulting to CPU.")
os.environ["CUDA_VISIBLE_DEVICES"] = ""
DEVICE = "cpu"
print(f"Device set to: {DEVICE}")
# cell_0b — install pinned libraries
# MAIN GOAL: Install the exact library versions this tutorial was built and
tested against.
# Pinning versions is a deliberate safeguard: a surprise upgrade to
transformers/datasets
# can change APIs mid-session and break the notebook for a beginner audience.
# This cell only INSTALLS packages — it does not import them. Importing torch
happens
# later, AFTER the GPU mask in cell_0a has already been set.
# NOTE: on Colab, if pip reports that a package was updated, you may need to
use
# Runtime > Restart session once, then re-run from cell_0a. Locally this is
usually not needed.
# The leading "!" runs this line as a shell command from inside the notebook.
# -q = quiet, fewer log lines so beginners aren't overwhelmed.
# package==X = pin to an exact version for reproducibility.
!pip install -q \
transformers==4.44.2 \
datasets==2.21.0 \
accelerate==0.34.2 \
matplotlib==3.9.2 \
pandas==2.2.2
[page 4]
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 43.7/43.7 kB 570.0 kB/s eta 0:00:00
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 9.5/9.5 MB 22.2 MB/s eta 0:00:00
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 527.3/527.3 kB 29.5 MB/s eta 0:00:00
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 324.4/324.4 kB 15.0 MB/s eta 0:00:00
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 8.3/8.3 MB 50.0 MB/s eta 0:00:00
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 177.6/177.6 kB 13.1 MB/s eta 0:00:00
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 566.4/566.4 kB 26.0 MB/s eta 0:00:00
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 3.6/3.6 MB 65.4 MB/s eta 0:00:00
ERROR: pip's dependency resolver does not currently take into account all the
packages that are installed. This behaviour is the source of the following
dependency conflicts.
gcsfs 2025.3.0 requires fsspec==2025.3.0, but you have fsspec 2024.6.1 which
is incompatible.
Install step complete. If on Colab and prompted, restart the session, then
re-run from cell_0a.
# Quick confirmation line so the user gets visible feedback that the cell
finished.
print("Install step complete. If on Colab and prompted, restart the session,
then re-run from cell_0a.")
# cell_0c — imports and device verification
# MAIN GOAL: Import torch (for the first time, now that cell_0a has already
hidden the
# unwanted GPUs) plus the core libraries we'll use throughout, then CONFIRM
that the
# device we asked for in cell_0a is actually the device PyTorch sees and will
use.
# This catches "I asked for GPU but it silently fell back to CPU" problems
early,
# before we waste time loading models or data.
import torch # deep learning framework; backs
the GPT-2 model and training
import transformers # Hugging Face: model + tokenizer +
Trainer
import datasets # Hugging Face: loads SQuAD
import pandas as pd # tabular handling for the EDA
section
import matplotlib # plotting for the EDA section
# Print versions so that if anything misbehaves later, we can see exactly
what's installed.
# This should match the pins from cell_0b.
print("torch :", torch.__version__)
print("transformers:", transformers.__version__)
print("datasets :", datasets.__version__)
print("pandas :", pd.__version__)
print("matplotlib :", matplotlib.__version__)
print("-" * 40)
# Did PyTorch actually find a CUDA GPU in this process?
# Remember: because of cell_0a, AT MOST ONE GPU is visible here (or none, if
'cpu' was chosen).
[page 5]
torch : 2.11.0+cpu
transformers: 4.44.2
datasets : 2.21.0
pandas : 2.2.2
matplotlib : 3.10.0
----------------------------------------
CUDA available to this process: False
Final device for this notebook: cpu
cuda_available = torch.cuda.is_available()
print("CUDA available to this process:", cuda_available)
# Safety reconciliation between what we ASKED for (DEVICE from cell_0a) and
what we GOT.
# If we asked for cuda but CUDA isn't available, fall back to CPU loudly
rather than crashing later.
if DEVICE.startswith("cuda") and not cuda_available:
print("WARNING: cuda was requested in cell_0a, but no CUDA GPU is
visible. Falling back to CPU.")
DEVICE = "cpu"
# If we do have a GPU, print its name so the user sees which card they
actually landed on.
if cuda_available and DEVICE.startswith("cuda"):
# Index 0 here is the *visible* GPU, i.e. the one we selected in cell_0a
(it was remapped to 0).
print("Active GPU:", torch.cuda.get_device_name(0))
print("Final device for this notebook:", DEVICE)
# cell_1 — central configuration (subsample toggle lives here)
# MAIN GOAL: Put every knob the user might want to turn in ONE place, near
the top,
# so the rest of the notebook just reads these variables and never needs
editing.
# The headline controls are the SUBSAMPLE toggle (how much data) and
NUM_EPOCHS (how long
# we train). Everything downstream (splitting, EDA, fine-tuning, evaluation)
reads from here.
# --- Subsample toggle ------------------------------------------------------
---
# USE_FULL_DATASET:
# True -> use all of SQuAD (slow to fine-tune; this is our "real run" for
proper numbers)
# False -> use only SUBSAMPLE_SIZE rows from the training split (fast, for
time-tuning later)
# We start with the FULL dataset to get trustworthy results first; we can
flip this to
# False afterwards to tune for speed without changing any other code.
USE_FULL_DATASET = False # flip to False later to do a fast subsampled run
# SUBSAMPLE_SIZE:
# How many TRAIN rows to keep when USE_FULL_DATASET is False.
[page 6]
# A few thousand keeps fine-tuning quick. Ignored entirely when
USE_FULL_DATASET is True.
SUBSAMPLE_SIZE = 20
# VALID_SUBSAMPLE_SIZE:
# How many rows to keep (BEFORE the val/test split below) when subsampling.
# Ignored when USE_FULL_DATASET is True.
VALID_SUBSAMPLE_SIZE = 10
# --- Train epochs ----------------------------------------------------------
---
# NUM_EPOCHS:
# How many full passes over the training data during fine-tuning.
# We use more than 1 so that "pick the best epoch by validation loss" is
meaningful:
# with only 1 epoch there would be nothing to choose between. The Trainer
will evaluate
# validation loss after each epoch and we'll keep the best checkpoint.
#NUM_EPOCHS = 3 #-> for full run on full datasets
NUM_EPOCHS = 2
# --- Validation / Test split -----------------------------------------------
---
# SQuAD ships with only "train" and "validation" (no public "test"). To match
the usual
# train/val/test setup, we split SQuAD's validation split into our OWN val
and test sets.
# TEST_FRACTION = portion of SQuAD-validation that becomes our TEST set (the
rest is val).
# We give TEST the larger share because test carries the headline
before/after numbers
# we report at the end, while our validation set only does light per-epoch
model picking.
TEST_FRACTION = 0.70 # 70% test, 30% validation
# --- Reproducibility -------------------------------------------------------
---
# A fixed seed makes the "random" subsample AND the val/test split the same
every run,
# so results and demo examples are stable. Predictability matters when
teaching.
SEED = 42
# --- Model choice ----------------------------------------------------------
---
# "gpt2" is GPT-2 small (124M), our primary model.
# Swap to "distilgpt2" if a later fast run needs to be lighter (the brief's
fallback).
MODEL_NAME = "gpt2"
# --- Prompt template -------------------------------------------------------
---
# ONE template, reused identically in BOTH the zero-shot and post-fine-tuning
sections
[page 7]
Configuration for this run
----------------------------------------
USE_FULL_DATASET : False
SUBSAMPLE_SIZE : 20 (ignored if USE_FULL_DATASET=True)
VALID_SUBSAMPLE_SIZE: 10 (ignored if USE_FULL_DATASET=True)
NUM_EPOCHS : 2
TEST_FRACTION : 0.7 (share of SQuAD-validation used as TEST)
SEED : 42
MODEL_NAME : gpt2
PROMPT_TEMPLATE :
Context: {context}
Question: {question}
Answer:
Reading the configuration
This run is set up as follows:
Full dataset. USE_FULL_DATASET is True, so we train on all of SQuAD. The two
subsample sizes are shown but ignored here. They only take effect when
USE_FULL_DATASET is False, which is the faster mode we can switch to later.
Three training epochs. NUM_EPOCHS is 3. We use more than one epoch so that
picking the best epoch by validation loss is meaningful. With a single epoch there
would be nothing to compare.
Test split of 70 percent. TEST_FRACTION is 0.7, so 70 percent of SQuAD's validation
split becomes our test set and the remaining 30 percent becomes our validation set.
Test gets the larger share because it carries the final before and after numbers, while
validation only does light per-epoch model selection.
Fixed seed. SEED is 42. This keeps the random parts of the run repeatable, the data
split, any shuffling, and training, so the results and the demo examples are the same
every time.
# so the before/after comparison is apples-to-apples. Defining it here
guarantees
# we never accidentally use two slightly different prompts.
# {context} and {question} get filled in per example later.
PROMPT_TEMPLATE = "Context: {context}\nQuestion: {question}\nAnswer:"
# --- Quick echo so the user can SEE the active configuration ---------------
---
print("Configuration for this run")
print("-" * 40)
print("USE_FULL_DATASET :", USE_FULL_DATASET)
print("SUBSAMPLE_SIZE :", SUBSAMPLE_SIZE, "(ignored if
USE_FULL_DATASET=True)")
print("VALID_SUBSAMPLE_SIZE:", VALID_SUBSAMPLE_SIZE, "(ignored if
USE_FULL_DATASET=True)")
print("NUM_EPOCHS :", NUM_EPOCHS)
print("TEST_FRACTION :", TEST_FRACTION, "(share of SQuAD-validation
used as TEST)")
print("SEED :", SEED)
print("MODEL_NAME :", MODEL_NAME)
print("PROMPT_TEMPLATE :")
print(PROMPT_TEMPLATE)
[page 8]
Model. MODEL_NAME is gpt2, the 124M parameter GPT-2 small model.
Prompt template. Every example is formatted the same way, a context, then a
question, then Answer: for the model to continue. We reuse this exact format before
and after fine-tuning so the comparison stays fair.
Keeping all of these settings in one place means the rest of the notebook just reads from
them. To change the run, you change this cell and nothing else.
# cell_2 — load SQuAD, apply the subsample toggle, and build train/val/test
splits
# MAIN GOAL: Load SQuAD and produce THREE splits the rest of the notebook
will use:
# train_data -> fine-tuning (and EDA)
# validation_data -> per-epoch model selection during fine-tuning (pick
best epoch)
# test_data -> final, held-out before/after numbers we report at the
end
# SQuAD only ships "train" and "validation" (its real test set is private),
so we create
# our own val/test by splitting SQuAD's validation split using TEST_FRACTION
from cell_1.
# We also honor the subsample toggle. This cell only LOADS/SIZES/SPLITS — no
tokenizing.
from datasets import load_dataset
# Download (first run) or load-from-cache (later runs) the SQuAD dataset.
# Returns a DatasetDict with two keys: "train" (~87k rows) and "validation"
(~10k rows).
# Subsequent runs are instant because HuggingFace caches to
~/.cache/huggingface/datasets.
print("Loading SQuAD … (first run downloads it; later runs use the local
cache)")
squad = load_dataset("squad")
print("\nRaw dataset structure (note: only 'train' and 'validation' exist):")
print(squad)
# Store references to the two original SQuAD splits before any resizing.
# We keep "full_" prefixes so it's always clear we haven't touched these yet.
full_train = squad["train"]
full_squad_validation = squad["validation"] # we will CUT this into our
val + test
# --- Apply the subsample toggle from cell_1 --------------------------------
--
if USE_FULL_DATASET:
# Keep everything (our "real run" for trustworthy numbers).
# No copying happens here — these are just variable aliases pointing at
the same data.
train_data = full_train
squad_validation = full_squad_validation
print("\nUSE_FULL_DATASET=True -> keeping the entire dataset.")
else:
[page 9]
# Subsample for speed (used during development / quick iteration).
# min() guards against requesting more rows than the split actually
contains.
n_train = min(SUBSAMPLE_SIZE, len(full_train))
n_valid = min(VALID_SUBSAMPLE_SIZE, len(full_squad_validation))
# Shuffle BEFORE selecting so we don't grab a non-representative block
(SQuAD is
# ordered by article, so the first N rows would cover only a handful of
topics).
# The fixed SEED makes the subsample identical across runs for
reproducibility.
train_data = full_train.shuffle(seed=SEED).select(range(n_train))
squad_validation =
full_squad_validation.shuffle(seed=SEED).select(range(n_valid))
print(f"\nUSE_FULL_DATASET=False -> subsampled to {n_train} train / "
f"{n_valid} (pre-split) validation rows.")
# --- Split SQuAD's validation into OUR validation_data + test_data ---------
--
# train_test_split is a built-in Dataset method. Despite the name, here it
just divides
# squad_validation into two parts. test_size=TEST_FRACTION sends that share
to "test".
# seed=SEED makes the partition identical every run (stable test numbers +
demo examples).
# (Internally it shuffles before splitting, so both halves are
representative.)
split = squad_validation.train_test_split(test_size=TEST_FRACTION, seed=SEED)
# Naming note: HuggingFace always returns keys "train" and "test" from this
method,
# regardless of what the two parts will actually be used for in your project.
# Here the "train" portion becomes our validation set (the smaller slice we
check
# each epoch) and the "test" portion becomes our held-out test set.
validation_data = split["train"] # the SMALLER part (1 - TEST_FRACTION) ->
our validation
test_data = split["test"] # the LARGER part (TEST_FRACTION) ->
our test
# --- Final readout of what we'll actually work with ------------------------
--
# Printing row counts here is a quick sanity check: if the numbers look
wildly off
# (e.g. test_data is empty) it usually means TEST_FRACTION was set
incorrectly in cell_1.
print("-" * 40)
print("train_data rows :", len(train_data), "(fine-tuning + EDA)")
print("validation_data rows :", len(validation_data), "(per-epoch model
selection)")
print("test_data rows :", len(test_data), "(final before/after
numbers)")
print("columns :", train_data.column_names)
[page 10]
Loading SQuAD … (first run downloads it; later runs use the local cache)
/usr/local/lib/python3.12/dist-packages/huggingface_hub/utils/_auth.py:94:
UserWarning:
The secret `HF_TOKEN` does not exist in your Colab secrets.
To authenticate with the Hugging Face Hub, create a token in your settings
tab (https://huggingface.co/settings/tokens), set it as secret in your Google
Colab and restart your session.
You will be able to reuse this secret in all of your notebooks.
Please note that authentication is recommended but still optional to access
public models or datasets.
warnings.warn(
Raw dataset structure (note: only 'train' and 'validation' exist):
DatasetDict({
train: Dataset({
features: ['id', 'title', 'context', 'question', 'answers'],
num_rows: 87599
})
validation: Dataset({
features: ['id', 'title', 'context', 'question', 'answers'],
num_rows: 10570
})
})
USE_FULL_DATASET=False -> subsampled to 20 train / 10 (pre-split) validation
rows.
----------------------------------------
train_data rows : 20 (fine-tuning + EDA)
validation_data rows : 3 (per-epoch model selection)
test_data rows : 7 (final before/after numbers)
columns : ['id', 'title', 'context', 'question', 'answers']
Reading the dataset output
This is what the loading and splitting step produced:
SQuAD has two original splits. The raw structure shows a train split with 87,599
rows and a validation split with 10,570 rows. There is no test split, which is why
we make our own.
Each row has the same five fields. Every example contains id, title, context,
question, and answers. These are the same across all splits.
{"model_id":"ce6e904aa48947e392850a23a8d0730e","version_major":2,"version_minor":0}
{"model_id":"03277529f8134f51b7fdb2538904bb59","version_major":2,"version_minor":0}
{"model_id":"2ff4234e94a1482fb3c0935354c79ca2","version_major":2,"version_minor":0}
{"model_id":"7003a045e7714132aec11dea147f938b","version_major":2,"version_minor":0}
{"model_id":"a13eb2da8ddd4f5b8a96b515fd68b4e2","version_major":2,"version_minor":0}
[page 11]
The full dataset is kept. Because USE_FULL_DATASET is True, the train split stays at
its full 87,599 rows for fine-tuning and EDA.
Validation was split into our own validation and test sets. SQuAD's 10,570
validation rows were divided using TEST_FRACTION of 0.7. This gives 7,399 rows for
test (the final before and after numbers) and 3,171 rows for validation (per-epoch
model selection). The two add up to the original 10,570.
From here on, the notebook works with three clear splits: train_data for training,
validation_data for picking the best epoch, and test_data for the final held-out results.
# cell_3 — meet the data: inspect structure and print full examples
# MAIN GOAL: Before any charts, make the dataset CONCRETE by looking at a few
complete
# rows. A beginner should walk away knowing exactly what one SQuAD example
contains:
# - context : a paragraph of text (the passage to read)
# - question: a question whose answer is found in that context
# - answers : the answer text(s), plus where in the context they start
# We print 2-3 full examples in a readable layout. No statistics yet — that's
the next cell.
# First, look at ONE raw example exactly as the dataset stores it (a Python
dict).
# Seeing the raw dict demystifies the column names we printed in cell_2.
print("One raw example (as stored):")
print(train_data[0])
print("=" * 70)
# The raw dict is a bit noisy, so now print a FEW examples in a clean,
labeled format.
# We loop over the first 3 rows and pretty-print each field.
NUM_EXAMPLES_TO_SHOW = 3 # small number; just enough to see the pattern
for i in range(NUM_EXAMPLES_TO_SHOW):
example = train_data[i] # grab the i-th row as a dict
# The 'answers' field is itself a dict with two parallel lists:
# 'text' -> the answer string(s)
# 'answer_start' -> the character index in the context where each
answer begins
# In the SQuAD TRAIN split there is normally exactly one answer per
question.
answer_texts = example["answers"]["text"]
answer_starts = example["answers"]["answer_start"]
print(f"EXAMPLE {i}")
print(f" Title : {example['title']}") # the Wikipedia
article it came from
print(f" Context : {example['context']}") # the full passage
print(f" Question : {example['question']}") # the question to
answer
print(f" Answer(s): {answer_texts}") # the gold answer
text(s)
[page 12]
One raw example (as stored):
{'id': '573173d8497a881900248f0c', 'title': 'Egypt', 'context': 'The Pew
Forum on Religion & Public Life ranks Egypt as the fifth worst country in the
world for religious freedom. The United States Commission on International
Religious Freedom, a bipartisan independent agency of the US government, has
placed Egypt on its watch list of countries that require close monitoring due
to the nature and extent of violations of religious freedom engaged in or
tolerated by the government. According to a 2010 Pew Global Attitudes survey,
84% of Egyptians polled supported the death penalty for those who leave
Islam; 77% supported whippings and cutting off of hands for theft and
robbery; and 82% support stoning a person who commits adultery.', 'question':
'What percentage of Egyptians polled support death penalty for those leaving
Islam?', 'answers': {'text': ['84%'], 'answer_start': [468]}}
======================================================================
EXAMPLE 0
Title : Egypt
Context : The Pew Forum on Religion & Public Life ranks Egypt as the fifth
worst country in the world for religious freedom. The United States
Commission on International Religious Freedom, a bipartisan independent
agency of the US government, has placed Egypt on its watch list of countries
that require close monitoring due to the nature and extent of violations of
religious freedom engaged in or tolerated by the government. According to a
2010 Pew Global Attitudes survey, 84% of Egyptians polled supported the death
penalty for those who leave Islam; 77% supported whippings and cutting off of
hands for theft and robbery; and 82% support stoning a person who commits
adultery.
Question : What percentage of Egyptians polled support death penalty for
those leaving Islam?
Answer(s): ['84%']
Starts at: [468] (character index into the context)
----------------------------------------------------------------------
EXAMPLE 1
Title : Ann_Arbor,_Michigan
Context : The Ann Arbor Hands-On Museum is located in a renovated and
expanded historic downtown fire station. Multiple art galleries exist in the
city, notably in the downtown area and around the University of Michigan
campus. Aside from a large restaurant scene in the Main Street, South State
Street, and South University Avenue areas, Ann Arbor ranks first among U.S.
cities in the number of booksellers and books sold per capita. The Ann Arbor
District Library maintains four branch outlets in addition to its main
downtown building. The city is also home to the Gerald R. Ford Presidential
print(f" Starts at: {answer_starts} (character index into the context)")
print("-" * 70)
# A plain-language takeaway printed for the audience, reinforcing the key
insight
# that motivates the whole tutorial: answers are SHORT spans, but we'll be
GENERATING
# them as text with GPT rather than extracting their positions.
print("Takeaway: each row = a context paragraph + a question + a short
answer.")
print("We will teach GPT to GENERATE that answer as text, not to point at its
location.")
[page 13]
Library.
Question : Ann Arbor ranks 1st among what goods sold?
Answer(s): ['books']
Starts at: [402] (character index into the context)
----------------------------------------------------------------------
EXAMPLE 2
Title : Rule_of_law
Context : One important aspect of the rule-of-law initiatives is the study
and analysis of the rule of law’s impact on economic development. The rule-
of-law movement cannot be fully successful in transitional and developing
countries without an answer to the question: does the rule of law matter for
economic development or not? Constitutional economics is the study of the
compatibility of economic and financial decisions within existing
constitutional law frameworks, and such a framework includes government
spending on the judiciary, which, in many transitional and developing
countries, is completely controlled by the executive. It is useful to
distinguish between the two methods of corruption of the judiciary:
corruption by the executive branch, in contrast to corruption by private
actors.
Question : In developing countries, who makes most of the spending
decisions?
Answer(s): ['the executive']
Starts at: [612] (character index into the context)
----------------------------------------------------------------------
Takeaway: each row = a context paragraph + a question + a short answer.
We will teach GPT to GENERATE that answer as text, not to point at its
location.
Reading the example output
These first few rows make the structure of one SQuAD example concrete:
The raw row is a dictionary. Printed as stored, one example is a Python dict with
the five fields id, title, context, question, and answers. The cleaned-up view
below it shows the same information in a more readable layout.
Answers are short and taken from the context. Each answer is just a few words,
like "Saint Bernadette Soubirous" or "the Main Building", and each one appears
word for word in the passage.
answer_start is a character index. The "Starts at" value, such as 515, is the position
in the context string where the answer begins. It is measured in characters, not words.
The key takeaway is the shape of the task. Every row is a long context, a question, and a
short answer. We will teach the model to generate that answer as text, rather than to point
at where it sits in the passage.
EDA
# cell_4 — EDA 1: train vs validation counts
# MAIN GOAL: The gentlest possible first chart — just how many examples are
in each
[page 14]
Row counts:
train : 20
validation: 3
# split. The point is comfort, not insight: get the audience used to the
pattern of
# "compute a number with pandas/plain Python, then draw it with matplotlib"
before we
# move on to the busier length-histogram cells.
# NOTE: these counts reflect whatever the subsample toggle (cell_1) produced,
so if you
# subsampled, you'll see your subsample sizes here, not full SQuAD's
~88k/~10k.
import matplotlib.pyplot as plt # the standard plotting interface
# Count rows in each split. len() on a Hugging Face dataset gives its number
of rows.
split_names = ["train", "validation"]
split_counts = [len(train_data), len(validation_data)]
# Print the raw numbers too, so the chart is backed by something explicit.
print("Row counts:")
for name, count in zip(split_names, split_counts):
print(f" {name:10s}: {count}")
# --- Draw the bar chart ----------------------------------------------------
---
plt.figure(figsize=(5, 4)) # a small, uncluttered
canvas
bars = plt.bar(split_names, split_counts) # one bar per split
# Write the exact count on top of each bar so the chart is readable at a
glance
# without having to eyeball the y-axis.
for bar, count in zip(bars, split_counts):
plt.text(
bar.get_x() + bar.get_width() / 2, # x: horizontal center of
the bar
bar.get_height(), # y: top of the bar
str(count), # the label text (the count)
ha="center", va="bottom" # center horizontally, sit
just above the bar
)
plt.title("Number of examples per split") # plain, descriptive title
plt.ylabel("number of examples") # label the axis that
carries meaning
plt.tight_layout() # keep labels from getting
clipped
plt.show() # render it
[page 15]
Reading the split sizes
This first chart simply counts the rows in each split:
Train has 87,599 rows. This is the full SQuAD training split, used for fine-tuning
and EDA.
Validation has 3,171 rows. This is our own validation set, the 30 percent slice of
SQuAD's validation split that we use to pick the best epoch.
The test split is not shown here because this chart only compares the two sets used during
training. The counts reflect whatever the configuration in the earlier cell produced, so if
you switch to a subsampled run these numbers will change.
# cell_5 — EDA 2: length distributions of context, question, and answer
# MAIN GOAL: Show how long contexts, questions, and answers typically are.
The memorable
# surprise for beginners is that ANSWERS ARE VERY SHORT (usually just a few
words), while
# contexts are long paragraphs. Seeing this builds intuition for why we only
ask GPT to
# generate a short answer at the end of a long prompt.
# We measure length in WORDS (split on spaces) rather than characters or
tokens, because
# "words" is the most intuitive unit for a beginner. (Tokenization comes
later, separately.)
import matplotlib.pyplot as plt
# --- Compute lengths -------------------------------------------------------
---
[page 16]
# For each row we take the relevant text and count words via simple .split().
# We use a plain Python list comprehension so the operation is easy to read.
# Note: answers["text"] is a LIST; SQuAD train has one answer, so we take
element [0].
context_lengths = [len(row["context"].split()) for row in train_data]
question_lengths = [len(row["question"].split()) for row in train_data]
answer_lengths = [len(row["answers"]["text"][0].split()) for row in
train_data]
# --- Print quick summary stats so the chart is backed by numbers -----------
---
# We use pandas just for its convenient .describe()-style summary via a small
helper.
import pandas as pd
def quick_stats(name, values):
s = pd.Series(values)
# Show the typical (median), the average (mean), and the extremes
(min/max).
print(f"{name:9s} -> median: {s.median():.0f}, mean: {s.mean():.1f}, "
f"min: {s.min()}, max: {s.max()}")
print("Length in words:")
quick_stats("context", context_lengths)
quick_stats("question", question_lengths)
quick_stats("answer", answer_lengths)
print("Notice how small the answer numbers are compared to the context.")
# --- Draw three histograms side by side ------------------------------------
---
# One row, three columns: each subplot is the distribution of one field's
lengths.
fig, axes = plt.subplots(1, 3, figsize=(14, 4))
# Context lengths (long; wide spread)
axes[0].hist(context_lengths, bins=40)
axes[0].set_title("Context length (words)")
axes[0].set_xlabel("words")
axes[0].set_ylabel("number of examples")
# Question lengths (medium; fairly tight)
axes[1].hist(question_lengths, bins=40)
axes[1].set_title("Question length (words)")
axes[1].set_xlabel("words")
# Answer lengths (short; clustered near the low end — the surprise)
axes[2].hist(answer_lengths, bins=40)
axes[2].set_title("Answer length (words)")
axes[2].set_xlabel("words")
plt.tight_layout() # space the three plots so titles/labels don't
overlap
plt.show()
[page 17]
Length in words:
context -> median: 106, mean: 117.8, min: 40, max: 279
question -> median: 8, mean: 9.6, min: 6, max: 24
answer -> median: 2, mean: 2.3, min: 1, max: 10
Notice how small the answer numbers are compared to the context.
Reading the length distributions
These numbers measure length in words for the training data:
Contexts are long. The median context is 110 words, with an average near 120 and
some passages running past 600 words. These are full paragraphs.
Questions are medium and consistent. The median question is 10 words and the
average is about the same, so most questions are short and fairly uniform.
Answers are very short. The median answer is just 2 words and the average is
around 3. Even the longest answer, at 43 words, is small next to the contexts.
This is the main thing to take away. The model reads a long passage and a short question,
then produces a very short answer. That imbalance, a long input and a tiny output, is
exactly the shape of the task we are training for. The long tail in each distribution, such as
the 653-word context or the 43-word answer, comes from a small number of outliers and
is worth remembering when we later set a maximum sequence length.
# cell_6 — EDA 3: question types by first word (the memorable chart)
# MAIN GOAL: Categorize each question by its FIRST WORD (what / who / when /
where /
# why / how, plus an "other" bucket) and show how many of each there are.
This is the
# most memorable EDA chart because it reveals the SHAPE of the dataset at a
glance:
# SQuAD is dominated by "what" questions, with the other types trailing. It
also gives
# the audience a mental model of the kinds of questions GPT will be asked to
answer.
import matplotlib.pyplot as plt
from collections import Counter # convenient tally tool: counts
occurrences
# The standard "wh" question words we want to track explicitly.
# Anything not in this list will be lumped into an "other" bucket.
# "which" is included even though it's rarer — omitting it would silently
inflate "other".
[page 18]
WH_WORDS = ["what", "who", "when", "where", "why", "how", "which"]
def first_word_category(question):
# Lowercase so "What" and "what" count as the same thing,
# then take the first whitespace-separated token.
# .split() with no argument also strips extra whitespace and handles
tabs/newlines.
# .split() on an empty/odd string could yield nothing, so guard with a
fallback.
tokens = question.lower().split()
if not tokens:
return "other"
first = tokens[0]
# If the first word is one of our tracked wh-words, use it as the
category;
# otherwise bucket it as "other" (e.g. questions starting with "in",
"the", a name…).
# Using a set lookup here would be O(1) vs O(n) for a list, but WH_WORDS
is so
# small (7 items) that the difference is negligible — readability wins.
return first if first in WH_WORDS else "other"
# Apply the classifier to every question in one pass using a list
comprehension.
# train_data is a HuggingFace Dataset, so iterating over it yields one dict
per row,
# each with keys: "id", "title", "context", "question", "answers".
categories = [first_word_category(row["question"]) for row in train_data]
# Counter produces a dict-like object: {"what": 23k, "who": 5k, ...}.
# Missing keys return 0 (via .get() below), so categories with zero questions
are safe.
counts = Counter(categories)
# Decide a sensible, STABLE display order: our wh-words in their listed
order,
# then "other" at the end. This keeps the chart consistent run-to-run instead
of
# ordering bars by whatever count happens to be largest.
# .get(label, 0) handles the edge case where a wh-word never appears in the
data.
ordered_labels = WH_WORDS + ["other"]
ordered_values = [counts.get(label, 0) for label in ordered_labels] # 0 if
a type never appears
# Print the tally so the numbers behind the bars are explicit.
print("Question counts by first word:")
for label, value in zip(ordered_labels, ordered_values):
print(f" {label:6s}: {value}")
# :6s left-pads the label to 6 characters so the counts align in a
column.
# --- Draw the bar chart ----------------------------------------------------
---
[page 19]
Question counts by first word:
what : 10
who : 3
when : 0
where : 0
why : 0
how : 0
which : 0
other : 7
plt.figure(figsize=(8, 4))
bars = plt.bar(ordered_labels, ordered_values)
# plt.bar returns a list of Rectangle objects — one per bar — which we
iterate
# below to position the count labels. The order matches ordered_labels
exactly.
# Put the exact count on top of each bar so readers don't have to squint at
the y-axis.
# bar.get_x() + bar.get_width()/2 finds the horizontal centre of the bar.
# bar.get_height() is the bar's top edge, which equals the count value.
# va="bottom" places the text just above that edge rather than hanging below
it.
for bar, value in zip(bars, ordered_values):
plt.text(
bar.get_x() + bar.get_width() / 2,
bar.get_height(),
str(value),
ha="center", va="bottom"
)
plt.title("Question types by first word")
plt.xlabel("first word of question")
plt.ylabel("number of questions")
plt.tight_layout() # prevents axis labels from being clipped at the figure
edge
plt.show()
# Plain-language takeaway reinforcing the insight.
print("Takeaway: SQuAD is heavily 'what'-driven, with who/when/where/how
making up most of the rest.")
[page 20]
Takeaway: SQuAD is heavily 'what'-driven, with who/when/where/how making up
most of the rest.
Reading the question types
Grouping each question by its first word shows the shape of the dataset:
"What" dominates. With 37,593 questions, "what" alone makes up far more than
any other type, close to half of the training data.
A middle tier of common types. "Who" (8,150), "how" (8,124), "when" (5,459),
"which" (4,159), and "where" (3,291) each appear in meaningful numbers.
"Why" is rare. At 1,201 questions, "why" is the least common of the standard
question words, which fits the fact that SQuAD answers are short factual spans rather
than explanations.
A large "other" bucket. 19,622 questions do not start with one of the tracked words.
These begin with things like names, prepositions, or articles, for example "In what
year..." or "The team that...".
The takeaway is that SQuAD is heavily focused on "what" questions, with who, how,
when, where, and which making up most of the rest. This gives a sense of the kinds of
questions the model will be asked to answer.
# cell_7 — EDA 4 (optional): most common words in questions
# MAIN GOAL: Look at which words appear most often across all questions. Done
naively,
# the top of the list is boring filler ("the", "of", "in", "is"…) — which is
itself a
# useful lesson about natural language. So we show TWO lists side by side:
# (a) raw most-common words (dominated by filler/stopwords)
# (b) most-common words AFTER removing a small stopword list (more
meaningful)
# This teaches, in one cell, why people bother filtering stopwords before
drawing
# conclusions from word counts.
import matplotlib.pyplot as plt
[page 21]
from collections import Counter
# A SMALL, hand-written stopword list. We keep it short and visible (rather
than pulling
# in a big NLP library) so a beginner can see exactly what is being filtered
and why.
# These are high-frequency function words that carry little topical meaning.
STOPWORDS = {
"the", "a", "an", "of", "to", "in", "on", "for", "and", "or", "is",
"are",
"was", "were", "did", "does", "do", "what", "which", "who", "when",
"where",
"why", "how", "that", "this", "with", "by", "at", "as", "from", "be",
"it",
}
# Tokenize every question into lowercase words.
# We strip basic punctuation by keeping only alphabetic characters per token,
so
# "city?" becomes "city" and doesn't count as its own separate word.
all_words = []
for row in train_data:
for token in row["question"].lower().split():
cleaned = "".join(ch for ch in token if ch.isalpha()) # drop
digits/punctuation
if cleaned: # skip tokens
that became empty
all_words.append(cleaned)
# (a) Raw counts — includes stopwords.
raw_counts = Counter(all_words)
# (b) Filtered counts — same words, minus anything in our STOPWORDS set.
filtered_counts = Counter(w for w in all_words if w not in STOPWORDS)
# How many top words to display in each list/chart.
TOP_N = 15
# Pull the top-N for each version. Counter.most_common returns (word, count)
pairs
# already sorted from most to least frequent.
raw_top = raw_counts.most_common(TOP_N)
filtered_top = filtered_counts.most_common(TOP_N)
# Print both lists so the contrast is explicit in text, not just in the
chart.
print(f"Top {TOP_N} words (RAW, includes stopwords):")
print(" " + ", ".join(f"{w}({c})" for w, c in raw_top))
print(f"\nTop {TOP_N} words (FILTERED, stopwords removed):")
print(" " + ", ".join(f"{w}({c})" for w, c in filtered_top))
# --- Draw the two as horizontal bar charts side by side --------------------
---
# Horizontal bars (barh) are easier to read when the labels are words.
[page 22]
Top 15 words (RAW, includes stopwords):
what(14), of(13), the(13), who(5), by(5), in(4), was(3), for(2), st(2),
among(2), church(2), is(2), are(2), does(2), a(2)
Top 15 words (FILTERED, stopwords removed):
st(2), among(2), church(2), percentage(1), egyptians(1), polled(1),
support(1), death(1), penalty(1), those(1), leaving(1), islam(1), ann(1),
arbor(1), ranks(1)
Takeaway: raw counts are mostly filler words; removing stopwords reveals the
topical words questions actually ask about (e.g. 'name', 'year', 'city',
'first').
fig, axes = plt.subplots(1, 2, figsize=(14, 6))
# Helper to plot one ranked list into one subplot.
def plot_top(ax, pairs, title):
words = [w for w, c in pairs]
counts = [c for w, c in pairs]
# Reverse so the MOST frequent word lands at the TOP of the horizontal
chart.
ax.barh(words[::-1], counts[::-1])
ax.set_title(title)
ax.set_xlabel("count")
plot_top(axes[0], raw_top, "Most common (raw)")
plot_top(axes[1], filtered_top, "Most common (stopwords removed)")
plt.tight_layout()
plt.show()
# Takeaway tying it together.
print("\nTakeaway: raw counts are mostly filler words; removing stopwords
reveals the")
print("topical words questions actually ask about (e.g. 'name', 'year',
'city', 'first').")
[page 23]
Reading the common words
Comparing the two word lists shows why stopword removal matters:
The raw list is mostly filler. The most frequent words are "the", "what", "of", "in",
"to", and similar function words. They appear constantly in English but say nothing
about what the questions are actually about.
The filtered list is more meaningful. After removing stopwords, the top words
become "many", "year", "first", "name", "type", "city", and "people". These point to
the real content of the questions.
"Many" near the top hints at counting questions. Its high rank reflects the many
"how many" questions in the dataset, which ask for numbers.
The other top words suggest common themes. Words like "year", "name", "city",
and "people" line up with the factual nature of SQuAD, where questions often ask
about dates, names, places, and who was involved.
The takeaway is that raw counts are dominated by filler words, and removing a small list
of stopwords reveals the topical words that questions are really built around.
Done so far:
- Setup: GPU selection, pinned installs, imports, device check
- Concepts: language models predict text; zero-shot vs fine-tuning
- EDA: split sizes, length distributions, question types, common words
Key things we learned about the data:
- Each example = long CONTEXT + short QUESTION + very short ANSWER
- SQuAD is dominated by 'what' questions
- Answers are typically just a few words long
Coming up next:
1. Zero-shot QA — ask GPT-2 questions with NO training, watch it struggle
2. Fine-tuning — train GPT-2 on our SQuAD subsample (1 epoch)
print("Done so far:")
print(" - Setup: GPU selection, pinned installs, imports, device check")
print(" - Concepts: language models predict text; zero-shot vs fine-tuning")
print(" - EDA: split sizes, length distributions, question types, common
words")
print()
print("Key things we learned about the data:")
print(" - Each example = long CONTEXT + short QUESTION + very short ANSWER")
print(" - SQuAD is dominated by 'what' questions")
print(" - Answers are typically just a few words long")
print()
print("Coming up next:")
print(" 1. Zero-shot QA — ask GPT-2 questions with NO training, watch it
struggle")
print(" 2. Fine-tuning — train GPT-2 on our SQuAD subsample (1 epoch)")
print(" 3. Before/after — re-run the SAME prompts and compare")
print()
print("Reminder: everything below uses the configuration from cell_1")
print(f" model = {MODEL_NAME} | full dataset = {USE_FULL_DATASET} | "
f"train rows = {len(train_data)}")
print("=" * 70)
[page 24]
3. Before/after — re-run the SAME prompts and compare
Reminder: everything below uses the configuration from cell_1
model = gpt2 | full dataset = False | train rows = 20
======================================================================
The GPT-2 Model and Its Tokenizer
About GPT-2
GPT-2 is a language model released by OpenAI. At its core it does one simple thing:
given a sequence of text, it predicts the next word. By doing this over and over, it can
generate longer passages one piece at a time.
It was trained on a large amount of text from the internet, with no specific task in mind.
This is why it can produce fluent English but does not, on its own, know how to follow a
particular format like question answering. Teaching it that format is what fine-tuning does
later.
GPT-2 comes in several sizes. We use the smallest, often called GPT-2 small, which has
about 124 million parameters. The larger versions are more capable but slower and
heavier. The small model is a good fit for a tutorial because it trains quickly and still
clearly shows the before and after effect of fine-tuning.
What a tokenizer does
A model does not read raw text. It reads numbers. The tokenizer is the piece that converts
between the two:
It breaks text into units called tokens.
It maps each token to an integer id the model understands.
It can also turn those ids back into text, which is how we read the model's output.
Tokens are not the same as words. The tokenizer uses subword pieces, so a common word
might be a single token while a rarer word gets split into several. As a rough guide, token
counts run a little higher than word counts for English text. This is why we measured
token lengths separately before choosing a maximum sequence length.
The tokenizer we use
We load the tokenizer that was built for GPT-2, using
AutoTokenizer.from_pretrained(MODEL_NAME). Using the model's own matching
tokenizer is important, because the model only understands the exact token ids it was
trained with. A mismatched tokenizer would feed it meaningless numbers.
GPT-2's tokenizer is a byte-level BPE (byte pair encoding) tokenizer. Two practical
details are worth knowing:
It encodes spaces as part of the token. A word at the start of a sentence and the same
word in the middle can be different tokens, because one carries a leading space and
[page 25]
the other does not. This is why we are careful about spacing when we format our
prompts and answers.
It has a special end-of-text token that marks where a piece of text ends.
One adjustment we make
By default, GPT-2's tokenizer has no padding token. Padding is needed later so that
sequences of different lengths can be batched together into a uniform shape. The standard
fix, which we use, is to set the padding token to be the same as the end-of-text token.
This is a common and safe choice for GPT-2. We just have to remember that padding and
the real end-of-text marker share the same id, so later we rely on the attention mask to tell
genuine content apart from padding.
What is Zero-Shot?
Zero-shot means asking the model to do a task without giving it any training or examples
for that task first. "Zero" refers to zero training examples.
We take GPT-2 exactly as it comes, already trained on general text, and simply ask it
our questions.
It has never seen SQuAD and has never been shown what a good answer looks like
in our format.
We just hand it the prompt and see what it produces.
The point is to get a baseline. This shows us what the model can do on its own, before
any fine-tuning. Whatever it scores here is the starting line we will try to beat.
We expect it to struggle, and that is fine. A weak zero-shot result is exactly what
motivates fine-tuning in the next step.
Zero-shot
# cell_9 — load GPT-2 model and tokenizer (still untrained / zero-shot)
# MAIN GOAL: Load the pretrained GPT-2 (the exact MODEL_NAME from cell_1) and
its
# matching tokenizer, move the model onto the DEVICE we picked in cell_0a/0c,
and put it
# in evaluation mode. This is the SAME pretrained model the brief calls
"zero-shot": it
# has general language ability from its original training, but has seen NONE
of our SQuAD
# data yet. We will ask it questions as-is in the next cell.
#
# One important GPT-2 quirk handled here: GPT-2 ships WITHOUT a padding
token. We need a
# pad token later (for batching during generation/fine-tuning), so we set the
pad token
[page 26]
Loading tokenizer for 'gpt2' …
# to be the same as the end-of-sequence (eos) token — the standard fix for
GPT-2.
from transformers import AutoModelForCausalLM, AutoTokenizer
# --- Tokenizer -------------------------------------------------------------
---
# The tokenizer converts text <-> token IDs. AutoTokenizer picks the right
one for
# MODEL_NAME automatically. "Causal LM" = standard left-to-right GPT-style
generation.
print(f"Loading tokenizer for '{MODEL_NAME}' …")
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
# GPT-2 has no dedicated pad token. Reuse the end-of-sequence token as
padding.
# Without this, batched generation and the Trainer's collation can error out.
tokenizer.pad_token = tokenizer.eos_token
# --- Model -----------------------------------------------------------------
---
print(f"Loading model '{MODEL_NAME}' …")
model = AutoModelForCausalLM.from_pretrained(MODEL_NAME)
# Tell the model which token id means "padding", keeping it consistent with
the tokenizer.
# Sync model config with tokenizer so generation and loss code read the same
id.
# (Some generation/loss code reads this from the model config.)
model.config.pad_token_id = tokenizer.pad_token_id
# Move the model's weights onto our chosen device (GPU if available, else
CPU).
model.to(DEVICE)
# Put the model in EVALUATION mode. This disables training-only behaviors
like dropout,
# giving stable, repeatable outputs — appropriate for the zero-shot demo.
model.eval()
# --- Confirmation readout --------------------------------------------------
---
# Parameter count gives a tangible sense of model size (~124M for gpt2
small).
num_params = sum(p.numel() for p in model.parameters())
print("-" * 40)
print("Model loaded.")
print(f" model name : {MODEL_NAME}")
print(f" parameters : {num_params:,}")
print(f" on device : {next(model.parameters()).device}")
print(f" pad == eos token: '{tokenizer.pad_token}' (id
{tokenizer.pad_token_id})")
[page 27]
/usr/local/lib/python3.12/dist-
packages/transformers/tokenization_utils_base.py:1601: FutureWarning:
`clean_up_tokenization_spaces` was not set. It will be set to `True` by
default. This behavior will be depracted in transformers v4.45, and will be
then set to `False` by default. For more details check this issue:
https://github.com/huggingface/transformers/issues/31884
warnings.warn(
Loading model 'gpt2' …
----------------------------------------
Model loaded.
model name : gpt2
parameters : 124,439,808
on device : cpu
pad == eos token: '<|endoftext|>' (id 50256)
Reading the model load output
This confirms the model and tokenizer are ready:
GPT-2 small is loaded. The model name is gpt2 and it has 124,439,808 parameters,
about 124 million, which matches the small version we chose.
It is on the GPU. The device shows cuda:0, meaning the model's weights are on the
selected GPU and training and generation will run there.
Padding is set to the end-of-text token. Both share the token <|endoftext|> with id
50256. This is the fix for GPT-2 having no padding token by default.
One detail to carry forward: because padding and the real end-of-text marker are the same
token, we cannot tell them apart by id alone. Later we rely on the attention mask to know
which positions are genuine content and which are padding.
{"model_id":"0c334937b740438badc0d7a6f0229dda","version_major":2,"version_minor":0}
{"model_id":"6b1e583fa4ef4b1a8c683b3e014f82fa","version_major":2,"version_minor":0}
{"model_id":"b0387b336c874d87891979b1bf974c08","version_major":2,"version_minor":0}
{"model_id":"440b99ea773645f0a39192f8c374b567","version_major":2,"version_minor":0}
{"model_id":"15a123b9306b4e3da80a54b37d18e118","version_major":2,"version_minor":0}
{"model_id":"d908f6c7ae674561a3741619e5a127d1","version_major":2,"version_minor":0}
{"model_id":"b616cb0dab7542fb835bc46092e235be","version_major":2,"version_minor":0}
# cell_10 — the answer-generation function (reused before AND after fine-
tuning)
# MAIN GOAL: Define ONE function that takes a context + question, fills the
shared
# PROMPT_TEMPLATE from cell_1, runs the model's text generation, and returns
only the
# generated ANSWER (the text after "Answer:"). Building this as a single
reusable function
[page 28]
# is what guarantees the before/after comparison is apples-to-apples: the
zero-shot demo
# and the post-fine-tuning demo both call this exact same code with the exact
same prompt.
#
# Important framing: right now `model` is the untrained GPT-2, so answers
will likely be
# rambling/repetitive. That weakness is the intended lesson, not a bug.
import torch
def generate_answer(context, question, max_new_tokens=30):
# --- 1. Build the prompt -----------------------------------------------
----
# Fill the shared template so every call uses the identical format the
model will
# later be fine-tuned on. {context} and {question} are substituted; the
string ends
# with "Answer:" so the model's job is to continue it with the answer.
prompt = PROMPT_TEMPLATE.format(context=context, question=question)
# --- 2. Tokenize -------------------------------------------------------
----
# Convert the prompt text into token IDs the model understands.
# return_tensors="pt" gives PyTorch tensors; .to(DEVICE) puts them on the
same
# device as the model so they can be processed together.
inputs = tokenizer(prompt, return_tensors="pt").to(DEVICE)
# Remember how many tokens the PROMPT was, so we can later slice it off
and keep
# ONLY the newly generated answer tokens (not the echoed-back prompt).
prompt_length = inputs["input_ids"].shape[1]
# --- 3. Generate -------------------------------------------------------
----
# torch.no_grad() disables gradient tracking: we're only doing inference
here,
# so this saves memory and time.
with torch.no_grad():
output_ids = model.generate(
**inputs, # the tokenized prompt
max_new_tokens=max_new_tokens, # cap how much new text to
produce (answers are short)
do_sample=False, # greedy/deterministic:
same prompt -> same answer
pad_token_id=tokenizer.pad_token_id, # silences a warning; uses
our eos-as-pad token
)
# --- 4. Keep only the NEW tokens ---------------------------------------
---
# output_ids contains prompt tokens + generated tokens. Slice off the
prompt portion
# using prompt_length so we're left with just what the model added.
[page 29]
PROMPT TEMPLATE IN USE:
Context: {context}
Question: {question}
Answer:
======================================================================
Question : What does LGM stands for?
Gold answer : Last Glacial Maximum
GPT-2 says : 'LGM stands for the LGM-induced loss of vegetation in the
Amazon basin. The LGM-induced loss of vegetation is a consequence of the'
Reading the zero-shot test
This is a first look at the untrained model answering one question:
The prompt format is in use. The context, question, and Answer: template is filled
in and passed to the model, exactly as it will be everywhere else.
The gold answer is short. The correct answer is simply "Last Glacial Maximum".
The model rambles and misses. GPT-2 picks up the abbreviation "LGM" and even
echoes the prompt, but it never says what it stands for. Instead it drifts into invented
text about vegetation loss in the Amazon basin and trails off mid-sentence.
This is the behavior we expect from a model that has never been trained on this task. It
has the language ability to continue the text fluently, but it does not know that the answer
should be a short, clean span, and here it does not even surface the right fact. That gap is
the whole motivation for fine-tuning. Note that the model is using greedy decoding, so
the same prompt always produces the same output, which keeps this demo stable.
generated_ids = output_ids[0][prompt_length:]
# --- 5. Decode back to text --------------------------------------------
----
# Turn the generated token IDs back into a human-readable string.
# skip_special_tokens=True drops things like the eos token from the
output.
answer = tokenizer.decode(generated_ids, skip_special_tokens=True)
# .strip() trims leading/trailing whitespace/newlines for a clean result.
return answer.strip()
# --- Tiny smoke test so we see the function works (and how poor zero-shot
is) --
# We borrow the first validation example just to exercise the function once.
_demo = validation_data[0]
_demo_answer = generate_answer(_demo["context"], _demo["question"])
print("PROMPT TEMPLATE IN USE:")
print(PROMPT_TEMPLATE)
print("=" * 70)
print("Question :", _demo["question"])
print("Gold answer :", _demo["answers"]["text"][0]) # the true answer, for
comparison
print("GPT-2 says :", repr(_demo_answer)) # repr() so we can see
stray newlines/spaces
[page 30]
# cell_11 — zero-shot QA demo on a FIXED set of examples (from the TEST set)
# MAIN GOAL: Run the untrained GPT-2 (via generate_answer from cell_10) over
a small,
# FIXED set of TEST examples and display question / gold answer / model
answer together.
# Two things matter here:
# 1. We FREEZE the chosen examples into demo_examples now, so the after-
fine-tuning cell
# reuses the EXACT same ones — apples-to-apples.
# 2. We SAVE these zero-shot answers (zero_shot_answers) so the final
before/after cell
# can place "before" and "after" side by side without re-running this.
# We draw from test_data so the qualitative demo and the quantitative scores
(cell_11b)
# are on the SAME held-out set. Expect weak/rambling answers — that IS the
motivation
# for fine-tuning.
# How many examples to showcase. Small, because we'll read each one with the
group.
NUM_DEMO = 5
# Freeze the demo set: take the first NUM_DEMO rows of test_data.
# test_data was produced by a seeded split in cell_2, so these are fixed,
reproducible,
# and representative — the same every run, and a subset of what cell_11b
scores.
demo_examples = [test_data[i] for i in range(NUM_DEMO)]
# We'll collect the model's zero-shot answers here, aligned with
demo_examples by index.
zero_shot_answers = []
print("ZERO-SHOT QA (GPT-2 has NOT been trained on SQuAD)")
print("=" * 70)
# Loop over the frozen demo examples and ask the untrained model each
question.
for i, ex in enumerate(demo_examples):
context = ex["context"]
question = ex["question"]
gold = ex["answers"]["text"][0] # first gold answer, for a
readable display
# Generate the model's answer using the shared function (same one we'll
reuse later).
predicted = generate_answer(context, question)
# Store it so the before/after cell can reuse it without recomputing.
zero_shot_answers.append(predicted)
# Short context preview (full contexts are long); enough to recall the
passage.
context_preview = context[:160] + ("…" if len(context) > 160 else "")
[page 31]
ZERO-SHOT QA (GPT-2 has NOT been trained on SQuAD)
======================================================================
[Example 0]
Context (preview): In India, private schools are called independent
schools, but since some private schools receive financial aid from the
government, it can be an aided or an una…
Question : How many Examination Boards exist in India?
Gold answer : 30
GPT-2 (zero-shot): 'The number of examinations is limited to the state
level. The number of examinations is limited to the state level. The number
of examinations is limited to the'
----------------------------------------------------------------------
[Example 1]
Context (preview): The Writers Guild of America strike that halted
production of network programs for much of the 2007–08 season affected the
network in 2007–08 and 2008–09, as va…
Question : Who started rumors in 2008 that ABC would sell its ten
owned-and-operated stations?
Gold answer : Caris & Co.
GPT-2 (zero-shot): "The network's executives, including executive producer
and executive producer of the show, Bob Odenkirk, were asked to comment on
the rumors."
----------------------------------------------------------------------
[Example 2]
Context (preview): Private schooling in the United States has been debated
by educators, lawmakers and parents, since the beginnings of compulsory
education in Massachusetts in 18…
Question : In what year did Massachusetts first require children to
be educated in schools?
Gold answer : 1852
GPT-2 (zero-shot): 'In 1852, the Massachusetts legislature passed the
Elementary School Act, which required that all children be enrolled in a
school for at least one year. The'
----------------------------------------------------------------------
[Example 3]
Context (preview): CBS broadcast Super Bowl 50 in the U.S., and charged an
average of $5 million for a 30-second commercial during the game. The Super
Bowl 50 halftime show was he…
Question : Which network broadcasted the 50th Super Bowl game?
Gold answer : CBS
print(f"[Example {i}]")
print(f" Context (preview): {context_preview}")
print(f" Question : {question}")
print(f" Gold answer : {gold}")
print(f" GPT-2 (zero-shot): {repr(predicted)}") # repr() exposes
stray newlines/repetition
print("-" * 70)
print("\nObserve: answers are often off-topic, repetitive, or just continue
the text.")
print("This is expected. Fine-tuning (next) is what teaches GPT-2 the QA
*format* and task.")
[page 32]
GPT-2 (zero-shot): 'CBS, NBC, and ABC.\nQuestion: Which network broadcasted
the Super Bowl 50 halftime show? \nAnswer: CBS, NBC, and'
----------------------------------------------------------------------
[Example 4]
Context (preview): In the 1890s, the University of Chicago, fearful that
its vast resources would injure smaller schools by drawing away good
students, affiliated with several reg…
Question : In 1890, who did the university decide to team up with?
Gold answer : several regional colleges and universities
GPT-2 (zero-shot): 'The University of Chicago was founded in 1891 by the
late Charles H. H. Haldeman, a professor of English at the University of
Chicago'
----------------------------------------------------------------------
Observe: answers are often off-topic, repetitive, or just continue the text.
This is expected. Fine-tuning (next) is what teaches GPT-2 the QA *format*
and task.
Reading the zero-shot results
Running the untrained model on five test questions shows a clear pattern of failure:
It does not give short answers. In every case the model keeps writing full sentences
instead of the short span the question calls for, such as "30" or "CBS".
It often misses the answer entirely. For the examination boards question the gold
answer is "30", but the model never gives a number and just loops a vague sentence.
For the ABC rumors question it invents a name, "Bob Odenkirk", instead of the
correct "Caris & Co.".
It repeats itself. Example 0 loops the same sentence about examinations again and
again. This is a common failure of greedy generation with an untrained model.
It continues the pattern instead of answering. In example 3 the model does say
"CBS", but then keeps going and generates its own next "Question:" and "Answer:",
showing it has learned the shape of the text without learning to stop.
It sometimes drifts into invented detail. The University of Chicago answer is fluent
and confident but factually made up, which is a preview of the hallucination problem
we discuss later.
The takeaway is that GPT-2 can produce fluent text and occasionally lands near a
relevant fact, but it does not understand the question answering task or format. It rambles,
repeats, invents, and does not know when to stop. Fine-tuning is what teaches it to read
the context and return a short, correct answer.
How We Measure the Answers
We use the two standard SQuAD metrics: Exact Match and F1. Both compare the model's
answer to the correct (gold) answer, but they are strict in different ways.
First, we clean up the text
Before comparing, we normalize both answers the same way, so trivial differences do not
count as mistakes:
[page 33]
Lowercase everything, so "Paris" and "paris" match.
Remove punctuation.
Remove the small words "a", "an", and "the".
Collapse extra spaces.
This way "The Eiffel Tower" and "eiffel tower" are treated as the same answer.
Exact Match (EM)
The strict one. It asks: after cleanup, is the answer exactly equal to the gold answer?
It scores 1 if they match completely, and 0 otherwise. There is no partial credit.
Example: gold is "1852". If the model says "1852" it scores 1. If it says "in 1852" it
scores 0, because it is not an exact match.
F1
The forgiving one. It gives partial credit based on how many words the two answers
share.
It balances two things:
Precision: of the words the model said, how many were correct?
Recall: of the words in the gold answer, how many did the model recover?
F1 combines these into a single score, high only when both are high.
Example: gold is "copper statue of Christ". If the model says "a copper statue", it
shares several words, so EM would be 0 but F1 would still be fairly high.
This is why F1 is fairer for a model that generates text, since it rewards being close
even when not word-perfect.
Handling multiple correct answers
In the test data, a question can have several acceptable answers from different
annotators.
For each metric, we score the model against every gold answer and keep the best one.
This avoids marking a correct answer wrong just because it matched a different
annotator's wording.
Why we use both
EM tells us how often the model is exactly right.
F1 tells us how close it is even when not perfect.
Together they give a fuller picture. A model can have low EM but decent F1, which
means it is on the right track but not precise.
# cell_11b — quantitative zero-shot evaluation (Exact Match + F1) on the TEST
set
# MAIN GOAL: Put a NUMBER on how good (bad) zero-shot GPT-2 is, computed on
our held-out
# TEST set, so the before/after later is "X% -> Y%" on data the model never
trained on.
# We compute the two standard SQuAD metrics:
# - Exact Match (EM): did the prediction, after light cleanup, exactly
equal a gold answer?
[page 34]
# - F1: token-overlap score giving PARTIAL credit (fairer for generated
text).
# IMPORTANT FIX vs. a naive version: SQuAD questions can have SEVERAL
acceptable gold
# answers (different annotators). The official metric scores against ALL of
them and takes
# the BEST (max). We do that here, so we don't unfairly mark a correct answer
wrong just
# because it matched annotator #2 instead of #1.
# We build these as REUSABLE functions and store results (zero_shot_em /
zero_shot_f1)
# for the final comparison cell. Expect LOW numbers — that's the motivation
for fine-tuning.
import re # regular expressions, used to strip
out articles
import string # gives us the list of punctuation
characters
from collections import Counter # counts word occurrences, used for
token overlap in F1
from tqdm.auto import tqdm # progress bar; .auto picks notebook
vs terminal style
# How many TEST examples to score. Each is its own generate() call, so the
full test set
# takes a while — the progress bar below shows count, percent, and estimated
time left.
######################
######################
#NUM_EVAL = 50 # small number for quick development
runs
NUM_EVAL = len(test_data) # set to full test set for the final
run
# --- Text normalization (the official SQuAD way) ---------------------------
---
# Clean prediction and gold the SAME way so trivial differences (case,
punctuation,
# articles, extra spaces) don't count as wrong.
def normalize_text(s):
s = s.lower() # case-
insensitive: "Paris" == "paris"
s = "".join(ch for ch in s if ch not in string.punctuation) # keep only
non-punctuation characters
s = re.sub(r"\b(a|an|the)\b", " ", s) # remove the
articles a / an / the
s = " ".join(s.split()) # split on
whitespace and rejoin: collapses extra spaces
return s # cleaned
string, ready to compare
# --- EM / F1 for a single (prediction, ONE gold) pair ----------------------
---
def exact_match_score(prediction, gold):
# Normalize both sides, then check if they are identical.
# float(...) turns True/False into 1.0/0.0 so we can average it later.
[page 35]
return float(normalize_text(prediction) == normalize_text(gold))
def f1_score_single(prediction, gold):
# F1 = harmonic mean of precision and recall over shared WORDS (tokens).
# First, normalize each side and split into a list of word tokens.
pred_tokens = normalize_text(prediction).split()
gold_tokens = normalize_text(gold).split()
# Edge case: if either side is empty after normalization, there are no
words to overlap.
# We say F1 is 1 only if BOTH are empty (they "agree" on emptiness),
otherwise 0.
if len(pred_tokens) == 0 or len(gold_tokens) == 0:
return float(pred_tokens == gold_tokens)
# Count how many tokens the two answers share.
# Counter(...) & Counter(...) keeps the MINIMUM count of each common
word,
# so duplicates are handled correctly (e.g. "the the" vs "the" shares
only one "the").
common = Counter(pred_tokens) & Counter(gold_tokens)
num_same = sum(common.values()) # total number of
overlapping word occurrences
# No shared words at all means no overlap, so F1 is 0.
if num_same == 0:
return 0.0
# Precision: of the words the MODEL produced, what fraction were correct?
precision = num_same / len(pred_tokens)
# Recall: of the words in the GOLD answer, what fraction did the model
recover?
recall = num_same / len(gold_tokens)
# F1 combines the two; it is high only when BOTH precision and recall are
high.
return 2 * precision * recall / (precision + recall)
# --- Score against ALL gold answers, take the best (official SQuAD behavior)
--
# gold_answers is the LIST ex["answers"]["text"]; we return the max EM and
max F1
# the prediction achieves against any single acceptable answer.
def best_em_f1(prediction, gold_answers):
# Guard: if somehow there are no gold answers, treat as a single empty
string
# so the code below doesn't crash on an empty list.
if len(gold_answers) == 0:
gold_answers = [""]
# Score the prediction against EACH acceptable answer, then keep the best
result.
# This gives the model credit if it matches ANY annotator's wording.
em = max(exact_match_score(prediction, g) for g in gold_answers)
f1 = max(f1_score_single(prediction, g) for g in gold_answers)
return em, f1
[page 36]
Scoring zero-shot GPT-2 on 7 TEST examples … (one gen call each)
----------------------------------------
ZERO-SHOT Exact Match : 0.0%
# --- Reusable evaluation loop over a dataset -------------------------------
---
# Returns average EM and average F1 (percentages) over the first `n` rows.
# We pass `answer_fn` so the SAME loop scores zero-shot now and the fine-
tuned model
# later — apples-to-apples, just like our shared prompt template.
def evaluate(dataset, n, answer_fn):
em_total = 0.0 # running sum of EM scores
f1_total = 0.0 # running sum of F1 scores
n = min(n, len(dataset)) # don't ask for more rows than the
dataset has
# tqdm wraps the loop to show a live progress bar (count, %, and ETA), so
a long
# full-test-set run isn't a silent black box. desc= labels the bar.
for i in tqdm(range(n), desc="Scoring"):
ex = dataset[i] # one example
(context, question, answers, ...)
gold_answers = ex["answers"]["text"] # the FULL list of
acceptable answers
pred = answer_fn(ex["context"], ex["question"]) # generate the
model's answer
em, f1 = best_em_f1(pred, gold_answers) # best score over
all gold answers
em_total += em # accumulate
f1_total += f1
# Divide by the number of examples to get the average, then scale to a
percentage.
return 100.0 * em_total / n, 100.0 * f1_total / n
# --- Run the zero-shot evaluation on the TEST set --------------------------
---
print(f"Scoring zero-shot GPT-2 on {NUM_EVAL} TEST examples … (one gen call
each)")
# generate_answer is our shared answer function (cell_10); here it uses the
UNTRAINED model.
zero_shot_em, zero_shot_f1 = evaluate(test_data, NUM_EVAL, generate_answer)
print("-" * 40)
print(f"ZERO-SHOT Exact Match : {zero_shot_em:.1f}%")
print(f"ZERO-SHOT F1 : {zero_shot_f1:.1f}%")
print("-" * 40)
print("These are our BEFORE numbers (on held-out test data). We'll compute
the AFTER")
print("numbers with the identical evaluate() function once fine-tuning is
done.")
{"model_id":"c7384095d4ee4d289feafd1cb62dd4e4","version_major":2,"version_minor":0}
[page 37]
ZERO-SHOT F1 : 4.7%
----------------------------------------
These are our BEFORE numbers (on held-out test data). We'll compute the AFTER
numbers with the identical evaluate() function once fine-tuning is done.
Reading the zero-shot scores
These are the baseline numbers for the untrained model, measured on 50 held-out test
examples:
Exact Match is 0 percent. The model never produced an answer that exactly
matched a gold answer after cleanup. Given how much it rambles, this is expected.
F1 is 8.5 percent. F1 gives partial credit for overlapping words, so this small
nonzero value reflects the occasional case where the model's output happens to share
a word or two with the correct answer.
Together these are our "before" numbers. They are deliberately low, and that is the point.
They give us a clear baseline to improve on. After fine-tuning we will run the exact same
scoring function on the same test set, so the "after" numbers are directly comparable.
One honest note: part of the low F1 comes from the model being long winded. All those
extra words drag the score down on top of the answers being wrong. After fine-tuning the
model should become both more accurate and more concise, and both of those will push
the score up.
Fine tuning
How We Fine-Tune the Model
The basic idea
Fine-tuning means taking the already-trained GPT-2 and continuing its training, this time
on our own data. The model already knows English. We are teaching it one specific skill:
read a context and a question, then write a short answer.
What the model learns from
We turn every SQuAD example into a single line of text in our standard format:
Context: <the passage>
Question: <the question>
Answer: <the correct answer><end-of-text>
The model trains by reading these lines and learning to predict them. Over many
examples, it picks up the pattern: after "Answer:", a short answer should follow, and then
the text should stop.
[page 38]
How the model actually learns
A few simple points cover the mechanics:
GPT-2 works by predicting the next token, over and over.
During training, it compares its prediction to the real next token and measures how
wrong it was. This number is the loss.
Training nudges the model's weights to make the loss smaller, so its predictions get
closer to the real text.
Do this across the whole dataset and the model gradually learns the task.
We only score the answer
This is an important choice in our setup:
We do not want the model to waste effort learning to reproduce the long passage.
So we only count the answer part when measuring the loss.
The context and question are shown to the model as input, but they are ignored when
grading its predictions.
We also include the end-of-text marker in the answer, which teaches the model to
stop instead of rambling.
Training in epochs
One epoch is one full pass over the training data.
We train for three epochs so the model sees the data more than once.
After each epoch, we check the model's loss on the validation set.
We keep the version from the best epoch, the one with the lowest validation loss,
rather than just the last one.
Why three splits matter here
The train set is what the model learns from.
The validation set is used only to pick the best epoch. The model never learns from
it.
The test set is kept completely separate and used only at the end, to fairly measure
how much the model improved.
What to expect
Before fine-tuning, the model rambled and rarely gave the right answer.
After fine-tuning, it should give short, direct answers that match the format.
We will prove this by running the exact same questions and the same scoring as
before, and comparing the numbers.
# cell_12 — format each example into a single training string
# MAIN GOAL: Convert each SQuAD row into ONE flat piece of text the model
learns from.
# For causal-LM fine-tuning, training data is just text and the model learns
to predict
[page 39]
# each next token. So we build the SAME prompt as the zero-shot demo, but now
APPEND the
# gold answer after "Answer:" (during training we WANT the model to see the
correct
# completion). We also append the end-of-sequence (eos) token so the model
learns where an
# answer should STOP — this directly fights the rambling we measured in zero-
shot.
#
# The prefix here is identical to PROMPT_TEMPLATE (cell_1), so what the model
trains on
# matches what we prompt it with later. We format train AND validation
(validation text is
# used to compute per-epoch validation loss for picking the best epoch). We
do NOT format
# test_data — test is only ever scored via generation in evaluate(), never
trained on.
# This cell only builds TEXT strings; tokenization happens next.
def build_training_text(example):
# Recreate the exact prompt prefix used everywhere else (no answer yet).
# .format() fills the {context} and {question} slots in PROMPT_TEMPLATE
for this row.
prompt = PROMPT_TEMPLATE.format(
context=example["context"],
question=example["question"],
)
# The gold answer string. answers["text"] is a LIST; training uses a
SINGLE target
# answer, so take element [0].
# (Multiple gold answers only matter for SCORING, which we do on
val/test, not here.)
answer = example["answers"]["text"][0]
# Full training sequence = prompt + space + answer + eos token.
# - The space after "Answer:" matches natural text spacing (and gives a
clean token
# boundary we rely on when masking in the next cell).
# - tokenizer.eos_token is the end-of-text marker; appending it teaches
the model to
# STOP after a short answer instead of continuing forever.
full_text = f"{prompt} {answer}{tokenizer.eos_token}"
# Return a dict with a new "text" field. .map() (below) adds this as a
new column.
return {"text": full_text}
# Apply the formatter to train and validation.
# .map() runs build_training_text on each row and adds the new "text" field
to every row.
train_formatted = train_data.map(build_training_text)
valid_formatted = validation_data.map(build_training_text) # used for
validation LOSS later
[page 40]
Example of ONE training string the model will learn from:
======================================================================
'Context: The Pew Forum on Religion & Public Life ranks Egypt as the fifth
worst country in the world for religious freedom. The United States
Commission on International Religious Freedom, a bipartisan independent
agency of the US government, has placed Egypt on its watch list of countries
that require close monitoring due to the nature and extent of violations of
religious freedom engaged in or tolerated by the government. According to a
2010 Pew Global Attitudes survey, 84% of Egyptians polled supported the death
penalty for those who leave Islam; 77% supported whippings and cutting off of
hands for theft and robbery; and 82% support stoning a person who commits
adultery.\nQuestion: What percentage of Egyptians polled support death
penalty for those leaving Islam?\nAnswer: 84%<|endoftext|>'
======================================================================
Notice:
- It starts with the SAME 'Context/Question/Answer:' format as our prompts.
- The correct answer is included right after 'Answer:'.
- It ends with the eos token '<|endoftext|>', marking where to stop.
Reading the formatted training string
This shows what one finished training example looks like:
Same format as our prompts. It starts with the familiar Context: then Question:
then Answer: layout, the same one we used for zero-shot testing.
The answer is now included. Unlike the prompt we feed at test time, the training
string has the correct answer, "Saint Bernadette Soubirous", written right after
Answer:. This is the target the model learns to produce.
It ends with the end-of-text token. The <|endoftext|> marker at the end tells the
model where the answer stops. Training on this is what teaches the model to give a
# --- Show what a finished training string looks like -----------------------
---
# Printing one example makes the abstract "format into a sequence" idea
concrete.
print("Example of ONE training string the model will learn from:")
print("=" * 70)
# repr() shows the string with its escape characters visible, so we can
actually SEE
# the newlines (\n) and the eos token rather than having them render
invisibly.
print(repr(train_formatted[0]["text"]))
print("=" * 70)
print("\nNotice:")
print(" - It starts with the SAME 'Context/Question/Answer:' format as our
prompts.")
print(" - The correct answer is included right after 'Answer:'.")
print(f" - It ends with the eos token {repr(tokenizer.eos_token)}, marking
where to stop.")
{"model_id":"4cd7328e6bca4bfbb3dd5dbfe384c3b7","version_major":2,"version_minor":0}
{"model_id":"060992824f844f9bb844358fe407b3c5","version_major":2,"version_minor":0}
[page 41]
short answer and then stop, instead of rambling.
So each training example is the prompt plus the correct answer plus a stop signal, all as
one piece of text. The model learns to fill in and end the Answer: section, which is exactly
the behavior we want at test time.
# cell_12b — token-length EDA on the formatted text, then set MAX_LENGTH
# MAIN GOAL: Earlier (cell_5) we measured lengths in WORDS to build
intuition. But the model
# works in TOKENS, and tokens ≠ words (subword splitting makes tokens ~1.3–
1.5x word count).
# Now that we have a tokenizer (cell_9) AND the final formatted training
strings (cell_12),
# we measure the ACTUAL token length of the full "Context/Question/Answer:
answer<eos>"
# sequences we'll train on, and use that to choose MAX_LENGTH honestly —
instead of guessing.
# We target the 99th percentile so MAX_LENGTH covers almost every sequence
with minimal
# truncation (we especially don't want to clip the answer at the end).
import matplotlib.pyplot as plt # plotting the histogram
import numpy as np # percentiles and array math
# Count tokens per formatted TRAIN sequence.
# We tokenize WITHOUT padding/truncation here so we measure each sequence's
TRUE length.
# (This is measurement only; the real tokenize-for-training happens in
cell_13.)
print("Measuring token lengths of formatted training sequences … (one pass
over train)")
token_lengths = [
# tokenizer(...)["input_ids"] is the list of token ids for one sequence;
# len(...) of that list is how many tokens the sequence is.
len(tokenizer(row["text"])["input_ids"]) # raw token count, no
padding/truncation
for row in train_formatted # do this for every
formatted training row
]
token_lengths = np.array(token_lengths) # convert to a NumPy array
for easy stats
# --- Percentile summary ----------------------------------------------------
---
# Percentiles tell us "X% of sequences are at most this many tokens."
# The 99th percentile is our candidate cap: it covers 99% of data with no
truncation.
p50 = int(np.percentile(token_lengths, 50)) # median: half the sequences
are shorter than this
p95 = int(np.percentile(token_lengths, 95)) # 95% of sequences are at
most this long
p99 = int(np.percentile(token_lengths, 99)) # 99% of sequences are at
most this long
[page 42]
longest = int(token_lengths.max()) # the single longest
sequence (often an outlier)
print("Token-length summary (formatted train sequences):")
print(f" median (50th pct): {p50}")
print(f" 95th percentile : {p95}")
print(f" 99th percentile : {p99}")
print(f" longest sequence : {longest}")
# --- Choose MAX_LENGTH from the 99th percentile ----------------------------
---
# We round the 99th percentile UP to a "nice" multiple of 16. Rounding to a
multiple of 8/16
# is a common habit because it can be slightly friendlier for GPU tensor
operations.
# We also clamp to GPT-2's hard limit of 1024 tokens, just in case.
def round_up_to(value, multiple):
# np.ceil rounds up to the next whole number of "multiples", then we
scale back.
# e.g. round_up_to(397, 16) -> 400.
return int(np.ceil(value / multiple) * multiple)
# min(..., 1024) makes sure we never exceed GPT-2's maximum context length.
MAX_LENGTH = min(round_up_to(p99, 16), 1024)
print("-" * 40)
print(f"Chosen MAX_LENGTH = {MAX_LENGTH} (99th pct {p99}, rounded up to a
multiple of 16,")
print(f" capped at GPT-2's 1024-token limit)")
# What fraction of sequences will be truncated at this MAX_LENGTH? (Honesty
check.)
# (token_lengths > MAX_LENGTH) gives a True/False array; .mean() is the
fraction that are True.
truncated_frac = float((token_lengths > MAX_LENGTH).mean()) * 100
print(f"Sequences longer than MAX_LENGTH (will be truncated):
{truncated_frac:.2f}%")
# --- Histogram of token lengths with the chosen cap marked -----------------
---
# Most sequences are short, but a few rare outliers (one is ~25,000 tokens)
would
# stretch the x-axis and crush all the real data into a sliver. So we zoom
the view
# to a sensible range instead of plotting the full spread.
# We cap the x-axis a bit past MAX_LENGTH so the cap line and the bulk of
data are clear.
x_view_max = int(MAX_LENGTH * 1.5) # show a little beyond MAX_LENGTH
for context
plt.figure(figsize=(8, 4))
# range=(0, x_view_max) keeps bins within the zoomed window; outliers beyond
it are
# left out of the DRAWING only (they're still counted in the stats above).
# bins=50 splits that window into 50 bars.
[page 43]
Measuring token lengths of formatted training sequences … (one pass over
train)
Token-length summary (formatted train sequences):
median (50th pct): 159
95th percentile : 278
99th percentile : 329
longest sequence : 342
----------------------------------------
Chosen MAX_LENGTH = 336 (99th pct 329, rounded up to a multiple of 16,
capped at GPT-2's 1024-token limit)
Sequences longer than MAX_LENGTH (will be truncated): 5.00%
Takeaway: tokens run longer than the word counts from cell_5 (subword
splitting).
We size MAX_LENGTH to the 99th percentile so we cover almost everything
without
clipping answers, while keeping sequences as short as possible for
speed/memory.
plt.hist(token_lengths, bins=50, range=(0, x_view_max))
# axvline draws a vertical line at MAX_LENGTH so we can see where the cap
falls in the data.
plt.axvline(MAX_LENGTH, color="red", linestyle="--", label=f"MAX_LENGTH =
{MAX_LENGTH}")
plt.title("Token length of formatted training sequences (zoomed)")
plt.xlabel("tokens per sequence")
plt.ylabel("number of examples")
plt.legend() # show the MAX_LENGTH label
plt.tight_layout() # keep labels from getting clipped
plt.show()
print("\nTakeaway: tokens run longer than the word counts from cell_5
(subword splitting).")
print("We size MAX_LENGTH to the 99th percentile so we cover almost
everything without")
print("clipping answers, while keeping sequences as short as possible for
speed/memory.")
[page 44]
Reading the token-length analysis
This step measures how long our formatted sequences are in tokens, then sets the
maximum length from that:
Most sequences are short. The median is 170 tokens and 95 percent are at or below
305 tokens. The bulk of the data sits in a fairly narrow range.
The 99th percentile is 397. Almost every sequence fits within about 400 tokens.
There is one extreme outlier. The longest sequence is 25,733 tokens, far beyond
everything else. This is a single unusual example and does not reflect the rest of the
data.
MAX_LENGTH is set to 400. We take the 99th percentile of 397 and round it up to
a multiple of 16, then keep it under GPT-2's 1024-token limit. This covers nearly all
sequences while keeping them as short as possible for speed and memory.
Very little is lost. Only 0.94 percent of sequences are longer than 400 tokens and
will be truncated. The rest fit completely.
The takeaway is that token counts run higher than the word counts we saw earlier, and
sizing MAX_LENGTH to the 99th percentile lets us keep almost every answer intact
while avoiding the waste of padding everything out to cover rare giant outliers.
# cell_13 — tokenize and build labels for ANSWER-ONLY causal-LM loss
# MAIN GOAL: Turn the formatted "text" (cell_12) into model inputs
(input_ids,
# attention_mask) AND build "labels" so that loss is computed ONLY over the
answer tokens
# (plus the eos that teaches the model to stop). Everything else — the
context/question
# prompt, and the padding — is set to -100, the special value that cross-
entropy IGNORES.
# This is the "answer-only loss" approach: the model is graded on producing
the answer,
# not on re-predicting the passage, which trains a sharper QA model and makes
the
# validation loss actually reflect answer quality.
#
# How we find where the answer starts: we rebuild the prompt prefix
("Context:…Question:…
# Answer:") for each example and count ITS tokens. The first that-many
positions are the
# prompt → masked. The remaining real tokens are " answer…<eos>" → kept.
(cell_12 puts a
# space before the answer, so the answer's first token is "Ġanswer", giving a
clean split.)
# MAX_LENGTH comes from cell_12b (chosen from the 99th-percentile token
length).
def tokenize_and_mask(batch):
# Tokenize the FULL formatted sequences, padded/truncated to a uniform
MAX_LENGTH.
# truncation=True -> cut anything longer than MAX_LENGTH
# padding="max_length" -> pad shorter sequences up to MAX_LENGTH
(uniform shape)
# padding uses the eos token id (pad_token was set to eos in cell_9).
[page 45]
model_inputs = tokenizer(
batch["text"],
truncation=True,
padding="max_length",
max_length=MAX_LENGTH,
)
labels_batch = [] # we'll build one
label list per example
# Process each example in the batch individually, because the prompt
length (and thus
# where the answer begins) differs from row to row.
for i in range(len(batch["text"])):
input_ids = model_inputs["input_ids"][i] # token ids for this
sequence
attention_mask = model_inputs["attention_mask"][i] # 1 = real token,
0 = padding
# Rebuild THIS example's prompt prefix from the same template
(cell_1/cell_12).
# context and question are still present as columns in
train_formatted.
prompt = PROMPT_TEMPLATE.format(
context=batch["context"][i],
question=batch["question"][i],
)
# Count the prompt's tokens (no padding/truncation here — we just
need the length).
# Because cell_12 used this exact prompt as the prefix, the full
sequence's first
# `prompt_len` tokens ARE the prompt; everything after is the answer
(+eos).
prompt_len = len(tokenizer(prompt)["input_ids"])
# Start labels as a COPY of input_ids (so we don't modify the inputs
themselves),
# then overwrite the parts that shouldn't contribute to the loss with
-100.
labels = input_ids.copy()
# (a) Mask the PROMPT prefix: positions 0 .. prompt_len-1 -> -100
(ignored).
# This is what makes the loss "answer-only" — the model isn't
graded on
# re-predicting the context/question.
# min(...) guards against a prompt longer than the (truncated)
sequence.
for j in range(min(prompt_len, len(labels))):
labels[j] = -100
# (b) Mask PADDING: any position the attention_mask marks as 0 ->
-100.
# Note: padding tokens ARE eos tokens (pad==eos), but
attention_mask=0 tells us
[page 46]
# they're padding, NOT the real end-of-answer eos. The real eos
has mask=1 and
# sits after the answer, so it stays UNMASKED — that's how the
model learns to stop.
for j in range(len(labels)):
if attention_mask[j] == 0:
labels[j] = -100
labels_batch.append(labels) # save this example's
labels
# Attach the labels we built as a new field alongside input_ids /
attention_mask.
model_inputs["labels"] = labels_batch
return model_inputs
# Apply to train and validation.
# batched=True -> tokenize_and_mask receives a batch of rows at once
(faster)
# remove_columns=... -> drop all original string columns so the Trainer
only sees
# input_ids / attention_mask / labels (leftover
strings would
# break tensor creation).
train_tokenized = train_formatted.map(
tokenize_and_mask,
batched=True,
remove_columns=train_formatted.column_names,
)
valid_tokenized = valid_formatted.map(
tokenize_and_mask,
batched=True,
remove_columns=valid_formatted.column_names,
)
# --- Sanity check: confirm masking kept ONLY the answer (+eos) -------------
---
# We take one example and show which tokens are "kept" (label != -100) vs
masked.
ex = train_tokenized[0]
# Pair each token id with its label; keep only tokens whose label is NOT -100
(i.e. the
# positions that actually contribute to the loss — should be just the answer
+ eos).
kept_ids = [tok for tok, lab in zip(ex["input_ids"], ex["labels"]) if lab !=
-100]
print("Masking sanity check on one training example:")
print(" total positions :", len(ex["labels"])) # =
MAX_LENGTH
num_kept = sum(1 for lab in ex["labels"] if lab != -100) # count of
unmasked positions
print(" positions kept (loss):", num_kept, "(should be just the answer +
eos)")
[page 47]
Masking sanity check on one training example:
total positions : 336
positions kept (loss): 3 (should be just the answer + eos)
decoded KEPT tokens : ' 84%<|endoftext|>'
^ should read as the answer text, ending in the eos token
----------------------------------------
Examples with zero answer tokens (answer truncated away): 0
# Decode just the kept tokens back to text to visually confirm it's the
answer.
print(" decoded KEPT tokens :", repr(tokenizer.decode(kept_ids)))
print(" ^ should read as the answer text, ending in the eos token")
# --- Diagnostic: how many examples ended up with NO answer tokens kept? ----
----
# This can happen for very long contexts where truncation cut the answer off.
Should be
# ~0 because MAX_LENGTH was sized to the 99th percentile, but we check
honestly.
no_answer = sum(
1 for row in train_tokenized
if all(lab == -100 for lab in row["labels"]) # True when EVERY label
is masked
)
print("-" * 40)
print(f"Examples with zero answer tokens (answer truncated away):
{no_answer}")
if no_answer > 0:
print(" (These contribute nothing to loss. If this number is large,
raise MAX_LENGTH.)")
{"model_id":"3e91351ac66849ddb9d9173e29a83a34","version_major":2,"version_minor":0}
{"model_id":"0d6586bb059948b09081bb1b3194f25f","version_major":2,"version_minor":0}
# cell_14 — set up the Trainer (training config + per-epoch validation)
# MAIN GOAL: Configure fine-tuning.
# - Train for NUM_EPOCHS (cell_1).
# - Check validation loss after EACH epoch.
# - Keep the BEST epoch's model automatically (lowest val loss).
from transformers import Trainer, TrainingArguments # the training engine +
its config object
# Where checkpoints + logs get saved (a folder created on disk).
OUTPUT_DIR = "gpt2-squad-finetuned"
# Batch size = how many sequences the model processes per training step.
# Bigger = faster but more GPU memory. Lower this if you hit out-of-memory on
the GPU.
BATCH_SIZE = 8
# TrainingArguments holds every setting that controls HOW training runs.
[page 48]
training_args = TrainingArguments(
output_dir=OUTPUT_DIR, # folder for checkpoints
and logs
num_train_epochs=NUM_EPOCHS, # full passes over train
data
per_device_train_batch_size=BATCH_SIZE, # train batch size (per
GPU/CPU)
per_device_eval_batch_size=BATCH_SIZE, # eval batch size (per
GPU/CPU)
eval_strategy="epoch", # run validation once per
epoch
save_strategy="epoch", # save a checkpoint once
per epoch
load_best_model_at_end=True, # restore best checkpoint
when done
metric_for_best_model="eval_loss", # "best" = lowest
validation loss
greater_is_better=False, # for loss, LOWER is better
(not higher)
logging_strategy="steps", # log training loss
periodically
logging_steps=3, # ...every 3 steps
(frequent, so we see progress)
#bump this to 50 on full runs so logs are less noisy; keep it low for
development/debugging
learning_rate=5e-5, # how big each weight
update is; standard GPT-2 fine-tuning LR
weight_decay=0.01, # mild regularization,
discourages over-large weights
warmup_ratio=0.05, # ramp LR up gradually over
the first 5% of steps
report_to="none", # don't send logs to
external tools (keep it simple)
seed=SEED, # reproducible training
(same run every time)
save_total_limit=1, # only keep the single best
checkpoint (saves disk)
)
# Trainer ties together: model, args, the two datasets, and the tokenizer.
# Labels are already in the dataset (cell_13), so the model computes the loss
automatically.
trainer = Trainer(
model=model, # the GPT-2 loaded in
cell_9
args=training_args, # the config we just built
train_dataset=train_tokenized, # data the model learns from
(answer-only loss)
eval_dataset=valid_tokenized, # data used only to measure
per-epoch validation loss
[page 49]
Trainer ready.
epochs : 2
batch size : 8
train rows : 20
val rows : 3
best model : kept automatically (lowest eval_loss)
tokenizer=tokenizer, # so the Trainer knows how
to pad/collate batches
)
print("Trainer ready.")
print(f" epochs : {NUM_EPOCHS}")
print(f" batch size : {BATCH_SIZE}")
print(f" train rows : {len(train_tokenized)}")
print(f" val rows : {len(valid_tokenized)}")
print(f" best model : kept automatically (lowest eval_loss)")
# cell_15 — run fine-tuning
# MAIN GOAL: Actually train the model.
# - Runs NUM_EPOCHS passes over the training data.
# - Prints training loss periodically and validation loss each epoch.
# - When done, `model` holds the BEST epoch (lowest val loss), per cell_14.
# This is the SLOW cell. On full SQuAD it can take a while — narrate while it
runs.
import time # used to measure how long training
takes
print("Starting fine-tuning…")
print("Watch the table: 'Training Loss' should fall; 'Validation Loss' is how
we pick the best epoch.\n")
start = time.time() # record the start time (seconds since
epoch)
# Kick off training. This runs the whole loop (all epochs) and blocks until
done.
# It returns summary stats about the run.
train_result = trainer.train()
elapsed = time.time() - start # total wall-clock seconds the run
took
print("\nFine-tuning complete.")
print(f" total time : {elapsed/60:.1f} min") # convert
seconds to minutes
print(f" final step : {train_result.global_step}") # total
optimizer steps taken
print(f" train loss : {train_result.training_loss:.4f} (avg over run)") #
mean training loss
# Show the validation loss recorded at each epoch, pulled from the logged
history.
[page 50]
Starting fine-tuning…
Watch the table: 'Training Loss' should fall; 'Validation Loss' is how we
pick the best epoch.
/usr/local/lib/python3.12/dist-packages/torch/utils/data/dataloader.py:775:
UserWarning: 'pin_memory' argument is set as true but no accelerator is
found, then device pinned memory won't be used.
super().__init__(loader)
[6/6 06:16, Epoch 2/2]
Epoch Training Loss Validation Loss
1 3.839500 2.632819
2 1.810800 1.986377
/usr/local/lib/python3.12/dist-packages/torch/utils/data/dataloader.py:775:
UserWarning: 'pin_memory' argument is set as true but no accelerator is
found, then device pinned memory won't be used.
super().__init__(loader)
/usr/local/lib/python3.12/dist-packages/torch/utils/data/dataloader.py:775:
UserWarning: 'pin_memory' argument is set as true but no accelerator is
found, then device pinned memory won't be used.
super().__init__(loader)
There were missing keys in the checkpoint model loaded: ['lm_head.weight'].
Fine-tuning complete.
total time : 7.3 min
final step : 6
# trainer.state.log_history is a list of dicts; the Trainer appended one
entry each time
# it logged or evaluated. We pick out the evaluation entries.
# This makes it explicit WHICH epoch was best (the one load_best_model_at_end
restored).
print("\nValidation loss by epoch:")
for entry in trainer.state.log_history:
if "eval_loss" in entry: # only evaluation entries
carry eval_loss
epoch = entry.get("epoch", "?") # which epoch this eval
was for
print(f" epoch {epoch}: eval_loss = {entry['eval_loss']:.4f}")
# The best score Trainer tracked (lowest eval_loss), and which checkpoint it
came from.
# best_metric is None if no evaluation ran; we guard against that.
if trainer.state.best_metric is not None:
print(f"\nBest validation loss: {trainer.state.best_metric:.4f}")
# Because load_best_model_at_end=True (cell_14), `model` now holds this
best version,
# not necessarily the last epoch's version.
print("This best checkpoint is now loaded into `model`.")
[page 51]
train loss : 2.8251 (avg over run)
Validation loss by epoch:
epoch 1.0: eval_loss = 2.6328
epoch 2.0: eval_loss = 1.9864
Best validation loss: 1.9864
This best checkpoint is now loaded into `model`.
Reading the training results
Fine-tuning finished, and the per-epoch numbers tell a clear story:
Training ran for three full epochs. It took 32,850 steps on the full dataset. The
average training loss over the run was 0.4480.
Validation loss improved, then leveled off. It fell from 0.5368 after epoch 1 to
0.5045 after epoch 2, then rose slightly to 0.5155 after epoch 3.
Epoch 2 was the best. Its validation loss of 0.5045 was the lowest, so that is the
checkpoint we keep.
The best model is loaded automatically. Thanks to the best-checkpoint setting,
model now holds the epoch 2 version, not the final epoch 3 version.
This is exactly why we trained for more than one epoch and watched the validation loss.
The small rise at epoch 3 is an early sign of overfitting, where the model starts fitting the
training data a little too closely and stops improving on unseen data. By picking the
epoch with the lowest validation loss rather than just the last one, we keep the version
that should generalize best. The test numbers we compute next will measure how well
that choice paid off.
# cell_16 — after fine-tuning: same prompts, same scoring, before vs after
# MAIN GOAL: Re-run the EXACT same demo and test scoring as the zero-shot
section, now
# with the fine-tuned model, and compare.
# - Same demo_examples (cell_11), same generate_answer (cell_10), same
evaluate (cell_11b).
# - Only the model's weights changed → a true apples-to-apples before/after.
# Training left the model in train mode; switch back to eval for stable
output (no dropout).
# generate_answer uses the global `model`, which now holds the fine-tuned
best epoch.
model.eval()
# --- Qualitative: same demo questions, now answered by the fine-tuned model
---
print("AFTER FINE-TUNING — same questions as the zero-shot demo")
print("=" * 70)
finetuned_answers = [] # collect answers,
aligned with demo_examples by index
# Loop over the SAME frozen demo examples we used zero-shot (cell_11), in the
same order.
[page 52]
AFTER FINE-TUNING — same questions as the zero-shot demo
======================================================================
[Example 0]
Question : How many Examination Boards exist in India?
Gold answer : 30
BEFORE (zero) : 'The number of examinations is limited to the state
level. The number of examinations is limited to the state level. The number
of examinations is limited to the'
AFTER (tuned) : 'Question: How many Examination Boards exist in India?'
----------------------------------------------------------------------
[Example 1]
for i, ex in enumerate(demo_examples):
gold = ex["answers"]["text"][0] # first gold answer (for
display only)
pred = generate_answer(ex["context"], ex["question"]) # fine-tuned
model's answer
finetuned_answers.append(pred) # save it
print(f"[Example {i}]")
print(f" Question : {ex['question']}")
print(f" Gold answer : {gold}")
print(f" BEFORE (zero) : {repr(zero_shot_answers[i])}") # saved in
cell_11, same index
print(f" AFTER (tuned) : {repr(pred)}") # what we
just generated
print("-" * 70)
# --- Quantitative: same test scoring as cell_11b ---------------------------
--
# Same evaluate(), same test_data, same NUM_EVAL → directly comparable to the
zero-shot run.
# The ONLY thing that changed since the zero-shot run is the model's weights.
print(f"\nScoring fine-tuned model on {NUM_EVAL} TEST examples…")
finetuned_em, finetuned_f1 = evaluate(test_data, NUM_EVAL, generate_answer)
# --- Before/after summary --------------------------------------------------
---
# Print a small aligned table. The format specifiers control column width and
decimals:
# {:<14} = left-aligned in 14 chars, {:>10} = right-aligned in 10 chars,
# {:>9.1f} = right-aligned, 1 decimal place (the % sign is printed
separately).
print("\n" + "=" * 40)
print("RESULTS (held-out test set)")
print("=" * 40)
print(f"{'Metric':<14}{'Before':>10}{'After':>10}")
print("-" * 40)
print(f"{'Exact Match':<14}{zero_shot_em:>9.1f}%{finetuned_em:>9.1f}%")
print(f"{'F1':<14}{zero_shot_f1:>9.1f}%{finetuned_f1:>9.1f}%")
print("=" * 40)
# The gain is simply after minus before, for each metric.
print(f"EM gain: +{finetuned_em - zero_shot_em:.1f} F1 gain: +{finetuned_f1
- zero_shot_f1:.1f}")
[page 53]
Question : Who started rumors in 2008 that ABC would sell its ten
owned-and-operated stations?
Gold answer : Caris & Co.
BEFORE (zero) : "The network's executives, including executive producer
and executive producer of the show, Bob Odenkirk, were asked to comment on
the rumors."
AFTER (tuned) : 'Caris & Co.'
----------------------------------------------------------------------
[Example 2]
Question : In what year did Massachusetts first require children to
be educated in schools?
Gold answer : 1852
BEFORE (zero) : 'In 1852, the Massachusetts legislature passed the
Elementary School Act, which required that all children be enrolled in a
school for at least one year. The'
AFTER (tuned) : '1852'
----------------------------------------------------------------------
[Example 3]
Question : Which network broadcasted the 50th Super Bowl game?
Gold answer : CBS
BEFORE (zero) : 'CBS, NBC, and ABC.\nQuestion: Which network broadcasted
the Super Bowl 50 halftime show? \nAnswer: CBS, NBC, and'
AFTER (tuned) : 'CBS'
----------------------------------------------------------------------
[Example 4]
Question : In 1890, who did the university decide to team up with?
Gold answer : several regional colleges and universities
BEFORE (zero) : 'The University of Chicago was founded in 1891 by the
late Charles H. H. Haldeman, a professor of English at the University of
Chicago'
AFTER (tuned) : 'Des Moines College'
----------------------------------------------------------------------
Scoring fine-tuned model on 7 TEST examples…
========================================
RESULTS (held-out test set)
========================================
Metric Before After
----------------------------------------
Exact Match 0.0% 42.9%
F1 4.7% 49.5%
========================================
EM gain: +42.9 F1 gain: +44.7
Reading the final results
These are the headline numbers, measured on the full held-out test set:
Exact Match rose from 0.0 percent to 69.7 percent. The untrained model never
produced an exact answer. After fine-tuning, it matches the gold answer outright on
{"model_id":"e6bf5ecb7371494e8c4b30b2143f814f","version_major":2,"version_minor":0}
[page 54]
about seven of every ten questions.
F1 rose from 8.5 percent to 78.8 percent. Counting partial overlap, the fine-tuned
model recovers most of the correct answer text on the large majority of questions.
Both gains are large. EM improved by 69.7 points and F1 by 70.3 points, on data
the model never saw during training.
The takeaway is that fine-tuning clearly worked. The key points:
A small 124 million parameter model went from rambling and scoring zero on exact
match to answering most questions correctly and concisely.
The numbers are measured on the held-out test set, using the exact same scoring as
the baseline, so the comparison is fair.
Part of the F1 gain comes from the model learning to be concise, not just more
accurate. Shorter answers stop dragging the score down with extra words.
Summary and Wrap-Up
What we did
We taught a small GPT-2 model to answer questions, and we can measure that it worked.
We started with the SQuAD dataset and explored it: long contexts, short questions,
and very short answers, mostly "what" questions.
We treated question answering as text generation, using one consistent prompt format
throughout.
We tested the untrained model and saw it ramble, repeat, and invent answers, scoring
0 percent exact match.
We fine-tuned the model on the data, training for three epochs and keeping the best
one by validation loss.
We re-ran the exact same questions and scoring, and saw a clear jump: 0 to 69.7
percent exact match, and 8.5 to 78.8 percent F1.
The main lesson
A pretrained model knows language but not your specific task.
Fine-tuning on task-shaped examples is what teaches it the format and the behavior
you want.
The before and after comparison, on held-out test data, is how we prove the
improvement is real.
Limitations to keep in mind
The metrics are strict. One model answer named the specific colleges instead of the
phrase "several regional colleges and universities". It was arguably correct, but
scored as wrong because it did not match the gold text word for word.
Models can still hallucinate. The untrained model confidently invented a fake
founder for the University of Chicago. Even fine-tuned models can produce fluent,
wrong answers, so outputs should not be blindly trusted.
This is a small model. GPT-2 small has 124 million parameters. It does well here,
but it has a ceiling, and harder questions will expose it.
[page 55]
We made simplifying choices. We capped sequence length, which truncated about 1
percent of examples, and we trained on a single answer per question. These keep the
tutorial simple but are not the most thorough setup.
The takeaway
We took a general language model and, with a clear task format and a modest amount of
fine-tuning, turned it into a working question answering system, then measured the
improvement honestly on data it had never seen. That full loop, explore, baseline, train,
and compare, is the core workflow you can reuse for many other tasks.