# samsum summarize v5 GC PDF

course: Module 4 — Generative AI & LLMs
module: Module-4-Generative-AI-LLMs
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-4-Generative-AI-LLMs/General/07_06_26_-_Materials/samsum_summarize_v5_GC_PDF.pdf
pages: 47

---
[page 1]
About the dataset: SAMSum
Before we load anything, it helps to know what we are actually working with. The whole
tutorial depends on this dataset, so a few minutes here makes everything later easier to
follow.
What it is
SAMSum is a collection of about 16,000 short conversations that look like text-message
or messenger chats. Each conversation comes paired with a short summary written by a
person. Our task for the whole tutorial is simple to state: given the conversation, produce
a summary close to the human one.
Here is the shape of a single example:
id: a unique identifier for the conversation. We do not use it for modeling, it is just a
label.
dialogue: the conversation itself, written as a series of lines like "Amanda: I baked
cookies. Do you want some?". This is the INPUT we feed the model.
summary: a short, third-person description of what the conversation was about. This
is the TARGET we want the model to learn to produce.
Where it came from and why that matters
SAMSum was created by a research group at Samsung and released alongside a 2019
academic paper. The conversations were not scraped from real private chats. Instead,
linguists who are fluent in English wrote them by hand, modeled on the kinds of
messages they send in everyday life. Because of this, the conversations cover a wide mix
of tones: casual, semi-formal, and formal, and they include things real chats contain, such
as slang, typos, and emoticons.
This origin matters for two reasons. First, the data is clean and consistent, because trained
people wrote it on purpose, which is part of why it is friendly for a first fine-tuning
project. Second, the summaries follow a clear convention: they are short, and they
describe the conversation in the third person ("Amanda baked cookies and will bring
some to Jerry"), rather than copying lines out of the chat. That convention is exactly the
behavior we will teach the model.
How it is structured
The conversations are spread fairly evenly across four groups based on how many back-
and-forth turns they contain: short ones with 3 to 6 turns, then 7 to 12, then 13 to 18, and
longer ones with 19 to 30. Most conversations, around three quarters of them, are
between just two people, while the rest involve three or more. This variety is useful,
because it means the model has to learn to summarize both quick exchanges and longer
group chats.

[page 2]
The dataset comes already split into three parts, and we keep these splits separate for
honest evaluation:
train (about 14,700 examples): what the model learns from.
validation (about 800 examples): checked during training to watch for overfitting.
test (about 800 examples): held back and only used at the end to measure how well
the model really does.
Why we chose it for this tutorial
SAMSum hits a sweet spot for learning. It is small enough to fine-tune quickly, even on
modest hardware, so you are not waiting hours to see a result. It is clean, so you spend
your time learning the method rather than fighting messy data. And it is intuitive: anyone
can read a chat and its summary and judge for themselves whether the model did a good
job, without needing any special domain knowledge. That last point is surprisingly
valuable, because it lets you trust your own eyes alongside the metrics.
One honest note on usage
SAMSum is released for research and learning under a non-commercial license. That is
perfect for a tutorial like this one, but it is worth knowing that you cannot simply drop
this dataset into a commercial product without checking the license terms first. Being
aware of the license attached to your data is a good habit to build early.
Enter physical GPU index (e.g. 0, 1, 2) or 'cpu': cpu
Device set to: cpu
# cell_0a — GPU mask setup before any torch import
# MAIN GOAL: Let the user pick which physical GPU to use (or CPU),
# then hide all other GPUs from this process via CUDA_VISIBLE_DEVICES.
# Must run before importing torch or any CUDA-using library.
import os
choice = input("Enter physical GPU index (e.g. 0, 1, 2) or 'cpu': 
").strip().lower()
if choice == "cpu":
    os.environ["CUDA_VISIBLE_DEVICES"] = ""
    DEVICE = "cpu"
elif choice.isdigit():
    os.environ["CUDA_VISIBLE_DEVICES"] = choice      # hide every other GPU
    DEVICE = "cuda:0"                                # selected GPU maps to 
cuda:0 internally
else:
    print("Invalid input. Defaulting to CPU.")
    os.environ["CUDA_VISIBLE_DEVICES"] = ""
    DEVICE = "cpu"
print(f"Device set to: {DEVICE}")

[page 3]
Installing Libraries
  Preparing metadata (setup.py) ... ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 
84.1/84.1 kB 2.8 MB/s eta 0:00:00
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 61.1/61.1 kB 4.0 MB/s eta 0:00:00
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 71.3/71.3 kB 3.4 MB/s eta 0:00:00
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 495.3/495.3 kB 11.7 MB/s eta 0:00:00
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 100.6/100.6 kB 5.1 MB/s eta 0:00:00
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 51.5/51.5 kB 2.8 MB/s eta 0:00:00
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 144.3/144.3 kB 10.2 MB/s eta 0:00:00
# cell_0b — Install all required libraries
# MAIN GOAL: Install every Python package this tutorial needs, in one place,
# BEFORE we import anything. Running installs first means the import cell
# (next) won't fail with "ModuleNotFoundError". The "-q" flag keeps the
# output quiet so the notebook stays readable.
# NOTE: In a fresh environment you only need to run this cell ONCE.
# If you restart the kernel later, the installed packages persist on disk,
# so you can usually skip re-running this.
!pip install -q \
    transformers \
    datasets \
    evaluate \
    rouge-score \
    bert-score \
    sentencepiece \
    accelerate \
    py7zr
# --- Why each package is here ---
# transformers  : the HuggingFace library that gives us the T5 model + 
tokenizer
# datasets      : one-line loading of the SAMSum dataset from the HuggingFace 
Hub
# evaluate      : modern wrapper that runs metrics (we use it to compute 
ROUGE)
# rouge-score   : the actual ROUGE implementation that "evaluate" calls under 
the hood
# bert-score    : our secondary, semantic-similarity metric
# sentencepiece : REQUIRED by the T5 tokenizer — T5 fails to load without it
# accelerate    : REQUIRED by the HuggingFace Trainer we use in Block 3
# py7zr         : SAMSum ships as a .7z archive; this lets "datasets" unpack 
it
print("All libraries installed.")
# cell_0c — Imports and device confirmation

[page 4]
PyTorch version : 2.11.0+cpu
Using device    : cpu
Importing Data
A note on how we are loading the data
In earlier work, the usual pattern was to have a CSV file sitting on your own computer
and read it into the notebook, often with something like pd.read_csv("myfile.csv").
The file lived locally, you pointed pandas at its path, and you got a table back.
# MAIN GOAL: Bring in every library we'll use across the whole tutorial, in 
one
# place, AND confirm that the device we picked back in cell_0a is really 
usable.
# This must run AFTER the GPU mask (cell_0a) and AFTER the installs 
(cell_0b).
import torch                       # core deep-learning library; powers the 
model + training
from datasets import load_dataset  # one call to pull SAMSum from the 
HuggingFace Hub
from transformers import (
    AutoTokenizer,                 # turns text into token IDs the model 
understands
    AutoModelForSeq2SeqLM,         # loads T5 in its text-in / text-out 
("seq2seq") form
)
import evaluate                    # runs our metrics (ROUGE now, more later)
# --- Reconcile our earlier device choice with reality ---
# cell_0a set DEVICE based ONLY on what you typed, before torch existed.
# Now we can actually ask torch whether a GPU is visible. If we asked for
# "cuda:0" but no GPU is available, we fall back to CPU so nothing crashes.
if DEVICE.startswith("cuda") and not torch.cuda.is_available():
    print("Warning: a GPU was requested, but torch sees none. Falling back to 
CPU.")
    DEVICE = "cpu"
# --- Report what we ended up with ---
print(f"PyTorch version : {torch.__version__}")
print(f"Using device    : {DEVICE}")
# If we're on GPU, print its name so you can confirm it's the card you 
expected.
# (Remember: because of the mask in cell_0a, your chosen GPU now shows as 
index 0.)
if DEVICE.startswith("cuda"):
    print(f"GPU name        : {torch.cuda.get_device_name(0)}")

[page 5]
This time we do something different. We call load_dataset("knkarthick/samsum"), and
the data is pulled from the Hugging Face Hub, which is an online repository of datasets,
rather than from a file on your machine. Here is what actually changes, and why it is
worth understanding.
Where the data lives
Before (local CSV): the data was a file you already had on disk. Nothing was
downloaded. If you moved to another computer, you had to bring the file with you.
Now (Hugging Face Hub): the data lives online. The first time you run the load, it is
downloaded and then cached on your machine, so later runs are fast and do not re-
download. Anyone running this notebook gets the exact same data without you
having to share a file.
What you get back
Before: pd.read_csv gives you a pandas DataFrame, a single table.
Now: load_dataset gives you a DatasetDict, which is a container holding multiple
splits at once. In our case it already contains the train, validation, and test sets as
separate pieces. With a local CSV you would normally have to split the data into
train and test yourself, and it is easy to do that inconsistently. Here the split is fixed
and standard, so everyone is training and testing on the same partition.
Why this approach is used here
Reproducibility: because the data comes from one named source, every person who
runs this notebook gets identical data. A local CSV can quietly differ from machine
to machine (someone edited a row, a different version got saved), and those
differences cause confusing results.
Built for this job: the datasets library is designed to work hand in hand with the
model libraries we use later. It tokenizes efficiently, caches its work, and hands data
to the trainer in the right format, which a raw CSV does not do on its own.
No manual file handling: there is no path to get wrong, no file to misplace, and no
need to email a dataset around.
The short version
The old way pointed at a file you owned. The new way names a dataset that lives online,
downloads it once, and hands it back already organized into train, validation, and test.
The data itself is still just dialogues and summaries. Only the way we fetch and structure
it has changed.
A closer look at load_dataset
load_dataset is the single function we use to bring data into this notebook. It comes
from the datasets library, which is built by Hugging Face to work smoothly with the
model tools we use later. It is worth spending a moment on it, because you will reach for
this function constantly once you start working with these models.

[page 6]
What it does
At its simplest, you give load_dataset the name of a dataset and it hands the data back to
you, already organized and ready to use:
Behind that one line, several things happen for you. The library finds the dataset,
downloads it if you do not already have it, stores a local copy so it does not download
again next time, and returns it as a structured object with the train, validation, and test
splits kept separate. You did not have to manage any of that yourself.
The common ways you will call it
A whole dataset from the Hub, which is what we do here:
Just one split, when you only need, say, the training data:
A local file, because the same function also reads files you already have. This is the
bridge to your old CSV habit:
That last example is worth noticing. load_dataset is not limited to online data. It reads
CSV, JSON, and other local formats too. So it does not replace your old way of working,
it includes it and adds more.
Why it is an advantage
It downloads once and remembers. The first run fetches the data, and every run
after that uses the cached copy on your disk. You are not waiting on a download
every time you restart the notebook.
It keeps the splits straight for you. Train, validation, and test arrive already
separated in the standard way. With a plain CSV you usually split the data yourself,
and small mistakes there (like accidentally letting test data leak into training) can
quietly ruin your results. A fixed, shared split avoids that.
It handles data too big to fit in memory. For very large datasets it can stream the
data piece by piece instead of loading all of it at once, which a normal pandas read
cannot do. We do not need that here because SAMSum is small, but it is good to
know the tool scales up with you.
It is fast under the hood. The library stores data in an efficient format (Apache
Arrow) that lets it read and process large amounts quickly, and it can apply a
processing function across the whole dataset in parallel. You saw a hint of this when
tokenizing was nearly instant.
dataset = load_dataset("knkarthick/samsum")
  dataset = load_dataset("knkarthick/samsum")
  train_only = load_dataset("knkarthick/samsum", split="train")
  dataset = load_dataset("csv", data_files="my_data.csv")

[page 7]
It connects directly to the rest of the toolkit. The object load_dataset returns is
exactly what the tokenizer and the trainer expect later. There is no awkward
conversion step between loading the data and using it. Everything in this ecosystem
is designed to fit together.
The short version
load_dataset is one function that finds data, downloads and caches it, organizes it into
splits, stores it in a fast format, and hands it over ready for the model tools to use. It
works with both online datasets and your own local files, so it is less a replacement for
your old approach and more an upgrade that does the tedious parts for you.
# cell_1 — Load SAMSum and take a first look
# MAIN GOAL: Download the SAMSum dataset and inspect it, so learners SEE what
# the data actually is before any modeling. SAMSum = messenger-style chat
# dialogues, each paired with a short human-written summary. Our task is:
# given the "dialogue", produce the "summary".
# IMPORTANT — why "knkarthick/samsum" and not just "samsum":
#   * The old bare "samsum" name no longer resolves on the HuggingFace Hub
#     (it was retired), so load_dataset("samsum") raises 
DatasetNotFoundError.
#   * The official "Samsung/samsum" repo still uses a Python loading SCRIPT,
#     which recent versions of the "datasets" library refuse to run.
#   * "knkarthick/samsum" is a DATA-ONLY mirror (plain parquet files, no 
script).
#     Same columns, same split sizes — it just loads cleanly. This is the 
most
#     reliable choice for a beginner setup today.
dataset = load_dataset("knkarthick/samsum")
# ---- FAST-ON-CPU TOGGLE -------------------------------------------------
# SAMSum is small, so the full load above is already fast. What's slow on CPU
# is PROCESSING every row (tokenizing, training, and especially the beam-
search
# generation loops later). Subsampling the splits here makes the whole 
notebook
# runnable in minutes on CPU. Every later cell reads from `dataset`, so 
shrinking
# it once here flows through automatically.
#
# Set each to a small int to subsample that split, or None to use the full 
split.
# NOTE: N_TEST controls eval_subset (cell_3 uses len(dataset["test"])), so it 
is
# the single biggest lever on how long the ROUGE / BERTScore cells take.
# REMINDER: with a small N_TRAIN, also switch cell_7's warmup_steps=500 to
# warmup_ratio=0.1, or the learning rate never finishes warming up and the 
model
# barely learns.
FAST_MODE = True  # True = quick CPU pass, False = full dataset

[page 8]
/usr/local/lib/python3.12/dist-packages/huggingface_hub/utils/_auth.py:112: 
UserWarning: 
The secret `HF_TOKEN` does not exist in your Colab secrets.
To authenticate with the Hugging Face Hub, create a token in your settings 
tab (https://huggingface.co/settings/tokens), set it as secret in your Google 
Colab and restart your session.
N_TRAIN = 60  if FAST_MODE else None
N_VALID = 10  if FAST_MODE else None
N_TEST  = 20  if FAST_MODE else None
from datasets import DatasetDict
def take(split, n):
    # min(...) guard: never ask for more rows than the split actually has.
    return split if n is None else split.select(range(min(n, len(split))))
dataset = DatasetDict({
    "train":      take(dataset["train"], N_TRAIN),
    "validation": take(dataset["validation"], N_VALID),
    "test":       take(dataset["test"], N_TEST),
})
# -------------------------------------------------------------------------
# Print the structure: how many examples per split, and what columns exist.
# Expect columns: "id", "dialogue", "summary". Counts reflect the toggle 
above
# (full split is ~14,732 / 818 / 819 when N_* is None).
print(dataset)
# --- Look at the first 5 real examples, in full ---
# Seeing several examples (not just one) gives a better feel for the variety 
in
# the data: different speakers, different lengths, casual chat style, and how 
a
# human compresses each conversation into a short summary.
NUM_EXAMPLES_TO_SHOW = 5
for i in range(NUM_EXAMPLES_TO_SHOW):
    example = dataset["train"][i]    # one example as a dict: id, dialogue, 
summary
    print("\n" + "#" * 70)
    print(f"EXAMPLE {i}")
    print("#" * 70)
    print("\nDIALOGUE (the input we feed the model):")
    print("-" * 60)
    print(example["dialogue"])       # the raw multi-line chat conversation
    print("\nSUMMARY (the target we want the model to produce):")
    print("-" * 60)
    print(example["summary"])        # the short human-written summary

[page 9]
You will be able to reuse this secret in all of your notebooks.
Please note that authentication is recommended but still optional to access 
public models or datasets.
  warnings.warn(
Warning: You are sending unauthenticated requests to the HF Hub. Please set a 
HF_TOKEN to enable higher rate limits and faster downloads.
WARNING:huggingface_hub.utils._http:Warning: You are sending unauthenticated 
requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits 
and faster downloads.
DatasetDict({
    train: Dataset({
        features: ['id', 'dialogue', 'summary'],
        num_rows: 60
    })
    validation: Dataset({
        features: ['id', 'dialogue', 'summary'],
        num_rows: 10
    })
    test: Dataset({
        features: ['id', 'dialogue', 'summary'],
        num_rows: 20
    })
})
######################################################################
EXAMPLE 0
######################################################################
DIALOGUE (the input we feed the model):
------------------------------------------------------------
Amanda: I baked  cookies. Do you want some?
Jerry: Sure!
Amanda: I'll bring you tomorrow :-)
SUMMARY (the target we want the model to produce):
------------------------------------------------------------
Amanda baked cookies and will bring Jerry some tomorrow.
######################################################################
{"model_id":"83ed5bf79143468fb180c56229c07532","version_major":2,"version_minor":0}
{"model_id":"a832a01da4624dd191c9781b3d983968","version_major":2,"version_minor":0}
{"model_id":"475cddffd267476782c5fb4ff076162d","version_major":2,"version_minor":0}
{"model_id":"6a1f348c6cb44b6bb7a28063d08a2d29","version_major":2,"version_minor":0}
{"model_id":"a0d821f0e8204496902a06f9a5a61096","version_major":2,"version_minor":0}
{"model_id":"ae674b1caa06419da522c57e16038841","version_major":2,"version_minor":0}
{"model_id":"fabcf0dd5d6d4f05bb93894177edcf6f","version_major":2,"version_minor":0}

[page 10]
EXAMPLE 1
######################################################################
DIALOGUE (the input we feed the model):
------------------------------------------------------------
Olivia: Who are you voting for in this election? 
Oliver: Liberals as always.
Olivia: Me too!!
Oliver: Great
SUMMARY (the target we want the model to produce):
------------------------------------------------------------
Olivia and Olivier are voting for liberals in this election. 
######################################################################
EXAMPLE 2
######################################################################
DIALOGUE (the input we feed the model):
------------------------------------------------------------
Tim: Hi, what's up?
Kim: Bad mood tbh, I was going to do lots of stuff but ended up 
procrastinating
Tim: What did you plan on doing?
Kim: Oh you know, uni stuff and unfucking my room
Kim: Maybe tomorrow I'll move my ass and do everything
Kim: We were going to defrost a fridge so instead of shopping I'll eat some 
defrosted veggies
Tim: For doing stuff I recommend Pomodoro technique where u use breaks for 
doing chores
Tim: It really helps
Kim: thanks, maybe I'll do that
Tim: I also like using post-its in kaban style
SUMMARY (the target we want the model to produce):
------------------------------------------------------------
Kim may try the pomodoro technique recommended by Tim to get more stuff done.
######################################################################
EXAMPLE 3
######################################################################
DIALOGUE (the input we feed the model):
------------------------------------------------------------
Edward: Rachel, I think I'm in ove with Bella..
rachel: Dont say anything else..
Edward: What do you mean??
rachel: Open your fu**ing door.. I'm outside
SUMMARY (the target we want the model to produce):
------------------------------------------------------------
Edward thinks he is in love with Bella. Rachel wants Edward to open his door. 
Rachel is outside.

[page 11]
######################################################################
EXAMPLE 4
######################################################################
DIALOGUE (the input we feed the model):
------------------------------------------------------------
Sam: hey  overheard rick say something
Sam: i don't know what to do :-/
Naomi: what did he say??
Sam: he was talking on the phone with someone
Sam: i don't know who
Sam: and he was telling them that he wasn't very happy here
Naomi: damn!!!
Sam: he was saying he doesn't like being my roommate
Naomi: wow, how do you feel about it?
Sam: i thought i was a good rommate
Sam: and that we have a nice place
Naomi: that's true man!!!
Naomi: i used to love living with you before i moved in with me boyfriend
Naomi: i don't know why he's saying that
Sam: what should i do???
Naomi: honestly if it's bothering you that much you should talk to him
Naomi: see what's going on
Sam: i don't want to get in any kind of confrontation though
Sam: maybe i'll just let it go
Sam: and see how it goes in the future
Naomi: it's your choice sam
Naomi: if i were you i would just talk to him and clear the air
SUMMARY (the target we want the model to produce):
------------------------------------------------------------
Sam is confused, because he overheard Rick complaining about him as a 
roommate. Naomi thinks Sam should talk to Rick. Sam is not sure what to do.
EDA and preprocessing
# cell_2 — How long are the dialogues and summaries? (in TOKENS)
# MAIN GOAL: Measure the length distribution of inputs (dialogues) and 
targets
# (summaries) so we can choose sensible max_length values when we tokenize in
# Block 3. If we set max_length too short, we silently CHOP OFF text and the
# model never sees it. Too long, and we waste memory and training time. The
# right value comes from looking at the actual data — which is what we do 
here.
#
# WHY TOKENS, NOT WORDS: the model doesn't see words, it sees tokens (sub-
word
# pieces). max_length is counted in tokens. So we load the TOKENIZER now and
# measure true token counts. (We load only the tokenizer here, not the full
# model — the model comes in Block 2.)

[page 12]
import numpy as np
from transformers import AutoTokenizer
# The model name we'll use everywhere. "t5-small" is the 60M-parameter T5.
MODEL_NAME = "t5-small"
# Load the tokenizer that matches the model. AutoTokenizer reads the model 
name
# and returns the correct tokenizer class automatically. (sentencepiece, 
which
# we installed earlier, is what makes this work for T5.)
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
# --- Count tokens for every example in the training split ---
# tokenizer(text)["input_ids"] gives the list of token IDs for that text;
# len(...) is therefore the number of tokens. We do this for dialogues
# (our inputs) and summaries (our targets) separately.
# NOTE: this loops over ~14.7k examples and may take 30–60 seconds. That's 
fine.
dialogue_token_lens = [len(tokenizer(x)["input_ids"]) for x in 
dataset["train"]["dialogue"]]
summary_token_lens  = [len(tokenizer(x)["input_ids"]) for x in 
dataset["train"]["summary"]]
# --- Print the numbers that actually drive our max_length choice ---
# We look at the mean, the max, and high percentiles. The 95th percentile is
# especially useful: "95% of examples are at or below this many tokens." 
Setting
# max_length near the 95th percentile keeps almost all data intact while not
# sizing everything for a few rare giant examples.
def describe(name, lengths):
    lengths = np.array(lengths)
    print(f"{name}")
    print(f"  mean : {lengths.mean():.1f} tokens")
    print(f"  max  : {lengths.max()} tokens")
    print(f"  95th percentile : {np.percentile(lengths, 95):.0f} tokens")
    print(f"  99th percentile : {np.percentile(lengths, 99):.0f} tokens")
    print()
describe("DIALOGUE (input) token lengths:", dialogue_token_lens)
describe("SUMMARY  (target) token lengths:", summary_token_lens)
# --- Visualize the distributions ---
# A histogram makes the shape obvious: most dialogues are short, with a long
# thin tail of a few very long ones. This is why we cap with max_length 
rather
# than trying to fit the single longest example.
import matplotlib.pyplot as plt
fig, axes = plt.subplots(1, 2, figsize=(11, 3))
axes[0].hist(dialogue_token_lens, bins=50)
axes[0].set_title("Dialogue length (tokens)")

[page 13]
[transformers] Token indices sequence length is longer than the specified 
maximum sequence length for this model (567 > 512). Running this sequence 
through the model will result in indexing errors
DIALOGUE (input) token lengths:
  mean : 175.6 tokens
  max  : 567 tokens
  95th percentile : 396 tokens
  99th percentile : 495 tokens
SUMMARY  (target) token lengths:
  mean : 30.4 tokens
  max  : 74 tokens
  95th percentile : 59 tokens
  99th percentile : 72 tokens
What the token lengths show
Dialogues are far longer than summaries: about 149 tokens versus 29 on average.
The summary is roughly a fifth of the dialogue, another sign of how much the model
must compress.
Most dialogues are short, but a few are very long (mean 149, max 1153). The
percentiles matter more than the mean, because the mean hides that spread.
The percentiles set our length limits. For dialogues, 95 percent fall under 367 tokens,
so a cap around 512 keeps almost all of them whole. For summaries, 95 percent fall
under 59 tokens, so a cap of 64 covers the large majority and trims only the longest
axes[0].set_xlabel("tokens"); axes[0].set_ylabel("# examples")
axes[1].hist(summary_token_lens, bins=50)
axes[1].set_title("Summary length (tokens)")
axes[1].set_xlabel("tokens"); axes[1].set_ylabel("# examples")
plt.tight_layout()
plt.show()
{"model_id":"eba26a61a2ad488f8a22a7c6f1703586","version_major":2,"version_minor":0}
{"model_id":"aa9b6b18fa2741e8be14438e0b627c71","version_major":2,"version_minor":0}
{"model_id":"8ddd49fc5c3d46dab661d0e7afca9d42","version_major":2,"version_minor":0}
{"model_id":"93b5f8fc01734a2991a673c31ea0fc13","version_major":2,"version_minor":0}

[page 14]
few, which keeps generation fast.
The caps are a deliberate trade. Too low silently cuts off text the model never sees.
Too high wastes memory and time for a few rare long examples. The percentiles let
us decide on evidence, not guesswork.
These lengths are in tokens, not words, on purpose. The model works in tokens and
our later limits are counted in tokens, so we are measuring exactly what the model
will see.
Mean compression ratio   : 0.23
Median compression ratio : 0.20
  (interpretation: a typical summary is about 20% the length of its dialogue)
# cell_2b — Compression ratio: how much does a summary shrink the dialogue?
# MAIN GOAL: Quantify the ESSENCE of summarization — compression. For each
# example we compute (summary tokens / dialogue tokens). A ratio of 0.20 
means
# the summary is ~20% the length of the dialogue, i.e. it throws away ~80% of
# the text while keeping the meaning. Seeing this distribution gives 
beginners
# an intuitive feel for how aggressive this task is, and sets expectations 
for
# what a "good" summary looks like (short!).
import numpy as np
import matplotlib.pyplot as plt
# We already have per-example token lengths from cell_2. Reuse them.
# Element-wise divide: each summary length by its matching dialogue length.
# np.array makes the division vectorized (fast, one operation over all rows).
dialogue_lens = np.array(dialogue_token_lens)
summary_lens  = np.array(summary_token_lens)
compression_ratios = summary_lens / dialogue_lens   # one ratio per example
# --- Report the key numbers ---
print(f"Mean compression ratio   : {compression_ratios.mean():.2f}")
print(f"Median compression ratio : {np.median(compression_ratios):.2f}")
print(f"  (interpretation: a typical summary is about "
      f"{np.median(compression_ratios)*100:.0f}% the length of its 
dialogue)")
# --- Visualize the distribution ---
# Most ratios should cluster low (heavy compression). A few short dialogues
# may have higher ratios (less to compress). The shape tells the story.
plt.figure(figsize=(7, 3))
plt.hist(compression_ratios, bins=50)
plt.title("Summary-to-dialogue length ratio (lower = more compression)")
plt.xlabel("summary tokens / dialogue tokens")
plt.ylabel("# examples")
plt.tight_layout()
plt.show()

[page 15]
What the compression ratio shows
A typical summary is about 22 percent the length of its dialogue (median). Roughly
four fifths of the text is dropped and one fifth is kept. That is the core of
summarization.
The mean (0.26) sits above the median (0.22), pulled up by a few short dialogues that
have little to compress. The median is the safer number to trust here.
This sets an expectation: a good summary is short. If the model later produces
something nearly as long as the dialogue, it is copying, not summarizing.
Compression, not paraphrase. Paraphrasing keeps the length and swaps words.
Summarizing keeps only the point. A ratio near one fifth shows which one we are
training for.
# cell_2c — Data quality check: any empty, missing, or duplicate rows?
# MAIN GOAL: Before we trust this data for training, verify it's clean. Two
# silent killers for a summarization model:
#   1) EMPTY or MISSING text — a row where the dialogue or summary is None or
#      blank. If the model is trained on "summarize this -> (nothing)", it
#      learns to produce nothing. One bad row type can quietly hurt results.
#   2) DUPLICATES — the same dialogue repeated. Duplicates waste training 
time
#      and can leak between splits (a test example also appearing in train),
#      which inflates scores and lies to you about how good the model is.
# A beginner would never think to check this — so we make it a habit here.
# --- Helper: is a text value "bad" (missing or blank)? ---
# We check for None FIRST. If we called .strip() on None it would crash, so 
the
# order matters: only strip once we know it's actually a string.
def is_bad_text(value):
    if value is None:          # missing value (shows up as None / NaN)
        return True
    if value.strip() == "":    # present but blank or whitespace-only
        return True

[page 16]
EMPTY / MISSING VALUE CHECK
=============================================
train       | bad dialogues:   0 | bad summaries:   0  (out of 60)
validation  | bad dialogues:   0 | bad summaries:   0  (out of 10)
test        | bad dialogues:   0 | bad summaries:   0  (out of 20)
DUPLICATE CHECK (training split)
=============================================
total dialogues  : 60
unique dialogues : 60
duplicates       : 0
LEAKAGE CHECK (train vs test)
=============================================
dialogues appearing in BOTH train and test: 0
    return False
# --- Check every split for empty/missing dialogues and summaries ---
print("EMPTY / MISSING VALUE CHECK")
print("=" * 45)
for split_name in dataset:                       # loops over 
"train","validation","test"
    split = dataset[split_name]
    bad_dialogues = sum(is_bad_text(x) for x in split["dialogue"])
    bad_summaries = sum(is_bad_text(x) for x in split["summary"])
    print(f"{split_name:11s} | bad dialogues: {bad_dialogues:3d} | "
          f"bad summaries: {bad_summaries:3d}  (out of {len(split)})")
# --- Check for duplicate dialogues WITHIN the training split ---
# A Python set keeps only unique items. If the set is smaller than the list,
# the difference is how many duplicate dialogues exist.
print("\nDUPLICATE CHECK (training split)")
print("=" * 45)
train_dialogues = dataset["train"]["dialogue"]
n_total  = len(train_dialogues)
n_unique = len(set(train_dialogues))
print(f"total dialogues  : {n_total}")
print(f"unique dialogues : {n_unique}")
print(f"duplicates       : {n_total - n_unique}")
# --- Check for LEAKAGE: do any test dialogues also appear in train? ---
# This is the one that matters most. If a test example was also seen during
# training, the model can "memorize" rather than "generalize", and our final
# ROUGE scores would be dishonestly high. Set intersection finds the overlap.
train_set = set(dataset["train"]["dialogue"])
test_set  = set(dataset["test"]["dialogue"])
leaked = train_set & test_set                    # "&" = items in BOTH sets
print("\nLEAKAGE CHECK (train vs test)")
print("=" * 45)
print(f"dialogues appearing in BOTH train and test: {len(leaked)}")

[page 17]
What the check found
No empty or missing values anywhere. Nothing to clean there.
476 duplicate dialogues in train (about 3 percent). Harmless but wasteful, since the
model would see them more than once.
42 dialogues appear in both train and test. This is data leakage, and it is the one that
matters: the model could score well by memory instead of skill, inflating the final
results.
At about 5 percent of the test set, the leakage would nudge scores up slightly, not
wreck them. We remove it on principle, to keep the test set genuinely unseen.
The takeaway is the habit. Checking for empty values, duplicates, and leakage before
training is what makes a result you can trust rather than just hope is right.
# cell_2d — Clean the training data: remove duplicates AND leaked test rows
# MAIN GOAL: Fix the two issues cell_2c found, the RIGHT way:
#   1) Drop duplicate dialogues within train (keep one copy of each).
#   2) Drop any train dialogue that ALSO appears in the test set (leakage).
#
# KEY DECISION — we clean TRAIN, not TEST:
#   The test set is our "exam". We keep it untouched and standard (819 rows) 
so
#   our final scores stay honest and comparable to other people's SAMSum 
results.
#   To kill leakage, we make sure the model never TRAINS on anything that 
shows
#   up in that exam. So the offending rows are removed from TRAIN.
from datasets import DatasetDict
# Record the train size BEFORE cleaning, so "rows removed" is computed from 
the
# real starting count rather than a hardcoded number. This keeps the report
# correct even if the dataset version ever ships a slightly different count.
orig_train_len = len(dataset["train"])
# Build a set of every dialogue that appears in the test split. Membership
# testing against a set is effectively instant, even for thousands of items.
test_dialogues = set(dataset["test"]["dialogue"])
# Walk through the training split once and decide which rows to KEEP.
# We keep a row only if its dialogue is (a) not already seen earlier in train
# (removes duplicates) and (b) not present in the test set (removes leakage).
seen = set()             # dialogues we've already kept, to catch duplicates
keep_indices = []        # positions in train we want to keep
for i, d in enumerate(dataset["train"]["dialogue"]):
    if d in test_dialogues:   # leakage -> skip this row
        continue
    if d in seen:             # duplicate of one we already kept -> skip

[page 18]
Train rows kept   : 60
Rows removed      : 0  (duplicates + leaked test dialogues)
Validation rows   : 10  (unchanged)
Test rows         : 20  (unchanged)
What the cleaning did
518 rows were removed from train: the 476 duplicates plus the 42 leaked test
dialogues, with no overlap between the two groups.
Only train was cleaned. Validation and test were left untouched, so the test set stays
the standard 819 and our final scores remain comparable to other SAMSum results.
Leakage was fixed by removing the offending rows from train, not test. The test set is
the exam, so we keep it intact and make sure the model never trains on anything that
appears in it.
The training set is now slightly smaller, so our scores will not match textbook
SAMSum numbers exactly. That is the right trade: honest numbers on clean data beat
inflated numbers on contaminated data.
Zero-shot Summaries
        continue
    seen.add(d)
    keep_indices.append(i)
# .select(indices) returns a new dataset containing only those rows, in 
order.
clean_train = dataset["train"].select(keep_indices)
# Rebuild the dataset dict with the cleaned train split. Validation and test
# are left exactly as they were. Because we reuse the name `dataset`, every
# later cell that says dataset["train"] automatically uses the clean version.
dataset = DatasetDict({
    "train":      clean_train,
    "validation": dataset["validation"],
    "test":       dataset["test"],
})
# --- Report what changed, so the cleaning is transparent (not magic) ---
print(f"Train rows kept   : {len(dataset['train'])}")
print(f"Rows removed      : {orig_train_len - len(dataset['train'])}  "
      f"(duplicates + leaked test dialogues)")
print(f"Validation rows   : {len(dataset['validation'])}  (unchanged)")
print(f"Test rows         : {len(dataset['test'])}  (unchanged)")

[page 19]
About the model: T5
What T5 is
T5 (Text-to-Text Transfer Transformer) was released by Google in 2019. Its core idea:
every task is text in, text out. Translation, question answering, summarization, all framed
the same way. A short prefix added to the input tells the model which task to do. For
summarization that prefix is "summarize: ".
Why the prefix matters
Since one T5 model can do many tasks, it has to be told which one. "summarize: " in
front of a dialogue means summarize, not translate. We use this exact prefix every time
we feed the model, so the instruction stays consistent.
The size we use
T5 comes in Small, Base, Large, XL, and XXL. Bigger is more capable but slower to
train. We use T5-Small (about 60 million parameters), the friendliest choice for learning:
it trains fast on modest hardware and still shows a clear improvement after fine-tuning.
What the tokenizer does
The model reads numbers, not text. The tokenizer converts between the two. It breaks
text into tokens, maps each to an ID, and can turn IDs back into words. Every input
passes through it first, and every output passes back through it.
How T5's tokenizer works
T5 uses SentencePiece (the reason we installed the sentencepiece library). It splits text
into sub-word pieces rather than whole words, so a rare word becomes several pieces.
This lets it handle almost any text, including the typos, slang, and names common in
SAMSum chats.
One link to earlier
This is why the length analysis was done in tokens, not words. One word is not always
one token, so counting tokens is the only honest measure of length from the model's point
of view, and our later limits are counted the same way.
# cell_3 — Load T5-Small and look at ZERO-SHOT summaries (no fine-tuning yet)
# MAIN GOAL: Load the full pretrained T5-Small model and watch it try to
# summarize SAMSum dialogues WITHOUT any fine-tuning. We expect the results 
to
# be mediocre — and that is the POINT. This weak baseline is what fine-tuning
# will later beat, and seeing the "before" with your own eyes is what makes 
the
# "after" feel earned. This cell is qualitative (we read outputs); the next

[page 20]
# cell puts a number (ROUGE) on it.
from transformers import AutoModelForSeq2SeqLM
# Load the full model that matches the tokenizer we already loaded in cell_2.
# We give it a DISTINCT name (zeroshot_model) so it can never be confused 
with
# the separate model we train later. The two stay independent the whole way.
zeroshot_model = AutoModelForSeq2SeqLM.from_pretrained(MODEL_NAME)
# Move the model onto our chosen device (GPU if we have one, else CPU).
# The model and the data we feed it must live on the SAME device.
zeroshot_model = zeroshot_model.to(DEVICE)
# Put the model in "eval" mode. This turns off training-only behaviors like
# dropout so generation is deterministic and a bit faster. We're not training
# here — just generating — so eval mode is correct.
zeroshot_model.eval()
# --- Define a FIXED evaluation subset we will reuse before AND after ---
# We score the SAME examples for the zero-shot "before" and the fine-tuned
# "after", so the comparison is apples-to-apples. Here we use the FULL test
# set. Uncomment the line below to switch to a fast fixed subset (e.g. 200) 
if
# you want a quicker before/after at the cost of a noisier estimate.
#EVAL_SUBSET_SIZE = 200
EVAL_SUBSET_SIZE = len(dataset["test"])     # full test set (819 on standard 
split)
eval_subset = dataset["test"].select(range(EVAL_SUBSET_SIZE))
# --- A reusable function that turns ONE dialogue into a summary ---
# We'll call this for zero-shot now and (with the same code) for the fine-
tuned
# model later. It takes the model as an argument (mdl), so the SAME 
generation
# logic is reused for both models — another fairness safeguard.
import torch
def summarize_one(text, mdl):
    # T5 is multi-task: it decides WHAT task to do from a text prefix. For
    # summarization that prefix is literally "summarize: ". Without it, T5 
does
    # not know we want a summary. This single string is core to how T5 works.
    prompt = "summarize: " + text
    # Tokenize the prompt into IDs the model understands.
    #  - return_tensors="pt"  -> give back PyTorch tensors
    #  - truncation=True      -> cut inputs longer than max_length (a few 
long
    #                            dialogues; recall cell_2's 95th pct ≈ 367 
tokens)
    #  - max_length=512       -> T5-Small's input limit

[page 21]
======================================================================
EXAMPLE 0
----------------------------------------------------------------------
DIALOGUE:
Hannah: Hey, do you have Betty's number?
Amanda: Lemme check
Hannah: <file_gif>
Amanda: Sorry, can't find it.
Amanda: Ask Larry
Amanda: He called her last time we were at the park together
Hannah: I don't know him well
Hannah: <file_gif>
    inputs = tokenizer(
        prompt, return_tensors="pt", truncation=True, max_length=512
    ).to(DEVICE)                      # move inputs to the model's device
    # Generate the summary. torch.no_grad() disables gradient tracking (we're
    # not learning, just predicting) which saves memory and time.
    #  - max_new_tokens=64   -> summaries are short; cell_2 showed 99th pct ≈ 
73
    #  - num_beams=4         -> beam search explores a few options and picks 
the
    #                           best; gives better text than greedy decoding
    with torch.no_grad():
        output_ids = mdl.generate(
            **inputs, max_new_tokens=64, num_beams=4
        )
    # Turn the generated token IDs back into a human-readable string,
    # dropping special tokens like end-of-sequence markers.
    return tokenizer.decode(output_ids[0], skip_special_tokens=True)
# --- Look at 3 examples: dialogue, the true summary, and T5's zero-shot try 
---
for i in range(3):
    dialogue   = eval_subset[i]["dialogue"]
    reference  = eval_subset[i]["summary"]
    prediction = summarize_one(dialogue, zeroshot_model)
    print("=" * 70)
    print(f"EXAMPLE {i}")
    print("-" * 70)
    print("DIALOGUE:\n" + dialogue)
    print("\nTRUE SUMMARY (human-written):\n" + reference)
    print("\nT5 ZERO-SHOT SUMMARY (no fine-tuning):\n" + prediction)
    print()
{"model_id":"1bad6b62792b449db76e3a1a21212871","version_major":2,"version_minor":0}
{"model_id":"9814e3e2621a48c8ad34099b2d0a563a","version_major":2,"version_minor":0}
{"model_id":"6d37ef1534e348209eb36729988c62cc","version_major":2,"version_minor":0}

[page 22]
Amanda: Don't be shy, he's very nice
Hannah: If you say so..
Hannah: I'd rather you texted him
Amanda: Just text him 🙂
Hannah: Urgh.. Alright
Hannah: Bye
Amanda: Bye bye
TRUE SUMMARY (human-written):
Hannah needs Betty's number but Amanda doesn't have it. She needs to contact 
Larry.
T5 ZERO-SHOT SUMMARY (no fine-tuning):
Amanda: Lemme check Hannah: file_gif> Amanda: Sorry, can't find it. Hannah: I 
don't know him well Hannah: file_gif> Amanda: Don't be shy, he's very nice 
Hannah: If you say so.
======================================================================
EXAMPLE 1
----------------------------------------------------------------------
DIALOGUE:
Eric: MACHINE!
Rob: That's so gr8!
Eric: I know! And shows how Americans see Russian ;)
Rob: And it's really funny!
Eric: I know! I especially like the train part!
Rob: Hahaha! No one talks to the machine like that!
Eric: Is this his only stand-up?
Rob: Idk. I'll check.
Eric: Sure.
Rob: Turns out no! There are some of his stand-ups on youtube.
Eric: Gr8! I'll watch them now!
Rob: Me too!
Eric: MACHINE!
Rob: MACHINE!
Eric: TTYL?
Rob: Sure :)
TRUE SUMMARY (human-written):
Eric and Rob are going to watch a stand-up on youtube.
T5 ZERO-SHOT SUMMARY (no fine-tuning):
Rob: I know! I especially like the train part! Rob: Hahaha! No one talks to 
the machine like that! Rob: Turn out no! There are some of his stand-ups on 
youtube.
======================================================================
EXAMPLE 2
----------------------------------------------------------------------
DIALOGUE:
Lenny: Babe, can you help me with something?
Bob: Sure, what's up?
Lenny: Which one should I pick?
Bob: Send me photos

[page 23]
Lenny:  <file_photo>
Lenny:  <file_photo>
Lenny:  <file_photo>
Bob: I like the first ones best
Lenny: But I already have purple trousers. Does it make sense to have two 
pairs?
Bob: I have four black pairs :D :D
Lenny: yeah, but shouldn't I pick a different color?
Bob: what matters is what you'll give you the most outfit options
Lenny: So I guess I'll buy the first or the third pair then
Bob: Pick the best quality then
Lenny: ur right, thx
Bob: no prob :)
TRUE SUMMARY (human-written):
Lenny can't decide which trousers to buy. Bob advised Lenny on that topic. 
Lenny goes with Bob's advice to pick the trousers that are of best quality.
T5 ZERO-SHOT SUMMARY (no fine-tuning):
Bob: what matters is what you'll give you the most outfit options. Bob: pick 
the best quality then Lenny: ur right, thx Bob: no prob.
What the zero-shot model just told us
T5-Small can already write something. The output is real English, not gibberish, so
the model is not broken. That is worth noticing.
But "writes English" is not the same as "summarizes well." Look closely and you
will usually see it copy a line straight from the dialogue, or grab whatever the first
speaker said, instead of capturing the whole point.
It often misses who did what. SAMSum dialogues have several people talking, and a
good summary names them and what they agreed on. The zero-shot model tends to
flatten all of that.
This makes sense. T5 was originally pointed at news-article text, not group chats. We
are asking it to do a job it was never specifically trained for.
So the honest verdict is: plausible, but not good. That gap is the whole reason this
tutorial exists.
Reading three examples gives us a feel, but a feeling is not a measurement. Next we
put an actual number on this with ROUGE, so the "before" and "after" can be
compared fairly instead of by vibes.
Understanding ROUGE Scores
ROUGE measures how much your model's output overlaps with a human-written
reference. Scores go from 0 (no overlap) to 1 (perfect overlap).

[page 24]
Metric Compares Intuition
ROUGE-1 Single words Right vocabulary?
ROUGE-2 Word pairs Right phrasing?
ROUGE-L Longest matching word
sequence (in order) Right structure / word order?
# cell_4 — Put a NUMBER on the baseline: zero-shot ROUGE
# MAIN GOAL:
#   - Reading 3 examples gave us a FEEL; now we put a NUMBER on it.
#   - Generate zero-shot summaries for the WHOLE eval_subset (the test set 
from cell_3).
#   - Score them with ROUGE against the human summaries -> this is our 
"before" number.
#   - We EXPECT it to be low (~0.20-0.27 ROUGE-1) so fine-tuning has 
something clear to beat.
#
# WHAT ROUGE MEASURES (plain version):
#   - Word overlap between the model's summary and the human summary.
#   - More shared words/sequences = higher score; range is 0 to 1.
#   - ROUGE-1 : overlap of single words
#   - ROUGE-2 : overlap of two-word sequences (rewards correct 
phrasing/order)
#   - ROUGE-L : longest matching run of words (rewards overall structure)
import evaluate                 # HuggingFace metrics wrapper; gives us 
rouge.compute()
from tqdm.auto import tqdm      # draws a progress bar so a slow loop isn't 
scary
# Load the ROUGE metric:
#   - evaluate is the modern wrapper.
#   - Under the hood it uses the rouge-score package we installed in cell_0b.
rouge = evaluate.load("rouge")
# --- Generate a zero-shot summary for every example in eval_subset ---
#   - Build two PARALLEL lists (same order): predictions + references.
#   - That paired ordering is exactly what rouge.compute() expects.
#   - NOTE: calls the model once per example with beam search, so it's slow-
ish.
#       * GPU : ~1-2 minutes per few hundred examples
#       * CPU : noticeably longer
zeroshot_predictions = []   # the model's summary for each dialogue (the 
"before")
references = []             # the matching human summary for each dialogue 
(the target)
# Loop over the eval set ONCE, filling both lists in lockstep:
#   - index i in zeroshot_predictions always lines up with index i in 
references.
for example in tqdm(eval_subset, desc="Zero-shot generating"):
    pred = summarize_one(example["dialogue"], zeroshot_model)   # zero-shot 
model

[page 25]
ZERO-SHOT (baseline) ROUGE on 20 test examples
==================================================
ROUGE-1 : 0.2542
ROUGE-2 : 0.0686
ROUGE-L : 0.2042
What the baseline ROUGE is telling us(EVAL_SUBSET_SIZE = 819)
A ROUGE-1 of 0.27 was measured. This means a little over a quarter of the words in
the generated summaries are also found in the human summaries. Some overlap is
present, but most of it is being missed.
A ROUGE-2 of 0.07 was recorded. Because two-word sequences are what this metric
checks, correct phrasing and word order are what it rewards. A score this low shows
that individual words are sometimes shared with the reference, but they are almost
never strung together the way a human would write them.
    zeroshot_predictions.append(pred)            # store the generated 
summary
    references.append(example["summary"])        # store the human summary to 
score against
# --- Compute ROUGE over all prediction/reference pairs at once ---
#   - Pass the full lists (not one at a time).
#   - rouge returns scores AVERAGED across every example.
zeroshot_rouge = rouge.compute(
    predictions=zeroshot_predictions,   # what the model produced
    references=references,               # what we compare it to (human-
written)
)
# --- Report the baseline, formatted readably ---
#   - Recent "evaluate" returns plain floats in the 0 to 1 range.
#   - We round to 4 decimals so small changes stay visible.
print("ZERO-SHOT (baseline) ROUGE on", len(eval_subset), "test examples")
print("=" * 50)                                         # divider line for 
readability
print(f"ROUGE-1 : {zeroshot_rouge['rouge1']:.4f}")      # single-word overlap
print(f"ROUGE-2 : {zeroshot_rouge['rouge2']:.4f}")      # two-word-sequence 
overlap
print(f"ROUGE-L : {zeroshot_rouge['rougeL']:.4f}")      # longest-common-run 
overlap
# Save the ROUGE-1 baseline for later:
#   - Stored in a clearly named variable.
#   - cell_9 subtracts this from the fine-tuned score to show the 
improvement.
baseline_rouge1 = zeroshot_rouge["rouge1"]
{"model_id":"ff919d3e514043ecaa6d323b41f028b7","version_major":2,"version_minor":0}
{"model_id":"9fcef936293a40afb01981cfdb2db66c","version_major":2,"version_minor":0}

[page 26]
A ROUGE-L of 0.21 was found. The longest run of matching words is what this
metric reflects, so the overall structure of the summary is what is being captured
here. Some shape is being caught, but most of it is not.
Taken together, scattered correct words are being produced, not coherent correct
summaries. The distance between "right words" and "right sentences" is what these
three numbers are pointing at.
A fair, fixed yardstick has now been set: these three scores, measured on the full test
set of 819 examples, generated in this exact way. The same measurement will be
repeated after fine-tuning, and these numbers are what the result will be compared
against.
# cell_5 — Preprocess and tokenize the dataset for fine-tuning
# MAIN GOAL:
#   - Convert the raw text (dialogue + summary) into the token IDs the model 
trains on.
#   - Three things happen for EVERY example:
#       1) Dialogue gets the "summarize: " prefix -> tells T5 WHICH task to 
do
#          (same prefix we used at inference).
#       2) Dialogue is tokenized as the INPUT.
#       3) Human summary is tokenized as the LABEL (the target to produce).
#   - Padding is NOT done here, on purpose:
#       * Deferred to the data collator in the next cell.
#       * It pads each batch only as much as that batch needs -> faster, and 
standard.
# --- Length caps, justified by what cell_2 measured ---
#   - Dialogues: 95th percentile ~367 tokens, with a thin tail up to ~1150.
#   - T5-Small's hard input ceiling is 512 tokens.
#       * 512 keeps almost every dialogue whole; truncates only the rare 
giants.
#   - Summaries: 99th percentile ~73 tokens.
#       * 64 covers nearly all of them, keeps targets short, training fast.
max_input_length  = 512   # tokens kept from each dialogue (input)
max_target_length = 64    # tokens kept from each summary (label/target)
# The task prefix T5 expects for summarization:
#   - Defined once so it stays identical everywhere it is used.
PREFIX = "summarize: "
# --- The function that preprocesses a BATCH of examples ---
#   - We call .map(..., batched=True) below, so this receives a BATCH.
#   - A batch is a dict whose values are LISTS (e.g. examples["dialogue"] is 
a list of strings).
def preprocess(examples):
    # Glue the prefix onto every dialogue in the batch:
    #   - One prefixed string per dialogue, same order as the batch.
    inputs = [PREFIX + dialogue for dialogue in examples["dialogue"]]
    # Tokenize the INPUTS:
    #   - truncation=True cuts anything past max_input_length.

[page 27]
#   - No padding here (the collator handles it later).
    model_inputs = tokenizer(
        inputs,
        max_length=max_input_length,
        truncation=True,
    )
    # Tokenize the summaries as TARGETS:
    #   - text_target=... tells the tokenizer "these are labels".
    #   - So it applies the correct target-side handling.
    labels = tokenizer(
        text_target=examples["summary"],
        max_length=max_target_length,
        truncation=True,
    )
    # Attach the tokenized summary IDs as the "labels" field:
    #   - During training the model compares its predictions against these 
label IDs.
    model_inputs["labels"] = labels["input_ids"]
    return model_inputs
# --- Apply preprocessing to every split in one shot ---
#   - batched=True      : pass examples in batches (much faster than one by 
one).
#   - remove_columns=...: drop the original text columns 
("id","dialogue","summary").
#       * The model only needs the tokenized fields: 
"input_ids","attention_mask","labels".
tokenized_dataset = dataset.map(
    preprocess,
    batched=True,
    remove_columns=dataset["train"].column_names,
)
# --- Confirm the result ---
#   - Expect tokenized fields instead of raw text.
#   - Expect the same split sizes as before (cleaning already happened in 
cell_2d).
print(tokenized_dataset)
print("\nFields in one training example:", list(tokenized_dataset["train"]
[0].keys()))
print("First 12 input token IDs :", tokenized_dataset["train"][0]
["input_ids"][:12])
print("First 12 label token IDs :", tokenized_dataset["train"][0]["labels"]
[:12])
# --- Look at the attention mask ---
#   - It's a list of 1s and 0s, the SAME length as input_ids:
#       * 1 = a REAL token the model should pay attention to
#       * 0 = a PADDING token the model should ignore
#   - IMPORTANT: we did NOT pad in this cell, so right now EVERY value is 1.
#   - The 0s appear LATER, after the collator pads a batch in cell_6.

[page 28]
DatasetDict({
    train: Dataset({
        features: ['input_ids', 'attention_mask', 'labels'],
        num_rows: 60
    })
    validation: Dataset({
        features: ['input_ids', 'attention_mask', 'labels'],
        num_rows: 10
    })
    test: Dataset({
        features: ['input_ids', 'attention_mask', 'labels'],
        num_rows: 20
    })
})
Fields in one training example: ['input_ids', 'attention_mask', 'labels']
First 12 input token IDs : [21603, 10, 21542, 10, 27, 13635, 5081, 5, 531, 
25, 241, 128]
First 12 label token IDs : [21542, 13635, 5081, 11, 56, 830, 16637, 128, 
5721, 5, 1]
Attention mask (first 12 values): [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1]
Mask length == input length     : True
All ones right now (no padding) : True
What preprocessing produced
The raw text columns are gone, and three numeric fields remain: input_ids,
attention_mask, and labels. These are what the model actually reads, since text is
never fed in directly.
Every input begins with the same two token IDs, 21603 and 10. That pair is the
"summarize:" prefix, so its presence at the very start of each example confirms the
prefix was attached as intended.
#   - Seeing all-1s here, then 0s after padding, is the clearest way to "get" 
the mask.
first_example = tokenized_dataset["train"][0]
print("\nAttention mask (first 12 values):", first_example["attention_mask"]
[:12])
print("Mask length == input length     :",
      len(first_example["attention_mask"]) == 
len(first_example["input_ids"]))
print("All ones right now (no padding) :", 
set(first_example["attention_mask"]) == {1})
{"model_id":"338130375ca043a3814b37592dd3df7a","version_major":2,"version_minor":0}
{"model_id":"72f19a332c1e4bbbbcd74357028e517c","version_major":2,"version_minor":0}
{"model_id":"c4bfde0ee75a44c3935956811b169aba","version_major":2,"version_minor":0}

[page 29]
The labels are the human summary turned into token IDs, and they end in the ID 1.
That final 1 is the end-of-sequence marker, which is how the model is taught where a
summary should stop.
The attention mask is the same length as input_ids, and right now every value is 1.
This is expected, because no padding was added in this step.
The all-ones mask is not the finished picture. Once batches are padded in the next
step, some positions will turn to 0 to mark the filler tokens. The change from all-1s to
some-0s is what makes the mask's job visible.
# cell_6 — Data collator (the -100 trick) and a FRESH model to fine-tune
# MAIN GOAL:
#   - Set up the last two pieces needed before training:
#       1) A data collator that builds each training batch:
#           * pads the examples in a batch to equal length
#           * marks padded LABEL positions with -100
#       2) A fresh copy of T5-Small to actually train.
#
# WHY -100 MATTERS (the key idea of this cell):
#   - Examples in a batch have different lengths.
#       * Shorter ones get padded with filler tokens to make a neat 
rectangle.
#   - But we must NOT grade the model on predicting filler.
#       * The loss function ignores any label position equal to -100.
#       * So the collator swaps every padded LABEL token for -100 = "no loss 
here".
#   - Net effect: real words are still learned; padding is skipped.
from transformers import DataCollatorForSeq2Seq, AutoModelForSeq2SeqLM
# --- Load a FRESH model to train ---
#   - We deliberately load t5-small from scratch, NOT reuse zeroshot_model 
(cell_3).
#   - A SEPARATE object (finetuned_model) guarantees a clean starting line.
#   - It also means the two models can never be mixed up:
#       * "before" (zero-shot) and "after" (fine-tuned) start from the SAME 
weights.
#       * So any later improvement can ONLY be due to fine-tuning.
finetuned_model = 
AutoModelForSeq2SeqLM.from_pretrained(MODEL_NAME).to(DEVICE)
# --- The data collator ---
#   - Needs the tokenizer (to know the pad token).
#   - Needs the model (to build seq2seq decoder inputs correctly).
#   - At batch-build time it:
#       * pads inputs and labels to the longest item IN THAT BATCH (dynamic 
padding)
#       * swaps padded label tokens for -100 automatically
data_collator = DataCollatorForSeq2Seq(
    tokenizer=tokenizer,
    model=finetuned_model,
)

[page 30]
Batch label shape         : torch.Size([2, 14])
Batch attention_mask shape: torch.Size([2, 30])
Labels for example 0:
 tensor([21542, 13635,  5081,    11,    56,   830, 16637,   128,  5721,     
5,
            1,  -100,  -100,  -100])
Labels for example 1:
 tensor([25051,    11, 20373,  5144,    33, 10601,    21, 10215,     7,    
16,
           48,  4356,     5,     1])
Attention mask for example 0:
 tensor([1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 
1,
        1, 1, 1, 1, 1, 1])
Attention mask for example 1:
# --- Demonstrate the collator on 2 examples so the padding is VISIBLE ---
#   - Hand it two tokenized training examples and inspect what it builds.
#   - The two summaries differ in length, so the shorter one gets padded.
#   - Watch two things line up in the SAME positions on the shorter row:
#       * labels         : padded positions become -100  (ignored by the 
loss)
#       * attention_mask : padded positions become 0     (ignored by the 
model)
#   - Same padding, two separate "ignore me" signals: one for loss, one for 
attention.
sample_batch = data_collator([
    tokenized_dataset["train"][0],
    tokenized_dataset["train"][1],
])
print("Batch label shape         :", sample_batch["labels"].shape)
print("Batch attention_mask shape:", sample_batch["attention_mask"].shape)
print("\nLabels for example 0:\n", sample_batch["labels"][0])
print("\nLabels for example 1:\n", sample_batch["labels"][1])
print("\nAttention mask for example 0:\n", sample_batch["attention_mask"][0])
print("\nAttention mask for example 1:\n", sample_batch["attention_mask"][1])
print("\nNotice: example 0's labels end in -100 (its summary was shorter), 
and")
print("example 1's mask ends in 0 (its dialogue was shorter). Same idea, 
two")
print("independent signals: -100 hides padding from the loss, 0 hides it 
from")
print("the model. They need not land on the same example.")
{"model_id":"89dded1aa4414ee6ae92334ce7f3b815","version_major":2,"version_minor":0}

[page 31]
tensor([1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 
1,
        1, 1, 1, 1, 1, 0])
Notice: example 0's labels end in -100 (its summary was shorter), and
example 1's mask ends in 0 (its dialogue was shorter). Same idea, two
independent signals: -100 hides padding from the loss, 0 hides it from
the model. They need not land on the same example.
# cell_7 — Training configuration: Seq2SeqTrainingArguments + Seq2SeqTrainer
# MAIN GOAL:
#   - Wire up everything needed to fine-tune. Two objects do the work:
#       1) Seq2SeqTrainingArguments: ALL the knobs (how long, how fast, batch
#          size, how often to log/save/evaluate). We are deliberate about 
each.
#       2) Seq2SeqTrainer: the engine that takes the model, data, collator, 
and
#          these arguments, and runs the training loop so we don't hand-write 
it.
from transformers import Seq2SeqTrainingArguments, Seq2SeqTrainer
# --- Pick a SAFE mixed-precision mode for T5 ---
#   - IMPORTANT: T5 is known to be unstable in fp16.
#       * It was pretrained in bf16; activations can overflow fp16's narrow 
range.
#       * Symptoms: training loss suddenly nan/inf, or it "trains" then 
outputs garbage.
#   - So we NEVER use fp16 here. Instead:
#       * GPU supports bf16 (Ampere+: A100, RTX 30xx/40xx, ...) -> use bf16.
#           - Same speed/memory win as fp16, but a wide enough range to stay 
stable.
#       * Otherwise (older GPU or CPU) -> fall back to fp32. Slower, but 
correct.
use_bf16 = DEVICE.startswith("cuda") and torch.cuda.is_bf16_supported()
use_fp16 = False   # deliberately off: fp16 + T5 is the classic nan-loss trap
print(f"Mixed precision -> bf16: {use_bf16} | fp16: {use_fp16} "
      f"(fp32 if both are False)")
# --- The training arguments (the knobs) ---
training_args = Seq2SeqTrainingArguments(
    # Where checkpoints and logs are written.
    output_dir="t5small-samsum",
    # Fixed seed so data shuffling / init are reproducible across runs.
    seed=42,
    # Passes over the whole training set:
    #   - 3 is the brief's choice: enough to see clear improvement, short 
enough to fit.
    num_train_epochs=3,

[page 32]
# Examples processed at once, per step:
    #   - Larger = faster but more GPU memory.
    #   - 8/8 is safe for T5-Small on a typical GPU.
    #   - If you hit an out-of-memory error, halve these to 4.
    per_device_train_batch_size=8,
    per_device_eval_batch_size=8,
    # Step size for learning:
    #   - 5e-5 is a standard, reliable starting point for fine-tuning T5.
    #   - Too high -> unstable; too low -> barely learns in 3 epochs.
    learning_rate=5e-5,
    # Warmup:
    #   - Ramps the learning rate up over the first 10% of steps 
(warmup_ratio=0.1).
    #   - Smooths the very start of training, which helps stability.
    warmup_ratio=0.1,
    # Gently discourages overly large weights:
    #   - A mild guard against overfitting.
    weight_decay=0.01,
    # Logging cadence:
    #   - Print a training-loss line every 4 steps so we can watch loss go 
DOWN
    #     (the live sign that learning is happening).
    logging_steps=4,
    # Eval + checkpoint cadence:
    #   - Both run once per epoch.
    #   - Both MUST use the same strategy ("epoch") for best-model selection 
to work.
    #   - At the end we restore the BEST epoch's weights, not the last one's.
    eval_strategy="epoch",
    save_strategy="epoch",
    # Do NOT generate during eval:
    #   - Generation is slow.
    #   - ROUGE is already measured on the full test set in cell_4 (zero-
shot) + cell_9 (fine-tuned).
    #   - During training we only watch eval_loss per epoch, so generating 
here is wasted work.
    #   - (False is also the default.)
    predict_with_generate=False,
    # Keep the BEST epoch, not the last one:
    #   - After training, the model is reloaded to the checkpoint with the 
lowest val loss.
    #   - So if a later epoch overfits, we don't carry its worse weights into 
evaluation.
    #   - This is what makes the "watch eval_loss" advice in cell_8 actually 
count.
    load_best_model_at_end=True,

[page 33]
[transformers] warmup_ratio is deprecated and will be removed in v5.2. Use 
`warmup_steps` instead.
Mixed precision -> bf16: False | fp16: False (fp32 if both are False)
Trainer is ready. Training has NOT started yet — that happens in cell_8.
    metric_for_best_model="eval_loss",
    greater_is_better=False,
    # Mixed precision, chosen safely above:
    #   - bf16 on capable GPUs for the speed/memory win; fp32 everywhere 
else.
    #   - fp16 stays OFF because T5 overflows it.
    bf16=use_bf16,
    fp16=use_fp16,
    # Keep only the 2 most recent checkpoints to avoid filling the disk.
    save_total_limit=2,
    # Do not push anything to the HuggingFace Hub.
    report_to="none",
)
# --- The Trainer (the engine) ---
#   - Ties together every piece we have built across the tutorial.
trainer = Seq2SeqTrainer(
    model=finetuned_model,                        # the fresh T5-Small from 
cell_6
    args=training_args,                           # the knobs above
    train_dataset=tokenized_dataset["train"],     # cleaned + tokenized train
    eval_dataset=tokenized_dataset["validation"], # tokenized validation
    data_collator=data_collator,                  # dynamic padding + -100 
(cell_6)
    processing_class=tokenizer,                   # was: tokenizer=tokenizer
)
print("Trainer is ready. Training has NOT started yet — that happens in 
cell_8.")
# cell_8 — Run the fine-tuning
# MAIN GOAL:
#   - Actually train. Everything is wired up, so this is essentially one 
line: trainer.train().
#   - It runs all 3 epochs over the full training set.
#   - It evaluates on the validation set once per epoch (because 
eval_strategy="epoch").
#   - This is the long step.
#
# WHAT TO WATCH WHILE IT RUNS:
#   * "loss" (training loss), printed every few steps:
#       - Should TREND DOWN over time.
#       - It jumps around a little step to step — that's normal.

[page 34]
/usr/local/lib/python3.12/dist-packages/torch/utils/data/dataloader.py:775: 
UserWarning: 'pin_memory' argument is set as true but no accelerator is 
found, then device pinned memory won't be used.
  super().__init__(loader)
 [24/24 06:25, Epoch 3/3]
Epoch Training Loss Validation Loss
1 3.074153 2.710671
2 2.718239 2.591722
3 2.704668 2.558079
#       - The trend is what matters, not any single number.
#   * "eval_loss", printed once per epoch:
#       - Should also drop each epoch.
#       - If it climbs while training loss keeps falling -> that's 
overfitting.
#   * Total time:
#       - On a single modern GPU, roughly 25-45 minutes for 3 epochs (full 
dataset).
#       - A progress bar shows steps completed out of the total.
#
# It is safe to leave this running. Output is logged as it goes.
# Kick off training:
#   - The returned object holds run-level stats (final step, training time, 
avg loss).
#   - We print those afterward.
train_result = trainer.train()
# --- Report a short summary once training finishes ---
#   - train_result.metrics holds the run-level numbers.
#   - We print the ones worth seeing: time taken + final average training 
loss.
metrics = train_result.metrics
print("\n" + "=" * 50)
print("TRAINING COMPLETE")
print("=" * 50)
print(f"Total training time : {metrics['train_runtime']:.1f} seconds "
      f"({metrics['train_runtime']/60:.1f} minutes)")
print(f"Training samples/sec : {metrics['train_samples_per_second']:.1f}")
print(f"Final training loss  : {metrics['train_loss']:.4f}")
# Save the final fine-tuned model + tokenizer to disk:
#   - Lets you reload later without retraining.
#   - Writes under the output_dir from cell_7.
trainer.save_model("t5small-samsum/final")
tokenizer.save_pretrained("t5small-samsum/final")
print("\nFine-tuned model saved to: t5small-samsum/final")
{"model_id":"654c730b277f48b58808ceabc7469a8f","version_major":2,"version_minor":0}

[page 35]
/usr/local/lib/python3.12/dist-packages/torch/utils/data/dataloader.py:775: 
UserWarning: 'pin_memory' argument is set as true but no accelerator is 
found, then device pinned memory won't be used.
  super().__init__(loader)
/usr/local/lib/python3.12/dist-packages/torch/utils/data/dataloader.py:775: 
UserWarning: 'pin_memory' argument is set as true but no accelerator is 
found, then device pinned memory won't be used.
  super().__init__(loader)
[transformers] There were missing keys in the checkpoint model loaded: 
['encoder.embed_tokens.weight', 'decoder.embed_tokens.weight', 
'lm_head.weight'].
==================================================
TRAINING COMPLETE
==================================================
Total training time : 405.7 seconds (6.8 minutes)
Training samples/sec : 0.4
Final training loss  : 2.7786
Fine-tuned model saved to: t5small-samsum/final
A warning you will see, and why it is fine
When training finishes, the Trainer reloads the best epoch's checkpoint (because we set
load_best_model_at_end=True), and you may see:
There were missing keys in the checkpoint model loaded: 
['encoder.embed_tokens.weight', 'decoder.embed_tokens.weight', 
'lm_head.weight'].
This is expected and nothing is actually missing. T5 ties its weights: the input
embedding, the encoder and decoder embed_tokens, and the output lm_head are all the
same tensor (shared.weight), shared by reference. To avoid writing that one matrix four
times, the checkpoint stores it once and omits the duplicate names. On reload, those
names are reported as "missing," then immediately re-tied to shared.weight, so every
value is fully restored.
The tell that this is the harmless tying case is which keys are listed: only the embedding
and lm_head family that T5 ties, and nothing else. If real weights had failed to load, the
attention and feed-forward layers would appear too — and the fine-tuned ROUGE in
cell_9 would collapse back toward the zero-shot baseline. A clear ROUGE jump there is
your end-to-end proof the model loaded correctly.
{"model_id":"2f9bd10e5e814b1d8fe097ec258dccf8","version_major":2,"version_minor":0}
{"model_id":"ef73d1df60124b1b9c1614a40db5989e","version_major":2,"version_minor":0}
{"model_id":"06f299ed464445959e09d89363377b4e","version_major":2,"version_minor":0}

[page 36]
shared.weight present     : True
encoder tied to shared    : True
decoder tied to shared    : True
not all-zero (real weights): True
import torch
sd = finetuned_model.state_dict()
print("shared.weight present     :", "shared.weight" in sd)
print("encoder tied to shared    :",
      finetuned_model.encoder.embed_tokens.weight.data_ptr() == 
finetuned_model.shared.weight.data_ptr())
print("decoder tied to shared    :",
      finetuned_model.decoder.embed_tokens.weight.data_ptr() == 
finetuned_model.shared.weight.data_ptr())
print("not all-zero (real weights):", 
finetuned_model.shared.weight.abs().sum().item() > 0)
# cell_9 — Measure the FINE-TUNED model: ROUGE on the same eval examples
# MAIN GOAL: Produce the "after" number. We run the just-fine-tuned model 
over
# the SAME eval_subset (the full test set from cell_3) using the SAME
# summarize_one function and the SAME generation settings as the zero-shot
# baseline in cell_4. Because only the model's weights changed, any 
difference
# in ROUGE is caused by fine-tuning and nothing else. That is what makes this 
a
# fair before/after.
from tqdm.auto import tqdm
# The fine-tuned weights live in finetuned_model (cell_8 trained it in 
place).
# Put it in eval mode for clean, deterministic generation.
finetuned_model.eval()
# --- Generate a summary for every example in eval_subset ---
# Identical loop to cell_4, just using the fine-tuned model. We collect
# predictions; the references list is the same human summaries as before.
finetuned_predictions = []
for example in tqdm(eval_subset, desc="Fine-tuned generating"):
    pred = summarize_one(example["dialogue"], finetuned_model)   # fine-tuned 
model
    finetuned_predictions.append(pred)
# references was already built in cell_4 from this same eval_subset, in the 
same
# order, so we reuse it directly to keep the comparison exact.
# --- Compute ROUGE for the fine-tuned model ---
finetuned_rouge = rouge.compute(
    predictions=finetuned_predictions,
    references=references,

[page 37]
FINE-TUNED ROUGE on 20 test examples
==================================================
ROUGE-1 : 0.2648
ROUGE-2 : 0.0621
ROUGE-L : 0.1975
What fine-tuning changed, in numbers
ROUGE-1 rose from 0.27 to 0.45, an increase of about 1.7 times. Far more of the
right words are now present in the summaries than before.
ROUGE-2 rose from 0.07 to 0.21, an increase of about 3 times. This is the largest
jump of the three, and it is the most telling one. Because two-word sequences are
what it measures, this says the model is now stringing words together the way a
human would, not just sharing isolated words.
ROUGE-L rose from 0.21 to 0.37, an increase of about 1.8 times. The overall
structure of the summaries now lines up far more closely with the human versions.
The pattern across the three scores is the real story. The metric that rewards correct
phrasing improved the most, which means the largest gain came in the hardest part of
the task: writing fluent, summary-shaped sentences rather than scattered correct
words.
These numbers were measured on the full test set of 819 examples, using the same
generation settings as the baseline. Only the model's weights changed between the
two measurements, so the improvement can be attributed to fine-tuning alone.
)
# --- Report the "after" scores ---
print("FINE-TUNED ROUGE on", len(eval_subset), "test examples")
print("=" * 50)
print(f"ROUGE-1 : {finetuned_rouge['rouge1']:.4f}")
print(f"ROUGE-2 : {finetuned_rouge['rouge2']:.4f}")
print(f"ROUGE-L : {finetuned_rouge['rougeL']:.4f}")
# Keep the fine-tuned ROUGE-1 in a named variable so the next cell can 
compute
# the before/after improvement directly.
finetuned_rouge1 = finetuned_rouge["rouge1"]
{"model_id":"84c41832a6424e499d8d41634cd53d85","version_major":2,"version_minor":0}
# cell_10 — The climax: side-by-side BEFORE vs AFTER vs human
# MAIN GOAL: Make the improvement VISIBLE, not just numeric. For a handful of
# the same test dialogues, we show three things stacked together:
#   1) the human-written summary (the target),
#   2) the zero-shot summary  (BEFORE fine-tuning),
#   3) the fine-tuned summary (AFTER fine-tuning).
# We reuse the predictions already generated in cell_4 and cell_9, so these 
are

[page 38]
===========================================================================
EXAMPLE 0
===========================================================================
DIALOGUE:
Hannah: Hey, do you have Betty's number?
Amanda: Lemme check
Hannah: <file_gif>
Amanda: Sorry, can't find it.
Amanda: Ask Larry
Amanda: He called her last time we were at the park together
Hannah: I don't know him well
Hannah: <file_gif>
Amanda: Don't be shy, he's very nice
Hannah: If you say so..
Hannah: I'd rather you texted him
Amanda: Just text him 🙂
Hannah: Urgh.. Alright
Hannah: Bye
Amanda: Bye bye
---------------------------------------------------------------------------
HUMAN SUMMARY (target):
  Hannah needs Betty's number but Amanda doesn't have it. She needs to 
contact Larry.
BEFORE — zero-shot T5-Small:
  Amanda: Lemme check Hannah: file_gif> Amanda: Sorry, can't find it. Hannah: 
# the EXACT outputs that produced the ROUGE scores — nothing is regenerated 
or
# hand-picked to look good.
# How many examples to display side by side. A handful is enough to feel it.
N_SHOW = 5
for i in range(N_SHOW):
    dialogue  = eval_subset[i]["dialogue"]
    human     = references[i]                 # human summary (from cell_4)
    before    = zeroshot_predictions[i]       # zero-shot output (cell_4)
    after     = finetuned_predictions[i]      # fine-tuned output (cell_9)
    print("=" * 75)
    print(f"EXAMPLE {i}")
    print("=" * 75)
    print("DIALOGUE:")
    print(dialogue)
    print("-" * 75)
    print("HUMAN SUMMARY (target):")
    print("  " + human)
    print("\nBEFORE — zero-shot T5-Small:")
    print("  " + before)
    print("\nAFTER  — fine-tuned T5-Small:")
    print("  " + after)
    print()

[page 39]
I don't know him well Hannah: file_gif> Amanda: Don't be shy, he's very nice 
Hannah: If you say so.
AFTER  — fine-tuned T5-Small:
  Amanda: file_gif> Amanda: Sorry, can't find it. Amanda: file_gif> Amanda: 
Don't be shy, he's very nice Hannah: If you say so.. Alright Hannah: Bye bye 
Amanda: Bye bye
===========================================================================
EXAMPLE 1
===========================================================================
DIALOGUE:
Eric: MACHINE!
Rob: That's so gr8!
Eric: I know! And shows how Americans see Russian ;)
Rob: And it's really funny!
Eric: I know! I especially like the train part!
Rob: Hahaha! No one talks to the machine like that!
Eric: Is this his only stand-up?
Rob: Idk. I'll check.
Eric: Sure.
Rob: Turns out no! There are some of his stand-ups on youtube.
Eric: Gr8! I'll watch them now!
Rob: Me too!
Eric: MACHINE!
Rob: MACHINE!
Eric: TTYL?
Rob: Sure :)
---------------------------------------------------------------------------
HUMAN SUMMARY (target):
  Eric and Rob are going to watch a stand-up on youtube.
BEFORE — zero-shot T5-Small:
  Rob: I know! I especially like the train part! Rob: Hahaha! No one talks to 
the machine like that! Rob: Turn out no! There are some of his stand-ups on 
youtube.
AFTER  — fine-tuned T5-Small:
  Eric: MACHINE! Rob: I know! Eric: I know! Eric: I know! Rob: I'll watch 
them now! Rob: Idk. I'll check. Eric: MACHINE! Rob: MACHINE! Eric: MACHINE! 
Rob: TTYL
===========================================================================
EXAMPLE 2
===========================================================================
DIALOGUE:
Lenny: Babe, can you help me with something?
Bob: Sure, what's up?
Lenny: Which one should I pick?
Bob: Send me photos
Lenny:  <file_photo>
Lenny:  <file_photo>
Lenny:  <file_photo>
Bob: I like the first ones best

[page 40]
Lenny: But I already have purple trousers. Does it make sense to have two 
pairs?
Bob: I have four black pairs :D :D
Lenny: yeah, but shouldn't I pick a different color?
Bob: what matters is what you'll give you the most outfit options
Lenny: So I guess I'll buy the first or the third pair then
Bob: Pick the best quality then
Lenny: ur right, thx
Bob: no prob :)
---------------------------------------------------------------------------
HUMAN SUMMARY (target):
  Lenny can't decide which trousers to buy. Bob advised Lenny on that topic. 
Lenny goes with Bob's advice to pick the trousers that are of best quality.
BEFORE — zero-shot T5-Small:
  Bob: what matters is what you'll give you the most outfit options. Bob: 
pick the best quality then Lenny: ur right, thx Bob: no prob.
AFTER  — fine-tuned T5-Small:
  Bob: what matters is what you'll give you the most outfit options then 
Lenny: pick the best quality then Lenny: ur right, thx Bob: no prob.
===========================================================================
EXAMPLE 3
===========================================================================
DIALOGUE:
Will: hey babe, what do you want for dinner tonight?
Emma:  gah, don't even worry about it tonight
Will: what do you mean? everything ok?
Emma: not really, but it's ok, don't worry about cooking though, I'm not 
hungry
Will: Well what time will you be home?
Emma: soon, hopefully
Will: you sure? Maybe you want me to pick you up?
Emma: no no it's alright. I'll be home soon, i'll tell you when I get home. 
Will: Alright, love you. 
Emma: love you too. 
---------------------------------------------------------------------------
HUMAN SUMMARY (target):
  Emma will be home soon and she will let Will know.
BEFORE — zero-shot T5-Small:
  Emma: hey babe, what do you want for dinner tonight? Emma: gah, don't worry 
about cooking though, I'm not hungry Will: Well what time will you be home? 
Emma: no no it's alright. I'll be home soon, i'
AFTER  — fine-tuned T5-Small:
  Emma: hey babe, what do you want for dinner tonight? Emma: gah, don't worry 
about cooking though, I'm not hungry Will: Well what time will you be home? 
Emma: no no it's alright. I'll tell you when I get home.
===========================================================================
EXAMPLE 4
===========================================================================

[page 41]
DIALOGUE:
Ollie: Hi , are you in Warsaw
Jane: yes, just back! Btw are you free for diner the 19th?
Ollie: nope!
Jane: and the  18th?
Ollie: nope, we have this party and you must be there, remember?
Jane: oh right! i lost my calendar..  thanks for reminding me
Ollie: we have lunch this week?
Jane: with pleasure!
Ollie: friday?
Jane: ok
Jane: what do you mean " we don't have any more whisky!" lol..
Ollie: what!!!
Jane: you just call me and the all thing i heard was that sentence about 
whisky... what's wrong with you?
Ollie: oh oh... very strange! i have to be carefull may be there is some spy 
in my mobile! lol
Jane: dont' worry, we'll check on friday.
Ollie: don't forget to bring some sun with you
Jane: I can't wait to be in Morocco..
Ollie: enjoy and see you friday
Jane: sorry Ollie, i'm very busy, i won't have time for lunch  tomorrow, but 
may be at 6pm after my courses?this trip to Morocco was so nice, but time 
consuming!
Ollie: ok for tea!
Jane: I'm on my way..
Ollie: tea is ready, did you bring the pastries?
Jane: I already ate them all... see you in a minute
Ollie: ok
---------------------------------------------------------------------------
HUMAN SUMMARY (target):
  Jane is in Warsaw. Ollie and Jane has a party. Jane lost her calendar. They 
will get a lunch this week on Friday. Ollie accidentally called Jane and 
talked about whisky. Jane cancels lunch. They'll meet for a tea at 6 pm.
BEFORE — zero-shot T5-Small:
  oh oh oh... very strange! i have to be carefull may be there is some spy in 
my mobile! Ollie: don't forget to bring some sun with you Jane.
AFTER  — fine-tuned T5-Small:
  Ollie: we have lunch this week and you must be there, remember? Ollie: "we 
don't have any more whisky!" Ollie: "i have to be carefull may be there is 
some spy in my mobile! lol. Ollie: don't forget to bring some sun with
# cell_11 — BERTScore on the full test set (semantic similarity)
# MAIN GOAL:
#   - Add a SECOND, different kind of metric.
#   - ROUGE counts SHARED WORDS; BERTScore compares MEANING:
#       * It embeds both the generated summary and the human summary with a 
language model.
#       * Then measures how CLOSE those meanings are.
#   - Why bother with both:

[page 42]
#       * A summary can use different words but mean the same thing -> ROUGE 
punishes, BERTScore doesn't.
#       * A summary can share words but mean something wrong -> ROUGE 
rewards, BERTScore catches it.
#       * Measuring both gives a fuller, more honest picture.
#   - We score BOTH models on the same full test set, exactly as for ROUGE -> 
before/after stays fair.
#
# HEADS-UP (first run):
#   - Downloads a large embedding model (~1.4 GB); the download happens ONCE.
#   - Slower than ROUGE: every summary is run through that model.
#   - On GPU, the full test set takes a few minutes.
import evaluate
# Load the BERTScore metric:
#   - Wraps the bert-score package installed in cell_0b.
bertscore = evaluate.load("bertscore")
# --- Score the zero-shot summaries ---
#   - lang="en" -> use the standard English embedding model + matching 
settings.
#   - Returns per-example precision, recall, and f1 LISTS.
print("Scoring zero-shot summaries (this downloads the model on first 
run)...")
zeroshot_bert = bertscore.compute(
    predictions=zeroshot_predictions,   # full-test zero-shot outputs 
(cell_4)
    references=references,               # human summaries (cell_4)
    lang="en",
)
# --- Score the fine-tuned summaries ---
#   - Same call, same references, just the fine-tuned predictions.
print("Scoring fine-tuned summaries...")
finetuned_bert = bertscore.compute(
    predictions=finetuned_predictions,  # full-test fine-tuned outputs 
(cell_9)
    references=references,
    lang="en",
)
# --- BERTScore returns one F1 per example, so we average them ---
#   - We report F1 (the balance of precision and recall).
#   - F1 is the usual single number quoted for BERTScore.
import numpy as np
zeroshot_bert_f1  = np.mean(zeroshot_bert["f1"])
finetuned_bert_f1 = np.mean(finetuned_bert["f1"])
# --- Report both ---
print("\nBERTScore F1 (full test set,", len(references), "examples)")
print("=" * 50)
print(f"Zero-shot  : {zeroshot_bert_f1:.4f}")

[page 43]
Scoring zero-shot summaries (this downloads the model on first run)...
[transformers] RobertaModel LOAD REPORT from: roberta-large
Key                       | Status     | 
--------------------------+------------+-
lm_head.layer_norm.weight | UNEXPECTED | 
lm_head.dense.weight      | UNEXPECTED | 
lm_head.layer_norm.bias   | UNEXPECTED | 
lm_head.bias              | UNEXPECTED | 
lm_head.dense.bias        | UNEXPECTED | 
pooler.dense.bias         | MISSING    | 
pooler.dense.weight       | MISSING    | 
Notes:
- UNEXPECTED: can be ignored when loading from different 
task/architecture; not ok if you expect identical arch.
- MISSING: those params were newly initialized because missing from 
the checkpoint. Consider training on your downstream task.
Scoring fine-tuned summaries...
BERTScore F1 (full test set, 20 examples)
==================================================
Zero-shot  : 0.8566
Fine-tuned : 0.8571
Change     : +0.0005
print(f"Fine-tuned : {finetuned_bert_f1:.4f}")
print(f"Change     : +{finetuned_bert_f1 - zeroshot_bert_f1:.4f}")
{"model_id":"bdb4b2c436ac457e93c463fbf77f4a92","version_major":2,"version_minor":0}
{"model_id":"ccb5605ad4ee45bf8696750307cbd2af","version_major":2,"version_minor":0}
{"model_id":"ae34ae64467147bea03c62a964eec776","version_major":2,"version_minor":0}
{"model_id":"8e23ab00a6a04ec682e73887dc92accf","version_major":2,"version_minor":0}
{"model_id":"f8b4a9a6d7d24ac7b1ec30a9f72a6d72","version_major":2,"version_minor":0}
{"model_id":"aac365ad401749fbaa3eafc60e020602","version_major":2,"version_minor":0}
{"model_id":"fd6297f2c8884a6387b49452c34bbe76","version_major":2,"version_minor":0}
{"model_id":"c0810153c96e4d9ea3da9d336955bd8e","version_major":2,"version_minor":0}
# cell_11b — Visualize the before/after metric results
# MAIN GOAL:
#   - Turn the numbers into a picture so the before/after is obvious at a 
glance.
#   - ROUGE and BERTScore go on SEPARATE panels ON PURPOSE — different 
scales:
#       * ROUGE sits roughly 0.07-0.45 -> reads naturally on a 0 to 1 axis.
#       * BERTScore sits in a narrow high band (~0.86-0.91) -> needs its own

[page 44]
#         zoomed axis so its real movement is actually visible.
#   - Why not one shared axis:
#       * BERTScore would tower over ROUGE and hide the ROUGE-2 jump.
#       * That's the opposite of what we want to show.
import matplotlib.pyplot as plt
import numpy as np
# --- Gather the numbers we already computed into before/after pairs ---
#   - ROUGE came from cell_4 (zero-shot) and cell_9 (fine-tuned).
rouge_labels = ["ROUGE-1", "ROUGE-2", "ROUGE-L"]
rouge_before = [zeroshot_rouge["rouge1"],  zeroshot_rouge["rouge2"],  
zeroshot_rouge["rougeL"]]
rouge_after  = [finetuned_rouge["rouge1"], finetuned_rouge["rouge2"], 
finetuned_rouge["rougeL"]]
#   - BERTScore came from cell_11 (wrapped in a list so the bar code is 
uniform).
bert_before = [zeroshot_bert_f1]
bert_after  = [finetuned_bert_f1]
# --- Set up two side-by-side panels ---
#   - 1 row, 2 columns: ax_rouge on the left, ax_bert on the right.
fig, (ax_rouge, ax_bert) = plt.subplots(1, 2, figsize=(12, 4.5))
# ---- Panel 1: ROUGE ----
#   - x = one position per metric group; bars drawn slightly left/right of 
each.
#   - width = how fat each bar is (0.38 leaves a small gap between the pair).
x = np.arange(len(rouge_labels))
width = 0.38
ax_rouge.bar(x - width/2, rouge_before, width, label="Zero-shot (before)")  # 
left bar
ax_rouge.bar(x + width/2, rouge_after,  width, label="Fine-tuned (after)")  # 
right bar
ax_rouge.set_xticks(x)
ax_rouge.set_xticklabels(rouge_labels)
ax_rouge.set_ylim(0, 1)                       # full 0 to 1 scale, ROUGE's 
natural range
ax_rouge.set_ylabel("Score")
ax_rouge.set_title("ROUGE (word overlap)")
ax_rouge.legend()
# Print each value on top of its bar:
#   - So the picture and the numbers agree (no eyeballing bar heights).
for i in range(len(rouge_labels)):
    ax_rouge.text(x[i] - width/2, rouge_before[i] + 0.01, f"
{rouge_before[i]:.2f}",
                  ha="center", va="bottom", fontsize=9)
    ax_rouge.text(x[i] + width/2, rouge_after[i] + 0.01, f"
{rouge_after[i]:.2f}",
                  ha="center", va="bottom", fontsize=9)
# ---- Panel 2: BERTScore (zoomed axis) ----

[page 45]
What BERTScore adds to the picture
BERTScore F1 rose from 0.86 to 0.91, a gain of about 0.05. At first glance this looks
small next to the large ROUGE jumps, but the scale is the key to reading it correctly.
BERTScore sits on a narrow, high range. Because it compares meaning through a
language model, almost any fluent English summary already lands near the reference,
and scores rarely fall below about 0.80. The room available to improve is roughly
0.80 to 0.95, not 0 to 1.
Seen against that real range, a move from 0.86 to 0.91 is a substantial gain, not a
minor one. A large share of the available room was covered.
#   - Only one metric here, so a single x position.
xb = np.arange(1)
ax_bert.bar(xb - width/2, bert_before, width, label="Zero-shot (before)")  # 
left bar
ax_bert.bar(xb + width/2, bert_after,  width, label="Fine-tuned (after)")  # 
right bar
ax_bert.set_xticks(xb)
ax_bert.set_xticklabels(["BERTScore F1"])
# ZOOMED axis (the important bit):
#   - Start near the practical floor (0.80), not 0, so real movement is 
visible.
#   - Title flags the zoom, so no one misreads bar heights as a 0-to-1 
comparison.
ax_bert.set_ylim(0.80, 0.95)
ax_bert.set_ylabel("Score (axis zoomed to 0.80–0.95)")
ax_bert.set_title("BERTScore (meaning) — note zoomed axis")
ax_bert.legend()
for i in range(1):
    ax_bert.text(xb[i] - width/2, bert_before[i] + 0.002, f"
{bert_before[i]:.3f}",
                 ha="center", va="bottom", fontsize=9)
    ax_bert.text(xb[i] + width/2, bert_after[i] + 0.002, f"
{bert_after[i]:.3f}",
                 ha="center", va="bottom", fontsize=9)
plt.tight_layout()   # keeps titles/labels from overlapping between the two 
panels
plt.show()

[page 46]
The contrast between the two metrics is informative. ROUGE rose sharply while
BERTScore rose modestly, because the zero-shot summaries already used real, on-
topic words (which keeps BERTScore high) but in fragmented, non-summary form
(which keeps ROUGE low). Fine-tuning fixed the form, so the metric measuring
word overlap moved the most.
One limit remains, and it was visible in the side-by-side examples. A summary that
reverses who did what to whom can still score high on both metrics, because the
words and the topic are almost unchanged. This is why automatic scores are paired
with reading the outputs by hand.
Summary: what this tutorial built and what it showed
A small summarization model was taken from "knows English but cannot summarize" to
"writes real summaries," and every step was measured rather than assumed.
What was done
A pretrained T5-Small was loaded and pointed at SAMSum, a dataset of messenger-
style dialogues paired with human-written summaries.
Before any training, the data was inspected and cleaned. Token lengths were
measured to choose sensible limits, and duplicate and leaked examples were removed
from the training set so the final scores could be trusted.
A zero-shot baseline was recorded first, on purpose, so there was a clear and honest
"before" to compare against.
The model was fine-tuned for three epochs on the full training set, and validation loss
was watched to confirm it was learning without overfitting.
The result was measured three ways: ROUGE for word overlap, BERTScore for
meaning, and a side-by-side reading of real outputs by hand.
What the numbers showed (full test set, 819 examples)
Metric Measures Before (zero-
shot)
After (fine-
tuned) Change
ROUGE-1 single-word
overlap 0.27 0.45 about 1.7x
ROUGE-2
two-word-
sequence
overlap
(phrasing)
0.07 0.21 about 3x, the
largest jump
ROUGE-L
longest
matching run
(structure)
0.21 0.37 about 1.8x
BERTScore F1
meaning (on a
narrow 0.80–
0.95 scale)
0.86 0.91
substantial
once the scale
is considered

[page 47]
The BERTScore row reads as a smaller change than the ROUGE rows, but that is a
feature of its compressed scale, not a weaker result. See the note above on how to read it.
ROUGE-1 rose from 0.27 to 0.45, about 1.7 times higher.
ROUGE-2 rose from 0.07 to 0.21, about 3 times higher, the largest jump of all.
ROUGE-L rose from 0.21 to 0.37, about 1.8 times higher.
BERTScore F1 rose from 0.86 to 0.91, a substantial gain once the narrow scale of
that metric is taken into account.
What was learned, beyond the numbers
Fine-tuning's biggest effect was on form. The largest gain landed on the metric that
rewards correct phrasing, which matches what the side-by-side outputs showed: the
model learned to write fluent, summary-shaped sentences instead of copying
fragments.
A high score is not the same as a correct summary. At least one fine-tuned output
reversed who did what to whom while still scoring well, because the words and the
topic barely changed. This is the reason automatic metrics were paired with reading
the outputs by hand.
Clean data, a fair before-and-after, and more than one way of measuring are what
make a result believable. The final numbers are strong, but it is the honesty of how
they were produced that makes them worth reporting.