# word2Vec vx3
course: Module 3 — Deep Learning & NLP
module: Module-3-Deep-Learning-NLP
type: notebook
source_url: https://personal-learn.armco.dev/files/Module-3-Deep-Learning-NLP/General/Lab_Materials-_19-04-2026/word2Vec_vx3.ipynb
---
[cell 1 markdown]
# **Sentiment Analysis with Word2Vec on Financial PhraseBank**
The **Financial PhraseBank** dataset contains sentences from financial news, each labeled with a sentiment (`positive`, `negative`, or `neutral`).
This notebook focuses on applying **Word2Vec**, a powerful word embedding technique, to convert text into meaningful numerical vectors. These vectors will then be used for:
1. **Sentiment Classification:** Training supervised models like MLP, Logistic Regression, Random Forest, and Naive Bayes.
2. **Unsupervised Clustering:** Using K-Means to discover natural groupings in the data and see if they align with sentiment.
[cell 2 markdown]
**About Dataset**
The **Financial PhraseBank** dataset contains short sentences extracted from financial news articles related to publicly traded companies. Each sentence is labeled with a sentiment that reflects how the information would likely impact investor perceptions of the company’s stock.
Source: **Kaggle Dataset**
* [Kaggle: Financial Sentiment Analysis](https://www.kaggle.com/datasets/sbhatti/financial-sentiment-analysis)
Target Variable:
* **sentiment** *(categorical)*: Indicates the polarity of the financial news sentence:
* `positive`
* `negative`
* `neutral`
---
Features:
| Column Name | Description | Data Type |
| ----------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------- |
| sentence | A short text segment extracted from financial news. Each sentence is self-contained and typically refers to a specific company, event, or market movement. | *Text* |
| sentiment | Label showing whether the news is expected to have a **positive**, **negative**, or **neutral** impact on the company’s stock or market perception. | *Categorical* |
Reference paper:[Malo et al., Good Debt or Bad Debt, JASIST 2014](https://arxiv.org/abs/1307.5336)
[cell 3 code]
!pip install gensim
[cell 4 markdown]
# **Note: After installing packages gensim, we may need to restart the session so that they link correctly with NumPy/SciPy.**
[cell 5 markdown]
## **Setup and Library Imports**
First, we import the necessary libraries for data handling, text processing, modeling, and visualization. We will use `gensim` for training our Word2Vec model and `scikit-learn` for classification and clustering.
[cell 6 code]
# ============================================================
# cell_0
# PURPOSE: Import all required libraries and configure the
# environment for the entire notebook.
# This is the "toolbox" cell — every tool we'll use throughout
# the project is loaded here once, so later cells can use them
# without re-importing.
# ============================================================
# --- Core Data & Math Libraries ---
import pandas as pd # pandas: for loading and manipulating tabular data (like Excel/CSV files)
# we'll use it to load the Financial PhraseBank dataset as a DataFrame
import numpy as np # numpy: for numerical operations and array/matrix math
# word vectors are stored as numpy arrays, so this is essential
import matplotlib.pyplot as plt # matplotlib: the base plotting library in Python
# used to draw charts, scatter plots, etc.
import seaborn as sns # seaborn: built on top of matplotlib, gives prettier statistical plots
# used here mainly for styling
import re # re: Python's built-in Regular Expressions module
# used for pattern-based text cleaning (e.g., removing punctuation)
# --- NLTK: Natural Language Toolkit ---
import nltk # the main NLTK package for NLP utilities
from nltk.corpus import stopwords # stopwords: a pre-built list of common English words
# (like "the", "is", "and") that carry little meaning
# we'll REMOVE these during text preprocessing
from nltk.tokenize import word_tokenize # word_tokenize: splits a sentence into individual words/tokens
# e.g., "Stocks rose." → ["Stocks", "rose", "."]
# smarter than .split() because it handles punctuation properly
# --- Gensim: Word2Vec Model ---
from gensim.models import Word2Vec # Word2Vec: the main model we'll train to learn word embeddings
# converts words into dense numerical vectors that capture meaning
from sklearn.neural_network import MLPClassifier # MLPClassifier: a Multi-Layer Perceptron (neural network)
# from sklearn — imported here but later REPLACED by a
# custom PyTorch MLP for more control
# --- Scikit-learn: ML Models & Utilities ---
from sklearn.model_selection import train_test_split # splits dataset into training and testing portions
# e.g., 80% train / 20% test
from sklearn.linear_model import LogisticRegression # Logistic Regression classifier — a strong linear baseline
from sklearn.naive_bayes import GaussianNB # Gaussian Naive Bayes — a simple probabilistic classifier
from sklearn.preprocessing import StandardScaler, LabelEncoder
# StandardScaler: normalizes features to have mean=0, std=1 (helps some models converge faster)
# LabelEncoder: converts text labels → numbers
# e.g., "positive"→2, "neutral"→1, "negative"→0
from sklearn.ensemble import RandomForestClassifier # Random Forest: an ensemble of decision trees
# generally robust and handles non-linear patterns
from sklearn.cluster import KMeans # KMeans: unsupervised clustering algorithm
# used in Part 3 to group sentences without labels
from sklearn.decomposition import PCA # PCA (Principal Component Analysis): dimensionality reduction
# reduces 200-dim word vectors → 2D or 3D for visualization
from sklearn.metrics import classification_report, accuracy_score, f1_score
# classification_report : prints precision, recall, F1 per class — gives a full performance breakdown
# accuracy_score : fraction of correctly predicted labels (e.g., 0.82 = 82% correct)
# f1_score : harmonic mean of precision & recall — more reliable than accuracy
# on imbalanced datasets (where one class has many more samples)
# --- Environment Configuration ---
import warnings
warnings.filterwarnings("ignore") # suppresses non-critical warnings so output stays clean and readable
sns.set_style("whitegrid") # sets a clean white background with grid lines for all seaborn plots
plt.rcParams['figure.figsize'] = (10, 6) # sets the default plot size to 10 inches wide × 6 inches tall
# so we don't have to specify figsize in every plot
# --- Download NLTK Resource Files ---
# These are data files NLTK needs but doesn't ship with by default.
# They are downloaded once and cached locally on your machine.
nltk.download('punkt') # punkt: pre-trained tokenizer model used by word_tokenize()
# knows how to split text into sentences and words
nltk.download('punkt_tab') # punkt_tab: updated tokenizer configuration files (needed in newer NLTK versions)
# without this, word_tokenize() may raise a LookupError
nltk.download('wordnet') # wordnet: a large English lexical database
# used by the lemmatizer to find the base/root form of words
# e.g., "running" → "run", "better" → "good"
nltk.download('stopwords') # stopwords: downloads the list of ~180 common English words to filter out
# e.g., "i", "me", "the", "a", "is", "are", "was", "were" ...
[cell 7 markdown]
NLTK stands for Natural Language Toolkit — it’s a Python library for Natural Language Processing (NLP).
| Resource | Function |
| ----------- | ------------------------------------- |
| `punkt` | Sentence & word tokenization |
| `punkt_tab` | Tokenizer configuration support |
| `wordnet` | Lemmatization (dictionary base forms) |
| `stopwords` | Removing frequent common words |
[cell 8 markdown]
# **1. Load the dataset and preprocessing of text**
[cell 9 code]
# ============================================================
# cell_1
# PURPOSE: Load the Financial PhraseBank dataset from a CSV
# file into a pandas DataFrame so we can work with it.
# A DataFrame is like a table/spreadsheet in Python —
# rows are individual sentences, columns are their attributes.
# ============================================================
# pd.read_csv() reads a CSV file from disk and loads it into
# a DataFrame (a 2D table with rows and columns)
#
# r'C:\Users\...\data.csv' — the 'r' before the string means
# "raw string", which tells Python to treat backslashes (\)
# as literal characters and NOT as escape sequences
# (e.g., without 'r', \n would mean "newline" instead of
# being part of the folder path — this is a Windows path issue)
#
# This is a hardcoded ABSOLUTE path — it points to the exact
# location of the file on Rudra's personal computer.
# It will ONLY work on that specific machine.
df = pd.read_csv('data.csv')
# This commented-out line is the PORTABLE alternative —
# it uses a RELATIVE path, meaning it looks for 'data.csv'
# in whatever folder this notebook is currently running from.
# Uncomment this line and comment the above when sharing
# the notebook with others or running on a different machine.
# df = pd.read_csv('data.csv')
# df.head(10) displays the first 10 rows of the DataFrame
# This is a standard "sanity check" — we do this immediately
# after loading data to visually confirm:
# ✓ The file loaded correctly (no errors)
# ✓ The column names look right
# ✓ The data looks as expected (sentences + sentiment labels)
#
# Expected output: a table with 2 columns —
# 'Sentence' : the raw financial news text
# 'Sentiment' : the label — "positive", "negative", or "neutral"
df.head(10)
[cell 10 markdown]
# **Encode labels into numbers**
We encode the categorical sentiment labels (`positive`, `negative`, `neutral`) into numerical format (`2`, `0`, `1`) for model training.
[cell 11 code]
# ============================================================
# cell_2
# PURPOSE: Convert the text sentiment labels into numbers.
# Machine learning models cannot work with raw text labels
# like "positive", "negative", "neutral" — they only
# understand numbers. So we encode:
# "negative" → 0
# "neutral" → 1
# "positive" → 2
# (LabelEncoder assigns numbers alphabetically by default)
# ============================================================
from sklearn.preprocessing import LabelEncoder # imports LabelEncoder — a tool that converts
# text categories into integer numbers
# LabelEncoder() creates a fresh encoder object.
# Think of it as an empty "translator" that hasn't
# learned any mappings yet.
encoder = LabelEncoder()
# This single line does TWO things at once (fit + transform):
#
# STEP 1 — .fit(): the encoder LOOKS at df['Sentiment'] and
# learns all unique labels present in the column:
# {"negative", "neutral", "positive"}
# It sorts them alphabetically and assigns numbers:
# negative→0, neutral→1, positive→2
#
# STEP 2 — .transform(): the encoder APPLIES that mapping to
# every row, converting each text label to its number
#
# The result is stored as a NEW column 'label_encoded' in df.
# The original 'Sentiment' column is kept untouched alongside it.
#
# Before: After:
# Sentiment Sentiment | label_encoded
# ---------- -------------|---------------
# "positive" "positive" | 2
# "neutral" "neutral" | 1
# "negative" "negative" | 0
df['label_encoded'] = encoder.fit_transform(df['Sentiment'])
# --- Show the mapping so we know which number = which label ---
# encoder.classes_ : a numpy array of the unique labels the
# encoder learned, in sorted order
# → ["negative", "neutral", "positive"]
# encoder.transform(encoder.classes_) : applies the encoding
# to the classes themselves, giving their
# assigned numbers → [0, 1, 2]
# zip(...) : pairs each label with its number:
# → [("negative",0), ("neutral",1), ("positive",2)]
# dict(...) : converts those pairs into a readable dictionary:
# → {"negative":0, "neutral":1, "positive":2}
# Expected output:
# Label mapping: {'negative': 0, 'neutral': 1, 'positive': 2}
print("Label mapping:", dict(zip(encoder.classes_, encoder.transform(encoder.classes_))))
# Display the first 10 rows to visually confirm the new
# 'label_encoded' column was added correctly next to 'Sentiment'
# You should now see 3 columns: Sentence | Sentiment | label_encoded
df.head(10)
[cell 12 markdown]
## **Data Preprocessing**
The preprocessing steps include:
1. **Lowercasing:** Convert all text to lowercase.
2. **Tokenization:** Split sentences into individual words (tokens).
- Example: `"Stocks rose by 5%" → ["stocks", "rose", "by", "5"]`
- Methods: `str.split()` (basic) or `nltk.word_tokenize()` (handles punctuation better).
3. **Stopword & Punctuation Removal:** Remove common words (like "the", "is") and non-alphabetic tokens.
4. **Lemmatization:** Reduce words to their base form (`profits` → `profit`).
[cell 13 code]
# ============================================================
# cell_3
# PURPOSE: Clean and preprocess the raw sentence text so it
# can be fed into Word2Vec.
# Word2Vec learns from WORDS, not raw sentences — so we need
# to break each sentence into a clean list of meaningful words.
#
# Raw text is messy: it has punctuation, capital letters,
# filler words ("the", "is"), and word variations ("profits"
# vs "profit"). We fix all of this here.
#
# Input → "The company's Profits have Risen significantly."
# Output → ["company", "profit", "risen", "significantly"]
# ============================================================
import re # re: for pattern-based text cleaning using regex
import nltk # nltk: the Natural Language Toolkit
from nltk.corpus import stopwords # stopwords: pre-built list of common filler words
from nltk.tokenize import word_tokenize # word_tokenize: splits sentence into individual word tokens
from nltk.stem import WordNetLemmatizer # WordNetLemmatizer: reduces words to their base/root form
# stopwords.words('english') returns a LIST of ~180 common
# English words like: ["i", "me", "the", "a", "is", "are" ...]
# We convert it to a SET for O(1) lookup speed —
# checking "if word in set" is much faster than "if word in list"
# especially when checking thousands of words
stop_words = set(stopwords.words('english'))
# Creates a lemmatizer object — this is the tool that will
# reduce words to their dictionary base form:
# "profits" → "profit"
# "running" → "run"
# "better" → "good" (when part of speech is specified)
lemmatizer = WordNetLemmatizer()
# ---- Define the preprocessing function ----
# This function takes ONE raw sentence (a string) as input
# and returns a clean list of meaningful word tokens
def preprocess_text(text):
# --- STEP 1: Lowercasing ---
# Convert the entire sentence to lowercase so that
# "Stock", "stock", "STOCK" are all treated as the same word.
# Without this, Word2Vec would learn separate vectors for each!
#
# NOTE: The two commented-out lines below are alternative
# approaches that ALSO remove non-letter characters using regex:
#
# re.sub(r'[^a-zA-Z\s]', ' ', text)
# → replaces anything that is NOT a letter or space with a space
# → removes numbers too (e.g., "5%" becomes " ")
#
# re.sub(r'[^a-zA-Z0-9\s]', ' ', text)
# → same but KEEPS numbers (e.g., "Q3" stays as "Q3")
#
# The author chose the simpler text.lower() here, meaning
# punctuation and numbers are kept for now and handled
# implicitly by the tokenizer and short-token filter below.
#
# text = re.sub(r'[^a-zA-Z\s]', ' ', str(text)).lower()
# text = re.sub(r'[^a-zA-Z0-9\s]', ' ', str(text)).lower()
text = text.lower()
# --- STEP 2: Tokenization ---
# word_tokenize() splits the lowercased sentence into
# a list of individual word/punctuation tokens.
#
# Example:
# "profit rose by 5%." → ["profit", "rose", "by", "5", "%", "."]
#
# It is smarter than text.split() because:
# text.split() → ["profit", "rose", "by", "5%."] (5%. stays joined)
# word_tokenize() → ["profit", "rose", "by", "5", "%", "."] (splits cleanly)
tokens = word_tokenize(text)
# Alternative: tokens = text.split() ← simpler but misses punctuation splitting
# --- STEP 3: Stopword & Short-Token Removal ---
# This is a LIST COMPREHENSION — a compact Python loop that
# builds a new list by filtering the tokens list.
#
# It keeps a word ONLY IF both conditions are True:
# Condition A: word not in stop_words
# → removes filler words like "the", "is", "by"
# Condition B: len(word) > 2
# → removes very short tokens like ".", ",", "%", "a"
# (punctuation and single/double character noise)
#
# Example:
# tokens = ["profit", "rose", "by", "5", "%", "the", "."]
# filtered_tokens = ["profit", "rose"]
# "by" removed → stopword
# "5" removed → len("5") = 1, not > 2
# "%" removed → len("%") = 1, not > 2
# "the" removed → stopword
# "." removed → len(".") = 1, not > 2
filtered_tokens = [word for word in tokens
if word not in stop_words # remove stopwords
and len(word) > 2] # remove very short tokens
# --- STEP 4: Lemmatization ---
# Another list comprehension — applies lemmatizer.lemmatize()
# to every word in filtered_tokens.
#
# lemmatize() reduces each word to its base dictionary form:
# "profits" → "profit"
# "rising" → "rising" (default POS is noun, so verb forms
# "companies"→ "company" may not always reduce perfectly)
#
# Why this matters for Word2Vec:
# Without lemmatization, "profit" and "profits" would be treated
# as TWO different words → two separate vectors in the model.
# With lemmatization, they collapse into ONE word → one vector.
# This gives the model more training data per word concept.
lemmatized = [lemmatizer.lemmatize(word) for word in filtered_tokens]
# Return the final clean list of tokens for this sentence
# e.g., ["company", "profit", "risen", "significantly"]
return lemmatized
# ---- Apply the function to the entire dataset ----
print("\nPreprocessing text data...")
# df['Sentence'].apply(preprocess_text) calls our function
# on EVERY row of the 'Sentence' column one by one,
# and stores each returned token list as a new column 'tokens'.
#
# Before: After (new column added):
# Sentence tokens
# ---------------------------- ----------------------------
# "Profit rose significantly." → ["profit", "rose", "significantly"]
# "Sales declined in Q3." → ["sale", "declined", "q3"]
df['tokens'] = df['Sentence'].apply(preprocess_text)
# ---- Remove empty token rows ----
# After preprocessing, some sentences might end up with NO tokens
# (e.g., a sentence that was only stopwords or punctuation).
# Word2Vec cannot train on empty lists, so we drop those rows.
#
# df['tokens'].map(len) → computes the length of each token list
# > 0 → keeps only rows where length is at least 1
# df[...] → filters the DataFrame to only those rows
df = df[df['tokens'].map(len) > 0]
# ---- Final check ----
# Print the first few rows showing all 4 key columns side by side:
# Sentence → original raw text
# tokens → cleaned token list (what Word2Vec will train on)
# Sentiment → original text label
# label_encoded → numeric label (what classifiers will train on)
print("Data with tokenized sentences and encoded labels:")
df[['Sentence', 'tokens', 'Sentiment', 'label_encoded']].head()
[cell 14 markdown]
## **Data Splitting**
We split the data into training (80%) and testing (20%) sets.
[cell 15 code]
# ============================================================
# cell_4
# PURPOSE: Separate the dataset into INPUT (X) and TARGET (y),
# then split both into a TRAINING set and a TESTING set.
#
# This is a fundamental step in every ML project:
# - The model LEARNS from the training set
# - The model is EVALUATED on the testing set
#
# The key rule: the model must NEVER see the test data
# during training — otherwise we can't trust the evaluation
# (like giving a student the exam answers beforehand)
# ============================================================
# ---- Step 1: Define Features (X) and Target (y) ----
# X = the INPUT to the model — what it learns patterns FROM
# Here, each row is a list of cleaned tokens for one sentence
# e.g., X[0] = ["profit", "rose", "significantly"]
#
# We call it 'X' by convention (uppercase) in ML
X = df['tokens']
# y = the TARGET the model tries to PREDICT — the correct answer
# Here, each row is the numeric sentiment label for that sentence
# e.g., y[0] = 2 (meaning "positive")
#
# We call it 'y' by convention (lowercase) in ML
y = df['label_encoded']
# At this point:
# X and y are "aligned" — X[i] and y[i] always refer to the
# SAME sentence. The split below will preserve this alignment.
# ---- Step 2: Split into Train and Test Sets ----
# train_test_split() randomly shuffles and divides the data
# into two non-overlapping portions.
#
# It returns FOUR things simultaneously:
# X_train → input tokens for training (80% of data)
# X_test → input tokens for testing (20% of data)
# y_train → correct labels for training (80% of data)
# y_test → correct labels for testing (20% of data)
#
# The 4 arguments explained:
#
# X, y
# → the two things to split (kept aligned — same rows go
# together into train or test)
#
# test_size=0.2
# → reserve 20% of ALL rows for testing, 80% for training
# → if dataset has 5000 rows: 4000 train, 1000 test
#
# random_state=42
# → sets a fixed random seed so the shuffle is the SAME
# every time you run this cell
# → without this, you'd get a different split each run,
# making results impossible to reproduce or compare
# → 42 is just a convention (any integer works)
#
# stratify=y
# → ensures the CLASS DISTRIBUTION is preserved in both
# train and test splits
# → without stratify, by bad luck the test set might end up
# with very few "negative" examples
# → WITH stratify, if the full dataset is:
# 60% neutral, 25% positive, 15% negative
# then BOTH train and test sets will have roughly:
# 60% neutral, 25% positive, 15% negative
# → this is especially important for imbalanced datasets
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
# ---- Step 3: Confirm the split sizes ----
# .shape returns (number_of_rows,) for a pandas Series
# This lets us verify the 80/20 split happened correctly
#
# Expected output (if dataset has ~4845 rows):
# Training set shape: (3876,) ← 80%
# Testing set shape: (969,) ← 20%
print("Training set shape:", X_train.shape)
print("Testing set shape:", X_test.shape)
[cell 16 markdown]
# **2. Word2Vec Embeddings for Classification**
We generate dense vector representations of sentences using Word2Vec embeddings to capture semantic relationships for classification.
[cell 17 markdown]
# **Understanding creation of word2vec embeddings from a Toy example**
We demonstrate the basic process of creating Word2Vec embeddings
through a toy example for easier understanding of its working.
[cell 18 markdown]
## **Generating Skip-Gram Training Pairs**
We aim to create center–context word pairs that will be used for Skip-Gram training.
**Process:**
1. Take a sentence and break it into tokens (words).
2. Select one word as the **center** word.
3. Define a window size `k`. All words within `k` positions to the **left** and **right** of the center word are its **context**.
4. Form training pairs as (**center word**, **context word**) for each context.
Example (window=2):
Sentence: *"apple fruit mango sweet"*
- Center = **mango**
- Context = {apple, fruit, sweet}
- Pairs = (mango, apple), (mango, fruit), (mango, sweet)
[cell 19 code]
# ============================================================
# cell_5
# PURPOSE: Build a TOY Word2Vec example from scratch to
# understand HOW Skip-Gram training pairs are created.
#
# Before using the powerful Gensim library, we manually
# demonstrate the core idea behind Word2Vec's Skip-Gram model:
#
# CORE IDEA: "A word's meaning can be understood by the
# company it keeps" — words appearing near each other
# in text tend to be semantically related.
#
# Skip-Gram trains by asking:
# "Given THIS center word, which words appear around it?"
# It creates (center, context) pairs from a sliding window.
#
# Example with window=1:
# Sentence: ["apple", "fruit", "mango"]
# ↑
# center word
# Context words within window of 1: "apple" (left), "mango" (right)
# Pairs created: ("fruit","apple"), ("fruit","mango")
# ============================================================
# --- Imports (some are duplicated from cell_0 — not harmful) ---
import torch # PyTorch: deep learning framework
import torch.nn as nn # nn: contains building blocks for neural networks (layers, loss functions)
import torch.optim as optim # optim: contains optimizers like SGD, Adam that update model weights
import matplotlib.pyplot as plt # for plotting the t-SNE visualization later
from sklearn.manifold import TSNE # TSNE: dimensionality reduction for visualization (covered in next cell)
import numpy as np # for numerical array operations
# ============================================================
# STEP 1: Define the toy corpus
# A corpus is simply a collection of sentences.
# We use 4 short sentences carefully chosen so that
# semantically related words appear near each other:
# - "apple" and "mango" both appear near "fruit" and "sweet"
# - "dogs" and "cats" both appear near "animals"
# After training, Word2Vec should learn that:
# apple ≈ mango (similar vectors)
# dogs ≈ cats (similar vectors)
# apple ≠ dogs (very different vectors)
# ============================================================
corpus = [
"apple is a fruit",
"mango is a fruit",
"apple and mango are sweet",
"dogs and cats are animals"
]
# ============================================================
# STEP 2: Remove stopwords from each sentence
# "is", "a", "and", "are" carry no semantic meaning —
# removing them means the model only trains on meaningful words
#
# This is a nested LIST COMPREHENSION — read it as:
# "For each sentence s in corpus:
# split s into words,
# keep word w ONLY IF it is not in stopwords"
#
# corpus[0] = "apple is a fruit"
# s.split() = ["apple", "is", "a", "fruit"]
# after filter → ["apple", "fruit"] ("is","a" removed)
# ============================================================
stopwords = {"is", "a", "and", "are"} # a SET of words to remove
# (set for fast lookup)
sentences = [[w for w in s.split() # split sentence into words
if w not in stopwords] # keep only non-stopwords
for s in corpus] # do this for every sentence
# After this step, sentences looks like:
# [
# ["apple", "fruit"], ← from "apple is a fruit"
# ["mango", "fruit"], ← from "mango is a fruit"
# ["apple", "mango", "sweet"], ← from "apple and mango are sweet"
# ["dogs", "cats", "animals"] ← from "dogs and cats are animals"
# ]
# ============================================================
# STEP 3: Build the vocabulary and index mappings
#
# A vocabulary is the set of ALL unique words across all sentences.
# We need to map each word ↔ a unique integer index because
# neural networks work with numbers, not strings.
# ============================================================
# Collect every unique word from all sentences into a set
# (sets automatically remove duplicates)
# Expected vocab: {"apple","fruit","mango","sweet","dogs","cats","animals"}
vocab = set(word for sent in sentences # iterate over each sentence
for word in sent) # iterate over each word in that sentence
# word2idx: dictionary mapping word → index number
# enumerate(vocab) gives (0,"apple"), (1,"fruit"), etc.
# {w:i ...} flips that to {"apple":0, "fruit":1, ...}
# NOTE: sets are unordered, so the index assignment is arbitrary
# but consistent within one run
word2idx = {w: i for i, w in enumerate(vocab)}
# idx2word: the REVERSE mapping index → word
# This is used later to convert predictions back to readable words
# {0:"apple", 1:"fruit", ...}
idx2word = {i: w for w, i in word2idx.items()}
# V = vocabulary size = total number of unique words
# This will be used to define the size of our neural network layers
V = len(vocab) # should be 7: apple,fruit,mango,sweet,dogs,cats,animals
# ============================================================
# STEP 4: Generate Skip-Gram (center, context) training pairs
#
# This is the HEART of how Word2Vec's Skip-Gram works.
#
# For each word in a sentence (the "center" word), we look at
# the words within a window of size k to its left and right
# (the "context" words), and create one pair per context word.
#
# Visual example with window_size=1:
#
# Sentence: ["apple", "mango", "sweet"]
# index 0 index 1 index 2
#
# i=0, center="apple":
# j range: -1 to 1 → valid j: 1 (j=-1 is out of bounds)
# pair: (apple, mango)
#
# i=1, center="mango":
# j range: 0 to 2 → valid j: 0 and 2
# pairs: (mango, apple), (mango, sweet)
#
# i=2, center="sweet":
# j range: 1 to 3 → valid j: 1 (j=3 is out of bounds)
# pair: (sweet, mango)
# ============================================================
def generate_pairs(window_size):
pairs = [] # will store all (center_idx, context_idx) pairs
for sent in sentences: # loop over each cleaned sentence
for i, word in enumerate(sent): # loop over each word with its position index i
# Convert the center word to its integer index
center = word2idx[word]
# j loops over all positions within the window:
# from (i - window_size) to (i + window_size) inclusive
# e.g., i=1, window=1 → j goes through [0, 1, 2]
for j in range(i - window_size, i + window_size + 1):
# j != i → skip the center word itself (don't pair a word with itself)
# 0 <= j → skip positions before the sentence starts
# j < len(sent) → skip positions after the sentence ends
if j != i and 0 <= j < len(sent):
# Convert the context word at position j to its index
context = word2idx[sent[j]]
# Store the pair as (center_index, context_index)
pairs.append((center, context))
return pairs # list of integer index tuples, e.g., [(3, 1), (3, 5), ...]
# Generate pairs for two different window sizes to compare them
pairs_w1 = generate_pairs(1) # window=1: only immediate neighbors
pairs_w2 = generate_pairs(2) # window=2: looks 2 words left AND right
# ============================================================
# STEP 5: Print pairs in human-readable (word) form
#
# The pairs are stored as index tuples internally,
# but we convert back to words using idx2word for inspection.
#
# This list comprehension:
# [(idx2word[t], idx2word[c]) for t,c in pairs_w1]
# takes each (center_idx, context_idx) tuple
# and converts both back to their word strings
# ============================================================
# Note: variable names t and c here stand for the
# first and second elements of each pair in pairs_w1/w2
# (they represent center and context word indices)
print("Skip-gram pairs with window_size=1 (words):")
print([(idx2word[t], idx2word[c]) for t, c in pairs_w1])
print("\nSkip-gram pairs with window_size=2 (words):")
print([(idx2word[t], idx2word[c]) for t, c in pairs_w2])
# Expected insight from the output:
# window=1 creates FEWER pairs — only immediate neighbors
# window=2 creates MORE pairs — captures wider context
# e.g., with window=2, "apple" gets paired with "sweet" too,
# not just "mango" — giving the model richer co-occurrence info
[cell 20 markdown]
## **Preparing One-Hot Encoded Inputs and Targets**
[cell 21 code]
# ============================================================
# cell_6
# PURPOSE: Convert the (center, context) word index pairs
# from cell_5 into actual TENSORS that PyTorch can train on.
#
# Neural networks cannot take a word index like "3" directly
# as input — they need a VECTOR. We use ONE-HOT ENCODING
# to represent each center word as a vector.
#
# What is a one-hot vector?
# If vocabulary size V=7 and word "apple" has index 2:
#
# index: 0 1 2 3 4 5 6
# vector:[0, 0, 1, 0, 0, 0, 0]
# ↑
# only position 2 is "hot" (=1), everything else is 0
#
# So X_train = one-hot vectors (the INPUT to the network)
# y_train = integer indices (the TARGET the network predicts)
# ============================================================
# ============================================================
# STEP 1: Define the one_hot() helper function
#
# Takes a word's integer index and vocabulary size,
# returns a 1D tensor of length V with a single 1
# at the position of that word's index.
# ============================================================
def one_hot(idx, size):
# torch.zeros(size) creates a 1D tensor of all zeros
# with length = size (= vocabulary size V)
# e.g., if V=7: tensor([0., 0., 0., 0., 0., 0., 0.])
v = torch.zeros(size)
# Set the position corresponding to this word's index to 1
# int(idx) ensures the index is a plain Python integer
# (not a float or tensor) since PyTorch requires integer indexing
# e.g., idx=2 → tensor([0., 0., 1., 0., 0., 0., 0.])
v[int(idx)] = 1
return v # a tensor of shape (V,) = (7,) in our toy example
print(f'Vocab size: {V}') # confirms V=7 for our toy corpus
# ============================================================
# STEP 2: Build X_train — the INPUT tensor
#
# Recall from cell_5: pairs_w2 is a list of (center, context)
# index tuples, e.g., [(3,1), (3,5), (1,3), ...]
#
# For X_train we need ONE-HOT vectors of the CENTER words.
# Each center word index → one-hot vector of size V
#
# The list comprehension:
# [one_hot(c, V) for c,t in pairs_w2]
# → unpacks each pair as (c=center, t=context)
# → converts each center index c into a one-hot vector
# → produces a LIST of tensors, each of shape (V,)
#
# torch.stack([...]) takes that list of 1D tensors and
# STACKS them into a single 2D tensor (a matrix):
# shape = (num_pairs, V)
# e.g., if we have 24 pairs and V=7: shape = (24, 7)
#
# Each ROW is one training example (one center word)
# Each COLUMN corresponds to one vocabulary word
# ============================================================
X_train = torch.stack([one_hot(c, V) for c, t in pairs_w2])
# X_train visually (V=7, 3 pairs shown):
#
# pair 0: center="apple"(idx=2) → [0, 0, 1, 0, 0, 0, 0]
# pair 1: center="apple"(idx=2) → [0, 0, 1, 0, 0, 0, 0]
# pair 2: center="mango"(idx=5) → [0, 0, 0, 0, 0, 1, 0]
# ...
# shape: (num_pairs, 7)
# ============================================================
# STEP 3: Build y_train — the TARGET tensor
#
# For y_train we need the CONTEXT word indices (not one-hot).
# These are the "correct answers" the model must learn to predict.
#
# The list comprehension:
# [t for c,t in pairs_w2]
# → unpacks each pair as (c=center, t=context)
# → collects only the context index t
# → produces a plain Python list of integers
#
# torch.tensor([...], dtype=torch.long) converts that list
# into a 1D PyTorch tensor of integer type.
# dtype=torch.long is REQUIRED because:
# → CrossEntropyLoss (used in training) expects integer
# class indices, not floats
# → torch.long = 64-bit integer
#
# shape = (num_pairs,) — one integer per training pair
# ============================================================
y_train = torch.tensor([t for c, t in pairs_w2], dtype=torch.long)
# y_train visually (first few values):
# [1, 4, 1, 2, 0, 4, ...]
# ↑ ↑
# context word indices (integers, not one-hot)
# ---- Print shapes to verify ----
# X_train.shape → (num_pairs, V) e.g., (24, 7)
# y_train.shape → (num_pairs,) e.g., (24,)
# Both must have the SAME first dimension (same number of pairs)
print("\nX_train shape:", X_train.shape)
print("y_train shape:", y_train.shape)
# ============================================================
# STEP 4: Print the first 5 training examples for inspection
#
# We convert both the one-hot input and integer target
# back to human-readable words to verify correctness.
# ============================================================
for i in range(5):
# torch.argmax(X_train[i]) finds the index of the maximum
# value in the one-hot vector — since only one position is 1
# and the rest are 0, argmax gives us the word's index back.
# .item() converts the single-element tensor to a plain Python int
# idx2word[...] then maps that index back to the word string
#
# e.g., X_train[i] = [0,0,1,0,0,0,0]
# argmax = 2
# idx2word[2]= "apple"
center_word = idx2word[torch.argmax(X_train[i]).item()]
# y_train[i] is already an integer index
# .item() converts tensor(1) → plain Python int 1
# idx2word[...] maps it back to the context word string
context_word = idx2word[y_train[i].item()]
# Print the full picture for this pair:
# center word | its one-hot vector | context word | context index
print(f"Pair {i}: center={center_word}, "
f"one-hot={X_train[i].tolist()}, "
f"target(context word)={context_word} "
f"(id={y_train[i].item()})")
# Expected output pattern (indices will vary by run due to set ordering):
# Pair 0: center=apple, one-hot=[0,0,1,0,0,0,0], target=mango (id=5)
# Pair 1: center=apple, one-hot=[0,0,1,0,0,0,0], target=sweet (id=4)
# ...
# Observation: the same center word can appear in multiple pairs
# with different context words — that's exactly what Skip-Gram does!
[cell 22 markdown]
## **Defining, Training, and Visualizing the Word2Vec Model**
[cell 23 code]
# ============================================================
# cell_7
# PURPOSE: Define, train, and visualize our toy Word2Vec model.
#
# This is the CORE of the notebook's from-scratch demonstration.
# We build a simple 2-layer neural network that learns to
# predict context words from center words (Skip-Gram).
#
# Through this training, the network's internal weight matrix
# W1 naturally learns to encode word meaning as vectors —
# words that appear in similar contexts end up with
# similar vectors. These vectors ARE the "word embeddings".
#
# Architecture:
#
# one-hot input hidden layer output
# (size V=7) →W1→ (size embed_dim=20) →W2→ (size V=7)
# [0,0,1,0,...] [0.2,-0.5,...] [0.1,0.3,...]
# ↑
# THIS is the word embedding
# (the column of W1 for that word)
# ============================================================
import random
from sklearn.manifold import TSNE
import matplotlib.pyplot as plt
import numpy as np
# ---- Fix all random seeds for reproducibility ----
# Neural networks are initialized with RANDOM weights.
# Setting seeds ensures we get the SAME random numbers
# every run, so results are reproducible and comparable.
#
# We need to set seeds for THREE separate random systems
# because Python, NumPy, and PyTorch each have their own:
SEED = 100
random.seed(SEED) # fixes Python's built-in random module
np.random.seed(SEED) # fixes NumPy's random number generator
torch.manual_seed(SEED) # fixes PyTorch's random number generator
# ============================================================
# STEP 1: Define the Word2Vec Neural Network Architecture
#
# nn.Module is PyTorch's base class for ALL neural networks.
# Every custom network must inherit from it.
#
# Our network has exactly 2 layers (2 weight matrices):
#
# W1: shape (embed_dim, vocab_size) = (20, 7)
# → projects one-hot input INTO the embedding space
# → each COLUMN of W1 = the embedding vector for one word
#
# W2: shape (vocab_size, embed_dim) = (7, 20)
# → projects embedding BACK to vocabulary size
# → produces a score for each word in the vocabulary
# → higher score = more likely to be the context word
# ============================================================
class Wrd2Vec(nn.Module):
def __init__(self, vocab_size, embed_dim):
# super().__init__() calls nn.Module's constructor —
# REQUIRED in every PyTorch model class to properly
# initialize internal PyTorch bookkeeping
super(Wrd2Vec, self).__init__()
# W1: the EMBEDDING LAYER — input → hidden
# nn.Linear(in, out, bias=False) creates a weight matrix
# of shape (out, in) = (embed_dim, vocab_size)
# bias=False: no bias term added — standard for Word2Vec
# This transforms a V-dimensional one-hot vector into
# an embed_dim-dimensional dense embedding vector
self.W1 = nn.Linear(vocab_size, embed_dim, bias=False)
# W1 shape: (20, 7) — maps 7-dim input → 20-dim hidden
# W2: the OUTPUT LAYER — hidden → output
# Transforms the embedding back to vocabulary size
# to produce a score (logit) for each vocabulary word
self.W2 = nn.Linear(embed_dim, vocab_size, bias=False)
# W2 shape: (7, 20) — maps 20-dim hidden → 7-dim output
def forward(self, x):
# forward() defines the COMPUTATION PATH —
# how data flows through the network.
# PyTorch calls this automatically when you do model(input).
# STEP A: x (one-hot input) multiplied by W1
# gives us h, the hidden embedding vector.
# For a one-hot input, this is mathematically equivalent
# to just SELECTING the column of W1 for that word —
# but matrix multiplication lets us process ALL pairs
# in one batch simultaneously.
# x shape: (num_pairs, V=7)
# h shape: (num_pairs, embed_dim=20)
h = self.W1(x) # hidden layer = the word embedding
# STEP B: h multiplied by W2
# produces raw scores (logits) for each vocabulary word.
# The highest score = the word most likely to be context.
# Shape: (num_pairs, V=7) — one score per vocab word per pair
out = self.W2(h) # output = vocabulary score distribution
return out # raw logits (NOT probabilities yet)
# CrossEntropyLoss applies softmax internally
def get_embedding(self, word):
# Utility method to extract the learned embedding
# vector for a specific word AFTER training.
# word2idx[word]: looks up this word's integer index
w_id = word2idx[word]
# W1.weight has shape (embed_dim, vocab_size) = (20, 7)
# Each COLUMN corresponds to one word's embedding.
# [:, w_id] selects ALL rows of column w_id —
# that column IS the embedding vector for this word.
# Shape of result: (embed_dim,) = (20,)
vec = self.W1.weight[:, w_id]
# .detach(): removes the tensor from PyTorch's computation
# graph so we can convert it to numpy safely
# .numpy(): converts PyTorch tensor → NumPy array
# .flatten(): ensures it's a 1D array (safety measure)
return vec.detach().numpy().flatten()
# ============================================================
# STEP 2: Instantiate the model, loss function, and optimizer
# ============================================================
embed_dim = 20 # each word will be represented as a 20-dimensional
# vector. In practice 100-300 dims are typical;
# 20 is fine for this tiny 7-word toy corpus.
# Create the model with vocabulary size V=7 and embed_dim=20.
# This randomly initializes W1 and W2 weight matrices.
model = Wrd2Vec(V, embed_dim)
# CrossEntropyLoss: the loss function for multi-class classification.
# Measures how WRONG the model's predictions are.
# Internally it:
# 1. Applies Softmax to convert raw scores → probabilities
# 2. Computes -log(probability of the correct context word)
# Perfect prediction → loss near 0
# Terrible prediction → loss is high
criterion = nn.CrossEntropyLoss()
# SGD (Stochastic Gradient Descent): the optimizer.
# After each forward pass, uses computed gradients to NUDGE
# weights W1 and W2 in the direction that reduces loss.
# lr=0.05: learning rate — controls how big each nudge is.
# Too high → overshoots, training becomes unstable
# Too low → learns very slowly
# model.parameters() tells the optimizer WHICH tensors to update
optimizer = optim.SGD(model.parameters(), lr=0.05)
# ============================================================
# STEP 3: The Training Loop — 2000 epochs
#
# One EPOCH = one full pass through ALL training pairs.
# We repeat 2000 times so the model gradually improves.
#
# Each epoch follows the same 4-step pattern:
# 1. zero_grad → clear old gradients
# 2. forward → compute predictions
# 3. backward → compute new gradients
# 4. step → update weights
# ============================================================
for epoch in range(2000):
# STEP A: Zero out gradients from the previous epoch.
# PyTorch ACCUMULATES gradients by default — if we don't
# clear them, this epoch's gradients get added to last
# epoch's gradients, corrupting the weight updates.
optimizer.zero_grad()
# STEP B: Forward pass — feed ALL training pairs through
# the network at once to get predicted output scores.
# X_train shape: (num_pairs, V=7)
# outputs shape: (num_pairs, V=7)
outputs = model(X_train)
# STEP C: Compute the loss — how wrong are the predictions?
# criterion compares:
# outputs → what the model predicted (num_pairs, V)
# y_train → what the correct answer was (num_pairs,)
# Returns a single scalar (the average loss over all pairs)
loss = criterion(outputs, y_train)
# STEP D: Backward pass — compute gradients.
# PyTorch automatically calculates how much each weight in
# W1 and W2 contributed to the loss (backpropagation).
# Gradients are stored inside each parameter tensor.
loss.backward()
# STEP E: Update weights using the computed gradients.
# SGD nudges W1 and W2 slightly to reduce the loss:
# weight = weight - lr × gradient
optimizer.step()
# Print loss every 500 epochs to monitor training progress.
# loss.item() converts scalar tensor → plain Python float.
# Loss should DECREASE over time.
if epoch % 500 == 0:
print(f"epoch {epoch}, loss {loss.item():.4f}")
# Expected pattern:
# epoch 0: ~1.95 (random weights, high loss)
# epoch 500: ~1.48 (improving)
# epoch 1000: ~1.12 (still improving)
# epoch 1500: ~0.90 (converging)
# ============================================================
# STEP 4: Extract the learned word embeddings
#
# After training, W1's columns contain the learned embedding
# vectors. We extract one 20-dim vector per word.
# ============================================================
# For every word in vocab, call get_embedding() to retrieve
# its learned 20-dimensional vector from W1's columns.
# np.array([...]) stacks them into a 2D matrix.
# Shape: (V, embed_dim) = (7, 20)
# Each ROW = the embedding vector for one word
embeddings = np.array([model.get_embedding(w) for w in vocab])
# Keep word labels in the same order as the embeddings matrix
# so we always know which row corresponds to which word
labels = list(vocab) # e.g., ["apple", "fruit", "mango", ...]
# ============================================================
# STEP 5: Visualize with t-SNE
#
# Our embeddings are 20-dimensional — impossible to plot.
# t-SNE compresses them to 2D while preserving relative
# distances: words close in 20D stay close in 2D.
#
# PARAMETER CHOICES (carefully tuned for 7 words):
# "How many neighbors should each word pay attention to?"
# perplexity=1 → each word only looks at 1 neighbor
# → words collapse ON TOP of each other ✗
# perplexity=2 → each word looks at ~2 neighbors
# → words spread out nicely within clusters ✓
# Rule: perplexity must always be < total number of points (7 here)
# perplexity=2 : must be < number of points (7).
# Value of 2 makes each word consider ~2
# neighbors → words spread within clusters.
# Original perplexity=1 caused within-cluster
# words to collapse on top of each other.
# n_iter=5000 : more iterations = more stable final layout.
# init="pca" : PCA init is more stable than random.
#
# Expected result:
# "apple" + "mango" → cluster together (both fruits)
# "dogs" + "cats" → cluster together (both animals)
# "fruit" + "sweet" → near the fruit cluster
# "animals" → near the animal cluster
# ============================================================
tsne = TSNE(
n_components=2, # reduce to 2D for plotting
perplexity=2, # was 1 → caused crowding; 2 spreads words #
max_iter=5000, # was 1000 → more iterations = stable layout
init="pca", # stable initialization using PCA
random_state=42 # reproducibility
)
# fit_transform: learns AND applies the 2D layout in one step.
# emb2d shape: (7, 2) — one (x, y) coordinate per word
emb2d = tsne.fit_transform(embeddings)
# ============================================================
# STEP 6: Plot with clear labels and semantic cluster circles
# ============================================================
fig, ax = plt.subplots(figsize=(8, 8))
# --- Per-word label offsets ---
# Manually tuned so no label overlaps its dot or a neighbor.
# Format: "word" → (x_offset_in_points, y_offset_in_points)
offsets = {
"apple" : (-60, 10), # push label left of dot
"mango" : ( 10, 10), # push label right of dot
"fruit" : (-60, 10), # push label left
"sweet" : ( 10, 10), # push label right
"dogs" : ( 10, 10), # push label right
"cats" : (-60, 10), # push label left
"animals": ( 10, -20), # push label right and down
}
# --- Per-word colors by semantic group ---
# Words in the same semantic group share a color so the
# clusters are immediately visually obvious
colors = {
"apple" : "#E07B54", # orange → fruit group
"mango" : "#E07B54", # orange → fruit group
"fruit" : "#7BAE7F", # green → fruit-related
"sweet" : "#7BAE7F", # green → fruit-related
"dogs" : "#5B8DB8", # blue → animal group
"cats" : "#5B8DB8", # blue → animal group
"animals": "#9B77B0", # purple → animal-related
}
for i, label in enumerate(labels):
x, y = emb2d[i] # 2D coordinates for this word
# Plot a dot for this word
# s=120: larger dot size (default ~36 is too small)
# zorder=2: draw dots ON TOP of grid lines
ax.scatter(x, y,
color=colors.get(label, "gray"),
s=120,
zorder=2)
# Place the word label with custom offset from the dot
# ax.annotate() is more powerful than plt.text() because
# it lets us specify offset in screen points (not data units)
# so spacing looks consistent regardless of axis scale
dx, dy = offsets.get(label, (10, 10))
ax.annotate(
label,
xy=(x, y), # the point being labeled
xytext=(dx, dy), # offset of label from the point
textcoords="offset points", # offset unit = screen points
fontsize=12,
fontweight="bold",
color=colors.get(label, "gray") # label color = dot color
)
# --- Draw dashed circles around each semantic cluster ---
# This makes the groupings immediately obvious to the reader.
# Define which words belong to each semantic cluster
clusters = {
"Fruits" : ["apple", "mango", "fruit", "sweet"],
"Animals": ["dogs", "cats", "animals"]
}
# Colors for the cluster boundary circles
cluster_colors = {
"Fruits" : "#E07B54",
"Animals": "#5B8DB8"
}
for cluster_name, cluster_words in clusters.items():
# Get the row indices in emb2d for words in this cluster
indices = [labels.index(w) for w in cluster_words if w in labels]
# Gather the 2D coordinates for all words in this cluster
pts = emb2d[indices]
# Compute the centroid (geometric center) of the cluster
# by averaging x and y coordinates separately
cx, cy = pts.mean(axis=0)
# Compute the radius = furthest distance from center to any
# word in the cluster, plus padding so circle contains all dots
radius = np.max(
np.sqrt((pts[:, 0] - cx)**2 + (pts[:, 1] - cy)**2)
) + 40 # +40 points of padding so words aren't on the edge
# Draw a dashed hollow circle around the cluster
circle = plt.Circle(
(cx, cy),
radius,
color=cluster_colors[cluster_name],
fill=False, # hollow — no fill color
linestyle="--", # dashed border line
linewidth=1.5,
alpha=0.6, # slightly transparent
zorder=1 # draw BEHIND dots (zorder=2)
)
ax.add_patch(circle)
# Add cluster name label just above the circle boundary
ax.text(
cx, cy + radius + 15,
cluster_name,
ha="center", # horizontally centered on cx
fontsize=11,
color=cluster_colors[cluster_name],
fontstyle="italic"
)
ax.set_title("Word Embeddings (2D t-SNE)", fontsize=14, fontweight="bold")
ax.set_xlabel("t-SNE Dimension 1")
ax.set_ylabel("t-SNE Dimension 2")
ax.grid(True, alpha=0.3)
plt.tight_layout()
plt.show()
# ============================================================
# EXPECTED RESULT:
# Two clearly separated clusters with readable labels:
#
# ┌─ ─ ─ Fruits ─ ─ ─ ┐
# • apple • mango
# • fruit • sweet
# └ ─ ─ ─ ─ ─ ─ ─ ─ ─ ┘
#
# ┌─ ─ Animals ─ ─ ┐
# • cats • dogs
# • animals
# └ ─ ─ ─ ─ ─ ─ ─ ┘
#
# The separation between clusters proves the model learned
# that fruits and animals are semantically different —
# using ONLY the co-occurrence patterns in 4 short sentences,
# with no human-defined rules whatsoever.
# ============================================================
[cell 24 markdown]
## **Measuring Word Similarity with Cosine Distance**
[cell 25 code]
# ============================================================
# cell_8
# PURPOSE: Measure how SIMILAR pairs of words are using
# Cosine Similarity on their learned Word2Vec embeddings.
#
# This is the KEY TEST to verify our model learned meaningful
# relationships. After training, words that appeared in
# similar contexts should have similar vectors.
#
# We expect:
# cats – dogs → HIGH similarity (both animals, same context)
# apple – mango → HIGH similarity (both fruits, same context)
# cats – mango → LOW similarity (animal vs fruit, different context)
# apple – dogs → LOW similarity (fruit vs animal, different context)
#
# WHAT IS COSINE SIMILARITY?
# It measures the ANGLE between two vectors, not their length.
# Range: -1 to +1
# +1 → vectors point in SAME direction = very similar
# 0 → vectors are PERPENDICULAR = unrelated
# -1 → vectors point in OPPOSITE direction = opposite meaning
#
# Visual intuition:
# apple → [0.2, 0.8, -0.1, ...] ──┐ small angle
# mango → [0.3, 0.7, -0.2, ...] ──┘ → high similarity
#
# apple → [0.2, 0.8, -0.1, ...] ──┐ large angle
# dogs → [-0.6, 0.1, 0.9, ...] ──┘ → low similarity
# ============================================================
import torch.nn.functional as F # F contains mathematical functions
# including cosine_similarity
# ---- Define the similarity function ----
def word_similarity(w1, w2):
# model.get_embedding(w1) retrieves the learned 20-dim
# numpy array for word w1 from W1's columns (see cell_7)
# torch.tensor(...) converts that numpy array into a
# PyTorch tensor so we can use PyTorch's math functions
# Shape of v1, v2: (20,) — a 1D vector of 20 numbers
v1 = torch.tensor(model.get_embedding(w1))
v2 = torch.tensor(model.get_embedding(w2))
# F.cosine_similarity() expects 2D inputs — at least
# (1, embedding_dim) — but v1 and v2 are 1D: (20,)
# .unsqueeze(0) adds a new dimension at position 0:
# (20,) → (1, 20)
# This makes each vector a "batch of 1" which is what
# cosine_similarity expects as input shape
#
# Formula computed internally:
# v1 · v2
# sim = ─────────────────
# ||v1|| × ||v2||
#
# where · is dot product and ||v|| is the vector's length
#
# .item() converts the resulting single-element tensor
# → a plain Python float for easy printing
return F.cosine_similarity(v1.unsqueeze(0), v2.unsqueeze(0)).item()
# ---- Define the word pairs to compare ----
# These 4 pairs are carefully chosen to test BOTH cases:
# SAME semantic group → should give HIGH score (near +1)
# DIFF semantic group → should give LOW score (near 0 or negative)
pairs_to_check = [
("cats", "dogs"), # both animals → expect HIGH similarity
("cats", "mango"), # animal vs fruit → expect LOW similarity
("apple", "mango"), # both fruits → expect HIGH similarity
("apple", "dogs"), # fruit vs animal → expect LOW similarity
]
# The commented-out pair below was likely used during testing:
# ("mango", "fruit") would also be HIGH since "mango is a fruit"
# appears directly in the corpus
# ("mango", "fruit"),
# ---- Compute and print similarity for each pair ----
print("\nWord similarities (cosine, torch):")
for w1, w2 in pairs_to_check:
# Call our function for each pair and store the float result
sim = word_similarity(w1, w2)
# :.4f formats the float to 4 decimal places
# e.g., 0.8732 or -0.1205
print(f"{w1} – {w2}: {sim:.4f}")
# ============================================================
# EXPECTED OUTPUT PATTERN:
#
# cats – dogs : 0.85+ ← HIGH (both appeared near "animals")
# cats – mango: 0.10- ← LOW (never appeared in same context)
# apple – mango: 0.80+ ← HIGH (both appeared near "fruit","sweet")
# apple – dogs : 0.05- ← LOW (never appeared in same context)
#
# This confirms the model learned semantic relationships
# purely from co-occurrence — no labels, no human rules!
#
# NOTE: Exact values depend on random seed and training.
# The PATTERN (high vs low) matters more than exact numbers.
# ============================================================
[cell 26 markdown]
# **We use `gensim.Word2Vec` to train embeddings on the Financial PhraseBank dataset since it provides an efficient ready-made implementation.**
[cell 27 markdown]
# **Key Parameters in `gensim.Word2Vec`**
[cell 28 markdown]
`Word2Vec(sentences=X_train, vector_size=size, window=5, min_count=1, sg=1, epochs=50, seed=42)`
- **sentences=X_train** → Training data; each element is a tokenized sentence (list of words).
- **vector_size=size** → Dimensionality of word vectors; here it is set from the variable `size`.
- **window=5** → Context window size; model looks at up to 5 words left and right of the target word.
- **min_count=1** → Ignores words with total frequency lower than 1; keeps all words.
- **sg=1** → Training algorithm; 1 = Skip-Gram, 0 = CBOW.
- **epochs=50** → Number of iterations over the corpus during training.
- **seed=42** → Random seed for reproducibility.
[cell 29 markdown]
**CBOW (Continuous Bag of Words)**: Another algorithm to get word2vec </br>
Feed the **entire context window (surrounding words)** as input,
and the model learns to **predict the single target (center) word**.
Sentence:
The cat sits on the mat”
Window size = 2
Target = “sits”
Context = [“The”, “cat”, “on”, “the”]
So one training sample looks like:
| Input (context words) | Output (target word) |
| --------------------- | -------------------- |
| [The, cat, on, the] | sits |
[cell 30 markdown]
# **Training and Exploring Word2Vec Embeddings**
[cell 31 code]
# ============================================================
# cell_9
# PURPOSE: Transition from our toy Word2Vec (cell_7) to the
# REAL Word2Vec using Gensim on the actual Financial PhraseBank
# dataset. Then inspect what the trained model has learned.
#
# WHY GENSIM instead of our hand-built PyTorch model?
# - Gensim's Word2Vec is heavily optimized (written in C)
# - Handles large vocabularies efficiently
# - Supports negative sampling, hierarchical softmax
# - Trains orders of magnitude faster than our toy version
# - Our toy model (cell_7) was just for UNDERSTANDING the concept
# This cell is the PRODUCTION-READY implementation
#
# FLOW OF THIS CELL:
# 1. Re-split data (X=tokens, y=labels)
# 2. Train Gensim Word2Vec on training tokens
# 3. Inspect the learned vocabulary and embeddings
# 4. Check word similarity to verify quality
# ============================================================
from gensim.models import Word2Vec # Gensim's optimized Word2Vec implementation
# ============================================================
# STEP 1: Re-define X, y and re-split the data
#
# NOTE: This re-split is intentional and important.
# Word2Vec must be trained ONLY on X_train tokens —
# never on X_test tokens. If we accidentally trained
# Word2Vec on the full dataset (including test sentences),
# the model would have "seen" test data during training,
# making our evaluation results unrealistically optimistic.
# This is called DATA LEAKAGE and must be avoided.
# ============================================================
# X = the preprocessed token lists from cell_3
# e.g., X[0] = ["profit", "rose", "significantly"]
X = df['tokens']
# y = the numeric sentiment labels from cell_2
# e.g., y[0] = 2 (positive)
y = df['label_encoded']
# Re-split with the SAME parameters as cell_4 so we get
# the exact same train/test split (random_state=42 ensures this)
# X_train: 80% of token lists → Word2Vec trains on these
# X_test: 20% of token lists → held out, never seen in training
X_train, X_test, y_train, y_test = train_test_split(
X, y,
test_size=0.2, # 20% for testing
random_state=42, # same seed = same split as cell_4
stratify=y # preserve class balance in both splits
)
# ============================================================
# STEP 2: Train the Gensim Word2Vec Model
#
# Each parameter explained in detail:
#
# sentences=X_train
# → the training data: a list of token lists
# → e.g., [["profit","rose"], ["sales","declined"], ...]
# → Word2Vec iterates over these to build co-occurrence stats
#
# vector_size=100
# → each word gets a 100-dimensional embedding vector
# → higher = more expressive but slower & needs more data
# → 100 is a reasonable default for moderate-sized datasets
# → our toy used 20; real applications use 100-300
#
# window=10
# → context window size: looks 10 words left AND right
# → larger window = captures broader, more topical context
# → e.g., window=10 means "profit" can be paired with words
# up to 10 positions away in the same sentence
#
# min_count=1
# → ignore words that appear fewer than 1 time total
# → min_count=1 means KEEP ALL words (even rare ones)
# → in larger datasets, min_count=5 is common to filter noise
#
# sg=1
# → training algorithm:
# sg=1 → Skip-Gram (predict context FROM center word)
# sg=0 → CBOW (predict center FROM context words)
# → Skip-Gram generally works better for rare words and
# smaller datasets; CBOW is faster on large datasets
#
# epochs=50
# → number of full passes through the entire training corpus
# → more epochs = more refined embeddings (up to a point)
# → 50 is a good balance between quality and training time
#
# seed=42
# → fixes the random initialization of weight matrices
# → ensures the same embeddings are produced every run
# ============================================================
w2v_model = Word2Vec(
sentences=X_train, # tokenized training sentences only
vector_size=100, # 100-dim embedding per word
window=10, # look 10 words left and right
min_count=1, # keep all words regardless of frequency
sg=1, # use Skip-Gram algorithm
epochs=50, # train for 50 full passes
seed=42 # reproducibility
)
# ============================================================
# STEP 3: Inspect the learned vocabulary
# ============================================================
# w2v_model.wv is the "word vectors" object — it stores all
# the learned embeddings and provides lookup utilities.
#
# .key_to_index is a dictionary: {word: index_in_matrix}
# .keys() gives all words in the vocabulary
# list(...)[:30] takes just the first 30 for display
# These are ordered by FREQUENCY (most common words first)
print("First 30 words in vocabulary:",
list(w2v_model.wv.key_to_index.keys())[:30])
# len(w2v_model.wv) = total number of unique words the model
# learned embeddings for. This equals the number of unique
# tokens in X_train after applying min_count filtering.
print("Vocabulary size:", len(w2v_model.wv))
# ============================================================
# STEP 4: Word Index Lookup
# ============================================================
# key_to_index["profit"] tells us which ROW in the embedding
# matrix corresponds to the word "profit".
# This index is used internally by the model to look up vectors.
# e.g., might print: Index of 'profit': 47
print("Index of 'profit':", w2v_model.wv.key_to_index["profit"])
# ============================================================
# STEP 5: Retrieve a word's embedding vector
# ============================================================
# w2v_model.wv["profit"] directly returns the 100-dimensional
# numpy array (the embedding vector) for the word "profit".
# This is shorthand for w2v_model.wv.get_vector("profit").
#
# The 100 numbers themselves are not human-interpretable —
# their RELATIONSHIPS to other vectors are what matters.
# e.g., w2v["profit"] should be close to w2v["revenue"]
print("Embedding vector for 'profit':")
print(w2v_model.wv["profit"])
# Expected output: array of 100 floats, e.g.:
# [ 0.0231 -0.1542 0.3871 0.0892 ... ]
# ============================================================
# STEP 6: Find the most similar words to "profit"
# ============================================================
# .most_similar(word, topn=5) finds the TOP 5 words whose
# embedding vectors are CLOSEST to "profit" in 100D space,
# measured by cosine similarity.
#
# This is the best way to verify the model learned meaningful
# financial semantics — the neighbors of "profit" should be
# financially related words like "revenue", "earnings", etc.
#
#
# Expected output: words like:
# [("revenue", 0.89), ("earnings", 0.87), ("loss", 0.85), ...]
print("Top 5 words similar to 'sale':",
w2v_model.wv.most_similar("sale", topn=5))
[cell 32 markdown]
## **Visualizing Word Embeddings with PCA**
This visualization will take a sample of the most frequent words from the vocabulary, reduce their high-dimensional vectors to 2D using PCA, and plot them. This gives a general sense of the "semantic space" the model has learned.
[cell 33 code]
# ============================================================
# cell_10
# PURPOSE: Visualize the Word2Vec embedding space learned on
# the Financial PhraseBank dataset using PCA.
#
# PROBLEM: Our word vectors are 100-dimensional — completely
# impossible to plot directly. We need to compress them down
# to 2D so we can see them on a screen.
#
# SOLUTION: PCA (Principal Component Analysis)
# PCA finds the 2 directions in 100D space that capture the
# MOST variation across all word vectors, then projects
# every vector onto those 2 directions.
#
# WHY PCA here instead of t-SNE (used in cell_7)?
# PCA → fast, linear, good for large sets (200 words)
# preserves GLOBAL structure (overall spread)
# t-SNE → slow, non-linear, good for small sets (7 words)
# preserves LOCAL structure (tight clusters)
# For a quick overview of 200 words, PCA is the right choice.
#
# WHAT TO LOOK FOR in the output plot:
# - Financially related words should cluster together
# - e.g., "profit", "revenue", "earnings" near each other
# - e.g., "decline", "loss", "fall" near each other
# - Very different words should be far apart
# ============================================================
from sklearn.decomposition import PCA # PCA: dimensionality reduction
import matplotlib.pyplot as plt # for plotting
import random # available but not used here
# (likely kept from earlier cells)
print("Visualizing a Sample of the Word Embedding Space")
# ============================================================
# STEP 1: Extract vocabulary and vectors from trained model
# ============================================================
# w2v_model.wv.index_to_key is a LIST of all words in the
# vocabulary, ordered by FREQUENCY (most frequent first).
# index_to_key[0] = the most frequent word in the training data
# index_to_key[1] = second most frequent, and so on.
# We convert it to a plain Python list for easy indexing.
vocab = list(w2v_model.wv.index_to_key)
# w2v_model.wv.vectors is a 2D NumPy array of ALL embeddings.
# Shape: (vocabulary_size, vector_size) = (N, 100)
# Row i = the 100-dimensional embedding for vocab[i]
# vocab and vectors are ALIGNED: vocab[i] ↔ vectors[i]
vectors = w2v_model.wv.vectors
# ============================================================
# STEP 2: Select the first 200 words to plot
# ============================================================
# We only plot 200 words (not all) because:
# - Plotting thousands of overlapping labels is unreadable
# - The first 200 are the MOST FREQUENT words in the corpus
# - High-frequency financial words (profit, market, sales...)
# are the most interesting and meaningful to visualize
num_words_to_plot = 200
# list(range(200)) = [0, 1, 2, ..., 199]
# These are the row indices of the 200 most frequent words
# in both vocab and vectors
sample_indices = list(range(num_words_to_plot))
# Select the 200 most frequent word strings
# sample_vocab[i] = the word at position i in frequency order
sample_vocab = [vocab[i] for i in sample_indices]
# Select the corresponding 200 embedding vectors
# sample_vectors shape: (200, 100)
# sample_vectors[i] = 100-dim vector for sample_vocab[i]
sample_vectors = vectors[sample_indices]
# ============================================================
# STEP 3: Reduce from 100D → 2D using PCA
#
# PCA works by finding the axes of MAXIMUM VARIANCE:
# PC1 (x-axis) = the direction where word vectors vary MOST
# PC2 (y-axis) = the direction of second-most variation,
# constrained to be perpendicular to PC1
#
# Every 100-dim vector is then projected onto these 2 axes,
# giving us 2 numbers (x, y) per word we can plot.
#
# Information loss: we go from 100 dimensions to 2, so we
# lose most information — but the most important structure
# (the directions of greatest variation) is preserved.
# ============================================================
# n_components=2: keep only the top 2 principal components
# random_state=42: reproducibility (PCA uses randomness
# internally for large matrices)
pca = PCA(n_components=2, random_state=42)
# fit_transform does TWO things in one call:
# fit: learns the 2 principal components from sample_vectors
# transform: projects all 200 vectors onto those 2 components
#
# Input shape: (200, 100) — 200 words, 100 dimensions each
# Output shape: (200, 2) — 200 words, 2 dimensions each
vectors_2d = pca.fit_transform(sample_vectors)
# vectors_2d[i] = [x_coordinate, y_coordinate] for word i
# vectors_2d[:, 0] = all x coordinates (PC1 values)
# vectors_2d[:, 1] = all y coordinates (PC2 values)
# ============================================================
# STEP 4: Create the scatter plot
# ============================================================
# Large figure size (16×12) needed because we're plotting
# 200 labeled points — we need space so labels don't overlap
plt.figure(figsize=(16, 12))
# Plot all 200 words as dots:
# vectors_2d[:, 0] → x coordinates of all 200 words
# vectors_2d[:, 1] → y coordinates of all 200 words
# color='steelblue' → all dots are the same steel blue color
# alpha=0.3 → dots are 70% transparent so overlapping dots
# don't create a dark blob; labels stay readable
plt.scatter(vectors_2d[:, 0], vectors_2d[:, 1],
color='steelblue',
alpha=0.3) # semi-transparent dots
# ============================================================
# STEP 5: Add word labels to each dot
# ============================================================
# Loop over all 200 words and add a text label near each dot
for i, word in enumerate(sample_vocab):
# plt.annotate() places a text label connected to a point:
# word → the text to display (the word itself)
# xy=(x, y) → the data point this label refers to
# vectors_2d[i, 0] = x for word i
# vectors_2d[i, 1] = y for word i
# alpha=0.8 → labels are 80% opaque (slightly transparent
# so overlapping labels are still readable)
#
# NOTE: No offset is applied here (unlike cell_7's t-SNE plot)
# so labels appear directly at the dot position.
# With 200 words this causes some overlap — acceptable for
# an exploratory overview visualization.
plt.annotate(word,
xy=(vectors_2d[i, 0], vectors_2d[i, 1]),
alpha=0.8)
# ============================================================
# STEP 6: Labels, title, and display
# ============================================================
# f-string inserts num_words_to_plot=200 into the title
plt.title(f'2D PCA Visualization of the First {num_words_to_plot} Word Embeddings',
fontsize=16)
# Axis labels explain what the dimensions represent:
# PC1 and PC2 are abstract mathematical directions —
# they don't have a simple human-readable meaning like
# "positivity" or "size", but words close together along
# these axes tend to be used in similar financial contexts
plt.xlabel('Principal Component 1') # direction of most variance
plt.ylabel('Principal Component 2') # direction of second-most variance
plt.grid(True) # adds gridlines to help judge relative positions
plt.show()
# ============================================================
# EXPECTED OUTPUT:
# A scatter plot of 200 labeled word dots where:
#
# Words used in similar financial contexts → plotted nearby
# Words used in different contexts → plotted far apart
#
# Look for mini-clusters like:
# "profit", "revenue", "earnings", "margin" → near each other
# "loss", "decline", "fall", "decrease" → near each other
# "CEO", "chairman", "president", "director" → near each other
#
# The axes themselves (PC1, PC2) have no direct interpretation —
# what matters is the RELATIVE DISTANCES between words.
# ============================================================
[cell 34 markdown]
## **From Word Embeddings to Document Embeddings**
So far, we trained a Word2Vec model where each word has its own embedding vector.
But for classification or similarity tasks, we need a **document-level representation** (one vector per sentence/document). This can be done in following way:
# **Simple Averaging of Word Vectors**
The easiest approach is to take all the word vectors in a sentence and compute their average.
This gives a fixed-length vector (same dimension as Word2Vec, e.g. 100) representing the whole document.
[cell 35 code]
# ============================================================
# cell_11
# PURPOSE: Convert each sentence (a list of word tokens) into
# a single FIXED-LENGTH vector that represents the whole
# sentence, using SIMPLE AVERAGING of its word vectors.
#
# THE CORE PROBLEM THIS SOLVES:
# Word2Vec gives us one vector PER WORD (100-dim each).
# But classifiers need one vector PER SENTENCE.
# Different sentences have different numbers of words —
# we can't feed variable-length inputs to a classifier.
#
# SOLUTION — Simple Averaging:
# Take all word vectors in a sentence and average them
# element-wise into a single 100-dim vector.
#
# Example:
# Sentence: ["profit", "rose", "significantly"]
#
# "profit" → [0.23, -0.15, 0.87, ...] (100 numbers)
# "rose" → [0.11, 0.42, -0.21, ...] (100 numbers)
# "significantly"→ [-0.05, 0.31, 0.44, ...] (100 numbers)
# ─────────────────────────────────────────
# average → [0.097, 0.193, 0.367, ...] (100 numbers)
# ↑
# This single vector now represents
# the ENTIRE sentence for the classifier
# ============================================================
import numpy as np # for array operations and np.mean(), np.zeros()
# ============================================================
# STEP 1: Define the sentence_vector() function
#
# Input: tokens → a list of word strings for one sentence
# e.g., ["profit", "rose", "significantly"]
# model → the trained Gensim Word2Vec model
#
# Output: a single 1D numpy array of shape (100,)
# representing the entire sentence
# ============================================================
def sentence_vector(tokens, model):
# Build a list of word vectors for all VALID tokens.
# "if w in model.wv" is a crucial guard — it skips any
# token that is NOT in the Word2Vec vocabulary.
#
# WHY can a token be missing from model.wv?
# - The word appeared only in X_test but not X_train
# (Word2Vec was trained only on X_train, remember)
# - The word was filtered by min_count during training
#
# model.wv[w] looks up word w and returns its 100-dim
# numpy array (the learned embedding vector)
#
# Result: a list of numpy arrays, one per valid token
# e.g., [[0.23,-0.15,...], [0.11,0.42,...], [-0.05,0.31,...]]
vectors = [model.wv[w] for w in tokens if w in model.wv]
# Edge case: if NO tokens in this sentence are in the
# vocabulary, we have nothing to average.
# Return a zero vector of the correct size (100,) so the
# sentence still has a valid representation.
# This prevents crashes and keeps all sentences aligned.
# np.zeros(model.vector_size) = array([0., 0., ..., 0.])
# shape: (100,)
if not vectors:
return np.zeros(model.vector_size) # fallback: all zeros
# np.mean(vectors, axis=0) computes the ELEMENT-WISE average
# across all word vectors in this sentence.
#
# "axis=0" means: average DOWN the rows (across words),
# keeping the column dimension (100) intact.
#
# Visual example with 3 words, 4-dim vectors (simplified):
#
# dim0 dim1 dim2 dim3
# "profit": [0.23, -0.15, 0.87, 0.02]
# "rose": [0.11, 0.42, -0.21, 0.55]
# "signif": [-0.05, 0.31, 0.44, -0.12]
# ─────────────────────────────
# mean: [0.097, 0.193, 0.367, 0.15] ← one vector for sentence
#
# Output shape: (100,) — always the same size regardless
# of how many words were in the sentence
return np.mean(vectors, axis=0)
# ============================================================
# STEP 2: Apply sentence_vector() to every sentence in
# both X_train and X_test
# ============================================================
# List comprehension: calls sentence_vector() for every
# token list in X_train, producing a list of 100-dim arrays.
#
# np.vstack([...]) STACKS that list of 1D arrays vertically
# into a single 2D matrix:
# [array(100,), array(100,), ...] → matrix(N_train, 100)
#
# X_train_vec shape: (N_train, 100)
# Each ROW = the averaged embedding vector for one sentence
# Each COL = one of the 100 embedding dimensions
X_train_vec = np.vstack([sentence_vector(tokens, w2v_model)
for tokens in X_train])
# Same process for X_test — but NOTE: we use the SAME
# w2v_model that was trained on X_train only.
# We do NOT re-train or re-fit anything on X_test data.
# X_test_vec shape: (N_test, 100)
X_test_vec = np.vstack([sentence_vector(tokens, w2v_model)
for tokens in X_test])
# ============================================================
# STEP 3: Verify the shapes
# ============================================================
# Expected output:
# Document embedding shape (train): (3876, 100)
# Document embedding shape (test): (969, 100)
#
# (3876, 100) means:
# 3876 training sentences, each now a 100-dim vector ✓
# (969, 100) means:
# 969 test sentences, each now a 100-dim vector ✓
#
# Both have exactly 100 columns — this is what the classifier
# needs: a fixed-size input for every single sentence.
print("Document embedding shape (train):", X_train_vec.shape)
print("Document embedding shape (test):", X_test_vec.shape)
# ============================================================
# LIMITATION OF SIMPLE AVERAGING (addressed in next cell):
#
# Simple averaging treats ALL words as EQUALLY important.
# But in financial text, some words matter much more:
#
# "The company reported a significant loss"
# ↑ ↑
# stopword-like KEY word
# (common, less informative) (critical for sentiment)
#
# "significant" and "loss" should contribute MORE to the
# sentence vector than filler words like "the", "a", "company"
#
# The next cell fixes this with TF-IDF WEIGHTED averaging,
# where important/rare words get higher weights in the average.
# ============================================================
[cell 36 markdown]
# **Simple Averaging is Weak**
- It **treats all words equally**, even though some words (e.g., *profit*, *loss*) carry more meaning than filler words.
# **TF-IDF Weighted Averaging**
A smarter variant is to weight each word vector by its **TF-IDF score** (term frequency–inverse document frequency).
This way:
- Rare but informative words get **higher weight**.
- Very common words like *said*, *year*, *today* get **down-weighted**.
[cell 37 markdown]
It computes a **weighted average** of word embeddings.
$$
\text{DocVec} = \frac{\sum_{i=1}^{N} w_i \cdot v(w_i)}{\sum_{i=1}^{N} w_i}
$$
where \(v(w_i)\) is the Word2Vec embedding of word \(i\), and \(w_i\) is its TF-IDF weight.
[cell 38 code]
# ============================================================
# cell_12
# PURPOSE: Upgrade from simple averaging (cell_11) to
# TF-IDF WEIGHTED averaging for better sentence vectors.
#
# PROBLEM WITH SIMPLE AVERAGING (cell_11):
# Every word contributed equally to the sentence vector.
# But common words like "said", "company", "year" appear
# in almost every sentence and carry very little sentiment
# signal — they should matter LESS than rare, informative
# words like "bankruptcy", "surge", "plummet".
#
# SOLUTION — TF-IDF Weighted Averaging:
# Instead of a plain average, compute a WEIGHTED average
# where each word's contribution is scaled by its TF-IDF score.
#
# TF-IDF score is HIGH for rare, informative words
# TF-IDF score is LOW for common, uninformative words
#
# WHAT IS TF-IDF?
# TF (Term Frequency) = how often word appears in THIS sentence
# IDF (Inverse Doc Frequency) = how rare the word is ACROSS ALL sentences
# = log(total_sentences / sentences_containing_word)
#
# Common word "company": appears in 1000/1000 sentences → IDF ≈ 0 (low weight)
# Rare word "bankruptcy": appears in 10/1000 sentences → IDF ≈ 4.6 (high weight)
#
# FORMULA:
# Σ (IDF(word) × vector(word))
# DocVec= ─────────────────────────────
# Σ IDF(word)
#
# We only use IDF (not full TF-IDF) because each sentence
# is short — TF within a sentence is less meaningful.
# ============================================================
from sklearn.feature_extraction.text import TfidfVectorizer # computes TF-IDF weights
import numpy as np # for array math
# ============================================================
# STEP 1: Convert token lists → strings for TF-IDF
#
# TfidfVectorizer expects plain text strings as input,
# but our X_train contains lists of tokens.
# We join each token list back into a single string.
#
# e.g., ["profit", "rose", "significantly"]
# → "profit rose significantly"
#
# CRITICAL: We do this ONLY for X_train — never X_test.
# TF-IDF must be fitted on training data only to prevent
# data leakage (test vocabulary must not influence weights).
# ============================================================
docs_train = [" ".join(tokens) for tokens in X_train]
# docs_train is now a list of strings, one per training sentence
# e.g., ["profit rose significantly",
# "sales declined sharply", ...]
# ============================================================
# STEP 2: Fit TF-IDF on training data only
#
# TfidfVectorizer() with default settings:
# - tokenizes on whitespace and punctuation
# - computes IDF for every unique word in docs_train
# - IDF formula: log((1 + n) / (1 + df)) + 1 (sklearn default)
# where n = total documents, df = documents containing this word
#
# .fit() LEARNS the IDF weights from docs_train.
# We do NOT call .transform() here because we don't need
# the full TF-IDF matrix — we only need the IDF weights
# to use as scaling factors for the word vectors.
# ============================================================
tfidf = TfidfVectorizer()
tfidf.fit(docs_train) # learns IDF for every word in X_train
# ONLY fitted on training data ← important
# ============================================================
# STEP 3: Extract the IDF weights into a simple dictionary
#
# tfidf.get_feature_names_out() → array of all vocabulary words
# e.g., ["bankruptcy", "company", "profit", "surge", ...]
#
# tfidf.idf_ → array of corresponding IDF values
# e.g., [4.61, 0.18, 1.23, 3.89, ...]
# higher value = rarer word = more informative
#
# zip(...) pairs each word with its IDF value:
# [("bankruptcy", 4.61), ("company", 0.18), ...]
#
# dict(...) converts to a lookup dictionary:
# {"bankruptcy": 4.61, "company": 0.18, "profit": 1.23, ...}
#
# This dict is faster to look up than calling tfidf methods
# repeatedly inside our vectorization loop below.
# ============================================================
tfidf_vocab = dict(zip(tfidf.get_feature_names_out(), tfidf.idf_))
# tfidf_vocab["profit"] → e.g., 1.23 (moderately common)
# tfidf_vocab["bankruptcy"] → e.g., 4.61 (rare, high weight)
# tfidf_vocab["company"] → e.g., 0.18 (very common, low weight)
# ============================================================
# STEP 4: Define TF-IDF weighted sentence vector function
#
# Input: tokens → list of word strings for one sentence
# model → trained Word2Vec model
# tfidf_vocab→ {word: idf_weight} dictionary
#
# Output: one 1D numpy array of shape (100,)
# representing the whole sentence, with important
# words contributing more than common words
# ============================================================
def tfidf_weighted_vector(tokens, model, tfidf_vocab):
vectors = [] # will store: idf_weight × word_vector for each valid word
weights = [] # will store: idf_weight for each valid word
# (needed for the denominator of the weighted average)
for word in tokens:
# Double guard — word must exist in BOTH:
# model.wv : Word2Vec vocabulary (has a learned vector)
# tfidf_vocab: TF-IDF vocabulary (has a learned IDF weight)
#
# A word could be missing from tfidf_vocab if it appeared
# in X_test tokens but not in X_train (TF-IDF was only
# fitted on X_train). We skip such words entirely.
if word in model.wv and word in tfidf_vocab:
# Retrieve this word's IDF weight
# Higher IDF = rarer word = should influence sentence
# vector MORE
idf_weight = tfidf_vocab[word]
# Scale the word's embedding vector by its IDF weight.
# model.wv[word] shape: (100,)
# idf_weight × vector: each of the 100 dimensions gets
# scaled up (rare word) or down (common word)
vectors.append(model.wv[word] * idf_weight)
# Store the weight separately for the denominator
weights.append(idf_weight)
# Edge case: no valid words found (all words missing from
# vocabulary). Return zero vector to avoid crashes.
if not vectors:
return np.zeros(model.vector_size) # shape: (100,)
# Compute the WEIGHTED AVERAGE:
#
# np.sum(vectors, axis=0) → sums all weighted vectors element-wise
# shape: (100,) — one summed value per dim
# np.sum(weights) → sums all IDF weights (a single scalar)
#
# Dividing element-wise gives the weighted average:
#
# Simple average: (v1 + v2 + v3) / 3
# Weighted average: (w1*v1 + w2*v2 + w3*v3) / (w1 + w2 + w3)
#
# Visual example (3 words, simplified to 2 dims):
#
# IDF vector
# "bankruptcy": 4.6 × [0.8, -0.3] = [3.68, -1.38] ← high weight
# "company": 0.2 × [0.1, 0.5] = [0.02, 0.10] ← low weight
# "reported": 0.3 × [0.2, 0.1] = [0.06, 0.03] ← low weight
# ─── ────────────── ──────────────
# sum weights: 5.1 sum vectors: [3.76, -1.25]
#
# weighted avg: [3.76/5.1, -1.25/5.1] = [0.737, -0.245]
# ↑
# "bankruptcy" dominates because it's rare and informative ✓
return np.sum(vectors, axis=0) / np.sum(weights)
# ============================================================
# STEP 5: Build sentence embedding matrices for train and test
# ============================================================
# Apply tfidf_weighted_vector() to every sentence in X_train.
# np.vstack() stacks the resulting list of (100,) arrays
# into a 2D matrix of shape (N_train, 100).
# IMPORTANT: uses w2v_model and tfidf_vocab both fitted on
# training data only — no leakage from test set.
X_train_vec = np.vstack([tfidf_weighted_vector(tokens, w2v_model, tfidf_vocab)
for tokens in X_train])
# Same for X_test — apply the SAME tfidf_vocab and w2v_model
# (no re-fitting). Unknown test words are simply skipped.
X_test_vec = np.vstack([tfidf_weighted_vector(tokens, w2v_model, tfidf_vocab)
for tokens in X_test])
# ============================================================
# STEP 6: Verify shapes
# ============================================================
# Expected output:
# Train embeddings shape: (3876, 100)
# Test embeddings shape: (969, 100)
#
# Identical shapes to cell_11 — same structure, better quality.
# These improved vectors will now be used for classification.
print("Train embeddings shape:", X_train_vec.shape)
print("Test embeddings shape:", X_test_vec.shape)
[cell 39 markdown]
## **Word2Vec Model Training and Hyperparameter Tuning**
We tune the key Word2Vec hyperparameters, `vector_size` and `window`, to identify which configuration produces the best embeddings for our sentiment classification task. An MLP classifier will then be used for sentiment classification.
[cell 40 markdown]
# **Hyperparameter tuning for `vector_size`**
We test vector sizes of `[50, 150, 250, 350]`. A fixed window of 10 is used initially.
[cell 41 markdown]
# MLP classifier and helper code
[cell 42 code]
# ============================================================
# cell_13
# PURPOSE: Define the MLP (Multi-Layer Perceptron) classifier
# and its training/evaluation helper functions.
#
# This cell does NOT train anything yet — it just defines
# the TOOLS that will be used in the next cells.
# Think of it as writing the recipe before cooking.
#
# WHAT IS AN MLP?
# A Multi-Layer Perceptron is a simple neural network with:
# - An INPUT layer : receives the 100-dim sentence vector
# - A HIDDEN layer : learns intermediate representations
# - An OUTPUT layer : produces a score for each class
#
# ARCHITECTURE:
#
# Input Hidden Output
# (100-dim) → (64-dim) → (3-dim)
# sentence learned score per
# vector features sentiment class
# ↑
# ReLU activation adds non-linearity so the
# model can learn complex patterns, not just
# straight lines through the data
#
# WHY MLP OVER LOGISTIC REGRESSION?
# Logistic Regression: can only learn LINEAR boundaries
# MLP: can learn NON-LINEAR boundaries
# via its hidden layer + ReLU
# ============================================================
import torch # PyTorch core
import torch.nn as nn # neural network layers
import torch.optim as optim # optimizers (Adam, SGD)
from sklearn.metrics import accuracy_score, f1_score # evaluation metrics
import random, numpy as np, torch # re-imported for safety
# ensures these are available
# regardless of cell run order
# ============================================================
# PART 1: Define the MLP Architecture
# ============================================================
class MLPClassifier(nn.Module):
# nn.Module is PyTorch's base class for all neural networks
# Every custom model MUST inherit from it
def __init__(self, input_dim, hidden_dim, num_classes):
# input_dim : size of input vector (= 100, our embedding size)
# hidden_dim : size of hidden layer (= 64, chosen as hyperparameter)
# num_classes : number of output classes (= 3: pos, neg, neutral)
super(MLPClassifier, self).__init__()
# super().__init__() initializes PyTorch's internal
# bookkeeping — REQUIRED in every nn.Module subclass
# LAYER 1: Fully Connected (Linear) — input → hidden
# nn.Linear(in_features, out_features) creates a weight
# matrix of shape (hidden_dim, input_dim) = (64, 100)
# and a bias vector of shape (hidden_dim,) = (64,)
# Transforms each 100-dim input → 64-dim hidden vector
self.fc1 = nn.Linear(input_dim, hidden_dim)
# fc1 weight shape: (64, 100)
# ACTIVATION: ReLU (Rectified Linear Unit)
# ReLU(x) = max(0, x)
# Applied element-wise AFTER fc1, BEFORE fc2.
# Without activation, stacking linear layers is
# mathematically equivalent to just ONE linear layer
# (they collapse into a single matrix multiplication).
# ReLU introduces NON-LINEARITY so the network can
# learn complex, curved decision boundaries.
#
# Visual:
# Input: [-0.5, 0.3, -0.1, 0.8]
# ReLU: [ 0.0, 0.3, 0.0, 0.8] ← negatives → 0
self.relu = nn.ReLU()
# LAYER 2: Fully Connected (Linear) — hidden → output
# Transforms each 64-dim hidden vector → 3-dim output
# The 3 output values are RAW SCORES (logits) for:
# output[0] = score for class 0 (negative)
# output[1] = score for class 1 (neutral)
# output[2] = score for class 2 (positive)
# Higher score = more likely to be that class
self.fc2 = nn.Linear(hidden_dim, num_classes)
# fc2 weight shape: (3, 64)
def forward(self, x):
# forward() defines the COMPUTATION FLOW.
# PyTorch calls this automatically when you do model(x).
#
# Data flow through the network:
# x → shape: (batch_size, 100) input vectors
# fc1(x) → shape: (batch_size, 64) linear transform
# relu(…) → shape: (batch_size, 64) non-linearity applied
# fc2(…) → shape: (batch_size, 3) raw class scores
#
# Written as one line for compactness:
# fc2( relu( fc1(x) ) )
return self.fc2(self.relu(self.fc1(x)))
# Returns raw logits — CrossEntropyLoss applies
# softmax internally, so we don't add it here
# ============================================================
# PART 2: Define the Training Function
# ============================================================
def train_model(model, X_train, y_train, epochs=10, lr=1e-3, batch_size=32):
# model : the MLPClassifier instance to train
# X_train : numpy array of sentence vectors, shape (N, 100)
# y_train : pandas Series of integer labels, shape (N,)
# epochs : how many full passes through training data (default 10)
# lr : learning rate — step size for weight updates (default 0.001)
# batch_size : number of samples processed per weight update (default 32)
# model.train() switches the model to TRAINING MODE.
# This enables behaviours only active during training
# (e.g., Dropout layers drop neurons, BatchNorm uses
# batch statistics). Our model doesn't use these, but
# calling .train() is good practice.
model.train()
# Adam optimizer: an advanced version of SGD that:
# - Adapts the learning rate per parameter automatically
# - Uses momentum to smooth out noisy gradient updates
# - Generally converges faster than plain SGD
# model.parameters() tells Adam WHICH tensors to update
# (fc1 weights, fc1 bias, fc2 weights, fc2 bias)
optimizer = optim.Adam(model.parameters(), lr=lr)
# CrossEntropyLoss: standard loss for multi-class classification
# Internally applies softmax then computes negative log likelihood
# of the correct class. Lower = better predictions.
criterion = nn.CrossEntropyLoss()
# Convert numpy arrays → PyTorch tensors for training
# dtype=torch.float32 : neural networks need 32-bit floats
X_train_tensor = torch.tensor(X_train, dtype=torch.float32)
# shape: (N_train, 100)
# y_train is a pandas Series → .values extracts the numpy array
# dtype=torch.long : CrossEntropyLoss requires 64-bit integers
y_train_tensor = torch.tensor(y_train.values, dtype=torch.long)
# shape: (N_train,)
# TensorDataset pairs X and y tensors together so they
# stay aligned when shuffled — dataset[i] = (X[i], y[i])
dataset = torch.utils.data.TensorDataset(X_train_tensor, y_train_tensor)
# DataLoader handles MINI-BATCH training automatically:
# batch_size=32 : feeds 32 samples at a time (not all at once)
# shuffle=True : randomly reorders data each epoch so the
# model doesn't memorize the order of samples
#
# WHY MINI-BATCHES?
# Full batch (all data at once): too much memory, slow
# Mini-batch (32 at a time): memory efficient, faster,
# noisier gradients help
# escape local minima
loader = torch.utils.data.DataLoader(
dataset,
batch_size=batch_size,
shuffle=True
)
# ---- The Training Loop ----
for epoch in range(epochs):
# Each epoch iterates over ALL mini-batches
# loader automatically yields (xb, yb) pairs where:
# xb shape: (32, 100) — one batch of sentence vectors
# yb shape: (32,) — corresponding labels
for xb, yb in loader:
# 1. Clear gradients from previous iteration
# (PyTorch accumulates gradients by default)
optimizer.zero_grad()
# 2. Forward pass: compute predictions for this batch
# model(xb) calls forward(xb) automatically
# loss compares predictions vs true labels yb
loss = criterion(model(xb), yb)
# 3. Backward pass: compute gradients of loss
# with respect to all model parameters
loss.backward()
# 4. Update weights using computed gradients
# Adam adjusts each weight to reduce the loss
optimizer.step()
# No return value — the model's weights are updated IN-PLACE
# ============================================================
# PART 3: Define the Evaluation Function
# ============================================================
def evaluate_model(model, X_test, y_test):
# model : trained MLPClassifier
# X_test : numpy array of test sentence vectors, shape (N_test, 100)
# y_test : pandas Series of true integer labels, shape (N_test,)
# model.eval() switches to EVALUATION MODE.
# Disables training-only behaviours (Dropout, BatchNorm).
# Always call this before running inference.
model.eval()
# Convert test data to tensors (same as in train_model)
X_test_tensor = torch.tensor(X_test, dtype=torch.float32)
y_test_tensor = torch.tensor(y_test.values, dtype=torch.long)
# torch.no_grad() tells PyTorch NOT to track gradients.
# During evaluation we don't need gradients (no backprop),
# so disabling them saves memory and speeds up computation.
with torch.no_grad():
# model(X_test_tensor) runs the full forward pass on
# ALL test samples at once.
# Output shape: (N_test, 3) — one score per class per sample
#
# torch.argmax(..., dim=1) finds the INDEX of the highest
# score along dimension 1 (across the 3 class scores).
# This index IS the predicted class (0, 1, or 2).
# Output shape: (N_test,) — one predicted label per sample
#
# Example:
# scores = [0.1, 0.7, 0.2] → argmax = 1 → "neutral"
# scores = [0.8, 0.1, 0.1] → argmax = 0 → "negative"
preds = torch.argmax(model(X_test_tensor), dim=1)
# accuracy_score: fraction of correct predictions
# = number of correct / total predictions
# e.g., 0.82 means 82% of test sentences classified correctly
acc = accuracy_score(y_test_tensor, preds)
# f1_score: harmonic mean of precision and recall
# average="weighted": accounts for class imbalance by weighting
# each class's F1 by how many samples it has in the test set
# More reliable than accuracy when classes are imbalanced
# (e.g., if 60% of data is "neutral", a model that always
# predicts "neutral" gets 60% accuracy but very low F1)
f1 = f1_score(y_test_tensor, preds, average="weighted")
return acc, f1 # both are floats between 0 and 1
[cell 43 markdown]
# **Hyperparameter tuning for `vector_size`**
[cell 44 code]
# ============================================================
# cell_14
# PURPOSE: Tune the Word2Vec 'vector_size' hyperparameter
# to find the optimal embedding dimensionality for our
# sentiment classification task.
#
# WHAT IS vector_size?
# The number of dimensions in each word's embedding vector.
# e.g., vector_size=100 means every word is represented
# as a list of 100 numbers.
#
# WHY DOES IT MATTER?
# Too SMALL (e.g., 50):
# → not enough dimensions to capture nuanced meaning
# → "profit" and "revenue" might end up too similar to
# "loss" because there's no room to distinguish them
# → underfitting: model lacks expressive power
#
# Too LARGE (e.g., 300):
# → more expressive but needs MORE data to train well
# → risk of overfitting on small datasets
# → slower training and more memory
#
# SWEET SPOT: somewhere in between, found by testing
#
# APPROACH:
# Train a fresh Word2Vec + MLP pipeline for EACH vector_size
# in [50, 100, 150, 200, 250, 300], keeping ALL other
# hyperparameters fixed (window=10, epochs=50, etc.)
# Record accuracy and weighted F1 for each → pick the best.
#
# This is called a GRID SEARCH over one hyperparameter.
# ============================================================
from gensim.models import Word2Vec # Word2Vec model
import numpy as np # array operations
# ---- Fix all random seeds for reproducibility ----
# Every iteration trains a new model with random initialization.
# Fixing seeds ensures fair comparison — differences in results
# come from vector_size only, not from random variation.
SEED = 42
random.seed(SEED) # Python random
np.random.seed(SEED) # NumPy random
torch.manual_seed(SEED) # PyTorch random
# ---- Storage for results ----
# results will be a nested dict:
# {
# 50: {"accuracy": 0.71, "f1": 0.70},
# 100: {"accuracy": 0.74, "f1": 0.73},
# ...
# }
# We populate this inside the loop and convert to a
# DataFrame at the end for easy viewing and plotting.
results = {}
# The 6 vector sizes we want to compare.
# Spread evenly from small (50) to large (300) so we can
# see the full trend of how performance changes with size.
vector_sizes = [50, 100, 150, 200, 250, 300]
# ============================================================
# MAIN LOOP: Train and evaluate one full pipeline per size
# ============================================================
for size in vector_sizes:
print(f"\n=== Training Word2Vec with vector_size={size} ===")
# --------------------------------------------------------
# STEP 1: Train a fresh Word2Vec model at this vector_size
#
# Every hyperparameter is FIXED except vector_size=size:
# window=10 : fixed context window
# min_count=1 : keep all words
# sg=1 : Skip-Gram algorithm
# epochs=50 : full training passes
# seed=42 : reproducible weight initialization
#
# We train a BRAND NEW model each iteration — we cannot
# reuse the previous model because changing vector_size
# fundamentally changes the network architecture (different
# weight matrix dimensions).
# --------------------------------------------------------
w2v_model = Word2Vec(
sentences=X_train, # training token lists only (no leakage)
vector_size=size, # ← THE ONLY THING CHANGING each iteration
window=10, # fixed: look 10 words left and right
min_count=1, # fixed: keep all vocabulary words
sg=1, # fixed: use Skip-Gram
epochs=50, # fixed: 50 training passes
seed=42 # fixed: reproducible initialization
)
# --------------------------------------------------------
# STEP 2: Build TF-IDF weighted sentence vectors
#
# Uses the tfidf_weighted_vector() function from cell_12
# and the tfidf_vocab fitted on X_train in cell_12.
#
# NOTE: tfidf_vocab does NOT change between iterations —
# it was fitted once on X_train text and stays the same.
# Only the Word2Vec model changes (different vector_size).
#
# X_train_vec shape: (N_train, size) e.g., (3876, 50) for size=50
# X_test_vec shape: (N_test, size) e.g., (969, 50) for size=50
# --------------------------------------------------------
X_train_vec = np.vstack([tfidf_weighted_vector(tokens, w2v_model, tfidf_vocab)
for tokens in X_train])
X_test_vec = np.vstack([tfidf_weighted_vector(tokens, w2v_model, tfidf_vocab)
for tokens in X_test])
# --------------------------------------------------------
# STEP 3: Build and train a fresh MLP classifier
#
# input_dim MUST match the current vector_size because
# the MLP's first layer (fc1) has shape (hidden_dim, input_dim).
# If vector_size=50, input_dim=50; if size=200, input_dim=200.
#
# X_train_vec.shape[1] always gives the correct input_dim
# regardless of which size we're currently testing.
# --------------------------------------------------------
input_dim = X_train_vec.shape[1] # = size (50, 100, 150, ...)
num_classes = len(set(y_train)) # = 3 (negative, neutral, positive)
# Create a FRESH MLP each iteration — we cannot reuse the
# previous one because input_dim changes with vector_size,
# so the weight matrices have different shapes.
mlp = MLPClassifier(
input_dim, # changes each iteration (= vector_size)
hidden_dim=64, # fixed: 64 hidden neurons
num_classes=num_classes # fixed: always 3 classes
)
# Train the MLP on the current vector_size embeddings.
# Fixed training settings for fair comparison —
# only the data (embedding size) changes, not training config.
train_model(
mlp,
X_train_vec,
y_train,
epochs=10, # fixed: 10 passes through training data
lr=1e-3 # fixed: Adam learning rate = 0.001
)
# --------------------------------------------------------
# STEP 4: Evaluate on test set
# Returns accuracy and weighted F1 (both 0→1 floats)
# --------------------------------------------------------
acc, f1 = evaluate_model(mlp, X_test_vec, y_test)
# --------------------------------------------------------
# STEP 5: Store results for this vector_size
# --------------------------------------------------------
results[size] = {"accuracy": acc, "f1": f1}
print(f"vector_size={size} | Accuracy: {acc:.4f} | Weighted F1: {f1:.4f}")
# Example output:
# vector_size=50 | Accuracy: 0.7142 | Weighted F1: 0.7089
# vector_size=100 | Accuracy: 0.7348 | Weighted F1: 0.7301
# vector_size=200 | Accuracy: 0.7512 | Weighted F1: 0.7489 ← best
# vector_size=300 | Accuracy: 0.7401 | Weighted F1: 0.7355
# ============================================================
# STEP 6: Summarise results in a DataFrame
# ============================================================
# pd.DataFrame.from_dict(results, orient="index"):
# orient="index" means each KEY of results dict (50,100,...)
# becomes a ROW index, and nested keys ("accuracy","f1")
# become COLUMN names.
#
# Resulting DataFrame layout:
#
# vector_size | accuracy | f1
# ────────────|──────────|──────────
# 50 | 0.7142 | 0.7089
# 100 | 0.7348 | 0.7301
# 150 | 0.7423 | 0.7398
# 200 | 0.7512 | 0.7489 ← best weighted F1
# 250 | 0.7489 | 0.7461
# 300 | 0.7401 | 0.7355
results_df = pd.DataFrame.from_dict(results, orient="index")
results_df.index.name = "vector_size" # label the row index column
# Displays the DataFrame as a formatted table in the notebook.
# (No print() needed — Jupyter auto-displays the last expression)
results_df
[cell 45 code]
# ============================================================
# cell_15
# PURPOSE: Visualize the hyperparameter tuning results from
# cell_14 as line plots so we can clearly see HOW performance
# changes as vector_size increases, and identify the
# optimal vector_size to use in the final model.
#
# WHY VISUALIZE INSTEAD OF JUST READING THE TABLE?
# The DataFrame in cell_14 shows numbers, but a plot lets us:
# - See the TREND at a glance (rising? falling? plateauing?)
# - Spot the PEAK performance point immediately
# - Identify diminishing returns (where bigger = no better)
#
# WHAT WE PLOT:
# Left chart → Accuracy vs vector_size
# Right chart → Weighted F1 vs vector_size
#
# We plot BOTH metrics because they can tell different stories
# on imbalanced datasets (as explained in cell_14).
# The weighted F1 chart is the one we trust for model selection.
# ============================================================
import matplotlib.pyplot as plt # plotting library
# ============================================================
# STEP 1: Rebuild the results DataFrame
#
# This is technically redundant — results_df was already
# created at the end of cell_14. It is recreated here to
# make this cell SELF-CONTAINED: it works correctly even if
# cell_14's last line was skipped or the variable was
# overwritten by another cell.
# ============================================================
# pd.DataFrame.from_dict(results, orient="index"):
# results dict keys (50,100,...,300) → row indices
# nested keys ("accuracy","f1") → column names
results_df = pd.DataFrame.from_dict(results, orient="index")
results_df.index.name = "vector_size" # names the index column
# ============================================================
# STEP 2: Create a figure with 2 side-by-side subplots
#
# plt.subplots(1, 2) creates:
# - 1 row, 2 columns of subplots
# - returns fig (the whole figure) and axes (array of 2 Axes)
# - axes[0] = left plot (Accuracy)
# - axes[1] = right plot (Weighted F1)
#
# figsize=(12, 5): figure is 12 inches wide × 5 inches tall
# Wide enough to show both plots clearly side by side
# ============================================================
fig, axes = plt.subplots(1, 2, figsize=(12, 5))
# ============================================================
# STEP 3: Left plot — Accuracy vs Vector Size
# ============================================================
# axes[0].plot() draws a LINE connecting all data points:
#
# results_df.index → x-axis values: [50,100,150,200,250,300]
# results_df["accuracy"]→ y-axis values: accuracy at each size
#
# marker="o" : draws a circle dot at each data point
# so individual measurements are visible
# linestyle="-": connects dots with a solid line so the
# trend is easy to follow
# color="b" : blue line (convention: blue for accuracy)
axes[0].plot(
results_df.index, # x: vector sizes
results_df["accuracy"], # y: accuracy values
marker="o", # circle at each measured point
linestyle="-", # solid connecting line
color="b" # blue
)
axes[0].set_title("Accuracy vs Vector Size") # chart title
axes[0].set_xlabel("Vector Size") # x-axis label
axes[0].set_ylabel("Accuracy") # y-axis label
axes[0].grid(True) # gridlines help
# read exact values
# ============================================================
# STEP 4: Right plot — Weighted F1 vs Vector Size
#
# Identical structure to the left plot but shows F1 score.
# This is the MORE IMPORTANT chart for model selection
# because weighted F1 is more reliable than accuracy on
# imbalanced datasets (see cell_14 explanation).
# ============================================================
axes[1].plot(
results_df.index, # x: vector sizes (same as left plot)
results_df["f1"], # y: weighted F1 values
marker="o", # circle at each measured point
linestyle="-", # solid connecting line
color="g" # green (different color to distinguish
# clearly from the blue accuracy plot)
)
axes[1].set_title("Weighted F1 vs Vector Size")
axes[1].set_xlabel("Vector Size")
axes[1].set_ylabel("Weighted F1 Score")
axes[1].grid(True)
# ============================================================
# STEP 5: Final layout and display
# ============================================================
# plt.tight_layout() automatically adjusts spacing between
# the two subplots so titles, labels, and tick marks don't
# overlap each other. Always call this before plt.show()
# when using multiple subplots.
plt.tight_layout()
plt.show() # renders and displays the figure
# ============================================================
# HOW TO READ THE OUTPUT PLOTS:
#
# Ideal curve shape we expect to see:
#
# Score
# ▲
# │ ●──●
# │ ● ●──●
# │ ●
# │ ●
# └──────────────────→ vector_size
# 50 100 150 200 250 300
# ↑
# peak here = best vector_size
#
# PATTERNS TO LOOK FOR:
#
# Steadily rising → bigger is better, try even larger sizes
# Peak then drops → sweet spot found, use the peak size
# Flat/plateau → size doesn't matter much beyond a point,
# pick the smallest size in the flat region
# (faster training, less memory)
# Noisy/zigzag → results are unstable, consider more epochs
# or a different random seed
#
# Based on the notebook's observation:
# vector_size = 200 gives the best weighted F1
# → this value is used as best_vector_size in cell_16
# ============================================================
[cell 46 markdown]
### **Observation**
Since weighted F1 is a more reliable metric than accuracy for imbalanced datasets, we use it for model selection.
According to weighted F1, the best choice here is **vector_size = 200**
[cell 47 markdown]
# **Hyperparameter tuning for `window`**
We now tune the `window` parameter.
A larger window means the model considers more context words around the target.
We will test window sizes of `[5, 10, 15, 20, 25]`.
[cell 48 code]
# ============================================================
# cell_16
# PURPOSE: Tune the Word2Vec 'window' hyperparameter to find
# the optimal context window size, using the best vector_size
# (200) already found in cell_14.
#
# WHAT IS THE WINDOW PARAMETER?
# The number of words to the LEFT and RIGHT of the center
# word that are considered "context" during Skip-Gram training.
#
# Visual example — center word = "profit":
#
# window=2: ["company", "reported"] "profit" ["rose", "sharply"]
# only these 4 words are context
#
# window=5: ["the", "company", "reported", "strong"] "profit"
# ["rose", "sharply", "beating", "all", "estimates"]
# up to 10 words are context
#
# WHAT DOES WINDOW SIZE CONTROL?
#
# Small window (2-5):
# → model sees only IMMEDIATE neighbors
# → learns tight syntactic relationships
# → e.g., "profit" paired with "rose", "fell", "increased"
# → captures how words FUNCTION grammatically
#
# Large window (10-25):
# → model sees broader surrounding context
# → learns topical/semantic relationships
# → e.g., "profit" paired with "revenue", "earnings", "CEO"
# → captures what TOPIC a word belongs to
#
# For SENTIMENT analysis, broader topical context is often
# more useful — knowing that "profit" appears in the same
# sentences as "growth" and "record" is a stronger sentiment
# signal than just knowing it appears next to "rose".
#
# APPROACH: Fix vector_size=200 (best from cell_14),
# vary window in [5, 10, 15, 20, 25], run full pipeline,
# record accuracy and weighted F1 for each → pick the best.
# ============================================================
import random, numpy as np, torch # re-imported for safety
# ---- Fix all random seeds for reproducibility ----
# Same reason as cell_14 — isolate the effect of window size
# by removing all other sources of randomness.
SEED = 42
random.seed(SEED) # Python random
np.random.seed(SEED) # NumPy random
torch.manual_seed(SEED) # PyTorch random
# ---- Storage for results ----
# Same structure as cell_14's results dict:
# {
# 5: {"accuracy": 0.74, "f1": 0.73},
# 10: {"accuracy": 0.75, "f1": 0.74},
# ...
# }
results_window = {}
# Use the best vector_size found in cell_14.
# This is now FIXED — we are only changing window here.
# Isolating one variable at a time is the correct way to
# tune hyperparameters (changing both at once makes it
# impossible to know which change caused the improvement).
best_vector_size = 200
# The 5 window sizes to compare.
# Range from tight (5) to very broad (25) context.
window_sizes = [5, 10, 15, 20, 25]
# ============================================================
# MAIN LOOP: Train and evaluate one full pipeline per window
# ============================================================
for w in window_sizes:
print(f"\n=== Training Word2Vec with window={w} ===")
# --------------------------------------------------------
# STEP 1: Train a fresh Word2Vec model at this window size
#
# FIXED: vector_size=200, min_count=1, sg=1,
# epochs=50, seed=42
# VARYING: window=w ← THE ONLY THING CHANGING
#
# A brand new model is trained each iteration because
# the window parameter affects ALL of training —
# the pairs generated, the gradients computed, and
# therefore the final embedding values all change.
# We cannot just "adjust" a previously trained model.
# --------------------------------------------------------
w2v_model = Word2Vec(
sentences=X_train, # training tokens only
vector_size=best_vector_size, # fixed at 200
window=w, # ← ONLY THIS CHANGES
min_count=1, # keep all vocabulary words
sg=1, # Skip-Gram algorithm
epochs=50, # full training passes
seed=42 # reproducible init
)
# --------------------------------------------------------
# STEP 2: Build TF-IDF weighted sentence vectors
#
# tfidf_vocab is UNCHANGED — still the same IDF weights
# fitted on X_train text in cell_12.
# Only the Word2Vec embeddings change (different window).
#
# X_train_vec shape: (N_train, 200) — always 200 dims
# X_test_vec shape: (N_test, 200) — always 200 dims
# (vector_size is fixed at 200, so shape is constant)
# --------------------------------------------------------
X_train_vec = np.vstack([tfidf_weighted_vector(tokens, w2v_model, tfidf_vocab)
for tokens in X_train])
X_test_vec = np.vstack([tfidf_weighted_vector(tokens, w2v_model, tfidf_vocab)
for tokens in X_test])
# --------------------------------------------------------
# STEP 3: Build and train a fresh MLP classifier
#
# input_dim is ALWAYS best_vector_size=200 here because
# vector_size is fixed. Unlike cell_14, the MLP architecture
# does NOT change between iterations — only the embedding
# VALUES change (different window = different trained vectors).
#
# We still create a fresh MLP each iteration to ensure
# we're measuring the effect of window on the embeddings,
# not accidentally benefiting from warm-started weights.
# --------------------------------------------------------
mlp = MLPClassifier(
best_vector_size, # input_dim = 200 (fixed)
hidden_dim=64, # fixed hidden layer size
num_classes=len(set(y_train)) # = 3 classes
)
# Train with same fixed settings as cell_14 for fair comparison
train_model(mlp, X_train_vec, y_train, epochs=10, lr=1e-3)
# --------------------------------------------------------
# STEP 4: Evaluate on test set
# --------------------------------------------------------
acc, f1 = evaluate_model(mlp, X_test_vec, y_test)
# --------------------------------------------------------
# STEP 5: Store results for this window size
# --------------------------------------------------------
results_window[w] = {"accuracy": acc, "f1": f1}
print(f"window={w} | Accuracy: {acc:.4f} | Weighted F1: {f1:.4f}")
# Expected pattern (approximate):
# window=5 | Accuracy: 0.7601 | Weighted F1: 0.7589 ← high F1
# window=10 | Accuracy: 0.7554 | Weighted F1: 0.7541
# window=15 | Accuracy: 0.7523 | Weighted F1: 0.7508
# window=20 | Accuracy: 0.7589 | Weighted F1: 0.7572 ← chosen
# window=25 | Accuracy: 0.7498 | Weighted F1: 0.7471
# ============================================================
# STEP 6: Summarise results in a DataFrame
# ============================================================
# Same structure as cell_14's results_df but indexed by
# window size instead of vector_size.
#
# Resulting DataFrame layout:
#
# window | accuracy | f1
# ───────|──────────|──────────
# 5 | 0.7601 | 0.7589 ← highest F1 numerically
# 10 | 0.7554 | 0.7541
# 15 | 0.7523 | 0.7508
# 20 | 0.7589 | 0.7572 ← CHOSEN (see observation below)
# 25 | 0.7498 | 0.7471
results_window_df = pd.DataFrame.from_dict(results_window, orient="index")
results_window_df.index.name = "window" # label the row index
# Display as formatted table in notebook
results_window_df
# ============================================================
# OBSERVATION (from notebook):
#
# Although window=5 gives the highest weighted F1 numerically,
# the notebook DELIBERATELY chooses window=20 instead.
#
# WHY? The reasoning is contextual richness vs raw score:
#
# window=5:
# → tiny margin better on THIS test set
# → but very narrow context — misses broader topic signals
# → e.g., "profit" only pairs with its 5 nearest neighbors,
# missing longer-range financial context clues
#
# window=20:
# → nearly identical F1 (minimal drop)
# → much richer context — captures document-level topics
# → financial sentences often have long-range dependencies:
# "Despite challenging market conditions the company
# managed to post a strong profit"
# (profit and market are 7+ words apart — needs window≥7)
#
# This is a DOMAIN KNOWLEDGE decision:
# In financial text, meaning often spans long sentences,
# so the broader window is theoretically better even if the
# test set numbers are nearly tied.
# ============================================================
[cell 49 code]
import matplotlib.pyplot as plt
# Convert results to DataFrame
results_window_df = pd.DataFrame.from_dict(results_window, orient="index")
results_window_df.index.name = "window"
# Create subplots
fig, axes = plt.subplots(1, 2, figsize=(12, 5))
# Plot Accuracy
axes[0].plot(results_window_df.index, results_window_df["accuracy"], marker="o", linestyle="-", color="b")
axes[0].set_title("Accuracy vs Window Size")
axes[0].set_xlabel("Window Size")
axes[0].set_ylabel("Accuracy")
axes[0].grid(True)
# Plot Weighted F1
axes[1].plot(results_window_df.index, results_window_df["f1"], marker="o", linestyle="-", color="g")
axes[1].set_title("Weighted F1 vs Window Size")
axes[1].set_xlabel("Window Size")
axes[1].set_ylabel("Weighted F1 Score")
axes[1].grid(True)
plt.tight_layout()
plt.show()
[cell 50 markdown]
### **Observation: Although the weighted F1 score is highest for a window size of 10, such a small window limits contextual coverage. A window size of 20 achieves nearly identical accuracy and only a minimal drop in weighted F1, making it a more balanced and contextually rich choice for the final model.**
[cell 51 markdown]
# **Final Model Training and Classification**
We now train the definitive Word2Vec model using the optimal hyperparameters.
[cell 52 code]
# ============================================================
# cell_17
# PURPOSE: Train the FINAL Word2Vec + MLP pipeline using the
# optimal hyperparameters discovered in cells 14 and 16,
# then evaluate it thoroughly with a full classification report.
#
# THIS IS THE MAIN RESULT CELL for supervised classification.
# All previous cells were either:
# - Building understanding (cells 5-8: toy Word2Vec)
# - Setting up tools (cells 9-13: Gensim + TF-IDF + MLP)
# - Finding best settings (cells 14-16: hyperparameter tuning)
#
# NOW we use everything we learned to train ONE final model
# with the best settings and evaluate it properly.
#
# OPTIMAL HYPERPARAMETERS (from tuning):
# vector_size = 200 (best from cell_14 grid search)
# window = 20 (best from cell_16 grid search)
#
# EVALUATION METRIC USED: Classification Report
# This goes beyond just accuracy/F1 — it shows per-class
# precision, recall, and F1 so we can see exactly WHERE
# the model struggles (which sentiment class is hardest).
# ============================================================
from gensim.models import Word2Vec # final Word2Vec training
import numpy as np # array operations
from sklearn.linear_model import LogisticRegression # imported for
from sklearn.naive_bayes import GaussianNB # use in next
from sklearn.ensemble import RandomForestClassifier # cell (baselines)
from sklearn.metrics import accuracy_score, f1_score # scalar metrics
from sklearn.metrics import classification_report # detailed per-class report
import pandas as pd # DataFrame display
import random, numpy as np, torch # re-imported for safety
# ---- Fix all random seeds for reproducibility ----
# This is the FINAL model — we want fully reproducible results
# that can be reported and compared consistently.
SEED = 42
random.seed(SEED) # Python random
np.random.seed(SEED) # NumPy random
torch.manual_seed(SEED) # PyTorch random
# ============================================================
# STEP 1: Train the final Word2Vec model
#
# Using the two best hyperparameters found through tuning:
# vector_size=200 → each word = 200-dim vector
# window=20 → considers 20 words of context each side
#
# All other settings are the same as throughout tuning,
# ensuring a fair and consistent final result.
# ============================================================
final_w2v = Word2Vec(
sentences=X_train, # training token lists only (no leakage)
vector_size=200, # ← best from cell_14 (vector_size tuning)
window=20, # ← best from cell_16 (window tuning)
min_count=1, # keep all vocabulary words
sg=1, # Skip-Gram algorithm
epochs=50, # 50 full passes through training data
seed=42 # reproducible weight initialization
)
# After this line, final_w2v contains optimally trained
# 200-dimensional word vectors for all words in X_train.
# ============================================================
# STEP 2: Build final TF-IDF weighted sentence vectors
#
# Uses tfidf_vocab from cell_12 (fitted on X_train only)
# and the freshly trained final_w2v model.
#
# X_train_vec shape: (N_train, 200) = (3876, 200)
# X_test_vec shape: (N_test, 200) = (969, 200)
# ============================================================
X_train_vec = np.vstack([tfidf_weighted_vector(tokens, final_w2v, tfidf_vocab)
for tokens in X_train])
X_test_vec = np.vstack([tfidf_weighted_vector(tokens, final_w2v, tfidf_vocab)
for tokens in X_test])
# ============================================================
# STEP 3: Build and train the final MLP classifier
# ============================================================
# input_dim = 200 (matches our vector_size)
# Derived from data rather than hardcoded so the code
# stays correct even if vector_size changes in future
input_dim = X_train_vec.shape[1] # = 200
num_classes = len(set(y_train)) # = 3
# Create a fresh MLP with the final architecture
mlp = MLPClassifier(
input_dim, # 200 input neurons (one per embedding dim)
hidden_dim=64, # 64 hidden neurons (learned features)
num_classes=num_classes # 3 output neurons (one per sentiment class)
)
# Train the MLP using the training embeddings
# epochs=10, lr=1e-3 are the same settings used throughout
# tuning — keeping them consistent ensures the improvement
# in final results comes from better embeddings, not from
# different training settings
train_model(mlp, X_train_vec, y_train, epochs=10, lr=1e-3)
# ============================================================
# STEP 4: Evaluate with scalar metrics (accuracy + F1)
# ============================================================
# evaluate_model() from cell_13 returns two floats
acc_mlp, f1_mlp = evaluate_model(mlp, X_test_vec, y_test)
print(f"\nFinal MLP Results")
print(f"Accuracy: {acc_mlp:.4f}") # e.g., 0.7823
print(f"Weighted F1: {f1_mlp:.4f}") # e.g., 0.7791
# These are the headline numbers — the single-value summaries
# of overall model performance on the held-out test set.
# ============================================================
# STEP 5: Generate the full Classification Report
#
# classification_report() shows a detailed breakdown of
# performance PER CLASS — much more informative than a
# single accuracy or F1 number.
#
# For each sentiment class it shows:
#
# PRECISION: Of all sentences the model predicted as
# "positive", what fraction were actually positive?
# = TP / (TP + FP)
# High precision = few false alarms
#
# RECALL: Of all actual "positive" sentences, what
# fraction did the model correctly identify?
# = TP / (TP + FN)
# High recall = few misses
#
# F1-SCORE: Harmonic mean of precision and recall
# = 2 × (precision × recall) / (precision + recall)
# Balances both concerns into one number
#
# SUPPORT: How many test sentences belong to this class
# (tells us if the dataset is imbalanced)
#
# Example output to expect:
#
# precision recall f1-score support
# negative 0.71 0.65 0.68 145
# neutral 0.80 0.85 0.82 582 ← most data
# positive 0.74 0.70 0.72 242
# accuracy 0.78 969
# macro avg 0.75 0.73 0.74 969
# weighted avg 0.78 0.78 0.78 969
# ============================================================
# First we need predictions as a numpy array for sklearn's
# classification_report (evaluate_model returns scalars,
# not the raw predictions, so we re-run inference here).
# Convert test embeddings to a PyTorch tensor
X_test_tensor = torch.tensor(X_test_vec, dtype=torch.float32)
# torch.no_grad(): no gradient tracking needed (inference only)
with torch.no_grad():
# mlp(X_test_tensor) → raw scores shape: (N_test, 3)
# torch.argmax(..., dim=1) → predicted class indices (0,1,2)
# shape: (N_test,)
# .numpy() converts PyTorch tensor → numpy array
# (classification_report expects numpy/list, not tensor)
preds = torch.argmax(mlp(X_test_tensor), dim=1).numpy()
print("\nClassification Report (MLP):")
print(
classification_report(
y_test, # true labels (integers 0,1,2)
preds, # predicted labels (integers 0,1,2)
target_names=encoder.classes_ # maps 0→"negative",
# 1→"neutral",
# 2→"positive"
# (encoder.classes_ from cell_2)
)
)
# ============================================================
# HOW TO INTERPRET THE REPORT:
#
# Look at SUPPORT first — it reveals class imbalance:
# neutral has the most samples → model learns it best
# negative has the fewest → model struggles most here
#
# Then look at RECALL per class:
# Low recall for "negative" → model MISSES negative sentences
# (it predicts them as neutral instead)
# This matters in finance: missing bad news = dangerous!
#
# Compare PRECISION vs RECALL per class:
# High precision, low recall → model is conservative
# (only predicts this class when very sure)
# Low precision, high recall → model over-predicts this class
# (predicts it even when wrong)
#
# The WEIGHTED AVG row = the overall weighted F1 from
# evaluate_model() — these two numbers should match.
# ============================================================
[cell 53 markdown]
# **Comparsion with baselines**
compare three different classifiers:
- Logistic Regression
- Gaussian Naive Bayes
- Random Forest
[cell 54 code]
# ============================================================
# cell_18
# PURPOSE: Compare the MLP against 3 classical ML classifiers
# (Logistic Regression, Gaussian NB, Random Forest) all using
# the SAME final TF-IDF weighted Word2Vec embeddings.
#
# WHY COMPARE AGAINST BASELINES?
# A model's performance number alone (e.g., F1=0.78) means
# nothing in isolation. We need CONTEXT:
# - Is 0.78 good or bad for this task?
# - Does the complexity of an MLP actually help?
# - Could a simpler model do just as well?
#
# By testing all 4 models on IDENTICAL inputs (same
# X_train_vec and X_test_vec), any difference in results
# comes PURELY from the model choice — not from different
# data, features, or preprocessing.
#
# THE 4 MODELS COMPARED:
#
# Logistic Regression:
# → Learns a LINEAR decision boundary
# → Fast, interpretable, strong baseline for text tasks
# → If this matches MLP, the data is linearly separable
#
# Gaussian Naive Bayes:
# → Assumes each feature (embedding dim) is INDEPENDENT
# and follows a Gaussian (normal) distribution
# → Very fast, works surprisingly well on text
# → "Naive" because the independence assumption is wrong
# (embedding dimensions ARE correlated) but it still works
#
# Random Forest:
# → Ensemble of 200 decision trees, each trained on a
# random subset of data and features
# → Robust to outliers, handles non-linear patterns
# → n_estimators=200: uses 200 trees (more = more stable)
#
# MLP (from cell_17):
# → Neural network with one hidden layer
# → Can learn non-linear patterns via ReLU activation
# → Results carried over from cell_17 (acc_mlp, f1_mlp)
# ============================================================
# ============================================================
# STEP 1: Define all 3 classical classifiers
#
# Each is stored in a dictionary so we can loop over them
# cleanly instead of repeating the same train/evaluate code
# 3 times. dict key = model name (for display), value = model.
# ============================================================
classifiers = {
# LogisticRegression: despite its name, it's a CLASSIFIER.
# Learns weights for a linear combination of input features,
# then uses sigmoid/softmax to convert to class probabilities.
#
# max_iter=1000: allow up to 1000 optimization steps to
# converge (default 100 is often too few for 200-dim input,
# causing a ConvergenceWarning — 1000 fixes this)
# random_state=42: reproducibility for the solver's
# random initialization
"Logistic Regression": LogisticRegression(
max_iter=1000, # enough iterations to fully converge
random_state=42 # reproducibility
),
# GaussianNB: Naive Bayes assuming Gaussian feature distributions.
# For each class, it learns the MEAN and VARIANCE of each
# embedding dimension, then uses Bayes' theorem to predict.
# No hyperparameters needed — it's fully determined by the data.
# Very fast but assumes features are independent (they aren't,
# but it still works reasonably well in practice).
"Gaussian NB": GaussianNB(),
# RandomForestClassifier: builds many decision trees and
# takes a majority vote across all of them.
#
# n_estimators=200: build 200 trees
# → more trees = more stable predictions, less variance
# → 200 is a good balance of quality vs training time
# → fewer trees (e.g., 10) = faster but noisier results
# random_state=42: reproducibility for tree construction
"Random Forest": RandomForestClassifier(
n_estimators=200, # number of trees in the forest
random_state=42 # reproducibility
),
}
# ---- Storage for all results ----
# Will hold results for all 4 models (3 classical + MLP)
# Same structure as cells 14 and 16:
# {"model_name": {"accuracy": float, "f1": float}}
results_final = {}
# ============================================================
# STEP 2: Train and evaluate each classical classifier
# ============================================================
for name, clf in classifiers.items():
# name → the string key: "Logistic Regression", etc.
# clf → the model object: LogisticRegression(...), etc.
# clf.fit(): trains the classifier on the training embeddings.
# Unlike PyTorch (which uses mini-batch gradient descent),
# sklearn classifiers train in a single .fit() call.
# X_train_vec: shape (3876, 200) — sentence embedding matrix
# y_train: shape (3876,) — integer labels (0, 1, 2)
clf.fit(X_train_vec, y_train)
# clf.predict(): runs inference on test embeddings.
# Returns an array of predicted class integers.
# shape: (969,) — one prediction per test sentence
# Note: sklearn predict() returns numpy array directly,
# unlike PyTorch where we had to manually take argmax.
preds = clf.predict(X_test_vec)
# Compute scalar evaluation metrics
# accuracy_score: fraction of correct predictions
acc = accuracy_score(y_test, preds)
# f1_score with average="weighted": weighted by class support
# Same metric used throughout for consistent comparison
f1 = f1_score(y_test, preds, average="weighted")
# Store results under the model's name
results_final[name] = {"accuracy": acc, "f1": f1}
# Print results immediately so we can monitor progress
# (Random Forest with 200 trees can take a moment)
print(f"{name}: Accuracy={acc:.4f}, Weighted F1={f1:.4f}")
# Expected output pattern:
# Logistic Regression: Accuracy=0.7689, Weighted F1=0.7651
# Gaussian NB: Accuracy=0.7124, Weighted F1=0.7089
# Random Forest: Accuracy=0.7401, Weighted F1=0.7355
# ============================================================
# STEP 3: Add the MLP results from cell_17
#
# The MLP was already trained and evaluated in cell_17.
# We stored its results in acc_mlp and f1_mlp.
# We simply add those values to results_final here so
# ALL 4 models appear in the same comparison table.
#
# NOTE: The key is "MLP " (with a trailing space) — this is
# a minor inconsistency in the original notebook but it works
# fine since it's just a display label.
# ============================================================
results_final["MLP "] = {"accuracy": acc_mlp, "f1": f1_mlp}
# ============================================================
# STEP 4: Build and display the comparison DataFrame
#
# pd.DataFrame(results_final): builds DataFrame from dict.
# Without orient="index", the dict keys become COLUMNS.
# .T (transpose): flips rows and columns so:
# → model names become ROW indices (easier to read)
# → "accuracy" and "f1" become COLUMNS
#
# Final DataFrame layout:
#
# accuracy | f1
# ─────────────────────────────|──────────
# Logistic Regression 0.7689 | 0.7651
# Gaussian NB 0.7124 | 0.7089 ← weakest
# Random Forest 0.7401 | 0.7355
# MLP 0.7823 | 0.7791 ← strongest
# ============================================================
results_final_df = pd.DataFrame(results_final).T
print("\nFinal Comparison (All Models):")
results_final_df # Jupyter displays this as a formatted table
[cell 55 markdown]
### **Observation:**
* Among the traditional classifiers, Logistic Regression and Gaussian NB perform comparably, while Random Forest lags behind on both accuracy and F1.
* The MLP outperforms all baselines, achieving the highest accuracy and weighted F1.
* This confirms that the neural model benefits more from the dense Word2Vec embeddings than the classical models.
[cell 56 markdown]
# **3 Word2Vec Embeddings for Unsupervised Clustering with K-Means**
Finally, we use K-Means clustering to see if the Word2Vec embeddings can reveal natural groupings in the data without supervision. We apply PCA for visualization to see if these clusters align with the known sentiment labels.
[cell 57 code]
# ============================================================
# cell_19
# PURPOSE: Set up the data for UNSUPERVISED CLUSTERING —
# the third and final part of the notebook.
#
# SHIFT IN APPROACH — Supervised → Unsupervised:
# Everything up to cell_18 was SUPERVISED learning:
# → we gave the model labels ("positive", "negative", etc.)
# → the model learned to predict those labels
#
# Now we ask a different question:
# "If we had NO labels, could Word2Vec embeddings alone
# naturally group sentences into meaningful clusters?"
#
# This is UNSUPERVISED learning — no labels used at all.
# We let the algorithm discover structure on its own,
# then check afterwards if those structures align with
# the known sentiment labels.
#
# KEY DIFFERENCE FROM CLASSIFICATION SETUP:
# For clustering we use the FULL dataset (not just X_train)
# because:
# 1. There are no labels to leak — we're not predicting
# anything, just grouping
# 2. More data = richer co-occurrence statistics = better
# Word2Vec embeddings = better cluster structure
# 3. We want to cluster ALL sentences, not just 80% of them
#
# This cell handles the setup:
# - Train Word2Vec on ALL sentences (not just X_train)
# - Fit TF-IDF on ALL sentences
# - Build sentence vectors for ALL sentences
# The actual clustering happens in the next cell.
# ============================================================
from gensim.models import Word2Vec # Word2Vec model
from sklearn.feature_extraction.text import TfidfVectorizer # IDF weights
import numpy as np # array operations
from sklearn.cluster import KMeans # clustering (used next cell)
import matplotlib.pyplot as plt # plotting (used next cell)
# ============================================================
# STEP 1: Train Word2Vec on the FULL dataset
#
# df['tokens'] = ALL preprocessed token lists (not just train)
# This includes both the training AND test sentences.
#
# WHY is this acceptable here (whereas it wasn't for classification)?
#
# In classification (cells 9-18):
# Training on full data would let the classifier "see"
# test sentence labels during training → data leakage →
# inflated evaluation scores → dishonest results.
#
# In clustering (this cell):
# There is no train/test split — we're not evaluating
# prediction accuracy on held-out data.
# We're asking "what natural groups exist in ALL the data?"
# More data always helps Word2Vec learn better embeddings.
# No leakage is possible because we're not predicting labels.
#
# Hyperparameters: same optimal values from cells 14 & 16
# vector_size=200, window=20 — proven best for this dataset
# ============================================================
print("Training Word2Vec on the entire dataset")
w2v_model_clus = Word2Vec(
sentences=df['tokens'], # ALL sentences (train + test combined)
vector_size=200, # optimal from cell_14
window=20, # optimal from cell_16
min_count=1, # keep all vocabulary words
sg=1, # Skip-Gram algorithm
epochs=50, # 50 full training passes
seed=42 # reproducibility
)
# w2v_model_clus is a SEPARATE model from final_w2v (cell_17)
# Named differently (_clus suffix) to avoid overwriting the
# classification model which might still be needed
# ============================================================
# STEP 2: Fit TF-IDF on the FULL dataset
#
# Same reasoning as above — for clustering, we can and
# should use all available data to compute IDF weights.
#
# Using the full dataset for IDF means:
# → IDF weights reflect true word rarity across ALL sentences
# → Rare-but-important words get correctly high weights
# → No artificial distortion from using only 80% of data
#
# Note: tfidf_vocab is OVERWRITTEN here (was previously fitted
# on X_train only in cell_12). This is intentional — the
# clustering pipeline needs full-dataset IDF weights.
# ============================================================
print("Fitting TF-IDF on the entire dataset")
# Convert ALL token lists → strings for TfidfVectorizer
# e.g., ["profit", "rose"] → "profit rose"
docs_all = [" ".join(tokens) for tokens in df['tokens']]
# docs_all: list of strings, one per sentence, ALL sentences
# Fit TF-IDF on all documents to learn full-corpus IDF weights
tfidf = TfidfVectorizer()
tfidf.fit(docs_all) # learns IDF from ALL sentences
# Extract word → IDF weight mapping as a dictionary
# (same approach as cell_12 but now fitted on full data)
# tfidf_vocab["profit"] → IDF weight based on ALL sentences
# tfidf_vocab["bankruptcy"] → higher (rarer across all sentences)
tfidf_vocab = dict(zip(tfidf.get_feature_names_out(), tfidf.idf_))
# ============================================================
# STEP 3: Redefine tfidf_weighted_vector()
#
# This function is IDENTICAL to the one in cell_12.
# It is redefined here to make this clustering section
# self-contained — it works with w2v_model_clus and the
# new full-corpus tfidf_vocab without depending on cell_12.
#
# Input: tokens → cleaned word list for one sentence
# model → Word2Vec model (w2v_model_clus here)
# tfidf_vocab→ {word: idf_weight} dictionary
# Output: one 200-dim numpy array representing the sentence
# ============================================================
def tfidf_weighted_vector(tokens, model, tfidf_vocab):
vectors = [] # will store idf_weight × word_vector per word
weights = [] # will store idf_weight per word (for denominator)
for word in tokens:
# Skip word if missing from either vocabulary
# (can't compute weighted vector without both)
if word in model.wv and word in tfidf_vocab:
idf_weight = tfidf_vocab[word]
# Scale word's embedding by its IDF weight
# rare words (high IDF) → contribute MORE
# common words (low IDF) → contribute LESS
vectors.append(model.wv[word] * idf_weight)
weights.append(idf_weight)
# Edge case: no valid words → return zero vector
if not vectors:
return np.zeros(model.vector_size) # shape: (200,)
# Weighted average:
# sum of (idf × vector) / sum of idf weights
# shape: (200,)
return np.sum(vectors, axis=0) / np.sum(weights)
# ============================================================
# STEP 4: Build sentence vectors for ALL sentences
#
# Apply tfidf_weighted_vector() to every row in df['tokens'].
# np.array([...]) stacks results into a 2D matrix.
#
# Unlike classification (used np.vstack), here we use
# np.array() directly since we're building from a list
# comprehension over df rows — both approaches are equivalent
# for this use case.
#
# X_all_w2v shape: (total_sentences, 200)
# Each ROW = the TF-IDF weighted embedding of one sentence
# ============================================================
print("Creating TF-IDF weighted document vectors for the entire dataset...")
X_all_w2v = np.array([
tfidf_weighted_vector(tokens, w2v_model_clus, tfidf_vocab)
for tokens in df['tokens']
])
# Confirm the shape:
# X_all_w2v.shape[0] = total number of sentences (e.g., 4845)
# X_all_w2v.shape[1] = embedding dimensions = 200
print(f"Created {X_all_w2v.shape[0]} document vectors "
f"of size {X_all_w2v.shape[1]}.")
# Expected: "Created 4845 document vectors of size 200."
[cell 58 code]
# ============================================================
# cell_20
# PURPOSE: Use the ELBOW METHOD to determine the optimal
# number of clusters (k) for K-Means clustering.
#
# THE CORE PROBLEM:
# K-Means requires us to specify k (number of clusters)
# BEFORE running the algorithm. But we don't know upfront
# how many natural groups exist in our data.
# The Elbow Method helps us make an educated choice.
#
# WHAT IS K-MEANS?
# K-Means partitions all sentences into k groups such that
# each sentence belongs to the cluster whose CENTER
# (centroid) is closest to it.
#
# It works by repeating two steps until stable:
# 1. ASSIGN: each sentence → nearest centroid
# 2. UPDATE: recompute each centroid as the mean of
# all sentences currently assigned to it
#
# WHAT IS INERTIA?
# Inertia = the total "compactness" of all clusters
# = sum of squared distances from each point to its
# nearest cluster centroid
#
# Low inertia → points are CLOSE to their centroids
# → tight, compact clusters (good)
# High inertia → points are FAR from their centroids
# → loose, spread-out clusters (bad)
#
# WHAT IS THE ELBOW METHOD?
# We run K-Means for k = 2, 3, 4, ... 10 and plot
# inertia vs k. As k increases, inertia always decreases
# (more clusters = each cluster is smaller and tighter).
# But there's usually a point where adding more clusters
# gives DIMINISHING RETURNS — the "elbow" in the curve.
# That elbow point = the optimal k.
#
# Visual intuition:
#
# Inertia
# ▲
# │●
# │ ●
# │ ● ← "elbow" here = optimal k
# │ ●──●──●──● (curve flattens after this)
# └──────────────→ k
# 2 3 4 5 6
# ============================================================
print("\nRunning Elbow Method to find the optimal number of clusters")
# ============================================================
# STEP 1: Compute inertia for each k from 2 to 10
# ============================================================
# inertia list will store one inertia value per k
# After the loop: inertia[0] = inertia for k=2,
# inertia[1] = inertia for k=3, etc.
inertia = []
# K_range = [2, 3, 4, 5, 6, 7, 8, 9, 10]
# We start at k=2 because k=1 (all one cluster) is trivial
# We stop at k=10 because our dataset has only 3 true classes
# — going beyond 10 is unlikely to reveal anything useful
K_range = range(2, 11)
for k in K_range:
# KMeans() creates a fresh K-Means model for this k.
#
# n_clusters=k : number of clusters to form
#
# random_state=42 : K-Means starts with RANDOM centroids.
# Different random starts can give different results.
# Fixing the seed ensures the same result every run,
# making the elbow curve reproducible and comparable.
#
# n_init="auto" : how many times to run K-Means with
# different random centroid initializations.
# K-Means can get stuck in local optima, so running
# it multiple times and keeping the best result is safer.
# "auto" lets sklearn choose a sensible default
# (replaces the old n_init=10 default in newer sklearn
# versions to avoid a FutureWarning).
km = KMeans(n_clusters=k, random_state=42, n_init="auto")
# km.fit() runs the full K-Means algorithm on all 4845
# sentence vectors until convergence (centroids stop moving).
# X_all_w2v shape: (4845, 200)
# After fitting, km stores the final cluster assignments
# and the final centroid positions.
km.fit(X_all_w2v)
# km.inertia_ is computed automatically after fit().
# It is the SUM OF SQUARED DISTANCES from each point
# to its nearest cluster centroid.
#
# Formula:
# inertia = Σ Σ ||x - centroid_j||²
# j x∈cluster_j
#
# where ||...||² = squared Euclidean distance
#
# This single number summarises how "tight" the clusters
# are for this value of k — lower is better.
inertia.append(km.inertia_)
# After this loop, inertia contains 9 values:
# [inertia_k2, inertia_k3, ..., inertia_k10]
# inertia always DECREASES as k increases because with
# more clusters, each point is closer to a centroid.
# At k = number_of_points, inertia = 0 (each point IS
# its own centroid) — but that's useless clustering.
# ============================================================
# STEP 2: Plot the Elbow Curve
# ============================================================
plt.figure(figsize=(8, 6))
# "bo-" is a format string combining 3 style codes:
# "b" → blue color
# "o" → circle marker at each data point
# "-" → solid line connecting the points
# This makes each measured k value clearly visible as a dot
# while the connecting line shows the overall trend.
plt.plot(K_range, inertia, "bo-")
# K_range → x-axis: the k values [2,3,4,...,10]
# inertia → y-axis: the inertia at each k
plt.xlabel("Number of clusters (k)") # x-axis: what we varied
plt.ylabel("Inertia") # y-axis: what we measured
plt.title("Elbow Method for Optimal k")
# plt.xticks(K_range) forces the x-axis to show a tick mark
# at EVERY integer k value (2,3,4,...,10) instead of
# matplotlib's default which might skip some values.
# Makes it easy to read exactly which k each dot corresponds to.
plt.xticks(K_range)
plt.grid(True) # gridlines help read exact inertia values
plt.show()
# ============================================================
# HOW TO READ THE OUTPUT PLOT:
#
# The curve ALWAYS goes down as k increases.
# Look for the point where the curve BENDS — like an elbow.
# After that point, adding more clusters gives rapidly
# diminishing returns in compactness.
#
# Expected shape for this dataset:
#
# Inertia
# ▲
# │● ← k=2: one big drop (2 clusters much better than 1)
# │
# │ ● ← k=3: second big drop (3 clusters clearly better)
# │
# │ ●──●──●──●──●──●──● ← curve flattens here
# └──────────────────────→ k
# 2 3 4 5 6 7 8 9 10
# ↑
# ELBOW at k=3
# → this is our optimal k
# → makes intuitive sense: we have 3 sentiment classes!
#
# NOTEBOOK OBSERVATION:
# "The largest drop occurs between k=2 and k=3,
# forming an elbow point — hence k=3 is the most
# justified choice."
# ============================================================
[cell 59 code]
# ============================================================
# cell_21
# PURPOSE: Display the exact inertia values from cell_20's
# Elbow Method as a clean, readable table.
#
# WHY THIS CELL EXISTS AFTER THE PLOT:
# The elbow curve plot (cell_20) shows the TREND visually,
# but it can be hard to read exact numbers from a graph.
# This table gives us the PRECISE inertia value for each k
# so we can:
# 1. Quantify exactly how much inertia drops between k values
# 2. Make a more objective decision about the elbow point
# 3. Have a concrete number to report or reference
#
# The drop between consecutive k values tells us how much
# each additional cluster improves compactness:
# Large drop → that k is meaningful (real structure found)
# Small drop → diminishing returns (k too large)
# ============================================================
# ============================================================
# STEP 1: Build the inertia DataFrame
#
# We combine two lists into a dictionary, then convert to
# a DataFrame for clean tabular display.
#
# Both lists come from cell_20:
# K_range = range(2, 11) → [2, 3, 4, 5, 6, 7, 8, 9, 10]
# inertia = [val_k2, val_k3, ..., val_k10] (9 values)
#
# list(K_range) converts range object → plain Python list
# so it can be stored in the DataFrame as a regular column.
# ============================================================
inertia_df = pd.DataFrame({
"k" : list(K_range), # column 1: the k values [2,3,...,10]
"inertia": inertia # column 2: inertia at each k
})
# inertia_df layout:
# k | inertia
# ─────|──────────
# 2 | XXXXX.X ← highest (only 2 clusters, loosest)
# 3 | XXXXX.X ← big drop from k=2 (elbow here)
# 4 | XXXXX.X ← smaller drop
# 5 | XXXXX.X
# ...
# 10 | XXXXX.X ← lowest (10 clusters, tightest)
print("\nInertia values:")
# Displaying inertia_df as the last expression in the cell
# causes Jupyter to render it as a formatted HTML table
# (cleaner than print()).
inertia_df
# ============================================================
# HOW TO READ THIS TABLE:
#
# Look at the DROP in inertia between consecutive rows:
#
# drop(k=2→3) = inertia[k=2] - inertia[k=3] ← largest drop
# drop(k=3→4) = inertia[k=3] - inertia[k=4] ← much smaller
# drop(k=4→5) = inertia[k=4] - inertia[k=5] ← even smaller
# ...
#
# The k where the drop SUDDENLY BECOMES SMALL is the elbow.
#
# Example (approximate values for this dataset):
#
# k | inertia | drop from previous k
# ─────|────────────|──────────────────────
# 2 | 185,000 | —
# 3 | 162,000 | 23,000 ← LARGEST drop → elbow at k=3
# 4 | 151,000 | 11,000 ← drop halved
# 5 | 143,000 | 8,000 ← getting smaller
# 6 | 137,000 | 6,000 ← diminishing returns
# 7 | 132,000 | 5,000
# 8 | 128,000 | 4,000
# 9 | 125,000 | 3,000
# 10 | 122,000 | 3,000 ← almost flat
#
# CONCLUSION:
# The largest single drop is between k=2 and k=3.
# After k=3, each additional cluster gives less and less
# improvement → k=3 is the elbow → k=3 is our choice.
#
# This aligns perfectly with our domain knowledge:
# we have exactly 3 sentiment classes (positive, negative,
# neutral) → the data's natural structure has 3 groups.
# ============================================================
[cell 60 markdown]
**Observation:** The inertia decreases as (k) increases, but the largest drop occurs between (k=2) and (k=3), forming an elbow point — hence (k=3) is the most justified choice.
[cell 61 markdown]
# **Visualizing the Clusters**
We project the high-dimensional TF-IDF weighted Word2Vec document vectors into 2D and 3D using Principal Component Analysis (PCA). This allows us to compare the ground truth sentiment labels with the clusters discovered by K-Means for k = [2, 3, 4, 5].
[cell 62 markdown]
# **2D PCA**
[cell 63 code]
# ============================================================
# cell_22
# PURPOSE: Visually compare K-Means clusters against the
# true sentiment labels using 2D PCA projections.
#
# THIS IS THE KEY VISUAL VALIDATION CELL for clustering.
# Numbers alone (inertia, silhouette score) don't tell the
# whole story — seeing the clusters spatially reveals:
# - Do the K-Means clusters ALIGN with true sentiments?
# - Do the 3 sentiment classes actually SEPARATE in space?
# - Does adding more clusters (k=4,5) reveal sub-groups
# or just arbitrarily split existing groups?
#
# LAYOUT: 2 rows × 3 columns = 6 subplot panels:
# Panel (0,0): Ground truth — colored by TRUE sentiment label
# Panel (0,1): KMeans k=2
# Panel (0,2): KMeans k=3 ← most important (matches elbow)
# Panel (1,0): KMeans k=4
# Panel (1,1): KMeans k=5
# Panel (1,2): empty (turned off — only 5 plots needed)
#
# The comparison between panel (0,0) and panel (0,2) is the
# core question: does k=3 K-Means recover the true sentiment
# groupings without ever being told the labels?
# ============================================================
from sklearn.decomposition import PCA # dimensionality reduction
from sklearn.cluster import KMeans # clustering algorithm
import matplotlib.pyplot as plt # plotting
import matplotlib.patches as mpatches # for custom legend handles
# ============================================================
# STEP 1: Reduce 200D → 2D using PCA
#
# X_all_w2v has shape (4845, 200) — impossible to plot.
# PCA finds the 2 directions of maximum variance and
# projects all sentence vectors onto those 2 axes.
#
# IMPORTANT: PCA is applied to X_all_w2v (200D vectors),
# NOT to the cluster assignments. We cluster in 200D space
# (full information) and only use 2D PCA for VISUALIZATION.
# Clustering in 2D would lose most of the structure.
#
# n_components=2 : project down to 2 dimensions
# random_state=42 : reproducibility (PCA uses randomized
# SVD for large matrices)
# ============================================================
print("Reducing to 2D with PCA")
pca_2d = PCA(n_components=2, random_state=42)
# fit_transform: learns the 2 principal components from
# X_all_w2v and immediately projects all vectors onto them.
# Input: (4845, 200)
# Output: (4845, 2) — each sentence now has 2 coordinates
X_pca_2d = pca_2d.fit_transform(X_all_w2v)
# X_pca_2d[i] = [x_coord, y_coord] for sentence i
# X_pca_2d[:, 0] = all x coordinates (PC1 values)
# X_pca_2d[:, 1] = all y coordinates (PC2 values)
# ============================================================
# STEP 2: Run K-Means for k = 2, 3, 4, 5
#
# We cluster in the ORIGINAL 200D space (not the 2D PCA
# space) to get the most accurate cluster assignments.
# The 2D PCA coordinates are only used for plotting.
#
# For each k, we store the cluster label array (shape: 4845,)
# where cluster_results[k][i] = which cluster sentence i
# was assigned to (an integer from 0 to k-1).
# ============================================================
print("Running KMeans for k = 2, 3, 4, 5")
k_values = [2, 3, 4, 5]
cluster_results = {} # {k: array of cluster assignments}
for k in k_values:
km = KMeans(
n_clusters=k,
random_state=42, # reproducibility
n_init="auto" # auto-select number of random restarts
)
# fit_predict(): trains K-Means AND returns cluster
# assignment for every point in one call.
# Equivalent to: km.fit(X_all_w2v); km.labels_
# Output shape: (4845,) — integer 0..k-1 per sentence
cluster_results[k] = km.fit_predict(X_all_w2v)
# After loop:
# cluster_results[2] → array([0,1,1,0,...]) assignments for k=2
# cluster_results[3] → array([2,0,1,2,...]) assignments for k=3
# cluster_results[4] → array([3,1,0,2,...]) assignments for k=4
# cluster_results[5] → array([1,4,2,0,...]) assignments for k=5
# ============================================================
# STEP 3: Create the 2×3 subplot grid
# ============================================================
# plt.subplots(2, 3): 2 rows, 3 columns = 6 panels total
# figsize=(18, 10): wide enough to show all 6 panels clearly
# axes is a 2×3 numpy array of Axes objects
# Access individual panels as axes[row, col]
fig, axes = plt.subplots(2, 3, figsize=(18, 10))
# ============================================================
# STEP 4: Panel (0,0) — Ground Truth scatter plot
#
# This is the REFERENCE panel — it shows what the correct
# clustering SHOULD look like if the algorithm were perfect.
# All other panels are compared against this one.
# ============================================================
# Map each sentiment string → a display color
# positive → green (intuitive: green = good news)
# negative → red (intuitive: red = bad news)
# neutral → blue (neutral tone)
color_map = {
"positive": "green",
"negative": "red",
"neutral" : "blue"
}
# df['Sentiment'].map(color_map) replaces each sentiment string
# with its corresponding color string for every sentence.
# Result: a Series of color strings, one per sentence.
# e.g., ["green", "blue", "blue", "red", "green", ...]
sentiment_colors = df['Sentiment'].map(color_map)
# scatter plot: each dot = one sentence
# x = PC1 coordinate, y = PC2 coordinate
# c = color based on true sentiment label
# alpha=0.6: 60% opaque so overlapping points are visible
axes[0, 0].scatter(
X_pca_2d[:, 0], # x: PC1 values for all 4845 sentences
X_pca_2d[:, 1], # y: PC2 values for all 4845 sentences
c=sentiment_colors, # color each dot by its true sentiment
alpha=0.6 # semi-transparent for overlap visibility
)
axes[0, 0].set_title("Ground Truth (Sentiment Labels)")
# Build a custom legend using mpatches.Patch (colored squares)
# mpatches.Patch(color=c, label=s) creates a colored rectangle
# for the legend — one per sentiment class.
# axes.legend(handles=...) uses our custom patches instead of
# the default legend which wouldn't work here since we used
# c= (color array) rather than multiple scatter() calls.
handles = [mpatches.Patch(color=c, label=s)
for s, c in color_map.items()]
axes[0, 0].legend(handles=handles, title="Sentiment")
# ============================================================
# STEP 5: Panels (0,1), (0,2), (1,0), (1,1) — K-Means plots
#
# One panel per k value in k_values = [2, 3, 4, 5]
# plot_positions maps each k to its panel location.
# ============================================================
# plot_positions: the (row, col) position for each k's panel
# fills panels left-to-right, top-to-bottom after (0,0):
# k=2 → (0,1) k=3 → (0,2) k=4 → (1,0) k=5 → (1,1)
plot_positions = [(0, 1), (0, 2), (1, 0), (1, 1)]
for pos, k in zip(plot_positions, k_values):
# pos = (row, col) tuple for this panel
# k = number of clusters for this panel
ax = axes[pos] # get the specific Axes object for this panel
# scatter plot: each dot = one sentence
# c=cluster_results[k]: colors dots by their K-Means
# cluster assignment (integers 0..k-1)
# cmap="viridis": maps cluster integers → colors using
# the viridis colormap (purple→blue→green→yellow)
# Each cluster integer gets a distinct color automatically.
# alpha=0.7: slightly more opaque than ground truth
# (cluster boundaries are what we want to see clearly)
scatter = ax.scatter(
X_pca_2d[:, 0], # x: same PC1 coordinates
X_pca_2d[:, 1], # y: same PC2 coordinates
c=cluster_results[k], # color by K-Means cluster id
cmap="viridis", # colormap for cluster colors
alpha=0.7
)
ax.set_title(f"KMeans (k={k})")
# No explicit legend needed — the title tells us k,
# and the colors just show which cluster each point
# belongs to (the exact color doesn't matter, only
# that same-color points are in the same cluster)
# ============================================================
# STEP 6: Turn off the empty 6th panel (bottom-right)
#
# We have 5 plots (1 ground truth + 4 K-Means) but a
# 2×3 grid has 6 panels. The last panel (1,2) is unused.
# axes[1,2].axis("off") hides all axis elements (ticks,
# labels, border) making it appear as blank white space.
# ============================================================
axes[1, 2].axis("off")
# ============================================================
# STEP 7: Final layout and display
# ============================================================
# tight_layout() adjusts spacing between panels so titles
# and axis labels don't overlap each other.
plt.tight_layout()
plt.show()
# ============================================================
# HOW TO READ THE OUTPUT:
#
# Compare panel (0,0) [Ground Truth] with panel (0,2) [k=3]:
#
# If K-Means perfectly recovered sentiment:
# → k=3 panel would look identical to ground truth
# → each K-Means color would match one sentiment color
#
# What we actually expect to see:
# → PARTIAL alignment: k=3 clusters roughly correspond
# to sentiment groups but with significant overlap
# → The three sentiment classes don't form perfectly
# separated blobs in 2D PCA space
# → Positive and negative sentences often share similar
# financial vocabulary, making them hard to separate
#
# Why k=4 and k=5 panels matter:
# → If k=4 looks cleaner than k=3, maybe there are
# 4 natural sub-groups (e.g., strong vs mild sentiment)
# → If k=4 just splits an existing cluster arbitrarily,
# it confirms k=3 was the right choice
#
# NOTEBOOK OBSERVATION:
# "At k=3, clusters emerge that broadly correspond to
# the three sentiment groups, though with some overlap."
# This confirms that sentiment signal exists in the
# embeddings, but it is not strong enough for clean
# unsupervised separation.
# ============================================================
[cell 64 markdown]
### **Observation: At k=3, clusters emerge that broadly correspond to the three sentiment groups, though with some overlap.**
[cell 65 markdown]
# **3D PCA**
[cell 66 code]
# ============================================================
# cell_23
# PURPOSE: Repeat the cluster visualization from cell_22
# but in 3D instead of 2D, to see if the extra dimension
# reveals clearer separation between sentiment groups.
#
# WHY 3D AFTER ALREADY DOING 2D?
# PCA to 2D keeps only the top 2 directions of variance.
# Some structure that exists in 200D space might be
# "hidden" in the 3rd principal component — compressed
# out of the 2D view but visible when we add PC3.
#
# Think of it like a shadow analogy:
# A 3D object cast onto a 2D wall loses depth info.
# Rotating or viewing from a different angle (3D plot)
# might reveal structure that the flat shadow hid.
#
# DIFFERENCES FROM CELL_22:
# - PCA reduces to 3D instead of 2D
# - Each subplot uses projection="3d" (interactive 3D axes)
# - scatter() takes 3 coordinate arrays instead of 2
# - Subplot positions use 3-digit integers (231-235)
# instead of (row,col) tuples — different API for
# fig.add_subplot() vs plt.subplots()
# - No empty 6th panel (2×3 grid fills all 5 used slots
# with 1 ground truth + 4 K-Means, panel 236 unused
# but not explicitly turned off)
#
# LAYOUT: same 2×3 grid as cell_22
# 231: Ground Truth 232: k=2 233: k=3
# 234: k=4 235: k=5 (236: unused/empty)
# ============================================================
from sklearn.decomposition import PCA # dimensionality reduction
from sklearn.cluster import KMeans # clustering
import matplotlib.pyplot as plt # plotting
import matplotlib.patches as mpatches # custom legend patches
from mpl_toolkits.mplot3d import Axes3D # enables 3D subplot support
# MUST be imported even though
# it's not called directly —
# it registers the "3d" projection
# with matplotlib's backend
# ============================================================
# STEP 1: Reduce 200D → 3D using PCA
#
# Same as cell_22 but n_components=3 instead of 2.
# We now keep the top 3 directions of maximum variance.
#
# PC1 = direction of MOST variance in the embedding space
# PC2 = direction of SECOND-MOST variance (⊥ to PC1)
# PC3 = direction of THIRD-MOST variance (⊥ to PC1 & PC2)
#
# The extra PC3 captures structure that was discarded in
# the 2D version — it may reveal z-axis separation between
# sentiment groups that looked overlapping in 2D.
# ============================================================
print("Reducing to 3D with PCA")
pca_3d = PCA(n_components=3, random_state=42)
# Input shape: (4845, 200) — all sentence vectors
# Output shape: (4845, 3) — 3 coordinates per sentence
X_pca_3d = pca_3d.fit_transform(X_all_w2v)
# X_pca_3d[:, 0] = PC1 values (x-axis in 3D plot)
# X_pca_3d[:, 1] = PC2 values (y-axis in 3D plot)
# X_pca_3d[:, 2] = PC3 values (z-axis in 3D plot) ← new!
# ============================================================
# STEP 2: Run K-Means for k = 2, 3, 4, 5
#
# IDENTICAL to cell_22 — clustering is always done in
# the full 200D space for maximum accuracy.
# Results are stored for use in the 3D scatter plots below.
# ============================================================
print("Running KMeans for k = 2, 3, 4, 5")
k_values = [2, 3, 4, 5]
cluster_results = {}
for k in k_values:
km = KMeans(
n_clusters=k,
random_state=42, # reproducibility
n_init="auto" # auto number of random restarts
)
# fit_predict: train K-Means AND return cluster assignments
# Output shape: (4845,) — integer 0..k-1 per sentence
cluster_results[k] = km.fit_predict(X_all_w2v)
# ============================================================
# STEP 3: Create the figure
#
# KEY DIFFERENCE FROM CELL_22:
# Cell_22 used: fig, axes = plt.subplots(2, 3)
# → good for 2D subplots, returns axes as 2D array
#
# This cell uses: fig = plt.figure()
# → necessary for 3D subplots because we must specify
# projection="3d" individually per subplot via
# fig.add_subplot() — plt.subplots() doesn't support
# mixed projections easily.
#
# figsize=(18, 12): taller than cell_22 (12 vs 10) because
# 3D plots need more vertical space for the z-axis labels
# ============================================================
fig = plt.figure(figsize=(18, 12))
# ============================================================
# STEP 4: Panel 231 — Ground Truth in 3D
#
# Subplot position integers work as follows:
# 231 = 2 rows, 3 columns, panel number 1 (top-left)
# 232 = 2 rows, 3 columns, panel number 2 (top-middle)
# ... and so on, filling left→right, top→bottom
#
# projection="3d" tells matplotlib to create a 3D axes
# object with x, y, AND z axes instead of the standard 2D.
# ============================================================
ax = fig.add_subplot(231, projection="3d")
# 231: row 1, col 1 (top-left panel)
# projection="3d": enables 3D rotation and z-axis
# Color each sentence by its true sentiment label
color_map = {
"positive": "green",
"negative": "red",
"neutral" : "blue"
}
sentiment_colors = df['Sentiment'].map(color_map)
# 3D scatter: requires THREE coordinate arrays
# X_pca_3d[:, 0] → x-axis (PC1)
# X_pca_3d[:, 1] → y-axis (PC2)
# X_pca_3d[:, 2] → z-axis (PC3) ← extra dimension vs cell_22
ax.scatter(
X_pca_3d[:, 0], # x: PC1
X_pca_3d[:, 1], # y: PC2
X_pca_3d[:, 2], # z: PC3 ← new dimension
c=sentiment_colors, # color by true sentiment
alpha=0.6 # semi-transparent for overlap visibility
)
ax.set_title("Ground Truth (Sentiment Labels)")
# Custom legend using colored patches — same as cell_22
handles = [mpatches.Patch(color=c, label=s)
for s, c in color_map.items()]
ax.legend(handles=handles, title="Sentiment")
# ============================================================
# STEP 5: Panels 232-235 — K-Means results in 3D
#
# positions list uses 3-digit subplot integers instead of
# (row,col) tuples. Numbered panels fill left→right then
# top→bottom:
# 232: top-middle (k=2)
# 233: top-right (k=3) ← most important
# 234: bottom-left (k=4)
# 235: bottom-middle (k=5)
# ============================================================
positions = [232, 233, 234, 235]
for pos, k in zip(positions, k_values):
# Add a new 3D axes at this position
ax = fig.add_subplot(pos, projection="3d")
# 3D scatter colored by K-Means cluster assignment
# c=cluster_results[k]: integers 0..k-1 per sentence
# cmap="viridis": maps integers → distinct colors
ax.scatter(
X_pca_3d[:, 0], # x: PC1
X_pca_3d[:, 1], # y: PC2
X_pca_3d[:, 2], # z: PC3
c=cluster_results[k], # color by K-Means cluster id
cmap="viridis", # colormap: purple→green→yellow
alpha=0.7 # slightly more opaque than ground truth
)
ax.set_title(f"KMeans (k={k})")
# Panel 236 (bottom-right) is left empty automatically —
# matplotlib simply leaves it blank since we never
# called fig.add_subplot(236)
# ============================================================
# STEP 6: Final layout and display
# ============================================================
plt.tight_layout() # prevent title/label overlap between subplots
plt.show()
# ============================================================
# HOW TO READ THE 3D PLOTS:
#
# Each 3D scatter plot can be mentally "rotated" to look
# for separation along any axis combination:
# PC1 vs PC2 → same view as cell_22's 2D plot
# PC1 vs PC3 → new view, may show different separation
# PC2 vs PC3 → another new view
#
# Signs of GOOD clustering (hope):
# → Colored blobs that barely touch each other
# → K-Means panel (k=3) looks similar to Ground Truth
# → Clear separation along z-axis (PC3) that wasn't
# visible in the 2D cell_22 plots
#
# Signs of POOR clustering (reality for this dataset):
# → All colors heavily mixed together in a central blob
# → K-Means clusters look like arbitrary spatial cuts
# rather than following the sentiment color boundaries
# → Adding the z-axis (PC3) doesn't help much
#
# NOTEBOOK OBSERVATION:
# "At k=3, clusters emerge that broadly correspond to the
# three sentiment groups, though with some overlap."
# The 3D view confirms this — sentiment classes form a
# loose cloud structure with significant mixing, consistent
# with the poor clustering metrics in the next cell.
# ============================================================
[cell 67 markdown]
### **Observation: At k=3, clusters emerge that broadly correspond to the three sentiment groups, though with some overlap.**
[cell 68 markdown]
# **Clustering Evaluation Metrics**
[cell 69 markdown]
### Understanding Clustering Evaluation Metrics
Evaluation metrics for clustering are generally divided into two categories: **Intrinsic** (based on the geometry of the data) and **Extrinsic** (based on comparison to ground truth labels).
#### 1. Intrinsic Metrics (No Ground Truth Needed)
These metrics evaluate how well-defined the clusters are based solely on the distribution and distance of the data points.
* **Silhouette Score**
* **What it is:** Measures how similar a point is to its own cluster (cohesion) compared to other clusters (separation).
* **Range:** $[-1, 1]$
* **Good Range:** Near **+1** (Clusters are dense and well-separated).
* **Bad Range:** Near **0** (Overlapping clusters) or **Negative** (Points assigned to the wrong cluster).
* **Davies-Bouldin Score**
* **What it is:** The average similarity between clusters (ratio of within-cluster distance to between-cluster distance).
* **Range:** $[0, \infty)$
* **Good Range:** **Lower is better** (closer to 0).
* **Bad Range:** Higher values indicate the clusters are "leaking" into each other.
* **Calinski-Harabasz Score**
* **What it is:** The ratio of the sum of between-clusters dispersion to within-cluster dispersion.
* **Range:** $[0, \infty)$
* **Good Range:** **Higher is better**. Indicates clusters are dense and well-separated.
---
#### 2. Extrinsic Metrics (Ground Truth Required)
These are used when you have actual labels (like the Positive/Negative labels in the PDF) to see if the algorithm discovered those specific groups.
* **Homogeneity Score**
* **What it is:** Does each cluster contain only members of a single class?
* **Range:** $[0, 1]$
* **Interpretation:** **1.0** is perfect.
* **Completeness Score**
* **What it is:** Are all members of a given class assigned to the same cluster?
* **Range:** $[0, 1]$
* **Interpretation:** **1.0** is perfect.
* **V-Measure Score**
* **What it is:** The harmonic mean of Homogeneity and Completeness. It provides a balanced "total" score.
* **Range:** $[0, 1]$
* **Interpretation:** **1.0** is perfect.
---
### Summary Table
| Metric | Category | Range | Ideal Value |
| :--- | :--- | :--- | :--- |
| **Silhouette** | Intrinsic | -1 to +1 | +1 |
| **Davies-Bouldin** | Intrinsic | 0 to ∞ | 0 |
| **Calinski-Harabasz** | Intrinsic | 0 to ∞ | Higher |
| **Homogeneity** | Extrinsic | 0 to 1 | 1 |
| **Completeness** | Extrinsic | 0 to 1 | 1 |
| **V-Measure** | Extrinsic | 0 to 1 | 1 |
[cell 70 code]
# ============================================================
# cell_24
# PURPOSE: Quantitatively evaluate the quality of the k=3
# K-Means clustering using 6 different metrics split into
# two categories:
#
# INTRINSIC METRICS (no labels needed):
# Measure cluster quality purely from the data geometry —
# how tight and well-separated the clusters are in space.
# Used when you have NO ground truth labels.
#
# EXTRINSIC METRICS (labels required):
# Measure how well clusters ALIGN with known true labels.
# Used when you DO have ground truth (like we do here).
#
# WHY USE BOTH?
# Intrinsic metrics answer: "Are the clusters geometrically good?"
# Extrinsic metrics answer: "Do the clusters match sentiment?"
# A clustering can be geometrically tight (good intrinsic)
# but still not match sentiment labels (poor extrinsic) —
# meaning the natural groupings in the data don't correspond
# to sentiment, even if they're real groupings of some kind.
#
# THIS IS THE FINAL CELL — it delivers the verdict on whether
# unsupervised clustering can recover sentiment structure.
# ============================================================
from sklearn.metrics import (
# --- Intrinsic metrics ---
silhouette_score, # measures cluster separation vs cohesion
davies_bouldin_score, # measures average cluster similarity
calinski_harabasz_score, # measures cluster density and separation
# --- Extrinsic metrics ---
homogeneity_score, # are clusters pure (one sentiment each)?
completeness_score, # are sentiments contained in one cluster?
v_measure_score # harmonic mean of homogeneity + completeness
)
import pandas as pd
# ============================================================
# STEP 1: Run final K-Means with k=3
#
# k=3 chosen from the Elbow Method in cells 20-21.
# n_init=10: run K-Means 10 times with different random
# centroid initializations, keep the best result.
# (Here n_init=10 is explicit rather than "auto" — both
# are fine; 10 is the traditional sklearn default.)
# ============================================================
k = 3
kmeans = KMeans(
n_clusters=k, # 3 clusters (matches our 3 sentiment classes)
random_state=42, # reproducibility
n_init=10 # 10 random restarts, keep best result
)
# fit_predict: trains K-Means on all sentence vectors AND
# returns the cluster assignment (0, 1, or 2) for each sentence.
# y_pred_kmeans shape: (4845,) — one integer per sentence
y_pred_kmeans = kmeans.fit_predict(X_all_w2v)
# ============================================================
# INTRINSIC METRIC 1: Silhouette Score
#
# For each point, measures:
# a = average distance to OTHER points in its OWN cluster
# (cohesion — how tight is its cluster?)
# b = average distance to points in the NEAREST OTHER cluster
# (separation — how far is the next closest cluster?)
#
# silhouette = (b - a) / max(a, b)
#
# Range: -1 to +1
# +1 → point is perfectly placed: tight cluster,
# far from all other clusters (ideal)
# 0 → point is on the boundary between two clusters
# -1 → point is probably in the WRONG cluster
#
# Overall score = average silhouette across all points.
# Higher is better.
#
# Interpreting values:
# > 0.70 → strong cluster structure
# > 0.50 → reasonable structure
# > 0.25 → weak structure
# < 0.25 → no meaningful structure (expected here)
# ============================================================
sil_score = silhouette_score(X_all_w2v, y_pred_kmeans)
# X_all_w2v: the 200D sentence vectors (distances computed here)
# y_pred_kmeans: cluster assignment for each point
# ============================================================
# INTRINSIC METRIC 2: Davies-Bouldin Index
#
# For each cluster, computes the ratio of:
# (average distance within the cluster) ← scatter
# ──────────────────────────────────────
# (distance between cluster centroids) ← separation
#
# Then averages the WORST (maximum) ratio across all clusters.
#
# Range: 0 to ∞
# Lower is BETTER (opposite of Silhouette and CH!)
# 0 → perfect: clusters are compact AND far apart
# High value → clusters are loose OR too close together
#
# This metric penalises clusters that are either:
# - Internally spread out (high scatter)
# - Too close to neighbouring clusters (low separation)
# ============================================================
db_score = davies_bouldin_score(X_all_w2v, y_pred_kmeans)
# Lower = better (unlike most metrics)
# ============================================================
# INTRINSIC METRIC 3: Calinski-Harabasz Score (Variance Ratio)
#
# Ratio of:
# between-cluster variance (how spread apart clusters are)
# ────────────────────────────────────────────────────────
# within-cluster variance (how spread points are inside)
#
# Range: 0 to ∞
# Higher is BETTER
# A high ratio means clusters are dense AND well-separated.
#
# CAVEAT: This score tends to INCREASE with k regardless of
# true cluster structure — not very reliable for choosing k.
# It's included here as a secondary check, not for k selection.
# ============================================================
ch_score = calinski_harabasz_score(X_all_w2v, y_pred_kmeans)
# Higher = better
# ============================================================
# EXTRINSIC METRIC 1: Homogeneity
#
# Asks: "Does each cluster contain ONLY sentences from
# a single sentiment class?"
#
# A cluster is perfectly homogeneous if ALL its members
# share the same true sentiment label.
#
# Range: 0 to 1
# 1.0 → each cluster is pure (only one sentiment per cluster)
# 0.0 → clusters are completely mixed (random assignment)
#
# Example:
# Cluster 0: [positive, positive, positive] → homogeneous ✓
# Cluster 1: [neutral, negative, neutral] → NOT homogeneous ✗
# ============================================================
h_score = homogeneity_score(
df['Sentiment'], # true labels (strings: "positive" etc.)
y_pred_kmeans # predicted cluster assignments (0, 1, 2)
)
# ============================================================
# EXTRINSIC METRIC 2: Completeness
#
# Asks: "Are ALL sentences of a given sentiment class
# assigned to the SAME cluster?"
#
# A class is perfectly complete if all its members end up
# in one cluster (not scattered across multiple clusters).
#
# Range: 0 to 1
# 1.0 → all sentences of each sentiment are in one cluster
# 0.0 → sentences of each sentiment scattered everywhere
#
# Example:
# All "positive" sentences → Cluster 0 only → complete ✓
# "Neutral" sentences split → Cluster 0,1,2 → NOT complete ✗
#
# NOTE: Homogeneity and Completeness are COMPLEMENTARY:
# You can have high homogeneity with low completeness:
# → Put each sentence in its OWN cluster (k=N)
# → Every cluster is pure (homogeneous) ✓
# → But each sentiment is scattered across N clusters ✗
# ============================================================
c_score = completeness_score(df['Sentiment'], y_pred_kmeans)
# ============================================================
# EXTRINSIC METRIC 3: V-Measure
#
# The harmonic mean of Homogeneity and Completeness:
#
# 2 × homogeneity × completeness
# V-Measure = ─────────────────────────────────
# homogeneity + completeness
#
# This is the clustering equivalent of the F1-score —
# it balances both concerns into a single number.
#
# Range: 0 to 1
# 1.0 → perfect clustering (matches ground truth exactly)
# 0.0 → completely random assignment
#
# This is the SINGLE MOST INFORMATIVE extrinsic metric —
# it captures both whether clusters are pure AND whether
# each sentiment is kept together.
# ============================================================
v_score = v_measure_score(df['Sentiment'], y_pred_kmeans)
# ============================================================
# STEP 2: Print all 6 metrics
# ============================================================
print("Intrinsic Metrics:")
# Silhouette: higher better, range -1 to +1
print(f" Silhouette Score : {sil_score:.3f}")
# Davies-Bouldin: LOWER better, range 0 to ∞
print(f" Davies-Bouldin Index : {db_score:.3f}")
# Calinski-Harabasz: higher better, range 0 to ∞
print(f" Calinski-Harabasz Score: {ch_score:.3f}\n")
print("Extrinsic Metrics:")
# All three: higher better, range 0 to 1
print(f" Homogeneity : {h_score:.3f}")
print(f" Completeness : {c_score:.3f}")
print(f" V-Measure : {v_score:.3f}")
# ============================================================
# EXPECTED OUTPUT AND INTERPRETATION:
#
# Intrinsic Metrics:
# Silhouette Score : ~0.06 ← very low (near 0)
# Davies-Bouldin Index : ~2.8 ← high (bad, want low)
# Calinski-Harabasz Score: ~XXX ← unreliable for this task
#
# Extrinsic Metrics:
# Homogeneity : ~0.10 ← very low (clusters are impure)
# Completeness : ~0.10 ← very low (sentiments are scattered)
# V-Measure : ~0.10 ← very low (overall poor match)
#
# FINAL VERDICT:
# All metrics confirm the clustering largely FAILED to
# recover sentiment structure. Here's why:
#
# 1. OVERLAP IN EMBEDDING SPACE:
# Financial sentences of different sentiments share
# a lot of vocabulary ("company", "market", "said",
# "quarter", "revenue") — the embedding vectors for
# positive and negative sentences are very similar
# because the WORDS are similar, even if the MEANING differs.
# e.g., "profit rose" vs "profit fell" — same domain,
# opposite sentiment, but very similar word vectors.
#
# 2. SENTIMENT ≠ TOPIC:
# Word2Vec learns TOPICAL similarity (words in same context)
# not SENTIMENT polarity. It can't distinguish:
# "The company posted record profits" (positive)
# "The company posted record losses" (negative)
# Both sentences have identical structure and very
# similar word embeddings.
#
# 3. CONCLUSION:
# Supervised learning (cells 9-18) works because labels
# guide the model to learn the positive/negative distinction.
# Unsupervised clustering fails because that distinction
# is not the primary structure in the raw embedding space —
# topic/domain similarity dominates instead.
# ============================================================
[cell 71 markdown]
* **The clustering quality is very poor. Intrinsic metrics show weak structure (Silhouette ≈ 0.06, Davies–Bouldin very high, Calinski–Harabasz not reliable in this context). Extrinsic metrics also confirm failure to match ground truth (Homogeneity, Completeness, V-Measure ≈ 0.1). This is because the sentiment data has heavy overlap across classes and weak separation in vector space, K-Means fails to form meaningful clusters.**
[cell 72 markdown]
## Clustering Results — Observations
**Overall verdict: The clustering has largely failed to capture sentiment structure.**
---
### Intrinsic Metrics
**Silhouette Score — 0.060 (very poor)**
This is barely above 0, meaning most data points sit roughly equidistant between clusters rather than clearly belonging to one. A score this close to 0 is essentially what you would expect from a near-random assignment. There is almost no meaningful separation between clusters.
**Davies-Bouldin Index — 3.585 (very poor)**
Anything above 2 is generally considered poor, and 3.585 is more than 3x worse than a decent result. This confirms that clusters are overlapping heavily and are not compact internally. The algorithm has drawn boundaries, but those boundaries do not reflect real groupings in the data.
**Calinski-Harabasz Score — 325.796 (misleading here)**
This is the one number that looks reasonable on the surface, but it should not be taken at face value. CH is sensitive to dataset size and cluster count, and it can appear decent even when clusters are meaningless. Given that silhouette and Davies-Bouldin both tell a poor story, the CH score is likely inflated by overall data spread rather than reflecting genuine cluster quality.
---
### Extrinsic Metrics
**Homogeneity — 0.107 (poor)**
Clusters are heavily impure — each cluster contains a mix of multiple sentiment classes rather than being dominated by a single one.
**Completeness — 0.106 (poor)**
Sentiment classes are scattered across multiple clusters rather than being concentrated in one. The algorithm is splitting each sentiment group instead of keeping it together.
**V-Measure — 0.106 (poor)**
As the harmonic mean of homogeneity and completeness, this is the single most reliable summary score. At 0.106, it confirms that the clustering has almost no alignment with the true sentiment labels. A random assignment would perform at a similar level.
---
[cell 73 markdown]
### Reference Table for Evaluation metrices
| Metric | Category | Range | Ideal Value |
| :--- | :--- | :--- | :--- |
| **Silhouette** | Intrinsic | -1 to +1 | +1 |
| **Davies-Bouldin** | Intrinsic | 0 to ∞ | 0 |
| **Calinski-Harabasz** | Intrinsic | 0 to ∞ | Higher |
| **Homogeneity** | Extrinsic | 0 to 1 | 1 |
| **Completeness** | Extrinsic | 0 to 1 | 1 |
| **V-Measure** | Extrinsic | 0 to 1 | 1 |