# texteda bow tfidf v1 (1)
course: Module 3 — Deep Learning & NLP
module: Module-3-Deep-Learning-NLP
type: notebook
source_url: https://personal-learn.armco.dev/files/Module-3-Deep-Learning-NLP/General/Lab_Materials-05.04.2026/texteda_bow_tfidf_v1_(1).ipynb
---
[cell 1 markdown]
**Financial PhraseBank Dataset**
The **Financial PhraseBank** dataset contains short sentences extracted from financial news articles related to publicly traded companies. Each sentence is labeled with a sentiment that reflects how the information would likely impact investor perceptions of the company’s stock.
In this study, we focus on **preliminary text data analysis**, which involves performing **exploratory data analysis (EDA)**, and applying **Bag-of-Words (BoW)** and **TF–IDF** representations for both **sentiment classification** and **unsupervised clustering**.
Source: **Kaggle Dataset**
* [Kaggle: Financial Sentiment Analysis](https://www.kaggle.com/datasets/sbhatti/financial-sentiment-analysis)
Target Variable:
* **sentiment** *(categorical)*: Indicates the polarity of the financial news sentence:
* `positive`
* `negative`
* `neutral`
---
Features:
| Column Name | Description | Data Type |
| ----------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------- |
| sentence | A short text segment extracted from financial news. Each sentence is self-contained and typically refers to a specific company, event, or market movement. | *Text* |
| sentiment | Label showing whether the news is expected to have a **positive**, **negative**, or **neutral** impact on the company’s stock or market perception. | *Categorical* |
Reference paper:[Malo et al., Good Debt or Bad Debt, JASIST 2014](https://arxiv.org/abs/1307.5336)
[cell 2 markdown]
# **Importing the necessary libraries for data manipulation, visualization, and text processing.**
[cell 3 code]
# Core libraries
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
import re
import string
# NLTK for text processing
import nltk
from nltk.corpus import stopwords
from nltk.stem import WordNetLemmatizer
from wordcloud import WordCloud
# Scikit-learn for feature extraction, modeling, and evaluation
from sklearn.model_selection import train_test_split
from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.naive_bayes import MultinomialNB
from sklearn.ensemble import RandomForestClassifier
from sklearn.cluster import KMeans
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE
from sklearn.metrics import (
classification_report, confusion_matrix, accuracy_score,f1_score,
silhouette_score, davies_bouldin_score, calinski_harabasz_score,
homogeneity_score, completeness_score, v_measure_score
)
import warnings
warnings.filterwarnings("ignore")
# Download NLTK resources
nltk.download('stopwords')
nltk.download('wordnet')
# Set plotting style
sns.set_style("whitegrid")
plt.rcParams['figure.figsize'] = (10, 6)
[cell 4 markdown]
# **Load the dataset**
[cell 5 code]
# Load the dataset
df = pd.read_csv('data.csv')
# Display the head
df.head(10)
[cell 6 markdown]
**Basic information about the dataset**
[cell 7 code]
df.info()
[cell 8 markdown]
**Check for missing values**
[cell 9 code]
print("Missing values in each column:\n", df.isnull().sum())
[cell 10 markdown]
The dataset is loaded successfully with 5842 entries and two columns: 'Sentence' and 'Sentiment'. There are no missing values, so we can proceed with the analysis.
[cell 11 code]
stop_words = set(ENGLISH_STOP_WORDS)
print(f"--- Sample Data by Class ---\n")
for sentiment in df['Sentiment'].unique():
sample = df[df['Sentiment'] == sentiment].head(5)
print(f"**Class: {sentiment.upper()}**\n")
for _, row in sample.iterrows():
print(f"Input: {row['Sentence']}")
print(f"Output: {row['Sentiment']}\n")
[cell 12 markdown]
## **1. Exploratory Data Analysis (EDA)**
[cell 13 markdown]
# **Sentiment Distribution**
First, let's examine the distribution of the three sentiment labels: `neutral`, `positive`, and `negative`. This is crucial to check for class imbalance, which can affect model training.
[cell 14 code]
# Calculate the value counts
sentiment_counts = df['Sentiment'].value_counts(normalize=True)
print(sentiment_counts*100)
# Plot the sentiment distribution
plt.figure(figsize=(8, 5))
sns.countplot(x='Sentiment', data=df, order=sentiment_counts.index, palette='viridis')
plt.title('Distribution of Sentiments', fontsize=16)
plt.xlabel('Sentiment', fontsize=12)
plt.ylabel('Number of Headlines', fontsize=12)
plt.xticks(fontsize=11)
plt.show()
[cell 15 markdown]
**Observation:**
The dataset is clearly imbalanced. The `neutral` class is the dominant class, with significantly more samples than the `positive` and `negative` classes combined.
[cell 16 markdown]
# **Text Characteristics: Sentence Length**
Analyzing the length of the headlines can provide insights into their structure. We'll look at the distribution of word count and character count.
[cell 17 code]
df['word_count'] = df['Sentence'].apply(lambda x: len(str(x).split())) # split on spaces and count the words
df['char_count'] = df['Sentence'].apply(len)
df[['Sentence', 'word_count', 'char_count']].head()
[cell 18 code]
# Plot sentence length distribution (in words)
plt.figure(figsize=(12, 5))
plt.subplot(1, 2, 1)
sns.histplot(df['word_count'], bins=30, kde=True, color='skyblue')
plt.title('Sentence Length Distribution (Word Count)')
plt.xlabel('Word Count')
plt.ylabel('Frequency')
plt.subplot(1, 2, 2)
sns.histplot(df['char_count'], bins=30, kde=True, color='salmon')
plt.title('Sentence Length Distribution (Character Count)')
plt.xlabel('Character Count')
plt.ylabel('Frequency')
plt.tight_layout()
plt.show()
[cell 19 markdown]
**Observations from Sentence Length Distribution:**
The histograms provide a clear picture of the typical length of the financial headlines in this dataset:
* **Word Count (Left Plot):** The distribution is right-skewed, meaning that most headlines are short, with a long tail of fewer, longer headlines.
* The vast majority of sentences contain between **10 and 30 words**.
* The peak of the distribution (the mode) is around **15 words per headline**.
* It's rare to see headlines longer than 40-50 words, indicating a concise writing style.
* **Character Count (Right Plot):** This plot mirrors the word count distribution, also showing a right skew.
* Most headlines fall within the **50 to 150 character** range.
* The peak frequency is for headlines around **75-100 characters** in length.
[cell 20 markdown]
# **word count per sentiment category**
[cell 21 code]
# Analyze sentence length per sentiment
plt.figure(figsize=(10, 6))
sns.boxplot(x='Sentiment', y='word_count', data=df, palette='Set3')
plt.title('Word Count Distribution per Sentiment')
plt.xlabel('Sentiment')
plt.ylabel('Word Count')
plt.show()
[cell 22 markdown]
**Observations:**
* Most headlines are between 10 and 30 words long.
* The distribution of word count is fairly similar across all three sentiments. There isn't a strong indication that sentence length alone is a powerful predictor of sentiment.
[cell 23 markdown]
# **Word Frequency Analysis**
[cell 24 code]
# Calculate sentence length (Character Count & Word Count)
df['char_count'] = df['Sentence'].apply(len)
df['word_count'] = df['Sentence'].apply(lambda x: len(str(x).split()))
# Function to clean text and count vocabulary (remove Stop words)
def get_top_words(text_series, top_n=10):
all_text = ' '.join(text_series).lower()
# Remove special characters, keep only a-z and $
all_text = re.sub(r'[^a-z\$]', ' ', all_text)
words = [w for w in all_text.split() if w not in ENGLISH_STOP_WORDS and len(w) > 1]
return Counter(words).most_common(top_n)
# ==========================================
# Display numerical statistics (Terminal Output)
# ==========================================
print("--- Basic Information ---")
print(f"Total number of records: {len(df)} sentences")
print(f"Number of missing values:\n{df.isnull().sum()}\n")
print("--- Class Distribution ---")
class_counts = df['Sentiment'].value_counts()
class_pct = df['Sentiment'].value_counts(normalize=True) * 100
for sentiment, count in class_counts.items():
print(f"Class {sentiment.capitalize()}: {count} sentences ({class_pct[sentiment]:.1f}%)")
print("\n--- Text Length ---")
print(f"Average length (characters): {df['char_count'].mean():.1f} characters")
print(f"Average length (words): {df['word_count'].mean():.1f} words")
print(f"Longest sentence: {df['word_count'].max()} words | Shortest: {df['word_count'].min()} words")
print("Average word length per class:")
print(df.groupby('Sentiment')['word_count'].mean().round(1))
print("\n--- Most Frequent Words (Top Words) ---")
print(f"Overall: {get_top_words(df['Sentence'], 8)}")
for sentiment in df['Sentiment'].unique():
print(f"Class {sentiment.capitalize()}: {get_top_words(df[df['Sentiment'] == sentiment]['Sentence'], 8)}")
[cell 25 code]
import matplotlib.pyplot as plt
def plot_top_words_horizontal(text_series, top_n=10, title="Top Words"):
top_words = get_top_words(text_series, top_n)
words = [w[0] for w in top_words]
counts = [w[1] for w in top_words]
plt.figure(figsize=(8, 6))
plt.barh(words, counts)
plt.title(title)
plt.xlabel("Frequency")
plt.ylabel("Words")
# Highest frequency on top (better visualization)
plt.gca().invert_yaxis()
# Add value labels
for i, v in enumerate(counts):
plt.text(v + 0.5, i, str(v), va='center')
plt.tight_layout()
plt.show()
# For Overall Top Words
plot_top_words_horizontal(
df['Sentence'],
top_n=10,
title="Top 10 Most Frequent Words (Overall)"
)
# For Top Words Per Sentiment Class
for sentiment in df['Sentiment'].unique():
plot_top_words_horizontal(
df[df['Sentiment'] == sentiment]['Sentence'],
top_n=10,
title=f"Top Words - {sentiment.capitalize()}"
)
[cell 26 markdown]
**N-gram Analysis**
N-grams are contiguous sequences of *n* items from a given sample of text. By analyzing the most common unigrams (1-grams), bigrams (2-grams), and trigrams (3-grams), we can identify key phrases associated with each sentiment.
[cell 27 code]
def get_top_ngrams(corpus, n, top_k=20):
"""
generates and returns the top k n-grams from a given corpus.
"""
# Use CountVectorizer to get n-gram counts
vec = CountVectorizer(ngram_range=(n, n), stop_words='english').fit(corpus)
bag_of_words = vec.transform(corpus)
sum_words = bag_of_words.sum(axis=0)
words_freq = [(word, sum_words[0, idx]) for word, idx in vec.vocabulary_.items()]
words_freq = sorted(words_freq, key=lambda x: x[1], reverse=True)
return words_freq[:top_k]
[cell 28 code]
# Get top unigrams for each sentiment
unigrams_pos = get_top_ngrams(df[df['Sentiment']=='positive']['Sentence'], 1)
unigrams_neg = get_top_ngrams(df[df['Sentiment']=='negative']['Sentence'], 1)
unigrams_neu = get_top_ngrams(df[df['Sentiment']=='neutral']['Sentence'], 1)
# Create dataframes for plotting
df_unigrams_pos = pd.DataFrame(unigrams_pos, columns=['Unigram', 'Frequency'])
df_unigrams_neg = pd.DataFrame(unigrams_neg, columns=['Unigram', 'Frequency'])
df_unigrams_neu = pd.DataFrame(unigrams_neu, columns=['Unigram', 'Frequency'])
# Plotting
fig, axes = plt.subplots(3, 1, figsize=(8, 14))
sns.barplot(x='Frequency', y='Unigram', data=df_unigrams_pos, ax=axes[0], palette='Greens_d')
axes[0].set_title('Top 20 Unigrams for Positive Sentiment')
sns.barplot(x='Frequency', y='Unigram', data=df_unigrams_neg, ax=axes[1], palette='Reds_d')
axes[1].set_title('Top 20 Unigrams for Negative Sentiment')
sns.barplot(x='Frequency', y='Unigram', data=df_unigrams_neu, ax=axes[2], palette='Blues_d')
axes[2].set_title('Top 20 Unigrams for Neutral Sentiment')
plt.tight_layout()
plt.show()
[cell 29 markdown]
**Unigram Analysis Observations:**
1. **Significant Overlap of Core Vocabulary:**
* A key finding is the substantial overlap of common financial terms across all three sentiments. Words like `eur`, `mn` (million), `sales`, `profit`, `company`, `year`, and `net` are prominent in positive, negative, and neutral headlines.
* This indicates that these core nouns are context-dependent; their sentiment is determined by the surrounding words (verbs, adjectives). For example, `profit` can appear in a positive context ("profit rose") or a negative one ("profit fell").
2. **Distinct Sentiment Indicators:**
* **Positive Sentiment:** While sharing the core vocabulary, the positive list contains clear indicators of growth and success. The word `rose` is a powerful and unambiguous positive signal. The presence of `operating` often precedes positive metrics like profit or income.
* **Negative Sentiment:** Similarly, the negative list contains strong, distinct indicators of poor performance. The words `loss` and `decreased` are clear signals of negative financial news.
* **Neutral Sentiment:** The neutral list is characterized by descriptive and factual language. The high frequency of `said` suggests a focus on reporting statements or quotes. Words like `company`, `group`, and `business` are more prominent here, indicating a focus on factual announcements rather than performance evaluation.
[cell 30 markdown]
# **Bi-gram Anlaysis**
[cell 31 code]
# Get top bigrams for each sentiment
bigrams_pos = get_top_ngrams(df[df['Sentiment']=='positive']['Sentence'], 2)
bigrams_neg = get_top_ngrams(df[df['Sentiment']=='negative']['Sentence'], 2)
bigrams_neu = get_top_ngrams(df[df['Sentiment']=='neutral']['Sentence'], 2)
# Create dataframes for plotting
df_bigrams_pos = pd.DataFrame(bigrams_pos, columns=['Bigram', 'Frequency'])
df_bigrams_neg = pd.DataFrame(bigrams_neg, columns=['Bigram', 'Frequency'])
df_bigrams_neu = pd.DataFrame(bigrams_neu, columns=['Bigram', 'Frequency'])
# Plotting
fig, axes = plt.subplots(3, 1, figsize=(8, 14))
sns.barplot(x='Frequency', y='Bigram', data=df_bigrams_pos, ax=axes[0], palette='Greens_d')
axes[0].set_title('Top 20 Bigrams for Positive Sentiment')
sns.barplot(x='Frequency', y='Bigram', data=df_bigrams_neg, ax=axes[1], palette='Reds_d')
axes[1].set_title('Top 20 Bigrams for Negative Sentiment')
sns.barplot(x='Frequency', y='Bigram', data=df_bigrams_neu, ax=axes[2], palette='Blues_d')
axes[2].set_title('Top 20 Bigrams for Neutral Sentiment')
plt.tight_layout()
plt.show()
[cell 32 markdown]
**Bigram Analysis Observations:**
**1. Core Financial Phrases (The Common Ground)**
Across all three sentiments, we see a strong presence of standard financial reporting phrases. Bigrams like:
* `net sales`
* `operating profit`
* `eur mn` (and `mn eur`)
* `corresponding period`
These phrases represent the common **subjects** of financial news. They are inherently neutral and their sentiment is determined entirely by the context in which they appear. Their high frequency explains the initial similarity of the plots.
**2. The Emergence of Sentiment-Carrying Phrases (The Key Differentiators)**
The crucial differences lie in the phrases that describe an **action or state**.
* **For Positive Sentiment:** The model can learn from unambiguous phrases that signal growth. We see the clear emergence of:
* `profit rose`
* `rose eur`
These bigrams combine a subject (`profit`, `eur`) with a positive action (`rose`), providing a powerful signal.
* **For Negative Sentiment:** The negative chart contains the direct counterparts to the positive signals. Key phrases include:
* `operating loss` (the direct opposite of `operating profit`)
* `decreased eur`
These combinations of a subject and a negative action are strong indicators of negative sentiment.
* **For Neutral Sentiment:** The neutral bigrams confirm their factual, non-evaluative nature. They are dominated by:
* **Reporting Phrases:** `company said`, `said today`
* **Proper Nouns & Entities:** `omx helsinki`, `alma media`, `stock exchange`
* **Structural Phrases:** `board directors`, `share capital`
These phrases focus on reporting events and facts rather than evaluating performance.
[cell 34 markdown]
# **Tri-gram Anlaysis**
[cell 35 code]
# Get top trigrams for each sentiment
trigrams_pos = get_top_ngrams(df[df['Sentiment']=='positive']['Sentence'], 3)
trigrams_neg = get_top_ngrams(df[df['Sentiment']=='negative']['Sentence'], 3)
trigrams_neu = get_top_ngrams(df[df['Sentiment']=='neutral']['Sentence'], 3)
# Create dataframes for plotting
df_trigrams_pos = pd.DataFrame(trigrams_pos, columns=['Trigram', 'Frequency'])
df_trigrams_neg = pd.DataFrame(trigrams_neg, columns=['Trigram', 'Frequency'])
df_trigrams_neu = pd.DataFrame(trigrams_neu, columns=['Trigram', 'Frequency'])
# Plotting
fig, axes = plt.subplots(3, 1, figsize=(8, 14))
sns.barplot(x='Frequency', y='Trigram', data=df_trigrams_pos, ax=axes[0], palette='Greens_d')
axes[0].set_title('Top 20 Trigrams for Positive Sentiment')
sns.barplot(x='Frequency', y='Trigram', data=df_trigrams_neg, ax=axes[1], palette='Reds_d')
axes[1].set_title('Top 20 Trigrams for Negative Sentiment')
sns.barplot(x='Frequency', y='Trigram', data=df_trigrams_neu, ax=axes[2], palette='Blues_d')
axes[2].set_title('Top 20 Trigrams for Neutral Sentiment')
plt.tight_layout()
plt.show()
[cell 36 markdown]
# **Analysis of Uppercase Words**
The use of uppercase words can sometimes indicate urgency or strong sentiment.
[cell 37 code]
# Function to count uppercase words
def count_uppercase_words(text):
return len(re.findall(r'\b[A-Z]{2,}\b', text))
# Apply the function to create a new column
df['uppercase_word_count'] = df['Sentence'].apply(count_uppercase_words)
# Plot the distribution of uppercase words per sentiment
plt.figure(figsize=(10, 6))
sns.barplot(x='Sentiment', y='uppercase_word_count', data=df, estimator=np.mean, ci=None)
plt.title('Average Uppercase Words per Sentence by Sentiment')
plt.xlabel('Sentiment')
plt.show()
[cell 38 markdown]
**Observations**
1. **Negative sentences contain the highest average number of uppercase words per sentence**, slightly more than positive ones. This suggests that negative financial news often emphasizes entities (e.g., company tickers, acronyms) more strongly.
2. **Neutral sentences have the fewest uppercase words on average**, indicating that they are less entity-driven and use fewer tickers/acronyms compared to sentiment-laden sentences.
[cell 39 code]
# Display sentences with a high count of uppercase words
print("Examples of headlines with high uppercase word count:")
df.sort_values(by='uppercase_word_count', ascending=False).head(10)
[cell 40 code]
import re
from collections import Counter
import matplotlib.pyplot as plt
import seaborn as sns
# Function to extract all uppercase words (length >= 2)
def extract_uppercase_words(text):
return re.findall(r'\b[A-Z]{2,}\b', text)
# Apply to entire corpus
all_uppercase_words = df['Sentence'].apply(extract_uppercase_words).sum()
# Count frequencies
word_freq = Counter(all_uppercase_words).most_common(20)
# Convert to DataFrame for plotting
upper_df = pd.DataFrame(word_freq, columns=['Word', 'Frequency'])
#print("Top 20 Uppercase Tokens (Tickers/Acronyms):")
#print(upper_df)
# Plot
plt.figure(figsize=(10,6))
sns.barplot(x='Frequency', y='Word', data=upper_df, palette="viridis")
plt.title("Top 20 Uppercase Words in Financial PhraseBank", fontsize=16)
plt.xlabel("Frequency")
plt.ylabel("Token")
plt.show()
[cell 41 markdown]
**Observations**
1. **EUR dominates overwhelmingly** as the most frequent uppercase token, appearing far more often than any other ticker or acronym, which highlights the dataset’s strong focus on Euro-denominated financial news.
2. Other frequent uppercase tokens include **company tickers (AAPL, TSLA, FB, UPM, SPY)** and **financial acronyms (USD, EPS, CEO, FTSE)**, showing that uppercase words are largely entity-driven and represent key financial instruments or organizations.
[cell 42 markdown]
# **Word Cloud Visualization**
[cell 43 code]
import matplotlib.pyplot as plt
from sklearn.feature_extraction.text import ENGLISH_STOP_WORDS
from wordcloud import WordCloud
# Use sklearn's stopwords for consistency
custom_stopwords = set(ENGLISH_STOP_WORDS)
# Function to generate wordcloud for a given sentiment
def generate_wordcloud(data, sentiment, ax):
text = " ".join(data[data['Sentiment'] == sentiment]['Sentence']).lower()
wordcloud = WordCloud(width=800, height=400,
background_color='white',
stopwords=custom_stopwords,
repeat=False ,
colormap='viridis').generate(text)
ax.imshow(wordcloud, interpolation='bilinear')
ax.set_title(f"{sentiment.capitalize()} Sentiment", fontsize=14)
ax.axis("off")
# Create 1x3 subplots
fig, axes = plt.subplots(3, 1, figsize=(8,17))
# Generate wordclouds for each sentiment
generate_wordcloud(df, 'positive', axes[0])
generate_wordcloud(df, 'negative', axes[1])
generate_wordcloud(df, 'neutral', axes[2])
plt.tight_layout()
plt.show()
[cell 44 markdown]
# **Observations**
**Common financial terms dominate across all sentiments** : words like *eur, mn, company, finnish, year, sales* appear prominently in positive, negative, and neutral texts, reflecting the dataset’s financial focus.
[cell 45 markdown]
#**2. Data preparation and vectorization and Sentiment Classification**
[cell 46 markdown]
Before feeding the text data into our models, we need to clean it. This process involves several steps to standardize the text. Our preprocessing pipeline will include:
1. **Lowercasing:** Convert all text to lowercase so that words like *Profit* and *profit* are treated the same.
2. **Handling Punctuation and Numbers:**
* Remove punctuation marks that do not contribute to meaning.
* Preserve finance-specific symbols such as `$` (stock tickers) and `%` (percent changes).
* Replace numbers with a placeholder token (e.g., `NUM`) so that patterns like “NUM%” or “NUM million” are retained without increasing the vocabulary.
3. **Removing Stopwords:** Eliminate common words (e.g., "the", "a", "is") that typically do not affect sentiment.
4. **Lemmatization:** Reduce words to their base form (e.g., "running" → "run") to treat different inflections of the same word uniformly.
[cell 47 code]
import re, string
from nltk.corpus import stopwords
from nltk.stem import WordNetLemmatizer
lemmatizer = WordNetLemmatizer()
stop_words = set(stopwords.words("english"))
def preprocess_text(text):
# 1. Lowercase
text = text.lower()
# 2. Remove punctuation but keep $ and %
text = re.sub(r"[^\w\s$%]", "", text)
# 3. Replace numbers with placeholder NUM
text = re.sub(r"\d+(\.\d+)?", "NUM", text)
# 4. Tokenize
tokens = text.split()
# 5. Remove stopwords
tokens = [word for word in tokens if word not in stop_words]
# 6. Lemmatize
tokens = [lemmatizer.lemmatize(word) for word in tokens]
return " ".join(tokens)
[cell 48 code]
# Apply the preprocessing function
df['cleaned_sentence'] = df['Sentence'].apply(preprocess_text)
[cell 49 code]
# Display the original vs. cleaned sentences to verify
print("Original vs. Cleaned Sentences:\n")
for index, row in df.head(5).iterrows():
print(f"Original: {row['Sentence']}")
print(f"Cleaned: {row['cleaned_sentence']}, {row['Sentiment']}\n")
[cell 50 markdown]
# **Train-Test Split**
We first split our data into training (80%) and testing (20%) sets. We use `stratify=y` to ensure that the proportion of sentiment classes is the same in both sets, which is important for our imbalanced dataset.
[cell 51 code]
# Define features (X) and target (y)
X = df['cleaned_sentence']
y = df['Sentiment']
# Split the data
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
print("Training set shape:", X_train.shape)
print("Testing set shape:", X_test.shape)
[cell 52 markdown]
# **Feature Constructions; Vectorization**
Machine learning models cannot process raw text. We need to convert our cleaned sentences into numerical representations (vectors). We will implement and compare two standard methods: Bag-of-Words and TF-IDF.
[cell 53 markdown]
# **Method 1: Bag-of-Words (BoW)**
The Bag-of-Words model represents text by counting the occurrence of each word, disregarding grammar and word order. We use `CountVectorizer` for this.
[cell 54 code]
# Initialize and fit CountVectorizer
bow_vectorizer = CountVectorizer(max_features=3000, ngram_range=(1, 2))
# Transform the training and testing data
X_train_bow = bow_vectorizer.fit_transform(X_train)
X_test_bow = bow_vectorizer.transform(X_test)
print("Shape of BoW training matrix:", X_train_bow.shape)
print("Shape of BoW testing matrix:", X_test_bow.shape)
[cell 55 code]
print("First training sentence:\n", X_train.iloc[0])
# BoW features for this sentence
row0 = X_train_bow[0]
nz_idx = row0.nonzero()[1]
tokens = bow_vectorizer.get_feature_names_out()
print("\nNon-zero tokens for this sentence:")
for i, v in zip(nz_idx, row0.data):
print(f"Pos {i} -> {tokens[i]} : {v}")
[cell 56 markdown]
# **Hyperparameter search for the optimum feature size for BoW**
* The search space is [100, 500, 1000, 3000])
* We used Logistic Regression for sentiment classification
[cell 57 code]
import matplotlib.pyplot as plt
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, f1_score
# Different feature sizes to test
feature_sizes = [100, 500, 1000, 3000]
results = []
for size in feature_sizes:
# Vectorize with given max_features
bow_vectorizer = CountVectorizer(max_features=size, ngram_range=(1, 1))
X_train_bow = bow_vectorizer.fit_transform(X_train)
X_test_bow = bow_vectorizer.transform(X_test)
# Train Logistic Regression
lr_bow = LogisticRegression(max_iter=1000, class_weight='balanced', random_state=42)
lr_bow.fit(X_train_bow, y_train)
y_pred = lr_bow.predict(X_test_bow)
# Metrics
acc = accuracy_score(y_test, y_pred)
f1 = f1_score(y_test, y_pred, average='weighted')
results.append((size, acc, f1))
print(f"Feature size={size} | Accuracy={acc:.3f} | F1 (weighted)={f1:.3f}")
# Convert to DataFrame
import pandas as pd
df_res = pd.DataFrame(results, columns=["Features", "Accuracy", "F1"])
# Plotting
fig, axes = plt.subplots(2, 2, figsize=(12, 8))
axes = axes.ravel()
for idx, row in df_res.iterrows():
size = row["Features"]
acc = row["Accuracy"]
f1 = row["F1"]
ax = axes[idx]
ax.bar(["Accuracy", "F1"], [acc, f1], color=["skyblue", "orange"])
ax.set_ylim(0, 1)
ax.set_title(f"Features = {size}")
for i, v in enumerate([acc, f1]):
ax.text(i, v + 0.01, f"{v:.3f}", ha="center")
plt.suptitle("Logistic Regression with Different BoW Feature Sizes", fontsize=14)
plt.tight_layout(rect=[0, 0, 1, 0.96])
plt.show()
[cell 58 markdown]
**Based on the results, we selected max_features = 3000 for further experiments.**
[cell 59 markdown]
# **Hyperparameter search for the optimum n-gram size with BOW**
* The search space is [(1,1), (1,2), (1,3), (1,4)]
* We used Logistic Regression for sentiment classification
[cell 60 code]
import matplotlib.pyplot as plt
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, f1_score
import pandas as pd
# Different n-gram ranges to test
ngram_ranges = [(1,1), (1,2), (1,3), (1,4)]
results = []
for ngram in ngram_ranges:
# Vectorize with given n-gram range
bow_vectorizer = CountVectorizer(max_features=3000, ngram_range=ngram)
X_train_bow = bow_vectorizer.fit_transform(X_train)
X_test_bow = bow_vectorizer.transform(X_test)
# Train Logistic Regression
lr_bow = LogisticRegression(max_iter=1000, class_weight='balanced', random_state=42)
lr_bow.fit(X_train_bow, y_train)
y_pred = lr_bow.predict(X_test_bow)
# Metrics
acc = accuracy_score(y_test, y_pred)
f1 = f1_score(y_test, y_pred, average='weighted')
results.append((ngram, acc, f1))
print(f"ngram_range={ngram} | Accuracy={acc:.3f} | F1 (weighted)={f1:.3f}")
# Convert to DataFrame for easier handling
df_res = pd.DataFrame(results, columns=["N-gram Range", "Accuracy", "F1"])
# Plotting
fig, axes = plt.subplots(2, 2, figsize=(12, 8))
axes = axes.ravel()
for idx, row in df_res.iterrows():
ngram = row["N-gram Range"]
acc = row["Accuracy"]
f1 = row["F1"]
ax = axes[idx]
ax.bar(["Accuracy", "F1"], [acc, f1], color=["skyblue", "orange"])
ax.set_ylim(0, 1)
ax.set_title(f"N-gram {ngram}")
for i, v in enumerate([acc, f1]):
ax.text(i, v + 0.01, f"{v:.3f}", ha="center")
plt.suptitle("Logistic Regression with Different N-gram Ranges", fontsize=14)
plt.tight_layout(rect=[0, 0, 1, 0.96])
plt.show()
[cell 61 markdown]
**Based on the results, we selected N-gram (1,1) for further experiments.**
[cell 62 markdown]
# **Final Sentiment Classification Model for Logistic Regression based on Bow**
[cell 63 code]
# Initialize and fit CountVectorizer
bow_vectorizer = CountVectorizer(max_features=3000, ngram_range=(1, 1))
# Transform the training and testing data
X_train_bow = bow_vectorizer.fit_transform(X_train)
X_test_bow = bow_vectorizer.transform(X_test)
print("--- Logistic Regression with BoW Features ---")
lr_bow = LogisticRegression(max_iter=1000, class_weight='balanced', random_state=42)
lr_bow.fit(X_train_bow, y_train)
y_pred_lr_bow = lr_bow.predict(X_test_bow)
print("Classification Report:\n")
print(classification_report(y_test, y_pred_lr_bow, digits=3))
[cell 64 markdown]
# **Method 2: TF-IDF**
Term Frequency-Inverse Document Frequency (TF-IDF) improves upon BoW by weighting words based on their importance. It assigns a higher weight to words that are frequent in a document but rare across the entire corpus.
[cell 65 code]
# Initialize and fit TfidfVectorizer
tfidf_vectorizer = TfidfVectorizer(max_features=3000, ngram_range=(1, 2))
# Transform the training and testing data
X_train_tfidf = tfidf_vectorizer.fit_transform(X_train)
X_test_tfidf = tfidf_vectorizer.transform(X_test)
print("Shape of TF-IDF training matrix:", X_train_tfidf.shape)
print("Shape of TF-IDF testing matrix:", X_test_tfidf.shape)
[cell 66 code]
print("First training sentence:\n", X_train.iloc[0])
# Btfidf features for this sentence
row0 = X_train_tfidf[0]
nz_idx = row0.nonzero()[1]
tokens = tfidf_vectorizer.get_feature_names_out()
print("\nNon-zero tokens for this sentence:")
for i, v in zip(nz_idx, row0.data):
print(f"Pos {i} -> {tokens[i]} : {v}")
[cell 67 markdown]
# **Hyperparameter search for the optimum feature size for Tfidf**
* The search space is [100, 500, 1000, 3000])
* We used Logistic Regression for sentiment classification
[cell 68 code]
import matplotlib.pyplot as plt
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, f1_score
import pandas as pd
# Different feature sizes to test
feature_sizes = [100, 500, 1000, 3000]
results = []
for size in feature_sizes:
# Vectorize with TF-IDF for given max_features
tfidf_vectorizer = TfidfVectorizer(max_features=size, ngram_range=(1, 1))
X_train_tfidf = tfidf_vectorizer.fit_transform(X_train)
X_test_tfidf = tfidf_vectorizer.transform(X_test)
# Train Logistic Regression
lr_tfidf = LogisticRegression(max_iter=1000, class_weight='balanced', random_state=42)
lr_tfidf.fit(X_train_tfidf, y_train)
y_pred = lr_tfidf.predict(X_test_tfidf)
# Metrics
acc = accuracy_score(y_test, y_pred)
f1 = f1_score(y_test, y_pred, average='weighted')
results.append((size, acc, f1))
print(f"Feature size={size} | Accuracy={acc:.3f} | F1 (weighted)={f1:.3f}")
# Convert to DataFrame
df_res = pd.DataFrame(results, columns=["Features", "Accuracy", "F1"])
# Plotting
fig, axes = plt.subplots(2, 2, figsize=(12, 8))
axes = axes.ravel()
for idx, row in df_res.iterrows():
size = row["Features"]
acc = row["Accuracy"]
f1 = row["F1"]
ax = axes[idx]
ax.bar(["Accuracy", "F1"], [acc, f1], color=["skyblue", "orange"])
ax.set_ylim(0, 1)
ax.set_title(f"TF-IDF Features = {size}")
for i, v in enumerate([acc, f1]):
ax.text(i, v + 0.01, f"{v:.3f}", ha="center")
plt.suptitle("Logistic Regression with Different TF-IDF Feature Sizes", fontsize=14)
plt.tight_layout(rect=[0, 0, 1, 0.96])
plt.show()
[cell 69 markdown]
**Based on the results, we selected max_features = 3000 for further experiments.**
[cell 70 markdown]
# **Hyperparameter search for the optimum n-gram size with Tfidf**
* The search space is [(1,1), (1,2), (1,3), (1,4)]
* We used Logistic Regression for sentiment classification
[cell 71 code]
import matplotlib.pyplot as plt
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, f1_score
import pandas as pd
# Different n-gram ranges to test
ngram_ranges = [(1,1), (1,2), (1,3), (1,4)]
results = []
for ngram in ngram_ranges:
# Vectorize with TF-IDF for given n-gram range
tfidf_vectorizer = TfidfVectorizer(max_features=3000, ngram_range=ngram)
X_train_tfidf = tfidf_vectorizer.fit_transform(X_train)
X_test_tfidf = tfidf_vectorizer.transform(X_test)
# Train Logistic Regression
lr_tfidf = LogisticRegression(max_iter=1000, class_weight='balanced', random_state=42)
lr_tfidf.fit(X_train_tfidf, y_train)
y_pred = lr_tfidf.predict(X_test_tfidf)
# Metrics
acc = accuracy_score(y_test, y_pred)
f1 = f1_score(y_test, y_pred, average='weighted')
results.append((ngram, acc, f1))
print(f"ngram_range={ngram} | Accuracy={acc:.3f} | F1 (weighted)={f1:.3f}")
# Convert to DataFrame
df_res = pd.DataFrame(results, columns=["N-gram Range", "Accuracy", "F1"])
# Plotting
fig, axes = plt.subplots(2, 2, figsize=(12, 8))
axes = axes.ravel()
for idx, row in df_res.iterrows():
ngram = row["N-gram Range"]
acc = row["Accuracy"]
f1 = row["F1"]
ax = axes[idx]
ax.bar(["Accuracy", "F1"], [acc, f1], color=["skyblue", "orange"])
ax.set_ylim(0, 1)
ax.set_title(f"TF-IDF ngram={ngram}")
for i, v in enumerate([acc, f1]):
ax.text(i, v + 0.01, f"{v:.3f}", ha="center")
plt.suptitle("Logistic Regression with Different TF-IDF N-gram Ranges", fontsize=14)
plt.tight_layout(rect=[0, 0, 1, 0.96])
plt.show()
[cell 72 markdown]
**Based on the results, we selected N-gram (1,2) for further experiments.**
[cell 73 markdown]
# **Final Sentiment Classification Model for Logistic Regression based on Tfidf**
[cell 74 code]
tfidf_vectorizer = TfidfVectorizer(max_features=3000, ngram_range=(1,2))
X_train_tfidf = tfidf_vectorizer.fit_transform(X_train)
X_test_tfidf = tfidf_vectorizer.transform(X_test)
print("--- Logistic Regression with TF-IDF Features ---")
lr_tfidf = LogisticRegression(max_iter=1000, class_weight='balanced', random_state=42)
lr_tfidf.fit(X_train_tfidf, y_train)
y_pred_lr_tfidf = lr_tfidf.predict(X_test_tfidf)
print("Classification Report:\n")
print(classification_report(y_test, y_pred_lr_tfidf, digits=3))
[cell 75 markdown]
# **Comparing with Multinomial Naive Bayes and Random Forest**
[cell 76 code]
from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer
from sklearn.naive_bayes import MultinomialNB
from sklearn.ensemble import RandomForestClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report
# BoW
bow_vectorizer = CountVectorizer(max_features=3000, ngram_range=(1, 1))
X_train_bow = bow_vectorizer.fit_transform(X_train)
X_test_bow = bow_vectorizer.transform(X_test)
# TF-IDF
tfidf_vectorizer = TfidfVectorizer(max_features=3000, ngram_range=(1, 2))
X_train_tfidf = tfidf_vectorizer.fit_transform(X_train)
X_test_tfidf = tfidf_vectorizer.transform(X_test)
print("\n--- Naive Bayes with BoW Features ---")
nb_bow = MultinomialNB()
nb_bow.fit(X_train_bow, y_train)
y_pred_nb_bow = nb_bow.predict(X_test_bow)
print(classification_report(y_test, y_pred_nb_bow, digits=3))
print("\n--- Naive Bayes with TF-IDF Features ---")
nb_tfidf = MultinomialNB()
nb_tfidf.fit(X_train_tfidf, y_train)
y_pred_nb_tfidf = nb_tfidf.predict(X_test_tfidf)
print(classification_report(y_test, y_pred_nb_tfidf, digits=3))
print("\n--- Random Forest with BoW Features ---")
rf_bow = RandomForestClassifier(n_estimators=100, random_state=42, class_weight="balanced")
rf_bow.fit(X_train_bow, y_train)
y_pred_rf_bow = rf_bow.predict(X_test_bow)
print(classification_report(y_test, y_pred_rf_bow, digits=3))
print("\n--- Random Forest with TF-IDF Features ---")
rf_tfidf = RandomForestClassifier(n_estimators=100, random_state=42, class_weight="balanced")
rf_tfidf.fit(X_train_tfidf, y_train)
y_pred_rf_tfidf = rf_tfidf.predict(X_test_tfidf)
print(classification_report(y_test, y_pred_rf_tfidf, digits=3))
[cell 77 code]
results = {
"Model": [
"Logistic Regression (BoW)", "Logistic Regression (TF-IDF)",
"Naive Bayes (BoW)", "Naive Bayes (TF-IDF)",
"Random Forest (BoW)", "Random Forest (TF-IDF)"
],
"Accuracy": [
accuracy_score(y_test, y_pred_lr_bow), accuracy_score(y_test, y_pred_lr_tfidf),
accuracy_score(y_test, y_pred_nb_bow), accuracy_score(y_test, y_pred_nb_tfidf),
accuracy_score(y_test, y_pred_rf_bow), accuracy_score(y_test, y_pred_rf_tfidf)
],
"F1-Score": [
f1_score(y_test, y_pred_lr_bow, average='weighted'), f1_score(y_test, y_pred_lr_tfidf, average='weighted'),
f1_score(y_test, y_pred_nb_bow, average='weighted'), f1_score(y_test, y_pred_nb_tfidf, average='weighted'),
f1_score(y_test, y_pred_rf_bow, average='weighted'), f1_score(y_test, y_pred_rf_tfidf, average='weighted')
]
}
df_results = pd.DataFrame(results).sort_values(by="F1-Score", ascending=False).reset_index(drop=True)
df_results
[cell 78 markdown]
**Observations**
1. **Naive Bayes with TF-IDF** achieves the highest accuracy (0.683) but its weighted F1-score (0.647) is lower than Naive Bayes (BoW) and Logistic Regression (TF-IDF), showing that while it predicts the majority class well, its balance across all classes is weaker.
2. **Naive Bayes (BoW) and Logistic Regression (TF-IDF)** show more balanced performance, with F1-scores around, making them more reliable choices compared to Random Forest, which consistently underperforms in both Accuracy and F1.
[cell 79 markdown]
#**3. Clustering with K-Means**
[cell 80 markdown]
# **Finding the optimal number of clusters using tfidf vectorization**
[cell 81 code]
import pandas as pd
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.cluster import KMeans
import matplotlib.pyplot as plt
# TF-IDF vectorization
tfidf = TfidfVectorizer(max_features=5000, stop_words='english', ngram_range=(1,2))
X_tfidf = tfidf.fit_transform(df['cleaned_sentence'])
# Elbow method to find optimal K
inertia = []
K_range = range(2, 11)
for k in K_range:
km = KMeans(n_clusters=k, random_state=42, n_init=10)
km.fit(X_tfidf)
inertia.append(km.inertia_) # within-cluster sum of squares
# Step 3: Plot elbow curve
plt.figure(figsize=(8,6))
plt.plot(K_range, inertia, 'bo-')
plt.xlabel("Number of clusters (k)")
plt.ylabel("Inertia (WCSS)")
plt.title("Elbow Method for Optimal k")
plt.show()
[cell 82 code]
k_values = list(range(2, 11))
df_elbow = pd.DataFrame({
"k": k_values,
"Inertia": inertia
})
# Add drop magnitude (difference from previous inertia)
df_elbow["Drop"] = df_elbow["Inertia"].shift(1) - df_elbow["Inertia"]
print(df_elbow)
[cell 83 markdown]
**The largest drop in inertia occurs when moving from k=2 to k=3**
[cell 84 markdown]
# **Visualising clusters for k=[2,3,4,5]**
[cell 85 code]
import matplotlib.pyplot as plt
from sklearn.decomposition import PCA
from sklearn.cluster import KMeans
# Apply PCA on TF-IDF
X_pca = PCA(n_components=2, random_state=42).fit_transform(X_tfidf.toarray())
# Run KMeans for k=2,3,4,5
k_values = [2, 3, 4, 5]
cluster_results = {}
for k in k_values:
km = KMeans(n_clusters=k, random_state=42, n_init=10)
cluster_results[k] = km.fit_predict(X_tfidf)
# 2×3 subplot
fig, axes = plt.subplots(2, 3, figsize=(15, 10))
# Ground Truth plot with legend
color_map = {"positive": "red", "negative": "blue", "neutral": "green"}
for sentiment, color in color_map.items():
mask = df['Sentiment'] == sentiment
axes[0, 0].scatter(
X_pca[mask, 0],
X_pca[mask, 1],
c=color,
label=sentiment,
alpha=0.5
)
axes[0, 0].set_title("Ground Truth (Sentiment Labels)")
axes[0, 0].legend(title="Sentiment")
# KMeans plots
for i, k in enumerate(k_values):
ax = axes[(i+1)//3, (i+1)%3] # shift index by 1 to leave [0,0] for ground truth
ax.scatter(
X_pca[:, 0],
X_pca[:, 1],
c=cluster_results[k],
cmap="tab10",
alpha=0.7
)
ax.set_title(f"KMeans (k={k})")
# Leave last subplot blank
axes[1, 2].axis("off")
plt.tight_layout()
plt.show()
[cell 86 markdown]
**Although the elbow method suggests that k=3 provides the most reasonable partition, the clusters are not cleanly separated in the PCA visualizations. This indicates that sentiment-based separation in the Financial PhraseBank dataset is not strongly captured by unsupervised clustering on BoW/TF-IDF features, making the results less convincing for sentiment grouping.**
[cell 87 code]
from sklearn.metrics import (
silhouette_score, davies_bouldin_score, calinski_harabasz_score,
homogeneity_score, completeness_score, v_measure_score
)
# Final KMeans with k=3
k = 3
kmeans = KMeans(n_clusters=k, random_state=42, n_init=10)
y_pred_kmeans = kmeans.fit_predict(X_tfidf)
# Intrinsic metrics
sil_score = silhouette_score(X_tfidf, y_pred_kmeans)
db_score = davies_bouldin_score(X_tfidf.toarray(), y_pred_kmeans)
ch_score = calinski_harabasz_score(X_tfidf.toarray(), y_pred_kmeans)
# Extrinsic metrics (need ground truth labels)
h_score = homogeneity_score(df['Sentiment'], y_pred_kmeans)
c_score = completeness_score(df['Sentiment'], y_pred_kmeans)
v_score = v_measure_score(df['Sentiment'], y_pred_kmeans)
# Print
print("Intrinsic Metrics:")
print(f" Silhouette Score: {sil_score:.3f}")
print(f" Davies-Bouldin Index: {db_score:.3f}")
print(f" Calinski-Harabasz Score: {ch_score:.3f}\n")
print("Extrinsic Metrics :")
print(f" Homogeneity: {h_score:.3f}")
print(f" Completeness: {c_score:.3f}")
print(f" V-Measure: {v_score:.3f}")
[cell 88 markdown]
* **The clustering results are very weak. The intrinsic metrics show almost no meaningful separation (Silhouette ≈ 0, very high Davies–Bouldin score, and low Calinski–Harabasz value), indicating that the clusters are not compact or well-separated. The extrinsic metrics are also near zero (Homogeneity, Completeness, and V-Measure all < 0.01), which confirms that the clusters do not align with the actual sentiment labels. This suggests that unsupervised clustering on simple TF-IDF features is not effective for sentiment-based grouping in this dataset.**
* **Moreover, the PCA visualization of the ground truth labels itself shows heavy overlap between positive, negative, and neutral samples, which further explains why the clusters fail to appear well separated.**
[cell 89 markdown]
# **Finding the optimal number of clusters using bow vectorization**
[cell 90 code]
# Vectorize the full dataset using BoW
bow_vectorizer = CountVectorizer(max_features=5000, stop_words="english", ngram_range=(1,2))
X_bow_full = bow_vectorizer.fit_transform(df['cleaned_sentence'])
pca_bow = PCA(n_components=2, random_state=42)
X_pca_bow = pca_bow.fit_transform(X_bow_full.toarray())
# Run KMeans for k=2,3,4,5
k_values = [2, 3, 4, 5]
cluster_results = {}
for k in k_values:
km = KMeans(n_clusters=k, random_state=42, n_init=10)
cluster_results[k] = km.fit_predict(X_bow_full)
# 2×3 subplot
fig, axes = plt.subplots(2, 3, figsize=(15, 10))
# Ground Truth plot with legend
color_map = {"positive": "red", "negative": "blue", "neutral": "green"}
for sentiment, color in color_map.items():
mask = df['Sentiment'] == sentiment
axes[0, 0].scatter(
X_pca_bow[mask, 0],
X_pca_bow[mask, 1],
c=color,
label=sentiment,
alpha=0.5
)
axes[0, 0].set_title("Ground Truth (Sentiment Labels)")
axes[0, 0].legend(title="Sentiment")
# KMeans plots
for i, k in enumerate(k_values):
ax = axes[(i+1)//3, (i+1)%3]
ax.scatter(
X_pca_bow[:, 0],
X_pca_bow[:, 1],
c=cluster_results[k],
cmap="tab10",
alpha=0.7
)
ax.set_title(f"KMeans (k={k})")
# Leave last subplot blank
axes[1, 2].axis("off")
plt.tight_layout()
plt.show()
[cell 91 markdown]
**Cluster separation is poor: Just like TF-IDF, BoW also fails to produce clearly separated clusters in PCA space.**
[cell 92 markdown]
# **Better Cluster Separation with Topic Categories dataset (NewsGroup dataset)**:
[cell 93 code]
import pandas as pd
import re
from sklearn.datasets import fetch_20newsgroups
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.cluster import KMeans
from sklearn.decomposition import PCA
import matplotlib.pyplot as plt
# Load dataset (3 categories)
categories = ['rec.sport.baseball', 'sci.space', 'talk.politics.mideast']
newsgroups = fetch_20newsgroups(
subset='train',
categories=categories,
remove=('headers','footers','quotes')
)
df_alt = pd.DataFrame({
'sentence': newsgroups.data,
'Label': newsgroups.target
})
df_alt['Category'] = df_alt['Label'].map(dict(enumerate(newsgroups.target_names)))
print(newsgroups.target_names)
[cell 94 code]
df_alt[:10]
[cell 95 code]
# TF-IDF vectorization
tfidf = TfidfVectorizer(stop_words='english', max_features=3000,ngram_range=(1,2))
X_tfidf = tfidf.fit_transform(df_alt['sentence'])
# KMeans clustering for k=3 (since we know 3 topics)
k = 3
km = KMeans(n_clusters=k, random_state=42, n_init=10)
clusters = km.fit_predict(X_tfidf)
# PCA to 2D for visualization
X_pca = PCA(n_components=2, random_state=42).fit_transform(X_tfidf.toarray())
# Plot ground truth vs clustering side by side
fig, axes = plt.subplots(1, 2, figsize=(14, 6))
# Ground truth labels
axes[0].scatter(X_pca[:,0], X_pca[:,1], c=df_alt['Label'], cmap="tab10", alpha=0.7)
axes[0].set_title("Ground Truth (Topics)")
axes[0].set_xlabel("PC1")
axes[0].set_ylabel("PC2")
# KMeans clustering result
axes[1].scatter(X_pca[:,0], X_pca[:,1], c=clusters, cmap="tab10", alpha=0.7)
axes[1].set_title("KMeans Clustering (k=3)")
axes[1].set_xlabel("PC1")
axes[1].set_ylabel("PC2")
plt.tight_layout()
plt.show()
[cell 96 markdown]
* Clustering works well on this subset of the 20 Newsgroups dataset because the categories (sports, space, politics) have distinct vocabularies, so TF-IDF highlights words that strongly separate the topics, and PCA+KMeans reveals clear cluster boundaries.
* In contrast, the Financial PhraseBank dataset with sentiment labels does not show good clustering because positive, negative, and neutral texts often share much of the same financial terminology, making their word distributions highly overlapping and harder to separate in an unsupervised setting.