# texteda bow tfidf v1 (1)
course: Module 3 — Deep Learning & NLP
module: Module-3-Deep-Learning-NLP
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-3-Deep-Learning-NLP/General/Lab_Materials-05.04.2026/texteda_bow_tfidf_v1_(1).pdf
pages: 55
---
[page 1]
texteda_bow_tfidf_v1
April 4, 2026
Financial PhraseBank Dataset
The Financial PhraseBank datasetcontainsshortsentencesextractedfromfinancialnewsarticles
related to publicly traded companies. Each sentence is labeled with a sentiment that reflects how
the information would likely impact investor perceptions of the company’s stock.
In this study, we focus on preliminary text data analysis , which involves performing ex-
ploratory data analysis (EDA) , and applyingBag-of-W ords (BoW)and TF–IDF represen-
tations for bothsentiment classification and unsupervised clustering.
Source: Kaggle Dataset
• Kaggle: Financial Sentiment Analysis
Target Variable:
• sentiment (categorical): Indicates the polarity of the financial news sentence:
– positive
– negative
– neutral
Features:
Column
Name Description
Data
Type
sentenceA short text segment extracted from financial news. Each sentence is
self-contained and typically refers to a specific company, event, or market
movement.
Text
sentiment Label showing whether the news is expected to have apositive, negative, or
neutral impact on the company’s stock or market perception.
Categorical
Reference paper:Malo et al., Good Debt or Bad Debt, JASIST 2014
# Importing the necessary libraries for data manipulation, visualization, and text
processing.
[1]: # Core libraries
import pandas as pd
import numpy as np
1
[page 2]
import matplotlib.pyplot as plt
import seaborn as sns
import re
import string
# NLTK for text processing
import nltk
from nltk.corpus import stopwords
from nltk.stem import WordNetLemmatizer
from wordcloud import WordCloud
# Scikit-learn for feature extraction, modeling, and evaluation
from sklearn.model_selection import train_test_split
from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.naive_bayes import MultinomialNB
from sklearn.ensemble import RandomForestClassifier
from sklearn.cluster import KMeans
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE
from sklearn.metrics import (
classification_report, confusion_matrix, accuracy_score,f1_score,
silhouette_score, davies_bouldin_score, calinski_harabasz_score,
homogeneity_score, completeness_score, v_measure_score
)
import warnings
warnings.filterwarnings("ignore")
# Download NLTK resources
nltk.download('stopwords')
nltk.download('wordnet')
# Set plotting style
sns.set_style("whitegrid")
plt.rcParams['figure.figsize'] = (10, 6)
[nltk_data] Downloading package stopwords to /root/nltk_data…
[nltk_data] Unzipping corpora/stopwords.zip.
[nltk_data] Downloading package wordnet to /root/nltk_data…
2
[page 3]
1 Load the dataset
[2]: # Load the dataset
df = pd.read_csv('data.csv')
# Display the head
df.head(10)
[2]: Sentence Sentiment
0 The GeoSolutions technology will leverage Bene… positive
1 $ESI on lows, down $1.50 to $2.50 BK a real po… negative
2 For the last quarter of 2010 , Componenta 's n… positive
3 According to the Finnish-Russian Chamber of Co… neutral
4 The Swedish buyout firm has sold its remaining… neutral
5 $SPY wouldn't be surprised to see a green close positive
6 Shell's $70 Billion BG Deal Meets Shareholder … negative
7 SSH COMMUNICATIONS SECURITY CORP STOCK EXCHANG… negative
8 Kone 's net sales rose by some 14 % year-on-ye… positive
9 The Stockmann department store will have a tot… neutral
Basic information about the dataset
[3]: df.info()
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 5842 entries, 0 to 5841
Data columns (total 2 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 Sentence 5842 non-null object
1 Sentiment 5842 non-null object
dtypes: object(2)
memory usage: 91.4+ KB
Check for missing values
[4]: print("Missing values in each column:\n", df.isnull().sum())
Missing values in each column:
Sentence 0
Sentiment 0
dtype: int64
The dataset is loaded successfully with 5842 entries and two columns: ‘Sentence’ and ‘Sentiment’.
There are no missing values, so we can proceed with the analysis.
[29]: stop_words = set(ENGLISH_STOP_WORDS)
print(f"--- Sample Data by Class ---\n")
for sentiment in df['Sentiment'].unique():
sample = df[df['Sentiment'] == sentiment].head(5)
3
[page 4]
print(f"**Class: {sentiment.upper()}**\n")
for _, row in sample.iterrows():
print(f"Input: {row['Sentence']}")
print(f"Output: {row['Sentiment']}\n")
--- Sample Data by Class ---
**Class: POSITIVE**
Input: The GeoSolutions technology will leverage Benefon 's GPS solutions by
providing Location Based Search Technology , a Communities Platform , location
relevant multimedia content and a new and powerful commercial model .
Output: positive
Input: For the last quarter of 2010 , Componenta 's net sales doubled to EUR131m
from EUR76m for the same period a year earlier , while it moved to a zero pre-
tax profit from a pre-tax loss of EUR7m .
Output: positive
Input: $SPY wouldn't be surprised to see a green close
Output: positive
Input: Kone 's net sales rose by some 14 % year-on-year in the first nine months
of 2008 .
Output: positive
Input: Circulation revenue has increased by 5 % in Finland and 4 % in Sweden in
2008 .
Output: positive
**Class: NEGATIVE**
Input: $ESI on lows, down $1.50 to $2.50 BK a real possibility
Output: negative
Input: Shell's $70 Billion BG Deal Meets Shareholder Skepticism
Output: negative
Input: SSH COMMUNICATIONS SECURITY CORP STOCK EXCHANGE RELEASE OCTOBER 14 , 2008
AT 2:45 PM The Company updates its full year outlook and estimates its results
to remain at loss for the full year .
Output: negative
Input: $SAP Q1 disappoints as #software licenses down. Real problem? #Cloud
growth trails $MSFT $ORCL $GOOG $CRM $ADBE https://t.co/jNDphllzq5
Output: negative
Input: $AAPL afternoon selloff as usual will be brutal. get ready to lose a ton
4
[page 5]
of money.
Output: negative
**Class: NEUTRAL**
Input: According to the Finnish-Russian Chamber of Commerce , all the major
construction companies of Finland are operating in Russia .
Output: neutral
Input: The Swedish buyout firm has sold its remaining 22.4 percent stake ,
almost eighteen months after taking the company public in Finland .
Output: neutral
Input: The Stockmann department store will have a total floor space of over
8,000 square metres and Stockmann 's investment in the project will have a price
tag of about EUR 12 million .
Output: neutral
Input: Viking Line has canceled some services .
Output: neutral
Input: Ahlstrom Corporation STOCK EXCHANGE ANNOUNCEMENT 7.2.2007 at 10.30 A
total of 56,955 new shares of Ahlstrom Corporation have been subscribed with
option rights under the company 's stock option programs I 2001 and II 2001 .
Output: neutral
1.1 1. Exploratory Data Analysis (EDA)
2 Sentiment Distribution
First, let’s examine the distribution of the three sentiment labels: neutral, positive, and
negative. This is crucial to check for class imbalance, which can affect model training.
[5]: # Calculate the value counts
sentiment_counts = df['Sentiment'].value_counts(normalize=True)
print(sentiment_counts*100)
# Plot the sentiment distribution
plt.figure(figsize=(8, 5))
sns.countplot(x='Sentiment', data=df, order=sentiment_counts.index,␣
↪palette='viridis')
plt.title('Distribution of Sentiments', fontsize=16)
plt.xlabel('Sentiment', fontsize=12)
plt.ylabel('Number of Headlines', fontsize=12)
plt.xticks(fontsize=11)
plt.show()
5
[page 6]
Sentiment
neutral 53.577542
positive 31.701472
negative 14.720986
Name: proportion, dtype: float64
Observation:
The dataset is clearly imbalanced. Theneutral class is the dominant class, with significantly more
samples than thepositive and negative classes combined.
3 Text Characteristics: Sentence Length
Analyzing the length of the headlines can provide insights into their structure. We’ll look at the
distribution of word count and character count.
[6]: df['word_count'] = df['Sentence'].apply(lambda x: len(str(x).split())) # split ␣
↪on spaces and count the words
df['char_count'] = df['Sentence'].apply(len)
df[['Sentence', 'word_count', 'char_count']].head()
[6]: Sentence word_count char_count
0 The GeoSolutions technology will leverage Bene… 32 218
6
[page 7]
1 $ESI on lows, down $1.50 to $2.50 BK a real po… 11 55
2 For the last quarter of 2010 , Componenta 's n… 39 193
3 According to the Finnish-Russian Chamber of Co… 20 128
4 The Swedish buyout firm has sold its remaining… 23 135
[7]: # Plot sentence length distribution (in words)
plt.figure(figsize=(12, 5))
plt.subplot(1, 2, 1)
sns.histplot(df['word_count'], bins=30, kde=True, color='skyblue')
plt.title('Sentence Length Distribution (Word Count)')
plt.xlabel('Word Count')
plt.ylabel('Frequency')
plt.subplot(1, 2, 2)
sns.histplot(df['char_count'], bins=30, kde=True, color='salmon')
plt.title('Sentence Length Distribution (Character Count)')
plt.xlabel('Character Count')
plt.ylabel('Frequency')
plt.tight_layout()
plt.show()
Observations from Sentence Length Distribution:
The histograms provide a clear picture of the typical length of the financial headlines in this dataset:
• W ord Count (Left Plot): The distribution is right-skewed, meaning that most headlines
are short, with a long tail of fewer, longer headlines.
– The vast majority of sentences contain between10 and 30 words .
– The peak of the distribution (the mode) is around15 words per headline .
– It’s rare to see headlines longer than 40-50 words, indicating a concise writing style.
7
[page 8]
• Character Count (Right Plot): This plot mirrors the word count distribution, also show-
ing a right skew.
– Most headlines fall within the50 to 150 character range.
– The peak frequency is for headlines around75-100 characters in length.
4 word count per sentiment category
[8]: # Analyze sentence length per sentiment
plt.figure(figsize=(10, 6))
sns.boxplot(x='Sentiment', y='word_count', data=df, palette='Set3')
plt.title('Word Count Distribution per Sentiment')
plt.xlabel('Sentiment')
plt.ylabel('Word Count')
plt.show()
Observations:
• Most headlines are between 10 and 30 words long.
• The distribution of word count is fairly similar across all three sentiments. There isn’t a
strong indication that sentence length alone is a powerful predictor of sentiment.
8
[page 9]
5 Word Frequency Analysis
[36]: # Calculate sentence length (Character Count & Word Count)
df['char_count'] = df['Sentence'].apply(len)
df['word_count'] = df['Sentence'].apply(lambda x: len(str(x).split()))
# Function to clean text and count vocabulary (remove Stop words)
def get_top_words(text_series, top_n=10):
all_text = ' '.join(text_series).lower()
# Remove special characters, keep only a-z and $
all_text = re.sub(r'[^a-z\$]', ' ', all_text)
words = [w for w in all_text.split() if w not in ENGLISH_STOP_WORDS and␣
↪len(w) > 1]
return Counter(words).most_common(top_n)
# ==========================================
# Display numerical statistics (Terminal Output)
# ==========================================
print("--- Basic Information ---")
print(f"Total number of records: {len(df)} sentences")
print(f"Number of missing values:\n{df.isnull().sum()}\n")
print("--- Class Distribution ---")
class_counts = df['Sentiment'].value_counts()
class_pct = df['Sentiment'].value_counts(normalize=True) * 100
for sentiment, count in class_counts.items():
print(f"Class {sentiment.capitalize()}: {count} sentences␣
↪({class_pct[sentiment]:.1f}%)")
print("\n--- Text Length ---")
print(f"Average length (characters): {df['char_count'].mean():.1f} characters")
print(f"Average length (words): {df['word_count'].mean():.1f} words")
print(f"Longest sentence: {df['word_count'].max()} words | Shortest:␣
↪{df['word_count'].min()} words")
print("Average word length per class:")
print(df.groupby('Sentiment')['word_count'].mean().round(1))
print("\n--- Most Frequent Words (Top Words) ---")
print(f"Overall: {get_top_words(df['Sentence'], 8)}")
for sentiment in df['Sentiment'].unique():
print(f"Class {sentiment.capitalize()}: {get_top_words(df[df['Sentiment']␣
↪== sentiment]['Sentence'], 8)}")
--- Basic Information ---
Total number of records: 5842 sentences
Number of missing values:
Sentence 0
9
[page 10]
Sentiment 0
word_count 0
char_count 0
uppercase_word_count 0
cleaned_sentence 0
dtype: int64
--- Class Distribution ---
Class Neutral: 3130 sentences (53.6%)
Class Positive: 1852 sentences (31.7%)
Class Negative: 860 sentences (14.7%)
--- Text Length ---
Average length (characters): 117.0 characters
Average length (words): 21.0 words
Longest sentence: 81 words | Shortest: 2 words
Average word length per class:
Sentiment
negative 19.3
neutral 22.1
positive 19.9
Name: word_count, dtype: float64
--- Most Frequent Words (Top Words) ---
Overall: [('eur', 1734), ('mn', 821), ('company', 810), ('profit', 569),
('sales', 562), ('finnish', 539), ('said', 516), ('net', 500)]
Class Positive: [('eur', 622), ('mn', 259), ('year', 218), ('sales', 216),
('profit', 203), ('company', 203), ('net', 197), ('said', 194)]
Class Negative: [('eur', 390), ('mn', 233), ('profit', 154), ('sales', 105),
('year', 100), ('net', 100), ('finnish', 91), ('operating', 89)]
Class Neutral: [('eur', 722), ('company', 527), ('mn', 329), ('finnish', 267),
('said', 258), ('million', 251), ('sales', 241), ('finland', 236)]
[39]: import matplotlib.pyplot as plt
def plot_top_words_horizontal(text_series, top_n=10, title="Top Words"):
top_words = get_top_words(text_series, top_n)
words = [w[0] for w in top_words]
counts = [w[1] for w in top_words]
plt.figure(figsize=(8, 6))
plt.barh(words, counts)
plt.title(title)
plt.xlabel("Frequency")
plt.ylabel("Words")
10
[page 11]
# Highest frequency on top (better visualization)
plt.gca().invert_yaxis()
# Add value labels
for i, v in enumerate(counts):
plt.text(v + 0.5, i, str(v), va='center')
plt.tight_layout()
plt.show()
# For Overall Top Words
plot_top_words_horizontal(
df['Sentence'],
top_n=10,
title="Top 10 Most Frequent Words (Overall)"
)
# For Top Words Per Sentiment Class
for sentiment in df['Sentiment'].unique():
plot_top_words_horizontal(
df[df['Sentiment'] == sentiment]['Sentence'],
top_n=10,
title=f"Top Words - {sentiment.capitalize()}"
)
11
[page 12]
12
[page 13]
13
[page 14]
14
[page 15]
N-gram Analysis
N-grams are contiguous sequences ofn items from a given sample of text. By analyzing the most
common unigrams (1-grams), bigrams (2-grams), and trigrams (3-grams), we can identify key
phrases associated with each sentiment.
[9]: def get_top_ngrams(corpus, n, top_k=20):
"""
generates and returns the top k n-grams from a given corpus.
"""
# Use CountVectorizer to get n-gram counts
vec = CountVectorizer(ngram_range=(n, n), stop_words='english').fit(corpus)
bag_of_words = vec.transform(corpus)
sum_words = bag_of_words.sum(axis=0)
words_freq = [(word, sum_words[0, idx]) for word, idx in vec.vocabulary_.
↪items()]
words_freq = sorted(words_freq, key=lambda x: x[1], reverse=True)
return words_freq[:top_k]
[10]: # Get top unigrams for each sentiment
unigrams_pos = get_top_ngrams(df[df['Sentiment']=='positive']['Sentence'], 1)
15
[page 16]
unigrams_neg = get_top_ngrams(df[df['Sentiment']=='negative']['Sentence'], 1)
unigrams_neu = get_top_ngrams(df[df['Sentiment']=='neutral']['Sentence'], 1)
# Create dataframes for plotting
df_unigrams_pos = pd.DataFrame(unigrams_pos, columns=['Unigram', 'Frequency'])
df_unigrams_neg = pd.DataFrame(unigrams_neg, columns=['Unigram', 'Frequency'])
df_unigrams_neu = pd.DataFrame(unigrams_neu, columns=['Unigram', 'Frequency'])
# Plotting
fig, axes = plt.subplots(3, 1, figsize=(8, 14))
sns.barplot(x='Frequency', y='Unigram', data=df_unigrams_pos, ax=axes[0],␣
↪palette='Greens_d')
axes[0].set_title('Top 20 Unigrams for Positive Sentiment')
sns.barplot(x='Frequency', y='Unigram', data=df_unigrams_neg, ax=axes[1],␣
↪palette='Reds_d')
axes[1].set_title('Top 20 Unigrams for Negative Sentiment')
sns.barplot(x='Frequency', y='Unigram', data=df_unigrams_neu, ax=axes[2],␣
↪palette='Blues_d')
axes[2].set_title('Top 20 Unigrams for Neutral Sentiment')
plt.tight_layout()
plt.show()
16
[page 17]
17
[page 18]
Unigram Analysis Observations:
1. Significant Overlap of Core V ocabulary:
• A key finding is the substantial overlap of common financial terms across all three sen-
timents. Words like eur, mn (million), sales, profit, company, year, and net are
prominent in positive, negative, and neutral headlines.
• Thisindicatesthatthesecorenounsarecontext-dependent; theirsentimentisdetermined
by the surrounding words (verbs, adjectives). For example, profit can appear in a
positive context (“profit rose”) or a negative one (“profit fell”).
2. Distinct Sentiment Indicators:
• Positive Sentiment: While sharing the core vocabulary, the positive list contains clear
indicators of growth and success. The wordrose is a powerful and unambiguous positive
signal. The presence ofoperating often precedes positive metrics like profit or income.
• Negative Sentiment: Similarly, the negative list contains strong, distinct indicators of
poor performance. The wordsloss and decreased are clear signals of negative financial
news.
• Neutral Sentiment: The neutral list is characterized by descriptive and factual lan-
guage. The high frequency ofsaid suggests a focus on reporting statements or quotes.
Words likecompany, group, andbusiness are more prominent here, indicating a focus
on factual announcements rather than performance evaluation.
6 Bi-gram Anlaysis
[11]: # Get top bigrams for each sentiment
bigrams_pos = get_top_ngrams(df[df['Sentiment']=='positive']['Sentence'], 2)
bigrams_neg = get_top_ngrams(df[df['Sentiment']=='negative']['Sentence'], 2)
bigrams_neu = get_top_ngrams(df[df['Sentiment']=='neutral']['Sentence'], 2)
# Create dataframes for plotting
df_bigrams_pos = pd.DataFrame(bigrams_pos, columns=['Bigram', 'Frequency'])
df_bigrams_neg = pd.DataFrame(bigrams_neg, columns=['Bigram', 'Frequency'])
df_bigrams_neu = pd.DataFrame(bigrams_neu, columns=['Bigram', 'Frequency'])
# Plotting
fig, axes = plt.subplots(3, 1, figsize=(8, 14))
sns.barplot(x='Frequency', y='Bigram', data=df_bigrams_pos, ax=axes[0],␣
↪palette='Greens_d')
axes[0].set_title('Top 20 Bigrams for Positive Sentiment')
sns.barplot(x='Frequency', y='Bigram', data=df_bigrams_neg, ax=axes[1],␣
↪palette='Reds_d')
axes[1].set_title('Top 20 Bigrams for Negative Sentiment')
18
[page 19]
sns.barplot(x='Frequency', y='Bigram', data=df_bigrams_neu, ax=axes[2],␣
↪palette='Blues_d')
axes[2].set_title('Top 20 Bigrams for Neutral Sentiment')
plt.tight_layout()
plt.show()
19
[page 20]
20
[page 21]
Bigram Analysis Observations:
1. Core Financial Phrases (The Common Ground)
Across all three sentiments, we see a strong presence of standard financial reporting phrases. Bi-
grams like: *net sales * operating profit * eur mn (and mn eur) *corresponding period
These phrases represent the commonsubjects of financial news. They are inherently neutral and
their sentiment is determined entirely by the context in which they appear. Their high frequency
explains the initial similarity of the plots.
2. The Emergence of Sentiment-Carrying Phrases (The Key Differentiators)
The crucial differences lie in the phrases that describe anaction or state .
• F or Positive Sentiment: The model can learn from unambiguous phrases that signal
growth. We see the clear emergence of:
– profit rose
– rose eur These bigrams combine a subject (profit, eur) with a positive action (rose),
providing a powerful signal.
• F or Negative Sentiment: The negative chart contains the direct counterparts to the pos-
itive signals. Key phrases include:
– operating loss (the direct opposite ofoperating profit)
– decreased eur These combinations of a subject and a negative action are strong indi-
cators of negative sentiment.
• F or Neutral Sentiment: The neutral bigrams confirm their factual, non-evaluative nature.
They are dominated by:
– Reporting Phrases: company said, said today
– Proper Nouns & Entities: omx helsinki, alma media, stock exchange
– Structural Phrases: board directors, share capital These phrases focus on re-
porting events and facts rather than evaluating performance.
[ ]:
7 Tri-gram Anlaysis
[12]: # Get top trigrams for each sentiment
trigrams_pos = get_top_ngrams(df[df['Sentiment']=='positive']['Sentence'], 3)
trigrams_neg = get_top_ngrams(df[df['Sentiment']=='negative']['Sentence'], 3)
trigrams_neu = get_top_ngrams(df[df['Sentiment']=='neutral']['Sentence'], 3)
# Create dataframes for plotting
df_trigrams_pos = pd.DataFrame(trigrams_pos, columns=['Trigram', 'Frequency'])
df_trigrams_neg = pd.DataFrame(trigrams_neg, columns=['Trigram', 'Frequency'])
df_trigrams_neu = pd.DataFrame(trigrams_neu, columns=['Trigram', 'Frequency'])
# Plotting
21
[page 22]
fig, axes = plt.subplots(3, 1, figsize=(8, 14))
sns.barplot(x='Frequency', y='Trigram', data=df_trigrams_pos, ax=axes[0],␣
↪palette='Greens_d')
axes[0].set_title('Top 20 Trigrams for Positive Sentiment')
sns.barplot(x='Frequency', y='Trigram', data=df_trigrams_neg, ax=axes[1],␣
↪palette='Reds_d')
axes[1].set_title('Top 20 Trigrams for Negative Sentiment')
sns.barplot(x='Frequency', y='Trigram', data=df_trigrams_neu, ax=axes[2],␣
↪palette='Blues_d')
axes[2].set_title('Top 20 Trigrams for Neutral Sentiment')
plt.tight_layout()
plt.show()
22
[page 23]
23
[page 24]
8 Analysis of Uppercase Words
The use of uppercase words can sometimes indicate urgency or strong sentiment.
[13]: # Function to count uppercase words
def count_uppercase_words(text):
return len(re.findall(r'\b[A-Z]{2,}\b', text))
# Apply the function to create a new column
df['uppercase_word_count'] = df['Sentence'].apply(count_uppercase_words)
# Plot the distribution of uppercase words per sentiment
plt.figure(figsize=(10, 6))
sns.barplot(x='Sentiment', y='uppercase_word_count', data=df, estimator=np.
↪mean, ci=None)
plt.title('Average Uppercase Words per Sentence by Sentiment')
plt.xlabel('Sentiment')
plt.show()
Observations
1. Negative sentences contain the highest average number of uppercase words per
24
[page 25]
sentence, slightly more than positive ones. This suggests that negative financial news often
emphasizes entities (e.g., company tickers, acronyms) more strongly.
2. Neutral sentences have the fewest uppercase words on average , indicating that they
are less entity-driven and use fewer tickers/acronyms compared to sentiment-laden sentences.
[14]: # Display sentences with a high count of uppercase words
print("Examples of headlines with high uppercase word count:")
df.sort_values(by='uppercase_word_count', ascending=False).head(10)
Examples of headlines with high uppercase word count:
[14]: Sentence Sentiment word_count \
4764 $VOLC THIS STOCK HAS A VERY VERY VERY VERY VER… positive 19
4921 YIT CORPORATION SEPT. 24 , 2007 at 13:30 CORPO… neutral 47
5173 STOCK EXCHANGE ANNOUNCEMENT 20 July 2006 1 ( 1… neutral 34
5456 KAUKO-TELKO LTD PRESS RELEASE 19.06.2007 AT 14… positive 45
7 SSH COMMUNICATIONS SECURITY CORP STOCK EXCHANG… negative 34
65 Most bullish stocks on Twitter during this dip… positive 20
3617 Breaking 52 week highs timing looks great now … positive 19
4839 SSH COMMUNICATIONS SECURITY CORP STOCK EXCHANG… neutral 34
4811 End Of Day Scan: Stochastic Overbought $JDST $… negative 17
5474 $FB SHIFTING ON THE 15 MINUTE HERE - SUSPECT T… positive 12
char_count uppercase_word_count
4764 108 19
4921 248 16
5173 198 13
5456 256 11
7 190 10
65 132 10
3617 126 10
4839 190 10
4811 120 10
5474 58 10
[15]: import re
from collections import Counter
import matplotlib.pyplot as plt
import seaborn as sns
# Function to extract all uppercase words (length >= 2)
def extract_uppercase_words(text):
return re.findall(r'\b[A-Z]{2,}\b', text)
# Apply to entire corpus
all_uppercase_words = df['Sentence'].apply(extract_uppercase_words).sum()
# Count frequencies
25
[page 26]
word_freq = Counter(all_uppercase_words).most_common(20)
# Convert to DataFrame for plotting
upper_df = pd.DataFrame(word_freq, columns=['Word', 'Frequency'])
#print("Top 20 Uppercase Tokens (Tickers/Acronyms):")
#print(upper_df)
# Plot
plt.figure(figsize=(10,6))
sns.barplot(x='Frequency', y='Word', data=upper_df, palette="viridis")
plt.title("Top 20 Uppercase Words in Financial PhraseBank", fontsize=16)
plt.xlabel("Frequency")
plt.ylabel("Token")
plt.show()
Observations
1. EUR dominates overwhelmingly as the most frequent uppercase token, appearing far
more often than any other ticker or acronym, which highlights the dataset’s strong focus on
Euro-denominated financial news.
2. Other frequent uppercase tokens includecompany tickers (AAPL, TSLA, FB, UPM,
SPY)and financial acronyms (USD, EPS, CEO, FTSE) , showing that uppercase words
are largely entity-driven and represent key financial instruments or organizations.
26
[page 27]
9 Word Cloud Visualization
[16]: import matplotlib.pyplot as plt
from sklearn.feature_extraction.text import ENGLISH_STOP_WORDS
from wordcloud import WordCloud
# Use sklearn's stopwords for consistency
custom_stopwords = set(ENGLISH_STOP_WORDS)
# Function to generate wordcloud for a given sentiment
def generate_wordcloud(data, sentiment, ax):
text = " ".join(data[data['Sentiment'] == sentiment]['Sentence']).lower()
wordcloud = WordCloud(width=800, height=400,
background_color='white',
stopwords=custom_stopwords,
repeat=False ,
colormap='viridis').generate(text)
ax.imshow(wordcloud, interpolation='bilinear')
ax.set_title(f"{sentiment.capitalize()} Sentiment", fontsize=14)
ax.axis("off")
# Create 1x3 subplots
fig, axes = plt.subplots(3, 1, figsize=(8,17))
# Generate wordclouds for each sentiment
generate_wordcloud(df, 'positive', axes[0])
generate_wordcloud(df, 'negative', axes[1])
generate_wordcloud(df, 'neutral', axes[2])
plt.tight_layout()
plt.show()
27
[page 28]
28
[page 29]
10 Observations
Common financial terms dominate across all sentiments : words like eur, mn, company,
finnish, year, sales appear prominently in positive, negative, and neutral texts, reflecting the
dataset’s financial focus.
#2. Data preparation and vectorization and Sentiment Classification
Before feeding the text data into our models, we need to clean it. This process involves several
steps to standardize the text. Our preprocessing pipeline will include:
1. Lowercasing: Convert all text to lowercase so that words likeProfit and profit are treated
the same.
2. Handling Punctuation and Numbers:
• Remove punctuation marks that do not contribute to meaning.
• Preserve finance-specific symbols such as$ (stock tickers) and% (percent changes).
• Replace numbers with a placeholder token (e.g.,NUM) so that patterns like “NUM%” or
“NUM million” are retained without increasing the vocabulary.
3. Removing Stopwords: Eliminate common words (e.g., “the”, “a”, “is”) that typically do
not affect sentiment.
4. Lemmatization: Reduce words to their base form (e.g., “running” → “run”) to treat differ-
ent inflections of the same word uniformly.
[17]: import re, string
from nltk.corpus import stopwords
from nltk.stem import WordNetLemmatizer
lemmatizer = WordNetLemmatizer()
stop_words = set(stopwords.words("english"))
def preprocess_text(text):
# 1. Lowercase
text = text.lower()
# 2. Remove punctuation but keep $ and %
text = re.sub(r"[^\w\s$%]", "", text)
# 3. Replace numbers with placeholder NUM
text = re.sub(r"\d+(\.\d+)?", "NUM", text)
# 4. Tokenize
tokens = text.split()
# 5. Remove stopwords
tokens = [word for word in tokens if word not in stop_words]
# 6. Lemmatize
tokens = [lemmatizer.lemmatize(word) for word in tokens]
return " ".join(tokens)
29
[page 30]
[18]: # Apply the preprocessing function
df['cleaned_sentence'] = df['Sentence'].apply(preprocess_text)
[19]: # Display the original vs. cleaned sentences to verify
print("Original vs. Cleaned Sentences:\n")
for index, row in df.head(5).iterrows():
print(f"Original: {row['Sentence']}")
print(f"Cleaned: {row['cleaned_sentence']}, {row['Sentiment']}\n")
Original vs. Cleaned Sentences:
Original: The GeoSolutions technology will leverage Benefon 's GPS solutions by
providing Location Based Search Technology , a Communities Platform , location
relevant multimedia content and a new and powerful commercial model .
Cleaned: geosolutions technology leverage benefon gps solution providing
location based search technology community platform location relevant multimedia
content new powerful commercial model, positive
Original: $ESI on lows, down $1.50 to $2.50 BK a real possibility
Cleaned: $esi low $NUM $NUM bk real possibility, negative
Original: For the last quarter of 2010 , Componenta 's net sales doubled to
EUR131m from EUR76m for the same period a year earlier , while it moved to a
zero pre-tax profit from a pre-tax loss of EUR7m .
Cleaned: last quarter NUM componenta net sale doubled eurNUMm eurNUMm period
year earlier moved zero pretax profit pretax loss eurNUMm, positive
Original: According to the Finnish-Russian Chamber of Commerce , all the major
construction companies of Finland are operating in Russia .
Cleaned: according finnishrussian chamber commerce major construction company
finland operating russia, neutral
Original: The Swedish buyout firm has sold its remaining 22.4 percent stake ,
almost eighteen months after taking the company public in Finland .
Cleaned: swedish buyout firm sold remaining NUM percent stake almost eighteen
month taking company public finland, neutral
11 Train-Test Split
Wefirstsplitourdataintotraining(80%)andtesting(20%)sets. Weuse stratify=ytoensurethat
the proportion of sentiment classes is the same in both sets, which is important for our imbalanced
dataset.
[20]: # Define features (X) and target (y)
X = df['cleaned_sentence']
y = df['Sentiment']
30
[page 31]
# Split the data
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
print("Training set shape:", X_train.shape)
print("Testing set shape:", X_test.shape)
Training set shape: (4673,)
Testing set shape: (1169,)
12 Feature Constructions; Vectorization
Machine learning models cannot process raw text. We need to convert our cleaned sentences into
numerical representations (vectors). We will implement and compare two standard methods: Bag-
of-Words and TF-IDF.
13 Method 1: Bag-of-Words (BoW)
The Bag-of-Words model represents text by counting the occurrence of each word, disregarding
grammar and word order. We useCountVectorizer for this.
[ ]: # Initialize and fit CountVectorizer
bow_vectorizer = CountVectorizer(max_features=3000, ngram_range=(1, 2))
# Transform the training and testing data
X_train_bow = bow_vectorizer.fit_transform(X_train)
X_test_bow = bow_vectorizer.transform(X_test)
print("Shape of BoW training matrix:", X_train_bow.shape)
print("Shape of BoW testing matrix:", X_test_bow.shape)
Shape of BoW training matrix: (4673, 3000)
Shape of BoW testing matrix: (1169, 3000)
[ ]: print("First training sentence:\n", X_train.iloc[0])
# BoW features for this sentence
row0 = X_train_bow[0]
nz_idx = row0.nonzero()[1]
tokens = bow_vectorizer.get_feature_names_out()
print("\nNon-zero tokens for this sentence:")
for i, v in zip(nz_idx, row0.data):
print(f"Pos {i} -> {tokens[i]} : {v}")
31
[page 32]
First training sentence:
national conciliator juhani salonius met party wednesday said far apart view
propose mediation
Non-zero tokens for this sentence:
Pos 1616 -> national : 1
Pos 1253 -> juhani : 1
Pos 1511 -> met : 1
Pos 1941 -> party : 1
Pos 2937 -> wednesday : 1
Pos 2328 -> said : 1
Pos 820 -> far : 1
Pos 2908 -> view : 1
Pos 2132 -> propose : 1
14 Hyperparameter search for the optimum feature size for BoW
• The search space is [100, 500, 1000, 3000])
• We used Logistic Regression for sentiment classification
[ ]: import matplotlib.pyplot as plt
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, f1_score
# Different feature sizes to test
feature_sizes = [100, 500, 1000, 3000]
results = []
for size in feature_sizes:
# Vectorize with given max_features
bow_vectorizer = CountVectorizer(max_features=size, ngram_range=(1, 1))
X_train_bow = bow_vectorizer.fit_transform(X_train)
X_test_bow = bow_vectorizer.transform(X_test)
# Train Logistic Regression
lr_bow = LogisticRegression(max_iter=1000, class_weight='balanced',␣
↪random_state=42)
lr_bow.fit(X_train_bow, y_train)
y_pred = lr_bow.predict(X_test_bow)
# Metrics
acc = accuracy_score(y_test, y_pred)
f1 = f1_score(y_test, y_pred, average='weighted')
results.append((size, acc, f1))
32
[page 33]
print(f"Feature size={size} | Accuracy={acc:.3f} | F1 (weighted)={f1:.3f}")
# Convert to DataFrame
import pandas as pd
df_res = pd.DataFrame(results, columns=["Features", "Accuracy", "F1"])
# Plotting
fig, axes = plt.subplots(2, 2, figsize=(12, 8))
axes = axes.ravel()
for idx, row in df_res.iterrows():
size = row["Features"]
acc = row["Accuracy"]
f1 = row["F1"]
ax = axes[idx]
ax.bar(["Accuracy", "F1"], [acc, f1], color=["skyblue", "orange"])
ax.set_ylim(0, 1)
ax.set_title(f"Features = {size}")
for i, v in enumerate([acc, f1]):
ax.text(i, v + 0.01, f"{v:.3f}", ha="center")
plt.suptitle("Logistic Regression with Different BoW Feature Sizes",␣
↪fontsize=14)
plt.tight_layout(rect=[0, 0, 1, 0.96])
plt.show()
Feature size=100 | Accuracy=0.480 | F1 (weighted)=0.505
Feature size=500 | Accuracy=0.594 | F1 (weighted)=0.611
Feature size=1000 | Accuracy=0.606 | F1 (weighted)=0.621
Feature size=3000 | Accuracy=0.641 | F1 (weighted)=0.651
33
[page 34]
Based on the results, we selected max_features = 3000 for further experiments.
15 Hyperparameter search for the optimum n-gram size with
BOW
• The search space is [(1,1), (1,2), (1,3), (1,4)]
• We used Logistic Regression for sentiment classification
[ ]: import matplotlib.pyplot as plt
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, f1_score
import pandas as pd
# Different n-gram ranges to test
ngram_ranges = [(1,1), (1,2), (1,3), (1,4)]
results = []
for ngram in ngram_ranges:
# Vectorize with given n-gram range
bow_vectorizer = CountVectorizer(max_features=3000, ngram_range=ngram)
X_train_bow = bow_vectorizer.fit_transform(X_train)
34
[page 35]
X_test_bow = bow_vectorizer.transform(X_test)
# Train Logistic Regression
lr_bow = LogisticRegression(max_iter=1000, class_weight='balanced',␣
↪random_state=42)
lr_bow.fit(X_train_bow, y_train)
y_pred = lr_bow.predict(X_test_bow)
# Metrics
acc = accuracy_score(y_test, y_pred)
f1 = f1_score(y_test, y_pred, average='weighted')
results.append((ngram, acc, f1))
print(f"ngram_range={ngram} | Accuracy={acc:.3f} | F1 (weighted)={f1:.3f}")
# Convert to DataFrame for easier handling
df_res = pd.DataFrame(results, columns=["N-gram Range", "Accuracy", "F1"])
# Plotting
fig, axes = plt.subplots(2, 2, figsize=(12, 8))
axes = axes.ravel()
for idx, row in df_res.iterrows():
ngram = row["N-gram Range"]
acc = row["Accuracy"]
f1 = row["F1"]
ax = axes[idx]
ax.bar(["Accuracy", "F1"], [acc, f1], color=["skyblue", "orange"])
ax.set_ylim(0, 1)
ax.set_title(f"N-gram {ngram}")
for i, v in enumerate([acc, f1]):
ax.text(i, v + 0.01, f"{v:.3f}", ha="center")
plt.suptitle("Logistic Regression with Different N-gram Ranges", fontsize=14)
plt.tight_layout(rect=[0, 0, 1, 0.96])
plt.show()
ngram_range=(1, 1) | Accuracy=0.641 | F1 (weighted)=0.651
ngram_range=(1, 2) | Accuracy=0.629 | F1 (weighted)=0.639
ngram_range=(1, 3) | Accuracy=0.620 | F1 (weighted)=0.630
ngram_range=(1, 4) | Accuracy=0.624 | F1 (weighted)=0.634
35
[page 36]
Based on the results, we selected N-gram (1,1) for further experiments.
16 Final Sentiment Classification Model for Logistic Regression
based on Bow
[ ]: # Initialize and fit CountVectorizer
bow_vectorizer = CountVectorizer(max_features=3000, ngram_range=(1, 1))
# Transform the training and testing data
X_train_bow = bow_vectorizer.fit_transform(X_train)
X_test_bow = bow_vectorizer.transform(X_test)
print("--- Logistic Regression with BoW Features ---")
lr_bow = LogisticRegression(max_iter=1000, class_weight='balanced',␣
↪random_state=42)
lr_bow.fit(X_train_bow, y_train)
y_pred_lr_bow = lr_bow.predict(X_test_bow)
print("Classification Report:\n")
print(classification_report(y_test, y_pred_lr_bow, digits=3))
--- Logistic Regression with BoW Features ---
36
[page 37]
Classification Report:
precision recall f1-score support
negative 0.277 0.378 0.319 172
neutral 0.749 0.682 0.714 626
positive 0.706 0.693 0.699 371
accuracy 0.641 1169
macro avg 0.577 0.584 0.578 1169
weighted avg 0.666 0.641 0.651 1169
17 Method 2: TF-IDF
Term Frequency-Inverse Document Frequency (TF-IDF) improves upon BoW by weighting words
based on their importance. It assigns a higher weight to words that are frequent in a document
but rare across the entire corpus.
[ ]: # Initialize and fit TfidfVectorizer
tfidf_vectorizer = TfidfVectorizer(max_features=3000, ngram_range=(1, 2))
# Transform the training and testing data
X_train_tfidf = tfidf_vectorizer.fit_transform(X_train)
X_test_tfidf = tfidf_vectorizer.transform(X_test)
print("Shape of TF-IDF training matrix:", X_train_tfidf.shape)
print("Shape of TF-IDF testing matrix:", X_test_tfidf.shape)
Shape of TF-IDF training matrix: (4673, 3000)
Shape of TF-IDF testing matrix: (1169, 3000)
[ ]: print("First training sentence:\n", X_train.iloc[0])
# Btfidf features for this sentence
row0 = X_train_tfidf[0]
nz_idx = row0.nonzero()[1]
tokens = tfidf_vectorizer.get_feature_names_out()
print("\nNon-zero tokens for this sentence:")
for i, v in zip(nz_idx, row0.data):
print(f"Pos {i} -> {tokens[i]} : {v}")
First training sentence:
national conciliator juhani salonius met party wednesday said far apart view
propose mediation
Non-zero tokens for this sentence:
37
[page 38]
Pos 1616 -> national : 0.3245414490446402
Pos 1253 -> juhani : 0.3879186184658618
Pos 1511 -> met : 0.3788978370765597
Pos 1941 -> party : 0.32186634801993047
Pos 2937 -> wednesday : 0.30634741367509555
Pos 2328 -> said : 0.17037105069902347
Pos 820 -> far : 0.3335622304339423
Pos 2908 -> view : 0.3712708750859625
Pos 2132 -> propose : 0.3536235559598483
18 Hyperparameter search for the optimum feature size for Tfidf
• The search space is [100, 500, 1000, 3000])
• We used Logistic Regression for sentiment classification
[ ]: import matplotlib.pyplot as plt
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, f1_score
import pandas as pd
# Different feature sizes to test
feature_sizes = [100, 500, 1000, 3000]
results = []
for size in feature_sizes:
# Vectorize with TF-IDF for given max_features
tfidf_vectorizer = TfidfVectorizer(max_features=size, ngram_range=(1, 1))
X_train_tfidf = tfidf_vectorizer.fit_transform(X_train)
X_test_tfidf = tfidf_vectorizer.transform(X_test)
# Train Logistic Regression
lr_tfidf = LogisticRegression(max_iter=1000, class_weight='balanced',␣
↪random_state=42)
lr_tfidf.fit(X_train_tfidf, y_train)
y_pred = lr_tfidf.predict(X_test_tfidf)
# Metrics
acc = accuracy_score(y_test, y_pred)
f1 = f1_score(y_test, y_pred, average='weighted')
results.append((size, acc, f1))
print(f"Feature size={size} | Accuracy={acc:.3f} | F1 (weighted)={f1:.3f}")
# Convert to DataFrame
df_res = pd.DataFrame(results, columns=["Features", "Accuracy", "F1"])
38
[page 39]
# Plotting
fig, axes = plt.subplots(2, 2, figsize=(12, 8))
axes = axes.ravel()
for idx, row in df_res.iterrows():
size = row["Features"]
acc = row["Accuracy"]
f1 = row["F1"]
ax = axes[idx]
ax.bar(["Accuracy", "F1"], [acc, f1], color=["skyblue", "orange"])
ax.set_ylim(0, 1)
ax.set_title(f"TF-IDF Features = {size}")
for i, v in enumerate([acc, f1]):
ax.text(i, v + 0.01, f"{v:.3f}", ha="center")
plt.suptitle("Logistic Regression with Different TF-IDF Feature Sizes",␣
↪fontsize=14)
plt.tight_layout(rect=[0, 0, 1, 0.96])
plt.show()
Feature size=100 | Accuracy=0.476 | F1 (weighted)=0.496
Feature size=500 | Accuracy=0.611 | F1 (weighted)=0.625
Feature size=1000 | Accuracy=0.635 | F1 (weighted)=0.649
Feature size=3000 | Accuracy=0.656 | F1 (weighted)=0.668
39
[page 40]
Based on the results, we selected max_features = 3000 for further experiments.
19 Hyperparameter search for the optimum n-gram size with Tfidf
• The search space is [(1,1), (1,2), (1,3), (1,4)]
• We used Logistic Regression for sentiment classification
[ ]: import matplotlib.pyplot as plt
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import accuracy_score, f1_score
import pandas as pd
# Different n-gram ranges to test
ngram_ranges = [(1,1), (1,2), (1,3), (1,4)]
results = []
for ngram in ngram_ranges:
# Vectorize with TF-IDF for given n-gram range
tfidf_vectorizer = TfidfVectorizer(max_features=3000, ngram_range=ngram)
X_train_tfidf = tfidf_vectorizer.fit_transform(X_train)
X_test_tfidf = tfidf_vectorizer.transform(X_test)
# Train Logistic Regression
lr_tfidf = LogisticRegression(max_iter=1000, class_weight='balanced',␣
↪random_state=42)
lr_tfidf.fit(X_train_tfidf, y_train)
y_pred = lr_tfidf.predict(X_test_tfidf)
# Metrics
acc = accuracy_score(y_test, y_pred)
f1 = f1_score(y_test, y_pred, average='weighted')
results.append((ngram, acc, f1))
print(f"ngram_range={ngram} | Accuracy={acc:.3f} | F1 (weighted)={f1:.3f}")
# Convert to DataFrame
df_res = pd.DataFrame(results, columns=["N-gram Range", "Accuracy", "F1"])
# Plotting
fig, axes = plt.subplots(2, 2, figsize=(12, 8))
axes = axes.ravel()
for idx, row in df_res.iterrows():
40
[page 41]
ngram = row["N-gram Range"]
acc = row["Accuracy"]
f1 = row["F1"]
ax = axes[idx]
ax.bar(["Accuracy", "F1"], [acc, f1], color=["skyblue", "orange"])
ax.set_ylim(0, 1)
ax.set_title(f"TF-IDF ngram={ngram}")
for i, v in enumerate([acc, f1]):
ax.text(i, v + 0.01, f"{v:.3f}", ha="center")
plt.suptitle("Logistic Regression with Different TF-IDF N-gram Ranges",␣
↪fontsize=14)
plt.tight_layout(rect=[0, 0, 1, 0.96])
plt.show()
ngram_range=(1, 1) | Accuracy=0.656 | F1 (weighted)=0.668
ngram_range=(1, 2) | Accuracy=0.663 | F1 (weighted)=0.674
ngram_range=(1, 3) | Accuracy=0.654 | F1 (weighted)=0.666
ngram_range=(1, 4) | Accuracy=0.649 | F1 (weighted)=0.662
Based on the results, we selected N-gram (1,2) for further experiments.
41
[page 42]
20 Final Sentiment Classification Model for Logistic Regression
based on Tfidf
[ ]: tfidf_vectorizer = TfidfVectorizer(max_features=3000, ngram_range=(1,2))
X_train_tfidf = tfidf_vectorizer.fit_transform(X_train)
X_test_tfidf = tfidf_vectorizer.transform(X_test)
print("--- Logistic Regression with TF-IDF Features ---")
lr_tfidf = LogisticRegression(max_iter=1000, class_weight='balanced',␣
↪random_state=42)
lr_tfidf.fit(X_train_tfidf, y_train)
y_pred_lr_tfidf = lr_tfidf.predict(X_test_tfidf)
print("Classification Report:\n")
print(classification_report(y_test, y_pred_lr_tfidf, digits=3))
--- Logistic Regression with TF-IDF Features ---
Classification Report:
precision recall f1-score support
negative 0.354 0.529 0.424 172
neutral 0.784 0.685 0.731 626
positive 0.699 0.687 0.693 371
accuracy 0.663 1169
macro avg 0.612 0.634 0.616 1169
weighted avg 0.694 0.663 0.674 1169
21 Comparing with Multinomial Naive Bayes and Random Forest
[ ]: from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer
from sklearn.naive_bayes import MultinomialNB
from sklearn.ensemble import RandomForestClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import classification_report
# BoW
bow_vectorizer = CountVectorizer(max_features=3000, ngram_range=(1, 1))
X_train_bow = bow_vectorizer.fit_transform(X_train)
X_test_bow = bow_vectorizer.transform(X_test)
# TF-IDF
tfidf_vectorizer = TfidfVectorizer(max_features=3000, ngram_range=(1, 2))
X_train_tfidf = tfidf_vectorizer.fit_transform(X_train)
42
[page 43]
X_test_tfidf = tfidf_vectorizer.transform(X_test)
print("\n--- Naive Bayes with BoW Features ---")
nb_bow = MultinomialNB()
nb_bow.fit(X_train_bow, y_train)
y_pred_nb_bow = nb_bow.predict(X_test_bow)
print(classification_report(y_test, y_pred_nb_bow, digits=3))
print("\n--- Naive Bayes with TF-IDF Features ---")
nb_tfidf = MultinomialNB()
nb_tfidf.fit(X_train_tfidf, y_train)
y_pred_nb_tfidf = nb_tfidf.predict(X_test_tfidf)
print(classification_report(y_test, y_pred_nb_tfidf, digits=3))
print("\n--- Random Forest with BoW Features ---")
rf_bow = RandomForestClassifier(n_estimators=100, random_state=42,␣
↪class_weight="balanced")
rf_bow.fit(X_train_bow, y_train)
y_pred_rf_bow = rf_bow.predict(X_test_bow)
print(classification_report(y_test, y_pred_rf_bow, digits=3))
print("\n--- Random Forest with TF-IDF Features ---")
rf_tfidf = RandomForestClassifier(n_estimators=100, random_state=42,␣
↪class_weight="balanced")
rf_tfidf.fit(X_train_tfidf, y_train)
y_pred_rf_tfidf = rf_tfidf.predict(X_test_tfidf)
print(classification_report(y_test, y_pred_rf_tfidf, digits=3))
--- Naive Bayes with BoW Features ---
precision recall f1-score support
negative 0.338 0.448 0.385 172
neutral 0.770 0.727 0.748 626
positive 0.706 0.666 0.685 371
accuracy 0.666 1169
macro avg 0.604 0.613 0.606 1169
43
[page 44]
weighted avg 0.686 0.666 0.675 1169
--- Naive Bayes with TF-IDF Features ---
precision recall f1-score support
negative 0.444 0.116 0.184 172
neutral 0.698 0.899 0.786 626
positive 0.681 0.582 0.628 371
accuracy 0.683 1169
macro avg 0.608 0.533 0.533 1169
weighted avg 0.655 0.683 0.647 1169
--- Random Forest with BoW Features ---
precision recall f1-score support
negative 0.233 0.203 0.217 172
neutral 0.663 0.760 0.708 626
positive 0.751 0.609 0.673 371
accuracy 0.630 1169
macro avg 0.549 0.524 0.533 1169
weighted avg 0.628 0.630 0.625 1169
--- Random Forest with TF-IDF Features ---
precision recall f1-score support
negative 0.226 0.174 0.197 172
neutral 0.663 0.781 0.717 626
positive 0.748 0.601 0.667 371
accuracy 0.635 1169
macro avg 0.545 0.519 0.527 1169
weighted avg 0.626 0.635 0.624 1169
[ ]: results = {
"Model": [
"Logistic Regression (BoW)", "Logistic Regression (TF-IDF)",
"Naive Bayes (BoW)", "Naive Bayes (TF-IDF)",
"Random Forest (BoW)", "Random Forest (TF-IDF)"
],
"Accuracy": [
44
[page 45]
accuracy_score(y_test, y_pred_lr_bow), accuracy_score(y_test,␣
↪y_pred_lr_tfidf),
accuracy_score(y_test, y_pred_nb_bow), accuracy_score(y_test,␣
↪y_pred_nb_tfidf),
accuracy_score(y_test, y_pred_rf_bow), accuracy_score(y_test,␣
↪y_pred_rf_tfidf)
],
"F1-Score": [
f1_score(y_test, y_pred_lr_bow, average='weighted'), f1_score(y_test,␣
↪y_pred_lr_tfidf, average='weighted'),
f1_score(y_test, y_pred_nb_bow, average='weighted'), f1_score(y_test,␣
↪y_pred_nb_tfidf, average='weighted'),
f1_score(y_test, y_pred_rf_bow, average='weighted'), f1_score(y_test,␣
↪y_pred_rf_tfidf, average='weighted')
]
}
df_results = pd.DataFrame(results).sort_values(by="F1-Score", ascending=False).
↪reset_index(drop=True)
df_results
[ ]: Model Accuracy F1-Score
0 Naive Bayes (BoW) 0.666382 0.674508
1 Logistic Regression (TF-IDF) 0.662960 0.674030
2 Logistic Regression (BoW) 0.640719 0.651308
3 Naive Bayes (TF-IDF) 0.683490 0.647174
4 Random Forest (BoW) 0.630453 0.624764
5 Random Forest (TF-IDF) 0.634731 0.624480
Observations
1. Naive Bayes with TF-IDF achieves the highest accuracy (0.683) but its weighted F1-score
(0.647) is lower than Naive Bayes (BoW) and Logistic Regression (TF-IDF), showing that
while it predicts the majority class well, its balance across all classes is weaker.
2. Naive Bayes (BoW) and Logistic Regression (TF-IDF) show more balanced per-
formance, with F1-scores around, making them more reliable choices compared to Random
Forest, which consistently underperforms in both Accuracy and F1.
#3. Clustering with K-Means
22 Finding the optimal number of clusters using tfidf vectorization
[ ]: import pandas as pd
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.cluster import KMeans
import matplotlib.pyplot as plt
45
[page 46]
# TF-IDF vectorization
tfidf = TfidfVectorizer(max_features=5000, stop_words='english',␣
↪ngram_range=(1,2))
X_tfidf = tfidf.fit_transform(df['cleaned_sentence'])
# Elbow method to find optimal K
inertia = []
K_range = range(2, 11)
for k in K_range:
km = KMeans(n_clusters=k, random_state=42, n_init=10)
km.fit(X_tfidf)
inertia.append(km.inertia_) # within-cluster sum of squares
# Step 3: Plot elbow curve
plt.figure(figsize=(8,6))
plt.plot(K_range, inertia, 'bo-')
plt.xlabel("Number of clusters (k)")
plt.ylabel("Inertia (WCSS)")
plt.title("Elbow Method for Optimal k")
plt.show()
46
[page 47]
[ ]: k_values = list(range(2, 11))
df_elbow = pd.DataFrame({
"k": k_values,
"Inertia": inertia
})
# Add drop magnitude (difference from previous inertia)
df_elbow["Drop"] = df_elbow["Inertia"].shift(1) - df_elbow["Inertia"]
print(df_elbow)
k Inertia Drop
0 2 5626.551513 NaN
1 3 5570.782532 55.768981
2 4 5531.914439 38.868093
3 5 5513.797712 18.116727
4 6 5494.885546 18.912166
5 7 5460.300293 34.585253
6 8 5449.479987 10.820306
7 9 5434.234743 15.245244
8 10 5430.639985 3.594757
The largest drop in inertia occurs when moving from k=2 to k=3
23 Visualising clusters for k=[2,3,4,5]
[ ]: import matplotlib.pyplot as plt
from sklearn.decomposition import PCA
from sklearn.cluster import KMeans
# Apply PCA on TF-IDF
X_pca = PCA(n_components=2, random_state=42).fit_transform(X_tfidf.toarray())
# Run KMeans for k=2,3,4,5
k_values = [2, 3, 4, 5]
cluster_results = {}
for k in k_values:
km = KMeans(n_clusters=k, random_state=42, n_init=10)
cluster_results[k] = km.fit_predict(X_tfidf)
# 2×3 subplot
fig, axes = plt.subplots(2, 3, figsize=(15, 10))
# Ground Truth plot with legend
color_map = {"positive": "red", "negative": "blue", "neutral": "green"}
for sentiment, color in color_map.items():
47
[page 48]
mask = df['Sentiment'] == sentiment
axes[0, 0].scatter(
X_pca[mask, 0],
X_pca[mask, 1],
c=color,
label=sentiment,
alpha=0.5
)
axes[0, 0].set_title("Ground Truth (Sentiment Labels)")
axes[0, 0].legend(title="Sentiment")
# KMeans plots
for i, k in enumerate(k_values):
ax = axes[(i+1)//3, (i+1)%3] # shift index by 1 to leave [0,0] for ground ␣
↪truth
ax.scatter(
X_pca[:, 0],
X_pca[:, 1],
c=cluster_results[k],
cmap="tab10",
alpha=0.7
)
ax.set_title(f"KMeans (k={k})")
# Leave last subplot blank
axes[1, 2].axis("off")
plt.tight_layout()
plt.show()
48
[page 49]
Although the elbow method suggests that k=3 provides the most reasonable parti-
tion, the clusters are not cleanly separated in the PCA visualizations. This indicates
that sentiment-based separation in the Financial PhraseBank dataset is not strongly
captured by unsupervised clustering on BoW/TF-IDF features, making the results
less convincing for sentiment grouping.
[ ]: from sklearn.metrics import (
silhouette_score, davies_bouldin_score, calinski_harabasz_score,
homogeneity_score, completeness_score, v_measure_score
)
# Final KMeans with k=3
k = 3
kmeans = KMeans(n_clusters=k, random_state=42, n_init=10)
y_pred_kmeans = kmeans.fit_predict(X_tfidf)
# Intrinsic metrics
sil_score = silhouette_score(X_tfidf, y_pred_kmeans)
db_score = davies_bouldin_score(X_tfidf.toarray(), y_pred_kmeans)
ch_score = calinski_harabasz_score(X_tfidf.toarray(), y_pred_kmeans)
# Extrinsic metrics (need ground truth labels)
h_score = homogeneity_score(df['Sentiment'], y_pred_kmeans)
c_score = completeness_score(df['Sentiment'], y_pred_kmeans)
49
[page 50]
v_score = v_measure_score(df['Sentiment'], y_pred_kmeans)
# Print
print("Intrinsic Metrics:")
print(f" Silhouette Score: {sil_score:.3f}")
print(f" Davies-Bouldin Index: {db_score:.3f}")
print(f" Calinski-Harabasz Score: {ch_score:.3f}\n")
print("Extrinsic Metrics :")
print(f" Homogeneity: {h_score:.3f}")
print(f" Completeness: {c_score:.3f}")
print(f" V-Measure: {v_score:.3f}")
Intrinsic Metrics:
Silhouette Score: 0.010
Davies-Bouldin Index: 6.243
Calinski-Harabasz Score: 92.579
Extrinsic Metrics :
Homogeneity: 0.007
Completeness: 0.009
V-Measure: 0.008
• The clustering results are very weak. The intrinsic metrics show almost no
meaningful separation (Silhouette 0, very high Davies–Bouldin score, and low
Calinski–Harabasz value), indicating that the clusters are not compact or well-
separated. The extrinsic metrics are also near zero (Homogeneity , Completeness,
and V-Measure all < 0.01), which confirms that the clusters do not align with
the actual sentiment labels. This suggests that unsupervised clustering on simple
TF-IDF features is not effective for sentiment-based grouping in this dataset.
• Moreover, the PCA visualization of the ground truth labels itself shows heavy
overlap between positive, negative, and neutral samples, which further explains
why the clusters fail to appear well separated.
24 Finding the optimal number of clusters using bow vectorization
[ ]: # Vectorize the full dataset using BoW
bow_vectorizer = CountVectorizer(max_features=5000, stop_words="english",␣
↪ngram_range=(1,2))
X_bow_full = bow_vectorizer.fit_transform(df['cleaned_sentence'])
pca_bow = PCA(n_components=2, random_state=42)
X_pca_bow = pca_bow.fit_transform(X_bow_full.toarray())
# Run KMeans for k=2,3,4,5
k_values = [2, 3, 4, 5]
cluster_results = {}
50
[page 51]
for k in k_values:
km = KMeans(n_clusters=k, random_state=42, n_init=10)
cluster_results[k] = km.fit_predict(X_bow_full)
# 2×3 subplot
fig, axes = plt.subplots(2, 3, figsize=(15, 10))
# Ground Truth plot with legend
color_map = {"positive": "red", "negative": "blue", "neutral": "green"}
for sentiment, color in color_map.items():
mask = df['Sentiment'] == sentiment
axes[0, 0].scatter(
X_pca_bow[mask, 0],
X_pca_bow[mask, 1],
c=color,
label=sentiment,
alpha=0.5
)
axes[0, 0].set_title("Ground Truth (Sentiment Labels)")
axes[0, 0].legend(title="Sentiment")
# KMeans plots
for i, k in enumerate(k_values):
ax = axes[(i+1)//3, (i+1)%3]
ax.scatter(
X_pca_bow[:, 0],
X_pca_bow[:, 1],
c=cluster_results[k],
cmap="tab10",
alpha=0.7
)
ax.set_title(f"KMeans (k={k})")
# Leave last subplot blank
axes[1, 2].axis("off")
plt.tight_layout()
plt.show()
51
[page 52]
Cluster separation is poor: Just like TF-IDF, BoW also fails to produce clearly sepa-
rated clusters in PCA space.
25 Better Cluster Separation with Topic Categories dataset
(NewsGroup dataset):
[ ]: import pandas as pd
import re
from sklearn.datasets import fetch_20newsgroups
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.cluster import KMeans
from sklearn.decomposition import PCA
import matplotlib.pyplot as plt
# Load dataset (3 categories)
categories = ['rec.sport.baseball', 'sci.space', 'talk.politics.mideast']
newsgroups = fetch_20newsgroups(
subset='train',
categories=categories,
remove=('headers','footers','quotes')
)
52
[page 53]
df_alt = pd.DataFrame({
'sentence': newsgroups.data,
'Label': newsgroups.target
})
df_alt['Category'] = df_alt['Label'].map(dict(enumerate(newsgroups.
↪target_names)))
print(newsgroups.target_names)
['rec.sport.baseball', 'sci.space', 'talk.politics.mideast']
[ ]: df_alt[:10]
[ ]: sentence Label \
0 Reply address: mark.prado@permanet.org\n\n > F… 1
1 \n\nYeah Valentine, how many rings does Clemen… 0
2 yeah,\n\nThey just tore down the Kmart near my… 1
3 I've been to three talks in the last month whi… 1
4 \nProbably not--he's just singing someone else… 2
5 \n\n Let us hope that the performance o… 1
6 Ethnocentric USian that I am, I've assumed tha… 1
7 \n\nWell, OBP is the most important offensive … 0
8 \nOne consideration to remember is that if you… 1
9 From Israel Line, Thursday, April 22, 1993:\n … 2
Category
0 sci.space
1 rec.sport.baseball
2 sci.space
3 sci.space
4 talk.politics.mideast
5 sci.space
6 sci.space
7 rec.sport.baseball
8 sci.space
9 talk.politics.mideast
[ ]: # TF-IDF vectorization
tfidf = TfidfVectorizer(stop_words='english',␣
↪max_features=3000,ngram_range=(1,2))
X_tfidf = tfidf.fit_transform(df_alt['sentence'])
# KMeans clustering for k=3 (since we know 3 topics)
k = 3
km = KMeans(n_clusters=k, random_state=42, n_init=10)
53
[page 54]
clusters = km.fit_predict(X_tfidf)
# PCA to 2D for visualization
X_pca = PCA(n_components=2, random_state=42).fit_transform(X_tfidf.toarray())
# Plot ground truth vs clustering side by side
fig, axes = plt.subplots(1, 2, figsize=(14, 6))
# Ground truth labels
axes[0].scatter(X_pca[:,0], X_pca[:,1], c=df_alt['Label'], cmap="tab10",␣
↪alpha=0.7)
axes[0].set_title("Ground Truth (Topics)")
axes[0].set_xlabel("PC1")
axes[0].set_ylabel("PC2")
# KMeans clustering result
axes[1].scatter(X_pca[:,0], X_pca[:,1], c=clusters, cmap="tab10", alpha=0.7)
axes[1].set_title("KMeans Clustering (k=3)")
axes[1].set_xlabel("PC1")
axes[1].set_ylabel("PC2")
plt.tight_layout()
plt.show()
• Clustering works well on this subset of the 20 Newsgroups dataset because the categories
(sports, space, politics) have distinct vocabularies, so TF-IDF highlights words that strongly
separate the topics, and PCA+KMeans reveals clear cluster boundaries.
• In contrast, the Financial PhraseBank dataset with sentiment labels does not show good
clustering because positive, negative, and neutral texts often share much of the same financial
terminology, making their word distributions highly overlapping and harder to separate in an
unsupervised setting.
54
[page 55]
[ ]:
55