# LSTM vf
course: Module 3 — Deep Learning & NLP
module: Module-3-Deep-Learning-NLP
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-3-Deep-Learning-NLP/General/Lab_Materials-22-03-2026/LSTM_vf.pdf
pages: 53
---
[page 1]
Dow Jones Index Dataset
The Dow Jones Index dataset contains weekly stock market data for 30 major companies in the Dow Jones Industrial Average. Each record
represents one week of trading activity for a company, including price movements, trading volume, and dividend information.
Although the dataset provides many features, in this study we focus on predicting the next week’s closing price using a Long Short-Term
Memory (LSTM) network, a special type of Recurrent Neural Network (RNN).
The dataset covers a limited time span: 6 months of 2011 (January–June) — which results in about 750 weekly records across 30 stocks.
Source: UCI Machine Learning Repository
UCI Dataset Page
Target Variable:
next_weeks_close (numeric): Closing price of the stock for the following week.
Features:
Column Name Description Data Type
quarter Fiscal quarter of data (1–4) Integer
stock Stock ticker symbol (e.g., AA , AXP ) Categorical
date Week-ending date of the record Date
open Stock price at market open (that week) Numeric (float)
high Highest stock price during the week Numeric (float)
low Lowest stock price during the week Numeric (float)
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 1/53
[page 2]
Column Name Description Data Type
close Stock price at market close (that week) Numeric (float)
volume Total trading volume during the week Numeric (float)
percent_change_price Percent change from opening to closing price Numeric (float)
percent_change_volume_over_last_wk Percent change in volume relative to the previous week Numeric (float)
previous_weeks_volume Trading volume of the previous week Numeric (float)
next_weeks_open Stock price at market open (next week) Numeric (float)
next_weeks_close Stock price at market close (next week) — Target Numeric
percent_change_next_weeks_price Percent change in price from next week’s open to close Numeric (float)
days_to_next_dividend Number of days until the next dividend is paid Integer
percent_return_next_dividend Percent return from the next dividend relative to stock price Numeric (float)
Load into pandas dataframe
displaying the head
import pandas as pd
# Load file with first row as header
df = pd.read_csv("dow_jones_index.data")
df.head()
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 2/53
[page 3]
quarter stock date open high low close volume percent_change_price percent_change_volume_over_last_wk p
0 1 AA 1/7/2011 $15.82 $16.72 $15.78 $16.42 239655616 3.79267 NaN
1 1 AA 1/14/2011 $16.71 $16.71 $15.64 $15.97 242963398 -4.42849 1.380223
2 1 AA 1/21/2011 $16.19 $16.38 $15.60 $15.79 138428495 -2.47066 -43.024959
3 1 AA 1/28/2011 $15.87 $16.63 $15.82 $16.13 151379173 1.63831 9.355500
4 1 AA 2/4/2011 $16.18 $17.39 $16.18 $17.14 154387761 5.93325 1.987452
df.info()
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 750 entries, 0 to 749
Data columns (total 16 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 quarter 750 non-null int64
1 stock 750 non-null object
2 date 750 non-null object
3 open 750 non-null object
4 high 750 non-null object
5 low 750 non-null object
6 close 750 non-null object
7 volume 750 non-null int64
8 percent_change_price 750 non-null float64
9 percent_change_volume_over_last_wk 720 non-null float64
10 previous_weeks_volume 720 non-null float64
11 next_weeks_open 750 non-null object
12 next_weeks_close 750 non-null object
13 percent_change_next_weeks_price 750 non-null float64
14 days_to_next_dividend 750 non-null int64
15 percent_return_next_dividend 750 non-null float64
dtypes: float64(5), int64(3), object(8)
memory usage: 93.9+ KB
For predicting next week’s closing price, we must avoid any feature that already leaks future information. In this dataset, the following
cannot be used:
Out[ ]:
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 3/53
[page 4]
next_weeks_open → tells us the stock’s opening price in the future week.
next_weeks_close → is literally the target variable, and is redundant since it equals the following row’s close , so it cannot be an
input feature.
percent_change_next_weeks_price → is computed using the next week’s open and close, so it directly encodes future
outcome.
There are missing values in some columns ( percent_change_volume_over_last_wk , previous_weeks_volume ), but since we are
not using these features for our time series modeling, we do not fill the missing values.
Preprocessing Steps
Clean price columns : Removed $ and converted open , high , low , close , next_weeks_open , next_weeks_close to
float.
Convert date column : Changed date to datetime format for proper time handling.
Sort for time series: Sorted by stock and date so data is ordered correctly for sequence modeling.
import pandas as pd
def preprocess_dow_jones(path):
# Load raw CSV
df = pd.read_csv(path)
# Remove $ and convert to float for price columns
money_cols = ["open", "high", "low", "close",
"next_weeks_open", "next_weeks_close"]
for col in money_cols:
df[col] = df[col].str.replace("$", "", regex=False).astype(float)
# Convert date column to datetime
df["date"] = pd.to_datetime(df["date"])
# Sort values for time-series use (per stock, by date)
df = df.sort_values(by=["stock", "date"]).reset_index(drop=True)
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 4/53
[page 5]
return df
df_v1 = preprocess_dow_jones("dow_jones_index.data")
df_v1.head()
quarter stock date open high low close volume percent_change_price percent_change_volume_over_last_wk previous_w
0 1 AA 2011-
01-07 15.82 16.72 15.78 16.42 239655616 3.79267 NaN
1 1 AA 2011-
01-14 16.71 16.71 15.64 15.97 242963398 -4.42849 1.380223
2 1 AA 2011-
01-21 16.19 16.38 15.60 15.79 138428495 -2.47066 -43.024959
3 1 AA 2011-
01-28 15.87 16.63 15.82 16.13 151379173 1.63831 9.355500
4 1 AA 2011-
02-04 16.18 17.39 16.18 17.14 154387761 5.93325 1.987452
df_v1.info()
Out[ ]:
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 5/53
[page 6]
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 750 entries, 0 to 749
Data columns (total 16 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 quarter 750 non-null int64
1 stock 750 non-null object
2 date 750 non-null datetime64[ns]
3 open 750 non-null float64
4 high 750 non-null float64
5 low 750 non-null float64
6 close 750 non-null float64
7 volume 750 non-null int64
8 percent_change_price 750 non-null float64
9 percent_change_volume_over_last_wk 720 non-null float64
10 previous_weeks_volume 720 non-null float64
11 next_weeks_open 750 non-null float64
12 next_weeks_close 750 non-null float64
13 percent_change_next_weeks_price 750 non-null float64
14 days_to_next_dividend 750 non-null int64
15 percent_return_next_dividend 750 non-null float64
dtypes: datetime64[ns](1), float64(11), int64(3), object(1)
memory usage: 93.9+ KB
Records per Stock
df_v1["stock"].value_counts()
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 6/53
[page 7]
count
stock
AA 25
AXP 25
BA 25
BAC 25
CAT 25
CSCO 25
CVX 25
DD 25
DIS 25
GE 25
HD 25
HPQ 25
IBM 25
INTC 25
JNJ 25
JPM 25
KO 25
KRFT 25
MCD 25
MMM 25
MRK 25
MSFT 25
PFE 25
Out[ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 7/53
[page 8]
count
stock
PG 25
T 25
TRV 25
UTX 25
VZ 25
WMT 25
XOM 25
dtype: int64
Distribution of Closing Prices for Each Stock
import seaborn as sns
import matplotlib.pyplot as plt
plt.figure(figsize=(12,6))
sns.boxplot(x="stock", y="close", data=df_v1)
plt.xticks(rotation=90)
plt.title("Distribution of Closing Prices by Company")
plt.show()
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 8/53
[page 9]
1. Most companies have weekly closing prices that stay within a narrow and stable range i.e. their stock prices do not fluctuate much.
2. A few companies like IBM and CAT not only have higher average closing prices but also show larger ups and downs, indicating greater
volatility.
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 9/53
[page 10]
Setting fix random seeds across libraries to ensure
reproducible results in experiments.
import random
import numpy as np
import torch
SEED = 42
np.random.seed(SEED)
torch.manual_seed(SEED)
random.seed(SEED)
Assumption
Since each company in the Dow Jones UCI dataset has only about 25 weeks of data, we obtain very few company-specific
training examples for LSTM models.
For simplicity and to focus on understanding how time series modeling works, we treat the closing prices in this dataset as if
they belonged to a single company’s continuous time series, even though in reality they are grouped by different companies.
LSTM Setup for Stock Prediction
Tuning the Optimal Sequence Length
The tuning is done with sequence lengths of 4,5, 6,7,8 — meaning we use the closing prices of the past 4,5, 6, 7, or 8 weeks to predict
the closing price of the next week, effectively making this a univariate time series since only one feature (closing price) is used as
input.
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 10/53
[page 11]
Key Functions & Approach
Creating Sequences: The core of time-series modeling is transforming the flat list of closing prices into overlapping sequences. The
create_sequences function creates input windows ( X ) of a specified seq_len and corresponding targets ( y ). The choice of
seq_len is a critical hyperparameter that determines how much past information the model uses.
Train–Test Split: The data is split chronologically (80% for training, 20% for testing) to mimic a real-world scenario where we predict
the future based on the past. Random shuffling is avoided as it would destroy the temporal order.
Data Scaling: Stock prices are scaled to a [0, 1] range using MinMaxScaler. This helps stabilize the training process and is suitable for
price data, which is always non-negative.
LSTM Model: We use an LSTM network, which is well-suited for capturing temporal dependencies. Unlike a simple RNN, an LSTM
contains memory cells and gates that allow it to learn and remember information over longer sequences, making it more robust
against issues like the vanishing gradient problem.
import pandas as pd
import numpy as np
import torch
import torch.nn as nn
import torch.optim as optim
from sklearn.preprocessing import MinMaxScaler
from sklearn.metrics import mean_squared_error, r2_score, mean_absolute_error
import matplotlib.pyplot as plt
# Prepare the data
data = df_v1["close"].values.reshape(-1, 1)
scaler = MinMaxScaler(feature_range=(0, 1))
data_scaled = scaler.fit_transform(data)
torch.manual_seed(SEED)
# Function to create sequences
def create_sequences(data, seq_len=5):
X, y = [], []
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 11/53
[page 12]
for i in range(len(data) - seq_len):
X.append(data[i:i+seq_len])
y.append(data[i+seq_len])
return np.array(X), np.array(y)
# LSTM Model Definition
class LSTMModel(nn.Module):
def __init__(self, input_size=1, hidden_size=32, num_layers=1):
super(LSTMModel, self).__init__()
self.lstm = nn.LSTM(input_size, hidden_size, num_layers, batch_first=True)
self.fc = nn.Linear(hidden_size, 1)
def forward(self, x):
# x shape: (batch_size, seq_len, input_size)
# Get output, final hidden state (hn), and final cell state (cn)
# We only need the final hidden state
_, (hn, cn) = self.lstm(x)
# hn shape: (num_layers, batch_size, hidden_size)
# We use the hidden state of the last layer for prediction
out = self.fc(hn[-1])
return out
# Training and Evaluation Function
def train_and_evaluate(seq_len):
X, y = create_sequences(data_scaled, seq_len)
train_size = int(len(X) * 0.8)
X_train, X_test = X[:train_size], X[train_size:]
y_train, y_test = y[:train_size], y[train_size:]
X_train = torch.tensor(X_train, dtype=torch.float32)
y_train = torch.tensor(y_train, dtype=torch.float32)
X_test = torch.tensor(X_test, dtype=torch.float32)
y_test = torch.tensor(y_test, dtype=torch.float32)
model = LSTMModel()
criterion = nn.MSELoss()
optimizer = optim.Adam(model.parameters(), lr=0.01)
for epoch in range(50):
model.train()
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 12/53
[page 13]
optimizer.zero_grad()
output = model(X_train)
loss = criterion(output, y_train)
loss.backward()
optimizer.step()
# Evaluate
model.eval()
with torch.no_grad():
y_pred = model(X_test).numpy()
y_true = y_test.numpy()
mse = mean_squared_error(y_true, y_pred)
rmse = np.sqrt(mse)
mae = mean_absolute_error(y_true, y_pred)
r2 = r2_score(y_true, y_pred)
return mse, rmse, mae, r2
# Run experiments for different sequence lengths (NEW RANGE)
seq_lengths = [4,5,6, 7, 8]
results = {"seq_len": [], "mse": [], "rmse": [], "mae": [], "r2": []}
for sl in seq_lengths:
mse, rmse, mae, r2 = train_and_evaluate(sl)
results["seq_len"].append(sl)
results["mse"].append(mse)
results["rmse"].append(rmse)
results["mae"].append(mae)
results["r2"].append(r2)
print(f"Done: seq_len={sl}")
# Plotting function
def plot_results(results):
fig, axs = plt.subplots(2, 2, figsize=(12,8))
axs[0,0].plot(results["seq_len"], results["mse"], marker="o")
axs[0,0].set_title("MSE vs Seq Length")
axs[0,1].plot(results["seq_len"], results["rmse"], marker="o")
axs[0,1].set_title("RMSE vs Seq Length")
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 13/53
[page 14]
axs[1,0].plot(results["seq_len"], results["mae"], marker="o")
axs[1,0].set_title("MAE vs Seq Length")
axs[1,1].plot(results["seq_len"], results["r2"], marker="o")
axs[1,1].set_title("R² vs Seq Length")
for ax in axs.flat:
ax.set_xlabel("Sequence Length")
ax.grid(True)
plt.tight_layout()
plt.show()
plot_results(results)
Done: seq_len=4
Done: seq_len=5
Done: seq_len=6
Done: seq_len=7
Done: seq_len=8
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 14/53
[page 15]
pd.DataFrame(results)
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 15/53
[page 16]
seq_len mse rmse mae r2
0 4 0.002816 0.053068 0.037217 0.831838
1 5 0.001768 0.042050 0.023763 0.893752
2 6 0.002663 0.051604 0.035606 0.839983
3 7 0.001425 0.037752 0.021150 0.914360
4 8 0.002358 0.048560 0.033145 0.858304
From both the graphs and the tabular results, it can be seen that a sequence length of 7 gives the best performance, achieving the lowest
error metrics (MSE, RMSE, MAE) and the highest R² score. This suggests that using the past 6 weeks of data provides the optimal amount
of context for the LSTM to predict the next week's price.
Tuning Hidden Size and Number of Layers in LSTM
Hidden Size: This hyperparameter controls the number of hidden units (or memory cells) in each LSTM layer. It defines the model's
capacity to learn complex patterns. A larger hidden size can capture more intricate dependencies but also increases the risk of
overfitting.
Number of Layers: This determines how many LSTM layers are stacked. Stacking layers allows the model to learn hierarchical temporal
representations, where each layer processes the output sequence of the previous one. Deeper models can capture more abstract
features but are slower to train.
After fixing seq_len at 7, we performed a grid search over hidden_sizes = [32, 64, 128,256] and num_layers = [1,
2,3] , resulting in a total of 12 combinations of hidden size and number of layers.
import pandas as pd
import numpy as np
import torch
import torch.nn as nn
import torch.optim as optim
from sklearn.preprocessing import MinMaxScaler
Out[ ]:
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 16/53
[page 17]
from sklearn.metrics import mean_squared_error, r2_score, mean_absolute_error
import matplotlib.pyplot as plt
torch.manual_seed(SEED)
# Prepare the data
data = df_v1["close"].values.reshape(-1, 1)
scaler = MinMaxScaler(feature_range=(0, 1))
data_scaled = scaler.fit_transform(data)
def create_sequences(data, seq_len=6):
X, y = [], []
for i in range(len(data) - seq_len):
X.append(data[i:i+seq_len])
y.append(data[i+seq_len])
return np.array(X), np.array(y)
# LSTM Model Definition
class LSTMModel(nn.Module):
def __init__(self, input_size=1, hidden_size=32, num_layers=1):
super(LSTMModel, self).__init__()
self.lstm = nn.LSTM(input_size, hidden_size, num_layers, batch_first=True)
self.fc = nn.Linear(hidden_size, 1)
def forward(self, x):
_, (hn, cn) = self.lstm(x)
# Use the hidden state of the last layer for prediction
out = self.fc(hn[-1])
return out
# Training and Evaluation Function
def train_and_evaluate(seq_len, hidden_size, num_layers):
X, y = create_sequences(data_scaled, seq_len)
train_size = int(len(X) * 0.8)
X_train, X_test = X[:train_size], X[train_size:]
y_train, y_test = y[:train_size], y[train_size:]
X_train = torch.tensor(X_train, dtype=torch.float32)
y_train = torch.tensor(y_train, dtype=torch.float32)
X_test = torch.tensor(X_test, dtype=torch.float32)
y_test = torch.tensor(y_test, dtype=torch.float32)
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 17/53
[page 18]
model = LSTMModel(hidden_size=hidden_size, num_layers=num_layers)
criterion = nn.MSELoss()
optimizer = optim.Adam(model.parameters(), lr=0.01)
for epoch in range(50):
model.train()
optimizer.zero_grad()
output = model(X_train)
loss = criterion(output, y_train)
loss.backward()
optimizer.step()
model.eval()
with torch.no_grad():
y_pred = model(X_test).numpy()
y_true = y_test.numpy()
mse = mean_squared_error(y_true, y_pred)
rmse = np.sqrt(mse)
mae = mean_absolute_error(y_true, y_pred)
r2 = r2_score(y_true, y_pred)
return mse, rmse, mae, r2
# Grid Search
seq_len = 7 #
hidden_sizes = [32, 64, 128, 256]
num_layers_list = [1, 2, 3]
results = {"hidden_size": [], "num_layers": [], "mse": [], "rmse": [], "mae": [], "r2": []}
for hs in hidden_sizes:
for nl in num_layers_list:
mse, rmse, mae, r2 = train_and_evaluate(seq_len, hs, nl)
results["hidden_size"].append(hs)
results["num_layers"].append(nl)
results["mse"].append(mse)
results["rmse"].append(rmse)
results["mae"].append(mae)
results["r2"].append(r2)
print(f"Done: hidden_size={hs}, num_layers={nl}")
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 18/53
[page 19]
#Plotting
def plot_tuning_results(df_results):
fig, axs = plt.subplots(2, 2, figsize=(12,8))
for nl in df_results["num_layers"].unique():
subset = df_results[df_results["num_layers"]==nl]
axs[0,0].plot(subset["hidden_size"], subset["mse"], marker="o", label=f"layers={nl}")
axs[0,1].plot(subset["hidden_size"], subset["rmse"], marker="o", label=f"layers={nl}")
axs[1,0].plot(subset["hidden_size"], subset["mae"], marker="o", label=f"layers={nl}")
axs[1,1].plot(subset["hidden_size"], subset["r2"], marker="o", label=f"layers={nl}")
axs[0,0].set_title("MSE vs Hidden Size")
axs[0,1].set_title("RMSE vs Hidden Size")
axs[1,0].set_title("MAE vs Hidden Size")
axs[1,1].set_title("R² vs Hidden Size")
for ax in axs.flat:
ax.set_xlabel("Hidden Size")
ax.legend()
ax.grid(True)
plt.tight_layout()
plt.show()
df_results = pd.DataFrame(results)
plot_tuning_results(df_results)
Done: hidden_size=32, num_layers=1
Done: hidden_size=32, num_layers=2
Done: hidden_size=32, num_layers=3
Done: hidden_size=64, num_layers=1
Done: hidden_size=64, num_layers=2
Done: hidden_size=64, num_layers=3
Done: hidden_size=128, num_layers=1
Done: hidden_size=128, num_layers=2
Done: hidden_size=128, num_layers=3
Done: hidden_size=256, num_layers=1
Done: hidden_size=256, num_layers=2
Done: hidden_size=256, num_layers=3
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 19/53
[page 20]
df_results
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 20/53
[page 21]
hidden_size num_layers mse rmse mae r2
0 32 1 0.002088 0.045696 0.026292 0.874528
1 32 2 0.002210 0.047010 0.029172 0.867206
2 32 3 0.002576 0.050750 0.029360 0.845239
3 64 1 0.001910 0.043705 0.028409 0.885221
4 64 2 0.001379 0.037135 0.020275 0.917138
5 64 3 0.001389 0.037267 0.016396 0.916547
6 128 1 0.002892 0.053781 0.036514 0.826200
7 128 2 0.002099 0.045814 0.025405 0.873879
8 128 3 0.002063 0.045422 0.028918 0.876025
9 256 1 0.001810 0.042547 0.026734 0.891223
10 256 2 0.002353 0.048509 0.032749 0.858603
11 256 3 0.017387 0.131859 0.115412 -0.044760
From both the graph and the tabular results, a hidden size of 64 with 2 hidden layers gives the best performance. While a single layer
performs well, stacking a second layer with a larger hidden size appears to capture more complex patterns, leading to lower error and a
higher R² score. This combination provides the best balance between model capacity and generalization for this dataset.
Final Model Training and Evaluation
Based on the hyperparameter tuning, we now train the final LSTM model using the optimal configuration:
Sequence Length: 7
Hidden Size: 64
Number of Layers: 2
We train for 100 epochs to ensure the model converges properly and then evaluate its performance on both the training and test sets.
Out[ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 21/53
[page 22]
import pandas as pd
import numpy as np
import torch
import torch.nn as nn
from sklearn.preprocessing import MinMaxScaler
import matplotlib.pyplot as plt
torch.manual_seed(SEED)
# data prep
data = df_v1["close"].values.reshape(-1, 1)
# Normalize 0-1 scaling
scaler = MinMaxScaler()
data_scaled = scaler.fit_transform(data)
def create_sequences(data, seq_len=6):
X, y = [], []
for i in range(len(data) - seq_len):
X.append(data[i:i+seq_len])
y.append(data[i+seq_len])
return np.array(X), np.array(y)
# Use optimal sequence length (ADJUST IF YOURS IS DIFFERENT)
SEQ_LEN = 7
X, y = create_sequences(data_scaled, SEQ_LEN)
# Train-test split (chronological)
train_size = int(len(X) * 0.8)
X_train, X_test = X[:train_size], X[train_size:]
y_train, y_test = y[:train_size], y[train_size:]
# Convert to torch tensors
X_train = torch.tensor(X_train, dtype=torch.float32)
y_train = torch.tensor(y_train, dtype=torch.float32)
X_test = torch.tensor(X_test, dtype=torch.float32)
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 22/53
[page 23]
y_test = torch.tensor(y_test, dtype=torch.float32)
# LSTM Model
class LSTMModel(nn.Module):
def __init__(self, input_size=1, hidden_size=32, num_layers=2):
super(LSTMModel, self).__init__()
self.lstm = nn.LSTM(input_size, hidden_size, num_layers, batch_first=True)
self.fc = nn.Linear(hidden_size, 1)
def forward(self, x):
# The LSTM returns the output sequence and the final hidden and cell states.
# We only need the final hidden state of the last layer to make a prediction.
_, (hn, cn) = self.lstm(x)
# hn is of shape (num_layers, batch_size, hidden_size)
out = self.fc(hn[-1])
return out
# Instantiate with optimal hyperparameters (ADJUST IF YOURS ARE DIFFERENT)
model = LSTMModel(hidden_size=128, num_layers=2)
# Train the model
criterion = nn.MSELoss()
optimizer = torch.optim.Adam(model.parameters(), lr=0.01)
EPOCHS = 100
for epoch in range(EPOCHS):
model.train()
optimizer.zero_grad()
output = model(X_train)
loss = criterion(output, y_train)
loss.backward()
optimizer.step()
if (epoch+1) % 10 == 0:
print(f"Epoch {epoch+1}/{EPOCHS}, Loss: {loss.item():.6f}")
# Evaluate the model
model.eval()
with torch.no_grad():
train_pred = model(X_train).numpy()
test_pred = model(X_test).numpy()
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 23/53
[page 24]
from sklearn.metrics import mean_squared_error, r2_score
import numpy as np
def evaluate_and_print(y_train_true, y_train_pred, y_test_true, y_test_pred):
def _metrics(y_true, y_pred):
mse = mean_squared_error(y_true, y_pred)
rmse = np.sqrt(mse)
r2 = r2_score(y_true, y_pred)
return mse, rmse, r2
train_mse, train_rmse, train_r2 = _metrics(y_train_true, y_train_pred)
print(f"Train MSE: {train_mse:.4f}, RMSE: {train_rmse:.4f}, R2 score: {train_r2:.4f}")
test_mse, test_rmse, test_r2 = _metrics(y_test_true, y_test_pred)
print(f"Test MSE: {test_mse:.4f}, RMSE: {test_rmse:.4f}, R2 score: {test_r2:.4f}")
# Inverse transform to get actual prices
train_pred = scaler.inverse_transform(train_pred)
y_train_real = scaler.inverse_transform(y_train.numpy())
test_pred = scaler.inverse_transform(test_pred)
y_test_real = scaler.inverse_transform(y_test.numpy())
evaluate_and_print(y_train_real, train_pred, y_test_real, test_pred)
Epoch 10/100, Loss: 0.028179
Epoch 20/100, Loss: 0.023059
Epoch 30/100, Loss: 0.013124
Epoch 40/100, Loss: 0.008520
Epoch 50/100, Loss: 0.006434
Epoch 60/100, Loss: 0.005307
Epoch 70/100, Loss: 0.004979
Epoch 80/100, Loss: 0.004789
Epoch 90/100, Loss: 0.004709
Epoch 100/100, Loss: 0.004673
Train MSE: 119.7040, RMSE: 10.9409, R2 score: 0.9025
Test MSE: 30.1466, RMSE: 5.4906, R2 score: 0.9293
Actual vs Predicted Closing Prices
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 24/53
[page 25]
plt.figure(figsize=(12,6))
plt.plot(y_test_real, label="Actual Close Price", color="blue")
plt.plot(test_pred, label="Predicted Close Price", color="red", linestyle="--")
plt.title("LSTM: Actual vs. Predicted Closing Prices on Test Set")
plt.legend()
plt.show()
Observations:
1. The predicted curve follows the actual trend closely, capturing both upward and downward movements with high fidelity.
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 25/53
[page 26]
2. The LSTM model successfully tracks the overall patterns, and even near sharp jumps or drops, the deviations are minimal,
demonstrating its strong predictive capability.
Comparison with MLP
To benchmark the LSTM's performance, we also trained a simple feedforward neural network (MLP) on the same time-series sequences.
The sequences were flattened to be used as input features for the MLP. To highlight the LSTM's ability to handle temporal data, we
deliberately use a smaller MLP with a hidden size of 16.
import torch
import torch.nn as nn
import numpy as np
import pandas as pd
from sklearn.preprocessing import MinMaxScaler
from sklearn.metrics import mean_squared_error, r2_score, mean_absolute_error
torch.manual_seed(SEED)
scaler = MinMaxScaler()
data_scaled = scaler.fit_transform(df_v1["close"].values.reshape(-1, 1))
def create_sequences(data, seq_len=6):
X, y = [], []
for i in range(len(data) - seq_len):
X.append(data[i:i+seq_len])
y.append(data[i+seq_len])
return np.array(X), np.array(y)
SEQ_LEN = 7
X, y = create_sequences(data_scaled, SEQ_LEN)
train_size = int(len(X) * 0.8)
X_train, X_test = X[:train_size], X[train_size:]
y_train, y_test = y[:train_size], y[train_size:]
# MLP
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 26/53
[page 27]
class FFN(nn.Module):
def __init__(self, input_size, hidden_size=16):
super(FFN, self).__init__()
self.fc1 = nn.Linear(input_size, hidden_size)
self.relu = nn.ReLU()
self.fc2 = nn.Linear(hidden_size, 1)
def forward(self, x):
return self.fc2(self.relu(self.fc1(x)))
def train_ffn(X_train, y_train, X_test, y_test, epochs=100):
input_size = X_train.shape[1]
model = FFN(input_size=input_size, hidden_size=16)
criterion = nn.MSELoss()
optimizer = torch.optim.Adam(model.parameters(), lr=0.01)
X_train_t = torch.tensor(X_train, dtype=torch.float32)
y_train_t = torch.tensor(y_train, dtype=torch.float32)
X_test_t = torch.tensor(X_test, dtype=torch.float32)
for epoch in range(epochs):
model.train()
optimizer.zero_grad()
output = model(X_train_t)
loss = criterion(output, y_train_t)
loss.backward()
optimizer.step()
model.eval()
with torch.no_grad():
train_pred = model(X_train_t).numpy()
test_pred = model(X_test_t).numpy()
return train_pred, test_pred
X_train_flat = X_train.reshape(X_train.shape[0], -1)
X_test_flat = X_test.reshape(X_test.shape[0], -1)
train_pred_ffn, test_pred_ffn = train_ffn(X_train_flat, y_train, X_test_flat, y_test)
# inverse transform for MLP
train_pred_ffn = scaler.inverse_transform(train_pred_ffn)
test_pred_ffn = scaler.inverse_transform(test_pred_ffn)
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 27/53
[page 28]
y_train_real_ffn = scaler.inverse_transform(y_train)
y_test_real_ffn = scaler.inverse_transform(y_test)
# RNN
class RNNModel(nn.Module):
def __init__(self, input_size=1, hidden_size=32, num_layers=1):
super(RNNModel, self).__init__()
self.rnn = nn.RNN(input_size, hidden_size, num_layers, batch_first=True)
self.fc = nn.Linear(hidden_size, 1)
def forward(self, x):
out, h = self.rnn(x)
out = self.fc(h[-1])
return out
def train_rnn(X_train, y_train, X_test, y_test, epochs=100):
model = RNNModel(input_size=1, hidden_size=32, num_layers=1)
criterion = nn.MSELoss()
optimizer = torch.optim.Adam(model.parameters(), lr=0.01)
X_train_t = torch.tensor(X_train, dtype=torch.float32)
y_train_t = torch.tensor(y_train, dtype=torch.float32)
X_test_t = torch.tensor(X_test, dtype=torch.float32)
for epoch in range(epochs):
model.train()
optimizer.zero_grad()
output = model(X_train_t)
loss = criterion(output, y_train_t)
loss.backward()
optimizer.step()
model.eval()
with torch.no_grad():
train_pred = model(X_train_t).numpy()
test_pred = model(X_test_t).numpy()
return train_pred, test_pred
train_pred_rnn, test_pred_rnn = train_rnn(X_train, y_train, X_test, y_test)
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 28/53
[page 29]
# inverse transform for RNN
train_pred_rnn = scaler.inverse_transform(train_pred_rnn)
test_pred_rnn = scaler.inverse_transform(test_pred_rnn)
y_train_real_rnn = scaler.inverse_transform(y_train)
y_test_real_rnn = scaler.inverse_transform(y_test)
# Evaluation
def evaluate(y_train, train_pred, y_test, test_pred, name):
def _metrics(y_true, y_pred):
mse = mean_squared_error(y_true, y_pred)
rmse = np.sqrt(mse)
mae = mean_absolute_error(y_true, y_pred)
r2 = r2_score(y_true, y_pred)
return mse, rmse, mae, r2
train_mse, train_rmse, train_mae, train_r2 = _metrics(y_train, train_pred)
test_mse, test_rmse, test_mae, test_r2 = _metrics(y_test, test_pred)
return {
"Model": name,
"Train MSE": train_mse, "Train RMSE": train_rmse, "Train MAE": train_mae, "Train R2": train_r2,
"Test MSE": test_mse, "Test RMSE": test_rmse, "Test MAE": test_mae, "Test R2": test_r2
}
results_list = []
results_list.append(evaluate(y_train_real_ffn, train_pred_ffn, y_test_real_ffn, test_pred_ffn, "MLP"))
results_list.append(evaluate(y_train_real_rnn, train_pred_rnn, y_test_real_rnn, test_pred_rnn, "RNN"))
results_list.append(evaluate(y_train_real, train_pred, y_test_real, test_pred, "LSTM"))
df_results = pd.DataFrame(results_list)
df_results[['Model','Test MSE','Test RMSE','Test MAE','Test R2']]
Model Test MSE Test RMSE Test MAE Test R2
0 MLP 33.464074 5.784814 2.628603 0.921511
1 RNN 31.067253 5.573801 2.399222 0.927133
2 LSTM 30.146622 5.490594 2.247842 0.929292
Out[ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 29/53
[page 30]
The results show that the performance of LSTM and RNN is comparable and they outperforms the MLP. This highlights the LSTM's/RNN
strength in modeling sequential data, as it can capture temporal dependencies that a simple MLP with flattened features misses.
Multi-Step Forecasting: Predicting the Next 4 Weeks
Here, the LSTM is adapted to perform multi-step forecasting. Instead of predicting just the next week’s price, the model will output
predictions for the next 4 weeks simultaneously. This is a sequence-to-sequence (Seq2Seq) task where the input is a sequence of past
prices and the output is a sequence of future prices.
The model uses an encoder-decoder architecture. The encoder (an LSTM) processes the input sequence and summarizes it into a context
vector (the final hidden and cell states). The decoder (another LSTM) then uses this context to generate the output sequence one step at a
time, feeding its own previous prediction back as input for the next step (autoregression).
import pandas as pd
import numpy as np
import torch
import torch.nn as nn
from sklearn.preprocessing import MinMaxScaler
from sklearn.metrics import mean_squared_error, r2_score, mean_absolute_error
# Prepare data
data = df_v1["close"].values.reshape(-1, 1)
scaler = MinMaxScaler()
data_scaled = scaler.fit_transform(data)
# Create sequences for multi-step prediction
def create_sequences_multi_step(data, seq_len=6, pred_horizon=4):
X, y = [], []
for i in range(len(data) - seq_len - pred_horizon + 1):
X.append(data[i:i+seq_len])
y.append(data[i+seq_len : i+seq_len+pred_horizon].flatten())
return np.array(X), np.array(y)
SEQ_LEN = 7
PRED_HORIZON = 4
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 30/53
[page 31]
X, y = create_sequences_multi_step(data_scaled, SEQ_LEN, PRED_HORIZON)
train_size = int(len(X) * 0.8)
X_train, X_test = X[:train_size], X[train_size:]
y_train, y_test = y[:train_size], y[train_size:]
X_train = torch.tensor(X_train, dtype=torch.float32)
y_train = torch.tensor(y_train, dtype=torch.float32)
X_test = torch.tensor(X_test, dtype=torch.float32)
y_test = torch.tensor(y_test, dtype=torch.float32)
# Seq2Seq LSTM model
class Seq2SeqLSTM(nn.Module):
def __init__(self, input_size=1, hidden_size=32, num_layers=2, pred_horizon=4):
super(Seq2SeqLSTM, self).__init__()
self.pred_horizon = pred_horizon
self.lstm = nn.LSTM(input_size, hidden_size, num_layers, batch_first=True)
self.fc = nn.Linear(hidden_size, 1)
def forward(self, x, y=None, train=False):
# Encoder
_, (h, c) = self.lstm(x)
# Start decoding with the last observed value
dec_input = x[:, -1:, :] # (batch, 1, input_size)
preds = []
for t in range(self.pred_horizon):
out, (h, c) = self.lstm(dec_input, (h, c))
y_pred = self.fc(out[:, -1, :]) # (batch, 1)
preds.append(y_pred)
if train and y is not None:
# ground truth as next input during training
dec_input = y[:, t].unsqueeze(1).unsqueeze(2) # (batch, 1, 1)
else:
# own prediction during inference
dec_input = y_pred.unsqueeze(1)
return torch.cat(preds, dim=1) # (batch, pred_horizon)
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 31/53
[page 32]
model = Seq2SeqLSTM(hidden_size=64, num_layers=2, pred_horizon=PRED_HORIZON)
# Training
criterion = nn.MSELoss()
optimizer = torch.optim.Adam(model.parameters(), lr=0.01)
EPOCHS = 100
for epoch in range(EPOCHS):
model.train()
optimizer.zero_grad()
output = model(X_train, y_train, train=True)
loss = criterion(output, y_train)
loss.backward()
optimizer.step()
if (epoch+1) % 10 == 0:
print(f"Epoch {epoch+1}/{EPOCHS}, Loss: {loss.item():.6f}")
# Evaluation
model.eval()
with torch.no_grad():
test_pred = model(X_test, train=False).numpy()
# Inverse transform
test_pred = scaler.inverse_transform(test_pred)
y_test_real = scaler.inverse_transform(y_test.numpy())
# Metrics
def evaluate_multi_step(y_true, y_pred):
mse = mean_squared_error(y_true, y_pred)
rmse = np.sqrt(mse)
mae = mean_absolute_error(y_true, y_pred)
r2 = r2_score(y_true, y_pred)
print(f"Test MSE: {mse:.4f}, RMSE: {rmse:.4f}, MAE: {mae:.4f}, R2: {r2:.4f}")
evaluate_multi_step(y_test_real, test_pred)
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 32/53
[page 33]
Epoch 10/100, Loss: 0.051932
Epoch 20/100, Loss: 0.028025
Epoch 30/100, Loss: 0.015227
Epoch 40/100, Loss: 0.010574
Epoch 50/100, Loss: 0.006908
Epoch 60/100, Loss: 0.005958
Epoch 70/100, Loss: 0.005513
Epoch 80/100, Loss: 0.005180
Epoch 90/100, Loss: 0.004904
Epoch 100/100, Loss: 0.004725
Test MSE: 86.5752, RMSE: 9.3046, MAE: 4.9510, R2: 0.7967
import matplotlib.pyplot as plt
fig, axs = plt.subplots(2, 2, figsize=(14,8))
# Week +1
axs[0,0].plot(y_test_real[:,0], label="Actual Week+1")
axs[0,0].plot(test_pred[:,0], label="Predicted Week+1", linestyle="--")
axs[0,0].legend()
axs[0,0].set_title("Week +1")
# Week +2
axs[0,1].plot(y_test_real[:,1], label="Actual Week+2")
axs[0,1].plot(test_pred[:,1], label="Predicted Week+2", linestyle="--")
axs[0,1].legend()
axs[0,1].set_title("Week +2")
# Week +3
axs[1,0].plot(y_test_real[:,2], label="Actual Week+3")
axs[1,0].plot(test_pred[:,2], label="Predicted Week+3", linestyle="--")
axs[1,0].legend()
axs[1,0].set_title("Week +3")
# Week +4
axs[1,1].plot(y_test_real[:,3], label="Actual Week+4")
axs[1,1].plot(test_pred[:,3], label="Predicted Week+4", linestyle="--")
axs[1,1].legend()
axs[1,1].set_title("Week +4")
plt.suptitle("Multi-step LSTM Prediction (Next 4 Weeks Closing Price)", fontsize=14)
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 33/53
[page 34]
plt.tight_layout(rect=[0, 0.03, 1, 0.95])
plt.show()
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 34/53
[page 35]
The performance clearly dropped in the multi-step setting compared to single-step prediction. This is expected, as predicting further into
the future is inherently more difficult. Errors accumulate in the autoregressive decoding process: a small error in the Week +1 prediction is
fed as input to predict Week +2, potentially causing a larger error, and so on.
Despite this, the model captures the overall trends well for all four weeks. As the prediction horizon increases from Week +1 to Week +4,
the deviation from the actual values becomes more pronounced, but the general direction is often correct.
LSTM with Multiple Features for Prediction
In this final experiment, we enhance the LSTM model by using multiple input features instead of just the closing price. We use a
multivariate time series including open , high , low , close , and volume (OHLCV) from the past 6 weeks to predict the next week’s
closing price. This allows the model to learn from richer, more contextual information about the stock's weekly behavior.
import pandas as pd
import numpy as np
import torch
import torch.nn as nn
from sklearn.preprocessing import MinMaxScaler
from sklearn.metrics import mean_squared_error, r2_score, mean_absolute_error
torch.manual_seed(SEED)
feature_cols = ["open", "high", "low", "close", "volume"]
data_multi = df_v1[feature_cols].values
target = df_v1["close"].values.reshape(-1, 1)
# Normalize features and target separately
scaler_X = MinMaxScaler()
scaler_y = MinMaxScaler()
data_scaled = scaler_X.fit_transform(data_multi)
target_scaled = scaler_y.fit_transform(target)
SEQ_LEN = 7 # Use the optimal sequence length from tuning
def create_sequences_multi_feature(X, y, seq_len=6):
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 35/53
[page 36]
Xs, ys = [], []
for i in range(len(X) - seq_len):
Xs.append(X[i:i+seq_len])
ys.append(y[i+seq_len])
return np.array(Xs), np.array(ys)
X, y = create_sequences_multi_feature(data_scaled, target_scaled, SEQ_LEN)
print("X shape:", X.shape)
print("y shape:", y.shape)
# Train-test split
train_size = int(len(X) * 0.8)
X_train, X_test = X[:train_size], X[train_size:]
y_train, y_test = y[:train_size], y[train_size:]
X_train = torch.tensor(X_train, dtype=torch.float32)
y_train = torch.tensor(y_train, dtype=torch.float32)
X_test = torch.tensor(X_test, dtype=torch.float32)
y_test = torch.tensor(y_test, dtype=torch.float32)
# LSTM model (multi-feature)
class LSTMModelMulti(nn.Module):
def __init__(self, input_size, hidden_size=32, num_layers=1):
super(LSTMModelMulti, self).__init__()
self.lstm = nn.LSTM(input_size, hidden_size, num_layers, batch_first=True)
self.fc = nn.Linear(hidden_size, 1)
def forward(self, x):
_, (hn, cn) = self.lstm(x)
# Use the final hidden state from the last layer for prediction
out = self.fc(hn[-1])
return out
num_features = X.shape[2]
# Use optimal hyperparameters
model = LSTMModelMulti(input_size=num_features, hidden_size=128, num_layers=2)
# Training
criterion = nn.MSELoss()
optimizer = torch.optim.Adam(model.parameters(), lr=0.01)
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 36/53
[page 37]
EPOCHS = 100
for epoch in range(EPOCHS):
model.train()
optimizer.zero_grad()
output = model(X_train)
loss = criterion(output, y_train)
loss.backward()
optimizer.step()
if (epoch+1) % 10 == 0:
print(f"Epoch {epoch+1}/{EPOCHS}, Loss: {loss.item():.6f}")
# Evaluation
model.eval()
with torch.no_grad():
train_pred = model(X_train).numpy()
test_pred = model(X_test).numpy()
# Inverse transform
train_pred = scaler_y.inverse_transform(train_pred)
y_train_real = scaler_y.inverse_transform(y_train.numpy())
test_pred = scaler_y.inverse_transform(test_pred)
y_test_real = scaler_y.inverse_transform(y_test.numpy())
def evaluate_and_print(y_train_true, y_train_pred, y_test_true, y_test_pred):
def _metrics(y_true, y_pred):
mse = mean_squared_error(y_true, y_pred)
rmse = np.sqrt(mse)
r2 = r2_score(y_true, y_pred)
mae=mean_absolute_error(y_true,y_pred)
return mse, rmse, r2,mae
train_mse, train_rmse, train_r2,train_mae = _metrics(y_train_true, y_train_pred)
print(f"Train MSE: {train_mse:.4f}, RMSE: {train_rmse:.4f}, MAE: {train_mae:.4f}, R2 score: {train_r2:.4f}")
test_mse, test_rmse, test_r2,test_mae = _metrics(y_test_true, y_test_pred)
print(f"Test MSE: {test_mse:.4f}, RMSE: {test_rmse:.4f}, MAE: {test_mae:.4f}, R2 score: {test_r2:.4f}")
evaluate_and_print(y_train_real, train_pred, y_test_real, test_pred)
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 37/53
[page 38]
X shape: (743, 7, 5)
y shape: (743, 1)
Epoch 10/100, Loss: 0.041550
Epoch 20/100, Loss: 0.010583
Epoch 30/100, Loss: 0.006510
Epoch 40/100, Loss: 0.005654
Epoch 50/100, Loss: 0.005153
Epoch 60/100, Loss: 0.004963
Epoch 70/100, Loss: 0.004830
Epoch 80/100, Loss: 0.004743
Epoch 90/100, Loss: 0.004685
Epoch 100/100, Loss: 0.004656
Train MSE: 119.2154, RMSE: 10.9186, MAE: 3.7665, R2 score: 0.9029
Test MSE: 30.9975, RMSE: 5.5675, MAE: 2.3487, R2 score: 0.9273
Using multiple input features led to a slight dip in performance, suggesting that additional variables may be redundant and the closing
price alone already captures most of the predictive signal.
Implementing LSTM from Scratch
An LSTM cell controls information flow using four gates:
Forget Gate ( ) → controls how much of the previous cell state ( ) to keep.
Input Gate ( ) → decides how much new candidate information should enter the cell.
Candidate State ( ) → new content that can be added into the cell state.
Output Gate ( ) → decides how much of the cell state influences the hidden state ( ).
Updates
import torch
import torch.nn as nn
ft ct−1
it
gt
ot ht
ct = ft ⊙ ct−1 + it ⊙ gt
ht = ot ⊙ tanh(ct)
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 38/53
[page 39]
import math
class SimpleLSTM(nn.Module):
def __init__(self, input_size, hidden_size):
super(SimpleLSTM, self).__init__()
self.input_size = input_size
self.hidden_size = hidden_size
# One weight matrix for input -> gates
self.W_x = nn.Parameter(torch.Tensor(input_size, 4 * hidden_size))
# One weight matrix for hidden -> gates
self.W_h = nn.Parameter(torch.Tensor(hidden_size, 4 * hidden_size))
# One bias vector
self.b = nn.Parameter(torch.Tensor(4 * hidden_size))
self.reset_parameters()
def reset_parameters(self):
# Initialize weights uniformly (like PyTorch does)
stdv = 1.0 / math.sqrt(self.hidden_size)
for param in self.parameters():
nn.init.uniform_(param, -stdv, stdv)
def forward(self, x, state=None):
"""
x: (batch, seq_len, input_size)
state: (h0, c0), each (batch, hidden_size)
"""
batch, seq_len, _ = x.size()
if state is None:
h_t = torch.zeros(batch, self.hidden_size, device=x.device)
c_t = torch.zeros(batch, self.hidden_size, device=x.device)
else:
h_t, c_t = state
outputs = []
for t in range(seq_len):
x_t = x[:, t, :] # (batch, input_size)
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 39/53
[page 40]
gates = x_t @ self.W_x + h_t @ self.W_h + self.b
i, f, g, o = gates.chunk(4, dim=1) # split into 4 parts
i = torch.sigmoid(i) # input gate
f = torch.sigmoid(f) # forget gate
g = torch.tanh(g) # candidate cell
o = torch.sigmoid(o) # output gate
c_t = f * c_t + i * g # new cell state
h_t = o * torch.tanh(c_t) # new hidden state
outputs.append(h_t.unsqueeze(1))
outputs = torch.cat(outputs, dim=1) # (batch, seq_len, hidden_size)
return outputs, (h_t, c_t)
# Example
lstm = SimpleLSTM(input_size=10, hidden_size=20)
x = torch.randn(3, 5, 10) # (batch=3, seq_len=5, input_size=10)
h0 = torch.randn(3, 20)
c0 = torch.randn(3, 20)
out, (hn, cn) = lstm(x, (h0, c0))
print(out.shape) # (3, 5, 20)
print(hn.shape) # (3, 20)
print(cn.shape) # (3, 20)
torch.Size([3, 5, 20])
torch.Size([3, 20])
torch.Size([3, 20])
Stacked LSTM
Manual stacking: just pass the output of one LSTM as the input to the next.
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 40/53
[page 41]
inp = torch.randn(3, 5, 10)
lstm1 = SimpleLSTM(10, 20)
lstm2 = SimpleLSTM(20, 20)
out1, (h1, c1) = lstm1(inp)
out2, (h2, c2) = lstm2(out1)
print(out1.shape)
print(out2.shape)
torch.Size([3, 5, 20])
torch.Size([3, 5, 20])
Define a StackedLSTM class. It creates multiple SimpleLSTM layers in a loop and stores them in a ModuleList . The forward
pass also loops through these layers. This makes it easy to stack many layers by just setting num_layers .
class StackedLSTM(nn.Module):
def __init__(self, input_size, hidden_size, num_layers):
super(StackedLSTM, self).__init__()
layers = []
for i in range(num_layers):
in_size = input_size if i == 0 else hidden_size
layers.append(SimpleLSTM(in_size, hidden_size))
self.layers = nn.ModuleList(layers)
def forward(self, x):
"""
x: (batch, seq_len, input_size)
"""
out = x
for layer in self.layers:
out, _ = layer(out) # each SimpleLSTM keeps batch-first
return out
# Example
inp = torch.randn(3, 5, 10) # (batch=3, seq_len=5, input_size=10)
stacked = StackedLSTM(10, 20, num_layers=2)
In [ ]:
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 41/53
[page 42]
out = stacked(inp)
print(out.shape) # (3, 5, 20)
torch.Size([3, 5, 20])
LSTM Forecaster with manually implemented LSTM (One-Step
Prediction)
LSTM with a linear head to do prediction
import torch
import torch.nn as nn
import math
class LSTMForecaster(nn.Module):
def __init__(self, input_size, hidden_size, output_size):
super(LSTMForecaster, self).__init__()
self.lstm = SimpleLSTM(input_size, hidden_size)
self.fc = nn.Linear(hidden_size, output_size) # prediction head
def forward(self, x, state=None):
"""
x: (batch, seq_len, input_size)
"""
outputs, (hn, cn) = self.lstm(x, state)
# Use last hidden state for one-step prediction
pred = self.fc(hn) # hn shape: (batch, hidden_size)
return pred
# ---------------- Example ----------------
torch.manual_seed(0)
model = LSTMForecaster(input_size=10, hidden_size=20, output_size=1)
# Dummy input: batch=3, seq_len=5, input_size=10
x = torch.randn(3, 5, 10)
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 42/53
[page 43]
# Dummy target: one-step prediction -> shape (batch, output_size)
target = torch.randn(3, 1)
# Forward pass
pred = model(x)
print("Prediction shape:", pred.shape) # (3, 1)
# Loss
criterion = nn.MSELoss()
loss = criterion(pred, target)
print("Loss:", loss.item())
# Backward pass
loss.backward()
# Print gradients
print("\nGradients:")
print("W_x.grad norm:", model.lstm.W_x.grad.norm().item())
print("W_h.grad norm:", model.lstm.W_h.grad.norm().item())
print("b.grad norm:", model.lstm.b.grad.norm().item())
print("fc.weight.grad norm:", model.fc.weight.grad.norm().item())
print("fc.bias.grad norm:", model.fc.bias.grad.norm().item())
Prediction shape: torch.Size([3, 1])
Loss: 0.19849489629268646
Gradients:
W_x.grad norm: 0.20284342765808105
W_h.grad norm: 0.04552901163697243
b.grad norm: 0.1242167130112648
fc.weight.grad norm: 0.2483113706111908
fc.bias.grad norm: 0.45420366525650024
Using manually implemented LSTM to do one step closing price
prediction
import pandas as pd
import numpy as np
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 43/53
[page 44]
import torch
import torch.nn as nn
from sklearn.preprocessing import MinMaxScaler
import matplotlib.pyplot as plt
data = df_v1["close"].values.reshape(-1, 1)
torch.manual_seed(SEED)
# Normalize 0-1 scaling
scaler = MinMaxScaler()
data_scaled = scaler.fit_transform(data)
def create_sequences(data, seq_len=6):
X, y = [], []
for i in range(len(data) - seq_len):
X.append(data[i:i+seq_len])
y.append(data[i+seq_len])
return np.array(X), np.array(y)
# Use optimal sequence length (ADJUST IF YOURS IS DIFFERENT)
SEQ_LEN = 7
X, y = create_sequences(data_scaled, SEQ_LEN)
print(X.shape,y.shape)
# Train-test split (chronological)
train_size = int(len(X) * 0.8)
X_train, X_test = X[:train_size], X[train_size:]
y_train, y_test = y[:train_size], y[train_size:]
# Convert to torch tensors
X_train = torch.tensor(X_train, dtype=torch.float32)
y_train = torch.tensor(y_train, dtype=torch.float32)
X_test = torch.tensor(X_test, dtype=torch.float32)
y_test = torch.tensor(y_test, dtype=torch.float32)
# Instantiate with optimal hyperparameters (ADJUST IF YOURS ARE DIFFERENT)
model = LSTMForecaster(input_size=1, hidden_size=128, output_size=1)
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 44/53
[page 45]
# Train the model
criterion = nn.MSELoss()
optimizer = torch.optim.Adam(model.parameters(), lr=0.01)
EPOCHS = 100
for epoch in range(EPOCHS):
model.train()
optimizer.zero_grad()
output = model(X_train)
#print(X_train.shape)
#print(output.shape,y_train.shape)
loss = criterion(output, y_train)
loss.backward()
optimizer.step()
if (epoch+1) % 10 == 0:
print(f"Epoch {epoch+1}/{EPOCHS}, Loss: {loss.item():.6f}")
# Evaluate the model
model.eval()
with torch.no_grad():
train_pred = model(X_train).numpy()
test_pred = model(X_test).numpy()
from sklearn.metrics import mean_squared_error, r2_score
import numpy as np
def evaluate_and_print(y_train_true, y_train_pred, y_test_true, y_test_pred):
def _metrics(y_true, y_pred):
mse = mean_squared_error(y_true, y_pred)
rmse = np.sqrt(mse)
r2 = r2_score(y_true, y_pred)
return mse, rmse, r2
train_mse, train_rmse, train_r2 = _metrics(y_train_true, y_train_pred)
print(f"Train MSE: {train_mse:.4f}, RMSE: {train_rmse:.4f}, R2 score: {train_r2:.4f}")
test_mse, test_rmse, test_r2 = _metrics(y_test_true, y_test_pred)
print(f"Test MSE: {test_mse:.4f}, RMSE: {test_rmse:.4f}, R2 score: {test_r2:.4f}")
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 45/53
[page 46]
evaluate_and_print(y_train, train_pred, y_test, test_pred)
(742, 8, 1) (742, 1)
Epoch 10/100, Loss: 0.025655
Epoch 20/100, Loss: 0.014062
Epoch 30/100, Loss: 0.010093
Epoch 40/100, Loss: 0.007516
Epoch 50/100, Loss: 0.006297
Epoch 60/100, Loss: 0.005653
Epoch 70/100, Loss: 0.005372
Epoch 80/100, Loss: 0.005267
Epoch 90/100, Loss: 0.005147
Epoch 100/100, Loss: 0.005047
Train MSE: 0.0050, RMSE: 0.0710, R2 score: 0.8949
Test MSE: 0.0013, RMSE: 0.0365, R2 score: 0.9201
LSTM for Classification: EEG Eye State
Here we use the EEG Eye State dataset to perform binary classification of eye state (open vs. closed) from EEG signals. We will compare the
performance of an LSTM against a standard MLP to see if the temporal nature of the EEG signal provides an advantage.
!pip install -q liac-arff
Preparing metadata (setup.py) ... done
Building wheel for liac-arff (setup.py) ... done
# Download the EEG Eye State dataset from the UCI repository
!wget https://archive.ics.uci.edu/ml/machine-learning-databases/00264/EEG%20Eye%20State.arff
In [ ]:
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 46/53
[page 47]
--2025-09-11 18:54:27-- https://archive.ics.uci.edu/ml/machine-learning-databases/00264/EEG%20Eye%20State.arff
Resolving archive.ics.uci.edu (archive.ics.uci.edu)... 128.195.10.252
Connecting to archive.ics.uci.edu (archive.ics.uci.edu)|128.195.10.252|:443... connected.
HTTP request sent, awaiting response... 200 OK
Length: unspecified
Saving to: ‘EEG Eye State.arff’
EEG Eye State.arff [ <=> ] 1.62M 1.79MB/s in 0.9s
2025-09-11 18:54:29 (1.79 MB/s) - ‘EEG Eye State.arff’ saved [1696428]
from scipy.io import arff
import pandas as pd
# Load ARFF
data, meta = arff.loadarff("EEG Eye State.arff")
df_eeg = pd.DataFrame(data)
# Decode byte strings and convert target to integer
df_eeg['eyeDetection'] = df_eeg['eyeDetection'].str.decode('utf-8').astype(int)
df_eeg.head()
AF3 F7 F3 FC5 T7 P7 O1 O2 P8 T8 FC6 F4 F8 AF4 eyeDe
0 4329.23 4009.23 4289.23 4148.21 4350.26 4586.15 4096.92 4641.03 4222.05 4238.46 4211.28 4280.51 4635.90 4393.85
1 4324.62 4004.62 4293.85 4148.72 4342.05 4586.67 4097.44 4638.97 4210.77 4226.67 4207.69 4279.49 4632.82 4384.10
2 4327.69 4006.67 4295.38 4156.41 4336.92 4583.59 4096.92 4630.26 4207.69 4222.05 4206.67 4282.05 4628.72 4389.23
3 4328.72 4011.79 4296.41 4155.90 4343.59 4582.56 4097.44 4630.77 4217.44 4235.38 4210.77 4287.69 4632.31 4396.41
4 4326.15 4011.79 4292.31 4151.28 4347.69 4586.67 4095.90 4627.69 4210.77 4244.10 4212.82 4288.21 4632.82 4398.46
Training a MLP for Comparison
In [ ]:
Out[ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 47/53
[page 48]
import torch
import torch.nn as nn
import torch.optim as optim
from torch.utils.data import DataLoader, TensorDataset
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, roc_auc_score
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
# Prepare data
X = df_eeg.drop('eyeDetection', axis=1).astype(float).values
y = df_eeg['eyeDetection'].astype(int).values
# Split
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
# Scale
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)
# Torch tensors
X_train_t = torch.tensor(X_train_scaled, dtype=torch.float32)
y_train_t = torch.tensor(y_train, dtype=torch.long)
X_test_t = torch.tensor(X_test_scaled, dtype=torch.float32)
y_test_t = torch.tensor(y_test, dtype=torch.long)
train_ds = TensorDataset(X_train_t, y_train_t)
train_loader = DataLoader(train_ds, batch_size=64, shuffle=True)
# MLP Model
class FFN(nn.Module):
def __init__(self, input_size, hidden_size=64):
super(FFN, self).__init__()
self.net = nn.Sequential(
nn.Linear(input_size, hidden_size), nn.ReLU(),
nn.Linear(hidden_size, hidden_size), nn.ReLU(),
nn.Linear(hidden_size, 2)
)
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 48/53
[page 49]
def forward(self, x): return self.net(x)
model = FFN(input_size=X_train_scaled.shape[1])
criterion = nn.CrossEntropyLoss()
optimizer = optim.Adam(model.parameters(), lr=0.001)
# Training loop
for epoch in range(10):
for xb, yb in train_loader:
preds = model(xb)
loss = criterion(preds, yb)
loss.backward()
optimizer.step()
optimizer.zero_grad()
print(f"Epoch {epoch+1}, Loss: {loss.item():.4f}")
# Evaluation
with torch.no_grad():
y_pred = model(X_test_t).argmax(dim=1).numpy()
y_proba = torch.softmax(model(X_test_t), dim=1)[:,1].numpy()
ffn_results = {
"Model": "MLP",
"Accuracy": accuracy_score(y_test, y_pred),
"Precision": precision_score(y_test, y_pred),
"Recall": recall_score(y_test, y_pred),
"F1": f1_score(y_test, y_pred),
"ROC-AUC": roc_auc_score(y_test, y_proba)
}
print(ffn_results)
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 49/53
[page 50]
Epoch 1, Loss: 0.7259
Epoch 2, Loss: 0.6026
Epoch 3, Loss: 0.5276
Epoch 4, Loss: 0.4598
Epoch 5, Loss: 0.4575
Epoch 6, Loss: 0.5671
Epoch 7, Loss: 0.3787
Epoch 8, Loss: 0.4822
Epoch 9, Loss: 0.4263
Epoch 10, Loss: 0.3190
{'Model': 'MLP', 'Accuracy': 0.8287716955941254, 'Precision': 0.8489932885906041, 'Recall': 0.7524163568773234, 'F
1': 0.7977926685061095, 'ROC-AUC': np.float64(0.9070960710980616)}
Using LSTM for Predicting Eye State
Here sequences are artificially created from the continuous EEG stream, so that after every 50 timesteps the model predicts the eye state
at the next timestep. This framing lets the LSTM learn temporal dependencies instead of treating each reading independently.
import numpy as np
import torch
import torch.nn as nn
import torch.optim as optim
from torch.utils.data import DataLoader, TensorDataset
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, roc_auc_score
# Data already split and scaled from previous cell
# Create sequences
SEQ_LEN = 50
def create_sequences_clf(X, y, seq_len=50):
Xs, ys = [], []
for i in range(len(X) - seq_len):
Xs.append(X[i:i+seq_len])
ys.append(y[i+seq_len])
return np.array(Xs), np.array(ys)
X_train_seq, y_train_seq = create_sequences_clf(X_train_scaled, y_train, SEQ_LEN)
X_test_seq, y_test_seq = create_sequences_clf(X_test_scaled, y_test, SEQ_LEN)
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 50/53
[page 51]
print("LSTM Train shape:", X_train_seq.shape, y_train_seq.shape)
print("LSTM Test shape :", X_test_seq.shape, y_test_seq.shape)
# Torch tensors
X_train_t = torch.tensor(X_train_seq, dtype=torch.float32)
y_train_t = torch.tensor(y_train_seq, dtype=torch.float32)
X_test_t = torch.tensor(X_test_seq, dtype=torch.float32)
y_test_t = torch.tensor(y_test_seq, dtype=torch.float32)
train_ds = TensorDataset(X_train_t, y_train_t)
train_loader = DataLoader(train_ds, batch_size=64, shuffle=True)
# LSTM for Binary Classification
class LSTMBinary(nn.Module):
def __init__(self, input_size, hidden_size=64, num_layers=2):
super(LSTMBinary, self).__init__()
self.lstm = nn.LSTM(input_size, hidden_size, num_layers, batch_first=True)
self.fc = nn.Linear(hidden_size, 1)
def forward(self, x):
_, (h, _) = self.lstm(x)
out = self.fc(h[-1]) # Use last layer's hidden state
return out.squeeze(1)
model = LSTMBinary(input_size=X_train_seq.shape[2])
criterion = nn.BCEWithLogitsLoss()
optimizer = optim.Adam(model.parameters(), lr=0.001)
# Training loop
for epoch in range(10):
for xb, yb in train_loader:
preds = model(xb)
loss = criterion(preds, yb)
loss.backward()
optimizer.step()
optimizer.zero_grad()
print(f"Epoch {epoch+1}, Loss: {loss.item():.4f}")
# Evaluation
with torch.no_grad():
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 51/53
[page 52]
y_proba = torch.sigmoid(model(X_test_t)).numpy()
y_pred = (y_proba >= 0.5).astype(int)
lstm_results = {
"Model": "LSTM",
"Accuracy": accuracy_score(y_test_seq, y_pred),
"Precision": precision_score(y_test_seq, y_pred),
"Recall": recall_score(y_test_seq, y_pred),
"F1": f1_score(y_test_seq, y_pred),
"ROC-AUC": roc_auc_score(y_test_seq, y_proba)
}
print(lstm_results)
LSTM Train shape: (11934, 50, 14) (11934,)
LSTM Test shape : (2946, 50, 14) (2946,)
Epoch 1, Loss: 0.6841
Epoch 2, Loss: 0.6902
Epoch 3, Loss: 0.7015
Epoch 4, Loss: 0.6984
Epoch 5, Loss: 0.6793
Epoch 6, Loss: 0.6778
Epoch 7, Loss: 0.7123
Epoch 8, Loss: 0.7222
Epoch 9, Loss: 0.7009
Epoch 10, Loss: 0.6966
{'Model': 'LSTM', 'Accuracy': 0.5515953835709436, 'Precision': 0.48, 'Recall': 0.03644646924829157, 'F1': 0.06774876
499647142, 'ROC-AUC': np.float64(0.5003397978831851)}
import pandas as pd
results_df = pd.DataFrame([ffn_results, lstm_results])
results_df
Model Accuracy Precision Recall F1 ROC-AUC
0 MLP 0.828772 0.848993 0.752416 0.797793 0.907096
1 LSTM 0.551595 0.480000 0.036446 0.067749 0.500340
In [ ]:
Out[ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 52/53
[page 53]
MLP gave better results, while the LSTM performed poorly. This shows that forcing sequential framing does not help for this
dataset and a plain MLP works better.
In [ ]:
13/09/2025, 03:11 LSTM_vf
file:///home/abdys/Downloads/LSTM_vf.html 53/53