# LSTM vf
course: Module 3 — Deep Learning & NLP
module: Module-3-Deep-Learning-NLP
type: notebook
source_url: https://personal-learn.armco.dev/files/Module-3-Deep-Learning-NLP/General/Lab_Materials-22-03-2026/LSTM_vf.ipynb
---
[cell 1 markdown]
# Dow Jones Index Dataset
The **Dow Jones Index** dataset contains weekly stock market data for 30 major companies in the Dow Jones Industrial Average. Each record represents one week of trading activity for a company, including price movements, trading volume, and dividend information.
Although the dataset provides many features, in this study we focus on **predicting the next week’s closing price** using a Long Short-Term Memory (LSTM) network, a special type of Recurrent Neural Network (RNN).
The dataset covers a limited time span: **6 months of 2011 (January–June)** — which results in about 750 weekly records across 30 stocks.
## Source: **UCI Machine Learning Repository**
- [UCI Dataset Page](https://archive.ics.uci.edu/dataset/312/dow+jones+index)
## Target Variable:
- **next_weeks_close** _(numeric)_: Closing price of the stock for the following week.
---
### Features:
| Column Name | Description | Data Type |
| ---------------------------------- | ------------------------------------------------------------- | ----------------- |
| quarter | Fiscal quarter of data (1–4) | _Integer_ |
| stock | Stock ticker symbol (e.g., `AA`, `AXP`) | _Categorical_ |
| date | Week-ending date of the record | _Date_ |
| open | Stock price at market open (that week) | _Numeric (float)_ |
| high | Highest stock price during the week | _Numeric (float)_ |
| low | Lowest stock price during the week | _Numeric (float)_ |
| close | Stock price at market close (that week) | _Numeric (float)_ |
| volume | Total trading volume during the week | _Numeric (float)_ |
| percent_change_price | Percent change from opening to closing price | _Numeric (float)_ |
| percent_change_volume_over_last_wk | Percent change in volume relative to the previous week | _Numeric (float)_ |
| previous_weeks_volume | Trading volume of the previous week | _Numeric (float)_ |
| next_weeks_open | Stock price at market open (next week) | _Numeric (float)_ |
| next_weeks_close | Stock price at market close (next week) — **Target** | **Numeric** |
| percent_change_next_weeks_price | Percent change in price from next week’s open to close | _Numeric (float)_ |
| days_to_next_dividend | Number of days until the next dividend is paid | _Integer_ |
| percent_return_next_dividend | Percent return from the next dividend relative to stock price | _Numeric (float)_ |
[cell 2 markdown]
# **Load into pandas dataframe**
[cell 3 markdown]
# displaying the head
[cell 4 code]
import pandas as pd
# Load file with first row as header
df = pd.read_csv("dow_jones_index.data")
df.head()
[cell 5 code]
df.info()
[cell 6 markdown]
For **predicting next week’s closing price**, we must avoid any feature that already leaks **future information**. In this dataset, the following cannot be used:
- **`next_weeks_open`** → tells us the stock’s opening price in the _future week_.
- **`next_weeks_close`** → is literally the _target variable_, and is redundant since it equals the following row’s `close`, so it cannot be an input feature.
- **`percent_change_next_weeks_price`** → is computed using the next week’s open and close, so it directly encodes future outcome.
[cell 7 markdown]
There are missing values in some columns ( `percent_change_volume_over_last_wk`, `previous_weeks_volume`), but since we are not using these features for our time series modeling, we do not fill the missing values.
[cell 8 markdown]
# **Preprocessing Steps**
- **Clean price columns** : Removed `$` and converted `open`, `high`, `low`, `close`, `next_weeks_open`, `next_weeks_close` to float.
- **Convert date column** : Changed `date` to datetime format for proper time handling.
- **Sort for time series**: Sorted by `stock` and `date` so data is ordered correctly for sequence modeling.
[cell 9 code]
import pandas as pd
def preprocess_dow_jones(path):
# Load raw CSV
df = pd.read_csv(path)
# Remove $ and convert to float for price columns
money_cols = ["open", "high", "low", "close",
"next_weeks_open", "next_weeks_close"]
for col in money_cols:
df[col] = df[col].str.replace("$", "", regex=False).astype(float)
# Convert date column to datetime
df["date"] = pd.to_datetime(df["date"])
# Sort values for time-series use (per stock, by date)
df = df.sort_values(by=["stock", "date"]).reset_index(drop=True)
return df
df_v1 = preprocess_dow_jones("dow_jones_index.data")
df_v1.head()
[cell 10 code]
df_v1.info()
[cell 11 markdown]
# **Records per Stock**
[cell 12 code]
df_v1["stock"].value_counts()
[cell 13 markdown]
# **Distribution of Closing Prices for Each Stock**
[cell 14 code]
import seaborn as sns
import matplotlib.pyplot as plt
plt.figure(figsize=(12,6))
sns.boxplot(x="stock", y="close", data=df_v1)
plt.xticks(rotation=90)
plt.title("Distribution of Closing Prices by Company")
plt.show()
[cell 15 markdown]
1. Most companies have weekly closing prices that stay within a narrow and stable range i.e. their stock prices do not fluctuate much.
2. A few companies like **IBM** and **CAT** not only have higher average closing prices but also show larger ups and downs, indicating greater volatility.
[cell 16 markdown]
# **Setting fix random seeds across libraries to ensure reproducible results in experiments.**
[cell 17 code]
import random
import numpy as np
import torch
SEED = 42
np.random.seed(SEED)
torch.manual_seed(SEED)
random.seed(SEED)
[cell 18 markdown]
# **Assumption**
[cell 19 markdown]
- **Since each company in the Dow Jones UCI dataset has only about 25 weeks of data, we obtain very few company-specific training examples for LSTM models.**
- **For simplicity and to focus on understanding how time series modeling works, we treat the closing prices in this dataset as if they belonged to a single company’s continuous time series, even though in reality they are grouped by different companies.**
[cell 20 markdown]
# **LSTM Setup for Stock Prediction**
[cell 21 markdown]
# Tuning the Optimal Sequence Length
- The tuning is done with sequence lengths of 4,5, 6,7,8 — meaning we use the closing prices of the past 4,5, 6, 7, or 8 weeks to predict the closing price of the next week, effectively making this a **univariate time series** since only one feature (closing price) is used as input.
[cell 22 markdown]
### **Key Functions & Approach**
- **Creating Sequences**: The core of time-series modeling is transforming the flat list of closing prices into overlapping sequences. The `create_sequences` function creates input windows (`X`) of a specified `seq_len` and corresponding targets (`y`). The choice of `seq_len` is a critical hyperparameter that determines how much past information the model uses.
- **Train–Test Split**: The data is split chronologically (80% for training, 20% for testing) to mimic a real-world scenario where we predict the future based on the past. Random shuffling is avoided as it would destroy the temporal order.
- **Data Scaling**: Stock prices are scaled to a [0, 1] range using **MinMaxScaler**. This helps stabilize the training process and is suitable for price data, which is always non-negative.
- **LSTM Model**: We use an LSTM network, which is well-suited for capturing temporal dependencies. Unlike a simple RNN, an LSTM contains memory cells and gates that allow it to learn and remember information over longer sequences, making it more robust against issues like the vanishing gradient problem.
[cell 23 code]
import pandas as pd
import numpy as np
import torch
import torch.nn as nn
import torch.optim as optim
from sklearn.preprocessing import MinMaxScaler
from sklearn.metrics import mean_squared_error, r2_score, mean_absolute_error
import matplotlib.pyplot as plt
# Prepare the data
data = df_v1["close"].values.reshape(-1, 1)
scaler = MinMaxScaler(feature_range=(0, 1))
data_scaled = scaler.fit_transform(data)
torch.manual_seed(SEED)
# Function to create sequences
def create_sequences(data, seq_len=5):
X, y = [], []
for i in range(len(data) - seq_len):
X.append(data[i:i+seq_len])
y.append(data[i+seq_len])
return np.array(X), np.array(y)
# LSTM Model Definition
class LSTMModel(nn.Module):
def __init__(self, input_size=1, hidden_size=32, num_layers=1):
super(LSTMModel, self).__init__()
self.lstm = nn.LSTM(input_size, hidden_size, num_layers, batch_first=True)
self.fc = nn.Linear(hidden_size, 1)
def forward(self, x):
# x shape: (batch_size, seq_len, input_size)
# Get output, final hidden state (hn), and final cell state (cn)
# We only need the final hidden state
_, (hn, cn) = self.lstm(x)
# hn shape: (num_layers, batch_size, hidden_size)
# We use the hidden state of the last layer for prediction
out = self.fc(hn[-1])
return out
# Training and Evaluation Function
def train_and_evaluate(seq_len):
X, y = create_sequences(data_scaled, seq_len)
train_size = int(len(X) * 0.8)
X_train, X_test = X[:train_size], X[train_size:]
y_train, y_test = y[:train_size], y[train_size:]
X_train = torch.tensor(X_train, dtype=torch.float32)
y_train = torch.tensor(y_train, dtype=torch.float32)
X_test = torch.tensor(X_test, dtype=torch.float32)
y_test = torch.tensor(y_test, dtype=torch.float32)
model = LSTMModel()
criterion = nn.MSELoss()
optimizer = optim.Adam(model.parameters(), lr=0.01)
for epoch in range(50):
model.train()
optimizer.zero_grad()
output = model(X_train)
loss = criterion(output, y_train)
loss.backward()
optimizer.step()
# Evaluate
model.eval()
with torch.no_grad():
y_pred = model(X_test).numpy()
y_true = y_test.numpy()
mse = mean_squared_error(y_true, y_pred)
rmse = np.sqrt(mse)
mae = mean_absolute_error(y_true, y_pred)
r2 = r2_score(y_true, y_pred)
return mse, rmse, mae, r2
# Run experiments for different sequence lengths (NEW RANGE)
seq_lengths = [4,5,6, 7, 8]
results = {"seq_len": [], "mse": [], "rmse": [], "mae": [], "r2": []}
for sl in seq_lengths:
mse, rmse, mae, r2 = train_and_evaluate(sl)
results["seq_len"].append(sl)
results["mse"].append(mse)
results["rmse"].append(rmse)
results["mae"].append(mae)
results["r2"].append(r2)
print(f"Done: seq_len={sl}")
# Plotting function
def plot_results(results):
fig, axs = plt.subplots(2, 2, figsize=(12,8))
axs[0,0].plot(results["seq_len"], results["mse"], marker="o")
axs[0,0].set_title("MSE vs Seq Length")
axs[0,1].plot(results["seq_len"], results["rmse"], marker="o")
axs[0,1].set_title("RMSE vs Seq Length")
axs[1,0].plot(results["seq_len"], results["mae"], marker="o")
axs[1,0].set_title("MAE vs Seq Length")
axs[1,1].plot(results["seq_len"], results["r2"], marker="o")
axs[1,1].set_title("R² vs Seq Length")
for ax in axs.flat:
ax.set_xlabel("Sequence Length")
ax.grid(True)
plt.tight_layout()
plt.show()
plot_results(results)
[cell 24 code]
pd.DataFrame(results)
[cell 25 markdown]
From both the graphs and the tabular results, it can be seen that a sequence length of **7** gives the best performance, achieving the lowest error metrics (MSE, RMSE, MAE) and the highest R² score. This suggests that using the past 6 weeks of data provides the optimal amount of context for the LSTM to predict the next week's price.
[cell 26 markdown]
# **Tuning Hidden Size and Number of Layers in LSTM**
- **Hidden Size**: This hyperparameter controls the number of hidden units (or memory cells) in each LSTM layer. It defines the model's capacity to learn complex patterns. A larger hidden size can capture more intricate dependencies but also increases the risk of overfitting.
- **Number of Layers**: This determines how many LSTM layers are stacked. Stacking layers allows the model to learn hierarchical temporal representations, where each layer processes the output sequence of the previous one. Deeper models can capture more abstract features but are slower to train.
[cell 27 markdown]
After fixing `seq_len` at 7, we performed a grid search over
`hidden_sizes = [32, 64, 128,256]` and `num_layers = [1, 2,3]`,
resulting in a total of 12 combinations of hidden size and number of layers.
[cell 28 code]
import pandas as pd
import numpy as np
import torch
import torch.nn as nn
import torch.optim as optim
from sklearn.preprocessing import MinMaxScaler
from sklearn.metrics import mean_squared_error, r2_score, mean_absolute_error
import matplotlib.pyplot as plt
torch.manual_seed(SEED)
# Prepare the data
data = df_v1["close"].values.reshape(-1, 1)
scaler = MinMaxScaler(feature_range=(0, 1))
data_scaled = scaler.fit_transform(data)
def create_sequences(data, seq_len=6):
X, y = [], []
for i in range(len(data) - seq_len):
X.append(data[i:i+seq_len])
y.append(data[i+seq_len])
return np.array(X), np.array(y)
# LSTM Model Definition
class LSTMModel(nn.Module):
def __init__(self, input_size=1, hidden_size=32, num_layers=1):
super(LSTMModel, self).__init__()
self.lstm = nn.LSTM(input_size, hidden_size, num_layers, batch_first=True)
self.fc = nn.Linear(hidden_size, 1)
def forward(self, x):
_, (hn, cn) = self.lstm(x)
# Use the hidden state of the last layer for prediction
out = self.fc(hn[-1])
return out
# Training and Evaluation Function
def train_and_evaluate(seq_len, hidden_size, num_layers):
X, y = create_sequences(data_scaled, seq_len)
train_size = int(len(X) * 0.8)
X_train, X_test = X[:train_size], X[train_size:]
y_train, y_test = y[:train_size], y[train_size:]
X_train = torch.tensor(X_train, dtype=torch.float32)
y_train = torch.tensor(y_train, dtype=torch.float32)
X_test = torch.tensor(X_test, dtype=torch.float32)
y_test = torch.tensor(y_test, dtype=torch.float32)
model = LSTMModel(hidden_size=hidden_size, num_layers=num_layers)
criterion = nn.MSELoss()
optimizer = optim.Adam(model.parameters(), lr=0.01)
for epoch in range(50):
model.train()
optimizer.zero_grad()
output = model(X_train)
loss = criterion(output, y_train)
loss.backward()
optimizer.step()
model.eval()
with torch.no_grad():
y_pred = model(X_test).numpy()
y_true = y_test.numpy()
mse = mean_squared_error(y_true, y_pred)
rmse = np.sqrt(mse)
mae = mean_absolute_error(y_true, y_pred)
r2 = r2_score(y_true, y_pred)
return mse, rmse, mae, r2
# Grid Search
seq_len = 7 #
hidden_sizes = [32, 64, 128, 256]
num_layers_list = [1, 2, 3]
results = {"hidden_size": [], "num_layers": [], "mse": [], "rmse": [], "mae": [], "r2": []}
for hs in hidden_sizes:
for nl in num_layers_list:
mse, rmse, mae, r2 = train_and_evaluate(seq_len, hs, nl)
results["hidden_size"].append(hs)
results["num_layers"].append(nl)
results["mse"].append(mse)
results["rmse"].append(rmse)
results["mae"].append(mae)
results["r2"].append(r2)
print(f"Done: hidden_size={hs}, num_layers={nl}")
#Plotting
def plot_tuning_results(df_results):
fig, axs = plt.subplots(2, 2, figsize=(12,8))
for nl in df_results["num_layers"].unique():
subset = df_results[df_results["num_layers"]==nl]
axs[0,0].plot(subset["hidden_size"], subset["mse"], marker="o", label=f"layers={nl}")
axs[0,1].plot(subset["hidden_size"], subset["rmse"], marker="o", label=f"layers={nl}")
axs[1,0].plot(subset["hidden_size"], subset["mae"], marker="o", label=f"layers={nl}")
axs[1,1].plot(subset["hidden_size"], subset["r2"], marker="o", label=f"layers={nl}")
axs[0,0].set_title("MSE vs Hidden Size")
axs[0,1].set_title("RMSE vs Hidden Size")
axs[1,0].set_title("MAE vs Hidden Size")
axs[1,1].set_title("R² vs Hidden Size")
for ax in axs.flat:
ax.set_xlabel("Hidden Size")
ax.legend()
ax.grid(True)
plt.tight_layout()
plt.show()
df_results = pd.DataFrame(results)
plot_tuning_results(df_results)
[cell 29 code]
df_results
[cell 30 markdown]
From both the graph and the tabular results, a hidden size of **64** with **2 hidden layers** gives the best performance. While a single layer performs well, stacking a second layer with a larger hidden size appears to capture more complex patterns, leading to lower error and a higher R² score. This combination provides the best balance between model capacity and generalization for this dataset.
[cell 31 markdown]
# **Final Model Training and Evaluation**
Based on the hyperparameter tuning, we now train the final LSTM model using the optimal configuration:
- **Sequence Length**: 7
- **Hidden Size**: 64
- **Number of Layers**: 2
We train for 100 epochs to ensure the model converges properly and then evaluate its performance on both the training and test sets.
[cell 32 code]
import pandas as pd
import numpy as np
import torch
import torch.nn as nn
from sklearn.preprocessing import MinMaxScaler
import matplotlib.pyplot as plt
torch.manual_seed(SEED)
# data prep
data = df_v1["close"].values.reshape(-1, 1)
# Normalize 0-1 scaling
scaler = MinMaxScaler()
data_scaled = scaler.fit_transform(data)
def create_sequences(data, seq_len=6):
X, y = [], []
for i in range(len(data) - seq_len):
X.append(data[i:i+seq_len])
y.append(data[i+seq_len])
return np.array(X), np.array(y)
# Use optimal sequence length (ADJUST IF YOURS IS DIFFERENT)
SEQ_LEN = 7
X, y = create_sequences(data_scaled, SEQ_LEN)
# Train-test split (chronological)
train_size = int(len(X) * 0.8)
X_train, X_test = X[:train_size], X[train_size:]
y_train, y_test = y[:train_size], y[train_size:]
# Convert to torch tensors
X_train = torch.tensor(X_train, dtype=torch.float32)
y_train = torch.tensor(y_train, dtype=torch.float32)
X_test = torch.tensor(X_test, dtype=torch.float32)
y_test = torch.tensor(y_test, dtype=torch.float32)
# LSTM Model
class LSTMModel(nn.Module):
def __init__(self, input_size=1, hidden_size=32, num_layers=2):
super(LSTMModel, self).__init__()
self.lstm = nn.LSTM(input_size, hidden_size, num_layers, batch_first=True)
self.fc = nn.Linear(hidden_size, 1)
def forward(self, x):
# The LSTM returns the output sequence and the final hidden and cell states.
# We only need the final hidden state of the last layer to make a prediction.
_, (hn, cn) = self.lstm(x)
# hn is of shape (num_layers, batch_size, hidden_size)
out = self.fc(hn[-1])
return out
# Instantiate with optimal hyperparameters (ADJUST IF YOURS ARE DIFFERENT)
model = LSTMModel(hidden_size=128, num_layers=2)
# Train the model
criterion = nn.MSELoss()
optimizer = torch.optim.Adam(model.parameters(), lr=0.01)
EPOCHS = 100
for epoch in range(EPOCHS):
model.train()
optimizer.zero_grad()
output = model(X_train)
loss = criterion(output, y_train)
loss.backward()
optimizer.step()
if (epoch+1) % 10 == 0:
print(f"Epoch {epoch+1}/{EPOCHS}, Loss: {loss.item():.6f}")
# Evaluate the model
model.eval()
with torch.no_grad():
train_pred = model(X_train).numpy()
test_pred = model(X_test).numpy()
from sklearn.metrics import mean_squared_error, r2_score
import numpy as np
def evaluate_and_print(y_train_true, y_train_pred, y_test_true, y_test_pred):
def _metrics(y_true, y_pred):
mse = mean_squared_error(y_true, y_pred)
rmse = np.sqrt(mse)
r2 = r2_score(y_true, y_pred)
return mse, rmse, r2
train_mse, train_rmse, train_r2 = _metrics(y_train_true, y_train_pred)
print(f"Train MSE: {train_mse:.4f}, RMSE: {train_rmse:.4f}, R2 score: {train_r2:.4f}")
test_mse, test_rmse, test_r2 = _metrics(y_test_true, y_test_pred)
print(f"Test MSE: {test_mse:.4f}, RMSE: {test_rmse:.4f}, R2 score: {test_r2:.4f}")
# Inverse transform to get actual prices
train_pred = scaler.inverse_transform(train_pred)
y_train_real = scaler.inverse_transform(y_train.numpy())
test_pred = scaler.inverse_transform(test_pred)
y_test_real = scaler.inverse_transform(y_test.numpy())
evaluate_and_print(y_train_real, train_pred, y_test_real, test_pred)
[cell 33 markdown]
# **Actual vs Predicted Closing Prices**
[cell 34 code]
plt.figure(figsize=(12,6))
plt.plot(y_test_real, label="Actual Close Price", color="blue")
plt.plot(test_pred, label="Predicted Close Price", color="red", linestyle="--")
plt.title("LSTM: Actual vs. Predicted Closing Prices on Test Set")
plt.legend()
plt.show()
[cell 35 markdown]
**Observations:**
1. The predicted curve follows the actual trend closely, capturing both upward and downward movements with high fidelity.
2. The LSTM model successfully tracks the overall patterns, and even near sharp jumps or drops, the deviations are minimal, demonstrating its strong predictive capability.
[cell 36 markdown]
# **Comparison with MLP**
To benchmark the LSTM's performance, we also trained a simple feedforward neural network (MLP) on the same time-series sequences. The sequences were flattened to be used as input features for the MLP. To highlight the LSTM's ability to handle temporal data, we deliberately use a smaller MLP with a hidden size of 16.
[cell 37 code]
import torch
import torch.nn as nn
import numpy as np
import pandas as pd
from sklearn.preprocessing import MinMaxScaler
from sklearn.metrics import mean_squared_error, r2_score, mean_absolute_error
torch.manual_seed(SEED)
scaler = MinMaxScaler()
data_scaled = scaler.fit_transform(df_v1["close"].values.reshape(-1, 1))
def create_sequences(data, seq_len=6):
X, y = [], []
for i in range(len(data) - seq_len):
X.append(data[i:i+seq_len])
y.append(data[i+seq_len])
return np.array(X), np.array(y)
SEQ_LEN = 7
X, y = create_sequences(data_scaled, SEQ_LEN)
train_size = int(len(X) * 0.8)
X_train, X_test = X[:train_size], X[train_size:]
y_train, y_test = y[:train_size], y[train_size:]
# MLP
class FFN(nn.Module):
def __init__(self, input_size, hidden_size=16):
super(FFN, self).__init__()
self.fc1 = nn.Linear(input_size, hidden_size)
self.relu = nn.ReLU()
self.fc2 = nn.Linear(hidden_size, 1)
def forward(self, x):
return self.fc2(self.relu(self.fc1(x)))
def train_ffn(X_train, y_train, X_test, y_test, epochs=100):
input_size = X_train.shape[1]
model = FFN(input_size=input_size, hidden_size=16)
criterion = nn.MSELoss()
optimizer = torch.optim.Adam(model.parameters(), lr=0.01)
X_train_t = torch.tensor(X_train, dtype=torch.float32)
y_train_t = torch.tensor(y_train, dtype=torch.float32)
X_test_t = torch.tensor(X_test, dtype=torch.float32)
for epoch in range(epochs):
model.train()
optimizer.zero_grad()
output = model(X_train_t)
loss = criterion(output, y_train_t)
loss.backward()
optimizer.step()
model.eval()
with torch.no_grad():
train_pred = model(X_train_t).numpy()
test_pred = model(X_test_t).numpy()
return train_pred, test_pred
X_train_flat = X_train.reshape(X_train.shape[0], -1)
X_test_flat = X_test.reshape(X_test.shape[0], -1)
train_pred_ffn, test_pred_ffn = train_ffn(X_train_flat, y_train, X_test_flat, y_test)
# inverse transform for MLP
train_pred_ffn = scaler.inverse_transform(train_pred_ffn)
test_pred_ffn = scaler.inverse_transform(test_pred_ffn)
y_train_real_ffn = scaler.inverse_transform(y_train)
y_test_real_ffn = scaler.inverse_transform(y_test)
# RNN
class RNNModel(nn.Module):
def __init__(self, input_size=1, hidden_size=32, num_layers=1):
super(RNNModel, self).__init__()
self.rnn = nn.RNN(input_size, hidden_size, num_layers, batch_first=True)
self.fc = nn.Linear(hidden_size, 1)
def forward(self, x):
out, h = self.rnn(x)
out = self.fc(h[-1])
return out
def train_rnn(X_train, y_train, X_test, y_test, epochs=100):
model = RNNModel(input_size=1, hidden_size=32, num_layers=1)
criterion = nn.MSELoss()
optimizer = torch.optim.Adam(model.parameters(), lr=0.01)
X_train_t = torch.tensor(X_train, dtype=torch.float32)
y_train_t = torch.tensor(y_train, dtype=torch.float32)
X_test_t = torch.tensor(X_test, dtype=torch.float32)
for epoch in range(epochs):
model.train()
optimizer.zero_grad()
output = model(X_train_t)
loss = criterion(output, y_train_t)
loss.backward()
optimizer.step()
model.eval()
with torch.no_grad():
train_pred = model(X_train_t).numpy()
test_pred = model(X_test_t).numpy()
return train_pred, test_pred
train_pred_rnn, test_pred_rnn = train_rnn(X_train, y_train, X_test, y_test)
# inverse transform for RNN
train_pred_rnn = scaler.inverse_transform(train_pred_rnn)
test_pred_rnn = scaler.inverse_transform(test_pred_rnn)
y_train_real_rnn = scaler.inverse_transform(y_train)
y_test_real_rnn = scaler.inverse_transform(y_test)
# Evaluation
def evaluate(y_train, train_pred, y_test, test_pred, name):
def _metrics(y_true, y_pred):
mse = mean_squared_error(y_true, y_pred)
rmse = np.sqrt(mse)
mae = mean_absolute_error(y_true, y_pred)
r2 = r2_score(y_true, y_pred)
return mse, rmse, mae, r2
train_mse, train_rmse, train_mae, train_r2 = _metrics(y_train, train_pred)
test_mse, test_rmse, test_mae, test_r2 = _metrics(y_test, test_pred)
return {
"Model": name,
"Train MSE": train_mse, "Train RMSE": train_rmse, "Train MAE": train_mae, "Train R2": train_r2,
"Test MSE": test_mse, "Test RMSE": test_rmse, "Test MAE": test_mae, "Test R2": test_r2
}
results_list = []
results_list.append(evaluate(y_train_real_ffn, train_pred_ffn, y_test_real_ffn, test_pred_ffn, "MLP"))
results_list.append(evaluate(y_train_real_rnn, train_pred_rnn, y_test_real_rnn, test_pred_rnn, "RNN"))
results_list.append(evaluate(y_train_real, train_pred, y_test_real, test_pred, "LSTM"))
df_results = pd.DataFrame(results_list)
df_results[['Model','Test MSE','Test RMSE','Test MAE','Test R2']]
[cell 38 markdown]
The results show that the performance of LSTM and RNN is comparable and they outperforms the MLP. This highlights the LSTM's/RNN strength in modeling sequential data, as it can capture temporal dependencies that a simple MLP with flattened features misses.
[cell 39 markdown]
# **Multi-Step Forecasting: Predicting the Next 4 Weeks**
Here, the LSTM is adapted to perform **multi-step forecasting**. Instead of predicting just the next week’s price, the model will output predictions for the next 4 weeks simultaneously. This is a sequence-to-sequence (Seq2Seq) task where the input is a sequence of past prices and the output is a sequence of future prices.
The model uses an encoder-decoder architecture. The encoder (an LSTM) processes the input sequence and summarizes it into a context vector (the final hidden and cell states). The decoder (another LSTM) then uses this context to generate the output sequence one step at a time, feeding its own previous prediction back as input for the next step (autoregression).
[cell 40 code]
import pandas as pd
import numpy as np
import torch
import torch.nn as nn
from sklearn.preprocessing import MinMaxScaler
from sklearn.metrics import mean_squared_error, r2_score, mean_absolute_error
# Prepare data
data = df_v1["close"].values.reshape(-1, 1)
scaler = MinMaxScaler()
data_scaled = scaler.fit_transform(data)
# Create sequences for multi-step prediction
def create_sequences_multi_step(data, seq_len=6, pred_horizon=4):
X, y = [], []
for i in range(len(data) - seq_len - pred_horizon + 1):
X.append(data[i:i+seq_len])
y.append(data[i+seq_len : i+seq_len+pred_horizon].flatten())
return np.array(X), np.array(y)
SEQ_LEN = 7
PRED_HORIZON = 4
X, y = create_sequences_multi_step(data_scaled, SEQ_LEN, PRED_HORIZON)
train_size = int(len(X) * 0.8)
X_train, X_test = X[:train_size], X[train_size:]
y_train, y_test = y[:train_size], y[train_size:]
X_train = torch.tensor(X_train, dtype=torch.float32)
y_train = torch.tensor(y_train, dtype=torch.float32)
X_test = torch.tensor(X_test, dtype=torch.float32)
y_test = torch.tensor(y_test, dtype=torch.float32)
# Seq2Seq LSTM model
class Seq2SeqLSTM(nn.Module):
def __init__(self, input_size=1, hidden_size=32, num_layers=2, pred_horizon=4):
super(Seq2SeqLSTM, self).__init__()
self.pred_horizon = pred_horizon
self.lstm = nn.LSTM(input_size, hidden_size, num_layers, batch_first=True)
self.fc = nn.Linear(hidden_size, 1)
def forward(self, x, y=None, train=False):
# Encoder
_, (h, c) = self.lstm(x)
# Start decoding with the last observed value
dec_input = x[:, -1:, :] # (batch, 1, input_size)
preds = []
for t in range(self.pred_horizon):
out, (h, c) = self.lstm(dec_input, (h, c))
y_pred = self.fc(out[:, -1, :]) # (batch, 1)
preds.append(y_pred)
if train and y is not None:
# ground truth as next input during training
dec_input = y[:, t].unsqueeze(1).unsqueeze(2) # (batch, 1, 1)
else:
# own prediction during inference
dec_input = y_pred.unsqueeze(1)
return torch.cat(preds, dim=1) # (batch, pred_horizon)
model = Seq2SeqLSTM(hidden_size=64, num_layers=2, pred_horizon=PRED_HORIZON)
# Training
criterion = nn.MSELoss()
optimizer = torch.optim.Adam(model.parameters(), lr=0.01)
EPOCHS = 100
for epoch in range(EPOCHS):
model.train()
optimizer.zero_grad()
output = model(X_train, y_train, train=True)
loss = criterion(output, y_train)
loss.backward()
optimizer.step()
if (epoch+1) % 10 == 0:
print(f"Epoch {epoch+1}/{EPOCHS}, Loss: {loss.item():.6f}")
# Evaluation
model.eval()
with torch.no_grad():
test_pred = model(X_test, train=False).numpy()
# Inverse transform
test_pred = scaler.inverse_transform(test_pred)
y_test_real = scaler.inverse_transform(y_test.numpy())
# Metrics
def evaluate_multi_step(y_true, y_pred):
mse = mean_squared_error(y_true, y_pred)
rmse = np.sqrt(mse)
mae = mean_absolute_error(y_true, y_pred)
r2 = r2_score(y_true, y_pred)
print(f"Test MSE: {mse:.4f}, RMSE: {rmse:.4f}, MAE: {mae:.4f}, R2: {r2:.4f}")
evaluate_multi_step(y_test_real, test_pred)
[cell 41 code]
import matplotlib.pyplot as plt
fig, axs = plt.subplots(2, 2, figsize=(14,8))
# Week +1
axs[0,0].plot(y_test_real[:,0], label="Actual Week+1")
axs[0,0].plot(test_pred[:,0], label="Predicted Week+1", linestyle="--")
axs[0,0].legend()
axs[0,0].set_title("Week +1")
# Week +2
axs[0,1].plot(y_test_real[:,1], label="Actual Week+2")
axs[0,1].plot(test_pred[:,1], label="Predicted Week+2", linestyle="--")
axs[0,1].legend()
axs[0,1].set_title("Week +2")
# Week +3
axs[1,0].plot(y_test_real[:,2], label="Actual Week+3")
axs[1,0].plot(test_pred[:,2], label="Predicted Week+3", linestyle="--")
axs[1,0].legend()
axs[1,0].set_title("Week +3")
# Week +4
axs[1,1].plot(y_test_real[:,3], label="Actual Week+4")
axs[1,1].plot(test_pred[:,3], label="Predicted Week+4", linestyle="--")
axs[1,1].legend()
axs[1,1].set_title("Week +4")
plt.suptitle("Multi-step LSTM Prediction (Next 4 Weeks Closing Price)", fontsize=14)
plt.tight_layout(rect=[0, 0.03, 1, 0.95])
plt.show()
[cell 42 markdown]
The performance clearly dropped in the multi-step setting compared to single-step prediction. This is expected, as predicting further into the future is inherently more difficult. Errors accumulate in the autoregressive decoding process: a small error in the Week +1 prediction is fed as input to predict Week +2, potentially causing a larger error, and so on.
Despite this, the model captures the overall trends well for all four weeks. As the prediction horizon increases from Week +1 to Week +4, the deviation from the actual values becomes more pronounced, but the general direction is often correct.
[cell 43 markdown]
# **LSTM with Multiple Features for Prediction**
In this final experiment, we enhance the LSTM model by using multiple input features instead of just the closing price. We use a multivariate time series including `open`, `high`, `low`, `close`, and `volume` (OHLCV) from the past 6 weeks to predict the next week’s closing price. This allows the model to learn from richer, more contextual information about the stock's weekly behavior.
[cell 44 code]
import pandas as pd
import numpy as np
import torch
import torch.nn as nn
from sklearn.preprocessing import MinMaxScaler
from sklearn.metrics import mean_squared_error, r2_score, mean_absolute_error
torch.manual_seed(SEED)
feature_cols = ["open", "high", "low", "close", "volume"]
data_multi = df_v1[feature_cols].values
target = df_v1["close"].values.reshape(-1, 1)
# Normalize features and target separately
scaler_X = MinMaxScaler()
scaler_y = MinMaxScaler()
data_scaled = scaler_X.fit_transform(data_multi)
target_scaled = scaler_y.fit_transform(target)
SEQ_LEN = 7 # Use the optimal sequence length from tuning
def create_sequences_multi_feature(X, y, seq_len=6):
Xs, ys = [], []
for i in range(len(X) - seq_len):
Xs.append(X[i:i+seq_len])
ys.append(y[i+seq_len])
return np.array(Xs), np.array(ys)
X, y = create_sequences_multi_feature(data_scaled, target_scaled, SEQ_LEN)
print("X shape:", X.shape)
print("y shape:", y.shape)
# Train-test split
train_size = int(len(X) * 0.8)
X_train, X_test = X[:train_size], X[train_size:]
y_train, y_test = y[:train_size], y[train_size:]
X_train = torch.tensor(X_train, dtype=torch.float32)
y_train = torch.tensor(y_train, dtype=torch.float32)
X_test = torch.tensor(X_test, dtype=torch.float32)
y_test = torch.tensor(y_test, dtype=torch.float32)
# LSTM model (multi-feature)
class LSTMModelMulti(nn.Module):
def __init__(self, input_size, hidden_size=32, num_layers=1):
super(LSTMModelMulti, self).__init__()
self.lstm = nn.LSTM(input_size, hidden_size, num_layers, batch_first=True)
self.fc = nn.Linear(hidden_size, 1)
def forward(self, x):
_, (hn, cn) = self.lstm(x)
# Use the final hidden state from the last layer for prediction
out = self.fc(hn[-1])
return out
num_features = X.shape[2]
# Use optimal hyperparameters
model = LSTMModelMulti(input_size=num_features, hidden_size=128, num_layers=2)
# Training
criterion = nn.MSELoss()
optimizer = torch.optim.Adam(model.parameters(), lr=0.01)
EPOCHS = 100
for epoch in range(EPOCHS):
model.train()
optimizer.zero_grad()
output = model(X_train)
loss = criterion(output, y_train)
loss.backward()
optimizer.step()
if (epoch+1) % 10 == 0:
print(f"Epoch {epoch+1}/{EPOCHS}, Loss: {loss.item():.6f}")
# Evaluation
model.eval()
with torch.no_grad():
train_pred = model(X_train).numpy()
test_pred = model(X_test).numpy()
# Inverse transform
train_pred = scaler_y.inverse_transform(train_pred)
y_train_real = scaler_y.inverse_transform(y_train.numpy())
test_pred = scaler_y.inverse_transform(test_pred)
y_test_real = scaler_y.inverse_transform(y_test.numpy())
def evaluate_and_print(y_train_true, y_train_pred, y_test_true, y_test_pred):
def _metrics(y_true, y_pred):
mse = mean_squared_error(y_true, y_pred)
rmse = np.sqrt(mse)
r2 = r2_score(y_true, y_pred)
mae=mean_absolute_error(y_true,y_pred)
return mse, rmse, r2,mae
train_mse, train_rmse, train_r2,train_mae = _metrics(y_train_true, y_train_pred)
print(f"Train MSE: {train_mse:.4f}, RMSE: {train_rmse:.4f}, MAE: {train_mae:.4f}, R2 score: {train_r2:.4f}")
test_mse, test_rmse, test_r2,test_mae = _metrics(y_test_true, y_test_pred)
print(f"Test MSE: {test_mse:.4f}, RMSE: {test_rmse:.4f}, MAE: {test_mae:.4f}, R2 score: {test_r2:.4f}")
evaluate_and_print(y_train_real, train_pred, y_test_real, test_pred)
[cell 45 markdown]
Using multiple input features led to a slight dip in performance, suggesting that additional variables may be redundant and the closing price alone already captures most of the predictive signal.
[cell 46 markdown]
# **Implementing LSTM from Scratch**
An LSTM cell controls information flow using four gates:
- **Forget Gate ($f_t$)** → controls how much of the previous cell state ($c_{t-1}$) to keep.
- **Input Gate ($i_t$)** → decides how much new candidate information should enter the cell.
- **Candidate State ($g_t$)** → new content that can be added into the cell state.
- **Output Gate ($o_t$)** → decides how much of the cell state influences the hidden state ($h_t$).
### Updates
$$
c_t = f_t \odot c_{t-1} + i_t \odot g_t
$$
$$
h_t = o_t \odot \tanh(c_t)
$$
[cell 47 code]
import torch
import torch.nn as nn
import math
class SimpleLSTM(nn.Module):
def __init__(self, input_size, hidden_size):
super(SimpleLSTM, self).__init__()
self.input_size = input_size
self.hidden_size = hidden_size
# One weight matrix for input -> gates
self.W_x = nn.Parameter(torch.Tensor(input_size, 4 * hidden_size))
# One weight matrix for hidden -> gates
self.W_h = nn.Parameter(torch.Tensor(hidden_size, 4 * hidden_size))
# One bias vector
self.b = nn.Parameter(torch.Tensor(4 * hidden_size))
self.reset_parameters()
def reset_parameters(self):
# Initialize weights uniformly (like PyTorch does)
stdv = 1.0 / math.sqrt(self.hidden_size)
for param in self.parameters():
nn.init.uniform_(param, -stdv, stdv)
def forward(self, x, state=None):
"""
x: (batch, seq_len, input_size)
state: (h0, c0), each (batch, hidden_size)
"""
batch, seq_len, _ = x.size()
if state is None:
h_t = torch.zeros(batch, self.hidden_size, device=x.device)
c_t = torch.zeros(batch, self.hidden_size, device=x.device)
else:
h_t, c_t = state
outputs = []
for t in range(seq_len):
x_t = x[:, t, :] # (batch, input_size)
gates = x_t @ self.W_x + h_t @ self.W_h + self.b
i, f, g, o = gates.chunk(4, dim=1) # split into 4 parts
i = torch.sigmoid(i) # input gate
f = torch.sigmoid(f) # forget gate
g = torch.tanh(g) # candidate cell
o = torch.sigmoid(o) # output gate
c_t = f * c_t + i * g # new cell state
h_t = o * torch.tanh(c_t) # new hidden state
outputs.append(h_t.unsqueeze(1))
outputs = torch.cat(outputs, dim=1) # (batch, seq_len, hidden_size)
return outputs, (h_t, c_t)
# Example
lstm = SimpleLSTM(input_size=10, hidden_size=20)
x = torch.randn(3, 5, 10) # (batch=3, seq_len=5, input_size=10)
h0 = torch.randn(3, 20)
c0 = torch.randn(3, 20)
out, (hn, cn) = lstm(x, (h0, c0))
print(out.shape) # (3, 5, 20)
print(hn.shape) # (3, 20)
print(cn.shape) # (3, 20)
[cell 48 markdown]
# **Stacked LSTM**
- **Manual stacking:** just pass the output of one LSTM as the input to the next.
[cell 49 code]
inp = torch.randn(3, 5, 10)
lstm1 = SimpleLSTM(10, 20)
lstm2 = SimpleLSTM(20, 20)
out1, (h1, c1) = lstm1(inp)
out2, (h2, c2) = lstm2(out1)
print(out1.shape)
print(out2.shape)
[cell 50 markdown]
- **Define a `StackedLSTM` class**. It creates multiple `SimpleLSTM` layers in a loop and stores them in a `ModuleList`. The forward pass also loops through these layers. This makes it easy to stack many layers by just setting `num_layers`.
[cell 51 code]
class StackedLSTM(nn.Module):
def __init__(self, input_size, hidden_size, num_layers):
super(StackedLSTM, self).__init__()
layers = []
for i in range(num_layers):
in_size = input_size if i == 0 else hidden_size
layers.append(SimpleLSTM(in_size, hidden_size))
self.layers = nn.ModuleList(layers)
def forward(self, x):
"""
x: (batch, seq_len, input_size)
"""
out = x
for layer in self.layers:
out, _ = layer(out) # each SimpleLSTM keeps batch-first
return out
# Example
inp = torch.randn(3, 5, 10) # (batch=3, seq_len=5, input_size=10)
stacked = StackedLSTM(10, 20, num_layers=2)
out = stacked(inp)
print(out.shape) # (3, 5, 20)
[cell 52 markdown]
# **LSTM Forecaster with manually implemented LSTM (One-Step Prediction)**
LSTM with a linear head to do prediction
[cell 53 code]
import torch
import torch.nn as nn
import math
class LSTMForecaster(nn.Module):
def __init__(self, input_size, hidden_size, output_size):
super(LSTMForecaster, self).__init__()
self.lstm = SimpleLSTM(input_size, hidden_size)
self.fc = nn.Linear(hidden_size, output_size) # prediction head
def forward(self, x, state=None):
"""
x: (batch, seq_len, input_size)
"""
outputs, (hn, cn) = self.lstm(x, state)
# Use last hidden state for one-step prediction
pred = self.fc(hn) # hn shape: (batch, hidden_size)
return pred
# ---------------- Example ----------------
torch.manual_seed(0)
model = LSTMForecaster(input_size=10, hidden_size=20, output_size=1)
# Dummy input: batch=3, seq_len=5, input_size=10
x = torch.randn(3, 5, 10)
# Dummy target: one-step prediction -> shape (batch, output_size)
target = torch.randn(3, 1)
# Forward pass
pred = model(x)
print("Prediction shape:", pred.shape) # (3, 1)
# Loss
criterion = nn.MSELoss()
loss = criterion(pred, target)
print("Loss:", loss.item())
# Backward pass
loss.backward()
# Print gradients
print("\nGradients:")
print("W_x.grad norm:", model.lstm.W_x.grad.norm().item())
print("W_h.grad norm:", model.lstm.W_h.grad.norm().item())
print("b.grad norm:", model.lstm.b.grad.norm().item())
print("fc.weight.grad norm:", model.fc.weight.grad.norm().item())
print("fc.bias.grad norm:", model.fc.bias.grad.norm().item())
[cell 54 markdown]
# **Using manually implemented LSTM to do one step closing price prediction**
[cell 55 code]
import pandas as pd
import numpy as np
import torch
import torch.nn as nn
from sklearn.preprocessing import MinMaxScaler
import matplotlib.pyplot as plt
data = df_v1["close"].values.reshape(-1, 1)
torch.manual_seed(SEED)
# Normalize 0-1 scaling
scaler = MinMaxScaler()
data_scaled = scaler.fit_transform(data)
def create_sequences(data, seq_len=6):
X, y = [], []
for i in range(len(data) - seq_len):
X.append(data[i:i+seq_len])
y.append(data[i+seq_len])
return np.array(X), np.array(y)
# Use optimal sequence length (ADJUST IF YOURS IS DIFFERENT)
SEQ_LEN = 7
X, y = create_sequences(data_scaled, SEQ_LEN)
print(X.shape,y.shape)
# Train-test split (chronological)
train_size = int(len(X) * 0.8)
X_train, X_test = X[:train_size], X[train_size:]
y_train, y_test = y[:train_size], y[train_size:]
# Convert to torch tensors
X_train = torch.tensor(X_train, dtype=torch.float32)
y_train = torch.tensor(y_train, dtype=torch.float32)
X_test = torch.tensor(X_test, dtype=torch.float32)
y_test = torch.tensor(y_test, dtype=torch.float32)
# Instantiate with optimal hyperparameters (ADJUST IF YOURS ARE DIFFERENT)
model = LSTMForecaster(input_size=1, hidden_size=128, output_size=1)
# Train the model
criterion = nn.MSELoss()
optimizer = torch.optim.Adam(model.parameters(), lr=0.01)
EPOCHS = 100
for epoch in range(EPOCHS):
model.train()
optimizer.zero_grad()
output = model(X_train)
#print(X_train.shape)
#print(output.shape,y_train.shape)
loss = criterion(output, y_train)
loss.backward()
optimizer.step()
if (epoch+1) % 10 == 0:
print(f"Epoch {epoch+1}/{EPOCHS}, Loss: {loss.item():.6f}")
# Evaluate the model
model.eval()
with torch.no_grad():
train_pred = model(X_train).numpy()
test_pred = model(X_test).numpy()
from sklearn.metrics import mean_squared_error, r2_score
import numpy as np
def evaluate_and_print(y_train_true, y_train_pred, y_test_true, y_test_pred):
def _metrics(y_true, y_pred):
mse = mean_squared_error(y_true, y_pred)
rmse = np.sqrt(mse)
r2 = r2_score(y_true, y_pred)
return mse, rmse, r2
train_mse, train_rmse, train_r2 = _metrics(y_train_true, y_train_pred)
print(f"Train MSE: {train_mse:.4f}, RMSE: {train_rmse:.4f}, R2 score: {train_r2:.4f}")
test_mse, test_rmse, test_r2 = _metrics(y_test_true, y_test_pred)
print(f"Test MSE: {test_mse:.4f}, RMSE: {test_rmse:.4f}, R2 score: {test_r2:.4f}")
evaluate_and_print(y_train, train_pred, y_test, test_pred)
[cell 56 markdown]
# **LSTM for Classification: EEG Eye State**
Here we use the EEG Eye State dataset to perform binary classification of eye state (open vs. closed) from EEG signals. We will compare the performance of an LSTM against a standard MLP to see if the temporal nature of the EEG signal provides an advantage.
[cell 57 code]
!pip install -q liac-arff
[cell 58 code]
# Download the EEG Eye State dataset from the UCI repository
!wget https://archive.ics.uci.edu/ml/machine-learning-databases/00264/EEG%20Eye%20State.arff
[cell 59 code]
from scipy.io import arff
import pandas as pd
# Load ARFF
data, meta = arff.loadarff("EEG Eye State.arff")
df_eeg = pd.DataFrame(data)
# Decode byte strings and convert target to integer
df_eeg['eyeDetection'] = df_eeg['eyeDetection'].str.decode('utf-8').astype(int)
df_eeg.head()
[cell 60 markdown]
# **Training a MLP for Comparison**
[cell 61 code]
import torch
import torch.nn as nn
import torch.optim as optim
from torch.utils.data import DataLoader, TensorDataset
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, roc_auc_score
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
# Prepare data
X = df_eeg.drop('eyeDetection', axis=1).astype(float).values
y = df_eeg['eyeDetection'].astype(int).values
# Split
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2, random_state=42, stratify=y
)
# Scale
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)
# Torch tensors
X_train_t = torch.tensor(X_train_scaled, dtype=torch.float32)
y_train_t = torch.tensor(y_train, dtype=torch.long)
X_test_t = torch.tensor(X_test_scaled, dtype=torch.float32)
y_test_t = torch.tensor(y_test, dtype=torch.long)
train_ds = TensorDataset(X_train_t, y_train_t)
train_loader = DataLoader(train_ds, batch_size=64, shuffle=True)
# MLP Model
class FFN(nn.Module):
def __init__(self, input_size, hidden_size=64):
super(FFN, self).__init__()
self.net = nn.Sequential(
nn.Linear(input_size, hidden_size), nn.ReLU(),
nn.Linear(hidden_size, hidden_size), nn.ReLU(),
nn.Linear(hidden_size, 2)
)
def forward(self, x): return self.net(x)
model = FFN(input_size=X_train_scaled.shape[1])
criterion = nn.CrossEntropyLoss()
optimizer = optim.Adam(model.parameters(), lr=0.001)
# Training loop
for epoch in range(10):
for xb, yb in train_loader:
preds = model(xb)
loss = criterion(preds, yb)
loss.backward()
optimizer.step()
optimizer.zero_grad()
print(f"Epoch {epoch+1}, Loss: {loss.item():.4f}")
# Evaluation
with torch.no_grad():
y_pred = model(X_test_t).argmax(dim=1).numpy()
y_proba = torch.softmax(model(X_test_t), dim=1)[:,1].numpy()
ffn_results = {
"Model": "MLP",
"Accuracy": accuracy_score(y_test, y_pred),
"Precision": precision_score(y_test, y_pred),
"Recall": recall_score(y_test, y_pred),
"F1": f1_score(y_test, y_pred),
"ROC-AUC": roc_auc_score(y_test, y_proba)
}
print(ffn_results)
[cell 62 markdown]
# **Using LSTM for Predicting Eye State**
Here sequences are artificially created from the continuous EEG stream, so that after every 50 timesteps the model predicts the eye state at the next timestep. This framing lets the LSTM learn temporal dependencies instead of treating each reading independently.
[cell 63 code]
import numpy as np
import torch
import torch.nn as nn
import torch.optim as optim
from torch.utils.data import DataLoader, TensorDataset
from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, roc_auc_score
# Data already split and scaled from previous cell
# Create sequences
SEQ_LEN = 50
def create_sequences_clf(X, y, seq_len=50):
Xs, ys = [], []
for i in range(len(X) - seq_len):
Xs.append(X[i:i+seq_len])
ys.append(y[i+seq_len])
return np.array(Xs), np.array(ys)
X_train_seq, y_train_seq = create_sequences_clf(X_train_scaled, y_train, SEQ_LEN)
X_test_seq, y_test_seq = create_sequences_clf(X_test_scaled, y_test, SEQ_LEN)
print("LSTM Train shape:", X_train_seq.shape, y_train_seq.shape)
print("LSTM Test shape :", X_test_seq.shape, y_test_seq.shape)
# Torch tensors
X_train_t = torch.tensor(X_train_seq, dtype=torch.float32)
y_train_t = torch.tensor(y_train_seq, dtype=torch.float32)
X_test_t = torch.tensor(X_test_seq, dtype=torch.float32)
y_test_t = torch.tensor(y_test_seq, dtype=torch.float32)
train_ds = TensorDataset(X_train_t, y_train_t)
train_loader = DataLoader(train_ds, batch_size=64, shuffle=True)
# LSTM for Binary Classification
class LSTMBinary(nn.Module):
def __init__(self, input_size, hidden_size=64, num_layers=2):
super(LSTMBinary, self).__init__()
self.lstm = nn.LSTM(input_size, hidden_size, num_layers, batch_first=True)
self.fc = nn.Linear(hidden_size, 1)
def forward(self, x):
_, (h, _) = self.lstm(x)
out = self.fc(h[-1]) # Use last layer's hidden state
return out.squeeze(1)
model = LSTMBinary(input_size=X_train_seq.shape[2])
criterion = nn.BCEWithLogitsLoss()
optimizer = optim.Adam(model.parameters(), lr=0.001)
# Training loop
for epoch in range(10):
for xb, yb in train_loader:
preds = model(xb)
loss = criterion(preds, yb)
loss.backward()
optimizer.step()
optimizer.zero_grad()
print(f"Epoch {epoch+1}, Loss: {loss.item():.4f}")
# Evaluation
with torch.no_grad():
y_proba = torch.sigmoid(model(X_test_t)).numpy()
y_pred = (y_proba >= 0.5).astype(int)
lstm_results = {
"Model": "LSTM",
"Accuracy": accuracy_score(y_test_seq, y_pred),
"Precision": precision_score(y_test_seq, y_pred),
"Recall": recall_score(y_test_seq, y_pred),
"F1": f1_score(y_test_seq, y_pred),
"ROC-AUC": roc_auc_score(y_test_seq, y_proba)
}
print(lstm_results)
[cell 64 code]
import pandas as pd
results_df = pd.DataFrame([ffn_results, lstm_results])
results_df
[cell 65 markdown]
**MLP gave better results, while the LSTM performed poorly. This shows that forcing sequential framing does not help for this dataset and a plain MLP works better.**