# MLP backprop DryBean

course: Module 3 — Deep Learning & NLP
module: Module-3-Deep-Learning-NLP
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-3-Deep-Learning-NLP/General/Lab_Materials-14-03-2026/MLP_backprop_DryBean.pdf
pages: 41

---
[page 1]
MLP_backprop_DryBean
March 12, 2026
1 MLP with PyTorch - Backpropagation on Dry Bean Dataset
This notebook trains a Multi-Layer Perceptron (MLP) in PyT orchto classify 7 types of dry
beans from shape measurements.
Dataset Dry Bean Dataset (UCI)
Samples 13,611
F eatures 16 geometric shape features
T ask Multi-class Classification — 7 bean types
Architecture Input(16) → Hidden(8, ReLU) → Output(7, Softmax)
F ramework PyT orch
Learning goals: - Understand how PyTorch builds a computation graph during the forward
pass - See how loss.backward() runs backpropagation automatically via the chain rule -
Watch weights and gradients evolve every 10 epochs - Compare manual backprop math to
PyTorch autograd output
1.1 Step 1: Imports
[1]: import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import matplotlib.gridspec as gridspec
import torch
import torch.nn as nn
import torch.optim as optim
import torch.nn.functional as F
from torch.utils.data import DataLoader, TensorDataset
from sklearn.preprocessing import LabelEncoder, StandardScaler
from sklearn.model_selection import train_test_split
# Reproducibility
torch.manual_seed(42)
np.random.seed(42)
1

[page 2]
DEVICE = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
print(f"PyTorch version : {torch.__version__}")
print(f" Device : {DEVICE}")
PyTorch version : 2.10.0+cpu
Device : cpu
1.2 Step 2: Load and Explore the Dataset
[2]: df = pd.read_excel("Dry_Bean_Dataset.xlsx")
print(f"Dataset Shape : {df.shape}")
print(f"Features (16) : {df.columns[:-1].tolist()}")
print(f"\nClass Distribution:")
vc = df['Class'].value_counts()
for label, count in vc.items():
print(f" {label:>10} : {count:>5} samples ({count/len(df)*100:.1f}%)")
df.head()
Dataset Shape : (13611, 17)
Features (16) : ['Area', 'Perimeter', 'MajorAxisLength', 'MinorAxisLength',
'AspectRation', 'Eccentricity', 'ConvexArea', 'EquivDiameter', 'Extent',
'Solidity', 'roundness', 'Compactness', 'ShapeFactor1', 'ShapeFactor2',
'ShapeFactor3', 'ShapeFactor4']
Class Distribution:
DERMASON : 3546 samples (26.1%)
SIRA : 2636 samples (19.4%)
SEKER : 2027 samples (14.9%)
HOROZ : 1928 samples (14.2%)
CALI : 1630 samples (12.0%)
BARBUNYA : 1322 samples (9.7%)
BOMBAY : 522 samples (3.8%)
[2]: Area Perimeter MajorAxisLength MinorAxisLength AspectRation \
0 28395 610.291 208.178117 173.888747 1.197191
1 28734 638.018 200.524796 182.734419 1.097356
2 29380 624.110 212.826130 175.931143 1.209713
3 30008 645.884 210.557999 182.516516 1.153638
4 30140 620.134 201.847882 190.279279 1.060798
Eccentricity ConvexArea EquivDiameter Extent Solidity roundness \
0 0.549812 28715 190.141097 0.763923 0.988856 0.958027
1 0.411785 29172 191.272750 0.783968 0.984986 0.887034
2 0.562727 29690 193.410904 0.778113 0.989559 0.947849
3 0.498616 30724 195.467062 0.782681 0.976696 0.903936
2

[page 3]
4 0.333680 30417 195.896503 0.773098 0.990893 0.984877
Compactness ShapeFactor1 ShapeFactor2 ShapeFactor3 ShapeFactor4 Class
0 0.913358 0.007332 0.003147 0.834222 0.998724 SEKER
1 0.953861 0.006979 0.003564 0.909851 0.998430 SEKER
2 0.908774 0.007244 0.003048 0.825871 0.999066 SEKER
3 0.928329 0.007017 0.003215 0.861794 0.994199 SEKER
4 0.970516 0.006697 0.003665 0.941900 0.999166 SEKER
[3]: # Visualise feature distributions per class
features_to_plot = ['Area', 'Perimeter', 'MajorAxisLength', 'MinorAxisLength',
'Eccentricity', 'roundness', 'Compactness', 'Solidity']
class_names_all = df['Class'].unique()
palette = plt.cm.Set2(np.linspace(0, 1, len(class_names_all)))
fig, axes = plt.subplots(4, 2, figsize=(14, 12))
for ax, feat in zip(axes.flat, features_to_plot):
for cls, col in zip(class_names_all, palette):
vals = df[df['Class'] == cls][feat].values
ax.hist(vals, bins=30, alpha=0.55, color=col, label=cls, density=True)
ax.set_title(feat, fontsize=10, fontweight='bold')
ax.set_xlabel('Value', fontsize=8)
ax.tick_params(labelsize=7)
ax.grid(True, alpha=0.2)
handles = [plt.Rectangle((0,0),1,1, fc=c) for c in palette]
fig.legend(handles, class_names_all, loc='lower center', ncol=7,
fontsize=10, bbox_to_anchor=(0.5, -0.04))
plt.suptitle('Feature Distributions by Bean Class',
fontsize=14, fontweight='bold')
plt.tight_layout()
plt.savefig('feature_distributions.png', dpi=500, bbox_inches='tight')
plt.show()
3

[page 4]
1.3 Step 3: Preprocess and Build PyTorch DataLoaders
[4]: # ---------- Features & Labels ----------
X_raw = df.drop(columns=['Class']).values.astype(np.float32) # (13611, 16)
y_raw = df['Class'].values
# ---------- Encode string labels → integers 0..6 ----------
le = LabelEncoder()
y_int = le.fit_transform(y_raw).astype(np.int64) # required by␣
↪CrossEntropyLoss
class_names = le.classes_
n_classes = len(class_names)
print(f"Classes ({n_classes}): {class_names}")
# ---------- Normalize features ----------
scaler = StandardScaler()
4

[page 5]
X_scaled = scaler.fit_transform(X_raw).astype(np.float32)
# ---------- Train / Test split ----------
X_train_np, X_test_np, y_train_np, y_test_np = train_test_split(
X_scaled, y_int, test_size=0.2, random_state=42, stratify=y_int
)
# ---------- Convert to PyTorch tensors ----------
# NOTE: CrossEntropyLoss in PyTorch expects raw logits (no softmax)
# and integer class labels — NOT one-hot vectors.
X_train = torch.tensor(X_train_np).to(DEVICE) # (N, 16) float32
X_test = torch.tensor(X_test_np).to(DEVICE)
y_train = torch.tensor(y_train_np).to(DEVICE) # (N,) int64
y_test = torch.tensor(y_test_np).to(DEVICE)
# ---------- DataLoaders ----------
train_loader = DataLoader(TensorDataset(X_train, y_train),
batch_size=256, shuffle=True)
test_loader = DataLoader(TensorDataset(X_test, y_test),
batch_size=256, shuffle=False)
print(f"\nX_train : {X_train.shape} | X_test : {X_test.shape}")
print(f"y_train : {y_train.shape} | y_test : {y_test.shape}")
print(f"\nTrain batches : {len(train_loader)} (batch_size=256)")
print(f"Test batches : {len(test_loader)}")
Classes (7): ['BARBUNYA' 'BOMBAY' 'CALI' 'DERMASON' 'HOROZ' 'SEKER' 'SIRA']
X_train : torch.Size([10888, 16]) | X_test : torch.Size([2723, 16])
y_train : torch.Size([10888]) | y_test : torch.Size([2723])
Train batches : 43 (batch_size=256)
Test batches : 11
1.4 Step 4: Define the MLP in PyTorch
Input Layer Hidden Layer Output Layer
(16 neurons) → (8 neurons, ReLU) → (7 neurons, Softmax)
Key PyT orch design choice for multiclass:
nn.CrossEntropyLoss = LogSoftmax + NLLLoss combined.
So the model outputs raw logits (no softmax in forward).
Softmax is only applied manually when we need probabilities (e.g., for display).
[5]: class MLP(nn.Module):
"""
MLP: Input(16) → Hidden(8, ReLU) → Output(7, raw logits)
5

[page 6]
PyTorch's nn.Linear(in_features, out_features) creates:
W: (out, in) b: (out,)
and computes: z = x @ W.T + b
We return raw logits from the output layer.
CrossEntropyLoss applies log-softmax + NLL internally.
"""
def __init__(self, n_input=16, n_hidden=8, n_output=7):
super(MLP, self).__init__()
self.hidden = nn.Linear(n_input, n_hidden) # W1:(8,16), b1:(8,)
self.output = nn.Linear(n_hidden, n_output) # W2:(7,8), b2:(7,)
self.relu = nn.ReLU()
def forward(self, x):
"""
Every operation here is recorded by PyTorch's autograd engine
into a dynamic computation graph.
Calling loss.backward() later traverses this graph in reverse.
"""
z1 = self.hidden(x) # z1 = x @ W1.T + b1
a1 = self.relu(z1) # a1 = ReLU(z1)
logits = self.output(a1) # logits = a1 @ W2.T + b2 (raw scores)
return logits # CrossEntropyLoss handles softmax internally
def predict_proba(self, x):
"""Returns softmax probabilities (for evaluation/display only)."""
return F.softmax(self.forward(x), dim=1)
# Instantiate
model = MLP(n_input=16, n_hidden=8, n_output=n_classes).to(DEVICE)
print("=" * 55)
print(" MLP Architecture")
print("=" * 55)
print(model)
print("=" * 55)
print(f"\nParameter breakdown:")
total = 0
for name, param in model.named_parameters():
print(f" {name:25s}: shape={str(param.shape):15s} count={param.numel()}")
total += param.numel()
print(f"\n Total trainable parameters: {total}")
=======================================================
MLP Architecture
6

[page 7]
=======================================================
MLP(
(hidden): Linear(in_features=16, out_features=8, bias=True)
(output): Linear(in_features=8, out_features=7, bias=True)
(relu): ReLU()
)
=======================================================
Parameter breakdown:
hidden.weight : shape=torch.Size([8, 16]) count=128
hidden.bias : shape=torch.Size([8]) count=8
output.weight : shape=torch.Size([7, 8]) count=56
output.bias : shape=torch.Size([7]) count=7
Total trainable parameters: 199
1.5 Step 5: How PyTorch Autograd Works
Before training, let’s see PyTorch building and traversing the computation graph on a single sample.
[6]: print("￿￿ PyTorch Autograd Demo — One Sample, One Forward+Backward ￿￿\n")
x_demo = X_train[:1] # shape (1, 16)
y_demo = y_train[:1] # shape (1,) — integer class label
model.zero_grad()
# ￿￿ Forward: PyTorch builds a computation graph ￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
logits_demo = model(x_demo) # (1, 7)
probs_demo = F.softmax(logits_demo, dim=1) # (1, 7)
loss_demo = nn.CrossEntropyLoss()(logits_demo, y_demo)
print(f"Input shape : {x_demo.shape}")
print(f"True class : {class_names[y_demo.item()]} (index {y_demo.
↪item()})")
print(f"\nOutput logits : {logits_demo.detach().numpy()[0].round(4)}")
print(f"After Softmax : ")
for cls, p in zip(class_names, probs_demo.detach().numpy()[0]):
marker = ' ← predicted' if p == probs_demo.max().item() else ''
marker += ' ← TRUE' if cls == class_names[y_demo.item()] else ''
print(f" {cls:>10}: {p:.4f}{marker}")
print(f"\nCross-Entropy Loss : {loss_demo.item():.4f}")
# ￿￿ Backward: one call — PyTorch applies the full chain rule ￿￿￿￿￿￿￿￿￿￿
loss_demo.backward()
print("\n￿￿ Gradients computed by loss.backward() ￿￿")
for name, param in model.named_parameters():
7

[page 8]
if param.grad is not None:
g = param.grad
print(f" ￿L/￿{name:20s}: shape={str(g.shape):12s} "
f"mean={g.mean().item():+.6f} norm={g.norm().item():.6f}")
print("\nPyTorch computed ALL gradients automatically — this is backpropagation!
↪")
print(" It applied the chain rule layer by layer, from output back to input.")
￿￿ PyTorch Autograd Demo — One Sample, One Forward+Backward ￿￿
Input shape : torch.Size([1, 16])
True class : SEKER (index 5)
Output logits : [ 0.0975 0.3665 0.0477 0.0888 -0.0231 -0.0878 0.2331]
After Softmax :
BARBUNYA: 0.1406
BOMBAY: 0.1840 ← predicted
CALI: 0.1337
DERMASON: 0.1393
HOROZ: 0.1246
SEKER: 0.1168 ← TRUE
SIRA: 0.1610
Cross-Entropy Loss : 2.1473
￿￿ Gradients computed by loss.backward() ￿￿
￿L/￿hidden.weight : shape=torch.Size([8, 16]) mean=+0.009082
norm=1.560571
￿L/￿hidden.bias : shape=torch.Size([8]) mean=+0.044977 norm=0.385947
￿L/￿output.weight : shape=torch.Size([7, 8]) mean=+0.000000
norm=0.651042
￿L/￿output.bias : shape=torch.Size([7]) mean=+0.000000 norm=0.955195
PyTorch computed ALL gradients automatically — this is backpropagation!
It applied the chain rule layer by layer, from output back to input.
1.6 Step 6: The PyTorch Training Loop
Every training step follows this exact 5-line recipe:
optimizer.zero_grad() # 1. Clear gradients from previous step
logits = model(X_batch) # 2. Forward pass → builds computation graph
loss = criterion(logits, y_batch)# 3. Compute loss
loss.backward() # 4. BACKPROPAGATION → fills .grad on all params
optimizer.step() # 5. Gradient descent: w ← w − lr × ￿L/￿w
8

[page 9]
[7]: def train_epoch(model, loader, criterion, optimizer):
"""Run one full training epoch. Returns avg loss and accuracy."""
model.train()
total_loss, correct, total = 0.0, 0, 0
for X_batch, y_batch in loader:
# 1. Clear stale gradients
optimizer.zero_grad()
# 2. Forward pass
logits = model(X_batch)
# 3. Compute Cross-Entropy loss
# (applies log-softmax + NLL internally)
loss = criterion(logits, y_batch)
# 4. Backpropagation — chain rule through all layers
loss.backward()
# 5. Gradient descent step
optimizer.step()
total_loss += loss.item() * X_batch.size(0)
preds = logits.argmax(dim=1)
correct += (preds == y_batch).sum().item()
total += X_batch.size(0)
return total_loss / total, correct / total
@torch.no_grad()
def evaluate(model, loader, criterion):
"""Evaluate on a DataLoader. Returns avg loss and accuracy."""
model.eval()
total_loss, correct, total = 0.0, 0, 0
for X_batch, y_batch in loader:
logits = model(X_batch)
loss = criterion(logits, y_batch)
total_loss += loss.item() * X_batch.size(0)
correct += (logits.argmax(dim=1) == y_batch).sum().item()
total += X_batch.size(0)
return total_loss / total, correct / total
print("Training and evaluation functions defined!")
9

[page 10]
Training and evaluation functions defined!
1.7 Step 7: Train for 50 Epochs - Weight Updates Every 10 Epochs
[8]: # ￿￿ Hyperparameters ￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
LEARNING_RATE = 0.01
EPOCHS = 50
PRINT_EVERY = 10
# ￿￿ Fresh model ￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
model = MLP(n_input=16, n_hidden=8, n_output=n_classes).to(DEVICE)
criterion = nn.CrossEntropyLoss() # softmax + NLL combined
optimizer = optim.SGD(model.parameters(), lr=LEARNING_RATE)
# ￿￿ History ￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
history = {
'train_loss': [], 'test_loss': [],
'train_acc' : [], 'test_acc' : [],
'W1_norm' : [], 'W2_norm' : [],
'dW1_norm' : [], 'dW2_norm' : [],
}
print("=" * 78)
print(f" Training | lr= {LEARNING_RATE} | hidden=8 | epochs= {EPOCHS} | ␣
↪SGD | 7 classes ")
print("=" * 78)
hdr = (f"{'Epoch':>6} | {'Train Loss':>11} | {'Train Acc':>10} "
f"| {'Test Loss':>10} | {'Test Acc':>9} | {'|W1| mean':>9} | {'|W2|␣
↪mean':>9}")
print(hdr)
print("-" * len(hdr))
for epoch in range(1, EPOCHS + 1):
train_loss, train_acc = train_epoch(model, train_loader, criterion,␣
↪optimizer)
test_loss, test_acc = evaluate(model, test_loader, criterion)
# Capture fresh gradients for logging
model.train()
Xb, yb = next(iter(train_loader))
optimizer.zero_grad()
criterion(model(Xb), yb).backward()
W1 = model.hidden.weight.data
W2 = model.output.weight.data
10

[page 11]
dW1 = model.hidden.weight.grad
dW2 = model.output.weight.grad
history['train_loss'].append(train_loss)
history['test_loss'].append(test_loss)
history['train_acc'].append(train_acc)
history['test_acc'].append(test_acc)
history['W1_norm'].append(W1.abs().mean().item())
history['W2_norm'].append(W2.abs().mean().item())
history['dW1_norm'].append(dW1.abs().mean().item() if dW1 is not None else␣
↪0.0)
history['dW2_norm'].append(dW2.abs().mean().item() if dW2 is not None else␣
↪0.0)
if epoch % PRINT_EVERY == 0 or epoch == 1:
print(f"{epoch:>6} | {train_loss:>11.4f} | {train_acc*100:>9.2f}% "
f"| {test_loss:>10.4f} | {test_acc*100:>8.2f}% "
f"| {W1.abs().mean().item():>9.5f} | {W2.abs().mean().item():>9.
↪5f}")
print("-" * len(hdr))
print(f"\n Final Test Accuracy : {test_acc*100:.2f}%")
==============================================================================
Training | lr=0.01 | hidden=8 | epochs=50 | SGD | 7 classes
==============================================================================
Epoch | Train Loss | Train Acc | Test Loss | Test Acc | |W1| mean | |W2|
mean
--------------------------------------------------------------------------------
--
1 | 1.9776 | 12.84% | 1.9100 | 32.32% | 0.13275 |
0.18332
10 | 1.0669 | 60.84% | 1.0421 | 61.62% | 0.17138 |
0.23996
20 | 0.7383 | 73.01% | 0.7292 | 73.41% | 0.20581 |
0.29645
30 | 0.5788 | 79.56% | 0.5737 | 79.80% | 0.23191 |
0.34259
40 | 0.4880 | 84.42% | 0.4846 | 83.95% | 0.25237 |
0.37788
50 | 0.4255 | 87.39% | 0.4224 | 87.73% | 0.26871 |
0.40681
--------------------------------------------------------------------------------
--
Final Test Accuracy : 87.73%
11

[page 12]
1.8 Step 8: Detailed Weight & Gradient Snapshots Every 10 Epochs
[9]: model2 = MLP(n_input=16, n_hidden=8, n_output=n_classes).to(DEVICE)
optimizer2 = optim.SGD(model2.parameters(), lr=LEARNING_RATE)
criterion2 = nn.CrossEntropyLoss()
feature_names = df.columns[:-1].tolist()
print("=" * 72)
print(" Detailed Weight & Gradient Snapshots (every 10 epochs)")
print("=" * 72)
for epoch in range(1, EPOCHS + 1):
train_loss, train_acc = train_epoch(model2, train_loader, criterion2,␣
↪optimizer2)
# Capture gradients
model2.train()
Xb, yb = next(iter(train_loader))
optimizer2.zero_grad()
criterion2(model2(Xb), yb).backward()
if epoch % PRINT_EVERY == 0:
W1 = model2.hidden.weight.data.cpu().numpy() # (8, 16)
b1 = model2.hidden.bias.data.cpu().numpy() # (8,)
W2 = model2.output.weight.data.cpu().numpy() # (7, 8)
dW1 = model2.hidden.weight.grad.cpu().numpy() # (8, 16)
dW2 = model2.output.weight.grad.cpu().numpy() # (7, 8)
print(f"\n{'￿'*72}")
print(f" EPOCH {epoch:>3} | Train Loss = {train_loss:.4f} "
f" Train Acc = {train_acc*100:.2f}%")
print(f"{'￿'*72}")
print(" [W1] Input → Hidden (all 8 hidden neurons) ")
for h in range(8):
w_row = W1[h] # weights from all 16 inputs into hidden neuron h
print(f" H{h+1}: mean={w_row.mean():+.5f} "
f"std={w_row.std():.5f} "
f"min={w_row.min():+.5f} "
f"max={w_row.max():+.5f} "
f"bias={b1[h]:+.5f}")
print("\n [W2] Hidden → Output (per class) ")
for c in range(n_classes):
w_row = W2[c] # weights from all 8 hidden neurons into output␣
↪class c
print(f" {class_names[c]:>10}: mean={w_row.mean():+.5f} "
12

[page 13]
f"std={w_row.std():.5f} "
f"min={w_row.min():+.5f} "
f"max={w_row.max():+.5f}")
print(f"\n Gradient norms → "
f"|￿L/￿W1| = {np.abs(dW1).mean():.7f} "
f"|￿L/￿W2| = {np.abs(dW2).mean():.7f}")
print(f" (shrinking gradient norms = network converging)")
print(f"\n{'='*72}")
print(" Weight inspection complete!")
========================================================================
Detailed Weight & Gradient Snapshots (every 10 epochs)
========================================================================
￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
EPOCH 10 | Train Loss = 1.1412 Train Acc = 62.16%
￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
[W1] Input → Hidden (all 8 hidden neurons)
H1: mean=+0.05895 std=0.19822 min=-0.41661 max=+0.38914 bias=+0.21557
H2: mean=+0.02719 std=0.16399 min=-0.23806 max=+0.29928 bias=-0.11836
H3: mean=-0.10189 std=0.17364 min=-0.43402 max=+0.20775 bias=+0.14226
H4: mean=-0.07534 std=0.21212 min=-0.39319 max=+0.46038 bias=+0.08811
H5: mean=-0.02611 std=0.20184 min=-0.33069 max=+0.47172 bias=+0.40088
H6: mean=-0.00412 std=0.17565 min=-0.29402 max=+0.28482 bias=-0.05124
H7: mean=-0.03610 std=0.15900 min=-0.25272 max=+0.29069 bias=-0.01221
H8: mean=-0.04740 std=0.14972 min=-0.39178 max=+0.16347 bias=-0.18755
[W2] Hidden → Output (per class)
BARBUNYA: mean=+0.00644 std=0.23636 min=-0.30417 max=+0.30184
BOMBAY: mean=-0.03849 std=0.19515 min=-0.29265 max=+0.24814
CALI: mean=+0.01855 std=0.20499 min=-0.25163 max=+0.43722
DERMASON: mean=+0.03797 std=0.44000 min=-0.44071 max=+0.74470
HOROZ: mean=+0.19353 std=0.23464 min=-0.29468 max=+0.50568
SEKER: mean=-0.11571 std=0.24340 min=-0.39656 max=+0.44458
SIRA: mean=-0.05299 std=0.29272 min=-0.37439 max=+0.45633
Gradient norms → |￿L/￿W1| = 0.0205141 |￿L/￿W2| = 0.0305316
(shrinking gradient norms = network converging)
￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
EPOCH 20 | Train Loss = 0.6609 Train Acc = 79.71%
￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
[W1] Input → Hidden (all 8 hidden neurons)
H1: mean=+0.07960 std=0.26467 min=-0.58064 max=+0.49165 bias=+0.29824
H2: mean=+0.06660 std=0.25657 min=-0.47305 max=+0.38051 bias=-0.04702
H3: mean=-0.12602 std=0.22160 min=-0.56985 max=+0.24837 bias=+0.25884
13

[page 14]
H4: mean=-0.08707 std=0.27506 min=-0.51031 max=+0.58343 bias=+0.20281
H5: mean=-0.05240 std=0.25175 min=-0.40518 max=+0.57885 bias=+0.58100
H6: mean=-0.01308 std=0.20577 min=-0.30326 max=+0.34330 bias=-0.01033
H7: mean=-0.02208 std=0.19881 min=-0.33015 max=+0.33657 bias=-0.01750
H8: mean=-0.05883 std=0.17304 min=-0.53323 max=+0.15004 bias=-0.15547
[W2] Hidden → Output (per class)
BARBUNYA: mean=-0.00011 std=0.31641 min=-0.37865 max=+0.39528
BOMBAY: mean=-0.05888 std=0.29220 min=-0.46390 max=+0.43669
CALI: mean=+0.02116 std=0.24582 min=-0.32670 max=+0.48632
DERMASON: mean=+0.02325 std=0.54476 min=-0.56768 max=+0.86265
HOROZ: mean=+0.17448 std=0.33733 min=-0.46178 max=+0.62442
SEKER: mean=-0.08321 std=0.40610 min=-0.47047 max=+0.86075
SIRA: mean=-0.02739 std=0.37475 min=-0.40522 max=+0.70136
Gradient norms → |￿L/￿W1| = 0.0107901 |￿L/￿W2| = 0.0185927
(shrinking gradient norms = network converging)
￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
EPOCH 30 | Train Loss = 0.4972 Train Acc = 85.11%
￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
[W1] Input → Hidden (all 8 hidden neurons)
H1: mean=+0.09717 std=0.30733 min=-0.68732 max=+0.54709 bias=+0.34241
H2: mean=+0.08831 std=0.29536 min=-0.55553 max=+0.41223 bias=-0.03204
H3: mean=-0.14127 std=0.24846 min=-0.64371 max=+0.29643 bias=+0.36233
H4: mean=-0.09544 std=0.31319 min=-0.56901 max=+0.67076 bias=+0.21951
H5: mean=-0.06932 std=0.27491 min=-0.44562 max=+0.60877 bias=+0.73101
H6: mean=-0.01594 std=0.23102 min=-0.34765 max=+0.38775 bias=+0.02160
H7: mean=-0.00940 std=0.23232 min=-0.38446 max=+0.37153 bias=-0.03316
H8: mean=-0.06978 std=0.19994 min=-0.66041 max=+0.13606 bias=-0.11923
[W2] Hidden → Output (per class)
BARBUNYA: mean=+0.00575 std=0.39674 min=-0.48534 max=+0.51507
BOMBAY: mean=-0.07235 std=0.35085 min=-0.58269 max=+0.52234
CALI: mean=+0.02119 std=0.29822 min=-0.41965 max=+0.53479
DERMASON: mean=+0.01969 std=0.60182 min=-0.62557 max=+0.99636
HOROZ: mean=+0.17137 std=0.40033 min=-0.57574 max=+0.69264
SEKER: mean=-0.08266 std=0.46292 min=-0.51748 max=+0.98134
SIRA: mean=-0.01369 std=0.42329 min=-0.43355 max=+0.86118
Gradient norms → |￿L/￿W1| = 0.0103913 |￿L/￿W2| = 0.0153864
(shrinking gradient norms = network converging)
￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
EPOCH 40 | Train Loss = 0.4126 Train Acc = 87.69%
￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
[W1] Input → Hidden (all 8 hidden neurons)
H1: mean=+0.10983 std=0.33503 min=-0.75663 max=+0.58081 bias=+0.37120
14

[page 15]
H2: mean=+0.10102 std=0.31940 min=-0.60543 max=+0.44900 bias=-0.02541
H3: mean=-0.15109 std=0.26383 min=-0.68487 max=+0.32094 bias=+0.44058
H4: mean=-0.10214 std=0.34170 min=-0.60759 max=+0.75148 bias=+0.19359
H5: mean=-0.08200 std=0.28549 min=-0.47058 max=+0.60933 bias=+0.85097
H6: mean=-0.01753 std=0.25023 min=-0.40798 max=+0.42970 bias=+0.03725
H7: mean=+0.00139 std=0.26006 min=-0.42295 max=+0.40018 bias=-0.04768
H8: mean=-0.07801 std=0.22481 min=-0.76452 max=+0.16716 bias=-0.09172
[W2] Hidden → Output (per class)
BARBUNYA: mean=+0.01289 std=0.46021 min=-0.58167 max=+0.62466
BOMBAY: mean=-0.08027 std=0.38995 min=-0.65965 max=+0.57580
CALI: mean=+0.01639 std=0.33932 min=-0.49431 max=+0.56855
DERMASON: mean=+0.02094 std=0.64287 min=-0.65484 max=+1.10657
HOROZ: mean=+0.17191 std=0.44318 min=-0.65784 max=+0.75614
SEKER: mean=-0.08141 std=0.49482 min=-0.54348 max=+1.04426
SIRA: mean=-0.01114 std=0.45243 min=-0.45654 max=+0.96723
Gradient norms → |￿L/￿W1| = 0.0062237 |￿L/￿W2| = 0.0105421
(shrinking gradient norms = network converging)
￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
EPOCH 50 | Train Loss = 0.3625 Train Acc = 89.09%
￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
[W1] Input → Hidden (all 8 hidden neurons)
H1: mean=+0.11911 std=0.35425 min=-0.80414 max=+0.61925 bias=+0.39210
H2: mean=+0.10997 std=0.33619 min=-0.63941 max=+0.47479 bias=-0.02036
H3: mean=-0.15790 std=0.27319 min=-0.70931 max=+0.33377 bias=+0.50226
H4: mean=-0.10693 std=0.36515 min=-0.63570 max=+0.82640 bias=+0.14866
H5: mean=-0.09202 std=0.28971 min=-0.48848 max=+0.59666 bias=+0.95173
H6: mean=-0.01900 std=0.26490 min=-0.45314 max=+0.46241 bias=+0.04181
H7: mean=+0.00954 std=0.28134 min=-0.47107 max=+0.42175 bias=-0.05998
H8: mean=-0.08319 std=0.24555 min=-0.84551 max=+0.21742 bias=-0.07282
[W2] Hidden → Output (per class)
BARBUNYA: mean=+0.01860 std=0.50777 min=-0.65398 max=+0.71093
BOMBAY: mean=-0.08653 std=0.41744 min=-0.71556 max=+0.61031
CALI: mean=+0.01093 std=0.36995 min=-0.54333 max=+0.59379
DERMASON: mean=+0.02398 std=0.67425 min=-0.67017 max=+1.19733
HOROZ: mean=+0.17345 std=0.47446 min=-0.71995 max=+0.80695
SEKER: mean=-0.07815 std=0.51668 min=-0.55820 max=+1.08821
SIRA: mean=-0.01300 std=0.47441 min=-0.47499 max=+1.05047
Gradient norms → |￿L/￿W1| = 0.0062122 |￿L/￿W2| = 0.0093403
(shrinking gradient norms = network converging)
========================================================================
Weight inspection complete!
15

[page 16]
1.9 Step 9: Convergence Plots
[10]: epochs_x = np.arange(1, EPOCHS + 1)
check_epochs = list(range(10, EPOCHS + 1, 10))
check_idx = [e - 1 for e in check_epochs]
COLORS = {'train': '#e63946', 'test': '#457b9d',
'W1': '#2a9d8f', 'W2': '#e9c46a',
'dW1': '#f4a261', 'dW2': '#264653'}
fig = plt.figure(figsize=(16, 12))
gs = gridspec.GridSpec(2, 2, hspace=0.42, wspace=0.35)
# ￿￿ Plot 1: Loss ￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
ax1 = fig.add_subplot(gs[0, 0])
ax1.plot(epochs_x, history['train_loss'], color=COLORS['train'], lw=2.5,␣
↪label='Train Loss')
ax1.plot(epochs_x, history['test_loss'], color =COLORS['test'], lw =2.5,␣
↪label='Test Loss', ls='--')
ax1.scatter([e for e in check_epochs], [history['train_loss'][i] for i in␣
↪check_idx],
color=COLORS['train'], s=80, zorder=5)
ax1.scatter([e for e in check_epochs], [history['test_loss'][i] for i in␣
↪check_idx],
color=COLORS['test'], s =80, zorder=5, marker='D')
ax1.set_title('Loss Convergence (Cross-Entropy)', fontsize=13,␣
↪fontweight='bold')
ax1.set_xlabel('Epoch'); ax1.set_ylabel('Cross-Entropy Loss')
ax1.legend(); ax1.grid(True, alpha=0.3)
ax1.set_xticks(range(0, EPOCHS + 1, 5))
# ￿￿ Plot 2: Accuracy ￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
ax2 = fig.add_subplot(gs[0, 1])
ax2.plot(epochs_x, [a*100 for a in history['train_acc']],␣
↪color=COLORS['train'], lw=2.5, label='Train Acc')
ax2.plot(epochs_x, [a*100 for a in history['test_acc']], color =COLORS['test'],␣
↪ lw=2.5, label='Test Acc', ls='--')
ax2.scatter([e for e in check_epochs], [history['train_acc'][i]*100 for i in␣
↪check_idx],
color=COLORS['train'], s=80, zorder=5)
ax2.scatter([e for e in check_epochs], [history['test_acc'][i]*100 for i in␣
↪check_idx],
color=COLORS['test'], s =80, zorder=5, marker='D')
ax2.set_title('Accuracy Convergence (7 Classes)', fontsize=13,␣
↪fontweight='bold')
ax2.set_xlabel('Epoch'); ax2.set_ylabel('Accuracy (%)')
ax2.legend(); ax2.grid(True, alpha=0.3)
16

[page 17]
ax2.set_xticks(range(0, EPOCHS + 1, 5))
# ￿￿ Plot 3: Weight Norms ￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
ax3 = fig.add_subplot(gs[1, 0])
ax3.plot(epochs_x, history['W1_norm'], color=COLORS['W1'], lw=2.5, label='|W1|␣
↪mean (input→hidden) ')
ax3.plot(epochs_x, history['W2_norm'], color=COLORS['W2'], lw=2.5, label='|W2|␣
↪mean (hidden→output) ', ls='--')
for e in check_epochs:
ax3.axvline(e, color='gray', lw=0.8, ls=':', alpha=0.7)
ax3.text(e + 0.2, max(history['W1_norm']) * 0.97, f'e{e}', fontsize=7,␣
↪color='gray')
ax3.set_title('Mean Absolute Weight Values', fontsize=13, fontweight='bold')
ax3.set_xlabel('Epoch'); ax3.set_ylabel('Mean |w|')
ax3.legend(); ax3.grid(True, alpha=0.3)
ax3.set_xticks(range(0, EPOCHS + 1, 5))
# ￿￿ Plot 4: Gradient Norms ￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
ax4 = fig.add_subplot(gs[1, 1])
ax4.semilogy(epochs_x, history['dW1_norm'], color=COLORS['dW1'], lw=2.5,␣
↪label='|￿L/￿W1| hidden grads ')
ax4.semilogy(epochs_x, history['dW2_norm'], color=COLORS['dW2'], lw=2.5,␣
↪label='|￿L/￿W2| output grads ', ls='--')
for e in check_epochs:
ax4.axvline(e, color='gray', lw=0.8, ls=':', alpha=0.7)
ax4.set_title('Gradient Norms (log scale)', fontsize=13, fontweight='bold')
ax4.set_xlabel('Epoch'); ax4.set_ylabel('Mean |gradient| (log)')
ax4.legend(); ax4.grid(True, alpha=0.3)
ax4.set_xticks(range(0, EPOCHS + 1, 5))
fig.suptitle(
'PyTorch MLP — Backpropagation Convergence on Dry Bean Dataset\n'
'Architecture: 16 → 8 (ReLU) → 7 (Softmax) | SGD | lr=0.05 ',
fontsize=14, fontweight='bold', y=1.01
)
plt.savefig('convergence_plots.png', dpi=130, bbox_inches='tight')
plt.show()
17

[page 18]
1.10 Step 11: Step-by-Step Backprop Trace - One Sample
[11]: sample_idx = 5
x_s = X_train[sample_idx:sample_idx+1] # (1, 16)
y_s = y_train[sample_idx:sample_idx+1] # (1,) integer label
true_class = class_names[y_s.item()]
print("=" * 68)
print(f" STEP-BY-STEP BACKPROPAGATION TRACE (sample #{sample_idx})")
print(f" True class : {true_class} (index {y_s.item()})")
print("=" * 68)
# Extract weights as numpy arrays
W1 = model.hidden.weight.data.cpu().numpy() # (8, 16)
b1 = model.hidden.bias.data.cpu().numpy() # (8,)
W2 = model.output.weight.data.cpu().numpy() # (7, 8)
b2 = model.output.bias.data.cpu().numpy() # (7,)
18

[page 19]
x_np = x_s.cpu().numpy() # (1, 16)
y_idx = y_s.item() # integer: 0..6
# ￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
# FORWARD PASS
# ￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
print("\n￿￿￿￿ FORWARD PASS ￿￿￿￿")
Z1 = x_np @ W1.T + b1 # (1, 8)
A1 = np.maximum(0, Z1) # ReLU → (1, 8)
print(f"Z1 (pre-ReLU) : {Z1[0].round(4)}")
print(f"A1 (post-ReLU) : {A1[0].round(4)}")
print(f" ({np.sum(Z1[0] <= 0)} of 8 neurons are dead/zero after ReLU)")
Z2 = A1 @ W2.T + b2 # (1, 7) raw logits
exp_z = np.exp(Z2 - Z2.max(axis=1, keepdims=True)) # numerically stable
A2 = exp_z / exp_z.sum(axis=1, keepdims=True) # softmax → (1, 7)
print(f"\nZ2 (logits) : {Z2[0].round(4)}")
print(f"\nA2 (softmax probabilities):")
for cls, p in zip(class_names, A2[0]):
marker = ' ← predicted' if p == A2[0].max() else ''
marker += ' ← TRUE' if cls == true_class else ''
print(f" {cls:>10}: {p:.4f}{marker}")
ce_loss = -np.log(A2[0, y_idx] + 1e-9)
print(f"\nCross-Entropy loss = -log(p_true) = -log({A2[0, y_idx]:.4f}) =␣
↪{ce_loss:.4f}")
# ￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
# BACKWARD PASS — manual chain rule
# ￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
print("\n￿￿￿￿ BACKWARD PASS (manual chain rule) ￿￿￿￿")
# ￿² = Softmax(Z2) − one_hot(y)
# This clean form comes from d(CrossEntropy)/dZ2 = A2 - Y
delta2 = A2.copy() # start from softmax output
delta2[0, y_idx] -= 1.0 # subtract 1 at the true class position (1, 7)
print(f"￿² = A2 − one_hot(y) [output layer error]: ")
for cls, d in zip(class_names, delta2[0]):
marker = ' ← true class (subtract 1 here)' if cls == true_class else ''
print(f" {cls:>10}: {d:+.4f}{marker}")
dW2 = delta2.T @ A1 # (7, 8)
db2 = delta2 # (1, 7)
print(f"\n￿L/￿W2 shape={dW2.shape} norm={np.abs(dW2).mean():.6f}")
19

[page 20]
# Backprop through W2, then through ReLU
delta1 = (delta2 @ W2) * (Z1 > 0).astype(float) # (1, 8)
dW1 = delta1.T @ x_np # (8, 16)
db1 = delta1 # (1, 8)
print(f"\n￿¹ (hidden layer error via chain rule):")
print(f" {delta1[0].round(6)}")
print(f"\n￿L/￿W1 shape={dW1.shape} norm={np.abs(dW1).mean():.6f}")
====================================================================
STEP-BY-STEP BACKPROPAGATION TRACE (sample #5)
True class : SIRA (index 6)
====================================================================
￿￿￿￿ FORWARD PASS ￿￿￿￿
Z1 (pre-ReLU) : [ 1.9509 -0.5945 0.6318 -0.8384 0.3991 1.7679 0.9613
-1.3167]
A1 (post-ReLU) : [1.9509 0. 0.6318 0. 0.3991 1.7679 0.9613 0. ]
(3 of 8 neurons are dead/zero after ReLU)
Z2 (logits) : [-1.1965 -2.7817 -0.9519 3.1604 -1.0133 -0.421 3.2097]
A2 (softmax probabilities):
BARBUNYA: 0.0060
BOMBAY: 0.0012
CALI: 0.0077
DERMASON: 0.4704
HOROZ: 0.0072
SEKER: 0.0131
SIRA: 0.4942 ← predicted ← TRUE
Cross-Entropy loss = -log(p_true) = -log(0.4942) = 0.7047
￿￿￿￿ BACKWARD PASS (manual chain rule) ￿￿￿￿
￿² = A2 − one_hot(y) [output layer error]:
BARBUNYA: +0.0060
BOMBAY: +0.0012
CALI: +0.0077
DERMASON: +0.4704
HOROZ: +0.0072
SEKER: +0.0131
SIRA: -0.5058 ← true class (subtract 1 here)
￿L/￿W2 shape=(7, 8) norm=0.103155
￿¹ (hidden layer error via chain rule):
[ 0.16589 -0. -0.469274 0. -0.137489 0.043978 0.493364
0. ]
20

[page 21]
￿L/￿W1 shape=(8, 16) norm=0.056394
1.11 Step 12: Final Evaluation — Per-Class Accuracy
[12]: model.eval()
with torch.no_grad():
all_preds = model(X_test).argmax(dim=1).cpu().numpy()
all_true = y_test.cpu().numpy()
final_acc = np.mean(all_preds == all_true)
print("=" * 55)
print(" Final Evaluation on Test Set")
print("=" * 55)
print(f" Overall Accuracy : {final_acc * 100:.2f}%\n")
print(f" {'Class':>10} | {'Correct':>7} | {'Total':>7} | {'Accuracy':>9}")
print(f" {'￿'*42}")
for i, cls in enumerate(class_names):
mask = all_true == i
n_total = mask.sum()
n_correct = (all_preds[mask] == i).sum()
acc_cls = n_correct / n_total if n_total > 0 else 0.0
print(f" {cls:>10} | {n_correct:>7} | {n_total:>7} | {acc_cls*100:>8.2f}%")
# Confusion matrix
print(f"\n Confusion Matrix (rows=true, cols=predicted):")
conf = np.zeros((n_classes, n_classes), dtype=int)
for t, p in zip(all_true, all_preds):
conf[t, p] += 1
header_str = f" {'':>10} " + " ".join(f"{c:>8}" for c in class_names)
print(header_str)
for i, cls in enumerate(class_names):
row = " ".join(f"{conf[i,j]:>8}" for j in range(n_classes))
print(f" {cls:>10} {row}")
print(f"\n{'='*55}")
print(" Architecture Summary")
print(f"{'='*55}")
print(f" Framework : PyTorch {torch.__version__}")
print(f" Input : 16 shape features")
print(f" Hidden : 8 neurons (ReLU) ")
print(f" Output : 7 neurons (Softmax via CrossEntropyLoss) ")
print(f" Params : {sum(p.numel() for p in model.parameters())}")
print(f" Optimizer : SGD | lr= {LEARNING_RATE}")
print(f" Epochs : {EPOCHS}")
21

[page 22]
print(f" Final Test Accuracy : {final_acc*100:.2f}%")
print(f"{'='*55}")
=======================================================
Final Evaluation on Test Set
=======================================================
Overall Accuracy : 87.73%
Class | Correct | Total | Accuracy
￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
BARBUNYA | 174 | 265 | 65.66%
BOMBAY | 100 | 104 | 96.15%
CALI | 304 | 326 | 93.25%
DERMASON | 671 | 709 | 94.64%
HOROZ | 366 | 386 | 94.82%
SEKER | 389 | 406 | 95.81%
SIRA | 385 | 527 | 73.06%
Confusion Matrix (rows=true, cols=predicted):
BARBUNYA BOMBAY CALI DERMASON HOROZ SEKER
SIRA
BARBUNYA 174 7 66 1 4 5
8
BOMBAY 0 100 4 0 0 0
0
CALI 12 0 304 0 8 2
0
DERMASON 0 0 0 671 0 12
26
HOROZ 0 0 6 4 366 0
10
SEKER 0 0 0 8 0 389
9
SIRA 0 0 5 100 21 16
385
=======================================================
Architecture Summary
=======================================================
Framework : PyTorch 2.10.0+cpu
Input : 16 shape features
Hidden : 8 neurons (ReLU)
Output : 7 neurons (Softmax via CrossEntropyLoss)
Params : 199
Optimizer : SGD | lr=0.01
Epochs : 50
Final Test Accuracy : 87.73%
=======================================================
22

[page 23]
1.12 Define MLP & Train with Full Weight History Recording
We record every weight, gradient, and delta at every epoch — giving us the full picture of
how backpropagation updates each individual connection.
[13]: class MLP(nn.Module):
def __init__(self, n_in=16, n_hidden=8, n_out=7):
super().__init__()
self.hidden = nn.Linear(n_in, n_hidden)
self.output = nn.Linear(n_hidden, n_out)
self.relu = nn.ReLU()
nn.init.zeros_(self.hidden.bias)
nn.init.zeros_(self.output.bias)
def forward(self, x):
return self.output(self.relu(self.hidden(x)))
EPOCHS = 50
LR = 0.05
model = MLP().to(DEVICE)
criterion = nn.CrossEntropyLoss()
optimizer = optim.SGD(model.parameters(), lr=LR)
# ￿￿ Storage: shape described at each key ￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
# W1_hist[epoch, neuron, feature] — (50, 8, 16)
# W2_hist[epoch, class, neuron ] — (50, 7, 8)
W1_hist = np.zeros((EPOCHS, 8, 16))
W2_hist = np.zeros((EPOCHS, 7, 8))
b1_hist = np.zeros((EPOCHS, 8))
b2_hist = np.zeros((EPOCHS, 7))
dW1_hist = np.zeros((EPOCHS, 8, 16)) # gradients
dW2_hist = np.zeros((EPOCHS, 7, 8))
delta_W1 = np.zeros((EPOCHS, 8, 16)) # actual weight change Δw = -lr * grad
delta_W2 = np.zeros((EPOCHS, 7, 8))
loss_hist = []
acc_hist_train = []
acc_hist_test = []
print('Training and recording every weight at every epoch...')
print(f'{"Epoch":>6} | {"Train Loss":>11} | {"Train Acc":>10} | {"Test Acc":
↪>9}')
print('-' * 46)
for epoch in range(EPOCHS):
# Save weights BEFORE update
W1_before = model.hidden.weight.data.cpu().numpy().copy()
23

[page 24]
# ￿￿ Training ￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
model.train()
ep_loss, correct, total = 0.0, 0, 0
for Xb, yb in train_loader:
optimizer.zero_grad()
loss = criterion(model(Xb), yb)
loss.backward()
optimizer.step()
ep_loss += loss.item() * Xb.size(0)
correct += (model(Xb).argmax(1) == yb).sum().item()
total += Xb.size(0)
# Capture gradients via one dedicated backward on full train set
model.train()
optimizer.zero_grad()
criterion(model(X_train), y_train).backward()
# ￿￿ Record everything ￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
W1_hist[epoch] = model.hidden.weight.data.cpu().numpy()
W2_hist[epoch] = model.output.weight.data.cpu().numpy()
b1_hist[epoch] = model.hidden.bias.data.cpu().numpy()
b2_hist[epoch] = model.output.bias.data.cpu().numpy()
dW1_hist[epoch] = model.hidden.weight.grad.cpu().numpy()
dW2_hist[epoch] = model.output.weight.grad.cpu().numpy()
delta_W1[epoch] = W1_hist[epoch] - W1_before # actual change this epoch
if epoch > 0:
delta_W2[epoch] = W2_hist[epoch] - W2_hist[epoch - 1]
# Test accuracy
model.eval()
with torch.no_grad():
test_acc = (model(X_test).argmax(1) == y_test).float().mean().item()
loss_hist.append(ep_loss / total)
acc_hist_train.append(correct / total)
acc_hist_test.append(test_acc)
if (epoch + 1) % 10 == 0 or epoch == 0:
print(f'{epoch+1:>6} | {ep_loss/total:>11.4f} | {correct/total*100:>9.
↪2f}% | {test_acc*100:>8.2f}%')
print(f'\nRecorded W1_hist shape: {W1_hist.shape} (epochs × neurons ×␣
↪features)')
print(f' Recorded W2_hist shape: {W2_hist.shape} (epochs × classes ×␣
↪neurons)')
24

[page 25]
Training and recording every weight at every epoch…
Epoch | Train Loss | Train Acc | Test Acc
----------------------------------------------
1 | 1.7456 | 40.46% | 60.78%
10 | 0.3764 | 89.63% | 89.79%
20 | 0.2536 | 91.87% | 91.55%
30 | 0.2269 | 92.40% | 91.92%
40 | 0.2170 | 92.47% | 92.21%
50 | 0.2115 | 92.55% | 92.21%
Recorded W1_hist shape: (50, 8, 16) (epochs × neurons × features)
Recorded W2_hist shape: (50, 7, 8) (epochs × classes × neurons)
1.13 Weight Trajectory — Every Connection in Hidden Layer
Each line = one weight connecting an input feature → hidden neuron .
One panel per hidden neuron (H1–H8). Shows how backprop gradually adjusts each weight.
[14]: epochs_x = np.arange(1, EPOCHS + 1)
FEAT_COLORS = ['#e6194b', '#3cb44b', '#0000ff', '#f58231', '#911eb4',␣
↪'#00ced1', '#aa6e28', '#000075', '#e63900', '#008000', '#9400d3', '#ff8c00',␣
↪'#dc143c', '#00868b', '#8b0000', '#006400']
feat_colors = FEAT_COLORS
fig, axes = plt.subplots(2, 4, figsize=(20, 9))
for h, ax in enumerate(axes.flat):
for f in range(16):
ax.plot(epochs_x, W1_hist[:, h, f],
color=feat_colors[f], lw=1.2, alpha=0.85,
label=feature_names[f])
ax.axhline(0, color='#555577', lw=0.7, ls='--')
ax.set_title(f'Hidden Neuron H{h+1}', fontsize=11, fontweight='bold',␣
↪color='#222222')
ax.set_xlabel('Epoch', fontsize=9)
ax.set_ylabel('Weight value', fontsize=9)
ax.grid(True)
# Mark epoch-10 checkpoints
for ck in range(10, EPOCHS + 1, 10):
ax.axvline(ck, color='#ff6b6b', lw=0.6, ls=':', alpha=0.6)
# Legend below the figure
handles = [plt.Line2D([0],[0], color=feat_colors[f], lw=2) for f in range(16)]
fig.legend(handles, feature_names,
loc='lower center', ncol=8, fontsize=8,
bbox_to_anchor=(0.5, -0.04), framealpha=0.3)
fig.suptitle('W1: Individual Weight Trajectories per Hidden Neuron\n'
25

[page 26]
'(Each line = one input feature → this neuron)',
fontsize=14, fontweight='bold', color='#222222', y=1.01)
plt.tight_layout()
plt.savefig('w1_weight_trajectories.png', dpi=130, bbox_inches='tight')
plt.show()
1.14 Gradient Magnitude per Neuron — How Hard Backprop is Pushing Each
Weight
[15]: fig, axes = plt.subplots(2, 4, figsize=(20, 9))
for h, ax in enumerate(axes.flat):
for f in range(16):
ax.plot(epochs_x, np.abs(dW1_hist[:, h, f]),
color=feat_colors[f], lw=1.2, alpha=0.85)
ax.set_title(f'|￿L/￿W1| — Neuron H{h+1}', fontsize=11, fontweight='bold',␣
↪color='#222222')
ax.set_xlabel('Epoch', fontsize=9)
ax.set_ylabel('|Gradient|', fontsize=9)
ax.grid(True)
for ck in range(10, EPOCHS + 1, 10):
ax.axvline(ck, color='#ff6b6b', lw=0.6, ls=':', alpha=0.6)
handles = [plt.Line2D([0],[0], color=feat_colors[f], lw=2) for f in range(16)]
fig.legend(handles, feature_names,
loc='lower center', ncol=8, fontsize=8,
bbox_to_anchor=(0.5, -0.04), framealpha=0.3)
26

[page 27]
fig.suptitle('W1: Gradient Magnitudes |￿L/￿w| per Hidden Neuron\n'
'(Decreasing gradients = converging)',
fontsize=14, fontweight='bold', color='#222222', y=1.01)
plt.tight_layout()
plt.savefig('w1_gradient_magnitudes.png', dpi=130, bbox_inches='tight')
plt.show()
1.15 Actual Weight Delta Δw per Neuron (What Backprop Actually Applies)
Δw = w_after − w_before = −lr × ￿L/￿w
This is the signed update — positive means the weight increased, negative means it decreased.
[16]: fig, axes = plt.subplots(2, 4, figsize=(20, 9))
for h, ax in enumerate(axes.flat):
for f in range(16):
ax.plot(epochs_x, delta_W1[:, h, f],
color=feat_colors[f], lw=1.0, alpha=0.85)
ax.axhline(0, color='#ffffff', lw=0.8, ls='--', alpha=0.5)
ax.set_title(f'Δw = −lr·￿L/￿w — Neuron H{h+1}', fontsize=10,
fontweight='bold', color='#222222')
ax.set_xlabel('Epoch', fontsize=9)
ax.set_ylabel('Weight change Δw', fontsize=9)
ax.grid(True)
for ck in range(10, EPOCHS + 1, 10):
ax.axvline(ck, color='#ff6b6b', lw=0.6, ls=':', alpha=0.6)
handles = [plt.Line2D([0],[0], color=feat_colors[f], lw=2) for f in range(16)]
fig.legend(handles, feature_names,
27

[page 28]
loc='lower center', ncol=8, fontsize=8,
bbox_to_anchor=(0.5, -0.04), framealpha=0.3)
fig.suptitle('W1: Weight Updates Δw Applied Each Epoch per Hidden Neuron\n'
'(+ve = weight increased, −ve = weight decreased)',
fontsize=14, fontweight='bold', color='#222222', y=1.01)
plt.tight_layout()
plt.savefig('w1_weight_deltas.png', dpi=130, bbox_inches='tight')
plt.show()
1.16 Output Layer W2 — Weight Trajectory per Class Neuron
Each panel = one output class neuron. Each line = its connection to one hidden neuron.
[17]: HIDDEN_COLORS = ['#e6194b', '#0000ff', '#3cb44b', '#ff8c00', '#9400d3',␣
↪'#00ced1', '#dc143c', '#228b22']
hidden_colors = HIDDEN_COLORS
fig, axes = plt.subplots(2, 4, figsize=(20, 9))
# 7 class panels + 1 combined summary
for c in range(7):
ax = axes.flat[c]
for h in range(8):
ax.plot(epochs_x, W2_hist[:, c, h],
color=hidden_colors[h], lw=1.8, alpha=0.9,
label=f'H{h+1}')
ax.axhline(0, color='#555577', lw=0.7, ls='--')
ax.set_title(f'Output: {class_names[c]}', fontsize=11,
28

[page 29]
fontweight='bold', color='#222222')
ax.set_xlabel('Epoch', fontsize=9)
ax.set_ylabel('Weight value', fontsize=9)
ax.grid(True)
for ck in range(10, EPOCHS + 1, 10):
ax.axvline(ck, color='#ff6b6b', lw=0.6, ls=':', alpha=0.6)
ax.legend(fontsize=7, ncol=4, loc='upper right')
# Panel 8: mean absolute W2 per class
ax8 = axes.flat[7]
CLASS_COLORS = ['#e6194b', '#0000ff', '#3cb44b', '#ff8c00', '#9400d3',␣
↪'#00ced1', '#dc143c']
class_palette = CLASS_COLORS
for c in range(7):
mean_abs = np.abs(W2_hist[:, c, :]).mean(axis=1)
ax8.plot(epochs_x, mean_abs, color=class_palette[c],
lw=2, label=class_names[c])
ax8.set_title('Mean |W2| per Output Class', fontsize=11,
fontweight='bold', color='#222222')
ax8.set_xlabel('Epoch', fontsize=9)
ax8.set_ylabel('Mean |weight|', fontsize=9)
ax8.legend(fontsize=8)
ax8.grid(True)
fig.suptitle('W2: Output Layer Weight Trajectories per Class\n'
'(Each line = hidden neuron → this class output)',
fontsize=14, fontweight='bold', color='#222222', y=1.01)
plt.tight_layout()
plt.savefig('w2_weight_trajectories.png', dpi=130, bbox_inches='tight')
plt.show()
29

[page 30]
1.17 W2 Gradient Magnitudes per Output Neuron
[18]: fig, axes = plt.subplots(2, 4, figsize=(20, 9))
for c in range(7):
ax = axes.flat[c]
for h in range(8):
ax.plot(epochs_x, np.abs(dW2_hist[:, c, h]),
color=hidden_colors[h], lw=1.8, alpha=0.9, label=f'H{h+1}')
ax.set_title(f'|￿L/￿W2| — {class_names[c]}', fontsize=11,
fontweight='bold', color='#222222')
ax.set_xlabel('Epoch', fontsize=9)
ax.set_ylabel('|Gradient|', fontsize=9)
ax.legend(fontsize=7, ncol=4, loc='upper right')
ax.grid(True)
for ck in range(10, EPOCHS + 1, 10):
ax.axvline(ck, color='#ff6b6b', lw=0.6, ls=':', alpha=0.6)
# Panel 8: all gradient norms overlaid
ax8 = axes.flat[7]
ax8.semilogy(epochs_x, np.abs(dW1_hist).mean(axis=(1,2)),
color='#a8dadc', lw=2.5, label='|￿L/￿W1| (hidden)')
ax8.semilogy(epochs_x, np.abs(dW2_hist).mean(axis=(1,2)),
color='#e63946', lw=2.5, ls='--', label='|￿L/￿W2| (output)')
ax8.set_title('Overall Gradient Norms (log)', fontsize=11,
fontweight='bold', color='#222222')
ax8.set_xlabel('Epoch', fontsize=9)
ax8.set_ylabel('Mean |gradient| (log)', fontsize=9)
ax8.legend(fontsize=9)
ax8.grid(True)
fig.suptitle('W2: Gradient Magnitudes |￿L/￿w| per Output Class\n'
'(Each line = gradient for the hidden→class connection)',
fontsize=14, fontweight='bold', color='#222222', y=1.01)
plt.tight_layout()
plt.savefig('w2_gradient_magnitudes.png', dpi=130, bbox_inches='tight')
plt.show()
30

[page 31]
W e record both W1 and W2 after every epoch.
[19]: torch.manual_seed(42)
np.random.seed(42)
DEVICE = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
# ￿￿ Load data ￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
df = pd.read_excel('Dry_Bean_Dataset.xlsx')
FEATURE_NAMES = df.columns[:-1].tolist()
X_raw = df.drop(columns=['Class']).values.astype(np.float32)
y_raw = df['Class'].values
le = LabelEncoder()
y_int = le.fit_transform(y_raw).astype(np.int64)
CLASS_NAMES = le.classes_
X_scaled = StandardScaler().fit_transform(X_raw).astype(np.float32)
X_tr, X_te, y_tr, y_te = train_test_split(
X_scaled, y_int, test_size=0.2, random_state=42, stratify=y_int
)
X_train = torch.tensor(X_tr).to(DEVICE)
y_train = torch.tensor(y_tr).to(DEVICE)
train_loader = DataLoader(TensorDataset(X_train, y_train), batch_size=256,␣
↪shuffle=True)
# ￿￿ Model ￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
class MLP(nn.Module):
def __init__(self):
super().__init__()
31

[page 32]
self.hidden = nn.Linear(16, 8)
self.output = nn.Linear(8, 7)
self.relu = nn.ReLU()
nn.init.zeros_(self.hidden.bias)
nn.init.zeros_(self.output.bias)
def forward(self, x):
return self.output(self.relu(self.hidden(x)))
model = MLP().to(DEVICE)
criterion = nn.CrossEntropyLoss()
optimizer = optim.SGD(model.parameters(), lr=0.05)
# ￿￿ Train 50 epochs — record W1 and W2 after every epoch ￿￿￿￿￿￿￿￿￿￿￿￿￿
EPOCHS = 50
# W1: PyTorch shape (8 neurons, 16 features) → W1_hist[epoch, neuron, feature]
# W2: PyTorch shape (7 classes, 8 neurons ) → W2_hist[epoch, class, neuron ]
W1_hist = np.zeros((EPOCHS, 8, 16))
W2_hist = np.zeros((EPOCHS, 7, 8))
loss_hist = []
for epoch in range(EPOCHS):
model.train()
ep_loss = 0.0
for Xb, yb in train_loader:
optimizer.zero_grad()
loss = criterion(model(Xb), yb)
loss.backward()
optimizer.step()
ep_loss += loss.item()
W1_hist[epoch] = model.hidden.weight.data.cpu().numpy() # (8, 16)
W2_hist[epoch] = model.output.weight.data.cpu().numpy() # (7, 8)
loss_hist.append(ep_loss / len(train_loader))
print('Training complete!')
print(f'W1_hist : {W1_hist.shape} (epochs x neurons x features)')
print(f'W2_hist : {W2_hist.shape} (epochs x classes x neurons)')
print(f'Final loss : {loss_hist[-1]:.4f}')
Training complete!
W1_hist : (50, 8, 16) (epochs x neurons x features)
W2_hist : (50, 7, 8) (epochs x classes x neurons)
Final loss : 0.2218
32

[page 33]
1.18 Choose Features & Classes to Inspect
[20]: # ￿￿ W1: 3 input features to inspect ￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
SELECTED_FEATURES = ['Area', 'Perimeter', 'Eccentricity']
FEAT_IDX = [FEATURE_NAMES.index(f) for f in SELECTED_FEATURES]
# ￿￿ W2: 3 output classes to inspect ￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿￿
SELECTED_CLASSES = ['BOMBAY', 'SEKER', 'DERMASON']
CLASS_IDX = [list(CLASS_NAMES).index(c) for c in SELECTED_CLASSES]
# ￿￿ Same high-contrast colour per hidden neuron in every plot ￿￿￿￿￿￿￿￿
NEURON_COLORS = [
'#e6194b', # H1 vivid red
'#0000ff', # H2 pure blue
'#3cb44b', # H3 vivid green
'#ff8c00', # H4 dark orange
'#9400d3', # H5 vivid purple
'#00ced1', # H6 dark turquoise
'#dc143c', # H7 crimson
'#228b22', # H8 forest green
]
epochs_x = np.arange(1, EPOCHS + 1)
print('W1 features :', SELECTED_FEATURES, ' indices:', FEAT_IDX)
print('W2 classes : ', SELECTED_CLASSES, ' indices:', CLASS_IDX)
W1 features : ['Area', 'Perimeter', 'Eccentricity'] indices: [0, 1, 5]
W2 classes : ['BOMBAY', 'SEKER', 'DERMASON'] indices: [1, 5, 3]
1.19 Helper — Reusable Plot Function
[21]: def plot_weight_grid(series_8, neuron_colors, epochs_x, ylabel_fn, suptitle,␣
↪savename):
"""
series_8 : list of 8 arrays shape (EPOCHS,) — one per hidden neuron
ylabel_fn : function of h (int) -> y-axis label string
"""
fig, axes = plt.subplots(2, 4, figsize=(18, 7), sharey=False)
fig.patch.set_facecolor('white')
for h, ax in enumerate(axes.flat):
w = series_8[h]
color = neuron_colors[h]
ax.fill_between(epochs_x, w, alpha=0.15, color=color)
ax.plot(epochs_x, w, color=color, lw=2.2, zorder=3)
33

[page 34]
# Dot + annotated value at every-10-epoch checkpoint
for ck in range(0, len(epochs_x), 10):
ax.scatter(ck + 1, w[ck], color=color, s=60,
zorder=5, edgecolors='black', linewidths=0.8)
ax.annotate(f'{w[ck]:.3f}',
xy=(ck + 1, w[ck]),
xytext=(0, 9), textcoords='offset points',
ha='center', fontsize=7.5, color='#111111',
bbox=dict(boxstyle='round,pad=0.15', fc='white',
ec=color, lw=0.8, alpha=0.85))
# Dashed init line, dotted final line
ax.axhline(w[0], color ='gray', lw =0.9, ls='--', alpha=0.6,
label=f'Init {w[0]:+.3f}')
ax.axhline(w[-1], color='#111111', lw=0.9, ls=':', alpha =0.85,
label=f'Final {w[-1]:+.3f}')
ax.set_title(f'H{h+1} Dw = {w[-1]-w[0]:+.3f}',
fontsize=11, fontweight='bold', color=color, pad=5)
ax.set_xlabel('Epoch', fontsize=9)
ax.set_ylabel(ylabel_fn(h), fontsize=8)
ax.legend(fontsize=7, loc='best')
ax.grid(True, alpha=0.35)
ax.set_xlim(1, len(epochs_x))
ax.xaxis.set_major_locator(ticker.MultipleLocator(10))
fig.suptitle(suptitle, fontsize=13, fontweight='bold', color='#111111')
plt.tight_layout()
plt.savefig(savename, dpi=130, bbox_inches='tight')
plt.show()
print('plot_weight_grid() ready')
plot_weight_grid() ready
1.20 W1 — Input → Hidden Layer
W1[neuron, feature] — weight connecting an input feature to a hidden neuron.
Backprop update rule: W1 -= lr * dL/dW1 where dL/dW1 = delta1.T @ X
One cell per feature. 8 panels per cell (one per hidden neuron).
34

[page 35]
1.20.1 W1 — F eature: Area
[22]: import matplotlib.ticker as ticker
fname = 'Area'
feat_idx = FEAT_IDX[0]
plot_weight_grid(
series_8 = [W1_hist[:, h, feat_idx] for h in range(8)],
neuron_colors= NEURON_COLORS,
epochs_x = epochs_x,
ylabel_fn = lambda h: f'w(Area -> H{h+1})',
suptitle = (
'W1 | Feature: "Area"\n'
'Each panel = one hidden neuron y = weight value after every epoch '
),
savename = 'W1_Area.png'
)
1.20.2 W1 — F eature: Perimeter
[23]: fname = 'Perimeter'
feat_idx = FEAT_IDX[1]
plot_weight_grid(
series_8 = [W1_hist[:, h, feat_idx] for h in range(8)],
neuron_colors= NEURON_COLORS,
epochs_x = epochs_x,
ylabel_fn = lambda h: f'w(Perimeter -> H{h+1})',
suptitle = (
'W1 | Feature: "Perimeter"\n'
'Each panel = one hidden neuron y = weight value after every epoch '
35

[page 36]
),
savename = 'W1_Perimeter.png'
)
1.20.3 W1 — F eature: Eccentricity
[24]: fname = 'Eccentricity'
feat_idx = FEAT_IDX[2]
plot_weight_grid(
series_8 = [W1_hist[:, h, feat_idx] for h in range(8)],
neuron_colors= NEURON_COLORS,
epochs_x = epochs_x,
ylabel_fn = lambda h: f'w(Eccentricity -> H{h+1})',
suptitle = (
'W1 | Feature: "Eccentricity"\n'
'Each panel = one hidden neuron y = weight value after every epoch '
),
savename = 'W1_Eccentricity.png'
)
36

[page 37]
1.21 W2 — Hidden → Output Layer
W2[class, neuron] — weight connecting a hidden neuron to an output class neuron.
Backprop update rule: W2 -= lr * dL/dW2 where dL/dW2 = delta2.T @ A1
One cell per class. 8 panels per cell (one per hidden neuron). Each panel shows how strongly that
hidden neuron ‘votes’ for this class over training.
1.21.1 W2 — Output Class: BOMBAY
[25]: cname = 'BOMBAY'
class_idx = CLASS_IDX[0]
plot_weight_grid(
series_8 = [W2_hist[:, class_idx, h] for h in range(8)],
neuron_colors= NEURON_COLORS,
epochs_x = epochs_x,
ylabel_fn = lambda h: f'w(H{h+1} -> BOMBAY)',
suptitle = (
'W2 | Output Class: "BOMBAY"\n'
'Each panel = one hidden neuron y = weight value after every epoch '
),
savename = 'W2_BOMBAY.png'
)
37

[page 38]
1.21.2 W2 — Output Class: SEKER
[26]: cname = 'SEKER'
class_idx = CLASS_IDX[1]
plot_weight_grid(
series_8 = [W2_hist[:, class_idx, h] for h in range(8)],
neuron_colors= NEURON_COLORS,
epochs_x = epochs_x,
ylabel_fn = lambda h: f'w(H{h+1} -> SEKER)',
suptitle = (
'W2 | Output Class: "SEKER"\n'
'Each panel = one hidden neuron y = weight value after every epoch '
),
savename = 'W2_SEKER.png'
)
38

[page 39]
1.21.3 W2 — Output Class: DERMASON
[27]: cname = 'DERMASON'
class_idx = CLASS_IDX[2]
plot_weight_grid(
series_8 = [W2_hist[:, class_idx, h] for h in range(8)],
neuron_colors= NEURON_COLORS,
epochs_x = epochs_x,
ylabel_fn = lambda h: f'w(H{h+1} -> DERMASON)',
suptitle = (
'W2 | Output Class: "DERMASON"\n'
'Each panel = one hidden neuron y = weight value after every epoch '
),
savename = 'W2_DERMASON.png'
)
1.22 One Neuron, Both Layers
Pick any hidden neuron and see: - T op row (W1): its weights from the 3 selected features -
Bottom row (W2): its weights to the 3 selected classes
Change NEURON_TO_INSPECT to 0–7 to explore different neurons.
[28]: NEURON_TO_INSPECT = 0 # change to 0-7
FEAT_COLORS_3 = ['#e6194b', '#0000ff', '#3cb44b']
CLASS_COLORS_3 = ['#ff8c00', '#9400d3', '#00ced1']
fig, axes = plt.subplots(2, 3, figsize=(18, 8))
fig.patch.set_facecolor('white')
39

[page 40]
# ￿￿ Top row: W1 — how this neuron receives from each feature ￿￿￿￿￿￿￿￿￿
for col, (fname, fidx, fc) in enumerate(
zip(SELECTED_FEATURES, FEAT_IDX, FEAT_COLORS_3)):
ax = axes[0, col]
w = W1_hist[:, NEURON_TO_INSPECT, fidx]
ax.fill_between(epochs_x, w, alpha=0.18, color=fc)
ax.plot(epochs_x, w, color=fc, lw=2.4)
ax.scatter(epochs_x, w, color=fc, s=20,
edgecolors='black', linewidths=0.4, zorder=4)
for ck in range(0, EPOCHS, 10):
ax.annotate(f'{w[ck]:.3f}', xy=(ck+1, w[ck]),
xytext=(0, 10), textcoords='offset points',
ha='center', fontsize=8.5, fontweight='bold',␣
↪color='#111111',
bbox=dict(boxstyle='round,pad=0.2', fc='white', ec=fc, lw=1.
↪1, alpha=0.9))
ax.axhline(w[0], color ='gray', lw =0.8, ls='--', alpha=0.6)
ax.axhline(w[-1], color='#111111', lw=0.8, ls=':', alpha =0.8)
ax.set_title(
f'W1: {fname} -> H{NEURON_TO_INSPECT+1}\n'
f'Dw = {w[-1]-w[0]:+.4f}',
fontsize=11, fontweight='bold', color=fc)
ax.set_xlabel('Epoch'); ax.grid(True, alpha=0.35)
ax.set_ylabel(f'w({fname} -> H{NEURON_TO_INSPECT+1})', fontsize=9)
ax.xaxis.set_major_locator(ticker.MultipleLocator(10))
# ￿￿ Bottom row: W2 — how this neuron votes for each class ￿￿￿￿￿￿￿￿￿￿￿￿
for col, (cname, cidx, cc) in enumerate(
zip(SELECTED_CLASSES, CLASS_IDX, CLASS_COLORS_3)):
ax = axes[1, col]
w = W2_hist[:, cidx, NEURON_TO_INSPECT]
ax.fill_between(epochs_x, w, alpha=0.18, color=cc)
ax.plot(epochs_x, w, color=cc, lw=2.4)
ax.scatter(epochs_x, w, color=cc, s=20,
edgecolors='black', linewidths=0.4, zorder=4)
for ck in range(0, EPOCHS, 10):
ax.annotate(f'{w[ck]:.3f}', xy=(ck+1, w[ck]),
xytext=(0, 10), textcoords='offset points',
ha='center', fontsize=8.5, fontweight='bold',␣
↪color='#111111',
bbox=dict(boxstyle='round,pad=0.2', fc='white', ec=cc, lw=1.
↪1, alpha=0.9))
ax.axhline(w[0], color ='gray', lw =0.8, ls='--', alpha=0.6)
ax.axhline(w[-1], color='#111111', lw=0.8, ls=':', alpha =0.8)
ax.set_title(
f'W2: H{NEURON_TO_INSPECT+1} -> {cname}\n'
40

[page 41]
f'Dw = {w[-1]-w[0]:+.4f}',
fontsize=11, fontweight='bold', color=cc)
ax.set_xlabel('Epoch'); ax.grid(True, alpha=0.35)
ax.set_ylabel(f'w(H{NEURON_TO_INSPECT+1} -> {cname})', fontsize=9)
ax.xaxis.set_major_locator(ticker.MultipleLocator(10))
fig.suptitle(
f'Hidden Neuron H{NEURON_TO_INSPECT+1} — Both Layers\n'
f'Top row: W1 (what it receives from features) | '
f'Bottom row: W2 (what it sends to output classes)',
fontsize=13, fontweight='bold', color='#111111'
)
plt.tight_layout()
plt.savefig(f'both_layers_H{NEURON_TO_INSPECT+1}.png', dpi=130,␣
↪bbox_inches='tight')
plt.show()
[28]:
41