# MNIST CNN PyTorch HandsOn (3)
course: Module 3 — Deep Learning & NLP
module: Module-3-Deep-Learning-NLP
type: notebook
source_url: https://personal-learn.armco.dev/files/Module-3-Deep-Learning-NLP/General/Lab_Materials-11-04-2026/MNIST_CNN_PyTorch_HandsOn_(3).ipynb
---
[cell 1 markdown]
# Question 1 — CNN From Scratch (PyTorch, MNIST)
## Objective
Build and optimize a **convolutional neural network (CNN)** to classify handwritten digits from the **MNIST** dataset using **PyTorch**.
Follow a hands-on, low-level approach where you write and understand each component.
Focus on tuning key hyperparameters such as the **learning rate**, **number of filters**, **kernel sizes**, **dropout**, **batch size**, and **activation functions**.
---
[cell 2 markdown]
## Problem Statement
The MNIST dataset consists of 70,000 images of handwritten digits (0–9). Your task is to build a **CNN** that can accurately classify these digits.
You will experiment with various hyperparameters, use **dropout** to prevent overfitting, and compare **activation functions** and **optimizers**.
### Dataset
- **Training Set**: 60,000 labeled images (28×28 grayscale).
- **Test Set**: 10,000 labeled images for evaluation.
- The dataset can be loaded via **torchvision.datasets.MNIST**.
[cell 3 markdown]
## Install & Imports
Below we install/verify required packages and import all dependencies.
If you're on Colab, PyTorch is usually pre-installed. We still show a fallback install cell for reproducibility.
[cell 4 code]
# (Optional) If running locally and you need to install:
# !pip install -q torch torchvision torchaudio matplotlib scikit-learn pandas
import os, math, time, random, json
import numpy as np
import torch
import torch.nn as nn
import torch.nn.functional as F
import torch.optim as optim
from torch.utils.data import DataLoader, random_split
from torchvision import datasets, transforms
import matplotlib.pyplot as plt
from sklearn.metrics import confusion_matrix, classification_report
import pandas as pd
device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
print('Using device:', device)
torch.manual_seed(42)
np.random.seed(42)
random.seed(42)
[cell 5 markdown]
## 1. Data Preprocessing (Text)
- **Normalization**: Scale pixel values to \([0,1]\) using `transforms.ToTensor()` (it already divides by 255).
- **Shape**: MNIST images are 1×28×28 (channel-first).
- **Train/Validation Split**: Reserve a part of the training set for validation (e.g., 54k/6k split).
- **Batching**: Use DataLoader with batch sizes like 32, 64, or 128.
- **Shuffling**: Shuffle the training set each epoch.
[cell 6 code]
# --- Data transforms ---
transform = transforms.Compose([
transforms.ToTensor(), # converts to [0,1] tensor of shape [1,28,28]
])
# --- Download & prepare datasets ---
root = './data'
train_full = datasets.MNIST(root=root, train=True, transform=transform, download=True)
test_ds = datasets.MNIST(root=root, train=False, transform=transform, download=True)
# --- Split train into train/val ---
val_size = 6000
train_size = len(train_full) - val_size
train_ds, val_ds = random_split(train_full, [train_size, val_size])
# --- Create DataLoaders ---
BATCH_SIZE = 64 # can be tuned later
train_loader = DataLoader(train_ds, batch_size=BATCH_SIZE, shuffle=True, num_workers=2, pin_memory=True)
val_loader = DataLoader(val_ds, batch_size=BATCH_SIZE, shuffle=False, num_workers=2, pin_memory=True)
test_loader = DataLoader(test_ds, batch_size=BATCH_SIZE, shuffle=False, num_workers=2, pin_memory=True)
print(f"Train: {len(train_ds)} samples | Val: {len(val_ds)} | Test: {len(test_ds)}")
[cell 7 markdown]
### (Optional) Visualize a few samples
[cell 8 code]
# Visualize a grid of random samples from the training dataset
def show_samples(dataset, n=10):
plt.figure(figsize=(10,2))
idxs = np.random.choice(len(dataset), size=n, replace=False)
for i, idx in enumerate(idxs):
if hasattr(dataset, 'dataset') and hasattr(dataset, 'indices'):
# random_split returns Subset objects; map index
img, label = dataset.dataset[dataset.indices[idx]]
else:
img, label = dataset[idx]
plt.subplot(1, n, i+1)
plt.imshow(img.squeeze(0), cmap='gray')
plt.title(str(label))
plt.axis('off')
plt.suptitle('Random MNIST samples')
plt.show()
show_samples(train_ds, n=10)
[cell 9 markdown]
## 2. Model Building (Text)
We define a standard CNN architecture:
- **Conv2d (1→C1, kernel=3, stride=1, padding=1)** → **ReLU** → **MaxPool(2×2)**
- **Conv2d (C1→C2, kernel=3, stride=1, padding=1)** → **ReLU** → **MaxPool(2×2)**
- **Flatten**
- **Dense/Linear** → **ReLU** → **Dropout(p)**
- **Dense/Linear** → **Softmax (inside CrossEntropyLoss)**
**Why padding=1 with kernel=3?** It preserves spatial size before pooling (28→28, then pooling halves to 14, then 7).
**Why ReLU?** It accelerates convergence and reduces vanishing gradients.
**Why Dropout?** It reduces overfitting by randomly dropping activations during training.
[cell 10 code]
class CNN(nn.Module):
def __init__(self, c1=16, c2=32, fc=128, dropout=0.2, activation='relu'):
super().__init__()
# --- Convolutional feature extractor ---
self.conv1 = nn.Conv2d(1, c1, kernel_size=3, stride=1, padding=1)
self.conv2 = nn.Conv2d(c1, c2, kernel_size=3, stride=1, padding=1)
self.pool = nn.MaxPool2d(kernel_size=2, stride=2)
# --- Activation function ---
if activation == 'relu':
self.act = nn.ReLU(inplace=True)
elif activation == 'leakyrelu':
self.act = nn.LeakyReLU(0.01, inplace=True)
elif activation == 'tanh':
self.act = nn.Tanh()
else:
raise ValueError("Unsupported activation. Choose from: 'relu', 'leakyrelu', 'tanh'")
# After two pools, 28 -> 14 -> 7; channels = c2
self.flatten_dim = c2 * 7 * 7
# --- Classifier head ---
self.fc1 = nn.Linear(self.flatten_dim, fc)
self.drop = nn.Dropout(p=dropout)
self.fc2 = nn.Linear(fc, 10)
def forward(self, x):
# Block 1
x = self.conv1(x)
x = self.act(x)
x = self.pool(x)
# Block 2
x = self.conv2(x)
x = self.act(x)
x = self.pool(x)
# Classifier
x = torch.flatten(x, 1)
x = self.fc1(x)
x = self.act(x)
x = self.drop(x)
x = self.fc2(x) # CrossEntropyLoss applies softmax internally
return x
def count_parameters(model):
return sum(p.numel() for p in model.parameters() if p.requires_grad)
# Instantiate a baseline model
model = CNN(c1=16, c2=32, fc=128, dropout=0.2, activation='relu').to(device)
print(model)
print('Trainable parameters:', count_parameters(model))
[cell 11 markdown]
## 3. Training Utilities (Text)
**Loss**: `CrossEntropyLoss` (includes `LogSoftmax`).
**Optimizers**: Try **Adam** or **SGD** with momentum.
**Scheduler**: Reduce LR on plateau for stability.
**Early Stopping**: Stop when validation loss doesn't improve for several epochs to avoid overfitting.
[cell 12 code]
class EarlyStopping:
"""Simple early-stopping based on validation loss."""
def __init__(self, patience=5, min_delta=0.0):
self.patience = patience
self.min_delta = min_delta
self.counter = 0
self.best = None
self.should_stop = False
def step(self, val):
if self.best is None or (self.best - val) > self.min_delta:
self.best = val
self.counter = 0
else:
self.counter += 1
if self.counter >= self.patience:
self.should_stop = True
def accuracy_from_logits(logits, y):
preds = torch.argmax(logits, dim=1)
return (preds == y).float().mean().item()
def train_one_epoch(model, loader, criterion, optimizer):
model.train()
total_loss, total_acc, n = 0.0, 0.0, 0
for x, y in loader:
x, y = x.to(device), y.to(device)
optimizer.zero_grad()
logits = model(x)
loss = criterion(logits, y)
loss.backward()
optimizer.step()
total_loss += loss.item() * x.size(0)
total_acc += (logits.argmax(1) == y).sum().item()
n += x.size(0)
return total_loss / n, total_acc / n
@torch.no_grad()
def evaluate(model, loader, criterion):
model.eval()
total_loss, total_acc, n = 0.0, 0.0, 0
for x, y in loader:
x, y = x.to(device), y.to(device)
logits = model(x)
loss = criterion(logits, y)
total_loss += loss.item() * x.size(0)
total_acc += (logits.argmax(1) == y).sum().item()
n += x.size(0)
return total_loss / n, total_acc / n
[cell 13 markdown]
## 4. Baseline Training (Text)
We train a baseline CNN with the following hyperparameters:
- Filters: (16, 32), FC: 128
- Dropout: 0.2
- Optimizer: Adam (lr=0.001)
- Batch size: 64
- Epochs: 10–15 with **early stopping** and **LR reduction on plateau**
[cell 14 code]
# --- Baseline hyperparameters ---
EPOCHS = 12
LEARNING_RATE = 1e-3
PATIENCE = 4
model = CNN(c1=16, c2=32, fc=128, dropout=0.2, activation='relu').to(device)
criterion = nn.CrossEntropyLoss()
optimizer = optim.Adam(model.parameters(), lr=LEARNING_RATE)
scheduler = optim.lr_scheduler.ReduceLROnPlateau(optimizer, mode='min', factor=0.5, patience=2, min_lr=1e-6, verbose=True)
early = EarlyStopping(patience=PATIENCE, min_delta=1e-4)
history = {'train_loss': [], 'train_acc': [], 'val_loss': [], 'val_acc': []}
for epoch in range(1, EPOCHS+1):
tr_loss, tr_acc = train_one_epoch(model, train_loader, criterion, optimizer)
va_loss, va_acc = evaluate(model, val_loader, criterion)
scheduler.step(va_loss)
early.step(va_loss)
history['train_loss'].append(tr_loss)
history['train_acc'].append(tr_acc)
history['val_loss'].append(va_loss)
history['val_acc'].append(va_acc)
print(f"Epoch {epoch:02d} | Train Loss: {tr_loss:.4f} Acc: {tr_acc:.4f} | Val Loss: {va_loss:.4f} Acc: {va_acc:.4f}")
if early.should_stop:
print('Early stopping triggered.')
break
# --- Plot learning curves ---
plt.figure(figsize=(10,4))
plt.plot(history['train_loss'], label='Train Loss')
plt.plot(history['val_loss'], label='Val Loss')
plt.xlabel('Epoch'); plt.ylabel('Loss'); plt.title('Loss Curves'); plt.legend(); plt.show()
plt.figure(figsize=(10,4))
plt.plot(history['train_acc'], label='Train Acc')
plt.plot(history['val_acc'], label='Val Acc')
plt.xlabel('Epoch'); plt.ylabel('Accuracy'); plt.title('Accuracy Curves'); plt.legend(); plt.show()
[cell 15 markdown]
## 5. Evaluation on Test Set (Text)
We now evaluate the trained model on the **test set**, compute the **confusion matrix**, and print the **classification report**.
[cell 16 code]
# --- Evaluate on test set ---
test_loss, test_acc = evaluate(model, test_loader, criterion)
print(f"Test Loss: {test_loss:.4f} | Test Acc: {test_acc:.4f}")
# --- Predictions for confusion matrix ---
model.eval()
y_true, y_pred = [], []
with torch.no_grad():
for x, y in test_loader:
x = x.to(device)
logits = model(x)
preds = logits.argmax(1).cpu().numpy().tolist()
y_pred.extend(preds)
y_true.extend(y.numpy().tolist())
cm = confusion_matrix(y_true, y_pred)
print("Classification Report:")
print(classification_report(y_true, y_pred, digits=4))
plt.figure(figsize=(6,6))
plt.imshow(cm)
plt.title('Confusion Matrix')
plt.xlabel('Predicted'); plt.ylabel('True')
plt.colorbar()
for i in range(10):
for j in range(10):
plt.text(j, i, cm[i,j], ha='center', va='center')
plt.show()
[cell 17 markdown]
## 6. Hyperparameter Tuning (Text)
Try several configurations and compare **validation accuracy** and **test accuracy**. You can vary:
- Convolution filters `(c1, c2)`
- Fully-connected width `fc`
- **Dropout** rate
- **Activation**: ReLU vs LeakyReLU vs Tanh
- **Optimizer**: Adam vs SGD with momentum
- **Learning rate** and **batch size**
> **Tip:** Keep epochs modest (e.g., 8–12) when sweeping multiple configs.
[cell 18 code]
def run_experiment(cfg, train_loader, val_loader, test_loader, max_epochs=10):
"""Run one training experiment with a config dict; return summary dict."""
model = CNN(c1=cfg['c1'], c2=cfg['c2'], fc=cfg['fc'], dropout=cfg['dropout'], activation=cfg['activation']).to(device)
criterion = nn.CrossEntropyLoss()
if cfg['optimizer'] == 'adam':
optimizer = optim.Adam(model.parameters(), lr=cfg['lr'])
else:
optimizer = optim.SGD(model.parameters(), lr=cfg['lr'], momentum=0.9, nesterov=True)
scheduler = optim.lr_scheduler.ReduceLROnPlateau(optimizer, mode='min', factor=0.5, patience=2, min_lr=1e-6, verbose=False)
early = EarlyStopping(patience=3, min_delta=1e-4)
hist = {'val_acc': []}
best_state = None
best_val = -1.0
for epoch in range(1, max_epochs+1):
tr_loss, tr_acc = train_one_epoch(model, train_loader, criterion, optimizer)
va_loss, va_acc = evaluate(model, val_loader, criterion)
scheduler.step(va_loss)
early.step(va_loss)
hist['val_acc'].append(va_acc)
# Keep best model by val_acc
if va_acc > best_val:
best_val = va_acc
best_state = {k: v.cpu().clone() for k, v in model.state_dict().items()}
# print(f"[{cfg['name']}] Epoch {epoch:02d}: train_acc={tr_acc:.4f}, val_acc={va_acc:.4f}")
if early.should_stop:
break
# Load best weights before test
model.load_state_dict(best_state)
test_loss, test_acc = evaluate(model, test_loader, criterion)
return {
'name': cfg['name'],
'config': cfg,
'val_acc_curve': hist['val_acc'],
'best_val_acc': best_val,
'test_acc': test_acc
}
# --- Example experiment grid (keep small for demo speed) ---
CONFIGS = [
{'name':'A_relu_16-32_fc128_do0.2_adam_1e-3', 'c1':16, 'c2':32, 'fc':128, 'dropout':0.2, 'activation':'relu', 'optimizer':'adam', 'lr':1e-3},
{'name':'B_relu_32-64_fc128_do0.2_adam_1e-3', 'c1':32, 'c2':64, 'fc':128, 'dropout':0.2, 'activation':'relu', 'optimizer':'adam', 'lr':1e-3},
{'name':'C_lrelu_16-32_fc256_do0.3_sgd_1e-2', 'c1':16, 'c2':32, 'fc':256, 'dropout':0.3, 'activation':'leakyrelu', 'optimizer':'sgd', 'lr':1e-2},
{'name':'D_tanh_32-64_fc128_do0.2_adam_1e-3', 'c1':32, 'c2':64, 'fc':128, 'dropout':0.2, 'activation':'tanh', 'optimizer':'adam', 'lr':1e-3},
]
# Optionally change batch size for experiments
def make_loader(ds, batch):
return DataLoader(ds, batch_size=batch, shuffle=True, num_workers=2, pin_memory=True)
EXP_BATCH = 64
train_loader_exp = make_loader(train_ds, EXP_BATCH)
val_loader_exp = make_loader(val_ds, EXP_BATCH)
test_loader_exp = make_loader(test_ds, EXP_BATCH)
results = []
for cfg in CONFIGS:
print(f"Running: {cfg['name']}")
res = run_experiment(cfg, train_loader_exp, val_loader_exp, test_loader_exp, max_epochs=10)
print(f" -> best_val_acc={res['best_val_acc']:.4f}, test_acc={res['test_acc']:.4f}")
results.append(res)
# --- Compare results ---
summary = pd.DataFrame([
{'Name': r['name'], 'Best Val Acc': r['best_val_acc'], 'Test Acc': r['test_acc']}
for r in results
]).sort_values('Test Acc', ascending=False).reset_index(drop=True)
summary
[cell 19 markdown]
### Compare Validation Accuracy Curves
[cell 20 code]
plt.figure(figsize=(12,5))
for r in results:
plt.plot(r['val_acc_curve'], label=r['name'])
plt.xlabel('Epoch'); plt.ylabel('Val Accuracy')
plt.title('Validation Accuracy by Configuration')
plt.legend()
plt.show()
[cell 21 markdown]
## 7. Save Best Model (Text)
You can save the **best-performing** model for later inference or deployment.
[cell 22 code]
# Save the best model among experiments
best_idx = int(np.argmax([r['test_acc'] for r in results])) if len(results) else None
if best_idx is not None:
best_name = results[best_idx]['name']
print('Best experiment:', best_name, 'with Test Acc =', results[best_idx]['test_acc'])
# For brevity, we save the baseline-trained model above:
torch.save(model.state_dict(), 'mnist_cnn_baseline.pt')
print('Saved baseline model to mnist_cnn_baseline.pt')
else:
print('No experiments were run; skipping save.')
[cell 23 markdown]
## 8. Discussion & Extensions (Text)
- **Padding & Stride**: With `kernel=3, padding=1`, the feature map size is preserved before pooling.
Pooling halves spatial dimensions (28→14→7). Try `stride=2` in convolution to be more aggressive.
- **Dropout**: Use 0.2–0.5 to regularize; too high may hurt accuracy.
- **Activations**: ReLU is a strong default; LeakyReLU helps avoid dead neurons; Tanh is smoother but slower.
- **Optimizers**: Adam usually converges faster; SGD with momentum can generalize well with tuned LR.
- **BatchNorm**: You can add after each convolution to stabilize training.
- **Early Stopping**: Already included via a simple callback.
- **Transfer to Fashion-MNIST or CIFAR-10**: For CIFAR-10, change input channels to 3 and consider deeper nets.
**Expected Outcomes**
With a tuned 2-Conv CNN, **98–99%** test accuracy on MNIST is achievable.