# Introduction to Deep Learning

course: Module 3 — Deep Learning & NLP
module: Module-3-Deep-Learning-NLP
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-3-Deep-Learning-NLP/General/Introduction_to_Deep_Learning.pdf
pages: 43

---
[page 1]
Introduction to Deep 
Learning

[page 2]
Agenda
● Introduction to Deep Learning
● Introduction to Artificial Neural Network
● Terminologies
● Artificial Neural Network
● Gradient Descent
● Activation functions
● Hyper parameter Tuning

[page 3]
• Usually the phrase is used to refer to deep neural 
networks, which are built on artificial networks loosely 
based on how the neurons in your brain work. 
• But unlike real neurons, which can connect with each 
other all across your brain, the nodes (or “neurons”) in 
these networks communicate across specific layers. 
• There’s an input layer that takes in information, an output 
layer that gives the response, and one or more “hidden” 
layers where the learning takes place. 
• Deep refers to the number of layers in a neural network. 
Deep Learning

[page 4]
How deep learning differs from machine learning (in general)
Simple way to remember it:
Machine learning → You help the model a lot
Deep learning → The model figures things out on its own
Deep learning is a subset of machine learning — not a replacement.
“Deep” in deep learning means learning happens through many layers, each building 
more abstract understanding from the one before.
Deep Learning
Machine Learning Deep Learning
Human designs features Model learns features itself
Works well on structured data Excels at images, audio, text
Needs more manual tuning Needs more data & compute

[page 5]
• Deep refers to the number of layers in a neural network. 
• Deep model has many layers stacked on top of each other
Deep Learning

[page 6]
Intuition: learning in stages
Think of it like a factory line for understanding data.
Example in case of images:
Early layers → detect edges and simple shapes
Middle layers → combine edges into parts (eyes, wheels)
Deep layers → recognize objects (faces, cars)
So the “depth” lets the model build understanding step by step.
Deep Learning

[page 7]
Why depth matters
Depth allows neural networks to:
• Learn complex patterns
• Represent hierarchies (simple → complex)
• Handle tasks like vision, speech, and language really well
That’s why modern systems for translation, image recognition, and chat are deep.
One-line definition
Deep learning = neural networks with many layers that learn hierarchical representations of data.
Deep Learning

[page 8]
A math-free analogy: Learning like a human
Imagine teaching someone to recognize a cat.
Shallow learning (not deep)
You give them rules:
• “If it has pointy ears”
• “If it has whiskers”
• “If it says meow”
This works… until the cat is fluffy, blurry, sideways, or cartoonish.
Deep Learning

[page 9]
Shallow learning (not deep)
You define rules: “If it has pointy ears”, “If it has whiskers”, “If it says meow” >>> This works… until the cat is fluffy,
blurry, sideways, or cartoonish.
Rule-based or shallow approaches only work in simple, familiar cases. They break when the situation changes 
even a little.
“Fluffy”
A rule like “cats have sharp outlines” fails when: long fur hides edges, the shape looks messy or uneven
So the model gets confused.
“Blurry”
If the image is: low resolution, motion-blurred, out of focus. Then rules like “cats have pointy ears” stop being 
reliable because the ears aren’t clearly visible. >>> So the model gets confused.
“Sideways”
If the cat is: lying down , rotated, viewed from a strange angle. Then rules based on position (“ears on top”, “tail 
at the back”) no longer hold. >>> So the model gets confused.
“Cartoonish”
If the cat is: a drawing, an emoji, a stylized animation, Then it breaks rules based on real -world appearance, 
even though humans still instantly know it’s a cat. >>> So the model gets confused.
Deep Learning

[page 10]
The core idea : Hand-made rules are brittle — they fail when reality doesn’t look exactly as expected.
Deep learning works better because it:
• doesn’t rely on fixed rules
• learns patterns that survive variation
• The model understands the overall concept of a cat, not just a checklist of visible parts or not just cat 
features
Deep learning
Instead, you let them look at thousands of cats and learn in stages:
• First, they notice lines and edges
• Then shapes and textures
• Then parts (ears, eyes, tails)
• Finally: “That’s a cat”
No rules given.
They build understanding layer by layer.
That layered learning is what “deep” means.
Simple methods work only for clean, ideal cases — deep learning handles the messy, real world.
Deep Learning

[page 11]
• Deep refers to the number of layers in a neural network. 
• Deep model has many layers stacked on top of each other
Deep Learning

[page 12]
Why deep learning needs so much data ?
Core idea : Deep models have many layers, and each layer must learn its part of reality.
More layers = more things to learn = more examples needed.
Analogy: learning to drive
Shallow learning
You memorize rules: “Stop at red lights”  and “Stay in your lane”
You don’t need many examples.
Deep learning
You learn by experience:
rain
night driving
traffic
weird intersections
bad drivers
To generalize, you need lots of real situations. Deep models learn like this — from experience, not rules.
Deep Learning

[page 13]
Why deep learning needs so much data ?
Data must cover:
• lighting changes
• angles
• sizes
• noise
• backgrounds
• styles (photos, drawings, emojis)
Each layer learns invariances (“this is still a cat even if…”), and that only comes from many examples.
Deep learning needs lots of data to see enough variations to learn stable concepts
Deep Learning

[page 14]
Why shallow models fail on Complex data , Example - Images
What images actually are?
An image is:
• millions of pixels
• each pixel just a number
• no built-in meaning
A shallow model sees: “Here are a lot of numbers.”
It does not see: “Oh, that’s an ear.”
Deep networks:
learn edges >>> combine edges into shapes >>> combine shapes into parts >>> combine parts into objects
Each layer solves a simpler sub-problem.
Shallow models try to do everything at once — and fail.
Deep learning needs large datasets because it learns complex, hierarchical representations directly from 
data, while shallow models fail on images because they lack the depth needed to build abstraction from 
raw pixels.
Deep Learning

[page 15]
Introduction to ANN
Artificial neural networks (ANNs) are biologically inspired computational networks.
Artificial Neural networks have been applied in diverse fields including:
• aerospace,
• automotive,
• banking, defense,
• electronics,
• entertainment,
• financial,
• insurance,
• manufacturing,
• medical,
• oil and gas,
• speech,
• securities,
• telecommunications,
• transportation,
• and environment
• …….

[page 16]
Applications of ANN
Prediction
Signal 
Processing
Image 
Processing
Pattern 
Recognition
Bioinformatics
ANN

[page 17]
Biological Basis of Artificial Neural Networks
Artificial neural networks are a technology
based on studies of the brain and nervous
system.
These networks emulate a biological neural
network
Specifically, ANN models simulate the
electrical activity of the brain and nervous
system.
Processing elements also known as either
a neuron or perceptron are connected to
other processing elements.

[page 18]
Artificial Neural Networks
The neurons are arranged in a layer or vector, with 
the output of one layer serving as the input to the 
next layer and possibly other layers. 
A neuron may be connected to the neurons in the 
subsequent layer, with these connections 
simulating the synaptic connections of the brain. 
Weighted data signals entering a neuron simulate 
the electrical excitation of a nerve cell and 
consequently the transference of information 
within the network or brain.

[page 19]
Artificial Neural Networks
● A given node takes the weighted sum of its 
inputs, and passes it through an activation 
function. 
● The output of the node, then becomes the input 
of another node in the next layer. 
● The signal flows from left to right, and the final 
output is calculated by performing this 
procedure for all the nodes. 
● Training a deep neural network means 
learning the weights associated with all the 
edges.

[page 20]
Artificial Neural Networks
● Without activation, Each neuron could simply 
outputs z, but that would be just a linear 
function, which makes it rather inflexible for 
modelling real-world data. 
● The equation for a given node looks as follows. 
The weighted sum of its inputs passed through a 
non-linear activation function. 
● It can be represented as a vector dot product, 
where n is the number of inputs for the node.

[page 21]
Artificial Neural Networks
● The bias term was omitted for simplicity. 
● Bias is an input to all the nodes and always has 
the value 1. 
● It also helps the model to train when all the input 
features are 0. 
(If this sounds complicated right now you can safely ignore the 
bias terms. For completeness, the above equation looks as 
follows with the bias included.)

[page 22]
Artificial Neural Networks
● So far we have described the forward pass, 
meaning given an input and weights how the 
output is computed. 
● After the training is complete, we only run the 
forward pass to make the predictions.
● But we first need to train our model to actually 
learn the weights

[page 23]
Artificial Neural Networks
The training procedure works as follows:
● Randomly initialize the weights for all the nodes.  
● For every training example, perform a forward 
pass using the current weights, and calculate the 
output of each node going from left to right. The 
final output is the value of the last node.
● Compare the final output with the actual target 
in the training data, and measure the error using 
a loss function.
● Perform a backwards pass from right to left and 
propagate the error to every individual node 
using backpropagation. 
● Calculate each weight’s contribution to the error, 
and adjust the weights accordingly using gradient 
descent. 
● Propagate the error gradients back starting from 
the last layer

[page 24]
Artificial Neural Networks
In the standard ML, the feed forward architecture 
is known as the multilayer perceptron (MLP) 
we can change the number of neurons in the output 
layer to match the number of outputs we want from 
our network.
The optimal number of hidden layers is the subject of 
much discussion, but it’s completely up to whoever 
builds the network.

[page 25]
Artificial Neural Networks
● There’s a lot going on already, even with the 
basic forward pass.
● To simplify and understand the intuition behind 
it:
Essentially what each layer of the ANN does is a 
non-linear transformation of the input from one 
vector space to another.
● Let’s use the ANN in Figure 1 above as an 
example. We have a 3-dimensional input 
corresponding to a vector in 3D space. We then 
pass it through two hidden layers with 4 nodes 
each. And the final output is a 1D vector or a 
scalar.

[page 26]
Artificial Neural Networks
Essentially what each layer of the ANN does is a non-linear 
transformation of the input from one vector space to another.
● Let’s use the ANN in Figure as an example. 
● We have a 3-dimensional input corresponding to 
a vector in 3D space. 
● We then pass it through two hidden layers with 4 
nodes each. 
● And the final output is a 1D vector or a scalar.

[page 27]
Artificial Neural Networks
● So if we visualize this as a sequence of vector 
transformations, we first map the 3D input to a 4D vector 
space, then we perform another transformation to a new 4D 
space, and the final transformation reduces it to 1D. 
● This is just a chain of matrix multiplications. 
● The forward pass performs these matrix dot products and 
applies the activation function element-wise to the result. The 
figure below only shows the weight matrices being used 
without the activations.

[page 28]
Artificial Neural Networks
● The input vector x has 1 row and 3 columns. 
● To transform it into a 4D space, we need to multiply it with 
a 3x4 matrix. Then to another 4D space, we multiply with 
a 4x4 matrix. And finally to reduce it to a 1D space, we use 
a 4x1 matrix.
● The dimensions of the matrices represent the input and 
output dimensions of a layer. The connection between a layer 
with 3 nodes and 4 nodes is a matrix multiplication using 
a 3x4 matrix.

[page 29]
Artificial Neural Networks
● These matrices represent the weights that define the ANN. 
● To make a prediction using the ANN on a given input, we only 
need to know these weights and the activation function (and 
the biases), nothing more. 
● We train the ANN via backpropagation to “learn” these 
weights.
● If we put everything together it looks like the figure below.

[page 30]
Artificial Neural Networks
● A fully connected layer between 3 nodes and 4 nodes is just a 
matrix multiplication of the 1x3 input vector (yellow nodes) 
with the 3x4 weight matrix W1. 
● The result of this dot product is a 1x4 vector represented as 
the blue nodes. We then multiply this 1x4 vector with 
a 4x4 matrix W2, resulting in a 1x4 vector, the green nodes. 
And finally a using a 4x1 matrix W3 we get the output.

[page 31]
Artificial Neural Networks
● We have omitted the activation function in the above figures 
for simplicity. In reality after every matrix multiplication, we 
apply the activation function to each element of the resulting 
matrix. More formally:
● The output of the matrix multiplications go through the 
activation function f. In case of the sigmoid function, this 
means taking the sigmoid of each element in the matrix.

[page 32]
Artificial Neural Networks
● We have omitted the activation function in the above figures 
for simplicity. In reality after every matrix multiplication, we 
apply the activation function to each element of the resulting 
matrix. More formally:
● The output of the matrix multiplications go through the 
activation function f. In case of the sigmoid function, this 
means taking the sigmoid of each element in the matrix.

[page 33]
Artificial Neural Networks - Need
● So far we talked about what deep models are and how they 
work, but why do we need to go deep in the first place?
● We saw that a layer of ANN just performs a non-linear 
transformation of its inputs from one vector space to 
another. 
● If we take a classification problem as an example, we want to 
separate out the classes by drawing a decision boundary. 
● The input data in its given form is not separable. 
● By performing non-linear transformations at each layer, we 
are able to project the input to a new vector space, and draw 
a complex decision boundary to separate the classes.

[page 34]
Artificial Neural Networks - Need
● Let’s visualize what we just described with a concrete 
example. 
● Given the following data we can see that it isn’t linearly 
separable.
● So we project it to a higher dimensional space by performing 
a non-linear transformation, and then it becomes linearly 
separable. The green hyperplane is the decision boundary.

[page 35]
Artificial Neural Networks - Need
● This is equivalent to drawing a complex decision boundary in the original input space
● So the main benefit of having a deeper model is being able to do more non-linear transformations 
of the input and drawing a more complex decision boundary.

[page 36]
Artificial Neural Networks - Need
● As a summary, ANNs are very flexible yet powerful deep learning models. 
● They are universal function approximators, meaning they can model any complex function. 
● There has been an incredible surge on their popularity recently due to a couple of reasons: ---
- huge increase in computational power especially GPUs and distributed training, 
- and vast amount of training data.

[page 37]
Artificial Neural Networks – Parameters 
Our machine has inputs and outputs, but how do we control 
what inputs create what outputs? 
That is, how do we change the neural network so certain inputs 
(say an image of an apple) give the correct outputs 
We can add “knobs” to our machine to control the output for a 
given input. 
In machine learning language, these “knobs” are called 
the parameters of a neural network. 
If we tune these knobs to the correct place, then for any input we 
can get the output that we want.
Going back to our apples and oranges example, if we give our 
machine an image of an apple but it tells us it’s an orange then we 
can go ahead and adjust the knobs of our machine (in other words, 
tune the parameters) until the machine tells us it sees an apple.
In essence, this is what it means to train a neural network and 
this is exactly what the backpropagation algorithm does.

[page 38]
Artificial Neural Networks – Cost 
● If we feed a neural network say with an image of an apple and it tells us it sees an orange, then 
the cost for that particular example would be high. 
● The term “cost” comes from the fact that if you can think of a neural network with a 
high cost and therefore many wrong answers as bad, or expensive or has high error
● Once we have a cost function and many training examples, we can then perform gradient 
descent to minimize the cost function by adjusting our parameters.
● Gradient descent is a way to find the minimum of a function. In the case of a neural 
network, the function that we want to minimize is the cost function. 
● Gradient descent does this by adjusting the parameters of the network such that we get a 
lower value from the cost function than before. 
● In a sense, gradient descent “moves” downhill whenever possible. And each time it moves 
downhill, the gradient descent “saves” its progress by updating the weights and biases in each 
neuron. Eventually, gradient descent will have found the very bottom of the cost function.

[page 39]
Artificial Neural Networks - Backpropagation
● When the neural network outputs the 
wrong answer, we find the slopes of the 
output layer first because it was the direct 
cause of the incorrect answer. 
● And since the output layer depends on the 
hidden layer 
● Eventually we’ll work your way back to the 
hidden layer closest to the input layer.

[page 40]
Artificial Neural Networks - Backpropagation
● Thus, we easily calculate the slopes of the 
last layer, and then the second to last layer, 
and end up working backwards until we 
reach the first, input layer. 
● This is the algorithm: “backpropagation.” 
● We calculate slopes by starting from the 
back and propagating our algorithm 
backwards through the neural network 
until we get all the slopes for gradient 
descent.

[page 41]
Activation Functions
● It may be defined as the extra force or effort applied over the input to 
obtain an exact output. In ANN, we can also apply activation functions 
over the input to get the exact output.
Threshold Function
 Piecewise Linear Function

[page 42]
Threshold Function
Activation function A = “activated” if Y > threshold else not
Alternatively, A = 1 if y> threshold, 0 otherwise.
What if you would want multiple such neurons to be connected to bring in 
more classes, class1, class 2, class 3 etc.
What will happen if more than 1 neuron is “activated”. All neurons will output a 1 ( 
from step function).?

[page 43]
Piece-wise Threshold Function
We can use linear function A=c.x This gives rise to Piece-wise Linear function.