# Linear Regression Explanation
course: Module 2 — Machine Learning Algorithms
module: Module-2-Machine-Learning-Algorithms
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-2-Machine-Learning-Algorithms/General/Linear_Regression_Explanation.pdf
pages: 26
---
[page 1]
Linear Regression
[page 2]
What is Regression?
Uses of Regression
Linear Regression
Types of Linear Regression
Mathematical demonstration
Outline
[page 3]
Regression Meaning in English Language
The English term "regression" comes from the Latin word regressus,
meaning "a return" or "going back."
The term regression was first used in a statistical sense by Sir Francis
Galton in the 19th century. He observed that:
• Tall parents tend to have tall children, but the children are usually not
as tall as the parents.
• Likewise, short parents tend to have short children, but not as short.
• He called this phenomenon “regression toward the mean” — the idea
that extreme traits tend to move closer to average values in the next
generation.
[page 4]
Regression in ML
Linear regression is a statistical and machine learning technique
used to model the relationship between a dependent variable
called target and one or more independent variables called
predictors or features.
Key Idea
It tries to fit a straight line in 2D or a hyperplane in higher
dimensions that best describes the relationship between variables.
[page 5]
Purpose - Regression
Predicting / forecasting values of the target variable
Understanding the strength and type (positive/negative) of relationships
between variables
The strength of relationships refers to how strongly two variables are
related to each other — that is how much change in one variable is
associated with change in another, can be measured by correlation
analysis or by R Squared
[page 6]
Assumptions - Regression
Linear regression relies on several assumptions:
• Linearity: Relationship between input and output is linear
• Independence: Observations are independent of each other
• Normal distribution of Errors: Constant variance of errors and errors
are normally distributed
[page 7]
R-squared (R²)
Explains how much of the variability in the dependent variable
is explained by the model
Range: 0 to 1
0.0 → model explains no variance in y
1.0 → model explains all the variance in y
Example:
R2 = 0.85 → 85% of the variation in the target is explained by the
predictors → strong relationship
[page 8]
R-squared (R²)
Explains how much of the variability in the dependent variable
is explained by the model
[page 9]
Major uses of regression analysis are:
➢ Predicting an effect
Ex. How much additional sale income will be generated for each 1000 dollar spent on
marketing.
➢Trend forecasting
Ex. What will be the price of gold in next six months.
Regression: Uses
[page 10]
Types of Linear Regression
[page 11]
➢ In simple linear regression, the dependent variable depends only on a single
independent variable.
➢ For simple linear regression, the form of the model is-
ෝ𝒚 = β0 + β1X
• ෝ𝒚 is a dependent/predicted variable.
• X is an independent variable.
• β0 and β1 are the regression coefficients.
• β0 is the intercept of the line.
• β1 is the slope of the line.
Simple Linear Regression
[page 12]
Understanding Linear Regression
[page 13]
Fit data with the best line which
"goes through" the points.
[page 14]
For each point the difference
between the forecasted point and
the actual observation is the
error
ෝ𝒚 = β0 + β1X
ෝ𝒚 − 𝐟𝐨𝐫𝐞𝐜𝐚𝐬𝐭ed /
predicted value
β0 - Intercept
β1 - Slope
Simple Linear Regression
[page 15]
➢ In multiple linear regression, the dependent variable depends on more than one
independent variables.
➢ For multiple linear regression, the form of the model is-
ෝ𝒚 = β0 + β1X1 + β2X2 + β3X3 + …… + βnXn
Here,
• ෝ𝒚 is a dependent variable.
• X1, X2, …., Xn are independent variables.
• β0, β1,…, βn are the regression coefficients.
Multiple Linear Regression
[page 16]
Cost Function – Mean Squared Error (MSE)
• To select the best fit line, we will consider a cost function, represented as follows:
𝑪𝒐𝒔𝒕 𝑭𝒖𝒏𝒄𝒕𝒊𝒐𝒏 = 𝟏
𝒎
𝒊=𝟏
𝒎
(ෝ𝒚𝒊 − 𝒚𝒊)𝟐
where m = number of data instances
ො𝑦𝑖 = forecasted value
𝑦𝑖 = actual value
• Minimize the cost function.
• The line with minimum error will be the best fit line.
[page 17]
Let’s understand this with an example:
ෝ𝒚 = β0 +β1X
β0 = 0
β1 = 1
ෝ𝒚 = X
for X=1, ො𝐲 =1
for X=2, ො𝐲 =2
for X=3, ො𝐲 =3
for X=4, ො𝐲 =4
Error =0
β0 = 0
β1 = 0.5
ෝ𝒚 = 0.5X
for X=1, ො𝐲 =0.5
for X=2, ො𝐲 =1.0
for X=3, ො𝐲 =1.5
for X=4, ො𝐲 =2
Error=0.9375
𝐶𝑜𝑠𝑡 𝑓𝑢𝑛𝑐𝑡𝑖𝑜𝑛 = 1
4 (1 − 1)2+(2 − 2)2+(3 − 3)2+(4 − 4)2
1 2 3 4
1 2 3 4
β1=1
β1=0.5
X
Y
Case 1:
Case 2:
Case 3: β0 = 0
β1 = 0.0
ෝ𝒚 = 0
Erroe=7.5
X Y
1 1
2 2
3 3
4 4
[page 18]
Linear Regression with Maths Intuition
x y
1 3
2 4
3 2
4 4
5 5
𝑦 = 𝑚𝑥 + 𝑐
[page 19]
Understanding Linear Regression
x y 𝒙 − ഥ𝒙 𝒚 − ഥ𝒚 (𝒙 − ഥ𝒙) 𝟐 (𝒙 − ഥ𝒙)(𝒚 − ഥ𝒚)
1 3 -2 -0.6 4 1.2
2 4 -1 0.4 1 -.04
3 2 0 -1.6 0 0
4 4 1 0.4 1 0.4
5 5 2 1.4 4 2.8
Mean=3 Mean=3.6 =10 =4
𝑚 = (𝒙 − ഥ𝒙)(𝒚 − ഥ𝒚)
(𝒙 − ഥ𝒙) 𝟐
𝑦 = 𝑚𝑥 + 𝑐
3.6 = 1.2 + 𝑐
𝑚 = 0.4
𝑐 = 2.4
[page 20]
Understanding Linear Regression
𝑦 = 𝑚𝑥 + 𝑐
𝑚 = 0.4
𝑐 = 2.4
𝑦 = 0.4𝑥 + 2.4
𝑦 = 0.4𝑥 + 2.4
[page 21]
Understanding Linear Regression
Actual values
Forecasted values
𝑦 = 0.4 ∗ 1 + 2.4
𝑦 = 0.4 ∗ 2 + 2.4
𝑦 = 0.4 ∗ 5 + 2.4
Regression line
𝑦 = 0.4𝑥 + 2.4
[page 22]
Error
error
error
error
error
Distance between actual
and forecasted values
[page 23]
Regression line
Calculation of R2
Distance between forecasted and mean
vs
Distance between actual and mean
𝑅2 = σ( ො𝑦 − ത𝑦)2
σ(𝑦 − ത𝑦)2
[page 24]
x y 𝒚 − ഥ𝒚 (𝒚 − ഥ𝒚) 𝟐 ො𝑦 ො𝑦 − ഥ𝒚 ( ො𝑦 − ഥ𝒚) 𝟐
1 3 -0.6 0.36 2.8 -0.8 0.64
2 4 0.4 0.16 3.2 -0.4 0.16
3 2 -1.6 2.56 3.6 0 0
4 4 0.4 0.16 4.0 0.4 0.16
5 5 1.4 1.96 4.4 0.8 0.64
Mean=3.6 =5.2 =1.6
Calculation of R2
𝑅2 = σ( ො𝑦 − ത𝑦)2
σ(𝑦 − ത𝑦)2 = 1.6
5.2
𝑅2 0.3
Distance between forecasted and mean
vs
Distance between actual and mean
[page 25]
Root Mean Square Error (RMSE): RMSE is the square root of the mean of
square of all errors. It is a negatively-oriented score, which means lower
values are better.
𝑅𝑀𝑆𝐸 = 1
𝑛
𝑖=1
𝑛
( ෝ𝑦𝑖 − 𝑦𝑖)2
Performance evaluation metrics
[page 26]
Mean Absolute Error (MAE): The mean absolute error of a model is the mean
of the absolute values of the individual forecasting errors. It is a negatively-
oriented score, which means lower values are better.
𝑀𝐴𝐸 = σ𝑖=1
𝑛 ෝ𝑦𝑖 − 𝑦𝑖
𝑛
Performance evaluation metrics