# Linear Regression Summary

course: Module 2 — Machine Learning Algorithms
module: Module-2-Machine-Learning-Algorithms
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-2-Machine-Learning-Algorithms/General/Linear_Regression_Summary.pdf
pages: 26

---
[page 1]
Linear Regression

[page 2]
 What is Regression?
 Uses of Regression
 Linear Regression
 Types of Linear Regression
 Mathematical demonstration
Outline

[page 3]
Regression Meaning in English Language
The English term "regression" comes from the Latin word regressus, 
meaning "a return" or "going back."
The term regression was first used in a statistical sense by Sir Francis 
Galton in the 19th century. He observed that:
•  Tall parents tend to have tall children, but the children are usually not 
as tall as the parents.
• Likewise, short parents tend to have short children, but not as short.
• He called this phenomenon “regression toward the mean” — the idea 
that extreme traits tend to move closer to average values in the next 
generation.

[page 4]
Regression in ML
Linear regression is a statistical and machine learning technique 
used to model the relationship between a dependent variable 
called target and one or more independent variables called 
predictors or features.
Key Idea
It tries to fit a straight line in 2D or a hyperplane in higher 
dimensions that best describes the relationship between variables.

[page 5]
Purpose - Regression
Predicting / forecasting values of the target variable
Understanding the strength and type (positive/negative) of relationships 
between variables
The strength of relationships refers to how strongly two variables are 
related to each other — that is how much change in one variable is 
associated with change in another, can be measured by correlation 
analysis or by R Squared

[page 6]
Assumptions - Regression
Linear regression relies on several assumptions:
• Linearity: Relationship between input and output is linear
• Independence: Observations are independent of each other
• Normal distribution of Errors: Constant variance of errors and errors 
are normally distributed

[page 7]
R-squared (R²)
Explains how much of the variability in the dependent variable 
is explained by the model
Range: 0 to 1
0.0 → model explains no variance in y
1.0 → model explains all the variance in y
Example:
R2 = 0.85 → 85% of the variation in the target is explained by the 
predictors → strong relationship

[page 8]
R-squared (R²)
Explains how much of the variability in the dependent variable 
is explained by the model

[page 9]
Major uses of regression analysis are:
➢ Predicting an effect
 Ex. How much additional sale income will be generated for each 1000 dollar spent on  
marketing.
➢Trend forecasting 
 Ex. What will be the price of gold in next six months.
Regression: Uses

[page 10]
Types of Linear Regression

[page 11]
➢ In simple linear regression, the dependent variable depends only on a single 
independent variable.
➢ For simple linear regression, the form of the model is-
ෝ𝒚 = β0 + β1X 
• ෝ𝒚 is a dependent/predicted variable.
• X is an independent variable.
• β0 and β1 are the regression coefficients.
• β0 is the intercept of the line.
• β1 is the slope of the line.
Simple Linear Regression

[page 12]
Understanding Linear Regression

[page 13]
Fit data with the best line which 
"goes through" the points.

[page 14]
For each point the difference 
between the forecasted point and 
the actual observation is the 
error
ෝ𝒚 = β0 + β1X
ෝ𝒚  − 𝐟𝐨𝐫𝐞𝐜𝐚𝐬𝐭ed / 
predicted value
β0  - Intercept
β1 - Slope
Simple Linear Regression

[page 15]
➢ In multiple linear regression, the dependent variable depends on more than one 
independent variables.
➢ For multiple linear regression, the form of the model is-
ෝ𝒚 = β0 + β1X1 + β2X2 + β3X3 + …… + βnXn
Here,
• ෝ𝒚 is a dependent variable.
• X1, X2, …., Xn are independent variables.
• β0, β1,…, βn are the regression coefficients.
Multiple Linear Regression

[page 16]
Cost Function – Mean Squared Error (MSE)
• To select the best fit line, we will consider a cost function, represented as follows:
𝑪𝒐𝒔𝒕 𝑭𝒖𝒏𝒄𝒕𝒊𝒐𝒏 = 𝟏
𝒎 ෍
𝒊=𝟏
𝒎
(ෝ𝒚𝒊 − 𝒚𝒊)𝟐
 where m = number of data instances
            ො𝑦𝑖 = forecasted value
            𝑦𝑖 = actual value
• Minimize the cost function.
• The line with minimum error will be the best fit line.

[page 17]
Let’s understand this with an example:
ෝ𝒚 = β0 +β1X
β0 = 0
               β1 = 1
ෝ𝒚 = X
for X=1, ො𝐲 =1
for X=2, ො𝐲 =2
for X=3, ො𝐲 =3
for X=4, ො𝐲 =4
Error =0
β0 = 0
           β1 = 0.5
 ෝ𝒚 = 0.5X
for X=1, ො𝐲 =0.5
for X=2, ො𝐲 =1.0
for X=3, ො𝐲 =1.5
for X=4, ො𝐲 =2
Error=0.9375
𝐶𝑜𝑠𝑡 𝑓𝑢𝑛𝑐𝑡𝑖𝑜𝑛 = 1
4 (1 − 1)2+(2 − 2)2+(3 − 3)2+(4 − 4)2
1 2 3 4
1 2 3 4
β1=1
β1=0.5
X
Y
Case 1:
Case 2:
Case 3: β0 = 0
           β1 = 0.0
ෝ𝒚 = 0
Erroe=7.5
X Y
1 1
2 2
3 3
4 4

[page 18]
Linear Regression with Maths Intuition
x y
1 3
2 4
3 2
4 4
5 5
𝑦 = 𝑚𝑥 + 𝑐

[page 19]
Understanding Linear Regression  
x y 𝒙 − ഥ𝒙 𝒚 − ഥ𝒚 (𝒙 − ഥ𝒙) 𝟐 (𝒙 − ഥ𝒙)(𝒚 − ഥ𝒚)
1 3 -2 -0.6 4 1.2
2 4 -1 0.4 1 -.04
3 2 0 -1.6 0 0
4 4 1 0.4 1 0.4
5 5 2 1.4 4 2.8
Mean=3 Mean=3.6  =10  =4
𝑚 = ෍ (𝒙 − ഥ𝒙)(𝒚 − ഥ𝒚) 
(𝒙 − ഥ𝒙) 𝟐
𝑦 = 𝑚𝑥 + 𝑐
3.6 = 1.2 + 𝑐
𝑚 = 0.4
𝑐 = 2.4

[page 20]
Understanding Linear Regression  
𝑦 = 𝑚𝑥 + 𝑐
𝑚 = 0.4
𝑐 = 2.4
𝑦 = 0.4𝑥 + 2.4
𝑦 = 0.4𝑥 + 2.4

[page 21]
Understanding Linear Regression  
Actual  values
Forecasted  values
𝑦 = 0.4 ∗ 1 + 2.4
𝑦 = 0.4 ∗ 2 + 2.4
𝑦 = 0.4 ∗ 5 + 2.4
Regression line
𝑦 = 0.4𝑥 + 2.4

[page 22]
Error
error
error
error
error
Distance between actual 
and forecasted values

[page 23]
Regression line
Calculation of R2
Distance between forecasted and mean
vs
Distance between actual and mean
𝑅2 = σ( ො𝑦 − ത𝑦)2
σ(𝑦 − ത𝑦)2

[page 24]
x y 𝒚 − ഥ𝒚 (𝒚 − ഥ𝒚) 𝟐 ො𝑦 ො𝑦 − ഥ𝒚 ( ො𝑦 − ഥ𝒚) 𝟐
1 3 -0.6 0.36 2.8 -0.8 0.64
2 4 0.4 0.16 3.2 -0.4 0.16
3 2 -1.6 2.56 3.6 0 0
4 4 0.4 0.16 4.0 0.4 0.16
5 5 1.4 1.96 4.4 0.8 0.64
Mean=3.6  =5.2  =1.6
Calculation of R2
𝑅2 = σ( ො𝑦 − ത𝑦)2
σ(𝑦 − ത𝑦)2 = 1.6
5.2
𝑅2  0.3
Distance between forecasted and mean
vs
Distance between actual and mean

[page 25]
Root Mean Square Error (RMSE): RMSE is the square root of the mean of  
square of all errors. It is a negatively-oriented score, which means lower 
values are better.
𝑅𝑀𝑆𝐸 = 1
𝑛 ෍
𝑖=1
𝑛
( ෝ𝑦𝑖 − 𝑦𝑖)2
Performance evaluation metrics

[page 26]
 Mean Absolute Error (MAE): The mean absolute error of a model is the mean 
of the absolute values of the individual forecasting errors. It is a negatively-
oriented score, which means lower values are better.
𝑀𝐴𝐸 = σ𝑖=1
𝑛 ෝ𝑦𝑖 − 𝑦𝑖
𝑛
Performance evaluation metrics