# numpypandas%2B
course: Academy — 65-GENAI-for-Engineers
module: Academy/65-GENAI-for-Engineers
type: notebook
source_url: https://personal-learn.armco.dev/files/Academy/65-GENAI-for-Engineers/folders/Live_Intervention_OOPs_and_Numpy_Folder/numpypandas%2B.ipynb
---
[cell 1 markdown]
Numpy
[cell 2 code]
import numpy as np #np is the alias name for numpy
[cell 3 code]
my_list=[1,5,7,9,4]
[cell 4 code]
type(my_list)
[cell 5 code]
#converting list into an array
arr=np.array(my_list)
[cell 6 code]
print(arr)
[cell 7 code]
type(arr) #nd means 'n' dimentional and here value of n=1 i.e. 1darray
[cell 8 code]
arr.shape
[cell 9 code]
#2d array is also called as multinested array
my_list_1=[5,7,19,44,45]
my_list_2=[17,19,14,45,4]
my_list_3=[9,4,11,4,7]
[cell 10 code]
arr_new=np.array([my_list_1,my_list_2,my_list_3])
[cell 11 code]
print(arr_new) #2d array
[cell 12 code]
arr_new.shape #here 3 is the number of rows and 5 is the number of columns
[cell 13 code]
arr_new.reshape(5,3)
[cell 14 code]
arr_new.reshape(5,4)
[cell 15 code]
#3D array
[cell 16 code]
#generating random numbers from numpy
[cell 17 code]
arr=np.arange(0,10,1)
[cell 18 code]
print(arr)
[cell 19 code]
arr_2=np.arange(0,10,2)
[cell 20 code]
print(arr_2)
[cell 21 code]
np.linspace(1,10,30) #here 30 denotes the total number of elements we want in the given range of 1-10
[cell 22 code]
np.linspace(1,10,50)
[cell 23 code]
np.linspace(1,10,80)
[cell 24 code]
np.random.rand(3,4) #here 3 is the number of rows and 4 is the number of columns
#by default the range is from 0 to 1 (1 being exclusive)
#the numbers generated will change everytime we run the code cell
[cell 25 code]
np.random.randint(1,20,4) #1 to 20 is the range and in the specified range, we want 4 elements
#the numbers generated will change everytime we run the code cell
[cell 26 markdown]
Pandas
[cell 27 markdown]
1. Load and read the data
2. Statistical inferences from the data
3. General inferences from the data
4. Creating dummy datasets from numpy and pandas both
[cell 28 code]
import pandas as pd #pd is the alias name for pandas
[cell 29 code]
#loading and reading the data
data=pd.read_csv("/content/tips.csv")
[cell 30 code]
data.head() #head() by default will show the first 5 rows of the dataset
[cell 31 code]
data.head(10) #here first 10 rows will be displayed
[cell 32 code]
data.head(15) #here first 15 rows will be displayed
[cell 33 code]
data.tail() #by default will show the last 5 rows of the dataset
[cell 34 code]
data.tail(10)
[cell 35 code]
data.info() #will display all the basic information about the dataset
[cell 36 code]
data.describe() #will show statistical inferences about the numerical columns in the dataset
[cell 37 code]
import numpy as np
[cell 38 code]
my_data=pd.DataFrame(np.arange(0,20).reshape(5,4),index=['row1','row2','row3','row4','row5'],columns=['col1','col2','col3','col4'])
[cell 39 code]
print(my_data)
[cell 40 code]
my_data.to_csv('sunday.csv')
[cell 41 code]
my_data.to_excel('sunday.xlsx')
[cell 42 code]
#fetching multiple values using .iloc[]
my_data.iloc[0:3,0:2]
[cell 43 code]
print(my_data)
[cell 44 code]
my_data.iloc[0:3,2:]
[cell 45 code]
my_data.iloc[:4,1:3]