# INTRODUCTION TO NLP

course: Module 3 — Deep Learning & NLP
module: Module-3-Deep-Learning-NLP
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-3-Deep-Learning-NLP/General/INTRODUCTION_TO_NLP.pdf
pages: 29

---
[page 1]
SPECIAL DATA - WORKING 
WITH TEXT DATA

[page 2]
AGENDA
Pre-processing
 Lowercasing
 Tokenization
 Stop-words Removal
 Normalization
 Stemming
 Lemmatization
 Noise Removal
 Part-of-speech tagging
Vector Space Model for Document Representation

[page 3]
INTRODUCTION TO NLP
● Natural language processing is an area of computer science that is integral to 
artificial intelligence. 
● Natural language is the way people communicate via speech and text in real life. 
● This includes everything from signs to instant messages and voice conversations. 
● Natural language is inconsistent, messy, and highly variable. 
● Computers were built to work with highly standardized and uniform data, so, 
originally couldn’t analyze natural language. 
● Natural Language Processing aims at the convergence of machine and human 
languages, and it seeks to enable computers to effectively process large amounts 
of data presented as natural language just as humans can.

[page 4]
TEXT PRE -PROCESSING
There are different ways to preprocess text. Some of the approaches:
•Lowercasing
•Tokenization
•Stop-words Removal
•Normalization
•Stemming
•Lemmatization
•Noise Removal
•Part-of-speech (POS) tagging

[page 5]
LOWERCASING / CASE -FOLDING
A common strategy is to do case-folding by reducing all letters to 
lower case. 
Often this is a good idea since otherwise it might lead to 
mismatches of the same words 
Example - it will allow instances of Automobile at the beginning of 
a sentence to match with a query of automobile.
It will also help in web searches when most users type in ferrari when 
they are interested in a Ferrari car.

[page 6]
LOWERCASING / CASE -FOLDING
Case folding can however equate words that might better be kept apart. 
Many proper nouns are derived from common nouns and so are 
distinguished only by case, including companies, such as –
•General Motors versus the word ‘general’,
•The Associated Press versus ‘associated’, 
•government organizations like - the Fed vs. fed
•and person names like – Bush versus bush,
•Black versus color black. 
Can lead to unintended query expansion with acronyms like - CAT to 
cat

[page 7]
RUNNING EXAMPLE
Input =['Starbucks is doing very well lately.',  'Overall, while it may seem 
there is already a Starbucks on every corner, Starbucks still has a lot of room 
to grow.’, 'They just began expansion into food products, which has been going 
quite well so far for them.’]
Output:
 ['starbucks is doing very well lately.', 'overall, while it may seem there is already a 
starbucks on every corner , starbucks still has a lot of room to grow.', 'they just 
began expansion into food products, which has been going quite well so far for 
them.']

[page 8]
EXPANDING CONTRACTIONS
 Contractions are words or combinations of words that are shortened by 
dropping letters and replacing them by an apostrophe, and removing them 
contributes to text standardization
There are different ways to expand contractions, but one of the most straight 
forward one is to create a dictionary of contractions with their corresponding 
expansions:
Example:
 The students aren’t willing to live in on-campus housing, and they’ll not visit their 
home-towns.
 The students are not willing to live in on-campus housing, and they will not visit their 
home-towns.


[page 9]
TOKENIZATION
Tokenization is a way of separating a piece of text into smaller units called tokens. Here, tokens can 
be either words, characters, or sub-words.
For example, consider the sentence: “Never give up”.
Word tokenization : Assuming space as a delimiter, the tokenization of the sentence results in 3 tokens – 
Never-give-up. 
Similarly, tokens can be either characters or sub-words. For example, let us consider “smarter”:
Character tokenization: s-m-a-r-t-e-r
Sub-word tokenization: smart-er (Not very common)
The simplest approach of tokenization is to split at the white spaces. Extra white spaces are also 
stripped during this step.
Example:
 The students aren’t willing to live in on-campus housing, and they will not visit their home-towns.
 Tokens = [‘The’, ‘students’, ‘aren’t’, ‘willing’, ‘to’, ‘live’, ‘in’, ‘on-campus’, ‘housing,’, ‘and’, ‘they’, ‘will’, ‘not’, ‘visit’, ‘their’, 
‘home-towns.’}

[page 10]
TOKENIZATION CONTD. 
•Tokenization is performed on the corpus to obtain tokens. 
•The tokens obtained are then used to prepare a vocabulary. 
•Vocabulary refers to the set of unique tokens in the corpus.
 
•A vocabulary can be constructed by considering each unique token in the corpus 
or by considering the top K Frequently Occurring Words.

[page 11]
RUNNING EXAMPLE
Input = ['starbucks is doing very well lately. overall, while it may seem 
there is already a starbucks on every corner , starbucks still has a lot of 
room to grow. they just began expansion into food products, which has 
been going quite well so far for them.'] 
Word Tokens:
[‘starbucks', 'is', 'doing', 'very', 'well', 'lately', '.']
[‘overall', ',', 'while', 'it', 'may', 'seem', 'there', 'is', 'already', 'a', ‘starbucks', 
'on', 'every', 'corner', ',', ‘starbucks', 'still', 'has', 'a', 'lot', 'of', 'room', 'to', 
'grow', '.']
[‘they', 'just', 'began', 'expansion', 'into', 'food', 'products', ',', 'which', 'has', 
'been', 'going', 'quite', 'well', 'so', 'far', 'for', 'them', '.']

[page 12]
WHICH TOKENIZATION TO USE ?
Drawbacks of Word Tokenization
One of the major issues with word tokens is dealing with Out Of Vocabulary (OOV) words. 
OOV words refer to the new words which are encountered at testing. These new words do not 
exist in the vocabulary. Hence, these methods fail in handling OOV words.
To solve this issue,  replace the rare words/simply add in training data withunknown tokens (UNK). 
This helps the model to learn the representation of OOV words in terms of UNK tokens
So, during test time, any word that is not present in the vocabulary will be mapped to a UNK 
token. This is how we can tackle the problem of OOV in word tokenizers.
The problem with this approach is that the entire information of the word is lost as we are 
mapping OOV to UNK tokens. And another issue is that every OOV word gets the same 
representation
Generally, pre-trained models are trained on a large volume of the text corpus leading to huge 
vocabularies and hence OOV words don’t usually occur
This opens the door to Character Tokenization.

[page 13]
WHICH TOKENIZATION TO USE ?
Character Tokenization
Character Tokenization splits text into a set of characters.
It overcomes the drawbacks we saw above about Word Tokenization.
Character Tokenizers handles OOV words coherently by preserving the information of the word as 
it breaks down the OOV word into characters and represents the word in terms of these characters
It also limits the size of the vocabulary to 26 since the vocabulary contains a unique set of 
characters
Drawbacks of Character Tokenization
Character tokens solve the OOV problem but the length of the input and output sentences increases 
rapidly as we are representing a sentence as a sequence of characters. 
As a result, it becomes challenging to learn the relationship between the characters to form 
meaningful words.

[page 14]
STOP-WORDS’ REMOVAL
Stop words are a set of commonly used words in a language. 
Examples of stop words in English are “a”, “the”, “is”, “are” and etc. 
By removing low information words from text, we can focus on the important words instead.
Punctuations can also be removed at this stage
Example of stop word removal.  All stop words are removed completely or replaced with a 
dummy character, X:
Input- The students aren’t willing to live in on-campus housing, and they will not 
visit their home-towns.
Output- X students X willing X live X on-campus housing, X they X X visit X 
home-towns.

[page 15]
RUNNING EXAMPLE
Input: 
 ['Starbucks', 'is', 'doing', 'very', 'well', 'lately', '.']
 ['Overall', ',', 'while', 'it', 'may', 'seem', 'there', 'is', 'already', 'a', 'Starbucks', 'on', 
'every', 'corner', ',', 'Starbucks', 'still', 'has', 'a', 'lot', 'of', 'room', 'to', 'grow', '.']
 ['They', 'just', 'began', 'expansion', 'into', 'food', 'products', ',', 'which', 'has', 'been', 
'going', 'quite', 'well', 'so', 'far', 'for', 'them', '.’]
Output:
 ['Starbucks', 'well', 'lately’]
 ['Overall’, 'may’, 'seem’, 'already’, 'Starbucks’, 'every’, 'corner’, 'Starbucks’, 'still’, 'lot’, 
'room’, 'grow’] 
 ['began’, 'expansion’, 'food’, 'products’, 'going’, 'quite’, 'well’, 'far' ]

[page 16]
TEXT NORMALIZATION
Natural language, as a human resource, tends to follow the inherent nature of its creator’s randomness. 
This means that, as we “produce” natural language, we imprint our random states to it. 
When we normalize text, we attempt to reduce its randomness, bringing it closer to a predefined 
“standard”. 
This helps us to reduce the amount of different information that the computer has to deal with, and 
therefore improves efficiency. 
The goal of is to reduce inflectional forms and sometimes derivationally related forms of a word to 
a common base form.
In linguistics, inflexion is a process of word formation in which a word is modified to express 
different grammatical categories such as tense, case, person, number, gender mood etc.
Normalization techniques : stemming and lemmatization
Example:
Mapping words such as “stopwords”, “stop-words” and “stop words” to just “stopword”.

[page 17]
STEMMING
Stemming refers to a crude heuristic process that chops off the ends of words in 
the hope of achieving this goal correctly most of the time.
Stemming is the process of reducing a word to its word stem by eliminating 
affixes  - suffixes, prefixes, infixes, circumfixes from a word.
The result? Stemming a word or sentence may result in words that are not actual 
words.
So the words “trouble”, “troubled”, “troubling” and “troubles” might actually be 
converted to “troubl” 
For example, the words “university” and “universe” that get reduced to “univers”.
(instead of trouble because the ends were just chopped off)

[page 18]
RUNNING EXAMPLE
Input:
 ['Starbucks', 'well', 'lately’],
 ['Overall’, 'may’, 'seem’, 'already’, 'Starbucks’, 'every’, 'corner’, 'Starbucks’, 'still’, 'lot’, 
'room’, 'grow’], 
 ['began’, 'expansion’, 'food’, 'products’, 'going’, 'quite’, 'well’, 'far’]
Output:
 ['starbuck well late',
  'overal may seem alread starbuck ever corner starbuck still lot room grow',
  'began expans food product go quit well far']

[page 19]
LEMMATIZATION
•   Lemmatization is very similar to stemming, but the goal is to remove 
inflections and map a word to its root form.
•   Stemming usually refers to a heuristic process that chops off the ends of 
words in the hope of achieving this goal correctly most of the time, and 
often includes the removal of derivational affixes.
•   Lemmatization usually refers to doing things properly by doing 
morphological analysis of words, normally aiming to remove 
inflectional endings only and to return the base or dictionary form of a 
word, which is known as the lemma . 
•   Lemmatizer is usually more sophisticated than stemming, since stemmers 
works on an individual word without knowledge of the context.

[page 20]
RUNNING EXAMPLE
Input:
 ['Starbucks', 'well', 'lately’],
 ['Overall’, 'may’, 'seem’, 'already’, 'Starbucks’, 'every’, 'corner’, 'Starbucks’, 'still’, 'lot’, 'room’, 
'grow’], 
 ['began’, 'expansion’, 'food’, 'products’, 'going’, 'quite’, 'well’, 'far’]
Output of Lemmatization:
 ['starbucks well lately',
 'overall may seem already starbucks every corner starbucks still lot room grow',
  'began expansion food product going quite well far']
Output of Stemming:
 starbuck well late',
  'overal may seem alread starbuck ever corner starbuck still lot room grow',
  'began expans food product go quit well far']

[page 21]
NOISY TEXT 
Noisy text is text with differences between the representation of the text and the intended, 
correct, or original text.
Idiomatic expressions, abbreviations, acronyms and business-specific lingo can all cause noisy text. 
It is particularly prevalent in the unstructured text found in blog posts, chat conversations, and SMS 
messages. Other potential causes include poor spelling and punctuation, typographical errors and 
poor translations  and speech recognition programs
Noise removal is about removing characters digits and pieces of text that can interfere with the 
text analysis. It is also highly domain dependent.
Example, “…trouble…”, “trouble9”, “trouble>”, and “1.trouble” are converted to “trouble”.

[page 22]
PART-OF-SPEECH (POS) TAGGING
POS tagging assigns every word with the corresponding part of speech.
POS tag list:
 JJ adjective 'big'
 JJR adjective, comparative 'bigger'
 JJS adjective, superlative 'biggest'
 NN noun, singular 'desk'
 NNS noun plural 'desks'
 NNP proper noun, singular 'Harrison’
 VB verb, base form take
 VBD verb, past tense took
 VBG verb, gerund/present participle taking
 etc…

[page 23]
RUNNING EXAMPLE
Input:
 ['starbucks well lately',
  'overall may seem already starbucks every corner 
starbucks still lot room grow',
  'began expansion food product going quite well 
far’]
Output:
 [('starbucks', 'NNS'), ('well', 'RB'), ('lately', 'RB’)]
 [('overall', 'JJ'), ('may', 'MD'), ('seem', 'VB'), 
('already', 'RB'), ('starbucks', 'NNS'), ('every', 
'DT'), ('corner', 'NN'), ('starbucks', 'NNS'), ('still', 
'RB'), ('lot', 'VBP'), ('room', 'NN'), ('grow', 'NN’)]
 [('they', 'PRP'), ('began', 'VBD'), ('expansion', 
'NN'), ('food', 'NN'), ('products', 'NNS'), ('going', 
'VBG'), ('quite', 'RB'), ('well', 'RB'), ('far', 'RB')]
POS Tag Meanings
• NNS → Noun, plural
(e.g., “starbucks”, “products”)
• NN → Noun, singular or mass
(e.g., “corner”, “room”, “food”, “expansion”)
• RB → Adverb
(e.g., “well”, “lately”, “already”, “still”, “quite”, “far”)
• JJ → Adjective
(e.g., “overall”)
• MD → Modal verb
(e.g., “may”)
• VB → Verb, base form
(e.g., “seem”)
• VBP → Verb, non-3rd person singular present
VBD → Verb, past tense
(e.g., “began”)
• VBG → Verb, gerund/present participle
(e.g., “going”)
• PRP → Personal pronoun
(e.g., “they”)
• DT → Determiner
(e.g., “every”)

[page 24]
VECTOR SPACE MODEL
● One of the critical components in Natural Language Processing 
(NLP) is to encode text information in a numerical format that can 
be fed into an NLP model. 
● This technique which represents words in a numerical vector 
space, is called Vector Space Modelling. 
● A typical vector space model  has a dimension of V×N
● Where V is number of documents and N is a size of unique 
vocabulary

[page 25]
COUNT VECTOR - BAG OF WORDS (BOW)
The bag-of-words model is a simple 
representation
In this model, A document (sentence, 
paragraph, etc.) is treated as a bag 
(multiset) of its words, ignoring:
 grammar
 word order
 syntax
This disregards word order but 
keeps multiplicity
The bag-of-words model is commonly 
used in document classification where 
the frequency of occurrence of each 
word is used as a feature

[page 26]
COUNT VECTOR - BAG OF WORDS (BOW)
(1) John likes to watch movies. Mary likes movies too. 
(2) Mary also likes to watch football games. 
BoW1 = {"John":1,"likes":2,"to":1,"watch":1,"movies":2,"Mary":1,"too":1}; 
BoW2 = {"Mary":1,"also":1,"likes":1,"to":1,"watch":1,"football":1,"games":1}; 
Each key is the word, and each value is the number of occurrences of that 
word in the given text document.
The order of elements is free, so, 
{"too":1,"Mary":1,"movies":2,"John":1,"watch":1,"likes":2,"to":1} is also 
equivalent to BoW1.  
The order of elements POS tagging and stemming are not inherently part 
of Bag of Words (BOW) 
BOW only cares about counting words (or frequencies) in documents.
It does not look at grammar, syntax, or meaning — just raw tokens.

[page 27]
COUNT VECTOR - BAG OF WORDS (BOW)
if another document is like a union of these two,
(3) John likes to watch movies. Mary likes movies too. Mary also likes to 
watch football games. 
BoW3 = 
{"John":1,"likes":3,"to":2,"watch":2,"movies":2,"Mary":2,"too":1,"also":1,"f
ootball":1,"games":1}; 
(1) John likes to watch movies. Mary likes movies too. 
(2) Mary also likes to watch football games. 
(1) [1, 2, 1, 1, 2, 1, 1, 0, 0, 0] 
(2) [0, 1, 1, 1, 0, 1, 0, 1, 1, 1]

[page 28]
COUNT VECTOR – N-GRAM MODEL
The Bag-of-words model is an orderless document representation - only the counts of words matter. 
As an alternative, the n-gram model can store this spatial information. 
Applying to the same example above, abigram model will parse the text into the following units and store 
the term frequency of each unit as before.
(1) John likes to watch movies. Mary likes movies too. 
(2) Mary also likes to watch football games. 
[ "John likes“, "likes to“, "to watch", "watch movies", "Mary likes", "likes movies", "movies too", ] 
(1) [ "John likes“:1, "likes to“:1, "to watch“:1, "watch movies“:1, "Mary likes“:1, "likes movies“:1, 
"movies too“:1, ] 
(2) ["Mary also“:1, “also likes“:1, "likes to“:1, "to watch“:1, "watch football“:1, "football games“:1]
Conceptually, we can view bag-of-word model as a special case of the n-gram model, with n=1. 
For n>1 the model is named n-gram where n denotes the number of grouped words

[page 29]
TF-IDF (TERM FREQUENCY -INVERSE DOCUMENT FREQUENCY)
Here the term-specific weights in the document vectors are products of local and global 
parameters. 
The weight vector for document d is  
The number of times a term occurs in a document is called its term frequency tf(i, j),
Term frequency adjusted for document length: