# INTRODUCTION TO NLP
course: Module 3 — Deep Learning & NLP
module: Module-3-Deep-Learning-NLP
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-3-Deep-Learning-NLP/General/INTRODUCTION_TO_NLP.pdf
pages: 29
---
[page 1]
SPECIAL DATA - WORKING
WITH TEXT DATA
[page 2]
AGENDA
Pre-processing
Lowercasing
Tokenization
Stop-words Removal
Normalization
Stemming
Lemmatization
Noise Removal
Part-of-speech tagging
Vector Space Model for Document Representation
[page 3]
INTRODUCTION TO NLP
● Natural language processing is an area of computer science that is integral to
artificial intelligence.
● Natural language is the way people communicate via speech and text in real life.
● This includes everything from signs to instant messages and voice conversations.
● Natural language is inconsistent, messy, and highly variable.
● Computers were built to work with highly standardized and uniform data, so,
originally couldn’t analyze natural language.
● Natural Language Processing aims at the convergence of machine and human
languages, and it seeks to enable computers to effectively process large amounts
of data presented as natural language just as humans can.
[page 4]
TEXT PRE -PROCESSING
There are different ways to preprocess text. Some of the approaches:
•Lowercasing
•Tokenization
•Stop-words Removal
•Normalization
•Stemming
•Lemmatization
•Noise Removal
•Part-of-speech (POS) tagging
[page 5]
LOWERCASING / CASE -FOLDING
A common strategy is to do case-folding by reducing all letters to
lower case.
Often this is a good idea since otherwise it might lead to
mismatches of the same words
Example - it will allow instances of Automobile at the beginning of
a sentence to match with a query of automobile.
It will also help in web searches when most users type in ferrari when
they are interested in a Ferrari car.
[page 6]
LOWERCASING / CASE -FOLDING
Case folding can however equate words that might better be kept apart.
Many proper nouns are derived from common nouns and so are
distinguished only by case, including companies, such as –
•General Motors versus the word ‘general’,
•The Associated Press versus ‘associated’,
•government organizations like - the Fed vs. fed
•and person names like – Bush versus bush,
•Black versus color black.
Can lead to unintended query expansion with acronyms like - CAT to
cat
[page 7]
RUNNING EXAMPLE
Input =['Starbucks is doing very well lately.', 'Overall, while it may seem
there is already a Starbucks on every corner, Starbucks still has a lot of room
to grow.’, 'They just began expansion into food products, which has been going
quite well so far for them.’]
Output:
['starbucks is doing very well lately.', 'overall, while it may seem there is already a
starbucks on every corner , starbucks still has a lot of room to grow.', 'they just
began expansion into food products, which has been going quite well so far for
them.']
[page 8]
EXPANDING CONTRACTIONS
Contractions are words or combinations of words that are shortened by
dropping letters and replacing them by an apostrophe, and removing them
contributes to text standardization
There are different ways to expand contractions, but one of the most straight
forward one is to create a dictionary of contractions with their corresponding
expansions:
Example:
The students aren’t willing to live in on-campus housing, and they’ll not visit their
home-towns.
The students are not willing to live in on-campus housing, and they will not visit their
home-towns.
[page 9]
TOKENIZATION
Tokenization is a way of separating a piece of text into smaller units called tokens. Here, tokens can
be either words, characters, or sub-words.
For example, consider the sentence: “Never give up”.
Word tokenization : Assuming space as a delimiter, the tokenization of the sentence results in 3 tokens –
Never-give-up.
Similarly, tokens can be either characters or sub-words. For example, let us consider “smarter”:
Character tokenization: s-m-a-r-t-e-r
Sub-word tokenization: smart-er (Not very common)
The simplest approach of tokenization is to split at the white spaces. Extra white spaces are also
stripped during this step.
Example:
The students aren’t willing to live in on-campus housing, and they will not visit their home-towns.
Tokens = [‘The’, ‘students’, ‘aren’t’, ‘willing’, ‘to’, ‘live’, ‘in’, ‘on-campus’, ‘housing,’, ‘and’, ‘they’, ‘will’, ‘not’, ‘visit’, ‘their’,
‘home-towns.’}
[page 10]
TOKENIZATION CONTD.
•Tokenization is performed on the corpus to obtain tokens.
•The tokens obtained are then used to prepare a vocabulary.
•Vocabulary refers to the set of unique tokens in the corpus.
•A vocabulary can be constructed by considering each unique token in the corpus
or by considering the top K Frequently Occurring Words.
[page 11]
RUNNING EXAMPLE
Input = ['starbucks is doing very well lately. overall, while it may seem
there is already a starbucks on every corner , starbucks still has a lot of
room to grow. they just began expansion into food products, which has
been going quite well so far for them.']
Word Tokens:
[‘starbucks', 'is', 'doing', 'very', 'well', 'lately', '.']
[‘overall', ',', 'while', 'it', 'may', 'seem', 'there', 'is', 'already', 'a', ‘starbucks',
'on', 'every', 'corner', ',', ‘starbucks', 'still', 'has', 'a', 'lot', 'of', 'room', 'to',
'grow', '.']
[‘they', 'just', 'began', 'expansion', 'into', 'food', 'products', ',', 'which', 'has',
'been', 'going', 'quite', 'well', 'so', 'far', 'for', 'them', '.']
[page 12]
WHICH TOKENIZATION TO USE ?
Drawbacks of Word Tokenization
One of the major issues with word tokens is dealing with Out Of Vocabulary (OOV) words.
OOV words refer to the new words which are encountered at testing. These new words do not
exist in the vocabulary. Hence, these methods fail in handling OOV words.
To solve this issue, replace the rare words/simply add in training data withunknown tokens (UNK).
This helps the model to learn the representation of OOV words in terms of UNK tokens
So, during test time, any word that is not present in the vocabulary will be mapped to a UNK
token. This is how we can tackle the problem of OOV in word tokenizers.
The problem with this approach is that the entire information of the word is lost as we are
mapping OOV to UNK tokens. And another issue is that every OOV word gets the same
representation
Generally, pre-trained models are trained on a large volume of the text corpus leading to huge
vocabularies and hence OOV words don’t usually occur
This opens the door to Character Tokenization.
[page 13]
WHICH TOKENIZATION TO USE ?
Character Tokenization
Character Tokenization splits text into a set of characters.
It overcomes the drawbacks we saw above about Word Tokenization.
Character Tokenizers handles OOV words coherently by preserving the information of the word as
it breaks down the OOV word into characters and represents the word in terms of these characters
It also limits the size of the vocabulary to 26 since the vocabulary contains a unique set of
characters
Drawbacks of Character Tokenization
Character tokens solve the OOV problem but the length of the input and output sentences increases
rapidly as we are representing a sentence as a sequence of characters.
As a result, it becomes challenging to learn the relationship between the characters to form
meaningful words.
[page 14]
STOP-WORDS’ REMOVAL
Stop words are a set of commonly used words in a language.
Examples of stop words in English are “a”, “the”, “is”, “are” and etc.
By removing low information words from text, we can focus on the important words instead.
Punctuations can also be removed at this stage
Example of stop word removal. All stop words are removed completely or replaced with a
dummy character, X:
Input- The students aren’t willing to live in on-campus housing, and they will not
visit their home-towns.
Output- X students X willing X live X on-campus housing, X they X X visit X
home-towns.
[page 15]
RUNNING EXAMPLE
Input:
['Starbucks', 'is', 'doing', 'very', 'well', 'lately', '.']
['Overall', ',', 'while', 'it', 'may', 'seem', 'there', 'is', 'already', 'a', 'Starbucks', 'on',
'every', 'corner', ',', 'Starbucks', 'still', 'has', 'a', 'lot', 'of', 'room', 'to', 'grow', '.']
['They', 'just', 'began', 'expansion', 'into', 'food', 'products', ',', 'which', 'has', 'been',
'going', 'quite', 'well', 'so', 'far', 'for', 'them', '.’]
Output:
['Starbucks', 'well', 'lately’]
['Overall’, 'may’, 'seem’, 'already’, 'Starbucks’, 'every’, 'corner’, 'Starbucks’, 'still’, 'lot’,
'room’, 'grow’]
['began’, 'expansion’, 'food’, 'products’, 'going’, 'quite’, 'well’, 'far' ]
[page 16]
TEXT NORMALIZATION
Natural language, as a human resource, tends to follow the inherent nature of its creator’s randomness.
This means that, as we “produce” natural language, we imprint our random states to it.
When we normalize text, we attempt to reduce its randomness, bringing it closer to a predefined
“standard”.
This helps us to reduce the amount of different information that the computer has to deal with, and
therefore improves efficiency.
The goal of is to reduce inflectional forms and sometimes derivationally related forms of a word to
a common base form.
In linguistics, inflexion is a process of word formation in which a word is modified to express
different grammatical categories such as tense, case, person, number, gender mood etc.
Normalization techniques : stemming and lemmatization
Example:
Mapping words such as “stopwords”, “stop-words” and “stop words” to just “stopword”.
[page 17]
STEMMING
Stemming refers to a crude heuristic process that chops off the ends of words in
the hope of achieving this goal correctly most of the time.
Stemming is the process of reducing a word to its word stem by eliminating
affixes - suffixes, prefixes, infixes, circumfixes from a word.
The result? Stemming a word or sentence may result in words that are not actual
words.
So the words “trouble”, “troubled”, “troubling” and “troubles” might actually be
converted to “troubl”
For example, the words “university” and “universe” that get reduced to “univers”.
(instead of trouble because the ends were just chopped off)
[page 18]
RUNNING EXAMPLE
Input:
['Starbucks', 'well', 'lately’],
['Overall’, 'may’, 'seem’, 'already’, 'Starbucks’, 'every’, 'corner’, 'Starbucks’, 'still’, 'lot’,
'room’, 'grow’],
['began’, 'expansion’, 'food’, 'products’, 'going’, 'quite’, 'well’, 'far’]
Output:
['starbuck well late',
'overal may seem alread starbuck ever corner starbuck still lot room grow',
'began expans food product go quit well far']
[page 19]
LEMMATIZATION
• Lemmatization is very similar to stemming, but the goal is to remove
inflections and map a word to its root form.
• Stemming usually refers to a heuristic process that chops off the ends of
words in the hope of achieving this goal correctly most of the time, and
often includes the removal of derivational affixes.
• Lemmatization usually refers to doing things properly by doing
morphological analysis of words, normally aiming to remove
inflectional endings only and to return the base or dictionary form of a
word, which is known as the lemma .
• Lemmatizer is usually more sophisticated than stemming, since stemmers
works on an individual word without knowledge of the context.
[page 20]
RUNNING EXAMPLE
Input:
['Starbucks', 'well', 'lately’],
['Overall’, 'may’, 'seem’, 'already’, 'Starbucks’, 'every’, 'corner’, 'Starbucks’, 'still’, 'lot’, 'room’,
'grow’],
['began’, 'expansion’, 'food’, 'products’, 'going’, 'quite’, 'well’, 'far’]
Output of Lemmatization:
['starbucks well lately',
'overall may seem already starbucks every corner starbucks still lot room grow',
'began expansion food product going quite well far']
Output of Stemming:
starbuck well late',
'overal may seem alread starbuck ever corner starbuck still lot room grow',
'began expans food product go quit well far']
[page 21]
NOISY TEXT
Noisy text is text with differences between the representation of the text and the intended,
correct, or original text.
Idiomatic expressions, abbreviations, acronyms and business-specific lingo can all cause noisy text.
It is particularly prevalent in the unstructured text found in blog posts, chat conversations, and SMS
messages. Other potential causes include poor spelling and punctuation, typographical errors and
poor translations and speech recognition programs
Noise removal is about removing characters digits and pieces of text that can interfere with the
text analysis. It is also highly domain dependent.
Example, “…trouble…”, “trouble9”, “trouble>”, and “1.trouble” are converted to “trouble”.
[page 22]
PART-OF-SPEECH (POS) TAGGING
POS tagging assigns every word with the corresponding part of speech.
POS tag list:
JJ adjective 'big'
JJR adjective, comparative 'bigger'
JJS adjective, superlative 'biggest'
NN noun, singular 'desk'
NNS noun plural 'desks'
NNP proper noun, singular 'Harrison’
VB verb, base form take
VBD verb, past tense took
VBG verb, gerund/present participle taking
etc…
[page 23]
RUNNING EXAMPLE
Input:
['starbucks well lately',
'overall may seem already starbucks every corner
starbucks still lot room grow',
'began expansion food product going quite well
far’]
Output:
[('starbucks', 'NNS'), ('well', 'RB'), ('lately', 'RB’)]
[('overall', 'JJ'), ('may', 'MD'), ('seem', 'VB'),
('already', 'RB'), ('starbucks', 'NNS'), ('every',
'DT'), ('corner', 'NN'), ('starbucks', 'NNS'), ('still',
'RB'), ('lot', 'VBP'), ('room', 'NN'), ('grow', 'NN’)]
[('they', 'PRP'), ('began', 'VBD'), ('expansion',
'NN'), ('food', 'NN'), ('products', 'NNS'), ('going',
'VBG'), ('quite', 'RB'), ('well', 'RB'), ('far', 'RB')]
POS Tag Meanings
• NNS → Noun, plural
(e.g., “starbucks”, “products”)
• NN → Noun, singular or mass
(e.g., “corner”, “room”, “food”, “expansion”)
• RB → Adverb
(e.g., “well”, “lately”, “already”, “still”, “quite”, “far”)
• JJ → Adjective
(e.g., “overall”)
• MD → Modal verb
(e.g., “may”)
• VB → Verb, base form
(e.g., “seem”)
• VBP → Verb, non-3rd person singular present
VBD → Verb, past tense
(e.g., “began”)
• VBG → Verb, gerund/present participle
(e.g., “going”)
• PRP → Personal pronoun
(e.g., “they”)
• DT → Determiner
(e.g., “every”)
[page 24]
VECTOR SPACE MODEL
● One of the critical components in Natural Language Processing
(NLP) is to encode text information in a numerical format that can
be fed into an NLP model.
● This technique which represents words in a numerical vector
space, is called Vector Space Modelling.
● A typical vector space model has a dimension of V×N
● Where V is number of documents and N is a size of unique
vocabulary
[page 25]
COUNT VECTOR - BAG OF WORDS (BOW)
The bag-of-words model is a simple
representation
In this model, A document (sentence,
paragraph, etc.) is treated as a bag
(multiset) of its words, ignoring:
grammar
word order
syntax
This disregards word order but
keeps multiplicity
The bag-of-words model is commonly
used in document classification where
the frequency of occurrence of each
word is used as a feature
[page 26]
COUNT VECTOR - BAG OF WORDS (BOW)
(1) John likes to watch movies. Mary likes movies too.
(2) Mary also likes to watch football games.
BoW1 = {"John":1,"likes":2,"to":1,"watch":1,"movies":2,"Mary":1,"too":1};
BoW2 = {"Mary":1,"also":1,"likes":1,"to":1,"watch":1,"football":1,"games":1};
Each key is the word, and each value is the number of occurrences of that
word in the given text document.
The order of elements is free, so,
{"too":1,"Mary":1,"movies":2,"John":1,"watch":1,"likes":2,"to":1} is also
equivalent to BoW1.
The order of elements POS tagging and stemming are not inherently part
of Bag of Words (BOW)
BOW only cares about counting words (or frequencies) in documents.
It does not look at grammar, syntax, or meaning — just raw tokens.
[page 27]
COUNT VECTOR - BAG OF WORDS (BOW)
if another document is like a union of these two,
(3) John likes to watch movies. Mary likes movies too. Mary also likes to
watch football games.
BoW3 =
{"John":1,"likes":3,"to":2,"watch":2,"movies":2,"Mary":2,"too":1,"also":1,"f
ootball":1,"games":1};
(1) John likes to watch movies. Mary likes movies too.
(2) Mary also likes to watch football games.
(1) [1, 2, 1, 1, 2, 1, 1, 0, 0, 0]
(2) [0, 1, 1, 1, 0, 1, 0, 1, 1, 1]
[page 28]
COUNT VECTOR – N-GRAM MODEL
The Bag-of-words model is an orderless document representation - only the counts of words matter.
As an alternative, the n-gram model can store this spatial information.
Applying to the same example above, abigram model will parse the text into the following units and store
the term frequency of each unit as before.
(1) John likes to watch movies. Mary likes movies too.
(2) Mary also likes to watch football games.
[ "John likes“, "likes to“, "to watch", "watch movies", "Mary likes", "likes movies", "movies too", ]
(1) [ "John likes“:1, "likes to“:1, "to watch“:1, "watch movies“:1, "Mary likes“:1, "likes movies“:1,
"movies too“:1, ]
(2) ["Mary also“:1, “also likes“:1, "likes to“:1, "to watch“:1, "watch football“:1, "football games“:1]
Conceptually, we can view bag-of-word model as a special case of the n-gram model, with n=1.
For n>1 the model is named n-gram where n denotes the number of grouped words
[page 29]
TF-IDF (TERM FREQUENCY -INVERSE DOCUMENT FREQUENCY)
Here the term-specific weights in the document vectors are products of local and global
parameters.
The weight vector for document d is
The number of times a term occurs in a document is called its term frequency tf(i, j),
Term frequency adjusted for document length: