# Word2Vec%20Summary%20%281%29

course: Module 3 — Deep Learning & NLP
module: Module-3-Deep-Learning-NLP
type: pdf
source_url: https://personal-learn.armco.dev/files/Module-3-Deep-Learning-NLP/General/Lab_Materials-_19-04-2026/Word2Vec%20Summary%20%281%29.pdf
pages: 19

---
[page 1]
Word Embeddings and 
Word2Vec

[page 2]
Word Embeddings and Word2Vec
Word2Vec is a shallow, two-layer neural networks which is trained to 
reconstruct linguistic contexts of words.
It takes as input a large corpus of words and produces a vector space, 
typically of several hundred dimensions, with each unique word in the 
corpus being assigned a corresponding vector in the space.
Word vectors are positioned in the vector space such that words that share 
common contexts in the corpus are located in close proximity to one another 
in the space.
Word2Vec is a particularly computationally-efficient predictive model for learning 
word embeddings from raw text.

[page 3]
Word Embeddings and Word2Vec
It comes in two flavors, the Continuous 
Bag-of-Words (CBOW) model and the 
Skip-Gram model.
Algorithmically, these models are similar.
Continuous Bag-of-Words (CBOW)
CBOW predicts target words (e.g. 
‘mat’) from the surrounding context 
words (‘the cat sits on the’).
Statistically, it has the effect that CBOW 
smoothes over a lot of the distributional 
information (by treating an entire 
context as one observation).
For the most part, this turns out to be a 
useful thing for smaller datasets.

[page 4]
Word Embeddings and Word2Vec
It comes in two flavors, the Continuous 
Bag-of-Words (CBOW) model and the 
Skip-Gram model.
Algorithmically, these models are similar.
Skip-Gram
Skip-gram predicts surrounding 
context words from the target words 
(inverse of CBOW).
Statistically, skip-gram treats each 
context-target pair as a new 
observation, and this tends to do better 
when we have larger datasets.

[page 5]
Word Embeddings and Word2Vec
How does Word2Vec produce word embeddings?
Word2Vec uses a trick you may have seen elsewhere in machine 
learning.
Word2Vec is a simple neural network with a single hidden layer, and 
like all neural networks, it has weights, and during training, its goal is 
to adjust those weights to reduce a loss function. 
However, Word2Vec is not going to be used for the task it was 
trained on, instead, we will just take its hidden weights, use them as 
our word embeddings, and toss the rest of the model

[page 6]
Word Embeddings and Word2Vec
How does Word2Vec produce word embeddings Contd.?
Another place where this trick is used is in unsupervised feature learning, 
An auto-encoder is trained to compress an input vector in the hidden layer and 
decompress it back to the original in the output layer. 
After it’s done training, the output layer (the decompression step) is stripped off 
and only the hidden layer is used since it has learned good features, 
It’s a trick for learning good image features without having labeled training data.
In a word2vec model, we take a large input vector, compress it down to a 
smaller dense vector and then instead of decompressing it back to the 
original input vector we output probabilities of target words.

[page 7]
Word Embeddings and Word2Vec
First of all, we cannot feed a 
word as string into a neural 
network.
Instead, we feed words as one-
hot vectors, which is basically a 
vector of the same length as 
the vocabulary, filled with 
zeros except at the index that 
represents the word we want to 
represent, which is assigned “1”.
The hidden layer is a standard 
fully-connected (Dense) layer 
whose weights are the word 
embeddings.
The output layer outputs 
probabilities for the target 
words from the vocabulary.

[page 8]
Word Embeddings and Word2Vec
The rows of the hidden layer weight matrix, 
are actually the word vectors (word 
embeddings) we want!
The hidden layer operates as a lookup table. 
The output of the hidden layer is just the “word 
vector” for the input word.
More concretely, if you multiply a 1 x 10,000 
one-hot vector by a 10,000 x 300 matrix, it will 
effectively just select the matrix row 
corresponding to the ‘1’.

[page 9]
Word Embeddings and Word2Vec
The rows of the hidden layer weight matrix, are actually the word vectors (word 
embeddings) we want!
The end goal of all of this is to learn this hidden layer weight matrix and then toss the output 
layer when we’re done!
The output layer is simply a softmax activation function
Technically, If different words are similar in context, then Word2Vec should have similar 
outputs when these words are passed as inputs, 
And to have a similar outputs, the computed word vectors (in the hidden layer) for these 
words have to be similar, 
Thus Word2Vec is motivated to learn similar word vectors for words in similar context.

[page 10]
Word Embeddings and Word2Vec
Word2Vec is able to capture multiple different degrees of similarity between words, such that semantic and 
syntactic patterns can be reproduced using vector arithmetic. 
Patterns such as “Man is to Woman as Brother is to Sister” can be generated through algebraic operations on 
the vector representations of these words such that the vector representation of “Brother” - ”Man” + ”Woman” 
produces a result which is closest to the vector representation of “Sister” in the model. 
Such relationships can be generated for a range of semantic relations (such as Country—Capital) as well as 
syntactic relations (e.g. present tense—past tense).

[page 11]
Word Embeddings and Word2Vec
A Word2vec model can be trained with negative 
sampling.
Traditionally, predictive models are trained using 
the maximum likelihood principle to maximize the 
probability of the next words given the previous 
words in terms of a softmax function over all the 
vocabulary words.
However, this training procedure is quite 
computationally expensive given a large vocabulary 
set, because we need to compute and normalize all 
vocabulary words at each training step. 
The same applies for Word2Vec, even though the 
network is shallow, just 2-layers, it is extremely 
wide, therefore, some way of reducing 
computations is required

[page 12]
Word Embeddings and Word2Vec
Negative sampling
Negative sampling reduces computation by 
sampling just N negative instances along 
with the target word instead of sampling the 
whole vocabulary.
Technically, negative sampling ignores most 
of the ‘0’ in the one-hot label word vector, 
and only propagates and updates the 
weights for the target and a few negative 
classes which were randomly sampled.
More concretely, negative sampling samples 
negative instances(words) along with the 
target word and minimizes the log-likelihood 
of the sampled negative instances while 
maximizing the log-likelihood of the target 
word.

[page 13]
Word Embeddings and Word2Vec
The Problem: Full Softmax is Expensive
Word2Vec (Skip-gram) wants to predict the probability of a context word given a 
center word. That requires a softmax over the entire vocabulary:

[page 14]
Word Embeddings and Word2Vec
Negative Sampling Idea
Instead of computing probabilities over the 
whole vocabulary, Word2Vec reframes it as 
a binary classification problem:
For each training pair 𝑤𝑐𝑒𝑛𝑡𝑒𝑟 𝑤𝑐𝑜𝑛𝑡𝑒𝑥𝑡  ,
treat it as a positive example (label = 1).
Then, sample a few "fake" words (negative 
samples) from the vocabulary and pair them 
with the same center word — these are 
negative examples (label = 0).
So the model learns:
Push embeddings of true pairs closer.
Push embeddings of randomly sampled 
(false) pairs apart.

[page 15]
Word Embeddings and Word2Vec
Why It Works Well
Efficiency: Only update a handful of words’ embeddings per step, not the entire 
vocabulary.
Quality: Negative sampling acts like contrastive learning — makes embeddings 
capture semantic differences.
Scalability: Works with very large corpora and vocabularies.
In short: Negative sampling turns Word2Vec into a set of binary classification 
tasks (real vs. noise pairs), avoiding the expensive full softmax and making 
training practical on huge datasets.

[page 16]
Word Embeddings and Word2Vec
How the negative samples are chosen?
The negative samples are chosen using a unigram distribution.
Essentially, the probability for selecting a word as a negative sample 
is related to its frequency, with more frequent words being more 
likely to be selected as negative samples.
Specifically, each word is given a weight equal to it’s frequency 
(word count) raised to the 3/4 power. 
The probability for a selecting a word is just it’s weight divided by the 
sum of weights for all words.

[page 17]
Word Embeddings and Word2Vec
Practical methodology
The use of different model parameters and different corpus sizes can greatly affect the quality of a 
word2vec model.
Accuracy can be improved in a number of ways, including 
the choice of model architecture (CBOW or Skip-Gram), 
increasing the training data set, 
increasing the number of vector dimensions, 
and increasing the window size of words considered by the algorithm. 
Each of these improvements comes with the cost of increased computational complexity and therefore 
increased model generation time.
In models using large corpora and a high number of dimensions, the skip-gram model yields the highest 
overall accuracy.
However, the CBOW is less computationally expensive and yields similar accuracy results.
Accuracy increases overall as the number of words used increase, and as the number of dimensions 
increases. Doubling the amount of training data results in an equivalent increase in computational 
complexity as doubling the number of vector dimensions.

[page 18]
Word Embeddings and Word2Vec
Practical methodology – How to reduce training time
Sub-sampling
Some frequent words often provide little information. 
Words with frequency above a certain threshold (e.g ‘a’, ‘an’ and ‘that’) may be subsampled to increase 
training speed and performance.
Also, common word pairs or phrases may be treated as single “words” to increase training speed.
Dimensionality
Quality of word embedding increases with higher dimensionality. 
However, after reaching some threshold, the marginal gain will diminish. 
Typically, the dimensionality of the vectors is set to be between 100 and 1,000.

[page 19]
Word Embeddings and Word2Vec
Practical methodology – How to reduce training time
Context window
The size of the context window determines how many words before and after a given word would be 
included as context words of the given word. 
According to the authors’ note, the recommended value is 10 for skip-gram and 5 for CBOW.
Here is an example of Skip-Gram with context window of size 2: