# Word2Vec Summary (1) course: Module 3 — Deep Learning & NLP module: Module-3-Deep-Learning-NLP type: pdf source_url: https://personal-learn.armco.dev/files/Module-3-Deep-Learning-NLP/General/Lab_Materials-_19-04-2026/Word2Vec_Summary_(1).pdf pages: 19 --- [page 1] Word Embeddings and Word2Vec [page 2] Word Embeddings and Word2Vec Word2Vec is a shallow, two-layer neural networks which is trained to reconstruct linguistic contexts of words. It takes as input a large corpus of words and produces a vector space, typically of several hundred dimensions, with each unique word in the corpus being assigned a corresponding vector in the space. Word vectors are positioned in the vector space such that words that share common contexts in the corpus are located in close proximity to one another in the space. Word2Vec is a particularly computationally-efficient predictive model for learning word embeddings from raw text. [page 3] Word Embeddings and Word2Vec It comes in two flavors, the Continuous Bag-of-Words (CBOW) model and the Skip-Gram model. Algorithmically, these models are similar. Continuous Bag-of-Words (CBOW) CBOW predicts target words (e.g. ‘mat’) from the surrounding context words (‘the cat sits on the’). Statistically, it has the effect that CBOW smoothes over a lot of the distributional information (by treating an entire context as one observation). For the most part, this turns out to be a useful thing for smaller datasets. [page 4] Word Embeddings and Word2Vec It comes in two flavors, the Continuous Bag-of-Words (CBOW) model and the Skip-Gram model. Algorithmically, these models are similar. Skip-Gram Skip-gram predicts surrounding context words from the target words (inverse of CBOW). Statistically, skip-gram treats each context-target pair as a new observation, and this tends to do better when we have larger datasets. [page 5] Word Embeddings and Word2Vec How does Word2Vec produce word embeddings? Word2Vec uses a trick you may have seen elsewhere in machine learning. Word2Vec is a simple neural network with a single hidden layer, and like all neural networks, it has weights, and during training, its goal is to adjust those weights to reduce a loss function. However, Word2Vec is not going to be used for the task it was trained on, instead, we will just take its hidden weights, use them as our word embeddings, and toss the rest of the model [page 6] Word Embeddings and Word2Vec How does Word2Vec produce word embeddings Contd.? Another place where this trick is used is in unsupervised feature learning, An auto-encoder is trained to compress an input vector in the hidden layer and decompress it back to the original in the output layer. After it’s done training, the output layer (the decompression step) is stripped off and only the hidden layer is used since it has learned good features, It’s a trick for learning good image features without having labeled training data. In a word2vec model, we take a large input vector, compress it down to a smaller dense vector and then instead of decompressing it back to the original input vector we output probabilities of target words. [page 7] Word Embeddings and Word2Vec First of all, we cannot feed a word as string into a neural network. Instead, we feed words as one- hot vectors, which is basically a vector of the same length as the vocabulary, filled with zeros except at the index that represents the word we want to represent, which is assigned “1”. The hidden layer is a standard fully-connected (Dense) layer whose weights are the word embeddings. The output layer outputs probabilities for the target words from the vocabulary. [page 8] Word Embeddings and Word2Vec The rows of the hidden layer weight matrix, are actually the word vectors (word embeddings) we want! The hidden layer operates as a lookup table. The output of the hidden layer is just the “word vector” for the input word. More concretely, if you multiply a 1 x 10,000 one-hot vector by a 10,000 x 300 matrix, it will effectively just select the matrix row corresponding to the ‘1’. [page 9] Word Embeddings and Word2Vec The rows of the hidden layer weight matrix, are actually the word vectors (word embeddings) we want! The end goal of all of this is to learn this hidden layer weight matrix and then toss the output layer when we’re done! The output layer is simply a softmax activation function Technically, If different words are similar in context, then Word2Vec should have similar outputs when these words are passed as inputs, And to have a similar outputs, the computed word vectors (in the hidden layer) for these words have to be similar, Thus Word2Vec is motivated to learn similar word vectors for words in similar context. [page 10] Word Embeddings and Word2Vec Word2Vec is able to capture multiple different degrees of similarity between words, such that semantic and syntactic patterns can be reproduced using vector arithmetic. Patterns such as “Man is to Woman as Brother is to Sister” can be generated through algebraic operations on the vector representations of these words such that the vector representation of “Brother” - ”Man” + ”Woman” produces a result which is closest to the vector representation of “Sister” in the model. Such relationships can be generated for a range of semantic relations (such as Country—Capital) as well as syntactic relations (e.g. present tense—past tense). [page 11] Word Embeddings and Word2Vec A Word2vec model can be trained with negative sampling. Traditionally, predictive models are trained using the maximum likelihood principle to maximize the probability of the next words given the previous words in terms of a softmax function over all the vocabulary words. However, this training procedure is quite computationally expensive given a large vocabulary set, because we need to compute and normalize all vocabulary words at each training step. The same applies for Word2Vec, even though the network is shallow, just 2-layers, it is extremely wide, therefore, some way of reducing computations is required [page 12] Word Embeddings and Word2Vec Negative sampling Negative sampling reduces computation by sampling just N negative instances along with the target word instead of sampling the whole vocabulary. Technically, negative sampling ignores most of the ‘0’ in the one-hot label word vector, and only propagates and updates the weights for the target and a few negative classes which were randomly sampled. More concretely, negative sampling samples negative instances(words) along with the target word and minimizes the log-likelihood of the sampled negative instances while maximizing the log-likelihood of the target word. [page 13] Word Embeddings and Word2Vec The Problem: Full Softmax is Expensive Word2Vec (Skip-gram) wants to predict the probability of a context word given a center word. That requires a softmax over the entire vocabulary: [page 14] Word Embeddings and Word2Vec Negative Sampling Idea Instead of computing probabilities over the whole vocabulary, Word2Vec reframes it as a binary classification problem: For each training pair 𝑤𝑐𝑒𝑛𝑡𝑒𝑟 𝑤𝑐𝑜𝑛𝑡𝑒𝑥𝑡 , treat it as a positive example (label = 1). Then, sample a few "fake" words (negative samples) from the vocabulary and pair them with the same center word — these are negative examples (label = 0). So the model learns: Push embeddings of true pairs closer. Push embeddings of randomly sampled (false) pairs apart. [page 15] Word Embeddings and Word2Vec Why It Works Well Efficiency: Only update a handful of words’ embeddings per step, not the entire vocabulary. Quality: Negative sampling acts like contrastive learning — makes embeddings capture semantic differences. Scalability: Works with very large corpora and vocabularies. In short: Negative sampling turns Word2Vec into a set of binary classification tasks (real vs. noise pairs), avoiding the expensive full softmax and making training practical on huge datasets. [page 16] Word Embeddings and Word2Vec How the negative samples are chosen? The negative samples are chosen using a unigram distribution. Essentially, the probability for selecting a word as a negative sample is related to its frequency, with more frequent words being more likely to be selected as negative samples. Specifically, each word is given a weight equal to it’s frequency (word count) raised to the 3/4 power. The probability for a selecting a word is just it’s weight divided by the sum of weights for all words. [page 17] Word Embeddings and Word2Vec Practical methodology The use of different model parameters and different corpus sizes can greatly affect the quality of a word2vec model. Accuracy can be improved in a number of ways, including the choice of model architecture (CBOW or Skip-Gram), increasing the training data set, increasing the number of vector dimensions, and increasing the window size of words considered by the algorithm. Each of these improvements comes with the cost of increased computational complexity and therefore increased model generation time. In models using large corpora and a high number of dimensions, the skip-gram model yields the highest overall accuracy. However, the CBOW is less computationally expensive and yields similar accuracy results. Accuracy increases overall as the number of words used increase, and as the number of dimensions increases. Doubling the amount of training data results in an equivalent increase in computational complexity as doubling the number of vector dimensions. [page 18] Word Embeddings and Word2Vec Practical methodology – How to reduce training time Sub-sampling Some frequent words often provide little information. Words with frequency above a certain threshold (e.g ‘a’, ‘an’ and ‘that’) may be subsampled to increase training speed and performance. Also, common word pairs or phrases may be treated as single “words” to increase training speed. Dimensionality Quality of word embedding increases with higher dimensionality. However, after reaching some threshold, the marginal gain will diminish. Typically, the dimensionality of the vectors is set to be between 100 and 1,000. [page 19] Word Embeddings and Word2Vec Practical methodology – How to reduce training time Context window The size of the context window determines how many words before and after a given word would be included as context words of the given word. According to the authors’ note, the recommended value is 10 for skip-gram and 5 for CBOW. Here is an example of Skip-Gram with context window of size 2: