Word2Vec: Skip-Gram Feedforward Architecture

Reading through papers on the Word2vec skip-gram model, I found myself confused on a fairly uninteresting point in the mechanics of the output layer. What was never made explicit enough (at least to me) is that the output layer returns the exact same output distribution for each context. To see why this must be true,… Read More Word2Vec: Skip-Gram Feedforward Architecture