word2vec Explained: deriving Mikolov et al.'s negative-sampling word-embedding method

word2vec Explained: deriving Mikolov et al.'s negative-sampling word-embedding method
复制标题

DOI:
--
复制
发表时间:
2014-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Yoav Goldberg;Omer Levy
Yoav Goldberg;Omer Levy
中科院分区:
其他
文献类型:
--
作者:
Yoav Goldberg;Omer Levy

文献摘要

被引文献

相似文献

Tomas Mikolov及其同事(此HTTPS URL)的Word2Vec软件最近获得了很多吸引力,并提供了最新的单词嵌入。该软件背后的学习模型在两篇研究论文中进行了描述。我们发现这些论文中对模型的描述有些神秘且难以遵循。虽然动机和演示对于神经网络的语言模型人群可能很明显,但我们不得不努力弄清楚方程背后的理由。本说明是为了解释Tomas Mikolov,Ilya Sutskever,Kai Chen,Greg Corrado和Jeffrey Dean的“单词和短语分布式表示及其组成性”中的方程式(4)(负抽样)。
The word2vec software of Tomas Mikolov and colleagues (this https URL ) has gained a lot of traction lately, and provides state-of-the-art word embeddings. The learning models behind the software are described in two research papers. We found the description of the models in these papers to be somewhat cryptic and hard to follow. While the motivations and presentation may be obvious to the neural-networks language-modeling crowd, we had to struggle quite a bit to figure out the rationale behind the equations. This note is an attempt to explain equation (4) (negative sampling) in "Distributed Representations of Words and Phrases and their Compositionality" by Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado and Jeffrey Dean.