Understanding Composition of Word Embeddings via Tensor Decomposition

Understanding Composition of Word Embeddings via Tensor Decomposition
复制标题

DOI:
--
复制
发表时间:
2019-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Abraham Frandsen;Rong Ge
Abraham Frandsen;Rong Ge
中科院分区:
其他
文献类型:
--
作者:
Abraham Frandsen;Rong Ge

文献摘要

相似文献

词嵌入是自然语言处理中的一个强有力的工具。本文考虑词的嵌入合成问题,给出两个词的向量表示,计算整个短语的一个向量。我们给出了一个能够捕捉词之间特定句法关系的生成模型。在我们的模型下,我们证明了三个词之间的相关性(由它们的PMI衡量)形成了一个张量,它具有近似的低阶Tucker分解。Tucker分解的结果给出了单词嵌入以及核心张量,该张量可以用于产生更好的单词嵌入的组合。我们还用实验对理论结果进行了补充,验证了我们的假设,并证明了新合成方法的有效性。
Word embedding is a powerful tool in natural language processing. In this paper we consider the problem of word embedding composition \--- given vector representations of two words, compute a vector for the entire phrase. We give a generative model that can capture specific syntactic relations between words. Under our model, we prove that the correlations between three words (measured by their PMI) form a tensor that has an approximate low rank Tucker decomposition. The result of the Tucker decomposition gives the word embeddings as well as a core tensor, which can be used to produce better compositions of the word embeddings. We also complement our theoretical results with experiments that verify our assumptions, and demonstrate the effectiveness of the new composition method.