Word Representations

Word Representations
复制标题

单词表示

DOI:
10.1007/978-981-13-0062-2_3
复制
发表时间:
2018
期刊:
Neural Representations of Natural Language
影响因子:
--
通讯作者:
Bennamoun
Bennamoun
中科院分区:
--
文献类型:
--
作者:
Lyndon White;R. Togneri;Wei Liu;Bennamoun

文献摘要

被引文献

相似文献

单词嵌入是将机器学习带入自然语言处理前沿的核心创新。本章讨论如何创建一个数字向量来捕捉单词的显著特征(例如,语义)。讨论从经典的语言建模问题开始。通过解决这个问题,使用基于神经网络的方法,创建了单词嵌入。技术,如CBOW和跳格模型(Word2vec),以及在将其与共同位置上的常见线性代数约简联系起来方面的更新进展,如所讨论的。本章还详细讨论了经常令人困惑的分层Softmax和负抽样技术。最后简要介绍了其他一些应用程序和相关技术。
Word embeddings are the core innovation that has brought machine learning to the forefront of natural language processing. This chapter discusses how one can create a numerical vector that captures the salient features (e.g. semantic meaning) of a word. Discussion begins with the classic language modelling problem. By solving this, using a neural network-based approach, word-embeddings are created. Techniques such as CBOW and skip-gram models (word2vec), and more recent advances in relating this to common linear algebraic reductions on co-locations as discussed. The chapter also includes a detailed discussion of the often confusing hierarchical softmax, and negative sampling techniques. It concludes with a brief look at some other applications and related techniques.