A neural probabilistic language model

A neural probabilistic language model
复制标题

DOI:
10.1162/153244303322533223
复制
发表时间:
2003-08-15
影响因子:
6
通讯作者:
Jauvin, C
Jauvin, C
中科院分区:
计算机科学3区
文献类型:
--
作者:
Bengio, Y;Ducharme, R;Jauvin, C

文献摘要

被引文献

相似文献

统计语言建模的一个目标是学习语言中单词序列的联合概率函数。这在本质上是困难的,因为维度的诅咒:将在其上测试模型的单词序列可能不同于在训练期间看到的所有单词序列。基于n元语法的传统但非常成功的方法通过连接在训练集中看到的非常短的重叠序列来获得泛化。我们建议通过学习单词的分布式表示来对抗维度诅咒,该表示允许每个训练句子向模型通知有关语义相邻句子的指数数量。该模型同时学习(1)每个单词的分布式表示以及(2)用这些表示表示的单词序列的概率函数。获得泛化是因为,如果以前从未见过的单词序列由与形成已见句子的单词相似(在具有附近表示的意义上)的单词组成,则该序列获得高概率。在合理的时间内训练如此大的模型(具有数百万个参数)本身就是一个巨大的挑战。我们报告了使用神经网络来实现概率函数的实验,在两个文本语料库上表明,所提出的方法显著改进了最先进的n元语法模型,并且所提出的方法允许利用较长的上下文。
A goal of statistical language modeling is to learn the joint probability function of sequences of words in a language. This is intrinsically difficult because of the curse of dimensionality: a word sequence on which the model will be tested is likely to be different from all the word sequences seen during training. Traditional but very successful approaches based on n-grams obtain generalization by concatenating very short overlapping sequences seen in the training set. We propose to fight the curse of dimensionality by learning a distributed representation for words which allows each training sentence to inform the model about an exponential number of semantically neighboring sentences. The model learns simultaneously (1) a distributed representation for each word along with (2) the probability function for word sequences, expressed in terms of these representations. Generalization is obtained because a sequence of words that has never been seen before gets high probability if it is made of words that are similar (in the sense of having a nearby representation) to words forming an already seen sentence. Training such large models (with millions of parameters) within a reasonable time is itself a significant challenge. We report on experiments using neural networks for the probability function, showing on two text corpora that the proposed approach significantly improves on state-of-the-art n-gram models, and that the proposed approach allows to take advantage of longer contexts.