Continuous space language models using restricted Boltzmann machines

Continuous space language models using restricted Boltzmann machines
复制标题

DOI:
--
复制
发表时间:
2012
期刊:
--
影响因子:
--
通讯作者:
J. Niehues;A. Waibel
J. Niehues;A. Waibel
中科院分区:
其他
文献类型:
--
作者:
J. Niehues;A. Waibel

文献摘要

被引文献

相似文献

提出了一种基于受限Boltzmann机器的统计机器翻译中连续空间语言模型的新方法。N元语法的概率是通过RBM的自由能计算的,而不是前馈神经网络。因此,计算速度要快得多,并且可以集成到翻译过程中,而不是仅在重新排序步骤中使用语言模型。此外,在语言模型中引入额外的单词因素是很简单的。我们观察到,如果包括自动生成的词类作为额外的单词因素,训练中的收敛速度会更快。我们在TED演讲的德语到英语和英语到法语的翻译任务上对基于RBM的语言模型进行了评估。我们没有取代传统的基于n语法的语言模型,而是在更重要但更小的领域数据上训练基于RBM的语言模型,并以对数线性的方式将它们组合在一起。使用这种方法,我们可以在翻译任务上显示出BLEU大约半个百分点的改进。
We present a novel approach for continuous space language models in statistical machine translation by using Restricted Boltzmann Machines (RBMs). The probability of an n-gram is calculated by the free energy of the RBM instead of a feedforward neural net. Therefore, the calculation is much faster and can be integrated into the translation process instead of using the language model only in a re-ranking step. Furthermore, it is straightforward to introduce additional word factors into the language model. We observed a faster convergence in training if we include automatically generated word classes as an additional word factor. We evaluated the RBM-based language model on the German to English and English to French translation task of TED lectures. Instead of replacing the conventional n-grambased language model, we trained the RBM-based language model on the more important but smaller in-domain data and combined them in a log-linear way. With this approach we could show improvements of about half a BLEU point on the translation task.