Compositional Morphology for Word Representations and Language Modelling

Compositional Morphology for Word Representations and Language Modelling
复制标题

DOI:
--
复制
发表时间:
2014-05
期刊:
--
影响因子:
--
通讯作者:
Jan A. Botha;Phil Blunsom
Jan A. Botha;Phil Blunsom
中科院分区:
其他
文献类型:
--
作者:
Jan A. Botha;Phil Blunsom

文献摘要

被引文献

相似文献

本文提出了一种可扩展的方法,将组合形态表示集成到基于向量的概率语言模型。我们的方法进行评估的上下文中的日志双线性语言模型,呈现适当有效的机器翻译解码器内的实施因素的词汇。我们进行了内在和外在的评估,展示了一系列语言的结果,这些语言表明我们的模型学习了在单词相似性任务上表现良好的形态表示,并导致困惑的大幅减少。当用于翻译成具有大词汇量的形态丰富的语言时,我们的模型相对于使用退避n-gram模型的基线系统获得了高达1.2个BLEU点的改进。
This paper presents a scalable method for integrating compositional morphological representations into a vector-based probabilistic language model. Our approach is evaluated in the context of log-bilinear language models, rendered suitably efficient for implementation inside a machine translation decoder by factoring the vocabulary. We perform both intrinsic and extrinsic evaluations, presenting results on a range of languages which demonstrate that our model learns morphological representations that both perform well on word similarity tasks and lead to substantial reductions in perplexity. When used for translation into morphologically rich languages with large vocabularies, our models obtain improvements of up to 1.2 BLEU points relative to a baseline system using back-off n-gram models.