An Evaluation of Two Vocabulary Reduction Methods for Neural Machine Translation

An Evaluation of Two Vocabulary Reduction Methods for Neural Machine Translation
复制标题

神经机器翻译的两种词汇缩减方法的评估

DOI:
--
复制
发表时间:
2018
期刊:
Conference of the Association for Machine Translation in the Americas
影响因子:
--
通讯作者:
Marcello Federico
Marcello Federico
中科院分区:
--
文献类型:
--
作者:
Duygu Ataman;Marcello Federico

文献摘要

被引文献

相似文献

神经机器翻译模型通常使用fi大小的词汇表进行训练,以控制计算复杂性和学习单词表示的质量。然而,这限制了模型的准确性和泛化能力,特别是对于形态丰富的语言,这些语言通常具有非常稀疏的词汇,其中包含fl选择或派生的单词形式中的稀有词汇。一些研究试图通过将单词分割成子词级别的表示并在这一级别上对翻译进行建模来克服这一问题。然而,最近的fi结果表明,如果这些方法在分词过程中中断单词结构,它们可能会导致语义或句法损失,并导致生成不准确的翻译。为了研究这一现象,我们对自然机器翻译中的两种非监督词汇量缩减方法进行了广泛的评价。LMVRST是著名的字节对编码,这是一种统计的子词切分方法,而第二种是基于语言的词汇量缩减,这是一种考虑了子词的形态特征的切分方法。我们在涉及英语和fi以及其他语言(阿拉伯语、捷克语、德语、意大利语和土耳其语)的十个翻译方向上比较了这两种方法,每种方法代表了不同的语系和形态类型。LMVR在大多数语言中都获得了明显更好的性能,表现出与测试语言的词汇稀疏性和词法复杂性成正比的增长。fi。
Neural machine translation (NMT) models are conventionally trained with fixed-size vocabularies to control the computational complexity and the quality of the learned word representations. This, however, limits the accuracy and the generalization capability of the models, especially for morphologically-rich languages, which usually have very sparse vocabularies containing rare inflected or derivated word forms. Some studies tried to overcome this problem by segmenting words into subword level representations and modeling translation at this level. However, recent findings have shown that if these methods interrupt the word structure during segmentation, they might cause semantic or syntactic losses and lead to generating inaccurate translations. In order to investigate this phenomenon, we present an extensive evaluation of two unsupervised vocabulary reduction methods in NMT. The first is the well-known byte-pair-encoding (BPE), a statistical subword segmentation method, whereas the second is linguistically-motivated vocabulary reduction (LMVR), a segmentation method which also considers morphological properties of subwords. We compare both approaches on ten translation directions involving English and five other languages (Arabic, Czech, German, Italian and Turkish), each representing a distinct language family and morphological typology. LMVR obtains significantly better performance in most languages, showing gains proportional to the sparseness of the vocabulary and the morphological complexity of the tested language.