Adaptation Data Selection using Neural Language Models: Experiments in Machine Translation

Adaptation Data Selection using Neural Language Models: Experiments in Machine Translation
复制标题

DOI:
--
复制
发表时间:
2013-08
期刊:
--
影响因子:
--
通讯作者:
Kevin Duh;Graham Neubig;Katsuhito Sudoh;Hajime Tsukada
Kevin Duh;Graham Neubig;Katsuhito Sudoh;Hajime Tsukada
中科院分区:
其他
文献类型:
--
作者:
Kevin Duh;Graham Neubig;Katsuhito Sudoh;Hajime Tsukada

文献摘要

被引文献

相似文献

数据选择是统计机器翻译领域自适应的有效方法。这个想法是使用在小的域内文本上训练的语言模型,从大的通用域语料库中选择相似的句子,然后将其合并到训练数据中。在以前的工作中已经证明了实质性的收获,这些工作采用了标准的ngram语言模型。在这里,我们探索使用神经语言模型进行数据选择。我们假设神经语言模型中单词的连续向量表示使它们比n-grams更有效地建模未知单词上下文,这在一般领域文本中很普遍。在对4种语言对(英语到德语、法语、俄语、西班牙语)的综合评估中,我们发现神经语言模型确实是数据选择的可行工具:虽然改进是不同的(即BLEU的增益为0.1到1.7),但它们可以快速训练小型域内数据,有时可以大大优于传统的n-grams。
Data selection is an effective approach to domain adaptation in statistical machine translation. The idea is to use language models trained on small in-domain text to select similar sentences from large general-domain corpora, which are then incorporated into the training data. Substantial gains have been demonstrated in previous works, which employ standard ngram language models. Here, we explore the use of neural language models for data selection. We hypothesize that the continuous vector representation of words in neural language models makes them more effective than n-grams for modeling unknown word contexts, which are prevalent in general-domain text. In a comprehensive evaluation of 4 language pairs (English to German, French, Russian, Spanish), we found that neural language models are indeed viable tools for data selection: while the improvements are varied (i.e. 0.1 to 1.7 gains in BLEU), they are fast to train on small in-domain data and can sometimes substantially outperform conventional n-grams.