Cross-lingual text similarity exploiting neural machine translation models

Cross-lingual text similarity exploiting neural machine translation models
复制标题

DOI:
10.1177/0165551520912676
复制
发表时间:
2020-03
影响因子:
2.4
通讯作者:
Kazuhiro Seki
Kazuhiro Seki
中科院分区:
计算机科学3区
文献类型:
--
作者:
Kazuhiro Seki

文献摘要

被引文献

相似文献

本文利用神经机器翻译模型研究跨语言文本相似度。基于机器翻译的一种直接方法是使用翻译文本,以便使问题单语化。另一种可能的方法是使用最近在相关工作中提出的机器翻译模型的中间状态,这可以避免翻译错误的传播。我们的目标是改进这两种方法独立,然后联合收割机的两种类型的信息,即,翻译和中间状态,在一个学习排名的框架来计算跨语言的文本相似度。为了评估我们的方法的有效性和普遍性,我们进行了实证实验,对英日和英印地语翻译语料库的跨语言句子检索任务。结果表明,我们使用翻译和中间状态的方法优于其他基于神经网络的方法,甚至可以与基于最先进的机器翻译系统的强大基线相媲美。
This article studies cross-lingual text similarity using neural machine translation models. A straightforward approach based on machine translation is to use translated text so as to make the problem monolingual. Another possible approach is to use intermediate states of machine translation models as recently proposed in the related work, which could avoid propagation of translation errors. We aim at improving both approaches independently and then combine the two types of information, that is, translations and intermediate states, in a learning-to-rank framework to compute cross-lingual text similarity. To evaluate the effectiveness and generalisability of our approach, we conduct empirical experiments on English–Japanese and English–Hindi translation corpora for a cross-lingual sentence retrieval task. It is demonstrated that our approach using translations and intermediate states outperforms other neural network–based approaches and is even comparable with a strong baseline based on a state-of-the-art machine translation system.