On Cross-Lingual Text Similarity Using Neural Translation Models

On Cross-Lingual Text Similarity Using Neural Translation Models
复制标题

DOI:
10.2197/ipsjjip.27.315
复制
发表时间:
2019
期刊:
J. Inf. Process.
影响因子:
--
通讯作者:
Kazuhiro Seki
Kazuhiro Seki
中科院分区:
其他
文献类型:
--
作者:
Kazuhiro Seki

文献摘要

被引文献

相似文献

准确计算两个不同语言文本之间的相似度在跨语言信息检索和跨语言文本挖掘/分析等应用中具有巨大的价值。本文基于神经网络对这一重要问题进行了研究。具体来说,我们的重点是神经机器翻译模型。在使用翻译模型时,我们特别注意的不是翻译本身,而是存储在翻译模型中的给定文本的中间状态。我们的假设是,中间状态捕捉输入文本的句法和语义意义,是一个很好的代表性的文本,避免不可避免的翻译错误。为了研究这一假设的有效性,我们研究了中间状态的效用及其在计算跨语言文本相似性方面的有效性,并与其他基于神经网络的文本分布式表示方法(包括基于单词和段落嵌入的方法)进行了比较。我们证明了使用中间状态的方法不仅优于这些方法,而且也是一个强大的机器推理为基础的。此外,它表明,中间状态和翻译文本的工作互补对方,尽管事实上,他们是从相同的NMT模型产生的。
Accurately computing the similarity between two texts written in different languages has tremendous value in many applications, such as cross-lingual information retrieval and cross-lingual text mining/analytics. This paper studies the important problem based on neural networks. Specifically, our focus is on the neural machine translation models. While translation models are utilized, we pay special attention not to the translation itself but to the intermediate states of given texts stored in the translation models. Our assumption is that the intermediate states capture the syntactic and semantic meaning of input texts and are a good representation of the texts, avoiding inevitable translation errors. To study the validity of the assumption, we investigate the utility of the intermediates states and their effectiveness in computing cross-lingual text similarity in comparison with other neural network-based distributed representations of texts, including word and paragraph embedding-based approaches. We demonstrate that an approach using the intermediate states outperforms not only these approaches but also a strong machine translation-based one. Furthermore, it is revealed that intermediate states and translated texts work complementarily each other despite the fact that they are generated from the same NMT models.