Investigating Multilingual NMT Representations at Scale

Investigating Multilingual NMT Representations at Scale
复制标题

DOI:
10.18653/v1/d19-1167
复制
发表时间:
2019-09
期刊:
--
影响因子:
--
通讯作者:
Sneha Kudugunta;Ankur Bapna;Isaac Caswell;N. Arivazhagan;Orhan Firat
Sneha Kudugunta;Ankur Bapna;Isaac Caswell;N. Arivazhagan;Orhan Firat
中科院分区:
其他
文献类型:
--
作者:
Sneha Kudugunta;Ankur Bapna;Isaac Caswell;N. Arivazhagan;Orhan Firat

文献摘要

被引文献

相似文献

多语言神经机器翻译(NMT)模型在迁移学习环境中取得了巨大的经验成功。然而,人们对这些黑箱表示的理解很少,它们的传输模式仍然难以捉摸。在这项工作中,我们试图使用奇异值典型相关分析(SVCCA)来理解大规模多语言NMT表示(103种语言),SVCCA是一种表示相似性框架,允许我们比较不同语言,层和模型的表示。我们的分析验证了几个经验结果和长期的直觉,并揭示了关于表征如何在多语言翻译模型中演变的新观察。我们从分析中得出三个主要结果,并对跨语言迁移学习产生影响:(i)不同语言的编码器表示基于语言相似性聚类,(ii)由编码器学习的源语言的表示依赖于目标语言,反之亦然,以及(iii)当对任意语言对进行微调时,高资源和/或语言相似语言的表示更鲁棒,这对于确定在零或几次激发设置中可以预期多少跨语言迁移是关键的。我们进一步将我们的发现与多语言NMT和迁移学习中现有的经验观察联系起来。
Multilingual Neural Machine Translation (NMT) models have yielded large empirical success in transfer learning settings. However, these black-box representations are poorly understood, and their mode of transfer remains elusive. In this work, we attempt to understand massively multilingual NMT representations (with 103 languages) using Singular Value Canonical Correlation Analysis (SVCCA), a representation similarity framework that allows us to compare representations across different languages, layers and models. Our analysis validates several empirical results and long-standing intuitions, and unveils new observations regarding how representations evolve in a multilingual translation model. We draw three major results from our analysis, with implications on cross-lingual transfer learning: (i) Encoder representations of different languages cluster based on linguistic similarity, (ii) Representations of a source language learned by the encoder are dependent on the target language, and vice-versa, and (iii) Representations of high resource and/or linguistically similar languages are more robust when fine-tuning on an arbitrary language pair, which is critical to determining how much cross-lingual transfer can be expected in a zero or few-shot setting. We further connect our findings with existing empirical observations in multilingual NMT and transfer learning.