On Learning Language-Invariant Representations for Universal Machine Translation

On Learning Language-Invariant Representations for Universal Machine Translation
复制标题

DOI:
--
复制
发表时间:
2020-07
期刊:
--
影响因子:
--
通讯作者:
Hao Zhao;Junjie Hu;Andrej Risteski
Hao Zhao;Junjie Hu;Andrej Risteski
中科院分区:
其他
文献类型:
--
作者:
Hao Zhao;Junjie Hu;Andrej Risteski

文献摘要

被引文献

相似文献

通用机器翻译的目标是学习在任何语言对之间进行翻译,给定所有语言对的\emph{一小部分}成对翻译文档的语料库。尽管令人印象深刻的实证结果和对大规模多语言模型的兴趣日益浓厚,但对这种通用机器翻译模型造成的翻译错误的理论分析只是新生的。在本文中,我们在一般情况下正式证明了这种努力的某些不可能性,并证明了在存在额外(但自然的)数据结构时的积极结果。对于前者,我们推导了多对多翻译设置中翻译误差的下界,这表明,如果没有对语言结构进行假设,任何旨在学习多个语言对之间共享句子表示的算法都必须在至少一个翻译任务上产生较大的翻译误差。对于后者,我们表明,如果语料库中的成对文档遵循自然的\emph{编码器-解码器}生成过程,我们可以期待一个自然的“泛化”概念:线性数量的语言对,而不是二次的,足以学习一个良好的表示。我们的理论还解释了语言对之间哪种类型的连接图更适合:就每个语言对所需的文档总数而言,路径较长的连接图会导致更差的样本复杂性。我们相信我们的理论见解和启示有助于未来通用机器翻译的算法设计。
The goal of universal machine translation is to learn to translate between any pair of languages, given a corpus of paired translated documents for \emph{a small subset} of all pairs of languages. Despite impressive empirical results and an increasing interest in massively multilingual models, theoretical analysis on translation errors made by such universal machine translation models is only nascent. In this paper, we formally prove certain impossibilities of this endeavour in general, as well as prove positive results in the presence of additional (but natural) structure of data. For the former, we derive a lower bound on the translation error in the many-to-many translation setting, which shows that any algorithm aiming to learn shared sentence representations among multiple language pairs has to make a large translation error on at least one of the translation tasks, if no assumption on the structure of the languages is made. For the latter, we show that if the paired documents in the corpus follow a natural \emph{encoder-decoder} generative process, we can expect a natural notion of ``generalization'': a linear number of language pairs, rather than quadratic, suffices to learn a good representation. Our theory also explains what kinds of connection graphs between pairs of languages are better suited: ones with longer paths result in worse sample complexity in terms of the total number of documents per language pair needed. We believe our theoretical insights and implications contribute to the future algorithmic design of universal machine translation.