Inducing Language-Agnostic Multilingual Representations

Inducing Language-Agnostic Multilingual Representations
复制标题

DOI:
10.18653/v1/2021.starsem-1.22
复制
发表时间:
2020-08
期刊:
--
影响因子:
--
通讯作者:
Wei Zhao;Steffen Eger;Johannes Bjerva;Isabelle Augenstein
Wei Zhao;Steffen Eger;Johannes Bjerva;Isabelle Augenstein
中科院分区:
其他
文献类型:
--
作者:
Wei Zhao;Steffen Eger;Johannes Bjerva;Isabelle Augenstein

文献摘要

被引文献

相似文献

跨语言表示有可能使NLP技术适用于世界上绝大多数语言。然而,它们目前需要大量的预训练语料库或类型学上相似的语言。在这项工作中,我们通过从多语言嵌入中去除语言身份信号来解决这些障碍。为此,我们研究了三种方法:(i)将目标语言(全部)的向量空间重新对齐到一个枢轴源语言;(ii)去除语言特定的手段和差异,从而产生更好的嵌入判别作为副产品;(3)通过去除词形缩略和句子重新排序来增加语言之间的输入相似性。我们评估了19种不同类型语言的XNLI和无参考MT评估。我们的研究结果揭示了这些方法的局限性——与矢量规范化不同,矢量空间重新对齐和文本规范化不能在编码器和语言之间实现一致的增益。然而,由于这些方法的加性效应,它们的组合在所有任务和语言中平均减少了8.9分(m-BERT)和18.2分(XLM-R)的跨语言迁移差距。
Cross-lingual representations have the potential to make NLP techniques available to the vast majority of languages in the world. However, they currently require large pretraining corpora or access to typologically similar languages. In this work, we address these obstacles by removing language identity signals from multilingual embeddings. We examine three approaches for this: (i) re-aligning the vector spaces of target languages (all together) to a pivot source language; (ii) removing language-specific means and variances, which yields better discriminativeness of embeddings as a by-product; and (iii) increasing input similarity across languages by removing morphological contractions and sentence reordering. We evaluate on XNLI and reference-free MT evaluation across 19 typologically diverse languages. Our findings expose the limitations of these approaches—unlike vector normalization, vector space re-alignment and text normalization do not achieve consistent gains across encoders and languages. Due to the approaches’ additive effects, their combination decreases the cross-lingual transfer gap by 8.9 points (m-BERT) and 18.2 points (XLM-R) on average across all tasks and languages, however.