From Zero to Hero: On the Limitations of Zero-Shot Language Transfer with Multilingual Transformers

From Zero to Hero: On the Limitations of Zero-Shot Language Transfer with Multilingual Transformers
复制标题

DOI:
10.18653/v1/2020.emnlp-main.363
复制
发表时间:
2020-11
期刊:
--
影响因子:
--
通讯作者:
Anne Lauscher;Vinit Ravishankar;Ivan Vulic;Goran Glavas
Anne Lauscher;Vinit Ravishankar;Ivan Vulic;Goran Glavas
中科院分区:
其他
文献类型:
--
作者:
Anne Lauscher;Vinit Ravishankar;Ivan Vulic;Goran Glavas

文献摘要

被引文献

相似文献

通过语言建模预训练的大规模多语言转换器(MMT)(例如,mBERT,XLM-R)已经成为NLP中零触发语言传输的默认范例,提供了无与伦比的传输性能。然而,目前的评估,验证他们的有效性转移(a)到语言与足够大的预训练语料库,(B)之间接近的语言。在这项工作中,我们分析了下游语言迁移与MMT的局限性,表明,很像跨语言的词嵌入,他们是在资源贫乏的情况下,并为遥远的语言有效性大大降低。我们的实验,包括三个较低级别的任务(词性标注,依存分析,NER)和两个高级别的任务(NLI,QA),经验相关的迁移性能与源语言和目标语言之间的语言接近度,但也与MMT预训练中使用的目标语言语料库的大小。最重要的是,我们证明了廉价的少数拍摄转移(即,在一些目标语言实例上进行额外的微调)是令人惊讶的有效的全面,使更多的研究工作超越了限制零拍摄条件。
Massively multilingual transformers (MMTs) pretrained via language modeling (e.g., mBERT, XLM-R) have become a default paradigm for zero-shot language transfer in NLP, offering unmatched transfer performance. Current evaluations, however, verify their efficacy in transfers (a) to languages with sufficiently large pretraining corpora, and (b) between close languages. In this work, we analyze the limitations of downstream language transfer with MMTs, showing that, much like cross-lingual word embeddings, they are substantially less effective in resource-lean scenarios and for distant languages. Our experiments, encompassing three lower-level tasks (POS tagging, dependency parsing, NER) and two high-level tasks (NLI, QA), empirically correlate transfer performance with linguistic proximity between source and target languages, but also with the size of target language corpora used in MMT pretraining. Most importantly, we demonstrate that the inexpensive few-shot transfer (i.e., additional fine-tuning on a few target-language instances) is surprisingly effective across the board, warranting more research efforts reaching beyond the limiting zero-shot conditions.