Zero-Shot Cross-Lingual Transfer with Meta Learning

Zero-Shot Cross-Lingual Transfer with Meta Learning
复制标题

DOI:
10.18653/v1/2020.emnlp-main.368
复制
发表时间:
2020-03
期刊:
ArXiv
影响因子:
--
通讯作者:
F. Nooralahzadeh;Giannis Bekoulis;Johannes Bjerva;Isabelle Augenstein
F. Nooralahzadeh;Giannis Bekoulis;Johannes Bjerva;Isabelle Augenstein
中科院分区:
其他
文献类型:
--
作者:
F. Nooralahzadeh;Giannis Bekoulis;Johannes Bjerva;Isabelle Augenstein

文献摘要

被引文献

相似文献

最近,学习在任务之间共享什么一直是一个非常重要的话题,因为知识的战略性共享已被证明可以改善下游任务的绩效。这对多语种应用程序尤其重要,因为世界上大多数语言都资源不足。在这里,我们考虑了当英语以外的语言几乎没有数据可用时,同时设置多个不同语言的训练模型。我们证明了这一具有挑战性的设置可以使用元学习来实现,在元学习中,除了训练源语言模型外,另一个模型还学习选择哪些训练实例对第一个最有利。我们在不同的自然语言理解任务(自然语言推理、问题回答)中使用标准监督、零命中跨语言以及少命中跨语言设置进行实验。我们广泛的实验设置证明了元学习在总共15种语言中的一致有效性。我们改进了零激发和少激发NLI(在多NLI和XNLI上)和QA(在MLQA数据集上)的最新水平。综合错误分析表明,语言之间类型特征的相关性可以部分解释通过元学习学到的参数共享是有益的。
Learning what to share between tasks has been a topic of great importance recently, as strategic sharing of knowledge has been shown to improve downstream task performance. This is particularly important for multilingual applications, as most languages in the world are under-resourced. Here, we consider the setting of training models on multiple different languages at the same time, when little or no data is available for languages other than English. We show that this challenging setup can be approached using meta-learning, where, in addition to training a source language model, another model learns to select which training instances are the most beneficial to the first. We experiment using standard supervised, zero-shot cross-lingual, as well as few-shot cross-lingual settings for different natural language understanding tasks (natural language inference, question answering). Our extensive experimental setup demonstrates the consistent effectiveness of meta-learning for a total of 15 languages. We improve upon the state-of-the-art for zero-shot and few-shot NLI (on MultiNLI and XNLI) and QA (on the MLQA dataset). A comprehensive error analysis indicates that the correlation of typological features between languages can partly explain when parameter sharing learned via meta-learning is beneficial.