Crosslingual Generalization through Multitask Finetuning

Crosslingual Generalization through Multitask Finetuning
复制标题

通过多任务微调进行跨语言泛化

DOI:
--
复制
发表时间:
2023
期刊:
Annual Meeting of the Association for Computational Linguistics
影响因子:
--
通讯作者:
Colin Raffel
Colin Raffel
中科院分区:
--
文献类型:
--
作者:
Niklas Muennighoff;Thomas Wang;Lintang Sutawika;Adam Roberts;Stella Biderman;Teven Le Scao;M Saiful Bari;Sheng Shen;Zheng;Hailey Schoelkopf;Xiangru Tang;Dragomir R. Radev;Alham Fikri Aji;Khalid Almubarak;Samuel Albanie;Zaid Alyafeai;Albert Webson;Edward Raff;Colin Raffel

文献摘要

被引文献

相似文献

多任务提示微调(MTF)已被证明可以帮助大型语言模型在零触发设置中推广到新任务,但到目前为止,MTF的探索主要集中在英语数据和模型上。我们将MTF应用于预训练的多语言BLOOM和mT5模型家族,以生成称为BLOOMZ和mT 0的微调变体。我们发现,用英语提示对英语任务进行大型多语言语言模型的微调,可以将任务生成到只出现在预训练语料库中的非英语语言。对带有英语提示的多语言任务进行微调,进一步提高了英语和非英语任务的性能,从而获得各种最先进的零触发结果。我们还研究了多语言任务的微调,这些任务的提示已经从英语机器翻译成与每个数据集的语言相匹配。我们发现,对这些机器翻译的提示进行培训,可以在相应语言的人类书面提示上获得更好的性能。令人惊讶的是,我们发现模型能够对他们从未有意见过的语言进行零射击泛化。我们推测模型正在学习更高级别的功能,这些功能与任务和语言无关。此外,我们还介绍了xP 3,这是一个由46种语言的监督数据集组成的组合,带有英语和机器翻译的提示。我们的代码、数据集和模型可在https://github.com/ bigscience-workshop/xmtf上免费获得。
Multitask prompted finetuning (MTF) has been shown to help large language models generalize to new tasks in a zero-shot setting, but so far explorations of MTF have focused on English data and models. We apply MTF to the pretrained multilingual BLOOM and mT5 model families to produce finetuned variants called BLOOMZ and mT0. We find finetuning large multilingual language models on English tasks with English prompts allows for task genrealization to non-English languages that appear only in the pretraining corpus. Finetuning on multilingual tasks with English prompts further improves performance on English and non-English tasks leading to various state-of-the-art zero-shot results. We also investigate finetuning on multilingual tasks with prompts that have been machine-translated from English to match the language of each dataset. We find training on these machine-translated prompts leads to better performance on human-written prompts in the respective languages. Surprisingly, we find models are capable of zero-shot generalization to tasks in languages they have never intentionally seen. We conjecture that the models are learning higher-level capabilities that are both task- and language-agnostic. In addition, we introduce xP3, a composite of supervised datasets in 46 languages with English and machine-translated prompts. Our code, datasets and models are freely available at https://github.com/ bigscience-workshop/xmtf.