AfroLM: A Self-Active Learning-based Multilingual Pretrained Language Model for 23 African Languages

AfroLM: A Self-Active Learning-based Multilingual Pretrained Language Model for 23 African Languages
复制标题

AfroLM:基于自主学习的 23 种非洲语言的多语言预训练语言模型

DOI:
--
复制
发表时间:
2022
期刊:
SUSTAINLP
影响因子:
--
通讯作者:
Chris C. Emezue
Chris C. Emezue
中科院分区:
--
文献类型:
--
作者:
Bonaventure F. P. Dossou;A. Tonja;Oreen Yousuf;Salomey Osei;Abigail Oppong;Iyanuoluwa Shode;Oluwabusayo Olufunke Awoyomi;Chris C. Emezue

文献摘要

被引文献

相似文献

近年来,多语言预训练语言模型因其在众多下游自然语言处理任务(NLP)中的出色表现而得到了广泛关注。然而,对这些大型多语种语言模型进行预训练需要大量的训练数据,这对于非洲语言来说是不可用的。主动学习是一种半监督学习算法,在这种算法中,模型始终如一地动态学习,以识别最有益的样本来训练自己,从而在下游任务中实现更好的优化和性能。此外,主动学习有效而实际地解决了现实世界的数据稀缺问题。尽管主动学习有很多好处,但在NLP的背景下,特别是在多语言模型预训练中,它很少得到考虑。在本文中,我们介绍了AfroLM,这是一个使用我们新颖的自主动学习框架从头开始对23种非洲语言(迄今为止最大的努力)进行预训练的多语言语言模型。在比现有基线小14倍的数据集上进行预训练,AfroLM在各种NLP下游任务(NER、文本分类和情感分析)上优于许多多语言预训练语言模型(AfriBERTa、XLMR-base、mBERT)。另外的域外情感分析实验表明,AfroLM能够很好地泛化各个域。我们在https://github.com/bonaventuredossou/MLM_AL上发布了代码源代码和框架中使用的数据集。
In recent years, multilingual pre-trained language models have gained prominence due to their remarkable performance on numerous downstream Natural Language Processing tasks (NLP). However, pre-training these large multilingual language models requires a lot of training data, which is not available for African Languages. Active learning is a semi-supervised learning algorithm, in which a model consistently and dynamically learns to identify the most beneficial samples to train itself on, in order to achieve better optimization and performance on downstream tasks. Furthermore, active learning effectively and practically addresses real-world data scarcity. Despite all its benefits, active learning, in the context of NLP and especially multilingual language models pretraining, has received little consideration. In this paper, we present AfroLM, a multilingual language model pretrained from scratch on 23 African languages (the largest effort to date) using our novel self-active learning framework. Pretrained on a dataset significantly (14x) smaller than existing baselines, AfroLM outperforms many multilingual pretrained language models (AfriBERTa, XLMR-base, mBERT) on various NLP downstream tasks (NER, text classification, and sentiment analysis). Additional out-of-domain sentiment analysis experiments show that AfroLM is able to generalize well across various domains. We release the code source, and our datasets used in our framework at https://github.com/bonaventuredossou/MLM_AL.