XLM-T: Multilingual Language Models in Twitter for Sentiment Analysis and Beyond

XLM-T: Multilingual Language Models in Twitter for Sentiment Analysis and Beyond
复制标题

DOI:
--
复制
发表时间:
2021-04
期刊:
--
影响因子:
--
通讯作者:
Francesco Barbieri;Luis Espinosa Anke;José Camacho-Collados
Francesco Barbieri;Luis Espinosa Anke;José Camacho-Collados
中科院分区:
其他
文献类型:
--
作者:
Francesco Barbieri;Luis Espinosa Anke;José Camacho-Collados

文献摘要

被引文献

相似文献

语言模型在当前的自然语言处理中无处不在,其多语言能力最近引起了相当大的关注。然而,当前的分析几乎完全集中在标准基准(的多语言变体)上,并且依赖于干净的预训练和特定于任务的语料库作为多语言信号。在本文中,我们介绍了 XLM-T,一种在 Twitter 中训练和评估多语言语言模型的模型。在本文中,我们提供:(1)一个新的强大的多语言基线,由一个 XLM-R(Conneau 等人,2020)模型组成,该模型针对 30 多种语言的数百万条推文进行了预训练,以及随后对目标任务进行微调的起始代码; (2) 一组八种不同语言的统一情感分析 Twitter 数据集以及在此数据集上训练的 XLM-T 模型。
Language models are ubiquitous in current NLP, and their multilingual capacity has recently attracted considerable attention. However, current analyses have almost exclusively focused on (multilingual variants of) standard benchmarks, and have relied on clean pre-training and task-specific corpora as multilingual signals. In this paper, we introduce XLM-T, a model to train and evaluate multilingual language models in Twitter. In this paper we provide: (1) a new strong multilingual baseline consisting of an XLM-R (Conneau et al. 2020) model pre-trained on millions of tweets in over thirty languages, alongside starter code to subsequently fine-tune on a target task; and (2) a set of unified sentiment analysis Twitter datasets in eight different languages and a XLM-T model trained on this dataset.