XLM-T: A Multilingual Language Model Toolkit for Twitter

XLM-T: A Multilingual Language Model Toolkit for Twitter
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
ArXiv
影响因子:
--
通讯作者:
Francesco Barbieri;Luis Espinosa Anke;José Camacho-Collados
Francesco Barbieri;Luis Espinosa Anke;José Camacho-Collados
中科院分区:
其他
文献类型:
--
作者:
Francesco Barbieri;Luis Espinosa Anke;José Camacho-Collados

文献摘要

被引文献

相似文献

语言模型在当前的NLP中无处不在,其多语言能力最近引起了相当大的关注。然而,目前的分析几乎完全集中在标准基准(的多语言变体)上,并依赖于干净的预训练和特定于任务的语料库作为多语言信号。在本文中,我们介绍了XLM-T,一个框架1使用和评估Twitter的多语言模型。该框架具有两个主要资产:(1)由XLM-R组成的强大的多语言基线(Conneau等人,2020)模型在超过30种语言的数百万条推文上进行了预训练,以及随后对目标任务进行微调的启动代码;以及(2)一组统一的艾德情绪分析Twitter数据集,采用八种不同的语言。这是一个模块化的框架,可以很容易地扩展到其他任务,以及与最近的努力,也旨在Twitter特定的数据集的同质化(Barbieri等人,2020年)。
Language models are ubiquitous in current NLP, and their multilingual capacity has recently attracted considerable attention. However, current analyses have almost exclusively focused on (multilingual variants of) standard benchmarks, and have relied on clean pre-training and task-specific corpora as multilingual signals. In this paper, we introduce XLM-T, a framework 1 for using and evaluating multilingual language models in Twitter. This framework features two main assets: (1) a strong multilingual baseline consisting of an XLM-R (Conneau et al., 2020) model pre-trained on millions of tweets in over thirty languages, alongside starter code to subsequently fine-tune on a target task; and (2) a set of unified sentiment analysis Twitter datasets in eight different languages. This is a modular framework that can easily be extended to additional tasks, as well as integrated with recent efforts also aimed at the homogenization of Twitter-specific datasets (Barbieri et al., 2020).