RoBERTa: A Robustly Optimized BERT Pretraining Approach

RoBERTa: A Robustly Optimized BERT Pretraining Approach
复制标题

DOI:
10.48550/arxiv.1907.11692
复制
发表时间:
2019-07-26
影响因子:
4.9
通讯作者:
Stoyanov, Veselin
Stoyanov, Veselin
中科院分区:
管理学3区
文献类型:
--
作者:
Liu, Yinhan;Ott, Myle;Stoyanov, Veselin

文献摘要

被引文献

相似文献

语言模型预训练带来了显着的性能提升,但不同方法之间的仔细比较是具有挑战性的。训练在计算上是昂贵的,通常在不同大小的私有数据集上进行,并且,正如我们将展示的那样,超参数的选择对最终结果有重大影响。我们提出了BERT预训练的复制研究(Devlin等人,2019年),仔细衡量了许多关键超参数和训练数据大小的影响。我们发现BERT明显训练不足,并且可以匹配或超过在它之后发布的每个模型的性能。我们最好的模型在GLUE,RACE和SQuAD上达到了最先进的结果。这些结果凸显了之前被忽视的设计选择的重要性,并引发了对最近报告的改进来源的质疑。我们发布模型和代码。1
Language model pretraining has led to significant performance gains but careful comparison between different approaches is challenging. Training is computationally expensive, often done on private datasets of different sizes, and, as we will show, hyperparameter choices have significant impact on the final results. We present a replication study of BERT pretraining (Devlin et al., 2019) that carefully measures the impact of many key hyperparameters and training data size. We find that BERT was significantly undertrained, and can match or exceed the performance of every model published after it. Our best model achieves state-of-the-art results on GLUE, RACE and SQuAD. These results highlight the importance of previously overlooked design choices, and raise questions about the source of recently reported improvements. We release our models and code.1