Pre-Training With Whole Word Masking for Chinese BERT

Pre-Training With Whole Word Masking for Chinese BERT
复制标题

DOI:
10.1109/taslp.2021.3124365
复制
发表时间:
2021-01-01
影响因子:
5.4
通讯作者:
Yang, Ziqing
Yang, Ziqing
中科院分区:
计算机科学2区
文献类型:
--
作者:
Cui, Yiming;Che, Wanxiang;Yang, Ziqing

文献摘要

被引文献

相似文献

Transformers 的双向编码器表示 (BERT) 在各种 NLP 任务中显示出惊人的改进,并且已提出其连续变体以进一步提高预训练语言模型的性能。在本文中,我们的目标是首先介绍中文 BERT 的全字掩码(wwm)策略,以及一系列中文预训练语言模型。然后我们还提出了一个简单但有效的模型,称为 MacBERT,它在几个方面改进了 RoBERTa。特别是,我们提出了一种新的掩蔽策略,称为 MLM 作为校正(Mac)。为了证明这些模型的有效性,我们创建了一系列中文预训练语言模型作为基线,包括 BERT、RoBERTa、ELECTRA、RBT 等。我们对十个中文 NLP 任务进行了广泛的实验,以评估创建的中文预训练语言模型以及提出的 MacBERT。实验结果表明,MacBERT 可以在许多 NLP 任务上实现最先进的性能,我们还用一些可能有助于未来研究的发现来消除细节。我们开源我们预先训练的语言模型,以进一步促进我们的研究社区。(1)
Bidirectional Encoder Representations from Transformers (BERT) has shown marvelous improvements across various NLP tasks, and its consecutive variants have been proposed to further improve the performance of the pre-trained language models. In this paper, we aim to first introduce the whole word masking (wwm) strategy for Chinese BERT, along with a series of Chinese pre-trained language models. Then we also propose a simple but effective model called MacBERT, which improves upon RoBERTa in several ways. Especially, we propose a new masking strategy called MLM as correction (Mac). To demonstrate the effectiveness of these models, we create a series of Chinese pre-trained language models as our baselines, including BERT, RoBERTa, ELECTRA, RBT, etc. We carried out extensive experiments on ten Chinese NLP tasks to evaluate the created Chinese pre-trained language models as well as the proposed MacBERT. Experimental results show that MacBERT could achieve state-of-the-art performances on many NLP tasks, and we also ablate details with several findings that may help future research. We open-source our pre-trained language models for further facilitating our research community.(1)