InforMask: Unsupervised Informative Masking for Language Model Pretraining

InforMask: Unsupervised Informative Masking for Language Model Pretraining
复制标题

DOI:
10.48550/arxiv.2210.11771
复制
发表时间:
2022-10
期刊:
Comput. Syst. Sci. Eng.
影响因子:
--
通讯作者:
Nafis Sadeq;Canwen Xu;Julian McAuley
Nafis Sadeq;Canwen Xu;Julian McAuley
中科院分区:
其他
文献类型:
--
作者:
Nafis Sadeq;Canwen Xu;Julian McAuley

文献摘要

被引文献

相似文献

掩蔽语言建模广泛用于预训练自然语言理解(NLU)的大型语言模型。然而,随机掩蔽是次优的,为所有令牌分配相等的掩蔽率。在本文中,我们提出了InforMask,一种新的无监督掩蔽策略,用于训练掩蔽语言模型。InforMask利用逐点互信息(PMI)来选择要屏蔽的信息量最大的令牌。我们进一步提出了两个优化InforMask,以提高其效率。通过一次性的预处理步骤,InforMask在事实回忆基准LAMA和问答基准SQuAD v1和v2上优于随机掩蔽和先前提出的掩蔽策略。
Masked language modeling is widely used for pretraining large language models for natural language understanding (NLU). However, random masking is suboptimal, allocating an equal masking rate for all tokens. In this paper, we propose InforMask, a new unsupervised masking strategy for training masked language models. InforMask exploits Pointwise Mutual Information (PMI) to select the most informative tokens to mask. We further propose two optimizations for InforMask to improve its efficiency. With a one-off preprocessing step, InforMask outperforms random masking and previously proposed masking strategies on the factual recall benchmark LAMA and the question answering benchmark SQuAD v1 and v2.