When BERT Plays the Lottery, All Tickets Are Winning

When BERT Plays the Lottery, All Tickets Are Winning
复制标题

DOI:
10.18653/v1/2020.emnlp-main.259
复制
发表时间:
2020-05
期刊:
--
影响因子:
--
通讯作者:
Sai Prasanna;Anna Rogers;Anna Rumshisky
Sai Prasanna;Anna Rogers;Anna Rumshisky
中科院分区:
其他
文献类型:
--
作者:
Sai Prasanna;Anna Rogers;Anna Rumshisky

文献摘要

被引文献

相似文献

NLP最近的大部分成功都归功于基于transformer的大型模型,如BERT(Devlin et al,2019)。然而,这些模型已被证明可以简化为更少数量的自我注意力头和层。我们从彩票假说的角度来考虑这一现象。对于微调BERT,我们表明,(a)有可能找到一个子网络的元素,实现与完整的模型的性能相当,(B)类似大小的子网络从模型的其余部分进行采样表现较差。然而,“坏”子网络可以单独微调,以实现比“好”子网络略差的性能,这表明预训练BERT中的大多数权重都是潜在有用的。我们还表明,“好”的子网络在不同的GLUE任务中有很大的不同,这为学习BERT在推理时实际使用的知识提供了可能性。
Much of the recent success in NLP is due to the large Transformer-based models such as BERT (Devlin et al, 2019). However, these models have been shown to be reducible to a smaller number of self-attention heads and layers. We consider this phenomenon from the perspective of the lottery ticket hypothesis. For fine-tuned BERT, we show that (a) it is possible to find a subnetwork of elements that achieves performance comparable with that of the full model, and (b) similarly-sized subnetworks sampled from the rest of the model perform worse. However, the "bad" subnetworks can be fine-tuned separately to achieve only slightly worse performance than the "good" ones, indicating that most weights in the pre-trained BERT are potentially useful. We also show that the "good" subnetworks vary considerably across GLUE tasks, opening up the possibilities to learn what knowledge BERT actually uses at inference time.