Revealing the Dark Secrets of BERT

Revealing the Dark Secrets of BERT
复制标题

DOI:
10.18653/v1/d19-1445
复制
发表时间:
2019-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Olga Kovaleva;Alexey Romanov;Anna Rogers;Anna Rumshisky
Olga Kovaleva;Alexey Romanov;Anna Rogers;Anna Rumshisky
中科院分区:
其他
文献类型:
--
作者:
Olga Kovaleva;Alexey Romanov;Anna Rogers;Anna Rumshisky

文献摘要

被引文献

相似文献

基于BERT的架构目前在许多自然语言处理(NLP)任务上取得了最先进的性能,但对于其成功的确切机制知之甚少。在当前的工作中,我们专注于对自注意力机制的解释,它是BERT的基本底层组件之一。利用GLUE任务的一个子集以及一组精心设计的感兴趣的特征,我们提出了一种方法,并对单个BERT头编码的信息进行了定性和定量分析。我们的研究结果表明,存在一组有限的注意力模式在不同的头之间重复出现,这表明整体模型存在过度参数化的问题。虽然不同的头始终使用相同的注意力模式,但它们对不同任务的性能影响各不相同。我们表明,手动禁用某些头中的注意力会导致在常规微调的BERT模型基础上性能得到提高。
BERT-based architectures currently give state-of-the-art performance on many NLP tasks, but little is known about the exact mechanisms that contribute to its success. In the current work, we focus on the interpretation of self-attention, which is one of the fundamental underlying components of BERT. Using a subset of GLUE tasks and a set of handcrafted features-of-interest, we propose the methodology and carry out a qualitative and quantitative analysis of the information encoded by the individual BERT’s heads. Our findings suggest that there is a limited set of attention patterns that are repeated across different heads, indicating the overall model overparametrization. While different heads consistently use the same attention patterns, they have varying impact on performance across different tasks. We show that manually disabling attention in certain heads leads to a performance improvement over the regular fine-tuned BERT models.