Losing Heads in the Lottery: Pruning Transformer Attention in Neural Machine Translation

Losing Heads in the Lottery: Pruning Transformer Attention in Neural Machine Translation
复制标题

DOI:
10.18653/v1/2020.emnlp-main.211
复制
发表时间:
2020-11
期刊:
--
影响因子:
--
通讯作者:
Maximiliana Behnke;Kenneth Heafield
Maximiliana Behnke;Kenneth Heafield
中科院分区:
其他
文献类型:
--
作者:
Maximiliana Behnke;Kenneth Heafield

文献摘要

被引文献

相似文献

注意力机制是Transformer体系结构的关键组件。最近的研究表明,大多数注意力头对他们的决定没有信心,可以修剪。然而,在训练模型之前删除它们会导致质量降低。在本文中,我们应用彩票假设在训练的早期阶段修剪头部。我们在机器翻译上的实验表明,在早期训练中,可以从transformer-big中删除多达四分之三的注意力头部,土耳其语→英语的BLEU平均变化为-0.1。修剪后的模型推理速度是原来的1.5倍,尽管代价是更长的训练时间。我们的方法是对其他方法的补充,例如教师-学生,英语→德语学生模型获得了额外10%的速度提升,去除了75%的编码器注意力和0.2 BLEU损失。
The attention mechanism is the crucial component of the transformer architecture. Recent research shows that most attention heads are not confident in their decisions and can be pruned. However, removing them before training a model results in lower quality. In this paper, we apply the lottery ticket hypothesis to prune heads in the early stages of training. Our experiments on machine translation show that it is possible to remove up to three-quarters of attention heads from transformer-big during early training with an average -0.1 change in BLEU for Turkish→English. The pruned model is 1.5 times as fast at inference, albeit at the cost of longer training. Our method is complementary to other approaches, such as teacher-student, with English→German student model gaining an additional 10% speed-up with 75% encoder attention removed and 0.2 BLEU loss.