Pay Less Attention with Lightweight and Dynamic Convolutions

Pay Less Attention with Lightweight and Dynamic Convolutions
复制标题

DOI:
--
复制
发表时间:
2019-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Felix Wu;Angela Fan;Alexei Baevski;Yann Dauphin;Michael Auli
Felix Wu;Angela Fan;Alexei Baevski;Yann Dauphin;Michael Auli
中科院分区:
其他
文献类型:
--
作者:
Felix Wu;Angela Fan;Alexei Baevski;Yann Dauphin;Michael Auli

文献摘要

被引文献

相似文献

自我注意是构建语言和图像生成模型的有用机制。它通过将每个元素与当前时间步进行比较来确定上下文元素的重要性。在本文中,我们证明了一个非常轻量级的卷积可以与报告的最佳自我注意结果竞争。接下来,我们介绍动态卷积,它比自注意更简单,更有效。我们仅基于当前时间步长预测单独的卷积核,以确定上下文元素的重要性。这种方法所需的操作数量与输入长度成线性关系,而自我注意力是二次的。在大规模机器翻译、语言建模和抽象摘要上的实验表明,动态卷积优于强自注意模型。在WMT'14英语-德语测试集上,动态卷积达到了29.7 BLEU的最新水平。
Self-attention is a useful mechanism to build generative models for language and images. It determines the importance of context elements by comparing each element to the current time step. In this paper, we show that a very lightweight convolution can perform competitively to the best reported self-attention results. Next, we introduce dynamic convolutions which are simpler and more efficient than self-attention. We predict separate convolution kernels based solely on the current time-step in order to determine the importance of context elements. The number of operations required by this approach scales linearly in the input length, whereas self-attention is quadratic. Experiments on large-scale machine translation, language modeling and abstractive summarization show that dynamic convolutions improve over strong self-attention models. On the WMT'14 English-German test set dynamic convolutions achieve a new state of the art of 29.7 BLEU.