Pay Attention to MLPs

Pay Attention to MLPs
复制标题

DOI:
--
复制
发表时间:
2021-05
期刊:
--
影响因子:
--
通讯作者:
Hanxiao Liu;Zihang Dai;David R. So;Quoc V. Le
Hanxiao Liu;Zihang Dai;David R. So;Quoc V. Le
中科院分区:
其他
文献类型:
--
作者:
Hanxiao Liu;Zihang Dai;David R. So;Quoc V. Le

文献摘要

被引文献

相似文献

Transformer已经成为深度学习中最重要的架构创新之一,并在过去几年中实现了许多突破。在这里,我们提出了一个简单的网络架构,gMLP,基于MLP与门控,并表明它可以执行以及变压器在关键的语言和视觉应用。我们的比较表明,自我注意力对视觉转换器来说并不重要,因为gMLP可以达到相同的精度。对于BERT,我们的模型在预训练困惑度上与Transformers不相上下,并且在一些下游NLP任务上更好。在gMLP性能较差的微调任务中,使gMLP模型大幅增大可以缩小与Transformer的差距。总的来说,我们的实验表明,gMLP可以在增加的数据和计算上与Transformers一样扩展。
Transformers have become one of the most important architectural innovations in deep learning and have enabled many breakthroughs over the past few years. Here we propose a simple network architecture, gMLP, based on MLPs with gating, and show that it can perform as well as Transformers in key language and vision applications. Our comparisons show that self-attention is not critical for Vision Transformers, as gMLP can achieve the same accuracy. For BERT, our model achieves parity with Transformers on pretraining perplexity and is better on some downstream NLP tasks. On finetuning tasks where gMLP performs worse, making the gMLP model substantially larger can close the gap with Transformers. In general, our experiments show that gMLP can scale as well as Transformers over increased data and compute.