Bayesian Attention Modules

Bayesian Attention Modules
复制标题

DOI:
--
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Xinjie Fan;Shujian Zhang;Bo Chen;Mingyuan Zhou
Xinjie Fan;Shujian Zhang;Bo Chen;Mingyuan Zhou
中科院分区:
其他
文献类型:
--
作者:
Xinjie Fan;Shujian Zhang;Bo Chen;Mingyuan Zhou

文献摘要

被引文献

相似文献

注意力模块作为简单有效的工具,不仅使深度神经网络在许多领域取得了最先进的结果,而且还增强了它们的可解释性。目前大多数模型使用确定性注意模块,由于其简单性和易于优化。另一方面,随机对应物尽管有潜在的好处,但不太受欢迎。主要原因是随机注意力经常引入优化问题或需要显著的模型改变。在本文中,我们提出了一个可扩展的随机版本的注意力,易于实现和优化。我们通过将可重参数化的分布归一化来构造单纯形约束的注意力分布,使训练过程可微。我们在贝叶斯框架中学习它们的参数,其中引入了数据相关的先验来进行正则化。我们将提出的随机注意力模块应用于各种基于注意力的模型,并应用于图节点分类、视觉问答、图像字幕、机器翻译和语言理解。我们的实验表明,所提出的方法带来了一致的改善,在相应的基线。
Attention modules, as simple and effective tools, have not only enabled deep neural networks to achieve state-of-the-art results in many domains, but also enhanced their interpretability. Most current models use deterministic attention modules due to their simplicity and ease of optimization. Stochastic counterparts, on the other hand, are less popular despite their potential benefits. The main reason is that stochastic attention often introduces optimization issues or requires significant model changes. In this paper, we propose a scalable stochastic version of attention that is easy to implement and optimize. We construct simplex-constrained attention distributions by normalizing reparameterizable distributions, making the training process differentiable. We learn their parameters in a Bayesian framework where a data-dependent prior is introduced for regularization. We apply the proposed stochastic attention modules to various attention-based models, with applications to graph node classification, visual question answering, image captioning, machine translation, and language understanding. Our experiments show the proposed method brings consistent improvements over the corresponding baselines.