Contextual Dropout: An Efficient Sample-Dependent Dropout Module

Contextual Dropout: An Efficient Sample-Dependent Dropout Module
复制标题

DOI:
--
复制
发表时间:
2021-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Xinjie Fan;Shujian Zhang;Korawat Tanwisuth;Xiaoning Qian;Mingyuan Zhou
Xinjie Fan;Shujian Zhang;Korawat Tanwisuth;Xiaoning Qian;Mingyuan Zhou
中科院分区:
其他
文献类型:
--
作者:
Xinjie Fan;Shujian Zhang;Korawat Tanwisuth;Xiaoning Qian;Mingyuan Zhou

文献摘要

被引文献

相似文献

Dropout已被证明是一个简单有效的模块,不仅可以规范深度神经网络的训练过程,还可以为预测提供不确定性估计。然而,不确定性估计的质量高度依赖于丢弃概率。由于其简单性,大多数当前模型在所有数据样本中使用相同的dropout分布。尽管在建模不确定性的灵活性方面有潜在的收益,但另一方面,样本相关的丢弃很少被探索,因为它经常遇到可扩展性问题或涉及非平凡的模型更改。在本文中,我们提出了一个有效的结构设计的上下文dropout作为一个简单的和可扩展的样本相关的dropout模块,它可以应用于各种各样的模型的代价只有轻微增加的内存和计算成本。我们学习的辍学概率与变分目标,兼容伯努利辍学和高斯辍学。我们将上下文丢弃模块应用于各种模型,并将其应用于图像分类和视觉问答,并展示了该方法在大规模数据集(如ImageNet和VQA 2.0)上的可扩展性。我们的实验结果表明,所提出的方法优于基线方法的准确性和质量的不确定性估计。
Dropout has been demonstrated as a simple and effective module to not only regularize the training process of deep neural networks, but also provide the uncertainty estimation for prediction. However, the quality of uncertainty estimation is highly dependent on the dropout probabilities. Most current models use the same dropout distributions across all data samples due to its simplicity. Despite the potential gains in the flexibility of modeling uncertainty, sample-dependent dropout, on the other hand, is less explored as it often encounters scalability issues or involves non-trivial model changes. In this paper, we propose contextual dropout with an efficient structural design as a simple and scalable sample-dependent dropout module, which can be applied to a wide range of models at the expense of only slightly increased memory and computational cost. We learn the dropout probabilities with a variational objective, compatible with both Bernoulli dropout and Gaussian dropout. We apply the contextual dropout module to various models with applications to image classification and visual question answering and demonstrate the scalability of the method with large-scale datasets, such as ImageNet and VQA 2.0. Our experimental results show that the proposed method outperforms baseline methods in terms of both accuracy and quality of uncertainty estimation.