The Flip Side of the Reweighted Coin: Duality of Adaptive Dropout and Regularization

The Flip Side of the Reweighted Coin: Duality of Adaptive Dropout and Regularization
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
--
影响因子:
--
通讯作者:
Daniel LeJeune;Hamid Javadi;Richard Baraniuk
Daniel LeJeune;Hamid Javadi;Richard Baraniuk
中科院分区:
其他
文献类型:
--
作者:
Daniel LeJeune;Hamid Javadi;Richard Baraniuk

文献摘要

相似文献

稀疏深度(神经)网络最成功的方法是在整个训练过程中自适应屏蔽网络权重的方法。通过检查线性情况下的这种掩蔽或丢失,我们通过所谓的“$\eta$-trick”揭示了这种自适应方法和正则化之间的二元性,该技巧将两者都视为迭代重新加权优化。我们证明,任何以单调方式适应权重的 dropout 策略都对应于有效的次二次正则化惩罚,因此会导致稀疏解。我们获得了几种流行的稀疏化策略的有效惩罚,这与稀疏优化中常用的经典惩罚非常相似。以变​​分 dropout 为案例研究,我们在深度网络稀疏化任务中展示了自适应 dropout 方法和经典方法之间类似的经验行为,验证了我们的理论。
Among the most successful methods for sparsifying deep (neural) networks are those that adaptively mask the network weights throughout training. By examining this masking, or dropout, in the linear case, we uncover a duality between such adaptive methods and regularization through the so-called"$\eta$-trick"that casts both as iteratively reweighted optimizations. We show that any dropout strategy that adapts to the weights in a monotonic way corresponds to an effective subquadratic regularization penalty, and therefore leads to sparse solutions. We obtain the effective penalties for several popular sparsification strategies, which are remarkably similar to classical penalties commonly used in sparse optimization. Considering variational dropout as a case study, we demonstrate similar empirical behavior between the adaptive dropout method and classical methods on the task of deep network sparsification, validating our theory.