Information Dropout: Learning Optimal Representations Through Noisy Computation

Information Dropout: Learning Optimal Representations Through Noisy Computation
复制标题

DOI:
10.1109/tpami.2017.2784440
复制
发表时间:
2018-12-01
影响因子:
23.6
通讯作者:
Soatto, Stefano
Soatto, Stefano
中科院分区:
计算机科学1区
文献类型:
--
作者:
Achille, Alessandro;Soatto, Stefano

文献摘要

被引文献

相似文献

深度学习中常用的交叉熵损失与最优表示的定义属性密切相关,但不能强制执行一些关键属性。我们表明,这可以通过添加正则化项来解决,这反过来又与在深度神经网络的激活中注入乘法噪声有关,其中的一个特殊情况是dropout的常见做法。我们证明了我们的正则化损失函数可以有效地使用信息Dropout最小化,信息Dropout是一种基于信息论原理的Dropout的推广,可以自动适应数据,并且可以更好地利用有限容量的架构。当任务是重建输入时,我们表明我们的损失函数产生变分自编码器作为特殊情况,从而提供了表征学习,信息论和变分推理之间的联系。最后,我们证明了我们可以通过强制因式先验来促进最优解纠缠表示的创建,这一事实在最近的工作中已经得到了经验观察。我们的实验验证了我们方法背后的理论直觉,我们发现Information Dropout实现了与二进制Dropout相当或更好的泛化性能,特别是在较小的模型上,因为它可以自动地使噪声适应网络的结构,以及测试样本。
The cross-entropy loss commonly used in deep learning is closely related to the defining properties of optimal representations, but does not enforce some of the key properties. We show that this can be solved by adding a regularization term, which is in turn related to injecting multiplicative noise in the activations of a Deep Neural Network, a special case of which is the common practice of dropout. We show that our regularized loss function can be efficiently minimized using Information Dropout, a generalization of dropout rooted in information theoretic principles that automatically adapts to the data and can better exploit architectures of limited capacity. When the task is the reconstruction of the input, we show that our loss function yields a Variational Autoencoder as a special case, thus providing a link between representation learning, information theory and variational inference. Finally, we prove that we can promote the creation of optimal disentangled representations simply by enforcing a factorized prior, a fact that has been observed empirically in recent work. Our experiments validate the theoretical intuitions behind our method, and we find that Information Dropout achieves a comparable or better generalization performance than binary dropout, especially on smaller models, since it can automatically adapt the noise to the structure of the network, as well as to the test sample.