On the information bottleneck theory of deep learning

On the information bottleneck theory of deep learning
复制标题

DOI:
10.1088/1742-5468/ab3985
复制
发表时间:
2019-12-01
影响因子:
2.4
通讯作者:
Cox, David D.
Cox, David D.
中科院分区:
物理与天体物理3区
文献类型:
--
作者:
Saxe, Andrew M.;Bansal, Yamini;Cox, David D.

文献摘要

被引文献

相似文献

深度神经网络的实际成功并没有与理论进展相匹配,理论进展可以令人满意地解释它们的行为。在这项工作中,我们研究了深度学习的信息瓶颈(IB)理论,该理论提出了三个具体的主张:首先,深度网络经历了两个不同的阶段,包括初始拟合阶段和随后的压缩阶段;其次,压缩阶段与深度网络的出色泛化性能有因果关系;第三,压缩阶段由于随机梯度下降的类似扩散的行为而发生。在这里,我们表明,这些主张在一般情况下都不成立,而是反映了在确定性网络中计算有限互信息度量的假设。当使用简单的分箱计算时,我们通过分析结果和模拟的组合证明,在先前的工作中观察到的信息平面轨迹主要是所采用的神经非线性的函数:双边饱和非线性如双曲正切在神经激活进入饱和状态时产生压缩相位,但线性激活函数和单侧饱和非线性(如广泛使用的ReLU)实际上并不如此。此外,我们发现压缩和泛化之间没有明显的因果关系:不压缩的网络仍然能够泛化,反之亦然。接下来,我们通过证明我们可以使用全批量梯度下降而不是随机梯度下降来复制IB结果,从而证明压缩阶段(当它存在时)并不是由训练中的随机性引起的。最后,我们表明,当输入域由任务相关和任务无关信息的子集组成时,隐藏表示确实会压缩任务无关信息,尽管有关输入的整体信息可能会随着训练时间单调增加,并且这种压缩与拟合过程同时发生,而不是在随后的压缩期间。
The practical successes of deep neural networks have not been matched by theoretical progress that satisfyingly explains their behavior. In this work, we study the information bottleneck (IB) theory of deep learning, which makes three specific claims: first, that deep networks undergo two distinct phases consisting of an initial fitting phase and a subsequent compression phase; second, that the compression phase is causally related to the excellent generalization performance of deep networks; and third, that the compression phase occurs due to the diffusion-like behavior of stochastic gradient descent. Here we show that none of these claims hold true in the general case, and instead reflect assumptions made to compute a finite mutual information metric in deterministic networks. When computed using simple binning, we demonstrate through a combination of analytical results and simulation that the information plane trajectory observed in prior work is predominantly a function of the neural nonlinearity employed: double-sided saturating nonlinearities like tanh yield a compression phase as neural activations enter the saturation regime, but linear activation functions and single-sided saturating nonlinearities like the widely used ReLU in fact do not. Moreover, we find that there is no evident causal connection between compression and generalization: networks that do not compress are still capable of generalization, and vice versa. Next, we show that the compression phase, when it exists, does not arise from stochasticity in training by demonstrating that we can replicate the IB findings using full batch gradient descent rather than stochastic gradient descent. Finally, we show that when an input domain consists of a subset of task-relevant and task-irrelevant information, hidden representations do compress the task-irrelevant information, although the overall information about the input may monotonically increase with training time, and that this compression happens concurrently with the fitting process rather than during a subsequent compression period.