Implicit Bias in Deep Linear Classification: Initialization Scale vs Training Accuracy

Implicit Bias in Deep Linear Classification: Initialization Scale vs Training Accuracy
复制标题

DOI:
--
复制
发表时间:
2020-07
期刊:
ArXiv
影响因子:
--
通讯作者:
E. Moroshko;Suriya Gunasekar;Blake E. Woodworth;J. Lee;N. Srebro;Daniel Soudry
E. Moroshko;Suriya Gunasekar;Blake E. Woodworth;J. Lee;N. Srebro;Daniel Soudry
中科院分区:
其他
文献类型:
--
作者:
E. Moroshko;Suriya Gunasekar;Blake E. Woodworth;J. Lee;N. Srebro;Daniel Soudry

文献摘要

被引文献

相似文献

我们提供了一个详细的渐近研究梯度流轨迹及其隐式优化偏差当最小化指数损失在“对角线性网络”。这是显示“内核”和非内核(“丰富”或“活动”)制度之间转换的最简单模型。我们展示了如何通过初始化规模和最小化训练损失之间的关系来控制过渡。我们的结果表明,梯度下降的一些极限行为只有在荒谬的训练精度(远远超过$10^{-100}$)时才会出现。此外,在合理的初始化尺度和训练精度下的隐式偏差更为复杂,不受这些限制的影响。
We provide a detailed asymptotic study of gradient flow trajectories and their implicit optimization bias when minimizing the exponential loss over "diagonal linear networks". This is the simplest model displaying a transition between "kernel" and non-kernel ("rich" or "active") regimes. We show how the transition is controlled by the relationship between the initialization scale and how accurately we minimize the training loss. Our results indicate that some limit behaviors of gradient descent only kick in at ridiculous training accuracies (well beyond $10^{-100}$). Moreover, the implicit bias at reasonable initialization scales and training accuracies is more complex and not captured by these limits.