Analyzing the Generalization Capability of SGLD Using Properties of Gaussian Channels

Analyzing the Generalization Capability of SGLD Using Properties of Gaussian Channels
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Hao Wang;Yizhe Huang;Rui Gao;F. Calmon
Hao Wang;Yizhe Huang;Rui Gao;F. Calmon
中科院分区:
其他
文献类型:
--
作者:
Hao Wang;Yizhe Huang;Rui Gao;F. Calmon

文献摘要

相似文献

优化是训练机器学习模型的关键组成部分,对其泛化有很大影响。在本文中,我们考虑一种特殊的优化方法——随机梯度朗之万动力学(SGLD)算法——并研究由 SGLD 训练的模型的泛化。我们通过将 SGLD 与信息和通信理论中发现的高斯通道联系起来,得出了一个新的泛化界限。我们的界限可以根据训练数据计算出来,并结合梯度的方差来量化损失景观的特定类型的“清晰度”。我们还考虑了一种与 SGLD 密切相关的算法,即差分私有 SGD (DP-SGD)。我们证明了 DP-SGD 的泛化能力可以通过迭代来增强。具体来说,如果 DP-SGD 算法输出最后一次迭代,同时隐藏其他迭代,则可以通过包含时间衰减因子来锐化我们的边界。这种衰减因子使得早期迭代对我们的界限的贡献随着时间的推移而减少,并且是通过强大的数据处理不等式(信息论的基本工具)建立的。我们通过数值实验证明了我们的界限,表明它可以预测真实泛化差距的行为。
Optimization is a key component for training machine learning models and has a strong impact on their generalization. In this paper, we consider a particular optimization method—the stochastic gradient Langevin dynamics (SGLD) algorithm—and investigate the generalization of models trained by SGLD. We derive a new generalization bound by connecting SGLD with Gaussian channels found in information and communication theory. Our bound can be computed from the training data and incorporates the variance of gradients for quantifying a particular kind of “sharpness” of the loss landscape. We also consider a closely related algorithm with SGLD, namely differentially private SGD (DP-SGD). We prove that the generalization capability of DP-SGD can be amplified by iteration. Specifically, our bound can be sharpened by including a time-decaying factor if the DP-SGD algorithm outputs the last iterate while keeping other iterates hidden. This decay factor enables the contribution of early iterations to our bound to reduce with time and is established by strong data processing inequalities—a fundamental tool in information theory. We demonstrate our bound through numerical experiments, showing that it can predict the behavior of the true generalization gap.