Low-Precision Stochastic Gradient Langevin Dynamics

Low-Precision Stochastic Gradient Langevin Dynamics
复制标题

DOI:
10.48550/arxiv.2206.09909
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Ruqi Zhang;A. Wilson;Chris De Sa
Ruqi Zhang;A. Wilson;Chris De Sa
中科院分区:
其他
文献类型:
--
作者:
Ruqi Zhang;A. Wilson;Chris De Sa

文献摘要

相似文献

虽然低精度优化已被广泛用于加速深度学习,但低精度采样在很大程度上仍未被探索。因此,采样在许多大规模场景中根本不可行,尽管它为神经网络的泛化和不确定性估计提供了显着的好处。在本文中,我们提供了低精度随机梯度朗之万动力学(SGLD)的第一次研究,表明其成本可以显着降低而不牺牲性能,由于其固有的能力,以处理系统噪声。我们证明了在强凸条件下,低精度SGLD与全精度梯度卷积的收敛性受量化误差的影响比SGD对应的收敛性小。为了进一步实现低精度梯度平滑,我们为SGLD开发了一个新的量化函数,它在每个更新步骤中都保留了方差。我们证明了低精度SGLD在各种深度学习任务中仅使用8位即可实现与全精度SGLD相当的性能。
While low-precision optimization has been widely used to accelerate deep learning, low-precision sampling remains largely unexplored. As a consequence, sampling is simply infeasible in many large-scale scenarios, despite providing remarkable benefits to generalization and uncertainty estimation for neural networks. In this paper, we provide the first study of low-precision Stochastic Gradient Langevin Dynamics (SGLD), showing that its costs can be significantly reduced without sacrificing performance, due to its intrinsic ability to handle system noise. We prove that the convergence of low-precision SGLD with full-precision gradient accumulators is less affected by the quantization error than its SGD counterpart in the strongly convex setting. To further enable low-precision gradient accumulators, we develop a new quantization function for SGLD that preserves the variance in each update step. We demonstrate that low-precision SGLD achieves comparable performance to full-precision SGLD with only 8 bits on a variety of deep learning tasks.