Stochastic Gradient Descent in Correlated Settings: A Study on Gaussian Processes

Stochastic Gradient Descent in Correlated Settings: A Study on Gaussian Processes
复制标题

DOI:
--
复制
发表时间:
2020
期刊:
AI Mag.
影响因子:
--
通讯作者:
Hao Chen;Lili Zheng;R. Kontar;Garvesh Raskutti
Hao Chen;Lili Zheng;R. Kontar;Garvesh Raskutti
中科院分区:
其他
文献类型:
--
作者:
Hao Chen;Lili Zheng;R. Kontar;Garvesh Raskutti

文献摘要

被引文献

相似文献

随机梯度下降(SGD)及其变体由于其泛化性能和内在的计算优势,已成为独立样本的大规模机器学习问题的首选算法。然而,随机梯度是具有相关样本的全梯度的有偏估计的事实导致缺乏对SGD在相关设置下如何表现的理论理解,并阻碍了其在这种情况下的使用。在本文中,我们专注于高斯过程(GP),并通过证明minibatch SGD收敛到全损失函数的临界点,并以O(1 K)的速率恢复模型超参数,直到统计误差项取决于minibatch大小,从而向前迈出了一步。对模拟和真实的数据集的数值研究表明,与最先进的GP方法相比,小批量SGD具有更好的泛化能力,同时减少了计算负担,并为GP开辟了一个新的、以前未探索的数据大小机制。
Stochastic gradient descent (SGD) and its variants have established themselves as the go-to algorithms for large-scale machine learning problems with independent samples due to their generalization performance and intrinsic computational advantage. However, the fact that the stochastic gradient is a biased estimator of the full gradient with correlated samples has led to the lack of theoretical understanding of how SGD behaves under correlated settings and hindered its use in such cases. In this paper, we focus on the Gaussian process (GP) and take a step forward towards breaking the barrier by proving minibatch SGD converges to a critical point of the full loss function, and recovers model hyperparameters with rate O( 1 K ) up to a statistical error term depending on the minibatch size. Numerical studies on both simulated and real datasets demonstrate that minibatch SGD has better generalization over state-of-the-art GP methods while reducing the computational burden and opening a new, previously unexplored, data size regime for GPs.