A Multi-Batch L-BFGS Method for Machine Learning

A Multi-Batch L-BFGS Method for Machine Learning
复制标题

DOI:
--
复制
发表时间:
2016-05
期刊:
ArXiv
影响因子:
--
通讯作者:
A. Berahas;J. Nocedal;Martin Takác
A. Berahas;J. Nocedal;Martin Takác
中科院分区:
其他
文献类型:
--
作者:
A. Berahas;J. Nocedal;Martin Takác

文献摘要

被引文献

相似文献

如何并行化随机梯度下降(SGD)方法一直是文献中关注的问题。在本文中,我们将重点放在批处理方法上,这些方法在每次迭代中使用相当大一部分训练集来促进并行性,并使用二阶信息。为了改进学习过程,我们采用多批处理方法,其中批处理在每次迭代中都会发生变化。这可能会造成困难,因为L-BFGS使用梯度差来更新Hessian近似,并且当使用不同的数据点计算这些梯度时,该过程可能不稳定。本文给出了如何在多批设置下进行稳定的拟牛顿更新,说明了该算法在分布式计算平台上的行为,并研究了其在凸和非凸情况下的收敛性。
The question of how to parallelize the stochastic gradient descent (SGD) method has received much attention in the literature. In this paper, we focus instead on batch methods that use a sizeable fraction of the training set at each iteration to facilitate parallelism, and that employ second-order information. In order to improve the learning process, we follow a multi-batch approach in which the batch changes at each iteration. This can cause difficulties because L-BFGS employs gradient differences to update the Hessian approximations, and when these gradients are computed using different data points the process can be unstable. This paper shows how to perform stable quasi-Newton updating in the multi-batch setting, illustrates the behavior of the algorithm in a distributed computing platform, and studies its convergence properties for both the convex and nonconvex cases.