Sample size selection in optimization methods for machine learning

Sample size selection in optimization methods for machine learning
复制标题

DOI:
10.1007/s10107-012-0572-5
复制
发表时间:
2012-08-01
影响因子:
2.7
通讯作者:
Wu, Yuchen
Wu, Yuchen
中科院分区:
数学2区
文献类型:
--
作者:
Byrd, Richard H.;Chin, Gillian M.;Wu, Yuchen

文献摘要

被引文献

相似文献

本文提出了一种在大规模机器学习问题的批式优化方法中使用不同样本量的方法。本文的第一部分讨论了函数和梯度评价中的动态样本选择的微妙问题。基于在批次梯度计算过程中获得的方差估计,我们提出了一个增加样本大小的准则。我们建立了一个关于梯度法总代价的复杂性界限。本文的第二部分描述了一种实用的牛顿法,它使用比计算函数和梯度更小的样本来计算海森向量积,并且还采用了动态抽样技术。论文的第三部分将重点转移到L(1)--旨在产生稀疏解的正则化问题上。我们提出了一种类牛顿方法,它包括两个阶段:识别零变量的(最小)梯度投影阶段和在自由变量中应用次采样Hessian牛顿迭代的子空间阶段。对语音识别问题的数值测试说明了该算法的性能。
This paper presents a methodology for using varying sample sizes in batch-type optimization methods for large-scale machine learning problems. The first part of the paper deals with the delicate issue of dynamic sample selection in the evaluation of the function and gradient. We propose a criterion for increasing the sample size based on variance estimates obtained during the computation of a batch gradient. We establish an complexity bound on the total cost of a gradient method. The second part of the paper describes a practical Newton method that uses a smaller sample to compute Hessian vector-products than to evaluate the function and the gradient, and that also employs a dynamic sampling technique. The focus of the paper shifts in the third part of the paper to L (1)-regularized problems designed to produce sparse solutions. We propose a Newton-like method that consists of two phases: a (minimalistic) gradient projection phase that identifies zero variables, and subspace phase that applies a subsampled Hessian Newton iteration in the free variables. Numerical tests on speech recognition problems illustrate the performance of the algorithms.