A Theoretical and Empirical Comparison of Gradient Approximations in Derivative-Free Optimization
A Theoretical and Empirical Comparison of Gradient Approximations in Derivative-Free Optimization
复制标题
DOI:
10.1007/s10208-021-09513-z
复制
发表时间:
2019-05
影响因子:
3
通讯作者:
A. Berahas;Liyuan Cao;K. Choromanski;K. Scheinberg
中科院分区:
文献类型:
--
作者:
A. Berahas;Liyuan Cao;K. Choromanski;K. Scheinberg
In this paper, we analyze several methods for approximating gradients of noisy functions using only function values. These methods include finite differences, linear interpolation, Gaussian smoothing, and smoothing on a sphere. The methods differ in the number of functions sampled, the choice of the sample points, and the way in which the gradient approximations are derived. For each method, we derive bounds on the number of samples and the sampling radius which guarantee favorable convergence properties for a line search or fixed step size descent method. To this end, we use the results in Berahas et al. (Global convergence rate analysis of a generic line search algorithm with noise, arXiv:1910.04055, 2019) and show how each method can satisfy the sufficient conditions, possibly only with some sufficiently large probability at each iteration, as happens to be the case with Gaussian smoothing and smoothing on a sphere. Finally, we present numerical results evaluating the quality of the gradient approximations as well as their performance in conjunction with a line search derivative-free optimization algorithm.