Linear interpolation gives better gradients than Gaussian smoothing in derivative-free optimization

Linear interpolation gives better gradients than Gaussian smoothing in derivative-free optimization
复制标题

在无导数优化中,线性插值比高斯平滑提供更好的梯度

DOI:
--
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
K. Scheinberg
K. Scheinberg
中科院分区:
--
文献类型:
--
作者:
A. Berahas;Liyuan Cao;K. Choromanski;K. Scheinberg

文献摘要

参考文献

被引文献

相似文献

在本文中,我们考虑无导数优化问题,其中的目标函数是光滑的,但计算有一定量的噪声,功能评估是昂贵的,没有导数信息。我们的动机是最近流行的强化学习中的策略优化问题[Choromaski et al. 2018; Fazel et al. 2018; Salimans et al. 2016],并且可以用公式表示为具有上述特征的无导数优化问题。在每一个这些作品的一些近似的梯度构造和(随机)梯度方法。在[Salimans et al. 2016]中,梯度信息沿沿着高斯方向聚合,而在[Choromaski et al. 2018]中,梯度信息沿沿着正交方向计算。我们提供了一个一阶线搜索方法的收敛速度分析,类似于文献中使用的,并推导出梯度近似,确保这种收敛的条件。然后,我们通过严格的方差分析和对强化学习任务的数值比较证明,[Salimans et al. 2016]中使用的高斯采样方法明显劣于[Choromaski et al. 2018]中使用的正交采样以及更一般的插值方法。
In this paper, we consider derivative free optimization problems, where the objective function is smooth but is computed with some amount of noise, the function evaluations are expensive and no derivative information is available. We are motivated by policy optimization problems in reinforcement learning that have recently become popular [Choromaski et al. 2018; Fazel et al. 2018; Salimans et al. 2016], and that can be formulated as derivative free optimization problems with the aforementioned characteristics. In each of these works some approximation of the gradient is constructed and a (stochastic) gradient method is applied. In [Salimans et al. 2016] the gradient information is aggregated along Gaussian directions, while in [Choromaski et al. 2018] it is computed along orthogonal direction. We provide a convergence rate analysis for a first-order line search method, similar to the ones used in the literature, and derive the conditions on the gradient approximations that ensure this convergence. We then demonstrate via rigorous analysis of the variance and by numerical comparisons on reinforcement learning tasks that the Gaussian sampling method used in [Salimans et al. 2016] is significantly inferior to the orthogonal sampling used in [Choromaski et al. 2018] as well as more general interpolation methods.
DOI: --
发表时间: 2018-01
影响因子: 8.7
作者:
Maryam Fazel;Rong Ge;S. Kakade;M. Mesbahi
通讯作者: Maryam Fazel;Rong Ge;S. Kakade;M. Mesbahi