A gradient sampling method with complexity guarantees for Lipschitz functions in high and low dimensions

A gradient sampling method with complexity guarantees for Lipschitz functions in high and low dimensions
复制标题

DOI:
--
复制
发表时间:
2021-12
期刊:
--
影响因子:
--
通讯作者:
Damek Davis;D. Drusvyatskiy;Y. Lee;Swati Padmanabhan;Guanghao Ye
Damek Davis;D. Drusvyatskiy;Y. Lee;Swati Padmanabhan;Guanghao Ye
中科院分区:
其他
文献类型:
--
作者:
Damek Davis;D. Drusvyatskiy;Y. Lee;Swati Padmanabhan;Guanghao Ye

文献摘要

被引文献

相似文献

Zhang等人对Goldstein的经典次梯度方法进行了新的改进,保证了最小化Lipschitz函数的效率为$O(\varepsilon^{-4})$。然而,他们的工作使用了一个非标准的亚梯度oracle模型,并且要求函数是方向可微的。在本文中,我们证明了这两个假设都可以通过简单地在算法的每一步中添加一个小的随机扰动来消除。所得到的方法适用于任何Lipschitz函数,其值和梯度可以在可微点处求值。此外,我们还提出了一种新的切割平面算法,该算法在低维情况下具有更好的效率:$O(d\varepsilon^{-3})$用于Lipschitz函数,$O(d\varepsilon^{-2})$用于弱凸函数。
Zhang et al. introduced a novel modification of Goldstein's classical subgradient method, with an efficiency guarantee of $O(\varepsilon^{-4})$ for minimizing Lipschitz functions. Their work, however, makes use of a nonstandard subgradient oracle model and requires the function to be directionally differentiable. In this paper, we show that both of these assumptions can be dropped by simply adding a small random perturbation in each step of their algorithm. The resulting method works on any Lipschitz function whose value and gradient can be evaluated at points of differentiability. We additionally present a new cutting plane algorithm that achieves better efficiency in low dimensions: $O(d\varepsilon^{-3})$ for Lipschitz functions and $O(d\varepsilon^{-2})$ for those that are weakly convex.