Provably Faster Algorithms for Bilevel Optimization

Provably Faster Algorithms for Bilevel Optimization
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
--
影响因子:
--
通讯作者:
Junjie Yang;Kaiyi Ji;Yingbin Liang
Junjie Yang;Kaiyi Ji;Yingbin Liang
中科院分区:
其他
文献类型:
--
作者:
Junjie Yang;Kaiyi Ji;Yingbin Liang

文献摘要

相似文献

Bilevel优化已被广泛应用于许多重要的机器学习应用中,例如超参数优化和元学习。最近,已经提出了几种基于动量的算法来更快地解决二聚体优化问题。但是,那些基于动量的算法并没有达到基于SGD的算法的$ \ Mathcal {\ widetilde o}(\ epsilon^{ - 2})$的计算复杂性。在本文中,我们提出了两种新算法进行双光线优化,其中第一种算法采用了基于动量的递归迭代,第二算法在嵌套环中采用递归梯度估计来降低方差。我们表明,这两种算法都达到了$ \ mathcal {\ widetilde o}(\ epsilon^{ - 1.5})$的复杂性,该算法以数量级的顺序优于所有现有算法。我们的实验验证了我们的理论结果,并证明了我们在高参数应用中算法的出色经验性能。
Bilevel optimization has been widely applied in many important machine learning applications such as hyperparameter optimization and meta-learning. Recently, several momentum-based algorithms have been proposed to solve bilevel optimization problems faster. However, those momentum-based algorithms do not achieve provably better computational complexity than $\mathcal{\widetilde O}(\epsilon^{-2})$ of the SGD-based algorithm. In this paper, we propose two new algorithms for bilevel optimization, where the first algorithm adopts momentum-based recursive iterations, and the second algorithm adopts recursive gradient estimations in nested loops to decrease the variance. We show that both algorithms achieve the complexity of $\mathcal{\widetilde O}(\epsilon^{-1.5})$, which outperforms all existing algorithms by the order of magnitude. Our experiments validate our theoretical results and demonstrate the superior empirical performance of our algorithms in hyperparameter applications.