BOME! Bilevel Optimization Made Easy: A Simple First-Order Approach

BOME! Bilevel Optimization Made Easy: A Simple First-Order Approach
复制标题

DOI:
10.48550/arxiv.2209.08709
复制
发表时间:
2022-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Mao Ye;B. Liu;S. Wright;Peter Stone;Qian Liu
Mao Ye;B. Liu;S. Wright;Peter Stone;Qian Liu
中科院分区:
其他
文献类型:
--
作者:
Mao Ye;B. Liu;S. Wright;Peter Stone;Qian Liu

文献摘要

相似文献

Bilevel optimization(BO)可用于解决各种重要的机器学习问题,包括但不限于超参数优化、元学习、连续学习和强化学习。传统的BO方法需要通过隐式微分的低级优化过程进行微分,这需要与Hessian矩阵相关的昂贵计算。最近一直在寻求BO的一阶方法,但迄今为止提出的方法对于大规模深度学习应用来说往往是复杂和不切实际的。在这项工作中,我们提出了一个简单的一阶BO算法,它只依赖于一阶梯度信息,不需要隐式微分,对于深度学习中的大规模非凸函数是实用和高效的。我们提供了非渐近收敛分析所提出的方法,非凸目标的稳定点,并提出实证结果,显示其上级的实际性能。
Bilevel optimization (BO) is useful for solving a variety of important machine learning problems including but not limited to hyperparameter optimization, meta-learning, continual learning, and reinforcement learning. Conventional BO methods need to differentiate through the low-level optimization process with implicit differentiation, which requires expensive calculations related to the Hessian matrix. There has been a recent quest for first-order methods for BO, but the methods proposed to date tend to be complicated and impractical for large-scale deep learning applications. In this work, we propose a simple first-order BO algorithm that depends only on first-order gradient information, requires no implicit differentiation, and is practical and efficient for large-scale non-convex functions in deep learning. We provide non-asymptotic convergence analysis of the proposed method to stationary points for non-convex objectives and present empirical results that show its superior practical performance.