Accelerating Optimization and Reinforcement Learning with Quasi Stochastic Approximation
Accelerating Optimization and Reinforcement Learning with Quasi Stochastic Approximation
复制标题
使用准随机逼近加速优化和强化学习
DOI:
10.23919/acc50511.2021.9482825
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Sean P. Meyn
中科院分区:
文献类型:
--
作者:
Shuhang Chen;Adithya M. Devraj;A. Bernstein;Sean P. Meyn
The paper sets out to obtain precise convergence rates for quasi-stochastic approximation (QSA), with applications to optimization and reinforcement learning. The main contributions are obtained for general nonlinear algorithms, under the assumption that there is a well defined linearization near the optimal parameter $\theta^{\ast}$, with Hurwitz linearization matrix $A^{\ast}$. Subject to stability of the algorithm (general conditions are surveyed in the paper): (i)If the algorithm gain is chosen as $a_{t}=g/(1+t)^{\rho}$ with $g > 0$ and $\rho\in(0,1)$, then a “finite-t” approximation is obtained \begin{equation*} a_{t}^{-1}\{\Theta_{t}-\theta^{\ast}\}=\bar{Y}+\Xi_{t}^{\mathrm{I}}+o(1) \end{equation*} where $\Theta_{t}$ is the parameter estimate, $\bar{Y}\in \mathbb{R}^{d}$ is a vector identified in the paper, and $\{\Xi_{t}^{\mathrm{I}}\}$ is bounded with zero mean. (ii)The approximation continues to hold with $a_{t}=g/(1+t)$ under the stronger assumption that $I+gA^{\ast}$ is Hurwitz. (iii)The Ruppert-Polyak averaging technique is extended to this setting, in which the estimates $\{\Theta_{t}\}$ are obtained using the gain in (i), and $\Theta_{t}^{\mathbf{RP}}$ is defined to be the running average. The convergence rate is $1/t$ if and only if $\bar{Y}=0$. (iv)The theory is illustrated with applications to gradient-free optimization, and policy gradient algorithms for reinforcement learning.
DOI:
--
发表时间:
2020-02
期刊:
ArXiv
影响因子:
--
作者:
Shuhang Chen;Adithya M. Devraj;A. Bušić;Sean P. Meyn
通讯作者:
Shuhang Chen;Adithya M. Devraj;A. Bušić;Sean P. Meyn
DOI:
10.23919/acc45564.2020.9147814
发表时间:
2019-09
期刊:
2020 American Control Conference (ACC)
影响因子:
--
作者:
Yue-Chun Chen;A. Bernstein;Adithya M. Devraj;Sean P. Meyn
通讯作者:
Yue-Chun Chen;A. Bernstein;Adithya M. Devraj;Sean P. Meyn
DOI:
10.1109/cdc40024.2019.9029247
发表时间:
2019
期刊:
Proceedings of the IEEE Conference on Decision Control
影响因子:
--
作者:
Bernstein, Andrey;Chen, Yue;Colombino, Marcello;Dall'Anese, Emiliano;Mehta, Prashant;Meyn, Sean
通讯作者:
Meyn, Sean