A Lyapunov characterization of robust policy optimization
A Lyapunov characterization of robust policy optimization
复制标题
鲁棒策略优化的李亚普诺夫表征
DOI:
10.1007/s11768-023-00163-w
复制
发表时间:
2023
影响因子:
1.4
通讯作者:
Jiang, Zhong-Ping
中科院分区:
文献类型:
--
作者:
Cui, Leilei;Jiang, Zhong-Ping
In this paper, we study the robustness property of policy optimization (particularly Gauss–Newton gradient descent algorithm which is equivalent to the policy iteration in reinforcement learning) subject to noise at each iteration. By invoking the concept of input-to-state stability and utilizing Lyapunov’s direct method, it is shown that, if the noise is sufficiently small, the policy iteration algorithm converges to a small neighborhood of the optimal solution even in the presence of noise at each iteration. Explicit expressions of the upperbound on the noise and the size of the neighborhood to which the policies ultimately converge are provided. Based on Willems’ fundamental lemma, a learning-based policy iteration algorithm is proposed. The persistent excitation condition can be readily guaranteed by checking the rank of the Hankel matrix related to an exploration signal. The robustness of the learning-based policy iteration to measurement noise and unknown system disturbances is theoretically demonstrated by the input-to-state stability of the policy iteration. Several numerical simulations are conducted to demonstrate the efficacy of the proposed method.
登录
查看更多内容
影响因子:
8.7
作者:
Maryam Fazel;Rong Ge;S. Kakade;M. Mesbahi
通讯作者:
Maryam Fazel;Rong Ge;S. Kakade;M. Mesbahi
DOI:
--
发表时间:
2021
期刊:
American Control Conference
影响因子:
--
作者:
Leilei Cui;Kaan Özbay;Zhong
通讯作者:
Zhong
DOI:
10.1016/j.automatica.2020.109035
发表时间:
2020-08
期刊:
Autom.
影响因子:
--
作者:
Bo Pang;Zhong-Ping Jiang;I. Mareels
通讯作者:
Bo Pang;Zhong-Ping Jiang;I. Mareels
影响因子:
6.8
作者:
Hesameddin Mohammadi;A. Zare;M. Soltanolkotabi;M. Jovanovi'c
通讯作者:
Hesameddin Mohammadi;A. Zare;M. Soltanolkotabi;M. Jovanovi'c
影响因子:
6.8
作者:
Bo Pang;Zhong-Ping Jiang
通讯作者:
Bo Pang;Zhong-Ping Jiang