Convex Programs and Lyapunov Functions for Reinforcement Learning: A Unified Perspective on the Analysis of Value-Based Methods

Convex Programs and Lyapunov Functions for Reinforcement Learning: A Unified Perspective on the Analysis of Value-Based Methods
复制标题

DOI:
10.23919/acc53348.2022.9867291
复制
发表时间:
2022-02
期刊:
2022 American Control Conference (ACC)
影响因子:
--
通讯作者:
Xing-ming Guo;B. Hu
Xing-ming Guo;B. Hu
中科院分区:
其他
文献类型:
--
作者:
Xing-ming Guo;B. Hu

文献摘要

被引文献

相似文献

基于值的方法在马尔可夫决策过程(MDP)和强化学习(RL)中发挥着重要作用。在本文中,我们提出了一个统一的控制理论框架,用于分析基于值的方法,如值计算(VC),值迭代(VI)和时间差(TD)学习(线性函数逼近)。建立在基于值的方法和动态系统之间的内在联系,我们可以直接使用现有的凸测试条件在控制理论中,以获得各种收敛结果的上述基于值的方法。这些测试条件是线性规划(LP)或半定规划(SDP)形式的凸规划,并且可以以简单的方式求解以构造李雅普诺夫函数。我们的分析揭示了反馈控制系统和RL算法之间的一些有趣的联系。我们希望这样的联系可以激发更多的工作在系统/控制理论和RL的交叉点。
Value-based methods play a fundamental role in Markov decision processes (MDPs) and reinforcement learning (RL). In this paper, we present a unified control-theoretic framework for analyzing valued-based methods such as value computation (VC), value iteration (VI), and temporal difference (TD) learning (with linear function approximation). Built upon an intrinsic connection between value-based methods and dynamic systems, we can directly use existing convex testing conditions in control theory to derive various convergence results for the aforementioned value-based methods. These testing conditions are convex programs in form of either linear programming (LP) or semidefinite programming (SDP), and can be solved to construct Lyapunov functions in a straightforward manner. Our analysis reveals some intriguing connections between feedback control systems and RL algorithms. It is our hope that such connections can inspire more work at the intersection of system/control theory and RL.