A MATHEMATICAL ANALYSIS OF ACTOR-CRITIC ARCHITECTURES FOR LEARNING OPTIMAL CONTROLS THROUGH INCREMENTAL DYNAMIC PROGRAMMING (cid:3)

A MATHEMATICAL ANALYSIS OF ACTOR-CRITIC ARCHITECTURES FOR LEARNING OPTIMAL CONTROLS THROUGH INCREMENTAL DYNAMIC PROGRAMMING (cid:3)
复制标题

通过增量动态编程学习最优控制的 Actor-Critic 架构的数学分析 (cid:3)

DOI:
--
复制
发表时间:
1990
期刊:
--
影响因子:
--
通讯作者:
Iii Leemon C. Baird
Iii Leemon C. Baird
中科院分区:
--
文献类型:
--
作者:
Ronald J. Williams;Iii Leemon C. Baird

文献摘要

被引文献

相似文献

将动态编程理论的要素与适合在线学习的功能相结合,形成了沃特金斯称为增量动态编程的方法。在这里,我们采用这种增量动态规划的观点,并获得了一些与理解演员批评学习系统的能力和局限性相关的初步数学结果。此类系统的示例包括 Samuel 的学习跳棋玩家、Hol-land 的水桶旅算法、Witten 的自适应控制器以及 Barto、Sutton 和 Anderson 的自适应启发式批评算法。这里特别强调的是在个体状态或状态-行动对的行为者和批评者的更新中完全异步的影响。主要结果是,虽然一般不能保证收敛到最佳性能,但在许多情况下可以保证这种收敛。
Combining elements of the theory of dynamic pro-grammingwith features appropriate for on-line learning has led to an approach Watkins has called incremental dynamic programming. Here we adopt this incremental dynamic programming point of view and obtain some preliminary mathematical results relevant to understanding the capabilities and limitations of actor-critic learning systems. Examples of such systems are Samuel's learning checker player, Hol-land's bucket brigade algorithm, Witten's adaptive controller, and the adaptive heuristic critic algorithm of Barto, Sutton, and Anderson. Particular emphasis here is on the e(cid:11)ect of complete asynchrony in the updating of the actor and the critic across individual states or state-action pairs. The main results are that, while convergence to optimal performance is not guaranteed in general, there are a number of situations in which such convergence is assured.