A MATHEMATICAL ANALYSIS OF ACTOR-CRITIC ARCHITECTURES FOR LEARNING OPTIMAL CONTROLS THROUGH INCREMENTAL DYNAMIC PROGRAMMING (cid:3)
A MATHEMATICAL ANALYSIS OF ACTOR-CRITIC ARCHITECTURES FOR LEARNING OPTIMAL CONTROLS THROUGH INCREMENTAL DYNAMIC PROGRAMMING (cid:3)
复制标题
通过增量动态编程学习最优控制的 Actor-Critic 架构的数学分析 (cid:3)
DOI:
--
复制
发表时间:
1990
期刊:
影响因子:
--
通讯作者:
Iii Leemon C. Baird
中科院分区:
文献类型:
--
作者:
Ronald J. Williams;Iii Leemon C. Baird
Combining elements of the theory of dynamic pro-grammingwith features appropriate for on-line learning has led to an approach Watkins has called incremental dynamic programming. Here we adopt this incremental dynamic programming point of view and obtain some preliminary mathematical results relevant to understanding the capabilities and limitations of actor-critic learning systems. Examples of such systems are Samuel's learning checker player, Hol-land's bucket brigade algorithm, Witten's adaptive controller, and the adaptive heuristic critic algorithm of Barto, Sutton, and Anderson. Particular emphasis here is on the e(cid:11)ect of complete asynchrony in the updating of the actor and the critic across individual states or state-action pairs. The main results are that, while convergence to optimal performance is not guaranteed in general, there are a number of situations in which such convergence is assured.