Transition Point Dynamic Programming
Transition Point Dynamic Programming
复制标题
过渡点动态规划
DOI:
--
复制
发表时间:
1993
期刊:
影响因子:
--
通讯作者:
P. Lawrence
中科院分区:
文献类型:
--
作者:
K. Buckland;P. Lawrence
Transition point dynamic programming (TPDP) is a memory-based, reinforcement learning, direct dynamic programming approach to adaptive optimal control that can reduce the learning time and memory usage required for the control of continuous stochastic dynamic systems. TPDP does so by determining an ideal set of transition points (TPs) which specify only the control action changes necessary for optimal control. TPDP converges to an ideal TP set by using a variation of Q-learning to assess the merits of adding, swapping and removing TPs from states throughout the state space. When applied to a race track problem, TPDP learned the optimal control policy much sooner than conventional Q-learning, and was able to do so using less memory.