Transition Point Dynamic Programming

Transition Point Dynamic Programming
复制标题

过渡点动态规划

DOI:
--
复制
发表时间:
1993
期刊:
Neural Information Processing Systems
影响因子:
--
通讯作者:
P. Lawrence
P. Lawrence
中科院分区:
--
文献类型:
--
作者:
K. Buckland;P. Lawrence

文献摘要

被引文献

相似文献

过渡点动态规划(Transition Point Dynamic Programming,TPDP)是一种基于记忆的强化学习直接动态规划方法,用于自适应最优控制,可以减少连续随机动态系统控制所需的学习时间和内存使用。TPDP通过确定一组理想的过渡点(TP)来实现这一点,这些过渡点仅指定最优控制所需的控制动作变化。TPDP收敛到一个理想的TP集,通过使用Q学习的变化,以评估整个状态空间的状态添加,交换和删除TP的优点。当应用于赛道问题时,TPDP比传统的Q学习更快地学习最优控制策略,并且能够使用更少的内存。
Transition point dynamic programming (TPDP) is a memory-based, reinforcement learning, direct dynamic programming approach to adaptive optimal control that can reduce the learning time and memory usage required for the control of continuous stochastic dynamic systems. TPDP does so by determining an ideal set of transition points (TPs) which specify only the control action changes necessary for optimal control. TPDP converges to an ideal TP set by using a variation of Q-learning to assess the merits of adding, swapping and removing TPs from states throughout the state space. When applied to a race track problem, TPDP learned the optimal control policy much sooner than conventional Q-learning, and was able to do so using less memory.