Online, Model-Free Motion Planning in Dynamic Environments: An Intermittent, Finite Horizon Approach with Continuous-Time Q-Learning

Online, Model-Free Motion Planning in Dynamic Environments: An Intermittent, Finite Horizon Approach with Continuous-Time Q-Learning
复制标题

DOI:
10.23919/acc45564.2020.9148047
复制
发表时间:
2020-07
期刊:
2020 American Control Conference (ACC)
影响因子:
--
通讯作者:
George P. Kontoudis;Zirui Xu;K. Vamvoudakis
George P. Kontoudis;Zirui Xu;K. Vamvoudakis
中科院分区:
其他
文献类型:
--
作者:
George P. Kontoudis;Zirui Xu;K. Vamvoudakis

文献摘要

被引文献

相似文献

提出了一种基于Q学习的动态演化环境下的在线运动学运动规划方法。该方法解决了有限时域连续时间最优控制问题,完全未知的系统动态。一个行动者-评论家结构采用沿着与以前的经验的缓冲区,近似的最优策略,并减轻学习信号的要求。该方法配备了终端状态评估,以实现快速导航。路径规划被分配给RRTX。一个障碍物增加和本地重新规划策略负责无碰撞导航。严格的李雅普诺夫为基础的证明,以保证闭环稳定的平衡点。我们评估的有效性的方法与模拟。
This paper presents an online kinodynamic motion planning scheme for dynamically evolving environments, by employing Q-learning. The methodology addresses the finite horizon continuous-time optimal control problem with completely unknown system dynamics. An actor-critic structure is employed along with a buffer of previous experiences, to approximate the optimal policy and alleviate the learning signal requirements. The methodology is equipped with a terminal state evaluation to achieve fast navigation. The path planning is assigned to the RRTX. An obstacle augmentation and a local re-planning strategy are responsible for collision-free navigation. Rigorous Lyapunov-based proofs are provided to guarantee closed-loop stability of the equilibrium point. We evaluate the efficacy of the methodology with simulations.