A new deep Q-learning method with dynamic epsilon adjustment and path planner assisted techniques for Turtlebot mobile robot

A new deep Q-learning method with dynamic epsilon adjustment and path planner assisted techniques for Turtlebot mobile robot
复制标题

DOI:
10.1117/12.2663695
复制
发表时间:
2023-06
期刊:
--
影响因子:
--
通讯作者:
W. Cheng;Zhengbin Ni;Xiangnan Zhong
W. Cheng;Zhengbin Ni;Xiangnan Zhong
中科院分区:
其他
文献类型:
--
作者:
W. Cheng;Zhengbin Ni;Xiangnan Zhong

文献摘要

相似文献

深度Q学习(DQL)方法在自主移动的机器人中已被证明是一个巨大的成功。然而,DQL的例程通常会产生不正确的代理行为(多个原地循环动作),伴随着长时间的训练片段,直到收敛。为了解决这个问题,该项目开发了新的技术,以改善模拟和物理实验中的DQL训练。具体地,集成了动态EpperimentAdjustment方法,以减少非理想代理行为的频率,从而提高控制性能(即,目标速率)。在物理训练过程中,设计了动态窗口法(DWA)全局路径规划器,使智能体能够在固定的训练次数内以较少的碰撞达到更多的目标。GMapping同时定位和地图(SLAM)方法也被应用于提供一个SLAM地图的路径规划器。实验结果表明,我们开发的方法可以显着提高在模拟和物理训练环境中的训练性能。
Deep Q-learning (DQL) method has been proven a great success in autonomous mobile robots. However, the routine of DQL can often yield improper agent behavior (multiple circling-in-place actions) that comes with long training episodes until convergence. To address such problem, this project develops novel techniques that improve DQL training in both simulations and physical experiments. Specifically, the Dynamic Epsilon Adjustment method is integrated to reduce the frequency of non-ideal agent behaviors and therefore improve the control performance (i.e., goal rate). A Dynamic Window Approach (DWA) global path planner is designed in the physical training process so that the agent can reach more goals with less collision within a fixed amount of episodes. The GMapping Simultaneous Localization and Mapping (SLAM) method is also applied to provide a SLAM map to the path planner. The experiment results demonstrate that our developed approach can significantly improve the training performance in both simulation and physical training environment.