A Reinforcement Learning-Based Strategy of Path Following for Snake Robots with an Onboard Camera.

A Reinforcement Learning-Based Strategy of Path Following for Snake Robots with an Onboard Camera.
复制标题

DOI:
10.3390/s22249867
复制
发表时间:
2022-12-15
期刊:
Sensors (Basel, Switzerland)
影响因子:
--
通讯作者:
Fang Y
Fang Y
中科院分区:
其他
文献类型:
--
作者:
Liu L;Guo X;Fang Y

文献摘要

参考文献

被引文献

相似文献

对于蛇机器人的路径跟踪,许多基于模型的控制器都表现出了很强的跟踪能力。然而,令人满意的表现往往依赖于精确的建模和简化的假设。此外,视觉感知对于自主闭环控制也是必不可少的,这使得蛇机器人的路径跟踪变得更加具有挑战性。为此,设计了一种新的基于强化学习的递阶控制框架,使带有车载摄像头的蛇机器人能够实现自主的自我定位和路径跟踪。具体地说,首先对路径跟踪策略进行分层训练,将RL算法和步态知识很好地结合起来。在此基础上,充分优化了训练效率,大大提高了控制策略的路径跟踪性能,无需任何额外的训练即可在实际的蛇形机器人上实现。随后,为了促进路径跟踪过程中的视觉自定位,在训练路径跟踪策略的奖励函数中增加了视觉定位稳定项,使蛇形机器人在运动过程中具有平稳的转向能力,从而保证了视觉定位的准确性,便于实际应用。对比仿真和实验结果表明,所提出的分层路径跟踪控制方法在收敛速度和跟踪精度方面具有较好的性能。
For path following of snake robots, many model-based controllers have demonstrated strong tracking abilities. However, a satisfactory performance often relies on precise modelling and simplified assumptions. In addition, visual perception is also essential for autonomous closed-loop control, which renders the path following of snake robots even more challenging. Hence, a novel reinforcement learning-based hierarchical control framework is designed to enable a snake robot with an onboard camera to realize autonomous self-localization and path following. Specifically, firstly, a path following policy is trained in a hierarchical manner, in which the RL algorithm and gait knowledge are well combined. On this basis, the training efficiency is sufficiently optimized, and the path following performance of the control policy is greatly improved, which can then be implemented on a practical snake robot without any additional training. Subsequently, in order to promote visual self-localization during path following, a visual localization stabilization item is added to the reward function that trains the path following strategy, which endows a snake robot with smooth steering ability during locomotion, thereby guaranteeing the accuracy of visual localization and facilitating practical applications. Comparative simulations and experimental results are illustrated to exhibit the superior performance of the proposed hierarchical path following the control method in terms of convergence speed and tracking accuracy.
DOI: 10.1109/tcyb.2021.3055519
发表时间: 2021-03-30
影响因子: 11.8
作者:
Cao, Zhengcai;Zhang, Dong;Zhou, MengChu
通讯作者: Zhou, MengChu
DOI: 10.1126/scirobotics.abc5986
发表时间: 2020-10-21
期刊: SCIENCE ROBOTICS
影响因子: 25
作者:
Lee, Joonho;Hwangbo, Jemin;Hutter, Marco
通讯作者: Hutter, Marco
DOI: 10.1177/0278364920987859
发表时间: 2021-04-01
影响因子: 9.2
作者:
Ibarz, Julian;Tan, Jie;Levine, Sergey
通讯作者: Levine, Sergey
DOI: 10.1109/tcst.2015.2467208
发表时间: 2016-05-01
影响因子: 4.8
作者:
Mohammadi, Alireza;Rezapour, Ehsan;Pettersen, Kristin Y.
通讯作者: Pettersen, Kristin Y.
使用强化学习的最优自主控制:调查
DOI: 10.1109/tnnls.2017.2773458
发表时间: 2018-06-01
影响因子: 10.4
作者:
Kiumarsi, Bahare;Vamvoudakis, Kyriakos G.;Lewis, Frank L.
通讯作者: Lewis, Frank L.