Obstacle avoidance of redundant manipulators using neural networks based reinforcement learning

Obstacle avoidance of redundant manipulators using neural networks based reinforcement learning
复制标题

DOI:
10.1016/j.rcim.2011.07.004
复制
发表时间:
2012-04-01
影响因子:
10.4
通讯作者:
Mogan, Gheorghe
Mogan, Gheorghe
中科院分区:
计算机科学1区
文献类型:
--
作者:
Duguleana, Mihai;Barbuceanu, Florin Grigore;Mogan, Gheorghe

文献摘要

被引文献

相似文献

本文提出了一种解决冗余机械臂操作任务中避障问题的新方法。所开发的解决方案基于一个使用Q -学习强化技术的双神经网络。Q -学习已应用于机器人技术中,用于实现无障碍导航或计算路径规划问题。大多数研究使用经典雅可比矩阵方法的变体,或通过最小化在已知环境中操作的机械臂的冗余分辨率来解决逆运动学和避障问题。试图使用神经网络解决逆运动学问题的研究人员往往只处理工作区域中存在的一个障碍物。本文重点关注在复杂未知环境中计算逆运动学和避障问题,工作区域中有多个障碍物。Q -学习与神经网络一起使用,以便在每个时刻规划和执行手臂运动。针对一般冗余运动学连杆链开发的算法已在PowerCube机械臂的特定案例上进行了测试。在真实机器人上实施该解决方案之前,模拟被集成在一个沉浸式虚拟环境中,以便更好地进行运动分析和更安全地进行测试。研究结果表明,所提出的方法具有良好的平均速度和令人满意的目标到达成功率。(C) 2011爱思唯尔有限公司。保留所有权利。
This paper proposes a new approach for solving the problem of obstacle avoidance during manipulation tasks performed by redundant manipulators. The developed solution is based on a double neural network that uses Q-learning reinforcement technique. Q-learning has been applied in robotics for attaining obstacle free navigation or computing path planning problems. Most studies solve inverse kinematics and obstacle avoidance problems using variations of the classical Jacobian matrix approach, or by minimizing redundancy resolution of manipulators operating in known environments. Researchers who tried to use neural networks for solving inverse kinematics often dealt with only one obstacle present in the working field. This paper focuses on calculating inverse kinematics and obstacle avoidance for complex unknown environments, with multiple obstacles in the working field. Q-learning is used together with neural networks in order to plan and execute arm movements at each time instant. The algorithm developed for general redundant kinematic link chains has been tested on the particular case of PowerCube manipulator. Before implementing the solution on the real robot, the simulation was integrated in an immersive virtual environment for better movement analysis and safer testing. The study results show that the proposed approach has a good average speed and a satisfying target reaching success rate. (C) 2011 Elsevier Ltd. All rights reserved.