Fuzzy Logic-Driven Variable Time-Scale Prediction-Based Reinforcement Learning for Robotic Multiple Peg-in-Hole Assembly

Fuzzy Logic-Driven Variable Time-Scale Prediction-Based Reinforcement Learning for Robotic Multiple Peg-in-Hole Assembly
复制标题

模糊逻辑驱动的基于可变时间尺度预测的机器人多钉孔装配强化学习

DOI:
10.1109/tase.2020.3024725
复制
发表时间:
2022-01-01
影响因子:
5.6
通讯作者:
Xu,Jing
Xu,Jing
中科院分区:
计算机科学1区
文献类型:
--
作者:
Hou,Zhimin;Li,Zhihu;Xu,Jing

文献摘要

相似文献

强化学习(RL)已越来越多地用于单个钉孔装配,其中装配技能是通过与装配环境的交互以类似于人类所采用的技能的方式来学习的。然而,现有的强化学习算法很难应用于多钉孔装配,因为装配环境要复杂得多,需要充分的探索,导致训练时间长,数据效率低。为此,本文重点研究了如何预测装配环境,以及如何将预测环境用于装配动作控制,以提高强化学习算法的数据效率。具体来说,首先,装配环境是准确预测的可变时间尺度预测(VTSP)定义为一般价值函数(GVF),减少了不必要的探索。其次,我们提出了一个模糊逻辑驱动的基于可变时间尺度预测的强化学习(FLDVTSP-RL)的装配动作控制,以提高RL算法的效率,其中预测的环境映射到建议的阻抗动作空间中的阻抗参数的模糊逻辑系统(FLS)作为行动基线。为了证明VTSP的有效性和FLDVTSP-RL方法的数据效率,建立了一个双钉孔组装实验;结果表明,FLDVTSP-deep Q-learning(DQN)与DQN相比,组装时间减少了约44%; FLDVTSP-deep deterministic policy gradient(DDPG)与DDPG相比,组装时间减少了约24%。从业人员注意事项-多销孔组件的复杂装配环境导致无法从力传感器准确识别接触状态。因此,需要基于接触状态识别来调整控制参数的基于接触模型的方法不能直接应用于这种复杂的环境中。最近,没有接触状态识别的强化学习(RL)方法最近引起了科学界的兴趣。然而,现有的强化学习方法仍然依赖于大量的探索和较长的训练时间,不能直接应用于现实世界的任务。本文的灵感来自于人类可以通过几次试验学习装配技能的方式,这种方式依赖于环境的可变时间尺度预测(VTSP)和优化的装配动作控制策略。我们提出的模糊逻辑驱动的基于可变时间尺度预测的强化学习(FLDVTSP-RL)可以分两步实现。首先,装配环境预测的VTSP定义为一般价值函数(GVF)。其次,装配动作控制是在一个阻抗动作空间中实现的,该空间具有由模糊逻辑系统(FLS)从预测环境映射的阻抗参数定义的基线。最后,进行了双钉孔组装实验,与深度Q学习(DQN)相比,FLDVTSP-DQN可以减少约44%的组装时间;与深度确定性策略梯度(DDPG)相比,FLDVTSP-DDPG可以减少约24%的组装时间。
Reinforcement learning (RL) has been increasingly used for single peg-in-hole assembly, where assembly skill is learned through interaction with the assembly environment in a manner similar to skills employed by human beings. However, the existing RL algorithms are difficult to apply to the multiple peg-in-hole assembly because the much more complicated assembly environment requires sufficient exploration, resulting in a long training time and less data efficiency. To this end, this article focuses on how to predict the assembly environment and how to use the predicted environment in assembly action control to improve the data efficiency of the RL algorithm. Specifically, first, the assembly environment is exactly predicted by a variable time-scale prediction (VTSP) defined as general value functions (GVFs), reducing the unnecessary exploration. Second, we propose a fuzzy logic-driven variable time-scale prediction-based reinforcement learning (FLDVTSP-RL) for assembly action control to improve the efficiency of the RL algorithm, in which the predicted environment is mapped to the impedance parameter in the proposed impedance action space by a fuzzy logic system (FLS) as the action baseline. To demonstrate the effectiveness of VTSP and the data efficiency of the FLDVTSP-RL methods, a dual peg-in-hole assembly experiment is set up; the results show that FLDVTSP-deep Q-learning (DQN) decreases the assembly time about 44% compared with DQN and FLDVTSP-deep deterministic policy gradient (DDPG) decreases the assembly time about 24% compared with DDPG. Note to Practitioners—The complicated assembly environment of the multiple peg-in-hole assembly results in a contact state that cannot be recognized exactly from the force sensor. Therefore, contact-model-based methods that require tuning of the control parameters based on the contact state recognition cannot be applied directly in this complicated environment. Recently, reinforcement learning (RL) methods without contact state recognition have recently attracted scientific interest. However, the existing RL methods still rely on numerous explorations and a long training time, which cannot be directly applied to real-world tasks. This article takes inspiration from the manner in which human beings can learn assembly skills with a few trials, which relies on the variable time-scale predictions (VTSPs) of the environment and the optimized assembly action control strategy. Our proposed fuzzy logic-driven variable time-scale prediction-based reinforcement learning (FLDVTSP-RL) can be implemented in two steps. First, the assembly environment is predicted by the VTSP defined as general value functions (GVFs). Second, assembly action control is realized in an impedance action space with a baseline defined by the impedance parameter mapped from the predicted environment by the fuzzy logic system (FLS). Finally, a dual peg-in-hole assembly experiment is conducted; compared with deep Q-learning (DQN), FLDVTSP-DQN can decrease the assembly time about 44%; compared with deep deterministic policy gradient (DDPG), FLDVTSP-DDPG can decrease the assembly time about 24%.