Distributed Continuous Control with Meta Learning on Robotic Arms

Distributed Continuous Control with Meta Learning on Robotic Arms
复制标题

DOI:
10.1109/smc.2018.00522
复制
发表时间:
2018-10
期刊:
2018 IEEE International Conference on Systems, Man, and Cybernetics (SMC)
影响因子:
--
通讯作者:
Kuan-Ting Chen;Sheng-De Wang
Kuan-Ting Chen;Sheng-De Wang
中科院分区:
其他
文献类型:
--
作者:
Kuan-Ting Chen;Sheng-De Wang

文献摘要

被引文献

相似文献

深度强化学习(Deep - Q-Learning, DQN)和策略梯度(Policy Gradient, PG)等方法被提出用于机械臂控制代理的训练。确定性深度策略梯度(Deterministic Deep Policy Gradient, DDPG)方法利用确定性策略代替随机策略,进一步简化了训练过程,提高了训练性能。强化学习从环境中获取奖励,并训练底层控制代理来完成任务。适当的奖励可以获得更好的表现和更短的训练时间,但需要领域知识和试错法来定义适当的奖励函数。在本文中,我们提出了一种基于DDPG,利用优先体验重播(PER),异步代理学习和元学习的方法。提出的元学习方法使用多个分布式学习器,称为工作者,从连续的先前状态和奖励中学习。通过对六自由度(IRB140)和七自由度(LBR iiwa 14 R820)机械臂进行仿真,训练控制主体在三维空间中到达随机目标。实验表明,本文提出的算法在任务成功率和训练速度上都优于带有专门奖励函数的DDPG算法。
Deep reinforcement learning has been proposed to train the control agent for robotic arms, such as Deep Q-Learning (DQN) and Policy Gradient (PG). The approach of Deterministic Deep Policy Gradient (DDPG) takes the advantage of deterministic policy instead of stochastic policy to further simplify the training process and improve the performance. Reinforcement Learning takes the reward from the environment and trains the underlying control agent to achieve the task. An appropriate reward will get better performance and shorter training time, but it requires the domain knowledge and the method of trial and error to define the appropriate reward function. In this paper, we proposed a method that is based on DDPG and makes use of Prioritized Experience Replay (PER), Asynchronous Agent Learning and Meta Learning. The proposed Meta Learning approach uses multiple distributed learners, called workers, to learn from consecutive previous states and rewards. Simulations are done on 6-DOF (IRB140) and 7-DOF (LBR iiwa 14 R820) robotic arms to train the control agents to reach random targets in the three dimension space. The experiments show that the algorithm we proposed is better than the algorithm using DDPG with specialized reward function on the task success rate and the training speed.