Deep reinforcement learning with smooth policy update: Application to robotic cloth manipulation

Deep reinforcement learning with smooth policy update: Application to robotic cloth manipulation
复制标题

DOI:
10.1016/j.robot.2018.11.004
复制
发表时间:
2019-02-01
影响因子:
4.3
通讯作者:
Matsubara, Takamitsu
Matsubara, Takamitsu
中科院分区:
计算机科学3区
文献类型:
--
作者:
Tsurumine, Yoshihisa;Cui, Yunduan;Matsubara, Takamitsu

文献摘要

被引文献

相似文献

深度强化学习(DRL),它可以学习复杂的策略,并将高维观测作为输入,例如,图像,已成功应用于各种任务。因此,将它们应用于机器人学习和执行日常活动(如洗涤和折叠衣服,烹饪和清洁)可能是合适的,因为这些任务对于通常需要(1)直接访问状态变量或(2)从感官输入中提取的精心设计的手工设计特征的非DRL方法来说是困难的。然而,将DRL应用于真实的机器人仍然非常具有挑战性,因为传统的DRL算法需要大量的训练样本进行学习,这在真实的机器人中是艰巨的。为了缓解这一困境,本文提出了两种有效的DRL算法:深度P网络(DPN)和决斗深度P网络(DDPN)。其核心思想是将联合收割机平滑策略更新的本质与深度神经网络自动特征提取的能力相结合,以更少的样本提升样本效率和学习稳定性。所提出的方法首先研究了机器人手臂达到任务的模拟,比较以前的DRL方法,并适用于两个真实的机器人布料操作任务:(1)翻转手帕和(2)折叠有限数量的样本的T恤。所有的结果表明,我们的方法优于以前的DRL方法。(C)2018作者由爱思唯尔公司出版
Deep Reinforcement Learning (DRL), which can learn complex policies with high-dimensional observations as inputs, e.g., images, has been successfully applied to various tasks. Therefore, it may be suitable to apply them for robots to learn and perform daily activities like washing and folding clothes, cooking, and cleaning since such tasks are difficult for non-DRL methods that often require either (1) direct access to state variables or (2) well-designed hand-engineered features extracted from sensory inputs. However, applying DRL to real robots remains very challenging because conventional DRL algorithms require a huge number of training samples for learning, which is arduous in real robots. To alleviate this dilemma, in this paper, we propose two sample efficient DRL algorithms: Deep P-Network (DPN) and Dueling Deep P-Network (DDPN). The core idea is to combine the nature of smooth policy update with the capability of automatic feature extraction in deep neural networks to enhance the sample efficiency and learning stability with fewer samples. The proposed methods were first investigated by a robot-arm reaching task in the simulation that compared previous DRL methods and applied to two real robotic cloth manipulation tasks: (1) flipping a handkerchief and (2) folding a t-shirt with a limited number of samples. All the results suggest that our method outperformed the previous DRL methods. (C) 2018 The Authors. Published by Elsevier B.V.