Deep Reinforcement Learning With Adversarial Training for Automated Excavation Using Depth Images

Deep Reinforcement Learning With Adversarial Training for Automated Excavation Using Depth Images
复制标题

DOI:
10.1109/access.2022.3140781
复制
发表时间:
2022
期刊:
影响因子:
3.9
通讯作者:
Takayuki Osa;M. Aizawa
Takayuki Osa;M. Aizawa
中科院分区:
计算机科学3区
文献类型:
--
作者:
Takayuki Osa;M. Aizawa

文献摘要

被引文献

相似文献

挖掘是施工期间最频繁执行的任务之一,通常会对操作人员造成危险。为了减少潜在风险并解决劳动力短缺的问题,挖掘自动化至关重要。虽然以前的研究已经取得了可喜的成果的基础上使用强化学习(RL)的自动挖掘,挖掘任务的RL的背景下的属性还没有得到充分的研究。在这项研究中,我们研究了Qt-Opt,这是连续动作空间的Q学习算法的变体,用于使用深度图像学习挖掘任务。受监督学习中虚拟对抗训练的启发,我们提出了一种正则化方法,该方法使用虚拟对抗样本来减少Q学习算法中Q值的高估。我们的研究结果表明,Qt-Opt是更采样效率比国家的最先进的演员-评论家的方法在我们的问题设置,我们验证了所提出的方法进一步提高了Qt-Opt的采样效率。我们的研究结果表明,多个最佳行动往往存在于挖掘过程中,政策表示的选择是令人满意的性能至关重要。
Excavation, which is one of the most frequently performed tasks during construction often poses danger to human operators. To reduce potential risks and address the problem of workforce shortage, automation of excavation is essential. Although previous studies have yielded promising results based on the use of reinforcement learning (RL) for automated excavation, the properties of excavation task in the context of RL have not been sufficiently investigated. In this study, we investigate Qt-Opt, which is a variant of Q-learning algorithms for continuous action space, for learning the excavation task using depth images. Inspired by virtual adversarial training in supervised learning, we propose a regularization method that uses virtual adversarial samples to reduce overestimation of Q-values in a Q-learning algorithm. Our results reveal that Qt-Opt is more sample-efficient than state-of-the-art actor-critic methods in our problem setting, and we verify that the proposed method further improves the sample efficiency of Qt-Opt. Our results demonstrate that multiple optimal actions often exist within the process of excavation and the choice of policy representation is crucial for satisfactory performance.