Efficiently Guiding Imitation Learning Algorithms with Human Gaze

Efficiently Guiding Imitation Learning Algorithms with Human Gaze
复制标题

DOI:
--
复制
发表时间:
2020-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Akanksha Saran;Ruohan Zhang;Elaine Schaertl Short;S. Niekum
Akanksha Saran;Ruohan Zhang;Elaine Schaertl Short;S. Niekum
中科院分区:
其他
文献类型:
--
作者:
Akanksha Saran;Ruohan Zhang;Elaine Schaertl Short;S. Niekum

文献摘要

被引文献

相似文献

人类的凝视是人类任务示范中的意图揭示信号。在这项工作中,我们使用人类示威者的凝视提示来增强最新的反向加固学习(IRL)和行为克隆(BC)算法的表现。我们提出了一种新颖的方法,用于以计算有效的方式利用凝视数据 - 作为辅助损失函数的一部分编码人的注意力,而无需在这些模型中添加任何其他可学习的参数,而无需在测试时凝视数据。辅助损失鼓励网络在人类目光固定的地区进行卷积激活。我们展示了如何使用我们的辅助注视损失(基于覆盖范围的目光损失或CGL)来增强任何现有的卷积体系结构,这些量子可以指导学习更好的奖励功能或政策。我们表明,我们提出的方法可以提高卑诗省和IRL方法在各种Atari游戏中的性能。我们还将使用模仿学习方法的两种基线方法进行比较。我们的方法的表现优于一种基线方法,称为“凝视调制的辍学”(GMD),并且与另一种使用目光作为网络输入的方法相媲美,从而增加了可学习参数的量。
Human gaze is known to be an intention-revealing signal in human demonstrations of tasks. In this work, we use gaze cues from human demonstrators to enhance the performance of state-of-the-art inverse reinforcement learning (IRL) and behavioral cloning (BC) algorithms. We propose a novel approach for utilizing gaze data in a computationally efficient manner --- encoding the human's attention as part of an auxiliary loss function, without adding any additional learnable parameters to those models and without requiring gaze data at test time. The auxiliary loss encourages a network to have convolutional activations in regions where the human's gaze fixated. We show how to augment any existing convolutional architecture with our auxiliary gaze loss (coverage-based gaze loss or CGL) that can guide learning toward a better reward function or policy. We show that our proposed approach improves performance of both BC and IRL methods on a variety of Atari games. We also compare against two baseline methods for utilizing gaze data with imitation learning methods. Our approach outperforms a baseline method, called gaze-modulated dropout (GMD), and is comparable to another method (AGIL) which uses gaze as input to the network and thus increases the amount of learnable parameters.