Detecting Incorrect Visual Demonstrations for Improved Policy Learning

Detecting Incorrect Visual Demonstrations for Improved Policy Learning
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Mostafa Hussein;M. Begum
Mostafa Hussein;M. Begum
中科院分区:
其他
文献类型:
--
作者:
Mostafa Hussein;M. Begum

文献摘要

相似文献

:仅从原始视频演示中学习任务是机器人视觉模仿学习研究的最新技术。这里隐含的假设是所有视频演示都显示执行任务的最佳/次优方式。如果事实并非如此怎么办?如果一个或多个视频显示执行任务的方式错误怎么办?从这种不正确的演示中学到的任务策略对于机器人和人类来说可能是不安全的。因此,在将视频演示交给策略学习算法之前分析其正确性非常重要。这是一项具有挑战性的任务,特别是由于状态空间非常大。本文提出了一种框架,可以自动检测由多个子任务组成的顺序任务的错误视频演示。我们分析演示池,以识别任务特征遵循“破坏性”序列的视频。我们分析熵来衡量这种破坏,并通过解决极小极大问题,为不正确的视频分配较差的权重。我们使用两个真实世界的视频数据集评估了该框架:我们使用 YuMi 机器人定制设计的泡茶和公开的 50-Salads 。实验结果表明,所提出的框架在检测不正确的视频演示方面是有效的,即使它们占演示集的 40% 也是如此。我们还表明,当从训练池中丢弃不正确的演示时,各种最先进的模仿学习算法可以学习更好的策略。
: Learning tasks only from raw video demonstrations is the current state of the art in robotics visual imitation learning research. The implicit assumption here is that all video demonstrations show an optimal/sub-optimal way of performing the task. What if that is not true? What if one or more videos show a wrong way of executing the task? A task policy learned from such incorrect demonstrations can be potentially unsafe for robots and humans. It is therefore important to analyze the video demonstrations for correctness before handing them over to the policy learning algorithm. This is a challenging task, especially due to the very large state space. This paper proposes a framework to autonomously detect incorrect video demonstrations of sequential tasks consisting of several sub-tasks. We analyze the demonstration pool to identify video(s) for which task-features follow a ‘disruptive’ sequence. We analyze entropy to measure this disruption and – through solving a minmax problem – assign poor weights to incorrect videos. We evaluated the framework with two real-world video datasets: our custom-designed Tea-Making with a YuMi robot and the publicly available 50-Salads . Experimental results show the effectiveness of the proposed framework in detecting incorrect video demonstrations even when they make up 40% of the demonstration set. We also show that various state-of-the-art imitation learning algorithms learn a better policy when incorrect demonstrations are discarded from the training pool.