Eliciting Compatible Demonstrations for Multi-Human Imitation Learning

Eliciting Compatible Demonstrations for Multi-Human Imitation Learning
复制标题

DOI:
10.48550/arxiv.2210.08073
复制
发表时间:
2022-10
期刊:
--
影响因子:
--
通讯作者:
Kanishk Gandhi;Siddharth Karamcheti;Madeline Liao;Dorsa Sadigh
Kanishk Gandhi;Siddharth Karamcheti;Madeline Liao;Dorsa Sadigh
中科院分区:
其他
文献类型:
--
作者:
Kanishk Gandhi;Siddharth Karamcheti;Madeline Liao;Dorsa Sadigh

文献摘要

相似文献

从人类提供的演示中模仿学习是机器人操作学习策略的强大方法。虽然模仿学习的理想数据集是同质和低方差的-反映了执行任务的单一最佳方法-但自然的人类行为具有很大的异质性,有几种最佳方法来演示任务。这种多模态对人类用户来说是无关紧要的,任务的变化表现为潜意识的选择;例如,向下伸手,然后穿过去抓住一个物体,而不是向前伸手,然后向下。然而,这种不匹配给交互式模仿学习带来了问题,在交互式模仿学习中,用户序列通过迭代地收集新的、可能相互冲突的演示来改进策略。为了解决这个问题的演示不兼容,这项工作设计了一种方法,1)测量一个新的演示给定的基本政策的兼容性,和2)积极引发更多的兼容性演示新用户。在两个模拟任务,需要长期的视野,灵巧的操作和现实世界的“食品电镀“任务与弗兰卡Panda手臂,我们表明,我们都可以通过事后过滤识别不兼容的演示,并应用我们的兼容性措施,积极引出新用户的兼容演示,从而提高任务成功率在模拟和真实的环境。
Imitation learning from human-provided demonstrations is a strong approach for learning policies for robot manipulation. While the ideal dataset for imitation learning is homogenous and low-variance -- reflecting a single, optimal method for performing a task -- natural human behavior has a great deal of heterogeneity, with several optimal ways to demonstrate a task. This multimodality is inconsequential to human users, with task variations manifesting as subconscious choices; for example, reaching down, then across to grasp an object, versus reaching across, then down. Yet, this mismatch presents a problem for interactive imitation learning, where sequences of users improve on a policy by iteratively collecting new, possibly conflicting demonstrations. To combat this problem of demonstrator incompatibility, this work designs an approach for 1) measuring the compatibility of a new demonstration given a base policy, and 2) actively eliciting more compatible demonstrations from new users. Across two simulation tasks requiring long-horizon, dexterous manipulation and a real-world"food plating"task with a Franka Emika Panda arm, we show that we can both identify incompatible demonstrations via post-hoc filtering, and apply our compatibility measure to actively elicit compatible demonstrations from new users, leading to improved task success rates across simulated and real environments.