Learning From Imperfect Demonstrations From Agents With Varying Dynamics

Learning From Imperfect Demonstrations From Agents With Varying Dynamics
复制标题

DOI:
10.1109/lra.2021.3068912
复制
发表时间:
2021-03
影响因子:
5.2
通讯作者:
Zhangjie Cao;Dorsa Sadigh
Zhangjie Cao;Dorsa Sadigh
中科院分区:
计算机科学2区
文献类型:
--
作者:
Zhangjie Cao;Dorsa Sadigh

文献摘要

相似文献

模仿学习使机器人能够从演示中学习。以前的模仿学习算法通常假设访问最优专家演示。然而,在许多现实世界的应用中,这种假设是有限的。大多数收集的演示不是最佳的,或者是由具有稍微不同动态的代理产生的。因此,我们解决了模仿学习的问题时,演示可以是次优的或从代理不同的动态。我们开发了一个由可行性分数和最优性分数组成的度量标准来衡量演示对模仿学习的有用程度。建议的分数可以从更多信息的演示中学习,而忽略不太相关的演示。我们在四个环境中的模拟和真实的机器人上的实验表明,改进的学习策略具有更高的预期回报。
Imitation learning enables robots to learn from demonstrations. Previous imitation learning algorithms usually assume access to optimal expert demonstrations. However, in many real-world applications, this assumption is limiting. Most collected demonstrations are not optimal or are produced by an agent with slightly different dynamics. We therefore address the problem of imitation learning when the demonstrations can be sub-optimal or be drawn from agents with varying dynamics. We develop a metric composed of a feasibility score and an optimality score to measure how useful a demonstration is for imitation learning. The proposed score enables learning from more informative demonstrations, and disregarding the less relevant demonstrations. Our experiments on four environments in simulation and on a real robot show improved learned policies with higher expected return.