Imitation Learning by Estimating Expertise of Demonstrators

Imitation Learning by Estimating Expertise of Demonstrators
复制标题

DOI:
--
复制
发表时间:
2022-02
期刊:
--
影响因子:
--
通讯作者:
M. Beliaev;Andy Shih;Stefano Ermon;Dorsa Sadigh;Ramtin Pedarsani
M. Beliaev;Andy Shih;Stefano Ermon;Dorsa Sadigh;Ramtin Pedarsani
中科院分区:
其他
文献类型:
--
作者:
M. Beliaev;Andy Shih;Stefano Ermon;Dorsa Sadigh;Ramtin Pedarsani

文献摘要

被引文献

相似文献

许多现有的模仿学习数据集是从多个示范者收集的,每个示范者在环境的不同部分具有不同的专业知识。然而,标准的模仿学习算法通常将所有演示者视为同质的,无论他们的专业知识如何,吸收任何次优演示者的弱点。在这项工作中,我们表明,在演示专家的无监督学习可以导致模仿学习算法的性能的一致提高。我们开发和优化一个联合模型在一个学习的政策和专业知识水平的示威者。这使我们的模型能够从最优行为中学习,并过滤掉每个演示者的次优行为。我们的模型学习一个单一的政策,甚至可以胜过最好的演示者,并可用于估计任何演示者在任何状态的专业知识。我们说明了我们的研究结果,真正的机器人连续控制任务,从Robomimic和离散环境,如MiniGrid和国际象棋,在21 $的23 $设置中的$21$的表现优于竞争方法,平均为$7\%$和高达$60\%$的改善方面的最终奖励。
Many existing imitation learning datasets are collected from multiple demonstrators, each with different expertise at different parts of the environment. Yet, standard imitation learning algorithms typically treat all demonstrators as homogeneous, regardless of their expertise, absorbing the weaknesses of any suboptimal demonstrators. In this work, we show that unsupervised learning over demonstrator expertise can lead to a consistent boost in the performance of imitation learning algorithms. We develop and optimize a joint model over a learned policy and expertise levels of the demonstrators. This enables our model to learn from the optimal behavior and filter out the suboptimal behavior of each demonstrator. Our model learns a single policy that can outperform even the best demonstrator, and can be used to estimate the expertise of any demonstrator at any state. We illustrate our findings on real-robotic continuous control tasks from Robomimic and discrete environments such as MiniGrid and chess, out-performing competing methods in $21$ out of $23$ settings, with an average of $7\%$ and up to $60\%$ improvement in terms of the final reward.