Learning Performance Graphs From Demonstrations via Task-Based Evaluations

Learning Performance Graphs From Demonstrations via Task-Based Evaluations
复制标题

DOI:
10.1109/lra.2022.3226072
复制
发表时间:
2022-04
影响因子:
5.2
通讯作者:
Aniruddh Gopinath Puranic;Jyotirmoy V. Deshmukh;S. Nikolaidis
Aniruddh Gopinath Puranic;Jyotirmoy V. Deshmukh;S. Nikolaidis
中科院分区:
计算机科学2区
文献类型:
--
作者:
Aniruddh Gopinath Puranic;Jyotirmoy V. Deshmukh;S. Nikolaidis

文献摘要

被引文献

相似文献

在机器人从演示中学习(LfD)的范例中,理解和评估演示行为在提取机器人控制策略中起着关键作用。在没有这些知识的情况下,机器人可能会推断出不正确的奖励函数,从而导致不期望的或不安全的控制策略。先前的工作已经使用了时序逻辑规范,由人类专家根据其重要性手动排名,从不完美/次优的演示中学习奖励函数。为了克服对专家排名的依赖,我们提出了一种新的算法,从演示中学习,提供的规格的性能图的形式的部分排序。通过各种实验,包括工业移动的机器人的模拟,我们表明,提取奖励功能与学习的图形结果在机器人的政策类似于那些手动指定的排序。我们还表明,在一个用户研究中,学习排序匹配的排序或排名的参与者在模拟驾驶域的演示。这些结果表明,我们可以准确地评估示范方面提供的任务规范,从一个小的不完善的数据集,以最小的专家输入。
In the paradigm of robot learning-from-demonstra tions (LfD), understanding and evaluating the demonstrated behaviors plays a critical role in extracting control policies for robots. Without this knowledge, a robot may infer incorrect reward functions that lead to undesirable or unsafe control policies. Prior work has used temporal logic specifications, manually ranked by human experts based on their importance, to learn reward functions from imperfect/suboptimal demonstrations. To overcome reliance on expert rankings, we propose a novel algorithm that learns from demonstrations, a partial ordering of provided specifications in the form of a performance graph. Through various experiments, including simulation of industrial mobile robots, we show that extracting reward functions with the learned graph results in robot policies similar to those generated with the manually specified orderings. We also show in a user study that the learned orderings match the orderings or rankings by participants for demonstrations in a simulated driving domain. These results show that we can accurately evaluate demonstrations with respect to provided task specifications from a small set of imperfect data with minimal expert input.