Learning rewards from exploratory demonstrations using probabilistic temporal ranking

Learning rewards from exploratory demonstrations using probabilistic temporal ranking
复制标题

DOI:
10.1007/s10514-023-10120-w
复制
发表时间:
2020-02
期刊:
影响因子:
3.5
通讯作者:
Michael Burke;Katie Lu;Daniel Angelov;Artūras Straižys;Craig Innes;Kartic Subr;S. Ramamoorthy
Michael Burke;Katie Lu;Daniel Angelov;Artūras Straižys;Craig Innes;Kartic Subr;S. Ramamoorthy
中科院分区:
计算机科学3区
文献类型:
--
作者:
Michael Burke;Katie Lu;Daniel Angelov;Artūras Straižys;Craig Innes;Kartic Subr;S. Ramamoorthy

文献摘要

相似文献

信息路径规划是机器人视觉伺服和主动视点选择的一种成熟方法,但通常假设合适的成本函数或目标状态是已知的。这项工作认为逆问题,任务的目标是未知的,奖励功能需要推断出从探索性的示例演示提供的演示,用于在下游信息路径规划政策。不幸的是,许多现有的奖励推理策略是不适合这类问题,由于探索性质的示范。在本文中,我们提出了一种替代方法来科普这些次优的,探索性的演示发生的问题。我们假设,在需要发现的任务中,任何演示的连续状态都越来越有可能与更高的奖励相关联,并使用此假设来生成基于时间的二进制比较结果,并推断支持这些排名的奖励函数,在概率生成模型下。我们形式化这个概率的时间rankingapproach,并表明它改进了现有的方法来执行奖励推理的自主超声扫描,一个新的应用程序的学习,从演示医学成像,同时也是在广泛的目标为导向的学习演示任务的价值。
Informative path-planning is a well established approach to visual-servoing and active viewpoint selection in robotics, but typically assumes that a suitable cost function or goal state is known. This work considers the inverse problem, where the goal of the task is unknown, and a reward function needs to be inferred from exploratory example demonstrations provided by a demonstrator, for use in a downstream informative path-planning policy. Unfortunately, many existing reward inference strategies are unsuited to this class of problems, due to the exploratory nature of the demonstrations. In this paper, we propose an alternative approach to cope with the class of problems where these sub-optimal, exploratory demonstrations occur. We hypothesise that, in tasks which require discovery, successive states of any demonstration are progressively more likely to be associated with a higher reward, and use this hypothesis to generate time-based binary comparison outcomes and infer reward functions that support these ranks, under a probabilistic generative model. We formalise thisprobabilistic temporal rankingapproach and show that it improves upon existing approaches to perform reward inference for autonomous ultrasound scanning, a novel application of learning from demonstration in medical imaging while also being of value across a broad range of goal-oriented learning from demonstration tasks.