Machine Teaching for Inverse Reinforcement Learning: Algorithms and Applications

Machine Teaching for Inverse Reinforcement Learning: Algorithms and Applications
复制标题

DOI:
10.1609/aaai.v33i01.33017749
复制
发表时间:
2018-05
期刊:
--
影响因子:
--
通讯作者:
Daniel S. Brown;S. Niekum
Daniel S. Brown;S. Niekum
中科院分区:
其他
文献类型:
--
作者:
Daniel S. Brown;S. Niekum

文献摘要

相似文献

反向强化学习(IRL)从演示中推断出奖励函数,允许策略改进和推广。然而,尽管最近对IRL的兴趣很大,但很少有工作来了解教一个特定的顺序决策任务所需的最小演示集。我们形式化的问题,找到最大的信息演示IRL作为一个机器教学的问题,其目标是找到所需的演示指定奖励等价类的演示的最小数量。我们扩展了以前的工作顺序决策任务的算法教学,通过减少集合覆盖问题,使一个有效的近似算法,用于确定一组最大信息量的演示。我们将我们提出的机器教学算法应用到两个新的应用程序:提供一个下限的查询数量需要学习的政策,使用主动IRL和开发一种新的IRL算法,可以更有效地学习信息演示比标准的IRL方法。
Inverse reinforcement learning (IRL) infers a reward function from demonstrations, allowing for policy improvement and generalization. However, despite much recent interest in IRL, little work has been done to understand the minimum set of demonstrations needed to teach a specific sequential decisionmaking task. We formalize the problem of finding maximally informative demonstrations for IRL as a machine teaching problem where the goal is to find the minimum number of demonstrations needed to specify the reward equivalence class of the demonstrator. We extend previous work on algorithmic teaching for sequential decision-making tasks by showing a reduction to the set cover problem which enables an efficient approximation algorithm for determining the set of maximallyinformative demonstrations. We apply our proposed machine teaching algorithm to two novel applications: providing a lower bound on the number of queries needed to learn a policy using active IRL and developing a novel IRL algorithm that can learn more efficiently from informative demonstrations than a standard IRL approach.