Bayesian Robust Optimization for Imitation Learning

Bayesian Robust Optimization for Imitation Learning
复制标题

DOI:
--
复制
发表时间:
2020-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Daniel S. Brown;S. Niekum;Marek Petrik
Daniel S. Brown;S. Niekum;Marek Petrik
中科院分区:
其他
文献类型:
--
作者:
Daniel S. Brown;S. Niekum;Marek Petrik

文献摘要

相似文献

模仿学习的主要挑战之一是确定当演示的状态分布之外时代理应该采取什么行动。逆强化学习(IRL)可以通过学习参数化的奖励函数来实现对新状态的泛化,但这些方法仍然面临真实奖励函数和相应最优策略的不确定性。现有的基于IRL的安全模仿学习方法使用最大最小框架来处理这种不确定性,该框架在对抗性奖励函数的假设下优化策略,而风险中性IRL方法可以优化平均值或MAP奖励函数的策略。虽然完全忽略风险可能会导致过于激进和不安全的策略,但完全对抗意义上的优化也是有问题的,因为它可能导致过于保守的策略,在实践中表现不佳。为了在这两个极端之间提供一个桥梁,我们提出了贝叶斯鲁棒优化模仿学习(BROIL)。BROIL利用贝叶斯奖励函数推理和用户特定的风险容忍度来有效地优化平衡预期回报和条件风险价值的稳健策略。我们的实证结果表明,BROIL提供了一种自然的方式来插值之间的回报最大化和风险最小化的行为,并优于现有的风险敏感和风险中性的逆强化学习算法。
One of the main challenges in imitation learning is determining what action an agent should take when outside the state distribution of the demonstrations. Inverse reinforcement learning (IRL) can enable generalization to new states by learning a parameterized reward function, but these approaches still face uncertainty over the true reward function and corresponding optimal policy. Existing safe imitation learning approaches based on IRL deal with this uncertainty using a maxmin framework that optimizes a policy under the assumption of an adversarial reward function, whereas risk-neutral IRL approaches either optimize a policy for the mean or MAP reward function. While completely ignoring risk can lead to overly aggressive and unsafe policies, optimizing in a fully adversarial sense is also problematic as it can lead to overly conservative policies that perform poorly in practice. To provide a bridge between these two extremes, we propose Bayesian Robust Optimization for Imitation Learning (BROIL). BROIL leverages Bayesian reward function inference and a user specific risk tolerance to efficiently optimize a robust policy that balances expected return and conditional value at risk. Our empirical results show that BROIL provides a natural way to interpolate between return-maximizing and risk-minimizing behaviors and outperforms existing risk-sensitive and risk-neutral inverse reinforcement learning algorithms.