Risk-sensitive inverse reinforcement learning via semi- and non-parametric methods

Risk-sensitive inverse reinforcement learning via semi- and non-parametric methods
复制标题

通过半参数和非参数方法进行风险敏感的逆强化学习

DOI:
--
复制
发表时间:
2017
期刊:
Int. J. Robotics Res.
影响因子:
--
通讯作者:
M. Pavone
M. Pavone
中科院分区:
--
文献类型:
--
作者:
Sumeet Singh;Jonathan Lacotte;Anirudha Majumdar;M. Pavone

文献摘要

被引文献

相似文献

关于逆强化学习(IRL)的文献通常假定人类采取行动是为了使成本函数的期望值最小化,即人类是风险中性的。然而,在实践中,人类往往远非风险中性。为了填补这一差距,本文的目标是设计一个风险敏感(RS)逆强化学习框架,以明确考虑人类的风险敏感性。为此,我们基于一致性风险度量提出了一类灵活的模型,这使我们能够捕捉从风险中性到最坏情况的整个风险偏好范围。我们提出了基于线性规划的高效非参数算法和基于最大似然的半参数算法,用于推断人类在丰富的静态和动态决策环境中的潜在风险度量和成本函数。所得到的方法在一个有10名人类参与者的模拟驾驶游戏中得到了验证。我们的方法能够以数据高效的方式推断和模拟从高度风险厌恶到风险中性的各种性质不同的驾驶风格。此外,风险敏感逆强化学习方法与风险中性模型的比较表明,风险敏感逆强化学习框架在定性和定量方面都更准确地捕捉到观察到的参与者行为,尤其是在可能发生碰撞等灾难性结果的场景中。
The literature on inverse reinforcement learning (IRL) typically assumes that humans take actions to minimize the expected value of a cost function, i.e., that humans are risk neutral. Yet, in practice, humans are often far from being risk neutral. To fill this gap, the objective of this paper is to devise a framework for risk-sensitive (RS) IRL to explicitly account for a human’s risk sensitivity. To this end, we propose a flexible class of models based on coherent risk measures, which allow us to capture an entire spectrum of risk preferences from risk neutral to worst case. We propose efficient non-parametric algorithms based on linear programming and semi-parametric algorithms based on maximum likelihood for inferring a human’s underlying risk measure and cost function for a rich class of static and dynamic decision-making settings. The resulting approach is demonstrated on a simulated driving game with 10 human participants. Our method is able to infer and mimic a wide range of qualitatively different driving styles from highly risk averse to risk neutral in a data-efficient manner. Moreover, comparisons of the RS-IRL approach with a risk-neutral model show that the RS-IRL framework more accurately captures observed participant behavior both qualitatively and quantitatively, especially in scenarios where catastrophic outcomes such as collisions can occur.