Inverse Risk-Sensitive Reinforcement Learning

Inverse Risk-Sensitive Reinforcement Learning
复制标题

DOI:
10.1109/tac.2019.2926674
复制
发表时间:
2017-03
影响因子:
6.8
通讯作者:
L. Ratliff;Eric V. Mazumdar
L. Ratliff;Eric V. Mazumdar
中科院分区:
计算机科学2区
文献类型:
--
作者:
L. Ratliff;Eric V. Mazumdar

文献摘要

被引文献

相似文献

这项工作解决了马尔可夫决策过程中的决策代理是风险敏感的逆强化学习的问题。特别是,一个风险敏感的强化学习算法的收敛保证,利用连贯的风险指标和模型的人的决策,其起源于行为心理学和经济学。风险敏感强化学习算法为基于梯度的反向强化学习算法提供了理论基础,该算法寻求最小化根据观察到的行为定义的损失函数。它示出的损失函数的梯度相对于模型参数是很好地定义和计算通过收缩映射参数。评估所提出的技术进行网格世界的例子,一个典型的基准问题。
This work addresses the problem of inverse reinforcement learning in Markov decision processes where the decision-making agent is risk-sensitive. In particular, a risk-sensitive reinforcement learning algorithm with convergence guarantees that makes use of coherent risk metrics and models of human decision-making which have their origins in behavioral psychology and economics is presented. The risk-sensitive reinforcement learning algorithm provides the theoretical underpinning for a gradient-based inverse reinforcement learning algorithm that seeks to minimize a loss function defined on the observed behavior. It is shown that the gradient of the loss function with respect to the model parameters is well defined and computable via a contraction map argument. Evaluation of the proposed technique is performed on a Grid World example, a canonical benchmark problem.