A survey of inverse reinforcement learning: Challenges, methods and progress

A survey of inverse reinforcement learning: Challenges, methods and progress
复制标题

DOI:
10.1016/j.artint.2021.103500
复制
发表时间:
2021-03-30
影响因子:
14.4
通讯作者:
Doshi, Prashant
Doshi, Prashant
中科院分区:
计算机科学2区
文献类型:
--
作者:
Arora, Saurabh;Doshi, Prashant

文献摘要

被引文献

相似文献

反向强化学习(IRL)是在给定其策略或观察到的行为的情况下推断代理的奖励函数的问题。类似于RL,IRL被认为是一个问题和一类方法。通过分类调查IRL中现存的文献,本文为机器学习的研究人员和从业者以及新的研究人员提供了全面的参考,以了解IRL的挑战并选择最适合手头问题的方法。该调查正式介绍了IRL问题沿着其核心挑战,如难以进行准确的推理和推广,其对先验知识的敏感性,以及解决方案复杂性与问题大小不成比例的增长。本文调查了大量的基础方法,这些方法按其目标的共性分组在一起,并阐述了这些方法如何减轻挑战。我们进一步讨论了传统IRL方法的扩展,用于处理不完美的感知,不完整的模型,学习多个奖励函数和非线性奖励函数。文章最后的调查与讨论的一些广泛的进展,在研究领域和目前开放的研究问题。(C)2021爱思唯尔有限公司版权所有。
Inverse reinforcement learning (IRL) is the problem of inferring the reward function of an agent, given its policy or observed behavior. Analogous to RL, IRL is perceived both as a problem and as a class of methods. By categorically surveying the extant literature in IRL, this article serves as a comprehensive reference for researchers and practitioners of machine learning as well as those new to it to understand the challenges of IRL and select the approaches best suited for the problem on hand. The survey formally introduces the IRL problem along with its central challenges such as the difficulty in performing accurate inference and its generalizability, its sensitivity to prior knowledge, and the disproportionate growth in solution complexity with problem size. The article surveys a vast collection of foundational methods grouped together by the commonality of their objectives, and elaborates how these methods mitigate the challenges. We further discuss extensions to the traditional IRL methods for handling imperfect perception, an incomplete model, learning multiple reward functions and nonlinear reward functions. The article concludes the survey with a discussion of some broad advances in the research area and currently open research questions. (C) 2021 Elsevier B.V. All rights reserved.