Human-Centered Reinforcement Learning: A Survey

Human-Centered Reinforcement Learning: A Survey
复制标题

以人为本的强化学习:一项调查

DOI:
10.1109/thms.2019.2912447
复制
发表时间:
2019-08-01
影响因子:
3.6
通讯作者:
He, Bo
He, Bo
中科院分区:
计算机科学3区
文献类型:
--
作者:
Li, Guangliang;Gomez, Randy;He, Bo

文献摘要

被引文献

相似文献

以人为中心的强化学习(RL),其中智能体学习如何从人类观察者提供的评估反馈中执行任务,近年来变得越来越流行。RL代理能够从人类反馈中学习的优势导致了对现实生活问题的适用性越来越大。本文描述了最先进的以人为本的强化学习算法,旨在成为开始以人为本的强化学习研究人员的起点。此外,本文的目的是提出一个全面的调查,在这一领域的最新突破,并提供参考最有趣和最成功的作品。本文从环境奖励的角度介绍了强化学习的概念,讨论了以人为本的强化学习的起源及其与传统强化学习的区别。然后,我们描述了人类评价反馈的不同解释,在过去的十年中产生了许多以人为中心的RL算法。此外,我们还描述了从人类评价反馈和环境奖励中学习的代理研究,以及提高以人为中心的RL的效率。最后,我们总结了应用领域的概述,并讨论了未来的工作和开放的问题。
Human-centered reinforcement learning (RL), in which an agent learns how to perform a task from evaluative feedback delivered by a human observer, has become more and more popular in recent years. The advantage of being able to learn from human feedback for a RL agent has led to increasing applicability to real-life problems. This paper describes the state-of-the-art human-centered RL algorithms and aims to become a starting point for researchers who are initiating their endeavors in human-centered RL. Moreover, the objective of this paper is to present a comprehensive survey of the recent breakthroughs in this field and provide references to the most interesting and successful works. After starting with an introduction of the concepts of RL from environmental reward, this paper discusses the origins of human-centered RL and its difference from traditional RL. Then we describe different interpretations of human evaluative feedback, which have produced many human-centered RL algorithms in the past decade. In addition, we describe research on agents learning from both human evaluative feedback and environmental rewards as well as on improving the efficiency of human-centered RL. Finally, we conclude with an overview of application areas and a discussion of future work and open questions.