Provably Feedback-Efficient Reinforcement Learning via Active Reward Learning

Provably Feedback-Efficient Reinforcement Learning via Active Reward Learning
复制标题

DOI:
10.48550/arxiv.2304.08944
复制
发表时间:
2023-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Dingwen Kong;Lin F. Yang
Dingwen Kong;Lin F. Yang
中科院分区:
其他
文献类型:
--
作者:
Dingwen Kong;Lin F. Yang

文献摘要

被引文献

相似文献

在强化学习(RL)中,适当的奖励函数对于指定任务至关重要。然而,在实践中,即使是为简单的任务设计正确的奖励函数也是极具挑战性的。人在循环(HiL)强化学习允许人类通过提供各种类型的反馈向强化学习代理传达复杂的目标。然而,尽管取得了巨大的经验成功,HiL强化学习通常需要来自人类老师的太多反馈,并且理论理解不足。在本文中,我们专注于从理论角度解决这个问题,旨在提供可证明的反馈高效算法框架,该框架采用人在环来指定给定任务的奖励。我们提供了一种基于主动学习的强化学习算法,该算法首先在不指定奖励函数的情况下探索环境,然后仅向人类教师询问一些状态-动作对任务的奖励。之后,该算法保证为高概率任务提供近乎最优的策略。我们证明,即使反馈中存在随机噪声,该算法也只接受对奖励函数的$\widetilde{O}(H{{\dim_{R}^2}})$查询,从而为任何$\epsilon>0$提供$\epsilon$ -最优策略。这里$H$是RL环境的范围,$\dim_{R}$指定了表示奖励函数的函数类的复杂性。相比之下,标准RL算法需要查询至少$\Omega(\operatorname{poly}(d, 1/\epsilon))$状态-动作对的奖励函数,其中$d$取决于环境转换的复杂性。
An appropriate reward function is of paramount importance in specifying a task in reinforcement learning (RL). Yet, it is known to be extremely challenging in practice to design a correct reward function for even simple tasks. Human-in-the-loop (HiL) RL allows humans to communicate complex goals to the RL agent by providing various types of feedback. However, despite achieving great empirical successes, HiL RL usually requires too much feedback from a human teacher and also suffers from insufficient theoretical understanding. In this paper, we focus on addressing this issue from a theoretical perspective, aiming to provide provably feedback-efficient algorithmic frameworks that take human-in-the-loop to specify rewards of given tasks. We provide an active-learning-based RL algorithm that first explores the environment without specifying a reward function and then asks a human teacher for only a few queries about the rewards of a task at some state-action pairs. After that, the algorithm guarantees to provide a nearly optimal policy for the task with high probability. We show that, even with the presence of random noise in the feedback, the algorithm only takes $\widetilde{O}(H{{\dim_{R}^2}})$ queries on the reward function to provide an $\epsilon$-optimal policy for any $\epsilon>0$. Here $H$ is the horizon of the RL environment, and $\dim_{R}$ specifies the complexity of the function class representing the reward function. In contrast, standard RL algorithms require to query the reward function for at least $\Omega(\operatorname{poly}(d, 1/\epsilon))$ state-action pairs where $d$ depends on the complexity of the environmental transition.