Reward Biased Maximum Likelihood Estimation for Reinforcement Learning

Reward Biased Maximum Likelihood Estimation for Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2020-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Akshay Mete;Rahul Singh;Xi Liu;P. Kumar
Akshay Mete;Rahul Singh;Xi Liu;P. Kumar
中科院分区:
其他
文献类型:
--
作者:
Akshay Mete;Rahul Singh;Xi Liu;P. Kumar

文献摘要

被引文献

相似文献

Kumar和Becker(1982)提出的基于报酬偏差的基于最大似然估计的自适应控制(RBMLE)原理是基于上置信限(UCB)方法(Lai和Robbins,1985)的另一种方法,该原理现在被称为“面对不确定性的乐观”(Auer et al.,2002)。它利用修正的最大似然估计,偏向于产生更高平均回报的马尔可夫决策过程(MDP)模型。然而,它的后悔性能从来没有被分析过早期的强化学习(RL(Sutton等人,1998))任务,涉及到未知MDP的最优控制。我们证明了它具有$O(\logT)$的学习遗憾,其中$T$是时间范围,类似于最新的算法。它为解决RL问题提供了一种替代的通用方法。
The principle of Reward-Biased Maximum Likelihood Estimate Based Adaptive Control (RBMLE) that was proposed in Kumar and Becker (1982) is an alternative approach to the Upper Confidence Bound Based (UCB) Approach (Lai and Robbins, 1985) for employing the principle now known as "optimism in the face of uncertainty" (Auer et al., 2002). It utilizes a modified maximum likelihood estimate, with a bias towards those Markov Decision Process (MDP) models that yield a higher average reward. However, its regret performance has never been analyzed earlier for reinforcement learning (RL (Sutton et al., 1998)) tasks that involve the optimal control of unknown MDPs. We show that it has a learning regret of $O(\log T )$ where $T$ is the time-horizon, similar to the state-of-art algorithms. It provides an alternative general purpose method for solving RL problems.