Variational Regret Bounds for Reinforcement Learning
Variational Regret Bounds for Reinforcement Learning
复制标题
强化学习的变分遗憾界限
DOI:
--
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
P. Auer
中科院分区:
文献类型:
--
作者:
Pratik Gajane;R. Ortner;P. Auer
We consider undiscounted reinforcement learning in Markov decision processes (MDPs) where both the reward functions and the state-transition probabilities may vary (gradually or abruptly) over time. For this problem setting, we propose an algorithm and provide performance guarantees for the regret evaluated against the optimal non-stationary policy. The upper bound on the regret is given in terms of the total variation in the MDP. This is the first variational regret bound for the general reinforcement learning setting.