Variational Regret Bounds for Reinforcement Learning

Variational Regret Bounds for Reinforcement Learning
复制标题

强化学习的变分遗憾界限

DOI:
--
复制
发表时间:
2019
期刊:
Conference on Uncertainty in Artificial Intelligence
影响因子:
--
通讯作者:
P. Auer
P. Auer
中科院分区:
--
文献类型:
--
作者:
Pratik Gajane;R. Ortner;P. Auer

文献摘要

被引文献

相似文献

我们在马尔可夫决策过程(mdp)中考虑不打折的强化学习,其中奖励函数和状态转移概率都可能随时间变化(逐渐或突然)。针对这一问题,我们提出了一种算法,并提供了针对最优非平稳策略评估后悔的性能保证。遗憾的上界是根据MDP的总变化给出的。这是一般强化学习设置的第一个变分遗憾边界。
We consider undiscounted reinforcement learning in Markov decision processes (MDPs) where both the reward functions and the state-transition probabilities may vary (gradually or abruptly) over time. For this problem setting, we propose an algorithm and provide performance guarantees for the regret evaluated against the optimal non-stationary policy. The upper bound on the regret is given in terms of the total variation in the MDP. This is the first variational regret bound for the general reinforcement learning setting.