Improved upper bounds on the expected error in constant step-size Q-learning

Improved upper bounds on the expected error in constant step-size Q-learning
复制标题

改进了恒定步长 Q 学习中预期误差的上限

DOI:
--
复制
发表时间:
2013
期刊:
American Control Conference
影响因子:
--
通讯作者:
R. Srikant
R. Srikant
中科院分区:
--
文献类型:
--
作者:
Carolyn L. Beck;R. Srikant

文献摘要

被引文献

相似文献

我们考虑应用于有限状态和动作空间的固定步长 Q 学习算法、折扣奖励马尔可夫决策问题 (MDP)。在之前的工作中,我们推导了 Q 值估计误差的一阶矩的界限,特别是误差无穷大范数的预期稳态值。本文和前一篇论文的目标都是在无限的时间范围内最大化奖励的贴现总和。然而,在我们之前的工作中,我们得出的界限仅在步长足够小(有时不切实际)时才成立。在本文中,我们提出了一个新的误差界限,与之前一样,随着步长变为零,该误差界限也变为零,但对于步长的所有值也有效。为了获得新的界限,我们将时间划分为帧,使得帧内未访问的某些状态的概率严格小于 1:然后通过在每一帧中对系统进行一次采样来找到我们的误差界限。
We consider fixed step-size Q-learning algorithms applied to finite state and action space, discounted reward Markov decision problems (MDPs). In previous work we derived a bound on the first moment of the Q-value estimation error, specifically on the expected steady-state value of the infinity norm of the error. The goal in both this paper, and the previous, is to maximize a discounted sum of rewards over an infinite time horizon. However, in our previous work, the bound we derived holds only when the step-size is sufficiently, and sometimes impractically, small. In this paper, we present a new error bound that, as before, goes to zero as the step-size goes to zero, but is also valid for all values of the step-size. To obtain the new bound, we divide time into frames such that the probability that there is some state that is not visited within the frame is strictly less than 1: Our error bound is then found by sampling the system one time in every frame.