Q-learning with Logarithmic Regret

Q-learning with Logarithmic Regret
复制标题

DOI:
--
复制
发表时间:
2020-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Kunhe Yang;Lin F. Yang;S. Du
Kunhe Yang;Lin F. Yang;S. Du
中科院分区:
其他
文献类型:
--
作者:
Kunhe Yang;Lin F. Yang;S. Du

文献摘要

被引文献

相似文献

本文提出了第一个非渐近结果,表明如果最优 $Q$ 函数中存在严格正的次优差距,则无模型算法可以实现情景表格强化学习的对数累积遗憾。我们证明了 [Jin 等人] 中研究的乐观 $Q$ 学习。 2018] 享有 ${\mathcal{O}}\left(\frac{SA\cdot \mathrm{poly}\left(H\right)}{\mathrm{gap}_{\min}}\log\left(SAT\right)\right)$ 累积后悔界限,其中 $S$ 是状态数,$A$ 是动作数,$H$ 是规划范围,$T$ 是总步数,并且$\mathrm{gap}_{\min}$ 是最小次优差距。该界限与 $S,A,T$ 方面的信息理论下限相匹配,最高可达 $\log\left(SA\right)$ 因子。我们进一步将分析扩展到折扣设置并获得类似的对数累积遗憾界限。
This paper presents the first non-asymptotic result showing that a model-free algorithm can achieve a logarithmic cumulative regret for episodic tabular reinforcement learning if there exists a strictly positive sub-optimality gap in the optimal $Q$-function. We prove that the optimistic $Q$-learning studied in [Jin et al. 2018] enjoys a ${\mathcal{O}}\left(\frac{SA\cdot \mathrm{poly}\left(H\right)}{\mathrm{gap}_{\min}}\log\left(SAT\right)\right)$ cumulative regret bound, where $S$ is the number of states, $A$ is the number of actions, $H$ is the planning horizon, $T$ is the total number of steps, and $\mathrm{gap}_{\min}$ is the minimum sub-optimality gap. This bound matches the information theoretical lower bound in terms of $S,A,T$ up to a $\log\left(SA\right)$ factor. We further extend our analysis to the discounted setting and obtain a similar logarithmic cumulative regret bound.