Model-free Reinforcement Learning in Infinite-horizon Average-reward Markov Decision Processes

Model-free Reinforcement Learning in Infinite-horizon Average-reward Markov Decision Processes
复制标题

DOI:
--
复制
发表时间:
2019-10
期刊:
--
影响因子:
--
通讯作者:
Chen-Yu Wei;Mehdi Jafarnia-Jahromi;Haipeng Luo;Hiteshi Sharma;R. Jain
Chen-Yu Wei;Mehdi Jafarnia-Jahromi;Haipeng Luo;Hiteshi Sharma;R. Jain
中科院分区:
其他
文献类型:
--
作者:
Chen-Yu Wei;Mehdi Jafarnia-Jahromi;Haipeng Luo;Hiteshi Sharma;R. Jain

文献摘要

相似文献

无模型强化学习被认为是记忆和计算效率高的,并且更适合于大规模问题。本文介绍了两种学习无限时域平均报酬马尔可夫决策过程(MDP)的无模型算法。第一个算法减少了问题的折扣奖励版本,并达到$\mathcal{O}(T^{2/3})$后悔后,$T$步骤,弱沟通的MDPs的最小假设。据我们所知,这是第一个无模型算法一般MDPs在这种情况下。第二个算法利用了最近的进步,在自适应算法的对抗多武装土匪和改进的遗憾,$\mathcal{O}(\sqrt{T})$,虽然有一个更强的遍历假设。这一结果显着改善了由Abbasi-Yadkori等人(2019 a)针对无限时域平均奖励设置中遍历MDP的唯一现有无模型算法实现的$\mathcal{O}(T^{3/4})$遗憾。
Model-free reinforcement learning is known to be memory and computation efficient and more amendable to large scale problems. In this paper, two model-free algorithms are introduced for learning infinite-horizon average-reward Markov Decision Processes (MDPs). The first algorithm reduces the problem to the discounted-reward version and achieves $\mathcal{O}(T^{2/3})$ regret after $T$ steps, under the minimal assumption of weakly communicating MDPs. To our knowledge, this is the first model-free algorithm for general MDPs in this setting. The second algorithm makes use of recent advances in adaptive algorithms for adversarial multi-armed bandits and improves the regret to $\mathcal{O}(\sqrt{T})$, albeit with a stronger ergodic assumption. This result significantly improves over the $\mathcal{O}(T^{3/4})$ regret achieved by the only existing model-free algorithm by Abbasi-Yadkori et al. (2019a) for ergodic MDPs in the infinite-horizon average-reward setting.