Concurrent PAC RL

Concurrent PAC RL
复制标题

并发 PAC 强化学习

DOI:
--
复制
发表时间:
2015
期刊:
AAAI Conference on Artificial Intelligence
影响因子:
--
通讯作者:
E. Brunskill
E. Brunskill
中科院分区:
--
文献类型:
--
作者:
Z. Guo;E. Brunskill

文献摘要

被引文献

相似文献

在许多现实情况下,决策者可能会并行地在许多单独的强化学习任务中做出决策,但关于并发强化学习的工作很少。在高效探索RL文献的基础上,我们引入了两个新的并发RL算法,并限制了它们的样本复杂度。我们表明,在一些温和的条件下,无论是当代理是已知的,在许多副本相同的MDP,当他们是不相同的,但采取从一个有限的集合,我们可以获得线性改善的样本复杂度不共享信息。这是非常令人兴奋的,因为线性加速是人们可能希望获得的最大值。我们的初步实验证实了这一结果,并显示了经验的好处。
In many real-world situations a decision maker may make decisions across many separate reinforcement learning tasks in parallel, yet there has been very little work on concurrent RL. Building on the efficient exploration RL literature, we introduce two new concurrent RL algorithms and bound their sample complexity. We show that under some mild conditions, both when the agent is known to be acting in many copies of the same MDP, and when they are not the same but are taken from a finite set, we can gain linear improvements in the sample complexity over not sharing information. This is quite exciting as a linear speedup is the most one might hope to gain. Our preliminary experiments confirm this result and show empirical benefits.