Scaling Up Reinforcement Learning through Targeted Exploration

Scaling Up Reinforcement Learning through Targeted Exploration
复制标题

通过有针对性的探索扩大强化学习

DOI:
--
复制
发表时间:
2011
期刊:
AAAI Conference on Artificial Intelligence
影响因子:
--
通讯作者:
Yoonsuck Choe
Yoonsuck Choe
中科院分区:
--
文献类型:
--
作者:
Timothy A. Mann;Yoonsuck Choe

文献摘要

被引文献

相似文献

最近的强化学习(RL)算法,如R-MAX,只(以高概率)做出少量糟糕的决策。在实践中,这些算法不能很好地扩展状态的数量的增长,因为算法花费太多的精力探索。我们介绍了一种RL算法State Targeted R-MAX(STAR-MAX),它探索了状态空间的一个子集,称为探索包络线。当R-MAX等于总状态空间时,STAR-MAX的行为与R-MAX相同。当k是状态空间的子集时,为了将探索保持在k内,需要恢复规则β。我们比较了现有的算法与我们的算法采用各种探索信封。通过适当地选择最小值,随着状态数量的增加,STAR-MAX的扩展性能远远优于现有的RL算法。我们的算法的一个可能的缺点是它依赖于一个很好的选择的β和β。然而,我们表明,一个有效的恢复规则β可以在线学习,并可以从演示中学习。我们还发现,与R-MAX相比,即使是随机采样的探索包络也可以提高累积奖励。我们希望这些结果能够为大规模问题中的RL带来更有效的方法。
Recent Reinforcement Learning (RL) algorithms, such as R-MAX, make (with high probability) only a small number of poor decisions. In practice, these algorithms do not scale well as the number of states grows because the algorithms spend too much effort exploring. We introduce an RL algorithm State TArgeted R-MAX (STAR-MAX) that explores a subset of the state space, called the exploration envelope ξ. When ξ equals the total state space, STAR-MAX behaves identically to R-MAX. When ξ is a subset of the state space, to keep exploration within ξ, a recovery rule β is needed. We compared existing algorithms with our algorithm employing various exploration envelopes. With an appropriate choice of ξ, STAR-MAX scales far better than existing RL algorithms as the number of states increases. A possible drawback of our algorithm is its dependence on a good choice of ξ and β. However, we show that an effective recovery rule β can be learned on-line and ξ can be learned from demonstrations. We also find that even randomly sampled exploration envelopes can improve cumulative rewards compared to R-MAX. We expect these results to lead to more efficient methods for RL in large-scale problems.