PAC Optimal Exploration in Continuous Space Markov Decision Processes

PAC Optimal Exploration in Continuous Space Markov Decision Processes
复制标题

DOI:
10.1609/aaai.v27i1.8678
复制
发表时间:
2013-06
期刊:
Proceedings of the AAAI Conference on Artificial Intelligence
影响因子:
--
通讯作者:
Jason Pazis;Ronald E. Parr
Jason Pazis;Ronald E. Parr
中科院分区:
其他
文献类型:
--
作者:
Jason Pazis;Ronald E. Parr

文献摘要

被引文献

相似文献

当前的探索算法可以分为两大类:启发式和PAC最优。虽然许多研究人员已经成功地使用了启发式方法,如ε贪婪探索,但这些方法缺乏正式的有限样本保证,可能需要大量的微调才能产生良好的结果。PAC最优探索算法,另一方面,提供了强有力的理论保证,但不适用于域的现实规模。本文的目标是弥合理论与实践之间的差距,通过引入C-PACE,算法,提供了强有力的理论保证,可应用于有趣的,连续的空间问题。
Current exploration algorithms can be classified in two broad categories: Heuristic, and PAC optimal. While numerous researchers have used heuristic approaches such as epsilon-greedy exploration successfully, such approaches lack formal, finite sample guarantees and may need a significant amount of fine-tuning to produce good results. PAC optimal exploration algorithms, on the other hand, offer strong theoretical guarantees but are inapplicable in domains of realistic size. The goal of this paper is to bridge the gap between theory and practice, by introducing C-PACE, an algorithm which offers strong theoretical guarantees and can be applied to interesting, continuous space problems.