Sample-based learning and search with permanent and transient memories

Sample-based learning and search with permanent and transient memories
复制标题

基于样本的学习和搜索,具有永久和短暂的记忆

DOI:
--
复制
发表时间:
2008
期刊:
International Conference on Machine Learning
影响因子:
--
通讯作者:
Martin Müller
Martin Müller
中科院分区:
--
文献类型:
--
作者:
David Silver;R. Sutton;Martin Müller

文献摘要

被引文献

相似文献

我们提出了一种强化学习架构Dyna-2,它包含基于样本的学习和基于样本的搜索,并且在学习和搜索过程中跨状态进行概括。我们将Dyna-2应用于高性能计算机围棋。在这个领域中,最成功的规划方法是基于基于样本的搜索算法,如UCT,其中状态被单独处理,最成功的学习方法是基于时间差学习算法,如Sarsa,其中使用线性函数逼近。在这两种情况下,都形成了价值函数的估计,但在第一种情况下,它是短暂的,每次移动后都会计算然后丢弃,而在第二种情况下,它更持久,在许多移动和游戏中慢慢积累。Dyna-2的想法是暂时的计划记忆和永久的学习记忆保持分离,但两者都基于线性函数近似,并且都由Sarsa更新。为了将Dyna-2应用于9 x9 Computer Go,我们在函数逼近器中使用了一百万个二进制特征,这些特征基于匹配棋盘小片段的模板。仅使用瞬态记忆,Dyna-2至少表现得和UCT一样好。使用两种记忆相结合,它显着优于UCT。我们基于Dyna-2的程序在Computer Go Online Server上获得了比任何手工制作或传统搜索程序更高的评级。
We present a reinforcement learning architecture, Dyna-2, that encompasses both sample-based learning and sample-based search, and that generalises across states during both learning and search. We apply Dyna-2 to high performance Computer Go. In this domain the most successful planning methods are based on sample-based search algorithms, such as UCT, in which states are treated individually, and the most successful learning methods are based on temporal-difference learning algorithms, such as Sarsa, in which linear function approximation is used. In both cases, an estimate of the value function is formed, but in the first case it is transient, computed and then discarded after each move, whereas in the second case it is more permanent, slowly accumulating over many moves and games. The idea of Dyna-2 is for the transient planning memory and the permanent learning memory to remain separate, but for both to be based on linear function approximation and both to be updated by Sarsa. To apply Dyna-2 to 9x9 Computer Go, we use a million binary features in the function approximator, based on templates matching small fragments of the board. Using only the transient memory, Dyna-2 performed at least as well as UCT. Using both memories combined, it significantly outperformed UCT. Our program based on Dyna-2 achieved a higher rating on the Computer Go Online Server than any handcrafted or traditional search based program.