On-line Policy Improvement using Monte-Carlo Search

On-line Policy Improvement using Monte-Carlo Search
复制标题

DOI:
--
复制
发表时间:
1996-12
期刊:
--
影响因子:
--
通讯作者:
G. Tesauro;Gregory R. Galperin
G. Tesauro;Gregory R. Galperin
中科院分区:
其他
文献类型:
--
作者:
G. Tesauro;Gregory R. Galperin

文献摘要

被引文献

相似文献

我们提出了一种蒙特 - 卡洛模拟算法,以改善自适应控制器的实时策略。在蒙特 - 卡洛模拟中,使用初始策略在模拟的每个步骤中做出决策,从统计上测量了每种可能动作的长期预期奖励。然后采取最大程度地提高预期奖励的动作,从而改善政策。我们的算法很容易并行,并且已在IBM SP1和SP2并行RISC超级计算机上实现。我们在将此算法应用于Backgammon领域时获得了有希望的初始结果。报告了各种初始策略的结果,从随机策略到TD-Gammon,这是一个极强的多层神经网络。在每种情况下,蒙特 - 卡洛算法在基本参与者的错误率下,大幅减少了5倍或更多。该算法在许多其他自适应控制应用程序中也可能有用,在这些应用程序中可以模拟环境。
We present a Monte-Carlo simulation algorithm for real-time policy improvement of an adaptive controller. In the Monte-Carlo simulation, the long-term expected reward of each possible action is statistically measured, using the initial policy to make decisions in each step of the simulation. The action maximizing the measured expected reward is then taken, resulting in an improved policy. Our algorithm is easily parallelizable and has been implemented on the IBM SP1 and SP2 parallel-RISC supercomputers. We have obtained promising initial results in applying this algorithm to the domain of backgammon. Results are reported for a wide variety of initial policies, ranging from a random policy to TD-Gammon, an extremely strong multi-layer neural network. In each case, the Monte-Carlo algorithm gives a substantial reduction, by as much as a factor of 5 or more, in the error rate of the base players. The algorithm is also potentially useful in many other adaptive control applications in which it is possible to simulate the environment.