Optimistic Temporal Difference Learning for 2048

Optimistic Temporal Difference Learning for 2048
复制标题

2048 年乐观时间差异学习

DOI:
10.1109/tg.2021.3109887
复制
发表时间:
2021
影响因子:
2.3
通讯作者:
I
I
中科院分区:
计算机科学3区
文献类型:
--
作者:
Hung Guei;Lung;I

文献摘要

参考文献

被引文献

相似文献

时间差分(TD)学习及其变体,如多阶段TD学习和时间相干性(TC)学习,已成功应用于2048游戏。这些方法依靠2048游戏环境的随机性进行探索。在本文中,我们提议采用乐观初始化(OI)来鼓励对2048游戏的探索,并通过实验表明学习质量得到了显著提高。这种方法将特征权重乐观地初始化为非常大的值。由于一旦访问状态权重往往会降低,智能体往往会探索那些未被访问或访问次数很少的状态。我们的实验表明,带有乐观初始化的TD和TC学习都显著提高了性能。结果,达到相同性能所需的网络规模显著减小。通过诸如期望极大值搜索、多阶段学习和方块降级技术等额外的调整,我们的设计达到了最先进的性能,即平均得分625377,达到32768个方块的比率为72%。此外,在足够大的测试中,达到65536个方块的比率为0.02%。
Temporal difference (TD) learning and its variants, such as multistage TD learning and temporal coherence (TC) learning, have been successfully applied to 2048. These methods rely on the stochasticity of the environment of 2048 for exploration. In this article, we propose to employ optimistic initialization (OI) to encourage exploration for 2048, and empirically show that the learning quality is significantly improved. This approach optimistically initializes the feature weights to very large values. Since weights tend to be reduced once the states are visited, agents tend to explore those states which are unvisited or visited few times. Our experiments show that both TD and TC learning with OI significantly improve the performance. As a result, the network size required to achieve the same performance is significantly reduced. With additional tunings such as expectimax search, multistage learning, and tile-downgrading technique, our design achieves the state-of-the-art performance, namely an average score of 625 377 and a rate of 72% reaching 32 768-tiles. In addition, for sufficiently large tests, 65 536-tiles are reached at a rate of 0.02%.
DOI: --
发表时间: 2006
期刊: --
影响因子: --
作者:
Iroon Polytechniou-
通讯作者: Iroon Polytechniou-