Mastering 2048 With Delayed Temporal Coherence Learning, Multistage Weight Promotion, Redundant Encoding, and Carousel Shaping

Mastering 2048 With Delayed Temporal Coherence Learning, Multistage Weight Promotion, Redundant Encoding, and Carousel Shaping
复制标题

通过延迟时间一致性学习、多级权重提升、冗余编码和轮播整形掌握 2048

DOI:
10.1109/tciaig.2017.2651887
复制
发表时间:
2016
影响因子:
2.3
通讯作者:
Wojciech Jaśkowski
Wojciech Jaśkowski
中科院分区:
计算机科学3区
文献类型:
--
作者:
Wojciech Jaśkowski

文献摘要

被引文献

相似文献

《2048》是一款引人入胜的单人非确定性视频益智游戏,由于其规则简单但玩法难以精通,近年来大受欢迎。由于《2048》可以很方便地嵌入到离散状态马尔可夫决策过程框架中,我们将其作为评估强化学习中现有方法和新方法的试验平台。为了开发一个强大的《2048》游戏程序,我们采用了带有系统的$n$元组网络的时序差分学习。我们表明,通过时间相干性学习、带有权重提升的多级函数逼近器、旋转木马式塑造以及冗余编码,这种基本方法可以得到显著改进。此外,我们演示了如何利用$n$元组网络的特性,通过延迟(衰减的)更新以及应用无锁乐观并行性来轻松利用多个CPU核心,从而提高学习过程的算法效率。通过这种方式,我们能够开发出迄今为止已知的最佳《2048》游戏程序,这证实了所引入的方法对于离散状态马尔可夫决策问题的有效性。
<italic>2048</italic> is an engaging single-player nondeterministic video puzzle game, which, thanks to the simple rules and hard-to-master gameplay, has gained massive popularity in recent years. As <italic>2048</italic> can be conveniently embedded into the discrete-state Markov decision processes framework, we treat it as a testbed for evaluating existing and new methods in reinforcement learning. With the aim to develop a strong <italic>2048</italic> playing program, we employ temporal difference learning with systematic <inline-formula><tex-math notation="LaTeX">$n$ </tex-math></inline-formula>-tuple networks. We show that this basic method can be significantly improved with temporal coherence learning, multistage function approximator with weight promotion, carousel shaping, and redundant encoding. In addition, we demonstrate how to take advantage of the characteristics of the <inline-formula> <tex-math notation="LaTeX">$n$</tex-math></inline-formula>-tuple network, to improve the algorithmic effectiveness of the learning process by delaying the (decayed) update and applying lock-free optimistic parallelism to effortlessly make advantage of multiple CPU cores. This way, we were able to develop the best known <italic>2048</italic> playing program to date, which confirms the effectiveness of the introduced methods for discrete-state Markov decision problems.