Fast and Epsilon-Optimal Discretized Pursuit Learning Automata

Fast and Epsilon-Optimal Discretized Pursuit Learning Automata
复制标题

DOI:
10.1109/tcyb.2014.2365463
复制
发表时间:
2015-10
影响因子:
11.8
通讯作者:
Junqi Zhang;Cheng Wang;Mengchu Zhou
Junqi Zhang;Cheng Wang;Mengchu Zhou
中科院分区:
计算机科学1区
文献类型:
--
作者:
Junqi Zhang;Cheng Wang;Mengchu Zhou

文献摘要

被引文献

相似文献

学习自动机(LA)是强化学习的有力工具。离散化追求LA是其中最受欢迎的一种。在迭代过程中,它的操作包括三个基本阶段:1)选择下一个动作; 2)找到最佳估计动作; 3)更新状态概率。然而,当动作的数量很大时,学习变得非常缓慢,因为每次迭代都要进行太多的更新。增加的更新主要来自第1和第3阶段。提出了一种新的具有保证ε-最优性的快速离散追踪LA,以执行阶段1和3,其计算复杂度与动作的数量无关。除了它的低计算复杂度,它实现了更快的收敛速度比经典的在平稳环境中运行时。本文可以促进LA向大规模行动导向领域的应用,这需要高效的强化学习工具,确保ε-最优性,快速收敛速度,每次迭代的计算复杂度低。
Learning automata (LA) are powerful tools for reinforcement learning. A discretized pursuit LA is the most popular one among them. During an iteration its operation consists of three basic phases: 1) selecting the next action; 2) finding the optimal estimated action; and 3) updating the state probability. However, when the number of actions is large, the learning becomes extremely slow because there are too many updates to be made at each iteration. The increased updates are mostly from phases 1 and 3. A new fast discretized pursuit LA with assured ε-optimality is proposed to perform both phases 1 and 3 with the computational complexity independent of the number of actions. Apart from its low computational complexity, it achieves faster convergence speed than the classical one when operating in stationary environments. This paper can promote the applications of LA toward the large-scale-action oriented area that requires efficient reinforcement learning tools with assured ε-optimality, fast convergence speed, and low computational complexity for each iteration.