Extending the Peak Bandwidth of Parameters for Softmax Selection in Reinforcement Learning

Extending the Peak Bandwidth of Parameters for Softmax Selection in Reinforcement Learning
复制标题

DOI:
10.1109/tnnls.2016.2558295
复制
发表时间:
2017-08
影响因子:
10.4
通讯作者:
Kazunori Iwata
Kazunori Iwata
中科院分区:
计算机科学1区
文献类型:
--
作者:
Kazunori Iwata

文献摘要

被引文献

相似文献

Softmax选择是强化学习中最流行的动作选择方法之一。尽管最近提出的各种方法在完全参数调整的情况下可能更有效,但是实现需要调整许多参数的复杂方法可能是困难的。因此,考虑到其实现和调优的成本节约,softmax选择仍然值得重新考虑。实际上,这种方法在实际中只要为环境设置一个适当的参数就足够了。本文的目的是改进该方法的变量设置,以扩展好参数的带宽,从而降低实现和参数整定的成本。为了实现这一点,我们利用马尔可夫决策过程中的渐近均分特性来扩展softmax选择的峰值带宽。使用各种情节的任务,我们表明,我们的设置是有效的,在扩展带宽,它产生了更好的政策,在稳定性。在一系列统计测试中定量评估带宽。
Softmax selection is one of the most popular methods for action selection in reinforcement learning. Although various recently proposed methods may be more effective with full parameter tuning, implementing a complicated method that requires the tuning of many parameters can be difficult. Thus, softmax selection is still worth revisiting, considering the cost savings of its implementation and tuning. In fact, this method works adequately in practice with only one parameter appropriately set for the environment. The aim of this paper is to improve the variable setting of this method to extend the bandwidth of good parameters, thereby reducing the cost of implementation and parameter tuning. To achieve this, we take advantage of the asymptotic equipartition property in a Markov decision process to extend the peak bandwidth of softmax selection. Using a variety of episodic tasks, we show that our setting is effective in extending the bandwidth and that it yields a better policy in terms of stability. The bandwidth is quantitatively assessed in a series of statistical tests.