A new criterion using information gain for action selection strategy in reinforcement learning

A new criterion using information gain for action selection strategy in reinforcement learning
复制标题

DOI:
10.1109/tnn.2004.828760
复制
发表时间:
2004-07
影响因子:
--
通讯作者:
Kazunori Iwata;K. Ikeda;H. Sakai
Kazunori Iwata;K. Ikeda;H. Sakai
中科院分区:
--
文献类型:
--
作者:
Kazunori Iwata;K. Ikeda;H. Sakai

文献摘要

被引文献

相似文献

在本文中,我们把收益序列看作是一个参数复合源的输出。利用这一事实,即源的编码率显示的信息量的回报,我们描述/spl lscr/学习算法的预测编码的思想的基础上,估计一个预期的信息增益有关的未来信息,并给出收敛证明的信息增益。利用信息增益,我们提出了一个新的标准,用于概率动作选择策略的比值/spl Ω/的回波损耗信息增益。在实验结果中,我们发现,我们的/spl Ω/-为基础的策略相比,传统的Q为基础的策略表现良好。
In this paper, we regard the sequence of returns as outputs from a parametric compound source. Utilizing the fact that the coding rate of the source shows the amount of information about the return, we describe /spl lscr/-learning algorithms based on the predictive coding idea for estimating an expected information gain concerning future information and give a convergence proof of the information gain. Using the information gain, we propose the ratio /spl omega/ of return loss to information gain as a new criterion to be used in probabilistic action-selection strategies. In experimental results, we found that our /spl omega/-based strategy performs well compared with the conventional Q-based strategy.