Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning

Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning
复制标题

DOI:
10.1023/a:1022672621406
复制
发表时间:
2004
期刊:
影响因子:
7.5
通讯作者:
Ronald J. Williams
Ronald J. Williams
中科院分区:
计算机科学3区
文献类型:
--
作者:
Ronald J. Williams

文献摘要

被引文献

相似文献

本文针对含有随机单元的连接型网络,提出了一类通用的联想强化学习算法。这些算法被称为加强算法,它们在立即加强任务和某些有限形式的延迟加强任务中,在沿着预期加强的梯度的方向上进行权重调整,并且它们这样做时没有显式地计算梯度估计,甚至没有存储可以计算这种估计的信息。给出了这些算法的具体例子,其中一些与某些现有的算法有着密切的关系,而另一些算法是新颖的,但本身就可能很有趣。还给出了这样的算法如何自然地与反向传播相结合的结果。最后,我们简要讨论了围绕这种算法的使用的一些额外问题,包括关于它们的限制行为的已知情况,以及可能用于帮助开发类似但潜在更强大的强化学习算法的进一步考虑。
This article presents a general class of associative reinforcement learning algorithms for connectionist networks containing stochastic units. These algorithms, called REINFORCE algorithms, are shown to make weight adjustments in a direction that lies along the gradient of expected reinforcement in both immediate-reinforcement tasks and certain limited forms of delayed-reinforcement tasks, and they do this without explicitly computing gradient estimates or even storing information from which such estimates could be computed. Specific examples of such algorithms are presented, some of which bear a close relationship to certain existing algorithms while others are novel but potentially interesting in their own right. Also given are results that show how such algorithms can be naturally integrated with backpropagation. We close with a brief discussion of a number of additional issues surrounding the use of such algorithms, including what is known about their limiting behaviors as well as further considerations that might be used to help develop similar but potentially more powerful reinforcement learning algorithms.