Experiments with Reinforcement Learning in Problems with Continuous State and Action Spaces

Experiments with Reinforcement Learning in Problems with Continuous State and Action Spaces
复制标题

DOI:
10.1177/105971239700600201
复制
发表时间:
1997-09
期刊:
影响因子:
1.6
通讯作者:
J. Santamaría;R. Sutton;A. Ram
J. Santamaría;R. Sutton;A. Ram
中科院分区:
计算机科学4区
文献类型:
--
作者:
J. Santamaría;R. Sutton;A. Ram

文献摘要

被引文献

相似文献

解决强化学习问题的一个关键因素是值函数。这个函数的目的是衡量任何给定状态的长期效用或价值。这个函数很重要,因为代理可以使用这个度量来决定下一步要做什么。当应用于具有连续状态和动作空间的系统时,强化学习中的一个常见问题是值函数必须在由实值变量组成的域上运行,这意味着它应该能够表示无限多个状态和动作对的值。由于这个原因,当最优策略的接近解不可用时,函数逼近器被用来表示值函数。在本文中,我们扩展了先前提出的强化学习算法,以便它可以与函数近似器一起使用,该近似器可以在状态和动作空间中概括个人经验的值。特别地,我们讨论了使用稀疏粗编码函数逼近器来表示值函数的好处,并详细描述了三种实现:小脑模型连接控制器、基于实例和基于案例。此外,我们还讨论了在状态和动作空间的不同区域具有不同分辨率的函数逼近器如何影响代理的性能和学习效率。我们提出了一种简单的模块化技术,可用于实现具有非均匀分辨率的函数逼近器,以便在状态和动作空间的重要区域中以更高的精度表示值函数。我们在双积分器和钟摆摆动系统中进行了大量的实验来证明所提出的想法。”
A key element in the solution of reinforcement learning problems is the value function. The purpose of this function is to measure the long-term utility or value of any given state. The function is important because an agent can use this measure to decide what to do next. A common problem in reinforcement learning when applied to systems having continuous states and action spaces is that the value function must operate with a domain consisting of real-valued variables, which means that it should be able to represent the value of infinitely many state and action pairs. For this reason, function approximators are used to represent the value function when a close-form solution of the optimal policy is not available. In this article, we extend a previously proposed reinforcement learning algorithm so that it can be used with function approximators that generalize the value of individual experiences across both state and action spaces. In particular, we discuss the benefits of using sparse coarse-coded function approximators to represent value functions and describe in detail three implementations: cerebellar model articulation controllers, instance-based, and case-based. Additionally, we discuss how function approximators having different degrees of resolution in different regions of the state and action spaces may influence the performance and learning efficiency of the agent. We propose a simple and modular technique that can be used to implement function approximators with nonuniform degrees of resolution so that the value function can be represented with higher accuracy in important regions of the state and action spaces. We performed extensive experiments in the double-integrator and pendulum swing-up systems to demonstrate the proposed ideas. '