Q-functionals for Value-Based Continuous Control

Q-functionals for Value-Based Continuous Control
复制标题

DOI:
10.1609/aaai.v37i7.26073
复制
发表时间:
2023-06
期刊:
--
影响因子:
--
通讯作者:
Bowen He;Sam Lobel;Sreehari Rammohan;Shangqun Yu;G. Konidaris
Bowen He;Sam Lobel;Sreehari Rammohan;Shangqun Yu;G. Konidaris
中科院分区:
其他
文献类型:
--
作者:
Bowen He;Sam Lobel;Sreehari Rammohan;Shangqun Yu;G. Konidaris

文献摘要

相似文献

我们提出了Q-泛函,这是一种用于连续控制深度强化学习的替代架构。我们的网络不是为状态-动作对返回单个值,而是将状态转换为一个函数,该函数可以对许多动作进行快速并行评估,从而使我们能够通过采样有效地选择高值动作。这与非策略连续控制的典型架构形成对比,在非策略连续控制中,策略网络的唯一目的是从Q函数中选择动作。我们将动作相关的Q函数表示为动作空间上的基函数(傅立叶,多项式等)的加权和,其中权重是状态相关的,并由Q函数网络输出。快速采样使得需要在Q函数上进行蒙特-卡罗积分的各种技术变得实用,并且使得除了简单的值最大化之外的动作选择策略成为可能。我们的框架的特点,描述了各种实现的Q-泛函,并展示了强大的性能上的一套连续控制任务。
We present Q-functionals, an alternative architecture for continuous control deep reinforcement learning. Instead of returning a single value for a state-action pair, our network transforms a state into a function that can be rapidly evaluated in parallel for many actions, allowing us to efficiently choose high-value actions through sampling. This contrasts with the typical architecture of off-policy continuous control, where a policy network is trained for the sole purpose of selecting actions from the Q-function. We represent our action-dependent Q-function as a weighted sum of basis functions (Fourier, Polynomial, etc) over the action space, where the weights are state-dependent and output by the Q-functional network. Fast sampling makes practical a variety of techniques that require Monte-Carlo integration over Q-functions, and enables action-selection strategies besides simple value-maximization. We characterize our framework, describe various implementations of Q-functionals, and demonstrate strong performance on a suite of continuous control tasks.