Action Branching Architectures for Deep Reinforcement Learning

Action Branching Architectures for Deep Reinforcement Learning
复制标题

DOI:
10.1609/aaai.v32i1.11798
复制
发表时间:
2017-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Arash Tavakoli;Fabio Pardo;Petar Kormushev
Arash Tavakoli;Fabio Pardo;Petar Kormushev
中科院分区:
其他
文献类型:
--
作者:
Arash Tavakoli;Fabio Pardo;Petar Kormushev

文献摘要

被引文献

相似文献

离散动作算法是深度强化学习最近取得的许多成功的核心。然而,将这些算法应用于高维动作任务需要处理可能动作的数量与动作维度的数量的组合增加。对于需要通过离散化对动作进行精细控制的连续动作任务,该问题进一步恶化。在本文中,我们提出了一种新的神经架构,具有一个共享的决策模块,然后是几个网络分支,每个动作维度一个。这种方法通过允许每个单独的动作维度具有一定程度的独立性,实现了网络输出数量与自由度数量的线性增加。为了说明这种方法,我们提出了一种新的代理,称为分支决斗Q网络(BDQ),作为决斗双深Q网络(决斗DDQN)的分支变体。我们评估我们的代理在一组具有挑战性的连续控制任务的性能。实证结果表明,所提出的代理规模优雅的环境与行动维度的增加,并表明在协调分布式行动分支的共享决策模块的意义。此外,我们表明,所提出的代理执行竞争力对国家的最先进的连续控制算法,深度确定性策略梯度(DDPG)。
Discrete-action algorithms have been central to numerous recent successes of deep reinforcement learning. However, applying these algorithms to high-dimensional action tasks requires tackling the combinatorial increase of the number of possible actions with the number of action dimensions. This problem is further exacerbated for continuous-action tasks that require fine control of actions via discretization. In this paper, we propose a novel neural architecture featuring a shared decision module followed by several network branches, one for each action dimension. This approach achieves a linear increase of the number of network outputs with the number of degrees of freedom by allowing a level of independence for each individual action dimension. To illustrate the approach, we present a novel agent, called Branching Dueling Q-Network (BDQ), as a branching variant of the Dueling Double Deep Q-Network (Dueling DDQN). We evaluate the performance of our agent on a set of challenging continuous control tasks. The empirical results show that the proposed agent scales gracefully to environments with increasing action dimensionality and indicate the significance of the shared decision module in coordination of the distributed action branches. Furthermore, we show that the proposed agent performs competitively against a state-of-the-art continuous control algorithm, Deep Deterministic Policy Gradient (DDPG).