SO(2)-Equivariant Reinforcement Learning

SO(2)-Equivariant Reinforcement Learning
复制标题

DOI:
10.48550/arxiv.2203.04439
复制
发表时间:
2022-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Dian Wang;R. Walters;Robert W. Platt
Dian Wang;R. Walters;Robert W. Platt
中科院分区:
其他
文献类型:
--
作者:
Dian Wang;R. Walters;Robert W. Platt

文献摘要

相似文献

等变神经网络在其卷积层的结构中强制对称,从而在学习等变或不变函数时大大提高了样本效率。这样的模型是适用于机器人操作学习,通常可以制定为一个旋转对称的问题。本文研究的背景下,$Q$-学习和actor-critic强化学习的同变模型架构。我们确定等变和不变的最佳$Q$-功能和最佳的政策,并提出等变DQN和SAC算法,利用这种结构。我们目前的实验表明,我们的等变版本的DQN和SAC可以显着更有效的样本比竞争算法的一类重要的机器人操作问题。
Equivariant neural networks enforce symmetry within the structure of their convolutional layers, resulting in a substantial improvement in sample efficiency when learning an equivariant or invariant function. Such models are applicable to robotic manipulation learning which can often be formulated as a rotationally symmetric problem. This paper studies equivariant model architectures in the context of $Q$-learning and actor-critic reinforcement learning. We identify equivariant and invariant characteristics of the optimal $Q$-function and the optimal policy and propose equivariant DQN and SAC algorithms that leverage this structure. We present experiments that demonstrate that our equivariant versions of DQN and SAC can be significantly more sample efficient than competing algorithms on an important class of robotic manipulation problems.