Temporal Disentanglement of Representations for Improved Generalisation in Reinforcement Learning

Temporal Disentanglement of Representations for Improved Generalisation in Reinforcement Learning
复制标题

DOI:
10.48550/arxiv.2207.05480
复制
发表时间:
2022-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Mhairi Dunion;Trevor A. McInroe;K. Luck;Josiah P. Hanna;Stefano V. Albrecht
Mhairi Dunion;Trevor A. McInroe;K. Luck;Josiah P. Hanna;Stefano V. Albrecht
中科院分区:
其他
文献类型:
--
作者:
Mhairi Dunion;Trevor A. McInroe;K. Luck;Josiah P. Hanna;Stefano V. Albrecht

文献摘要

被引文献

相似文献

强化学习(RL)代理通常不能很好地概括在训练期间没有观察到的状态空间中的环境变化。这个问题对于基于图像的RL来说尤其成问题,在RL中,只改变一个变量,例如背景颜色,就会改变图像中的许多像素。改变的像素可能会导致代理对图像的潜在表征发生剧烈变化,导致学习到的策略失败。为了学习更稳健的表示,我们引入了时间解缠(TED),这是一种自我监督的辅助任务,利用RL观测的顺序性质导致解缠的图像表示。我们发现,与最先进的表征学习方法相比,使用TED作为辅助任务的RL算法通过持续训练更快地适应环境变量的变化。由于TED强制实施了一种分离的表示结构,我们的实验也表明,用TED训练的策略更好地泛化到与任务无关的变量的不可见值(例如背景颜色)以及影响最优策略的变量的不可见值(例如目标位置)。
Reinforcement Learning (RL) agents are often unable to generalise well to environment variations in the state space that were not observed during training. This issue is especially problematic for image-based RL, where a change in just one variable, such as the background colour, can change many pixels in the image. The changed pixels can lead to drastic changes in the agent's latent representation of the image, causing the learned policy to fail. To learn more robust representations, we introduce TEmporal Disentanglement (TED), a self-supervised auxiliary task that leads to disentangled image representations exploiting the sequential nature of RL observations. We find empirically that RL algorithms utilising TED as an auxiliary task adapt more quickly to changes in environment variables with continued training compared to state-of-the-art representation learning methods. Since TED enforces a disentangled structure of the representation, our experiments also show that policies trained with TED generalise better to unseen values of variables irrelevant to the task (e.g. background colour) as well as unseen values of variables that affect the optimal policy (e.g. goal positions).