Does Self-supervised Learning Really Improve Reinforcement Learning from Pixels?

Does Self-supervised Learning Really Improve Reinforcement Learning from Pixels?
复制标题

DOI:
10.48550/arxiv.2206.05266
复制
发表时间:
2022-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Xiang Li;Jinghuan Shang;Srijan Das;M. Ryoo
Xiang Li;Jinghuan Shang;Srijan Das;M. Ryoo
中科院分区:
其他
文献类型:
--
作者:
Xiang Li;Jinghuan Shang;Srijan Das;M. Ryoo

文献摘要

被引文献

相似文献

我们研究了自监督学习(SSL)是否可以改善像素的在线强化学习(RL)。我们扩展了对比强化学习框架(例如,CURL),共同优化SSL和RL损失,并对各种自监督损失进行大量实验。我们的观察结果表明,现有的RL SSL框架在使用相同数量的数据和增强时,仅利用图像增强,无法带来比基线有意义的改进。我们进一步进行进化搜索,以找到RL的多个自监督损失的最佳组合,但发现即使这样的损失组合也未能有意义地优于仅利用精心设计的图像增强的方法。在包括真实世界机器人环境在内的多个不同环境中对这些方法进行评估后,我们确认没有任何一种自监督丢失或图像增强方法可以主导所有环境,并且当前SSL和RL联合优化的框架是有限的。最后,我们对多个因素进行了消融研究,并展示了用不同方法学习的表征的性质。
We investigate whether self-supervised learning (SSL) can improve online reinforcement learning (RL) from pixels. We extend the contrastive reinforcement learning framework (e.g., CURL) that jointly optimizes SSL and RL losses and conduct an extensive amount of experiments with various self-supervised losses. Our observations suggest that the existing SSL framework for RL fails to bring meaningful improvement over the baselines only taking advantage of image augmentation when the same amount of data and augmentation is used. We further perform evolutionary searches to find the optimal combination of multiple self-supervised losses for RL, but find that even such a loss combination fails to meaningfully outperform the methods that only utilize carefully designed image augmentations. After evaluating these approaches together in multiple different environments including a real-world robot environment, we confirm that no single self-supervised loss or image augmentation method can dominate all environments and that the current framework for joint optimization of SSL and RL is limited. Finally, we conduct the ablation study on multiple factors and demonstrate the properties of representations learned with different approaches.