Domain Adaptation In Reinforcement Learning Via Latent Unified State Representation

Domain Adaptation In Reinforcement Learning Via Latent Unified State Representation
复制标题

DOI:
10.1609/aaai.v35i12.17251
复制
发表时间:
2021-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Jinwei Xing;Takashi Nagata;Kexin Chen;Xinyun Zou;E. Neftci;J. Krichmar
Jinwei Xing;Takashi Nagata;Kexin Chen;Xinyun Zou;E. Neftci;J. Krichmar
中科院分区:
其他
文献类型:
--
作者:
Jinwei Xing;Takashi Nagata;Kexin Chen;Xinyun Zou;E. Neftci;J. Krichmar

文献摘要

相似文献

尽管深度强化学习(RL)最近取得了成功,但领域适应仍然是一个悬而未决的问题。尽管强化学习智能体的泛化能力对于深度强化学习的现实应用至关重要,但零样本策略迁移仍然是一个具有挑战性的问题,因为即使是微小的视觉变化也可能使经过训练的智能体在新任务中完全失败。为了解决这个问题,我们提出了一个两阶段的 RL 代理,它首先在第一阶段学习跨多个域一致的潜在统一状态表示(LUSR),然后在第二阶段基于 LUSR 在一个源域中进行 RL 训练。 LUSR 的跨域一致性允许从源域获取的策略推广到其他目标域,而无需额外的训练。我们首先在具有定制操作的赛车游戏变体中展示我们的方法,然后在 CARLA 中进行验证,CARLA 是一个具有更复杂和真实视觉观察的自动驾驶模拟器。我们的结果表明,这种方法可以在相关 RL 任务中实现最先进的域适应性能,并且优于基于潜在表示的 RL 和图像到图像翻译的现有方法。
Despite the recent success of deep reinforcement learning (RL), domain adaptation remains an open problem. Although the generalization ability of RL agents is critical for the real-world applicability of Deep RL, zero-shot policy transfer is still a challenging problem since even minor visual changes could make the trained agent completely fail in the new task. To address this issue, we propose a two-stage RL agent that first learns a latent unified state representation (LUSR) which is consistent across multiple domains in the first stage, and then do RL training in one source domain based on LUSR in the second stage. The cross-domain consistency of LUSR allows the policy acquired from the source domain to generalize to other target domains without extra training. We first demonstrate our approach in variants of CarRacing games with customized manipulations, and then verify it in CARLA, an autonomous driving simulator with more complex and realistic visual observations. Our results show that this approach can achieve state-of-the-art domain adaptation performance in related RL tasks and outperforms prior approaches based on latent-representation based RL and image-to-image translation.