Leveraging Fully Observable Policies for Learning under Partial Observability

Leveraging Fully Observable Policies for Learning under Partial Observability
复制标题

DOI:
10.48550/arxiv.2211.01991
复制
发表时间:
2022-11
期刊:
--
影响因子:
--
通讯作者:
Hai V. Nguyen;Andrea Baisero;Dian Wang;Chris Amato;Robert W. Platt
Hai V. Nguyen;Andrea Baisero;Dian Wang;Chris Amato;Robert W. Platt
中科院分区:
其他
文献类型:
--
作者:
Hai V. Nguyen;Andrea Baisero;Dian Wang;Chris Amato;Robert W. Platt

文献摘要

被引文献

相似文献

由于缺乏可观测的状态信息,部分可观测域中的强化学习具有挑战性。值得庆幸的是,在模拟器中离线学习这些状态信息通常是可能的。特别是,我们提出了一种部分可观察的强化学习方法,该方法在离线训练期间使用完全可观察的策略(我们称之为状态专家)来提高在线性能。基于软Actor-Critic(SAC),我们的代理平衡执行类似于状态专家的行动,并在部分可观测性下获得高回报。我们的方法可以利用完全可观察的策略进行探索,并且部分领域是完全可观察的,同时仍然能够在部分可观察性下学习。在六个机器人领域,我们的方法优于纯模仿,纯强化学习,这两种类型的顺序或并行组合,以及最近在相同设置中的最先进的方法。一个成功的政策转移到一个物理机器人在操作任务从像素显示我们的方法的实用性,在学习感兴趣的政策下的部分可观测性。
Reinforcement learning in partially observable domains is challenging due to the lack of observable state information. Thankfully, learning offline in a simulator with such state information is often possible. In particular, we propose a method for partially observable reinforcement learning that uses a fully observable policy (which we call a state expert) during offline training to improve online performance. Based on Soft Actor-Critic (SAC), our agent balances performing actions similar to the state expert and getting high returns under partial observability. Our approach can leverage the fully-observable policy for exploration and parts of the domain that are fully observable while still being able to learn under partial observability. On six robotics domains, our method outperforms pure imitation, pure reinforcement learning, the sequential or parallel combination of both types, and a recent state-of-the-art method in the same setting. A successful policy transfer to a physical robot in a manipulation task from pixels shows our approach's practicality in learning interesting policies under partial observability.