Unbiased Asymmetric Reinforcement Learning under Partial Observability

Unbiased Asymmetric Reinforcement Learning under Partial Observability
复制标题

DOI:
10.5555/3535850.3535857
复制
发表时间:
2021-05
期刊:
--
影响因子:
--
通讯作者:
Andrea Baisero;Chris Amato
Andrea Baisero;Chris Amato
中科院分区:
其他
文献类型:
--
作者:
Andrea Baisero;Chris Amato

文献摘要

相似文献

在部分可观察的强化学习中,离线训练可以访问在线训练和/或执行期间不可用的潜在信息,例如系统状态。非对称的行动者-批评者方法通过基于状态的批评者训练基于历史的策略来利用这些信息。然而,许多非对称方法缺乏理论基础,并且仅在有限的域上进行评估。我们研究了使用基于状态的批评者的非对称行动者-批评者方法的理论,并揭示了破坏常见变体有效性并限制其解决部分可观察性的能力的基本问题。我们提出了一个无偏的非对称演员批评的变体,它能够利用状态信息,同时保持理论上的合理性,保持政策梯度定理的有效性,并在训练过程中引入无偏和相对较低的方差。表现出显着的部分可观测性域进行的实证评估证实了我们的分析,表明无偏不对称的演员-评论家收敛到更好的政策和/或更快的对称和有偏不对称基线。
In partially observable reinforcement learning, offline training gives access to latent information which is not available during online training and/or execution, such as the system state. Asymmetric actor-critic methods exploit such information by training a history-based policy via a state-based critic. However, many asymmetric methods lack theoretical foundation, and are only evaluated on limited domains. We examine the theory of asymmetric actor-critic methods which use state-based critics, and expose fundamental issues which undermine the validity of a common variant, and limit its ability to address partial observability. We propose an unbiased asymmetric actor-critic variant which is able to exploit state information while remaining theoretically sound, maintaining the validity of the policy gradient theorem, and introducing no bias and relatively low variance into the training process. An empirical evaluation performed on domains which exhibit significant partial observability confirms our analysis, demonstrating that unbiased asymmetric actor-critic converges to better policies and/or faster than symmetric and biased asymmetric baselines.