CEIP: Combining Explicit and Implicit Priors for Reinforcement Learning with Demonstrations

CEIP: Combining Explicit and Implicit Priors for Reinforcement Learning with Demonstrations
复制标题

DOI:
10.48550/arxiv.2210.09496
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Kai Yan;A. Schwing;Yu-Xiong Wang
Kai Yan;A. Schwing;Yu-Xiong Wang
中科院分区:
其他
文献类型:
--
作者:
Kai Yan;A. Schwing;Yu-Xiong Wang

文献摘要

相似文献

虽然强化学习在密集奖励环境中得到了广泛的应用,但用稀疏奖励训练自主代理仍然具有挑战性。为了解决这个困难,以前的工作已经显示出有希望的结果时,不仅使用特定于任务的演示,但任务不可知的,虽然有些相关的演示。在大多数情况下,可用的演示被提炼成一个隐含的先验,通常通过一个单一的深网表示。可以查询的数据库形式的显式先验也显示出令人鼓舞的结果。为了更好地从现有的演示中受益,我们开发了一种方法来结合显式和隐式先验(CEIP)的联合收割机。CEIP利用多个隐式先验以并行的标准化流的形式形成单个复杂先验。此外,CEIP使用了一个有效的显式检索和前推机制,以条件的隐式先验知识。在三个具有挑战性的环境中,我们发现所提出的CEIP方法,以改善先进的技术。
Although reinforcement learning has found widespread use in dense reward settings, training autonomous agents with sparse rewards remains challenging. To address this difficulty, prior work has shown promising results when using not only task-specific demonstrations but also task-agnostic albeit somewhat related demonstrations. In most cases, the available demonstrations are distilled into an implicit prior, commonly represented via a single deep net. Explicit priors in the form of a database that can be queried have also been shown to lead to encouraging results. To better benefit from available demonstrations, we develop a method to Combine Explicit and Implicit Priors (CEIP). CEIP exploits multiple implicit priors in the form of normalizing flows in parallel to form a single complex prior. Moreover, CEIP uses an effective explicit retrieval and push-forward mechanism to condition the implicit priors. In three challenging environments, we find the proposed CEIP method to improve upon sophisticated state-of-the-art techniques.