Learning from Sparse Offline Datasets via Conservative Density Estimation

Learning from Sparse Offline Datasets via Conservative Density Estimation
复制标题

DOI:
10.48550/arxiv.2401.08819
复制
发表时间:
2024-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Zhepeng Cen;Zuxin Liu;Zitong Wang;Yi-Fan Yao;Henry Lam;Ding Zhao
Zhepeng Cen;Zuxin Liu;Zitong Wang;Yi-Fan Yao;Henry Lam;Ding Zhao
中科院分区:
其他
文献类型:
--
作者:
Zhepeng Cen;Zuxin Liu;Zitong Wang;Yi-Fan Yao;Henry Lam;Ding Zhao

文献摘要

相似文献

离线强化学习(RL)为从预先收集的数据集中学习策略提供了一个有希望的方向,而不需要与环境进行进一步的交互。然而,现有的方法很难处理分布外(OOD)外推错误,特别是在奖励稀少或数据稀缺的情况下。在本文中,我们提出了一种新的训练算法,称为保守密度估计(CDE),它通过显式地对状态-动作占用平稳分布施加约束来解决这一挑战。CDE通过解决边际重要性抽样中的支持失配问题,克服了现有方法的局限性,如平稳分布校正法。我们的方法在D4RL基准测试上获得了最先进的性能。值得注意的是,在奖励稀少或数据不足的挑战性任务中,CDE的表现始终优于基线,这表明了我们的方法在解决离线RL中的外推误差问题方面的优势。
Offline reinforcement learning (RL) offers a promising direction for learning policies from pre-collected datasets without requiring further interactions with the environment. However, existing methods struggle to handle out-of-distribution (OOD) extrapolation errors, especially in sparse reward or scarce data settings. In this paper, we propose a novel training algorithm called Conservative Density Estimation (CDE), which addresses this challenge by explicitly imposing constraints on the state-action occupancy stationary distribution. CDE overcomes the limitations of existing approaches, such as the stationary distribution correction method, by addressing the support mismatch issue in marginal importance sampling. Our method achieves state-of-the-art performance on the D4RL benchmark. Notably, CDE consistently outperforms baselines in challenging tasks with sparse rewards or insufficient data, demonstrating the advantages of our approach in addressing the extrapolation error problem in offline RL.