Breaking the Sample Size Barrier in Model-Based Reinforcement Learning with a Generative Model

Breaking the Sample Size Barrier in Model-Based Reinforcement Learning with a Generative Model
复制标题

DOI:
10.1287/opre.2023.2451
复制
发表时间:
2020-05
期刊:
Oper. Res.
影响因子:
--
通讯作者:
Gen Li;Yuting Wei;Yuejie Chi;Yuantao Gu;Yuxin Chen
Gen Li;Yuting Wei;Yuejie Chi;Yuantao Gu;Yuxin Chen
中科院分区:
其他
文献类型:
--
作者:
Gen Li;Yuting Wei;Yuejie Chi;Yuantao Gu;Yuxin Chen

文献摘要

被引文献

相似文献

本文研究了现代强化学习,样本效率的一个核心问题,并在解决理想主义方案方面取得了进步,该场景假设使用生成模型或模拟器。尽管有大量的先前工作解决了这个问题,但尚未确定样本复杂性和统计准确性之间的权衡。特别是,所有先前的结果都遭受严重的样本障碍,因为他们所声称的统计保证只有在样本量超过一定的阈值时才能保证。当前的论文克服了这一障碍并完全解决了这个问题。更具体地说,我们在任何给定的目标准确性水平上建立了基于模型的方法的最小值。据我们所知,这项工作提供了第一个最小值的最佳选择,可以保证适应整个样本范围的范围(除此之外,找到有意义的政策是从理论上讲是信息)。
This paper studies a central issue in modern reinforcement learning, the sample efficiency, and makes progress toward solving an idealistic scenario that assumes access to a generative model or a simulator. Despite a large number of prior works tackling this problem, a complete picture of the trade-offs between sample complexity and statistical accuracy has yet to be determined. In particular, all prior results suffer from a severe sample size barrier in the sense that their claimed statistical guarantees hold only when the sample size exceeds some enormous threshold. The current paper overcomes this barrier and fully settles this problem; more specifically, we establish the minimax optimality of the model-based approach for any given target accuracy level. To the best of our knowledge, this work delivers the first minimax-optimal guarantees that accommodate the entire range of sample sizes (beyond which finding a meaningful policy is information theoretically infeasible).