Adaptive Discretization for Model-Based Reinforcement Learning

Adaptive Discretization for Model-Based Reinforcement Learning
复制标题

基于模型的强化学习的自适应离散化

DOI:
--
复制
发表时间:
2020
期刊:
Neural Information Processing Systems
影响因子:
--
通讯作者:
C. Yu
C. Yu
中科院分区:
--
文献类型:
--
作者:
Sean R. Sinclair;Tianyu Wang;Gauri Jain;Siddhartha Banerjee;C. Yu

文献摘要

参考文献

被引文献

相似文献

在大的(潜在连续的)状态-动作空间中,我们引入自适应离散化技术来设计有效的基于模型的情节强化学习算法。我们的算法是基于乐观的一步值迭代扩展,以保持空间的自适应离散化。从理论上讲,我们为我们的算法提供了最坏情况下的后悔界限,与最先进的基于模型的算法相比是有竞争力的;此外,我们的界限是通过模块证明技术获得的,这可能会扩展到在问题上加入额外的结构。 从实现的角度来看,由于保持了状态和动作空间的更有效划分,我们的算法具有更低的存储和计算需求。我们通过几个典型控制问题的实验来说明这一点,实验表明,我们的算法在更快的收敛速度和更低的内存使用量方面明显优于固定离散化算法。有趣的是,我们根据经验观察到,虽然基于固定离散化的模型算法的性能远远优于无模型的算法,但这两种算法的性能与自适应离散化相当。
We introduce the technique of adaptive discretization to design efficient model-based episodic reinforcement learning algorithms in large (potentially continuous) state-action spaces. Our algorithm is based on optimistic one-step value iteration extended to maintain an adaptive discretization of the space. From a theoretical perspective, we provide worst-case regret bounds for our algorithm, which are competitive compared to the state-of-the-art model-based algorithms; moreover, our bounds are obtained via a modular proof technique, which can potentially extend to incorporate additional structure on the problem. From an implementation standpoint, our algorithm has much lower storage and computational requirements, due to maintaining a more efficient partition of the state and action spaces. We illustrate this via experiments on several canonical control problems, which shows that our algorithm empirically performs significantly better than fixed discretization in terms of both faster convergence and lower memory usage. Interestingly, we observe empirically that while fixed-discretization model-based algorithms vastly outperform their model-free counterparts, the two achieve comparable performance with adaptive discretization.
DOI: 10.1145/3366703
发表时间: 2019-10
期刊: Proceedings of the ACM on Measurement and Analysis of Computing Systems
影响因子: --
作者:
Sean R. Sinclair;Siddhartha Banerjee;C. Yu
通讯作者: Sean R. Sinclair;Siddhartha Banerjee;C. Yu
DOI: 10.1145/3299873
发表时间: 2019-08-01
期刊: JOURNAL OF THE ACM
影响因子: 2.5
作者:
Kleinberg, Robert;Slivkins, Aleksandrs;Upfal, Eli
通讯作者: Upfal, Eli