Finite Time Guarantees for Continuous State MDPs with Generative Model

Finite Time Guarantees for Continuous State MDPs with Generative Model
复制标题

DOI:
10.1109/cdc42340.2020.9303840
复制
发表时间:
2020-12
期刊:
2020 59th IEEE Conference on Decision and Control (CDC)
影响因子:
--
通讯作者:
Hiteshi Sharma;R. Jain
Hiteshi Sharma;R. Jain
中科院分区:
其他
文献类型:
--
作者:
Hiteshi Sharma;R. Jain

文献摘要

相似文献

在本文中,我们提出了在线经验值学习(ONEVaL),一个“在线”的强化学习算法的连续MDP是“准无模型”(需要一个生成/模拟模型,但不是模型本身),可以计算接近最优的政策,并配备了非渐近性能保证,包括规定所需的样本复杂度为指定的性能界限。该算法依赖于使用“完全”随机化策略,该策略将生成β混合样本轨迹。它还依赖于RKHS中的随机函数近似,用于任意小的函数近似误差,以及通过来自生成模型的下一状态的几个样本对下一状态的值的“经验”估计。在一些基准问题上,我们证明了它的良好的数值性能。我们注意到,该算法不需要超参数调整,并且对于似乎困扰深度RL算法的其他问题也具有鲁棒性。
In this paper, we present Online Empirical Value Learning (ONEVaL), an ‘online’ reinforcement learning algorithm for continuous MDPs that is ‘quasi-model-free’ (needs a generative/simulation model but not the model per se) that can compute nearly-optimal policies and comes with nonasymptotic performance guarantees including prescriptions on required sample complexity for specified performance bounds. The algorithm relies on use of a ‘fully’ randomized policy that will generate a β-mixing sample trajectory. It also relies on randomized function approximation in an RKHS for arbitrarily small function approximation error, and an ‘empirical’ estimate of value from the next state by several samples of the next state from the generative model. We demonstrate its’ good numerical performance on some benchmark problems. We note that the algorithm requires no hyper-parameter tuning, and is also robust to other concerns that seem to plague Deep RL algorithms.