Finite Time Guarantees for Continuous State MDPs with Generative Model
Finite Time Guarantees for Continuous State MDPs with Generative Model
复制标题
DOI:
10.1109/cdc42340.2020.9303840
复制
发表时间:
2020-12
期刊:
影响因子:
--
通讯作者:
Hiteshi Sharma;R. Jain
中科院分区:
文献类型:
--
作者:
Hiteshi Sharma;R. Jain
In this paper, we present Online Empirical Value Learning (ONEVaL), an ‘online’ reinforcement learning algorithm for continuous MDPs that is ‘quasi-model-free’ (needs a generative/simulation model but not the model per se) that can compute nearly-optimal policies and comes with nonasymptotic performance guarantees including prescriptions on required sample complexity for specified performance bounds. The algorithm relies on use of a ‘fully’ randomized policy that will generate a β-mixing sample trajectory. It also relies on randomized function approximation in an RKHS for arbitrarily small function approximation error, and an ‘empirical’ estimate of value from the next state by several samples of the next state from the generative model. We demonstrate its’ good numerical performance on some benchmark problems. We note that the algorithm requires no hyper-parameter tuning, and is also robust to other concerns that seem to plague Deep RL algorithms.