A Finite-Time Analysis of Q-Learning with Neural Network Function Approximation

A Finite-Time Analysis of Q-Learning with Neural Network Function Approximation
复制标题

DOI:
--
复制
发表时间:
2019-12
影响因子:
5.8
通讯作者:
Pan Xu;Quanquan Gu
Pan Xu;Quanquan Gu
中科院分区:
医学2区
文献类型:
--
作者:
Pan Xu;Quanquan Gu

文献摘要

相似文献

具有神经网络函数逼近的Q学习(简称神经Q学习)是最流行的深度强化学习算法之一。尽管它的经验成功,神经Q学习的非渐近收敛速度仍然几乎是未知的。在本文中,我们提出了一种神经Q学习算法的有限时间分析,其中数据由马尔可夫决策过程生成,动作值函数由深度ReLU神经网络近似。我们证明了神经Q学习找到最优策略的$O(1/\sqrt{T})$收敛速度,如果神经函数逼近器是充分overparameterized,其中$T$是迭代次数。据我们所知,我们的结果是第一个有限时间分析下的非独立同分布神经Q学习。数据假设
Q-learning with neural network function approximation (neural Q-learning for short) is among the most prevalent deep reinforcement learning algorithms. Despite its empirical success, the non-asymptotic convergence rate of neural Q-learning remains virtually unknown. In this paper, we present a finite-time analysis of a neural Q-learning algorithm, where the data are generated from a Markov decision process and the action-value function is approximated by a deep ReLU neural network. We prove that neural Q-learning finds the optimal policy with $O(1/\sqrt{T})$ convergence rate if the neural function approximator is sufficiently overparameterized, where $T$ is the number of iterations. To our best knowledge, our result is the first finite-time analysis of neural Q-learning under non-i.i.d. data assumption.