Sample Complexity and Overparameterization Bounds for Temporal-Difference Learning With Neural Network Approximation

Sample Complexity and Overparameterization Bounds for Temporal-Difference Learning With Neural Network Approximation
复制标题

DOI:
10.1109/tac.2023.3234234
复制
发表时间:
2021-03
影响因子:
6.8
通讯作者:
Semih Cayci;Siddhartha Satpathi;Niao He;F. I. R. Srikant
Semih Cayci;Siddhartha Satpathi;Niao He;F. I. R. Srikant
中科院分区:
计算机科学2区
文献类型:
--
作者:
Semih Cayci;Siddhartha Satpathi;Niao He;F. I. R. Srikant

文献摘要

相似文献

在这篇文章中,我们研究了一般状态空间上基于神经网络的值函数逼近的时间差(TD)学习的动力学,即神经TD学习。我们考虑两个实际使用的算法,投影自由和最大范数正则神经TD学习,并建立这些算法的第一收敛界。从我们的研究结果中一个有趣的观察是,最大范数正则化可以显着提高TD学习算法的样本复杂度和overparameterization方面的性能。在这项工作中的结果依赖于一个李雅普诺夫漂移分析的网络参数作为一个停止和控制的随机过程。
In this article, we study the dynamics of temporal-difference (TD) learning with neural network-based value function approximation over a general state space, namely, neural TD learning. We consider two practically used algorithms, projection-free and max-norm regularized neural TD learning, and establish the first convergence bounds for these algorithms. An interesting observation from our results is that max-norm regularization can dramatically improve the performance of TD learning algorithms in terms of sample complexity and overparameterization. The results in this work rely on a Lyapunov drift analysis of the network parameters as a stopped and controlled random process.