Going Beyond Linear RL: Sample Efficient Neural Function Approximation

Going Beyond Linear RL: Sample Efficient Neural Function Approximation
复制标题

DOI:
--
复制
发表时间:
2021-07
期刊:
--
影响因子:
--
通讯作者:
Baihe Huang;Kaixuan Huang;S. Kakade;Jason D. Lee;Qi Lei;Runzhe Wang;Jiaqi Yang
Baihe Huang;Kaixuan Huang;S. Kakade;Jason D. Lee;Qi Lei;Runzhe Wang;Jiaqi Yang
中科院分区:
其他
文献类型:
--
作者:
Baihe Huang;Kaixuan Huang;S. Kakade;Jason D. Lee;Qi Lei;Runzhe Wang;Jiaqi Yang

文献摘要

相似文献

由Q函数的神经网络近似支持的深度强化学习(RL)已经取得了巨大的经验成功。传统上,RL的理论主要集中在线性函数逼近(或更高维)方法上,而对于Q函数的神经网络逼近的非线性RL却知之甚少。这是本文工作的重点,我们研究了两层神经网络的函数逼近(同时考虑了RELU和多项式激活函数)。我们的第一个结果是在两层神经网络的完备性下的生成模型设置下的一个计算和统计上有效的算法。我们的第二个结果考虑了这种设置,但仅在神经网络函数类的可实现性下。这里,假设动力学是确定性的,样本复杂性在代数维度中线性扩展。在所有情况下,我们的结果都比线性(或更难以捉摸的维度)方法所能获得的结果有显著的改进。
Deep Reinforcement Learning (RL) powered by neural net approximation of the Q function has had enormous empirical success. While the theory of RL has traditionally focused on linear function approximation (or eluder dimension) approaches, little is known about nonlinear RL with neural net approximations of the Q functions. This is the focus of this work, where we study function approximation with two-layer neural networks (considering both ReLU and polynomial activation functions). Our first result is a computationally and statistically efficient algorithm in the generative model setting under completeness for two-layer neural networks. Our second result considers this setting but under only realizability of the neural net function class. Here, assuming deterministic dynamics, the sample complexity scales linearly in the algebraic dimension. In all cases, our results significantly improve upon what can be attained with linear (or eluder dimension) methods.