Non-Asymptotic Analysis for Two Time-scale TDC with General Smooth Function Approximation

Non-Asymptotic Analysis for Two Time-scale TDC with General Smooth Function Approximation
复制标题

DOI:
--
复制
发表时间:
2021-04
期刊:
--
影响因子:
--
通讯作者:
Yue Wang;Shaofeng Zou;Yi Zhou
Yue Wang;Shaofeng Zou;Yi Zhou
中科院分区:
其他
文献类型:
--
作者:
Yue Wang;Shaofeng Zou;Yi Zhou

文献摘要

被引文献

相似文献

带梯度校正的时差学习(TDC)是一种用于强化学习中策略评估的双时间尺度算法。该算法最初是在线性函数逼近下提出的,后来被推广到一般光滑函数逼近下。在[bhatnagar2009收敛]中建立了一般光滑函数逼近下的按政策设置的渐近收敛,然而,由于非线性和双时间尺度更新结构、非凸目标函数和时变投影到切平面上的挑战,有限样本分析仍然没有得到解决。在这篇文章中,我们发展了新的技巧来显式刻画具有I.I.D.或马尔可夫样本的一般非策略设置的有限样本误差界,并证明了它的收敛速度为$\数学O(1/\SQRTT)$(高达$\数学O(\logT)$)。我们的方法可以广泛地应用于具有一般光滑函数逼近的基于值的强化学习算法。
Temporal-difference learning with gradient correction (TDC) is a two time-scale algorithm for policy evaluation in reinforcement learning. This algorithm was initially proposed with linear function approximation, and was later extended to the one with general smooth function approximation. The asymptotic convergence for the on-policy setting with general smooth function approximation was established in [bhatnagar2009convergent], however, the finite-sample analysis remains unsolved due to challenges in the non-linear and two-time-scale update structure, non-convex objective function and the time-varying projection onto a tangent plane. In this paper, we develop novel techniques to explicitly characterize the finite-sample error bound for the general off-policy setting with i.i.d.\ or Markovian samples, and show that it converges as fast as $\mathcal O(1/\sqrt T)$ (up to a factor of $\mathcal O(\log T)$). Our approach can be applied to a wide range of value-based reinforcement learning algorithms with general smooth function approximation.