Target-Based Temporal Difference Learning

Target-Based Temporal Difference Learning
复制标题

DOI:
--
复制
发表时间:
2019-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Donghwan Lee;Niao He
Donghwan Lee;Niao He
中科院分区:
其他
文献类型:
--
作者:
Donghwan Lee;Niao He

文献摘要

相似文献

目标网络的使用一直是最近用于强化学习的深度Q学习算法的流行和关键组成部分,但从理论方面知之甚少。在这项工作中,我们介绍了一个新的家庭的目标为基础的时间差(TD)学习算法,并提供理论分析其收敛性。与标准TD学习不同,基于目标的TD算法保持两个独立的学习参数-目标变量和在线变量。特别是,我们引入了该家族中的三个成员,称为平均TD,双TD和周期TD,其中目标变量通过平均,对称或周期方式更新,反映了深度Q学习实践中使用的技术。我们建立了平均TD和双TD的渐近收敛性分析和周期TD的有限样本分析。此外,我们还提供了一些模拟结果显示潜在的上级收敛这些基于目标的TD算法相比,标准的TD学习。虽然这项工作的重点是线性函数近似和策略评估设置,但我们认为这是对目标网络的深度Q学习变体的理论理解的有意义的一步。
The use of target networks has been a popular and key component of recent deep Q-learning algorithms for reinforcement learning, yet little is known from the theory side. In this work, we introduce a new family of target-based temporal difference (TD) learning algorithms and provide theoretical analysis on their convergences. In contrast to the standard TD-learning, target-based TD algorithms maintain two separate learning parameters-the target variable and online variable. Particularly, we introduce three members in the family, called the averaging TD, double TD, and periodic TD, where the target variable is updated through an averaging, symmetric, or periodic fashion, mirroring those techniques used in deep Q-learning practice. We establish asymptotic convergence analyses for both averaging TD and double TD and a finite sample analysis for periodic TD. In addition, we also provide some simulation results showing potentially superior convergence of these target-based TD algorithms compared to the standard TD-learning. While this work focuses on linear function approximation and policy evaluation setting, we consider this as a meaningful step towards the theoretical understanding of deep Q-learning variants with target networks.