Finite-Sample Analysis of Decentralized Temporal-Difference Learning with Linear Function Approximation

Finite-Sample Analysis of Decentralized Temporal-Difference Learning with Linear Function Approximation
复制标题

线性函数逼近的分散式时差学习的有限样本分析

DOI:
--
复制
发表时间:
2019-11
期刊:
arXiv
影响因子:
--
通讯作者:
Z Yang
Z Yang
中科院分区:
其他
文献类型:
--
作者:
J Sun;G Wang;GB Giannakis;Q Yang;Z Yang

文献摘要

参考文献

相似文献

受多智能体强化学习(MARL)在网络机器人、蜂群无人机和传感器网络等工程应用中的新兴应用的推动,我们在完全去中心化的环境中研究策略评估问题,使用时差(TD)学习和线性函数逼近来处理实践中的大状态空间。一组代理的目标是通过与邻居交换本地估计,从共享环境中观察到的本地私人奖励中协作学习给定策略的价值函数。尽管它们简单且使用广泛,但我们对这种去中心化 TD 学习算法的理论理解仍然有限。现有结果是基于 i.i.d 获得的。数据样本,或通过施加“额外”投影步骤来控制马尔可夫观察引起的“梯度”偏差。在本文中,我们提供了独立同分布下完全去中心化 TD(0) 学习的有限样本分析。以及马尔可夫样本,并证明所有局部估计线性收敛到最优值的一个小邻域。由此产生的误差界限是同类中的第一个——从某种意义上说,它们在最实际的假设下成立——这是通过新颖的多步骤李亚普诺夫分析实现的。
Motivated by the emerging use of multi-agent reinforcement learning (MARL) in engineering applications such as networked robotics, swarming drones, and sensor networks, we investigate the policy evaluation problem in a fully decentralized setting, using temporal-difference (TD) learning with linear function approximation to handle large state spaces in practice. The goal of a group of agents is to collaboratively learn the value function of a given policy from locally private rewards observed in a shared environment, through exchanging local estimates with neighbors. Despite their simplicity and widespread use, our theoretical understanding of such decentralized TD learning algorithms remains limited. Existing results were obtained based on i.i.d. data samples, or by imposing an `additional' projection step to control the `gradient' bias incurred by the Markovian observations. In this paper, we provide a finite-sample analysis of the fully decentralized TD(0) learning under both i.i.d. as well as Markovian samples, and prove that all local estimates converge linearly to a small neighborhood of the optimum. The resultant error bounds are the first of its type---in the sense that they hold under the most practical assumptions ---which is made possible by means of a novel multi-step Lyapunov analysis.
DOI: 10.1007/978-93-86279-38-5
发表时间: 2008-09
期刊: ZAMM - Journal of Applied Mathematics and Mechanics / Zeitschrift für Angewandte Mathematik und Mechanik
影响因子: --
作者:
V. Borkar
通讯作者: V. Borkar
DOI: 10.1007/978-3-319-51204-4_7
发表时间: 2016
期刊: --
影响因子: --
作者:
E. Yanmaz;M. Quaritsch;S. Yahyanejad;B. Rinner;H. Hellwagner;C. Bettstetter
通讯作者: E. Yanmaz;M. Quaritsch;S. Yahyanejad;B. Rinner;H. Hellwagner;C. Bettstetter
DOI: 10.1109/tnn.1998.712192
发表时间: 1998
期刊: IEEE Trans. Neural Networks
影响因子: --
作者:
R. S. Sutton;A. Barto
通讯作者: R. S. Sutton;A. Barto
DOI: --
发表时间: 2018-06
期刊: ArXiv
影响因子: --
作者:
Hoi-To Wai;Zhuoran Yang;Zhaoran Wang;Mingyi Hong
通讯作者: Hoi-To Wai;Zhuoran Yang;Zhaoran Wang;Mingyi Hong
DOI: 10.1109/tsg.2019.2951769
发表时间: 2019-04
影响因子: 9.6
作者:
Qiuling Yang;Gang Wang;A. Sadeghi;G. Giannakis;Jian Sun-
通讯作者: Qiuling Yang;Gang Wang;A. Sadeghi;G. Giannakis;Jian Sun-