Federated Temporal Difference Learning with Linear Function Approximation under Environmental Heterogeneity

Federated Temporal Difference Learning with Linear Function Approximation under Environmental Heterogeneity
复制标题

DOI:
10.48550/arxiv.2302.02212
复制
发表时间:
2023-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Han Wang;A. Mitra;Hamed Hassani;George Pappas;James Anderson
Han Wang;A. Mitra;Hamed Hassani;George Pappas;James Anderson
中科院分区:
其他
文献类型:
--
作者:
Han Wang;A. Mitra;Hamed Hassani;George Pappas;James Anderson

文献摘要

相似文献

我们通过考虑政策评估问题来启动对环境异质性的联合加强学习的研究。我们的设置涉及$ n $代理与共享相同状态和行动空间但其奖励功能和状态过渡内核不同的环境交互的代理。假设代理可以通过中央服务器进行通信,我们问:交换信息是否会加快评估共同政策的过程?为了回答这个问题,我们提供了具有线性函数近似的联合时间差异(TD)学习算法的首次全面有限时间分析,同时考虑了Markovian采样,代理环境中的异质性以及多个本地更新以节省通信。我们的分析至关重要地依赖于几种新颖的成分:(i)在TD固定点上得出扰动范围,这是代理基础马尔可夫决策过程(MDP)中异质性的函数; (ii)引入虚拟MDP,以密切近似联合TD算法的动力学; (iii)使用虚拟MDP与联合优化建立明确的连接。将这些碎片整合在一起,我们严格地证明在低杂种性方面,交换模型估计会导致代理数量的线性收敛加速。
We initiate the study of federated reinforcement learning under environmental heterogeneity by considering a policy evaluation problem. Our setup involves $N$ agents interacting with environments that share the same state and action space but differ in their reward functions and state transition kernels. Assuming agents can communicate via a central server, we ask: Does exchanging information expedite the process of evaluating a common policy? To answer this question, we provide the first comprehensive finite-time analysis of a federated temporal difference (TD) learning algorithm with linear function approximation, while accounting for Markovian sampling, heterogeneity in the agents' environments, and multiple local updates to save communication. Our analysis crucially relies on several novel ingredients: (i) deriving perturbation bounds on TD fixed points as a function of the heterogeneity in the agents' underlying Markov decision processes (MDPs); (ii) introducing a virtual MDP to closely approximate the dynamics of the federated TD algorithm; and (iii) using the virtual MDP to make explicit connections to federated optimization. Putting these pieces together, we rigorously prove that in a low-heterogeneity regime, exchanging model estimates leads to linear convergence speedups in the number of agents.