Distributed Policy Gradient with Heterogeneous Computations for Federated Reinforcement Learning

Distributed Policy Gradient with Heterogeneous Computations for Federated Reinforcement Learning
复制标题

DOI:
10.1109/ciss56502.2023.10089771
复制
发表时间:
2023-03
期刊:
2023 57th Annual Conference on Information Sciences and Systems (CISS)
影响因子:
--
通讯作者:
Ye Zhu;Xiaowen Gong
Ye Zhu;Xiaowen Gong
中科院分区:
其他
文献类型:
--
作者:
Ye Zhu;Xiaowen Gong

文献摘要

相似文献

联邦学习(FL)在过去几年中的快速发展最近激发了联邦强化学习(FRL),其中多个强化学习(RL)代理协作学习共同的决策策略,而无需与环境交换原始交互数据。在本文中,我们考虑了一个一般的FRL框架,其中代理与不同的环境进行交互,具有相同的状态和动作空间,但不同的奖励和动态。受代理通常具有异构计算能力这一事实的启发,我们提出了一种用于FRL的联邦异构策略梯度(FedHPG)算法,其中代理可以使用不同数量的数据轨迹(即,批量大小)以及用于它们各自的PG算法的不同数量的局部计算迭代。我们的特征性能界的学习精度FedHPG,这表明它实现了学习精度的样本复杂度为O$(1/2),这与现有的RL算法的性能相匹配。结果还显示了局部迭代次数和迭代批量大小对学习精度的影响。在此基础上,将FedHPG算法扩展为异构策略梯度方差约简(FedHPGVR)算法,并分析了该算法的收敛性。实验结果验证了理论结果的基准RL任务。
The rapid advances in federated learning (FL) in the past few years have recently inspired federated reinforcement learning (FRL), where multiple reinforcement learning (RL) agents collaboratively learn a common decision-making policy without exchanging their raw interaction data with their environments. In this paper, we consider a general FRL framework where agents interact with different environments with identical state and action spaces but different rewards and dynamics. Motivated by the fact that agents often have heterogeneous computation capabilities, we propose a Federated Heterogeneous Policy Gradient (FedHPG) algorithm for FRL, where agents can use different numbers of data trajectories (i.e., batch sizes) and different numbers of local computation iterations for their respective PG algorithms. We characterize performance bounds for the learning accuracy of FedHPG, which shows that it achieves a learning accuracy ∊ with sample complexity of $O$ (1/∊2), which matches the performance of existing RL algorithms. The results also show the impacts of local iteration numbers and batch sizes for iteration on the learning accuracy. We also extend FedHPG to heterogeneous policy gradient variance reduction (FedHPGVR) algorithm based on the variance reduction method, and analyze the convergence of this algorithm. The theoretical results are verified empirically for benchmark RL tasks.