Federated Ensemble Model-Based Reinforcement Learning in Edge Computing

Federated Ensemble Model-Based Reinforcement Learning in Edge Computing
复制标题

边缘计算中基于联邦集成模型的强化学习

DOI:
10.1109/tpds.2023.3264480
复制
发表时间:
2021-09
影响因子:
5.3
通讯作者:
Jin Wang;Jia Hu;Jed Mills;G. Min;Ming Xia;N. Georgalas
Jin Wang;Jia Hu;Jed Mills;G. Min;Ming Xia;N. Georgalas
中科院分区:
计算机科学2区
文献类型:
--
作者:
Jin Wang;Jia Hu;Jed Mills;G. Min;Ming Xia;N. Georgalas

文献摘要

相似文献

联邦学习(FL)是一种保护隐私的分布式机器学习范式,它可以在地理上分布和异构的设备之间进行协作训练,而无需收集它们的数据。联邦强化学习(FRL)是将FL扩展到监督学习模型之外的一种学习方法,用于处理边缘计算系统中的顺序决策问题。然而,现有的FRL算法直接将无模型RL与FL联合收割机结合,往往导致样本复杂度高,缺乏理论保证。为了应对这些挑战,我们提出了一种新的FRL算法,有效地结合了基于模型的RL和集成知识蒸馏到FL的第一次。具体来说,我们利用FL和知识蒸馏为客户创建一个动态模型集合,然后仅使用集合模型来训练策略,而不与环境进行交互。此外,我们从理论上证明了所提出的算法的单调改进是有保证的。大量的实验结果表明,我们的算法获得了更高的采样效率相比,经典的无模型FRL算法在具有挑战性的连续控制基准环境下的边缘计算设置。结果还突出了异构客户端数据和本地模型更新步骤对FRL性能的显著影响,验证了从我们的理论分析中获得的见解。
Federated learning (FL) is a privacy-preserving distributed machine learning paradigm that enables collaborative training among geographically distributed and heterogeneous devices without gathering their data. Extending FL beyond the supervised learning models, federated reinforcement learning (FRL) was proposed to handle sequential decision-making problems in edge computing systems. However, the existing FRL algorithms directly combine model-free RL with FL, thus often leading to high sample complexity and lacking theoretical guarantees. To address the challenges, we propose a novel FRL algorithm that effectively incorporates model-based RL and ensemble knowledge distillation into FL for the first time. Specifically, we utilise FL and knowledge distillation to create an ensemble of dynamics models for clients, and then train the policy by solely using the ensemble model without interacting with the environment. Furthermore, we theoretically prove that the monotonic improvement of the proposed algorithm is guaranteed. The extensive experimental results demonstrate that our algorithm obtains much higher sample efficiency compared to classic model-free FRL algorithms in the challenging continuous control benchmark environments under edge computing settings. The results also highlight the significant impact of heterogeneous client data and local model update steps on the performance of FRL, validating the insights obtained from our theoretical analysis.