Federated Deep Reinforcement Learning for the Distributed Control of NextG Wireless Networks

Federated Deep Reinforcement Learning for the Distributed Control of NextG Wireless Networks
复制标题

DOI:
10.1109/dyspan53946.2021.9677132
复制
发表时间:
2021-12
期刊:
2021 IEEE International Symposium on Dynamic Spectrum Access Networks (DySPAN)
影响因子:
--
通讯作者:
Peyman Tehrani;Francesco Restuccia;M. Levorato
Peyman Tehrani;Francesco Restuccia;M. Levorato
中科院分区:
其他
文献类型:
--
作者:
Peyman Tehrani;Francesco Restuccia;M. Levorato

文献摘要

被引文献

相似文献

下一代(NextG)网络预计将支持要求苛刻的触觉互联网应用,如增强现实和联网自动驾驶汽车。尽管最近的创新带来了更大链路容量的希望,但它们对环境的敏感性和不稳定的性能违背了传统的基于模型的控制原理。零接触数据驱动方法可以提高网络适应当前操作条件的能力。诸如强化学习(RL)算法之类的工具可以仅基于观察的历史来构建最优控制策略。具体而言,使用深度神经网络(DNN)作为预测器的深度RL(DRL)已被证明即使在复杂环境和高维输入下也能实现良好的性能。然而,DRL模型的训练需要大量的数据,这可能会限制其对底层环境不断变化的统计数据的适应性。此外,无线网络本质上是分布式系统,其中集中式DRL方法将需要过多的数据交换,而完全分布式方法可能导致较慢的收敛速率和性能下降。在本文中,为了解决这些挑战,我们提出了一种用于DRL的联邦学习(FL)方法,我们将其称为联邦DRL(F-DRL),其中基站(BS)通过仅共享模型的权重而不是训练数据来协作训练嵌入式DNN。我们评估了两个不同版本的F-DRL,价值和政策的基础上,并显示其实现的上级性能相比,分布式和集中式DRL。
Next Generation (NextG) networks are expected to support demanding tactile internet applications such as augmented reality and connected autonomous vehicles. Whereas recent innovations bring the promise of larger link capacity, their sensitivity to the environment and erratic performance defy traditional model-based control rationales. Zero-touch data-driven approaches can improve the ability of the network to adapt to the current operating conditions. Tools such as reinforcement learning (RL) algorithms can build optimal control policy solely based on a history of observations. Specifically, deep RL (DRL), which uses a deep neural network (DNN) as a predictor, has been shown to achieve good performance even in complex environments and with high dimensional inputs. However, the training of DRL models require a large amount of data, which may limit its adaptability to ever-evolving statistics of the underlying environment. Moreover, wireless networks are inherently distributed systems, where centralized DRL approaches would require excessive data exchange, while fully distributed approaches may result in slower convergence rates and performance degradation. In this paper, to address these challenges, we propose a federated learning (FL) approach to DRL, which we refer to federated DRL (F-DRL), where base stations (BS) collaboratively train the embedded DNN by only sharing models’ weights rather than training data. We evaluate two distinct versions of F-DRL, value and policy based, and show the superior performance they achieve compared to distributed and centralized DRL.