How to Attack and Defend NextG Radio Access Network Slicing With Reinforcement Learning

How to Attack and Defend NextG Radio Access Network Slicing With Reinforcement Learning
复制标题

DOI:
10.1109/ojvt.2022.3229229
复制
发表时间:
2021-01
影响因子:
6.4
通讯作者:
Yi Shi;Y. Sagduyu;T. Erpek;M. C. Gursoy
Yi Shi;Y. Sagduyu;T. Erpek;M. C. Gursoy
中科院分区:
--
文献类型:
--
作者:
Yi Shi;Y. Sagduyu;T. Erpek;M. C. Gursoy

文献摘要

被引文献

相似文献

本文中,在下一代(NextG)无线接入网络中考虑了用于网络切片的强化学习(RL),其中基站(gNodeB)为用户设备的请求分配资源块(RB),并旨在随着时间的推移最大化所接受请求的总奖励。基于对抗性机器学习,引入了一种新颖的无线攻击来操纵 RL 算法并破坏 NextG 网络切片。对手观察频谱并构建自己的基于强化学习的代理模型,该模型根据能量预算选择要堵塞的 RB,其目标是最大化由于堵塞 RB 导致的失败请求数量。通过干扰 RB,攻击者减少了 RL 算法的奖励。由于该奖励被用作更新 RL 算法的输入,因此即使对手停止干扰,性能也无法恢复。这种攻击是根据恢复时间和(最大和总)奖励损失来评估的,并且它被证明比基准(随机和近视)干扰攻击更有效。引入了不同的反应式和主动式防御方案,例如一旦检测到攻击就暂停 RL 算法的更新,在 RL 的决策过程中引入随机性以误导对手的学习过程,或者操纵反馈(NACK)机制使对手无法获得可靠的信息,以表明在提高 RL 算法的奖励方面,防御 NextG 网络切片针对这种攻击是可行的。
In this paper, reinforcement learning (RL) for network slicing is considered in next generation (NextG) radio access networks, where the base station (gNodeB) allocates resource blocks (RBs) to the requests of user equipments and aims to maximize the total reward of accepted requests over time. Based on adversarial machine learning, a novel over-the-air attack is introduced to manipulate the RL algorithm and disrupt NextG network slicing. The adversary observes the spectrum and builds its own RL based surrogate model that selects which RBs to jam subject to an energy budget with the objective of maximizing the number of failed requests due to jammed RBs. By jamming the RBs, the adversary reduces the RL algorithm's reward. As this reward is used as the input to update the RL algorithm, the performance does not recover even after the adversary stops jamming. This attack is evaluated in terms of both the recovery time and the (maximum and total) reward loss, and it is shown to be much more effective than benchmark (random and myopic) jamming attacks. Different reactive and proactive defense schemes such as suspending the RL algorithm's update once an attack is detected, introducing randomness to the decision process in RL to mislead the learning process of the adversary, or manipulating the feedback (NACK) mechanism such that the adversary may not obtain reliable information are introduced to show that it is viable to defend NextG network slicing against this attack, in terms of improving the RL algorithm's reward.