Mobile Edge Computing Against Smart Attacks with Deep Reinforcement Learning in Cognitive MIMO IoT Systems

Mobile Edge Computing Against Smart Attacks with Deep Reinforcement Learning in Cognitive MIMO IoT Systems
复制标题

移动边缘计算通过认知 MIMO 物联网系统中的深度强化学习抵御智能攻击

DOI:
10.1007/s11036-020-01572-w
复制
发表时间:
2020-07-02
影响因子:
3.8
通讯作者:
Liu, Yun
Liu, Yun
中科院分区:
计算机科学4区
文献类型:
--
作者:
Ge, Songyang;Lu, Beiling;Liu, Yun

文献摘要

被引文献

相似文献

在无线物联网(IoT)系统中,多输入多输出(MIMO)和认知无线电(CR)技术通常被引入到移动的边缘计算(MEC)结构中,以提高频谱效率和传输可靠性。然而,这种基于CR的MIMO物联网系统将遭受来自无线环境的各种智能攻击,即使是物联网系统中的MEC服务器也不够安全,容易受到这些攻击。在本文中,我们研究了认知MIMO物联网系统中的安全通信问题,该系统包括主用户(PU),次用户(SU),智能攻击者和多个MEC服务器。我们的系统设计的目标是优化SU的效用,包括其效率和安全性。SU将选择在CR场景中未被PU占用的空闲MEC服务器,并且通过以适当的发射功率卸载这些任务来向服务器分配其计算任务的适当卸载速率。在这样的CR IoT系统中,攻击者将选择一种类型的智能攻击。在此基础上,提出了两种基于深度强化学习的资源分配策略,分别是基于Dyna架构和优先级扫描的边缘服务器选择(DPESS)策略和基于Deep Q网络的边缘服务器选择(DESS)策略,以寻求在不需要信道状态信息(CSI)的情况下实现效用最大化的最优策略。具体地,由于通过利用经验重放技术和随机梯度下降(SGD)训练的卷积神经网络(CNN),DESS方案的收敛速度显著提高。此外,理论上推导了所提出的两个方案的纳什均衡和存在条件的MEC对策模型的智能攻击。与传统的Q学习算法相比,所提出的DPESS和DESS方案可以提高SU的平均效用和保密容量。数值模拟也被提出来验证我们的建议在效率和安全性方面的更好的性能,包括更高的收敛速度的DESS策略。
In wireless Internet of Things (IoT) systems, the multi-input multi-output (MIMO) and cognitive radio (CR) techniques are usually involved into the mobile edge computing (MEC) structure to improve the spectrum efficiency and transmission reliability. However, such a CR based MIMO IoT system will suffer from a variety of smart attacks from wireless environments, even the MEC servers in IoT systems are not secure enough and vulnerable to these attacks. In this paper, we investigate a secure communication problem in a cognitive MIMO IoT system comprising of a primary user (PU), a secondary user (SU), a smart attacker and several MEC servers. The target of our system design is to optimize utility of the SU, including its efficiency and security. The SU will choose an idle MEC server that is not occupied by the PU in the CR scenario, and allocates a proper offloading rate of its computation tasks to the server, by unloading such tasks with proper transmit power. In such a CR IoT system, the attacker will select one type of smart attacks. Then two deep reinforcement learning based resource allocation strategies are proposed to find an optimal policy of maximal utility without channel state information(CSI), one of which is the Dyna architecture and Prioritized sweeping based Edge Server Selection (DPESS) strategy, and the other is the Deep Q-network based Edge Server Selection (DESS) strategy. Specifically, the convergence speed of the DESS scheme is significantly improved due to the trained convolutional neural network (CNN) by utilizing the experience replay technique and stochastic gradient descent (SGD). In addition, the Nash equilibrium and existence conditions of the proposed two schemes are theoretically deduced for the modeled MEC game against smart attacks. Compared with the traditional Q-learning algorithm, the average utility and secrecy capacity of the SU can be improved by the proposed DPESS and DESS schemes. Numerical simulations are also presented to verify the better performance of our proposals in terms of efficiency and security, including the higher convergence speed of the DESS strategy.