Adaptive and Robust Routing With Lyapunov-Based Deep RL in MEC Networks Enabled by Blockchains

Adaptive and Robust Routing With Lyapunov-Based Deep RL in MEC Networks Enabled by Blockchains
复制标题

DOI:
10.1109/jiot.2020.3034601
复制
发表时间:
2021-02
影响因子:
10.6
通讯作者:
Zirui Zhuang;Jingyu Wang;Q. Qi;J. Liao;Zhu Han
Zirui Zhuang;Jingyu Wang;Q. Qi;J. Liao;Zhu Han
中科院分区:
计算机科学1区
文献类型:
--
作者:
Zirui Zhuang;Jingyu Wang;Q. Qi;J. Liao;Zhu Han

文献摘要

被引文献

相似文献

物联网的最新发展带来了大量及时敏感和爆发的数据流。同样,对存储,计算和通信的联合优化需要多频道边缘计算框架。自适应网络控制已使用深度加固学习(RL)探索,但是它不足以解决爆发网络流量流,尤其是当网络流量模式随时间变化时。我们在环境中以时间变化的链接延迟为Lyapunov优化问题制定路由控制。我们确定当包括传播延迟时,优化性能和建模准确性之间存在权衡。我们提出了一种基于基于RL(DRL)的新型自适应网络路由方法来解决上述问题。 Lyapunov优化技术用于减少Lyapunov漂移的上限,从而改善了网络系统中的排队稳定性。通过使用Markovian到达过程对网络流量模式进行建模,我们表明可以将网络路由问题建模为Markov决策过程和基于价值的RL方法来解决它们。我们使用经过的时间共识机制证明基于区块链的协议,以确保可信赖的网络统计信息信息交换为路由框架。实验结果表明,所提出的方法可以学习路由策略并适应不断变化的环境。所提出的方法在多个设置中优于基线背压方法,并且收敛速度比现有方法更快。此外,DRL模块可以有效地学习对长期Lyapunov漂移和惩罚功能的更好估计,从而在积压大小,端到端延迟,信息时代和吞吐量方面提供了卓越的结果。此外,基于区块链的网络统计交换可以提供针对恶意节点的路由框架。此外,所提出的模型在各种拓扑下都可以很好地表现,因此可以在一般情况下使用。
The most recent development of the Internet of Things brings massive timely sensitive and bursty data flows. Also, joint optimization on storage, computation, and communication is in need for multiaccess edge computing frameworks. The adaptive network control has been explored using deep reinforcement learning (RL), but it is not sufficient for bursty network traffic flows, especially when the network traffic pattern may change over time. We formulate the routing control in an environment with time-variant link delays as a Lyapunov optimization problem. We identify that there is a tradeoff between optimization performance and modeling accuracy when the propagation delays are included. We propose a novel deep RL (DRL)-based adaptive network routing method to tackle the issues mentioned above. A Lyapunov optimization technique is used to reduce the upper bound of the Lyapunov drift, improving queuing stability in networked systems. By modeling the network traffic pattern using the Markovian arrival process, we show that network routing problems can be modeled as Markov decision processes and value-iteration-based RL methods can be used to solve them. We design a blockchain-based protocol using proof of elapsed time consensus mechanism to ensure a trustworthy network statistics information exchange for the routing framework. Experiment results show that the proposed method can learn a routing policy and adapt to the changing environment. The proposed method outperforms the baseline backpressure method in multiple settings and converges faster than existing methods. Moreover, the DRL module can effectively learn a better estimation of the long-term Lyapunov drift and penalty functions, providing superior results in terms of the backlog size, end-to-end latency, age of information, and throughput. Furthermore, the blockchain-based network statistics exchange can provide the routing framework against malicious nodes. In addition, the proposed model performs well under various topologies, and thus can be used in general cases.