iRAF: A Deep Reinforcement Learning Approach for Collaborative Mobile Edge Computing IoT Networks

iRAF: A Deep Reinforcement Learning Approach for Collaborative Mobile Edge Computing IoT Networks
复制标题

iRAF:用于协作移动边缘计算物联网网络的深度强化学习方法

DOI:
10.1109/jiot.2019.2913162
复制
发表时间:
2019-08-01
影响因子:
10.6
通讯作者:
Hu, Jianhao
Hu, Jianhao
中科院分区:
计算机科学1区
文献类型:
--
作者:
Chen, Jienan;Chen, Siyu;Hu, Jianhao

文献摘要

被引文献

相似文献

近年来,随着人工智能(AI)的发展,数据驱动的人工智能方法在解决复杂问题方面表现出了惊人的性能,以支持物联网(IoT)世界的海量资源消耗和延迟敏感的服务。在本文中,我们提出了一种智能资源分配框架(iRAF)来解决协作移动边缘计算(CoMEC)网络的复杂资源分配问题。 iRAF的核心是一种多任务深度强化学习算法,用于根据网络状态和任务特征(例如边缘服务器和设备的计算能力、通信信道质量、资源利用率和服务的延迟要求等)做出资源分配决策。所提出的iRAF可以自动学习网络环境并生成资源分配决策,以通过自播放训练最大化延迟和功耗方面的性能。 iRAF 成为自己的老师:训练深度神经网络(DNN)以自监督学习的方式预测 iRAF 的资源分配动作,其中训练数据是从蒙特卡罗树搜索(MCTS)算法的搜索过程中生成的。 MCTS 的一个主要优点是,它将模拟未来的轨迹,从根状态开始,通过评估奖励值来获得最佳动作。数值结果表明,与贪婪搜索和基于深度 Q 学习的方法相比,我们提出的 iRAF 在服务延迟性能方面分别实现了 59.27% 和 51.71% 的改进。
Recently, as the development of artificial intelligence (AI), data-driven AI methods have shown amazing performance in solving complex problems to support the Internet of Things (IoT) world with massive resource-consuming and delay-sensitive services. In this paper, we propose an intelligent resource allocation framework (iRAF) to solve the complex resource allocation problem for the collaborative mobile edge computing (CoMEC) network. The core of iRAF is a multitask deep reinforcement learning algorithm for making resource allocation decisions based on network states and task characteristics, such as the computing capability of edge servers and devices, communication channel quality, resource utilization, and latency requirement of the services, etc. The proposed iRAF can automatically learn the network environment and generate resource allocation decision to maximize the performance over latency and power consumption with self-play training. iRAF becomes its own teacher: a deep neural network (DNN) is trained to predict iRAF's resource allocation action in a self-supervised learning manner, where the training data is generated from the searching process of Monte Carlo tree search (MCTS) algorithm. A major advantage of MCTS is that it will simulate trajectories into the future, starting from a root state, to obtain a best action by evaluating the reward value. Numerical results show that our proposed iRAF achieves 59.27% and 51.71% improvement on service latency performance compared with the greedy-search and the deep Q-learning-based methods, respectively.