Joint Offloading and Resource Allocation in Mobile Edge Computing Systems: An Actor-Critic Approach

Joint Offloading and Resource Allocation in Mobile Edge Computing Systems: An Actor-Critic Approach
复制标题

DOI:
10.1109/glocom.2018.8647593
复制
发表时间:
2018-12
期刊:
2018 IEEE Global Communications Conference (GLOBECOM)
影响因子:
--
通讯作者:
Zhicai Zhang;F. Richard Yu;Fang Fu;Qiao Yan;Zhouyang Wang
Zhicai Zhang;F. Richard Yu;Fang Fu;Qiao Yan;Zhouyang Wang
中科院分区:
其他
文献类型:
--
作者:
Zhicai Zhang;F. Richard Yu;Fang Fu;Qiao Yan;Zhouyang Wang

文献摘要

被引文献

相似文献

将计算密集型任务从用户设备(UE)卸载到移动边缘计算(MEC)服务器是一种很有前途的增强UE计算能力的技术。然而,MEC会带来额外的能量消耗和时延,这促使在移动网络中部署具有MEC的能量收集(EH)微蜂窝网络。由于这种网络的复杂性,有效地为UE分配资源是具有挑战性的。本文研究了具有MEC的能量采集(EH)小蜂窝网络中的卸载决策、无线和计算资源分配问题。与已有文献不同的是,我们的研究重点是通过最大限度地增加分流任务的数量来提高移动运营商的收入,同时降低能量消耗和时延。此外,在MEC服务器端创建队列,将未执行的任务存储在一个时隙中,作为效用函数的惩罚,以避免严重的延迟。考虑到队列长度的变化、小型基站(SBSS)和下行信道的EH电池的状态,将上述问题建模为马尔可夫决策过程(MDP)。针对MDP中的状态和动作是无限的,提出了一种具有资格跟踪的在线、按策略的行动者-批评者算法。仿真结果表明,与策略梯度算法和Q学习算法相比,该算法具有更好的性能。
Offloading computationally intensive tasks from user equipments (UEs) to mobile edge computing (MEC) servers is a promising technique to boost up the computational capacity of UEs. However, MEC will incur extra energy consumption and time delays, which motivates the deployment of energy harvesting (EH) small cell networks with MEC in mobile networks. Due to the complexity of such networks, it is challenging to effectively allocate resources for UEs. In this paper, we investigate the offloading decision, wireless and computational resources allocation problem in energy harvesting (EH) small cell networks with MEC. Different from existing literatures, our research focuses on improving mobile operators' revenue by maximizing the amount of the offloaded tasks while decreasing the energy expenditure and time-delays. Besides, queues are created at the MEC server side to store the un-executed tasks in a time slot, which is used as a punishment in our utility function to avoid serious delay. Considering the varying lengths of queues, the states of EH-batteries of small base stations (SBSs) and down-link channels, the above problem is modeled as a Markov decision process (MDP). Since the states and actions in the MDP are infinite, an online and on-policy actor-critic with eligibility traces algorithm is proposed to resolve the problem. Simulation results show the proposed algorithm has superior performances compared with the policy-gradient algorithm and Q- learning.