Performance Optimization in Mobile-Edge Computing via Deep Reinforcement Learning

Performance Optimization in Mobile-Edge Computing via Deep Reinforcement Learning
复制标题

DOI:
10.1109/vtcfall.2018.8690980
复制
发表时间:
2018-03
期刊:
2018 IEEE 88th Vehicular Technology Conference (VTC-Fall)
影响因子:
--
通讯作者:
Xianfu Chen;Honggang Zhang;Celimuge Wu;S. Mao;Yusheng Ji;M. Bennis
Xianfu Chen;Honggang Zhang;Celimuge Wu;S. Mao;Yusheng Ji;M. Bennis
中科院分区:
其他
文献类型:
--
作者:
Xianfu Chen;Honggang Zhang;Celimuge Wu;S. Mao;Yusheng Ji;M. Bennis

文献摘要

被引文献

相似文献

为了提高移动的设备的计算体验的质量,移动边缘计算(MEC)通过在紧密接近的无线电接入网络内提供计算能力而成为有前景的范例。然而,MEC系统的计算卸载策略的设计仍然具有挑战性。具体地,是在本地移动终端执行到达的计算任务还是卸载任务以用于云执行,应该以更智能的方式适应环境动态。在本文中,我们认为MEC的一个代表性的移动的用户在超密集的网络,其中多个基站(BS)之一,可以选择计算卸载。将求解最优计算卸载策略的问题建模为马尔可夫决策过程,以最小化长期开销为目标,根据移动的用户与基站之间的信道质量、能量队列状态以及任务队列状态进行卸载决策.为了打破状态空间高维的诅咒,我们提出了一种基于深度Q网络的策略计算卸载算法,在没有动态统计先验知识的情况下学习最优策略。数值实验结果表明,与基准策略相比,该算法在平均代价上有了显著的提高。
To improve the quality of computation experience for mobile devices, mobile-edge computing (MEC) is emerging as a promising paradigm by providing computing capabilities within radio access networks in close proximity. Nevertheless, the design of computation offloading policies for a MEC system remains challenging. Specifically, whether to execute an arriving computation task at local mobile device or to offload a task for cloud execution should adapt to the environmental dynamics in a smarter manner. In this paper, we consider MEC for a representative mobile user in an ultra dense network, where one of multiple base stations (BSs) can be selected for computation offloading. The problem of solving an optimal computation offloading policy is modelled as a Markov decision process, where our objective is to minimize the long-term cost and an offloading decision is made based on the channel qualities between the mobile user and the BSs, the energy queue state as well as the task queue state. To break the curse of high dimensionality in state space, we propose a deep $Q$-network-based strategic computation offloading algorithm to learn the optimal policy without having a priori knowledge of the dynamic statistics. Numerical experiments provided in this paper show that our proposed algorithm achieves a significant improvement in average cost compared with baseline policies.