Learning-Based Offloading of Tasks with Diverse Delay Sensitivities for Mobile Edge Computing

Learning-Based Offloading of Tasks with Diverse Delay Sensitivities for Mobile Edge Computing
复制标题

DOI:
10.1109/globecom38437.2019.9013498
复制
发表时间:
2019-12
期刊:
2019 IEEE Global Communications Conference (GLOBECOM)
影响因子:
--
通讯作者:
Tianyu Zhang;Yi-Han Chiang;C. Borcea;Yusheng Ji
Tianyu Zhang;Yi-Han Chiang;C. Borcea;Yusheng Ji
中科院分区:
其他
文献类型:
--
作者:
Tianyu Zhang;Yi-Han Chiang;C. Borcea;Yusheng Ji

文献摘要

相似文献

不断发展的移动的应用需要越来越多的计算资源来平滑用户体验,有时还需要满足延迟要求。因此,由于计算能力和电池寿命的限制,移动的设备(MD)逐渐难以及时完成所有任务。为了科普这个问题,人们创建了移动的边缘计算(MEC)系统,以帮助附近边缘服务器上的MD处理任务。现有的工作致力于解决MEC任务卸载问题,包括具有简单延迟约束的任务卸载问题,但是大多数工作忽略了最后期限约束和延迟敏感任务的共存(即,任务的不同延迟敏感性)。在本文中,我们提出了一种基于行动者-批评者的深度强化学习(ADRL)模型,该模型考虑了不同的延迟敏感性,并自适应地卸载任务,以最小化由截止日期受限任务的截止日期错过和延迟敏感任务的延迟引起的总惩罚。我们使用由任务的不同延迟敏感性组成的真实的数据集来训练ADRL模型。我们的仿真结果表明,所提出的解决方案优于几个启发式算法的总惩罚,它也保持其性能增益在不同的系统设置。
The ever-evolving mobile applications need more and more computing resources to smooth user experience and sometimes meet delay requirements. Therefore, mobile devices (MDs) are gradually having difficulties to complete all tasks in time due to the limitations of computing power and battery life. To cope with this problem, mobile edge computing (MEC) systems were created to help with task processing for MDs at nearby edge servers. Existing works have been devoted to solving MEC task offloading problems, including those with simple delay constraints, but most of them neglect the coexistence of deadline-constrained and delay- sensitive tasks (i.e., the diverse delay sensitivities of tasks). In this paper, we propose an actor-critic based deep reinforcement learning (ADRL) model that takes the diverse delay sensitivities into account and offloads tasks adaptively to minimize the total penalty caused by deadline misses of deadline-constrained tasks and the lateness of delay-sensitive tasks. We train the ADRL model using a real data set that consists of the diverse delay sensitivities of tasks. Our simulation results show that the proposed solution outperforms several heuristic algorithms in terms of total penalty, and it also retains its performance gains under different system settings.