Reliability-Aware Online Scheduling for DNN Inference Tasks in Mobile-Edge Computing

Reliability-Aware Online Scheduling for DNN Inference Tasks in Mobile-Edge Computing
复制标题

DOI:
10.1109/jiot.2023.3243266
复制
发表时间:
2023-07
影响因子:
10.6
通讯作者:
Huirong Ma;Rui Li;Xiaoxi Zhang;Zhi Zhou;Xu Chen
Huirong Ma;Rui Li;Xiaoxi Zhang;Zhi Zhou;Xu Chen
中科院分区:
计算机科学1区
文献类型:
--
作者:
Huirong Ma;Rui Li;Xiaoxi Zhang;Zhi Zhou;Xu Chen

文献摘要

相似文献

移动边缘计算 (MEC) 被广泛认为是一种很有前途的技术,通过利用边缘服务器 (ES) 执行邻近的深度神经网络 (DNN) 推理任务,为资源有限的物联网 (IoT) 设备提供人工智能 (AI) 功能。然而,在未知的系统动态(例如ES的可用性不确定)的情况下,在网络边缘调度DNN推理任务可能会失败,从而难以保证物联网设备的可靠服务。为了克服这一挑战,我们提出了一种 MEC 中 DNN 推理任务的可靠性感知在线调度方案,通过利用在线反馈和离线数据来学习 ES 的不确定可用性,以最大限度地提高 DNN 推理任务的推理精度和服务可靠性(即系统范围内处理的 DNN 推理任务的数量)。我们首先将可靠性感知 DNN 推理任务调度问题表述为一种新颖的约束组合多臂老虎机 (CMAB) 问题。然后,通过将Lyapunov优化技术、强盗学习、近似子模最大化和历史数据有机地结合起来,我们设计了一种带有强盗学习(RTBL)算法的可靠性感知任务调度方案来解决这个问题。不幸的是,即使准确预测了系统的不确定性,任务调度问题仍然是 NP 难问题。因此,为了解决这个问题,我们设计了一种基于调度问题子模性的高级近似算法,该算法获得了接近最优的解决方案并提供了令人满意的性能保证。最后,我们进行严格的理论分析和比赛驱动的模拟来展示RTBL的出色表现。
Mobile-edge computing (MEC) is widely envisioned as a promising technique for provisioning artificial intelligence (AI) capability for resource-limited Internet of Things (IoT) devices by leveraging edge servers (ESs) for executing deep neural network (DNN) inference tasks in proximity. However, scheduling DNN inference tasks at the network edge under unknown system dynamics (e.g., uncertain availability of ESs) may suffer from failures, making it difficult to guarantee reliable services for the IoT device. To overcome this challenge, we propose a reliability-aware online scheduling scheme for DNN inference tasks in MEC by leveraging both online feedback and offline data to learn the uncertain availability of ESs to maximize both the inference accuracy and service reliability of DNN inference tasks (i.e., the number of DNN inference tasks processed during the system span). We first formulate the reliability-aware DNN inference tasks scheduling problem as a novel constrained combinatorial multiarmed bandit (CMAB) problem. Then by integrating the Lyapunov optimization technique, bandit learning, approximated submodular maximization, and historical data organically, we design a reliability-aware task scheduling scheme with a bandit learning (RTBL) algorithm to solve this problem. Unfortunately, even with an accurate prediction of the system uncertainties, the task scheduling problem is still NP-hard. To deal with it, we, therefore, design an advanced approximation algorithm based on the submodularity of the scheduling problem which obtains a near-optimal solution and provides a satisfactory performance guarantee. Finally, we conduct rigorous theoretical analysis and race-driven simulations to show RTBL’s brilliant performance.