Autodidactic Neurosurgeon: Collaborative Deep Inference for Mobile Edge Intelligence via Online Learning

Autodidactic Neurosurgeon: Collaborative Deep Inference for Mobile Edge Intelligence via Online Learning
复制标题

DOI:
10.1145/3442381.3450051
复制
发表时间:
2021-02
期刊:
Proceedings of the Web Conference 2021
影响因子:
--
通讯作者:
Letian Zhang;Lixing Chen;Jie Xu
Letian Zhang;Lixing Chen;Jie Xu
中科院分区:
其他
文献类型:
--
作者:
Letian Zhang;Lixing Chen;Jie Xu

文献摘要

被引文献

相似文献

深度学习(DL)的最新突破导致了许多智能移动的应用和服务的出现,但同时也对资源受限的移动的设备提出了前所未有的计算挑战。本文在资源受限的移动终端和功能强大的边缘服务器之间构建了一个协同深度推理系统,旨在加入设备上处理和计算卸载的能力。该系统的基本思想是将深度神经网络(DNN)划分为运行在移动终端上的前端部分和运行在边缘服务器上的后端部分,关键挑战是如何定位最佳划分点以最小化端到端的推理延迟。与现有的DNN分区工作不同,DNN分区严重依赖于专用的离线分析阶段来搜索最佳分区点,我们的系统具有内置的在线学习模块,称为自动教学神经外科医生(ANS),可以自动学习最佳分区点。因此,ANS是能够密切关注系统环境的变化,产生新的知识,自适应决策。ANS的核心是一种新的上下文Bandit学习算法μLinUCB,它不仅具有可证明的理论学习性能保证,而且是超轻量的,易于实现。我们实现了我们的系统上的视频流对象检测测试平台,以验证ANS的设计和评估其性能。实验表明,ANS在跟踪系统变化和减少端到端的推理延迟方面显着优于最先进的基准测试。
Recent breakthroughs in deep learning (DL) have led to the emergence of many intelligent mobile applications and services, but in the meanwhile also pose unprecedented computing challenges on resource-constrained mobile devices. This paper builds a collaborative deep inference system between a resource-constrained mobile device and a powerful edge server, aiming at joining the power of both on-device processing and computation offloading. The basic idea of this system is to partition a deep neural network (DNN) into a front-end part running on the mobile device and a back-end part running on the edge server, with the key challenge being how to locate the optimal partition point to minimize the end-to-end inference delay. Unlike existing efforts on DNN partitioning that rely heavily on a dedicated offline profiling stage to search for the optimal partition point, our system has a built-in online learning module, called Autodidactic Neurosurgeon (ANS), to automatically learn the optimal partition point on-the-fly. Therefore, ANS is able to closely follow the changes of the system environment by generating new knowledge for adaptive decision making. The core of ANS is a novel contextual bandit learning algorithm, called μLinUCB, which not only has provable theoretical learning performance guarantee but also is ultra-lightweight for easy real-world implementation. We implement our system on a video stream object detection testbed to validate the design of ANS and evaluate its performance. The experiments show that ANS significantly outperforms state-of-the-art benchmarks in terms of tracking system changes and reducing the end-to-end inference delay.