JointDNN: An Efficient Training and Inference Engine for Intelligent Mobile Cloud Computing Services

JointDNN: An Efficient Training and Inference Engine for Intelligent Mobile Cloud Computing Services
复制标题

DOI:
10.1109/tmc.2019.2947893
复制
发表时间:
2021-02-01
影响因子:
7.9
通讯作者:
Pedram, Massoud
Pedram, Massoud
中科院分区:
计算机科学2区
文献类型:
--
作者:
Eshratifar, Amir Erfan;Abrishami, Mohammad Saeed;Pedram, Massoud

文献摘要

被引文献

相似文献

深度学习模型正在许多移动的智能应用程序中部署。智能个人助理、自动汽车和智能家居服务等端侧服务通常采用移动的上的简单本地模型或云上的复杂远程模型。然而,最近的研究表明,在移动的和云之间划分DNN计算可以增加延迟和能源效率。在本文中,我们提出了一个有效的,自适应的,实用的引擎,JointDNN,用于移动终端和云之间的协同计算DNN在推理和训练阶段。JointDNN不仅为移动的端提供了一种高效节能的DNN查询方法,而且与仅使用云的方法相比,通过减少其工作负载和通信量,使云服务器受益。给定DNN架构,我们研究了在移动终端上处理某些层和在云服务器上处理某些层的效率。我们在层粒度上为DNN中的前向和后向传播提供了优化配方,它可以适应移动的电池限制和云服务器负载约束以及服务质量。与现状方法相比,JointDNN在查询DNN的延迟和移动的能耗上分别减少了18倍和32倍。
Deep learning models are being deployed in many mobile intelligent applications. End-side services, such as intelligent personal assistants, autonomous cars, and smart home services often employ either simple local models on the mobile or complex remote models on the cloud. However, recent studies have shown that partitioning the DNN computations between the mobile and cloud can increase the latency and energy efficiencies. In this paper, we propose an efficient, adaptive, and practical engine, JointDNN, for collaborative computation between a mobile device and cloud for DNNs in both inference and training phase. JointDNN not only provides an energy and performance efficient method of querying DNNs for the mobile side but also benefits the cloud server by reducing the amount of its workload and communications compared to the cloud-only approach. Given the DNN architecture, we investigate the efficiency of processing some layers on the mobile device and some layers on the cloud server. We provide optimization formulations at layer granularity for forward- and backward-propagations in DNNs, which can adapt to mobile battery limitations and cloud server load constraints and quality of service. JointDNN achieves up to 18 and 32 times reductions on the latency and mobile energy consumption of querying DNNs compared to the status-quo approaches, respectively.