Head Network Distillation: Splitting Distilled Deep Neural Networks for Resource-Constrained Edge Computing Systems

Head Network Distillation: Splitting Distilled Deep Neural Networks for Resource-Constrained Edge Computing Systems
复制标题

DOI:
10.1109/access.2020.3039714
复制
发表时间:
2020-11
期刊:
影响因子:
3.9
通讯作者:
Yoshitomo Matsubara;Davide Callegaro;S. Baidya;M. Levorato;Sameer Singh
Yoshitomo Matsubara;Davide Callegaro;S. Baidya;M. Levorato;Sameer Singh
中科院分区:
计算机科学3区
文献类型:
--
作者:
Yoshitomo Matsubara;Davide Callegaro;S. Baidya;M. Levorato;Sameer Singh

文献摘要

被引文献

相似文献

随着深度神经网络(DNN)模型复杂性的增加,其在移动的设备上的部署变得越来越具有挑战性,特别是在图像分类等复杂视觉任务中。许多最近的贡献旨在产生与移动的设备的有限计算能力相匹配的紧凑模型,或者将这种繁重的模型的执行卸载到网络边缘处的具有计算能力的设备-边缘服务器。在本文中,我们建议修改DNN模型的结构和训练过程,以实现早期网络层的网络内压缩。我们的训练过程源于知识蒸馏,这是一种传统上用于构建小型学生模型以模仿大型教师模型输出的技术。在这里,我们采用这种想法来获得积极的压缩,同时保持准确性。我们的结果表明,我们的方法对于在复杂数据集上训练的最先进的模型是有效的,并且可以扩展参数区域,其中边缘计算是一个可行且有利的选择。此外,我们证明,在许多设置的实际利益,我们减少了推理时间,如MobileNet v2在移动终端执行的专门模型,同时提高准确性。
As the complexity of Deep Neural Network (DNN) models increases, their deployment on mobile devices becomes increasingly challenging, especially in complex vision tasks such as image classification. Many of recent contributions aim either to produce compact models matching the limited computing capabilities of mobile devices or to offload the execution of such burdensome models to a compute-capable device at the network edge – the edge servers. In this paper, we propose to modify the structure and training process of DNN models for complex image classification tasks to achieve in-network compression in the early network layers. Our training process stems from knowledge distillation, a technique that has been traditionally used to build small – student – models mimicking the output of larger – teacher – models. Here, we adopt this idea to obtain aggressive compression while preserving accuracy. Our results demonstrate that our approach is effective for state-of-the-art models trained over complex datasets, and can extend the parameter region in which edge computing is a viable and advantageous option. Additionally, we demonstrate that in many settings of practical interest we reduce the inference time with respect to specialized models such as MobileNet v2 executed at the mobile device, while improving accuracy.