Neuron Manifold Distillation for Edge Deep Learning

Neuron Manifold Distillation for Edge Deep Learning
复制标题

DOI:
10.1109/iwqos52092.2021.9521267
复制
发表时间:
2021-06
期刊:
2021 IEEE/ACM 29th International Symposium on Quality of Service (IWQOS)
影响因子:
--
通讯作者:
Zeyi Tao;Qi Xia;Qun Li
Zeyi Tao;Qi Xia;Qun Li
中科院分区:
其他
文献类型:
--
作者:
Zeyi Tao;Qi Xia;Qun Li

文献摘要

相似文献

尽管深度神经网络在各种目标检测任务中显示出非凡的能力,但由于其高昂的计算成本,其在资源受限的设备或嵌入式系统上的部署非常具有挑战性。模型分割、剪枝或量化等方法的使用都是以准确性损失为代价的。最近提出的知识蒸馏(knowledge distillation, KD)旨在将模型知识从一个训练有素的模型(教师)转移到一个更小、更快的模型(学生),这可以显著降低计算成本、内存使用和延长电池寿命。在这项工作中,我们提出了一种新的神经元流形蒸馏(NMD),其中学生模型不仅模仿教师的输出激活,而且还学习教师的特征几何结构。我们的方法产生了高质量、紧凑和轻量级的学生模型。我们在多个数据集上进行了不同蒸馏配置的综合实验,并且所提出的方法证明了蒸馏模型在精度和速度权衡方面的一致改进。
Although deep neural networks show their extraordinary power in various object detection tasks, it is very challenging for them to be deployed on resource constrained devices or embedded systems due to their high computational cost. Efforts such as model partition, pruning or quantization have been used at an expense of accuracy loss. Recently proposed knowledge distillation (KD) aims at transferring model knowledge from a well-trained model (teacher) to a smaller and faster model (student), which can significantly reduce the computational cost, memory usage, and prolong the battery lifetime. In this work, we propose a novel neuron manifold distillation (NMD), where the student models not only imitate teacher’s output activations, but also learn the feature geometry structure of the teacher. Our approach produces a high-quality, compact, and lightweight student model. We conduct comprehensive experiments with different distillation configurations over multiple datasets, and the proposed method demonstrates a consistent improvement in accuracy-speed trade-offs for the distilled model.