A Privacy-Preserving-Oriented DNN Pruning and Mobile Acceleration Framework

A Privacy-Preserving-Oriented DNN Pruning and Mobile Acceleration Framework
复制标题

DOI:
10.1145/3386263.3407650
复制
发表时间:
2020-03
期刊:
Proceedings of the 2020 on Great Lakes Symposium on VLSI
影响因子:
--
通讯作者:
Yifan Gong;Zheng Zhan;Z. Li;Wei Niu;Xiaolong Ma;Wenhao Wang;Bin Ren;Caiwen Ding;X. Lin;Xiaolin Xu;Yanzhi Wang
Yifan Gong;Zheng Zhan;Z. Li;Wei Niu;Xiaolong Ma;Wenhao Wang;Bin Ren;Caiwen Ding;X. Lin;Xiaolin Xu;Yanzhi Wang
中科院分区:
其他
文献类型:
--
作者:
Yifan Gong;Zheng Zhan;Z. Li;Wei Niu;Xiaolong Ma;Wenhao Wang;Bin Ren;Caiwen Ding;X. Lin;Xiaolin Xu;Yanzhi Wang

文献摘要

相似文献

已经提出了深度神经网络(DNN)的权重修剪以满足移动的边缘设备的有限存储和计算能力。然而,以前的修剪方法主要集中在减少模型大小和/或提高性能,而不考虑用户数据的隐私。为了减轻这种担忧,我们提出了一个面向隐私保护的修剪和移动的加速框架,它不需要私有训练数据集。在算法层次上,提出了一种基于交替方向乘法(ADMM)的系统权重剪枝方法,利用随机生成的合成数据迭代求解每一层的基于模式的剪枝问题.此外,编译器级别的相应优化被用于设备上的推理加速。使用该框架,用户可以避免非专家的耗时修剪过程,并直接受益于压缩模型。实验结果表明,所提出的框架优于三个最先进的端到端DNN框架,即,TensorFlow-Lite、TVM和MNN,加速分别高达4.2倍、2.5倍和2.0倍,几乎没有精度损失,同时保护数据隐私。
Weight pruning of deep neural networks (DNNs) has been proposed to satisfy the limited storage and computing capability of mobile edge devices. However, previous pruning methods mainly focus on reducing the model size and/or improving performance without considering the privacy of user data. To mitigate this concern, we propose a privacy-preserving-oriented pruning and mobile acceleration framework that does not require the private training dataset. At the algorithm level of the proposed framework, a systematic weight pruning technique based on the alternating direction method of multipliers (ADMM) is designed to iteratively solve the pattern-based pruning problem for each layer with randomly generated synthetic data. In addition, corresponding optimizations at the compiler level are leveraged for inference accelerations on devices. With the proposed framework, users could avoid the time-consuming pruning process for non-experts and directly benefit from compressed models. Experimental results show that the proposed framework outperforms three state-of-art end-to-end DNN frameworks, i.e., TensorFlow-Lite, TVM, and MNN, with speedup up to 4.2×, 2.5×, and 2.0×, respectively, with almost no accuracy loss, while preserving data privacy.