SpENCNN: Orchestrating Encoding and Sparsity for Fast Homomorphically Encrypted Neural Network Inference

SpENCNN: Orchestrating Encoding and Sparsity for Fast Homomorphically Encrypted Neural Network Inference
复制标题

DOI:
--
复制
发表时间:
2023
期刊:
--
影响因子:
--
通讯作者:
Ran Ran-Ran;Xinwei Luo;Wei Wang;Tao Liu;Gang Quan;Xiaolin Xu;Caiwen Ding;Wujie Wen
Ran Ran-Ran;Xinwei Luo;Wei Wang;Tao Liu;Gang Quan;Xiaolin Xu;Caiwen Ding;Wujie Wen
中科院分区:
其他
文献类型:
--
作者:
Ran Ran-Ran;Xinwei Luo;Wei Wang;Tao Liu;Gang Quan;Xiaolin Xu;Caiwen Ding;Wujie Wen

文献摘要

相似文献

同态加密(HE)是一种很有前途的技术,可以保护公共云上机器学习即服务(MLaaS)的客户端数据隐私。然而,HE操作可能比明文的对应操作慢几个数量级,从而导致过高的推理延迟,严重阻碍了HE的实用性。在本文中,我们提出了一个基于HE的快速神经网络(NN)推理框架-SpENCNN建立在HE操作感知模型稀疏性和单指令多数据(SIMD)友好的数据包装的协同设计,以改善NN推理延迟。特别是,我们首先开发了一种基于数据大小和密文大小的加密感知的HE组卷积技术,该技术可以在不同的组之间划分通道,然后通过新颖的组交织编码将它们编码为相同的密文,从而大大减少HE卷积中的加密校验操作的数量。我们进一步定制了HE友好的子块权重修剪,以减少昂贵的基于HE的卷积操作。我们的实验表明,SpENCNN可以分别为LeNet,VGG-5,HEFNet和ResNet-20实现8.37 ×,12.11 ×,19.26 ×和1.87 ×的整体加速,精度损失可以忽略不计。我们的代码可以在https://github上公开获取。网站/
Homomorphic Encryption (HE) is a promising technology to protect clients’ data privacy for Machine Learning as a Service (MLaaS) on public clouds. However, HE operations can be orders of magnitude slower than their counterparts for plaintexts and thus result in prohibitively high inference latency, seriously hindering the practicality of HE. In this paper, we propose a HE-based fast neural network (NN) inference framework–SpENCNN built upon the co-design of HE operation-aware model sparsity and the single-instruction-multiple-data (SIMD)-friendly data packing, to improve NN inference latency. In particular, we first develop an encryption-aware HE-group convolution technique that can partition channels among different groups based on the data size and ciphertext size, and then encode them into the same ciphertext by novel group-interleaved encoding, so as to dramatically reduce the number of bottlenecked operations in HE convolution. We further tailor a HE-friendly sub-block weight pruning to reduce the costly HE-based convolution operation. Our experiments show that SpENCNN can achieve overall speedups of 8.37 × , 12.11 × , 19.26 × , and 1.87 × for LeNet, VGG-5, HEFNet, and ResNet-20 respectively, with negligible accuracy loss. Our code is publicly available at https://github. com/