AQ2PNN: Enabling Two-party Privacy-Preserving Deep Neural Network Inference with Adaptive Quantization

AQ2PNN: Enabling Two-party Privacy-Preserving Deep Neural Network Inference with Adaptive Quantization
复制标题

DOI:
10.1145/3613424.3614297
复制
发表时间:
2023-10
期刊:
2023 56th IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
通讯作者:
Yukui Luo;Nuo Xu;Hongwu Peng;Chenghong Wang;Shijin Duan;Kaleel Mahmood;Wujie Wen;Caiwen Ding;Xiaolin Xu
Yukui Luo;Nuo Xu;Hongwu Peng;Chenghong Wang;Shijin Duan;Kaleel Mahmood;Wujie Wen;Caiwen Ding;Xiaolin Xu
中科院分区:
其他
文献类型:
--
作者:
Yukui Luo;Nuo Xu;Hongwu Peng;Chenghong Wang;Shijin Duan;Kaleel Mahmood;Wujie Wen;Caiwen Ding;Xiaolin Xu

文献摘要

相似文献

机器学习作为一项服务的日益普遍性(MLAAS)可以使广泛的应用程序提出了许多安全性和隐私问题。但是,最新的(SOTA)2pc-dnn技术是针对CPU和CPU+GPU等传统的指令架构(ISA)系统的定制。 2pc-dnns由于缺乏兼容的算法和硬件加速器,仍然很大程度上出乎意料。为了减轻SOTA解决方案的瓶颈并填补了现有的研究空白,研究了在现场可编程的门阵列上的2pc-dnns的构建,我们引入了量子范围,我们引入了ADNN-Endnn formations。 PGA。 AQ2PNN引入了一种创新的2PC-RELU方法,以替换Yao的乱码电路(GC)。 FPGA,以调整数据位DNN在密码域中的不同层,从而在没有损害DNN性能的情况下降低了跨越的沟通开销,例如,我们使用Resnet18,Resnet50和vgg16的DNN体系结构彻底评估AQ2PNN。沟通减少开销提高了25%,提高了能源效率26.3倍,并且可比较甚至优越的吞吐量和准确性。
The growing prevalence of Machine Learning as a Service (MLaaS) enables a wide range of applications but simultaneously raises numerous security and privacy concerns. A key issue involves the potential privacy exposure of involved parties, such as the customer’s input data and the vendor’s model. Consequently, two-party computing (2PC) has emerged as a promising solution to safeguard the privacy of different parties during deep neural network (DNN) inference. However, the state-of-the-art (SOTA) 2PC-DNN techniques are tailored explicitly to traditional instruction set architecture (ISA) systems like CPUs and CPU+GPU. This reliance on ISA systems significantly constrains their energy efficiency, as these architectures typically employ 32- or 64-bit instruction sets. In contrast, the possibilities of harnessing dynamic and adaptive quantization to build high-performance 2PC-DNNs remain largely unexplored due to the lack of compatible algorithms and hardware accelerators.To mitigate the bottleneck of SOTA solutions and fill the existing research gaps, this work investigates the construction of 2PC-DNNs on field programmable gate arrays (FPGAs). We introduce AQ2PNN, an end-to-end framework that effectively employs adaptive quantization schemes to develop high-performance 2PC-DNNs on FPGAs. From an algorithmic perspective, AQ2PNN introduces an innovative 2PC-ReLU method to replace Yao’s Garbled Circuits (GC). Regarding hardware, AQ2PNN employs an extensive set of building blocks for linear operators, non-linear operators, and a specialized Oblivious Transfer (OT) module for secure data exchange, respectively. These algorithm-hardware co-designed modules extremely utilize the fine-grained reconfigurability of FPGAs, to adapt the data bit-width of different DNN layers in the ciphertext domain, thereby reducing communication overhead between parties without compromising DNN performance, such as accuracy. We thoroughly assess AQ2PNN using widely adopted DNN architectures, including ResNet18, ResNet50, and VGG16, all trained on ImageNet and producing quantized models. Experimental results demonstrate that AQ2PNN outperforms SOTA solutions, achieving significantly reduced communication overhead by 25%, improved energy efficiency by 26.3×, and comparable or even superior throughput and accuracy.