Characterization of MPC-based Private Inference for Transformer-based Models

Characterization of MPC-based Private Inference for Transformer-based Models
复制标题

DOI:
10.1109/ispass55109.2022.00025
复制
发表时间:
2022-05
期刊:
2022 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS)
影响因子:
--
通讯作者:
Yongqin Wang;G. Suh;Wenjie Xiong;Benjamin Lefaudeux;Brian Knott;M. Annavaram;Hsien-Hsin S. Lee
Yongqin Wang;G. Suh;Wenjie Xiong;Benjamin Lefaudeux;Brian Knott;M. Annavaram;Hsien-Hsin S. Lee
中科院分区:
其他
文献类型:
--
作者:
Yongqin Wang;G. Suh;Wenjie Xiong;Benjamin Lefaudeux;Brian Knott;M. Annavaram;Hsien-Hsin S. Lee

文献摘要

相似文献

在这项工作中,我们提供了一个深入的表征研究运行Transformer模型与安全多方计算(MPC)的性能开销。MPC是一种加密框架,用于在存在不可信计算节点的情况下保护模型和输入数据隐私。我们的表征研究表明,Transformers为基于MPC的私有机器学习推理带来了几个性能挑战。首先,Transformers广泛依赖“softmax”功能。虽然softmax函数在非私有执行中相对便宜,但softmax在MPC推理运行时占主导地位,消耗了总推理运行时的50%。进一步的研究表明,计算最大值(为softmax提供数值稳定性所需)是延迟增加的关键原因。其次,MPC依赖于近似作为softmax计算的一部分的非线性函数,并且窄的动态范围使得在保持精度的同时优化softmax变得非常困难。最后,与CNN不同,基于Transformer的NLP模型使用大型嵌入表将输入单词转换为嵌入向量。对这些嵌入表的访问可能会泄露输入;因此,需要对嵌入访问模式进行额外的混淆以保证输入隐私。隐藏地址访问的一种方法是将嵌入表查找转换为矩阵乘法。然而,这种天真的方法显着增加了MPC推理运行时间。然后,我们应用张量训练(TT)分解,一个有损压缩技术表示嵌入表,并评估其性能嵌入查找。我们使用详细的实验显示性能改进和相应的模型精度的影响之间的权衡。
In this work, we provide an in-depth characterization study of the performance overhead for running Transformer models with secure multi-party computation (MPC). MPC is a cryptographic framework for protecting both the model and input data privacy in the presence of untrusted compute nodes. Our characterization study shows that Transformers introduce several performance challenges for MPC-based private machine learning inference. First, Transformers rely extensively on “softmax” functions. While softmax functions are relatively cheap in a non-private execution, softmax dominates the MPC inference runtime, consuming up to 50% of the total inference runtime. Further investigation shows that computing the maximum, needed for providing numerical stability to softmax, is a key culprit for the increase in latency. Second, MPC relies on approximating non-linear functions that are part of the softmax computations, and the narrow dynamic ranges make optimizing softmax while maintaining accuracy quite difficult. Finally, unlike CNNs, Transformer-based NLP models use large embedding tables to convert input words into embedding vectors. Accesses to these embedding tables can disclose inputs; hence, additional obfuscation for embedding access patterns is required for guaranteeing the input privacy. One approach to hide address accesses is to convert an embedding table lookup into a matrix multiplication. However, this naive approach increases MPC inference runtime significantly. We then apply tensor-train (TT) decomposition, a lossy compression technique for representing embedding tables, and evaluate its performance on embedding lookups. We show the trade-off between performance improvements and the corresponding impact on model accuracy using detailed experiments.