Defending against model extraction attacks with physical unclonable function

Defending against model extraction attacks with physical unclonable function
复制标题

DOI:
10.1016/j.ins.2023.01.102
复制
发表时间:
2023-05
期刊:
Inf. Sci.
影响因子:
--
通讯作者:
Dawei Li;Di Liu;Ying Guo;Yangkun Ren;Jieyu Su;Jianwei Liu
Dawei Li;Di Liu;Ying Guo;Yangkun Ren;Jieyu Su;Jianwei Liu
中科院分区:
其他
文献类型:
--
作者:
Dawei Li;Di Liu;Ying Guo;Yangkun Ren;Jieyu Su;Jianwei Liu

文献摘要

相似文献

机器学习模型,特别是深度神经网络(DNN)模型,在商业活动中有着广泛而有价值的应用。训练用于商业用途的深度学习模型需要大量的私有数据、专家知识和计算资源。这种训练模型的巨大商业价值引起了攻击者的注意。攻击者可以通过反复查询目标模型以获取所请求样本的输出来构建数据集,然后在此数据集上训练与目标模型功能相似的替代模型。本文提出了一种基于物理不可克隆函数(PUF)的黑盒模型提取攻击防御方案。我们在用户端部署PUF,在服务提供者端部署相应的PUF模型,以确保只有合法用户才能获得正确的模型预测。我们的实验结果表明,通过选择合适的模糊提取器阈值,合法用户可以恢复99.5%以上的预测结果,而服务提供商的额外计算开销很少。在对攻击者最有利的情况下进行模型提取攻击,得到的替代模型预测精度仅为10%左右,验证了所提方案的有效性。与现有的防御措施相比,我们的方案不仅有效地防止了黑盒模型提取攻击,而且保证了合法用户预测服务的准确性不受影响。
Machine learning models, especially deep neural network (DNN) models, have widespread and valuable applications in business activities. Training a deep learning model for commercial use requires plenty of private data, expert knowledge, and computing resources. The huge commercial value of such trained models has attracted the attention of attackers. Attackers can construct a dataset by repeatedly querying the target model for the output of the requested samples and then train a substitute model on this dataset that functions similarly to the target model. In this paper, we propose a defense scheme based on physical unclonable function (PUF) against such black-box model extraction attacks. We deploy a PUF on the user side and the corresponding PUF model on the service provider side to ensure that only legitimate users can obtain the correct model predictions. Our experimental results show that by choosing a suitable fuzzy extractor thresholdd, legitimate users can recover more than 99.5% of the prediction results with a little additional computational overhead to the service provider. We perform a model extraction attack in the most favorable case for the attacker, and the prediction accuracy of the obtained substitute model is only about 10%, which demonstrates the effectiveness of our proposed scheme. Compared to existing defenses, our scheme not only effectively prevents black-box model extraction attacks but also ensures that the accuracy of the prediction service for legitimate users is not affected.