On the Feasibility of Training-time Trojan Attacks through Hardware-based Faults in Memory

On the Feasibility of Training-time Trojan Attacks through Hardware-based Faults in Memory
复制标题

DOI:
10.1109/host54066.2022.9840266
复制
发表时间:
2022-06
期刊:
2022 IEEE International Symposium on Hardware Oriented Security and Trust (HOST)
影响因子:
--
通讯作者:
Kunbei Cai;Zhenkai Zhang;F. Yao
Kunbei Cai;Zhenkai Zhang;F. Yao
中科院分区:
其他
文献类型:
--
作者:
Kunbei Cai;Zhenkai Zhang;F. Yao

文献摘要

相似文献

训练时木马攻击一直是可能篡改深度学习模型完整性的主要安全威胁之一。现有的木马攻击要么需要毒害训练数据集,要么依赖于对训练过程的控制。在本文中,我们研究了在训练时利用基于硬件的故障攻击在深度神经网络(DNN)中引入木马的实用性。具体来说,我们考虑使用 rowhammer 攻击向量进行基于内存的故障注入。我们提出了一种新的攻击框架,攻击者在训练期间向 DNN 模型的特征图注入故障。我们研究了特征图中位翻转的影响,并得出了位翻转策略,该策略使受害者模型能够将扰动的特征图模式与目标标签相关联,而不影响正常输入的预测。我们进一步提出了一种输入触发识别算法,该算法在推理时获取木马模型的触发模式。我们的评估表明,我们的攻击可以木马DNN模型,并且攻击成功率非常高。我们的工作强调了了解基于硬件的故障攻击在机器学习训练中的影响的重要性。
Training-time trojan attacks have been one of the major security threats that can tamper integrity of deep learning models. Existing trojan attacks either require poisoning of the training dataset or depend on control of the training process. In this paper, we investigate the practicality of leveraging hardware-based fault attacks to introduce trojan in deep neural networks (DNNs) at training time. Specifically, we consider a memory-based fault injection using the rowhammer attack vector. We propose a new attack framework where the adversary injects faults to the feature map of DNN models during training. We investigate the impact of bit flips in feature maps and derive a bit flip strategy that enables the victim model to associate a perturbed feature map pattern with a target label without impacting the prediction of normal inputs. We further propose an input trigger identification algorithm that obtains the trigger pattern for the trojaned model at inference time. Our evaluation shows that our attack can trojan DNN models with very high attack success rate. Our work highlights the importance of understanding the impact of hardware-based fault attacks in machine learning training.