Towards Understanding and Enhancing Robustness of Deep Learning Models against Malicious Unlearning Attacks

Towards Understanding and Enhancing Robustness of Deep Learning Models against Malicious Unlearning Attacks
复制标题

DOI:
10.1145/3580305.3599526
复制
发表时间:
2023-08
期刊:
Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining
影响因子:
--
通讯作者:
Wei Qian-;Chenxu Zhao;Wei Le;Meiyi Ma;Mengdi Huai
Wei Qian-;Chenxu Zhao;Wei Le;Meiyi Ma;Mengdi Huai
中科院分区:
其他
文献类型:
--
作者:
Wei Qian-;Chenxu Zhao;Wei Le;Meiyi Ma;Mengdi Huai

文献摘要

被引文献

相似文献

鉴于丰富的数据的可用性,深度学习模型在过去十年中得到了发展并变得无处不在。在实践中,由于许多不同的原因(例如隐私、可用性和保真度),个人也希望经过训练的深度模型忘记一些特定数据。受此推动,机器遗忘(也称为选择性数据遗忘)得到了深入研究,其目的是消除在遗忘过程中任何特定训练样本对训练模型的影响。然而,人们通常将机器遗忘方法作为值得信赖的基本工具,很少对其可靠性产生任何怀疑。事实上,机器遗忘的作用日益重要,使得深度学习模型容易受到恶意攻击的风险。为了更好地了解深度学习模型在恶意环境中的性能,我们认为研究深度学习模型对在遗忘过程中发生的恶意遗忘攻击的鲁棒性至关重要。为了弥补这一差距,在本文中,我们首先证明恶意遗忘攻击对深度学习系统的安全构成巨大威胁。具体来说,我们提出了一类广泛的恶意遗忘攻击,其中恶意制作的遗忘请求触发深度学习模型以高度可控和可预测的方式对目标样本做出错误行为。此外,为了提高深度学习模型的鲁棒性,我们还提出了一种通用的防御机制,旨在根据有效的恶意取消学习请求对未学习模型的梯度影响来识别和取消学习。此外,还进行了理论分析来分析所提出的方法。对真实世界数据集的大量实验验证了深度学习模型对恶意遗忘攻击的脆弱性以及所引入的防御机制的有效性。
Given the availability of abundant data, deep learning models have been advanced and become ubiquitous in the past decade. In practice, due to many different reasons (e.g., privacy, usability, and fidelity), individuals also want the trained deep models to forget some specific data. Motivated by this, machine unlearning (also known as selective data forgetting) has been intensively studied, which aims at removing the influence that any particular training sample had on the trained model during the unlearning process. However, people usually employ machine unlearning methods as trusted basic tools and rarely have any doubt about their reliability. In fact, the increasingly critical role of machine unlearning makes deep learning models susceptible to the risk of being maliciously attacked. To well understand the performance of deep learning models in malicious environments, we believe that it is critical to study the robustness of deep learning models to malicious unlearning attacks, which happen during the unlearning process. To bridge this gap, in this paper, we first demonstrate that malicious unlearning attacks pose immense threats to the security of deep learning systems. Specifically, we present a broad class of malicious unlearning attacks wherein maliciously crafted unlearning requests trigger deep learning models to misbehave on target samples in a highly controllable and predictable manner. In addition, to improve the robustness of deep learning models, we also present a general defense mechanism, which aims to identify and unlearn effective malicious unlearning requests based on their gradient influence on the unlearned models. Further, theoretical analyses are conducted to analyze the proposed methods. Extensive experiments on real-world datasets validate the vulnerabilities of deep learning models to malicious unlearning attacks and the effectiveness of the introduced defense mechanism.