Revisiting Data-Free Knowledge Distillation with Poisoned Teachers

Revisiting Data-Free Knowledge Distillation with Poisoned Teachers
复制标题

DOI:
10.48550/arxiv.2306.02368
复制
发表时间:
2023-06
期刊:
--
影响因子:
--
通讯作者:
Junyuan Hong;Yi Zeng;Shuyang Yu;L. Lyu;R. Jia;Jiayu Zhou
Junyuan Hong;Yi Zeng;Shuyang Yu;L. Lyu;R. Jia;Jiayu Zhou
中科院分区:
其他
文献类型:
--
作者:
Junyuan Hong;Yi Zeng;Shuyang Yu;L. Lyu;R. Jia;Jiayu Zhou

文献摘要

相似文献

无数据知识蒸馏 (KD) 有助于将知识从预训练模型(称为教师模型)转移到较小的模型(称为学生模型),而无需访问用于训练教师模型的原始训练数据。然而,无数据 KD 中所需的合成数据或分发外 (OOD) 数据的安全性在很大程度上是未知的且尚未得到充分探索。在这项工作中,我们首先努力揭示无数据 KD 的安全风险。不受信任的预训练模型。然后,我们提出了反后门无数据 KD(ABD),这是第一个无数据 KD 方法的插件防御方法,以减少潜在后门被转移的机会。我们根据经验评估了我们提出的 ABD 在减少转移的后门知识方面的有效性,同时保持与普通 KD 兼容的下游性能。我们预计这项工作将成为警告和减轻无数据 KD 中潜在后门的里程碑。代码发布于 https://github.com/illidanlab/ABD。
Data-free knowledge distillation (KD) helps transfer knowledge from a pre-trained model (known as the teacher model) to a smaller model (known as the student model) without access to the original training data used for training the teacher model. However, the security of the synthetic or out-of-distribution (OOD) data required in data-free KD is largely unknown and under-explored. In this work, we make the first effort to uncover the security risk of data-free KD w.r.t. untrusted pre-trained models. We then propose Anti-Backdoor Data-Free KD (ABD), the first plug-in defensive method for data-free KD methods to mitigate the chance of potential backdoors being transferred. We empirically evaluate the effectiveness of our proposed ABD in diminishing transferred backdoor knowledge while maintaining compatible downstream performances as the vanilla KD. We envision this work as a milestone for alarming and mitigating the potential backdoors in data-free KD. Codes are released at https://github.com/illidanlab/ABD.