Undistillable: Making A Nasty Teacher That CANNOT teach students

Undistillable: Making A Nasty Teacher That CANNOT teach students
复制标题

DOI:
--
复制
发表时间:
2021-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Haoyu Ma;Tianlong Chen;Ting-Kuei Hu;Chenyu You;Xiaohui Xie;Zhangyang Wang
Haoyu Ma;Tianlong Chen;Ting-Kuei Hu;Chenyu You;Xiaohui Xie;Zhangyang Wang
中科院分区:
其他
文献类型:
--
作者:
Haoyu Ma;Tianlong Chen;Ting-Kuei Hu;Chenyu You;Xiaohui Xie;Zhangyang Wang

文献摘要

相似文献

知识蒸馏(Knowledge Distillation, KD)是一种广泛使用的技术,用于将知识从预训练的教师模型转移到(通常更轻量级的)学生模型。然而,在某些情况下,这种技术更多的是一种诅咒而不是祝福。例如,KD带来了暴露知识产权(ip)的潜在风险:即使经过训练的机器学习模型在“黑箱”中发布(例如,作为可执行软件或没有开源代码的api), KD仍然可以通过模仿输入输出行为来复制它。为了防止这种不必要的KD影响,本文引入并研究了一个称为“讨厌的老师”的概念:一个经过特殊训练的教师网络,其性能几乎与正常教师相同,但会显著降低通过模仿学习的学生模型的性能。我们提出了一种简单而有效的算法来构建讨厌的老师,称为自我破坏知识蒸馏。具体来说,我们的目标是使“讨厌的老师”的输出与正常的预训练网络的输出之间的差异最大化。在多个数据集上的大量实验表明,我们的方法对标准KD和无数据KD都是有效的,首次为模型所有者提供了理想的KD免疫。我们希望我们的初步研究能够引起人们对这一具有社会和法律重要性的新现实问题的更多认识和兴趣。
Knowledge Distillation (KD) is a widely used technique to transfer knowledge from pre-trained teacher models to (usually more lightweight) student models. However, in certain situations, this technique is more of a curse than a blessing. For instance, KD poses a potential risk of exposing intellectual properties (IPs): even if a trained machine learning model is released in 'black boxes' (e.g., as executable software or APIs without open-sourcing code), it can still be replicated by KD through imitating input-output behaviors. To prevent this unwanted effect of KD, this paper introduces and investigates a concept called Nasty Teacher: a specially trained teacher network that yields nearly the same performance as a normal one, but would significantly degrade the performance of student models learned by imitating it. We propose a simple yet effective algorithm to build the nasty teacher, called self-undermining knowledge distillation. Specifically, we aim to maximize the difference between the output of the nasty teacher and a normal pre-trained network. Extensive experiments on several datasets demonstrate that our method is effective on both standard KD and data-free KD, providing the desirable KD-immunity to model owners for the first time. We hope our preliminary study can draw more awareness and interest in this new practical problem of both social and legal importance.