DisGUIDE: Disagreement-Guided Data-Free Model Extraction

DisGUIDE: Disagreement-Guided Data-Free Model Extraction
复制标题

DOI:
10.1609/aaai.v37i8.26150
复制
发表时间:
2023-06
期刊:
2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Jonathan Rosenthal;Eric Enouen;H. Pham;Lin Tan
Jonathan Rosenthal;Eric Enouen;H. Pham;Lin Tan
中科院分区:
其他
文献类型:
--
作者:
Jonathan Rosenthal;Eric Enouen;H. Pham;Lin Tan

文献摘要

相似文献

对机器学习作为服务(MLAAS)系统的最新模型抢劫攻击已转向无数据的方法,显示了窃取经过难以访问数据的模型的可行性。但是,由于提取的模型的精度较低,并且对攻击模型的查询数量很大,因此这些攻击是无效或有限的。高查询成本使此类技术对于每个查询收费的在线MLAA系统不可行。与先前的无数据提取技术相比,我们创建了一种新颖的方法来获得更高的准确性和查询效率。具体来说,我们介绍了一种新颖的生成器培训方案,该方案最大化了两个试图在攻击下复制模型的克隆模型之间的分歧损失。这种损失加上多样性损失和经验重播,使发电机能够生成更好的实例来训练克隆模型。我们对流行数据集CIFAR-10和CIFAR-100的评估表明,我们的方法分别提高了最终模型准确性高达3.42%和18.48%。实现先前技术状态所需的平均查询数量最多减少了64.95%。我们希望这将促进对此类攻击的可行无数据模型提取和防御措施的未来工作。
Recent model-extraction attacks on Machine Learning as a Service (MLaaS) systems have moved towards data-free approaches, showing the feasibility of stealing models trained with difficult-to-access data. However, these attacks are ineffective or limited due to the low accuracy of extracted models and the high number of queries to the models under attack. The high query cost makes such techniques infeasible for online MLaaS systems that charge per query. We create a novel approach to get higher accuracy and query efficiency than prior data-free model extraction techniques. Specifically, we introduce a novel generator training scheme that maximizes the disagreement loss between two clone models that attempt to copy the model under attack. This loss, combined with diversity loss and experience replay, enables the generator to produce better instances to train the clone models. Our evaluation on popular datasets CIFAR-10 and CIFAR-100 shows that our approach improves the final model accuracy by up to 3.42% and 18.48% respectively. The average number of queries required to achieve the accuracy of the prior state of the art is reduced by up to 64.95%. We hope this will promote future work on feasible data-free model extraction and defenses against such attacks.