MAZE: Data-Free Model Stealing Attack Using Zeroth-Order Gradient Estimation

MAZE: Data-Free Model Stealing Attack Using Zeroth-Order Gradient Estimation
复制标题

DOI:
10.1109/cvpr46437.2021.01360
复制
发表时间:
2020-05
期刊:
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
S. Kariyappa;A. Prakash;Moinuddin K. Qureshi
S. Kariyappa;A. Prakash;Moinuddin K. Qureshi
中科院分区:
其他
文献类型:
--
作者:
S. Kariyappa;A. Prakash;Moinuddin K. Qureshi

文献摘要

被引文献

相似文献

高质量的机器学习(ML)模型通常被公司视为有价值的知识产权(MS)攻击允许Blackbox访问ML模型的广告,以使用目标模型的预测来复制其功能但是,最佳现有MS攻击无法产生高临界克隆,而无需访问目标数据集或需要查询目标的代表性数据集在本文中,我们表明,防止对目标数据集的访问,我们提出了一个模型,我们提出了一个无数据的模型,该模型使用Zeroth-rorde梯度估计来窃取攻击在先前的工作中,迷宫仅使用使用通用模型创建的合成数据来执行MSS。甚至0.90×至0.99倍,甚至均超过了依赖部分数据(JBDA,克隆精度为0.13×to 0.69 x)的攻击,并且在替代数据上(仿制,克隆,克隆精度为0.52×x 0.52×至0.97 x)。在局部数据设置中的迷宫,并开发迷宫-PD,从而产生接近目标分布的综合数据进一步提高了克隆的准确性(0.97倍至1.0倍)攻击需要2×-24×。
High quality Machine Learning (ML) models are often considered valuable intellectual property by companies. Model Stealing (MS) attacks allow an adversary with blackbox access to a ML model to replicate its functionality by training a clone model using the predictions of the target model for different inputs. However, best available existing MS attacks fail to produce a high-accuracy clone without access to the target dataset or a representative dataset necessary to query the target model. In this paper, we show that preventing access to the target dataset is not an adequate defense to protect a model. We propose MAZE – a data-free model stealing attack using zeroth-order gradient estimation that produces high-accuracy clones. In contrast to prior works, MAZE uses only synthetic data created using a generative model to perform MS.Our evaluation with four image classification models shows that MAZE provides a normalized clone accuracy in the range of 0.90× to 0.99×, and outperforms even the recent attacks that rely on partial data (JBDA, clone accuracy 0.13× to 0.69×) and on surrogate data (KnockoffNets, clone accuracy 0.52× to 0.97×). We also study an extension of MAZE in the partial-data setting, and develop MAZE-PD, which generates synthetic data closer to the target distribution. MAZE-PD further improves the clone accuracy (0.97× to 1.0×) and reduces the query budget required for the attack by 2×-24×.