The vulnerability of learning to adversarial perturbation increases with intrinsic dimensionality

The vulnerability of learning to adversarial perturbation increases with intrinsic dimensionality
复制标题

DOI:
10.1109/wifs.2017.8267651
复制
发表时间:
2017-12
期刊:
2017 IEEE Workshop on Information Forensics and Security (WIFS)
影响因子:
--
通讯作者:
L. Amsaleg;J. Bailey;Dominique Barbe;S. Erfani;Michael E. Houle;Vinh Nguyen;Miloš Radovanović
L. Amsaleg;J. Bailey;Dominique Barbe;S. Erfani;Michael E. Houle;Vinh Nguyen;Miloš Radovanović
中科院分区:
其他
文献类型:
--
作者:
L. Amsaleg;J. Bailey;Dominique Barbe;S. Erfani;Michael E. Houle;Vinh Nguyen;Miloš Radovanović

文献摘要

相似文献

最近的研究表明,机器学习系统,包括最先进的深度神经网络,很容易受到对抗性攻击。通过向输入对象添加难以察觉的对抗性噪声,分类器很可能被欺骗将修改后的对象分配给任何所需的类。还观察到这些对抗性样本在不同模型中具有良好的泛化性。对对抗样本性质的完整理解尚未出现。为了实现这一目标,我们提出了一种新颖的理论结果,将学习的对抗性脆弱性与数据的内在维度正式联系起来。特别是,我们的研究表明,随着局部固有维度 (LID) 的增加,1-NN 分类器变得越来越容易被颠覆。我们表明,在预期中,通过添加随着 LID 增加而减少的噪声量,可以将测试点的 k 近邻转换为其 1 近邻。我们还对 LID 对合成数据和真实数据的对抗性扰动的影响进行了实验验证,并讨论了我们的结果对通用分类器的影响。
Recent research has shown that machine learning systems, including state-of-the-art deep neural networks, are vulnerable to adversarial attacks. By adding to the input object an imperceptible amount of adversarial noise, it is highly likely that the classifier can be tricked into assigning the modified object to any desired class. It has also been observed that these adversarial samples generalize well across models. A complete understanding of the nature of adversarial samples has not yet emerged. Towards this goal, we present a novel theoretical result formally linking the adversarial vulnerability of learning to the intrinsic dimensionality of the data. In particular, our investigation establishes that as the local intrinsic dimensionality (LID) increases, 1-NN classifiers become increasingly prone to being subverted. We show that in expectation, a k-nearest neighbor of a test point can be transformed into its 1-nearest neighbor by adding an amount of noise that diminishes as the LID increases. We also provide an experimental validation of the impact of LID on adversarial perturbation for both synthetic and real data, and discuss the implications of our result for general classifiers.