Provable Robustness against Wasserstein Distribution Shifts via Input Randomization

Provable Robustness against Wasserstein Distribution Shifts via Input Randomization
复制标题

DOI:
--
复制
发表时间:
2023
期刊:
ArXiv
影响因子:
--
通讯作者:
Aounon Kumar;Alexander Levine;T. Goldstein;S. Feizi
Aounon Kumar;Alexander Levine;T. Goldstein;S. Feizi
中科院分区:
其他
文献类型:
--
作者:
Aounon Kumar;Alexander Levine;T. Goldstein;S. Feizi

文献摘要

相似文献

机器学习中经认证的鲁棒性主要集中于输入分布中每个样本具有固定攻击预算的对抗性扰动。在这项工作中,我们提出了在数据分布有界 Wasserstein 偏移下模型准确性的可证明的鲁棒性保证。我们证明,在变换空间内随机化模型输入的简单过程对于该变换下的分布变化具有鲁棒性。我们的框架允许特定于数据的扰动大小在输入分布中的不同点上变化,并且足够通用以包括固定大小的扰动。我们的证书为 Wasserstein 球内围绕原始分布的输入分布的任何变化(自然或对抗)提供了模型性能的保证下限。我们应用我们的技术来证明对图像自然(非对抗性)变换(例如颜色偏移、色调偏移以及亮度和饱和度变化)的鲁棒性。在输入图像中清晰可见的变化下,我们为鲁棒模型获得了强有力的性能保证。我们的实验通过证明鲁棒模型准确性的认证下限高于分布变化下无防御模型的经验准确性来证明我们的证书的非空性。我们还展示了针对对抗性攻击的可证明的分布式鲁棒性。此外,我们的结果还意味着在所谓的“不可学习”数据集上训练的模型的性能有保证的下限(硬度结果),这些数据集已被毒害以干扰模型训练。我们证明了稳健模型的性能保证保持不变
Certified robustness in machine learning has primarily focused on adversarial perturbations with a fixed attack budget for each sample in the input distribution. In this work, we present provable robustness guarantees on the accuracy of a model under bounded Wasserstein shifts of the data distribution. We show that a simple procedure that randomizes the input of the model within a transformation space is provably robust to distributional shifts under that transformation. Our framework allows the datum-specific perturbation size to vary across different points in the input distribution and is general enough to include fixed-sized perturbations as well. Our certificates produce guaranteed lower bounds on the performance of the model for any shift (natural or adversarial) of the input distribution within a Wasserstein ball around the original distribution. We apply our technique to certify robustness against natural (non-adversarial) transformations of images such as color shifts, hue shifts, and changes in brightness and saturation. We obtain strong performance guarantees for the robust model under clearly visible shifts in the input images. Our experiments establish the non-vacuousness of our certificates by showing that the certified lower bound on a robust model’s accuracy is higher than the empirical accuracy of an undefended model under a distribution shift. We also show provable distributional robustness against adversarial attacks. Moreover, our results also imply guaranteed lower bounds (hardness result) on the performance of models trained on so-called “unlearnable” datasets that have been poisoned to interfere with model training. We show that the performance of a robust model is guaranteed to remain