Characterizing the Optimal 0-1 Loss for Multi-class Classification with a Test-time Attacker

Characterizing the Optimal 0-1 Loss for Multi-class Classification with a Test-time Attacker
复制标题

DOI:
10.48550/arxiv.2302.10722
复制
发表时间:
2023-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Sihui Dai;Wen-Luan Ding;A. Bhagoji;Daniel Cullina;Ben Y. Zhao;Haitao Zheng;Prateek Mittal
Sihui Dai;Wen-Luan Ding;A. Bhagoji;Daniel Cullina;Ben Y. Zhao;Haitao Zheng;Prateek Mittal
中科院分区:
其他
文献类型:
--
作者:
Sihui Dai;Wen-Luan Ding;A. Bhagoji;Daniel Cullina;Ben Y. Zhao;Haitao Zheng;Prateek Mittal

文献摘要

相似文献

找到对对抗性样本鲁棒的分类器对于它们的安全部署至关重要。因此,在给定威胁模型下针对给定数据分布确定最佳可能分类器的鲁棒性并将其与最先进的训练方法所实现的鲁棒性进行比较是一种重要的诊断工具。在本文中,我们找到了可实现的信息理论上的损失在任何离散数据集上的多类分类器的测试时攻击者的存在下的下限。我们提供了一个通用的框架,找到最佳的0-1损失,围绕从数据和对抗性约束的冲突超图的建设。我们进一步定义了攻击者-分类器游戏的其他变体,这些变体比成熟的超图构造更有效地确定最佳损失范围。我们的评估表明,第一次,分析的差距,以最佳的鲁棒性分类器在多类设置的基准数据集。
Finding classifiers robust to adversarial examples is critical for their safe deployment. Determining the robustness of the best possible classifier under a given threat model for a given data distribution and comparing it to that achieved by state-of-the-art training methods is thus an important diagnostic tool. In this paper, we find achievable information-theoretic lower bounds on loss in the presence of a test-time attacker for multi-class classifiers on any discrete dataset. We provide a general framework for finding the optimal 0-1 loss that revolves around the construction of a conflict hypergraph from the data and adversarial constraints. We further define other variants of the attacker-classifier game that determine the range of the optimal loss more efficiently than the full-fledged hypergraph construction. Our evaluation shows, for the first time, an analysis of the gap to optimal robustness for classifiers in the multi-class setting on benchmark datasets.