Excess Capacity and Backdoor Poisoning

Excess Capacity and Backdoor Poisoning
复制标题

DOI:
--
复制
发表时间:
2021-09
期刊:
--
影响因子:
--
通讯作者:
N. Manoj;Avrim Blum
N. Manoj;Avrim Blum
中科院分区:
其他
文献类型:
--
作者:
N. Manoj;Avrim Blum

文献摘要

相似文献

后门数据中毒攻击是一种对抗性攻击,其中攻击者将几个带水印的、错误标记的训练样本注入到训练集中。水印不会影响模型在典型数据上的测试时性能;然而,模型在水印示例上可靠地出错。为了更好地了解后门数据中毒攻击的基础,我们提出了一个正式的理论框架,在此框架内,人们可以讨论后门数据中毒攻击的分类问题。然后,我们使用它来分析围绕这些攻击的重要统计和计算问题。在统计方面,我们确定了一个参数,我们称之为记忆能力,它捕捉了学习问题对后门攻击的内在脆弱性。这使我们能够讨论几个自然学习问题对后门攻击的鲁棒性。我们的研究结果有利于攻击者提出明确的后门攻击的建设,我们的鲁棒性结果表明,一些自然的问题设置不能产生成功的后门攻击。从计算的角度来看,我们证明了在某些假设下,对抗训练可以检测到训练集中后门的存在。然后,我们表明,在类似的假设下,两个密切相关的问题,我们称之为后门过滤和强大的推广几乎是等价的。这意味着设计可以识别训练集中的水印示例的算法是渐进必要和充分的,以便获得既能很好地推广到看不见的数据又能对后门鲁棒的学习算法。
A backdoor data poisoning attack is an adversarial attack wherein the attacker injects several watermarked, mislabeled training examples into a training set. The watermark does not impact the test-time performance of the model on typical data; however, the model reliably errs on watermarked examples. To gain a better foundational understanding of backdoor data poisoning attacks, we present a formal theoretical framework within which one can discuss backdoor data poisoning attacks for classification problems. We then use this to analyze important statistical and computational issues surrounding these attacks. On the statistical front, we identify a parameter we call the memorization capacity that captures the intrinsic vulnerability of a learning problem to a backdoor attack. This allows us to argue about the robustness of several natural learning problems to backdoor attacks. Our results favoring the attacker involve presenting explicit constructions of backdoor attacks, and our robustness results show that some natural problem settings cannot yield successful backdoor attacks. From a computational standpoint, we show that under certain assumptions, adversarial training can detect the presence of backdoors in a training set. We then show that under similar assumptions, two closely related problems we call backdoor filtering and robust generalization are nearly equivalent. This implies that it is both asymptotically necessary and sufficient to design algorithms that can identify watermarked examples in the training set in order to obtain a learning algorithm that both generalizes well to unseen data and is robust to backdoors.