SPECTRE: Defending Against Backdoor Attacks Using Robust Covariance Estimation

SPECTRE: Defending Against Backdoor Attacks Using Robust Covariance Estimation
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
J. Hayase;Weihao Kong;Ragahv Somani;Sewoong Oh
J. Hayase;Weihao Kong;Ragahv Somani;Sewoong Oh
中科院分区:
其他
文献类型:
--
作者:
J. Hayase;Weihao Kong;Ragahv Somani;Sewoong Oh

文献摘要

相似文献

现代机器学习越来越需要对来自多个来源的大量数据进行训练,但并非所有数据都是可信的。一个特别值得关注的场景是,当一小部分中毒数据被攻击者指定的艾德水印触发时,会改变训练模型的行为。这样一个受损的模型将不被注意到,因为该模型是准确的,否则。已经有一些有希望的尝试使用这种模型的中间表示来将损坏的示例与干净的示例分开。然而,这些防御措施只有在中毒样本的特定光谱特征足够大以进行检测时才有效。现有的防御措施无法抵御各种攻击。我们提出了一种新的防御算法,使用强大的协方差估计,以放大损坏的数据的频谱特征。这种防御提供了一个干净的模型,完全消除了后门,即使在以前的方法没有希望检测到中毒的例子的政权。2
Modern machine learning increasingly requires training on a large collection of data from multiple sources, not all of which can be trusted. A particularly concerning scenario is when a small fraction of poisoned data changes the behavior of the trained model when triggered by an attacker-specified watermark. Such a compromised model will be deployed unnoticed as the model is accurate otherwise. There have been promising attempts to use the intermediate representations of such a model to separate corrupted examples from clean ones. However, these defenses work only when a certain spectral signature of the poisoned examples is large enough for detection. There is a wide range of attacks that cannot be protected against by the existing defenses. We propose a novel defense algorithm using robust covariance estimation to amplify the spectral signature of corrupted data. This defense provides a clean model, completely removing the backdoor, even in regimes where previous methods have no hope of detecting the poisoned examples. 2