When Does Machine Learning FAIL? Generalized Transferability for Evasion and Poisoning Attacks

When Does Machine Learning FAIL? Generalized Transferability for Evasion and Poisoning Attacks
复制标题

DOI:
--
复制
发表时间:
2018-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Octavian Suciu;R. Marginean;Yigitcan Kaya;Hal Daumé;Tudor Dumitras
Octavian Suciu;R. Marginean;Yigitcan Kaya;Hal Daumé;Tudor Dumitras
中科院分区:
其他
文献类型:
--
作者:
Octavian Suciu;R. Marginean;Yigitcan Kaya;Hal Daumé;Tudor Dumitras

文献摘要

被引文献

相似文献

针对机器学习系统的攻击代表了一个日益增长的威胁,最近提出的大量攻击凸显了这一点。然而,攻击通常会对对手的知识和能力做出不切实际的假设。为了系统地评估这种威胁,我们提出了FAIL攻击者模型,该模型沿着四个维度描述了对手的知识和控制。FAIL模型允许我们考虑范围广泛的较弱的对手,这些对手的控制有限,对所使用的特征、学习算法和训练实例的知识不完整。在此框架内,我们评估了已知逃避攻击的一般可转移性,并设计了StingRay,这是一种广泛适用的针对性中毒攻击-它适用于4种机器学习应用程序,这些应用程序使用3种不同的学习算法,并且可以绕过2种现有防御。我们的评估提供了对中毒和逃避样本在模型之间的可转移性的更深入的见解,并为研究针对这种威胁的防御提供了有希望的方向。
Attacks against machine learning systems represent a growing threat as highlighted by the abundance of attacks proposed lately. However, attacks often make unrealistic assumptions about the knowledge and capabilities of adversaries. To evaluate this threat systematically, we propose the FAIL attacker model, which describes the adversary's knowledge and control along four dimensions. The FAIL model allows us to consider a wide range of weaker adversaries that have limited control and incomplete knowledge of the features, learning algorithms and training instances utilized. Within this framework, we evaluate the generalized transferability of a known evasion attack and we design StingRay, a targeted poisoning attack that is broadly applicable---it is practical against 4 machine learning applications, which use 3 different learning algorithms, and it can bypass 2 existing defenses. Our evaluation provides deeper insights into the transferability of poison and evasion samples across models and suggests promising directions for investigating defenses against this threat.