Systematic Testing of the Data-Poisoning Robustness of KNN

Systematic Testing of the Data-Poisoning Robustness of KNN
复制标题

DOI:
10.1145/3597926.3598129
复制
发表时间:
2023-07
期刊:
Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis
影响因子:
--
通讯作者:
Yannan Li;Jingbo Wang;Chao Wang
Yannan Li;Jingbo Wang;Chao Wang
中科院分区:
其他
文献类型:
--
作者:
Yannan Li;Jingbo Wang;Chao Wang

文献摘要

相似文献

数据中毒旨在通过污染基于机器学习的软件组件的训练集来更改其测试输入的预测结果,从而危及其安全。现有的判定数据中毒稳健性的方法要么准确性较差,要么运行时间较长,更重要的是,它们只能验证一些真正健壮的情况,但当验证失败时,它们仍然没有定论。换句话说,他们不能伪造真正不可靠的案例。为了克服这一局限性,我们提出了一种基于系统测试的方法,该方法可以对一种广泛使用的监督学习技术k-近邻(KNN)进行篡改和数据中毒稳健性证明。我们的方法比基线枚举法更快更准确,由于在抽象域中采用了新颖的过度近似分析,快速缩小了搜索空间,并在具体域中进行了系统的测试,找到了实际的违规行为。我们已经在一组监督学习数据集上对我们的方法进行了评估。我们的结果表明,该方法的性能明显优于最新技术,并且可以确定KNN预测结果在大多数测试输入下的数据中毒稳健性。
Data poisoning aims to compromise a machine learning based software component by contaminating its training set to change its prediction results for test inputs. Existing methods for deciding data-poisoning robustness have either poor accuracy or long running time and, more importantly, they can only certify some of the truly-robust cases, but remain inconclusive when certification fails. In other words, they cannot falsify the truly-non-robust cases. To overcome this limitation, we propose a systematic testing based method, which can falsify as well as certify data-poisoning robustness for a widely used supervised-learning technique named k-nearest neighbors (KNN). Our method is faster and more accurate than the baseline enumeration method, due to a novel over-approximate analysis in the abstract domain, to quickly narrow down the search space, and systematic testing in the concrete domain, to find the actual violations. We have evaluated our method on a set of supervised-learning datasets. Our results show that the method significantly outperforms state-of-the-art techniques, and can decide data-poisoning robustness of KNN prediction results for most of the test inputs.