Semi-Targeted Model Poisoning Attack on Federated Learning via Backward Error Analysis

Semi-Targeted Model Poisoning Attack on Federated Learning via Backward Error Analysis
复制标题

DOI:
10.1109/ijcnn55064.2022.9891990
复制
发表时间:
2022-03
期刊:
2022 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
通讯作者:
Yuwei Sun;H. Ochiai;Jun Sakuma
Yuwei Sun;H. Ochiai;Jun Sakuma
中科院分区:
其他
文献类型:
--
作者:
Yuwei Sun;H. Ochiai;Jun Sakuma

文献摘要

被引文献

相似文献

针对联邦学习的模型中毒攻击是通过破坏边缘模型侵入整个系统,导致机器学习模型出现故障。这些受损的模型被篡改以执行对手期望的行为。特别地,我们考虑了一种半目标的情况,其中源类是预先确定的,而目标类不是。目标是使全局分类器对源类的数据进行错误分类。虽然已经采用标签翻转等方法将有毒参数注入到联邦学习中,但研究表明,它们的性能通常是类敏感的,随着所应用的目标类的不同而变化。通常情况下,当攻击转移到不同的目标职业时,攻击会变得不那么有效。为了克服这一挑战,我们提出了攻击距离感知攻击(ADA),通过在特征空间中找到优化的目标类来增强中毒攻击。此外,我们研究了一个更具挑战性的情况,即对手对客户数据的先验知识有限。为了解决这一问题,ADA基于后向误差分析,从共享模型参数中推断出潜在特征空间中不同类别之间的成对距离。我们通过在三种不同的图像分类任务中改变攻击频率因子对ADA进行了广泛的实证评估。结果,在最具挑战性的情况下,ADA成功地将攻击性能提高了1.8倍,攻击频率为0.01。
Model poisoning attacks on federated learning intrude in the entire system via compromising an edge model, resulting in malfunctioning of machine learning models. Such compromised models are tampered with to perform adversary-desired behaviors. In particular, we considered a semi-targeted situation where the source class is predetermined however the target class is not. The goal is to cause the global classifier to misclassify data of the source class. Though approaches such as label flipping have been adopted to inject poisoned parameters into federated learning, it has been shown that their performances are usually class-sensitive varying with different target classes applied. Typically, an attack can become less effective when shifting to a different target class. To overcome this challenge, we propose the Attacking Distance-aware Attack (ADA) to enhance a poisoning attack by finding the optimized target class in the feature space. Moreover, we studied a more challenging situation where an adversary had limited prior knowledge about a client's data. To tackle this problem, ADA deduces pair-wise distances between different classes in the latent feature space from shared model parameters based on the backward error analysis. We performed extensive empirical evaluations on ADA by varying the factor of attacking frequency in three different image classification tasks. As a result, ADA succeeded in increasing the attack performance by 1.8 times in the most challenging case with an attacking frequency of 0.01.