CoPur: Certifiably Robust Collaborative Inference via Feature Purification

CoPur: Certifiably Robust Collaborative Inference via Feature Purification
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
J. Liu
J. Liu
中科院分区:
其他
文献类型:
--
作者:
J. Liu

文献摘要

相似文献

协作推理利用由不同代理提供的不同特征(例如,传感器)进行更准确的推断。一个常见的设置是每个代理将其嵌入的特征而不是原始数据发送到融合中心(FC)进行联合预测。在这种情况下,我们考虑推理阶段的攻击时,一小部分代理商受到损害。受损代理不会将嵌入式功能发送到FC,或者发送任意嵌入式功能。为了解决这个问题,我们提出了一个通过特征纯化(CoPur)的可验证的鲁棒协同推理框架,通过利用特征向量上对抗扰动的块稀疏性质,以及嵌入特征之间的冗余(假设整体特征位于底层的低维流形上)。我们从理论上表明,所提出的特征纯化方法可以鲁棒地恢复真实的特征向量,尽管对抗腐败和/或不完整的观察。我们还提出并测试了一种无针对性的分布式特征翻转攻击,该攻击对模型、训练数据、标签以及其他代理持有的特征都是不可知的,并且被证明在攻击最先进的防御方面是有效的。在ExtraSensory和NUS-WIDE数据集上的实验表明,CoPur在针对目标和非目标对抗性攻击的鲁棒性方面明显优于现有防御。
Collaborative inference leverages diverse features provided by different agents (e.g., sensors) for more accurate inference. A common setup is where each agent sends its embedded features instead of the raw data to the Fusion Center (FC) for joint prediction. In this setting, we consider inference phase attacks when a small fraction of agents is compromised. The compromised agent either does not send embedded features to the FC or sends arbitrary embedded features. To address this, we propose a certifiably robust COllaborative inference framework via feature PURification (CoPur), by leveraging the block-sparse nature of adversarial perturbations on the feature vector, as well as redundancy across the embedded features (by assuming the overall features lie on an underlying lower dimensional manifold). We theoretically show that the proposed feature purification method can robustly recover the true feature vector, despite adversarial corruptions and/or incomplete observations. We also propose and test an untargeted distributed feature-flipping attack, which is agnostic to the model, training data, label, as well as features held by other agents, and is shown to be effective in attacking state-of-the-art defenses. Experiments on ExtraSensory and NUS-WIDE datasets show that CoPur significantly outperforms existing defenses in terms of robustness against targeted and untargeted adversarial attacks.