Aspis: Robust Detection for Distributed Learning

Aspis: Robust Detection for Distributed Learning
复制标题

DOI:
10.1109/isit50566.2022.9834813
复制
发表时间:
2021-08
期刊:
2022 IEEE International Symposium on Information Theory (ISIT)
影响因子:
--
通讯作者:
Konstantinos Konstantinidis;A. Ramamoorthy
Konstantinos Konstantinidis;A. Ramamoorthy
中科院分区:
其他
文献类型:
--
作者:
Konstantinos Konstantinidis;A. Ramamoorthy

文献摘要

相似文献

最先进的机器学习模型经常在大规模分布式群集上进行培训。至关重要的是,当某些计算设备表现出异常(拜占庭)行为并将任意结果返回到参数服务器(PS)时,可能会损害此类系统。这种行为可能归因于许多原因,包括系统故障和精心策划的攻击。现有工作表明鲁棒的聚集和/或计算冗余,以减轻扭曲梯度的影响。但是,当对手知道任务任务并可以明智地选择受到攻击的工人以造成最大损害时,这些方案中的大多数都是无效的。我们提出的方法ASPI使用基于子集的作业将梯度计算分配给工人,该分配允许对工人的行为进行多次一致性检查。通过PS进行适当构造的图表中计算出的梯度和集团找到的检查,可以有效地检测和排除训练中的对手。我们证明了在弱和强烈攻击下ASPI的拜占庭的弹性保证,并在各种培训方案上广泛评估了该系统,并且与CIFAR-10数据集中的许多最新方法相比,准确性提高了约30%的准确性以及减少损坏梯度的比例从16%到99%。
State-of-the-art machine learning models are routinely trained on large-scale distributed clusters. Crucially, such systems can be compromised when some of the computing devices exhibit abnormal (Byzantine) behavior and return arbitrary results to the parameter server (PS). This behavior may be attributed to a plethora of reasons, including system failures and orchestrated attacks. Existing work suggests robust aggregation and/or computational redundancy to alleviate the effect of distorted gradients. However, most of these schemes are ineffective when an adversary knows the task assignment and can choose the attacked workers judiciously to induce maximal damage. Our proposed method Aspis assigns gradient computations to workers using a subset-based assignment which allows for multiple consistency checks on the behavior of a worker. Examination of the calculated gradients and clique-finding in an appropriately constructed graph by the PS allows for efficient detection and exclusion of adversaries from the training. We prove the Byzantine resilience guarantees of Aspis under weak and strong attacks and extensively evaluate the system on various training scenarios and demonstrate an improvement of about 30% in accuracy compared to many state-of-the-art approaches on the CIFAR-10 dataset as well as reduction of the fraction of corrupted gradients ranging from 16% to 99%.