DETOX: A Redundancy-based Framework for Faster and More Robust Gradient Aggregation

DETOX: A Redundancy-based Framework for Faster and More Robust Gradient Aggregation
复制标题

DOI:
--
复制
发表时间:
2019-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Shashank Rajput;Hongyi Wang;Zachary B. Charles;Dimitris Papailiopoulos
Shashank Rajput;Hongyi Wang;Zachary B. Charles;Dimitris Papailiopoulos
中科院分区:
其他
文献类型:
--
作者:
Shashank Rajput;Hongyi Wang;Zachary B. Charles;Dimitris Papailiopoulos

文献摘要

相似文献

为了提高分布式训练对最坏情况或拜占庭节点故障的恢复能力,最近的几种方法已经用稳健的聚合方法取代了梯度平均。此类技术可能具有较高的计算成本,通常是计算节点数量的二次方,并且仅具有有限的鲁棒性保证。其他方法改为使用冗余来保证鲁棒性,但只能容忍有限数量的拜占庭故障。在这项工作中,我们提出了 DETOX,一种拜占庭弹性分布式训练框架,它将算法冗余与鲁棒聚合相结合。 DETOX 分两个步骤运行,一个是过滤步骤,使用有限的冗余来显着减少拜占庭节点的影响,另一个是分层聚合步骤,可以与任何最先进的鲁棒聚合方法一起使用。我们从理论上证明,这会导致鲁棒性的大幅提高,并且每次迭代的运行时间与计算节点的数量几乎呈线性关系。我们在各种大规模机器学习任务中对真实的分布式设置进行了广泛的实验,表明 DETOX 比许多最先进的拜占庭弹性方法带来了数量级的准确性和加速改进。
To improve the resilience of distributed training to worst-case, or Byzantine node failures, several recent approaches have replaced gradient averaging with robust aggregation methods. Such techniques can have high computational costs, often quadratic in the number of compute nodes, and only have limited robustness guarantees. Other methods have instead used redundancy to guarantee robustness, but can only tolerate limited number of Byzantine failures. In this work, we present DETOX, a Byzantine-resilient distributed training framework that combines algorithmic redundancy with robust aggregation. DETOX operates in two steps, a filtering step that uses limited redundancy to significantly reduce the effect of Byzantine nodes, and a hierarchical aggregation step that can be used in tandem with any state-of-the-art robust aggregation method. We show theoretically that this leads to a substantial increase in robustness, and has a per iteration runtime that can be nearly linear in the number of compute nodes. We provide extensive experiments over real distributed setups across a variety of large-scale machine learning tasks, showing that DETOX leads to orders of magnitude accuracy and speedup improvements over many state-of-the-art Byzantine-resilient approaches.