Defending Against Saddle Point Attack in Byzantine-Robust Distributed Learning

Defending Against Saddle Point Attack in Byzantine-Robust Distributed Learning
复制标题

DOI:
--
复制
发表时间:
2018-06
期刊:
--
影响因子:
--
通讯作者:
Dong Yin;Yudong Chen;K. Ramchandran;P. Bartlett
Dong Yin;Yudong Chen;K. Ramchandran;P. Bartlett
中科院分区:
其他
文献类型:
--
作者:
Dong Yin;Yudong Chen;K. Ramchandran;P. Bartlett

文献摘要

被引文献

相似文献

我们研究了鲁棒的分布式学习,涉及最小化非凸损失函数的鞍点。我们考虑拜占庭设置,其中一些工人机器具有异常甚至任意和对抗性行为。在这种情况下,拜占庭机器可能会在远离任何真正的局部最小值的鞍点附近创建假的局部最小值,即使使用了鲁棒的梯度估计器。我们开发ByzantinePGD,一个强大的一阶算法,可以证明逃脱鞍点和假的局部极小值,并收敛到一个近似的真正的局部极小,具有低迭代复杂度。作为副产品,我们给出了一个更简单的算法和分析,在通常的非拜占庭设置逃逸鞍点。我们进一步讨论了三个强大的梯度估计,可用于ByzantinePGD,包括中位数,修剪均值,迭代滤波。我们在具体的统计环境中的性能特点,并认为他们的近最优的低维和高维制度。
We study robust distributed learning that involves minimizing a non-convex loss function with saddle points. We consider the Byzantine setting where some worker machines have abnormal or even arbitrary and adversarial behavior. In this setting, the Byzantine machines may create fake local minima near a saddle point that is far away from any true local minimum, even when robust gradient estimators are used. We develop ByzantinePGD, a robust first-order algorithm that can provably escape saddle points and fake local minima, and converge to an approximate true local minimizer with low iteration complexity. As a by-product, we give a simpler algorithm and analysis for escaping saddle points in the usual non-Byzantine setting. We further discuss three robust gradient estimators that can be used in ByzantinePGD, including median, trimmed mean, and iterative filtering. We characterize their performance in concrete statistical settings, and argue for their near-optimality in low and high dimensional regimes.