DSAG: A Mixed Synchronous-Asynchronous Iterative Method for Straggler-Resilient Learning

DSAG: A Mixed Synchronous-Asynchronous Iterative Method for Straggler-Resilient Learning
复制标题

DOI:
10.1109/tcomm.2022.3227286
复制
发表时间:
2021-11
影响因子:
8.3
通讯作者:
A. Severinson;E. Rosnes;S. E. Rouayheb;A. G. Amat
A. Severinson;E. Rosnes;S. E. Rouayheb;A. G. Amat
中科院分区:
计算机科学2区
文献类型:
--
作者:
A. Severinson;E. Rosnes;S. E. Rouayheb;A. G. Amat

文献摘要

相似文献

我们考虑离散弹性学习。在许多以前的工作中,例如,在编码计算文献中,离散被建模为随机延迟,这些延迟在工作人员之间独立且相同地分布。然而,在许多实际情况下,一个特定的工人可能会在很长一段时间内挣扎。我们提出了一个延迟模型来捕捉这种行为,并通过在Microsoft Azure、Amazon Web Services (AWS)和一个小型本地集群上收集的痕迹来证实。在此基础上,提出了基于随机平均梯度法(SAG)的同步-异步混合迭代优化方法DSAG。我们还提出了一种动态负载平衡策略,以进一步减少掉队工人的影响。我们评估了DSAG的主成分分析,将其作为大型基因组数据集的有限和优化问题,以及在AWS上由100个工作人员组成的集群上进行逻辑回归,并发现DSAG比SAG快50%左右,比编码计算方法快两倍以上,对于我们考虑的特定场景。
We consider straggler-resilient learning. In many previous works, e.g., in the coded computing literature, straggling is modeled as random delays that are independent and identically distributed between workers. However, in many practical scenarios, a given worker may straggle over an extended period of time. We propose a latency model that captures this behavior and is substantiated by traces collected on Microsoft Azure, Amazon Web Services (AWS), and a small local cluster. Building on this model, we propose DSAG, a mixed synchronous-asynchronous iterative optimization method, based on the stochastic average gradient (SAG) method, that combines timely and stale results. We also propose a dynamic load-balancing strategy to further reduce the impact of straggling workers. We evaluate DSAG for principal component analysis, cast as a finite-sum optimization problem, of a large genomics dataset, and for logistic regression on a cluster composed of 100 workers on AWS, and find that DSAG is up to about 50% faster than SAG, and more than twice as fast as coded computing methods, for the particular scenario that we consider.