Understanding Synchronization Costs for Distributed ML on Transient Cloud Resources

Understanding Synchronization Costs for Distributed ML on Transient Cloud Resources
复制标题

DOI:
10.1109/ic2e.2019.00029
复制
发表时间:
2019-06
期刊:
2019 IEEE International Conference on Cloud Engineering (IC2E)
影响因子:
--
通讯作者:
Pradeep Ambati;David E. Irwin;Prashant J. Shenoy;Lixin Gao;Ahmed Ali-Eldin;Jeannie R. Albrecht
Pradeep Ambati;David E. Irwin;Prashant J. Shenoy;Lixin Gao;Ahmed Ali-Eldin;Jeannie R. Albrecht
中科院分区:
其他
文献类型:
--
作者:
Pradeep Ambati;David E. Irwin;Prashant J. Shenoy;Lixin Gao;Ahmed Ali-Eldin;Jeannie R. Albrecht

文献摘要

被引文献

相似文献

云平台经常执行并行批处理应用程序,比如分布式机器学习(ML),其中包含大量的同步屏障。这些屏障会阻止任何任务在所有任务都到达某一指定点之前继续前进,这会使应用程序的性能降低到最慢的“掉队”任务的性能水平,从而显著降低应用程序的性能。为了解决这个问题,研究人员提出了许多缓解掉队任务的技术,包括推测性地重新执行掉队任务以及对严格的屏障语义进行各种放宽。虽然这些技术提高了并行应用程序的性能,但它们在重新执行任务或等待所浪费的资源方面产生了成本。重要的是,这些成本在针对专用资源的先前研究中往往是隐含的,但在云环境中则变得明确,因为云会按精细的时间间隔对资源收费。此外,在云平台中,不同技术之间的成本差异会加剧,因为云对临时资源的收费要低得多,而这些临时资源在很大范围内实际上会产生一种概率性的性能。虽然临时资源的低标价很有吸引力,但资源回收会增加掉队任务的频率和严重程度,这会降低并行作业的性能并增加总体执行成本。为了更好地理解同步成本,我们针对不同的缓解掉队任务的技术开发了简单的分析模型,并比较了它们在按需资源和临时资源上的成本和性能。我们的分析表明:i)与按需服务器相比,临时服务器提供了复杂的权衡,并且由于其概率性性能,尽管价格有很大折扣,但可能会导致更高的总体成本;ii)缓解掉队任务这一经过充分研究的常见方法,在使用会导致频繁且严重的掉队任务的临时服务器时效果较差;iii)一种最近的灵活同步方法提供了最佳的成本和性能。
Cloud platforms often execute parallel batch applications, such as distributed machine learning (ML), that include numerous synchronization barriers. These barriers, which prevent any task from advancing beyond a specified point until all tasks have reached that point, significantly degrade application performance by reducing it to that of the slowest "straggler" task. To address the problem, researchers have proposed numerous straggler mitigation techniques, including speculatively re-executing straggler tasks and various relaxations of strict barrier semantics. While these techniques improve parallel application performance, they incur a cost in terms of the resources wasted re-executing tasks or waiting. Importantly, these costs, which are often implicit in prior work that targets dedicated resources, become explicit in the cloud, which charges for resources at fine-grained intervals. In addition, the cost difference between techniques is exacerbated in cloud platforms, since they charge substantially less for transient resources that effectively yield a probabilistic performance across a wide range. While transient resources' low list price is attractive, revocations increase the frequency and severity of stragglers, which decreases parallel job performance and increases overall execution cost. To better understand the cost of synchronization, we develop simple analytical models of different straggler mitigation techniques and compare their cost and performance on on-demand and transient resources. Our analysis shows that i) transient servers offer complex tradeoffs compared to on-demand servers, and can result in higher overall costs despite their highly discounted price due to their probabilistic performance; ii) common approaches to straggler mitigation, which is a well-studied problem, are less effective using transient servers that cause frequent and severe stragglers; and iii) a recent approach to flexible synchronization offers the best cost and performance.