Machine Learning on Volatile Instances: Convergence, Runtime, and Cost Tradeoffs

Machine Learning on Volatile Instances: Convergence, Runtime, and Cost Tradeoffs
复制标题

DOI:
10.1109/tnet.2021.3112082
复制
发表时间:
2022-02
期刊:
IEEE/ACM Transactions on Networking
影响因子:
--
通讯作者:
Xiaoxi Zhang;Jianyu Wang;Li-Feng Lee;T. Yang;Akansha Kalra;Gauri Joshi;Carlee Joe-Wong
Xiaoxi Zhang;Jianyu Wang;Li-Feng Lee;T. Yang;Akansha Kalra;Gauri Joshi;Carlee Joe-Wong
中科院分区:
其他
文献类型:
--
作者:
Xiaoxi Zhang;Jianyu Wang;Li-Feng Lee;T. Yang;Akansha Kalra;Gauri Joshi;Carlee Joe-Wong

文献摘要

相似文献

由于今天机器学习中使用的神经网络模型和训练数据集的庞大规模,必须通过在多个工作节点上划分梯度评估等任务来分布随机梯度下降(SGD)。然而,运行分布式SGD可能会非常昂贵,因为它可能需要专门的计算资源,如gpu,用于很长一段时间。我们提出了具有成本效益的策略来利用易变的云实例,这些实例比标准实例便宜,但可能被更高优先级的工作负载中断。据我们所知,这项工作是第一次量化活动工作节点数量的变化(作为抢占的结果)如何影响SGD收敛和训练模型的时间。通过理解实例的抢占概率、准确性和训练时间之间的权衡,我们能够获得在易变实例(如Amazon EC2 spot实例和其他可抢占的云实例)上配置分布式SGD作业的实用策略。实验结果表明,我们的策略以较低的成本获得了良好的训练效果。
Due to the massive size of the neural network models and training datasets used in machine learning today, it is imperative to distribute stochastic gradient descent (SGD) by splitting up tasks such as gradient evaluation across multiple worker nodes. However, running distributed SGD can be prohibitively expensive because it may require specialized computing resources such as GPUs for extended periods of time. We propose cost-effective strategies to exploit volatile cloud instances that are cheaper than standard instances, but may be interrupted by higher priority workloads. To the best of our knowledge, this work is the first to quantify how variations in the number of active worker nodes (as a result of preemption) affect SGD convergence and the time to train the model. By understanding these trade-offs between preemption probability of the instances, accuracy, and training time, we are able to derive practical strategies for configuring distributed SGD jobs on volatile instances such as Amazon EC2 spot instances and other preemptible cloud instances. Experimental results show that our strategies achieve good training performance at substantially lower cost.