Horus: Interference-Aware and Prediction-Based Scheduling in Deep Learning Systems

Horus: Interference-Aware and Prediction-Based Scheduling in Deep Learning Systems
复制标题

DOI:
10.1109/tpds.2021.3079202
复制
发表时间:
2022-01
影响因子:
5.3
通讯作者:
Gingfung Yeung;Damian Borowiec;Renyu Yang;A. Friday;R. Harper;Peter Garraghan
Gingfung Yeung;Damian Borowiec;Renyu Yang;A. Friday;R. Harper;Peter Garraghan
中科院分区:
计算机科学2区
文献类型:
--
作者:
Gingfung Yeung;Damian Borowiec;Renyu Yang;A. Friday;R. Harper;Peter Garraghan

文献摘要

被引文献

相似文献

为了加速深度学习(DL)模型的训练,利用配备硬件加速器(如GPU)的机器集群来减少执行时间。需要最先进的资源管理器来提高GPU利用率并最大化吞吐量。虽然在同一GPU上共同定位DL作业已被证明是有效的,但这可能会导致干扰,从而导致减速。在这篇文章中,我们提出了Horus:一个干扰感知和基于预测的DL系统资源管理器。Horus主动预测异构DL作业的GPU利用率,从DL模型的计算图功能推断,消除了在线分析和隔离保留GPU的需要。通过微基准测试和跨异构GPU硬件的作业协同定位组合,我们将GPU利用率确定为通用代理指标,以确定良好的布局决策,与当前的方法相比,这些方法保留隔离的GPU来执行在线分析,并直接测量每个独特的提交作业的GPU利用率。我们的方法促进了高资源利用率和最大完工时间的减少;通过真实世界的实验和大规模跟踪驱动的模拟,我们证明了Horus在GPU资源利用率方面优于其他DL资源管理器高达61.5%,最大完工时间减少23.7- 30.7%,作业等待时间减少68.3%。
To accelerate the training of Deep Learning (DL) models, clusters of machines equipped with hardware accelerators such as GPUs are leveraged to reduce execution time. State-of-the-art resource managers are needed to increase GPU utilization and maximize throughput. While co-locating DL jobs on the same GPU has been shown to be effective, this can incur interference causing slowdown. In this article we propose Horus: an interference-aware and prediction-based resource manager for DL systems. Horus proactively predicts GPU utilization of heterogeneous DL jobs extrapolated from the DL model’s computation graph features, removing the need for online profiling and isolated reserved GPUs. Through micro-benchmarks and job co-location combinations across heterogeneous GPU hardware, we identify GPU utilization as a general proxy metric to determine good placement decisions, in contrast to current approaches which reserve isolated GPUs to perform online profiling and directly measure GPU utilization for each unique submitted job. Our approach promotes high resource utilization and makespan reduction; via real-world experimentation and large-scale trace driven simulation, we demonstrate that Horus outperforms other DL resource managers by up to 61.5 percent for GPU resource utilization, 23.7–30.7 percent for makespan reduction and 68.3 percent in job wait time reduction.