Sia: Heterogeneity-aware, goodput-optimized ML-cluster scheduling

Sia: Heterogeneity-aware, goodput-optimized ML-cluster scheduling
复制标题

DOI:
10.1145/3600006.3613175
复制
发表时间:
2023-10
期刊:
Proceedings of the 29th Symposium on Operating Systems Principles
影响因子:
--
通讯作者:
Suhas Jayaram Subramanya;Daiyaan Arfeen;Shouxu Lin;Aurick Qiao;Zhihao Jia;G. Ganger
Suhas Jayaram Subramanya;Daiyaan Arfeen;Shouxu Lin;Aurick Qiao;Zhihao Jia;G. Ganger
中科院分区:
其他
文献类型:
--
作者:
Suhas Jayaram Subramanya;Daiyaan Arfeen;Shouxu Lin;Aurick Qiao;Zhihao Jia;G. Ganger

文献摘要

相似文献

Sia调度器可以有效地将异构深度学习(DL)集群资源分配给弹性资源自适应作业。尽管一些最近的出版物解决了一个方面或另一个方面(例如,异构性或资源自适应性),没有解决所有问题,并且即使在没有组合调度问题的全部复杂性的情况下,大多数问题也不能很好地扩展到大集群和/或重工作负载。Sia引入了一种新的调度公式,可以扩展到搜索空间大小,并有意将作业及其配置与GPU类型和数量相匹配,同时适应集群负载和作业组合随时间的变化。Sia还引入了一种低配置开销的方法来引导(针对每个新作业)用于评估可能的资源分配的吞吐量模型,并且它是第一个支持混合并行作业弹性扩展的集群调度器。广泛的评估表明,Sia的性能优于最先进的制造商。例如,即使在相对较小的44到64 GPU集群上,混合使用三种GPU类型,Sia也可以将平均作业完成时间(JCT)减少30- 93%,将第99百分位JCT和完工时间减少28- 95%,并将来自3个真实环境的工作负载的GPU使用时间减少12- 55%。额外的实验表明,Sia可扩展到至少2000个GPU的集群,提供了更好的公平性,并且对调度器参数设置不太敏感。
The Sia scheduler efficiently assigns heterogeneous deep learning (DL) cluster resources to elastic resource-adaptive jobs. Although some recent schedulers address one aspect or another (e.g., heterogeneity or resource-adaptivity), none addresses all and most scale poorly to large clusters and/or heavy workloads even without the full complexity of the combined scheduling problem. Sia introduces a new scheduling formulation that can scale to the search-space sizes and intentionally match jobs and their configurations to GPU types and counts, while adapting to changes in cluster load and job mix over time. Sia also introduces a low-profiling-overhead approach to bootstrapping (for each new job) throughput models used to evaluate possible resource assignments, and it is the first cluster scheduler to support elastic scaling of hybrid parallel jobs. Extensive evaluations show that Sia outperforms state-of-the-art schedulers. For example, even on relatively small 44- to 64-GPU clusters with a mix of three GPU types, Sia reduces average job completion time (JCT) by 30--93%, 99th percentile JCT and makespan by 28--95%, and GPU hours used by 12--55% for workloads derived from 3 real-world environments. Additional experiments demonstrate that Sia scales to at least 2000-GPU clusters, provides improved fairness, and is not over-sensitive to scheduler parameter settings.