HIOS: Hierarchical Inter-Operator Scheduler for Real-Time Inference of DAG-Structured Deep Learning Models on Multiple GPUs

HIOS: Hierarchical Inter-Operator Scheduler for Real-Time Inference of DAG-Structured Deep Learning Models on Multiple GPUs
复制标题

DOI:
10.1109/cluster52292.2023.00016
复制
发表时间:
2023-10
期刊:
2023 IEEE International Conference on Cluster Computing (CLUSTER)
影响因子:
--
通讯作者:
Turja Kundu;Tong Shu
Turja Kundu;Tong Shu
中科院分区:
其他
文献类型:
--
作者:
Turja Kundu;Tong Shu

文献摘要

相似文献

实时科学应用中支持神经网络的数据分析对推理延迟提出了严格的要求。同时,最近的深度学习(DL)模型设计趋势是用多个分支取代单个分支,以实现高预测精度和鲁棒性,这使得算子间并行化成为改善推理延迟的有效方法。然而,现有的用于推理加速的算子间并行化技术主要集中在单个 GPU 的利用率优化上。随着输入样本的数据量和深度学习模型的规模不断增长,单个GPU的有限资源不足以支持大型算子的并行执行。为了打破这个限制,我们研究了多个 GPU 之间以及每个 GPU 内的混合操作间并行性。在本文中,我们设计并实现了一种分层操作符间调度器(HIOS),可以自动将大型操作符分配到不同的 GPU 上,并将小型操作符分组在同一 GPU 中以并行执行。特别是,我们提出了一种新颖的调度算法,名为 HIOS-LP,它包括通过迭代最长路径(LP)映射的 GPU 算子并行化和基于滑动窗口的 GPU 算子并行化。除了广泛的模拟结果之外,现代卷积神经网络基准测试的实验表明,我们的 HIOS-LP 在实际系统中的性能比最先进的操作员间调度算法 IOS 高出 17%。
Neural-network-enabled data analysis in real-time scientific applications imposes stringent requirements on inference latency. Meanwhile, recent deep learning (DL) model design trends to replace a single branch with multiple branches for high prediction accuracy and robustness, which makes inter-operator parallelization become an effective approach to improve inference latency. However, existing inter-operator parallelization techniques for inference acceleration are mainly focused on utilization optimization in a single GPU. With the data size of an input sample and the scale of a DL model ever-growing, the limited resource of a single GPU is insufficient to support the parallel execution of large operators. In order to break this limitation, we study hybrid inter-operator parallelism both among multiple GPUs and in each GPU. In this paper, we design and implement a hierarchical inter-operator scheduler (HIOS) to automatically distribute large operators onto different GPUs and group small operators in the same GPU for parallel execution. Particularly, we propose a novel scheduling algorithm, named HIOS-LP, which consists of inter-GPU operator parallelization through iterative longest-path (LP) mapping and intra-GPU operator parallelization based on a sliding window. In addition to extensive simulation results, experiments with modern convolutional neural network benchmarks demonstrate that our HIOS-LP outperforms the state-of-the-art inter-operator scheduling algorithm IOS by up to 17% in real systems.