COLTI: Towards Concurrent and Co-located DNN Training and Inference

COLTI: Towards Concurrent and Co-located DNN Training and Inference
复制标题

COLTI:迈向并发和同地 DNN 训练和推理

DOI:
10.1145/3588195.3595940
复制
发表时间:
2023
期刊:
ACM
影响因子:
--
通讯作者:
Rafique, M. Mustafa
Rafique, M. Mustafa
中科院分区:
--
文献类型:
--
作者:
Mobin, Jaiaid;Maurya, Avinash;Rafique, M. Mustafa

文献摘要

参考文献

相似文献

深度学习模型广泛用于各种领域,例如,科学模拟、预测和建模。然而,训练这些密集网络是计算和内存密集型的,通常需要图形处理单元(GPU)等加速器。虽然这种DNN工作负载消耗了有限的板载高带宽内存(HBM)的大部分,但它们通常未充分利用GPU计算资源。在这种情况下,可以利用GPU上的空闲计算资源来运行挂起的作业,这些作业可以(1)容纳在剩余的HBM上,或者(2)可以与其他并发工作负载共享存储器资源。然而,最先进的工作负载调度器和DNN运行时并没有设计用于利用HBM托管来提高资源利用率和吞吐量。在这项工作中,我们提出了CORTI,它引入了一套新技术,通过在内存受限的GPU设备上协同定位DNN训练和推理来解决上述挑战。我们对PyTorch框架中实现的三种不同DNN模型的初步评估表明,完工时间和内存利用率分别提高了37%和40%。
Deep learning models are extensively used in a wide range of domains, e.g., scientific simulations, predictions, and modeling. However, training these dense networks is both compute and memory intensive, and typically requires accelerators such as Graphics Processing Units (GPUs). While such DNN workloads consume a major proportion of the limited onboard high-bandwidth memory (HBM), they typically underutilize the GPU compute resources. In such scenarios, the idle compute resources on the GPU can be leveraged to run pending jobs that can either be (1) accommodated on the remainder HBM, or (2) can share memory resources with other concurrent workloads. However, state-of-the-art workload schedulers and DNN runtimes are not designed to leverage HBM co-location to improve resource utilization and throughput. In this work, we propose COLTI, which introduces a set of novel techniques to solve the aforementioned challenges by co-locating DNN training and inference on memory-constrained GPU devices. Our preliminary evaluations of three different DNN models implemented in the PyTorch framework demonstrate up to 37% and 40% improvement in makespan and memory utilization, respectively.
DOI: 10.1145/3419111.3421284
发表时间: 2020-10
期刊: Proceedings of the 11th ACM Symposium on Cloud Computing
影响因子: --
作者:
Aditya Dhakal;Sameer G. Kulkarni;K. Ramakrishnan
通讯作者: Aditya Dhakal;Sameer G. Kulkarni;K. Ramakrishnan