COLTI: Towards Concurrent and Co-located DNN Training and Inference
COLTI: Towards Concurrent and Co-located DNN Training and Inference
复制标题
COLTI:迈向并发和同地 DNN 训练和推理
DOI:
10.1145/3588195.3595940
复制
发表时间:
2023
期刊:
影响因子:
--
通讯作者:
Rafique, M. Mustafa
中科院分区:
文献类型:
--
作者:
Mobin, Jaiaid;Maurya, Avinash;Rafique, M. Mustafa
Deep learning models are extensively used in a wide range of domains, e.g., scientific simulations, predictions, and modeling. However, training these dense networks is both compute and memory intensive, and typically requires accelerators such as Graphics Processing Units (GPUs). While such DNN workloads consume a major proportion of the limited onboard high-bandwidth memory (HBM), they typically underutilize the GPU compute resources. In such scenarios, the idle compute resources on the GPU can be leveraged to run pending jobs that can either be (1) accommodated on the remainder HBM, or (2) can share memory resources with other concurrent workloads. However, state-of-the-art workload schedulers and DNN runtimes are not designed to leverage HBM co-location to improve resource utilization and throughput. In this work, we propose COLTI, which introduces a set of novel techniques to solve the aforementioned challenges by co-locating DNN training and inference on memory-constrained GPU devices. Our preliminary evaluations of three different DNN models implemented in the PyTorch framework demonstrate up to 37% and 40% improvement in makespan and memory utilization, respectively.
DOI:
10.1145/3419111.3421284
发表时间:
2020-10
期刊:
Proceedings of the 11th ACM Symposium on Cloud Computing
影响因子:
--
作者:
Aditya Dhakal;Sameer G. Kulkarni;K. Ramakrishnan
通讯作者:
Aditya Dhakal;Sameer G. Kulkarni;K. Ramakrishnan