GSLICE: controlled spatial sharing of GPUs for a scalable inference platform

GSLICE: controlled spatial sharing of GPUs for a scalable inference platform
复制标题

DOI:
10.1145/3419111.3421284
复制
发表时间:
2020-10
期刊:
Proceedings of the 11th ACM Symposium on Cloud Computing
影响因子:
--
通讯作者:
Aditya Dhakal;Sameer G. Kulkarni;K. Ramakrishnan
Aditya Dhakal;Sameer G. Kulkarni;K. Ramakrishnan
中科院分区:
其他
文献类型:
--
作者:
Aditya Dhakal;Sameer G. Kulkarni;K. Ramakrishnan

文献摘要

被引文献

相似文献

对基于云的推理服务的需求不断增加,需要使用图形处理单元(GPU)。 )帮助通过将动态的GPU资源分配和管理框架纳入挑战,以最大程度地提高性能和资源利用率。资源分配和批处理方案可以说明网络流量特征,同时还将推理潜伏期保持在服务水平以下GSLICE。 100μs)与默认的MP和Tensorrt相比,GSLICE提高了GPU的利用效率60--800% 2--13倍的骨料吞吐量改善。
The increasing demand for cloud-based inference services requires the use of Graphics Processing Unit (GPU). It is highly desirable to utilize GPU efficiently by multiplexing different inference tasks on the GPU. Batched processing, CUDA streams and Multi-process-service (MPS) help. However, we find that these are not adequate for achieving scalability by efficiently utilizing GPUs, and do not guarantee predictable performance. GSLICE addresses these challenges by incorporating a dynamic GPU resource allocation and management framework to maximize performance and resource utilization. We virtualize the GPU by apportioning the GPU resources across different Inference Functions (IFs), thus providing isolation and guaranteeing performance. We develop self-learning and adaptive GPU resource allocation and batching schemes that account for network traffic characteristics, while also keeping inference latencies below service level objectives. GSLICE adapts quickly to the streaming data's workload intensity and the variability of GPU processing costs. GSLICE provides scalability of the GPU for IF processing through efficient and controlled spatial multiplexing, coupled with a GPU resource re-allocation scheme with near-zero (< 100μs) downtime. Compared to default MPS and TensorRT, GSLICE improves GPU utilization efficiency by 60--800% and achieves 2--13X improvement in aggregate throughput.