ScaleServe: a scalable multi-GPU machine learning inference system and benchmarking suite

ScaleServe: a scalable multi-GPU machine learning inference system and benchmarking suite
复制标题

ScaleServe:可扩展的多 GPU 机器学习推理系统和基准测试套件

DOI:
10.1145/3530390.3532735
复制
发表时间:
2022
期刊:
Proceedings of the 14th Workshop on General Purpose Processing Using GPU
影响因子:
--
通讯作者:
Wong, Daniel
Wong, Daniel
中科院分区:
--
文献类型:
--
作者:
Jahanshahi, Ali;Chow, Marcus;Wong, Daniel

文献摘要

相似文献

我们提出了ScaleServe,这是一个可扩展的多GPU机器学习推理系统,它(1)构建在端到端的开源软件堆栈上,(2)与硬件供应商无关,(3)采用模块化组件设计,为用户提供轻松修改和扩展各种配置旋钮的能力。ScaleServe还提供了来自推理服务器不同层的详细性能指标,使设计人员能够准确定位瓶颈。我们通过在8-GPU服务器上执行计算机视觉和自然语言处理等多项机器学习任务,展示了ScaleServe的服务可伸缩性。ResNet152的性能结果表明,ScaleServe能够在多GPU平台上很好地扩展。
We present, ScaleServe, a scalable multi-GPU machine learning inference system that (1) is built on an end-to-end open-sourced software stack, (2) is hardware vendor-agnostic, and (3) is designed with modular components to provide users with ease to modify and extend various configuration knobs. ScaleServe also provides detailed performance metrics from different layers of the inference server which allow designers to pinpoint bottlenecks.We demonstrate ScaleServe's serving scalability with several machine learning tasks including computer vision and natural language processing on an 8-GPU server. The performance results for ResNet152 shows that ScaleServe is able to scale well on a multi-GPU platform.