SplitRPC: A {Control + Data} Path Splitting RPC Stack for ML Inference Serving

SplitRPC: A {Control + Data} Path Splitting RPC Stack for ML Inference Serving
复制标题

DOI:
10.1145/3589974
复制
发表时间:
2023-05
期刊:
Proceedings of the ACM on Measurement and Analysis of Computing Systems
影响因子:
--
通讯作者:
Adithya Kumar;A. Sivasubramaniam;T. Zhu
Adithya Kumar;A. Sivasubramaniam;T. Zhu
中科院分区:
其他
文献类型:
--
作者:
Adithya Kumar;A. Sivasubramaniam;T. Zhu

文献摘要

被引文献

相似文献

智能编译器和运行时系统驱动的硬件加速器的采用越来越多,使ML服务民主化,并大幅减少了其执行时间。这促使我们将注意力转移到分布式设置下的这些ML服务,并在将加速器提供服务时由RPC机制(“ RPC税”)施加的间接费用。多年来设计的RPC实施是隐式假设主机CPU服务的请求,我们专注于将这些作品扩展到基于加速器的服务。尽管最新提出要求智能执行此任务的建议对于简单的内核是合理的,但是服务复杂的ML模型需要更细微的视图来优化数据路径和控制/编排这些加速器。我们对当今的商品网络接口卡(NIC)进行了编程,以将控制和数据路径分开,以有效地传输控制,同时有效地将有效载荷传输到加速器。与将这些路径捆绑在一起的统一方法相反,我们设计和实施了SplitRPC的灵活性 - 一种控制 +数据路径优化了用于ML推理服务的RPC机制。 SplitRPC允许我们在加速器上优化数据a,同时允许CPU保持完整的编排功能。我们在商品和智能机构上实现了SplitRPC,并演示了运行不同编译器/运行时系统的基于GPU的ML服务如何受益。对于使用不同的推理运行时间服务的各种ML模型,我们证明了SplitRPC可以有效地最大程度地减少RPC税,同时在不需要昂贵的智能设备的情况下,在现有内核旁路方法方面提供了吞吐量和潜伏期的显着增长。
The growing adoption of hardware accelerators driven by their intelligent compiler and runtime system counterparts has democratized ML services and precipitously reduced their execution times. This motivates us to shift our attention to efficiently serve these ML services under distributed settings and characterize the overheads imposed by the RPC mechanism ('RPC tax') when serving them on accelerators. The RPC implementations designed over the years implicitly assume the host CPU services the requests, and we focus on expanding such works towards accelerator-based services. While recent proposals calling for SmartNICs to take on this task are reasonable for simple kernels, serving complex ML models requires a more nuanced view to optimize both the data-path and the control/orchestration of these accelerators. We program today's commodity network interface cards (NICs) to split the control and data paths for effective transfer of control while efficiently transferring the payload to the accelerator. As opposed to unified approaches that bundle these paths together, limiting the flexibility in each of these paths, we design and implement SplitRPC - a control + data path optimizing RPC mechanism for ML inference serving. SplitRPC allows us to optimize the datapath to the accelerator while simultaneously allowing the CPU to maintain full orchestration capabilities. We implement SplitRPC on both commodity NICs and SmartNICs and demonstrate how GPU-based ML services running different compiler/runtime systems can benefit. For a variety of ML models served using different inference runtimes, we demonstrate that SplitRPC is effective in minimizing the RPC tax while providing significant gains in throughput and latency over existing kernel by-pass approaches, without requiring expensive SmartNIC devices.