Grus: Enabling Latency SLOs for GPU-Accelerated NFV Systems

Grus: Enabling Latency SLOs for GPU-Accelerated NFV Systems
复制标题

DOI:
10.1109/icnp.2018.00025
复制
发表时间:
2018-09
期刊:
2018 IEEE 26th International Conference on Network Protocols (ICNP)
影响因子:
--
通讯作者:
Zhilong Zheng;J. Bi;Haiping Wang;Chen Sun;Heng Yu;Hongxin Hu;K. Gao;Jianping Wu
Zhilong Zheng;J. Bi;Haiping Wang;Chen Sun;Heng Yu;Hongxin Hu;K. Gao;Jianping Wu
中科院分区:
其他
文献类型:
--
作者:
Zhilong Zheng;J. Bi;Haiping Wang;Chen Sun;Heng Yu;Hongxin Hu;K. Gao;Jianping Wu

文献摘要

被引文献

相似文献

图形处理单元(GPU)最近已被开发作为硬件加速器,以提高网络功能虚拟化(NFV)的性能。然而,当多个网络功能(NF)共同位于同一机器中时,GPU加速的NFV系统遭受显著的延迟变化,这阻止了运营商支持延迟服务水平目标(SLO)。现有的研究努力,以解决这个问题,只能保证有限数量的SLO的资源利用效率非常低。在本文中,我们提出了Grus框架,以支持GPU加速的NFV系统中的延迟SLO。Grus深入分析了延迟变化的来源,并提出了三个设计原则:(1)需要动态设置批量大小以限制CPU中的数据包延迟;(2)需要通过PCI-E进行数据传输的重新排序机制以保证停滞时间;(3)需要最大化GPU中的并发性以避免NF执行等待时间。在这些原则的指导下,Grus由两个逻辑层组成,包括基础设施层和调度层。基础架构层配备了一个CPU内可重新排序的工作池,可以调整队列大小和数据包传输顺序,以及GPU内可控并发执行器,以提供最大化的并发性。调度层运行启发式算法来执行准确和快速的调度,以保证基于我们的预测模型的SLO。我们已经实现了Grus的原型。广泛的评估表明,Grus可以显着减少延迟变化,并满足比最先进的解决方案多4.5倍的SLO条款。
Graphics Processing Unit (GPU) has been recently exploited as a hardware accelerator to improve the performance of Network Function Virtualization (NFV). However, GPU-accelerated NFV systems suffer from significant latency variation when multiple network functions (NFs) are co-located in the same machine, which prevents operators from supporting latency Service Level Objectives (SLOs). Existing research efforts to address this problem can only guarantee a limited number of SLOs with very low resource utilization efficiency. In this paper, we present the Grus framework to support latency SLOs in GPU-accelerated NFV systems. Grus thoroughly analyzes the sources of latency variation and proposes three design principles: (1) dynamic batch size setting is needed to bound packet batching latency in CPU; (2) a reordering mechanism for data transfer over PCI-E is required to guarantee the stalling time; and (3) maximizing concurrency in GPU is necessary to avoid NF execution waiting time. Guided by the principles, Grus consists of two logical layers including an infrastructure layer and a scheduling layer. The infrastructure layer is equipped with an in-CPU Reorder-able Worker Pool that could adjust batching size and packet transfer order, and in-GPU Controllable Concurrent Executors to provide maximized concurrency. The scheduling layer runs a heuristic algorithm to perform accurate and fast scheduling to guarantee SLOs based on our prediction models. We have implemented a prototype of Grus. Extensive evaluations demonstrate that Grus can significantly reduce latency variation and satisfy 4.5× more SLO terms than state-of-the-art solutions.