Aequitas: Admission Control for Performance-Critical RPCs in Datacenters

Aequitas: Admission Control for Performance-Critical RPCs in Datacenters
复制标题

Aequitas:数据中心中性能关键型 RPC 的准入控制

DOI:
10.1145/3544216.3544271
复制
发表时间:
2022
期刊:
ACM SIGCOMM
影响因子:
--
通讯作者:
Vahdat, Amin
Vahdat, Amin
中科院分区:
--
文献类型:
--
作者:
Zhang, Yiwen;Kumar, Gautam;Dukkipati, Nandita;Wu, Xian;Jha, Priyaranjan;Chowdhury, Mosharaf;Vahdat, Amin

文献摘要

相似文献

随着分散存储和微服务架构的日益普及,高扇出和扇入远程过程调用(RPC)现在产生了现代数据中心中的大部分流量。虽然网络在RPC性能中起着至关重要的作用,但由于RPC特性的广泛变化,传统的流量分类类别无法充分捕获其重要性。因此,满足服务水平目标(SLO),特别是对性能关键(PC)的RPC,仍然具有挑战性。我们提出Aequitas,一个分布式的发送端驱动的准入控制方案,使用商品的加权公平竞争(WFQ),以保证RPC级的SLO。在存在网络过载的情况下,它通过限制允许进入任何给定QoS的流量并降低其余流量的级别来强制执行群集范围的RPC延迟SLO。我们的分析和经验表明,这个简单的计划运作良好。当网络需求超过规定容量时,Aequitas实现的延迟SLO比第99.9页的最新拥塞控制低3.8倍,并且与pFabric、Qjump、D3、PDQ和Homa相比,允许多达2倍多的PCRPC满足SLO。我们的整个生产部署的结果显示延迟改善了10%。
With the increasing popularity of disaggregated storage and microservice architectures, high fan-out and fan-in Remote Procedure Calls (RPCs) now generate most of the traffic in modern datacenters. While the network plays a crucial role in RPC performance, traditional traffic classification categories cannot sufficiently capture their importance due to wide variations in RPC characteristics. As a result, meeting service-level objectives (SLOs), especially for performance-critical (PC) RPCs, remains challenging.We present Aequitas, a distributed sender-driven admission control scheme that uses commodity Weighted-Fair Queuing (WFQ) to guarantee RPC-level SLOs. In the presence of network overloads, it enforces cluster-wide RPC latency SLOs by limiting the amount of traffic admitted into any given QoS and downgrading the rest. We show analytically and empirically that this simple scheme works well. When the network demand spikes beyond provisioned capacity, Aequitas achieves a latency SLO that is 3.8× lower than the state-of-art congestion control at the 99.9th-pand admits up to 2× morePCRPCs meeting SLO when compared with pFabric, Qjump, D3, PDQ, and Homa. Results in our fleetwide production deployment show a 10% latency improvement.