Cerebros: Evading the RPC Tax in Datacenters

Cerebros: Evading the RPC Tax in Datacenters
复制标题

Cerebros:逃避数据中心的 RPC 税

DOI:
10.1145/3466752.3480055
复制
发表时间:
2021
期刊:
MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture
影响因子:
--
通讯作者:
Falsafi, Babak
Falsafi, Babak
中科院分区:
--
文献类型:
--
作者:
Pourhabibi, Arash;Sutherland, Mark;Daglis, Alexandros;Falsafi, Babak

文献摘要

参考文献

被引文献

相似文献

新兴的微服务范式将在线服务分解为细粒度的软件模块,经常使用远程过程调用(RPC)在数据中心网络上进行通信。网络堆栈的持续发展暴露了RPC层本身是一个瓶颈,我们发现它占微服务总执行周期的40%-90%。我们分解了构成生产RPC层的底层模块,并基于先前的证据证明,CPU只能对此类任务进行有限的改进,要求转向硬件以消除RPC层作为微服务性能的限制因素。尽管最近提出的加速器可以有效地处理RPC层的一部分,但它们的总体好处受到不必要的CPU参与的限制,这是因为加速器被设计为CPU控制下的协处理器。相反,我们展示了最终消除RPC层瓶颈需要由连接NIC的硬件加速器执行RPC层的所有模块。我们介绍了Cerebros,这是一个专用的RPC处理器,它执行ApacheThrift RPC层,并充当NIC和运行在CPU上的微服务之间的中间阶段。我们使用DeathStarBitch微服务套件进行的评估显示,Cerebros将RPC层的CPU周期减少了37-×,每个微服务请求的总周期减少了1.8-14倍。
The emerging paradigm of microservices decomposes online services into fine-grained software modules frequently communicating over the datacenter network, often using Remote Procedure Calls (RPCs). Ongoing advancements in the network stack have exposed the RPC layer itself as a bottleneck, that we show accounts for 40–90% of a microservice’s total execution cycles. We break down the underlying modules that comprise production RPC layers and demonstrate, based on prior evidence, that CPUs can only expect limited improvements for such tasks, mandating a shift to hardware to remove the RPC layer as a limiter of microservice performance. Although recently proposed accelerators can efficiently handle a portion of the RPC layer, their overall benefit is limited by unnecessary CPU involvement, which occurs because the accelerators are architected as co-processors under the CPU’s control. Instead, we show that conclusively removing the RPC layer bottleneck requires all of the RPC layer’s modules to be executed by a NIC-attached hardware accelerator. We introduce Cerebros, a dedicated RPC processor that executes the Apache Thrift RPC layer and acts as an intermediary stage between the NIC and the microservice running on the CPU. Our evaluation using the DeathStarBench microservice suite shows that Cerebros reduces the CPU cycles spent in the RPC layer by 37–64 ×, yielding a 1.8–14 × reduction in total cycles expended per microservice request.
Optimus Prime:加速服务器中的数据转换
DOI: 10.1145/3373376.3378501
发表时间: 2020
期刊: Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子: --
作者:
Arash Pourhabibi Zarandi;Siddharth Gupta;H. Kassir;Mark Sutherland;Zilu Tian;M. Drumond;B. Falsafi;Christoph E. Koch
通讯作者: Christoph E. Koch
用于内存机架规模计算的众核网络接口
DOI: 10.1145/2749469.2750415
发表时间: 2015
期刊: 2015 ACM/IEEE 42nd Annual International Symposium on Computer Architecture (ISCA)
影响因子: --
作者:
Alexandros Daglis;Stanko Novakovic;Edouard Bugnion;B. Falsafi;Boris Grot
通讯作者: Boris Grot
Boomerang:用于控制流交付的无元数据架构
DOI: 10.1109/hpca.2017.53
发表时间: 2017
期刊: 2017 IEEE International Symposium on High Performance Computer Architecture (HPCA)
影响因子: --
作者:
Rakesh Kumar;Cheng;Boris Grot;V. Nagarajan
通讯作者: V. Nagarajan
DOI: 10.1145/3173162.3173178
发表时间: 2018-03
期刊: Proceedings of the Twenty-Third International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子: --
作者:
Rakesh Kumar;Boris Grot;V. Nagarajan
通讯作者: Rakesh Kumar;Boris Grot;V. Nagarajan
Dagger:具有近内存可重构网卡的云微服务中高效快速的 RPC
DOI: 10.1145/3445814.3446696
发表时间: 2021
期刊: 26th ACM Interna-tional Conference on Architectural Support for Programming Languages andOperating Systems (ASPLOS ’21
影响因子: --
作者:
Lazarev, Nikita;Xiang, Shaojie;Adit, Neil;Zhang, Zhiru;Delimitrou, Christina
通讯作者: Delimitrou, Christina