Dagger: efficient and fast RPCs in cloud microservices with near-memory reconfigurable NICs

Dagger: efficient and fast RPCs in cloud microservices with near-memory reconfigurable NICs
复制标题

Dagger:具有近内存可重构网卡的云微服务中高效快速的 RPC

DOI:
10.1145/3445814.3446696
复制
发表时间:
2021
期刊:
26th ACM Interna-tional Conference on Architectural Support for Programming Languages andOperating Systems (ASPLOS ’21
影响因子:
--
通讯作者:
Delimitrou, Christina
Delimitrou, Christina
中科院分区:
--
文献类型:
--
作者:
Lazarev, Nikita;Xiang, Shaojie;Adit, Neil;Zhang, Zhiru;Delimitrou, Christina

文献摘要

参考文献

被引文献

相似文献

云服务从整体设计到微服务的持续转变对高效和高性能的数据中心网络堆栈产生了很高的需求,这些堆栈针对细粒度的工作负载进行了优化。基于软件栈和外围设备的商用网络系统在传递小消息时会带来很高的开销。我们提出了Dagger,一个基于FPGA的云RPC硬件加速结构,其中加速器通过可配置的内存互连与主机处理器紧密耦合。Dagger的三个关键设计原则是:(1)将整个RPC堆栈卸载到基于FPGA的NIC,(2)利用内存互连而不是PCIe总线作为与主机CPU的接口,以及(3)使加速结构可重新配置,因此它可以适应微服务的各种需求。我们表明,这些原则的组合显着提高了云RPC系统的效率和性能,同时保持其通用性。与高度优化的软件栈和使用专用RDMA适配器的系统相比,Dagger实现了1.3 - 3.8倍的每核RPC吞吐量。它还可以在4个CPU内核上使用8个线程扩展到84 Mrps,同时保持最先进的µ s级尾部延迟。我们还展示了大型第三方应用程序,如memcached和云母KVS,可以很容易地移植到Dagger上,只需对其代码库进行最小的更改,将其中位数和尾部KVS访问延迟分别降低到2.8 - 3.5 us和5.4 - 7.8 us。最后,我们通过使用一个实现航班登机服务的8层应用程序对其进行评估,证明了Dagger对于具有不同线程模型的多层端到端微服务是有益的。
The ongoing shift of cloud services from monolithic designs to mi- croservices creates high demand for efficient and high performance datacenter networking stacks, optimized for fine-grained work- loads. Commodity networking systems based on software stacks and peripheral NICs introduce high overheads when it comes to delivering small messages. We present Dagger, a hardware acceleration fabric for cloud RPCs based on FPGAs, where the accelerator is closely-coupled with the host processor over a configurable memory interconnect. The three key design principle of Dagger are: (1) offloading the entire RPC stack to an FPGA-based NIC, (2) leveraging memory interconnects instead of PCIe buses as the interface with the host CPU, and (3) making the acceleration fabric reconfigurable, so it can accommodate the diverse needs of microservices. We show that the combination of these principles significantly improves the efficiency and performance of cloud RPC systems while preserving their generality. Dagger achieves 1.3 − 3.8× higher per-core RPC throughput compared to both highly-optimized software stacks, and systems using specialized RDMA adapters. It also scales up to 84 Mrps with 8 threads on 4 CPU cores, while maintaining state-of- the-art µs-scale tail latency. We also demonstrate that large third- party applications, like memcached and MICA KVS, can be easily ported on Dagger with minimal changes to their codebase, bringing their median and tail KVS access latency down to 2.8 − 3.5 us and 5.4 − 7.8 us, respectively. Finally, we show that Dagger is beneficial for multi-tier end-to-end microservices with different threading models by evaluating it using an 8-tier application implementing a flight check-in service.
Optimus Prime:加速服务器中的数据转换
DOI: 10.1145/3373376.3378501
发表时间: 2020
期刊: Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子: --
作者:
Arash Pourhabibi Zarandi;Siddharth Gupta;H. Kassir;Mark Sutherland;Zilu Tian;M. Drumond;B. Falsafi;Christoph E. Koch
通讯作者: Christoph E. Koch
DOI: --
发表时间: 2019
期刊: --
影响因子: --
作者:
Ming Liu;Simon Peter;A. Krishnamurthy;P. Phothilimthana
通讯作者: Ming Liu;Simon Peter;A. Krishnamurthy;P. Phothilimthana
DOI: --
发表时间: 2018-10
期刊: --
影响因子: --
作者:
P. Phothilimthana;Ming Liu;Antoine Kaufmann;Simon Peter;Rastislav Bodík;T. Anderson
通讯作者: P. Phothilimthana;Ming Liu;Antoine Kaufmann;Simon Peter;Rastislav Bodík;T. Anderson