Receiver-Driven RDMA Congestion Control by Differentiating Congestion Types in Datacenter Networks

Receiver-Driven RDMA Congestion Control by Differentiating Congestion Types in Datacenter Networks
复制标题

DOI:
10.1109/icnp52444.2021.9651938
复制
发表时间:
2021-11
期刊:
2021 IEEE 29th International Conference on Network Protocols (ICNP)
影响因子:
--
通讯作者:
Jiao Zhang;Jiaming Shi;Xiaolong Zhong;Zirui Wan;Yuxing Tian;Tian Pan;Tao Huang
Jiao Zhang;Jiaming Shi;Xiaolong Zhong;Zirui Wan;Yuxing Tian;Tian Pan;Tao Huang
中科院分区:
其他
文献类型:
--
作者:
Jiao Zhang;Jiaming Shi;Xiaolong Zhong;Zirui Wan;Yuxing Tian;Tian Pan;Tao Huang

文献摘要

被引文献

相似文献

数据中心应用程序的开发导致需要具有微秒延迟的端到端通信。因此,RDMA在数据中心网络中变得流行,以减轻由传统软件网络堆栈的缓慢处理速度引起的延迟。然而,现有的RDMA拥塞控制机制在同时实现高吞吐量和低延迟方面远非最佳,或者需要额外的网络内功能支持。在本文中,通过利用观察,大多数拥塞发生在数据中心网络的最后一跳,我们提出了RCC,接收器驱动的快速拥塞控制机制,结合显式分配和迭代窗口调整的RDMA网络。首先,我们提出了一种网络拥塞判别方法,将拥塞分为最后一跳拥塞和网络内拥塞两种类型。然后,提出了一种显式窗口分配机制来解决最后一跳拥塞问题,使RTP能够在一个RTT内收敛到合适的发送速率。对于网络拥塞,提出了一种基于PID的迭代延迟窗口调整方案,以实现快速收敛和接近零的排队延迟。RCC不需要额外的网络支持,并且对硬件实现友好。在我们的评估中,RCC的总体平均FCT(流程完成时间)比Homa、DSPass、DCQCN、TIMELY和HPCC好4~79%。
The development of datacenter applications leads to the need for end-to-end communication with microsecond latency. As a result, RDMA is becoming prevalent in datacenter networks to mitigate the latency caused by the slow processing speed of the traditional software network stack. However, existing RDMA congestion control mechanisms are either far from optimal in simultaneously achieving high throughput and low latency or in need of additional in-network function support. In this paper, by leveraging the observation that most congestion occurs at the last hop in datacenter networks, we propose RCC, a receiver-driven rapid congestion control mechanism for RDMA networks that combines explicit assignment and iterative window adjustment. Firstly, we propose a network congestion distinguish method to classify congestions into two types, last-hop congestion and innetwork congestion. Then, an Explicit Window Assignment mechanism is proposed to solve the last-hop congestion, which enables senders to converge to a proper sending rate in one-RTT. For in-network congestion, a PID-based iterative delay-based window adjustment scheme is proposed to achieve fast convergence and near-zero queuing latency. RCC does not need additional innetwork support and is friendly to hardware implementation. In our evaluation, the overall average FCT (Flow Completion Time) of RCC is 4~79% better than Homa, ExpressPass, DCQCN, TIMELY, and HPCC.