Dart: Divide and Specialize for Fast Response to Congestion in RDMA-Based Datacenter Networks

Dart: Divide and Specialize for Fast Response to Congestion in RDMA-Based Datacenter Networks
复制标题

DOI:
10.1109/tnet.2019.2961671
复制
发表时间:
2018-05
期刊:
IEEE/ACM Transactions on Networking
影响因子:
--
通讯作者:
Jiachen Xue;M. Chaudhry;Balajee Vamanan;T. N. Vijaykumar;Mithuna Thottethodi
Jiachen Xue;M. Chaudhry;Balajee Vamanan;T. N. Vijaykumar;Mithuna Thottethodi
中科院分区:
其他
文献类型:
--
作者:
Jiachen Xue;M. Chaudhry;Balajee Vamanan;T. N. Vijaykumar;Mithuna Thottethodi

文献摘要

被引文献

相似文献

尽管与TCP相比,远程直接存储器访问(RDMA)承诺显著减少数据中心网络延迟(例如,10 $\times$),在存在incast的情况下的端到端拥塞控制是一个挑战。针对拥塞问题的全部一般性,先前的方案依赖于缓慢的迭代收敛到适当的发送速率(例如,时间需要50个RTT)。几篇论文已经表明,即使在超额预订的数据中心网络中,大多数拥塞也发生在接收器处。因此,我们提出了一个划分和专门的方法,称为飞镖,隔离接收器拥塞的常见情况下,并进一步细分为更简单的空间本地化和更难的空间分散的情况下,剩余的网络拥塞。对于接收器拥塞,我们提出了直接分配的发送速率(DASR),其中接收器为$n$n指示每个发送器削减其速率的一个因素$n$,收敛在只有一个RTT。对于空间定位的情况,Dart通过添加用于有序流偏转(IOFD)的新颖交换机硬件来提供快速(在一个RTT下)响应,因为RDMA不允许先前负载平衡方案所依赖的分组重新排序。对于不常见的空间分散的情况,Dart福尔斯回落到DCQCN。小规模测试床测量和大规模模拟分别显示,Dart实现了比InfiniBand、TIMELY和DCQCN低60%(2.5 $\times$)和79%(4.8 $\times$)的第99个百分点延迟,以及相似和高58%的吞吐量。
Though Remote Direct Memory Access (RDMA) promises to reduce datacenter network latencies significantly compared to TCP (e.g., 10 $\times$ ), end-to-end congestion control in the presence of incasts is a challenge. Targeting the full generality of the congestion problem, previous schemes rely on slow, iterative convergence to the appropriate sending rates (e.g., TIMELY takes 50 RTTs). Several papers have shown that even in oversubscribed datacenter networks most congestion occurs at the receiver. Accordingly, we propose a divide-and-specialize approach, called Dart, which isolates the common case of receiver congestion and further subdivides the remaining in-network congestion into the simpler spatially-localized and the harder spatially-dispersed cases. For receiver congestion, we propose direct apportioning of sending rates (DASR) in which a receiver for $n$ senders directs each sender to cut its rate by a factor of $n$ , converging in only one RTT. For the spatially-localized case, Dart provides fast (under one RTT) response by adding novel switch hardware for in-order flow deflection (IOFD) because RDMA disallows packet reordering on which previous load balancing schemes rely. For the uncommon spatially-dispersed case, Dart falls back to DCQCN. Small-scale testbed measurements and at-scale simulations, respectively, show that Dart achieves 60% (2.5 $\times$ ) and 79% (4.8 $\times$ ) lower $99^{th}$ -percentile latency, and similar and 58% higher throughput than InfiniBand, and TIMELY and DCQCN.