Layer 1

Layer 1
复制标题

第 1 层

DOI:
--
复制
发表时间:
--
期刊:
影响因子:
--
通讯作者:
Hwan Song
Hwan Song
中科院分区:
--
文献类型:
--
作者:
Raj Joshi;Chai Song;Xin Zhe Khooi;Nishant Budhdev;Ayush Mishra;Mun Choon Chan;Ben Leong;Hwan Song

文献摘要

被引文献

相似文献

由于链路损坏而导致的数据包丢失是大型仓库规模数据中心的一个主要问题。当前禁用损坏链接的最先进方法是不够的,因为实际上,由于容量限制,无法禁用所有损坏链接。在本文中,我们表明,在亚 RTT 时间尺度上实现链路本地重传以完全掩盖传输端点的损坏数据包丢失是可行的。我们的系统 LinkGuardian 采用一系列技术来 (i) 保持较低的数据包缓冲区要求,(ii) 在不使用超时的情况下从尾部数据包丢失中恢复,以及 (iii) 保持数据包顺序。我们在 Intel Tofino 交换机上实施 LinkGuardian,结果表明,对于丢失率为 10-3 的 100G 链路,LinkGuardian 可以将丢失率降低多达 6 个数量级,同时有效链路速度仅降低 8%。通过消除尾部数据包丢失,LinkGuardian 将 TCP 和 RDMA 的第 99.9 个百分点的流完成时间 (FCT) 分别提高了 51 倍和 66 倍。最后,我们还表明,在数据中心网络的环境中,简单的无序重传通常足以显着减轻短 TCP 流的损坏数据包丢失的影响。
Packet loss due to link corruption is a major problem in large warehouse-scale datacenters. The current state-of-the-art approach of disabling corrupting links is not adequate because, in practice, all the corrupting links cannot be disabled due to capacity constraints. In this paper, we show that, it is feasible to implement link-local retransmission at sub-RTT timescales to completely mask corruption packet losses from the transport endpoints. Our system, LinkGuardian, employs a range of techniques to (i) keep the packet buffer requirement low, (ii) recover from tail packet losses without employing timeouts, and (iii) preserve packet ordering. We implement LinkGuardian on the Intel Tofino switch and show that for a 100G link with a loss rate of 10-3, LinkGuardian can reduce the loss rate by up to 6 orders of magnitude while incurring only 8% reduction in effective link speed. By eliminating tail packet losses, LinkGuardian improves the 99.9th percentile flow completion time (FCT) for TCP and RDMA by 51x and 66x respectively. Finally, we also show that in the context of datacenter networks, simple outof-order retransmission is often sufficient to significantly mitigate the impact of corruption packet loss for short TCP flows.