HULA: Scalable Load Balancing Using Programmable Data Planes

HULA: Scalable Load Balancing Using Programmable Data Planes
复制标题

DOI:
10.1145/2890955.2890968
复制
发表时间:
2016-03
期刊:
Proceedings of the Symposium on SDN Research
影响因子:
--
通讯作者:
N. Katta;Mukesh M. Hira;Changhoon Kim;Anirudh Sivaraman;J. Rexford
N. Katta;Mukesh M. Hira;Changhoon Kim;Anirudh Sivaraman;J. Rexford
中科院分区:
其他
文献类型:
--
作者:
N. Katta;Mukesh M. Hira;Changhoon Kim;Anirudh Sivaraman;J. Rexford

文献摘要

被引文献

相似文献

数据中心网络使用多根拓扑(例如,树叶-脊柱、胖树)来提供大的对分带宽。这些拓扑在很大程度上使用多路径,并且需要数据平面负载平衡机制来有效利用其对分带宽。典型的负载均衡机制是等成本多路径路由(ECMP),它将流量均匀地分布在多条路径上。受ECMP缺点的启发,拥塞感知负载均衡技术如COMA已经被开发出来。这些技术有两个局限性。首先,由于交换机内存有限,它们只能在边缘交换机上维持少量的拥塞跟踪状态,并且不能扩展到大型拓扑。其次,因为它们是在定制硬件中实现的,所以不能在现场进行修改。本文提出了一种数据平面负载均衡算法Hula,它克服了这两个局限性。首先,不是让叶交换机跟踪到目的地的所有路径上的拥塞,而是每个草裙草裙交换机跟踪通过相邻交换机到目的地的最佳路径的拥塞。其次,我们为新兴的可编程开关设计草裙拉,并在P4中对其进行编程,以演示草裙舞可以在这种可编程芯片组上运行,而不需要定制硬件。我们在模拟中对草裙草裙进行了广泛的评估,结果表明,它在平均流完成时间方面优于可扩展扩展COMA(50%负载时为1.6倍,90%负载时为3倍)。
Datacenter networks employ multi-rooted topologies (e.g., Leaf-Spine, Fat-Tree) to provide large bisection bandwidth. These topologies use a large degree of multipathing, and need a data-plane load-balancing mechanism to effectively utilize their bisection bandwidth. The canonical load-balancing mechanism is equal-cost multi-path routing (ECMP), which spreads traffic uniformly across multiple paths. Motivated by ECMP's shortcomings, congestion-aware load-balancing techniques such as CONGA have been developed. These techniques have two limitations. First, because switch memory is limited, they can only maintain a small amount of congestion-tracking state at the edge switches, and do not scale to large topologies. Second, because they are implemented in custom hardware, they cannot be modified in the field. This paper presents HULA, a data-plane load-balancing algorithm that overcomes both limitations. First, instead of having the leaf switches track congestion on all paths to a destination, each HULA switch tracks congestion for the best path to a destination through a neighboring switch. Second, we design HULA for emerging programmable switches and program it in P4 to demonstrate that HULA could be run on such programmable chipsets, without requiring custom hardware. We evaluate HULA extensively in simulation, showing that it outperforms a scalable extension to CONGA in average flow completion time (1.6 x at 50% load, 3 x at 90% load).