Zero Downtime Release: Disruption-free Load Balancing of a Multi-Billion User Website

Zero Downtime Release: Disruption-free Load Balancing of a Multi-Billion User Website
复制标题

零停机发布:数十亿用户网站的无中断负载平衡

DOI:
10.1145/3387514.3405885
复制
发表时间:
2020
期刊:
and protocols for computer communication
影响因子:
--
通讯作者:
Benson, Theophilus A.
Benson, Theophilus A.
中科院分区:
--
文献类型:
--
作者:
Naseer, Usama;Niccolini, Luca;Pant, Udip;Frindell, Alan;Dasineni, Ranjeeth;Benson, Theophilus A.

文献摘要

参考文献

被引文献

相似文献

现代网络基础设施已经发展成为一个复杂的有机体,以满足数十亿用户的性能和可用性需求。代码升级、错误修复和安全更新等频繁发布已成为常态。数以百万计的全球分布式基础设施组件(包括服务器和负载平衡器)频繁重启,从每天多次到每周多次。然而,每次发布都会带来中断的可能性,因为它可能导致集群容量减少,干扰大规模运行的组件之间复杂的交互,并通过终止连接来干扰最终用户。受支持的服务和协议的规模和异构性使挑战进一步复杂化。在本文中,我们利用端到端网络基础设施的不同组件来防止或掩盖面对发布时的任何中断。零停机时间发布是Facebook使用的一系列机制,用于保护最终用户免受任何中断,在全球发布更新时保持集群容量和基础设施的健壮性。我们的评估表明,当大量生产服务器和代理重新启动时,这些机制可以防止任何显著的集群容量下降,并最大限度地减少不同服务(特别是TCP、HTTP和发布/订阅)的中断。
Modern network infrastructure has evolved into a complex organism to satisfy the performance and availability requirements for the billions of users. Frequent releases such as code upgrades, bug fixes and security updates have become a norm. Millions of globally distributed infrastructure components including servers and load-balancers are restarted frequently from multiple times per-day to per-week. However, every release brings possibilities of disruptions as it can result in reduced cluster capacity, disturb intricate interaction of the components operating at large scales and disrupt the end-users by terminating their connections. The challenge is further complicated by the scale and heterogeneity of supported services and protocols.In this paper, we leverage different components of the end-to-end networking infrastructure to prevent or mask any disruptions in face of releases. Zero Downtime Release is a collection of mechanisms used at Facebook to shield the end-users from any disruptions, preserve the cluster capacity and robustness of the infrastructure when updates are released globally. Our evaluation shows that these mechanisms prevent any significant cluster capacity degradation when a considerable number of productions servers and proxies are restarted and minimizes the disruption for different services (notably TCP, HTTP and publish/subscribe).
大规模快速发展:在 Facebook 部署 IETF QUIC 的经验
DOI: --
发表时间: 2018
期刊: Proceedings of the Workshop on the Evolution, Performance, and Interoperability of QUIC
影响因子: --
作者:
S. Iyengar
通讯作者: S. Iyengar
优化 UDP 以进行内容交付:GSO、节奏和零复制
DOI: --
发表时间: 2020
期刊:
影响因子: --
作者:
K. Albrecht;J. Callhoff;Matthias Schneider;A. Zink
通讯作者: A. Zink
Facebook 和 开源
DOI: --
发表时间: 2013
期刊:
影响因子: --
作者:
Силаков Денис Владимирович
通讯作者: Силаков Денис Владимирович
不断发展的数据中心中基于风险的网络变化规划
DOI: --
发表时间: 2019
期刊: Symposium on Operating Systems Principles
影响因子: --
作者:
Omid Alipourfard;Jiaqi Gao;Jérémie Koenig;Chris Harshaw;Amin Vahdat;Minlan Yu
通讯作者: Minlan Yu