Thinking about Availability in Large Service Infrastructures

Thinking about Availability in Large Service Infrastructures
复制标题

大型服务基础设施可用性的思考

DOI:
10.1145/3102980.3102983
复制
发表时间:
2017
期刊:
Proceedings of the 16th Workshop on Hot Topics in Operating Systems
影响因子:
--
通讯作者:
B. Welch
B. Welch
中科院分区:
--
文献类型:
--
作者:
J. Mogul;R. Isaacs;B. Welch

文献摘要

被引文献

相似文献

随着互联网、Web 和云计算的兴起,我们开始依赖一系列复杂的相互依赖的在线服务。云托管服务(例如 NetFlix 或 SnapChat)的成功运行可能依赖于数十个底层分布式系统,其中一些系统又相互依赖。虽然大多数人(青少年除外)并不认为 SnapChat 是一项至关重要的服务,但医疗保健等系统过去的失败却是如此。 gov [12, 33] 已经产生了现实世界的后果。随着传统企业计算迁移到云端,云提供商面临着提供更多“9”可用性的压力。一些提供商的“无责备事后分析”文化强调了对端到端可用性方法的需求[5,第 1 章]。 15]。然而,学习成功和有害做法的模式本身并不能导致制定能够抵御意外情况的原则。与此同时,研究界并没有像关注分布式共识和状态机复制等各种单点解决方案那样,对这些端到端可用性问题给予同样的关注,特别是对于大规模基础设施(参见第 5 节)。我们观察到,系统设计人员很难定义适合大型基础设施的整体可用性目标,然后再次努力将这些目标转换为组件服务的目标。在本文中,我们描述了为可用性提供精确且有原则的定义所面临的挑战(第 3 节)、以与我们已经学会的考虑安全性大致相同的方式思考可用性的可能性(第 4 节),以及设计高可用性基础设施的一些一般想法(第 6 节)。我们做
With the rise of the Internet, the Web, and cloud computing, we have come to depend on a complex stack of interdependent online services. Successful operation of a cloud-hosted service, such as NetFlix or SnapChat, can depend on dozens of underlying distributed systems, some of which in turn depend on each other. While most people (excepting teenagers) do not view SnapChat as a life-critical service, past failures of systems such as Healthcare. gov [12, 33] have had real-world consequences. Cloud providers are under pressure to deliver more “nines” of availability as traditional enterprise computing moves to the cloud.The need for an end-to-end approach to availability is highlighted by the “blameless postmortem” culture of some providers [5, Ch. 15]. However, learning the patterns of successful and harmful practices does not in itself lead to the creation of principles that can defend against the unexpected. Meanwhile, the research community has not given the same attention to these end-to-end availability issues, especially for large-scale infrastructures, as it has to various point solutions, such as distributed consensus and state-machine replication (see § 5). We have observed that system designers struggle to define overall availability goals suitable for large infrastructures, and then struggle again to convert these to goals for component services. In this paper, we describe the challenges of providing a precise and principled definition for availability (§ 3), the possibility of thinking about availability in much the same way that we have learned to think about security (§ 4), and some general ideas for designing highly-available infrastructures (§ 6). We do