Availability in Globally Distributed Storage Systems

Availability in Globally Distributed Storage Systems
复制标题

DOI:
--
复制
发表时间:
2010-10
期刊:
--
影响因子:
--
通讯作者:
D. Ford;François Labelle;Florentina I. Popovici;M. Stokely;Van-Anh Truong;L. Barroso;C. Grimes;
D. Ford;François Labelle;Florentina I. Popovici;M. Stokely;Van-Anh Truong;L. Barroso;C. Grimes;
中科院分区:
其他
文献类型:
--
作者:
D. Ford;François Labelle;Florentina I. Popovici;M. Stokely;Van-Anh Truong;L. Barroso;C. Grimes;

文献摘要

被引文献

相似文献

高可用性的云存储通常通过构建在商用服务器和磁盘驱动器集群之上的复杂、多层分布式系统来实现。要在包括软件、硬件、网络连接和电源问题在内的大量故障源中实现高性能和可用性,需要复杂的管理、负载平衡和恢复技术。虽然对存储系统的单个组件(如磁盘驱动器)进行了相对丰富的故障研究,但迄今为止关于大型基于云的存储服务的总体可用性行为的报道相对较少。我们对b谷歌的主要存储基础设施进行了为期一年的广泛研究,并提出了统计模型,以进一步了解多种设计选择的影响,如数据放置和复制策略,从而描述了云存储系统的可用性属性。通过这些模型,我们比较了在各种系统参数下的数据可用性,并给出了在我们的车队中观察到的真实故障模式。
Highly available cloud storage is often implemented with complex, multi-tiered distributed systems built on top of clusters of commodity servers and disk drives. Sophisticated management, load balancing and recovery techniques are needed to achieve high performance and availability amidst an abundance of failure sources that include software, hardware, network connectivity, and power issues. While there is a relative wealth of failure studies of individual components of storage systems, such as disk drives, relatively little has been reported so far on the overall availability behavior of large cloudbased storage services. We characterize the availability properties of cloud storage systems based on an extensive one year study of Google's main storage infrastructure and present statistical models that enable further insight into the impact of multiple design choices, such as data placement and replication strategies. With these models we compare data availability under a variety of system parameters given the real patterns of failures observed in our fleet.