Impact of data placement on resilience in large-scale object storage systems

Impact of data placement on resilience in large-scale object storage systems
复制标题

数据放置对大规模对象存储系统弹性的影响

DOI:
--
复制
发表时间:
2016
期刊:
IEEE Conference on Mass Storage Systems and Technologies
影响因子:
--
通讯作者:
C. Carothers
C. Carothers
中科院分区:
--
文献类型:
--
作者:
P. Carns;K. Harms;John Jenkins;Misbah Mubarak;R. Ross;C. Carothers

文献摘要

被引文献

相似文献

分布式对象存储体系结构已成为大数据、云和HPC计算中高性能存储的事实标准。使用商用硬件来降低成本的对象存储部署通常使用对象复制作为实现数据弹性的一种方法。然而,对于拥有数千台服务器和数十亿对象的系统来说,故障后修复对象副本是一项艰巨的任务,而且在现实系统中大规模评估此类场景越来越困难。如果不及时修复对象,弹性和可用性都会受到影响。在这项工作中,我们利用高保真离散事件仿真模型来研究具有数千台服务器、数十亿个对象和PB级数据的大规模对象存储系统上的副本重建。我们评估了著名的对象放置算法CRUSH的行为,并确定了聚合重建性能受对象放置策略限制的配置场景。在确定了这一瓶颈的根本原因之后,我们提出了对CRUSH的增强以及其上的使用策略,以支持可伸缩的副本重建。我们使用这些方法在1,024个节点的商用存储系统上演示了410GiB/S的模拟聚合重建速率(在预测的理想线性扩展的5%以内)。根据存储在系统上的数据的特征,我们还发现了重建性能中的一个意外现象。
Distributed object storage architectures have become the de facto standard for high-performance storage in big data, cloud, and HPC computing. Object storage deployments using commodity hardware to reduce costs often employ object replication as a method to achieve data resilience. Repairing object replicas after failure is a daunting task for systems with thousands of servers and billions of objects, however, and it is increasingly difficult to evaluate such scenarios at scale on real-world systems. Resilience and availability are both compromised if objects are not repaired in a timely manner. In this work we leverage a high-fidelity discrete-event simulation model to investigate replica reconstruction on large-scale object storage systems with thousands of servers, billions of objects, and petabytes of data. We evaluate the behavior of CRUSH, a well-known object placement algorithm, and identify configuration scenarios in which aggregate rebuild performance is constrained by object placement policies. After determining the root cause of this bottleneck, we then propose enhancements to CRUSH and the usage policies atop it to enable scalable replica reconstruction. We use these methods to demonstrate a simulated aggregate rebuild rate of 410 GiB/s (within 5% of projected ideal linear scaling) on a 1,024-node commodity storage system. We also uncover an unexpected phenomenon in rebuild performance based on the characteristics of the data stored on the system.