Reliability mechanisms for very large storage systems

Reliability mechanisms for very large storage systems
复制标题

DOI:
10.1109/mass.2003.1194851
复制
发表时间:
2003-04
期刊:
20th IEEE/11th NASA Goddard Conference on Mass Storage Systems and Technologies, 2003. (MSST 2003). Proceedings.
影响因子:
--
通讯作者:
Qin Xin;E. L. Miller;T. Schwarz;D. Long;S. Brandt;W. Litwin
Qin Xin;E. L. Miller;T. Schwarz;D. Long;S. Brandt;W. Litwin
中科院分区:
其他
文献类型:
--
作者:
Qin Xin;E. L. Miller;T. Schwarz;D. Long;S. Brandt;W. Litwin

文献摘要

被引文献

相似文献

可靠性和可用性在由数千个独立存储设备构建的大规模存储系统中越来越重要。大型系统必须能够承受单个组件的故障;在具有数千个磁盘的系统中,即使是不常见的故障也可能发生在某些设备中。我们关注两种类型的错误:不可恢复的读取错误和驱动器故障。我们讨论了检测和恢复这些错误的机制,介绍了改进的技术来检测磁盘读取错误和快速恢复磁盘故障。我们表明,简单的RAID不能保证足够的可靠性,我们的分析探讨系统可用性和存储效率之间的权衡其他方案。根据我们的数据,我们认为双向镜像对于大多数大型存储系统应该是足够的。对于那些需要非常高的可靠性,我们建议使用三向镜像或与RAID相结合的镜像。
Reliability and availability are increasingly important in large-scale storage systems built from thousands of individual storage devices. Large systems must survive the failure of individual components; in systems with thousands of disks, even infrequent failures are likely in some device. We focus on two types of errors: nonrecoverable read errors and drive failures. We discuss mechanisms for detecting and recovering from such errors, introducing improved techniques for detecting errors in disk reads and fast recovery from disk failure. We show that simple RAID cannot guarantee sufficient reliability; our analysis examines the tradeoffs among other schemes between system availability and storage efficiency. Based on our data, we believe that two-way mirroring should be sufficient for most large storage systems. For those that need very high reliability, we recommend either three-way mirroring or mirroring combined with RAID.