SSD failures in the field: symptoms, causes, and prediction models

SSD failures in the field: symptoms, causes, and prediction models
复制标题

DOI:
10.1145/3295500.3356172
复制
发表时间:
2019-11
期刊:
Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
J. Alter;Ji Xue;Alma Dimnaku;E. Smirni
J. Alter;Ji Xue;Alma Dimnaku;E. Smirni
中科院分区:
其他
文献类型:
--
作者:
J. Alter;Ji Xue;Alma Dimnaku;E. Smirni

文献摘要

被引文献

相似文献

近年来,固态硬盘(SSD)因其速度和能效而成为高性能数据中心的主要产品。在这项工作中,我们研究了来自Google数据中心的30,000个驱动器的故障特征,时间跨度为六年。我们的工作量的特点,导致故障的条件,并说明其根本原因不同于共同的期望,但仍然难以辨别。特别是,我们研究故障事件,导致人工干预的维修过程。我们观察到高水平的婴儿死亡率和婴儿和非婴儿失败之间的差异的特点。我们开发了几个机器学习失败预测模型,这些模型被证明是惊人的准确,实现了高召回率和低误报率。这些模型的使用不仅仅是简单的预测,因为它们可以帮助我们解决导致故障的工作负载特征之间的复杂交互,并从监控的症状中识别故障的根本原因。
In recent years, solid state drives (SSDs) have become a staple of high-performance data centers for their speed and energy efficiency. In this work, we study the failure characteristics of 30,000 drives from a Google data center spanning six years. We characterize the workload conditions that lead to failures and illustrate that their root causes differ from common expectation but remain difficult to discern. Particularly, we study failure incidents that result in manual intervention from the repair process. We observe high levels of infant mortality and characterize the differences between infant and non-infant failures. We develop several machine learning failure prediction models that are shown to be surprisingly accurate, achieving high recall and low false positive rates. These models are used beyond simple prediction as they aid us to untangle the complex interaction of workload characteristics that lead to failures and identify failure root causes from monitored symptoms.