Lifespan and Failures of SSDs and HDDs: Similarities, Differences, and Prediction Models

Lifespan and Failures of SSDs and HDDs: Similarities, Differences, and Prediction Models
复制标题

DOI:
10.1109/tdsc.2021.3131571
复制
发表时间:
2023-01
影响因子:
7.3
通讯作者:
Riccardo Pinciroli;Lishan Yang;J. Alter;E. Smirni
Riccardo Pinciroli;Lishan Yang;J. Alter;E. Smirni
中科院分区:
计算机科学2区
文献类型:
--
作者:
Riccardo Pinciroli;Lishan Yang;J. Alter;E. Smirni

文献摘要

相似文献

数据中心的停机时间通常主要是由IT设备故障引起的。存储设备是数据中心中最常出现故障的组件。我们对构成数据中心典型存储的硬盘驱动器(HDD)和固态硬盘(SSD)进行了一项比较研究。利用来自Backblaze数据集中同一制造商的10万个不同型号硬盘驱动器的6年现场数据以及来自谷歌数据中心的3种型号的3万个固态硬盘的6年现场数据,我们描述了导致故障的工作负载条件。我们说明它们的根本故障原因与普遍预期不同,而且仍然难以辨别。对于硬盘驱动器,我们观察到新硬盘和旧硬盘在故障方面没有太大差异。相反,可以根据磁头定位所花费的时间对硬盘进行区分来辨别故障。对于固态硬盘,我们观察到较高的早期故障率,并描述了早期故障和非早期故障之间的差异。我们开发了几种机器学习故障预测模型,这些模型被证明具有惊人的准确性,实现了高召回率和低误报率。这些模型不仅仅用于简单预测,因为它们帮助我们理清导致故障的工作负载特征的复杂相互作用,并从监测到的症状中确定故障的根本原因。
Data center downtime typically centers around IT equipment failure. Storage devices are the most frequently failing components in data centers. We present a comparative study of hard disk drives (HDDs) and solid state drives (SSDs) that constitute the typical storage in data centers. Using six-year field data of 100,000 HDDs of different models from the same manufacturer from the Backblaze dataset and six-year field data of 30,000 SSDs of three models from a Google data center, we characterize the workload conditions that lead to failures. We illustrate that their root failure causes differ from common expectations and that they remain difficult to discern. For the case of HDDs we observe that young and old drives do not present many differences in their failures. Instead, failures may be distinguished by discriminating drives based on the time spent for head positioning. For SSDs, we observe high levels of infant mortality and characterize the differences between infant and non-infant failures. We develop several machine learning failure prediction models that are shown to be surprisingly accurate, achieving high recall and low false positive rates. These models are used beyond simple prediction as they aid us to untangle the complex interaction of workload characteristics that lead to failures and identify failure root causes from monitored symptoms.