Be SMART, Save I/O: A Probabilistic Approach to Avoid Uncorrectable Errors in Storage Systems

Be SMART, Save I/O: A Probabilistic Approach to Avoid Uncorrectable Errors in Storage Systems
复制标题

DOI:
10.1109/cluster51413.2022.00038
复制
发表时间:
2022-09
期刊:
2022 IEEE International Conference on Cluster Computing (CLUSTER)
影响因子:
--
通讯作者:
Md. Arifuzzaman;M. Bhuiyan;Mehmet Gümüs;Engin Arslan
Md. Arifuzzaman;M. Bhuiyan;Mehmet Gümüs;Engin Arslan
中科院分区:
其他
文献类型:
--
作者:
Md. Arifuzzaman;M. Bhuiyan;Mehmet Gümüs;Engin Arslan

文献摘要

相似文献

无声数据损坏对存储系统中数据的完整性构成重大风险。尽管纠错码 (ECC) 可以恢复大多数此类错误,但其中有不可忽略的部分逃脱了 ECC,称为不可纠正错误 (UE)。尽管这种情况很少见,但存储系统规模的不断扩大和 I/O 速率的快速增长将 UE 之间的平均时间从几个月缩短到了几个小时。然而,与磁盘故障不同,UE 很难高精度预测,因此很难采取主动措施。在本文中,我们介绍了一种部署 UE 缓解策略的概率方法,该策略可以捕获 UE 的很大一部分,同时将系统开销保持在可容忍的范围内。为了实现这一目标,我们首先估计 I/O 操作暴露给 UE 的概率,并找到采用 UE 避免策略可以显着减少 UE 暴露的磁盘的最小子集。我们通过大量模拟证明,当使用所提出的概率模型来实现写入验证策略来检测 UE 并从中恢复时,可以通过 1% 的读取开销来避免超过 50% 的写入触发 UE,并且可以通过不到 3.5% 的读取开销来缓解超过 70% 的 UE。我们进一步测量了生产 Lustre 和 GPFS 文件系统中产生的读取开销对写入性能的影响,并验证了我们的发现,即可以避免超过一半的 UE,同时使写入 I/O 整体降低不到 0.9%。
Silent data corruption poses a significant risk to the integrity of data in storage systems. Although error correction codes (ECC) can recover the majority of such errors, a non-negligible portion of them escape ECC, referred as uncorrectable errors (UEs). Despite being rare in nature, increasing scale of storage systems and fast-growing I/O rates decreased the mean time between UEs from months to hours. Yet, unlike disk failures, UEs are hard to predict with high precision, making it difficult to adopt proactive measures. In this paper, we introduce a probabilistic approach to deploy UE mitigation strategies that can capture significant portion of UE while keeping the system overhead at a tolerable range. To achieve this, we first estimate the probability of I/O operations to be exposed to UEs and find a minimum subset of disks for which employing UE avoidance strategies can lead to significant decrease in UE exposure. We demonstrate through extensive simulations that when the proposed probabilistic model is used to implement write verification strategy to detect and recover from UEs, more than 50% of all write-triggered UEs can be avoided with 1% read overhead, and more than 70% of UEs can be mitigated with less than 3.5% read overhead. We further measure the impact of incurred read overhead on write performance in production Lustre and GPFS file systems and validate our findings that more than half of UEs can be avoided while degrading write I/O throughout by less than 0.9%.