MACORD: Online Adaptive Machine Learning Framework for Silent Error Detection

MACORD: Online Adaptive Machine Learning Framework for Silent Error Detection
复制标题

DOI:
10.1109/cluster.2017.128
复制
发表时间:
2017-09
期刊:
2017 IEEE International Conference on Cluster Computing (CLUSTER)
影响因子:
--
通讯作者:
Omer Subasi;S. Di;Prasanna Balaprakash;O. Unsal;Jesús Labarta;A. Cristal;S. Krishnamoorthy;F. Cappello
Omer Subasi;S. Di;Prasanna Balaprakash;O. Unsal;Jesús Labarta;A. Cristal;S. Krishnamoorthy;F. Cappello
中科院分区:
其他
文献类型:
--
作者:
Omer Subasi;S. Di;Prasanna Balaprakash;O. Unsal;Jesús Labarta;A. Cristal;S. Krishnamoorthy;F. Cappello

文献摘要

被引文献

相似文献

未来高性能计算(HPC)系统的资源容量(如计算核心、内存和存储)不断增加,可能会显著增加可靠性风险。静默数据损坏(SDC)或静默错误是损坏HPC执行结果的主要来源之一。与故障停止错误不同,SDC可能是有害和危险的,因为它们不能被硬件检测到。为了弥补这一点,我们提出了一个在线的机器学习为基础的沉默的数据CORruption检测框架(简称MACORD)检测SDC在HPC应用程序。在我们的研究中,我们全面研究了多种机器学习算法的预测能力,并使检测器能够在运行时自动选择最适合的算法,以适应数据动态。因为它只需要空间特征(即,通过将当前时间步中每个数据点的相邻数据值)添加到训练数据中,我们的学习框架表现出低内存开销(小于1%)。基于真实世界科学应用/基准的实验表明,我们的框架可以提高检测灵敏度(即,召回率高达99%。同时,在大多数情况下,假阳性率被限制在0.1%,这是一个数量级的改进相比,最新的国家的最先进的空间技术。
Future high-performance computing (HPC) systems with ever-increasing resource capacity (such as compute cores, memory and storage) may significantly increase the risks on reliability. Silent data corruptions (SDCs) or silent errors are among the major sources that corrupt HPC execution results. Unlike fail-stop errors, SDCs can be harmful and dangerous in that they cannot be detected by hardware. To remedy this, we propose an online MAchine-learning-based silent data CORruption Detection framework (abbreviated as MACORD) for detecting SDCs in HPC applications. In our study, we comprehensively investigate the prediction ability of a multitude of machine-learning algorithms and enable the detector to automatically select the best-fit algorithms at runtime to adapt to the data dynamics. Because it takes only spatial features (i.e., neighboring data values for each data point in the current time step) into the training data, our learning framework exhibits low memory overhead (less than 1%). Experiments based on real-world scientific applications/benchmarks show that our framework can elevate the detection sensitivity (i.e., recall) up to 99%. Meanwhile the false positive rate is limited to 0.1% in most cases, which is one order of magnitude improvement compared with the latest state-of-the-art spatial technique.