Efficient detection of silent data corruption in HPC applications with synchronization-free message verification

Efficient detection of silent data corruption in HPC applications with synchronization-free message verification
复制标题

通过免同步消息验证有效检测 HPC 应用程序中的静默数据损坏

DOI:
10.1007/s11227-021-03892-4
复制
发表时间:
2021-06
期刊:
The Journal of Supercomputing
影响因子:
--
通讯作者:
Depei Qian
Depei Qian
中科院分区:
其他
文献类型:
--
作者:
Guozhen Zhang;Yi Liu;Hailong Yang;Depei Qian

文献摘要

参考文献

相似文献

当今,高性能计算(HPC)正向亿亿级时代迈进。然而,以位翻转为表现形式的静默数据损坏(SDC)会给科学计算带来灾难性的后果,严重威胁到大规模HPC的可靠性。最常用的解决SDC的方法是基于模块冗余,通常需要通过同步和在程序执行期间执行额外的消息传输和比较来保持副本之间的执行进度一致。虽然这些方法可以检测具有高召回率的SDC,但它们可能会引入显著的性能开销,甚至大规模地停止执行进度。据我们所知,本文提出了第一个解决方案的SDC检测不需要同步和副本之间的额外的消息传输。它将消息日志记录与创新的异步消息比较机制相结合,该机制使用专门的服务例程(数据分析服务,DAS)来执行进度比较,而不会干扰目标程序的执行。此外,我们的解决方案采用了分布式并行架构来执行DAS,并利用一个创新的引用机制,基于单一的非确定性事件,以保证一致的执行不同的副本。我们实现了一个用户级的原型,称为同步免费SDC检测(SFSD)。在天河二号超级计算机上的实验结果表明,SFSD是有效的检测SDC,与低性能的开销(10%以内)和可以接受的召回率。此外,SFSD表现出良好的可扩展性时,应用于大规模的程序执行。
Nowadays, high-performance computing (HPC) is stepping forward to exascale era. However, silent data corruption (SDC) behaved as bit-flipping can cause disastrous consequences for scientific computation, which jeopardizes the reliability of HPC at large scale. The most commonly used methods to address SDC are based on modular redundancy, which usually requires keeping execution progress consistent between replicas by synchronization and performing additional message transmission and comparison during program execution. Although such methods can detect SDC with high recall, they can introduce significant performance overhead and even stall the execution progress at a large scale. To our knowledge, this paper proposes the first solution of SDC detection without requiring synchronization and additional message transmission between replicas. It combines message logging with an innovative asynchronous message comparison mechanism, which uses specialized service routines (Data-Analytic-Service, DAS) to perform progress comparison without interfering target program execution. Besides, our solution adopts a distributed parallel architecture to perform DAS and utilizes an innovative reference mechanism based on single non-deterministic event to guarantee the consistent execution of different replicas. We implemented a user-level prototype, termed as synchronization-free SDC detection (SFSD). The experimental results on the Tianhe-2 supercomputer show that SFSD is effective in detecting SDC, with low-performance overhead (within 10%) and an acceptable recall rate. Moreover, SFSD exhibits good scalability when applied to large-scale program executions.
DOI: 10.1145/2802658.2802665
发表时间: 2015-09
期刊: Proceedings of the 22nd European MPI Users' Group Meeting
影响因子: --
作者:
L. Bautista-Gomez;F. Cappello
通讯作者: L. Bautista-Gomez;F. Cappello
DOI: 10.1109/cluster.2017.128
发表时间: 2017-09
期刊: 2017 IEEE International Conference on Cluster Computing (CLUSTER)
影响因子: --
作者:
Omer Subasi;S. Di;Prasanna Balaprakash;O. Unsal;Jesús Labarta;A. Cristal;S. Krishnamoorthy;F. Cappello
通讯作者: Omer Subasi;S. Di;Prasanna Balaprakash;O. Unsal;Jesús Labarta;A. Cristal;S. Krishnamoorthy;F. Cappello
DOI: 10.1109/sc.2012.13
发表时间: 2012-11
期刊: 2012 International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子: --
作者:
Vilas Sridharan;Dean Liberty
通讯作者: Vilas Sridharan;Dean Liberty
DOI: 10.1145/2063384
发表时间: 2011-11
期刊: Proceedings of 2011 International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子: --
作者:
Scott A. Lathrop;J. Costa;W. Kramer
通讯作者: Scott A. Lathrop;J. Costa;W. Kramer
DOI: 10.1145/2687651
发表时间: 2015-01
期刊: ACM Transactions on Architecture and Code Optimization (TACO)
影响因子: --
作者:
Leo Porter;M. Laurenzano;Ananta Tiwari;Adam Jundt;W. A. Ward;R. Campbell;L. Carrington
通讯作者: Leo Porter;M. Laurenzano;Ananta Tiwari;Adam Jundt;W. A. Ward;R. Campbell;L. Carrington