Improving performance of iterative methods by lossy checkponting

Improving performance of iterative methods by lossy checkponting
复制标题

通过有损检查改善迭代方法的性能

DOI:
--
复制
发表时间:
2018
期刊:
IEEE International Symposium on High-Performance Parallel Distributed Computing
影响因子:
--
通讯作者:
F. Cappello
F. Cappello
中科院分区:
--
文献类型:
--
作者:
Dingwen Tao;S. Di;Xin Liang;Zizhong Chen;F. Cappello

文献摘要

被引文献

相似文献

迭代法是求解大型稀疏线性系统的常用方法,是许多现代科学模拟的基本操作。大规模迭代方法在并行运行大量队列时,必须定期对动态变量进行检查点,以避免不可避免的故障停止错误,这需要快速的I/O系统和较大的存储空间。为此,显著减少检查点开销对于改进迭代方法的整体性能是至关重要的。我们的贡献是四倍的。(1)我们提出了一种新的有损检查点方案,利用有损压缩器可以显著提高迭代方法的检查点性能。(2)为了保证在有损检查点方案下性能的提高,我们建立了有损检查点性能模型,并从理论上推导了有损检查点中由于数据失真而导致的额外迭代次数的上界。(3)分析了有损检查点(即有损检查点文件导致的额外迭代次数)对多种迭代方法的影响。(4)在2,048核的高性能计算环境下,使用著名的科学计算包PETSc和最先进的检查点/重启工具包,评估了具有最优检查点间隔的有损检查点方案。实验表明,在存在系统故障的情况下,我们优化的有损检查点方案可以显著降低迭代方法的容错开销,与传统检查点相比降低23% ~ 70%,与无损压缩检查点相比降低20% ~ 58%。
Iterative methods are commonly used approaches to solve large, sparse linear systems, which are fundamental operations for many modern scientific simulations. When the large-scale iterative methods are running with a large number of ranks in parallel, they have to checkpoint the dynamic variables periodically in case of unavoidable fail-stop errors, requiring fast I/O systems and large storage space. To this end, significantly reducing the checkpointing overhead is critical to improving the overall performance of iterative methods. Our contribution is fourfold. (1) We propose a novel lossy checkpointing scheme that can significantly improve the checkpointing performance of iterative methods by leveraging lossy compressors. (2) We formulate a lossy checkpointing performance model and derive theoretically an upper bound for the extra number of iterations caused by the distortion of data in lossy checkpoints, in order to guarantee the performance improvement under the lossy checkpointing scheme. (3) We analyze the impact of lossy checkpointing (i.e., extra number of iterations caused by lossy checkpointing files) for multiple types of iterative methods. (4) We evaluate the lossy checkpointing scheme with optimal checkpointing intervals on a high-performance computing environment with 2,048 cores, using a well-known scientific computation package PETSc and a state-of-the-art checkpoint/restart toolkit. Experiments show that our optimized lossy checkpointing scheme can significantly reduce the fault tolerance overhead for iterative methods by 23%∼70% compared with traditional checkpointing and 20%∼58% compared with lossless-compressed checkpointing, in the presence of system failures.