Fault Tolerant One-sided Matrix Decompositions on Heterogeneous Systems with GPUs

Fault Tolerant One-sided Matrix Decompositions on Heterogeneous Systems with GPUs
复制标题

DOI:
10.1109/sc.2018.00071
复制
发表时间:
2018-11
期刊:
SC18: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
Jieyang Chen;Hongbo Li;Sihuan Li;Xin Liang;Panruo Wu;Dingwen Tao;Kaiming Ouyang;Yuanlai Liu;Kai Zhao;Qiang Guan;Zizhong Chen
Jieyang Chen;Hongbo Li;Sihuan Li;Xin Liang;Panruo Wu;Dingwen Tao;Kaiming Ouyang;Yuanlai Liu;Kai Zhao;Qiang Guan;Zizhong Chen
中科院分区:
其他
文献类型:
--
作者:
Jieyang Chen;Hongbo Li;Sihuan Li;Xin Liang;Panruo Wu;Dingwen Tao;Kaiming Ouyang;Yuanlai Liu;Kai Zhao;Qiang Guan;Zizhong Chen

文献摘要

被引文献

相似文献

目前基于算法的容错(ABFT)方法在具有GPU的异构系统上用于单侧矩阵分解存在以下局限性:(1)它们不能提供足够的保护,因为它们中的大多数仅在一维上维护校验和;(2)它们的校验方案由于冗余校验和验证而效率不高;(3)它们不能保护PCIe通信;以及(4)基于特殊类型的矩阵乘法的校验和计算效率很低。通过克服上述限制,我们设计了一个有效的ABFT方法提供更强的保护单侧矩阵分解方法在异构系统。首先,我们通过在两个维度中使用校验和来提供全矩阵保护。其次,我们的检查计划是更有效的优先校验和验证根据矩阵运算的敏感性软错误。第三,我们通过重新排序校验和验证和分解步骤来保护PCIe通信。第四,通过更好地利用GPU,我们将校验和计算速度提高了1.7倍。
Current algorithm-based fault tolerance (ABFT) approach for one-sided matrix decomposition on heterogeneous systems with GPUs have following limitations: (1) they do not provide sufficient protection as most of them only maintain checksum in one dimension; (2) their checking scheme is not efficient due to redundant checksum verifications; (3) they fail to protect PCIe communication; and (4) the checksum calculation based on a special type of matrix multiplication is far from efficient. By overcoming the above limitations, we design an efficient ABFT approach providing stronger protection for one-sided matrix decomposition methods on heterogeneous systems. First, we provide full matrix protection by using checksums in two dimensions. Second, our checking scheme is more efficient by prioritizing the checksum verification according to the sensitivity of matrix operations to soft errors. Third, we protect PCIe communication by reordering checksum verifications and decomposition steps. Fourth, we accelerate the checksum calculation by 1.7x via better utilizing GPUs.