Improving Energy Saving of One-Sided Matrix Decompositions on CPU-GPU Heterogeneous Systems

Improving Energy Saving of One-Sided Matrix Decompositions on CPU-GPU Heterogeneous Systems
复制标题

DOI:
10.1145/3572848.3577496
复制
发表时间:
2023-01
期刊:
Proceedings of the 28th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming
影响因子:
--
通讯作者:
Jieyang Chen;Xin Liang;Kai Zhao;H. Sabzi;L. Bhuyan;Zizhong Chen
Jieyang Chen;Xin Liang;Kai Zhao;H. Sabzi;L. Bhuyan;Zizhong Chen
中科院分区:
其他
文献类型:
--
作者:
Jieyang Chen;Xin Liang;Kai Zhao;H. Sabzi;L. Bhuyan;Zizhong Chen

文献摘要

相似文献

单方面的密度矩阵分解(例如,Cholesky,LU和QR)是许多不同领域的科学计算中的关键组成部分-GPU异质系统通常用于基质分解,在这项工作中,我们旨在进一步改善CPU-GPU异质系统上单基质分解的能源节省。我们首先构建了一个基于算法的容错保护超频技术(ABFT-OC),以使我们能够利用可靠的超频进行关键基质分解操作。 ,这可以巧妙地结合ABFT-OC和DVF提供的能力,以最大程度地提高节能和维护性能和可靠性。与当前的最佳节能优化方法相比,最多可节省11.7%的能源,而无需降解,最多可提供14.1%的能量×延迟2。至1.43×性能提高,而无需花费额外的能量。
One-sided dense matrix decompositions (e.g., Cholesky, LU, and QR) are the key components in scientific computing in many different fields. Although their design has been highly optimized for modern processors, they still consume a considerable amount of energy. As CPU-GPU heterogeneous systems are commonly used for matrix decompositions, in this work, we aim to further improve the energy saving of onesided matrix decompositions on CPU-GPU heterogeneous systems. We first build an Algorithm-Based Fault Tolerance protected overclocking technique (ABFT-OC) to enable us to exploit reliable overclocking for key matrix decomposition operations. Then, we design an energy-saving matrix decomposition framework, Bi-directional Slack Reclamation (BSR), that can intelligently combine the capability provided by ABFT-OC and DVFS to maximize energy saving and maintain performance and reliability. Experiments show that BSR is able to save up to 11.7% more energy compared with the current best energy saving optimization approach with no performance degradation and up to 14.1% Energy×Delay2 reduction. Also, BSR enables the Pareto efficient performance-energy trade-off, which is able to provide up to 1.43× performance improvement without costing extra energy.