Adaptive Precision Block-Jacobi for High Performance Preconditioning in the Ginkgo Linear Algebra Software

Adaptive Precision Block-Jacobi for High Performance Preconditioning in the Ginkgo Linear Algebra Software
复制标题

Ginkgo 线性代数软件中用于高性能预处理的自适应精度 Block-Jacobi

DOI:
--
复制
发表时间:
2021
影响因子:
2.7
通讯作者:
E. Quintana
E. Quintana
中科院分区:
计算机科学3区
文献类型:
--
作者:
Goran Flegar;H. Anzt;T. Cojean;E. Quintana

文献摘要

被引文献

相似文献

在数值算法中使用混合精度是加速科学应用的一种有前途的策略。特别地,在高端GPU(图形处理单元)中采用用于低精度算术的专用硬件和数据格式已经激发了旨在仔细降低工作精度以便加速计算的许多努力。对于其性能受存储器带宽约束的算法,在存储器访问之前(和之后)压缩其数据的想法受到了相当大的关注。一个想法是存储一个近似运算符-像一个预条件-在低于工作精度,希望不会影响算法输出。我们实现了第一个高性能的自适应精度块Jacobi预处理器,它选择的精度格式用于存储预处理器数据的飞行,考虑到个别预处理器块的数值特性。我们实现了自适应块Jacobi预处理器作为Ginkgo线性代数库中的生产就绪功能,不仅考虑了IEEE标准的一部分精度格式,而且还考虑了优化指数和有效数长度的自定义格式,以适应预处理器块的特性。在最先进的GPU加速器上运行的实验表明,我们的实现提供了有吸引力的运行时间节省。
The use of mixed precision in numerical algorithms is a promising strategy for accelerating scientific applications. In particular, the adoption of specialized hardware and data formats for low-precision arithmetic in high-end GPUs (graphics processing units) has motivated numerous efforts aiming at carefully reducing the working precision in order to speed up the computations. For algorithms whose performance is bound by the memory bandwidth, the idea of compressing its data before (and after) memory accesses has received considerable attention. One idea is to store an approximate operator–like a preconditioner–in lower than working precision hopefully without impacting the algorithm output. We realize the first high-performance implementation of an adaptive precision block-Jacobi preconditioner which selects the precision format used to store the preconditioner data on-the-fly, taking into account the numerical properties of the individual preconditioner blocks. We implement the adaptive block-Jacobi preconditioner as production-ready functionality in the Ginkgo linear algebra library, considering not only the precision formats that are part of the IEEE standard, but also customized formats which optimize the length of the exponent and significand to the characteristics of the preconditioner blocks. Experiments run on a state-of-the-art GPU accelerator show that our implementation offers attractive runtime savings.