Recovering single precision accuracy from Tensor Cores while surpassing the FP32 theoretical peak performance

Recovering single precision accuracy from Tensor Cores while surpassing the FP32 theoretical peak performance
复制标题

DOI:
10.1177/10943420221090256
复制
发表时间:
2022-03
期刊:
The International Journal of High Performance Computing Applications
影响因子:
--
通讯作者:
Hiroyuki Ootomo;Rio Yokota
Hiroyuki Ootomo;Rio Yokota
中科院分区:
其他
文献类型:
--
作者:
Hiroyuki Ootomo;Rio Yokota

文献摘要

被引文献

相似文献

张量内核是NVIDIA图形处理器上的混合精度矩阵-矩阵乘法单元,在安培架构上的理论峰值性能超过300TFLOP/S。张量核是为了响应机器学习对密集矩阵乘法的高要求而开发的。然而,在科学计算中的许多应用,例如迭代求解器的预条件和低精度的傅立叶变换,都可以利用这些张量核心。为了计算张量核上的矩阵乘法,我们需要将输入矩阵转换为半精度,这导致了精度的损失。为了避免这种情况,我们可以使用附加的半精度变量在转换中保留尾数损失,并使用它们来校正矩阵-矩阵乘法的精度。即使进行了此修正,使用张量核也比使用FP32 SIMT核产生更高的吞吐量。然而,这种方法本身的校正能力是有限的,所产生的精度不能与在FP32SIMT核上进行矩阵乘法的精度相匹配。我们针对这一问题,开发了一种使用张量核的高精度、高性能和低功耗的矩阵-矩阵乘法实现,该实现与FP32 SIMT核的精度完全匹配,同时获得了优异的吞吐量。该实现基于NVIDIA的Cutlass。我们发现,实现这一精度的关键是如何在校正计算中处理张量芯内部的舍入和下溢概率。在NVIDIA A100图形处理器上,使用FF32张量核,在有限指数范围内实现了51TFLOP/S,在全指数范围内实现了33TFLOP/S,超过了理论上19.5TFLOP/S的峰值性能。
Tensor Core is a mixed-precision matrix–matrix multiplication unit on NVIDIA GPUs with a theoretical peak performance of more than 300 TFlop/s on Ampere architectures. Tensor Cores were developed in response to the high demand of dense matrix multiplication from machine learning. However, many applications in scientific computing such as preconditioners for iterative solvers and low-precision Fourier transforms can exploit these Tensor Cores. To compute a matrix multiplication on Tensor Cores, we need to convert input matrices to half-precision, which results in loss of accuracy. To avoid this, we can keep the mantissa loss in the conversion using additional half-precision variables and use them for correcting the accuracy of matrix–matrix multiplication. Even with this correction, the use of Tensor Cores yields higher throughput compared to FP32 SIMT Cores. Nevertheless, the correcting capability of this method alone is limited, and the resulting accuracy cannot match that of a matrix multiplication on FP32 SIMT Cores. We address this problem and develop a high accuracy, high performance, and low power consumption matrix–matrix multiplication implementation using Tensor Cores, which exactly matches the accuracy of FP32 SIMT Cores while achieving superior throughput. The implementation is based on NVIDIA’s CUTLASS. We found that the key to achieving this accuracy is how to deal with the rounding inside Tensor Cores and underflow probability during the correction computation. Our implementation achieves 51 TFlop/s for a limited exponent range using FP16 Tensor Cores and 33 TFlop/s for full exponent range of FP32 using TF32 Tensor Cores on NVIDIA A100 GPUs, which outperforms the theoretical FP32 SIMT Core peak performance of 19.5 TFlop/s.