Leveraging the bfloat16 Artificial Intelligence Datatype For Higher-Precision Computations

Leveraging the bfloat16 Artificial Intelligence Datatype For Higher-Precision Computations
复制标题

利用 bfloat16 人工智能数据类型进行更高精度的计算

DOI:
--
复制
发表时间:
2019
期刊:
IEEE Symposium on Computer Arithmetic
影响因子:
--
通讯作者:
A. Heinecke
A. Heinecke
中科院分区:
--
文献类型:
--
作者:
G. Henry;P. T. P. Tang;A. Heinecke

文献摘要

被引文献

相似文献

近年来,具有较低精确乘法和较高精确累积的融合 - 添加剂(FMA)单位已被证明对机器学习/人工智能应用有用,最著名的是由于其极端的计算强度,在培训深神经网络中。与经典的IEEE-754 32位(FP32)和64位(FP64)算术相比,这些降低的精度算术自然可以与它们的缩短宽度相比,这可以自然加速。所有主要硬件供应商的共同策略是积极进一步提高其性能。在深度学习中发现了一个特殊的FMA操作,该操作在fp32中累积时乘以两个BF16数字,其中BF16是具有IEEE FP32数值范围的16位浮点数据类型,但精度为8位。在本文中,我们研究了该FMA单元以潜在的绩效增益和对准确性的影响来实施更高精确的矩阵例程。我们演示了如何使用分解为多个较小的数据类型来组装高精度结果,从而利用FMA单元的较高精度积累。我们首先证明了向量内部产品和自然扩展的计算,可以通过在几个BF16数字中分解FP32数字,然后进行适当的计算来实现矩阵矩阵产品,然后进行适当的计算,这些计算可以适应与标准FP32计算相比的动态范围和保留精度高达5.2倍加速。此外,我们检查了以残余形式制定的线性方程的解决方案,该方程允许迭代改进。我们证明,获得的解决方案可与FP64在大量线性系统条件数字下提供的解决方案相媲美。
In recent years fused-multiply-add (FMA) units with lower-precision multiplications and higher-precision accumulation have proven useful in machine learning/artificial intelligence applications, most notably in training deep neural networks due to their extreme computational intensity. Compared to classical IEEE-754 32 bit (FP32) and 64 bit (FP64) arithmetic, these reduced precision arithmetic can naturally be sped up disproportional to their shortened width. The common strategy of all major hardware vendors is to aggressively further enhance their performance disproportionately. One particular FMA operation that multiplies two BF16 numbers while accumulating in FP32 has been found useful in deep learning, where BF16 is the 16-bit floating point datatype with IEEE FP32 numerical range but 8 significant bits of precision. In this paper, we examine the use this FMA unit to implement higher-precision matrix routines in terms of potential performance gain and implications on accuracy. We demonstrate how a decomposition into multiple smaller datatypes can be used to assemble a high-precision result, leveraging the higher precision accumulation of the FMA unit. We first demonstrate that computations of vector inner products and by natural extension, matrix-matrix products can be achieved by decomposing FP32 numbers in several BF16 numbers followed by appropriate computations that can accommodate the dynamic range and preserve accuracy compared to standard FP32 computations, while projecting up to 5.2x speed-up. Furthermore, we examine solution of linear equations formulated in the residual form that allows for iterative refinement. We demonstrate that the solution obtained to be comparable to those offered by FP64 under a large range of linear system condition numbers.