FP8 Formats for Deep Learning

FP8 Formats for Deep Learning
复制标题

用于深度学习的 FP8 格式

DOI:
10.48550/arxiv.2209.05433
复制
发表时间:
2022
期刊:
ArXiv
影响因子:
--
通讯作者:
Hao Wu
Hao Wu
中科院分区:
--
文献类型:
--
作者:
P. Micikevicius;Dusan Stosic;N. Burgess;Marius Cornea;P. Dubey;R. Grisenthwaite;Sangwon Ha;A. Heinecke;Patrick Judd;John Kamalu;Naveen Mellempudi;S. Oberman;M. Shoeybi;Michael Siu;Hao Wu

文献摘要

被引文献

相似文献

FP8是加速深度学习训练推断的自然发展,超出了现代处理器中常见的16位格式。和3位Mantissa)和E5M2(5位指数和2位Mantissa)。范围是通过未代表五个范围的范围来扩展的,并且在NANS上只有一个Mantissa位图案涵盖了主要的现代神经网络体系结构 - CNN,RNN和基于变压器的模型,使所有超级参数从16位基线培训课程中保持不变训练实验包括大型,最多175B参数,我们还检查了使用16位格式训练的语言模型的FP8训练后量化。
FP8 is a natural progression for accelerating deep learning training inference beyond the 16-bit formats common in modern processors. In this paper we propose an 8-bit floating point (FP8) binary interchange format consisting of two encodings - E4M3 (4-bit exponent and 3-bit mantissa) and E5M2 (5-bit exponent and 2-bit mantissa). While E5M2 follows IEEE 754 conventions for representatio of special values, E4M3’s dynamic range is extended by not representing infinities and having only one mantissa bit-pattern for NaNs. We demonstrate the efficacy of the FP8 format on a variety of image and language tasks, effectively matching the result quality achieved by 16-bit training sessions. Our study covers the main modern neural network architectures - CNNs, RNNs, and Transformer-based models, leaving all the hyperparameters unchanged from the 16-bit baseline training sessions. Our training experiments include large, up to 175B parameter, language models. We also examine FP8 post-training-quantization of language models trained using 16-bit formats that resisted fixed point int8 quantization.