Supporting Mixed-domain Mixed-precision Matrix Multiplication within the BLIS Framework

Supporting Mixed-domain Mixed-precision Matrix Multiplication within the BLIS Framework
复制标题

支持 BLIS 框架内的混合域混合精度矩阵乘法

DOI:
10.1145/3402225
复制
发表时间:
2021
影响因子:
2.7
通讯作者:
Geijn, Robert A.
Geijn, Robert A.
中科院分区:
计算机科学3区
文献类型:
--
作者:
Zee, Field G.;Parikh, Devangi N.;Geijn, Robert A.

文献摘要

参考文献

被引文献

相似文献

我们的方法内的一般矩阵乘法(gemm)操作的BLAS-like库实例化软件框架,从而每个矩阵operandA,B,和C可能被存储为单精度或双精度真实的或复杂的值,实现混合数据类型的支持的问题。另一个因素的复杂性,从而矩阵产品和积累被允许发生在一个精度不同的存储精度eitherAorB,也进行了讨论。我们首先将问题分解为正交维度,分别考虑域的混合和混合精度。支持所有组合的矩阵操作数存储在真实的或复杂的域映射出枚举的情况下,并描述了实现方法。支持存储和计算精度的所有组合是通过在计算的关键阶段(根据需要在打包和/或累加期间)对矩阵进行类型转换来处理的。还记录了几个可选的优化。在56核马尔维尔ThunderX 2和52核英特尔至强白金上收集的性能结果表明,高性能基本上得到了保留,不可避免的类型转换指令会导致适度的速度下降。混合数据类型的实现证实了组合的棘手性是避免的,该框架只依赖于两个汇编微内核来实现128个数据类型组合。
We approach the problem of implementing mixed-datatype support within the general matrix multiplication (gemm) operation of the BLAS-like Library Instantiation Software framework, whereby each matrix operandA,B, andCmay be stored as single- or double-precision real or complex values. Another factor of complexity, whereby the matrix product and accumulation are allowed to take place in a precision different from the storage precisions of eitherAorB, is also discussed. We first break the problem into orthogonal dimensions, considering the mixing of domains separately from mixing precisions. Support for all combinations of matrix operands stored in either the real or complex domain is mapped out by enumerating the cases and describing an implementation approach for each. Supporting all combinations of storage and computation precisions is handled by typecasting the matrices at key stages of the computation—during packing and/or accumulation, as needed. Several optional optimizations are also documented. Performance results gathered on a 56-core Marvell ThunderX2 and a 52-core Intel Xeon Platinum demonstrate that high performance is mostly preserved, with modest slowdowns incurred from unavoidable typecast instructions. The mixed-datatype implementation confirms that combinatorial intractability is avoided, with the framework relying on only two assembly microkernels to implement 128 datatype combinations.
DOI: 10.1137/16m108968x
发表时间: 2016-07
期刊: SIAM J. Sci. Comput.
影响因子: --
作者:
D. Matthews
通讯作者: D. Matthews
DOI: --
发表时间: 2017
影响因子: 2.7
作者:
F. V. Zee;T. Smith
通讯作者: T. Smith
DOI: 10.1145/1377603.1377607
发表时间: 2008-07-01
影响因子: 2.7
作者:
Goto, Kazushige;Van De Geijn, Robert
通讯作者: Van De Geijn, Robert
阻尼响应理论的准能量公式。
DOI: --
发表时间: 2009
影响因子: 4.4
作者:
K. Kristensen;J. Kauczor;T. Kjaergaard;P. Jørgensen
通讯作者: P. Jørgensen