Implementing High-performance Complex Matrix Multiplication via the 3m and 4m Methods

Implementing High-performance Complex Matrix Multiplication via the 3m and 4m Methods
复制标题

通过 3m 和 4m 方法实现高性能复数矩阵乘法

DOI:
--
复制
发表时间:
2017
影响因子:
2.7
通讯作者:
T. Smith
T. Smith
中科院分区:
计算机科学3区
文献类型:
--
作者:
F. V. Zee;T. Smith

文献摘要

被引文献

相似文献

在本文中,我们探讨了复杂的矩阵乘法的实现,我们首先要识别与常规方法相关的各种挑战,该方法需要精心编写的内核,以最低的水平(即汇编语言)实现复杂的算术然后着手开发一种复杂矩阵多路复用的方法,该方法避免了对复杂内核的需求。在基本线性代数子程序和类似Blas的图书馆实例化软件(BLIS)之类的库中,允许内核开发人员将精力集中在更少和更简单的内核上。 4M公式 - 具有多种变体,所有这些都仅依赖于实际矩阵乘法内核。 “诱导的”方法并观察到组装级方法实际上沿着算法变体的4M频谱驻留在BLIS框架中。由于本质上固有的挑战,因此更稳定(因此广泛适用)4M方法的性能受到限制。
In this article, we explore the implementation of complex matrix multiplication. We begin by briefly identifying various challenges associated with the conventional approach, which calls for a carefully written kernel that implements complex arithmetic at the lowest possible level (i.e., assembly language). We then set out to develop a method of complex matrix multiplication that avoids the need for complex kernels altogether. This constraint promotes code reuse and portability within libraries such as Basic Linear Algebra Subprograms and BLAS-Like Library Instantiation Software (BLIS) and allows kernel developers to focus their efforts on fewer and simpler kernels. We develop two alternative approaches—one based on the 3m method and one that reflects the classic 4m formulation—each with multiple variants, all of which rely only on real matrix multiplication kernels. We discuss the performance characteristics of these “induced” methods and observe that the assembly-level method actually resides along the 4m spectrum of algorithmic variants. Implementations are developed within the BLIS framework, and testing on modern hardware confirms that while the less numerically stable 3m method yields the fastest runtimes, the more stable (and thus widely applicable) 4m method’s performance is somewhat limited due to implementation challenges that appear inherent in nature.