Optimizations of H-matrix-vector Multiplication for Modern Multi-core Processors

Optimizations of H-matrix-vector Multiplication for Modern Multi-core Processors
复制标题

DOI:
10.1109/cluster51413.2022.00056
复制
发表时间:
2022-09
期刊:
2022 IEEE International Conference on Cluster Computing (CLUSTER)
影响因子:
--
通讯作者:
Tetsuya Hoshino;Akihiro Ida;T. Hanawa
Tetsuya Hoshino;Akihiro Ida;T. Hanawa
中科院分区:
其他
文献类型:
--
作者:
Tetsuya Hoshino;Akihiro Ida;T. Hanawa

文献摘要

相似文献

层次矩阵(h -矩阵)可以鲁棒逼近边界元法(BEM)中出现的密集矩阵。为了加快边界元中线性系统的求解速度,必须加快迭代线性解算器中矩阵-向量乘法的运算速度。然而,加速方法通常是针对密集或稀疏矩阵开发的,很少有关于分层矩阵向量乘法(HiMV)的报道。HiMV算法产生了大量的矩阵-向量乘法,这方面还没有得到充分的讨论。因此,艾滋病毒的效率尚未达到其潜力。本文讨论了现代多核cpu的HiMV优化方法:一种用于高效内存访问的h矩阵存储方法,一种在解向量约简操作中避免写争用的方法,一种线程间负载平衡方法,以及一种用于缓存效率的阻塞和子矩阵排序方法。我们证明了这些优化显著提高了现代基于cpu的超级计算机的性能。相对于密集矩阵向量乘法(DGEMV)的目标性能,在A64FX、AMD EPYC和Intel Xeon Cascade Lake处理器上的单插槽执行时,HiMV的失败率分别达到84.8%、100.7%和98.7%。对于具有高速高带宽内存的A64FX来说,内存性能和缓存效率的优化尤为重要。
Hierarchical matrices (H-matrices) can robustly approximate the dense matrices that appear in the boundary element method (BEM). To accelerate the solving of linear systems in the BEM, we must speed up the matrix-vector multiplication in the iterative linear solver. However, speed-up approaches are usually developed for dense or sparse matrices, and are rarely reported for hierarchical matrix-vector multiplication (HiMV). The HiMV algorithm generates a large number of matrix-vector multiplications, which have not been sufficiently discussed. Therefore, the efficiency of HiMV has not reached its potential. This paper discusses optimization methodologies of HiMV for modern multi-core CPUs: an H-matrix storage method for efficient memory access, a method that avoids write contentions during reduction operations on the solution vector, an inter-thread load-balancing method, and blocking and sub-matrix sorting methods for cache efficiency. We demonstrate that these optimizations significantly improve the performance of modern CPU-based supercomputers. Relative to the target performance of dense matrix-vector multiplication (DGEMV), the HiMV flops reached 84.8%, 100.7%, and 98.7% during single-socket execution on the A64FX, AMD EPYC, and Intel Xeon Cascade Lake processors, respectively. Optimization of memory performance and cache efficiency is especially important for the A64FX with high-speed high-bandwidth memory.