Fast Symmetric Eigenvalue Decomposition via WY Representation on Tensor Core

Fast Symmetric Eigenvalue Decomposition via WY Representation on Tensor Core
复制标题

通过张量核心上的 WY 表示进行快速对称特征值分解

DOI:
10.1145/3572848.3577516
复制
发表时间:
2023
期刊:
ACM
影响因子:
--
通讯作者:
Wu, Panruo
Wu, Panruo
中科院分区:
--
文献类型:
--
作者:
Zhang, Shaoshuai;Shah, Ruchi;Ootomo, Hiroyuki;Yokota, Rio;Wu, Panruo

文献摘要

参考文献

相似文献

对称特征值分解(EVD)是一种基本的分析和数值工具,在许多科学领域中使用。在性能方面,最先进的算法通常是两阶段三对角化方法。两阶段三对角化的第一阶段称为连续带约简(SBR),它将对称矩阵约简为带形式,其计算成本通常占主导地位。当使用Tensor Core(专用矩阵计算加速器)来加速昂贵的EVD时,由于矩阵计算的不利形状,传统的基于ZY表示的方法导致次优性能。在本文中,我们提出了一种使用WY表示而不是ZY表示的新方法(详见第3.2节),它可以提供局部性和并行性的更好组合,从而在Tensor Cores上表现得更好。实验上,所提出的方法可以带来高达3.7倍的加速比在SBR和2.3倍,在整个EVD相比,国家的最先进的实现。
Symmetric eigenvalue decomposition (EVD) is a fundamental analytic and numerical tool used in many scientific areas. The state-of-the-art algorithm in terms of performance is typically the two-stage tridiagonalization method. The first stage in the two-stage tridiagonalization is called successive band reduction (SBR), which reduces a symmetric matrix to a band form, and its computational cost usually dominates. When Tensor Core (specialized matrix computational accelerator) is used to accelerate the expensive EVD, the conventional ZY-representation-based method results in suboptimal performance due to unfavorable shapes of the matrix computations. In this paper, we propose a new method that uses WY representation instead of ZY representation (see Section 3.2 for details), which can provide a better combination of locality and parallelism so as to perform better on Tensor Cores. Experimentally, the proposed method can bring up to 3.7x speedup in SBR and 2.3x in the entire EVD compared to state-of-the-art implementations.
TensorCore GPU 上的基本线性代数运算
DOI: --
发表时间: 2020
期刊: ACM SIGPLAN Symposium on Scala
影响因子: --
作者:
Shaoshuai Zhang;Vivek Karihaloo;Panruo Wu
通讯作者: Panruo Wu
DOI: 10.1137/070699895
发表时间: 2008
期刊: SIAM J. Matrix Anal. Appl.
影响因子: --
作者:
R. Byers;Hongguo Xu
通讯作者: Hongguo Xu
使用聚合细粒度和内存感知内核并行简化对称特征值问题的压缩形式
DOI: 10.1145/2063384.2063394
发表时间: 2011
期刊: 2011 International Conference for High Performance Computing, Networking, Storage and Analysis (SC)
影响因子: --
作者:
A. Haidar;H. Ltaief;J. Dongarra
通讯作者: J. Dongarra
用于大规模 GPU 系统的格子玻尔兹曼
DOI: --
发表时间: 2011
期刊: International Conference on Parallel Computing
影响因子: --
作者:
A. Gray;Alistair Hart;A. Richardson;K. Stratford
通讯作者: K. Stratford
GPU 张量核心上的混合精度 LU 分解:减少数据移动和内存占用
DOI: --
发表时间: 2023
期刊: The international journal of high performance computing applications
影响因子: --
作者:
Florent Lopez;Théo Mary
通讯作者: Théo Mary